Your SEO plugin has probably offered to build you one.
An llms.txt file is a plain list of your important pages, written for software instead of for people. The idea is that an AI tool reads the list and takes your best work rather than guessing.
It is a sensible idea. What is missing is anyone confirming they read it.
Google has now put its position in writing, and the answer is short.
What an llms.txt file actually is
It is a markdown file at the root of your site listing the pages you want a language model to use. The author who proposed it, Jeremy Howard, describes it as a proposal to standardise "on using an /llms.txt file to provide information to help agents use a website."
The format is loose, and only one part of it is required.
- Where it goes: The spec says it "can be placed at the site root, or at any path within it, covering the pages under that path."
- What it must contain: "An H1 with the name of the project or site. This is the only required section."
- What it is for: The reasoning given is that "agents are best served by concise, expert-level information gathered in a single, accessible location."
- What it is not: The page calls itself a proposal. Nothing on it claims that any AI company has agreed to the idea.
That last point matters more than the format does. A proposal becomes a standard when the software on the other end starts honouring it, and that is the part to check.
The file was first published in September 2024 and was last revised in August 2026, so the idea has had two years to be picked up.
What Google says about llms.txt
Google answered this in its own documentation in June 2026, and the wording leaves no room to argue.
You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them.
That note went into the documentation on 15 June 2026.
The same page then deals with the obvious follow-up, which is whether publishing one could hurt. Google's guide to AI features says it is "completely fine" to keep such a file for other systems, and that doing so "will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them."
That sentence settles more than the obvious question.
- Which surfaces does it cover? The guide is about AI Overviews and AI Mode as well as the blue links, so the answer reaches all of them.
- Is there a risk? No. The file is neither a ranking factor nor a penalty, so nobody selling you one is lying about that half.
Google also notes that it may crawl and index many kinds of file besides HTML, and that doing so gives none of them special treatment. A file being fetched once is not the same as a file being used, which is worth remembering when a report shows you a hit.
Do ChatGPT, Claude and Perplexity read it?
None of them has published anything saying so. Their crawler documentation names the agents they send and the file those agents obey, and in all three cases that file is robots.txt.
| Company | Crawlers named in its docs | File the docs say it obeys | Says it reads llms.txt |
|---|---|---|---|
| OpenAI | OAI-SearchBot, GPTBot, ChatGPT-User, OAI-AdsBot | robots.txt | No |
| Anthropic | ClaudeBot, Claude-User, Claude-SearchBot | robots.txt | No |
| Perplexity | PerplexityBot, Perplexity-User | robots.txt | No |
All three name the same file.
Anthropic's crawler documentation puts it plainly, saying its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt." Perplexity's crawler page recommends that site owners allow its bot "in your site's robots.txt file" and says nothing about any other file.
There is a genuine source of confusion here, and it explains a lot of the enthusiasm.
All three companies publish an llms.txt file for their own developer documentation. Publishing one is the opposite of reading one. A firm looking at those files and concluding the companies must therefore consume them has the arrow pointing backwards.
What the server logs show
Somebody measured it. Ahrefs published a study of llms.txt files in June 2026 built on server-log data from 137,210 domains, covering requests made during May 2026.
Two findings matter, and both need their denominator said out loud.
- Around 28% of those 137,210 domains publish an llms.txt file, which is roughly 38,000 sites.
- Of the domains that have a valid file, 97% saw no requests for it at all during that month.
So the file is common and the requests are not. The small share of traffic that did arrive was mostly automated, and the study attributes 96% of it to bots of every kind rather than to AI systems specifically.
Almost nothing asked for the file. That covers one month on one panel of sites, so treat it as a strong signal instead of a permanent law, and note that it agrees with what the crawler documentation already said.
If you want to know what is reaching your own pages, your logs answer that better than any study, and the same reading tells you how much of your traffic is bots.
Why your plugin is offering to make one anyway
Because it costs the plugin nothing and it sounds like progress. Search for the term and the completions fill with plugin names, which tells you where the demand is being manufactured.
The offer is not dishonest. A generated file is accurate, cheap and harmless, and Google has confirmed it carries no penalty.
The problem is what it displaces.
- It looks like the work: A partner who has ticked the AI box stops asking the question that would have found the real gap.
- It is measured by its existence: Nobody reports on whether the file was fetched, because the plugin cannot see that.
- It arrives instead of a page: The same hour spent writing one clear answer on your own site produces something an engine can quote.
An engine cites a page because the page answered a question well. That is the whole mechanism behind getting cited by AI tools, and no index file shortcuts it.
The file that AI crawlers really do read
The file they all name is robots.txt.
Every crawler in the table above names it, and it is the one file with the power to stop an AI system reading your site at all.
| File | Who reads it | What it changes |
|---|---|---|
| robots.txt | Every named AI crawler and Googlebot | Whether a bot may fetch your pages |
| Your XML sitemap | Search engines | Which URLs get found and how fast |
| llms.txt | No system has published that it does | Nothing that has been demonstrated |
The risk runs the wrong way from what most owners expect. Security plugins and hosting firewalls block AI agents by default fairly often, so a firm can publish an index file for crawlers it is quietly turning away at the door.
Check the live file at your own domain followed by /robots.txt, then confirm your XML sitemaps are listed in it. Those two lines do more than any new file will.
Should your firm publish one?
Publish it if it takes five minutes, and expect nothing.
The honest case for the file is that it is free, it cannot hurt you, and some future tool might adopt the idea after all.
The dishonest case is the one being sold, which prices it as AI visibility work.
- Fair reason to add it: Your platform generates and updates it automatically at no cost.
- Fair reason to skip it: Somebody wants a fee, or wants to hand-maintain it every month.
- Bad reason to add it: A report promises AI traffic from it.
- Better use of the same hour: Write the one page a client actually asks about, and check that crawlers can reach it.
Track what happens instead of assuming. Google now reports where its AI features showed your pages, which is what Search Console's AI report exists for, and the tools that watch other engines are covered in our look at AI visibility tools.
Where that leaves the file
An llms.txt file is a reasonable proposal with no confirmed reader. Google has said in writing that Search ignores it, three AI companies document robots.txt and not this, and the logs from May 2026 show almost nobody asking for the files that exist.
- Add it if it is free: Let the platform generate it and move on.
- Fix robots.txt first: A blocked crawler beats a missing index file every time.
- Spend the budget on pages: An answer worth quoting is still the only thing that gets quoted.
The firms showing up in AI answers this year got there by being the clearest source on a question somebody asked. If you want to know which of your pages a crawler can actually reach, that is where our free site audit starts.
Frequently Asked Questions
Does it have to sit at the domain root?
The spec allows either the root or a subfolder, where a copy covers the pages beneath it. Nothing has been shown to fetch either location, so the placement question stays theoretical.
Will adding one hurt our rankings?
Google states the file neither helps nor harms visibility in Search. The cost is the attention it takes from work that does move something.
Is this the same as blocking AI from our content?
Blocking happens in robots.txt, where a disallow line stops a named agent fetching your pages. An index file makes no claim about permission and grants nothing.
Our agency charges a monthly fee to maintain it. Is that normal?
A recurring fee to hand-edit a file with no documented reader is worth questioning. Ask what evidence of fetches they can show from your own server logs.
Does an llms.txt file replace a sitemap?
Sitemaps are read by search engines and feed real discovery, so keep yours. The two files serve different audiences, and only one of those audiences has confirmed it is listening.

