AI SEARCH & AEO · September 2026 · ~10 min read
llms.txt: what a 137,000 site study found
An llms.txt file does nothing for AI search visibility. Ahrefs checked roughly 137,000 sites and found 97% of published llms.txt files were never fetched by anything at all. Google states its Search systems ignore them. If an agency is charging you to build one, they are charging for a file nobody reads.
On this page
- 01What is llms.txt supposed to do?
- 02What did the study actually measure?
- 03What is the line item actually worth?
- 04What does Google say about it?
- 05So what actually gets a local business into AI answers?
- 06Should I delete the one I already have?
- 07What to do instead
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Been quoted for AI search work?
That is a blunt opening, so here is the evidence behind it, along with what to do instead.
01What is llms.txt supposed to do?
The idea came from Jeremy Howard of Answer.AI and fast.ai, proposed in 2024 at llmstxt.org. A plain text file at the root of your site giving a language model a tidy map of your content. A robots.txt for models, roughly. The intent was reasonable and it was specific: an orientation index, so a model reading your documentation does not have to crawl everything.
Note what is not in that description. Nothing about rankings, citations, or visibility. The SEO industry attached the visibility claim later. The proposal never made it, and the person who wrote the proposal has never claimed it.
Worth being precise about one more thing: despite the filename, llms.txt controls nothing and blocks nothing. It is not a directive. It is a suggestion that nobody is obliged to read, and as it turns out, almost nobody does.
02What did the study actually measure?
Ahrefs took every domain in its Web Analytics product that received traffic in May 2026, 137,210 sites. It checked each root for a live llms.txt, verified the file was real Markdown returning a 200 rather than a soft error page, then examined every request made to llms.txt paths across the entire population, split by response code and user agent.
That method matters. It is not a survey of who publishes one. It is a log study of whether anything asks for them.
The findings:
- 28% of those domains publish an llms.txt file. Ahrefs discloses that as an upper bound, because its customers skew toward businesses that pay attention to SEO.
- 97% of those files received zero requests in May 2026. No bots, no humans, nothing at all.
- Of the traffic that did occur, 96% was bots and 4% humans, and 77% of those bots were not AI tools. SEO audit tools 21.7%, unidentified agents 14.9%, general crawlers 13.1%, technology profilers 11.6%.
- All four AI categories combined came to 19.5%. AI retrieval bots, the ones that actually produce citations in answers, were 1.1%, of which OpenAI's search crawler was 0.74%.
- Slackbot fetched llms.txt files more often than PerplexityBot did. Somebody pasting a link in a work chat generated more traffic to these files than the search product they were built for.
- Training crawlers fetched them roughly five times more often than retrieval bots did, which is the opposite of the direction that would help you.
The most telling result is a separate one. Ahrefs looked at requests for llms.txt files that did not exist and found the AI bot share was zero. Of all requests returning a 404, 98% were human. No AI system goes looking for the file speculatively. Publishing one does not put you on any list, and not publishing one does not remove you from any.
Ahrefs sells SEO software, so this is vendor research, and the company attaches its own caveat which is worth reproducing: fetched does not mean read, so every figure in the study is a ceiling rather than a floor. The real numbers are smaller than the ones above.
03What is the line item actually worth?
Price it against your own logs. The arithmetic takes ten minutes.
The quote. Say an agency wants $850 to build and install the file, plus $95 a month to maintain it. Year one: $850 plus $1,140, which is $1,990.
The measurement. Search ninety days of your raw access logs for requests to /llms.txt. Count them. Split them by user agent. Verify the interesting ones by IP range.
Scenario A, which the study says is the likely one. Zero requests in ninety days. Cost per useful fetch is undefined, because there were no fetches. You paid $1,990 for an object nothing asked for.
Scenario B, the good case. Four requests in ninety days, all from SEO audit tools and a technology profiler. That is $497.50 per fetch, by bots that cannot cite you. Better than zero, and still not a purchase.
One thing you cannot do, and I want to be precise about it. Ahrefs published a per-site figure, 97% receiving nothing, and a per-request figure, 1.1% of requests coming from retrieval bots. Those two cannot be multiplied into a probability for your site, because the study does not publish the joint distribution and the sites receiving traffic are not a random sample of the sites publishing files. Anyone who hands you a combined odds number has made it up.
What you can do is count your own. That is a real number about your business, and it costs nothing.
Then compare the alternative. The same $1,990 against category work on your Business Profile. Darren Shaw's study of 1.8 million Google Business Profiles across 4,209 categories found the exact match category outranks every adjacent one, with the pattern repeating for almost every category in the dataset, and primary category has been the highest scoring individual local ranking factor in every edition of Whitespark's expert survey. One of those two purchases has a large-sample finding behind it.
04What does Google say about it?
Google's own documentation on optimizing for its generative AI features addresses this directly, in a section titled Mythbusting:
"You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them... Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them."
Google is a self-interested narrator on plenty of topics. But on this one its position matches independent log data from a completely separate source, which is about as settled as anything gets in this field.
John Mueller of Google has put it more plainly still: "AFAIK none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it)." Read that as an instruction. He is telling you where to check.
There is one genuine inconsistency and you should know it before someone uses it on you. Google Search says skip the file. Chrome shipped a Lighthouse audit that checks for one. Pressed on the contradiction by Lily Ray, Mueller said the audit is "not done for search" and described llms.txt as a "temporary crutch, perhaps to save some tokens" for AI coding tools reading developer documentation.
That is the honest scope of the thing. If you publish a large developer documentation set and want coding assistants to find their way around it more cheaply, there is a narrow case. If you run a restaurant, a clinic or a contracting business, there is none.
The same documentation makes a broader point worth absorbing: optimizing for Google's generative features is optimizing for search. There is no separate discipline and no special file. Google's own guidance on AI optimization is short and worth reading in full, because most of what gets sold as AI search strategy contradicts it.
05So what actually gets a local business into AI answers?
The unglamorous answer is the same work that gets you into regular search results, because AI answers are built on those results.
For a local business, three inputs matter more than any file you can publish:
Your Google Business Profile. Google states plainly that its generative responses can include information about local businesses, and that a Business Profile helps your services appear in both AI responses and ordinary search results. This is the closest thing to official confirmation available, and it points at something you already control.
Your reviews. Whitespark's practitioner testing of AI Mode local answers in May 2026 found reviews and unstructured citations among the heaviest inputs, with ratings, review volumes, pricing and hours pulled from Business Profiles or comparable sources. That is vendor observation rather than a controlled test, and it points at the same asset Google does. How reviews feed AI answers is worth understanding on its own, because it changes how you prioritize them.
Mentions on other sites. Not directory listings, but the unglamorous stuff. A neighborhood blog, a local news piece, a community roundup, an association page. These turn up as citations in local AI answers far more often than most businesses expect.
None of it is exotic. It is the ordinary local visibility work, which is why understanding how the local pack is assembled remains the foundation under all of it.
06Should I delete the one I already have?
No. If your platform generated one automatically, and several do including Wix, leave it alone. Google says it neither helps nor harms, so removing it accomplishes nothing.
Two cautions if you keep it, and the first is not hypothetical.
Treat it as an untrusted input surface. One research crawler in the Ahrefs dataset identifies itself as `prompt-injection-survey/1.0`. Somebody is systematically probing llms.txt files as a prompt injection vector, on the reasonable theory that agents are built to trust a file whose stated purpose is orientation. Version-control the file, restrict who can edit it, and keep it to plain links with nothing instruction-shaped in it. If your platform generates it automatically, read what it generated.
And know that a file linked from nowhere is a file nobody finds. Agents fetch these when directed to them, not by going looking. The zero-404 finding above is the proof.
07What to do instead
Spend the hour you would have spent on llms.txt on any of these:
1. Fix your Google Business Profile categories, which are the strongest lever most businesses leave untouched after setup and the one with the largest study behind it. 2. Check whether your business details agree across the web. Inconsistent facts confuse the systems assembling answers about you. 3. Get one new review this week, and then keep going. Asking without sounding desperate is a solved problem. 4. Look at your own server logs to see which AI crawlers actually reach your site. That is real first-party data, unlike any visibility score you can buy, and it is the method both Google and Ahrefs point you at.
If you want to see what AI search has done to your click-through rate before deciding how much to care about any of this, that is a separate and more useful question.
Be honest with yourself
When you do not need this
If nobody has offered to sell you an llms.txt file, you can stop thinking about this entirely. There is no action to take.
If your business does not show up in ordinary Google results yet, AI visibility is not your problem. AI answers are built on search results. Fix the search results first and the rest follows.
If you publish a large set of developer documentation and want coding assistants to navigate it cheaply, the narrow original use case applies to you and this article does not. That is a documentation decision, not a marketing one.
And if an agency has already built you one, you have not been harmed. You paid for something with no effect, which is worth knowing for the next thing they propose.
Sources
- Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read," Linehan and Guan. June 2026, updated August 2026. 137,210 domains with traffic in May 2026, every request to llms.txt paths analysed by user agent and response code. Vendor research: Ahrefs sells SEO software, and states its own figures are ceilings.
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Source of the Mythbusting quotation, the four named ineffective tactics, and the Business Profile statement. First party platform documentation.
- llmstxt.org, the original proposal by Jeremy Howard. Answer.AI and fast.ai, 2024. Source of the file's actual stated purpose as an orientation index. Primary source for what it was designed to do.
- Darren Shaw, "What 1.8 million Google Business Profiles tell us about local SEO success," Search Engine Land. 25 August 2026. 1.8 million profiles across 4,209 categories. Cited as the alternative use of the same budget. Vendor research with a very large sample.
Related reading
- Structured data for AI: what is required and what is not. The other file-shaped thing you are being sold, separated into the part that works and the part that does not.
- When AEO is just SEO with a new name. How to price the rest of the proposal once you have removed this line.
- AI crawler user agents and what your server logs reveal. The method behind the ten minute check above, done properly.
- Google Business Profile categories: the single strongest lever you control. Where the money goes instead, with the 1.8 million profile study behind it.
Been quoted for AI search work?
Send the proposal to eric@seod.com and I will tell you which parts are real and which are being sold on vagueness. If you can also send ninety days of log lines for /llms.txt, I will tell you exactly what fetched it. No charge, no pitch, and I will say so plainly if the proposal is good.
I do this because the category is full of confident claims that fall apart when you check them, and a small business owner has no reasonable way to tell the difference. I have done the checking. You may as well have it.
Or keep reading more on AI search and AEO.