AI SEARCH & AEO · September 2026 · ~10 min read
Should you block AI crawlers
For most local businesses, no. Blocking removes you from systems your customers are already using, and it protects almost nothing, because your hours, address, services and reviews live on platforms you do not control. Blocking is a defensible choice for publishers whose written work is the product. It is rarely right for a local service business.
On this page
- 01Does blocking a crawler actually remove you from AI answers?
- 02What does each directive actually control?
- 03What is an AI opt-out service worth?
- 04Who has a genuine case for blocking?
- 05What does blocking cost a local business?
- 06What is the failure most businesses actually have?
- 07What to do this week
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Not sure what you are currently blocking?
The question usually arrives with some heat attached. Someone read that models are trained on the open web without permission, felt reasonably annoyed, and went looking for the switch.
The annoyance is legitimate. The switch does something different from what most people think it does.
01Does blocking a crawler actually remove you from AI answers?
Mostly not, and there is a large dataset on it.
BuzzStream, working with Citation Labs' XOFU tracking tool, examined 4 million citations across 3,600 prompts covering ChatGPT, Gemini, Google AI Overviews and Google AI Mode across ten industries, published in March 2026. It then read the robots.txt of every site being cited.
| Bot blocked in robots.txt | Blocking sites still cited |
|---|---|
| ChatGPT-User, live retrieval | 70.6% |
| OAI-SearchBot, search index | 82.4% |
| GPTBot, training | 88.2% |
| Google-Extended | 92.3% |
By citation volume, roughly 70% of all ChatGPT citations in that dataset came from sites blocking ChatGPT's retrieval bots, and roughly 95% came from sites blocking training bots. Two concrete cases: cnbc.com blocks ChatGPT-User, GPTBot and OAI-SearchBot at the same time and appeared 1,298 times. yahoo.com blocks Google-Extended and appeared roughly 30,000 times.
Nobody has a conclusive explanation, and BuzzStream does not claim one. Only 15% of the cited publications predated ChatGPT's launch, which weakens the "indexed before you blocked" theory. Around 70% of the dataset also blocks CCBot, which weakens the Common Crawl backdoor theory. Robots.txt is voluntary in the first place. And some retrieval systems lift citations straight out of search results, which your robots.txt has no authority over at all. BuzzStream sells link building software, so label it as vendor research with a disclosed method.
One combination the table does not support. It is tempting to multiply 70.6% by 82.4% to get the odds that a site blocking both is still cited. Do not. Those are two different populations of sites measured against two different directives, the study does not publish the intersection, and the product would be a number with no source. Read each row as its own finding.
02What does each directive actually control?
This is where most of the confusion lives, and the names are actively misleading.
GPTBot governs OpenAI's training crawl. It does not control ChatGPT search citations. 88.2% of sites blocking it were cited anyway.
OAI-SearchBot governs OpenAI's search index. It is the closest thing to "remove me from ChatGPT search," and 82.4% of the sites blocking it were still cited.
ChatGPT-User governs live, user-triggered fetches. Blocking it breaks the case where a customer pastes your URL into ChatGPT and asks about you, which is a use you probably want.
Google-Extended is the one people get most wrong. Google describes it as a robots.txt token, not a crawler with its own user agent. There is no Gemini crawler. Google grounds Gemini on content fetched by ordinary Googlebot. The token controls whether already-indexed pages may be used for training and grounding in Gemini. It does not remove you from AI Overviews or AI Mode.
To limit AI Overviews and AI Mode, Google names a different set of controls: `nosnippet`, `data-nosnippet`, `max-snippet` and `noindex`. Those are the same blunt instruments that strip your ordinary Search snippet. There is no way to opt out of one without paying for it in the other.
03What is an AI opt-out service worth?
Price it against what the directives can be shown to do.
The quote. An agency offers "AI content protection" at $1,200 to configure, plus $150 a month to monitor. Year one: $3,000.
What it buys, at best. Four robots.txt directives. On the measured evidence, the least effective of them leaves 88.2% of blocking sites still cited and the most effective leaves 70.6%.
Read the flip side carefully, because this is where people overclaim in the other direction. Yes, 29.4%derived of ChatGPT-User blockers were not cited. That is not a removal rate. Those sites might never have been cited regardless, and BuzzStream did not measure them before and after. The study shows blocking is not sufficient. It does not show how much it helps, and nobody has published that number.
Now the Google side, which is the part with a real price tag. The only controls Google names for its AI features are the snippet controls. Suppose you apply `nosnippet` sitewide. Only 7.9% of local searches trigger an AI Overview at all. You would be removing your snippet from the other 92.1%derived of local searches to influence the 7.9%. Run that on your own Search Console impressions rather than mine, but the ratio is the point.
So the honest scorecard. You pay $3,000 for four directives with an unmeasured effect, and the one control with a documented effect costs you ordinary search snippets across the board. Price the line item at zero and keep the $3,000.
04Who has a genuine case for blocking?
Publishers, first, and their case is real. US organic Google traffic between June 2025 and June 2026 fell roughly 50% at USA Today, 25% at CNN, 23% at Politico and more than 85% at Business Insider. When a summary of your writing substitutes for a visit to your page, the trade genuinely is different.
Membership sites and paid research, second. If people pay for access to the content, having it summarized for free is direct substitution.
Businesses with a specific licensing position, third. If you intend to license your archive and want a clean claim, blocking is part of that posture.
Notice what these have in common. The content is the product. For a plumber, a dentist, a restaurant or an agency, the content is an advertisement for the product. Advertisements do not usually want fewer distribution channels. And your exposure is a different shape: the wider zero-click trend is real, with SparkToro's panel putting 68.01% of US Google searches ending without a click in 2026, but a local business was never living mostly on organic clicks in the first place.
05What does blocking cost a local business?
Eligibility, mostly, and it is easy to overpay by accident.
The subtler cost is being absent from the moment of comparison. When someone asks an assistant which of three shops to call, the answer is assembled from whatever is retrievable. If your site is not retrievable, the answer gets assembled from your competitors' sites and whatever third parties say about you.
You also lose your own measurement. Blocked crawlers stop appearing in your logs, and you have removed the one deterministic signal you had. Whether you can see any of it downstream is a separate problem, since tracking referral traffic from AI assistants is harder than tracking anything else.
And blocking your own site does nothing about the information that actually answers local questions. Your name, address, phone, hours, categories, photos, services and reviews sit on your Business Profile and on review platforms. You cannot hide a local business from AI search by editing a file on your own server. The same asymmetry shows up in reviews, where each platform sets its own rules regardless of your preference, which is why Yelp's review solicitation rules differ from Google's.
Before deciding, get an honest read on what the exposure is currently worth. What AI Overviews did to click-through rates is the number that should drive this, not the annoyance.
06What is the failure most businesses actually have?
The opposite one, and it is invisible until you look.
The common, expensive, unintentional version of this problem is a CDN or security layer silently returning a 403 to OAI-SearchBot, PerplexityBot or Claude's crawler. Cloudflare, Sucuri and security plugins all do this under default or aggressive settings. Nobody chose it, nobody was told, and it does not appear anywhere in robots.txt.
So check server logs, not robots.txt. Robots.txt tells you what you asked for. Logs tell you what happened. If a retrieval bot is being turned away at the edge, you have accidentally bought the opt-out you were deciding whether to purchase, and you are not being cited while believing you are.
While you are auditing what you control, be skeptical of adjacent busywork. Bulk directory submissions are frequently citation building that is mostly theatre, and blocking crawlers can become the same kind of activity, something that feels decisive and changes nothing. If a tool told you your AI visibility dropped after a robots.txt change, verify before reacting, because AI visibility tools vary enormously in what they actually measure.
07What to do this week
1. Open your robots.txt and read it. Many sites already block crawlers nobody at the business chose to block, inherited from a template or a plugin. 2. Pull thirty days of server logs and check whether OAI-SearchBot, PerplexityBot and Claude's crawler are getting 200s or 403s. Verify by IP range, not by the user agent string. 3. Check your CDN and security plugin bot rules. That is where the accidental block lives. 4. Write down what you are actually trying to prevent, in one sentence. If the answer is "being summarized," check whether the summarizable information is even on your site. 5. If you decide to block, block by named token from the operator's own documentation, and record the date so you can connect it to any later change. 6. Do not touch snippet controls unless you accept that you are opting out of ordinary search results at the same time.
Be honest with yourself
When you do not need this
If you sell a local service and your content exists to bring in customers, you do not need to block anything. Spend the hour on your Business Profile instead.
If you have no website, there is nothing to block and the question does not apply to you.
If you are blocking out of principle rather than economics, that is a legitimate reason and I am not going to argue you out of it. Just do it knowingly, with the cost written down, rather than because an article implied it was a protective measure. On the measured evidence, it mostly is not.
Sources
- BuzzStream, "Does blocking AI bots stop citations?" Vince Nero. 19 March 2026. 4 million citations across 3,600 prompts, tracked with Citation Labs' XOFU tool across ChatGPT, Gemini, AI Overviews and AI Mode in ten industries. Source of the blocking table, the volume shares and the two named cases. Vendor research: BuzzStream sells link building software.
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Published May 2026, updated 10 July 2026. Source of the Google-Extended description, the absence of a Gemini crawler, and the snippet controls that do affect AI features. First party platform documentation.
- Google Search Central, list of Google crawlers and user-triggered fetchers. Reference for which tokens are crawlers and which are not. First party.
- OpenAI, bots and crawlers documentation. Source of the distinct purposes of GPTBot, OAI-SearchBot and ChatGPT-User, and of the published IP ranges used to verify them. First party platform documentation.
- Ahrefs, "What Triggers AI Overviews? 86 Factors and 146 Million SERPs Analyzed," Ryan Law. November 2025, desktop SERPs sampled September 2025. Source of the 7.9% local trigger rate used in the snippet arithmetic. Vendor research: Ahrefs sells SEO software.
- SparkToro with Similarweb, "In 2026, less than one third of Google searches still send a click". US panel, January to April 2026. Source of the zero-click figure and the publisher traffic declines. Panel estimates of Google search volume are contested, including by Google itself, so use the direction rather than the decimal.
Related reading
- AI crawler user agents and what your server logs reveal. How to run the log check in step two properly, with IP verification.
- What Google's own documentation says about optimizing for AI. The primary source for every claim here about what Google's controls do.
- Reviews as an input to AI answers. The asset your robots.txt has no authority over, and the one that carries most of your local exposure.
- Getting into AI answers when your business has no website. The clearest illustration of why blocking your own domain protects so little.
Not sure what you are currently blocking?
Send your domain to eric@seod.com and I will read your robots.txt and tell you which crawlers you are blocking today, which of those are search-facing versus training-only, and whether any of it looks accidental. If you can send thirty days of log lines too, I will tell you which retrieval bots are getting 403s from your CDN. No charge, no pitch.
I offer this because half the sites I look at are blocking something nobody at the business ever decided to block, and the accidental blocks are the expensive ones.
Otherwise, keep reading the AI search library.