ANALYTICS & DASHBOARDS · September 2026 · ~11 min read
Measuring AI search visibility honestly
You can count AI crawler visits in your server logs, and you can estimate mention rates by asking the same question many times and recording the spread. You cannot know your AI visibility as a single number. These systems are non-deterministic, so any score presented without a sampling method is an estimate wearing the costume of a fact.
On this page
- 01What can actually be measured?
- 02How bad is the variation, exactly?
- 03What can only be estimated, and how do I estimate it properly?
- 04What does one round of sampling actually produce?
- 05Why do share of voice scores not work?
- 06What should I do instead of chasing a score?
- 07What to do this week
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Questions about AI visibility?
This is the fastest moving category in marketing and the least measurable, which is a combination that attracts vendors. A dashboard with a confident number on it sells better than an honest range, so confident numbers are what the market supplies.
The position worth holding is simple. Know what you can measure, know what you can only estimate, and know what nobody can measure. Then say which is which, out loud, in your own reporting.
01What can actually be measured?
One thing, fully and deterministically: crawler activity in your own server logs.
When an AI system fetches a page from your site, your server records it. That entry is first party, complete, and not sampled. It tells you which pages are being retrieved, how often, and by which agent. Nothing else in this category is that solid.
The method has one requirement people skip. Verify crawlers by IP range against the published ranges from each provider, not by the user agent string. User agent strings are trivially spoofed, and a meaningful share of traffic claiming to be an AI crawler is not. Server log analysis for AI crawler traffic is worth doing properly for exactly this reason, because the unverified version produces a number that is confidently wrong.
Logs also catch the failure nobody looks for. The common, invisible problem is not that you failed to block AI crawlers, it is that your CDN or security plugin is silently returning a 403 to OAI-SearchBot, PerplexityBot or ClaudeBot. That is invisible in robots.txt and obvious in a log file.
What log data tells you: whether your content is being fetched at all, which pages, and whether that changed after you published something. What it does not tell you: whether any of it was used in an answer. Retrieval is not citation. Hold both facts at once.
There is a second measurable, from June 2026. Google Search Central announced Search Generative AI performance reports in Search Console, giving impressions within generative AI features by page, country, device and date. Clicks, click-through rate, position and the user's prompt are not available, and AI Mode is not separated from AI Overviews. It is a separate view of the same dataset rather than a new one, so do not add it to your organic totals.
Referral traffic from AI assistants is partially measurable. Some assistants send identifiable referrers, some do not. Count what appears and label it incomplete.
02How bad is the variation, exactly?
Worse than most people assume, and it has now been measured properly.
Researchers at the University of St. Gallen published "Don't Measure Once" in April 2026. They collected daily results from four engines, ChatGPT, Gemini, Google AI Mode and Perplexity, across four verticals with eight prompts each, over a 45 day window, and separately ran ten repeats of identical prompts on the same day.
Across 4,044 consecutive day pairs, only 34% to 42% of cited sources overlapped between two consecutive days. Roughly two thirds of sources change overnight. Brand mentions were more stable at 45% to 59%, and still highly variable.
The same day result is the one that ends the argument. Identical prompt, same day, repeated runs: source overlap of 32% to 43%. ChatGPT was the least stable engine in the set. Nothing about the world changed between those runs. The system simply does not return the same answer twice.
There is a mechanical reason underneath it. Language model inference is not deterministic even at temperature zero, because of batching, kernel scheduling and floating point ordering. Thinking Machines Lab ran an identical prompt 1,000 times at temperature zero and got 80 distinct completions. A separate variance decomposition of 12,933 brand responses attributed 34.8% of total variance to within-prompt resampling alone.
One more finding changes how you count. In the St. Gallen data, ChatGPT activated web search on only about 42.2% of runs, meaning 57.8% of runs returned no citations at all. Any measurement that does not record whether search actually fired is measuring a mixture of two different things.
03What can only be estimated, and how do I estimate it properly?
Whether you get mentioned in answers. That requires repeated sampling and it produces a range, not a score.
Build a prompt set of questions a real customer would ask. Not keywords. Actual sentences, twenty to forty of them, covering the ways someone might describe their problem. The St. Gallen team found per-prompt overlap ranging from below 0.2 to above 0.8 within a single campaign, which means monitoring one or two prompts measures those prompts' quirks rather than your visibility.
Run each prompt at least seven times, on the same day, for each engine you care about. Seven is not arbitrary: it is the point at which the standard error of a brand visibility rate drops below 0.10 in their data. Eight or more if you care about which sources get cited rather than whether your name appears.
Report on a two to four week rolling aggregate. Never week over week. And prefer brand level measures to source level ones, because brand mentions are the more stable of the two.
The variation is large enough that you should be slow to react to it. Deciding whether a change in that rate means anything is the same problem as telling signal from noise anywhere else in analytics, except that here the underlying process is genuinely random rather than merely noisy.
04What does one round of sampling actually produce?
Here is the arithmetic from a single engine. Run it with your own prompt file.
The instrument. 20 prompts, 7 runs each, one assistant, one location. That is 140 runs.
The raw result. Your business appeared in 31 of them.
Headline rate. 31 of 140 is 22%derived. That is the number a vendor would put on a tile.
Now apply the denominator that matters. At the St. Gallen web search rate of about 42.2%, roughly 59 derived of those 140 runs would have run a live search at all. The rest answered from the model with no citations. Against 59 search-enabled runs, 31 appearances is 53%derived.
Same data, two defensible numbers, more than double apart. Which one is right depends entirely on the question. If you want to know how often a customer sees you, use 22%. If you want to know whether your content wins when the system actually looks, use 53%. A score with no stated denominator is not a measurement, and this is why.
What to report. Appeared in 31 of 140 runs across 20 prompts on one assistant, of which 59 ran a live search. Rolling four week window. That sentence is longer than a score and it is the only version anybody can check.
05Why do share of voice scores not work?
Because they add two different populations together, and because nobody has the denominator.
The first problem is measured. Kevin Indig's Growth Memo analysis found that 62% of the time, when an AI system uses your content and links to it, the answer text does not name your brand. So link based citation tracking and name based mention tracking are counting two different populations. A product that reports one "AI share of voice" figure has either picked one and called it both, or added them, and neither is a number you can act on.
The second problem is structural. Share requires a denominator, and no provider publishes total query volume by topic. Any claim about your share of AI answers is an estimated base multiplied by an estimated inclusion rate. Two estimates multiplied is not a measurement.
Two more things nobody can give you. Revenue attributable to AI assistants, because most of that traffic arrives without a traceable referrer and the customer who asked an assistant on Tuesday and typed your name on Friday looks like direct traffic. And whether a specific piece of content caused a specific mention, because there is no experiment available to you that isolates it.
Say all of that plainly on your dashboard, in the unknown column. A dashboard that labels its own uncertainty is more useful than one where every tile carries the same false confidence.
Keep the size of the prize in view while you do it. Sterling Sky reports that for one of its largest multi-location clients, ChatGPT traffic went from 0.1% of Google traffic in 2025 to 2% in 2026. A twentyfold growth rate on a tiny absolute number, and still only 22% of what Bing sends that same client.
06What should I do instead of chasing a score?
The work that helps AI visibility is the work that helps search, which is Google's own stated position. Their documentation puts it directly: "From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO."
Google also states that a page needs to be indexed and eligible to be shown with a snippet, and that there are no additional technical requirements. Structured data is not required for generative features. Google explicitly names llms.txt, artificial content chunking, and rewriting content for AI systems as ineffective.
The llms.txt data supports that flatly. Ahrefs examined all 137,210 domains in its web analytics product with traffic in May 2026. 28% published an llms.txt file, and 97% of those files received zero requests that month from any bot or human. Zero AI bots requested files that did not exist, so publishing one puts you on no list and skipping it removes you from none. Ahrefs sells SEO software, and its own caveat is worth keeping: fetched does not mean read, so every figure in that study is a ceiling.
So the honest playbook is unglamorous. Be indexed. Answer real questions clearly and early on the page. Keep your business facts consistent everywhere. Do the operational things that make a business findable, including the ones that have nothing to do with AI, like whether a mobile visitor can tap your phone number once an assistant sends them.
07What to do this week
Get access to your server logs, or ask your host for them. That is the first and most valuable step, and for most sites it takes one email. While you are in there, check for 403 responses to AI crawler user agents.
Write twenty prompts in customer language. Save them in a file with today's date. That file is now your instrument, and it only has value if it stays unchanged.
Run each prompt seven times on one assistant. Record raw results in a spreadsheet, including whether the assistant ran a live search. You will be surprised by the variation, and that surprise is the lesson.
Add an unknown column to whatever you report internally, and put AI share of answers in it.
Then treat the cost of all this like any other line you control. Prime cost is the number that matters in a restaurant because it is the one you can move; AI visibility spend deserves the same test of whether you can move it before you fund it.
Be honest with yourself
When you do not need this
If your customers are walkin, referral, or repeat, and nobody in your market is asking an assistant for a recommendation yet, measuring this is a hobby. Check again in six months.
If you have not yet done the basics, an accurate profile, a site that loads, pages that answer questions, do those first. AI systems draw on the same underlying index, so the basics are the intervention.
And if a vendor is offering to sell you an AI visibility score with no stated sampling method, you do not need that product. Ask how many runs, from where, with what variance, and whether links and unlinked mentions are counted separately. The absence of an answer is your answer.
Sources
- Schulte, Bleeker and Kaufmann, University of St. Gallen, "Don't Measure Once: Measuring Visibility in AI Search (GEO)," arXiv:2604.07585. 8 April 2026. Four engines, four verticals, a 45 day window of daily runs plus ten same day repeats. Academic, not vendor research. Source of the overlap figures, the seven run prescription, the rolling window rule and the web search activation rate.
- Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read". 137,210 domains with traffic in May 2026, every request to llms.txt paths across the population. Vendor research: Ahrefs sells SEO software. Its own caveat is that fetched does not mean read.
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Platform operator documentation. Source of the quoted position and of the list of tactics Google names as ineffective.
- Sterling Sky, "The State of Local SEO in 2026". Agency field data. Source of the client level ChatGPT traffic share and the Bing comparison.
- Kevin Indig, Growth Memo, on citation without naming. Source of the finding that 62% of the time an AI answer links a source without stating the brand name. Practitioner analysis.
Related reading
- Non-determinism: why your AI ranking changes every time you check. The mechanism underneath the variation figures above, explained without the statistics.
- AI visibility tools compared honestly. What to ask each vendor about run counts, prompt sets and confidence intervals before you buy one.
- Tracking referral traffic from AI assistants. The partial measurement, and how much of it goes missing before it reaches your analytics.
- AI crawler user agents and what your server logs reveal. The reference list for the only fully deterministic measurement in this category.
Questions about AI visibility?
Email me at eric@seod.com with three questions you think a customer would actually ask an assistant to find a business like yours. I will run each one seven times, record what comes back, and send you the raw results with the variation intact so you can see how much it moves.
I do this personally. You will get the unedited output, including the runs where nobody relevant appeared, because that is the part vendors leave out.
More on measurement sits in the analytics library.