AI SEARCH & AEO · September 2026 · ~10 min read
Measuring AI visibility without fooling yourself
You cannot measure AI visibility with a single number. These systems are non-deterministic, so the same prompt returns different answers to different people at different moments. Honest measurement means repeated sampling of a fixed prompt set, plus your own server logs. Any product reporting one AI visibility score is showing you an average dressed as a fact.
On this page
- 01Why can't you just ask ChatGPT what it says about you?
- 02What does a valid sample actually look like?
- 03What is a monthly AI visibility report actually worth?
- 04What is the only fully deterministic thing you can measure?
- 05Which AI numbers in circulation are made up?
- 06What should a monthly report actually contain?
- 07What to do this week
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Questions about the number on your dashboard?
This matters because measurement decides spending. If your dashboard says your score went from 41 to 58, somebody is going to renew a retainer on that. If the number is noise, you just bought noise twice.
I spent sixteen years reading P&Ls where a number that looked precise was covering a counting problem underneath. Same shape here. The number is confident. The counting is not.
01Why can't you just ask ChatGPT what it says about you?
Because you are taking one sample from a distribution and treating it as the value.
There is a measured floor under this that no vendor can engineer around. Thinking Machines Lab ran an identical prompt 1,000 times at temperature zero and got 80 distinct completions, because batching, kernel scheduling and floating point ordering all move the result. A separate variance decomposition of 12,933 brand responses attributed 34.8% of the total variance to resampling the same prompt alone. Before location, account state or the time of day enters the picture, a third of the movement is the machine talking to itself.
The retrieval layer adds more. Researchers at the University of St. Gallen ran four assistants daily across four verticals for a 45 day window in early 2026, and separately ran identical prompts ten times inside a single day. Between two consecutive days, only 34% to 42% of cited sources overlapped. Same day, same prompt, the overlap was 32% to 43%. Roughly two thirds of the sources behind an answer change overnight.
That is not a bug you can report. It is how these systems work, and why your AI ranking changes every time you check it is worth understanding before you interpret any result.
The practical consequence is blunt. A screenshot proves nothing. It proves that on one run, at one moment, one system said one thing. An agency that sends you a screenshot of your business appearing in an AI answer has shown you a coin landing heads.
02What does a valid sample actually look like?
Repeated runs of a fixed prompt set, recorded as frequencies rather than yes or no. The St. Gallen paper puts numbers on what "repeated" has to mean:
- At least seven runs per prompt per day for brand visibility, which is where the standard error falls below 0.10. At least eight if you care which sources were cited.
- Report on a two to four week rolling aggregate, never week over week. The standard error of a detection rate drops below 0.10 at ten days and below 0.05 at twenty four.
- Use a large, varied prompt set. Per-prompt source overlap inside a single campaign ranged from below 0.2 to above 0.8. Two prompts measure those two prompts.
Build the prompt set from how customers actually talk. Not "best dentist San Jose" but the sentences people type into an assistant: who should I see for a cracked molar near Willow Glen, is there a dentist open Saturday in Honolulu, which dentist takes my plan and does not book six weeks out.
Then run each prompt on a schedule and record which businesses get named. After enough runs you have something real: you appear in roughly four of ten runs for this prompt, one of ten for that one, never for the third. It has a range, it moves for reasons, and it can be tracked.
You will also learn who is being cited alongside you, which is often the more useful output. Knowing which sources AI assistants actually pull for local queries turns a vanity check into a work list.
One warning from doing this repeatedly. You will occasionally see a competitor named that does not exist, or exists at an address nobody works from. AI answers inherit whatever is in the underlying local data, so fake listings in the map pack get laundered into confident prose. Report them, then carry on.
03What is a monthly AI visibility report actually worth?
Price it against the sample it contains. The arithmetic takes fifteen minutes and you can run it on any quote.
The quote. An agency adds an AI visibility report at $600 a month. Year one is $7,200.
Ask for the run log. Not the score, the log. Say the answer comes back as 25 prompts, checked once a week, one run each. That is 100 runs a month, 1,200 runs a year, which is $6 per run (derived).
Now price the same thing at the research standard. Twenty five prompts at seven runs a day over a 28 day rolling window is 4,900 runs a month. The delivered sample is 100 of 4,900, or about 2%derived of what the paper prescribes for the claim being made.
Then read the movement. The score went from 41 to 58. At one run per prompt, each prompt contributes a zero or a one. A 17 point move across 25 prompts is about four prompts flipping (derived). Four coin flips, reported as a trend, invoiced monthly.
One thing you cannot do, and I want to be precise about it. The same study reports that ChatGPT activated web search on only about 42.2% of runs, and separately that day to day source overlap runs 34% to 42%. It is tempting to multiply those into a single probability that a given source survives to tomorrow's answer. Do not. They are rates over different denominators, the study does not publish the joint distribution, and the product would be a number with no defined meaning. Report each one next to what it describes.
What you can do is count your own runs. That is a real number about your business and it costs nothing but time.
04What is the only fully deterministic thing you can measure?
Your server logs.
Every time an AI crawler fetches a page on your site, it leaves a record. That record is first-party, complete, and not an estimate. It is the one part of this subject where you can say a thing happened rather than probably happened.
The important detail is method. Identify crawlers by IP range, not by user agent string. A user agent is a text field that anything can type. Operators publish their IP ranges precisely so you can verify.
What logs cannot tell you is whether a fetch turned into a citation, or a citation into a customer. They tell you retrieval happened. That is less than you want and more than any score you can buy.
There is one official source of impressions now, and its limits are the point. Google launched Search Generative AI performance reports in Search Console in June 2026, giving impressions, pages, countries, devices and dates for its generative features. Clicks, click-through rate, position and the user's prompt are not available, and AI Mode is not separated from AI Overviews. No third-party product has better access than the platform's own reporting.
05Which AI numbers in circulation are made up?
Two get quoted constantly and neither survives a trace.
The first is that only 45% of businesses winning the Google map pack are recommended by AI assistants. The second is that just 12% of ChatGPT's citations match URLs on Google's first page. Both circulate in vendor blogs. Neither carries a sample size, a prompt set, a date range or a method, and no primary study produces either figure.
Look at what they have in common. Both are precise, both are about a quantity nobody can observe from outside these companies, and both are exactly the kind of number a business owner cannot check. That combination is the signature.
Google supplies the sentence to use on anything in that family: "Be wary of third-party tools that promise ranking success or claim to use 'internal' Google metrics. No third-party tool has access to our internal ranking or AI systems."
The test is not whether a number sounds plausible. It is whether you can name the sample. If a figure about AI citation rates arrives without a population, a prompt list and a date, treat it as decoration. That test also disqualifies a lot of things I would like to be able to tell you, which is the honest cost of applying it consistently.
06What should a monthly report actually contain?
Four things, and none of them is a single index number.
First, prompt-level appearance rates with the sample size stated. "Named in 6 of 20 runs" beats "score: 62" because it can be checked.
Second, crawler activity from your logs, by bot and by page, month over month.
Third, referral sessions from assistant domains, with the caveat that most AI-influenced visits never carry a referrer at all.
Fourth, a list of what changed on your side. Pages published, profile edits, reviews earned, mentions gained. Without this column the other three are unreadable, because you cannot attribute movement you did not cause.
If a report contains a score and nothing else, ask what the denominator is.
07What to do this week
1. Write twenty prompts in your customers' actual words. Not keywords. Sentences. 2. Run five of them, three times each, and record who gets named. That will teach you more than any tool demo. 3. Ask your host or developer for thirty days of raw access logs. If you cannot get them, that is worth knowing now. 4. Take whatever AI visibility number you pay for and ask two questions: what prompts, and how many runs. Run the cost-per-run arithmetic above on the answer. 5. Check whether your property has the Search Console generative AI view yet, and remember it reports impressions only. 6. Stop chasing markup as a measurement fix. Whether schema is required for AI answers has a documented answer, and it is not the lever people think.
Be honest with yourself
When you do not need this
If you are not spending money on AI search, you do not need to measure it. Curiosity is fine. A tracking program is premature.
If your business does not appear in ordinary local results yet, measure that instead. AI answers are assembled from search results, so the earlier fix is upstream and cheaper.
If you cannot commit to a schedule, do not start. A sporadic sample produces confident conclusions from noise, and those conclusions cost money.
If you run a small single-location business with steady word of mouth and a full book, this is a quarterly glance, not a monthly report. Regulated practices have the strongest case for watching closely, because an assistant repeating something inaccurate about a clinic is a different order of problem than a wrong closing time, and even responding to a patient review has rules about what you may confirm.
Sources
- Schulte, Bleeker and Kaufmann, University of St. Gallen, "Don't Measure Once: Measuring Visibility in AI Search (GEO)," arXiv:2604.07585. 8 April 2026. Four engines, four verticals, a 45 day window plus ten same-day repeated runs. Source of the run-count prescriptions, the overlap ranges and the 42.2% web-search activation rate. Academic, primary.
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Published May 2026, updated 10 July 2026. Source of the warning about third-party tools claiming internal Google metrics. First party platform documentation.
- Google Search Console help, performance report documentation. Reference for the dimensions Search Console publishes. First party. The Search Generative AI performance reports were announced by Google Search Central on 3 June 2026; that announcement is cited here without a link because the exact post URL is not recorded in our source library.
- Ahrefs, "How to Track AI Overviews: Mentions, Citations, Click Loss, and the Traffic Google Won't Show You". Method reference for isolating AI-shaped queries in Search Console rather than buying a score. Vendor research: Ahrefs sells SEO software.
- The "45% of map-pack winners are recommended by AI" and "12% of ChatGPT citations match Google page one" figures are listed deliberately without links. They circulate in vendor blogs with no traceable methodology, and this library records them only to refute them.
Related reading
- AI visibility tools compared honestly. The four categories of product and the seven questions that separate an instrument from a dashboard.
- AI crawler user agents and what your server logs reveal. How to do the only deterministic measurement here properly, with IP verification.
- Tracking referral traffic from AI assistants. What happens downstream of a citation, and why the number you see is a floor.
- When AEO is just SEO with a new name. How to price the rest of the proposal once the measurement line is sorted.
Questions about the number on your dashboard?
Send a screenshot of your AI visibility dashboard to eric@seod.com and I will tell you what the number is built from, whether the sampling behind it is defensible, and which part of it you can safely ignore. If you can also get the run log behind it, I will run the cost-per-run arithmetic above and send you the result. No charge and no pitch attached.
I do this because these dashboards are designed to be read, not audited, and most owners have no way to tell a measurement from a decoration.
Otherwise there is more on AI search and AEO here.