Skip to main content

AI SEARCH & AEO · September 2026 · ~11 min read

AI visibility tools compared honestly

Most AI visibility tools do one of four things: sample prompts and count mentions, monitor brand mentions across the web, track AI Overview presence in search results, or read your server logs. Only the last is deterministic. The other three are estimates, and the quality of an estimate is the sample size and prompt list you are not shown.

I am not going to rank named products. I have not audited them all, and a ranking of tools I have not tested would be the kind of confident nonsense this library exists to counter.

What I can give you is the four categories, what each can honestly claim, and the questions that separate a real instrument from a dashboard. Those questions work on any vendor, including ones that did not exist when this was written.

01

What are the four kinds of tool?

TypeWhat it doesWhat it can honestly claim
Prompt samplersRun a prompt list against assistants on a schedule, record who is namedAn appearance rate for those prompts, at that sample size
Mention monitorsWatch the web for your brand name in textWhere you are being written about
Overview trackersCheck whether search results show an AI Overview for tracked keywordsFeature presence and which sources were cited
Log analysersRead your server logs for crawler activityThat specific crawlers fetched specific pages, verified

Most commercial products bundle two or three of these and present the result as one score. The bundling is where the honesty goes. A composite index made of two estimates and a count is not more accurate than its parts, it is just harder to argue with.

A number you cannot decompose is a number you cannot check.

02

What sample size does an honest estimate actually require?

More than any tool advertises, and the number comes from a paper rather than a vendor.

Researchers at the University of St. Gallen ran four engines daily across four verticals over a 45 day window in early 2026, plus ten repeated same-day runs of identical prompts. Their prescriptions are specific and you should hold any vendor to them:

  • At least seven runs per prompt per day for brand visibility, which is where the standard error drops below 0.10. At least eight when you care about which sources were cited.
  • Report on a two to four week rolling aggregate. Never week over week. The standard error of a detection rate falls below 0.10 at ten days and below 0.05 at twenty four.
  • Use a large, varied prompt set. Per-prompt source overlap inside a single campaign ranged from below 0.2 to above 0.8. Monitoring two prompts measures those two prompts' quirks, not your visibility.
  • Prefer brand-level to source-level measurement, because source overlap is the noisier of the two.

Stability also differs by engine, which means a single blended score across four engines is averaging four different amounts of noise. Measured as source overlap, ChatGPT came in at 0.233, Perplexity 0.282, Google AI Mode 0.318 and Gemini 0.505. ChatGPT, the engine everyone asks about, is the least stable one to measure.

There is a floor under all of this that no vendor can engineer around. Large language model inference is not deterministic even at temperature zero, because of batching, kernel scheduling and floating point ordering. Thinking Machines Lab ran an identical prompt 1,000 times at temperature zero and got 80 distinct completions. A separate variance decomposition of 12,933 brand responses attributed 34.8% of total variance to repeated sampling of the same prompt alone.

03

What does the research-grade standard cost you?

Run the arithmetic before you decide whether to buy, build, or skip.

The full standard. Twenty prompts at seven runs each is 140 runs a day. Over a 28 day rolling window that is 3,920 runs. By hand at forty seconds a run, that is roughly 43 hours a month. Nobody with a business to run is doing that manually, which is the actual argument for a tool.

Now discount the sample. In the same study, ChatGPT activated web search on only about 42.2% of runs, meaning 57.8% returned no citations at all. Apply that to your 3,920 runs and about 1,654 of them involved retrieval. If a tool does not filter for whether search fired, more than half of what you paid to collect is counting nothing. That is question one for any vendor.

Then price the alternative. A tool quoted at $199 a month is $2,388 a year. Substitute the real quote. The decision is not whether $2,388 is a lot of money. It is whether 43 hours or $2,388 changes something you would otherwise do.

And the cheap middle path. Ten prompts, seven runs each, one day a week for four weeks is 280 runs a month, about three hours by hand. The paper does not endorse it, because it samples one day in seven rather than continuously. It is still enough to see a direction over a quarter, and it costs nothing. Call it a compromise when you report it.

04

Which one should you buy first?

For most local businesses, none of them, and I will defend that.

If you have to start somewhere, start with the log analyser, because it measures something that definitely happened. You may already have this for free depending on your host and analytics setup.

Prompt samplers are the most genuinely useful of the estimating tools, provided you can see and edit the prompt list. A sampler running its own generic keyword list is measuring a market you do not compete in. A sampler running twenty sentences your customers actually say is measuring your business.

Mention monitors are worth it if you intend to act on mentions. They are a work list generator, not a scoreboard.

Overview trackers matter most if you have real organic traffic exposure. If your business runs on the map and the phone, the feature presence data is interesting rather than decisive.

One finding is worth knowing before you pick an engine to obsess over. Citation concentration in the same study, measured as a Gini coefficient, averaged 0.715 across engines, meaning a small number of sources take most of the citations. Google AI Mode was the most concentrated at 0.782. Perplexity was the least at 0.671, which makes it the most winnable surface for a small site and the least representative one to build a score around.

05

What questions should you ask before paying?

Seven, and any vendor worth buying from answers all seven without flinching.

What is the exact prompt list, and can I edit it. How many runs per prompt per period. Do you filter out runs where the engine did not perform a web search. What location and account state do the runs use. How do you identify crawlers, by user agent string or by verified IP range. What exactly goes into the composite score and in what weights. What does the tool do when the answer names me without linking me.

The crawler question is the fastest filter. Identification by user agent string means counting self-descriptions, and anything can type any string. Verification by IP range is the only fully deterministic method available.

The score question is the second fastest. If nobody at the company will tell you the weights, the score is a brand asset, not a metric.

Google's own guidance supplies the third filter, in a sentence written for exactly this situation: "Be wary of third-party tools that promise ranking success or claim to use 'internal' Google metrics. No third-party tool has access to our internal ranking or AI systems."

Cross-check anything a tool tells you against what Google's own documentation says about optimizing for AI. A tool recommending special files, artificial chunking, or content rewritten for machines is recommending things Google names as ineffective, which tells you what its model of the world is built on.

06

Why is "AI share of voice" usually a broken number?

Because it adds two things that are not the same thing, and the gap between them has been measured.

Kevin Indig's Growth Memo analysis found that 62% of the time, when an AI system uses your content and links to you, it does not say your brand name in the answer text.

Read what that does to a dashboard. A link-based tracker sees a citation and counts it. A name-based tracker reads the answer, finds no brand mention, and counts nothing. The two methods are measuring different populations, and a composite that adds them produces a number with no defined meaning. Ask a vendor which one their score is built on. If the answer is both, ask how they are weighted. If there is no answer, the score is decoration.

The same problem has an official version. Google's Search Generative AI performance reports in Search Console give impressions, pages, countries, devices and dates. Clicks, click-through rate, position and the user's prompt are not available, and AI Mode is not separated from AI Overviews. No third-party tool has better access than the platform's own reporting, so any product claiming a click number or a position inside an AI answer is generating it rather than measuring it.

07

What does no tool do?

None of them tell you whether a citation produced a customer. That chain, from mention to visit to inquiry to sale, is broken in several places and no vendor has fixed it. Clicks from AI Overviews arrive in your analytics as ordinary organic traffic with no separate channel, which is a gap in the data rather than a gap in the product.

None of them see the whole surface. A large share of what gets said about a local business comes from a Business Profile, review platforms and third-party pages, which is why a business with no website still shows up in AI answers and why a site-centric tool has a blind spot the size of your actual visibility.

None of them fully model retrieval. Your one question becomes several searches behind the scenes, and query fan-out means the pages that get pulled in are often not the ones you targeted. A tool tracking your target keyword is watching the wrong door.

None of them can give you a rank, because there is no rank to give.

And none of them will tell you to spend the money elsewhere, which is frequently the correct advice.

08

What to do this week

1. Write your own prompt list, twenty sentences in customer language, before you evaluate any tool. It is the thing you are actually buying. 2. Run five of those prompts ten times each by hand. Now you have a baseline and a sense of the noise, for free. 3. Run the cost arithmetic above with the real quote in front of you, and answer the only question that matters: what would you do differently at each possible value of the number. 4. Send the seven questions above to any vendor you are considering. Judge the replies, not the demo. 5. Get your server logs, or find out you cannot. Either answer is useful. 6. If the tool budget is under real scrutiny, put it into your Business Profile categories instead, which remain the strongest local lever most businesses leave untouched after setup.

And keep the review flow running while you deliberate, because a quiet month costs you more than a missing dashboard. A negative review beats no new reviews is an uncomfortable rule, and it is the one that governs.

Be honest with yourself

When you do not need this

If you are not yet visible in ordinary local search, do not buy a tool to tell you that. Fix the ranking. AI answers are built on those results.

If you are a single-location business with a full book, a quarterly manual check is proportionate. A subscription is not.

If you would not change anything based on the number, do not buy the number. That question kills most of these purchases, and it should.

If your agency includes an AI visibility dashboard in the retainer, you are already paying for it. Ask them the seven questions instead of buying a second one.

Sources

Related reading

12

Considering a tool right now?

Tell me which AI visibility tool you are evaluating at eric@seod.com and I will send you the seven questions phrased for that vendor, plus what I would want to see in each answer before paying. If you have already had the demo, send me the deck and I will tell you what the headline number is made of. No charge, no pitch.

I offer this because the demos are good and the methods are usually undisclosed, and an owner has no fair way to tell those apart.

Otherwise, keep reading the AI search library.

Call Eric Email Eric