Skip to main content

AI SEARCH & AEO · September 2026 · ~10 min read

Non-determinism: why your AI ranking changes every time you check

Your AI ranking changes every time you check because these systems are non-deterministic. The same prompt does not return the same answer twice, even from one account on one day. There is no fixed position to hold. What exists is a probability of being named, and it is visible only by running a prompt many times and counting.

This single property breaks most of what the industry is currently selling, which is why it gets so little airtime.

If there is no stable position, there is no rank to report. If there is no rank, a dashboard showing your AI rank as 4 is showing you one draw from a distribution and calling it a measurement. That is not a small caveat. It is the difference between a number and a guess with a font.

01

What causes the variation?

Several things stacked on top of each other, and you control none of them.

The generation step itself. Language model inference is not deterministic even at temperature zero, because of batching, kernel scheduling and floating point ordering in the hardware. Thinking Machines Lab ran an identical prompt 1,000 times at temperature zero and got 80 distinct completions. A separate variance decomposition of 12,933 brand responses attributed 34.8% of the total variance to resampling the same prompt alone. A third of the movement happens before anything about you enters the picture.

Retrieval varies too. Google's AI features use query fan-out, which it describes as a set of concurrent, related queries generated by the model. The set of sub-queries generated is not fixed, so the pool of sources being summarized is not fixed either. Query fan-out changes what gets pulled in on any given run, which is a mechanism worth understanding on its own.

Then context. Your location, your account history, your device, the time of day, and whatever the system knows about the session. Two people in the same building can get different answers.

And underneath all of it, the products change constantly. Models are updated, retrieval is tuned, features roll out to some users before others. You are measuring a moving system with a moving instrument.

02

How much does it actually move?

More than almost anyone expects, and it has been measured properly.

Researchers at the University of St. Gallen ran four assistants daily across four verticals over a 45 day window in early 2026, and separately ran identical prompts ten times inside a single day. Two findings carry the article.

Day to day, across 4,044 consecutive-day pairs, only 34% to 42% of cited sources overlapped between two consecutive days. Roughly 65% of the sources behind an answer change overnight. Brand mentions are steadier, at 45% to 59%, and still highly variable.

Same day, same prompt, ten runs: source overlap of 32% to 43%, brand overlap of 33% to 59%. Nothing changed in the world between those runs. The instrument moved that much on its own.

Stability differs sharply by engine. Measured as source overlap, ChatGPT came in at 0.233, Perplexity at 0.282, Google AI Mode at 0.318 and Gemini at 0.505. The engine everyone asks about is the least stable one to measure.

One more finding that quietly invalidates a lot of tracking: ChatGPT activated web search on only about 42.2% of runs. The other 57.8% returned no citations at all. If a measurement does not record whether search actually fired, more than half of it is counting nothing.

03

Does that mean AI rankings are random?

No, and this is the part people overcorrect on.

Non-deterministic is not the same as random. The variation sits around a real underlying tendency. A business that is well documented, consistently described and frequently reviewed gets named more often. A business that is not gets named less often. Run enough samples and the pattern is stable even though no single run is.

Think of it the way you would think of covers on a Tuesday. You cannot predict tonight. You can predict the month, and if the month moves, something real changed.

The unit of measurement is the rate, not the result. Named in 7 of 20 runs is a fact about your business. Named once in a screenshot is a fact about that screenshot.

This also explains why appearing and then disappearing does not mean you were penalized. If you were at four in ten before and four in ten now, nothing happened. You just looked twice.

There is also a structural reason a small business is not competing on a level field inside any single run. Citation concentration across the four engines averaged a Gini coefficient of 0.715, meaning a small number of sources take most of the citations. Google AI Mode was the most concentrated at 0.782. Perplexity was the least at 0.671, which makes it the most winnable surface for a small site and the least representative one to build a score around.

04

Can anyone actually rank you first in ChatGPT?

Because there is no rank in ChatGPT to be first in.

There is no ordered list of businesses with positions that persist between queries. There is a generated answer, assembled from a source set that, on ChatGPT specifically, shares only about 23% of its members with the source set from an identical prompt run the same day. A claim to hold position one inside that is a claim about an object that does not exist.

The softer version of the same claim is more common and does more damage: "your AI visibility score went up 5 points this week." Run it against the sampling arithmetic below and it evaporates.

Two things you may not do with these numbers, and vendors do both.

You may not convert an appearance rate into a rank. Named in 7 of 20 runs is a frequency. It is not seventh place, and it is not a ranking that can be improved to sixth.

And you may not multiply rates from different denominators into a combined probability. The temptation here is to take the 42.2% web-search activation rate and the 32% to 43% same-day source overlap and produce odds that a given source appears twice in a row. Those are rates over different populations, the study does not publish the joint distribution, and the product would mean nothing. Report each next to what it measures.

05

How many runs before you believe anything?

The paper gives numbers, and they are the most useful thing in it.

  • At least seven runs per prompt per day for brand visibility, which is where the standard error falls below 0.10. At least eight if you care which sources were cited.
  • Report on a two to four week rolling aggregate, never week over week. The standard error of a detection rate falls below 0.10 at ten days and below 0.05 at twenty four.
  • Use a large, varied prompt set. Per-prompt source overlap inside a single campaign ranged from below 0.2 to above 0.8. Two prompts measure those two prompts' idiosyncrasies.
  • Prefer brand-level to source-level measurement, because source overlap is the noisier of the two.
06

What does a 5 point weekly move actually cost you?

Run the arithmetic before you renew anything on one.

The situation. A retainer at $2,500 a month, $30,000 a year. The dashboard reports an AI visibility score of 40 last week and 45 this week, and the renewal conversation is scheduled for Thursday.

The noise band. At seven runs per prompt per day, the standard error of an appearance rate is about 0.10. Two standard errors either side of 0.40 is a band running from roughly 0.20 to 0.60. The new reading of 0.45 sits comfortably inside it. A move from 40 to 45 is statistically indistinguishable from no change.

The window problem, which is worse. The paper puts the standard error below 0.05 only at 24 days of observation. A weekly report gives you seven of those days, which is about 29%derived of the observation window the claim requires. The reporting cadence is structurally incapable of supporting the sentence printed on the cover.

The decision cost. You are being asked to commit $30,000 on a reading that cannot distinguish 40 from 45, produced on a cadence that cannot reach the precision the reading implies.

What to ask for instead. The same money, reported as a 28 day rolling appearance rate per prompt, with the run count printed next to it. That is a number you can act on, and it will move less often, which is the point rather than a defect.

07

What should you do when a vendor sends you a screenshot?

Ask three questions. How many times was this prompt run. What was the exact prompt. What location and account state was used.

If the answers are one, unclear, and unspecified, you have been shown a coin landing heads. That is not dishonesty in every case. Plenty of people genuinely do not know that a single query is not evidence. But it decides how much weight the screenshot deserves, which is none.

The same discipline belongs in your regular reporting. What a monthly local SEO report should actually contain is sample sizes, methods and a record of what you changed, not a wall of numbers with no denominators.

Expect the answer to differ by assistant too, and be selective. Which surfaces matter for you depends on your category, in the same way that which review platforms actually matter varies by industry. A blended score across four engines is averaging four different amounts of noise, and the blend is dominated by the least stable one.

It is also worth checking whether the underlying advice survives this lens. Guidance built on the idea of a fixed AI ranking usually arrives bundled with tactics that do nothing, such as publishing a special file for AI systems, and a study of roughly 137,000 sites found 97% of those files are never fetched by anything.

08

What to do this week

1. Pick five prompts a customer would type and run each ten times. Record the businesses named each run. 2. Convert that into rates. Write them down with the date and the exact prompt. That is your baseline. 3. Record whether the assistant actually searched the web on each run. Runs with no citations are not measuring retrieval. 4. Stop reacting to single checks. Delete the habit of asking an assistant about yourself on a Tuesday and adjusting strategy on a Wednesday. 5. Ask any tool vendor you pay how many runs sit behind their number, then read how AI visibility tools compare once you ask that question. 6. Change one thing on your side this month and re-measure next month. One thing, or the measurement is wasted.

Before building a whole program around this, be honest about how much of it is genuinely new, because a lot of AEO is SEO with a new name, and that is not always a criticism.

Be honest with yourself

When you do not need this

If you are not measuring AI visibility at all, you do not need to solve its measurement problem. Do the ordinary local work and check in quarterly.

If you cannot commit to running prompts repeatedly on a schedule, do not start. A sporadic sample is worse than nothing, because it produces confident conclusions from noise and those conclusions cost money.

If a single check showed you missing and you felt a jolt of panic, that is the feeling this article exists to defuse. Run it ten more times before you do anything.

Sources

Related reading

12

Want a real distribution instead of a screenshot?

Send me one prompt you keep checking, exactly as you type it, at eric@seod.com. I will run it twenty times, record who gets named in each run, note which runs actually searched the web, and send you back the frequency table. No charge, no pitch, and the table is usually more interesting than whether you appeared.

I offer this because the difference between a screenshot and a distribution is the difference between anxiety and information, and twenty runs costs me very little.

Otherwise, keep reading the AI search library.

Call Eric Email Eric