Skip to main content

ANALYTICS & DASHBOARDS · September 2026 · ~11 min read

When a metric moves, telling signal from noise

Run four checks before you react. How large is the base, how far outside the normal range is the move, did anything change in the tracking, and did anything change in the world. If the base is small or the move sits inside the usual spread, it is noise. Most movements are noise.

The cost of getting this wrong runs both ways. React to noise and you change something that was working. Ignore signal and you lose a month before noticing a real break.

The way to tell them apart is not intuition. It is knowing what your normal range looks like, which almost nobody establishes in advance and which takes about an hour to establish once.

01

What is my normal range and how do I find it?

Take the last twelve to twenty-four periods of the metric. Weeks or months, whichever you review.

Find the highest and the lowest. That spread is your observed range, and anything landing inside it is not news. It is the metric doing what it has always done.

You can go one step further without any statistics. Sort the values, ignore the top two and the bottom two, and use what remains as your usual band. A move outside that band is worth a look. A move inside it is not.

Write the band down and put it on the dashboard next to the number. A tile that shows the current value against its historical range answers the signal question automatically, and it stops the meeting conversation before it starts.

Do this for every metric you review regularly. It takes an hour once and it saves an argument a month.

02

What does that look like with real numbers?

Here is a full worked band. Use your own twelve months.

The series. Monthly inquiries for the last twelve months: 41, 52, 47, 38, 55, 49, 44, 61, 39, 46, 50, 43.

Observed range. Lowest 38, highest 61. Anything between those two has already happened at least once.

Trimmed band. Drop the top two, 61 and 55, and the bottom two, 38 and 39. What remains runs 41 to 52. That is your usual band, and it took ninety seconds.

The average, for context. The twelve months sum to 565, so the mean is 47.

Now test three months against it. A month of 54 sits above the band but inside the observed range, so wait one more period. A month of 45 is squarely normal, so nothing happens. A month of 37 is below the band and below every month in the series, so it earns the four checks.

Size the 37 honestly before anyone panics. It is ten inquiries below the mean, which is 21%derived down. That is worth investigating, and it is still one month of a twelve month series. If the next month lands at 44, the 37 was a bad month and not a break.

Write the band on the tile. Inquiries 37, usual band 41 to 52, twelve month range 38 to 61. Anybody reading that tile now has the answer to the only question they were going to ask.

03

What are the four checks, in order?

Run them in this sequence, because the cheap ones eliminate most cases.

One, base size. How many events is this built on. Eleven leads after eight is three leads. Any metric with a base you could count on your hands will swing wildly every period with no cause. Below a few dozen events, stop here and wait.

Two, range. Is the new value outside the band you just calculated. If it is inside, you are looking at normal variation.

Three, tracking. Did anything change in how the number is collected. A website change, a new consent banner, a tag removed during a redesign, a filter added, a new booking tool. Tracking breaks are the single most common cause of dramatic metric moves, and they are almost always mistaken for real changes first. Check this before you look for a business explanation.

Four, the world. Weather, a holiday shifting weeks, a competitor opening or closing, a road closure, a price change, a staffing gap. Ask what was happening in the building. Marketing gets blamed for road works constantly.

Only after all four survive does a movement earn an investigation.

04

Which metrics move for boring reasons most often?

Some are noisy by nature and should be read gently.

Conversion rate at low traffic. The denominator is small, so this bounces. There is no visitor count that switches that off, which is the part the familiar rules get wrong. Kohavi, Deng, Longbotham and Xu, writing in the KDD 2014 proceedings, put the requirement where it belongs: minimum sample size follows from the metric's variance and from the sensitivity you want, meaning the size of the change you are trying to detect. The standard two proportion version is 16 × p × (1 - p) ÷ d², where p is your current rate, d is the absolute improvement you want to catch, and the 16 is 2(1.96 + 0.84)² rounded, 1.96 for 95% confidence and 0.84 for 80% power. At a 3% conversion rate, catching a 20% relative improvement, which would be a large win locally, takes about 12,900 sessions per variant (derived). Most local businesses are more than an order of magnitude below that, which means a conversion rate move in a single month cannot be separated from chance. Measuring conversion honestly when traffic is small means reporting the counts alongside the rate and refusing to read a single period.

Average position in search. Personalisation and location mean this varies without anything changing on your side.

Bounce and engagement metrics. These shift with traffic mix. More social traffic looks like worse engagement even when nothing got worse.

Direct traffic. A catchall bucket that absorbs whatever the tools cannot classify. It moves when tagging changes, not when behaviour does.

Call volume. Lumpy for most small businesses, and easily distorted if you have not separated answered from missed. It is also strongly patterned by daypart in a way that looks like performance and is not. Revmo AI's analysis of 12,091 restaurant call recordings found fast casual greeting consistency dropping 38 points from lunch to dinner and wait times rising 31% across the same shift change. A month with a different mix of lunch and dinner volume will move your call metrics with nothing having improved or degraded. Counting calls properly as conversions is what makes this metric readable at all.

Any AI visibility number. Covered in full below, because it is the extreme case and it is currently the most mis-read number in marketing reporting.

05

Why is a five point move in an AI visibility score always noise?

Because of how those scores are built, and the arithmetic is now published.

Researchers at the University of St. Gallen ran daily prompts across ChatGPT, Gemini, Google AI Mode and Perplexity over a 45 day window, plus ten repeats of identical prompts on the same day. Across 4,044 consecutive day pairs, only 34% to 42% of cited sources overlapped between two consecutive days. Same prompt, same day, repeated runs shared 32% to 43% of sources.

The instrument is unstable because the process is. Language model inference is not deterministic even at temperature zero, thanks to batching, kernel scheduling and floating point ordering. Thinking Machines Lab ran one prompt 1,000 times at temperature zero and got 80 distinct completions.

Now the arithmetic that matters for your report. The St. Gallen paper puts the standard error of a brand visibility rate below 0.10 at seven runs per prompt per day. It puts the standard error of a detection rate below 0.10 only at ten days of observation, and below 0.05 at twenty four. So a visibility score that moved from 40% to 45% week over week is statistically indistinguishable from no change. The paper's own instruction is to report on a two to four week rolling aggregate and never week over week.

That is the whole debunk of the most common line in an AI reporting deck. When a dashboard says the score went up five points this week, the honest reading is that nothing has been demonstrated. Ask three questions: how many runs per prompt, over what window, and what is the confidence interval. If a tool cannot tell you its n, it is selling you noise with a chart on it.

The same discipline applies to any number produced by sampling rather than counting, which is more of your reporting than you would like.

06

What does a real signal look like?

It has three properties, and it usually has all three.

It persists. A real change holds across multiple periods rather than spiking and returning. One bad week is weather. Four bad weeks is a problem.

It shows up in more than one place. Fewer inquiries and fewer calls and a lower conversion rate together is a signal. One of the three alone is often noise.

It has a plausible cause you can point at. Not a story invented afterwards, but something specific with a date. A change went live, a competitor launched, a page dropped out of the index.

One caution on the third property. Two credible measurements of the same real effect can disagree by a wide margin, and that does not make either one wrong. Ahrefs measured the click impact of AI Overviews at a 34.5% decrease in one cut and 58% in a later one, while SparkToro's panel work puts it around 60%. Same phenomenon, different baselines, dates and query mixes. When you attach a cause to a movement, quote a range rather than a point, or the next person to measure it will appear to contradict you.

When all three properties line up, act. When one does, wait another period. Waiting is almost always cheaper than reacting, and it is the discipline that separates operators who read data well from those who chase it.

Restaurant operators have the strongest version of this instinct already. Reading point of sale data for something other than sales means looking at void patterns, daypart shifts, and item mix to find the cause underneath the number, rather than reacting to the number itself.

07

What to do this week

Pick your three most watched metrics. Pull twelve periods of history for each and calculate the observed range and the trimmed band.

Add the band to your dashboard as a small note under each number. Current value, usual band, full range. That is the whole build.

Write the four checks on a card and keep it where the review happens. Base, range, tracking, world.

Then verify your tracking is currently working. Submit a test form, tap your own phone number, and confirm both appear. Doing this monthly catches the most common cause of fake signals. The small set of GA4 reports worth keeping makes that check quick.

Make the whole thing part of an existing routine. A monthly review habit that survives busy season is what turns these checks into something you actually run rather than something you agree with.

Be honest with yourself

When you do not need this

If the movement is in a metric you were not going to act on anyway, do not investigate it. Curiosity is expensive on a busy week.

If the cause is already obvious and documented, you closed for three days, skip the checks and write the note.

And if your business is small enough that you personally see every customer, you already know why the number moved. The framework is for people who cannot see the floor.

Sources

Related reading

11

Questions about a number that moved?

Email me at eric@seod.com with the metric, the before and after values, and roughly when it changed. Three lines is plenty. I will tell you which of the four checks I would run first and whether the movement is large enough to be worth your afternoon.

I answer these myself, and the honest answer is often that it is noise and you can stop thinking about it. That is a useful thing to be told.

More on measurement sits in the analytics library.

Call Eric Email Eric