Skip to main content

ANALYTICS & DASHBOARDS · September 2026 · ~10 min read

Setting a baseline before you change anything

A baseline is a dated record of where every number stood before you changed something, captured in a form you cannot revise later. Without one you can never prove a change worked, only argue about it. Capturing it takes about ninety minutes and it is the highest return ninety minutes in marketing.

Almost every dispute between a business and an agency is a baseline problem. Six months in, traffic is higher, the owner does not feel busier, and neither side can produce what the numbers looked like in month zero. The conversation becomes memory against memory, and memory reliably favours whoever is more confident.

The same thing happens without a vendor involved. You redesign the website, sales feel flat, and you have no idea whether the redesign hurt, helped, or did nothing while something else moved.

01

What exactly goes into a baseline?

Everything you plan to be judged on, plus the things that will be blamed later.

Outcome numbers. Revenue by month for the last two years if you have it, one year at minimum. Transaction count. Average ticket. New versus returning customers if your system distinguishes them.

Demand numbers. Inquiries per month, calls received and calls answered, form submissions, booking volume. Split by source if you can, and if you cannot, that is itself a finding.

Visibility numbers. Search Console clicks and impressions for the last sixteen months, exported. Rankings for your main terms, with the location you checked from written down. Review count and average rating per platform, on the date you checked.

Site numbers. Sessions by channel for twelve months. Conversion rate on your main pages. Page speed scores for your top three pages.

Cost numbers. Spend by channel per month. Anything you pay monthly that is supposed to produce customers.

Export the raw data where the tool allows it. Screenshots are acceptable and better than nothing, but a CSV survives a tool changing its interface, and interfaces change.

02

Why does the baseline have to come before the change?

Because after the change, the data you need is either gone or contaminated.

Some of it genuinely disappears. Analytics platforms cap retention. Ad platforms archive old campaigns. A website rebuild deletes the pages whose performance you wanted to compare against, along with the tracking that measured them.

The rest gets contaminated by knowledge. Once you know the outcome, you will select comparison periods that support it. That is not dishonesty, it is how people read data. The only defence is committing to the comparison period before you know the answer.

Write down what you expect to happen and by when, on the same day you capture the baseline. A prediction recorded in advance is the difference between measurement and storytelling. It also protects you in the other direction, because a change that underperforms your own written expectation is much harder for a vendor to reframe.

03

What does a written prediction actually buy you?

It buys you the ability to read a result correctly afterwards, and there is published arithmetic behind that claim.

Kohavi, Deng, Longbotham and Xu, generalising in the KDD 2014 proceedings from thousands of controlled experiments at Amazon, Booking.com, LinkedIn and Microsoft, work through the Bayesian version explicitly. With a standard significance threshold and standard power, a statistically significant result has an 89% chance of being a true positive if roughly one in three of your ideas actually works, which is the average Microsoft reports for itself. If only one idea in 500 works, that same significant result has a 3.1% chance of being real. The evidence did not change. Only the honest prior did.

That is the whole case for the prediction card, and here is what it looks like with numbers in it.

Capture day. Trailing twelve months of qualified inquiries: mean 46 a month, lowest 38, highest 61.

The card, written and dated before anything ships. "The new booking flow goes live 14 March. I expect qualified inquiries to average 55 a month across April and May, and I will judge it on the two months together, not on April alone. If it lands under 50 I will roll it back."

Size the claim against the history before you commit to it. 55 against a mean of 46 is nine inquiries, and the trailing high was 61. A single month of 55 proves nothing, because 55 has already happened without any change at all. Two consecutive months averaging 55 has not. That is why the card says two months.

The result. April lands at 57. May lands at 44. The two-month average is 50.5 derived, which is under the line you wrote. The change did not meet its own prediction, and you knew that in twenty seconds instead of arguing about it in June.

What happens without the card. April gets screenshotted, the 57 becomes the headline, May becomes "seasonal," and the flow stays. The KDD 2014 authors set the base rate for this honestly: at Bing, most experiments fail, and the ones that succeed move key metrics by 0.1% to 1.0% once diluted to overall impact, with roughly 10% to 20% of ideas succeeding at all. Your prior should be low, and writing it down is how you keep it low when the first month looks good.

04

How far back should the baseline go?

Twelve months minimum, twenty-four if your business has a season.

One month is not a baseline. It is a data point, and it might be a strange one. A single month contains a holiday, a weather event, a competitor closing, or one large customer.

If your business swings with the calendar, you need the full cycle. Otherwise you will compare your slow season to your busy one and conclude something false in either direction.

For anything that involves attribution, note the model in use at the time. Attribution settings change silently and change your history with them. If you already know how you are splitting credit across channels, record that method alongside the numbers.

05

Why can I not just compare myself to an industry average later?

Because for the metrics you care about most, there is no current published average, and the last one anybody published no longer measures the same thing.

The clearest case is the Google Business Profile. The only large-sample study of profile performance is BrightLocal's, published 15 July 2019, covering 45,264 local businesses across 36 industries with data from September 2017 to December 2018. It reports a median view-to-action conversion rate of 4.68%, medians of 1,009 searches and 1,260 views per listing per month, 59 total actions, and a 16% direct to 84% discovery split. BrightLocal sells local SEO software and drew the sample from its own audit-tool users, so it skews toward businesses that already care. It has never been refreshed.

And it is no longer reproducible, which is the part that matters for a baseline. Google has since removed Photo Insights, removed the direct versus discovery breakdown entirely, and redefined Views as a deduplicated unique-visitor count, where a person can only be counted once a day. The 2018 denominator and the 2026 denominator are different quantities. Any comparison of your current profile against that 4.68% is arithmetic across two definitions.

Two lessons follow, and they are the reason this article exists. Benchmark against your own trailing ninety days and your own named local competitors, never against a published average. And export raw, dated numbers now, because the platform that redefined Views did not ask anyone first and did not backfill the old definition. Owning the export is the only protection, which is the same reason your data should be something you can take with you.

06

What gets left out of a baseline, and why does that matter?

Two things, and they matter more than they look.

The first is anything you can only estimate. AI assistant visibility, brand awareness, word of mouth. These are real, they influence revenue, and they resist first-party measurement. Say so in the baseline document rather than substituting a vendor's score for them, because measuring AI search visibility honestly means naming what cannot be measured rather than filling the gap with a confident number.

There is a hard timing point buried in that one. Google's generative-AI performance data in Search Console began accumulating in May 2026 with no historical backfill, and it carries impressions, pages, countries, devices and dates only. No clicks, no click-through rate, no position, no prompts. You cannot go back and create an AI baseline for a period before the data existed. If you want one, the only move is to start exporting it monthly today. And if you want a mention-rate baseline, it has to be sampled rather than read: the University of St. Gallen work on AI answer variability recommends at least seven runs of each prompt per day and reporting on a two to four week rolling window, because a single run is not a measurement.

The second omission is operating context. What was happening in the business that month. Staffing gaps, a menu change, a supplier problem, a road closure. Marketing takes the blame for these constantly, and nobody can defend against it a year later without notes.

Restaurant operators have a version of this already. Comps and voids tell you what happened on the floor in a way the sales line never will, and the same principle applies here. Record the conditions, not only the result.

The connection between your marketing data and your transaction data is where the baseline gets genuinely useful. If your point of sale data can be joined to your marketing data, do that before the baseline rather than after.

07

What to do this week

Block ninety minutes. Create one folder, dated today.

Export twelve to twenty-four months from each source: Search Console, Analytics, your ad accounts, your booking or point of sale system, your call log. Raw exports, not summaries.

Screenshot anything that will not export. Rankings, review counts, profile insights.

Write the prediction card. What you are about to change, what you expect it to do, by when, over how many periods, and the number that would make you reverse it. Sign and date it. Put it in the folder.

Then, before you spend anything on more traffic, check the conversion side. If traffic is fine and inquiries are not, the baseline will show you that on day one and save you the campaign budget.

Be honest with yourself

When you do not need this

If the change you are making is small, reversible, and cheap, the baseline can be one screenshot. Do not build a project around swapping a headline.

If you are launching something entirely new with no history, there is no baseline to take. Capture the launch conditions instead, and accept that the first year is your baseline for the second.

And if you already have a stable dashboard with a year of history in it, you have a baseline. Date a copy of it and move on.

Sources

Related reading

11

Questions about your baseline?

Email me at eric@seod.com with the date you plan to launch or change something, and one line on what it is. I will send back the specific list of exports and screenshots to take before that date, tailored to your business type rather than a generic checklist.

It takes me a few minutes and there is nothing attached to it. Send it even if the change is next week.

The rest of the analytics library picks up from there.

Call Eric Email Eric