CHOOSING & WORKING WITH AN AGENCY · September 2026 · ~11 min read
How to read an agency's case studies critically
Look for three things: the starting number, the time period, and what else changed during it. A result without a baseline is not a result, a result without a date range cannot be judged, and a result achieved while the client also doubled their ad spend belongs partly to the ad spend. Most case studies are missing at least one.
On this page
- 01What is missing from almost every case study?
- 02How likely is it that the tactic caused the result?
- 03Which case study result should you refuse outright?
- 04Which metrics get chosen because they look good?
- 05Is this result even relevant to me?
- 06What should I ask about a case study?
- 07What does a good case study look like?
- 08What to do this week
- 09When you do not need this
- 10Sources
- 11Related reading
- 12Questions about a case study you are weighing?
This is not an accusation of dishonesty. Case studies are written by marketers who want to show the best true version of a story, and the best true version tends to leave out the messy parts. Your job is to ask for the messy parts.
We retired one of our own for exactly this reason. It claimed applications had tripled, and when someone asked what the number was before, the honest answer was that nobody had written it down. The change was real. The claim was not defensible. It came down.
01What is missing from almost every case study?
The starting number, most often.
"Grew to 400 calls a month" means nothing without knowing whether they started at 380 or at 40. When a case study gives you only the ending figure, that is usually because the beginning figure makes the story smaller.
Here is the arithmetic, and it takes ten seconds. From 380 to 400 is an increase of 5.3%derived. From 40 to 400 is an increase of 900%derived. Same headline, same ending number, two results that are nothing like each other. Write the ending figure on a page, then write both plausible starting figures underneath, and you will see immediately why the beginning was left out.
Second most often, the time period. Results over eighteen months and results over three months are different products at different prices, and a case study that says "after our work" is hiding the duration for a reason.
Third, what else was happening. New location, a competitor closing, a seasonal peak, a PR mention, a large increase in paid budget. Any of those can carry a result that gets attributed to the agency.
Ask what else changed during the period. It is one sentence, it is not aggressive, and the quality of the answer will tell you more than the case study did.
02How likely is it that the tactic caused the result?
Lower than the case study implies, and there is published arithmetic for this that nobody in marketing quotes.
Kohavi, Deng, Longbotham and Xu published "Seven Rules of Thumb for Web Site Experimenters" at KDD 2014, generalising from thousands of controlled experiments run at Microsoft, Amazon, LinkedIn and Booking. Two findings from it should change how you read any results story.
First, real wins are small. At Bing, most experiments failed, and the ones that succeeded improved key metrics by 0.1% to 1.0% once diluted to overall impact. Their reported success rate on ideas was about 10% to 20%. A case study describing a transformation is describing something that almost never happens at organisations with enormous testing capacity.
Second, the arithmetic on believing a result. The paper works it through directly. With a five percent false positive rate and eighty percent power, if one in three ideas is genuinely good, then a statistically significant result has an 89% chance of being real. If only one in 500 ideas is genuinely good, that same significant result has a 3.1% chance of being real.
You can rerun this on your own reading. Ask yourself how many of the tactics you have been pitched over the years actually worked. If your honest answer is closer to one in twenty than one in three, then a single impressive case study should move you very little, because the prior is doing more work than the evidence.
The single most useful line in the paper for reading a results story is its warning about published case studies. The authors ask whether the test was peer reviewed, whether it was properly run, whether there were outliers, and note that they have seen tests published where the significance threshold was not even met.
03Which case study result should you refuse outright?
The one about button colour, and it is worth knowing by name because it is still in circulation.
The claim is that red call to action buttons beat green, and that it was A/B tested. The underlying test was real, published by Joshua Porter. Kohavi and his co-authors name it in KDD 2014 as their example of a result that does not generalise, on the plain observation that if it were a general result, more sites would have red buttons. Their phrasing is that they do not believe it replicates well.
The same paper names a second one: the advice to use the word "free," from a different popular source. Their heading for the whole section is that your mileage will vary.
This matters far beyond buttons. A case study is a single test in a single business at a single moment, and the industry treats it as a technique. When an agency says a tactic works because it worked for a client, they are making the red button claim in a new costume. Ask how many times they have run it and how many times it worked. A vendor who has run something eleven times and seen it work seven times is telling you something real. A vendor with one story is telling you a story.
Where this breaks down, honestly: some case studies describe a fix rather than a tactic. Correcting a business category, recovering a suspended profile, or repairing tracking that had been broken for a year is not a hypothesis being tested. Those transfer, because the mechanism is deterministic. Ask which kind you are looking at.
04Which metrics get chosen because they look good?
The ones furthest from money.
Impressions grow when Google shows you more often, including for searches you do not want. Keywords ranked grows when you publish anything at all. Sessions grow when a page gets shared, or when a bot pattern changes. All three can climb during a quarter when the phone got quieter.
Ask for the metrics that sit next to revenue. Calls answered. Forms submitted. Appointments booked. Jobs sold. If those are not in the case study, ask why, and accept a real reason. Sometimes the client would not share revenue data, which is legitimate and should be stated rather than covered with a proxy.
A proxy metric can also point the wrong way, and the same paper documents it. Microsoft Office Online tested a redesigned page and measured clicks on revenue generating links as a stand in for purchases. The new page produced a 64% reduction in clicks per user. It also showed the price, so the users who did click were better qualified and converted at a much higher rate. Read as a case study, that page was a disaster. Read as a business result, it was not.
AI visibility claims deserve the same scrutiny and get less of it, because the measurement is genuinely unstable between runs. It is easy to fool yourself when measuring AI visibility, which means it is also easy to build a case study out of noise without meaning to. One detail worth knowing when you are handed one: the St. Gallen measurement study found citation concentration varies sharply by engine, with Google AI Mode the most concentrated surface and Perplexity the least. A case study built on Perplexity wins is a case study built on the easiest available surface, which does not make it false and does make it less impressive than it sounds.
05Is this result even relevant to me?
Check the business model before you check the numbers.
A storefront and a business that travels to customers rank through different mechanisms, and service area businesses face constraints a storefront never encounters. In Whitespark's 2026 expert survey, having an address showing on the profile rather than operating as a service area business scored 176, high in the top ten, and setting service areas larger than about two hours of driving was added as a suspension risk. A case study from a retail shop tells a mobile locksmith very little.
Check the competitive density. Results in a town with three competitors do not transfer to a city with sixty.
Check the starting position. Taking a business from invisible to present is a different job from taking a business from fourth to first. The second is harder and slower.
Check the client's own capacity. A result driven partly by a client who answered every request in an hour is not repeatable with a client who takes two weeks. That is worth knowing about yourself, honestly, before you buy.
06What should I ask about a case study?
Five questions, in one email, all reasonable.
What was the number before you started. What was the date range. What else changed in that period. Is this client still with you, and if not, why. Can I speak with them.
That last one is where case studies get tested. A case study you can call is worth ten you cannot, and reference calls give you something usable if you ask past the script.
If the client cannot be named for confidentiality reasons, that is a fair answer, and it should come with an offer of a different reference in the same industry.
07What does a good case study look like?
Boring, specific, and slightly unflattering in one place.
It states where the business started with real numbers. It states the period. It says what was actually done, in enough detail that you could check a few items yourself. It names something that did not work, because in any nine month engagement something did not.
It tells you what the arrangement was. Whether the client had inhouse help, how much of their time it took, and how often they met, because contact frequency shapes results more than most case studies admit.
And it says how the work was staffed. Whether it was one practitioner or a team, and whether any of it was subcontracted. That detail is treated as a secret and it should not be, since the choice between an offshore team, a local shop, and a freelancer changes what you can expect.
08What to do this week
Take the case study that impressed you most and write three numbers in the margin. Start value, end value, elapsed months. Whichever one you cannot fill in is your question. Then run the two percentage calculations above with the best and worst plausible starting numbers.
Send the five questions. Judge the reply on speed and specificity, not on polish.
Then ask the question that every agency with a portfolio should find awkward, and we are not exempt. How many engagements have you run in total, how many are on your website, and which one did you take down? SEOD removed a case study because a tripling claim had no recorded baseline, so we can answer the last part. We cannot claim our published work is a random sample, and neither can anyone else. Every portfolio you will ever read is selected by the person selling to you, which means the ratio of engagements run to case studies published is a real and answerable number that almost nobody volunteers.
Then look for the case study they did not publish. Ask for an engagement that did not go well and what they learned. The willingness to describe it plainly is the closest thing to watching them handle a bad month with you.
Be honest with yourself
When you do not need this
If you already have a strong referral from a business you trust in your own industry, the case studies are decoration. Call the referral and skip the reading.
If you are buying a small defined project, this level of scrutiny costs more than the project. Check that they have done the same task before and get on with it.
And if an agency simply does not publish case studies, do not treat that as a flag on its own. Plenty of good practitioners work under confidentiality, or never got around to writing any because they were busy. Ask for references instead.
Sources
- Kohavi, Deng, Longbotham and Xu, "Seven Rules of Thumb for Web Site Experimenters," KDD 2014. Peer reviewed conference paper drawing on thousands of controlled experiments at Microsoft, Amazon, LinkedIn and Booking. Source of the effect size range, the 89% and 3.1% posterior calculation, the Office Online result and the red button warning. Desktop web scale data, explicitly not local business data.
- Schulte, Bleeker and Kaufmann, University of St. Gallen, "Don't Measure Once: Measuring Visibility in AI Search," arXiv, 8 April 2026. Four engines over a 45 day window plus same day repeats. Source of the citation concentration comparison across engines. Academic preprint, primary research.
- Whitespark, "Local Search Ranking Factors," Darren Shaw, 6 November 2025. Forty seven experts scoring 187 factors. Source of the service area business scores. Vendor research, and expert opinion rather than test data.
Related reading
- The claims agencies make that do not survive checking. The same discipline applied to a statistic instead of a results story.
- Questions to ask before hiring a marketing agency. Where the case study questions sit inside a full first conversation.
- Signs your current agency has stopped doing the work. How the same reporting habits show up once you are the client in the story.
- Setting expectations for the first 90 days. What a realistic version of the case study timeline looks like from inside.
Questions about a case study you are weighing?
Email me at eric@seod.com with the case study, or a link to it, and I will send back the list of questions it does not answer. Just the questions, for you to ask them. I will not comment on whether to hire the agency and I will not put myself forward as an alternative.
There is more on hiring and working with agencies here as well.