ANALYTICS & DASHBOARDS · September 2026 · ~10 min read
Small sample sizes and how to avoid fooling yourself
A metric moving on a small base is usually noise. Eleven leads after eight is three leads, not a thirty-eight percent improvement. Before you act on any change, ask how many events it is built on, and if the answer is under a few dozen, treat it as a hint and wait for more data rather than a result.
On this page
- 01Why do small numbers produce such dramatic percentages?
- 02Where does the "you need 10,000 visitors to test" rule come from?
- 03How do I work out the number for my own site?
- 04What should I do instead of testing?
- 05Where does the small-sample rule mislead in the other direction?
- 06How do I stop fooling myself in ordinary reporting?
- 07What to do this week
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Questions about whether your numbers mean anything?
This is the most expensive misunderstanding in small business marketing. Owners change a page, see a better week, keep the change, and never learn it did nothing. Or they cut a channel after a bad fortnight that was inside the normal range all along.
Both mistakes come from reading percentage change on a base too small to carry it.
01Why do small numbers produce such dramatic percentages?
Because the percentage divides by the base, and a small base makes every ordinary fluctuation look enormous.
Go from four inquiries to six and you have a fifty percent increase. You also have two inquiries, which could be one person who called twice and one who found you through a friend. Nothing about your marketing changed.
Now consider the reverse. Six back to four is a thirty-three percent decline and it will get somebody blamed.
Any number small enough to count on your fingers will swing by large percentages every period, forever, with no cause. That is not a measurement problem you can fix with a better tool. It is arithmetic.
The practical rule: read absolute change first, percentage second. Write both. If the absolute change would not have registered as remarkable in conversation, the percentage is not telling you something new.
There is a formal version of the same instinct, and it comes from people who ran experiments at a scale nobody reading this ever will. Kohavi, Deng, Longbotham and Xu, in the KDD 2014 proceedings, name it Twyman's law: "Any figure that looks interesting or different is usually wrong!" Their examples run to hundreds of thousands and often millions of users per experiment, and they still treat a striking number as a bug report first and a finding second.
02Where does the "you need 10,000 visitors to test" rule come from?
A blog post. That is the whole answer, and it is worth knowing before somebody quotes it at you or you quote it at yourself.
The figure surfaces as reference [29] in that same KDD 2014 paper, where it resolves to Neil Patel, "11 Obvious A/B Tests You Should Try," QuickSprout, 14 January 2013. The four authors, from Microsoft and LinkedIn, cite it and immediately qualify it in the same sentence: "Our advice in previous articles is that you need 'thousands' of users in an experiment; Neil Patel [29] suggests 10,000 monthly visitors, but the guidance should be refined to the metrics of interest."
A 2013 blog post, promoted to an industry threshold by thirteen years of repetition. It is not a research finding, it has no sample and no method, and it is quoted constantly by agencies explaining why they can or cannot test something for you.
What the paper says instead is the sentence to build on: "Formulas for minimum sample size given the metric's variance and sensitivity (the amount of change one wants to detect) provide one lower bound." Sample size is a function of how noisy your metric is and how small a change you want to catch. It is not a traffic number, and two metrics on the same page can differ by more than a factor of twenty.
Their own worked table for Bing makes that concrete. A revenue-per-user metric with a skewness coefficient of 17.9 needed 114k users per variant to detect a 4.4% change. The same metric with outliers capped, skewness 5.2, needed 9.7k for 10.5%. Sessions per user, skewness 3.6, needed 4.70k for 5.4%. Same company, same site, same week. The metric decided the answer.
03How do I work out the number for my own site?
Worked example
With one formula from the paper and a calculator. Here is the whole derivation.
The authors give a rule of thumb for when the arithmetic behind a test is trustworthy at all: the minimum observations per variant is 355 multiplied by the square of the metric's skewness coefficient, recommended whenever skewness is above 1. That is a precondition rather than a full power calculation, so read the result as a floor and not as a finish line.
Step one, your conversion rate. Say your money page converts at 3%. A visitor either converts or does not, so the skewness of that measurement is fixed by the rate itself, and at 3% it works out to about 5.5 derived.
Step two, square it and multiply. 5.5 squared is 30.3, and 355 times 30.3 is roughly 10,760 sessions per variant (derived).
Step three, divide by your traffic. At 620 sessions a month split evenly between two variants, each variant collects 310 a month. 10,760 divided by 310 is about 35 months (derived). Nearly three years to satisfy the floor, before you have detected anything.
Step four, and this is the part that makes the point. Rerun it for a page converting at 10% instead of 3%. Skewness drops to about 2.7 derived, squared is 7.3, times 355 is roughly 2,590 sessions per variant (derived), which at 310 a month is about eight months (derived).
Same site, same traffic, same test, and the requirement fell by roughly three quarters because the metric changed. That is the entire argument. Anyone who answers "how much traffic do I need" without asking which metric you plan to move has not done the arithmetic.
Notice also where the folk rule came from. A 3% binary conversion metric lands near ten thousand, which is why the number feels right to people. It is right for one metric at one rate, by coincidence, not as a law.
What should I do instead of testing?
Make changes that are obviously right rather than marginally different.
The tests that need statistics are the small ones. Button colour, headline phrasing, three words in a form label. Those produce differences too small for your volume to detect, which means you will never know, which means the effort is wasted. The KDD 2014 authors are blunt about the size of real effects even at Bing: most experiments fail, and the successes move key metrics by 0.1% to 1.0% once diluted to overall impact. A local business cannot see a one percent change in anything.
The changes worth making at low volume are the ones you could defend without any data. A phone number that is tappable. A form with three fields instead of nine. A page that says what you sell above the fold. A booking flow that works on a phone.
That last one carries more weight than any test you could run. Mobile conversion problems that desktop testing never finds are usually not subtle. They are a broken element, a form that will not submit, a button under the keyboard. You do not need a sample size to justify fixing a thing that does not work.
Use qualitative evidence at low volume. Watch five people use the site. Read twenty inquiry messages. Listen to ten calls. Small samples are unreliable for measuring rates and excellent for finding problems, and those are different jobs.
The honest alternative to a test is a before and after over a long window, compared against the same period a year earlier, with the change and the date written down in advance. It is weaker evidence and it is what you actually have. Call it what it is.
05Where does the small-sample rule mislead in the other direction?
Three places, because "wait for more data" is not always the right answer either.
When the cost of being wrong is trivial. If a change is free and reversible, waiting three years for significance is the expensive option. Ship it, watch loosely, revert if it feels wrong.
When the sample is small because something broke. A lead count that halved is not a small-sample artifact if the form stopped submitting. Check the tracking before you invoke noise, because a broken tag produces exactly the shape of a real decline.
When you are looking for a problem rather than measuring a rate. Five session recordings cannot tell you your conversion rate. They can absolutely tell you that the phone number is not tappable on an iPhone. Do not apply a rate-measurement standard to a problem-finding activity.
06How do I stop fooling myself in ordinary reporting?
Four habits, and they cost nothing.
Always show the base. Every percentage on every report gets its absolute number beside it. Half of all bad conclusions die here.
Compare like periods. Not this month against last month. Seasonality distorts adjacent month comparisons enough to reverse the apparent direction of a trend, and it will do so consistently in the same months every year.
Use rolling averages for anything volatile. A four week or twelve week rolling figure smooths the swings that were never signal. Watch the rolling line, not the weekly points.
Write your prediction down first. Before a change, record what you expect and by when. A change that beats a written expectation is evidence. A change explained afterwards is a story. That is the whole reason capturing a baseline comes before anything else.
The reason these fail is never that people disagree with them. It is that they are extra steps in a busy week, which is the same reason SOPs get ignored by week two. Build them into the report template so nobody has to remember.
07What to do this week
Take your last three monthly reports. Find every percentage. Write the absolute numbers next to them.
Count how many of your conclusions survive. In most businesses it is fewer than half.
Pick your main conversion metric and calculate its monthly count. If it is under about thirty, accept now that you will not be able to measure small improvements, and redirect that effort into fixing obvious problems.
Run the derivation above once, with your own conversion rate and your own monthly sessions. Keep the answer written down, because it is the reply to every future proposal that starts with "we should test."
Add a rolling twelve week average to your main chart.
And make sure you can still do all this next year. Rolling averages and year over year comparisons need history, which means owning your data and being able to take it with you is what makes any of this possible after a vendor change.
Be honest with yourself
When you do not need this
If your business genuinely does high volume, thousands of conversions a month, the small sample problem is not yours. Test properly and enjoy it.
If you are making a decision that is cheap and reversible, skip the analysis. Change it, watch loosely, change it back if it feels wrong. The cost of being wrong is the only thing that justifies the cost of being sure.
And if you are looking at your own operation daily and can see what is happening, your eyes are a better instrument than a small sample. Trust them and stop building charts.
Sources
- Kohavi, Deng, Longbotham and Xu, "Seven Rules of Thumb for Web Site Experimenters," KDD 2014. Peer reviewed conference paper generalising from thousands of controlled experiments at Amazon, Booking.com, LinkedIn and Microsoft properties, most involving millions of users. Source of the sample-size rule, the skewness formula, the Bing metric table, the effect-size range and Twyman's law. The authors state their own limit: these are principles, not provably correct, and this is desktop web-scale data rather than local business data.
- Neil Patel, "11 Obvious A/B Tests You Should Try," QuickSprout, 14 January 2013. Cited only as the traceable origin of the 10,000-monthly-visitor rule, via reference [29] of the paper above. A blog post, with no sample and no method.
- Baymard Institute, "Checkout Optimization: 5 Ways to Minimize Form Fields in Checkout," 26 June 2024. Vendor research: Baymard sells UX benchmarking. Cited as an example of a directional principle worth acting on without a test of your own, and limited to e-commerce checkout rather than local lead forms.
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Platform operator documentation, including Google's warning to be wary of third-party tools that claim ranking success or internal Google metrics. Relevant here because the same credulity that accepts a folk threshold accepts a vendor score.
Related reading
- When a metric moves, telling signal from noise. The four checks to run in the moment, once you accept the base is small.
- Reading a monthly marketing report critically. Where somebody else's percentages hide their bases, and the three questions to send back.
- Measuring conversion honestly when traffic is small. What to report instead of a conversion rate when the denominator cannot carry one.
- Does A/B testing work for most local businesses?. The same arithmetic applied to a specific purchase decision you may be facing.
Questions about whether your numbers mean anything?
Email me at eric@seod.com with two numbers: your monthly website visitors and your monthly leads or bookings. That is all I need. I will run the derivation above on your figures and tell you the smallest change you could realistically detect, roughly how long it would take, and whether a test anyone has proposed to you is worth running.
I answer these myself. If the answer is that you cannot measure what you were hoping to measure, that is what you will get, along with what I would do instead.
More on measurement sits in the analytics library.