WEBSITE CONVERSION · September 2026 · ~11 min read
Measuring conversion honestly when your traffic is small
At small traffic, a monthly conversion rate is mostly noise. Count absolute leads instead of percentages, widen the window to a quarter, and compare against the same quarter last year. A business producing 9 to 24 leads a month cannot detect a small change, and the arithmetic that proves it is four lines long.
On this page
- 01Why is my conversion rate jumping around every month?
- 02How many visitors does a real answer actually need?
- 03Where does the 10,000 visitor rule come from?
- 04What should I measure instead of a monthly percentage?
- 05How wide does the window need to be?
- 06What do I do about the report my agency sends?
- 07What to do this week
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Questions about reading your own numbers?
This is the part nobody in marketing wants to say to a small business, because the entire reporting industry depends on monthly numbers that appear to mean something.
I came to this from restaurants, where the same trap exists. One slow Tuesday is not a trend. Two slow Tuesdays is not a trend either. Reacting to every wobble is how operators end up changing four things a month and never learning which one worked.
01Why is my conversion rate jumping around every month?
Because you are dividing a small number by another small number, and both wobble.
Take a site with 500 visits and 15 leads. Next month, 480 visits and 11 leads. The rate looks like it dropped meaningfully. What actually happened is that four fewer people called, and any of a dozen ordinary things explains four people. A holiday, weather, one commercial client who called instead of filling in the form.
Then somebody hunts for a cause and makes a change. The next month lands higher, because small numbers drift back toward their average, and the change takes credit it did not earn. Now you believe something untrue.
Small numbers move for reasons that have nothing to do with your website. Treating that movement as signal is how bad decisions get made confidently.
The people running the largest experimentation programs in the world are blunter about this than any agency will be. Ron Kohavi, Alex Deng, Roger Longbotham and Ya Xu, writing from Microsoft and LinkedIn in a 2014 KDD conference paper drawn from thousands of controlled experiments at Amazon, LinkedIn and Bing, invoke Twyman's law: any figure that looks interesting or different is usually wrong. At Bing they treated a surprisingly good result as a bug report, not a win.
02How many visitors does a real answer actually need?
Here is the arithmetic, in full, so you never have to take anyone's word for it again.
You are comparing two proportions. The standard sample size formula for that comparison, at the conventional 5% significance level and 80% power, is:
visitors per variant = 16 x p x (1 minus p) divided by the square of the difference you want to detect
The 16 is not arbitrary and you can rebuild it. It is two times the square of 1.96 plus 0.84, the two z values for 5% two sided significance and 80% power. That is 2 x 2.80 x 2.80, which is 15.68, rounded to 16. The p is your current conversion rate. The difference is the size of the improvement you want the test to be able to see, in percentage points.
Now put your own numbers in. Say you convert at 3%, so p x (1 minus p) is 0.0291, and 16 x 0.0291 is 0.4656. Divide that by the square of the improvement you want to detect.
| Improvement you want to detect | New rate | Visitors per variant | Conversions per variant | Months at 500 visits |
|---|---|---|---|---|
| A fifth better | 3.6%derived | 12,933 | 388 | 52 |
| Half again better | 4.5%derived | 2,069 | 62 | 8 |
| Double | 6% | 517 | 16 | 2 |
Read the middle column. To catch a realistic improvement, a fifth better than what you have now, you need roughly 12,933 visitors in each version. Two versions means about 25,900 visitors. At 500 visits a month that is over four years for one question.
The threshold is not a traffic number. It is a function of your conversion rate and how small a change you want to see. That is why nobody can honestly hand you one figure without asking what you are trying to detect.
03Where does the 10,000 visitor rule come from?
This one is worth knowing, because you will be quoted it.
The claim that a valid split test needs around 10,000 monthly visitors is everywhere in conversion optimization, usually stated as though it came from research. It did not. It surfaces as reference 29 in that same Kohavi, Deng, Longbotham and Xu paper, and resolves to a QuickSprout blog post by Neil Patel titled "11 Obvious A/B Tests You Should Try," published 14 January 2013. The four authors cite it and qualify it in the same sentence: the guidance, they write, should be refined to the metrics of interest.
A blog post from 2013, promoted to an industry threshold by twelve years of repetition. That is the whole trail.
The related folk figure, roughly 100 conversions per variant, has the same problem and you can now check it yourself. Run the formula backwards. One hundred conversions at a 3% rate is about 3,333 visitors per variant, and 0.4656 divided by 3,333 gives a detectable difference of about 1.2 percentage points. That is an improvement of roughly 39%derived over your starting rate. The "100 conversions" rule is only valid if you are hunting for a nearly 40% improvement, and almost nothing on a website produces one.
The honest version of the rule is that the popular thresholds are roughly right for enormous effects and badly wrong for realistic ones. The arithmetic above is the thing to trust. The round numbers are folklore.
For scale on what a realistic effect looks like, the same paper reports that at Bing most experiments fail, and those that succeed move key metrics by 0.1% to 1.0% once diluted to overall impact. Roughly one in 500 clears the bar of large, replicable, positive impact.
04What should I measure instead of a monthly percentage?
Four things, all of them countable and none of them requiring a sample size.
The absolute count of real inquiries. Fifteen leads this month, eleven last month. A count is honest about being small in a way that a percentage is not. "Down four" invites the right question. "Down three points" invites panic.
The count of qualified inquiries. People you could actually serve, in your area, for work you do. This is the number that connects to revenue, and it is often flat while total inquiries swing.
Booked jobs, and revenue from them. This is the only number that pays anyone.
And direct observations that do not need statistics at all. A dead tap in a session recording, a field where people abandon, a question five callers asked. Those are facts about your site, not estimates, and they are true whether you have 300 visits or 30,000.
Be careful comparing pages against each other too, since your homepage and a campaign page get different visitors with different intent. A homepage and a landing page are doing different jobs, and their conversion rates are not comparable numbers.
05How wide does the window need to be?
Wide enough that ordinary variation stops dominating.
A week is meaningless at this volume. A month is barely readable and should be treated as a count, not a rate. A quarter against the previous quarter is the smallest comparison worth acting on. Year over year for the same quarter is better still, because it removes the seasonality that ruins most small business comparisons.
Rolling averages help. A three month rolling count of leads smooths the noise and still shows a real trend when one exists.
There is a published discipline for this from an unexpected place. A 2026 University of St. Gallen paper by Julius Schulte, Malte Bleeker and Philipp Kaufmann worked out how many observations a stochastic system needs before a movement means anything, and reached two rules that transfer directly: take enough measurements that the standard error is smaller than the effect you want to claim, and report on a rolling two to four week aggregate rather than week over week. Their subject is AI search visibility. The arithmetic is the same.
The exception is a total failure. If the form stops delivering, leads go to zero and you see it immediately. That is why a monthly check of the count still matters even though the monthly rate does not. You are watching for a cliff, not a slope.
06What do I do about the report my agency sends?
Read it for the counts and ignore the percentages.
If a report claims a conversion rate improvement from a change made three weeks ago, on a few hundred visits, that claim cannot be supported. It is not always dishonest. Usually nobody did the arithmetic on what their own numbers can carry.
Ask two questions. How many conversions is this percentage based on, and what size of improvement would this sample be able to detect. A straight answer is worth a lot. A vague one tells you what the report is for.
Ask for absolute numbers beside every rate, and a year over year comparison once a quarter.
Be careful with benchmark comparisons too. When a report says your rate is below industry average, ask where the average came from. The most quoted figure, an all industry median of 8.18%, is from LocaliQ and WordStream's 2026 search advertising benchmarks, and it is the median conversion rate of paid search campaigns across thousands of Google Ads and Microsoft Ads accounts. It is a campaign level national median and was never a benchmark for a page. LocaliQ sells advertising services.
The same discipline applies to spend. If a competitor starts bidding on your brand name, your branded traffic and its conversion rate both move, and reading that as a website problem sends you rebuilding a page that was never at fault.
07What to do this week
Start a simple spreadsheet outside your analytics tool. One row per month, four columns: sessions, total inquiries, qualified inquiries, booked jobs. Backfill as far as your records allow.
Run the sample size formula once with your own conversion rate. Write down the number of visitors a realistic test would need and how many months that is at your traffic. Keep that number somewhere you can find it, because it ends every future argument about testing in one line.
Stop reporting a weekly conversion rate to yourself. Delete that dashboard view if you have one.
Set a quarterly review date and put it in the calendar. That is when you compare, not before.
Then spend the time you just freed on observation rather than measurement. Watch sessions, read the questions your last twenty leads asked, and check the first screen of your site on a phone, because what a first-time visitor understands in five seconds is something you can evaluate directly without waiting for data.
If your site has grown into dozens of thinly visited pages, that also makes measurement harder, and it is worth asking how many pages a small business site actually needs before adding more.
Be honest with yourself
When you do not need this
If you have real volume, several thousand conversions a year, you can read monthly rates and you should. This article is written for businesses below that line.
If your traffic is almost entirely paid and steady, ad platform data gives you more usable numbers, though the same caution about small conversion counts still applies.
And if you are not making decisions from the numbers anyway, do not build a measurement habit for its own sake. Count leads monthly, watch for the cliff, and put your attention on the work. That is a legitimate choice, and it is better than a dashboard nobody acts on. If you are unsure whether your inquiries or your visibility is the constraint, the gap between healthy traffic and flat inquiries is the place to start.
Sources
- Kohavi, Deng, Longbotham and Xu, "Seven Rules of Thumb for Web Site Experimenters," KDD 2014. Peer reviewed conference paper from Microsoft and LinkedIn. Source of Twyman's law, the 0.1% to 1.0% effect range, the one in 500 figure, and the provenance of the 10,000 visitor rule. Its examples run to hundreds of thousands or millions of users. Desktop web data, not local business data.
- Patel, Neil, "11 Obvious A/B Tests You Should Try," QuickSprout, 14 January 2013, archived copy. Recorded as reference 29 in the paper above, and the origin of the 10,000 visitor threshold. A blog post, not a study. The live URL refuses programmatic requests, so the archive and the paper's reference list are the record.
- The sample size formula is standard two proportion power analysis, attributable to no vendor. The constant 16 is two times the square of 1.96 plus 0.84. Every figure computed from it here can be rebuilt on paper.
- Schulte, Bleeker and Kaufmann, University of St. Gallen, "Don't Measure Once: Measuring Visibility in AI Search," arXiv:2604.07585, 8 April 2026. Source of the standard error and rolling window discipline. The paper measures AI search visibility, not website conversion, and is cited here for its measurement arithmetic only.
- LocaliQ and WordStream, "Search Advertising Benchmarks for Every Industry," 2026 edition, Stephanie Heitman. Source of the 8.18% figure. Last updated 1 June 2026. These are campaign level national medians for paid search, not page benchmarks. The 2026 sample size is not published; the 2025 edition disclosed 16,446 US campaigns with a minimum of 64 per category. Vendor research. LocaliQ sells advertising services.
Related reading
- Why A/B testing does not work for most local businesses. The same arithmetic aimed at a specific proposal, with the questions to ask a vendor.
- When a metric moves, telling signal from noise. A practical rule for deciding whether last month's dip deserves a meeting.
- Setting a baseline before you change anything. The step that makes before and after comparisons defensible when a controlled test is out of reach.
- Session recordings: what to look for in the first thirty. The measurement that works at any traffic level, because it is observation rather than inference.
Questions about reading your own numbers?
Email me at eric@seod.com with your monthly sessions and your monthly inquiry count for the last three months, however rough. I will run the sample size formula with your actual conversion rate, tell you which of those movements are real and which are noise, and set the measurement window that would give you a readable answer at your volume.
One email, no call, and it takes me a few minutes. If your numbers are large enough to read monthly, I will tell you that and you can ignore everything above.
There is more on conversion in the library.