Skip to main content

WEBSITE CONVERSION · September 2026 · ~11 min read

Why A/B testing does not work for most local businesses

Sample size depends on your conversion rate and on how small a change you want to catch, not on a visitor count. At 3%, a 20% improvement takes about 12,900 sessions per variant (derived). A local business at 300 to 800 visits a month is years away. Split tests at that volume are theatre.

This is the least popular sentence in conversion optimization, because the entire category is built on the idea that testing is what serious people do. Testing is what serious people with traffic do. Those are different businesses.

I am not arguing that testing is a scam in general. At real volume it is the best tool available. I am arguing that most local service businesses, restaurants, clinics, and trades do not have the traffic to use it, and that selling it to them anyway is a way to bill monthly for a number that means nothing.

01

How much traffic does a valid A/B test actually need?

Two things decide it, and neither is a visitor count: your conversion rate, which sets the variance, and the size of the improvement you want to see. That is the framing in Kohavi, Deng, Longbotham and Xu's KDD 2014 paper, written from thousands of controlled experiments at Microsoft, LinkedIn, Amazon and Booking.com: "Formulas for minimum sample size given the metric's variance and sensitivity... provide one lower bound."

The formula for a conversion rate, with every part visible: sessions per variant = 16 × p × (1 - p) ÷ d², where p is your current rate and d the absolute improvement you want to catch. The 16 is 2(1.96 + 0.84)² rounded up, 1.96 for 95% confidence and 0.84 for 80% power. At 3%, catching a 20% relative improvement means d = 0.006 and about 12,900 sessions per variant (derived).

Divide by the conversion rate and almost everything cancels, leaving the version to memorise: conversions per variant is roughly 16 ÷ r², where r is the relative lift. A 20% lift costs about 390 conversions per variant (derived). Backwards, the familiar hundred conversions per variant detects a lift of about 39%derived and nothing smaller. Nothing you change on a page moves conversion by two fifths.

The 10,000 visitors figure is not a research finding. It appears as reference [29] in that same paper, resolving to Neil Patel, "11 Obvious A/B Tests You Should Try," QuickSprout, 14 January 2013. The authors cite it and qualify it in the same sentence: "the guidance should be refined to the metrics of interest." A blog post with no sample and no method, promoted to a threshold by thirteen years of repetition.

At 500 monthly visits and 3%, that is seven or eight leads per variant a month, so 390 conversions take about 52 months (derived).

A test that takes four years to answer one question is not a test. It is a subscription.

And nothing holds still for four years. Season, ad spend, prices, a competitor opening, Google changing what it sends you. By the time the number arrives it is measuring a business that no longer exists.

02

What does the math look like with my own numbers?

Here is the calculation, in one table you can rebuild on paper. Monthly sessions times conversion rate gives leads. Halve it for a two-way split. Divide 390 by that figure, the per variant cost of a 20% lift (derived). That is months to a single answer.

Monthly sessionsConversion rateLeads a monthPer variantMonths to see a 20% lift (derived)
3003%94.587
5003%157.552
8003%241233
8008.18%653311
10,0003%300150under 3

That fourth row is worth sitting with. The 8.18% is the median conversion rate across all industries in LocaliQ and WordStream's 2026 search advertising benchmarks. It is a paid search rate, which runs higher than a typical site-wide rate, and LocaliQ sells advertising services, so treat it as a generous ceiling. Even at that rate, 800 sessions a month buys you one answered question a year.

Now add what the table leaves out. Three variants instead of two stretches the clock by half again. An inconclusive test, the most common outcome, costs the full window and returns nothing. And you rarely get a clean site for the next test, because you changed something else meanwhile.

Write your own row down. If the last column is over three, split testing is not a tool you own this year.

03

What happens when you run a test anyway?

You get a result. That is the trap.

At low volume, normal week-to-week variation is larger than any difference the test is trying to find. One corporate catering inquiry, one bad weather week, one holiday, and version B is ahead. None of that has anything to do with the button color.

So the report says version B is winning, you implement version B, and next month goes back to normal because there was never a difference. Then you are told the gain was offset by seasonality. That explanation is unfalsifiable, which is why it is used.

The measurement discipline that would prevent this is well described, just not in conversion optimization. A 2026 University of St. Gallen paper on measuring visibility in AI search worked out how many observations a stochastic system needs before a movement means anything, and landed on two rules that transfer directly: take enough runs that the standard error is smaller than the effect you are claiming, and report on a rolling two to four week aggregate rather than week over week. Its blunt version, applied to its own subject, is that a score moving from 40% to 45% in a week is statistically indistinguishable from no change. Different domain, identical arithmetic. If the movement is smaller than the noise, the movement is not evidence.

The real cost is attention. Every month spent watching a meaningless comparison is a month not spent on changes that obviously need making. Nobody needs a split test to know the phone number should be tappable, or that mobile behavior your desktop browser never shows you is where the breakage lives.

04

Why do the tools still show a winner?

Because most testing tools will happily declare significance on tiny samples if you let them, and because people stop tests when they like the answer.

Watch a small test day by day and one version is always ahead. It flips constantly. If you end the test at the moment your preferred version is up, you have not measured anything. You have measured your own timing. Statisticians call this optional stopping, and it is the single easiest way to manufacture a result that will not repeat.

There is a second mechanism, and the formula contains it. Sample size scales with the inverse square of the effect, so halving the effect quadruples the sample. Catching a doubling of your rate needs about 16 conversions per variant (derived). Catching a 5% relative improvement needs about 6,200 derived. Real button and headline changes live at the small end, which is why "we will test the copy" is arithmetic nobody checked.

There is also a selection effect in who talks about testing. Case studies come from companies with enormous traffic, because those are the only places a test resolves. The advice gets written for them, then repeated down the chain until a dentist with 400 monthly visits is pitched a testing program.

Ask any vendor pitching monthly testing one question. How many conversions per variant will this test need, and how many months is that at my volume. A straight answer ends the conversation honestly. An answer about learnings and iteration tells you what you needed to know.

05

What should I do instead of testing?

Three things, all of which produce answers in days rather than quarters.

Watch real sessions. Thirty recordings of actual visitors will show you dead taps, rage clicks, confusion at a specific field, and the exact point people leave. What to look for in the first thirty recordings is a short list, and none of it requires a sample size.

Analyze form abandonment. Where people start and stop tells you which field is doing damage. That is a direct observation, not a statistical inference, so small numbers are still informative. The one directional principle worth importing here comes from the Baymard Institute, which has tracked checkout flows since 2012 and concluded that the number of form fields matters far more to performance than the number of steps. Baymard sells UX research and its data is ecommerce checkout, not a plumber's quote form, so use the direction and not a number.

Read what your actual leads asked. Pull the last twenty inquiries and calls and write down every question that came up. Every recurring question is something your website failed to answer. That list is the highest quality conversion research a small business can get, and it is free.

Then make obvious changes and watch direction over quarters, not weeks. Not proof. Direction. Wider measurement windows are the price of small numbers, and comparing this quarter to last quarter is more honest than comparing this Tuesday to last Tuesday.

It also helps to stop treating every page the same way. A homepage and a landing page have different jobs, and separating them gives you a page you can change freely without risking the one that carries your search traffic.

06

What to do this week

Work out your own numbers before anyone quotes you. Monthly visits times conversion rate gives monthly leads. Halve it. Divide 390 by that, or 16 ÷ r² for a different target lift r. That is how many months one test would take.

Write that number down. If it is over three, testing is not available to you this year, and you can stop feeling behind about it.

Then spend the time you just freed on the three alternatives above. Start with the twenty most recent inquiries, because it takes an hour and it usually rewrites your homepage for you.

If you are choosing a platform right now, keep this in mind rather than paying for testing features you cannot use. The honest comparison between website builders and a custom build rarely turns on testing tools at this size.

Be honest with yourself

When you do not need this

If your traffic runs into tens of thousands of sessions a month, ignore everything here and test properly. Multi-location brands, ecommerce, and busy franchise groups clear the arithmetic and should use it.

If you run high-volume paid campaigns, ad-level testing is a different exercise and it can resolve faster, because impressions are plentiful even when site traffic is not. Some channels sidestep the question entirely, since Local Services Ads work on a different model than a landing page test.

If you run several locations on one website, you can pool traffic across them and reach a usable sample that no single location would produce. The same is true of an email or SMS list, where sends are plentiful and subject line tests resolve in days.

And there is one case where a before and after genuinely reads at small volume: when the change is total rather than incremental. A form that never delivered, a number that never dialed, a page that did not load on iOS. Going from zero to something is visible immediately, because you are not measuring a lift, you are measuring a repair.

If your website is obviously broken, do not test it. Fix it. Testing a broken page against a slightly less broken page is a slow way to arrive where you already knew you were going.

Sources

Related reading

10

Questions about a testing proposal?

Email me at eric@seod.com with the split testing proposal you were quoted and your rough monthly visits. I will calculate how many months that test would need to reach significance at your traffic level, and send you the number in one line.

If it clears, I will tell you it clears and you should sign it. Most of the time it does not, and knowing that before you commit to a retainer is worth the email.

There is more on conversion here.

Call Eric Email Eric