Detecting a small improvement needs thousands of visitors per version. Below a few hundred conversions a month, most tests will not conclude, and judgement plus sequential changes is the honest alternative.

The uncomfortable arithmetic

Split testing is presented as available to everybody, and the mathematics underneath it is indifferent to how much you would like an answer.

Detecting a difference requires enough events that the difference stands out from ordinary random variation.

The smaller the difference you are hoping to find, the more events you need, and the relationship is steep rather than proportional.

Halving the size of the effect you want to detect roughly quadruples the traffic required.

Rough numbers

Precise figures depend on your starting conversion rate and how confident you want to be, and the shape is what matters.

To detect a large improvement, something like a half increase in conversions, you might need a few hundred conversions per version.

To detect a modest one, in the region of a fifth, you are into the low thousands of visitors per version.

To detect a small one, a few percent, you need tens of thousands per version, which is simply outside what most business sites see in a year.

Online calculators will give you a specific figure in a minute, and running one before a test rather than after is the entire point.

What this means for a typical site

Most small business sites are in the first two rows, where the honest conclusion is that formal testing will rarely conclude anything.

Stopping early is the common failure

More damaging than not testing at all, because it produces confident wrong answers.

A test running for a week shows version B ahead by a wide margin. It looks decisive and the temptation to stop is considerable.

Early in a test, results swing dramatically on very few events, and a lead that looks large is frequently noise that has not yet averaged out.

Checking repeatedly and stopping when you like the answer converts a test into a search for a favourable moment, which will always eventually arrive.

Decide the sample size and duration before starting, and do not stop early because it looks good.

Run for whole weeks

A small rule that removes a large distortion.

Behaviour differs substantially by day. A business audience converts on weekdays; a consumer one may do better at weekends.

A test running ten days includes one weekend and part of another, which weights the result by whichever days happened to be included.

Run in multiples of seven days, and preferably at least two full weeks, so both versions see the same mix.

Avoid periods containing a holiday, a campaign, or anything else unrepresentative, since a test cannot separate your change from a bank holiday.

A worked example

A supplier with about four hundred visits a month to their quote page wanted to test two versions.

Their conversion rate was around six percent, meaning roughly twenty-four enquiries a month across both versions.

A calculator indicated that detecting a fifth improvement would need several thousand visitors per version.

At their traffic that was somewhere over two years, by which point the business and the market would both have changed.

They stopped planning the test and did something else: made the changes they were reasonably confident about, one at a time, and tracked enquiries quarterly against the previous year.

That is not a controlled experiment and they did not describe it as one. It was the honest option available at their volume, and it produced decisions rather than a test that would never have concluded.

What to do instead at low volume

Testing is not the only way to improve a page, and at small scale it is not the best one.

Watch five people use the site. Usability problems are visible with a handful of observers and need no statistics at all, because you are seeing a fault rather than measuring a rate.

Fix the things that are plainly wrong: broken forms, buried phone numbers, unreadable text, missing prices. None of these needs proof.

Ask the people who did not buy what stopped them, which produces specific reasons rather than percentages.

And make changes on judgement, in sequence, recording what you did, comparing over quarters rather than weeks.

Test bigger things

The other adjustment available when traffic is limited.

If you can only detect large differences, test changes capable of producing large differences.

A different button colour will not move anything detectably at your volume. Publishing prices where you previously published none might. Adding a phone option to a form-only page might. Rewriting a page around a completely different proposition might.

Small refinements are for sites with the traffic to see them. At low volume, spend your testing capacity on questions where the answer could be substantial.

The counter-case

None of this means small sites should ignore measurement.

Large effects are detectable at modest volumes, and a change that doubles enquiries will be visible without any calculator.

Testing can also be worthwhile on a page with disproportionate traffic even when the site is small, since one popular page may carry most of the visits.

And there is value in the discipline itself. A business that decides in advance what it expects, and checks afterwards, makes better decisions than one that does neither, whether or not any individual result reaches statistical certainty.

The error is not testing at low volume. It is running an underpowered test, reading a result, and treating it as established.

Before running a test

  1. Count conversions per month, not visits.
  2. Use a calculator to find the sample needed.
  3. Divide by your monthly rate to get the duration.
  4. If it exceeds two months, do not run it.
  5. Fix the obvious things instead, without testing.
  6. Run in whole weeks and do not stop early.
  7. Test large changes, not refinements.

Step three takes two minutes and prevents most of the wasted testing small businesses do.

Not having enough data to learn anything is covered in a budget too small to learn anything.


Frequently asked questions

How much traffic do I need to run a split test?

It depends on the size of difference you want to detect. A large improvement might need a few hundred conversions per version; a few percent needs tens of thousands of visitors.

Why does detecting small differences need so much traffic?

Because the requirement rises steeply. Halving the size of the effect you want to find roughly quadruples the traffic needed to find it.

Can I stop a test when one version is clearly ahead?

No. Early results swing on very few events, and checking repeatedly until you like the answer turns a test into a search for a favourable moment.

How long should a test run?

In multiples of seven days, at least two full weeks, so both versions see the same mix of weekdays and weekends. Avoid holidays and campaign periods.

What should I do instead at low volume?

Watch five people use the site, fix what is plainly broken, ask people who did not buy what stopped them, and make changes in sequence comparing quarterly.

Is testing pointless for a small site?

No. Large effects are detectable at modest volumes, and the discipline of predicting then checking improves decisions. The error is treating an underpowered result as established.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Planning a split test?

Count monthly conversions, run a calculator, divide. If the answer is over two months, do something else.

Start a Conversation