Sample size is the number of people, sessions or other observations included in a test or analysis. In marketing it usually means how many visitors each version of an A/B test needs before you can trust that a difference between them is real rather than luck.
How sample size works
The sample a test needs depends on four inputs, and you should fix all four before the test starts:
- Baseline conversion rate. How often the current page converts. Low rates need bigger samples, because conversions are rare events.
- Minimum detectable effect. The smallest improvement worth detecting. Small effects need very large samples.
- Significance level. How strict you are about false winners. A 95% confidence level is the common convention.
- Power. How likely the test is to spot a real effect of that size. 80% is the usual convention.
A worked example using the standard formula for comparing two conversion rates: if a landing page converts at 3% and you want to detect a rise to 3.6% (a 20% relative improvement), you need roughly 13,900 visitors per version, about 27,800 in total. If the page converts at 10% and you look for the same 20% lift, to 12%, the requirement falls to roughly 3,800 per version. Halving the effect you want to detect roughly quadruples the sample.
The count that really matters is conversions, not visitors. A test with 10,000 visitors and 30 enquiries on each side is a small test, whatever the visitor number suggests.
Why it matters
Most UK small business websites have less traffic than testing articles assume. A service page with 3,000 visits a month, converting at 3%, would take around nine months to reach the 27,800 visitors in the example above. If that traffic came from paid search at £2 a click, buying it would cost over £55,000.
Tests that stop short produce false winners. Early in any A/B test the numbers swing wildly, and one version often looks far ahead after a few days purely by chance. Roll out that “winner” and the improvement quietly disappears, or the change makes things worse. The same applies to ad platforms that declare a winning ad after a handful of clicks.
Common mistakes
- Starting a test without deciding the sample size, then stopping the moment the testing tool reports statistical significance. Checking repeatedly and stopping at the first good reading inflates false positives.
- Splitting limited traffic across four or five variants, which multiplies the sample needed.
- Ending a test mid-week. Behaviour on a Monday morning differs from a Saturday evening, so run whole weeks.
- Slicing the finished result into small segments, such as mobile users in Scotland, and treating each slice as its own finding.
- Counting sessions when the decision is about people. One person visiting five times is not five independent observations.
How to act on it
Before you build a variant, put your baseline rate and the smallest lift you care about into a sample size calculation, then divide by your weekly traffic to see how long the test would take. Plan for at least two full weeks, and longer around bank holidays or seasonal peaks.
If the answer is many months, change the plan rather than shortening the test. Test bolder changes that could produce a big effect, test on your highest-traffic page, or skip formal testing and make the change on the strength of customer research and usability sessions, then watch the trend. When the test ends, check the result in my free A/B test significance calculator before rolling anything out.
Planning tests that your traffic can support is part of how I design and improve landing pages for Google and Facebook ads.
