Statistical significance is a way of judging whether a difference you have measured, such as one landing page converting better than another, is likely to be real or could easily have happened by chance. A result is called significant when, given the amount of data collected, chance alone would rarely produce a gap that large.
How statistical significance works
In an A/B test you start from the assumption that the two versions perform the same, and then ask how surprising your data would be if that were true. The answer is expressed as a p-value. A p-value of 0.03 means that if there were truly no difference, you would see a gap at least this big about 3% of the time. Most marketing tests set the threshold at 0.05 before they start, which is what tools mean when they say “95% confidence”.
A worked example shows why volume matters. Suppose a version of a booking page gets 78 enquiries from 2,000 visitors (3.9%) while the original gets 60 from 2,000 (3.0%). That looks like a 30% improvement. Run a standard two-proportion test and the p-value comes out at roughly 0.12, so the result is not significant: a gap that size turns up fairly often between two identical pages. With ten times the traffic and the same rates, the result would be overwhelmingly significant.
Three ideas sit alongside significance and are worth deciding before a test begins:
- Sample size: how many visitors or conversions each version needs.
- Minimum detectable effect: the smallest improvement worth finding, which drives the sample size.
- Statistical power: the chance the test detects a real effect of that size, commonly set at 80%.
Why it matters
Many UK small business websites have modest traffic, which means tests often run out of data long before they can prove anything. Without a significance check it is easy to declare a winner on a few dozen conversions, roll out a change that does nothing, and credit it for a rise that was really seasonal. The same applies to Google Ads and Meta experiments, where the platforms report confidence levels that deserve the same scrutiny.
Significance also protects budgets. A decision to rebuild every service page around a “winning” layout should rest on evidence that would survive a second test, not on one lucky fortnight.
Common mistakes
- Stopping the moment the tool shows 95%. Checking repeatedly and stopping at the first good reading, known as peeking, inflates the false positive rate well beyond 5%.
- Reading 95% confidence as “95% chance the variant is better”. That is not what the number means.
- Confusing significant with important. A tiny but real lift may not repay the cost of the change.
- Testing many variants at once and celebrating whichever crosses the line, without adjusting for the extra comparisons.
- Running partial weeks, so weekday and weekend visitors are unevenly represented.
How to act on it
Before launching a test, write down the metric, the smallest lift you care about and the sample size that requires, then fix the duration in whole weeks. Let it run to the end before reading the result. When it finishes, check the numbers in an A/B test significance calculator and look at the confidence interval as well as the headline.
If your traffic is too low to reach significance in a sensible time, test bolder changes, test higher up the funnel where there are more events, or use qualitative research instead. Deciding what is worth testing and how to judge it is a regular part of my performance marketing work.
