Conversion and UX

Statistical Significance

Also called significance level, confidence level, stat sig

A test result is statistically significant when the difference is unlikely to be explained by random chance alone.

Quick facts: Statistical Significance

Category
Conversion and UX
Also called
significance level, confidence level, stat sig
Level
Intermediate
Affects
A/B test decisions, landing page changes, ad experiments, budget allocation
Where to see it
A/B test significance calculators, GA4, Google Ads experiments, Meta A/B tests, testing platforms such as VWO or Optimizely
In this article4
  1. How statistical significance works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

Statistical significance is a way of judging whether a difference you have measured, such as one landing page converting better than another, is likely to be real or could easily have happened by chance. A result is called significant when, given the amount of data collected, chance alone would rarely produce a gap that large.

How statistical significance works

In an A/B test you start from the assumption that the two versions perform the same, and then ask how surprising your data would be if that were true. The answer is expressed as a p-value. A p-value of 0.03 means that if there were truly no difference, you would see a gap at least this big about 3% of the time. Most marketing tests set the threshold at 0.05 before they start, which is what tools mean when they say “95% confidence”.

A worked example shows why volume matters. Suppose a version of a booking page gets 78 enquiries from 2,000 visitors (3.9%) while the original gets 60 from 2,000 (3.0%). That looks like a 30% improvement. Run a standard two-proportion test and the p-value comes out at roughly 0.12, so the result is not significant: a gap that size turns up fairly often between two identical pages. With ten times the traffic and the same rates, the result would be overwhelmingly significant.

Three ideas sit alongside significance and are worth deciding before a test begins:

  • Sample size: how many visitors or conversions each version needs.
  • Minimum detectable effect: the smallest improvement worth finding, which drives the sample size.
  • Statistical power: the chance the test detects a real effect of that size, commonly set at 80%.

Why it matters

Many UK small business websites have modest traffic, which means tests often run out of data long before they can prove anything. Without a significance check it is easy to declare a winner on a few dozen conversions, roll out a change that does nothing, and credit it for a rise that was really seasonal. The same applies to Google Ads and Meta experiments, where the platforms report confidence levels that deserve the same scrutiny.

Significance also protects budgets. A decision to rebuild every service page around a “winning” layout should rest on evidence that would survive a second test, not on one lucky fortnight.

Common mistakes

  • Stopping the moment the tool shows 95%. Checking repeatedly and stopping at the first good reading, known as peeking, inflates the false positive rate well beyond 5%.
  • Reading 95% confidence as “95% chance the variant is better”. That is not what the number means.
  • Confusing significant with important. A tiny but real lift may not repay the cost of the change.
  • Testing many variants at once and celebrating whichever crosses the line, without adjusting for the extra comparisons.
  • Running partial weeks, so weekday and weekend visitors are unevenly represented.

How to act on it

Before launching a test, write down the metric, the smallest lift you care about and the sample size that requires, then fix the duration in whole weeks. Let it run to the end before reading the result. When it finishes, check the numbers in an A/B test significance calculator and look at the confidence interval as well as the headline.

If your traffic is too low to reach significance in a sensible time, test bolder changes, test higher up the funnel where there are more events, or use qualitative research instead. Deciding what is worth testing and how to judge it is a regular part of my performance marketing work.

Do and do not

Do

  • Calculate the sample size and fix the test length before you start
  • Run tests in whole weeks and read the result only at the end
  • Report the confidence interval alongside the headline result

Do not

  • Stop a test the moment it first shows 95% confidence
  • Treat a significant result as proof that the change is worth the cost
  • Run many variants and pick the winner without adjusting for the extra comparisons

Questions people ask about this

Does 95% significance mean my new version is 95% likely to be better?

No. It means that if the two versions really performed the same, a difference as large as the one you saw would occur less than 5% of the time. It says nothing directly about how likely the variant is to be better, or by how much. The confidence interval around the difference is a more useful guide to the likely size of the effect.

How long should an A/B test run?

Long enough to reach the sample size you calculated in advance, and always in whole weeks so every day of the week is represented. For most small business sites that means at least two weeks, often longer. Ending early because the numbers look good is the most common way tests produce false winners.

Can I test with low website traffic?

You can, but you need to adjust your expectations. Test large, meaningful changes rather than button colours, measure an earlier step with more volume such as clicks to the form, and accept that some questions are better answered by watching five real users than by a split test that would take a year.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.