Free tools

A/B test significance calculator

Enter visitors and conversions for two versions of a page, ad or email and see whether the difference is likely to be real or just chance. For UK marketers who want to call a test on evidence, with the p-value and confidence interval explained in plain English.

How to use the calculator

  1. Take visitors and conversions for both versions from the same date range and the same source: the testing tool’s own report, or GA4 filtered to the test.
  2. Enter version A, the page or ad as it is now, and version B, the one with the change.
  3. Choose a confidence level. 95% is the usual choice. 90% is reasonable for a cheap change you can reverse in a minute; 99% is worth the extra traffic when the change is expensive to roll out or hard to undo.
  4. Run the check and read the verdict together with the notes underneath it, never on its own.

Visitors can be sessions, users or ad clicks, as long as both versions use the same unit. Conversions should be the action you actually care about, such as an enquiry form sent, a call or an order, rather than a click on the button that leads to it.

Reading the result

The verdict tells you whether the gap between the two conversion rates is big enough, for the traffic you have, that chance is an unlikely explanation. That is all statistical significance means. It says nothing about whether the gap is large enough to matter to your business.

  • Conversion rate A and B: conversions divided by visitors for each version.
  • Relative change: how far B’s rate sits above or below A’s, as a share of A. A move from 2.0% to 2.5% shows as +25%.
  • p-value: how often a gap at least this big would appear if both versions really converted at the same rate. Below 0.05 passes at 95% confidence. A p-value is not the probability that B is better, which is the most common misreading I come across.
  • Confidence interval: the range of plausible values for the real difference, in percentage points (B’s rate minus A’s). When the confidence interval runs from below zero to above it, the data still allows B to be worse.

The figures already in the boxes make a useful example. 120 conversions from 5,000 visitors against 150 from 5,000 is 2.40% against 3.00%, a 25% relative lift that looks decisive in a monthly report. The p-value is about 0.064, so it passes at 90% and fails at 95%, and the 95% interval runs from just below zero to about 1.2 points. My reading: B is probably better, the size of the gain is very uncertain, and I would not rebuild three other pages on the strength of it.

Fix the end date before you start

The calculation assumes you look once, at the end. Checking every morning and stopping the first day the verdict turns significant is known as peeking, and it produces false winners far more often than the confidence level suggests, because every look gives ordinary daily noise another chance to cross the line.

Before a test goes live I write down three things: the smallest improvement worth acting on (the minimum detectable effect), the sample size that improvement needs, and the date the test ends. I run whole weeks, since weekday and weekend behaviour often differ, and I keep tests away from bank holidays, Black Friday or a large email send unless that period is what is being tested.

When a split test is not worth running

Traffic decides more than ideas do. A page with a few hundred visitors and a handful of enquiries a month can only detect an enormous difference in any sensible timeframe, which is why the calculator warns when there are fewer than 100 conversions in total. On a site like that I would rather fix what a landing page checklist turns up (a slow page, a vague headline, a form that asks too much) and judge the change over a longer period.

Where there is enough traffic, test changes big enough to change behaviour: the offer, the promise in the headline, the length of the form, the page an ad sends people to. Button colours rarely earn the traffic they use up.

What this calculator cannot tell you

  • More than two versions. Comparing several variants with the control one at a time multiplies the chance of a false winner. Test fewer versions, or use a tool that corrects for multiple comparisons.
  • Value. It counts conversions, not what they are worth. A version that wins more enquiries of lower quality, or more orders at a smaller basket, can still lose money.
  • Bad data. If conversion tracking double-counts or misses events on one version, the maths is right and the answer is wrong. Check tracking before the test starts, not after it ends.
  • Platform tests. Google Ads campaign experiments and Meta’s A/B test feature split the audience and report confidence themselves. This page is useful as a second opinion on their exported numbers, and for landing page tests run outside the ad platforms.
  • What happens after rollout. An early win can fade once returning visitors stop reacting to something new, so keep watching the winner for a few weeks after it goes live.

Next step

When tests keep coming back inconclusive, the page is often the problem rather than the variant. I build and test landing pages for Google Ads and Facebook campaigns, and I am glad to look over a test plan or a result before you act on it: send me the numbers and both versions.

Frequently asked questions

How many visitors does an A/B test need?

It depends on your current conversion rate and the smallest improvement you care about. The lower the rate and the smaller the change, the more traffic you need, and halving the change you want to detect roughly quadruples the sample. Work the sample size out before the test starts; this calculator checks the result once it ends.

What should I do with a test that is not significant?

Decide whether the difference would matter even if it were real. If it would not, keep whichever version is simpler or cheaper to maintain and move on to a bolder idea. Extending a test again and again until it crosses the line is just peeking with extra steps. My glossary entry on how A/B testing works covers how to plan the next one.

Can I use this for email subject line tests?

Yes. Use emails delivered as visitors and clicks or orders as conversions for each version. I would avoid judging on opens, because Apple Mail Privacy Protection loads images automatically and inflates open counts, so clicks are the more trustworthy signal.

Ready to talk about your project?

A straight answer about what would move the numbers, and a written proposal if we are a fit.