A confidence interval is a range around a measured figure that shows how far the true value could plausibly be from what you observed, given how much data you have. If a landing page converted 50 of 1,250 visitors, the measured rate is 4.0%, but the 95% confidence interval runs from roughly 2.9% to 5.1%: with a sample that size, any true rate in that range could easily have produced what you saw.
How a confidence interval works
Every conversion rate, click-through rate or average order value you look at is measured on a sample, namely the visitors who happened to arrive during that period. Another month would give a slightly different figure. The interval puts a margin of error on the number, and its width depends on how many observations you have and how much they vary.
The confidence level, usually 95%, describes the method rather than this one result. If you ran the same test many times and built an interval each time, about 95% of those intervals would contain the true value. In everyday use, treat it as the range of outcomes that should not surprise you.
More data narrows the range, but slowly. The margin shrinks with the square root of the sample size, so you need roughly four times as many visitors to halve the width of the interval. That is why smaller sites find it so hard to prove small improvements.
Why it matters
In an A/B test, the interval that matters is the one around the difference between the two versions. Suppose version A converts 80 of 2,000 visitors (4.0%) and version B converts 96 of 2,000 (4.8%). B looks 20% better, but the 95% interval for the difference runs from about −0.5 to +2.1 percentage points. Because that range includes zero, the test has not shown B is better at all; it could plausibly be slightly worse.
Reading results this way protects your budget. A UK business with modest traffic will regularly see differences that look large and are mostly noise, and rolling out a new page, ad or offer on the strength of a headline percentage is how teams end up “improving” things that make no difference. The interval also shows how big the effect could realistically be, which a single p-value or a “95% significant” badge does not.
Common mistakes
- Reporting a single figure. A conversion rate quoted with no range implies a precision the data does not have.
- Checking whether two separate intervals overlap. Overlapping intervals do not prove there is no difference. Work out the interval for the difference itself, or use a calculator that does.
- Stopping when it first looks good. Checking repeatedly and stopping the moment the interval clears zero, known as peeking, makes false wins far more likely. Fix the sample size before you start.
- Treating 95% as a promise. Roughly one in twenty tests of a change that does nothing will still produce an interval that excludes zero, purely by chance.
- Forgetting the business case. An interval from +0.1 to +0.3 percentage points may be real and still not worth the cost of the change.
How to act on it
Before a test, use your expected traffic and current conversion rate to estimate how long it must run to detect the smallest change worth having. After it, enter the numbers in the A/B test significance calculator and look at the range as well as the verdict. If the interval includes zero, you have not found a winner. If it sits entirely above zero but its lower end is too small to matter, think twice before rebuilding anything.
For paid traffic, landing pages are often where this discipline pays for itself, because every visitor has a cost. A landing page built for an ad campaign can be tested in clearly different versions, with the sample size planned before launch rather than guessed afterwards.
