Analytics and Tracking

Confidence Interval

The range a true figure, such as a conversion rate, probably sits within, given how much data the measurement is based on.

Quick facts: Confidence Interval

Category
Analytics and Tracking
Level
Advanced
Affects
A/B test decisions, conversion rate reporting, budget decisions
Where to see it
A/B test significance calculators, Google Ads experiments, spreadsheets
In this article4
  1. How a confidence interval works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

A confidence interval is a range around a measured figure that shows how far the true value could plausibly be from what you observed, given how much data you have. If a landing page converted 50 of 1,250 visitors, the measured rate is 4.0%, but the 95% confidence interval runs from roughly 2.9% to 5.1%: with a sample that size, any true rate in that range could easily have produced what you saw.

How a confidence interval works

Every conversion rate, click-through rate or average order value you look at is measured on a sample, namely the visitors who happened to arrive during that period. Another month would give a slightly different figure. The interval puts a margin of error on the number, and its width depends on how many observations you have and how much they vary.

The confidence level, usually 95%, describes the method rather than this one result. If you ran the same test many times and built an interval each time, about 95% of those intervals would contain the true value. In everyday use, treat it as the range of outcomes that should not surprise you.

More data narrows the range, but slowly. The margin shrinks with the square root of the sample size, so you need roughly four times as many visitors to halve the width of the interval. That is why smaller sites find it so hard to prove small improvements.

Why it matters

In an A/B test, the interval that matters is the one around the difference between the two versions. Suppose version A converts 80 of 2,000 visitors (4.0%) and version B converts 96 of 2,000 (4.8%). B looks 20% better, but the 95% interval for the difference runs from about −0.5 to +2.1 percentage points. Because that range includes zero, the test has not shown B is better at all; it could plausibly be slightly worse.

Reading results this way protects your budget. A UK business with modest traffic will regularly see differences that look large and are mostly noise, and rolling out a new page, ad or offer on the strength of a headline percentage is how teams end up “improving” things that make no difference. The interval also shows how big the effect could realistically be, which a single p-value or a “95% significant” badge does not.

Common mistakes

  • Reporting a single figure. A conversion rate quoted with no range implies a precision the data does not have.
  • Checking whether two separate intervals overlap. Overlapping intervals do not prove there is no difference. Work out the interval for the difference itself, or use a calculator that does.
  • Stopping when it first looks good. Checking repeatedly and stopping the moment the interval clears zero, known as peeking, makes false wins far more likely. Fix the sample size before you start.
  • Treating 95% as a promise. Roughly one in twenty tests of a change that does nothing will still produce an interval that excludes zero, purely by chance.
  • Forgetting the business case. An interval from +0.1 to +0.3 percentage points may be real and still not worth the cost of the change.

How to act on it

Before a test, use your expected traffic and current conversion rate to estimate how long it must run to detect the smallest change worth having. After it, enter the numbers in the A/B test significance calculator and look at the range as well as the verdict. If the interval includes zero, you have not found a winner. If it sits entirely above zero but its lower end is too small to matter, think twice before rebuilding anything.

For paid traffic, landing pages are often where this discipline pays for itself, because every visitor has a cost. A landing page built for an ad campaign can be tested in clearly different versions, with the sample size planned before launch rather than guessed afterwards.

Do and do not

Do

  • Report a range alongside every test result
  • Plan the sample size before the test starts
  • Ask whether the low end of the range is worth acting on

Do not

  • Stop a test the first time it looks significant
  • Compare two separate intervals for overlap
  • Treat 95% confidence as certainty

Questions people ask about this

How is a confidence interval related to statistical significance?

They come from the same calculation. A result is significant at the 5% level when the 95% confidence interval for the difference does not include zero. The interval tells you more, because it also shows how large or small the true effect could be. See statistical significance for the other half of the picture.

Should I use 90% or 95% confidence?

Use 95% when a wrong decision would be expensive or hard to reverse, such as a site redesign. A 90% level gives a narrower interval and reaches a verdict sooner, at the price of more false wins, which can be acceptable for cheap changes you can easily undo, such as ad copy. Choose the level before the test starts, never after seeing the data.

Can I work out a confidence interval in a spreadsheet?

Yes, for a conversion rate. Work out the rate p, then the standard error as the square root of p × (1 − p) ÷ n, where n is the number of visitors, and add and subtract 1.96 times that figure for a 95% interval. This shortcut becomes unreliable when there are very few conversions, roughly under ten, so use a dedicated calculator then.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.