Analytics and Tracking

Peeking

Checking an A/B test repeatedly before it reaches its planned sample size and stopping as soon as the result looks significant.

Quick facts: Peeking

Category
Analytics and Tracking
Level
Intermediate
Affects
A/B test results, landing page decisions, conversion rate work
Where to see it
A/B testing platforms, significance calculators, sample size calculators
In this article4
  1. How peeking works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

Peeking is the habit of checking an A/B test while it is still running and stopping it the moment one version looks like a significant winner. It feels efficient, but it quietly breaks the statistics behind the test, so many of the “winners” you declare will be nothing more than chance.

How peeking works

Most simple test calculators use a fixed-horizon method. You decide the sample size in advance, run the test until you reach it, then check the result once. The familiar 95% confidence threshold means that, if there is really no difference between the versions, you will wrongly declare a winner about one time in twenty.

That one-in-twenty promise only holds if you look once. Results in a live test wander up and down as visitors arrive, especially in the early days when numbers are small. Every time you check, you give that random wandering another chance to cross the significance line. Check daily for three weeks and stop at the first crossing, and your real chance of a false winner is far higher than the 5% you think you are accepting.

Picture a dental practice testing two versions of its booking page. On day three, version B is ahead by 42 bookings to 22 and the tool flags it as significant. The test is stopped and B goes live. Over the next two months bookings settle back to where they were, because the early gap was a lucky run rather than a real effect. Nothing about that run was unusual; with small daily numbers, swings like it are common.

Why it matters

UK small and medium-sized sites rarely have the traffic for fast tests, which makes the temptation to stop early stronger. The cost is not only a wasted test. False winners get built into the site, the team learns the wrong lessons about what customers respond to, and later tests are designed on top of a mistaken result.

It also damages trust in testing itself. When a celebrated winner fails to move real revenue, people conclude that testing does not work, when the problem was how the result was read.

Common mistakes

  • No sample size set in advance. Without one, there is no defined end point, so every look becomes a decision point.
  • Stopping on a significant result but continuing on a non-significant one. This asymmetry is the core of the problem.
  • Running for less than a full week. Behaviour on a Saturday differs from a Tuesday. Short tests capture only part of the weekly cycle.
  • Trusting a tool’s banner without knowing its method. Some testing tools use sequential or Bayesian statistics designed for continuous monitoring; others do not. Check which yours uses before treating its live significance figure as final.

How to act on it

Before the test starts, write down the smallest uplift worth detecting, the sample size per variant and the minimum run time in whole weeks. Agree with everyone involved that the result is read once, at that point, unless something is broken.

It is fine to look during the test for problems: a variant not loading, tracking failing, a sharp fall in conversions that suggests a bug. Looking for errors is not peeking; acting on an apparent win is. When the planned sample is reached, check the outcome with an A/B test significance calculator and record it, whatever it says.

If you genuinely need the option to stop early, use a tool built for sequential testing and accept that it needs a little more data overall. And if the traffic is too low to reach a sensible sample in a reasonable time, test bigger changes rather than smaller ones. I build that kind of planning into testing as part of performance marketing work.

Do and do not

Do

  • Fix the sample size and run time before the test starts
  • Run tests in whole weeks
  • Look during the test only to catch broken variants

Do not

  • Stop a fixed-horizon test the first time it looks significant
  • Treat a live significance banner as final without knowing the tool's method
  • Build later tests on an early, unconfirmed winner

Questions people ask about this

Is it ever acceptable to stop an A/B test early?

Yes, if a variant is broken, is clearly harming sales, or you are using a sequential testing method designed for early stopping. What you should not do is stop a fixed-horizon test just because it crossed the significance line ahead of schedule. Record the reason whenever a test ends before its planned date.

Does peeking matter if the uplift is very large?

Large effects are less likely to be pure chance, but early results overstate effects more often than they understate them. A huge gap on day three frequently shrinks with more data. Letting the test finish tells you how big the real effect is, which matters when you are forecasting revenue from it.

How long should an A/B test run on a small UK site?

Long enough to reach the sample size your calculation requires, and never less than one full week, ideally two, to cover weekday and weekend behaviour. On low-traffic sites that can mean several weeks or more. If the required time is unrealistic, test a bolder change or measure a step with more volume, such as clicks to the booking page.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.