Peeking is the habit of checking an A/B test while it is still running and stopping it the moment one version looks like a significant winner. It feels efficient, but it quietly breaks the statistics behind the test, so many of the “winners” you declare will be nothing more than chance.
How peeking works
Most simple test calculators use a fixed-horizon method. You decide the sample size in advance, run the test until you reach it, then check the result once. The familiar 95% confidence threshold means that, if there is really no difference between the versions, you will wrongly declare a winner about one time in twenty.
That one-in-twenty promise only holds if you look once. Results in a live test wander up and down as visitors arrive, especially in the early days when numbers are small. Every time you check, you give that random wandering another chance to cross the significance line. Check daily for three weeks and stop at the first crossing, and your real chance of a false winner is far higher than the 5% you think you are accepting.
Picture a dental practice testing two versions of its booking page. On day three, version B is ahead by 42 bookings to 22 and the tool flags it as significant. The test is stopped and B goes live. Over the next two months bookings settle back to where they were, because the early gap was a lucky run rather than a real effect. Nothing about that run was unusual; with small daily numbers, swings like it are common.
Why it matters
UK small and medium-sized sites rarely have the traffic for fast tests, which makes the temptation to stop early stronger. The cost is not only a wasted test. False winners get built into the site, the team learns the wrong lessons about what customers respond to, and later tests are designed on top of a mistaken result.
It also damages trust in testing itself. When a celebrated winner fails to move real revenue, people conclude that testing does not work, when the problem was how the result was read.
Common mistakes
- No sample size set in advance. Without one, there is no defined end point, so every look becomes a decision point.
- Stopping on a significant result but continuing on a non-significant one. This asymmetry is the core of the problem.
- Running for less than a full week. Behaviour on a Saturday differs from a Tuesday. Short tests capture only part of the weekly cycle.
- Trusting a tool’s banner without knowing its method. Some testing tools use sequential or Bayesian statistics designed for continuous monitoring; others do not. Check which yours uses before treating its live significance figure as final.
How to act on it
Before the test starts, write down the smallest uplift worth detecting, the sample size per variant and the minimum run time in whole weeks. Agree with everyone involved that the result is read once, at that point, unless something is broken.
It is fine to look during the test for problems: a variant not loading, tracking failing, a sharp fall in conversions that suggests a bug. Looking for errors is not peeking; acting on an apparent win is. When the planned sample is reached, check the outcome with an A/B test significance calculator and record it, whatever it says.
If you genuinely need the option to stop early, use a tool built for sequential testing and accept that it needs a little more data overall. And if the traffic is too low to reach a sensible sample in a reasonable time, test bigger changes rather than smaller ones. I build that kind of planning into testing as part of performance marketing work.
