Test duration is how long an A/B test runs before you read the result. It should be decided before the test starts, from the number of visitors you need and the number you actually get, rather than chosen by watching the results until they look good.
How test duration works
Duration comes from two numbers. The first is the sample size: how many visitors each version needs for a difference of the size you care about to show up reliably. That depends on your current conversion rate, the minimum detectable effect you want to spot, and the confidence level you choose. The second is your traffic: how many eligible visitors reach the page each day.
A worked example, using illustrative figures rather than data from any real site: a landing page converts at 3% and you want to detect a lift to 3.6% (a 20% relative improvement) at 95% confidence and 80% power. That needs roughly 14,000 visitors per version, about 28,000 in total. At 1,000 visitors a day, the test needs four weeks. At 200 a day, it needs more than four months, which tells you the test is not practical as designed.
Those figures change sharply with the effect size. Halve the lift you want to detect and the visitors needed roughly quadruple, which is why small tweaks are so expensive to test on modest traffic.
Once you have a figure, round it up to whole weeks. Behaviour differs between weekdays and weekends, so a test that runs Monday to Thursday captures a different audience from one that covers a full week.
Why it matters
Ending a test early is the most common way to get a false winner. Results swing a lot in the first few days, and if you check daily and stop as soon as one version looks ahead, you will often crown a change that does nothing. This is known as peeking. Running far too long has its own cost: cookies expire, visitors return in the other version, and seasonal changes creep in.
UK calendars add their own noise. Bank holiday weekends, month-end paydays, school holidays, Black Friday and the run-up to Christmas all change who visits and how ready they are to buy. A test that straddles one of those is comparing different audiences as much as different pages.
Common mistakes
- Stopping when the tool first says significant. Decide the duration in advance and read the result at the end.
- Testing tiny changes on low traffic. A button colour test on a page with 100 visits a day may need a year.
- Partial weeks. Stopping on a Wednesday skews the sample towards weekday visitors.
- Ignoring the novelty effect. Returning visitors may click something simply because it is new; the effect fades in later weeks.
- Changing the test midway. Editing a variant or the traffic split resets what you are measuring.
How to act on it
Before launching, write down the conversion you are measuring, your baseline rate, the smallest lift worth acting on and your daily traffic. Work out the sample size, divide by traffic and round up to full weeks. If the answer is longer than about six to eight weeks, test a bolder change or a higher-traffic page instead. When the planned duration ends, check the result with an A/B test significance calculator and record it, win or lose.
Planning tests that your traffic can actually support is part of my PPC landing page design service.
