Analytics and Tracking

p-value

Also called p value

A number showing how likely a test result at least this extreme would be if the change made no real difference.

Quick facts: p-value

Category
Analytics and Tracking
Also called
p value
Level
Advanced
Affects
A/B test decisions, landing page changes, conversion rate optimisation, reporting
Where to see it
A/B testing platforms, online significance calculators, spreadsheets, GA4 data exported for analysis
In this article5
  1. How a p-value works
  2. What a p-value does not tell you
  3. Why it matters
  4. Common mistakes
  5. How to act on it

A p-value is a number between 0 and 1 that tells you how surprising your test result would be if the change you made had no real effect at all. In marketing it comes up in A/B tests: a small p-value means the difference between two versions is unlikely to be pure chance, while a large one means chance could easily explain it.

How a p-value works

Every test starts with a dull assumption, called the null hypothesis: the new version performs exactly the same as the old one. The p-value answers one question about that assumption. If it were true, how often would random variation alone produce a gap at least as big as the one you saw?

Take a worked example. A landing page converts 2.0% of 3,000 visitors, and a new version converts 2.6% of another 3,000. That looks like a healthy 30% improvement. Run the standard two-proportion test and the p-value comes out at roughly 0.12, meaning a gap this size would turn up about one time in eight even if both pages were identical. Keep the same conversion rates but double the traffic to 6,000 visitors each, and the p-value drops to about 0.03.

The convention in most testing tools is to call a result statistically significant when the p-value is below 0.05. That threshold is a habit, not a law of nature. It means accepting roughly a one-in-twenty chance of declaring a winner when nothing really changed.

What a p-value does not tell you

This is where most misreadings happen. A p-value of 0.03 does not mean there is a 97% chance the new version is better. It does not tell you how big the improvement is, and it says nothing about whether the improvement is worth having. A tiny, commercially pointless lift can produce a very small p-value on a site with huge traffic. For the likely size of the effect, look at the confidence interval instead.

Why it matters

Most UK small businesses do not have the traffic to run many tests, so each one carries weight. Declaring a false winner means you roll out a change that does nothing, or does harm, and you stop looking for the real problem. Ignoring a genuine winner because nobody understood the numbers wastes the effort that went into the test.

Understanding the p-value also protects you from confident reports. When an agency or software tool announces a winning headline after four days and 200 visitors, the p-value, and how it was reached, is the first thing to ask about.

Common mistakes

  • Peeking. Checking the p-value every day and stopping the moment it dips below 0.05 makes false winners far more likely, because random swings cross the line sooner or later.
  • Too many comparisons. Test five versions against each other, or twenty metrics, and something will look significant by chance alone.
  • Reading 0.06 as “nearly significant” proof. It is an inconclusive result.
  • Ignoring practical size. A significant 0.1 percentage point lift may not cover the cost of building the change.
  • Running tests that could never finish because traffic is too low for the effect you hope to find.

How to act on it

Before a test starts, write down the sample size you need for the smallest improvement worth detecting, the single metric that decides the winner, and the threshold you will use. Then let the test run to that size, ideally in whole weeks so weekday and weekend behaviour are both included, and read the p-value once at the end.

Report the result as a range, not a verdict: the likely lift, the confidence interval and the p-value together. If the result is inconclusive, say so, and decide whether a bigger change is worth testing next.

When I build and test landing pages for paid campaigns, I agree these rules with the client before launch, so nobody is tempted to call a winner early.

Do and do not

Do

  • Fix the sample size, metric and threshold before the test starts
  • Report the confidence interval alongside the p-value
  • Run tests in whole weeks

Do not

  • Stop a test the first time the p-value drops below 0.05
  • Read a p-value as the chance the variant is better
  • Test many variants or metrics without adjusting for it

Questions people ask about this

Is a p-value of 0.05 good enough?

It is the most common threshold, and fine for low-risk changes such as a headline or button wording. For an expensive or hard-to-reverse change, such as a new checkout or a price rise, a stricter threshold like 0.01 is sensible. Whatever you choose, set it before the test starts, not after you see the result.

Why does my testing tool show 'probability to beat' instead of a p-value?

Some tools use Bayesian statistics, which answer a different question: how likely the new version is to be better, given the data. It is not the same thing as a p-value and the numbers are not interchangeable. Both approaches still suffer if you stop early or test too many variants.

Can I get a meaningful p-value with low website traffic?

Only for large effects. As a rough guide, spotting a 10% relative lift on a page converting at 2% takes tens of thousands of visitors per version, far more than most small sites see in a month. If traffic is low, test bigger, bolder changes, run tests for longer, or use qualitative research such as user interviews and session recordings instead.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.