Statistical power

Definition
Statistical power is the probability a test will detect a real effect when one genuinely exists, rather than missing it entirely.

Why it matters

The danger is the false negative. A test on modest traffic comes back "no significant difference". The team concludes the change did nothing and drops it. In fact the sample was too small to see a real improvement. The test wasted the time it ran, and it also taught a wrong lesson that people then act on. This is the quiet problem for small businesses, where traffic is limited and tests often run underpowered.

How to apply it

  • Run a power calculation before the test starts. Many free calculators ask for the current conversion rate, the smallest lift worth detecting and the target power, then give the sample size needed.
  • Use 80 percent power as a common working target, not a guarantee.
  • If the answer is months of traffic, test a bigger, bolder change. Large effects need far less data than small ones.
  • Fix the sample size or run length in advance and do not stop early because the first few days look good or bad.
  • Record what each test actually needed so the next estimate is better.

What it is

Every test can go wrong in two ways. It can show a difference that is not real, which statistical significance guards against. Or it can fail to show a difference that is real. Power measures protection against the second error. A test with 80 percent power will spot a true effect of the expected size four times out of five.

Power rises with sample size and with the size of the effect being chased. It falls when the baseline rate is very low, because rare events produce few data points.

Common mistakes

  • Calling an underpowered test "no effect". A flat result from too little traffic means the test was inconclusive. Report it that way.
  • Calculating power after the test. Work out the needed sample size before launch. Doing it afterwards cannot rescue a test that was too small.
  • Chasing tiny lifts on low traffic. A change worth a few tenths of a percentage point needs far more visitors than a bold redesign. Match the ambition of the test to the traffic you have.
  • Stopping early. Ending a test when the numbers look good or bad changes its error rates. Fix the run length first.
  • Treating 80 percent as a law. It is a working convention. A costly decision may justify higher power.
  • Testing several variants at once on small traffic. Each extra variant divides the sample and lowers power for every comparison.
Worked example

Suppose an online bookshop wants to test a new checkout button. Its conversion rate is 2 per cent, and the team wants to detect a lift to 2.5 per cent. A power calculation at 80 per cent power says that needs about 14,000 checkout visitors per version. At 2,000 visits a week, the test would run for around fourteen weeks. A bolder layout that might lift conversion to 3 per cent needs only about 3,800 visitors per version, so it finishes in roughly four weeks.

VWO runs the split test, and the team writes the sample size into the test plan before launch, together with the rule for ending it. Power is the chance of spotting a real lift of the chosen size, so 80 per cent is a working target, not a guarantee. The first test needed fourteen weeks, so the bolder layout went into the next round.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Statistical significance

    The check against false positives, which power complements.

  2. Article

    Sample size

    The main lever that raises power.

  3. Article

    Confidence interval

    Shows the range of plausible effects once a test has enough data.

  4. Article

    A/B testing

    The practice where power matters most.

Where it shows up