← Glossary
GLOSSARY · Measurement & analysis

Statistical Power

The probability of detecting a real effect when one exists — a small sample misses effects that are there

In one line

Statistical power is the probability that a test detects a real effect when one exists. With low power, effects that are genuinely there get missed.

Why it matters

A non-significant result has two possible explanations: there really is no effect, or there is one and the sample was too small to see it.

Reporting the second as "no effect" kills working initiatives. The burden of proof is asymmetric — claiming no effect requires showing the test would have caught an effect of the size that matters.

Decide before you start

Power comes from sample size, target effect size, and variability. So the order is to decide how small a difference you care about, then derive the sample needed from that.

This matters most on low-conversion campaigns, where both the holdout size and the run length have to grow.

Go deeper

Sample planning and decision rules are covered in A/B testing.

Frequently asked questions

Does not significant mean there is no effect?
No. With low power, a real effect goes undetected. 'Not significant' means 'not yet distinguishable', and claiming no effect requires separate evidence that the test could have detected one.
What determines power?
Sample size, the size of the effect you want to detect, and the variability of the data. Detecting smaller differences requires disproportionately more sample, so decide up front how small a difference matters.
What power level is standard?
80% is the conventional target. That means a real effect is detected four times out of five, and the fifth run comes back non-significant despite the effect existing. A non-significant result from a test that never had 80% power is weak grounds for claiming no effect.
Related:A/B Testing: Sample Size, Significance, and Decision RulesAdvertising Uplift: Measure Net Lift With a Holdout Test