Statistical Power Analysis: Sizing Experiments Before You Run Them

Codeayan Team · Oct 9, 2026 · 0 Views

A team ships a checkout redesign to 5,000 visitors per arm. Conversion reads 4.3% against 4.0%, the difference is not significant, and the verdict is “no effect”. Nobody asks whether 5,000 visitors could have detected a lift that small. They could not. At that sample size the test had roughly a 12% chance of reaching significance even though the lift was real.

Short answer

Statistical power analysis links four quantities: significance level, power, effect size and sample size. Fix any three and the fourth is determined. Before a test you fix the first three, usually 0.05, 0.80 and the smallest effect worth shipping, then solve for sample size. A conversion rate moving from 4.0% to 4.5% needs about 25,531 users per arm.

What do you need before running a statistical power analysis?

Gather five inputs. The hard one is the minimum detectable effect, because it is a business judgement, not a statistic. A 0.5-point lift on a ₹1,200 basket is worth ₹6 per session. If ₹6 would not change the decision, do not size the test to find it.

  • A baseline rate or standard deviation taken from real historical data, not a guess.
  • A minimum detectable effect: the smallest lift that would change what you do.
  • A significance level (normally 0.05) and a power target (normally 0.80, Cohen’s convention).
  • A decision on one-sided versus two-sided testing, made before any data arrives.
  • The allocation ratio between arms, equal unless you have a reason.

Power is one minus the Type II error rate, so a 0.80 target accepts a one-in-five chance of missing a real effect. The relationship between p-values and Type I and II errors explains why both rates cannot be driven to zero at once. Skipping this step is common. Button and colleagues (Nature Reviews Neuroscience, 2013) estimated median power across neuroscience studies at 21%, so a typical study would miss a true effect of the assumed size about four times in five.

How do you calculate sample size in Python?

Use statsmodels. Convert the two proportions into an effect size, then ask NormalIndPower to solve for the missing parameter. Conversion counts follow a binomial distribution, and the central limit theorem is what lets a z-test approximate them at this scale.

  1. Convert the rates to an effect size. proportion_effectsize(0.045, 0.040), from the statsmodels power module, returns Cohen’s h, about 0.0248.
  2. Set alpha and power. Use 0.05 and 0.80 unless a missed win costs far more than a false alarm.
  3. Leave nobs1 unset. The solve_power documentation requires exactly one parameter to be None, and that is the one returned.
  4. Round up. That gives 25,531 per arm and 51,062 in total.
Python: Sample Size for a Conversion Test
1import math
2from statsmodels.stats.power import NormalIndPower
3from statsmodels.stats.proportion import proportion_effectsize
4
5baseline, target = 0.040, 0.045  # today's rate vs the lift worth shipping
6es = proportion_effectsize(target, baseline)  # Cohen's h
7n = NormalIndPower().solve_power(effect_size=es, alpha=0.05, power=0.80, ratio=1.0)
8print(round(es, 4), math.ceil(n))  # 0.0248 25531
9
10for t in (0.043, 0.045, 0.050):
11    h = proportion_effectsize(t, baseline)
12    print(t, math.ceil(NormalIndPower().solve_power(h, alpha=0.05, power=0.80)))
13# 0.043 69358
14# 0.045 25531
15# 0.05 6726

At 8,000 checkout sessions a day, 51,062 sessions take about 6.4 days. Run it for two full weeks anyway, so weekday and weekend behaviour both land in each arm.

How do you check the sample size is right?

Simulate it. Draw thousands of fake experiments at the planned sample size, run the real test on each, and count rejections. Where a true lift exists, the rejection rate should sit near your power target. Where none exists, it should sit near alpha.

Python: Simulating Power and False Positives
1import numpy as np
2from statsmodels.stats.proportion import proportions_ztest
3
4rng = np.random.default_rng(11)
5m, runs = 25531, 2000
6
7def rejection_rate(p_a, p_b):
8    hits = 0
9    for _ in range(runs):
10        a, b = rng.binomial(m, p_a), rng.binomial(m, p_b)
11        _, p = proportions_ztest([b, a], [m, m])
12        hits += p < 0.05
13    return hits / runs
14
15print(rejection_rate(0.040, 0.045))  # power: 0.8085
16print(rejection_rate(0.040, 0.040))  # false positives: 0.0515

Halve the sample and the same lift is detected only about half the time (a computed power of roughly 51%). That is the cost of cutting the test short.

Which mistakes quietly wreck a power calculation?

Most errors in a statistical power analysis change the answer by tens of percent, and each one looks harmless in a planning doc. Sample size scales with the inverse square of the effect, so a small cut to the detectable lift is expensive.

Planning choiceUsers per armCost or caveat
Baseline plan
4.0% to 4.5%, alpha 0.05, power 0.80
25,531 Reference point.
Power 0.90 34,179 About 34% more traffic.
Alpha 0.01 37,989 About 49% more traffic.
Detect 4.3% instead 69,358 2.7 times the traffic.
One-sided test 20,111 21% fewer users, valid only if declared before launch.

One more trap sits outside the table. The calculation assumes a single look at the data. Checking daily and stopping at the first p below 0.05 pushes the real false positive rate well above 5%, so either hold the analysis until the planned sample is in, or switch to a sequential design. And never run the calculation after the fact using the observed effect. That “post-hoc power” is a restatement of the p-value and tells you nothing new.

  • Statistical power analysis ties together alpha, power, effect size and sample size, so any three determine the fourth.
  • The minimum detectable effect is a business decision, and it drives the cost because sample size scales with its inverse square.
  • A 4.0% to 4.5% conversion lift at alpha 0.05 and power 0.80 needs 25,531 users per arm, while 5,000 per arm gives only about 24% power for that lift.
  • Simulating the planned test confirms that rejection rates land near the power target when a lift exists and near alpha when it does not.
  • Peeking, post-hoc power and a one-sided test chosen after launch each invalidate the calculation you did beforehand.

Conclusion

Run the statistical power analysis before the experiment is approved, and write the sample size and the stopping date into the plan. Once the numbers are in, the A/B testing best practices guide covers how to read the result without fooling yourself.

Frequently Asked Questions

What is statistical power analysis?

Statistical power analysis is a calculation that connects four quantities: significance level, power, effect size and sample size. Fixing any three determines the fourth. Researchers usually run it before data collection to find the sample size needed to detect a chosen effect with an acceptable chance of success.

What power level should I use for an experiment?

Use 0.80 as the default, a convention going back to Jacob Cohen. It accepts a one-in-five chance of missing a real effect. Raise it to 0.90 when a missed win is costly, but expect roughly a third more sample, since power and sample size rise together.

What is the minimum detectable effect?

The minimum detectable effect is the smallest true difference your test is designed to find at the chosen power. It should come from a business judgement about what change would alter a decision, not from the data. Halving it roughly quadruples the required sample size.

What happens if my A/B test is underpowered?

You will usually miss real effects, and the significant results you do get tend to overstate the true effect size. An underpowered test that shows no significant difference is inconclusive, not evidence of no effect, because it could not have detected a modest lift in the first place.

Can I calculate power after the experiment finishes?

Not usefully. Power computed from the observed effect is a direct function of the p-value, so it adds no information. If a result is not significant, report the confidence interval instead, and plan the next test with a power analysis based on the smallest effect that matters.