A/B Test Significance Calculator
Did your variant win, or did you get lucky? Enter visitors and conversions for each arm. The calculator returns the p-value, the relative lift, the confidence interval for the difference, and the smallest lift your sample was able to detect.
Worked example: 300 conversions from 10,000 control visitors (3.0%) against 360 from 10,000 variant visitors (3.6%) is a +20% relative lift with p = 0.018, significant at 95%. The 95% confidence interval for the difference runs from +0.10 to +1.10 percentage points.
How the calculation works
The calculator runs a two-sided, two-proportion z-test with a pooled standard error, the same test most A/B platforms report. z = (p_B − p_A) ÷ √(p̄(1 − p̄)(1/n_A + 1/n_B)), where p̄ is the pooled conversion rate. The p-value is the probability of seeing a difference at least this large if the variant truly had no effect.
The confidence interval uses the unpooled standard error of the difference. If the interval excludes zero, the result is significant at the chosen level. Its width tells you more than the p-value: a significant result with an interval from +0.1% to +40% says the lift is real but tells you almost nothing about its size.
The calculator also reports the smallest relative lift your sample could detect with 80% power. If your observed lift is far below it, a non-significant result means too little data, not proof of no effect.
How to read the result
Significant with a narrow interval: you have a usable estimate. Plan on the low end of the interval, not the point estimate.
Significant with a wide interval: the direction is probably right; the size is not known. Expect the realized lift after launch to be smaller than the test showed, because winners selected by significance overstate their effect.
Not significant: check the detectable lift. If it is larger than any lift you could realistically expect, the test was underpowered from the start. Size the next one with the sample size calculator before launch.
Mistakes that invalidate the result
Checking the p-value daily and stopping at the first value under 0.05. Each look is another chance for noise to cross the line; ten looks push the real false-positive rate near 20%.
Testing many metrics or segments and reporting the one that won. With 20 comparisons, one will usually clear p < 0.05 by chance.
Ignoring a sample ratio mismatch. If the traffic split is off from what you configured, the assignment is broken and the p-value means nothing. Run the SRM checker first.
Frequently asked questions
What p-value means an A/B test is significant?
By convention, p < 0.05 at 95% confidence. That threshold only holds if you fixed the sample in advance and read the result once. Repeated peeking makes p < 0.05 far easier to reach by chance.
Is a statistically significant result always worth shipping?
No. Significance says the difference is unlikely to be zero. The confidence interval tells you how large it plausibly is. If the low end of the interval would not cover the cost of the change, the result does not justify the rollout.
Why is my big lift not significant?
Large lifts on small samples carry wide uncertainty. Compare your observed lift with the detectable lift the calculator shows. If the sample could only detect a 40% lift, a 15% observed lift is indistinguishable from noise.
One-tailed or two-tailed test?
This calculator uses a two-tailed test, which detects harm as well as improvement. One-tailed tests halve the evidence required and are rarely justified for product decisions, because a variant that hurts conversion matters.
Can I use this for revenue per visitor?
No. This test is for conversion rates (yes/no outcomes). Revenue per visitor is a continuous, heavily skewed metric and needs a t-test or bootstrap on per-visitor values.