Can I Stop My A/B Test Yet?
Your dashboard says 97% confidence on day four. Can you call it? Enter the sample you planned, where the test is now, and how many times you have looked. The checker tells you whether stopping is defensible, how long is left, and what all that checking has done to your error rate.
Worked example: a 5% significance test checked 5 times carries a real false-positive rate of about 14%; checked 10 times, about 19%.
Why peeking breaks A/B tests
A 95% confidence threshold allows a 5% false-positive rate for one look at a pre-planned sample. Every additional look gives random noise another chance to cross the line. Armitage, McPherson and Rowe showed in 1969 that five looks at 5% significance push the real false-positive rate to about 14%, and ten looks to about 19%.
Early in a test, conversion rates swing widely. A variant that looks 30% better on day three often settles near zero by day twenty. Stopping on the swing locks in noise and calls it a winner.
How the checker decides
Planned sample reached and at least one full week run: stop and read the result, whatever it shows.
Sample not reached: keep running. The checker estimates the days left from your traffic so far.
Early stop: if the test has run a full week and the z-score clears an O'Brien-Fleming-type boundary, z ≥ z₁₋α/₂ ÷ √(share of planned sample), the evidence is strong enough for a pre-planned early stop. At 25% of the sample that boundary sits near z = 3.9; at 50%, near 2.8. This only holds if early stopping was part of the plan before launch.
Sample ratio mismatch: if the arms are unbalanced beyond chance, stop reading the result and fix the assignment first.
How to stop peeking without flying blind
Fix the sample before launch with a sample size calculator, and write the stopping rule down.
Monitor guardrails, not the primary metric. Errors, page speed and revenue crashes justify an emergency stop; a promising lift does not.
If you need the option to stop early, plan for it: use a sequential method with a defined number of looks and boundaries, and accept a slightly larger maximum sample in exchange.
Frequently asked questions
When can I stop an A/B test?
When it has reached the sample size you calculated before launch and has run at least one full week, ideally in whole weeks. Stopping earlier is only valid with a sequential design planned in advance.
What is peeking in A/B testing?
Repeatedly checking significance while a fixed-sample test is running and stopping when it crosses the threshold. It inflates the false-positive rate well above the nominal 5%.
My test hit 95% significance early. Is it a winner?
Probably not yet. Early significance in a fixed-sample test is common and often reverses. Unless the result clears a much stricter early-stopping boundary and you planned for early stopping, keep running to the planned sample.
What if my test will take months?
Then the planned effect is too small for your traffic. Test a bolder change, move the test to a higher-traffic step, or accept that this question cannot be answered with an A/B test at your volume.