A/B Test Sample Size Calculator

How many visitors does your A/B test need before the result means anything? Enter your baseline conversion rate and the smallest lift worth detecting. The calculator returns the sample per variant, the total sample, and how long the test has to run.

Worked example: at a 3% baseline, detecting a 10% relative lift (3% to 3.3%) at 95% significance and 80% power needs 53,211 visitors per variant, 106,422 in total, or 54 days at 2,000 visitors a day.

How the calculation works

The calculator uses the standard two-sided, two-proportion z-test, the same model behind most A/B testing platforms. It asks one question: how many visitors per variant do you need so that, if the true lift is at least your minimum detectable effect, the test has the chosen probability (power) of detecting it at the chosen significance level?

Per-variant sample size = (z₁₋α/₂·√(2p̄(1−p̄)) + z₁₋β·√(p₁(1−p₁) + p₂(1−p₂)))² ÷ (p₂ − p₁)², where p₁ is the baseline rate, p₂ the rate you want to detect, and p̄ their average.

With more than one challenger, the significance level is split across comparisons (Bonferroni correction). Each extra variant raises the sample every arm needs, as well as adding an arm.

How to choose the inputs

Baseline conversion rate: use the control experience's rate for the exact metric and audience in the test, measured over a normal period of at least two weeks. A sitewide average applied to a single page overstates or understates the sample.

Minimum detectable effect (MDE): the smallest lift that would change a business decision, not the lift you hope for. Halving the MDE roughly quadruples the sample. If the sample it produces is out of reach, test a bolder change or a higher-traffic step of the funnel.

Significance (α) and power: 5% significance and 80% power are the common defaults. Raise power to 90% when a missed winner is expensive; tighten significance when a false winner would drive a large, hard-to-reverse investment.

Mistakes that invalidate the result

Stopping the test the first time it looks significant. Repeated peeking inflates the false-positive rate far beyond 5%. Fix the sample in advance and read the result once it is reached.

Running for less than one full week. Weekday and weekend traffic convert differently; run in whole-week increments even if the sample is reached sooner.

Changing traffic allocation, targeting, or the variant mid-test. Any of these breaks the comparison the sample size was calculated for.

Frequently asked questions

What sample size do I need for an A/B test?

It depends on your baseline conversion rate and the smallest lift you need to detect. At a 3% baseline, detecting a 10% relative lift (3.0% to 3.3%) with 95% significance and 80% power takes about 53,000 visitors per variant. Detecting a 20% lift takes about 14,000.

How long should an A/B test run?

Divide the total sample by the daily visitors entering the test, then round up to whole weeks, with a minimum of one to two full weeks so weekday and weekend behavior are both represented.

What is minimum detectable effect?

The smallest true difference between control and variant that the test is designed to detect reliably. Smaller effects need larger samples; the relationship is roughly inverse-square.

Should I use relative or absolute MDE?

Relative MDE (e.g. +10% of the baseline) is easier to compare across pages with different conversion rates. Absolute MDE (e.g. +0.3 percentage points) is clearer when you are reasoning about revenue. The calculator supports both.

What if my site does not have enough traffic?

Test bigger, bolder changes so the realistic effect is larger; test at a higher-traffic step of the funnel; or use a higher-volume proxy metric that is validated against revenue. Running an underpowered test and reading it anyway produces confident-looking noise.

Book a consultation with Sinfa · [email protected]