Sample Ratio Mismatch (SRM) Checker
You configured a 50/50 split and got 50.9/49.1. Is that noise, or is your experiment broken? Enter the visitors in each arm and the split you planned. The checker runs a chi-square goodness-of-fit test and tells you whether the result can be trusted at all.
Worked example: a planned 50/50 test with 50,912 control and 49,088 variant visitors gives chi-square = 33.3 and p = 8.0e-9, a clear sample ratio mismatch. The result should not be trusted.
Why sample ratio mismatch matters
A sample ratio mismatch means the arms received a different share of traffic than you assigned. Random assignment lands very close to the planned split once samples are large. When it does not, something is filtering users out of one arm unevenly: a redirect that fails on some browsers, a bot filter, a slow variant that loses impatient visitors, or a tracking call that fires in only one version.
Whatever the cause, the users who remain in each arm are no longer comparable. The conversion difference you measure can come from who got counted rather than what they saw. More data or a better statistical method cannot rescue a test with SRM; you have to diagnose it and rerun.
How the check works
The checker computes the expected count for each arm from the total and your planned split, then the chi-square statistic Σ (observed − expected)² ÷ expected, with degrees of freedom equal to arms minus one.
It flags SRM at p < 0.0005, the strict threshold large experimentation platforms use because the check runs on every test and false alarms are costly. Values between 0.0005 and 0.01 earn a warning: investigate before relying on the result.
Common causes and where to look
Redirect tests: the variant URL loads slower or fails for some users. Compare arms by browser and device.
Bot and internal-traffic filtering applied after assignment. Check whether the filter treats arms differently.
Tracking that depends on the variant: an analytics tag missing from one template, or an event that fires before the page change in one arm only.
Assignment changes mid-test: allocation edited, targeting changed, or a variant paused and resumed.
Frequently asked questions
What is sample ratio mismatch?
A statistically significant gap between the share of users each A/B test arm received and the share you configured. It signals a bug in assignment or tracking, which makes the test result untrustworthy.
How big a difference counts as SRM?
It depends on sample size. 5,050 vs 4,950 is normal noise. 505,000 vs 495,000 is a serious mismatch with p far below 0.0005. Test it rather than eyeballing percentages.
Can I still use a test with SRM?
No. Find the cause, fix it and rerun. Reweighting or trimming the larger arm does not repair the bias, because you do not know which users went missing.
Should I check SRM on every test?
Yes, before reading any metric. Large experimentation programs report that SRM shows up in a meaningful share of tests, and it is a common reason trustworthy-looking results fail to replicate.