The decision criteria problem in experimentation
By Mo Touzani · · Experimentation
Most teams running A/B tests know when their results reach statistical significance. They check the dashboard, they see the p-value cross 0.05, and they feel a sense of resolution. What most of them have not defined is what to do next. This missing step is responsible for more wasted experimentation investment than any other single failure in applied testing programs.
The pre-specification gap
Statistical significance tells you that the observed difference between variants is unlikely to be due to random chance at a given probability threshold. It tells you nothing about whether the observed difference is large enough to matter, whether the change should be shipped, or what the team should do if the primary metric improves but a guardrail metric degrades.
Decision criteria answer these questions. They specify, before the test runs, what each possible result means and what action follows it. Without decision criteria, every test result is open to interpretation after the fact. Positive results get shipped. Negative results get rationalized or retested. Inconclusive results get shelved. None of these decisions was made by the evidence. They were made by the person looking at the evidence and deciding what it meant in the moment.
The research on confirmation bias in experimentation is well established. People interpret ambiguous evidence in ways consistent with their prior beliefs. An experimentation program that relies on post-hoc interpretation of results will systematically produce decisions that reflect the preferences of the decision-makers, not the results of the tests. Pre-specification is the structural safeguard against this.
What decision criteria specify
A decision criterion has four components: the metric that will be evaluated, the threshold that constitutes a meaningful result, the statistical confidence required, and the action that follows each possible outcome. For a primary conversion metric, this might look like: if the conversion rate improvement exceeds 2 percent relative lift at 95 percent confidence, ship; if the improvement is statistically significant but below 2 percent, evaluate implementation cost against expected revenue; if the result is a statistically significant decrease, do not ship and document the hypothesis as disproved.
Guardrail metrics require their own criteria. A guardrail metric is one that the test is required not to harm, even if the primary metric improves. Customer support contact rates, return rates, and revenue per customer are common guardrails in conversion optimization programs. A test that improves checkout conversion by 5 percent but increases return rates by 8 percent is not a success. Without a pre-specified guardrail threshold, this conclusion requires a judgment call after the fact, and that call will often go the wrong way.
The specific thresholds matter less than the fact that they are specified in advance. A team that decides before running a test that a 2 percent lift is the minimum useful result and holds to that standard has a more defensible and consistent decision process than a team that decides what "good enough" means after seeing the data.
The peeking problem
Sequential testing, where analysts check results as data accumulates and stop the test when a significant result appears, inflates the false positive rate substantially above the nominal significance level. Running a test at p less than 0.05 with a stated 5 percent false positive assumption, and checking results daily from day 10 onward, can produce effective false positive rates of 20 to 30 percent or higher depending on the number of interim checks performed.
The practical consequence: a large fraction of the "positive" results in programs that allow informal peeking are not real effects. They are random fluctuations that crossed the significance threshold during an interim check. The team ships the change, the expected lift does not materialize in post-ship monitoring, and the team attributes the miss to external factors rather than to the measurement error that produced the false positive.
Either commit to a fixed-horizon test that is only analyzed at the pre-specified endpoint, or use a sequential testing method designed to allow interim analysis without inflating the false positive rate. Informal peeking with a fixed-sample-size test is not a third option.
Handling inconclusive results
An inconclusive result from a well-powered test is real information. It means the true effect, if it exists, is smaller than the minimum detectable effect the test was designed to find. That finding is useful: it rules out the hypothesis that the change produces a practically significant improvement at the tested effect size. It does not mean the test was a failure. It means the hypothesis was tested and the evidence does not support it.
Most teams do not treat inconclusive results this way. They retest with a different variant (compounding the multiple-testing problem), extend the test window retroactively (inflating the false positive rate), or shelve the hypothesis without documenting what was learned. None of these responses treats the null result as the information it is.
Proper treatment of inconclusive results: document the hypothesis, the test design, the effect size tested for, and the null finding. Move on without retesting the same hypothesis unless there is a substantive reason to believe the original test design was flawed. An experimentation program that documents null results is building an organizational knowledge base. One that discards them is running the same ineffective tests repeatedly.
The framework for decision-grade experimentation
- Write the decision criteria document before writing the test brief. For each test, define in writing: the primary metric, the minimum effect size that constitutes a meaningful result, the statistical significance threshold, the guardrail metrics and their acceptable ranges, and the action that follows each possible outcome. This document should be reviewed by relevant stakeholders before the test is built. Anything agreed upon after results are available is not a decision criterion. It is a rationalization.
- Run only fixed-horizon tests or use a sequential framework designed for early stopping. Commit to a test duration at the start and do not end the test early based on interim results unless you are using a sequential testing method that controls the false positive rate under interim analysis. The test duration should be determined by the power calculation, not by the planning calendar or the pressure of a team waiting for results.
- Treat every null result as a finding about the hypothesis space. A well-powered test that returns an inconclusive result has produced information: the change does not produce a practically significant improvement at the tested effect size. Document this finding with the same rigor used for positive results. Include it in the experimentation log. Reference it when similar hypotheses are proposed in future planning cycles. The organizations that learn from null results accumulate knowledge faster than those that discard them.
What changes when decision criteria are in place
Experimentation programs with pre-specified decision criteria ship fewer changes and learn more from each test. The constraint is productive. When the team knows in advance what a positive result looks like and has committed to acting on it, the definition of a good experiment shifts from "a test that produces a significant p-value" to "a test that answers the question we designed it to answer."
This is the culture shift that separates high-performing experimentation programs from high-volume ones. Volume is easy to produce. Decision-grade evidence is harder. The teams that consistently generate evidence of the second type are the ones that do the disciplined work of specifying what they need from each test before they ask the test to deliver it.