Why most A/B tests don't change anything

By Mo Touzani · · Experimentation

Most A/B tests are designed to generate permission, not evidence. The team has a preferred direction, they run a test to validate it, and the test is considered successful when it confirms what they already believed. The result is an experimentation program that moves fast, ships constantly, and learns almost nothing.

The design flaw that comes before the experiment

Most experiments fail at the hypothesis stage, not the analysis stage. A hypothesis is a falsifiable claim about why a change will produce a specific outcome. "Adding a progress bar to the checkout flow will increase purchase completion by reducing uncertainty about the steps remaining" is a hypothesis. "Let us test a cleaner layout" is not. The second formulation has no prediction to evaluate. Whatever result comes back is compatible with it.

When there is no real hypothesis, there is no real learning. A positive result tells you the cleaner layout worked without telling you why. A negative result tells you it did not without telling you what to try next. The test generated a decision but not an understanding. Next quarter, the team runs another test from the same pool of aesthetics-based intuitions, and the cycle repeats.

The question that should precede every experiment: what specifically do we expect to happen, for what reason, and what result would change our minds? If the team cannot answer all three before the test runs, the experiment is not ready to run.

Statistical power and the test you never ran

Underpowered tests are the most common and least discussed failure mode in applied experimentation. A test is underpowered when the sample size is too small to detect the effect size that would matter to the business. The test runs, it returns a null result, the team concludes the change did not work, and moves on. The correct interpretation: the test was not capable of detecting a meaningful effect even if one existed.

The industry has largely solved for false positives. Teams know about p-values and statistical significance. The false negative problem gets far less attention. A company that runs 50 tests per quarter, each powered to detect a 20 percent relative lift when the true effects are in the 3 to 5 percent range, will produce mostly null results from tests that were never able to answer the question they were supposed to answer.

Power calculations are not bureaucracy. They are the mechanism by which a team confirms that a test, if run, is capable of producing a usable answer. Running a test without a power calculation is spending engineering and product resources on an experiment that was designed, structurally, to tell you nothing.

What most teams are testing

The vast majority of A/B tests in production environments test cosmetic or copy changes to pages with already-high conversion rates: button color, headline copy, image placement. These are real variables with real effects. They are also the wrong place to look for meaningful lift in most growth situations.

A 10 percent improvement in the conversion rate of a page that converts 60 percent of visitors produces a smaller absolute revenue impact than a 10 percent improvement in the conversion rate of a page that converts 8 percent of visitors. The pages with the most room for improvement are rarely the pages getting the most test traffic, because those pages are typically in the acquisition funnel where traffic volumes are lower and test cycles take longer.

Teams running the most tests are often optimizing the parts of the funnel that least need it. Teams running the fewest tests are often sitting on the highest-impact opportunities with no measurement infrastructure to act on them. Both problems stem from the same source: a testing culture that prioritizes test volume over test value.

The pre-specification problem

Most teams decide what a test result means after they see the result. This is the single most damaging practice in applied experimentation, and it is nearly universal. A positive result gets shipped. A negative result gets reframed as a learning or retested with different parameters. An inconclusive result gets shelved. None of these responses was specified in advance.

Pre-specification means deciding, before the test runs, what action follows each possible outcome. If the primary metric shows a statistically significant lift above the minimum detectable effect, ship. If it shows a statistically significant decrease, do not ship and document the hypothesis as disproved. If the result is inconclusive, define in advance what "inconclusive" means and what comes next.

The reason pre-specification matters: without it, every outcome is compatible with every prior belief. Teams that ship positive results and retest negative ones are not running experiments. They are running approval processes with extra steps.

The framework for high-conviction testing

  1. Write the hypothesis before writing the test spec. A valid experiment hypothesis names the change, the expected outcome, the mechanism by which the change produces that outcome, and the metric that will evaluate it. If your hypothesis does not include a mechanism, you do not have a hypothesis yet. Reject any test brief that arrives without one.
  2. Calculate the required sample size before building the test. Determine the minimum detectable effect that would justify shipping the change, given implementation cost and the opportunity cost of the traffic it requires. Use that effect size to calculate the sample size needed for 80 percent statistical power at your chosen significance threshold. If the required sample size exceeds what your traffic supports within a reasonable window, the test is not ready to run.
  3. Pre-specify what each result means and what action follows it. Before launch, document what constitutes a positive result, a negative result, and an inconclusive result. Assign an action to each outcome. Do not revisit the categorization after unblinding. A test program where decision criteria are set before results are seen is a learning system. One where they are set after is not.

What changes when you run experiments this way

High-conviction experimentation programs run fewer tests and learn more from each one. The constraint is not discipline for its own sake. It is recognition that underpowered, poorly hypothesized tests are expensive. They consume engineering time, product attention, and traffic that could have been allocated to a test capable of answering a real question.

The goal is not a high test velocity. The goal is a high rate of knowledge production per unit of organizational investment. Those are different optimization targets, and most experimentation programs are optimizing for the wrong one.

Book a consultation with Sinfa · [email protected]