Incrementality testing for companies that don't have Netflix's budget
By Mo Touzani · · Marketing Measurement
Incrementality testing has a reputation problem. The perception is that it requires a large media budget, a dedicated data science team, and the kind of randomization infrastructure that only platform-scale companies can build. That perception is wrong. The core of incrementality testing is a controlled experiment, and controlled experiments are more accessible than the industry has led most companies to believe.
What incrementality means
An incrementality test measures the additional revenue, conversions, or customers that a marketing activity produced above what would have happened without it. The word "above" is doing the important work in that sentence. Incrementality is not about the revenue that a channel touched or the conversions that appeared in its attribution window. It is specifically about the revenue that would not have occurred without the channel.
Most marketing measurement systems do not measure this. They measure presence, proximity, and correlation. They tell you where buyers were and what they saw before converting. They cannot tell you, without a controlled experiment, which part of that activity was causal.
The practical consequence: a channel with high attributed revenue and low incremental revenue is spending money to claim credit for conversions that would have happened anyway. Scaling that channel costs money and produces less new revenue than the attribution report suggests. The gap between attributed revenue and incremental revenue is the cost of the measurement error.
The geo-holdout test: accessible incrementality for real budgets
The most practical incrementality testing method for companies without platform-level randomization infrastructure is the geo-holdout test. The basic design: divide your target market into geographic units (DMAs, cities, or postal code groups), randomly assign a subset of those units to a control condition where the channel being tested is withheld or reduced, run the test for a defined period, and measure the revenue difference between test and control regions.
The inputs required are modest. You need geographic revenue data at the unit level, which most companies with any e-commerce or lead tracking infrastructure already have. You need the ability to suppress or modify channel activity by geography, which is available in most major advertising platforms. You need a test period long enough to accumulate sufficient conversions to detect the effect size you are testing for.
The output is a clean, defensible estimate of the incremental revenue per dollar spent in the channel being tested. It is not perfect. Geographic units are not perfectly comparable, and external events can produce noise during the test window. But it is causally grounded in a way that no attribution model can match, and it is within reach of a company spending $50,000 per month on paid media.
Designing the test to produce a usable result
The minimum viable geo-holdout test requires four things: geographic units with sufficient conversion volume to detect your minimum detectable effect within your test window, a random assignment process that balances observable pre-test characteristics between test and control groups, a clean suppression condition for the control group, and a pre-specified analysis plan that defines how the result will be interpreted before data is collected.
The effect size question is where most teams go wrong. The tendency is to design the test to detect whatever effect the model predicts, rather than the effect that would justify the channel investment at current spend levels. These are different numbers. If a 5 percent incremental lift would justify the current spend, the test should be powered to detect 5 percent. Designing the test to detect 20 percent and then concluding from a null result that the channel is working at sub-20-percent lift is a common and expensive error.
The test window is determined by the time required to accumulate the necessary conversions in the control group, not by the quarter schedule or the planning cycle. Calendar pressure is the most common reason geo-holdout tests produce inconclusive results.
Reading the results without overfitting them
A geo-holdout test produces a point estimate and a confidence interval. The point estimate is the measured incremental lift. The confidence interval tells you the range of plausible true effects, given the statistical noise in the data. Both numbers matter.
The error most teams make: treating the point estimate as the true effect and building scaling decisions on it directly. If the test shows a 12 percent incremental lift with a 90 percent confidence interval spanning 4 to 20 percent, the true incremental lift is somewhere in that range, not specifically at 12 percent. Scaling spend based on the point estimate without accounting for the interval is a systematic source of overconfidence.
The appropriate use of a geo-holdout result: it is one calibration point in an ongoing measurement program. Run the test, incorporate the result into the channel ROI estimate, update the model, and schedule the next test for a different channel or a different spend level. Incrementality testing is not a one-time exercise that definitively answers the question. It is a measurement practice that progressively reduces uncertainty about which channels are producing genuine revenue growth.
The framework for getting started
- Start with the channel where you are most uncertain about incremental value. The channels that deserve the first incrementality test are not the ones with the highest attributed revenue. They are the ones where the gap between attributed and incremental revenue is most likely to be large. Retargeting, brand keyword bidding, and affiliate channels are consistent candidates. The first test should answer the question that would produce the largest budget reallocation if the answer came back differently than expected.
- Use pre-test data to validate geographic comparability before randomizing. Before running the test, compare the pre-test conversion rates and revenue trends across your proposed geographic units. If the units that will serve as control regions look systematically different from the test regions before the intervention, the test result will be contaminated by that baseline difference. Matching on pre-test trends is not optional.
- Commit to the test window before seeing any data. Define the test duration based on the power calculation and do not extend or shorten it based on early results. Early results in geo-holdout tests are especially noisy. Adjusting the window after seeing early data is the equivalent of changing the hypothesis after observing the outcome. Commit to the design before the test starts.
The real advantage of incrementality testing
Companies that build an ongoing incrementality testing practice know something their competitors operating on attribution data do not: which channel spend is producing new revenue versus claiming credit for revenue that was already coming. This is a structural informational advantage, and it compounds over time.
The advantage is not that the tests are always right. It is that the tests are causally grounded in a way that attribution models are not. When a test result and an attribution report disagree, the test result is correct. Applied consistently, that discipline produces better allocation decisions with every budget cycle.