Your attribution model has never run a real experiment
By Mo Touzani · · Marketing Measurement
The industry spent five years diagnosing the wrong problem. Every attribution postmortem after iOS 14.5 pointed at privacy regulation as the villain. That was a convenient story. The actual story is that marketing measurement was broken long before Apple changed its tracking defaults, and the privacy shift turned on the lights in a room where nothing had ever been measured correctly.
The problem that predates privacy
Most attribution models have never produced a causal measurement in their operational history. They ingest historical spend data, find periods where advertising activity and revenue moved in the same direction, assign weights, and call the output attribution. Correlation with a confidence interval is still correlation. The model does not know whether the advertising caused the revenue or whether both were responding to the same underlying factor.
When platforms lost their deterministic identity graphs after iOS 14.5, modeled attribution rushed into the gap. Suddenly the industry had two sets of estimates that disagreed with each other and no principled method to resolve the conflict. The uncomfortable realization: both had always been estimates. The disagreement did not create the problem. It made the problem impossible to ignore.
This is not a data crisis. It is a methodology crisis. The measurement problem predates cookies, pixels, and device fingerprinting by decades. Running a controlled experiment to isolate the causal contribution of a marketing channel is hard. Inferring causation from observational data is comforting. The industry chose comfort, and the bill is now being presented.
To know whether your spend caused the outcome, you need an experiment. There is no statistical shortcut that replaces a controlled test.
DeepCausalMMM and the confidence trap
DeepCausalMMM is an interesting piece of engineering. The architecture addresses real limitations of linear regression and older Bayesian marketing mix models: neural networks that learn channel interdependencies, saturation curves that flex rather than linearize, train-test gaps that narrow toward 90 percent holdout accuracy on published benchmarks. The technical progress is real.
The framing around that progress is where the trouble starts. Better architecture on an observational dataset does not produce causal measurement. It produces more confident observational measurement. Confidence and accuracy are not the same thing, and the difference matters enormously when the output is driving budget decisions.
The training data contains a time series of spend across channels alongside revenue outcomes, with no experimental variation, no geo holdouts, and no channel blackouts used as controls. The model learns the correlations embedded in that history with more nuance than a linear regression would. But the correlations it learns are still correlations. A more sophisticated model pointed at observational data fails more convincingly, not more accurately.
A board will look at tighter uncertainty intervals and narrower confidence bands and interpret them as evidence that the model is right. Those intervals describe fit to historical patterns. They say nothing about what would happen if you cut a channel, doubled another, or entered a new market. Those are causal questions. They require causal evidence.
What channel ROI numbers measure
One published DeepCausalMMM case study reports a channel with marginal ROI near 0.13 sitting alongside another channel near 0.73. A six-to-one spread. That number will move budget in a board meeting. It probably should not.
What the number represents: the estimated marginal return for that specific company, at those specific spend levels, during the period captured in the training data, under the competitive conditions that prevailed at the time. Change any of those variables and the number changes. The model has no mechanism to account for this. It can only report what the historical pattern looked like.
Channel ROI figures from marketing mix models are local measurements, not universal rankings. The return on paid search at $200,000 per month is a structurally different figure from its return at $800,000 per month. Saturation curves are nonlinear. A channel that appears inefficient at current spend levels may be generating demand that converts through organic channels downstream. The model attributes that conversion to whatever channel touched it most visibly in the training data. It cannot run the counterfactual.
Attribution models, including sophisticated MMM variants, structurally undercount brand-building channels and overcount direct-response channels. Direct-response creates observable signals in the training data. Brand effects diffuse across time and resist clean attribution. Cutting what the model flags as weak and scaling what the model flags as strong is a hypothesis. It needs to be treated like one.
Treating a static ROI figure as a durable strategic truth is how companies reallocate significant budget based on historical correlations and then spend two quarters asking why revenue did not respond as the model predicted. The model predicted what it observed in the past. The past does not repeat under identical conditions.
The framework for budget decisions
- Ask what percentage of the underlying data came from controlled experiments. Before acting on any model output, ask one question: how much of the training data came from geo-holdout tests, randomized channel experiments, or incrementality studies versus uncontrolled historical spend? If the answer is zero or unclear, the model is producing directional pattern-matching, not causal measurement. Treat its outputs as hypotheses that need a test design, not conclusions that need a budget adjustment.
- Treat every channel ROI number as a local estimate with a short shelf life. Channel-level ROI figures from any model are snapshots of what performed at a specific scale, during a specific period, under specific market conditions. Those conditions change when your spend changes, when a competitor enters or exits, or when your creative mix rotates. A material budget reallocation should specify in advance the expected revenue impact and the timeline for evaluating it. That creates an objective basis for determining whether the model was right, rather than a post-hoc narrative built around whatever happened.
- Build the right benchmark for your actual cost structure before optimizing against it. AI-native companies, usage-based SaaS businesses, and companies with high variable acquisition costs operate under different margin structures than the benchmarks that dominate growth conversations. A 3:1 LTV-to-CAC ratio on 52 percent gross margins is a structurally different business from the same ratio on 80 percent margins. Optimizing channel mix against a benchmark designed for a different cost structure produces precise movement in the wrong direction. Calibrate the target first.
What the next two years reward
DeepCausalMMM is better architecture. That is straightforwardly true. For a team that has accepted the constraints of observational measurement and is working within them deliberately, the tool is an upgrade. Better architecture extracts more signal from existing data, surfaces nonlinearities that linear models miss, and produces richer output for hypothesis generation.
The companies that win on paid growth over the next two years are not going to be the ones running the most sophisticated models on unvalidated data. They are going to be the ones that treat experimental validation as a standing operational discipline: geo-holdout tests run alongside MMM output, incrementality studies that challenge the model's channel rankings before budget moves, and results used to calibrate whether the model is tracking reality before it is given authority over allocation.
Marketing measurement has a confidence problem, not a sophistication problem. The current generation of tools makes it easier than ever to produce authoritative-looking output from observational data. The organizations that understand the difference between confidence and accuracy will use those tools to generate hypotheses and run experiments to test them. The rest will allocate budget with precision toward the wrong channels, at the wrong scale, and remain puzzled when the numbers refuse to move.