Your AI search ROI number is 75% something else

By Mo Touzani · · Marketing Measurement

Ask ChatGPT to recommend a tool and it names a handful of brands. If yours is one of them, a buyer just met you, and your analytics recorded nothing. No impression, no click, no UTM. When that buyer returns, direct traffic or branded search takes the credit. Teams already spend real money on GEO (editing content so AI answers mention you) and GEM (paying for sponsored placements inside those answers). Ask what either returns and you get a share-of-voice chart. A paper posted to arXiv on September 10, 2026 by Masahiro Kato, Daiki Honma and Taka Kato offers the most rigorous answer so far. Read past the abstract and it carries a message GEO vendors will not enjoy.

How Generative Marketing Mix Modeling works

The authors start from a precise gap: your marketing data does not record how often AI answers name you, or whether anyone notices. Their framework, Generative Marketing Mix Modeling (GMMM), builds that missing input from four parts. Occurrence: how often answers to a cluster of questions contain your brand, estimated by asking each AI system the same questions repeatedly. Question counts: how often people in each market ask those questions. System shares: how those questions split across ChatGPT, Gemini, Perplexity and the rest. Notice: the chance a user registers your name, which the authors propose estimating with a user study.

Multiply the four and you get the expected number of noticed mentions per market and period. GMMM then applies standard MMM carryover and saturation curves and estimates the effect of switching GEO on or off across a complete sequence of periods. Paid placements get the same treatment, with spend and delivered placements replacing organic answers. The identification section is careful work: it states the conditions under which the effect can be recovered, and when it cannot.

What 2,240 real answers show

The authors asked GPT-5.6 Luna and GPT-4o 56 product recommendation questions, 28 in English and 28 in Japanese, 20 times each. Neither the system prompt nor the questions named the target brand, the note-taking tool Glasp. Glasp appeared in 33.8% of GPT-5.6 Luna answers and 27.8% of GPT-4o answers. The ranking flipped by language: GPT-4o led in English (35.4% versus 28.6%), while GPT-5.6 Luna led by a wide margin in Japanese (38.9% versus 20.2%). When GPT-4o did mention Glasp, it named it first 60.5% of the time, against 39.4% for GPT-5.6 Luna.

The practical lesson lands before any modeling starts. A single AI visibility score hides more than it shows. Your presence depends on the model, the language and the use case. If a vendor hands you one number, ask which questions, in which language, on which models.

The finding buried in Table 4

The authors simulated five markets, six product clusters each, over 36 periods, with GEO rolled out in stages, and checked whether their estimators recovered the true effect. The headline looks reassuring: the total GEO effect came back with roughly 17% error. Then look inside the total. GMMM splits GEO into two paths. One runs through more noticed mentions in AI answers. The other runs through everything else the same content changes do, such as lifting conventional search traffic and landing-page conversion.

In the simulation, the everything-else path made up 66% to 75% of the total effect. The AI-answer path, the part GEO is sold on, carried 46% to 48% error. Even an oracle that knew the true transformation parameters missed it by 31% to 35%. The total only looked accurate because the two component errors ran in opposite directions and cancelled, with a correlation near -0.6. The authors say it plainly: the total effect can conceal larger errors in its two model components.

If you improve your FAQ pages, documentation and product content, most of the measurable payoff in this setup arrives through channels you already measure. The AI-specific slice is the noisiest number in the model.

The real-world test the authors refused to oversell

An appendix analyzes a public referral dataset from Glasp's own repository, covering July 2025 to May 2026, comparing treated and control sessions around a GEO change made in late December 2025. With December 30 as the start date, referral traffic rose by a factor of 1.84. Move the start date to December 16 and the ratio drops to 1.18, indistinguishable from zero effect. Move it to January 6 and it climbs to 2.40. Adjust for a bot-filtering change in March and it falls to about 1.34, with an interval that includes no effect.

A placebo check found swings this large before the change too. The authors conclude the analysis does not by itself establish a treatment effect. That is honest science. Expect a vendor deck to quote a 1.8x AI referral lift and skip the next paragraph.

What the paper gets right

Scale errors matter less than you would fear. A formal result in the paper shows that if your notice probability or question volume is off by the same factor everywhere, the saturation curve absorbs the error and the effect estimate holds. Only errors that vary across markets, periods or question types distort the answer. You do not need a perfect notice estimate. You need a consistent one.

Two more lessons travel beyond GEO. Most of the accuracy gain in the simulations came from estimating carryover and saturation from data instead of fixing them at default values, which applies to every MMM you run. And the authors are explicit that causal interpretation comes from how treatment is assigned. A staggered rollout of content changes across product lines or regions gives any model something to learn from. One sitewide update does not.

Where it falls short for a real company

The simulations hand the estimators the hardest inputs. Question counts and system shares were known, and every cell had 50 notice observations. No AI platform publishes prompt volumes today, and few companies have run a notice study. The simulations also omit established media. In your data, paid search changes and content changes often move together, and the paper notes that this kind of overlap can make their effects impossible to separate. The design assumes a structured rollout across many units, and every real data point concerns a single brand, Glasp, with referral data from Glasp's own repository.

What to do this quarter

  1. Track AI presence by model, language and use case. Run a fixed question set monthly, at least 10 repetitions per question per model, and record mention rate and rank. Report the breakdown, never a blended score.
  2. Stagger your content changes and log the dates. Roll FAQ and documentation updates out to some product lines or regions first. That single design decision makes any later analysis, GMMM or a simpler comparison, far more credible.
  3. Measure the ordinary channels first. In the paper's simulation, most of GEO's effect flowed through conventional search and conversion. Check those before crediting AI answers with anything.
  4. Run a small, consistent notice test. Show 30 to 50 customers real AI answers that mention you and record whether they register your name. Keep the method fixed; consistency matters more than precision.
  5. Make vendors split the lift. Ask any GEO vendor how much of their claimed lift comes through AI answers versus everything else, and the error on each part. No answer, no budget.

The bottom line

This is the most serious measurement framework for AI search yet, and its own results argue for caution. The AI-answer slice of GEO is real, small and hard to estimate. The content work GEO requires earns its keep mostly in channels you already measure. Fund better content because it improves search and conversion. Treat any AI-specific ROI figure as unproven until someone shows you its error bars. Source: Kato, Honma and Kato, Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact, arXiv:2609.11915 (2026).

Book a consultation with Sinfa · [email protected]