Soku AI
All blog posts

Marketing Mix Modeling for Paid Media Teams

August 12, 2026 · 16 min read

Soku Team

Soku Team

Marketing Mix Modeling for Paid Media Teams

Marketing mix modeling has been sold to performance teams twice in five years: first as the answer to iOS 14, then as the answer to cookie deprecation. Both times the pitch was the same — user-level tracking is dying, so model the aggregate instead — and both times a lot of teams bought a model that could not answer the question they were actually asking.

MMM is a genuinely good tool. It is also a tool with a narrow, specific job, a large minimum data requirement, and a failure mode where it returns a precise-looking number that is wrong in a direction you cannot detect from inside the model. This is a working guide to which of those you are dealing with.

What MMM actually is

Strip away the vendor language and marketing mix modeling is a regression on aggregate time-series data. You take a history of periods — usually weekly, sometimes daily — and for each period you record what you spent on each channel, what else was happening (price, promotion, seasonality, distribution, competitor activity, the weather if you sell umbrellas), and what outcome you got. Then you fit a model that explains the outcome as a function of the inputs.

The output is a set of coefficients: roughly, "each additional 1,000 spent on this channel in this period was associated with this much incremental outcome." From those you derive contribution, marginal return, and — if you push the model that far — a recommended reallocation.

Two things follow immediately from that description, and almost every MMM misunderstanding traces back to ignoring one of them.

It never sees a person. MMM has no user IDs, no click IDs, no cookies and no device graph. This is its headline advantage — it is completely unaffected by ATT, cookie deprecation, consent rates, ad blockers or walled gardens — and it is also its hard ceiling. A model that never observes an individual cannot tell you anything about an individual. It cannot build an audience, it cannot tell you which creative to cut, and it cannot tell you what to do on Tuesday.

It is correlational, fitted to whatever variation exists in your history. MMM does not run an experiment. It observes that spend moved and outcomes moved and estimates a relationship. If your spend never varied, there is no variation to learn from, and the model will still return a number.

The question MMM is honestly for

MMM answers budget allocation questions at the channel level over a strategic horizon. Concretely, it is the right tool for:

  • How should we split next quarter's budget across channels?
  • Is our fifth channel contributing anything, or is it absorbing budget that would work harder elsewhere?
  • What is the marginal return on the next 100,000 into Meta, given we are already spending 400,000?
  • How much of last quarter's revenue would have happened anyway with no paid media at all?
  • Is brand spend, which produces no trackable conversions, doing anything?

That last one matters more than it gets credit for. MMM is the only mainstream method that treats untrackable spend — TV, out-of-home, podcast reads, sponsorships, most brand activity — on the same footing as trackable spend. If a meaningful share of your budget is in channels that cannot produce a click, no click-based method can compare them, and MMM is not competing with your attribution stack so much as covering the part it structurally cannot reach.

Here is what MMM is not for, stated flatly because this is where teams waste quarters:

  • Creative decisions. MMM cannot see creative. If your model has one Meta coefficient, it is averaging across every ad you ran.
  • Daily or weekly optimisation. A model refreshed monthly, fitted on weekly data, with wide confidence intervals, is not an operating signal.
  • Audience or targeting decisions. No user-level data, no audience conclusions.
  • Anything within a channel. MMM says "Meta." It does not say "Advantage+ versus manual", "prospecting versus retargeting", or "this campaign."

If the question you are trying to answer sits in that second list, MMM will not answer it, and buying one will not change that.

The data requirement, honestly

This is the section most MMM sales conversations skip, and it is the one that decides whether you should be doing this at all.

You need enough periods. The rough working floor is two to three years of weekly data — 104 to 156 observations. That is not arbitrary. You are estimating a coefficient per channel, plus adstock and saturation parameters per channel, plus seasonality, plus baseline and control variables. A five-channel model can easily carry 25–40 parameters. Fitting 30 parameters to 50 observations does not produce a weak answer; it produces an overfitted one that looks excellent in-sample and is worthless out of sample.

You need variation in spend. This is the requirement people miss. If you have spent almost exactly 100,000 a month on Meta for three years, your model has no information about what happens at 60,000 or 160,000. It will still fit a coefficient. That coefficient is an extrapolation dressed as an estimate, and its confidence interval will be enormous — which you will only notice if your vendor reports confidence intervals at all.

You need the non-media variables. Price changes, promotions, stockouts, distribution changes, PR events, competitor launches, seasonality. Omit a variable that moved with your spend and its effect gets attributed to the channel. The classic case: you spend more in Q4, you also discount in Q4, the model has no discount variable, and paid social gets credit for the discount. The number will look plausible and be badly wrong.

A practical sizing test before you commit: count your weeks of clean history, count your channels, and ask whether spend on each channel has moved by at least ±30% at some point. If you have 60 weeks, eight channels and flat budgets, no vendor and no open-source library will rescue you. Fix the data first — and in the meantime, run geo experiments, which need far less history.

Adstock and saturation, in plain arithmetic

Two transformations do most of the real work in an MMM, and understanding them is what lets you challenge a model's output rather than accept it.

Adstock encodes the fact that advertising has memory. Money spent this week keeps working next week, decayingly. The simple geometric form is:

effect_this_week = spend_this_week + λ × effect_last_week

where λ is the decay rate between 0 and 1. A λ of 0.3 means about 30% of last week's effect carries over — fast decay, typical of lower-funnel search. A λ of 0.8 means a long tail, more typical of TV or brand video.

This one parameter has an outsized effect on the conclusion. A channel with a high fitted λ accumulates credit across many weeks and will look strong. If λ is fitted freely on thin data, it can drift to a value that flatters a channel for reasons that have nothing to do with the channel. Ask what λ was fitted for each channel, and whether it is plausible against how the channel actually works. A search coefficient with an adstock half-life of two months should stop the meeting.

Saturation encodes diminishing returns — the tenth thousand does less than the first. It is usually a Hill or exponential curve mapping adstocked spend to effect. Saturation is where the actual budget recommendation comes from: reallocation advice is the model saying you are further up the flat part of one curve than another.

The catch is the same as before, in a sharper form: the curve is only observed where you have spent. If your Meta spend has ranged over 80,000–120,000 a month, the model has fitted the curve over that band. Its opinion about 300,000 a month is extrapolation. A recommendation to triple a channel's budget is almost always a statement about curve shape outside the observed range, and it should be treated as a hypothesis to test with a real budget increase, not an instruction to execute.

Why MMM disagrees with your ad platforms — and why that is correct

Your MMM will attribute fewer conversions to Meta than Meta does. This is not a bug and it is not evidence that either is broken. They are answering different questions.

Meta answers: of the conversions that happened, how many came from people who interacted with a Meta ad within the attribution window? That is a correlational, platform-observed, self-reported number, and it counts people who would have converted anyway.

MMM answers: how much of the total outcome moved because spend moved? That is an incrementality question, and it excludes people who would have converted anyway.

The gap between them is mostly baseline demand your platforms are claiming credit for. Retargeting is the extreme case: platform-reported ROAS on retargeting is routinely spectacular and MMM-estimated incrementality routinely is not, because retargeting reaches people who already showed intent.

Three practical consequences:

  1. Never sum platform-reported conversions across channels and compare the total to actuals. You will double-count, and the overcount is largest exactly where the platforms overlap most.
  2. Do not calibrate MMM to match platform numbers. Some vendors will offer this. It defeats the purpose — you are re-introducing the bias you built the model to avoid.
  3. Do calibrate MMM to experiments. This is the legitimate version, and it is the single highest-leverage thing you can do to an MMM.

Calibrate with experiments, or do not trust the coefficients

An MMM fitted on observational data alone is a set of educated guesses with confidence intervals. An MMM anchored to experimental results is a measurement system.

The mechanism is straightforward. Run a geo holdout: split comparable regions into test and control, turn a channel off or scale it sharply in test, hold control steady, and measure the difference in outcome. That gives you a causal estimate of incremental effect for that channel, in that period, at that spend level — with no modelling assumptions at all.

Then use it two ways. Before fitting, as a prior that pulls the coefficient toward the experimentally observed value. After fitting, as a check: if the model says a channel returns 4x and a clean geo test says 1.5x, the model is wrong and you now know it.

The reason this matters so much is that an MMM cannot validate itself. In-sample fit statistics measure how well the model explains history, and an overfitted model explains history beautifully. Holdout validation on time-series data helps but is weak when you only have 150 observations. External experimental evidence is the only thing that genuinely tests whether the coefficients mean anything.

If you are choosing between spending a quarter standing up an MMM and spending a quarter running three clean geo experiments, and you have fewer than two years of varied history, run the experiments. They answer narrower questions, and they answer them correctly.

Four ways MMM produces confident wrong answers

Omitted variable bias. Something you did not model moved with your spend, and its effect landed on a channel. Promotions and price are the usual culprits, seasonality the most common in retail. The output looks entirely normal.

Collinearity across channels. If you always raise Meta and TikTok budgets together, the model cannot separate them. It will still hand back two coefficients. Those coefficients are not individually meaningful, and their split can flip between refits with no change in the underlying business. Watch for a channel's contribution swinging wildly between monthly refreshes — that is the signature.

Extrapolating beyond observed spend. Covered above, and worth repeating because it is the most common way a model produces an actively harmful recommendation.

Treating the model as an operating system. MMM refreshes monthly or quarterly and answers strategic questions. Wiring it into weekly budget decisions applies a slow, wide-interval, channel-level signal to a fast, narrow, campaign-level problem. The correct architecture is layered: MMM for the quarterly channel split, experiments to validate and to answer specific causal questions, and platform data plus your own analytics for the daily work inside a channel. We covered how those layers fit together in cross-channel marketing attribution, and how MMM ranks against the alternatives by required volume in multi-touch attribution models compared.

Build, buy, or open source

Open source — Meta's Robyn, Google's Meridian, PyMC-Marketing. Free, transparent, and genuinely capable. The cost is that you need someone who can defend the modelling choices: priors, adstock parameterisation, saturation form, validation strategy. If nobody on the team can explain why a particular adstock decay was chosen, open source becomes a black box with extra steps.

Vendors — you are buying data pipelines, a refresh cadence, benchmarks and someone to sit in the meeting. The two questions that separate good from bad: do you report confidence intervals on every coefficient, and how do you calibrate to experiments? A vendor that reports point estimates with no intervals is selling false precision. A vendor with no experimental calibration story is selling a regression.

Building in-house only makes sense at large, sustained spend with a dedicated analyst. The modelling is the easy part; the data engineering — clean, consistent, weekly spend and outcome history across every channel with all the control variables — is where the effort actually goes, and it does not go away after launch.

An honest decision rule

Do MMM if: you have two-plus years of weekly history, spend has genuinely varied, you run more than three or four channels, a meaningful share of budget is in untrackable media, and you can run geo experiments to calibrate.

Do not do MMM if: you have under a year of data, budgets have been flat, you are on one or two channels, or what you actually want is a creative, audience or weekly-optimisation answer. In every one of those cases the money is better spent on incrementality testing and on getting your first-party conversion data clean — which is also what makes a future MMM possible.

MMM is a strategic instrument. Used for the quarterly channel split, calibrated against experiments, and reported with its uncertainty intact, it is the best available answer to a question nothing else can answer. Used as a weekly dashboard, it is an expensive way to be confidently wrong.

Related Tools

Related Use Cases

Relevant Reads

We use essential cookies to operate and secure Soku. With your permission, we also use optional analytics and advertising cookies to measure usage and campaigns. You can change your choice at any time. Privacy Policy