By the Odins team · Updated October 2026 · About 9 minutes
A marketing mix model is only as good as the data under it. Most MMM projects that stall do so on data, not on statistics: TV spend sitting in an agency spreadsheet, channel names that change every quarter, sales figures that arrive by email. This guide covers what data marketing mix modelling needs, where to get it, how to structure it and which problems ruin a model before it starts.
The short answer
You need four things, over the same period and at the same time grain:
- Marketing spend by channel, digital and offline, ideally daily or weekly.
- The outcome you want to explain: sales, orders, applications, sign-ups or paid-out loans.
- What else moves that outcome: prices, promotions, product launches, distribution changes, seasonality and one-off shocks.
- A structure that puts all of it in the same hierarchy of markets, products and channels.
Two years of weekly history is a comfortable floor, because the model needs to see each season at least twice. Three years is better. Six months can work if strong priors do more of the work, with wider uncertainty as the price.
The model does not read any of these rows on its own. It learns from how they move together: sales rising a week after a TV burst, search spend that tracks a promotion, a dip that lines up with an outage rather than with anything marketing did. That is why the four have to sit on one timeline, and why a gap in any of them costs more than it looks.
1. Digital channels: connect, do not export
Search, social, video, display, programmatic and affiliate platforms all have APIs. The reliable way to get their data is a direct, read-only connection that pulls the full history once and then refreshes every day. Spreadsheet exports work for a pilot, but they break the moment someone changes a column or forgets a month.
Pull more detail than you think you need. With campaign and ad-level data you can split a large channel later, for example Meta awareness versus Meta conversion, without going back to the source. The model then decides how granular it can be based on how much was spent in each part.
Odins maintains connections to Google Ads, Meta, TikTok, Snapchat, LinkedIn, Microsoft Ads, programmatic platforms and affiliate networks, among 600+ integrations, and builds a new connector when a customer uses a platform with an API that is not covered yet. See the connector overview.
2. Offline media: get the files the agency already has
TV, radio, outdoor, print, influencers and podcasts have no API. This is where most teams get stuck, and it is also where MMM adds the most, because these channels are invisible to click-based attribution.
The data you need usually exists already, in your media agency's systems. Ask for two things per channel: the invoice, which gives you the cost, and the spot or airing report, which gives you the timing. For TV, a spot report with the time each spot aired and the cost per spot is ideal. For radio, outdoor and print, a weekly or monthly spend breakdown is enough.
When you send the request, be specific, or you will get a summary that cannot be modelled:
- Everything they hold, going back at least three years.
- The raw export from their system, not a reformatted summary.
- Whether figures are gross or net, and whether agency fees and production costs are included.
- Whether the numbers are booked or actually invoiced.
- Which market and currency each line belongs to.
Send the request yourself
Agencies answer to their client, not to your vendor. With a cooperative agency, the history can be in place within a day. Odins then sets up a pipeline for the recurring files, so offline spend lands in the same structure as digital without a monthly chore on your side. More on this in offline media data.
3. Sales and outcome data: automate it, and model what matters
The outcome feed is the one teams most often leave manual. Automate it from day one. Daily is best, weekly works, and monthly is the minimum. Fresher data means the model can be checked against reality sooner, and forecasts can be followed up as the month unfolds.
Choose the outcome that matters to the business, not the one that is easiest to count. A lender should model paid-out loan volume, not application clicks. A subscription business should model new subscriptions, not site visits. If the outcome happens days or weeks after the marketing, the delay can be set in the model so cause and effect line up.
Remember that the model only sees what is in the outcome data. If a large share of sales happens in stores, through wholesale or through advisers, that data has to be included, or the model will understate what marketing does. Which outcome to pick for your kind of business is covered in MMM by business model.
4. Structure: one hierarchy for every source
Raw data from twenty sources uses twenty naming conventions. Before anything can be compared, it has to land in one structure. A practical order is market first, then product, then segment, and then a media hierarchy that goes as deep as the data allows: search, then Google, then Performance Max; TV, then TV 2, then the individual spot.
The point is that everything lands in the right place in one detailed hierarchy, online and offline together. Once that exists, cross-channel reporting is a by-product, and the model reads from the same dataset the marketing team looks at. There is one version of the numbers instead of three.
5. The data problems that ruin models
Most of these are invisible in a dashboard and obvious in a model. The model cannot tell a missing week from a week with no spend, or a renamed campaign from a new one, so each of them has to be caught before training.
| The problem | What the model does with it | The fix |
|---|---|---|
| Mixed gross and net figures across channels | Makes one channel look cheaper than it is, and the budget moves toward it. | State gross or net per source, and convert once, in the pipeline. |
| Missing weeks | Reads them as zero spend, and credits that week's sales to the baseline or to another channel. | Flag gaps and fill them from the source, not with averages. |
| Renamed campaigns and changed tracking | Breaks a channel's history into pieces that look like separate channels. | Map old and new names to the same node in the hierarchy. |
| Too little spend in a channel | Returns the prior, with a wide range, and calls it an estimate. | Group small channels with similar ones rather than forcing a separate estimate. |
| Always-on activity with no variation | Cannot separate it from the baseline, so its effect is unknown either way. | Vary the spend on purpose, or test it, before asking the model about it. |
| Known shocks that are not flagged | Reads an outage or a storm as a marketing effect, and forecasts it forward. | Mark the dates as events so the model does not learn from them. |
6. What arrives when
A realistic plan runs on three clocks. Digital channels connect on day one, because access is a link and an approval. Offline history depends on your agency and lands in the first weeks. The sales feed is a few days of work once the right person at your end is involved. Reporting on the structured data is useful from about week four, and the first model follows at six to eight weeks.
7. Give the data back to your own warehouse
Once marketing data is collected and structured, it is useful well beyond the model. Ask your vendor to deliver the cleaned dataset into your own warehouse, such as BigQuery or Snowflake, so your BI tools read from the same numbers. Some Odins customers use the platform for this alone: one structured dataset of all marketing spend and results, digital and offline side by side, in their own hierarchy. The model can then be added when budget and volume justify it.
Why this matters for vendor choice
A vendor that keeps the structured dataset inside its own product has you by the data, not by the model. Ask who owns the cleaned data and whether it can leave. The wider list of questions is in how to choose an MMM vendor.
Frequently asked questions
We only have two years of data. Is that enough?
Yes. Two years of weekly data is the comfortable floor for most businesses, because the model sees each season twice. With less history, priors carry more of the result and the uncertainty ranges are wider.
How do we get TV and radio data out of our media agency?
Ask for the invoice and the spot or airing report per channel, as raw exports, at least three years back, with gross or net, fees, booked or invoiced, and market stated explicitly. Send the request yourself, then have your vendor set up a recurring pipeline.
Which MMM vendors need the least preparation from us?
Managed platforms, where the vendor connects the sources, handles offline files and structures the data. With Odins, the work on your side is mainly granting read access to ad accounts, asking your agency for offline data and connecting the sales feed.
Do we need daily data, or is weekly enough?
Weekly is enough for the model. Daily is better for the sales feed, because it lets you check the forecast as the month unfolds instead of waiting for month end. Digital platforms deliver daily through their APIs anyway; offline channels are weekly or monthly by nature.
Can the cleaned data go into our own warehouse?
With Odins, yes. The structured dataset can flow to your own warehouse and you own it.
For how the data is used once it is in place, read our complete guide to marketing mix modelling, or see how Odins handles data integration.
-1.png?width=1132&height=292&name=Asset%201@2x%20(12)-1.png)