3rd-party validated lift for retail media campaigns.
Northbend Snacks, unit sales. In-store screens, 6-week flight.
at a 95% confidence interval
Counterfactual built from 124 matched stores that ran no media.
VALIDATED BYSeurat AnalyticsIllustrative figures. Not a client result.
Example output with analyzed store count and confidence range shown, produced by an independent analytics provider instead of the media seller. Lift calculations based on geo-targeted exposure + constructed counterfactual: in-store screens, forecourt media, and onsite retail media that ran in some stores and not others.
How we calculate it, in three steps
Take the outcome at the register
Unit and dollar sales of the advertised item and of its category, at store level, weekly. Not an impression, not a click, not a matched loyalty ID.
Build the comparison from stores that ran no media
Stores in the same chain with no exposure during the flight, weighted so that together they reproduce the exposed stores' sales path in the weeks before the campaign.
Difference the two
What the exposed stores did over the flight, minus what the matched stores did over the same weeks. The gap is the lift, and it is reported with a confidence interval rather than as a single number.
Media buyers have changed what they ask for
Roughly 70% of advertisers now name incrementality as the most important measure of a retail media investment, and over 70% say it is the hardest thing to measure. Just over half cite the lack of a common standard as the single biggest barrier to investment.1
The largest retail media networks have already responded. Several now put an outside firm's name on their lift results and promote this in their sales material. One of the largest states plainly that most networks “grade their own homework,” and uses 3rd-party validation as evidence that it does not. That sets the standard every other network gets measured against, no matter its size.
A network's own reporting is not wrong. It is just harder to trust. A number produced in-house settles the question for the ad network but leaves the advertiser uncertain.
1 Association of National Advertisers. Incrementality figures from the ANA Retail Media Survey (January 2024); the standardization barrier from ANA Retail Media Measurement Standardization (August 2026).
What a validated number does for the business
In an RFP
Against a larger network, an outside validated number closes the one gap you cannot close with reach.
At renewal
It answers the question a retail partner actually asks, which is whether the screens grew the category or moved share along the shelf.
In the rate card
Measured campaigns become a tier you can price above unmeasured inventory.
In planning
The detection floor tells you which campaigns are large enough to produce a claim worth publishing, before you commit to one.
A result fit for publication
- A lift estimate with the store counts and the range printed beside it.Not a single number with nothing under it.
- The comparison group named and fixed before the outcome window closes.Which stores stood in for what the exposed stores would have sold had the campaign never run, written down and dated while the result is still unknown.
- A one-page summary written to be forwarded.To a brand, a retail partner or an investor, without needing you in the room to explain it.
- A ‘Validated by Seurat’ mark you can put on it.Your campaign, your client, your story. Our name on the method and the findings.
- The full method, available to anyone who asks.Including anyone who expects to disagree with the result.
The comparison group that makes causal attribution possible is already there. It just wasn't analyzed.
Retail media has an advantage most channels do not: exposure separates by store, so a comparison group already exists. Nobody picked it at random, though. Stores with screens are larger, newer, better sited, or further along a rollout schedule, so comparing them raw measures the difference between the stores rather than the difference the campaign made.
So we weight the unexposed stores until they reproduce the exposed stores' pre-campaign sales path, then difference the two over the flight. Where the rollout was phased, we use an estimator built for staggered adoption. The assumption doing the work is that the two groups would have kept moving together, and we test it rather than assert it.
Screens went into the bigger, newer, busier stores. Nobody picked them at random.
Comparing them raw measures the difference between the stores, not the campaign.
We weight the unexposed stores to reproduce the exposed stores' pre-campaign path, then difference the two.
The specification →
The estimator is difference-in-differences on weekly store-level sales, with store and week fixed effects and standard errors clustered at the store. Where the rollout was phased, we use a staggered-adoption estimator rather than two-way fixed effects, following Callaway and Sant'Anna (Journal of Econometrics, 2021) and the decomposition in Goodman-Bacon (Journal of Econometrics, 2021). Where the exposed set is small, we weight the unexposed stores to reproduce the pre-period path, following Abadie, Diamond and Hainmueller (Journal of the American Statistical Association, 2010).
The outcome definition, the category definition, the control store list, the pre-period length and the flight window are fixed in writing and dated before the outcome window closes. The output is a point estimate with a confidence interval, not a single number.
Three things from you: the flight dates, the store list, and weekly sales.
| What you provide | What it enables in the analysis | |
|---|---|---|
| Flight dates | Must have | Fixes the before and after windows, the same way for every store. |
| Store list, exposed and unexposed | Must have | Separates the two groups and sets the pool the comparison is built from. |
| Weekly sales by store | Must have | The outcome itself. The lift is measured here. |
| Item-level sales alongside the category | Optional | Roughly halves the smallest lift we can find, by measuring the brand against its own category rather than against store traffic. |
| Promotion calendar by store-week | Optional | Rules out the item having been on price promotion in the same stores that ran the screens. |
| Screen hours by store | Optional | Powers the dose test: whether measured lift rises with how much exposure each store got. |
No impression count, no click, no conversion, and no ad-server or ad-platform-reported metric enters the model at any point, including yours. There is no input for either side to adjust.
The detection floor calculator prices what each optional input is worth before you ask your retail partner for it.
What size lift could your own network detect?
Every design has a floor: the smallest lift it could have found at all, set before the read by your store counts, flight length and how volatile the category is. It is the only way to read a null correctly, because a flight that comes back with no measurable lift either produced nothing or produced something too small for this design to see. This is a quick version of the calculator that finds yours.
Below this, the read returns a wide range rather than a negative. That is a different finding, and worth knowing before you commission anything.
Assumes twice as many comparison stores as exposed ones, and 26 weeks of pre-period history.
You see the results first.
- You see every result before anyone else.A read that does not clear is a private finding, and nothing goes out that you have not approved.
- We test for effects that cannot exist.Every read runs against dates and categories where nothing should be found. If the method finds something there, you hear it from us.
- We will not claim more than the design supports.That is what the confidence interval in the findings report is for.
- The fee is fixed before the read.It does not move with the direction or size of what we find.
A number that could only ever come back positive is not one a brand's analyst will accept.
Does the method find effects in weeks when nothing ran?
It must not. A read that fails this test is not delivered.
The detail →
We run the identical estimator on quiet windows, with the same exposed and control stores and no campaign in the window. The model should find nothing. If it reports lift on quiet weeks it is picking up noise, or the two store groups are drifting apart for reasons that have nothing to do with advertising, and any number it produces on a real flight is worthless.
Does it find effects in a category that was never advertised?
Also no. This is the parallel-trends assumption, tested rather than asserted.
The detail →
We run the same estimator on a placebo category: a category sold in the same stores, over the same weeks, with no exposure in the flight. If it moves, we are catching a traffic shock, a weather effect, or a promotion running alongside, rather than anything the screens did. This is the assumption the method depends on most and the one most likely to be violated in practice, so we test it rather than assume it.
Were the two groups moving together before the campaign?
They have to be, and we publish the chart that shows whether they were.
The detail →
Before the flight, the exposed stores and the weighted control stores should track each other. We plot the pre-period side by side and publish it with the read. A visible gap opening before the campaign started is a failed match, and we say so instead of reporting a number the match can't support.
Does the effect scale with how much exposure each store got?
This is the test that matters most.
The detail →
Stores do not receive identical exposure. Screen-hours, loop frequency and impressions delivered vary across the network, sometimes by a lot. So we can ask whether measured lift rises with exposure per store. If it does, confounding becomes very hard to sustain as an explanation: a competing cause would have to be correlated with screen-hours store by store, timed to the flight, and proportional in magnitude at each one. Coincidence does not usually arrive with a gradient attached. This is the same logic epidemiology uses when randomization is not available.
Whether the campaign grew the category, or moved it across the aisle.
With store-level sales for the advertised item and for its whole category, we can ask where the lift came from. If the brand rose and the category rose with it, the campaign brought demand into the store. If the brand rose and the category did not, it moved share from something else on the same shelf. Those are different numbers, they are worth different amounts, and they matter to different parties: the brand cares about the first reading, the retailer about the second.
Counting the sales that followed an exposure can't tell the two apart, because some of those sales would have landed at the register had the ad never run. That is why attribution numbers get defended by the party that produced them and discounted by the party that received them, and why both sides often have trouble settling the argument.
Incremental. The campaign brought demand into the store.
Substitution. The campaign moved share from something else on the same shelf.
What “Validated by Seurat” means

Seurat Analytics
Three things stand behind it. You can check all three.
Which stores form the comparison group, the outcome definition, the category definition, the pre-period length, the estimator and the flight window are all written down and dated while the result is still unknown. The falsification tests are named before they are run, so a test cannot be dropped for producing an inconvenient answer. Choosing any of it after seeing what each version produces is the single failure that makes most observational retail media results worth nothing.
It is not a share of measured lift, not contingent on a positive result, and not renewable on the strength of one. This is the same rule that governs a fairness opinion in a merger, for the same reason. A fee that moves with the answer has bought the answer.
Anyone who wants to check a read can have the specification and the code it ran on, including anyone who expects to disagree with the result. With the retailer's permission that extends to the data itself.
Seurat's approach vs. other methods.
What each one is for, and how far each one travels.
| Method | Where the number comes from | What you need to run it | What it tells a brand | Travels to |
|---|---|---|---|---|
| Network attribution | Your ad server and the retailer's loyalty file. | Nothing extra. It is already running. | Which sales followed an exposure. | Inside your own reporting. |
| Retailer closed-loop measurement | The retailer's POS and loyalty data. | A retail partner willing to run it. | Whether exposed shoppers bought more than similar unexposed shoppers. | To that retailer's brand partners. |
| Designed in-store test programs | The retailer's POS, through a test built before the flight. | The design in advance, and stores deliberately held back. | What to roll out chain-wide next year. | Widely, when there was time to build it. |
| Panel and mix modeling firms | A proprietary panel, or the advertiser's spend history. | An advertiser that already retains one. | How a brand should allocate across channels. | Inside that brand. |
| Seurat matched-store read | Store-level POS, on a specification fixed before the outcome window closes. | Flight dates, a store list, and weekly sales. | Lift, with the store count and the range beside it. | To brands, retail partners and investors, because the party reporting it did not sell the media. |
When Seurat Analytics is the right tool
- Optimizing loops and dayparts?Use the ad server. That is what it is for.
- Choosing what to roll out chain-wide next year, with time before the flight?Design the test.
- The flight already ran?This is the read. A holdout you did not build cannot be recovered after the fact, and we can still settle the campaign from the sales data.
- A brand, retail partner or investor wants a number that did not come from the seller?That is the other reason to reach for it.
Selling creator media, or any campaign whose audience is distributed by interest rather than by market? That exposure does not separate by store, so it gets a different read: the public marketing data read.