NEW RESEARCH79% of shoppers say the same chain is noticeably better or worse store to store.Get the full report, free →

New research: why the same chain wins in one store and leaks in another.Get the report →

Read B · For in-store, forecourt, and onsite retail media

3rd-party validated lift for retail media campaigns.

MATCHED-STORE READILLUSTRATIVE

Northbend Snacks, unit sales. In-store screens, 6-week flight.

+4.8%lift

at a 95% confidence interval

Counterfactual built from 124 matched stores that ran no media.

EXPOSED STORES62
COMPARISON STORES124
FLIGHT6 weeks
SMALLEST DETECTABLE2.6%
Quiet weeks: no effect found
Placebo category: no effect found
VALIDATED BYSeurat Analytics

Illustrative figures. Not a client result.

Example output with analyzed store count and confidence range shown, produced by an independent analytics provider instead of the media seller. Lift calculations based on geo-targeted exposure + constructed counterfactual: in-store screens, forecourt media, and onsite retail media that ran in some stores and not others.

Covered by this read
In-store screensForecourt mediaOnsite retail media

How we calculate it, in three steps

01

Take the outcome at the register

Unit and dollar sales of the advertised item and of its category, at store level, weekly. Not an impression, not a click, not a matched loyalty ID.

02

Build the comparison from stores that ran no media

Stores in the same chain with no exposure during the flight, weighted so that together they reproduce the exposed stores' sales path in the weeks before the campaign.

03

Difference the two

What the exposed stores did over the flight, minus what the matched stores did over the same weeks. The gap is the lift, and it is reported with a confidence interval rather than as a single number.

Why now

Media buyers have changed what they ask for

Roughly 70% of advertisers now name incrementality as the most important measure of a retail media investment, and over 70% say it is the hardest thing to measure. Just over half cite the lack of a common standard as the single biggest barrier to investment.1

The largest retail media networks have already responded. Several now put an outside firm's name on their lift results and promote this in their sales material. One of the largest states plainly that most networks “grade their own homework,” and uses 3rd-party validation as evidence that it does not. That sets the standard every other network gets measured against, no matter its size.

A network's own reporting is not wrong. It is just harder to trust. A number produced in-house settles the question for the ad network but leaves the advertiser uncertain.

1 Association of National Advertisers. Incrementality figures from the ANA Retail Media Survey (January 2024); the standardization barrier from ANA Retail Media Measurement Standardization (August 2026).

What it changes

What a validated number does for the business

In an RFP

Against a larger network, an outside validated number closes the one gap you cannot close with reach.

At renewal

It answers the question a retail partner actually asks, which is whether the screens grew the category or moved share along the shelf.

In the rate card

Measured campaigns become a tier you can price above unmeasured inventory.

In planning

The detection floor tells you which campaigns are large enough to produce a claim worth publishing, before you commit to one.

What you get

A result fit for publication

  • A lift estimate with the store counts and the range printed beside it.Not a single number with nothing under it.
  • The comparison group named and fixed before the outcome window closes.Which stores stood in for what the exposed stores would have sold had the campaign never run, written down and dated while the result is still unknown.
  • A one-page summary written to be forwarded.To a brand, a retail partner or an investor, without needing you in the room to explain it.
  • A ‘Validated by Seurat’ mark you can put on it.Your campaign, your client, your story. Our name on the method and the findings.
  • The full method, available to anyone who asks.Including anyone who expects to disagree with the result.
The method

The comparison group that makes causal attribution possible is already there. It just wasn't analyzed.

Retail media has an advantage most channels do not: exposure separates by store, so a comparison group already exists. Nobody picked it at random, though. Stores with screens are larger, newer, better sited, or further along a rollout schedule, so comparing them raw measures the difference between the stores rather than the difference the campaign made.

So we weight the unexposed stores until they reproduce the exposed stores' pre-campaign sales path, then difference the two over the flight. Where the rollout was phased, we use an estimator built for staggered adoption. The assumption doing the work is that the two groups would have kept moving together, and we test it rather than assert it.

The network
4 ran screens5 ran no media

Screens went into the bigger, newer, busier stores. Nobody picked them at random.

Compared raw
exposed storesunexposed storesstore difference

Comparing them raw measures the difference between the stores, not the campaign.

Weighted to match
campaign startsmatched herethe lift

We weight the unexposed stores to reproduce the exposed stores' pre-campaign path, then difference the two.

A matched-store read
store salescampaign starts
sales at exposed storesmatched unexposed stores
Illustrative. The estimate is the shaded area: what happened, minus the version of events where the campaign never ran.
The specification →

The estimator is difference-in-differences on weekly store-level sales, with store and week fixed effects and standard errors clustered at the store. Where the rollout was phased, we use a staggered-adoption estimator rather than two-way fixed effects, following Callaway and Sant'Anna (Journal of Econometrics, 2021) and the decomposition in Goodman-Bacon (Journal of Econometrics, 2021). Where the exposed set is small, we weight the unexposed stores to reproduce the pre-period path, following Abadie, Diamond and Hainmueller (Journal of the American Statistical Association, 2010).

The outcome definition, the category definition, the control store list, the pre-period length and the flight window are fixed in writing and dated before the outcome window closes. The output is a point estimate with a confidence interval, not a single number.

Three things from you: the flight dates, the store list, and weekly sales.

What you provideWhat it enables in the analysis
Flight datesMust haveFixes the before and after windows, the same way for every store.
Store list, exposed and unexposedMust haveSeparates the two groups and sets the pool the comparison is built from.
Weekly sales by storeMust haveThe outcome itself. The lift is measured here.
Item-level sales alongside the categoryOptionalRoughly halves the smallest lift we can find, by measuring the brand against its own category rather than against store traffic.
Promotion calendar by store-weekOptionalRules out the item having been on price promotion in the same stores that ran the screens.
Screen hours by storeOptionalPowers the dose test: whether measured lift rises with how much exposure each store got.
Flight datesMust haveFixes the before and after windows, the same way for every store.
Store list, exposed and unexposedMust haveSeparates the two groups and sets the pool the comparison is built from.
Weekly sales by storeMust haveThe outcome itself. The lift is measured here.
Item-level sales alongside the categoryOptionalRoughly halves the smallest lift we can find, by measuring the brand against its own category rather than against store traffic.
Promotion calendar by store-weekOptionalRules out the item having been on price promotion in the same stores that ran the screens.
Screen hours by storeOptionalPowers the dose test: whether measured lift rises with how much exposure each store got.

No impression count, no click, no conversion, and no ad-server or ad-platform-reported metric enters the model at any point, including yours. There is no input for either side to adjust.

The detection floor calculator prices what each optional input is worth before you ask your retail partner for it.

What size lift could your own network detect?

Every design has a floor: the smallest lift it could have found at all, set before the read by your store counts, flight length and how volatile the category is. It is the only way to read a null correctly, because a flight that comes back with no measurable lift either produced nothing or produced something too small for this design to see. This is a quick version of the calculator that finds yours.

40
4 weeks
What kind of category is the ad for?
Smallest lift you could detect4.1%

Below this, the read returns a wide range rather than a negative. That is a different finding, and worth knowing before you commission anything.

Assumes twice as many comparison stores as exposed ones, and 26 weeks of pre-period history.

Commitments

You see the results first.

  • You see every result before anyone else.A read that does not clear is a private finding, and nothing goes out that you have not approved.
  • We test for effects that cannot exist.Every read runs against dates and categories where nothing should be found. If the method finds something there, you hear it from us.
  • We will not claim more than the design supports.That is what the confidence interval in the findings report is for.
  • The fee is fixed before the read.It does not move with the direction or size of what we find.

A number that could only ever come back positive is not one a brand's analyst will accept.

Does the method find effects in weeks when nothing ran?

It must not. A read that fails this test is not delivered.

The detail →

We run the identical estimator on quiet windows, with the same exposed and control stores and no campaign in the window. The model should find nothing. If it reports lift on quiet weeks it is picking up noise, or the two store groups are drifting apart for reasons that have nothing to do with advertising, and any number it produces on a real flight is worthless.

Does it find effects in a category that was never advertised?

Also no. This is the parallel-trends assumption, tested rather than asserted.

The detail →

We run the same estimator on a placebo category: a category sold in the same stores, over the same weeks, with no exposure in the flight. If it moves, we are catching a traffic shock, a weather effect, or a promotion running alongside, rather than anything the screens did. This is the assumption the method depends on most and the one most likely to be violated in practice, so we test it rather than assume it.

Were the two groups moving together before the campaign?

They have to be, and we publish the chart that shows whether they were.

The detail →

Before the flight, the exposed stores and the weighted control stores should track each other. We plot the pre-period side by side and publish it with the read. A visible gap opening before the campaign started is a failed match, and we say so instead of reporting a number the match can't support.

Does the effect scale with how much exposure each store got?

This is the test that matters most.

The detail →

Stores do not receive identical exposure. Screen-hours, loop frequency and impressions delivered vary across the network, sometimes by a lot. So we can ask whether measured lift rises with exposure per store. If it does, confounding becomes very hard to sustain as an explanation: a competing cause would have to be correlated with screen-hours store by store, timed to the flight, and proportional in magnitude at each one. Coincidence does not usually arrive with a gradient attached. This is the same logic epidemiology uses when randomization is not available.

Incrementality

Whether the campaign grew the category, or moved it across the aisle.

With store-level sales for the advertised item and for its whole category, we can ask where the lift came from. If the brand rose and the category rose with it, the campaign brought demand into the store. If the brand rose and the category did not, it moved share from something else on the same shelf. Those are different numbers, they are worth different amounts, and they matter to different parties: the brand cares about the first reading, the retailer about the second.

Counting the sales that followed an exposure can't tell the two apart, because some of those sales would have landed at the register had the ad never run. That is why attribution numbers get defended by the party that produced them and discounted by the party that received them, and why both sides often have trouble settling the argument.

The category grew
beforeaftertaller totalcategory salesadvertised brandrest of category

Incremental. The campaign brought demand into the store.

Share moved along the shelf
beforeaftersame totalcategory salesadvertised brandrest of category

Substitution. The campaign moved share from something else on the same shelf.

Exposure-based attribution reports the same number in either case.
Validation

What “Validated by Seurat” means

Validated by

Seurat Analytics

Three things stand behind it. You can check all three.

  • Which stores form the comparison group, the outcome definition, the category definition, the pre-period length, the estimator and the flight window are all written down and dated while the result is still unknown. The falsification tests are named before they are run, so a test cannot be dropped for producing an inconvenient answer. Choosing any of it after seeing what each version produces is the single failure that makes most observational retail media results worth nothing.

  • It is not a share of measured lift, not contingent on a positive result, and not renewable on the strength of one. This is the same rule that governs a fairness opinion in a merger, for the same reason. A fee that moves with the answer has bought the answer.

  • Anyone who wants to check a read can have the specification and the code it ran on, including anyone who expects to disagree with the result. With the retailer's permission that extends to the data itself.

The field

Seurat's approach vs. other methods.

What each one is for, and how far each one travels.

MethodWhere the number comes fromWhat you need to run itWhat it tells a brandTravels to
Network attributionYour ad server and the retailer's loyalty file.Nothing extra. It is already running.Which sales followed an exposure.Inside your own reporting.
Retailer closed-loop measurementThe retailer's POS and loyalty data.A retail partner willing to run it.Whether exposed shoppers bought more than similar unexposed shoppers.To that retailer's brand partners.
Designed in-store test programsThe retailer's POS, through a test built before the flight.The design in advance, and stores deliberately held back.What to roll out chain-wide next year.Widely, when there was time to build it.
Panel and mix modeling firmsA proprietary panel, or the advertiser's spend history.An advertiser that already retains one.How a brand should allocate across channels.Inside that brand.
Seurat matched-store readStore-level POS, on a specification fixed before the outcome window closes.Flight dates, a store list, and weekly sales.Lift, with the store count and the range beside it.To brands, retail partners and investors, because the party reporting it did not sell the media.
Network attribution
Where the number comes fromYour ad server and the retailer's loyalty file.
What you need to run itNothing extra. It is already running.
What it tells a brandWhich sales followed an exposure.
Travels toInside your own reporting.
Retailer closed-loop measurement
Where the number comes fromThe retailer's POS and loyalty data.
What you need to run itA retail partner willing to run it.
What it tells a brandWhether exposed shoppers bought more than similar unexposed shoppers.
Travels toTo that retailer's brand partners.
Designed in-store test programs
Where the number comes fromThe retailer's POS, through a test built before the flight.
What you need to run itThe design in advance, and stores deliberately held back.
What it tells a brandWhat to roll out chain-wide next year.
Travels toWidely, when there was time to build it.
Panel and mix modeling firms
Where the number comes fromA proprietary panel, or the advertiser's spend history.
What you need to run itAn advertiser that already retains one.
What it tells a brandHow a brand should allocate across channels.
Travels toInside that brand.
Seurat matched-store read
Where the number comes fromStore-level POS, on a specification fixed before the outcome window closes.
What you need to run itFlight dates, a store list, and weekly sales.
What it tells a brandLift, with the store count and the range beside it.
Travels toTo brands, retail partners and investors, because the party reporting it did not sell the media.

When Seurat Analytics is the right tool

  • Optimizing loops and dayparts?Use the ad server. That is what it is for.
  • Choosing what to roll out chain-wide next year, with time before the flight?Design the test.
  • The flight already ran?This is the read. A holdout you did not build cannot be recovered after the fact, and we can still settle the campaign from the sales data.
  • A brand, retail partner or investor wants a number that did not come from the seller?That is the other reason to reach for it.
Get started