The public marketing data read.
Not geo-targetable: creator content, and any campaign that reaches an audience by interest rather than by market. No campaign viewers missed it based on location, so there is no geo-based comparison group waiting to be used. We build the comparison from historical data instead.
Take a public outcome
Branded search demand for the advertiser. When someone watches a creator segment and later goes looking for the brand by name, that search is a real behavioral signal, and it is publicly measurable.
Build a control set
Trends that move the way the brand's demand moves for reasons that have nothing to do with the campaign. Category demand, seasonal patterns, adjacent search behavior.
Forecast the world without the campaign
We fit on the weeks before the campaign, forecast forward, and compare the forecast to what actually happened. The gap is the lift.
The Counterfactual Edge: We build the version of events where the campaign never ran, then measure the gap between prediction and reality.
Every lift number is a comparison against something that did not happen. Most measurement gets that comparison by holding out an audience or a set of markets, then watching what the untreated group does. That requires designing the test before the campaign runs, and it requires audiences you can cleanly separate.
Creator advertising gives you neither. The campaign has usually already run by the time anyone asks what it did, and creator audiences do not split along geographic lines, because a creator's viewers are distributed by interest rather than by market. Geo-based holdouts, the standard tool, do not cleanly apply.
So we build the comparison from historical data.
We need one thing from you: the upload date.
That's it! The outcome data is public, the comparison trends are public, and no impression count, view count, click, conversion, or ad-platform-reported metric enters the model at any point. That is what lets us stay an independent observer. There is no input for either side to adjust.
The specification →
The estimator is a Bayesian structural time-series model, following Brodersen et al. (2015), implemented in Google's open-source CausalImpact package. The state specification includes a local level component, weekly seasonality, and regression on the donor pool. We fit on a defined pre-period, forecast into the post-period, and report the posterior distribution of the difference. The output is a point estimate with a credible interval, not a single number.
The full specification for any read we produce, including the donor pool, the pre-period length, and the state components, is published before the outcome window closes.
We run the tests built to break our own result.
An honest account of this method has to start with its weakness. This is an observational design, not a randomized experiment. Nothing was assigned at random. That is a real limitation, and the marketing science literature has documented cases where observational methods failed to recover the answer a randomized test later produced.
The response to that is not to claim otherwise. It is to show the work that rules out the alternatives.
Every read we produce includes three falsification tests. They are run before we report anything. On the public research we publish ourselves, the results go out either way. On your engagement, you see them first.
Does the method find effects when nothing happened?
It must not. A read that fails this test is not delivered.
The detail →
We run the identical estimator on dates when no campaign ran. Same brand, same controls, same specification, arbitrary intervention dates. The model should find nothing. If it reports lift on quiet weeks, the estimator is picking up noise or misfitting the seasonal structure, and any result it produces on a real campaign date is worthless.
Does it find effects in trends that could not have been affected?
Also no. This is the assumption the method depends on most.
The detail →
We run the same estimator on comparison trends that have no plausible exposure to the campaign. If they show lift, the model is capturing something moving across the whole category rather than something the campaign did. This also tests the assumption the method depends on most: the controls must be unaffected by the intervention. The CausalImpact documentation flags this explicitly, and it is the assumption most likely to be violated in practice. We test it rather than assert it.
Does the effect scale with the size of the campaign?
This is the test that matters most.
A single before-and-after result is weak evidence. Plenty of things could have moved demand that week. But run the estimator across many campaigns of different sizes and you can ask whether the effect scales with reach. If it does, a competing explanation has to be correlated with reach, timed to each individual upload, and proportional in magnitude to each one. Coincidence does not usually arrive with a gradient attached.
Stronger still is our ability to derive a marketing effectiveness threshold. Where lift appears only above a certain level of reach and is not statistically significant below it, we have a sharp break rather than a smooth trend. A confounder that drifts gradually cannot produce a break like that, so competing explanations get much harder to sustain.
This is the logic epidemiology relies on when randomization is not available: a dose-response relationship is evidence that's needed when you can't run an experiment.
What the design could and could not have detected
Separate from whether the effect is real is the question of how small an effect this design could have found at all.
Before the read, we compute the detection floor: given the campaign's reach, the variance in the outcome data, and the length of the pre-period, the smallest effect the model could have distinguished from noise. We publish that number alongside the estimate.
This matters for reading a null result correctly. If a campaign comes back with no measurable lift, there are two possible reasons, and they are not the same. The campaign may have produced nothing. Or it may have produced something smaller than this design could see. We tell you which one you are looking at.
A read that does not clear is a private finding.
You see every result before anyone else does, and nothing goes out that you have not approved. The fee is fixed and agreed before the read, so it does not move with the direction or size of what we find. A number that could only ever come back positive is not one a brand's analyst will accept.
Our Commitment to Transparent Analytical Methods
Four things, and you can check all four.
Upload dates and view counts are public. Branded search demand is public. No ad platform metric is ingested and no advertiser first-party data is requested. There is no access anyone can withdraw and no cooperation anyone can withhold, which is why a read can be produced whether or not the parties like the direction it is heading.
It is not a share of measured lift, not contingent on a positive result, and not renewable on the strength of one. This is the same rule that governs a fairness opinion in a merger, for the same reason. A fee that moves with the answer has bought the answer.
The control set, the pre-period length, the state components and the intervention date are written down and dated while the result is still unknown. Choosing the specification after seeing what each one produces is the failure this rules out, and it is the failure that makes most observational marketing results worth nothing.
Public inputs and a published specification are only worth something if somebody can actually run them. Anyone who wants to check a read can have the code and the data it ran on, including anyone who expects to disagree with the result.
Seurat's approach vs. other methods.
The comparison set, including where we come up short.
| Method | Data comes from | Paid by | Identification | Built to answer |
|---|---|---|---|---|
| Ad platform lift studies | The ad platform | The advertiser, run by the ad platform | Randomized holdout. Strong. | Did exposure move survey metrics inside that ad platform |
| Ad-server 3rd parties | Ad-server pixels | The advertiser or publisher | Randomized holdout and synthetic control. Strong. | Did the ad drive tracked conversions |
| Geo experiment platforms | Advertiser's own stack, plus ad platform APIs | The advertiser | Geo randomization. Strong, where it applies. | How should the advertiser allocate across channels |
| Retail panel firms | Proprietary panel | The advertiser | Panel and mix modeling. Varies. | Did creator activity move retail sales |
| Influencer marketing platforms | Social and ad platform APIs, plus their own tooling | The advertiser | Engagement reporting and attribution. Not causal. | Which creators performed on reported metrics |
| Seurat Channel Read | Public data, or retailer sales under a published specification. One date from you. | The media seller. Fixed fee, not outcome-contingent. | Observational counterfactual with published falsification tests. Admittedly weaker than randomization. | Is there an independent number a planner can defend |
But when you can't run an experiment, or don't have access to all the data, we're the answer
Every other method in the chart above needs something from a party with a stake in the answer. Ad platform data, advertiser pixel access, an advertiser's media stack, or a proprietary panel that cannot be inspected from outside. Each of those is fine on its own terms. Each also means the number cannot be produced, or checked, without the cooperation of someone who wants it to come out a particular way.
Ours only needs a starting date.
It runs backward
We can settle campaigns that already ran, with no test designed in advance. Nothing else on this list can do that, because a holdout you did not build cannot be recovered after the fact.
It cannot be tilted
On a public marketing data read, even if we wanted to produce a flattering number there is no input to adjust: the outcome data is public, the controls are public, and the specification is published before the window closes.
It can be checked
Channel Reads run entirely on public data, so anyone holding the published specification can reproduce the read and get our number, or fail to. That is the point of the signature. Engagements built on a client's own data are confidential and are not published; there the equivalent is that you see every specification we ran, not only the one we led with.
When to reach for Seurat Analytics
A randomized holdout gives stronger identification than we do, because random assignment removes confounding by construction while ours removes it by argument and by testing. If you can design the test in advance, design it.
This is not a replacement for the tools above. If you are choosing which creators to renew, use the pixel and the survey. If you are allocating budget across an advertiser's whole mix, that is a mix model, and it should be. This is for the moment when someone outside the transaction asks whether the channel works, and the honest answer requires that the person answering has nothing to gain from the result.
One channel, read from public data, with the brand withheld.
Everything below was measured independently by Seurat Analytics. None of it came from the advertiser or the ad platform.
A direct-to-consumer apparel brand ran paid integrations inside creator content between June 2024 and August 2026. We found 46 of them across 5 channels from public sources alone. 19 had 21 or more days clear of any other paid placement for the same brand, so their pre-campaign baseline is not contaminated by an overlapping campaign. Those 19 are the read.
What the average hides
Pooled across all 19 integrations the effect is 3.4%, with a margin of error wider than the effect itself. That is not a finding. It is the absence of one.
Averaging campaigns large enough to work together with campaigns too small to work produces a number that describes neither, and it is the most common way a real creator effect gets reported as no effect at all. 6 of the 19 integrations here were too small to move consumer demand. Reported as one average, they take the other 13 down with them.
How long it keeps working
Half the measured effect is still present a week after the upload. About a quarter is still present after two weeks. The weekly persistence is 0.52, a persistence normally reserved for television, and almost nothing in this market models creator content that way.
This matters more than the headline lift number. A model that retires the effect after a couple of days finds roughly a sixth of what the channel actually contributed. The under-counting is not a measurement error anyone made. It is an assumption nobody revisited.
Independent by construction
- Upload dates and view countsPublic
- Branded search demandPublic
- Ad platform metricsNone ingested
- Advertiser first-party dataNone requested
What it implies for buying
Creator advertising works, but only at weight. The threshold is not a rounding artifact. Below it the estimator finds nothing at all, and above it the effect is large and statistically clean. That argues for concentration over scatter: fewer, larger placements clear a bar that one-off executions cannot reach, and a budget spread evenly across small placements can spend a whole year below the line where anything happens.
Why this one is anonymous
We withheld the brand because this read was produced unsolicited, and naming a company that never asked to be measured is a different decision from measuring it. The trade-off is real: without the brand, you cannot reproduce this particular number yourself. The method still holds on any brand you care to name, and a read on one of yours is checkable end to end.
Selling in-store, forecourt, or onsite retail media instead? That exposure separates by store, so it gets a different read: the matched-store read.