NEW RESEARCH79% of shoppers say the same chain is noticeably better or worse store to store.Get the full report, free →

New research: why the same chain wins in one store and leaks in another.Get the report →

How it works

The claim, the mechanism, the math.

The claim

Every lever on the board was measured at stores like yours that already did it, adjusted for the ways your store is different, and priced in dollars with a confidence grade. After you pull a lever, we check what it actually did against our estimate and show you whether it landed inside the range.

Why a ranking can't answer this

The industry's reference dataset sorts operators into deciles on store operating profit and reports what the top decile looks like. It is the closest thing the channel has to an answer. It has 2 problems that no amount of additional data fixes.

It ranks on the outcome. Sort firms by profit, then report what correlates with sitting at the top, and you recover the things that travel with scale, like bigger stores, more foodservice and more staff, whether or not any of them caused the profit. Read the same tables on a per-transaction basis and the ordering can invert, with the bottom of the ranking landing above the middle. That is what a descriptive ranking does when the variable it sorted on is not the variable you asked about.

It reports averages, not spreads. Company means, with no store-level distribution published anywhere in it. So it can describe a good operator and cannot locate a good store.

A ranking tells you what winners look like. It cannot tell you what would happen if you changed something. Everything below is the difference between those 2 sentences.

The mechanism

The read. Two shapes of evidence go into the library.

When your stores make the change. A store ran the lever. We build that store's twin from stores like it that did not run the lever. Seurat compares each store with its twin: what that store would have sold without the change. The difference between the 2 is the effect.

When something changes near your store. The change happened near a store rather than in it. A competitor opened, a concept clustered on the corridor, a site closed. The effect on the nearby store is itself a measurement, and it sizes local demand without that store having done anything at all.

What we look for
Weekly sales at one store: actual vs. its twin
$42K$46K$50K$54K$58KWk 1Wk 4Wk 7Wk 10Wk 13Wk 16something changedTHE GAPgrowing to ~$7K/week
Actual salesThe store's twinThe gap

The match. A result measured somewhere else is only useful if the somewhere else resembles you. So matched stores are screened and adjusted to your store profile: format, region, traffic base, urbanicity, category mix, competitive density.

Match quality is not a footnote. It sets the width of the range and it sets the grade. The closer the match, the surer the grade.

That is the trade the library makes. Evidence from other operators is what makes an answer available at all. The matching and the grading are what make it yours.

Reading several stores together. When several stores have run the same lever, we read them together rather than picking one. Each is weighted by how well it matches you and by how strong its own read is.

Agreement narrows the range and raises the grade. Disagreement widens it and lowers it, and we say so. Sometimes the surest number comes from 5 decent matches read together rather than 1 excellent match read alone.

Pooling does 2 things at once. It narrows the range and raises the grade, and it makes any individual store unreadable in the output. A lever that only 1 store has ever run cannot be reported without effectively reporting that store, so it does not enter the library until enough stores have run it that none of them can be read back out.

5 MATCHED READS1 read alone5 reads together

The fingerprint. Five indexed bars, 0–100 against the stores where the reads were initially measured, name the matching variables. The fingerprint picks the stores like this one that the reads draw on, and explains the ranking.

Where each number came from.

Measuredobserved at matched stores that ran the lever.
Transportedmeasured there, adjusted to this fingerprint; always a range with a midpoint.
Simulatedmodeled where no matched-store read exists yet; field-sourced levers default here.
Validatedsettled at the client's own stores, rendered against the original estimate, including inside-the-range landings.
$8K$16Kmidpoint $12K
Transported: always a range with a midpoint

Every number ships with a range and a grade. Outside the noise range, it is real. Inside it, "we cannot separate this from noise" is itself a finding and you will be told that in those words.

We specify what a believable answer looks like before we run our analysis. On a single category, a promotion or merchandising change that lands between 3 and 8% is a good result. Most shoppers buy what they came for, and the share of any category's purchases that a promotional message can actually move is small. A read that comes back far outside that range will get a specification review before we share it with you.

A second marker on each row says how you would know it happened at all. Public means it can be confirmed from outside your business, in listings, posted hours, store photos or review text. POS means only your own sales data will show it. A few rows have a persistence flag: the lever is real, but nothing in any feed reports whether it is still being done.

Checking our estimate after you act. After you act, we grade our own accuracy. What we said a lever would do, against what it did. That is the discipline that keeps a library honest, and it is the reason the grades mean something.

$3.4K$5.1K
measured $4.1K vs. the $3.4–5.1K estimate. Inside the range ✓
A worked example

Take EV charging, because it is the cleanest case. No mid-market operator has enough installed history to learn from their own stores, so every useful number about what a charger does to inside sales was measured somewhere else, at sites that installed earlier. Read the sites that installed against comparable sites that did not, over the same weeks. Pool the sites that agree. Match the result to the profile of the site being considered. Report the range and the grade.

That is the whole library in one category.

Why now

The method went free. Google open-sourced CausalImpact and put the core engine in everyone's hands.

The compute got cheap. Models that needed a cluster in 2015 run on ordinary hardware now.

The outside signals connect. Weather, fuel prices, local events, and category trends are commercially connectable to store-level sales.

What used to take a university research team and months of compute now runs in an afternoon. That is why an operator with 30 stores can have what only the largest chains could buy 10 years ago.

The math

The methods have names you can look up: difference-in-differences, synthetic control, Bayesian structural time series. Google open-sourced the core as CausalImpact. Published, peer-reviewed, nothing proprietary.

Agreement across reads narrows the range and raises the grade. Disagreement widens it and lowers it, and we say so.

Weather, fuel prices, category price drift and the calendar are stripped out before we credit a change. Price drift matters more than it sounds: a category can take more dollars while selling fewer units, so a read on dollars alone can credit a change for inflation it did not cause.

Running designed tests today?No test design, no held-out stores →

We can get further without your data than you would expect. The dollar column is where that stops.

You give usYou get
Store-to-Store ReadNothingWhich of your stores are underperforming, and what your customers say is wrong with them
Calibration ReadOne data exportWhat the changes you've already made actually contributed to in-store sales
Lever Board2 years of sales dataEvery proven lever ranked and suggested for each of your stores, with its lift contribution evaluated
Remodel Sequence Read2 years of sales data and your capital planWhich sites on your capital plan should get the money first, which shouldn't get it at all, and what would have to change for the rest
MonitoringNothing newVerdicts as the reads settle, plus where a lever is not showing up in the data the way it should, which usually means it is not being executed the way it was specified