The claim, the mechanism, the math.
Every lever on the board was measured at stores like yours that already did it, adjusted for the ways your store is different, and priced in dollars with a confidence grade. After you pull a lever, we check what it actually did against our estimate and show you whether it landed inside the range.
The industry's reference dataset sorts operators into deciles on store operating profit and reports what the top decile looks like. It is the closest thing the channel has to an answer. It has 2 problems that no amount of additional data fixes.
It ranks on the outcome. Sort firms by profit, then report what correlates with sitting at the top, and you recover the things that travel with scale, like bigger stores, more foodservice and more staff, whether or not any of them caused the profit. Read the same tables on a per-transaction basis and the ordering can invert, with the bottom of the ranking landing above the middle. That is what a descriptive ranking does when the variable it sorted on is not the variable you asked about.
It reports averages, not spreads. Company means, with no store-level distribution published anywhere in it. So it can describe a good operator and cannot locate a good store.
A ranking tells you what winners look like. It cannot tell you what would happen if you changed something. Everything below is the difference between those 2 sentences.
The read. Two shapes of evidence go into the library.
When your stores make the change. A store ran the lever. We build that store's twin from stores like it that did not run the lever. Seurat compares each store with its twin: what that store would have sold without the change. The difference between the 2 is the effect.
When something changes near your store. The change happened near a store rather than in it. A competitor opened, a concept clustered on the corridor, a site closed. The effect on the nearby store is itself a measurement, and it sizes local demand without that store having done anything at all.
The match. A result measured somewhere else is only useful if the somewhere else resembles you. So matched stores are screened and adjusted to your store profile: format, region, traffic base, urbanicity, category mix, competitive density.
Match quality is not a footnote. It sets the width of the range and it sets the grade. The closer the match, the surer the grade.
That is the trade the library makes. Evidence from other operators is what makes an answer available at all. The matching and the grading are what make it yours.
Reading several stores together. When several stores have run the same lever, we read them together rather than picking one. Each is weighted by how well it matches you and by how strong its own read is.
Agreement narrows the range and raises the grade. Disagreement widens it and lowers it, and we say so. Sometimes the surest number comes from 5 decent matches read together rather than 1 excellent match read alone.
Pooling does 2 things at once. It narrows the range and raises the grade, and it makes any individual store unreadable in the output. A lever that only 1 store has ever run cannot be reported without effectively reporting that store, so it does not enter the library until enough stores have run it that none of them can be read back out.
The fingerprint. Five indexed bars, 0–100 against the stores where the reads were initially measured, name the matching variables. The fingerprint picks the stores like this one that the reads draw on, and explains the ranking.
Where each number came from.
Every number ships with a range and a grade. Outside the noise range, it is real. Inside it, "we cannot separate this from noise" is itself a finding and you will be told that in those words.
We specify what a believable answer looks like before we run our analysis. On a single category, a promotion or merchandising change that lands between 3 and 8% is a good result. Most shoppers buy what they came for, and the share of any category's purchases that a promotional message can actually move is small. A read that comes back far outside that range will get a specification review before we share it with you.
A second marker on each row says how you would know it happened at all. Public means it can be confirmed from outside your business, in listings, posted hours, store photos or review text. POS means only your own sales data will show it. A few rows have a persistence flag: the lever is real, but nothing in any feed reports whether it is still being done.
Checking our estimate after you act. After you act, we grade our own accuracy. What we said a lever would do, against what it did. That is the discipline that keeps a library honest, and it is the reason the grades mean something.
Take EV charging, because it is the cleanest case. No mid-market operator has enough installed history to learn from their own stores, so every useful number about what a charger does to inside sales was measured somewhere else, at sites that installed earlier. Read the sites that installed against comparable sites that did not, over the same weeks. Pool the sites that agree. Match the result to the profile of the site being considered. Report the range and the grade.
That is the whole library in one category.
The method went free. Google open-sourced CausalImpact and put the core engine in everyone's hands.
The compute got cheap. Models that needed a cluster in 2015 run on ordinary hardware now.
The outside signals connect. Weather, fuel prices, local events, and category trends are commercially connectable to store-level sales.
What used to take a university research team and months of compute now runs in an afternoon. That is why an operator with 30 stores can have what only the largest chains could buy 10 years ago.
The methods have names you can look up: difference-in-differences, synthetic control, Bayesian structural time series. Google open-sourced the core as CausalImpact. Published, peer-reviewed, nothing proprietary.
Agreement across reads narrows the range and raises the grade. Disagreement widens it and lowers it, and we say so.
Weather, fuel prices, category price drift and the calendar are stripped out before we credit a change. Price drift matters more than it sounds: a category can take more dollars while selling fewer units, so a read on dollars alone can credit a change for inflation it did not cause.
Running designed tests today?No test design, no held-out stores →
We can get further without your data than you would expect. The dollar column is where that stops.
| You give us | You get | |
|---|---|---|
| Store-to-Store Read | Nothing | Which of your stores are underperforming, and what your customers say is wrong with them |
| Calibration Read | One data export | What the changes you've already made actually contributed to in-store sales |
| Lever Board | 2 years of sales data | Every proven lever ranked and suggested for each of your stores, with its lift contribution evaluated |
| Remodel Sequence Read | 2 years of sales data and your capital plan | Which sites on your capital plan should get the money first, which shouldn't get it at all, and what would have to change for the rest |
| Monitoring | Nothing new | Verdicts as the reads settle, plus where a lever is not showing up in the data the way it should, which usually means it is not being executed the way it was specified |