The claim, the mechanism, the math.
Every move on the board was measured at matched stores that already ran it, transported to your store's fingerprint, and priced in dollars with a confidence grade. Pull a lever and the board calculates the read — then settles its own estimate, inside the range or not, on the record.
The read. Two shapes of evidence go into the library.
Adoption reads. A store ran the move. We build the version of that store that did not run it, assembled from comparable stores that did not move, and the difference between the two is the effect.
Exposure reads. The move happened near a store rather than in it. A competitor opened, a concept clustered on the corridor, a site closed. The effect on the nearby store is itself a measurement, and it sizes local demand without that store having done anything at all.
The match. A result measured somewhere else is only useful if the somewhere else resembles you. So matched stores are screened and adjusted to your store profile: format, traffic base, urbanicity, category mix, competitive density.
Match quality is not a footnote. It sets the width of the range and it sets the grade. The closer the match, the surer the grade.
That is the trade the library makes. Evidence from other operators is what makes an answer available at all. The matching and the grading are what make it yours.
The pool. When several stores have run the same move, we read them together rather than picking one. Each is weighted by how well it matches you and by how strong its own read is.
Agreement narrows the range and raises the grade. Disagreement widens it and lowers it, and we say so. Sometimes the surest number comes from five decent matches read together rather than one excellent match read alone.
Pooling does two things at once. It narrows the range and raises the grade, and it makes any individual store unreadable in the output. A move that only one store has ever run cannot be reported without effectively reporting that store, so it does not enter the library until enough stores have run it that none of them can be read back out.
The fingerprint. Five indexed bars, 0–100 against the twin-store benchmark, name the matching variables. The fingerprint selects the twin pool and explains the ranking.
Where each number came from.
Every number ships with a range and a grade. Outside the noise range, it is real. Inside it, "we cannot separate this from noise" is itself a finding and you will be told that in those words.
The settlement. After you act, we grade our own accuracy. What we said a move would do, against what it did. That is the discipline that keeps a library honest, and it is the reason the grades mean something.
Take EV charging, because it is the cleanest case. No mid-market operator has enough installed history to learn from their own stores, so every useful number about what a charger does to inside sales was measured somewhere else, at sites that installed earlier. Read the sites that installed against comparable sites that did not, over the same weeks. Pool the sites that agree. Match the result to the profile of the site being considered. Report the range and the grade.
That is the whole library in one category.
The method went free. Google open-sourced CausalImpact and put the core engine in everyone's hands.
The compute got cheap. Models that needed a cluster in 2015 run on ordinary hardware now.
The outside signals connect. Weather, fuel prices, local events, and category trends are commercially connectable to store-level sales.
What used to take a university research team and months of compute now runs in an afternoon. That is why an operator with thirty stores can have what only the largest chains could buy ten years ago.
The methods have names you can look up: difference-in-differences, synthetic control, Bayesian structural time series. Google open-sourced the core as CausalImpact. Published, peer-reviewed, nothing proprietary.
Agreement across reads narrows the range and raises the grade. Disagreement widens it and lowers it, and we say so.
Weather, fuel prices, and the calendar are stripped out before we credit a move.
Running designed tests today?No test design, no held-out stores →