NEW RESEARCH79% of shoppers say the same chain is noticeably better or worse store to store.Get the full report, free →

New research: why the same chain wins in one store and leaks in another.Get the report →

Calibration Read

You've been running experiments for years. Nobody wrote them down.

Nothing in a chain ever rolls out cleanly. The coffee club reached 22 stores between January and May. The fresh section went into 11 of 42. The remodels went in a sequence, set by supplier capacity and by which stores had the room.

That staggering is the thing that makes this possible. Every time part of your chain got something and part of it didn't, you ran a test without meaning to. The stores that hadn't had it yet are the comparison group. Seurat compares each store with its twin: what that store would have sold without the change.

So we're not asking you to start testing. We're reading the tests that already happened. We take 3 to 5 of them and settle the argument. Your changes, your stores, your sales data. Each read comes back with a lift estimate, a range around it, and a grade for how much weight the read can hold.

The tests you already ran

Illustrative, using the same 42-store proxy chain as the worked set below.

WhenWhereWhat ranBig enough to readWhat it was worth
Jan to May 202522 of 42Coffee clubYesUnits +9.4%, dollars +2.1%
Apr 202511 of 42Fresh sectionYesNet foodservice +$25K per store
Mar to Jun 202514 of 42Trial-size cart railYesNo lift. Came off the board
Aug 2024 to Feb 20259 of 42Light remodelThe counter, yes. The seating, no+$28K per store, all of it the counter
From Sep 202416 storesSecond morning shiftYes+$24K in sales, and $8K short once the shift was paid for
Jan 20251 storeNew microwaveNoToo small to read at 1 store, and it always will be
Aug 2024Store 112Nobody knowsYesSales fell 6% and stayed there
The rows nobody put there on purpose

The last row is the one to look at. That's a change we found in the sales history that nobody planned and nobody has explained.

A store that drops 6% and holds there reads as a soft comp for 12 months. It gets written off as a tough market or a competitor opening. Then it disappears, because after a year the comparison is post-drop against post-drop, so every report you have says the store is flat and fine. At a store doing $6,000 a day inside, that's about $131,000 a year, permanently, and no report will ever raise it again.

Those rows are usually the most valuable thing on the page.

Why this comes before the Board

The Lever Board ranks and suggests the levers you haven't pulled yet. It cannot do that accurately until it knows which ones you already made and how those landed.

A Calibration Read is that first pass. It also settles two things needed to make the Board accurate that can only come from your own data: what a point of lift is actually worth at your sales volumes, and how small a lift is detectable within your sales data at all.

The fee credits toward any full Lever Board engagement started within 6 months.

Five reads at a 42-store chain

Heartland Markets is our proxy chain: 42 stores in eastern Iowa, fuel plus kitchen, the same chain behind the Lever Board demo. Every figure below is illustrative. The format is exactly what you would receive.

The cart rail deployment, in full

What ran. A sub-dollar rack at the checkout cart rail, holding trial sizes of items sold in full size elsewhere in the store. 14 stores took it between 3 March and 19 June 2025. 28 stores never got one.

What it was compared against. The 28 stores that ran nothing, weighted so that together they reproduce the 14 stores' sales path across the 26 weeks before each install. Installs landed on different dates, so each store is read against its own before and after.

What else was going on. Several of these stores were also remodeled, or got the fresh section, inside the same window. Every other dated change in the chain goes into the read as a control, and store-weeks where 2 changes overlap are held out rather than credited to either one. A chain that only ever did one thing at a time would be easier to read and would not exist.

Then the part that decides everything.

The cart rail can be read three ways. All three are defensible, and they nest inside each other.

Outcome one: sales of the trial items themselves.+$1,840 per store per year. Grade B.Real, but its contribution is negligible. Nobody installs a fixture to sell $1,840 of candy.
Outcome two: sales of those same items in full size, elsewhere in the store.-$2,100 per store per year. Grade B.The trial sizes cannibalized sales from the full-size facing three feet away.
Outcome three: both together, which is the net effect on those categories.-$260 per store per year, on a confidence interval running from -$7,500 to +$7,000. Grade B.At 14 stores with the categories narrowed, the smallest lift this read can detect is about $7,300 per store per year. The estimate sits well inside that, so the result is not statistically significant.

What that means. Same fixture, same 14 stores, same weeks. One outcome says the rack sells something. One says it took more than it added. The third is the one the fixture was installed to affect, and it shows no lift: at this size, the cart rail did not do measurably better than breaking even.

This is why the outcome gets written down and dated before the read runs, not chosen afterward from whichever version looks best. Picking the outcome after seeing what each one produces discredits the analysis.

What it settled. We had this lever on Heartland's board at about $12K per store per year, an estimate transported from stores elsewhere that ran it. The measured interval at Heartland's own stores tops out at +$7,000. $12K sits outside that, so the estimate does not hold here and the row comes off the board.

That is the loop working as designed. The read cannot prove the cart rail is worth nothing. It can prove it is not worth $12K at these stores, which is enough to stop spending on it.

But there's still hope for that cart rail:

Everything above reads the cart rail as a fixture that has to pay for its own space. Change that assumption and the read changes with it.

  • Suppliers fund the samples. If the trial sizes arrive at no cost, the rack no longer has to cover the space it sits in.
  • The question moves. It stops being whether the category grew and becomes whether discovery brings people back.
  • So does the read. Repeat visit frequency over a longer window, not category sales over 26 weeks.
  • Worth running. It is a different lever from the one above, and we would scope it as a different read.

The verdict on the board is against the self-funded version. The supplier-funded version has not been tested and is not ruled out.

What ranLight remodel: online-order pickup counter and seatingWhen it ranAug 2024 to Feb 2025, phasedStores9 of 42What we readTotal inside sales, foodservice attach rate, traffic during seating hoursWhat came back+$28K per store per year, all of it on the pickup counter. Seating: no measurable liftGradeB
What ranCoffee clubWhen it ranJan to May 2025, phasedStores22 of 42What we readDispensed beverage units and dispensed beverage dollars, read separatelyWhat came backUnits +9.4%, dollars +2.1%GradeA
What ranSecond morning employee, Mon to FriWhen it ranFrom Sep 2024Stores16, read on 12What we readMorning daypart foodservice and dispensed beverageWhat came backSales +$24K per store per year. It still lost about $8K a store once the shift was paid forGradeB

Every row here is Measured: read at these stores, on the chain's own sales data. Nothing on this page is transported from anywhere else. That is the difference between a Calibration Read and a Lever Board.

What the grades mean

Every read on this page was measured. The grade says how much weight the result can hold.

AEnough stores that executed the change, a clean staggered rollout, and enough comparison stores that tracked closely before the change. The estimated lift effect is clearly bigger than a normal week's ups and downs (statistically significant).
BA sound read with a wider confidence interval. Fewer stores, or a noisier comparison, or an effect closer to the limit of what the data can resolve with statistical significance.
CReported but not relied on. The effect sits at or below what this data can detect at a statistically significant level, or the comparison group is compromised.

A grade is about the quality of the read, not the change. A grade A result can still say a change did nothing, but it would be a confident assertion of that fact.

Why we read changes that ran at several stores

One store cannot answer this

This is the arithmetic that shapes everything above, and it is worth 60 seconds.

A store doing $6,000 a day inside, with typical week-to-week variation, has a read threshold of about $55,000 a year. A change smaller than that disappears into the swing between a good week and a bad one, so no method can pull it out.

Almost nothing you do at a single store is worth $55,000. Most are worth $6,000 to $30,000. So at one store, on its own, almost nothing you do is measurable. That is not a limit of our method. It is a limit of the data, and it applies to anyone.

Two things we do to fix it:

Pool the stores that ran it. Reading 14 stores together instead of one lowers the threshold by roughly a factor of 4. More stores, lower threshold.

Narrow the outcome. Read the categories the change could plausibly touch, in the hours it could touch them, rather than the whole store all day. Roughly another halving.

Together they take a $55,000 read threshold down to about $7,000, and that is where the reads on this page live:

ReadStores pooledOutcome narrowedSmallest lift detectable
Coffee club22Yes$5.8K per store per year
Sub-$1 trial-size cart rail14Yes$7.3K
Second morning employee12Yes$7.9K
Fresh section11Yes$8.3K
Light remodel9No, it moves the whole store$18.3K

This is why the menu below asks whether a change ran at several stores rather than one, and why a change rolled out everywhere on a single day is the hardest thing in retail to read. It left nothing to compare against.

What lands on your desk

One page per read. What ran, when, where, and what it was compared against, the lift estimate with its confidence interval and grade, and which outside factors were removed before we credited the change with anything.

A front page. Which reads cleared, which came back below the read threshold, and what each one is worth at your sales volumes.

A back page. What your data can and cannot support going forward, including the smallest lift detectable within your own sales data. This is what prices a Board engagement accurately.

At least one null, if there is one. A change that showed no lift, or that grew sales without covering its own cost, is reported in exactly those words. It is a finding, not a failure to find one.

Nothing goes anywhere. These are reads of your own business on your own data. They are yours.

What we remove before crediting a change

Weather, fuel price swings, the promotion calendar, seasonality, and any other dated change running at the same store in the same weeks. Whatever is left after all of that is what the change did.

A note on the currency

Every dollar figure we report is incremental sales, not gross profit, because margins are yours and they vary by store. Where a change has a cost attached, as the morning shift does, we show the gross profit conversion next to the sales figure so both sit in the same currency. The second morning employee is the example: it grew sales by $24K a store and still lost about $8K once the shift was paid for.

What we need, and what we never ask for
What we need
  • Category sales by store, by week, going 2 years back. Units are necessary. Dollars are preferred alongside them, because some questions are only answerable with dollars.
  • The dates each change went live, store by store.
  • Store attributes: format, location type, remodel dates.
  • Optional and worth having: promotion calendar by store-week, and item-level detail on the affected categories. Each one roughly halves the smallest lift we can detect.
What we never ask for
  • No individual receipts or transaction-level data
  • No loyalty, card, or payment data
  • No personal data of any kind
  • No cameras, no app data, no tracking

One export, in whatever shape your back office produces it. If it needs three files and a call with your POS vendor, tell us and we will narrow the ask instead.

Here is what we would read at a chain like yours

These are the changes that come back inconclusive most often when you just look at raw sales over time. Accordingly, they're the ones that usually surprise teams when our methods unearth real insight about their performance. Pick 3 to 5 of these, or tell us what you actually ran.

0 selected

Where this sits

We can get further without your data than you would expect. The dollar column is where that stops.

You give usYou get
Store-to-Store ReadNothingWhich of your stores are underperforming, and what your customers say is wrong with them
Calibration ReadOne data exportWhat the changes you've already made actually contributed to in-store sales
Lever Board2 years of sales dataEvery proven lever ranked and suggested for each of your stores, with its lift contribution evaluated
Remodel Sequence Read2 years of sales data and your capital planWhich sites on your capital plan should get the money first, which shouldn't get it at all, and what would have to change for the rest
MonitoringNothing newVerdicts as the reads settle, plus where a lever is not showing up in the data the way it should, which usually means it is not being executed the way it was specified

About $85 per store per month, for one quarter. We recommend running a Calibration Read across your whole chain rather than a subset, because the reads need enough stores that made each change and enough that did not to compare them against. A 30-store chain runs roughly $7,600 across those three months, billed monthly. 20 business days from the data landing, and it credits in full toward any full Lever Board engagement started within 6 months.