What is the difference between time saved and value created?
TallyEye is Brackwater Vision's camera and cart system that recounts a warehouse bin without a person walking to it. Standish Fulfillment runs it across sixty thousand bin locations, including a dense corner holding premium eyewear and small accessories. Xanthe Norrington owns what Brackwater tells a customer's leadership every quarter. Grazyna Bexworth runs inventory control at Standish. The week before one quarterly review, a comment three lines long sat right next to the number everyone was about to celebrate.
- Report value created, not time saved, the moment the two disagree.Why: hours only measure speed. Value measures whether the counting still finds what matters, and leadership needs the second one first.
- Before any hours-saved number reaches a deck, check the discrepancy catch rate by SKU value band, not one blended average.Why: a single floor-wide catch rate can hide a bad number inside the loudest, cheapest bins.
- Weight the review threshold by dollar value at risk per bin, not one flat confidence cut-off for the whole floor.Why: the camera's confidence and its actual accuracy split apart on dense, mixed bins, and a flat line can't see that split.
- Keep funding the manual check on high-value bins, even though it adds hours instead of cutting them.Why: it is currently the only thing catching what the flat threshold misses.
- Leave the flat threshold alone on the low-value bulk bins.Why: a missed miscount there costs a few dollars and gets caught at the next count anyway. Slowing it down wastes the hours the tool is supposed to give back.
- Don't cut a program just because it never shows up as hours saved.Why: that optimizes for the mistake that's cheap to spot and ignores the one that's expensive to miss.
How to answer this, stage by stage
Nobody is grading whether you can define two terms correctly. They're grading whether you know which one to trust when they disagree, and whether you'd actually say so before the deck goes out.
Let's learn
Every quarter, five counters at Standish Fulfillment walked all sixty thousand bin locations by hand. Barcode scanner in one hand, clipboard in the other. Thirty six seconds a bin, on average: scan the label, count what's actually sitting there, key in a match or flag a mismatch. That's six hundred hours across the team, every ninety days, before a single order gets picked any faster because of it.
TallyEye is Brackwater Vision's tool for exactly this job. Ceiling cameras and a rail mounted scanning cart photograph every bin, day and night, and check what they see against Standish's warehouse system. When the camera's own confidence in a count drops below a set line, a person gets sent to go look. Above that line, the system just logs it and moves on.
With TallyEye running, the manual six hundred hours drops to about ninety five: only the bins the camera itself isn't sure about ever get a second look. That's roughly five hundred and five hours back to the team every quarter, close to three people's whole working month.
Here's the turn. Those five hundred hours were never the problem. The problem was what didn't change. About eight percent of Standish's bins hold premium eyewear and small accessories, packed close together because that section fills fast: forty two percent of the floor's entire inventory value sitting inside a sliver of the space. On the other ninety two percent, standard cartons, one SKU a bin, TallyEye catches ninety seven of every hundred real miscounts. On that eight percent, stacked bins with three or four small SKUs sharing a shelf, it catches seventy eight.
Here's why the split happens, and it's not a hardware problem. TallyEye's confidence score was tuned mostly on photos of single-SKU bins, because Brackwater's first customers were bulk goods warehouses where every bin usually holds one thing. On a bin with three overlapping SKUs, the model can still come back fairly confident, and be wrong far more often, because it never saw enough stacked, mixed bins to learn what its own confidence should actually mean there. A high number and a correct count are not the same thing, and on this eight percent of the floor, they come apart.
At its worst, that split cost Standish a phantom four hundred units on a bestselling sunglasses line, sitting in the system as in stock when the shelf held far less. TallyEye's normal threshold never flagged it. It was Grazyna's team, running its own manual check on the floor's highest value SKUs regardless of what the camera's confidence said, that caught it three weeks before a launch weekend that line was expected to carry about a hundred and eighty thousand dollars of sales.
What I would leave alone: the flat threshold, on the other ninety two percent of the floor. A missed miscount on a case of work gloves costs a few dollars and gets caught at the next count anyway. Slowing that down with extra review would waste exactly the hours TallyEye is supposed to give back.
The lesson: a time-saved number only tells you the counting got faster. It never tells you whether the counting still finds what matters. Those are two different questions, and a warehouse only learns it's been answering just one of them the week a shelf marked in stock turns out to be empty.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a slide that was completely honest about hours could still tell its reader the wrong story about safety.
Xanthe Norrington has owned Brackwater Vision's customer reporting dashboard for three years, the screen that turns a camera's raw counts into the slide a customer's leadership actually reads. She built the confidence threshold that decides which bin gets a second look, back when TallyEye had four customers, all of them bulk goods warehouses.
Standish Fulfillment signed on eighteen months ago. For the first year, TallyEye worked exactly the way it was built to. Standish's floor was mostly standard cartons then, boots and jackets, one SKU a bin. Evander Fairbrother, who runs the DC floor, watched his team's manual counting hours fall from six hundred a quarter to under a hundred, and he put that number in front of his own VP every quarter like a trophy.
It thinned in three beats, and none of them looked like a mistake. Beat one: Standish won a new line of premium eyewear and small electronics accessories, added to the floor over a few months, packed into a dense section because that's where the open racking was. Beat two: TallyEye's flag rate on that section stayed exactly where it always was, the same flat threshold as everywhere else, so nobody watching the dashboard saw anything change at all. Beat three: Grazyna's inventory control team quietly started running its own manual check on the floor's top value SKUs, a program nobody outside her team had approved or funded as its own line item, because she didn't trust a flag rate that never seemed to notice the section had gotten more valuable.
The trigger wasn't a stockout. It was a comment, three lines long, that Grazyna left on an internal ticket the week before Standish's quarterly review: a phantom four hundred units, found on a top SKU, caught by her team's own check three weeks before a launch weekend. Xanthe found it by accident, scrolling past the ticket while pulling numbers for that same quarter's slide.
On that slide, right next to Grazyna's comment, sat the number Evander had already approved sending up: five hundred and five hours saved, the best quarter yet. Nobody had connected the two. The comment thread had been read by three people. The hours-saved tile was about to be read by a room full of executives deciding whether to expand the contract.
Here's what makes this the real question, not a bigger camera or a smarter model. Standish's floor didn't have two problems. It had one number that measured a different thing than everyone assumed it measured. Hours saved counted how fast a bin got checked. It never asked whether the check was any good on the bins that actually held the floor's money.
The decision that opened the door went back to a short meeting, the month Brackwater first set TallyEye's confidence threshold, years before Standish ever signed. Someone asked whether the flag line should move depending on what a bin held. The answer was no: a flat number was simpler to explain to a customer, and back then every bin on every one of Brackwater's four customers held roughly the same kind of thing. It was the sensible call, for the floors that existed then.
Run the quarter again with one change: the flag threshold weighted by dollar value at risk, not one flat line for the whole floor. The same eight percent of bins, holding forty two percent of the floor's value, gets a lower bar to clear before a person looks. Standish's manual hours land at about a hundred and fifty five, not ninety five, sixty hours more than the flat threshold cost. But the phantom four hundred units gets caught by the system itself, in the ordinary flow of a Tuesday, not by a side program riding on Grazyna's own initiative and a comment nobody read in time.
One design let a dashboard tell a true story about speed and a false one about safety, on the same slide, in the same font. The other ties the flag line to what a bin is actually worth, so the slide can't say one thing while the floor is doing another.
What Xanthe would tell herself, back in that first meeting about the threshold: a flat number isn't neutral just because it's simple. It's a bet that every bin is worth roughly the same to get wrong, and that bet gets more expensive exactly as fast as a customer's floor gets more valuable.
PICK, or the two numbers that can't both lead the deck
Not a definitions quiz. PICK is what forces a real answer to which honest number you'd actually stake a renewal conversation on, instead of reporting whichever one is easiest to compute.
Three things worth stating directly, since this is where the real judgment sits. Brackwater considered, and rejected, a second fix: run a full manual recount on the whole eyewear section every month, and stop trusting the camera there at all. It lost, because it hands back roughly the same hundred and fifty hours a month the camera was supposed to save, for just eight percent of the floor, which defeats half of what Standish bought TallyEye to do. The AI-specific failure worth naming is confident wrongness by SKU class: the model's confidence score was calibrated mostly on single-SKU images, so on a stacked, mixed bin it can report a normal-looking confidence number while actually being wrong far more often, because it never learned what "unsure" should really look like on that kind of photo. The guardrail isn't a better camera, it's weighting the review threshold by the bin's own dollar value, so a moderately confident miscount on a four-hundred-dollar SKU gets a person's eyes even when the same confidence score on a four-dollar SKU wouldn't. And the trade-off is real and accepted on purpose: the weighted threshold costs sixty more review hours a quarter, a smaller hours-saved number, in exchange for catching the miscounts that actually threaten a stockout, rather than waiting one to two more model versions for accuracy alone to close the gap.
And if you want to be sure it really works, try it somewhere else
Same four letters, a garment cutting floor instead of a warehouse, and this time the hidden mistake isn't a phantom unit count. It's a shade of dye.
WeaveGuard is Kolvenbach Machine Vision's tool: cameras mounted over the cutting line photograph every fabric panel after it's cut, checking for tears, weave flaws, and dye-lot mismatches before a panel gets sewn into a garment. Fyodor Deveraux runs quality at Ingledew Textiles, which cuts fabric for premium outerwear across twelve lines.
WeaveGuard catches ninety six of every hundred structural defects, tears and weave flaws, the category it was mostly trained on. On dye-lot mismatches, a purely color-based defect, it catches seventy one. A run of twelve hundred premium jackets nearly shipped with a mismatched dye lot on the lining, caught only because Fyodor's team runs its own manual color check on every high-margin line, a program that adds inspection time rather than cutting it. Left unshipped, that run would likely have triggered a retailer chargeback and a delisting review worth about ninety five thousand dollars.
Same rank, different lever: at Standish the lever was a section that got more valuable. At Ingledew it's a defect type the model was simply never shown much of. Different cause, same asymmetry: whatever the model's own training never weighted heavily, its confidence score can't be trusted to weight heavily either.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the position: value created leads whenever it disagrees with time saved, because time saved can't see a blind spot the model itself doesn't know it has.
Cost: no budget this quarter to retrain the model on more mixed-bin or dye-lot photos. The cheap fix is a threshold weighted by what's at risk, not a new model.
The model got better, for real: say the high-value bin miss rate falls to five percent overnight, a genuinely better model. The kill criteria is now met. Hours saved can lead the deck again, but only because the evidence, not the hope, says so.
Where people run it wrong.
They report the number that's easiest to compute, hours times a wage, and never check it against a stratified baseline.
They treat a blended, floor-wide accuracy number as proof nothing's hiding underneath it.
They cut a program that "doesn't save time," without asking what it's quietly catching instead.
How to use it live. Ask the split question before naming either number: "is this a speed metric or a safety metric, because a flat threshold and a fast count protect two completely different things." That buys the room to actually answer, instead of picking whichever number sounds better in the room.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"If the model's the real problem, why not just fix the model instead of reporting around it?" Response: the model fix takes version cycles, six of them here and still above the kill line. The weighted threshold is the answer for tomorrow, not next year, and the report has to be honest in the meantime.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.
- #7 What ROI argument works for an internal AI tool with no revenue line?