Describe how you would measure the ROI of an eval and monitoring investment.
Flywheel reads sensor data off gym equipment and tells a maintenance team which machine to check next, before it breaks. Coppice Analytics builds it. Vaultline Fitness runs it across 260 locations and about 14,000 machines. Yara Solano owns the call at Coppice on what Flywheel's eval and monitoring budget actually buys. Declan Marsh runs facilities for Vaultline. One version upgrade taught both of them that a model can get better and worse at the same time, on two different machines, and the bill for the difference lands on someone who never even opened the app.
- Measure ROI as harm prevented, not accuracy gained.Why: a model can get more accurate overall and still get much worse for one group, and an overall accuracy number will never show you which.
- Give the eval set a floor per group, not just a total case count.Why: a set that's 90 percent one group will always look healthy on average, even while the other group quietly fails.
- Split every drift alert the same way you split the eval set.Why: one blended monitoring number hides a regression exactly as well as one blended eval number does.
- Gate every new model version behind a shadow period before it goes live for everyone.Why: shadow scoring against the live model, on real traffic, catches a regression days before a real failure catches it for you.
- Track both costs, every time a regression actually happens.Why: the gap between what a golden-set catch cost and what a field catch cost is the real ROI number, not a guess.
- Don't put a well-covered group through the same slow gate.Why: a shadow gate on a group the eval set already tests well just adds delay with nothing left to catch.
How to answer this, stage by stage
Nobody is grading whether you can say "we'd run evals and monitor it." They're grading whether you can turn that into a real number, and whether you know exactly which group would get hurt first if you skipped it.
Let's learn
What does it mean when a maintenance tool gets better on average and much worse for one kind of machine, in the same release?
Flywheel reads vibration, motor current, and cable tension off gym equipment and tells a maintenance team which machine to check next, before it fails on its own.
Before Flywheel, Vaultline serviced every machine on a fixed 90-day calendar, needed or not. Strength equipment repairs were logged on paper, never digitized. With Flywheel, every machine gets a daily risk score from 0 to 100, built from live sensor data. Coppice tests every new model version against a golden set first: 600 real failure cases pulled from Vaultline's own repair history.
Coppice shipped version 4 to fix a real complaint: treadmills were throwing too many false alarms, and technicians were tired of driving out to a machine that turned out to be fine. Version 4 fixed that. On the golden set as a whole, its miss rate barely changed, 6.1 percent to 6.8 percent. Nobody flagged it.
Here's the turn. That "barely changed" number was hiding something. The golden set was 540 treadmill and elliptical cases and only 60 cable and pulley machine cases, because cardio was the only equipment type with digitized failure history when Coppice built it. Split the same number by equipment class, and version 4's miss rate on cardio actually improved, from 6 percent down to 3. Its miss rate on strength equipment went from 7 percent up to 41.
Nobody saw that split, because nobody was watching it separately. Version 4 shipped to all 14,000 machines that same week.
What it costs, at its worst: a cable crossover machine at one Vaultline location had a fraying pulley cable. Under version 3, that pattern would have scored around 74, flagged for service inside two days. Under version 4, the same physical machine scored 22: no action needed. Nobody checked it. Nineteen days later, the cable snapped mid-use and a member was hurt badly enough to need stitches.
What I would leave alone: the shadow gate doesn't need to slow down every release. A change to how Flywheel formats its daily report, or a small tweak to the technician's queue screen, doesn't need two weeks of shadow scoring. Save the gate for anything that touches the risk score itself.
The lesson: an eval set is a promise about which mistakes you'll catch. That promise is only as good as what the set actually contains, and a set built once, from whatever data happened to exist that day, quietly stops being a safety net for whatever the fleet grows into next.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a golden set that was 90 percent one machine type was the actual bug, not the model underneath it.
Yara Solano has run product for Flywheel for three years, since before Coppice had a single strength-equipment sensor kit built. She can read a miss-rate table the way some people read a weather map: past the average, straight to whichever row is moving.
Coppice shipped Flywheel to Vaultline eighteen months ago, cardio equipment first: treadmills, ellipticals, bikes. The golden set Yara's team built at launch, 600 cases pulled from Vaultline's own digitized repair history, was almost entirely cardio, because cardio was the only equipment type with a repair history to build one from. For over a year, that was fine. Every new model version ran against the golden set before it shipped, the blended miss rate held steady around 6 percent, and Yara signed off without a second look. She called it a clean release, every time.
Vaultline added sensors to its strength equipment in early 2025: cable machines, leg presses, pulley towers. Flywheel's fleet went from all cardio to 35 percent strength equipment inside about eight months. Yara's team never rebuilt the golden set to match. It thinned in three beats, and none of them looked careless. Beat one: the strength-equipment rollout belonged to a different pod, and nobody on Yara's side was told the golden set hadn't grown with it. Beat two: version 4 was built to fix a real, loud complaint, treadmills throwing too many false alarms, and the fix worked so well that the blended miss rate barely moved. Beat three: a blended number that barely moves reads as a clean release, so version 4 shipped to the whole fleet the same week it passed the golden set.
The trigger wasn't a dashboard turning red. Nineteen days after version 4 shipped, a member at a Vaultline location, mid-set on a cable crossover machine, felt the handle jerk hard as the frayed pulley cable gave out. She needed stitches in her forearm. Declan Marsh, who runs facilities for Vaultline, pulled the machine's sensor history that same afternoon. The vibration signature had been climbing for weeks. Under version 3, a pattern like that would have scored around 74 and gone straight into the service queue. Under version 4, it had scored 22 every single day since the cable started fraying: no action needed.
Nobody ever built a step where a member, or even a technician, could ask Flywheel to double check a machine that felt off. The queue just moved from a sensor reading to a score to a list, and whatever the score said, that was the end of it.
Yara ran the full re-check that night. Version 4's miss rate on cardio machines really had improved, 6 percent down to 3. Its miss rate on strength equipment had gone from 7 percent to 41. The blended number, 6.1 to 6.8, had made both of those look like rounding error.
The decision that opened the door traced back to a fifteen-minute call the week Coppice first shipped Flywheel to strength equipment. Someone had asked whether the golden set needed new cases for the new machine types before the next model release. The answer was "not yet, cardio is still 90 percent of the fleet." At the time, that was true. Nobody set a date to revisit it once it stopped being true.
Run the same nineteen days again, with the fix Yara would put in place. The golden set now carries a floor of at least 150 cases per equipment class, not a blended total. Version 4 fails that check immediately: a 41 percent miss rate on strength equipment against a 2 percent bar, caught in an afternoon of automated testing, before a single real machine sees it. The fix costs about $3,200 in a data scientist's time and pushes the release back a week. If it had somehow slipped past the golden set anyway, the drift alert, now split the same way, would have fired eleven days in, when version 4's strength-equipment scores started diverging hard from version 3's shadow scores on the exact same machines.
One design let a blended average decide whether a model was safe to ship everywhere at once. The other lets the group with the least coverage decide, on purpose, because that's the one with nowhere else to get caught.
What Yara would tell herself, back on that fifteen-minute call: "not yet" is not a decision, it's a decision with no date on it. The golden set needed an owner and a trigger for when to grow, not a one-time build.
GUARD, or how to price a golden set as insurance instead of overhead
Not a fairness lecture bolted onto a testing plan. GUARD is what forces you to name who gets hurt before you're allowed to talk about accuracy at all.
And if you want to be sure it really works, try it somewhere else
Same five letters, a credit union's underwriting model instead of a gym fleet, and this time the hidden harm isn't a frayed cable. It's a denied loan nobody explains.
Havenmark Credit Union runs an internal model that scores loan applications for approval and pricing. Anders Voight leads credit risk there. Havenmark's golden set for that model was built almost entirely from its largest loan category, auto loans, because that's where years of outcome data already existed. Small-business loans, a newer product line, had barely 40 labeled cases in the same set.
Same rank, different lever, mapped straight onto GUARD: the group affected is every small-business applicant scored by a version they never see, plus Havenmark's own risk team trusting one blended default-prediction accuracy number across every loan type. The harm lands unevenly on small-business applicants, because their category barely has 40 golden-set cases against auto's several thousand. An applicant denied, or offered a worse rate, after a silent model swap has no way to know the model changed and no way to ask for the old version's read on their file. The fix is a golden set with a real floor per loan category, a drift alert split the same way, and a two-week shadow period before any new version prices a single real loan. And you'd detect it working by tracking the miss rate by loan category, not blended, and pricing the gap between a golden-set catch and a real bad-loan or discrimination complaint caught later.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: split the eval set and the monitoring by loan category, gate new versions behind a shadow period, price the ROI as the gap between a same-day catch and a field catch.
Cost: no budget this quarter to build 40 more labeled cases per category. Start with the two categories that carry the most real risk, not the two easiest to label.
The model got better, for real: say accuracy on auto loans jumps 10 points overnight. That's exactly when to check the small-business split hardest, because a big win on the majority category is precisely what a blended number will use to hide a loss somewhere else.
Where people run it wrong.
They watch one blended accuracy number and call the release clean.
They build the golden set once, at launch, and never revisit it as the product's own categories change.
They treat "we'd catch it in production eventually" as good enough, without ever pricing what "eventually" actually costs.
How to use it live. Ask the split question before naming a fix: "Is the golden set actually representative of every group this model touches now, or just the group it touched when we built it?" That buys you the room to actually answer, instead of guessing at a number.
Three things worth stating directly, since this is where the real judgment sits. The alternative Yara's team considered, and rejected, was pouring the fix into a bigger, better cardio training set, since that's where the volume and the loud complaint both were. It lost, because more cardio data makes the blended number look even healthier while doing nothing for the class that was actually failing. The AI-specific failure worth naming is silent distribution shift: version 4 was never wrong about treadmills, it generalized a treadmill-shaped fix onto a motor type it had barely seen, and nothing forced it to say it wasn't sure. The guardrail is the stratified golden set plus the split drift alert, so a regression on an underrepresented group can't hide inside an improving average. And the trade-off is real: a two-week shadow gate delays every release that touches the risk score, even the safe ones, and a golden set with a real floor per group costs real data-scientist time up front. That delay and that cost get accepted on purpose, because the alternative is finding out from a member's forearm instead of a report.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the underrepresented group genuinely doesn't have enough real failure history to build a golden set from yet?" Response: that's exactly what the shadow gate is for, score it silently against the live model until enough real cases build up, instead of shipping blind to it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.