Describe the second-order costs of an AI feature that leadership usually forgets.
Kerf reads a CAD or STL file before it prints and scores how likely the job is to finish clean. Tolerance Labs builds it. Thorncrest Manufacturing runs it across forty printers doing on-demand parts for industrial and medical clients. Anneliese Feng owns what Kerf's ROI number says back at Tolerance Labs. Dalton Bristow runs the print floor at Thorncrest. One overnight batch on a printer Kerf barely knew taught both of them that a feature's savings and its real cost can live in two places that never talk to each other.
- Tag every failed print with the model's original score and the combo's completed-print count.Why: this is the fix the whole gap turns on. Without it, no cost ever links back to the feature that cleared it.
- Track cost-to-serve weekly, split by combo, never blended shopwide.Why: a blended number, 3.1 percent, hid a 22 percent problem sitting right inside it.
- Route any combo under 150 completed prints through mandatory human review, no matter the score.Why: that's exactly where the model's real evidence is thin, and nothing on the screen shows that.
- Name the groups paying for this inside the ROI review itself, not just on leadership's dashboard.Why: the floor team and the support team, not leadership, absorb this cost first.
- Give floor staff a real field to flag a failure against the feature that cleared it.Why: without a lever, the people who see the real cost every week can't get it counted anywhere.
- Leave the global pass threshold alone for combos the model already knows well.Why: tightening it everywhere slows the ninety percent of the floor that isn't the problem, to fix the part that is.
How to answer this, stage by stage
Nobody is grading whether you can list "support tickets" and "maintenance" as hidden costs. They're grading whether you can name the one cost that's specific to a model being wrong quietly, and say exactly who pays for it.
Let's learn
What happens when an AI checker clears a file on a printer-material pairing it's only seen finish about forty prints?
Kerf reads a CAD or STL file before it goes to print and looks for the things that ruin a job: walls too thin to hold, geometry with no clear inside or outside, an overhang with nothing under it. Then it gives the file a predicted success score, built from every finished print Tolerance Labs has on record for that exact printer and material.
Before Kerf, a Thorncrest tech eyeballed every file by hand, about twenty five minutes each, and still watched one job in eight fail partway through a print nobody could fully explain. On a printer and material combo Kerf knows well, that drops hard: under five minutes a file, and under two failures in a hundred.
Here's the turn. Those remaining failures on well-known combos are not the real problem. The real problem is a printer-material combo Kerf has barely seen. A brand new resin on a brand new printer looks, on Kerf's screen, exactly like a combo it has scored ten thousand times. Same number, same green light.
What it costs, at its worst: Thorncrest added a large-format resin printer, the Aurelius 900, running a new engineering resin called Amberlyn 60. An overnight batch of six medical-device housings on that combo scored 94. Four of the six delaminated partway up the wall, a failure mode nobody had fed Kerf enough examples of yet. About $432 in resin, fourteen machine hours, and a rush shipment to make the client's deadline, all gone.
What I would leave alone: for combos with real history behind them, Kerf's single score is doing exactly what it should. Routing all of that through a human check too would slow the ninety percent of the floor that isn't the problem, to fix the ten percent that is.
The lesson: an ROI number built only from what a model stopped is a number that can never see what it let through. And the group that pays for what it let through is never the group reading that number.
Now here is the same thing as a story
Read the short version above when you're actually in the interview chair. Read this one when you want to feel exactly what forty finished prints buys you, next to four thousand.
Dalton Bristow can tell you which way a warped part will pull before the print is even finished, just from how the first layer laid down. He's run the print floor at Thorncrest Manufacturing for six years, back when every file got checked by eye and every overnight run got a walk-by at midnight, whether it needed one or not.
Kerf arrived about two years ago. For most of that time it was the best thing to happen to Dalton's night shift. Upload a file, get a score back in under a minute, and if it cleared north of 90 he'd start the print and go check something else. The failure rate on cleared jobs held under two percent, month after month. He stopped doing the midnight walk-by on anything Kerf had cleared. Then he stopped opening the file preview at all before hitting start. Why would he. The number had never once lied to him.
Thorncrest added the Aurelius 900, a large-format resin printer, about three months before any of this, running a new engineering resin called Amberlyn 60. Kerf scored files on that combo the same way it scored everything else: same number, same green light, same screen. Dalton fed it files the same way too. Nobody told him not to.
The trigger wasn't a dashboard. It was the newest tech on the floor, maybe three weeks in, looking over Dalton's shoulder at two files side by side. "Why does the new resin part have the same score as the nylon one we've run a thousand times?" Dalton didn't have an answer. He'd never once thought to ask it himself.
That question sent Dalton back through six weeks of teardown photos. He found the batch from three weeks earlier: six medical device housings, all Amberlyn 60 on the Aurelius 900, Kerf score 94, run overnight. Four of the six had delaminated partway up the wall, a failure mode nobody on the model side had seen enough of yet to teach Kerf to watch for. About $432 in resin, gone. Fourteen hours of machine time, gone. A rush shipment to make the client's deadline, gone too, at extra cost.
Nobody had flagged it as a Kerf problem. It went into the shop's operations log as a print failure, same bucket as a power blip or a bad batch of resin from the supplier. Nothing in that log ever asked what Kerf had said about the file first.
Dalton pulled three more incidents out of the same six weeks once he knew what to look for, all on the same new combo, none of them tagged, together worth about $2,200. Anneliese Feng, who owns Kerf's ROI number back at Tolerance Labs, had never seen any of it. Not because she'd hidden it. Because nobody had ever built a way for it to reach her.
The decision that opened the door traced back to a short call, months earlier, when Tolerance Labs first wired Kerf's scores into Thorncrest's shop system. Someone asked whether the failure log should also record Kerf's score at the time of print, so the two could be compared later. The answer was, we can join them by timestamp if we ever need to. Nobody ever built that join.
Run the same six weeks again with the fix in place. Every failed print now carries Kerf's original score and the combo's finished-print count at teardown. Any file on a combo under 150 finished prints goes to a human check first, no matter the score. The doomed batch never runs unwatched: a tech catches the delamination risk on the second housing, twenty minutes in, and pulls the job. Cost: about fifteen minutes of a tech's time, instead of four ruined parts and a rush shipment.
One design let a clean-looking number decide whether an unfamiliar resin got to run overnight, unwatched. The other lets the combo's own thin history decide that for itself.
What I'd tell myself, back on that short call: "we can join them later" is not a decision, it's a decision with nobody's name on it. And the floor pays for that gap every week nobody closes it.
GUARD, or finding the bill that never got mailed
Not a list of things that could go wrong. GUARD is what forces you to say who pays, before you're allowed to talk about savings at all.
And if you want to be sure it really works, try it somewhere else
Same five letters, a crop scanner instead of a print checker, and the hidden bill lands on a field instead of a build plate.
Pemberwick Growers is a small farming co-op. Its members use an app called Leafline to photograph crop leaves and get a disease-risk score before deciding whether to spray. Leafline's model trained mostly on the co-op's older, common crop varieties, because that's where years of labeled photos already existed. A newer, drought-hardy variety the co-op started planting last season has barely a few hundred labeled photos behind it.
Same rank, different lever, mapped straight onto GUARD: the groups are the grower working the new variety's fields, who never sees a confidence range, just a clean score, and Pemberwick's own agronomy team, led by Delia Bonsu, who reports "sprays avoided" to the co-op's board without knowing which of those avoided sprays sat on thin data. The harm lands hardest on growers of the new variety, because a few hundred labeled photos can't calibrate a score the way years of data can. A grower who loses part of a field to a missed outbreak has no way to prove Leafline cleared that field, and no lever to get that loss counted against the app's numbers, it's just called a bad season. The fix is tagging every disease report with whatever score Leafline gave that field earlier, and routing any variety under a real photo floor through an agronomist's own check before the score decides anything alone. And you'd detect it working by tracking crop loss per feature clearance, split by variety, not blended across the whole co-op.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to it: tag the score against the outcome, route thin-data varieties through a human check, track loss by variety instead of a co-op-wide average.
Cost: no budget to label more photos this season. Start with whichever variety carries the most planted acres, not whichever is easiest to label.
The model got better, for real: say overall disease-catch accuracy jumps ten points after a retrain. That's exactly when to check the thin-data variety hardest, because a strong average is precisely what a blended number will use to hide a variety still running blind.
Where people run it wrong.
They watch one co-op-wide accuracy number and call the rollout a win.
They build the training set once, at launch, and never revisit it as new varieties get planted.
They treat "we'd notice a bad season eventually" as good enough, without ever pricing what "eventually" actually costs.
How to use it live. Ask the coverage question before naming a fix: "is the training data actually spread across every group this model scores now, or just the group it scored when it launched?" That buys you the room to answer instead of guessing.
Three things worth stating directly, since the real judgment sits here. The alternative Tolerance Labs could have taken instead of routing new combos through review was simply tightening Kerf's pass threshold everywhere, flagging more borderline files across the whole shop. It lost, because it would have slowed roughly ninety percent of jobs, the ones running on combos Kerf already knows well, to fix a risk concentrated in the combos it doesn't. The AI-specific failure worth naming is distribution shift on a cold-start combo: Kerf was never wrong about resin in general, it generalized a pattern learned from other combos onto a printer and material pairing it had barely seen finish a print, and nothing forced it to say how thin that evidence actually was. The guardrail is the 150-print floor plus the split cost-to-serve number, so a blind spot like that shows up as a number instead of a surprise. And the trade-off is real: routing new combos through mandatory review adds review time to exactly the jobs that can least afford to wait, and that cost gets accepted on purpose, because the alternative is finding out from a teardown instead of a report.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a new combo genuinely doesn't have enough prints yet to build real history from?" Response: that's exactly what the 150-print floor and mandatory review are for, route it through a person until the combo earns enough finished prints to trust the score alone.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.