How would you decide between a model that is right 90 percent of the time and one that is right 85 percent but says when it is unsure?
Five points of accuracy sound like a lot to give up. They are cheap next to a machine that breaks with no warning at all.
- Pick the 85 percent model that flags its own uncertain calls.Why: a wrong call you can spot is worth far less than a wrong call you can't.
- Confirm the flagged pile actually catches most of the real misses, on a golden set, before shipping.Why: a model that flags at random gives you the review cost with none of the safety.
- Route flagged calls to a person whose review time beats the model's shortest warning.Why: a flag nobody checks in time does nothing at all.
- Leave the plain, higher-scoring model in place where a wrong call costs almost nothing.Why: paying for review time only pays off where a miss is expensive.
- Watch the flagged share over time, not just accuracy.Why: if the review pile grows past what one person can clear, the plan breaks even though the model looks the same on paper.
- Set a kill rule: no downstream review step at all means raw accuracy wins instead.Why: a flag with nobody to read it is decoration, not a safety net.
How to answer this, stage by stage
Nobody's grading whether you know both numbers matter. They're grading whether you'll commit to one, and whether you can price the gap instead of just naming it.
Let's learn
Here is the part nobody asks first about a tool that predicts machine failures. Not how often it's right. What happens to the times it's wrong.
Wearsight is a small AI tool bolted onto the sensors at Cinderforge Steelworks, a plant that runs stamping presses, rolling mills, and air compressors around the clock. It reads vibration and heat off each machine and predicts which one is about to fail, so a crew can fix it on a Tuesday instead of on a stretcher on a Friday.
Before Wearsight, the plant serviced every machine on a fixed calendar. Every ninety days, whether it needed it or not. That caught some problems early and missed others completely. The plant still lost about eleven machines a year to a breakdown nobody saw coming, each one costing a shift or more of lost output.
Two candidate versions of Wearsight ran side by side for a season before either one got the job for real. Version one is right ninety times out of a hundred, checked against six hundred forty past failure events an engineer had already confirmed by hand. Version two is right eighty five times out of a hundred on that same set of six hundred forty. But version two does one more thing. On every call, it also says how sure it is, and marks the shaky ones "needs a look."
Here is the turn. Version one's sixty four wrong calls, out of six hundred forty, look exactly like its five hundred seventy six right ones. Nobody on the floor can tell which is which by looking at the screen.
At its worst, this costs a plant a press that seizes with no warning, because the model that said "no action needed" that morning was one of the sixty four it got wrong, and there was nothing on the screen to say so.
What I would leave alone: a rubber conveyor roller costs forty dollars and gets swapped either way, whether the model is confident about it or not. Nobody needs a flagged review queue for a part that cheap. Spend the extra care where a miss is expensive, not everywhere at once.
The lesson: two models are almost never at the same true cost just because their headline numbers sit close together. Ask where the wrong ten percent goes before you ask how big the ten percent is.
Now here is the same thing as a story
Read this when you want to feel why the number mattered, not just know that it did.
Halvor Aduba has kept the presses running at Cinderforge for eleven years. Before any tool watched a single sensor, he could hear a bearing going bad from the far end of the shop floor, three weeks before it seized. He timed himself once, for fun, and it held every time.
Wearsight arrived the spring before last, running quietly in the background: two candidate versions, watching every machine, logging what they would have predicted without acting on any of it. For most of that season it was a comfortable thing to have running. Halvor checked the logs on Friday afternoons, saw both versions agreeing with each other and with him, and moved on to the rest of his week.
Through the summer he checked the logs less. The two versions agreed on almost every call, and both agreed with what he already knew. There didn't seem to be much left to watch for.
Then came a Tuesday in September, and it wasn't even about Wearsight at first.
Halvor was walking past the number three stamping press on his way to get coffee, the one Wearsight's log had marked "no action needed" that morning, both versions in agreement. He caught, for half a second, a rattle under the stamping cycle that didn't belong there. He stopped the press. Inside, a bearing was one shift away from letting go, the kind of failure that would have taken the press down for four days and cost the plant a rush part flown in from three states over.
That's the part that stuck with him. Not that the model was wrong. Ninety percent right means something is going to be wrong sometimes. What stuck was that there was nothing anywhere in Wearsight's output that would have told him to go check that press, if he hadn't been walking past it that particular morning.
He pulled the logs that night. Version one had called it "no action needed," flat, the same tone it used for a machine with ten good years left. Version two had also called it "no action needed," but underneath, a second field nobody had been reading said the vibration pattern didn't match anything in its training data well, and its own confidence was low. Version two had known it wasn't sure. Nobody had built anything that put that number in front of a person.
The old decision went back months, to when the season of shadow testing was designed. Someone wrote "pick whichever version scores higher on the six hundred forty case backtest" into the plan, back when both versions only output a plain yes or no. Nobody rewrote that plan once version two started attaching a confidence field, because on the leaderboard, it just looked like a worse model.
Run the same September morning again, with version two live and its low-confidence flag actually wired to a screen a person watches. The press still gets marked "needs a look" instead of "no action needed." A technician clears the flagged queue at seven that morning, forty minutes before Halvor would have walked past with his coffee anyway. Same bearing, same three weeks of warning it should have given and didn't. This time somebody reads the flag on purpose, not by accident.
One design bet everything on a number holding still forever. The other built in a place for the model to say "check me" before it needed a lucky coffee break to get caught.
What I'd tell myself, back when that evaluation plan got written: the day a model can say "I'm not sure," the question stops being which one scores higher, and starts being what you do with the ones it flags. Nobody asked that question. That's on the plan, not on either model.
PICK, the four questions that pick the model for you
This isn't a story about a lucky coffee break. It's PICK, run on a tradeoff, using a golden set instead of a gut feeling.
Three things worth stating plainly, since this is where the real judgment sits. The alternative Cinderforge almost took was leaving version one in production and asking someone to manually double check every single prediction, all six hundred forty a year, the same rigor either way. It lost because Halvor is one person watching a hundred forty machines. There's no shift long enough to hand check that many calls with real care, and a check done in a hurry catches nothing more than the flag would have caught for free. The AI specific failure worth naming by name is a kind of confident wrongness: the model treating a vibration pattern it has never actually seen as if it were a familiar one, and answering in the same flat tone either way. The guardrail is the confidence field itself, wired to a screen a person actually watches, plus a monthly audit that checks not just accuracy but whether the flagged pile really does hold most of the real misses. A model that flags at random would still pass an accuracy check and leave Halvor exactly as blind as before. That guardrail is not free. It costs a few thousand dollars a year in review time, and it trades away five points of raw accuracy that would have looked good on a slide. It buys back the days of downtime version one's silent misses would have cost over the same year. And the bar for shipping version two isn't zero misses, a model watching a rattling machine can't promise that. It's this: on a rolling five hundred case audit, the flagged pile has to catch at least eighty percent of the real misses and stay under half of all calls, checked every quarter, not assumed once and forgotten.
And if you want to be sure it really works, try it somewhere else
Same four letters, an insurance claims desk instead of a factory floor, with nothing about machines anywhere in sight.
Northfen Mutual runs an AI tool called Lossglass that reads photos and repair shop notes for a car accident claim, and estimates what the repair should cost, so an adjuster knows roughly what to expect before they open the file. Coen Vasteras runs claims operations there.
P, position. Ship the calibrated model, the one that flags a claim "needs a second look" when it's guessing outside its comfort zone, even though it scores five points lower on the yearly accuracy check.
I, impact. The plain model's wrong estimates land the same way every time, confident and silent, whether it's overpaying a shop by thousands or underpaying a driver who then fights the claim for weeks. The calibrated model's flagged claims cost Coen's team about fifteen minutes each to check by hand.
C, cost asymmetry. Out of nine hundred claims a year, the plain model gets ninety wrong, all silent, at roughly twenty four hundred dollars each in overpay or reopened claims. The calibrated model gets about a hundred thirty wrong, but most land in the flagged pile. Only about twenty slip through looking confident.
K, kill criteria. If Northfen ever paid claims straight from Lossglass with nobody reading the flagged pile at all, the flag would do nothing, and the higher raw score would be the safer pick instead.
Swap the trigger and it still runs.
Speed: an interviewer gives you ninety seconds. Skip straight to the position and the one number, the flagged pile catches most of the real misses.
Cost: there's no budget yet for a reviewer to clear the flagged queue. Don't skip the flag anyway, hand it to whoever already reviews the highest dollar claims, until a dedicated reviewer exists.
The model got better, for real: say next quarter's retrain lifts the plain model to ninety three percent. That's still a different claim than "we know which ones it's wrong about." A model that improves on average can keep the exact same blind spot the whole time.
Where people run it wrong.
They read the higher accuracy number and stop there, without asking what happens to the gap.
They build the flag and never wire it to a person who actually reads it, so it exists on paper only.
They let the flagged pile grow past what one reviewer can clear in a shift, and never notice review time creeping past the point where it still helps.
How to use it live. Say the tradeoff out loud before picking a side: "the real question isn't which model is more often right, it's which one's mistakes I can actually find." That line buys you a beat to think, and shows the interviewer you know what a PICK question is actually testing.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't adding a review step just slower and more expensive, no matter what?" Response: it costs a few thousand dollars a year in review time. It saves far more than that in the downtime the plain model would have missed silently. The trade only looks free standing if you don't price the alternative.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?