CaseAdvancedQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #15

How would you decide between a model that is right 90 percent of the time and one that is right 85 percent but says when it is unsure?

Five points of accuracy sound like a lot to give up. They are cheap next to a machine that breaks with no warning at all.

The direct answer
Pick the model that is right 85 percent of the time and says when it is unsure. The 90 percent model's mistakes look exactly like its correct calls, so nobody downstream can tell which one to double check. The 85 percent model puts most of its mistakes in a pile marked "check this one," so a person catches them before real damage happens.
Do this, in order
  1. Pick the 85 percent model that flags its own uncertain calls.Why: a wrong call you can spot is worth far less than a wrong call you can't.
  2. Confirm the flagged pile actually catches most of the real misses, on a golden set, before shipping.Why: a model that flags at random gives you the review cost with none of the safety.
  3. Route flagged calls to a person whose review time beats the model's shortest warning.Why: a flag nobody checks in time does nothing at all.
  4. Leave the plain, higher-scoring model in place where a wrong call costs almost nothing.Why: paying for review time only pays off where a miss is expensive.
  5. Watch the flagged share over time, not just accuracy.Why: if the review pile grows past what one person can clear, the plan breaks even though the model looks the same on paper.
  6. Set a kill rule: no downstream review step at all means raw accuracy wins instead.Why: a flag with nobody to read it is decoration, not a safety net.

How to answer this, stage by stage

Nobody's grading whether you know both numbers matter. They're grading whether you'll commit to one, and whether you can price the gap instead of just naming it.

1
Scope it to one plant, one tool, two real candidates
Say it like this
"Let's make this real. Cinderforge Steelworks runs about a hundred forty machines. Wearsight watches their sensors and predicts which one's about to fail. Halvor Aduba is the reliability engineer picking which version goes into production."
Why this works
A tradeoff answered in the abstract turns into "it depends." One plant, one decision, makes it a call you can defend.
2
Say your structure out loud
Say it like this
"I'm going to pick first, then show you why the accuracy gap isn't the real risk, then tell you what would make me pick the other one instead."
Why this works
Tells the interviewer you have a plan for a tradeoff question, instead of drifting between both sides until time runs out.
3
Give the position, cold, before any story
Say it like this
"I'd ship the 85 percent model, the one that says 'I'm not sure' on the calls it's shaky about. Not the 90 percent one."
Why this works
A PICK question tests whether you can commit. A reader who stops here already knows exactly what you'd do.
4
Name who feels each kind of error
Say it like this
"With the quiet 90 percent model, every wrong call looks exactly like a right one. Halvor's crew has no way to know which ten out of a hundred to double check. With the flagged model, most of the wrong calls land in a pile marked 'check this one,' and a person clears that pile before a machine actually goes down."
Why this works
Names both people the tradeoff is testing, not just a score on a slide.
5
Find the cost asymmetry, with a real number
Say it like this
"Out of six hundred forty calls a year, the 90 percent model gets sixty four wrong, all silent. The 85 percent model gets ninety six wrong, but eighty one of those land in the flagged pile. Only fifteen slip through looking confident. Fifteen quiet misses cost a lot less than sixty four."
Why this works
Turns "which model is better" into a number you can check, not a feeling about which sounds safer.
6
Give the kill criteria
Say it like this
"I'd flip my pick if nobody downstream could actually act on a flag. A machine watched by no one, or a review queue so backed up that nothing gets checked before it fails anyway. Then the flag is just decoration, and I'd take the higher raw score instead."
Why this works
Shows the pick is a real judgment call, not a rule you'd apply blindly everywhere.
7
Close on a testable bar, not a promise
Say it like this
"The bar I'd hold it to: on a rolling five hundred case audit, the flagged pile has to catch at least eighty percent of the real misses, and stay under half of all calls, or review turns into checking almost everything anyway."
Why this works
Closes with a number an interviewer can push on, which is what they're actually listening for.
If you remember one thing A wrong answer nobody can spot costs more than a wrong answer that raises its own hand. Five points of raw accuracy is a cheap price to pay for that.

Let's learn

Here is the part nobody asks first about a tool that predicts machine failures. Not how often it's right. What happens to the times it's wrong.

Wearsight is a small AI tool bolted onto the sensors at Cinderforge Steelworks, a plant that runs stamping presses, rolling mills, and air compressors around the clock. It reads vibration and heat off each machine and predicts which one is about to fail, so a crew can fix it on a Tuesday instead of on a stretcher on a Friday.

Before Wearsight, the plant serviced every machine on a fixed calendar. Every ninety days, whether it needed it or not. That caught some problems early and missed others completely. The plant still lost about eleven machines a year to a breakdown nobody saw coming, each one costing a shift or more of lost output.

Knowledge spark: what does "calibrated" mean here? A calibrated model doesn't just guess. It also guesses how good its own guess is, and that second number actually tracks reality. When it says "not sure," it really is wrong far more often than when it says "confident." A model can be calibrated and still less accurate on paper than a model with no such number at all.

Two candidate versions of Wearsight ran side by side for a season before either one got the job for real. Version one is right ninety times out of a hundred, checked against six hundred forty past failure events an engineer had already confirmed by hand. Version two is right eighty five times out of a hundred on that same set of six hundred forty. But version two does one more thing. On every call, it also says how sure it is, and marks the shaky ones "needs a look."

Five more wrong calls is not the story. Not being able to find them is.

Here is the turn. Version one's sixty four wrong calls, out of six hundred forty, look exactly like its five hundred seventy six right ones. Nobody on the floor can tell which is which by looking at the screen.

640 golden-set calls a year, sorted by what happens to the wrong ones
64 wrong, none flagged 15 wrong, silent 81 wrong, flagged Version one, 90% Version two, 85%
Right callWrong, but flaggedWrong, and silent
Both bars carry the same total, six hundred forty calls. What moves is where the wrong ones land. Version one hides all sixty four inside the right-looking bar. Version two catches eighty one of its ninety six, leaving only fifteen quiet misses.

At its worst, this costs a plant a press that seizes with no warning, because the model that said "no action needed" that morning was one of the sixty four it got wrong, and there was nothing on the screen to say so.

The decision that mattered Cinderforge's evaluation plan, written months earlier, said the higher accuracy model wins the bake off. That was fine when both candidates gave a plain yes or no with nothing else attached. It stopped being fine the moment one of them could also say "I'm not sure," and nobody rewrote the plan to ask what that was worth.

What I would leave alone: a rubber conveyor roller costs forty dollars and gets swapped either way, whether the model is confident about it or not. Nobody needs a flagged review queue for a part that cheap. Spend the extra care where a miss is expensive, not everywhere at once.

The lesson: two models are almost never at the same true cost just because their headline numbers sit close together. Ask where the wrong ten percent goes before you ask how big the ten percent is.

Now here is the same thing as a story

Read this when you want to feel why the number mattered, not just know that it did.

Halvor Aduba has kept the presses running at Cinderforge for eleven years. Before any tool watched a single sensor, he could hear a bearing going bad from the far end of the shop floor, three weeks before it seized. He timed himself once, for fun, and it held every time.

Wearsight arrived the spring before last, running quietly in the background: two candidate versions, watching every machine, logging what they would have predicted without acting on any of it. For most of that season it was a comfortable thing to have running. Halvor checked the logs on Friday afternoons, saw both versions agreeing with each other and with him, and moved on to the rest of his week.

Through the summer he checked the logs less. The two versions agreed on almost every call, and both agreed with what he already knew. There didn't seem to be much left to watch for.

Then came a Tuesday in September, and it wasn't even about Wearsight at first.

Halvor was walking past the number three stamping press on his way to get coffee, the one Wearsight's log had marked "no action needed" that morning, both versions in agreement. He caught, for half a second, a rattle under the stamping cycle that didn't belong there. He stopped the press. Inside, a bearing was one shift away from letting go, the kind of failure that would have taken the press down for four days and cost the plant a rush part flown in from three states over.

He never would have caught it from the screen. He caught it because he happened to be walking past with a coffee in his hand.

That's the part that stuck with him. Not that the model was wrong. Ninety percent right means something is going to be wrong sometimes. What stuck was that there was nothing anywhere in Wearsight's output that would have told him to go check that press, if he hadn't been walking past it that particular morning.

He pulled the logs that night. Version one had called it "no action needed," flat, the same tone it used for a machine with ten good years left. Version two had also called it "no action needed," but underneath, a second field nobody had been reading said the vibration pattern didn't match anything in its training data well, and its own confidence was low. Version two had known it wasn't sure. Nobody had built anything that put that number in front of a person.

The old decision went back months, to when the season of shadow testing was designed. Someone wrote "pick whichever version scores higher on the six hundred forty case backtest" into the plan, back when both versions only output a plain yes or no. Nobody rewrote that plan once version two started attaching a confidence field, because on the leaderboard, it just looked like a worse model.

Run the same September morning again, with version two live and its low-confidence flag actually wired to a screen a person watches. The press still gets marked "needs a look" instead of "no action needed." A technician clears the flagged queue at seven that morning, forty minutes before Halvor would have walked past with his coffee anyway. Same bearing, same three weeks of warning it should have given and didn't. This time somebody reads the flag on purpose, not by accident.

One design bet everything on a number holding still forever. The other built in a place for the model to say "check me" before it needed a lucky coffee break to get caught.

What I'd tell myself, back when that evaluation plan got written: the day a model can say "I'm not sure," the question stops being which one scores higher, and starts being what you do with the ones it flags. Nobody asked that question. That's on the plan, not on either model.

PICK, the four questions that pick the model for you

This isn't a story about a lucky coffee break. It's PICK, run on a tradeoff, using a golden set instead of a gut feeling.

PPosition. Your pick, in one sentence, before any reasoning.
Version two, the eighty five percent model that flags what it isn't sure about. Not the ninety percent model with no such flag.
Interviewers are testing whether you'll commit. "It depends" fails a PICK question before it even starts.
IImpact. Who feels each kind of error, and in what.
Halvor's crew feels version two's error as review time, about fifteen minutes checking a flagged reading. The plant feels version one's error as a machine that seizes with zero warning, because nothing on screen ever told anyone to look closer.
Naming both people, not just the model's score, is what separates a real answer from a leaderboard read.
CCost asymmetry. Which error is cheap and visible, which is hidden and expensive.
Version one's sixty four wrong calls a year are silent and expensive: some are missed failures that turn into days of downtime. Version two's ninety six wrong calls are mostly cheap and visible: eighty one sit in a flagged pile that costs review time, only fifteen slip through quiet and confident.
The heart of PICK: go after the error that's hidden and expensive, not the one that's merely more frequent.
KKill criteria. What evidence would flip the pick.
If a machine's flagged predictions had no reviewer downstream at all, or the review queue ran so far behind that a check always arrived after the failure window closed, the flag would do nothing. At that point the higher raw score is the safer pick.
This is what separates a confident answer from a stubborn one: naming what would actually change your mind.
Hand sketched comparison titled Who feels the miss. Left, a gauge icon labeled flagged uncertain, caption tech spends 10 minutes checking one reading. Right, a tilted scale icon labeled silent and wrong, caption press fails on the floor with no warning at all.
The two kinds of wrong, drawn side by side. One costs a coffee break's worth of a technician's time. The other costs a press.
Hours to clear a flagged review, by machines per reliability engineer
10 hrs 0 4 hrs, fastest warning Wearsight gives 140 machines per engineer, now 40 120 220
Review time, by staffing levelCinderforge today
Today's point sits under the four hour line, so a flagged case usually still gets checked before the fastest failure it warns about. Stretch staffing much past a hundred sixty machines per engineer and the line crosses. Past that point a flag arrives too late to matter, and this pick would flip.

Three things worth stating plainly, since this is where the real judgment sits. The alternative Cinderforge almost took was leaving version one in production and asking someone to manually double check every single prediction, all six hundred forty a year, the same rigor either way. It lost because Halvor is one person watching a hundred forty machines. There's no shift long enough to hand check that many calls with real care, and a check done in a hurry catches nothing more than the flag would have caught for free. The AI specific failure worth naming by name is a kind of confident wrongness: the model treating a vibration pattern it has never actually seen as if it were a familiar one, and answering in the same flat tone either way. The guardrail is the confidence field itself, wired to a screen a person actually watches, plus a monthly audit that checks not just accuracy but whether the flagged pile really does hold most of the real misses. A model that flags at random would still pass an accuracy check and leave Halvor exactly as blind as before. That guardrail is not free. It costs a few thousand dollars a year in review time, and it trades away five points of raw accuracy that would have looked good on a slide. It buys back the days of downtime version one's silent misses would have cost over the same year. And the bar for shipping version two isn't zero misses, a model watching a rattling machine can't promise that. It's this: on a rolling five hundred case audit, the flagged pile has to catch at least eighty percent of the real misses and stay under half of all calls, checked every quarter, not assumed once and forgotten.

And if you want to be sure it really works, try it somewhere else

Same four letters, an insurance claims desk instead of a factory floor, with nothing about machines anywhere in sight.

Northfen Mutual runs an AI tool called Lossglass that reads photos and repair shop notes for a car accident claim, and estimates what the repair should cost, so an adjuster knows roughly what to expect before they open the file. Coen Vasteras runs claims operations there.

P, position. Ship the calibrated model, the one that flags a claim "needs a second look" when it's guessing outside its comfort zone, even though it scores five points lower on the yearly accuracy check.
I, impact. The plain model's wrong estimates land the same way every time, confident and silent, whether it's overpaying a shop by thousands or underpaying a driver who then fights the claim for weeks. The calibrated model's flagged claims cost Coen's team about fifteen minutes each to check by hand.
C, cost asymmetry. Out of nine hundred claims a year, the plain model gets ninety wrong, all silent, at roughly twenty four hundred dollars each in overpay or reopened claims. The calibrated model gets about a hundred thirty wrong, but most land in the flagged pile. Only about twenty slip through looking confident.
K, kill criteria. If Northfen ever paid claims straight from Lossglass with nobody reading the flagged pile at all, the flag would do nothing, and the higher raw score would be the safer pick instead.

Estimated yearly cost, plain model versus calibrated model
$216,000/yr $54,300/yr Lossglass v1, plain Lossglass v2, calibrated
Silent overpay and reopened claimsReview time plus the few misses that slip through
Same shape as Wearsight, different numbers. The calibrated model costs less overall even with more raw mistakes, because most of them get caught before they cost anything.

Swap the trigger and it still runs.
Speed: an interviewer gives you ninety seconds. Skip straight to the position and the one number, the flagged pile catches most of the real misses.
Cost: there's no budget yet for a reviewer to clear the flagged queue. Don't skip the flag anyway, hand it to whoever already reviews the highest dollar claims, until a dedicated reviewer exists.
The model got better, for real: say next quarter's retrain lifts the plain model to ninety three percent. That's still a different claim than "we know which ones it's wrong about." A model that improves on average can keep the exact same blind spot the whole time.

Where people run it wrong.
They read the higher accuracy number and stop there, without asking what happens to the gap.
They build the flag and never wire it to a person who actually reads it, so it exists on paper only.
They let the flagged pile grow past what one reviewer can clear in a shift, and never notice review time creeping past the point where it still helps.

How to use it live. Say the tradeoff out loud before picking a side: "the real question isn't which model is more often right, it's which one's mistakes I can actually find." That line buys you a beat to think, and shows the interviewer you know what a PICK question is actually testing.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a question that asks you to pick between two options with an uneven cost?
Tap to flip
ANSWER
PICK: state your position first, name who feels each kind of error, find the cost asymmetry between them, then give kill criteria for when you'd flip.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Halvor Aduba, reliability engineer at Cinderforge Steelworks for eleven years, who could hear a bad bearing across the shop floor before any tool ever watched a sensor.
3 · THE POSITION
What did this answer pick, and why, in one line?
Tap to flip
ANSWER
Wearsight version two, right eighty five percent of the time, because its wrong calls mostly land in a pile marked "check this one" instead of hiding among the right ones.
4 · THE ASYMMETRY
What's the cost asymmetry here?
Tap to flip
ANSWER
A flagged wrong call costs a technician about fifteen minutes to check. A silent wrong call can cost days of downtime, because nothing on screen ever told anyone to look.
5 · THE KILL CRITERIA
What would flip this pick to the other model?
Tap to flip
ANSWER
If nobody downstream would ever read the flag, or the review queue backs up past the model's shortest warning, the flag stops helping and raw accuracy wins instead.
6 · THE NUMBER
Fill in the blank: out of 640 calls a year, version one gets ___ wrong, all silent. Version two gets 96 wrong, but only ___ of those slip through looking confident.
Tap to flip
ANSWER
64 wrong, all silent, for version one. Only 15 slip through confident for version two, because 81 of its 96 wrong calls land in the flagged pile.
7 · THE REPLAY
Same September morning, new design, what changes?
Tap to flip
ANSWER
The press gets marked "needs a look" instead of "no action needed." A technician clears the flag at seven that morning, forty minutes before Halvor's lucky coffee walk would have caught it anyway.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and who runs it?
Tap to flip
ANSWER
Lossglass, the repair-cost estimator at Northfen Mutual, run by Coen Vasteras. Same PICK letters, same shape, on insurance claims instead of machine failures.

Check yourself Score: 0 / 0

True or false
1. True or false: because version two gets more calls wrong overall, it is the riskier model to put into production.
  • True
  • False
Show hint
Check the golden-set chart. Same total, different split of where the wrong calls land.
Show answer
False. Raw wrongness count isn't what matters, where the wrongness lands is. Version two's misses mostly get flagged and caught before they cost anything.
Multiple choice
2. Why can't Halvor's crew tell which of version one's calls to double check?
  • A. They never got trained on the tool.
  • B. Version one gives every call the same confident tone, whether it's right or wrong.
  • C. Version one runs too slowly for anyone to review its output.
  • D. The plant removed the maintenance review queue entirely.
Show hint
Look at the September story, what Halvor found sitting in the logs that night.
Show answer
B. Version one has no second field saying how sure it is. "No action needed" looks identical whether it's right or one of the sixty four times it's wrong.
Fill in the blank
3. Out of 640 calls a year, version two flags about ___ of them as needs a look, and ___ of its 96 wrong calls fall inside that flagged pile.
Show hint
Look at the stacked bar chart in "Let's learn" and the K step's cost breakdown.
Show answer
About 256 calls (40 percent) get flagged, and 81 of the 96 wrong calls sit inside that flagged pile. That's what makes the flag worth building, most of the real misses actually show up inside it.
Short answer, name the rejected alternative
4. What alternative did Cinderforge almost take instead of shipping version two, and why did it lose?
Show hint
Look at the opening of the PICK recap's closing paragraph.
Show answer
Model answer: Keep version one in production and manually double check every single prediction, all 640 a year. It lost because Halvor is one person watching 140 machines. There's no shift long enough to hand check that volume with real care, and a rushed check catches nothing more than the flag would have caught for free.
Short answer, apply it yourself
5. Pick an AI tool you use yourself. Name a case where you'd rather it be right a little less often, if it told you when to double check it.
Show hint
Think of a tool that's usually right but never says when it's guessing off stale or unfamiliar information.
Show answer
Model answer: A navigation app that's almost always right but never flags when it's routing off a map that hasn't been updated. On the rare wrong turn, a "low confidence, road may have changed" flag beats the same confident voice it uses every other time.
Multiple choice
6. What would make the plain, higher-accuracy model the better pick instead?
  • A. If the plant hired more reliability engineers.
  • B. If nobody downstream ever actually read the flagged pile before a machine failed.
  • C. If the accuracy gap grew from five points to ten.
  • D. If the flagged pile only ever caught real misses and never flagged a correct call by mistake.
Show hint
Reread the K step, kill criteria, in the PICK recap.
Show answer
B. A flag only helps if someone downstream can act on it in time. With no reviewer or no time left before failure, the flag does nothing, and the higher raw score becomes the safer choice.
Before you close the answer
Why this works
Tests whether you'll commit to one side of a tradeoff and defend it with a real cost, or dodge into "it depends." Most candidates chase the higher accuracy number and never ask where the wrongness goes.
Follow-up traps
"What if the flagged pile turns out to be wrong half the time, just noise?" Response: then it isn't doing its job, and you'd catch that on the golden set before shipping, not after. The eighty percent catch rate is exactly the number that rules this out.

"Isn't adding a review step just slower and more expensive, no matter what?" Response: it costs a few thousand dollars a year in review time. It saves far more than that in the downtime the plain model would have missed silently. The trade only looks free standing if you don't price the alternative.
If pressed
Version two's confidence field isn't a second model bolted on. It's a distance score, how far this reading sits from anything in its training data, cheap to compute and added no real cost to how fast Wearsight runs in production.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more