ConceptAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #13
What is the design response to a model that is well-calibrated versus poorly calibrated?
GUARD the floor worker who can't check the machine's math, only trust it or not
Corrigan Bottling runs six filling lines at a contract beverage plant. Adaeze Umeh is the plant's Reliability Engineer. Ferroscope is the AI vendor tool that watches motor and bearing sensors and gives a number: how likely a part is to fail in the next 48 hours.
The direct answer
Don't design one confidence display for the model. Design two, and let each line earn its own. A line with a checked, verified number gets to show the plain percentage. A line with no track record, or one that's drifting, loses the percentage entirely and shows the raw evidence instead, plus a real measured hit rate for that specific line, until it earns the number back.
Do this, in order
Gate the raw percentage behind a per-line calibration check, not one global setting.Why: a number that's honest on Line 1 can be a coin flip on Line 5, and one shared display can't tell the operator which is which.
Replace the hidden number with the actual evidence: which sensor tripped, by how much.Why: an operator can judge raw evidence with their own experience; they can't judge a percentage they have no way to check.
Track a real, rolling hit rate per line and show that instead of the model's self-reported score.Why: the model's own number describes its training data. The rolling hit rate describes this actual line.
Alert the team automatically when predicted-vs-actual drifts past a set gap.Why: waiting for a floor incident to reveal bad calibration means the discovery costs someone their morning, or worse.
Leave the well-calibrated lines alone.Why: over-correcting a trustworthy number wastes the trust it actually earned.
How to answer this, stage by stage
The interviewer isn't checking whether you know what "calibration" means. They're checking whether you'd treat one confidence number the same on every line, or notice that the same digit can mean two different things.
Step 1
Scope it to one real plant
Say it like this
"I'll answer this for Corrigan Bottling, a contract co-packer running six filling lines, using Ferroscope's predictive-maintenance model."
Why this works
Stops "design response to calibration" from turning into a lecture on statistics.
Step 2
Name your structure
Say it like this
"I'll use GUARD. Groups affected, where the harm lands unequally, who can't push back, the actual design change, and how you'd detect it slipping."
Why this works
Tells the interviewer you have a method before you start telling a story.
Step 3
Name who can't check the number
Say it like this
"A night-shift operator can't run a statistics study on the model. They only get to trust the number on the screen or ignore it. That's the whole decision they're given."
Why this works
This is GUARD's hardest move: naming the person with no way to inspect what they're being asked to trust.
Step 4
Give the one decision
Say it like this
"Don't show the same style of number everywhere. Show the plain percentage only where it's been checked against real outcomes. Everywhere else, show the evidence and a real local hit rate instead."
Why this works
This is the direct answer, said in one breath, before any story backs it up.
Step 5
Prove it with the near miss, compressed
Say it like this
"On one of the newer lines, Ferroscope's flag read a low, reassuring number the night before a bearing seized. Nobody escalated, because the screen looked calm. It was caught by smell, not by the model."
Why this works
Turns "poor calibration is risky" into one specific, checkable failure, not a vague warning.
Step 6
Say what you'd measure after shipping
Say it like this
"I'd track the gap between what the model claims and what actually happens, per line, every week, and alert automatically once that gap crosses about twenty points."
Why this works
Shows you think about the display's own health, not just the display's design.
Step 7
Close on the one line
Say it like this
"A confidence number is only as honest as the checking behind it. Show it plainly where it's earned. Take it away where it isn't, and show your work instead."
Why this works
Restates the direct answer, ready for a live follow-up question.
Let's learn
What happens when a machine tells you the truth on day one and starts lying to you, politely, eight months later?
Ferroscope reads vibration and heat sensors on bottling-line motors and tells the team how likely a bearing is to fail soon. Before it existed, Corrigan Bottling serviced every motor on a flat 90-day schedule, whether it needed it or not, and still ate 14 unplanned stoppages a year, about six hours each. On the two original lines, once Ferroscope launched, unplanned stoppages fell to about 2 a year, each one flagged days ahead.
The number never changed shape. What it was actually measuring did.
Here's the turn: eight months later, three new lines were added, running slightly different motors Ferroscope had never seen in its training data. The model kept producing the same kind of number, a clean percentage, on those lines too. Nobody had checked whether that number still meant what it used to mean.
Ferroscope's stated confidence vs. its real hit rate, by line group
Same stated confidence, two very different realities. On the new lines the number was closer to a coin flip than a fact.
At its worst, this doesn't just mean a few false alarms. A confident-looking, low number can also tell an operator a bearing is probably fine when it isn't, which is worse than any noisy alarm, because nothing on screen tells them to double-check it.
The decision I would take back
At launch, the team picked one confidence-display style, a plain percentage badge, for every line, because at the time there was only one, well-tested line to show it on. That made sense with a single line and two years of matching data behind it. It stopped making sense the day new lines were added with no calibration history of their own, and the badge kept showing up looking exactly as confident as before.
What I would leave alone: Lines 1 and 2 don't need this fix at all. Their number has two years of real outcomes behind it, and it's still right about as often as it claims. Hiding a number that's actually earned its trust would just slow the team down for no reason.
The lesson: a model's honesty isn't a fixed trait, it's a property of the specific slice of the world it's being asked about right now. A design has to be able to say "I haven't checked this one yet" instead of always sounding equally sure.
Now here is the same thing as a story
The short version above is what you'd say out loud in the interview. Read this one for how close the actual near miss came.
Adaeze Umeh can tell, from the sound of a filling line alone, which motor is about to give her trouble. Nine years on the floor will do that. When Ferroscope went in on Lines 1 and 2, she watched it catch two real bearing failures in its first quarter that the old 90-day schedule would have missed completely.
The good months were quiet ones. By 6 a.m. most Tuesdays, Adaeze's team had already cleared the overnight flag list, usually one or two low numbers worth a glance, nothing urgent. The floor trusted the screen the way they trusted a coworker who was almost always right.
Knowledge spark: what does "calibrated" actually mean?
A model is calibrated when its own number matches reality: if it says "80% likely," that thing should actually happen about 8 times out of 10, checked against real cases. A model can be very accurate on average and still badly calibrated on a slice of data it's never really learned, like a brand new line.
Eight months in, three new filling lines came online, using a slightly different motor mount Ferroscope had never trained against. The flags kept appearing, worded exactly the same way. Nobody on the floor had a reason to treat them differently. Why would they?
Four flags, same stated confidence. Only two of them had actually earned it.
On a Thursday night, a bearing on Line 5 threw off a reading. Ferroscope's flag came back low, a reassuring number that read like "probably fine, check it next week." The old 90-day schedule would have had a technician's hands on that exact bearing within four days regardless. Nobody escalated. The screen looked calm.
The bearing on Line 5 seized fourteen hours later. Nobody got hurt. It was caught by a technician smelling hot grease across the room, not by the model that was supposed to be watching for exactly this.
Adaeze's plant ops director pulled ten flagged predictions at random the following week, from three different lines, and checked each one against what had actually happened. On the two original lines, the numbers held up close to what they claimed. On the two newest lines, of the flags stating 80% confidence or higher, only about 38% had actually come true. The badge had been quietly wrong for months, and it had never once said so.
Adaeze's team could go check the model's math after the fact. The operator on shift at 2 a.m. never could.
Adaeze's redesign split the display in two. Verified lines kept their plain number. Unverified or drifting lines lost the number and got the raw evidence instead: which sensor tripped, by how much, and a real hit rate built from that specific line's own history, however thin it still was.
Same digits on screen. One had two years of proof behind it. The other had none.
GUARD, the whole page in one lookNot a fairness lecture. GUARD is what tells you which number is allowed to sound confident.
G
Groups. Who's actually affected.
Adaeze's team, who can pull records and re-check the model later, and the operator on shift, who only sees the flag once.
Sets up who has power over the decision, and who doesn't.
U
Unequal. Where the harm lands hardest.
The newest lines, the ones the model was never trained on, carry the worst miscalibration. Whoever is on shift there absorbs the most risk.
The harm isn't spread evenly across the plant; it concentrates exactly where the model knows least.
A
Ability to contest. Who can't push back.
A floor operator has no way to independently check whether "80% confidence" is real. They can only defer to it or override it on a hunch.
This is the hardest step and the answer to the question: the design has to give that person something they can actually judge, not a number they can't.
R
Reduce. The actual design change.
Show the plain percentage only on lines with a checked history. Everywhere else, show the sensor evidence and a real local hit rate instead.
Turns "be careful with the number" into a concrete rule about what appears on screen and when.
D
Detect. How you'd know it's slipping.
A rolling, per-line check of predicted-versus-actual, alerting automatically once the gap passes about twenty points.
Catches the next drift before a floor incident has to do the catching instead.
One question decides the whole display. Everything else follows from the answer.
Gap between Ferroscope's stated confidence and Line 5's real hit rate, week by week
The gap crossed the concern line five weeks before the near miss on Line 5. Nobody was watching it, because nothing before this redesign tracked it at all.
One box in the middle was empty the whole time. Nothing there ever asked whether the number could be trusted yet.
The recap, one line per letter: groups is the team that can check versus the operator who can't, unequal is the newest lines carrying the worst risk, ability to contest is the operator with no way to verify the number themselves, reduce is splitting the display by calibration status, and detect is a rolling per-line accuracy check with an automatic alert.
And if you want to be sure it really works, try it somewhere elseSame five letters, a crop-disease detector on a farm instead of bearing sensors on a bottling line. A different subject carries the risk this time.
Ridgeback Growers Cooperative uses an AI tool that scans leaf photos and estimates how likely a crop is infected with a fungal blight, so an agronomist named Zuriel Wanjiru can decide which fields need spraying this week, not next. The tool launched trained on the cooperative's three original crop varieties and was solidly calibrated on all three. Then the co-op added two new varieties this season, grown from different seed stock the model had never seen.
Mapped onto GUARD: groups is the co-op's agronomy team, who can walk a field and double-check by eye, versus the smaller growers who lease equipment and trust the app's number outright because they can't afford the agronomist's time. Unequal is that the new varieties, again, carry the worst miscalibration: stated high-risk flags on the three original varieties were confirmed 82% of the time, but on the two new ones that fell to about 51%, close to a coin flip. Ability to contest is the leaseholder growers, who have no field history of their own to compare the number against. Reduce is the same fix: show the plain infection-probability number only for the three original varieties, and for the two new ones, show the raw leaf symptoms the model actually detected, plus how many total scans of that variety exist so far.
Same shape of gap. A different subject carries it this time: the grower with no agronomist on staff.
The old decision here isn't quite the same one. Ridgeback's team hadn't picked one universal display style out of convenience, they'd made a packaging decision: the app charged growers per scan, which pushed leaseholders to save their scans for the crops they were most worried about, meaning the new varieties, exactly the ones with the least calibration data, got scanned the most by the people least able to double-check the result.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "split the display by whether the number has been checked; show it plain where it's earned, show the evidence instead where it isn't," and stop there.
Cost: there's no budget to rebuild the display right now. Say so, and start by just logging predicted-versus-actual per line for free; that alone tells you where the real risk sits before you spend anything on the fix.
The model gets better, for real: if Ferroscope is retrained and a new line's calibration checks out clean for a full quarter, that line graduates back to showing the plain number, the same rule running in reverse.
Where people run it wrong.
They treat "the model is usually accurate" as the same thing as "the model is calibrated everywhere."
They let a single global display setting quietly outlive the one line it was actually tested on.
They wait for an incident to reveal drift instead of tracking the predicted-versus-actual gap on their own.
How to use it live. When someone asks how you'd handle a well- or poorly-calibrated model, ask back: has this exact slice of the model's job actually been checked against real outcomes recently? Let the honest answer, not the model's confidence, decide what the screen is allowed to say.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a risk/safety question about calibration, and what does each letter do?
Tap to flip
ANSWER
GUARD: groups, unequal harm, ability to contest, reduce, detect. It finds who can't check the model's math and designs around that gap.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Adaeze Umeh, Reliability Engineer at Corrigan Bottling, who can tell a failing motor by its sound after nine years on the floor.
3 · WHO CAN'T PUSH BACK
In GUARD's terms, who is the subject with no way to contest the number here?
Tap to flip
ANSWER
The floor operator on shift. They can only trust the flag or override it on a hunch; they have no way to independently check whether the percentage is real.
4 · THE DESIGN CHANGE
What's the actual reduce-step design decision this answer lands on?
Tap to flip
ANSWER
Show the plain percentage only on lines with a checked calibration history. Everywhere else, show the raw sensor evidence and a real local hit rate instead.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Picking one universal confidence-display style at launch, when only one, well-tested line existed to show it on, and never revisiting that choice as new lines were added.
6 · THE NUMBER
Fill in the blank: on the two new lines, of flags stating 80% confidence or higher, only about ___ percent actually came true.
Tap to flip
ANSWER
About 38 percent. Close to a coin flip, on a number that looked exactly as confident as the well-calibrated lines' 78 percent.
7 · THE REPLAY
Same near miss, new design, split display by calibration status. What changes?
Tap to flip
ANSWER
Line 5's flag no longer shows a falsely reassuring percentage; it shows raw sensor evidence and a thin, honestly-labeled local history, prompting a manual check instead of silence.
8 · CROSS PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which product, and who's the subject who can't push back there?
Tap to flip
ANSWER
Ridgeback Growers Cooperative's crop-blight detector. The subject is the leaseholder grower with no agronomist on staff and no field history of their own to compare against.
Check yourself Score: 0 / 0
Multiple choice
1. Why is a floor operator the person GUARD calls "unable to contest" in this story?
A. They aren't trained to read the sensor dashboard.
B. They have no way to independently check whether the model's stated confidence is actually accurate.
C. Company policy forbids them from questioning the AI.
D. They work the night shift, when nobody else is around.
Show hint
Look at the "ability to contest" step.
Show answer
B. They can defer to the number or override it on a hunch, but they have no way to check it against reality themselves.
True or false
2. True or false: Ferroscope's stated confidence on the new lines was actually less accurate than on the original lines, even though the numbers looked the same on screen.
True
False
Show hint
Look at the grouped bar chart.
Show answer
True. Both showed "80% confidence," but the real hit rate was 78% on original lines and only 38% on new ones.
Fill in the blank
3. Fill in the blank: at Ridgeback Growers, stated high-risk flags on original crop varieties were confirmed 82% of the time. On the new varieties, that fell to about ___ percent.
Show hint
Look at the Ridgeback Growers numbers in Section 4.
Show answer
51 percent. Barely better than guessing, on the exact crops the least-resourced growers leaned on most.
Short answer, apply it yourself
4. Think of an app you use that shows you a confidence number or a star rating. Has it ever been checked against real outcomes for your specific situation, or are you trusting it the same way everywhere?
Show hint
Ask whether the number was ever verified for a case like yours specifically, not just on average.
Show answer
Model answer: Most people trust it evenly everywhere, exactly the gap this answer is built to close.
Short answer, why no middle setting
5. Why wasn't "just tell operators to be a little more careful with the new lines" a real fix?
Show hint
Look at the disqualified fixes in the reduce step.
Show answer
Model answer: "Be more careful" gives the operator no new information. They still can't tell a trustworthy 80% from an untrustworthy one without a design change that shows the difference.
Short answer, where it wouldn't matter
6. Name a part of the Ferroscope relationship where this same calibration scrutiny genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Lines 1 and 2. Two years of matching outcomes means that number has actually earned the right to sound confident.
Before you close the answer
Why this works
Tests whether you treat "the model is confident" and "the model is trustworthy" as the same thing, and whether you can name who actually pays when they aren't.
Follow-up traps
"Isn't hiding the number just as risky as showing a wrong one?" Response: no, because hiding it comes with the evidence and a real local hit rate; the operator loses a fake precision, not real information.
"What if a line never builds up enough history to unlock the number?" Response: then it keeps showing evidence indefinitely; that's the honest outcome, not a bug in the design.
If pressed
The rolling calibration check used a minimum sample size of 25 flags per line before it would trust its own drift number, since a gap measured on 4 or 5 cases is noise, not a real signal.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.