CaseIntermediateDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #12
How should the interface change when the model's confidence is very low?
FLIPS delegation flip: she took the whole job back, when only one line needed a flag
Meridian Regional Hospital's radiology department uses ScanDraft, an AI tool that drafts a first-pass report from a scan before a resident and attending review it. Dr. Esperanza Cuartas is the senior attending. Bram Vandersteen is the resident she supervises most closely.
The direct answer
A low confidence score should not just show a smaller number. It should change the shape of the screen: the sentence itself has to read as hedged, not confident, the section has to sit inside a visibly different state, and the report can't finalize until a named reviewer takes one explicit action on that specific section. A quiet number next to confident-sounding prose gets skipped exactly when it matters most.
Change it in this order
Give very low confidence its own visibly different interface state, not just a smaller number.Why: a subtle score sitting beside otherwise normal text is easy to skim straight past.
Change the language itself, from a confident sentence to a hedged one.Why: a reader defaults to the wording next to them, and confident wording beats a small badge every time.
Require one explicit action on the flagged section before the report can finalize.Why: makes the low-confidence state impossible to scroll past by accident.
Route the flagged section to a specifically named second reviewer.Why: a flag that goes back to the same person who already reviewed it isn't really a second set of eyes.
Track whether the flag actually changes what a reviewer does, monthly.Why: a flag that never changes an outcome is decoration, not a safety feature.
How to answer this, stage by stage
The interviewer isn't testing whether you know confidence scores exist. They're testing whether you know a number alone doesn't change anyone's behavior.
Stage 1
Scope it to one real screen
Say it like this
"I'll answer this for ScanDraft at Meridian Regional Hospital, specifically the moment its confidence on a finding drops very low, and what Dr. Cuartas and her resident Bram actually see when that happens."
Why this works
Keeps the answer from becoming a general opinion about confidence scores in the abstract.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit that formed, identify the flip a bad design causes, pinpoint the old decision behind it, and show the replay with the fix in place."
Why this works
Shows this is a perturbation question, what changes when confidence swings low, not a static design brief.
Stage 3
Name the habit a flat design builds
Say it like this
"When every draft reads equally confident, a reviewer stops looking for the uncertain ones specifically. They review everything at the same speed, because nothing on the page tells them to slow down."
Why this works
Sets up why a flat interface is dangerous well before the incident, not just after it.
Stage 4
Give the interface change, before any reasoning
Say it like this
"Very low confidence gets hedged language, a distinct visual state, and a required explicit sign-off from a named second reviewer, not just a lower number sitting beside otherwise normal prose."
Why this works
This is the direct answer, and it names a change to the screen, not a promise to "add more review."
Stage 5
Prove it with the compressed near miss
Say it like this
"An ambiguous mass finding drafted at low confidence read exactly like every routine fracture read Bram had signed off that week. He didn't slow down, because nothing told him to. Dr. Cuartas caught it days later, by luck, on a second look."
Why this works
Turns "low confidence should change the interface" into a specific, believable failure.
Stage 6
Say what you'd measure to know the flag is working
Say it like this
"I'd track how often a flagged section actually gets caught by the named reviewer before sign-off, not just how often the flag appears. If the catch rate doesn't rise, the flag isn't doing anything."
Why this works
Shows the fix gets checked against real behavior, not assumed to work because it shipped.
Stage 7
Close on the one line
Say it like this
"A confidence number that doesn't change how the text reads or what a reviewer has to do is just decoration. Very low confidence should change the screen itself, not just the score printed on it."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Six months of good, then the snap
Every morning, Dr. Esperanza Cuartas walked the reading room floor before the residents arrived, checking which studies had come in overnight. For years, that meant she'd personally read every image herself, about 18 minutes a study, roughly 20 studies a day.
Five steps, and only one of them is genuinely hard to find.
ScanDraft drafts a preliminary report in under 90 seconds. Once Bram and the other residents started reviewing and editing those drafts instead of reading every image cold, Dr. Cuartas's own co-review time per study fell to about 5 minutes.
Studies where a resident caught an ambiguous finding without being prompted, by report design
Same residents, same ambiguous cases. The flag alone nearly quadrupled how often the hard finding got caught.
Here's the turn: the co-review time Dr. Cuartas got back was never the real story. The real problem was a draft that read exactly as confident on an ambiguous mass as it did on a routine fracture, so nothing told a resident which one actually needed a slower read.
A dial has room for nuance. A flip only has room for a person to notice, or not.
At its worst, a flat report design doesn't just cost an extra hour of careful reading. It lets an ambiguous, genuinely uncertain finding slide through as if it were routine, delaying a diagnosis by weeks until something else forces a second look.
The decision I would take back
ScanDraft's report card rendered every finding in the same font, the same confident tone, whether the model was highly certain or barely guessing. That made sense at launch, when confidence was usually high and every single draft got a careful, unhurried co-review regardless. It stopped making sense once residents, trusting the drafts, started reviewing faster.
What I would leave alone: a clean, high-confidence fracture read doesn't need this same visible hedging. Slowing every routine read down to match the ambiguous ones would waste the exact time the tool was built to save.
The lesson: a confidence number sitting quietly beside confident-sounding prose doesn't protect anyone. The prose has to change too, or the number is just decoration nobody has time to read.
Now here is the same thing as a story
The short version above is what you'd say defending this redesign to Meridian's chief of radiology. Read this one for how the actual near miss unfolded.
Bram had been under Dr. Cuartas's supervision for eight months, and the arrangement worked the way good training is supposed to: he drafted, she reviewed closely at first, then more lightly as his judgment proved itself case after case. ScanDraft made both of their jobs faster without changing that trust.
Knowledge spark: why would a model be unsure about one finding and confident about another?
A clean fracture on an x-ray has sharp, well-defined edges the model has seen thousands of times. An ambiguous soft-tissue mass can look like several different things depending on faint shading a model is genuinely less certain about. The model isn't being careless on the hard case. It's being honest that the hard case is actually hard, if the interface lets that honesty show.
Six months in, Bram was reviewing ScanDraft's drafts quickly and confidently, and Dr. Cuartas's spot-checks kept coming back clean. Nothing about the reports themselves ever looked different from one case to the next, so neither of them had reason to slow down on any particular one.
Six good months hid a habit that had already changed. Month eight is where the flat design ran out of luck.
The trigger was a single case: an ambiguous mass on a chest scan, drafted by ScanDraft at genuinely low confidence, worded in the same clean, declarative sentences as every clear fracture read that week. Bram signed off in the time he'd normally spend, since nothing on the page suggested this one deserved longer.
Bram didn't skip a step. The report gave him no reason to believe this case was any different from the twenty routine ones before it.
Dr. Cuartas caught it four days later, on an unrelated follow-up read of the same patient, not because any part of the system had flagged it. When she pulled ScanDraft's internal confidence log, the mass finding had scored well below the model's usual range, a fact that never surfaced anywhere Bram could see it.
Four missing signals. Any one of them would have told Bram to look twice.
Here's the flip: Dr. Cuartas didn't just tighten the review process. She pulled every delegation back, insisting on personally reading every scan herself again before any resident touched it, the same as before ScanDraft existed. Two people were now doing the work of one, and Bram's training stalled exactly where it had been going well.
The model's miss rate barely moved. What changed is whether anyone could see which three needed a second look.
The real fix wasn't more review. Meridian rebuilt ScanDraft's report card: any finding below a set confidence threshold renders in hedged language, "findings are equivocal, recommend independent read," inside a visibly bordered section, and the report can't finalize until a named second reader, not just Bram, explicitly signs that specific section.
The five steps, if you want to remember itNot a general lesson about trusting AI less. FLIPS is what finds the exact moment delegation stopped being safe.
F
Find the person. Whose morning is this?
Dr. Esperanza Cuartas, a senior attending eight months into delegating first-pass reads to her resident, Bram.
Grounds the whole answer in one real supervision relationship, not a generic "the reviewer."
L
Locate the habit. What did she stop doing because it worked?
She stopped personally reading every scan herself, trusting Bram's review of ScanDraft's drafts after six clean months.
The habit is the delegation itself, working exactly as intended, right up until it wasn't safe to.
I
Identify the flip. What verb snaps?
Handed the task down, to took it all back. Not a policy change, a full reversal of who reads every scan.
The hardest step, and the one this whole answer turns on: a delegation flip, not a verification flip.
P
Pinpoint the old decision. Which choice only made sense before?
Rendering every finding in the same confident tone regardless of the model's actual certainty, reasonable when every draft got a careful, unhurried review anyway.
A small, specific interface choice, not a decision about AI in general.
S
Show the replay. Same bad day, new design.
The same ambiguous mass now renders hedged and bordered, forcing a named second reader's explicit sign-off, so Bram notices the difference without needing to be told to slow down.
Delegation stays intact. Nobody has to take the whole job back to catch the one case that needed it.
Only one case type sits in the corner that actually needs a different screen.
The recap, one line per letter: find the person is Dr. Cuartas and Bram's supervision relationship, locate the habit is her trust in his review replacing her own independent read, identify the flip is delegation reversing into full takeover, pinpoint the old decision is the uniform confident-sounding report design, and show the replay is the hedged, bordered, second-reader flag that lets delegation survive the exact case that used to end it.
And if you want to be sure it really works, try it somewhere elseSame five steps, a contract-review team instead of a reading room. A different flip family breaks the second story.
Vexley & Corbind, a mid-size law firm, uses ClauseKey, an AI tool that drafts plain-language summaries of contract clauses for paralegals to review before an attorney signs off. Odele Frantzen is a paralegal there. Mapped onto FLIPS, but landing on a different family than Dr. Cuartas's story: find the person is Odele, locate the habit is feeding ClauseKey full, messy real contracts the way she always had.
The flip here is a pre-editing flip, not a delegation flip. After a few uncomfortable misses on long, tangled indemnification clauses, ClauseKey's uniformly confident-sounding summaries gave Odele no way to tell which clause needed a closer read. Rather than take the whole review back the way Dr. Cuartas did, Odele started quietly trimming and simplifying the messiest clauses herself before ever uploading a contract, sanitizing exactly the hard, ambiguous language that was the whole reason the clause was risky in the first place.
The old decision behind it: ClauseKey's team decided not to expose which specific clauses it was least sure about, since "every clause gets the same thorough treatment" sounded like a selling point. That held up fine on short, clean contracts. It stopped holding up once contracts got long enough that difficulty genuinely varied clause to clause.
Same four missing signals, a contract clause standing in for a radiology finding this time.
Contracts pre-trimmed by paralegals before upload, per week, before and after ClauseKey's low-confidence flag
Once ClauseKey flagged its own weak spots, paralegals stopped needing to hide the hard clauses from it beforehand.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "hedge the language, border the section, require a named second sign-off, at low confidence," and stop.
Cost: there's no engineering time this quarter to build a full second-reviewer routing system. Say so honestly, and ship the hedged language and visible border alone first, since that costs far less than the routing logic and still stops a reviewer from skimming past.
The model gets better, for real: if ScanDraft's confidence calibration improves and low-confidence cases get rarer, that's still not a reason to remove the flag. It's a reason to trust it more, since a rarer flag is even easier to miss if it ever goes quiet entirely.
Where people run it wrong.
They add a small confidence percentage but leave the surrounding sentence reading just as confident as ever.
They flag the section but route it back to the exact same reviewer who already approved it once.
They treat a near miss as a reason to remove all delegation, instead of fixing the one interface gap that caused it.
How to use it live. When someone asks how the interface should change at low confidence, ask yourself: if I only changed the number and nothing else on the page, would anyone's behavior actually change? If the honest answer is no, the number was never the fix.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Delegation flip: handed the task down, then took it back. A senior person reclaims work a junior person was handling well, once quality feels uncertain.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dr. Esperanza Cuartas, senior radiologist at Meridian Regional Hospital, supervising resident Bram Vandersteen.
3 · THE HABIT
What did Dr. Cuartas stop doing because it worked?
Tap to flip
ANSWER
Personally reading every scan herself, trusting Bram's review of ScanDraft's drafts after six clean months of spot-checks.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Delegating first-pass reads to Bram, versus reading every single scan herself again. No middle setting, she went from one to the other entirely.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Rendering every ScanDraft finding in the same confident tone, regardless of the model's actual certainty, made sense back when every draft got a careful review anyway.
6 · THE NUMBER
Fill in the blank: with the flagged design in place, residents caught an ambiguous finding without prompting in ___ percent of cases, versus 22 percent unflagged.
Tap to flip
ANSWER
84 percent. The flag alone nearly quadrupled the catch rate on the exact cases that needed a slower read.
7 · THE REPLAY
Same ambiguous mass finding, new report design. What changes?
Tap to flip
ANSWER
It renders hedged, inside a bordered section, and requires a named second reader's sign-off, so Bram notices without Dr. Cuartas ever having to take the whole delegation back.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and which flip family shows up instead?
Tap to flip
ANSWER
Vexley & Corbind's ClauseKey contract-review tool. There, the flip is pre-editing, not delegation: paralegal Odele Frantzen started manually simplifying messy clauses before upload instead of taking the whole review back.
Check yourself Score: 0 / 0
True or false
1. True or false: Dr. Cuartas's fix to the near miss was to have residents check ScanDraft's confidence score more carefully every time.
True
False
Show hint
Look at the P and S steps.
Show answer
False. "Check more carefully" is a dial, not a fix. The real change was hedged language, a bordered section, and a required second reader's sign-off.
Multiple choice
2. Why does this answer treat "Dr. Cuartas took every delegation back" as the flip, rather than "Bram missed a finding"?
A. Bram's mistake isn't relevant to the story.
B. A flip is a human behavior with two settings and no middle. Delegating versus taking the job entirely back is the actual behavior that snapped.
C. Hospital policy requires attendings to take back delegation after any error.
D. The model's confidence score itself is the flip.
Show hint
Look at "the flip is the model's behavior" as a common mistake, per the I step's failure modes.
Show answer
B. The flip has to be a person's action with two settings, not the model's number moving or a single missed case.
Fill in the blank
3. Fill in the blank: at Vexley & Corbind, paralegals manually pre-trimmed 18 contracts a week before ClauseKey's flag shipped, falling to about ___ a week by week eight after it shipped.
Show hint
Look at the line chart in Section 4.
Show answer
3 a week. Once ClauseKey flagged its own weak spots, paralegals no longer needed to hide the hard clauses from it first.
Short answer, apply it yourself
4. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse, or a little less clear about its own uncertainty?
Show hint
Think of something you've stopped double-checking because it's usually right.
Show answer
Model answer: Most people can name at least one tool they've quietly stopped verifying, exactly the habit this method is built to find.
Short answer, why no middle setting
5. Why couldn't Bram have just "reviewed a little more carefully" instead of the interface needing to change at all?
Show hint
Look at the I step's failure modes in flips-method.md's spirit: a dial isn't a flip.
Show answer
Model answer: Nothing on the page told him which case needed more care. Reviewing everything more carefully would have erased the time savings ScanDraft was built to provide.
Short answer, where it wouldn't matter
6. Name a case type in this same reading room where this low-confidence redesign genuinely wouldn't change anything.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A clean, high-confidence fracture read. It stays fast and unhedged, since slowing it down would waste the time the tool exists to save.
Before you close the answer
Why this works
Tests whether you understand that a confidence score is inert until it changes the surrounding text, the visual state, and a required action, and whether you can name the specific human relationship, delegation, that a flat design puts at risk.
Follow-up traps
"Wouldn't a simple color-coded badge next to the score be enough?" Response: a badge next to otherwise confident prose still competes with the wording; a reader defaults to the sentence, so the sentence itself has to change too.
"Isn't pulling all delegation back the safest response to a near miss?" Response: safest in the short term, but it doubles the workload and stalls the resident's training. The actual fix is narrower: flag the specific gap, not remove the whole arrangement.
If pressed
The rebuilt ScanDraft also logs which named second reviewer actually opened and acted on each flagged section, not just whether the flag appeared, so a chief of radiology can tell the difference between a flag that changed a decision and one that got rubber-stamped.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.