ConceptIntermediateDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #1
When should you show a confidence score to a user, and when should you hide it?
PICK show the raw number to whoever reads it forty times a day, never to whoever reads it once
Cadence Health runs a telehealth triage line. Renata Cabral is the triage line lead, an RN who has worked the overnight queue for eight years. UrgencySense is the AI tool that scores every incoming call for how urgent it looks, before a nurse or a caller ever reads a word of it.
The direct answer
Show the raw score to the trained person who sees hundreds of them and can feel what a 62 means. Hide it from the person who will see it exactly once, and give them a plain sentence and one action instead. A number only means something to someone who has enough of them to compare it against.
Do this, in order
Show the raw score only to the repeat operator, never the one-time user.Why: a nurse builds a felt sense of the scale across hundreds of calls; a caller gets one shot at reading it correctly, cold.
Replace the number with one plain sentence and one action for the one-time user.Why: "call back if new symptoms start" tells a caller what to do; "71%" tells them nothing they can act on.
Track how often the trained operator clears a case with no added note.Why: even the audience that should trust the number can start rubber-stamping it, and that's the earliest sign it's doing harm.
Set a kill line: if unexplained clears cross a real threshold, force a note before the case can close.Why: a number nobody has to explain becomes a number nobody actually reads.
Never let "we said we'd be transparent" freeze the decision in place.Why: a launch-day promise about honesty is not the same thing as a promise to print a bare percentage forever.
How to answer this, stage by stage
Nobody's grading whether you know what a confidence score is. They're grading whether you can tell the difference between two audiences looking at the exact same number.
Stage 1
Scope it to one real product and two audiences
Say it like this
"I'll answer this for Cadence Health's UrgencySense, an AI urgency score used by both the triage nurse who reads it live and the caller who reads it once, in an after-visit text."
Why this works
Stops "show or hide" from becoming a single yes-or-no answer when the real answer depends on who's asking.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my pick before any reasoning. Impact, who feels each choice and in what units. Cost asymmetry, which cost is cheap and which one is hidden. Kill criteria, what would change my mind."
Why this works
Shows the interviewer you're about to commit to a side, not list pros and cons forever.
Stage 3
State the position, before any reasoning
Say it like this
"Show the raw score to the nurse. Hide it from the caller, and replace it with a plain sentence and one action."
Why this works
This is the direct answer, said plainly, before the story does any convincing.
Stage 4
Name who feels each choice
Say it like this
"Hide the number from the nurse and her whole shift slows down, every call gets the same scrutiny. Show the raw number to a caller and one of them reads '71% non-emergency' as permission to wait out real chest pain."
Why this works
Both sides of the tradeoff get a real person and a real cost, not a hand-wave.
Stage 5
Prove the asymmetry with the near miss
Say it like this
"A caller with chest pain and a fainting spell got a text reading '71% likely non-emergency, monitor at home.' She waited nine hours overnight before her family drove her in. The cost of hiding the number from a nurse is a slower shift. The cost of showing it to her was almost her life."
Why this works
Turns "one cost is worse than the other" into a specific, checkable gap instead of a vague warning.
Stage 6
Say the kill line
Say it like this
"If nurses start clearing calls above 80 with no note attached, for two months running, that tells me the number is now doing harm even for the trained audience, and I'd force a note before the case can close."
Why this works
Shows the position isn't fixed forever, there's a real signal that would change it.
Stage 7
Close on the one line
Say it like this
"Show the number to whoever sees enough of them to know what it means. Hide it from whoever only gets one look, and give them a sentence instead."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Let's learn
Here is what happens when the same score gets shown to two very different people.
Before UrgencySense, Renata's team triaged every call by hand, working from a caller's own description over the phone, taking about six minutes per call to judge urgency. UrgencySense reads the same description and scores it in under two seconds, and cut the average time to a first triage decision to under ninety seconds.
The whole question turns on what happens between the second box and the fourth one, and who else gets to read that middle number.
Here's the turn: the extra speed was never the risk. The risk showed up the day Cadence Health decided the same raw score that helped Renata should also be printed, unexplained, on the text message a caller reads alone at home.
Cost of each choice, in a real bad week
One bar is minutes. The other is hours before real care. Those units are not the same size, and that's the whole cost asymmetry.
One box is an inconvenience. The other one is a person deciding whether to go to the hospital.
At its worst, an unexplained number shown to the wrong audience doesn't just confuse someone. It replaces a real judgment call with a false sense of official permission to wait.
The decision I would take back
Cadence Health's founding pitch to patients was full transparency: "you'll always see exactly how sure our tool is." That promise got kept literally, printing the raw percentage on every after-visit text. It made sense at launch, when it read as an honest, modern thing to say and nobody had data yet on how a caller alone at 11pm would actually read a bare number.
What I would leave alone: the weekly aggregate accuracy report Renata's team reviews together. Nobody's making a solo, high-stakes call off that number in the moment, so a raw percentage there does real good with none of the risk.
The lesson: a number isn't dangerous by itself. It's dangerous the moment it reaches someone who has no other numbers to measure it against.
Now here is the same thing as a story
The short version above is what you'd say defending this call in front of Cadence's clinical safety board. Read this one for how close the near miss actually came.
Renata Cabral had run the overnight triage queue at Cadence Health for eight years, and could tell a real emergency from a scared-but-fine caller before they'd finished their second sentence. The night in question started the way most did: steady, a little slow after midnight.
A caller, a woman in her fifties, described chest tightness and a brief fainting spell earlier that evening. UrgencySense scored the case 71, meaning "likely non-emergency," based mostly on how she'd phrased the chest tightness. The fainting spell, mentioned almost as an aside, didn't move the score much.
Knowledge spark: why would fainting barely move the score at all?
A model like this learns urgency mostly from the words people use most often in past emergency and non-emergency calls. Fainting is rarer in its training data than chest tightness, so it carries less weight in the score, even though a clinician would treat it as a serious flag on its own.
Because the score came back above the routing cutoff, the case never reached Renata's live queue at all. Instead, an automated text went straight to the caller: "71% likely non-emergency. Monitor at home for now. Call back if new symptoms start."
Same number, same screen, two completely different amounts of practice reading it.
She didn't wait because a doctor told her to. She waited because a percentage sounded official enough to trust more than the fainting spell she'd actually felt.
She waited. At 8:10 the next morning, unable to get out of bed, her family drove her to the emergency room, where she was found to be in the middle of a slow cardiac event. She survived, with more damage done than an earlier catch would have caused.
Renata's team pulled the case afterward and found something worse than one bad night: nurses reviewing the aggregate queue had quietly started clearing any call scoring above 80 with no note at all, for months, because the number had never once been wrong in a way anyone had caught.
Four different things on one screen. Before this near miss, Cadence had only ever designed the first one carefully.
PICK, in one screenNot a debate about whether numbers are good or bad. PICK is what tells you who gets to see this one.
P
Position. The pick, before any reasoning.
Show the raw score to the trained repeat operator. Hide it from the one-time user, and give them a plain sentence and one action instead.
This is the hardest step and the direct answer to the question, said in one breath.
I
Impact. Who feels each kind of error.
Hiding it from the nurse costs her thirty-five extra minutes a shift. Showing it to a caller cost nine hours of delay on a real cardiac event.
Names both sides in real units, not a vague "confuses users."
C
Cost asymmetry. Which cost is hidden.
The nurse's cost is visible and cheap, a slower shift you can see on a dashboard. The caller's cost is rare, hidden, and can be a life.
This is what makes it a real tradeoff instead of a coin flip.
K
Kill criteria. What would change the pick.
If nurses start clearing high scores with no note for two months running, the number is now doing harm even for the trained audience too.
Separates a confident pick from a stubborn one.
Any one of these three is worth checking before the next bad night forces the question.
Share of high-scoring calls nurses cleared with no added note, by month
The line crossed the kill line in month five. Nobody was watching for it, because the number had never been wrong in a way anyone had noticed yet.
The recap, one line per letter: position is show the nurse, hide it from the caller, impact is thirty-five minutes against nine hours, cost asymmetry is visible-and-cheap against hidden-and-rare-and-severe, and kill criteria is a real threshold on unexplained clears, not a vague promise to keep watching.
And if you want to be sure it really works, try it somewhere elseSame four letters, a credit union's wire-fraud alert instead of a triage line. A different hidden cost breaks the second story.
Coldwater Federal Credit Union uses SentryScore, an AI tool that scores outgoing wire transfers for how likely they are to be fraud, before a fraud analyst or the customer sending the money ever sees it. Priyanka Deshmukh is the fraud risk analyst who reviews flagged transfers each morning. Mapped onto PICK: position is show the raw fraud score to Priyanka, who reviews hundreds of transfers a month, and hide it from the customer at the counter, replacing it with a plain hold notice and a callback number; impact is Priyanka losing queue-sorting speed if the score is hidden from her, against a customer reading a low fraud score as false reassurance that a scam wire is safe to send.
The nurse and the fraud analyst land in the same corner. The caller and the wire-transfer sender land in the same one too, on a completely different product.
The old decision here isn't a transparency promise, it's a different reversal: Coldwater's onboarding flow originally showed customers a real-time "risk meter" during the transfer itself, added because it tested well as reassuring in a demo. It stopped making sense the day a customer sending money to a romance-scam contact saw a moderate risk score, read it as "probably fine," and kept clicking through.
The number this reversal turned on
With the risk meter shown to the customer, only 6 percent of flagged fraudulent transfers got stopped before completion. Once Coldwater hid the meter and replaced it with a plain hold message and a callback number, that rose to 41 percent, nearly seven times as many, from customers no longer reading a moderate score as reassurance.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show the raw score to whoever sees enough of them to read it right, hide it from whoever only gets one look," and stop.
Cost: there's no engineering time to build a second, simplified view for one-time users this quarter. Say so honestly, and start with the cheapest fix, removing the raw number from patient-facing text entirely, before building the nicer plain-language version.
The model gets better, for real: if UrgencySense's accuracy genuinely improves, that's still not a reason to show the raw score to a one-time caller. A better model deserves more trust from the nurse who can tell when it's wrong, not a bigger number thrown at someone with no way to judge it at all.
Where people run it wrong.
They treat "transparency" as always meaning "show the raw number," instead of asking who's actually reading it.
They design the number once and never check whether the trained audience has quietly started trusting it too much.
They wait for a bad night to notice the drift, instead of watching the clear-with-no-note rate every month.
How to use it live. When someone asks whether to show a score, ask back: how many of these will this exact person see before they have to act on one? One tells you to hide it. Hundreds tells you to show it.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits an "A or B" tradeoff question like showing or hiding a score?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It commits to a side, then finds the hidden cost that makes the pick real.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Renata Cabral, triage line lead at Cadence Health, who has run the overnight triage queue for eight years.
3 · THE POSITION
What's the actual pick this answer commits to?
Tap to flip
ANSWER
Show the raw score to the trained repeat operator. Hide it from the one-time user and give them a plain sentence and one action instead.
4 · THE COST ASYMMETRY
Which of the two costs is cheap and visible, and which is hidden and severe?
Tap to flip
ANSWER
Hiding the score from the nurse costs a visible thirty-five minutes a shift. Showing it to a caller is rare, hidden, and can cost hours on a real emergency.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Cadence Health's launch-day promise of full transparency, kept literally by printing the raw percentage on every patient-facing text with no plain-language version behind it.
6 · THE NUMBER
Fill in the blank: the caller with chest pain and a fainting spell was scored ___ percent likely non-emergency.
Tap to flip
ANSWER
71 percent. She waited nine hours before her family drove her to the emergency room.
7 · THE KILL CRITERIA
What real signal would change this answer's pick?
Tap to flip
ANSWER
If nurses clear calls above 80 with no note for two months running, crossing the 20 percent kill line, that's the number doing harm even for the trained audience.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what changes?
Tap to flip
ANSWER
Coldwater Federal Credit Union's SentryScore wire-fraud score. There, hiding the risk meter from the customer and replacing it with a plain hold message stopped seven times as many fraudulent transfers as showing a number they couldn't calibrate.
Check yourself Score: 0 / 0
Multiple choice
1. Why does this answer show the raw score to the nurse but hide it from the caller?
A. Nurses are legally required to see raw scores.
B. The nurse sees hundreds of scores and can calibrate what they mean; the caller sees one, cold, with no way to judge it.
C. Callers are not smart enough to read a percentage.
D. Raw scores cost more to display to patients than to staff.
Show hint
Look at the Position and Impact steps.
Show answer
B. Repetition is what makes a number readable. One look, cold, is not enough exposure to calibrate against.
True or false
2. True or false: the caller in the story waited to seek care because a nurse personally reviewed her case and told her to monitor at home.
True
False
Show hint
Look at the routing cutoff in Section 2.
Show answer
False. Her case scored above the routing cutoff and went straight to an automated text, never reaching Renata's live queue at all.
Fill in the blank
3. Fill in the blank: at Coldwater Federal Credit Union, hiding the fraud risk meter from the customer raised the fraudulent-transfer stop rate from 6 percent to about ___ percent.
Show hint
Look at the bar chart in Section 4.
Show answer
41 percent. Nearly seven times as many fraudulent transfers were stopped once the number was replaced with a plain hold message.
Short answer, apply it yourself
4. Think of an app that shows you a score or a percentage. Do you see it often enough to know what a "good" one looks like, or only once?
Show hint
Ask how many times you've seen that exact number before this one.
Show answer
Model answer: Most one-time scores (a credit-check result, a home valuation estimate) get read cold, with no personal baseline, which is exactly the audience this answer says to hide the raw number from.
Short answer, why the middle ground fails
5. Why not just show every user the raw score, plus a short explainer paragraph about how to read it?
Show hint
Look at the near miss and how little time a caller has to read an explainer at 11pm.
Show answer
Model answer: A one-time, high-stress reader will not stop to absorb an explainer. The plain sentence and single action have to do the whole job in the moment, or they don't help at all.
Short answer, where it wouldn't matter
6. Name a place in the Cadence Health product where showing the raw number genuinely causes no harm.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The weekly aggregate accuracy report the nursing team reviews together. Nobody is making a solo, urgent decision off that number in the moment.
Before you close the answer
Why this works
Tests whether you can commit to a real position on "show or hide" instead of saying "it depends," and whether you can name the specific hidden cost that makes one option genuinely worse than the other.
Follow-up traps
"Isn't hiding information from users paternalistic?" Response: the nurse isn't losing information, she still gets the raw score; the caller is getting a version she can actually act on alone, which is more respectful of what she's able to use in the moment, not less.
"What if the caller wants to see the raw number anyway?" Response: let a details view exist behind a tap for anyone who wants it, but the default, first-glance message stays plain-language, since a rare wrong reading during an emergency is not a fair trade for satisfying curiosity.
If pressed
The redesigned patient-facing text no longer prints a number at all. It reads "based on what you told us, this can likely wait until your regular doctor's hours, but call back right away if you feel worse, faint again, or notice new chest pain," with the actual score kept internal for Renata's team to audit against outcomes later.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.