ConceptAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #11
Explain when a binary confident-or-not display beats a graded one.
GUARD the reader without training reads the number as a fact, not a probability
Fenwick General Hospital's pharmacy uses Interax, an AI tool that scores possible drug interactions on a graded 0-to-100 risk scale. Priya Ilangovan is a staff pharmacist there. The same score, at launch, appeared both on her dashboard and on the after-visit summary handed straight to patients.
The direct answer
Binary beats graded whenever the person reading it can't interrogate it, or act on the nuance a number implies. Show pharmacists the full graded score, because they're trained to weigh it and have the power to escalate. Show patients a binary flag instead: call your pharmacist, or no action needed, with one plain reason. A patient reading "41 percent" has no way to know whether that means relax or call immediately, and no lever to find out except guessing.
Design it in this order
Ask who reads this number and what they can actually do about it.Why: the same score means something completely different to someone who can escalate versus someone who can only guess.
Give the trained reader the graded score.Why: a pharmacist can weigh 41 percent against a specific patient's history, which is exactly what the number is for.
Give the untrained reader a binary flag with one plain reason.Why: a patient can't verify a percentage, only decide whether to call, so give them a decision, not a statistic.
Build a real escalation path behind the binary flag, not just the flag itself.Why: a flag with nowhere to go is the same as no flag at all.
Track how patients act on the binary flag, monthly, not just whether pharmacists like the graded view.Why: the whole point is protecting the reader who has no way to check the design's work themselves.
How to answer this, stage by stage
Nobody's grading whether you know graded scores are more precise. They're grading whether you know precision only helps the reader who can actually use it.
Stage 1
Scope it to one real screen
Say it like this
"I'll answer this for Interax, a drug-interaction scoring tool at Fenwick General Hospital, comparing Priya Ilangovan's pharmacist dashboard to the after-visit summary a patient takes home."
Why this works
Keeps "binary versus graded" from becoming an abstract UI opinion with no real stakes behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's affected. Unequal, where the harm actually lands. Ability to contest, who can push back and who can't. Reduce, the concrete design change. Detect, how I'd know it's happening in production."
Why this works
Shows a repeatable way to reason about power, not just a personal preference for simple UI.
Stage 3
Name both people, plainly
Say it like this
"There's Priya, who can read a 41 percent score, weigh it against a specific patient's kidney function, and call the prescriber. And there's the patient, who reads the same number at home with nobody to ask and no way to check it."
Why this works
This is GUARD's core move: naming the operator and the subject on the same page, not just the design.
Stage 4
Give the decision, before any reasoning
Say it like this
"Keep the graded score for the pharmacist. Replace it with a binary flag and a plain reason for the patient. Precision only helps the reader who's trained to act on it."
Why this works
This is the direct answer, stated as a decision instead of a general observation about numbers.
Stage 5
Prove it with the near miss
Say it like this
"A patient with declining kidney function saw '41 percent interaction risk' on her printed summary, read it as less than half, and didn't call. Her kidney function made that specific 41 percent far more dangerous than the number alone suggested."
Why this works
Turns "graded scores can mislead" into a specific, checkable failure instead of a general worry.
Stage 6
Say how you'd detect it going wrong
Say it like this
"Track how often patients call in on a genuinely risky interaction, before and after the binary flag ships. If the call rate on real risk doesn't rise, the flag isn't doing its job."
Why this works
Shows this isn't a one-time design fix, it's something you'd keep watching in production.
Stage 7
Close on the one line
Say it like this
"Binary beats graded whenever the reader can't push back on the number. Give the score to the person who can weigh it, and give a decision to the person who can only guess."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Here is what happens when the same number means two different things
Before Interax, before Priya's shift began, was a case where anything, from off-brand pain relievers to a supplement a patient forgot to mention, could combine badly. Fenwick's pharmacists cross-checked interactions by hand against a printed reference, about 6 minutes a prescription, across roughly 90 prescriptions a shift.
Same score, two very different amounts of power to act on it.
Interax scores a possible interaction in under a second, on a 0-to-100 scale, and displayed that same number both on Priya's dashboard and, at launch, on the after-visit summary every patient took home. The design team called it consistency. It was actually two different audiences reading one number meant for only one of them.
Patients who called their pharmacist about a genuinely risky interaction, by summary design
Same underlying risk, same patients. The plain flag got six times more of the people who actually needed to call.
Here's the turn: the graded score wasn't wrong. It was accurate, well-calibrated, and genuinely useful, to the one reader trained to weigh it. The problem was handing that same precision to a reader with no way to check it, no training to interpret it, and no lever to ask a follow-up question before deciding.
The after-visit summary sits in the one corner a graded score should never be handed to: novice reader, high stakes.
At its worst, a patient reads a real, accurate risk score, misjudges what it means for their own specific situation, and doesn't call anyone, because the interface gave them a number to interpret instead of a decision to make.
The decision I would take back
Interax showed the same graded score everywhere, to keep the information "consistent and transparent" across every audience. That's a reasonable transparency principle for a dashboard used by trained staff. It stops being reasonable the moment the audience includes someone with no training and no lever to act on the number beyond guessing.
What I would leave alone: Priya's own dashboard doesn't need this same simplification. She's trained to read a graded score, has the standing to escalate, and loses real information if the score gets flattened to a binary flag for her too.
The lesson: a number isn't neutral just because it's accurate. Who reads it, and what they can do next, decides whether that same number helps or quietly misleads.
Now here is the same thing as a story
The short version above is what you'd say defending this design to Fenwick's chief pharmacy officer. Read this one for how close the real incident came.
Before Interax, the best part of Priya's shift was never having to explain a printout to a worried patient at the counter. After Interax, that became the thing she explained most, patients holding a summary with a number on it, asking her what it actually meant.
The first months were fine. Pharmacists loved the graded score, using it to triage a shift's worth of prescriptions by how much attention each one needed. Nobody had thought much yet about the same number sitting on the sheet a patient carried out the door.
Knowledge spark: why would a "41 percent" interaction score mean something worse for one patient than another?
A risk score is built from general patterns across many patients. It doesn't automatically know that this specific patient has declining kidney function, which changes how their body clears a drug and can turn a moderate interaction into a dangerous one. The number is honest about the general case. It says nothing about the exception sitting in front of it.
A woman in her sixties, managing declining kidney function, was prescribed a new blood thinner. Interax scored the interaction with her existing medication at 41 percent, printed plainly on her after-visit summary. She read it, reasoned that less than half wasn't so bad, and went home without calling.
The gap in the third box is the whole story. Nothing lived there for the patient to use.
The score was accurate. It had no way of knowing her kidneys made 41 percent mean something much closer to dangerous than the number let on.
She was hospitalized four days later with a bleeding complication, caught in time, but only because a family member noticed the symptoms and brought her in. When Priya's team reviewed the case, the score itself hadn't failed. The design handing that score to someone with no training to weigh it against her own condition had.
Four small facts replace one number nobody outside the pharmacy could actually weigh.
Fenwick didn't remove the graded score. They split the display: Priya's dashboard kept the full 0-to-100 scale with contributing factors listed. The patient summary now shows one of two things: "No action needed" or "Call your pharmacist before your next dose," with a single plain-language reason underneath either one.
GUARD, said plainly, no moralizingNot a lecture on health literacy. GUARD is what tells you exactly which reader the precision is actually for.
G
Groups. Who's affected by this design.
Priya, the pharmacist who reads and acts on the score, and the patient, who reads the same number with no training and no lever.
Names both people on the same page, not just the interface between them.
U
Unequal. Where the harm actually lands.
A patient with an unusual condition, like declining kidney function, gets hurt by a general-case number that doesn't know their specific exception.
The harm isn't evenly spread. It concentrates on exactly the patients a general score describes worst.
A
Ability to contest. Who can push back, and who can't.
Priya can escalate, call the prescriber, or order more monitoring. The patient can only guess what 41 percent means and decide whether to call, with no way to verify either choice.
The hardest step, and the one this whole answer turns on: precision only helps the reader who can interrogate it.
R
Reduce. The concrete design change.
Keep the graded score for pharmacists. Replace it with a binary flag and one plain reason on anything a patient takes home.
A specific product decision, not a policy memo about health literacy.
Either sign, alone, is worth switching a screen to a binary flag instead of arguing over exact thresholds.
D
Detect. How you'd know it's happening in production.
Track patient callback rate on genuinely risky interactions monthly, and watch for any case where a scored-but-unflagged situation later needed real intervention.
Turns "we fixed the display" into something you keep checking, not a one-time redesign.
The question is never "which display is better." It's always "who's reading it, and what can they actually do next."
The recap, one line per letter: groups is naming Priya and the patient as two separate readers with two different amounts of power, unequal is the harm concentrating on patients whose exceptions a general score can't see, ability to contest is the whole reason binary wins for patients, reduce is the actual split-display fix, and detect is watching real callback behavior, not just whether the redesign looks cleaner.
And if you want to be sure it really works, try it somewhere elseSame five letters, a county probation office instead of a hospital pharmacy. A different kind of unequal breaks the second story.
Calder County Probation uses Reliant, an AI tool that scores a defendant's recidivism risk on a graded percentage, shown to both the supervising officer and the defendant during case review. Deacon Marsh, a probation officer there, ran the same GUARD analysis on Reliant's shared display. Mapped onto GUARD: groups is the officer, who has real discretion to adjust supervision conditions, and the defendant, who receives a supervision plan built partly from the score; unequal is that a defendant with a borderline score, say 58 percent, gets the least clear treatment of anyone, since the number reads as almost arbitrary and gives them nothing concrete to respond to.
The ability-to-contest step is where this story diverges hardest from Priya's. A defendant can formally appeal a supervision decision. They cannot appeal the number itself, since there's no defined process for challenging how a percentage was calculated, only for challenging what was done because of it. The reduce step: Reliant now shows officers the full graded score with contributing factors, and shows defendants a binary "standard supervision" or "enhanced supervision, here's why" with the specific named factors, not a percentage.
Swap "call your pharmacist" for "here's who to talk to about your supervision plan." The shape of the fix travels.
Defendant appeals citing confusion over their risk score, monthly, before and after the binary supervision display
Most of those appeals were never really about the supervision decision. They were about a number nobody had given the defendant a way to understand.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "graded for the trained reader who can act on it, binary for the reader who can only guess," and stop.
Cost: there's no engineering time this sprint to build a second display. Say so honestly, and hide the graded score from patient-facing views entirely first, since removing a misleading precision costs less than building a new one.
The model gets better, for real: if Interax's scoring genuinely gets more accurate, that's still not a reason to show patients the raw number. It's a reason to trust the binary flag's threshold even more, since a better model makes the flag itself more reliable.
Where people run it wrong.
They treat "showing the same number everywhere" as fairness, when it actually hands precision to a reader who can't use it.
They add a binary flag but leave the graded score visible right next to it, so patients read the number anyway.
They build the flag but skip the escalation path behind it, leaving someone with nowhere real to go when it fires.
How to use it live. When someone asks whether binary or graded is better, ask back: can the person reading this number push back on it, ask a follow-up, or act on the nuance? If not, binary wins, no matter how accurate the graded score is.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a risk, safety, or fairness question like this one?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It names the operator and the subject, and asks who can actually push back.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Ilangovan, staff pharmacist at Fenwick General Hospital, using Interax's graded drug-interaction score.
3 · THE TWO READERS
Name the operator and the subject in this story.
Tap to flip
ANSWER
Priya, who can weigh the score and escalate it, is the operator. The patient reading the same score at home, with no training and no lever, is the subject.
4 · THE HARD STEP
Which GUARD step is hardest here, and why?
Tap to flip
ANSWER
Ability to contest. A patient can't verify or question a percentage, only guess whether to call, which is exactly why binary beats graded for that reader.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Showing the same graded score to both pharmacists and patients, called consistency at the time, without asking whether both readers could actually act on it.
6 · THE NUMBER
Fill in the blank: with the graded score shown, only ___ percent of patients with a genuinely risky interaction called their pharmacist.
Tap to flip
ANSWER
11 percent, versus 68 percent once the binary flag replaced the raw score on the patient summary.
7 · THE FIX, VERIFIED
Same 41 percent interaction, split display instead of one shared score. What changes?
Tap to flip
ANSWER
Priya still sees the full graded score to weigh against the patient's kidney function. The patient sees "Call your pharmacist before your next dose" instead of a number she has to interpret alone.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's different about who can contest the number?
Tap to flip
ANSWER
Calder County Probation's Reliant recidivism tool. There, a defendant can appeal a supervision decision, but there's no process to appeal the risk number itself, which is what the reduce step had to fix.
Check yourself Score: 0 / 0
Multiple choice
1. Why does this answer recommend a binary flag for patients instead of the graded score pharmacists see?
A. Patients find percentages confusing to read on a printed page.
B. A patient has no training to weigh the number and no lever to check it, unlike a pharmacist who can escalate.
C. Binary displays are required by hospital regulation.
D. Graded scores are always less accurate than binary flags.
Show hint
Look at the Ability to contest step.
Show answer
B. Precision only helps a reader who can interrogate and act on it. A patient at home has neither.
True or false
2. True or false: this answer recommends removing the graded score entirely from Interax, for every user.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Priya's own dashboard keeps the full graded score. Only the patient-facing summary switches to binary.
Fill in the blank
3. Fill in the blank: at Calder County Probation, defendant appeals citing confusion over their risk score fell from 22 a month to ___ a month after the binary supervision display shipped.
Show hint
Look at the line chart in Section 4.
Show answer
5 a month. Most of those appeals were really about an unclear number, not the supervision decision itself.
Short answer, apply it yourself
4. Think of a score or percentage you've been shown by an app or a form, with no expert nearby to explain it. Did you know what to actually do with it?
Show hint
Ask whether you had any way to push back on or verify the number.
Show answer
Model answer: Most people can recall a number they had to guess the meaning of, which is exactly the gap a binary flag closes.
Short answer, why no middle setting
5. Why wouldn't just adding a short explanation next to the graded percentage have been enough for the patient summary?
Show hint
Look at the follow-up trap about adding context to the number.
Show answer
Model answer: A number next to an explanation still asks the reader to weigh both against their own case. A decision, call or don't, doesn't ask them to weigh anything.
Short answer, where it wouldn't matter
6. Name a part of Interax's design where this same binary-versus-graded caution genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Priya's own dashboard. She's trained to read a graded score and has the standing to act on it, so simplifying it there would remove real information for no benefit.
Before you close the answer
Why this works
Tests whether you understand that a display choice isn't just about clarity or aesthetics, it's about who holds the power to act on a number, and whether you can name that power imbalance plainly instead of hiding behind a general debate about simplicity.
Follow-up traps
"Couldn't you just add a short explanation next to the graded percentage instead?" Response: an explanation next to a number still asks the reader to weigh both against their own situation; a binary flag asks nothing, it just tells them what to do.
"Doesn't hiding the graded score from patients feel paternalistic?" Response: the score isn't hidden, it's replaced with something more useful to that reader, a decision instead of a statistic they have no training to interpret.
If pressed
The binary flag's threshold isn't fixed at one number company-wide. It adjusts per patient based on factors like kidney or liver function already in the chart, so the same interaction can trigger the flag for one patient and not another, even though a pharmacist would see the exact same graded score for both.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.