CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #21

Design a confidence display for a medical information feature.

SPARK design the anchor before the model has its own bad night

Verity Health runs Ask Verity, a symptom-information feature a patient can open any time of day or night. Emeka Sorensen has worked warehouse loading docks for nine years, and has used Ask Verity a handful of times before, for his kids' fevers and his own strains.

The direct answer
Never show urgency as a single label or number alone. Show which specific symptoms are driving the read, what information is missing, and hard-code certain symptom combinations, like chest tightness with breathlessness and a family history of early heart disease, to always recommend care now, no matter what the overall score says.
Do this, in order
  1. Show which reported symptoms are actually driving the urgency read, not just the label.Why: a color or a word with nothing under it invites someone to stop thinking the moment they see it.
  2. Hard-code red-flag symptom combinations to always recommend care now.Why: some combinations are dangerous enough that they should never depend on an aggregate score being calibrated correctly that night.
  3. Tell the user exactly what information, if added, would change the read.Why: it turns a flat answer into an invitation to describe more, instead of a door that closes.
  4. Never show a specific diagnosis or a percentage risk of one named disease.Why: a precise-sounding number invites false confidence on a case that is genuinely ambiguous.
  5. Track whether users sought the recommended care within the advised window, split by true emergency outcome.Why: the cases that matter most rarely generate a complaint, since the person who needed care sooner has no way to know it.
  6. Leave the underlying symptom-matching model alone for common, low-stakes presentations.Why: this redesign is about what gets shown, not about the accuracy of the model behind it.

How to answer this, stage by stage

Nobody is grading whether your screen looks calming. They're grading whether it survives the one night the model's own confidence is wrong.

Stage 1
Scope it to one person, one moment
Say it like this
"I'll design this for Ask Verity, a symptom feature Verity Health runs, for the moment a patient is deciding, alone, at night, whether to wait until morning."
Why this works
Keeps "design a confidence display" from turning into a screen with no real decision behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how this decision happens today. Payoff, the habit I want it to build. Anchor, the one design decision. Risk, what breaks it. Keep out, what I'm not building yet."
Why this works
Shows a real design method instead of a list of screen ideas.
Stage 3
Ground the anchor in what happens today, without it
Say it like this
"Today, a worried patient calls a nurse triage line, waits ten or fifteen minutes on hold, describes the symptom, and gets a spoken read, thorough, but slow, and it backs up badly during a busy night."
Why this works
Proves the anchor is solving a real, specific gap, not a hypothetical one.
Stage 4
Give the anchor, the one decision
Say it like this
"Never show urgency as a bare label. Show the specific symptoms driving it, what's missing, and hard-code certain combinations to always say seek care now, regardless of the aggregate score."
Why this works
This is the direct answer, said as a concrete build decision instead of a design philosophy.
Stage 5
Prove the anchor survives its own risk
Say it like this
"The first time the model's aggregate confidence undersells a genuinely dangerous, specific combination, chest tightness with breathlessness and a family heart history, the hard floor catches it anyway. It never depends on the score being right that night."
Why this works
Answers the question the interviewer is actually asking underneath: what happens the first time you're wrong.
Stage 6
Close on the one line
Say it like this
"A confidence label is a summary, not an argument. Show the reasoning under it, and never let a good aggregate score override a known dangerous combination."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Ask Verity is a feature a national health information service runs. A patient describes their symptoms and gets a read on how urgent their situation might be.

Before Ask Verity, a patient called a nurse triage line, waited ten or fifteen minutes on hold during a busy night, and got a real person's careful judgment, thorough, but slow, and lines backed up badly during flu season.

Hand sketched flow diagram titled Today without Ask Verity. Five boxes: Notice a symptom, Call the triage line, Wait on hold highlighted, Describe it to a nurse, Get a plain instruction.
Ask Verity replaced the second and third boxes with something instant. It hadn't yet replaced the honesty of the fourth.

Now Ask Verity gives an instant urgency read, day or night, no wait at all.

Here's the turn: for most cases, the label alone was enough, so people stopped reading past the color, because they never really needed to. Then a case comes along where the label's aggregate confidence underplays one dangerous, specific symptom combination, and by the time anyone reads the actual reasoning behind it, if reasoning was ever shown at all, the moment to act sooner has already passed.

Building the urgency score for Emeka's case
100 50 0 70, high urgency line Symptoms alone, 40 Red flag combo, +35 Missing info, +10 Final score, 85 of 100
Symptoms alone landed under the line. The named combination is what pushed it into "seek care now," and the old design never showed any of these three pieces.

At its worst, someone with a genuinely serious symptom combination stays home overnight because a single word, "Medium," was allowed to carry the entire decision.

Hand sketched comparison titled The day the label was wrong. Left, a grey gauge icon labeled Old design, caption Medium urgency, no reasoning shown. Right, a red scale icon labeled New design, caption High urgency, red flag named.
Same night, same symptoms. Only what the screen was willing to say out loud changed.
The decision I would take back Ask Verity's early design showed only the urgency label and its color up front, with the actual reasoning available on a second screen behind a "why" link nobody had reason to click. That made sense when the label alone was right often enough that clicking through felt like extra work for no benefit. It stopped making sense the night the label's aggregate score undersold one dangerous, specific combination.

What I would leave alone: the underlying symptom-matching model doesn't need a rebuild for the common, low-stakes cases, a cold, a mild sprain. Nothing here questions its accuracy on those.

The lesson: a confidence label is a summary, not an argument. The moment a summary is asked to carry a decision on its own, someone will trust it past where it was ever meant to go.

Now here is the same thing as a story

The short version above is what you'd say defending this design to Verity Health's clinical safety board. Read this one for how quietly the trust built up.

Ask Verity lives on a phone, and at 11pm on a Tuesday, Emeka Sorensen's phone was the only thing between him and a decision about his own chest.

Emeka has worked loading docks for nine years and knows the difference between a pulled muscle and something worse, most of the time. He'd used Ask Verity a handful of times before, for a kid's fever, a sprained wrist, and it always gave a fast, clear answer that matched what a nurse would have said.

Knowledge spark: what does an "urgency band" actually mean? Most symptom-checking tools group their read into a small number of bands, like Low, Medium, and High urgency, instead of naming a specific illness. That's usually the right call, since a symptom checker often can't diagnose a specific condition reliably. But a band with nothing else attached asks the reader to trust a single word for a decision that might genuinely be life or death.

Over those months, Emeka stopped reading past the colored band at the top of the screen, since it had never once been wrong for him before. There wasn't a specific bad case that changed things, more a slow trust that built up quietly, three visits, then five, each one confirming the color was all he needed.

After an unusually long shift, Emeka felt tightness across his chest and some shortness of breath climbing the stairs to his apartment. He described it to Ask Verity: tightness, mild breathlessness, and mentioned, almost as an aside in his own words, that his father had had a heart attack at fifty-two. The app returned "Medium urgency, monitor for 24 hours," the same band he'd seen plenty of times before for a strained back.

Hand sketched timeline titled The night Emeka's chest tightened. Five milestones: Symptoms start 10 40 pm, App says Medium 24 hour 10 52 pm highlighted, Still tight 3 am, Calls a friend's number 3 05 am, Doctor confirms fine 4 10 am.
Four hours between a label he trusted and the moment he finally called someone.
The band did not lie. It just never said that a family history of early heart disease, paired with those two symptoms, is exactly the combination that should never wait 24 hours.

Emeka went to bed, planning to call his doctor in the morning if it hadn't passed. It hadn't passed by 3am, and worried, he finally called a number a friend had once mentioned, an after-hours line. It turned out to be strain and stress, and a close look by an actual doctor confirmed his heart was fine. But it took a friend's memory of a phone number, not Ask Verity's design, to get him checked sooner than his own morning plan would have.

Hand sketched labeled parts diagram titled What's in the redesigned urgency read. A document icon at the center labeled Urgency Read, with four callouts around it: urgency band, symptoms driving it, what's missing, red flag floor.
The old screen only ever showed the first of these four. Emeka never got to see the other three.

With the redesigned display, Ask Verity now shows "High urgency, seek care now" the moment the reported combination of chest tightness, shortness of breath, and early family heart history appears together, regardless of what the aggregate score alone would have said, and it shows exactly why: "your description includes chest tightness, breathlessness, and a family history of early heart disease. Together, these should be checked tonight, not in the morning." Run the same night forward: Emeka calls his doctor's after-hours line by 11:15pm instead of waiting until 3am.

The old design asked Emeka to trust a color. The new one shows him the actual sentence a nurse would have said out loud.

I built the label to be fast and simple, on purpose. It took watching how easily "Medium" can undersell a real risk to see that fast and simple isn't the same as honest.

SPARK, in one screenNot a lecture on making a screen feel reassuring. SPARK is what tells you which decision the whole design actually hangs on.

S
Situation. How this happens today, without the tool.
A worried patient calls a nurse triage line at night, waits ten or fifteen minutes, describes the symptom, and gets a spoken read.
Grounds the whole design in a real gap, not a hypothetical one.
P
Payoff. The habit this should build.
Reading the specific reasoning under an uncertain band, and describing more when asked, instead of glancing at a color and stopping there.
Names the actual behavior change the design is trying to produce.
A
Anchor. The one decision everything hangs on.
Show which symptoms drive the read, what's missing, and hard-code red-flag combinations to always say seek care now, regardless of the aggregate score.
This is the hardest step and the answer to the question: a concrete, arguable design decision.
R
Risk. What breaks the first time it's wrong.
The first time the model's aggregate confidence underrates a genuinely dangerous, specific combination. The hard floor is what survives that, since it never depends on the score being right.
Proves the anchor was designed against its own failure, not just described.
K
Keep out. What we won't build, day one.
No specific diagnosis, no percentage risk of a named disease. A precise-sounding number invites false confidence on a genuinely ambiguous case.
Shows judgment about what stays out, not just a wish list of what's in.
Hand sketched quadrant titled Where urgency bands earn a hard floor. Axes symptom specificity from vague to named combination, and risk if missed from low to high. Common cold symptoms and sprained wrist sit low on risk. Chest tightness alone sits in the middle. Tightness plus breathlessness plus family history sits high on both axes.
Only the top-right corner ever needs a hard floor rule. Everywhere else, the ordinary aggregate score is fine.

The recap, one line per letter: situation is a nurse triage line and a wait, payoff is teaching patients to read the reasoning instead of just the color, anchor is the redesigned display with its hard floor, risk is the night the aggregate score alone would have underrated a real combination, and keep out is holding back a specific diagnosis or a percentage on day one.

Hand sketched icon list titled What we deliberately did not build day one. Three items: a box icon labeled No specific diagnosis given, a gauge icon labeled No percent risk of one disease, a funnel icon labeled No treatment plan past the urgency band.
Each of these three is a real feature someone will ask for eventually. None of them belongs in the first version.

And if you want to be sure it really works, try it somewhere elseSame five letters, a home-insurance claim assistant instead of a symptom checker. A different anchor, a different hard floor.

ClaimClear is a home-insurance assistant that gives a homeowner filing a storm-damage claim a read on how likely it is to be approved. Lucia Bertrand filed a claim through it after a windstorm cracked her roof. Mapped onto SPARK: situation is a homeowner today, calling an adjuster and waiting days for a callback with no sense of where the claim stands; payoff is the habit to build, describing damage specifics fully instead of stopping the moment a likelihood number looks good.

The anchor here is structurally the same idea, aimed at a different failure: show which specific claim details are driving the approval-likelihood estimate, and hard-code certain causes, like pre-existing roof wear the policy explicitly excludes, to always require additional documentation, regardless of how high the overall approval-likelihood score reads. The risk ClaimClear's team designed against was a homeowner seeing "78% likely approved" and skipping the extra inspection photos that would have caught an exclusion clause early, only to have the claim denied weeks later once an adjuster finally looked closely.

Hand sketched labeled parts diagram titled What's in the redesigned urgency read, reused here for ClaimClear's approval read. Center document icon labeled Urgency Read, with urgency band, symptoms driving it, what's missing, and red flag floor around it, relabeled for claim approval likelihood.
Swap "symptoms" for "claim details" and "red flag floor" for "excluded cause floor." The same four parts still do the same job.
Claims later denied on appeal, with and without an early exclusion-cause flag
30% 15 0 22% Score alone, no flag 6% Early exclusion flag shown
Same claims, same policies. Naming the specific excluded cause early, instead of just a likelihood percentage, is what closed most of the gap.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show the reasoning behind the band, and hard-code known dangerous combinations to override the score," and stop.
Cost: there's no clinical panel available yet to define every red-flag combination. Say so honestly, and start with the small number of combinations already well established in emergency medicine, like this one, before expanding the list.
The model gets better, for real: if Ask Verity's underlying accuracy genuinely improves and false "Low urgency" reads become rare, that's still not a reason to remove the hard floor, a rare miss on a dangerous combination is exactly the one moment the floor exists to catch.

Where people run it wrong.
They treat a clean-looking single label as more trustworthy than a fuller explanation, when it's really just less checkable.
They let the aggregate score override a known dangerous combination, instead of the other way around.
They wait for a bad outcome to notice a design gap, instead of asking upfront what happens the one time the model is wrong.

How to use it live. When someone asks you to design a confidence display for anything that matters, ask yourself one question first: what happens the one night the model's aggregate score is confidently wrong? Design the anchor to survive that night, not just the ordinary ones.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "design a confidence display for X" question?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Ground the anchor in what happens today, then prove it survives being wrong.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Emeka Sorensen, a warehouse worker of nine years who had used Ask Verity a handful of times before, always with an answer that matched a nurse's judgment.
3 · THE SITUATION
How does this decision get made today, without Ask Verity?
Tap to flip
ANSWER
A patient calls a nurse triage line, waits ten to fifteen minutes on hold, describes the symptom, and gets a spoken read.
4 · THE ANCHOR
What's the one design decision this answer hangs on?
Tap to flip
ANSWER
Show which symptoms drive the read, what's missing, and hard-code red-flag combinations to always recommend care now, regardless of the aggregate score.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Showing only the urgency label and color up front, with the real reasoning hidden behind a "why" link nobody had reason to click.
6 · THE NUMBER
Fill in the blank: symptoms began at 10:40pm and Ask Verity said "Medium urgency, monitor for 24 hours" at 10:52pm. Emeka finally called someone at about ___.
Tap to flip
ANSWER
About 3:05am. Roughly four hours passed between the label and the first real check, all inside the window the label had told him was fine to wait.
7 · THE REPLAY
Same night, redesigned display. What changes?
Tap to flip
ANSWER
Ask Verity flags High urgency immediately, naming the exact combination, and Emeka calls the after-hours line by 11:15pm instead of 3am.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the anchor there?
Tap to flip
ANSWER
ClaimClear, a home-insurance claims assistant. The anchor shows which claim details drive the approval-likelihood estimate, with a hard floor that always demands more documentation for known excluded causes.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old design decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Showing only the label and color, with reasoning hidden behind a "why" link. It made sense while the label alone was right often enough that clicking through felt like unneeded extra work.
Multiple choice
2. Why does the redesigned display always flag chest tightness plus breathlessness plus a family history of early heart disease as High urgency, regardless of the aggregate score?
  • A. That combination is required by health regulation to always be flagged.
  • B. Some combinations are dangerous enough that they shouldn't depend on the aggregate score being calibrated correctly that night.
  • C. It makes the app appear more cautious to reviewers.
  • D. Family history always overrides every other symptom in the model.
Show hint
Look at the Risk step.
Show answer
B. The hard floor exists precisely so a bad night for the aggregate model can't quietly undersell a known dangerous combination.
True or false
3. True or false: the redesigned Ask Verity display gives users a specific percentage risk of a heart attack.
  • True
  • False
Show hint
Look at the Keep out step.
Show answer
False. It deliberately avoids a specific diagnosis or a percentage risk of one named disease, since that invites false precision on an ambiguous case.
Fill in the blank
4. Fill in the blank: Emeka's symptoms started at 10:40pm, and he finally called an after-hours line at about ___.
Show hint
Look at the timeline diagram.
Show answer
3:05am. The redesigned display would have gotten him there by about 11:15pm instead.
Short answer, apply it yourself
5. Think of a time an app or a person gave you a one-word or one-number read on something that mattered, with no reasoning attached. Would you have acted differently if you'd seen the reasoning?
Show hint
Ask whether you ever clicked a "why" or "learn more" link, or just trusted the headline.
Show answer
Model answer: Most people can recall at least one case where the reasoning, if they'd seen it, would have changed what they did next, the exact gap Emeka's story shows.
Short answer, where it wouldn't matter
6. Name a case in Ask Verity where this same red-flag-floor scrutiny genuinely doesn't need to apply.
Show hint
Look at the quadrant diagram's lower-left corner.
Show answer
Model answer: A mild cold with no other symptoms. There's no known dangerous combination hiding inside it, so the ordinary aggregate score is fine on its own.
Before you close the answer
Why this works
Tests whether you'll design the confidence display around what a person actually does next, or just make the label look calmer and call the job done.
Follow-up traps
"Isn't a hard-coded red-flag floor just an if-statement pretending to be a model?" Response: yes, deliberately. Some combinations are dangerous enough that they should never depend on a probabilistic score being calibrated correctly that specific night.

"What if a user describes their symptoms inaccurately and the floor never triggers?" Response: that's exactly what "what's missing" is for, the display asks for the specific missing detail instead of silently assuming the aggregate score already has everything it needs.
If pressed
Ask Verity's red-flag floor list is reviewed by an actual clinical panel on a fixed schedule, not engineering alone, since adding or missing one real symptom combination on that list is a clinical judgment, not a modeling one.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more