CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #4
Design the UI for a feature that is 70 percent confident in its answer.
SPARK give the 70 percent case its own card, with the evidence attached, not a badge to trust or ignore
Thistle Grove Primary Care is a family clinic. Amara Solanke is a nurse practitioner there, managing a panel of about two hundred patients. LabLens is the AI tool that reads each patient's lab history and flags concerning trends, like a slow decline in kidney function, before Amara opens the chart herself.
The direct answer
Give the 70 percent case its own middle tier, called "worth a look," not a copy of the confident flag and not silence. Show the actual two data points that triggered it, right on the card, so Amara is looking at real numbers, not trusting a badge. One tap snoozes it, one tap escalates it. The evidence is what stops her from either rubber-stamping it or learning to ignore it.
Do this, in order
Build a middle tier, "worth a look," separate from a confirmed flag.Why: collapsing a 95 percent case and a 70 percent case into the same red banner erases the exact distinction the question is asking about.
Show the raw trend, not a percentage, on the card itself.Why: two real lab values across two visits let Amara judge it herself; a number alone asks her to trust a badge she can't inspect.
Give one-tap snooze and one-tap escalate, nothing heavier.Why: a middle-confidence case needs a fast decision, not a form to fill out before she can move to the next chart.
Track how often "worth a look" gets opened at all, every week.Why: this is the earliest sign the tier is turning into wallpaper, before anyone misses a real trend.
Leave the fine-grained percentage, and patient-facing scores, out of day one entirely.Why: three tiers is a decision a clinician can act on fast; a slider or a raw number to patients solves a problem nobody asked for yet.
How to answer this, stage by stage
Nobody's grading whether you know that 70 percent is "medium confidence." They're grading whether your design still works the day it's wrong in either direction.
Stage 1
Scope it to one real feature and one real user
Say it like this
"I'll design this for LabLens, a lab-trend tool used by Amara Solanke, a nurse practitioner with a panel of two hundred patients, flagging a specific case: a slow decline in kidney function that the model is 70 percent sure is real."
Why this works
Stops "design the UI" from turning into an abstract slider mockup with no real user behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how she does it today. Payoff, the habit I want to build. Anchor, the one design decision. Risk, what breaks first if I'm wrong. Keep out, what I won't build yet."
Why this works
Tells the interviewer you're about to commit to one concrete decision, not brainstorm a feature list.
Stage 3
Reframe what the question is really asking
Say it like this
"This isn't really 'how do I show 70 percent.' It's 'how do I make sure the middle case doesn't get treated like the confident case or the no-case.' Those are the two ways this fails."
Why this works
Separates a real design answer from a list of confidence-display ideas.
Stage 4
Give the anchor, before any reasoning
Say it like this
"A middle tier called 'worth a look,' showing the two actual lab values that triggered it, with one tap to snooze and one tap to escalate. No percentage on the card at all."
Why this works
This is the direct answer, one concrete screen decision, before the story does any convincing.
Stage 5
Prove it survives the day it's wrong
Say it like this
"Before this, every trend, whether the model was 55 percent sure or 95 percent sure, showed up as the same red banner. A colleague of Amara's started skipping the pile entirely once it got long, and one banner she skipped turned out to be a real, worsening trend caught two months later than it should have been."
Why this works
Turns "the middle case needs its own design" into a specific, checkable near miss.
Stage 6
Say what you'd leave out on day one
Say it like this
"I wouldn't build a fine-grained confidence slider, and I wouldn't show any of this to patients yet. Three tiers, evidence on the card, is enough to act on fast. Anything finer just adds a number nobody asked for."
Why this works
Shows restraint instead of a feature list dressed up as thoroughness.
Stage 7
Close on the one line
Say it like this
"Give the middle case its own tier and its own evidence, so nobody has to trust a badge, they can look at the actual numbers and decide."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Let's learn
Here's what a 70 percent case actually needs, and what happens when it doesn't get it.
Before LabLens, Amara reviewed lab trends by pulling up each patient's history and comparing values by eye, catching a slow decline only when she had time to look closely, about fifteen minutes per chart on the ones she flagged for review. LabLens now scans every result overnight and surfaces likely trends in seconds, cutting her review time on flagged charts to about four minutes.
The second box is where a busy morning quietly drops a borderline case. That's exactly the case LabLens should be built to protect.
Here's the turn: the extra speed was never the risk. The risk showed up the day LabLens's launch design put every trend, clear or borderline, into the exact same red banner, with nothing in the screen to tell Amara's team which ones actually deserved a slower look.
Share of flagged trends opened before acting, by week, under the single-banner design
By week eight, two out of three flags went unopened. The 70 percent cases and the 95 percent cases were drowning each other out, identically dressed.
Four parts on one card. The raw trend chart is the part that does the real work.
At its worst, a middle-confidence case buried under an identical banner to every confirmed case doesn't just get skipped once. It teaches a clinician to stop trusting the whole pile, confirmed cases included.
The decision I would take back
LabLens's launch design merged every trend concern, whether the model was 55 percent sure or 95 percent sure, into one flag type and one red banner. That made sense at launch, when flag volume was low enough that every one got opened anyway. It stopped making sense once volume grew and nothing on the screen told Amara's team which flags actually needed a closer look.
What I would leave alone: LabLens's confirmed, high-confidence flags don't need a redesign. They already get opened and acted on fast, because there's little ambiguity left to sit with.
The lesson: the case you're least sure about is the one that most needs its own design, not a smaller font on the same banner as everything else.
Now here is the same thing as a story
The short version above is what you'd say pitching this redesign to Thistle Grove's clinical lead. Read this one for how the near miss actually happened.
Amara Solanke had worked as a nurse practitioner for six years, and could read a lab trend cold before her coffee finished brewing. A colleague two doors down had joined the same panel-review process a year into LabLens's rollout, when flag volume was already climbing.
By her third month, that colleague had started treating the flag pile the way most people treat a full inbox: skim the top few, let the rest wait for a quieter day. Every flag, mild or serious, looked exactly the same, one red banner reading "Trend detected: review recommended."
Knowledge spark: why would a clinician stop opening flags at all?
When every signal looks identical regardless of how strong the evidence behind it is, a person can't tell which ones are worth the time. Once that happens enough times without consequence, skipping starts to feel safe, right up until the one time it isn't.
One skipped banner belonged to a patient whose kidney function had been declining slowly across three visits, a pattern the model had flagged at 70 percent confidence, mixed enough that a careful look would have caught it clearly. It sat unopened for eight weeks until a routine follow-up visit caught the same trend, now further along than it needed to be.
Whichever way the design fails, the raw evidence sitting on the card is what gives someone a chance to catch it anyway.
Nothing about the trend was hidden. It was sitting in the chart the whole time. It just looked exactly like every case that didn't need a second look at all.
The redesign that followed gave the 70 percent case its own tier, called "worth a look," showing the two actual lab values that triggered it right on the card. Amara's team could glance at real numbers instead of trusting a badge, and open rates on the middle tier held steady for months after.
The redesign shipped one month after the drop was first noticed. Six months later, the open rate had held, not just recovered briefly.
SPARK, in one screenNot a mockup of every possible screen state. SPARK is what tells you which one decision actually matters.
S
Situation. How the job gets done today.
Amara compares lab values by eye across visits, catching trends only when she has time to look closely at each chart.
Grounds the design in a real workflow, not a blank screen.
P
Payoff. The habit worth building.
A clinician who checks the borderline cases at least as carefully as the confirmed ones, not a clinician who trusts a badge or tunes it out entirely.
Names what "success" actually looks like months from now.
A
Anchor. The one design decision.
A separate "worth a look" tier showing the raw trend on the card, with one-tap snooze and one-tap escalate.
This is the hardest step and the direct answer to the question.
R
Risk. What breaks first if it's wrong.
Either the middle tier gets rubber-stamped like a confirmed flag, or it gets ignored like noise. The raw evidence on the card is what catches both failure directions.
Proves the anchor survives being wrong, not just being right.
K
Keep out. What not to build yet.
No fine-grained percentage slider, no patient-facing confidence score, no automatic scheduling off a middle-tier flag.
Shows judgment about what the design doesn't need on day one.
Three outcomes, not a sliding scale. A clinician can act on three tiers fast; a percentage asks her to do the sorting herself.
Restraint here is what keeps the anchor simple enough for a clinician to act on between patients.
The recap, one line per letter: situation is Amara reading trends by eye today, payoff is a clinician who checks borderline cases as carefully as confirmed ones, anchor is a separate "worth a look" tier with raw evidence on the card, risk is the same evidence catching both over-trust and abandonment, and keep out is no slider and no patient-facing score, not yet.
And if you want to be sure it really works, try it somewhere elseSame five letters, a warehouse's return-fraud flag instead of a lab trend. A different day-it's-wrong breaks the second story.
Falkirk Distribution Center uses ReturnLens, an AI tool that flags customer returns likely to be fraudulent, before a processor opens the package. Marisol Vega is the returns-processing lead who reviews the flagged queue each shift. Mapped onto SPARK: situation is Marisol currently opening every returned package by hand to check it against the order; payoff is a team that spends real inspection time on the returns that deserve it, not an average inspection time that looks efficient on a dashboard.
The anchor carries over almost unchanged: a "worth a look" tier showing the actual evidence, photos of the returned item's condition next to the original order description, with one tap to clear and one tap to hold for manual review. The risk works differently here. At Thistle Grove, the worst case was a missed medical trend. At Falkirk, the worst case is a genuine customer, not a fraud attempt at all, getting held and refunded late over a middle-confidence flag that a five-second look at the photos would have cleared.
Same four parts. "The raw trend chart" becomes "the item photo and order history" on this product.
Genuine customers wrongly held for review, single-tier design vs three-tier redesign
Showing the actual evidence on the card, instead of a single flag banner, cut wrongly-held genuine customers by six times.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a separate 'worth a look' tier, raw evidence on the card, one tap to snooze or escalate," and stop.
Cost: there's no time this sprint to build a full evidence view. Say so honestly, and ship the three-tier labeling alone first, since the tier split matters more than the evidence detail on day one.
The model gets better, for real: if LabLens's accuracy genuinely improves, the middle tier should shrink in volume over time, not disappear. A better model still meets a genuinely ambiguous case sometimes, and that case still deserves its own design.
Where people run it wrong.
They design one screen for "confident" and call anything else an edge case, instead of designing the middle case on purpose.
They show a raw percentage and assume it explains itself, instead of showing the evidence behind it.
They skip testing whether the middle tier still gets opened months later, and find out only after a real miss.
How to use it live. When someone asks you to design for a specific confidence number, ask yourself: what does this case look like if the person using it stops trusting it, or starts trusting it too much? Design so the card survives both.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "design the UI for X" case question like this one?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It runs the method forward, ending on one concrete design decision that has to survive being wrong.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Amara Solanke, a nurse practitioner at Thistle Grove Primary Care, managing a panel of about two hundred patients.
3 · THE PAYOFF
What habit is this design actually trying to build?
Tap to flip
ANSWER
A clinician who checks borderline cases as carefully as confirmed ones, instead of trusting a badge blindly or tuning the middle tier out entirely.
4 · THE ANCHOR
What's the one concrete design decision this answer commits to?
Tap to flip
ANSWER
A separate "worth a look" tier showing the raw trend data on the card itself, with one-tap snooze and one-tap escalate, no percentage shown.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
LabLens's launch design merging every trend concern, 55 percent confident or 95 percent confident, into one identical red banner.
6 · THE NUMBER
Fill in the blank: under the single-banner design, the share of flags opened before acting fell from 90 percent to ___ percent by week eight.
Tap to flip
ANSWER
33 percent. Two out of three flags went unopened once every flag looked identical.
7 · THE RISK STEP
What's the one design element that catches the anchor being wrong in either direction?
Tap to flip
ANSWER
The raw evidence shown on the card. Whether someone over-trusts or ignores the tier label, the real numbers are still sitting right there to catch a second look.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and how does the risk step change?
Tap to flip
ANSWER
Falkirk Distribution Center's ReturnLens fraud flag. There, the worst-case risk flips: a genuine customer getting wrongly held, not a missed medical trend, but the same evidence-on-the-card anchor still catches it.
Check yourself Score: 0 / 0
Short answer, name the anchor
1. State the one concrete design decision this answer commits to for the 70 percent case.
Show hint
Look at the direct answer and the Anchor step.
Show answer
Model answer: A separate "worth a look" tier, showing the raw trend data on the card, with one-tap snooze and one-tap escalate, no bare percentage shown.
Multiple choice
2. Why did Amara's colleague start skipping flags entirely, according to the story?
A. She didn't believe the model worked at all.
B. Every flag looked identical regardless of how strong the evidence behind it was, so there was no way to tell which ones mattered.
C. The clinic told her to stop opening flags to save time.
D. LabLens stopped working for two weeks.
Show hint
Look at the knowledge spark in Section 2.
Show answer
B. With no visible distinction between a mild and a serious case, skipping the pile started to feel safe.
True or false
3. True or false: this answer recommends showing patients their own AI confidence percentage on day one.
True
False
Show hint
Look at the Keep Out step.
Show answer
False. That's explicitly one of the things left for later, not built on day one.
Fill in the blank
4. Fill in the blank: at Falkirk, the three-tier redesign cut wrongly-held genuine customers from 12 percent down to about ___ percent.
Show hint
Look at the bar chart in Section 4.
Show answer
2 percent. Roughly a six-times drop, from showing the actual photo evidence instead of a single flag banner.
Short answer, apply it yourself
5. Think of a notification or alert you get regularly. Does it ever tell you how sure it is, or does everything look equally urgent?
Show hint
Compare a spam filter's "maybe junk" folder to a single inbox with no filtering at all.
Show answer
Model answer: Most alert systems use one urgency level for everything, which is exactly the "merged tiers" mistake this answer says to design against.
Short answer, where it wouldn't matter
6. Name a part of LabLens where this three-tier redesign genuinely isn't needed.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The confirmed, high-confidence flags. They already get opened and acted on quickly, since there's little ambiguity left for a clinician to sit with.
Before you close the answer
Why this works
Tests whether you can turn a bare confidence number into one concrete, inspectable screen decision, and whether that decision still holds up the day a clinician either over-trusts or ignores it.
Follow-up traps
"What if 'worth a look' becomes its own kind of noise over time?" Response: track the open rate on that tier specifically every week, the same signal that caught the original problem, and tighten the trigger threshold if it starts drifting down again.
"Why not just show the number and trust clinicians to interpret it?" Response: a bare number asks every reader to build their own private calibration; showing the actual two data points lets each person judge the real evidence instead of trusting an abstraction.
If pressed
The final "worth a look" card also logs which specific values triggered the tier, so if a clinician escalates it, the record shows exactly what the model saw, not just a timestamp and a badge, which mattered the first time a flagged case ended up in a formal chart review.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.