CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #6
How would you design an interface that encourages verification without being annoying?
LEAD the reflex click is the number that warns you, weeks before a real miss does
Fentridge Pharmacy is a busy retail pharmacy. Naomi Petrosyan is the pharmacist on duty most afternoons, fourteen years behind the counter. InteractGuard is the AI tool built into the dispensing terminal that flags possible drug interactions before a prescription gets filled.
The direct answer
Match the friction to the actual severity, instead of giving every alert the same click-to-dismiss box. Minor interactions get a quiet inline badge nobody has to act on. Serious ones lock the screen until a real reason gets chosen, not just a click. And track how fast alerts get dismissed, because that speed is the number that tells you verification has become a reflex, weeks before a real miss does.
Do this, in order
Tier the friction by real severity, not by whether an interaction was flagged at all.Why: treating a minor reminder and a genuine bleeding risk identically is what teaches people to click through both the same way.
Lock high-severity alerts until a real reason is chosen, not just a button pressed.Why: a click can happen without reading; picking a specific reason from a short list can't.
Log how fast every alert gets dismissed, every time, from day one.Why: without this, there's no way to see the reflex forming before it costs someone something.
Set a real threshold on the fast-dismiss rate that triggers a design review.Why: a metric nobody acts on is decoration; this one needs a number that actually changes something.
Never "fix" alert fatigue by quietly raising the threshold for what counts as an alert.Why: fewer alerts looks like less annoyance and is actually fewer real interactions caught.
How to answer this, stage by stage
Nobody's grading whether you can list ten notification patterns. They're grading whether your design has a way to know it's failing before a real miss proves it.
Stage 1
Scope it to one real terminal and one real pharmacist
Say it like this
"I'll design this for Fentridge Pharmacy's InteractGuard, and Naomi Petrosyan, who reviews interaction alerts on nearly every prescription she fills."
Why this works
Stops "an interface that verifies without annoying" from becoming an abstract notification-design essay.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the real outcome. Early signal, the number that moves first. Abuse, how it gets gamed. Decision, what I'd do at each threshold."
Why this works
Signals you're answering the design question with a way to measure whether it's actually working.
Stage 3
Name the real outcome
Say it like this
"The outcome isn't 'fewer complaints about pop-ups.' It's genuinely harmful interactions caught before a prescription goes out the door."
Why this works
Grounds "not annoying" in the actual thing the interface exists to protect, not user comfort alone.
Stage 4
Give the design decision, before any reasoning
Say it like this
"Tier the friction. A quiet badge for minor interactions, one tap and one chosen reason for moderate ones, a locked screen requiring a real reason for anything genuinely severe."
Why this works
This is the direct answer, a concrete screen decision, before the story does any convincing.
Stage 5
Prove it with the near miss
Say it like this
"A blood thinner and a new antibiotic, a genuinely serious combination, got the same one-click box as thirty trivial reminders that shift. It got dismissed in under a second. A tech caught the actual risk by chance, twenty minutes later, before pickup."
Why this works
Turns "identical friction is dangerous" into a specific, checkable near miss instead of a vague warning about pop-up fatigue.
Stage 6
Say the early signal and the threshold
Say it like this
"The number that would have warned us weeks earlier is the share of alerts dismissed in under two seconds. If that crosses forty percent for two weeks straight, that's not efficiency, that's a design review."
Why this works
Shows the design isn't just a one-time decision, it has a real number watching it after launch.
Stage 7
Close on the one line
Say it like this
"Match the friction to the real risk, and watch how fast people click through, because that speed tells you the truth before a real miss does."
Why this works
Restates the direct answer in one breath, ready for a follow-up.
Let's learn
Here is what happens when every safety alert asks for the exact same amount of attention.
Before InteractGuard, Naomi's team checked interactions by hand, cross-referencing a printed reference guide, taking about ninety seconds per prescription. InteractGuard now checks every prescription in real time, cutting that to under five seconds when nothing's flagged.
The third box is where this whole answer lives. Before the redesign, every prescription landed in the same identical box there.
Here's the turn: the extra speed was never the risk. The risk showed up the day every flagged interaction, trivial or genuinely dangerous, produced the exact same one-click box, and nothing about the screen told anyone which ones actually mattered.
Share of interaction alerts dismissed within two seconds, by month
By month twelve, more than half of all alerts were dismissed too fast to have been read. Nobody was tracking this number until the near miss forced the question.
Four parts. Before the redesign, every severity got the same second part, and nothing at all like the fourth.
At its worst, an identical alert box for every severity doesn't just get annoying. It trains a careful person to stop reading any of them, right up until the one time that costs a patient something real.
The decision I would take back
InteractGuard's original design logged whether an alert was dismissed, but never how fast. That made sense at launch, when getting the alert to fire correctly was the whole project. It stopped making sense once nobody had any way to see the fast-dismiss rate climbing, because the one number that would have flagged the problem simply wasn't being kept.
What I would leave alone: the plain informational notes InteractGuard shows for routine dosage reminders, like "take with food," don't need any friction at all. Adding a click requirement there would just be more noise with nothing gained.
The lesson: an alert that asks for the same attention every time isn't really asking for attention at all, it's asking to be dismissed.
Now here is the same thing as a story
The short version above is what you'd say pitching this redesign to Fentridge's pharmacy director. Read this one for how close the near miss actually came.
Naomi Petrosyan had worked the counter at Fentridge for fourteen years, and could spot a risky drug combination from memory before InteractGuard ever existed. A new pharmacy tech, a few weeks into the job, asked her one afternoon why everyone clicked through the safety pop-ups so quickly. Naomi didn't have a good answer, and the question stuck with her.
That same week, a prescription came through for a patient already on a blood thinner, paired with a new antibiotic known to raise bleeding risk when combined. InteractGuard flagged it, the exact same one-click box it used for thirty routine reminders that shift, "take with food," "avoid alcohol," "may cause drowsiness."
Knowledge spark: why would a serious interaction get the same box as a minor one?
Many interaction-checking systems flag anything above a low statistical threshold, without separating "worth a gentle note" from "worth stopping the workflow." If the interface doesn't do that separating visually, every flag ends up looking equally urgent, and equally skippable.
Naomi dismissed it in under a second, the same reflex she'd built over months of identical boxes. A tech, double-checking a different detail on the same order twenty minutes later, happened to notice the interaction listed in the chart and flagged it before the prescription left the pharmacy.
Seventeen minutes between the reflex click and the accidental catch. Nothing in the design made that catch anything but luck.
Naomi didn't miss it because she stopped caring. She missed it because the screen never once told her this box was different from the thirty before it.
Fentridge's redesign tiered the alerts by real severity, and added a mandatory read-time lock with a chosen reason for anything genuinely serious. A comparable case the following month triggered the locked screen, and Naomi caught it herself, on the first pass, no lucky double-check required.
Three tiers, not one box for everything. The friction now matches what's actually at stake.
LEAD, in one screenNot a UX checklist for reducing pop-up fatigue. LEAD is what tells you the design is failing before a real miss proves it.
L
Link. The real outcome at stake.
Genuinely harmful interactions caught before a prescription leaves the pharmacy, not a lower count of pop-up complaints.
Keeps "not annoying" from becoming the whole goal on its own.
E
Early signal. The number that moves first.
Share of alerts dismissed within two seconds, climbing from 8 percent to 52 percent across a year, well before any real miss occurred.
This is the hardest step and the actual answer to the question underneath the design.
A
Abuse. How this gets gamed.
A team under pressure to reduce complaints could quietly raise the alert threshold, showing fewer pop-ups while catching fewer real interactions too.
Explains why "fewer alerts" alone is a dangerous way to define success.
D
Decision. What you'd do at each threshold.
Cross a 40 percent fast-dismiss rate for two weeks running, and it triggers a mandatory tiering review, not just a note to "be more careful."
Turns the metric into a real action, not a dashboard nobody checks.
The fast-dismiss rate was ringing for months. Nobody had built a way to hear it.
Any one of these three quietly makes the numbers look better while making the pharmacy less safe.
The recap, one line per letter: link is real interactions caught, not fewer complaints, early signal is the fast-dismiss rate climbing for months before the near miss, abuse is a team quietly raising the alert threshold to look less annoying, and decision is a real, pre-agreed trigger on the fast-dismiss rate, not a vague promise to watch it.
And if you want to be sure it really works, try it somewhere elseSame four letters, a legal translation firm's clause-review flag instead of a pharmacy terminal. A different abuse breaks the second story.
Pemberton Language Services uses ClauseFlag, an AI tool that flags contract clauses where a translation carries real ambiguity, nudging a human reviewer to verify before a translated contract ships to a client. Grant Sutherland is the reviewer who checks flagged clauses each day. Mapped onto LEAD: link is mistranslations that actually cause a contract dispute later, not a lower count of flagged clauses; early signal is the share of flagged clauses approved in under one second, without the reviewer opening the original text to compare.
The abuse here works differently than at Fentridge. Instead of a reviewer clicking through fast, the team under deadline pressure quietly started suppressing flags on clauses matching common boilerplate language, on the theory that boilerplate is rarely ambiguous. That "improved" the flag count and the fast-approval rate at the same time, while hiding the rare case where boilerplate-looking language actually carried a real, contract-specific ambiguity.
Same four parts. "Severity" here means how much a mistranslation could actually cost in a dispute, not a drug interaction's harm.
Suppressing boilerplate flags looked like progress on paper. It was actually nine real misses a quarter, hiding behind a quieter dashboard.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "tier the friction by real severity, and track how fast alerts get dismissed as the early warning," and stop.
Cost: there's no engineering time this sprint to build a full read-time lock. Say so honestly, and ship the dismiss-speed logging first, since knowing the problem exists costs far less than fixing it, and it's the prerequisite for fixing it at all.
The model gets better, for real: if InteractGuard's accuracy genuinely improves and false alarms drop, that's still not a reason to loosen the read-time lock on severe cases, a better model still occasionally flags something serious, and that case still deserves the friction it's earned.
Where people run it wrong.
They treat every flagged item as equally worth interrupting someone for, instead of tiering by what's actually at stake.
They measure success by fewer complaints about annoying alerts, instead of by real incidents actually caught.
They never log how fast something gets dismissed, so the reflex has nowhere to show up until it's already cost someone something.
How to use it live. When someone asks for verification without annoyance, ask yourself: what's the number that would show people clicking through on reflex, weeks before that reflex costs anyone anything? Build that number in from day one.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "design without annoying" case that also needs a way to measure success?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It pairs a concrete design decision with the leading number that proves the design is still working.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Naomi Petrosyan, a pharmacist at Fentridge Pharmacy, fourteen years behind the counter.
3 · THE LINK
What's the real outcome this design is protecting, in plain terms?
Tap to flip
ANSWER
Genuinely harmful drug interactions caught before a prescription leaves the pharmacy, not a lower count of pop-up complaints.
4 · THE EARLY SIGNAL
What's the leading indicator this answer says to track?
Tap to flip
ANSWER
The share of alerts dismissed within two seconds, too fast to have been read, tracked every month.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
InteractGuard's original design logging whether an alert was dismissed, but never how fast, which meant the fast-dismiss rate could climb for months completely unseen.
6 · THE NUMBER
Fill in the blank: the share of alerts dismissed within two seconds rose from 8 percent in month one to ___ percent by month twelve.
Tap to flip
ANSWER
52 percent. More than half of all alerts were being dismissed too fast to have been genuinely read.
7 · THE DECISION STEP
What real threshold would trigger a design review, according to this answer?
Tap to flip
ANSWER
The fast-dismiss rate crossing 40 percent for two weeks running, which triggers a mandatory tiering review rather than a vague reminder to be careful.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what different abuse shows up there?
Tap to flip
ANSWER
Pemberton Language Services' ClauseFlag translation-review tool. There, the team suppressed flags on boilerplate-looking clauses to reduce reviewer fatigue, hiding nine genuinely ambiguous clauses a quarter behind a quieter dashboard.
Check yourself Score: 0 / 0
Short answer, name the design decision
1. State the one concrete interface decision this answer commits to.
Show hint
Look at the direct answer and Stage 4 of the walkthrough.
Show answer
Model answer: Tier the alert friction by real severity: a quiet badge for minor cases, one tap and a chosen reason for moderate ones, a locked screen requiring a real reason for genuinely severe ones.
Multiple choice
2. Why is the fast-dismiss rate the "early signal" here, rather than the near miss itself?
A. It's the easiest number for a pharmacy to report to regulators.
B. It had been climbing for months before the near miss ever happened, warning of the reflex before it cost anyone anything.
C. It's required by pharmacy licensing rules.
D. It's the only number InteractGuard was capable of tracking.
Show hint
Look at the line chart and the Early Signal step.
Show answer
B. The rate climbed from 8 percent to 52 percent over a year, well before the near miss made the problem visible any other way.
True or false
3. True or false: this answer recommends giving every flagged interaction, minor or severe, the same amount of friction to keep the design simple.
True
False
Show hint
Look at the Link and Decision steps.
Show answer
False. Identical friction for every severity is exactly the design flaw this answer traces back and fixes.
Fill in the blank
4. Fill in the blank: at Pemberton, suppressing flags on boilerplate clauses led to ___ genuinely ambiguous clauses missed per quarter, versus 1 when flags were kept but tiered.
Show hint
Look at the bar chart in Section 4.
Show answer
9 clauses per quarter. A quieter dashboard looked like progress while hiding nine real misses behind it.
Short answer, apply it yourself
5. Think of an app that shows you frequent alerts or confirmations. Do you actually read them anymore, or click through on reflex?
Show hint
Think of a permissions prompt, a terms-of-service checkbox, or a delivery confirmation.
Show answer
Model answer: Most repeated, identical confirmations get reflex-clicked, exactly the pattern this answer says to design against by tiering friction to real stakes.
Short answer, where it wouldn't matter
6. Name a part of InteractGuard's output where adding friction genuinely isn't needed.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Routine dosage reminders like "take with food." They carry no real risk, so a click requirement there would just add noise.
Before you close the answer
Why this works
Tests whether you can design an interface decision and pair it with a real number that proves the design still works months later, instead of a one-time UI idea with no way to detect its own failure.
Follow-up traps
"Won't a locked screen for severe cases just get annoying too, over time?" Response: only if severe cases are common; the whole point of tiering is that the highest-friction tier stays rare enough to keep meaning something, unlike a single tier applied to everything.
"What if pharmacists just get faster at clicking through the locked screen too?" Response: that's exactly what the fast-dismiss tracking is built to catch, if read-time on the locked tier starts falling, that's the signal to redesign the lock itself, not evidence the problem is solved.
If pressed
Fentridge's final design also required the chosen reason on a severe alert to route into the patient's own record, not just clear the screen, so a supervising pharmacist reviewing the chart later could see exactly which reason was picked, not just that the alert was dismissed.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.