ConceptAdvancedDesigning for Uncertainty & Trust / Human-in-the-loop product design / #16

Explain when a human checkpoint is theatre rather than control.

FLIPS an interaction check that got very good, and very quiet

Cinderford Home Pharmacy fills prescriptions by mail and runs an AI model that flags likely drug interactions before a pharmacist confirms and dispenses. Naledi Dlamini has verified those flags for four years, and used to open the evidence behind every single one before she clicked confirm.

The direct answer
A checkpoint has become theatre the moment the time spent on it stops varying with how serious the case actually is. If a routine flag and a rare, dangerous one both get the same two-second glance before confirm, a click is happening, but nothing is being checked. Real control shows up as attention that scales with risk. Theatre looks identical no matter what's underneath it.
Do this, in order
  1. Measure whether review time actually varies with the flag's real risk, not just whether a click happens.Why: a checkpoint that takes the same two seconds every time, easy or dangerous, is a click, not a check.
  2. Put the pause back where a merged, one-click design removed it.Why: collapsing "view evidence" and "confirm" into one action is exactly what turns a real check into a reflex.
  3. Route low-confidence or unusual flags through a genuinely different, slower path.Why: the rare case is precisely the one a worn-smooth habit is least likely to catch.
  4. Track the open-and-read rate on the evidence panel, not just the confirm rate.Why: a confirm rate can look perfectly healthy right up until the morning it wasn't.
  5. Leave the checkpoint alone anywhere the AI is genuinely near-perfect and the miss cost is low.Why: not every check is worth defending. Some really can be automated away.

How to answer this, stage by stage

Nobody is grading whether you can define "rubber stamping." They're grading whether you can name the exact measurement that tells theatre and control apart.

Stage 1
Scope it to one checkpoint
Say it like this
"I'll ground this in a mail-order pharmacy, where an AI flags likely drug interactions and a pharmacist confirms before dispensing."
Why this works
Keeps "theatre versus control" from staying an abstract phrase.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay. This is the over-trust version, where the model getting better is what causes the flip."
Why this works
Signals that this isn't the usual "model gets worse" story, it's the inverted one.
Stage 3
Name the habit that faded
Say it like this
"Naledi used to open the evidence panel behind every flag before confirming. As the model kept being right, she opened it less, then rarely, then not at all."
Why this works
Shows the habit fading as a rational response, not carelessness.
Stage 4
Identify the flip, the hard step
Say it like this
"The flip is opens-and-reads versus never-opens. There's no stable middle. Once she stopped opening it, the confirm click took the same two seconds no matter what was behind it."
Why this works
This is the direct answer to the question, stated as a real behavior with two settings.
Stage 5
Pinpoint the old decision
Say it like this
"Cinderford's software merged 'view evidence' and 'confirm' into one click, back when alert volume was high and mostly noise. That made sense then. It stopped making sense once the alerts got rare and real."
Why this works
Names a specific, reasonable-at-the-time decision, not a vague policy failure.
Stage 6
Show the replay
Say it like this
"Split the click back into two steps. The evidence panel has to actually open before confirm activates. Same flag, same Tuesday, and this time she reads it, and catches the supplement the model never saw in the chart."
Why this works
Ends on a countable result, not a vague "much safer."
Stage 7
Close on the one line
Say it like this
"A checkpoint that takes the same two seconds whether the case is routine or the kind that kills someone was never really checking anything."
Why this works
Restates the direct answer, sharp enough to survive a follow-up.

Let's learn

Say Cinderford Home Pharmacy fills prescriptions by mail. An AI model reads every new order against the patient's medication history and flags anything that looks like a risky interaction, before a pharmacist confirms and it gets dispensed.

Before the model, a pharmacist checked every single new prescription against the patient's chart by hand. Thorough, and slow, several minutes each.

Now the model flags about sixty interactions a week out of thousands of orders, and a pharmacist reviews each flag before confirming.

Hand sketched comparison titled Small move, big snap. Left panel, a gauge icon labeled Slow fade, caption opens it a little less each month. Right panel, a box icon labeled Flat, then gone, caption stops opening it at all.
The fade looks gradual from the inside. From the outside, it has exactly two settings.

Here's the turn: the flags themselves were never the problem. The model stayed accurate. What changed was what a pharmacist did with a flag once opening the evidence behind it stopped feeling necessary.

Percent of flagged interactions where the evidence panel was actually opened before confirming
100% 50 0 Month 1, 95% Month 8, 34% Month 13, 4%
Nobody decided to stop checking. It happened a little at a time, and the confirm rate never once looked unhealthy.

At its worst, a genuinely rare interaction, the kind the model itself flags with less confidence, gets the same reflexive click as the sixty routine ones that week. The checkpoint is still there. It just isn't checking anything anymore.

Average seconds spent before confirming, common flag versus rare high-risk flag
4 sec 2 0 2.1 sec Common flag 2.3 sec Rare, high-risk flag
A tenth of a second of difference. That gap is the whole story, and it's nearly nothing.
The decision I would take back Cinderford's software team merged "view the evidence" and "confirm the flag" into one click, back when alert volume was high and mostly false alarms on common drug pairs. That made the workflow faster, and it made sense at the time. It stopped making sense once the model's precision improved and the remaining flags became rare but genuinely dangerous, because the one click that used to force at least a glance at the evidence became a single reflexive tap.

What I would leave alone: refill confirmations for medications with no interaction history at all don't need any of this. There's nothing risky hiding behind those, and a fast, light checkpoint there is honestly fine.

The lesson: a checkpoint that never changes its own pace, no matter what's behind the flag, was never really a checkpoint. It was a click wearing one.

Now here is the same thing as a story

The short version above is what you'd say to a pharmacy board reviewing the interaction workflow. Read this one for the two seconds that almost cost someone.

Naledi Dlamini has verified drug interaction flags at Cinderford for four years. She can picture the mechanism behind most common interactions without looking anything up, from memory alone.

Hand sketched timeline titled Fourteen months of habit thinning. Four milestones: Opens every flag month one, Opens most flags month four, Opens rarely month eight, Stops opening entirely month thirteen highlighted.
No single month looks alarming. The whole line does.

When the interaction model launched, she opened the evidence panel behind every flag, read the mechanism, and confirmed or escalated. It always checked out. Over about a year, without any one moment marking the change, she opened it less. Then rarely. Then, without quite noticing when it happened, not at all.

A pharmacist new to the mail-order side, shadowing her one afternoon, watched her clear a stretch of flags in seconds each and said, half-curious, "You don't even open those anymore, do you?" Naledi laughed it off. She wasn't wrong, though, and the comment sat with her longer than she expected.

Hand sketched metaphor scene titled Switch, not dial. Left, a gauge icon labeled DIAL, caption many settings, gradual. Right, a box icon labeled SWITCH, caption two settings, no middle.
Naledi's checking never had a dial. It only ever had two positions.

That same week, a flag came through for a returning patient starting a new antidepressant while already on a painkiller, a pairing the model scored as medium confidence, lower than its usual all-clear. Her hand was already moving toward confirm, the same two seconds as the sixty routine flags that week, when the comment from days earlier snagged something, and she opened the evidence instead.

The panel showed what the model's training data hadn't: the patient's intake notes mentioned a supplement, the same one that pushes that exact pairing into genuinely dangerous territory, information sitting in a free-text field the model wasn't weighing. She held the fill and called the prescriber.

The checkpoint had been there the entire time. For thirteen months, it just hadn't been checking anything.

With the redesigned screen, the evidence panel has to actually open, a real two-second minimum view, before the confirm button activates at all. Run the same week forward with that in place: the medium-confidence flag can't be waved through in a reflex, because the reflex itself no longer has anywhere to land.

I let the click get faster because faster looked like the system working. It took a new colleague's offhand question, and a supplement two lines down in a chart, to see that a fast click and a real check had quietly stopped being the same thing.

FLIPS, the over-trust version, in one screenNot the usual "model got worse" story. This is what happens when it gets better.

Hand sketched icon list titled The five letters. Five items: F find the person, L locate the habit, I identify the flip, P pinpoint the old decision, S show the replay.
The same five steps, whichever direction the model's accuracy is moving.
F
Find the person.
Naledi Dlamini, four years verifying interaction flags at a mail-order pharmacy.
Grounds the answer in one real screen, not a policy.
L
Locate the habit.
She opened the evidence panel behind every flag, then most, then rarely, then never, as the model kept being right.
Shows the fade as rational, not careless.
I
Identify the flip. Over-trust family.
Opens-and-reads versus never-opens, no stable middle. This is what turns a real check into theatre.
The hardest step, and the direct answer to the question.
P
Pinpoint the old decision.
Merging "view evidence" and "confirm" into one click, back when alert volume was high and mostly noise.
A specific, reasonable decision that stopped fitting reality.
S
Show the replay.
The evidence panel must actually open before confirm activates. Same week, the medium-confidence flag gets a real look, and the supplement gets caught.
Ends on a countable, specific catch, not a vague improvement.

The recap, one line per letter: find is Naledi at her screen, locate is the evidence-panel habit thinning to nothing, identify is the over-trust flip with no stable middle, pinpoint is the one-click merge that made sense before precision improved, and show is a re-split checkpoint that catches the case the old one would have waved through.

And if you want to be sure it really works, try it somewhere elseA different flip family entirely, a legal translation firm instead of a pharmacy.

Meridian Legal Translations uses an AI model to draft first-pass translations of contracts, which a human translator then reviews before delivery. Emil Rutkowski has translated legal documents there for eight years.

Mapped to a different family, the concealment flip: when the AI drafts started, Emil marked every delivered document "AI-assisted, reviewed by Emil," proud of the speed. A few embarrassing early drafts, an AI mistranslation of a liability clause that a client's own lawyer caught, made that label feel like a confession instead of a badge. He quietly stopped marking anything, delivering AI-assisted drafts as fully his own work. Nobody flagged errors to the team anymore, since nobody was owning up to using the tool in the first place, and the firm's own error-tracking data went dark exactly when it needed to see the most.

Hand sketched flow diagram titled The concealment flip. Four boxes: AI draft used, one bad draft embarrasses him, stops disclosing AI use highlighted, feedback loop goes dark.
A completely different flip family. The same theatre problem shows up anyway, one step further upstream.
Hand sketched quadrant titled Which checkpoints need a real second look. X axis how rare the case is, Y axis how bad a miss would be. Items placed: routine refill common low stakes, rare interaction rare high stakes highlighted, contract boilerplate common low stakes, liability clause rare high stakes highlighted.
The upper right corner is where a fast, uniform click is the most dangerous kind of theatre.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a checkpoint is theatre when review time stops varying with real risk, whatever the reason the person stopped truly looking," and stop.
Cost: there's no budget to redesign the confirm flow this quarter. Say so honestly, and start by simply logging open-versus-not-opened on the evidence panel, since that number alone will show whether the checkpoint is still real.
The model gets better, for real: if accuracy keeps improving, the temptation to skip the check only grows, which is exactly why this failure mode gets more likely, not less, as a model gets better.

Where people run it wrong.
They watch the confirm rate and call the checkpoint healthy, when a rubber-stamped confirm looks identical to a genuine one on that number alone.
They assume a checkpoint that catches nothing this month is a checkpoint that's working perfectly, instead of asking whether it's still capable of catching anything.
They treat "the click still happens" as proof of oversight, when the click was never the thing doing the checking.

How to use it live. When someone asks about theatre versus control, ask one question first: does this checkpoint take longer on the case that deserves it? If the honest answer is no, you've found the theatre, before you even need the story.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust. Checks sometimes, then stops checking at all. Fires when the model gets better, not worse, and rare errors ship unseen.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Naledi Dlamini, a pharmacist at Cinderford Home Pharmacy, four years verifying AI-flagged drug interactions.
3 · THE HABIT
What did she stop doing because it kept working?
Tap to flip
ANSWER
Opening the evidence panel behind a flag before confirming it. It kept checking out, so she opened it less and less.
4 · THE FLIP
What's the two-setting switch here?
Tap to flip
ANSWER
Opens-and-reads the evidence versus never-opens it. No stable middle setting exists once the habit fully thins out.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging "view evidence" and "confirm" into one click, made sense when alert volume was high and mostly noise, wrong once flags got rare and real.
6 · THE NUMBER
Fill in the blank: the evidence-panel open rate fell from 95 percent in month one to ___ percent by month thirteen.
Tap to flip
ANSWER
4 percent. The confirm rate itself never once looked unhealthy while this happened.
7 · THE REPLAY
Same medium-confidence flag, redesigned screen. What changes?
Tap to flip
ANSWER
Confirm can't activate until the evidence panel actually opens, so the flag gets a real look and the missed supplement gets caught before the fill.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Meridian Legal Translations' contract review. The concealment flip: disclosing AI use, then quietly hiding it after one embarrassing draft.

Check yourself Score: 0 / 0

Multiple choice
1. What actually turns Naledi's checkpoint into theatre?
  • A. The AI model's accuracy dropped over the thirteen months.
  • B. Cinderford stopped tracking the confirm rate entirely.
  • C. Review time stopped varying with how risky the flag actually was, so a routine and a rare case got the same two-second glance.
  • D. The pharmacy hired a less experienced pharmacist.
Show hint
Look at the direct answer.
Show answer
C. Real control shows up as attention that scales with risk. Theatre looks the same no matter what's underneath it.
True or false
2. True or false: a healthy-looking confirm rate proves a checkpoint is still functioning as real oversight.
  • True
  • False
Show hint
Look at "where people run it wrong."
Show answer
False. A rubber-stamped confirm looks identical to a genuine one on that number alone. The open-and-read rate is what actually reveals the difference.
Fill in the blank
3. Fill in the blank: Naledi's habit thinned over about ___ months before she stopped opening the evidence panel entirely.
Show hint
Look at the timeline diagram.
Show answer
Thirteen. No single month looked alarming on its own. The whole line, watched over a year, did.
Short answer, apply it yourself
4. Think of a confirmation click you do reflexively now, cookie banners, terms of service, an approval step at work. What would make that click a real check again?
Show hint
Think about what would have to slow you down before you could click through.
Show answer
Model answer: Something that forces the actual content to be seen before the click activates, the same fix as Cinderford's re-split screen, rather than trusting that the click itself means attention was paid.
Short answer, where it wouldn't matter
5. Name a place in Cinderford's system where a fast, unread confirm click genuinely isn't a problem.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Refill confirmations for medications with no interaction history at all. There's no real risk hiding behind those, so a light, fast checkpoint is genuinely fine there.
Short answer, the number question
6. If the model's confidence scoring were removed entirely, so every flag looked identical, would the same fix, requiring the panel to open before confirming, still help? Why or why not?
Show hint
Think about what the fix actually forces to happen, independent of the confidence score.
Show answer
Model answer: Yes. The fix doesn't depend on the confidence score at all, it forces a look at the evidence on every flag, which is what catches information the model never had, like a supplement in free-text notes.
Before you close the answer
Why this works
Tests whether you can name a measurable difference between real oversight and a click that only looks like it, instead of just asserting one exists.
Follow-up traps
"Isn't forcing the panel open just going to slow everyone down for no reason?" Response: it only forces a real, brief look, not a full manual re-check, and the alternative is a checkpoint that catches nothing while looking exactly like one that does.

"Couldn't a pharmacist just open the panel and still not really read it?" Response: possibly, which is why the open-and-read rate should be paired with a periodic spot check against her notes, the same audit discipline any leading-indicator metric needs.
If pressed
Cinderford's real fix also logs how many seconds the evidence panel stayed open before confirm, since a panel that opens and closes in under a second is functionally the same reflex wearing a new click.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more