Explain when a human checkpoint is theatre rather than control.
Cinderford Home Pharmacy fills prescriptions by mail and runs an AI model that flags likely drug interactions before a pharmacist confirms and dispenses. Naledi Dlamini has verified those flags for four years, and used to open the evidence behind every single one before she clicked confirm.
- Measure whether review time actually varies with the flag's real risk, not just whether a click happens.Why: a checkpoint that takes the same two seconds every time, easy or dangerous, is a click, not a check.
- Put the pause back where a merged, one-click design removed it.Why: collapsing "view evidence" and "confirm" into one action is exactly what turns a real check into a reflex.
- Route low-confidence or unusual flags through a genuinely different, slower path.Why: the rare case is precisely the one a worn-smooth habit is least likely to catch.
- Track the open-and-read rate on the evidence panel, not just the confirm rate.Why: a confirm rate can look perfectly healthy right up until the morning it wasn't.
- Leave the checkpoint alone anywhere the AI is genuinely near-perfect and the miss cost is low.Why: not every check is worth defending. Some really can be automated away.
How to answer this, stage by stage
Nobody is grading whether you can define "rubber stamping." They're grading whether you can name the exact measurement that tells theatre and control apart.
Let's learn
Say Cinderford Home Pharmacy fills prescriptions by mail. An AI model reads every new order against the patient's medication history and flags anything that looks like a risky interaction, before a pharmacist confirms and it gets dispensed.
Before the model, a pharmacist checked every single new prescription against the patient's chart by hand. Thorough, and slow, several minutes each.
Now the model flags about sixty interactions a week out of thousands of orders, and a pharmacist reviews each flag before confirming.
Here's the turn: the flags themselves were never the problem. The model stayed accurate. What changed was what a pharmacist did with a flag once opening the evidence behind it stopped feeling necessary.
At its worst, a genuinely rare interaction, the kind the model itself flags with less confidence, gets the same reflexive click as the sixty routine ones that week. The checkpoint is still there. It just isn't checking anything anymore.
What I would leave alone: refill confirmations for medications with no interaction history at all don't need any of this. There's nothing risky hiding behind those, and a fast, light checkpoint there is honestly fine.
The lesson: a checkpoint that never changes its own pace, no matter what's behind the flag, was never really a checkpoint. It was a click wearing one.
Now here is the same thing as a story
The short version above is what you'd say to a pharmacy board reviewing the interaction workflow. Read this one for the two seconds that almost cost someone.
Naledi Dlamini has verified drug interaction flags at Cinderford for four years. She can picture the mechanism behind most common interactions without looking anything up, from memory alone.
When the interaction model launched, she opened the evidence panel behind every flag, read the mechanism, and confirmed or escalated. It always checked out. Over about a year, without any one moment marking the change, she opened it less. Then rarely. Then, without quite noticing when it happened, not at all.
A pharmacist new to the mail-order side, shadowing her one afternoon, watched her clear a stretch of flags in seconds each and said, half-curious, "You don't even open those anymore, do you?" Naledi laughed it off. She wasn't wrong, though, and the comment sat with her longer than she expected.
That same week, a flag came through for a returning patient starting a new antidepressant while already on a painkiller, a pairing the model scored as medium confidence, lower than its usual all-clear. Her hand was already moving toward confirm, the same two seconds as the sixty routine flags that week, when the comment from days earlier snagged something, and she opened the evidence instead.
The panel showed what the model's training data hadn't: the patient's intake notes mentioned a supplement, the same one that pushes that exact pairing into genuinely dangerous territory, information sitting in a free-text field the model wasn't weighing. She held the fill and called the prescriber.
With the redesigned screen, the evidence panel has to actually open, a real two-second minimum view, before the confirm button activates at all. Run the same week forward with that in place: the medium-confidence flag can't be waved through in a reflex, because the reflex itself no longer has anywhere to land.
I let the click get faster because faster looked like the system working. It took a new colleague's offhand question, and a supplement two lines down in a chart, to see that a fast click and a real check had quietly stopped being the same thing.
FLIPS, the over-trust version, in one screenNot the usual "model got worse" story. This is what happens when it gets better.
The recap, one line per letter: find is Naledi at her screen, locate is the evidence-panel habit thinning to nothing, identify is the over-trust flip with no stable middle, pinpoint is the one-click merge that made sense before precision improved, and show is a re-split checkpoint that catches the case the old one would have waved through.
And if you want to be sure it really works, try it somewhere elseA different flip family entirely, a legal translation firm instead of a pharmacy.
Meridian Legal Translations uses an AI model to draft first-pass translations of contracts, which a human translator then reviews before delivery. Emil Rutkowski has translated legal documents there for eight years.
Mapped to a different family, the concealment flip: when the AI drafts started, Emil marked every delivered document "AI-assisted, reviewed by Emil," proud of the speed. A few embarrassing early drafts, an AI mistranslation of a liability clause that a client's own lawyer caught, made that label feel like a confession instead of a badge. He quietly stopped marking anything, delivering AI-assisted drafts as fully his own work. Nobody flagged errors to the team anymore, since nobody was owning up to using the tool in the first place, and the firm's own error-tracking data went dark exactly when it needed to see the most.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a checkpoint is theatre when review time stops varying with real risk, whatever the reason the person stopped truly looking," and stop.
Cost: there's no budget to redesign the confirm flow this quarter. Say so honestly, and start by simply logging open-versus-not-opened on the evidence panel, since that number alone will show whether the checkpoint is still real.
The model gets better, for real: if accuracy keeps improving, the temptation to skip the check only grows, which is exactly why this failure mode gets more likely, not less, as a model gets better.
Where people run it wrong.
They watch the confirm rate and call the checkpoint healthy, when a rubber-stamped confirm looks identical to a genuine one on that number alone.
They assume a checkpoint that catches nothing this month is a checkpoint that's working perfectly, instead of asking whether it's still capable of catching anything.
They treat "the click still happens" as proof of oversight, when the click was never the thing doing the checking.
How to use it live. When someone asks about theatre versus control, ask one question first: does this checkpoint take longer on the case that deserves it? If the honest answer is no, you've found the theatre, before you even need the story.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't a pharmacist just open the panel and still not really read it?" Response: possibly, which is why the open-and-read rate should be paired with a periodic spot check against her notes, the same audit discipline any leading-indicator metric needs.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Human-in-the-loop product design
- #1 When should a human be required to approve an AI action rather than merely able to?
- #2 Design the review interface for a human checking 200 AI-generated outputs an hour.
- #3 Explain how review fatigue undermines a human-in-the-loop design.
- #4 What is the difference between human-in-the-loop and human-on-the-loop?
- #5 How do you decide which cases get routed to a human?
- #6 Describe a confidence-based routing policy and its failure mode.