ConceptIntermediateDesigning for Uncertainty & Trust / Human-in-the-loop product design / #9

Explain the risk of a human reviewer who rubber-stamps everything.

FLIPSthe over-trust flip, and the click that replaced a whole ritual

Anchorhold Health Plan uses an AI tool that recommends approving or denying prior-authorization requests, the paperwork a doctor files before an insurer will pay for a procedure. Graciela Dumitrescu is the nurse reviewer who signs off on every recommendation before it becomes the plan's final answer to a patient.

The direct answer
The real risk isn't that the reviewer gets lazy. It's that the review screen was built to reward not looking: the AI's recommendation loads pre-selected, so clicking submit without reading anything produces the exact same outcome as a careful read. Once that's true, rubber-stamping becomes the fast, safe, sensible choice, and the errors that get through are rare, large, and signed by a human name.
Do this, in order
  1. Remove the pre-selected default so no action is the same as any action.Why: as long as clicking nothing agrees with the AI by default, a rubber stamp and a real review look identical on the screen and in the log.
  2. Require a one-line reason before either approving or denying.Why: writing even one sentence forces a reviewer to actually locate the clinical detail that justifies the call, not just recognize a familiar layout.
  3. Track override rate by request type, watching for it drifting toward zero.Why: a rate that keeps falling with no real gain in model accuracy is the earliest sign the review has become theater.
  4. Run a monthly, independently sampled clinical audit against the reviewer's own decisions.Why: a rubber stamp and a genuinely well-calibrated trust look the same from the reviewer's own numbers. Only an outside sample tells them apart.
  5. Never blame the reviewer by name in the fix.Why: Graciela did the fast, sensible thing the screen invited her to do every single time. The fix belongs to the screen, not a performance review.
  6. Leave alone request types where the AI's call and the nurse's call have never once disagreed under audit.Why: some requests really are unambiguous, and treating every low override rate as suspicious wastes the friction budget on cases that don't need it.

How to answer this, stage by stage

Nobody is grading whether you can say the word "rubber-stamp." They're grading whether you can find the exact design choice that makes it the rational move.

Stage 1
Ground it in one review seat
Say it like this
"I'll ground this in a real seat: a health plan's prior-authorization queue, where a nurse reviewer signs off on an AI's approve-or-deny recommendation before it goes to the patient."
Why this works
Stops the answer from becoming a lecture about complacency in the abstract.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a method for finding the exact behavior change, not a vague warning about human nature.
Stage 3
Name the habit before it broke
Say it like this
"Graciela used to open the AI's reasoning panel and read the clinical notes on every request before deciding. That habit is what a good review actually is."
Why this works
Establishes real competence, so the later loss actually lands.
Stage 4
Name the exact flip, not a feeling
Say it like this
"The flip is opening the reasoning panel and reading it, versus clicking submit on whatever's pre-selected without opening anything. There's no middle setting between those two."
Why this works
This is the direct answer to the question, stated as a two-setting switch, not a mood.
Stage 5
Name the decision you'd take back
Say it like this
"The screen pre-selects the AI's recommendation by default. That made sense when the point was cutting clicks for a nurse handling eighty requests a day. It stopped making sense once no-action and agreement became indistinguishable."
Why this works
Names a specific, reversible product decision, not a vague call for "more diligence."
Stage 6
Prove it with a real day, replayed
Say it like this
"Same denial, no pre-selected default, a one-line reason required. Graciela has to actually choose, and choosing forces her to look for the fact that would justify it. She catches the mismatch before it ships."
Why this works
Shows the fix is a real design change, not a promise to try harder.
Stage 7
Say where this exact risk wouldn't apply
Say it like this
"For request types where the AI and the nurse have never once disagreed under audit, a fast, low-friction approval isn't a red flag. Some cases really are unambiguous."
Why this works
Shows judgment instead of treating every fast decision as evidence of a rubber stamp.
Stage 8
Close on the one line
Say it like this
"A reviewer who rubber-stamps everything isn't the risk. A screen that makes rubber-stamping and real reviewing produce the same outcome is the risk, and that's a design choice, not a character flaw."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Picture a queue of five hundred prior-authorization requests a day across a health plan's nursing team, and one screen that loads the same way every single time.

Anchorhold's tool reads each request, a doctor asking to be paid for a procedure, and recommends approve or deny based on the plan's coverage rules and the patient's chart. Before the tool, a nurse read every request cold and made the call herself, slower, but every decision came from an actual read.

Hand sketched icon list titled The five letters. Five items: F find the person whose morning is this, L locate the habit what they stopped doing, I identify the flip the verb that snaps, P pinpoint the old decision, S show the replay same day new design.
The five-letter memory aid. The I step, in red-orange, is the one worth sitting with.

Now the tool's recommendation loads onto the screen already selected. Graciela can open a panel showing the reasoning behind it, or she can just click submit. Both take about the same three seconds.

Here's the turn: the danger was never that Graciela stopped caring. It's that the screen made not-looking and looking produce the exact same outcome, so there was never a moment where she consciously chose to stop reviewing. The habit just thinned until there was nothing left of it.

Override rate on AI prior-authorization recommendations, by month
10% 5 0 Month 1, 9% Month 6, 2% Month 9, under 1%
Nobody was watching this line alone. It fell steadily for nine months before the internal audit found anything at all.

At its worst, an entire request category gets waved through untouched for months, and a patient's legitimate procedure gets denied on a technicality nobody actually read closely enough to catch.

Hand sketched comparison titled Small move, big snap. Left panel, an amber gauge icon labeled Gradual drift, caption override rate creeps down. Right panel, a red box icon labeled The snap, caption stops opening the panel at all.
The drift is gradual. The actual flip, opening the panel or not, has no middle setting at all.
The decision I would take back Anchorhold's review screen pre-selects the AI's recommendation as the default state of the approve-or-deny toggle. That made sense at launch, when the whole point was cutting clicks for a nurse working through hundreds of requests a shift. It stopped making sense once doing nothing and agreeing became the same action, with no way for the log to tell them apart.

What I would leave alone: routine refill requests for a medication a patient has already been approved for twice before have never once produced a disagreement under audit. A fast approval there isn't a rubber stamp, it's a correct read of a genuinely simple case.

The lesson: a screen that rewards not-looking will get not-looking, no matter how conscientious the person sitting in front of it is.

Now here is the same thing as a story

The short version above is what you'd say defending this fix to Anchorhold's clinical operations director. Read this one for how the ritual actually thinned out.

Graciela Dumitrescu has reviewed prior-authorization requests for nine years, and she used to be the nurse other reviewers asked to double-check a denial that felt wrong, the one who could spot a coverage rule misapplied to a patient's actual history.

Knowledge spark: what is the over-trust flip? Most flips happen when a tool gets worse and a person starts checking more. This one runs the other way. When a tool's calls keep looking right, a person's checking quietly thins from every case, to some cases, to none, and the errors that slip through are rarer and worse.

In her first months with the tool, Graciela opened the reasoning panel on every request, sometimes catching a coverage rule the AI had applied to the wrong plan tier. As months passed and the recommendations kept matching what she'd have decided anyway, she started clicking submit on the pre-selected call without opening the panel at all, first on the requests that looked routine, then on nearly everything.

Hand sketched metaphor scene titled A dial you assumed, a switch that's real. Left, a grey gauge icon labeled Dial, caption how much to check. Right, a red box icon labeled Switch, caption check or dont.
The design assumed a dial, a little more or less checking. What actually existed was a switch with no setting in between.

There was no single afternoon where this changed. No two-in-a-row moment, no bad case that scared her into a new habit. It built up slowly enough that Graciela herself couldn't have named the week the panel stopped opening.

Hand sketched timeline titled How the habit thinned. Four milestones: Launch default pre-selected, Override 9 percent to 5 percent month 3, Under 2 percent month 6 highlighted, Internal audit month 9.
Nine months between the default going live and anyone outside the queue finding out what it had done.

An internal clinical audit in month nine sampled a hundred and fifty recent denials and found eleven that should have been approved on the patient's own chart, the exact kind of misapplied rule Graciela used to catch on sight.

The denials weren't wrong because Graciela stopped caring. They were wrong because a screen that made looking and not-looking produce the same result had quietly made looking optional.

With the redesigned screen, no default is pre-selected, and a one-line reason is required before either approving or denying. Run the same nine months forward: the requirement to actually choose forces Graciela to locate the fact that justifies the call on every request, the eleven misapplied denials get caught in the months they would have happened, and the audit finds a residual rate the clinical team can actually defend.

The old screen asked "did a decision get logged." The new one asks "did a person actually decide."

I built the pre-selected default first because it was the fastest way to cut clicks for a busy queue. It took an internal audit and eleven misapplied denials to see that the fastest click and the real review had become the same click.

The five steps, if you want to remember itNot a lecture on complacency. FLIPS is what finds the exact click that replaced the ritual.

F
Find the person. Whose morning is this.
Graciela Dumitrescu, a prior-authorization nurse reviewer with nine years on the job.
Grounds the flip in a specific, competent person, not a persona.
L
Locate the habit. What she stopped doing.
Opening the AI's reasoning panel and reading the clinical justification before deciding, on every request.
Names the exact not-doing, not a general drop in effort.
I
Identify the flip. The verb that snaps.
Opening the panel and reading it, versus clicking submit on the pre-selected recommendation without opening anything at all. No middle setting.
The hardest step, and the direct answer to the question.
P
Pinpoint the old decision.
Pre-selecting the AI's recommendation as the toggle's default state, which made agreement and inaction indistinguishable.
Names a specific, reversible design choice, not a vague call for vigilance.
S
Show the replay.
No default pre-selected, a required one-line reason. The same eleven misapplied denials get caught in the months they happen, not nine months late.
Ends in a countable result, not a promise to try harder.
Hand sketched labeled parts diagram titled What the redesigned screen needs. Center document icon labeled Approval Screen, four callouts: no pre-set default, justification required, sampled audit, override log.
None of these four assume Graciela is careless. They assume the screen was doing the deciding for her, and stop letting it.

The recap, one line per letter: find is Graciela, nine years reviewing prior authorizations; locate is opening the reasoning panel before deciding; identify is the panel opening or not, with nothing in between; pinpoint is the pre-selected default that made agreement and inaction the same click; and show is the redesigned screen catching the same eleven denials in real time instead of nine months late.

And if you want to be sure it really works, try it somewhere elseA different flip family, a freight claims adjuster instead of a prior-authorization nurse. Not over-trust this time, but the same missing friction.

Vantport Logistics built an AI tool that scores freight damage claims and recommends a payout amount. Ottoline Reyes adjusts claims and, once payouts started tracking close to what she'd have calculated by hand, began accepting the recommended amount on nearly every claim without pulling the shipment's own inspection photos.

This one isn't over-trust exactly, it's the scope flip: instead of checking every claim in full, Ottoline started running the recommended payout against her memory of the shipment's general condition rather than the actual photos and paperwork, a smaller, cheaper unit of checking that felt like the same job but wasn't.

Hand sketched flow diagram titled Vantports freight claims chain. Four boxes: AI scores claim, Ottoline opens file, Clicks the default highlighted, Payout issued.
A different industry, a different flip family, the same missing step where a real check used to sit.
Freight claims later found to be overpaid, quarterly audit sample, before and after requiring photo review
16 8 0 14 Before photo review 2 After photo review
Same quarterly sample size both times. Requiring the photo to actually open closed most of the gap.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the screen made agreeing and not-checking the same click, so remove the default and require a reason," and stop.
Cost: there's no budget to redesign the whole approval screen this quarter. Say so honestly, and start by simply removing the pre-selected default, the cheapest single change, since the required-reason field can follow later.
The model gets better, for real: even if the AI's recommendations keep improving, the missing friction is still a risk, because a better model makes the fast click feel even more justified, not less.

Where people run it wrong.
They treat a low override rate as proof the model is excellent, without ever checking whether anyone's still actually looking.
They fix this by asking the reviewer to "be more careful," which is a request, not a design change, and it doesn't survive a busy Tuesday.
They wait for an external audit or a regulator's letter to reveal the problem, instead of building a leading signal that would have caught it months earlier.

How to use it live. When someone asks about a reviewer who rubber-stamps everything, ask one question first: does the screen make doing nothing and doing the job produce the same outcome? If it does, that's the actual risk, and it's fixable without ever mentioning the reviewer's name.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
The over-trust flip: checking sometimes to checking not at all, usually fired by the model looking better, or a screen that rewards not looking.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Graciela Dumitrescu, a prior-authorization nurse reviewer at Anchorhold Health Plan, nine years on the job.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Opening the AI's reasoning panel and reading the clinical justification before deciding on a request.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Opening the panel and reading it, versus clicking submit on the pre-selected call without opening anything at all.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pre-selecting the AI's recommendation as the default toggle state, which made agreeing and doing nothing the exact same click.
6 · THE NUMBER
Fill in the blank: the internal audit sampled 150 denials and found ___ that should have been approved.
Tap to flip
ANSWER
Eleven. Just over seven percent of the sample, all cases Graciela used to catch on sight before the habit thinned.
7 · THE REPLAY
Same nine months, redesigned screen. What changes?
Tap to flip
ANSWER
No pre-set default plus a required one-line reason forces Graciela to locate the justifying fact on every request, catching the eleven misapplied denials in real time instead of nine months late.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Vantport Logistics' freight claims tool. The scope flip: checking the full photo record shrinks down to checking from memory alone.

Check yourself Score: 0 / 0

True or false
1. True or false: the main risk in this story is that Graciela became careless or stopped caring about her job.
  • True
  • False
Show hint
Look at the direct answer and "the lesson."
Show answer
False. The risk is the screen's design: pre-selecting the AI's call made not-looking and looking produce the same outcome. Graciela made the fast, sensible choice every single time.
Multiple choice
2. Why is "add more training on paying attention" a weak fix for this exact problem?
  • A. Training costs too much money to run.
  • B. It's a request for more effort, not a design change, and it won't survive a screen that still rewards not looking.
  • C. Nurses are not allowed to attend training sessions.
  • D. The AI model itself needs the training, not the reviewer.
Show hint
Look at "where people run it wrong."
Show answer
B. A dial turned up, like a training reminder, doesn't change what the screen itself rewards. The fix has to change the click, not the person's intentions.
Fill in the blank
3. Fill in the blank: override rate on prior-authorization recommendations started at 9 percent and fell under ___ percent by month nine.
Show hint
Look at the line chart in Section 1.
Show answer
1 percent. A drop of nearly nine-fold over nine months, with no matching jump in the AI model's own accuracy.
Short answer, apply it yourself
4. Think of any approval or sign-off screen you've used, expense reports, terms and conditions, a software update. Where does clicking through without reading produce the exact same result as reading carefully?
Show hint
Think about any checkbox or button that's already checked or highlighted for you before you touch anything.
Show answer
Model answer: Most people can name a pre-checked box, a pre-selected shipping option, or a default-accepted setting where skipping the read changes nothing about what happens next.
Short answer, where it wouldn't matter
5. Name a prior-authorization request type at Anchorhold where a fast approval isn't actually a warning sign.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Routine refills for a medication a patient has already been approved for twice before. Audits have never found a disagreement there, so a fast approval reflects a genuinely simple case.
Short answer, the number question
6. If the audit had sampled 150 denials and found only 1 misapplied instead of 11, would removing the pre-selected default still be worth the added friction? Why or why not?
Show hint
Weigh the cost of the added click against the size and stakes of the harm.
Show answer
Model answer: Likely still yes for prior authorizations specifically, since a single misapplied denial can mean a delayed or denied medical procedure, a harm severe enough that even a lower error rate is worth the extra few seconds per request.
Before you close the answer
Why this works
Tests whether you can locate the exact design choice that makes a rubber stamp the rational move, instead of moralizing about human attention.
Follow-up traps
"Couldn't you just fire or retrain the reviewer instead?" Response: the next reviewer in that same seat, facing the same pre-selected default, would drift the same way, because the incentive lives in the screen, not the person.

"Won't requiring a written reason just slow down every single request, including the easy ones?" Response: pair it with the "leave alone" list, unambiguous request types where audits confirm the AI and the nurse always agree, and skip the friction there.
If pressed
Anchorhold's actual fix also randomizes which of the two toggle positions, approve or deny, appears on the left each time, since a fixed layout alone was found to create its own new muscle-memory click pattern within about six weeks.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more