ConceptAdvancedDesigning for Uncertainty & Trust / Human-in-the-loop product design / #3

Explain how review fatigue undermines a human-in-the-loop design.

FLIPS the over-trust flip, the one that fires on good news

Penwright & Cole LLP uses ClauseScan, a tool that reads an incoming contract during due diligence and flags clauses worth a lawyer's second look: indemnification, liability caps, assignment rights. Elin Kowalczyk is a junior associate who does the first pass on flagged clauses, and whose habit of actually reading them quietly changed underneath her, weeks before anyone noticed.

The direct answer
Review fatigue doesn't make a person check less carefully. It makes them stop checking at all, in one quiet snap, once a flag has been right often enough that opening it feels like a waste of the ten seconds it takes. A human-in-the-loop design that assumes a person will always open the file underneath the summary has already assumed away the exact failure it exists to catch.
Do this, in order
  1. Design for the day the reviewer stops opening the file, not the day they read every word.Why: this is the actual failure mode. A design built only for a careful reader is built for a version of the job that quietly stops existing.
  2. Mark which flags are genuinely unusual, not just which flags exist.Why: a flat, equally-confident summary on every flag is what makes a genuine outlier invisible inside a long, calm list.
  3. Track how often a flag actually gets opened, not just how often it gets cleared.Why: a clearance rate looks identical whether someone read the clause or just clicked past the tag.
  4. Force open a small, disguised sample of routine-looking flags on a fixed schedule.Why: this is the only way to know a reviewer's real open rate, not just their reported one.
  5. Leave genuinely boilerplate clauses, like a standard governing-law line, safe to skim from the summary alone.Why: not every flag deserves the same suspicion. The fix is for the rare outlier, not for punishing routine work.

How to answer this, stage by stage

Nobody is grading whether you know the phrase "review fatigue." They're grading whether you can name the exact moment it stops being a phrase and starts being a missed clause.

Stage 1
Ground it in one reviewer, one tool
Say it like this
"I'll use a law firm's contract review tool, ClauseScan, that flags risky clauses for a junior associate to check before a partner sees the file."
Why this works
Keeps "review fatigue" from staying an abstract phrase for the rest of the answer.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a method for finding the exact behavior that breaks, not a lecture on fatigue in general.
Stage 3
Name the habit the tool built
Say it like this
"Early on, Elin opened the full clause behind every single flag. That habit is the actual product. The time it saves is a side effect."
Why this works
Frames the habit as rational, not careless, which is what keeps the story honest.
Stage 4
Name the flip, not a feeling
Say it like this
"This isn't Elin trusting the tool a little more each week. It's a snap: opens every flag, then one week later, opens almost none. Two settings, no real middle."
Why this works
This is the direct answer, and the hardest step: a verb, not a mood.
Stage 5
Name the design decision that made it possible
Say it like this
"Every flag got the same calm, complete-sounding one-line summary, whether the clause behind it was routine or genuinely strange. Nothing on screen ever asked to be doubted."
Why this works
Names a real product decision, not "she should have read more carefully."
Stage 6
Say what wouldn't need fixing
Say it like this
"Genuinely boilerplate clauses, a standard governing-law line, are fine to clear from the summary alone almost every time. I wouldn't force a full open on those."
Why this works
Shows judgment, not a blanket rule that treats every flag as equally dangerous.
Stage 7
Prove the fix with a countable replay
Say it like this
"With a novelty marker on the flag, the unusual assignment clause stands out from the calm rows around it. Elin opens it the same afternoon it's flagged, not five weeks later in an audit."
Why this works
Ends on something countable, a same-day catch instead of a five-week gap.
Stage 8
Close on the one line
Say it like this
"Fatigue doesn't lower the bar someone clears. It removes the bar entirely, all at once, and a design that never notices that has already failed."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Here is what happens when a tool works fine for months, then fails in a way nobody wrote down as a rule.

ClauseScan reads an incoming contract during due diligence and flags the clauses worth a lawyer's second look, indemnification, liability caps, who a contract's rights transfer to if the company gets sold. Before it, an associate at Penwright & Cole read every page of every contract in a deal's document set by hand, about forty-five minutes a contract, ten contracts a realistic day's work.

Hand sketched flow diagram titled A contract, before ClauseScan. Four boxes: Contract arrives, Associate reads every page, Clause noted by hand, Filed after full read.
Slower, but every page passed under a person's eyes at least once.

Now ClauseScan flags roughly one clause in seven across a document set, and an associate like Elin reviews around 150 flagged clauses a day instead of reading every page of every contract, clearing a deal's whole document set in the time it used to take to read ten contracts start to finish.

Here's the turn: the handful of routine flags she waves through fast were never the danger. The danger is that her habit of opening the full clause behind a flag didn't fade gradually. It snapped, in about a week, from doing it every time to almost never, because the flag had been right, calmly and consistently, for too long to keep questioning.

Elin's full-context open rate on flagged clauses, by week of the due-diligence sprint
100% 50 0 Week 1, 95% Week 3, 41% Week 4, 19%
This wasn't a gentle slide. Most of the drop happened inside eight working days.

At its worst, a genuinely unusual clause, one ClauseScan correctly flagged, slides past unread because its one-line summary sounded exactly as calm as the hundred and fifty routine ones around it, and the client signs a deal carrying a risk nobody actually looked at.

Hand sketched comparison titled Small move big snap. Left, a green gauge icon labeled Rises slowly, caption a few clauses skipped. Right, a burgundy scale icon labeled Snaps flat then jumps, caption most clauses skipped.
The habit didn't dial down evenly. It held, then dropped, in about a week.
The decision I would take back Every ClauseScan flag got the same calm, complete-sounding one-line summary, whether the clause underneath was routine or genuinely unusual. That made sense when volume was low enough that associates opened every flag's full context anyway. It stopped making sense once trust and volume together made the one-liner become the entire review.

What I would leave alone: genuinely boilerplate clauses, a standard governing-law line with no unusual terms, really are fine to clear from the summary alone almost every time. This fix is for the rare outlier, not a blanket demand to reread everything twice.

The lesson: a flag that's always been right gets read like it will always be right. That is exactly the day it needs the closest look, and exactly the day nobody gives it one.

Now here is the same thing as a story

The short version above is what you'd say explaining this to Penwright & Cole's practice group. Read this one for how quietly the habit actually thinned.

Elin Kowalczyk has been a junior associate at Penwright & Cole for three years, and partners hand her the first pass on merger indemnification language specifically because she once caught a clause that would have made a client liable for a seller's pre-existing tax debt, buried in dense cross-reference language nobody else had time to trace.

Knowledge spark: why is this called an over-trust flip, not a verification flip? A verification flip is someone starting to check everything after being burned. An over-trust flip runs the other way: it fires when things have been going well, when a flag keeps turning out routine, and the person's checking quietly stops, not because they got lazy, but because checking kept confirming nothing was wrong.

When a large acquisition's due-diligence sprint began, Elin opened the full clause text behind every flag ClauseScan raised, printing the day's clause summaries onto a single sheet she annotated by hand in the margins as she went. In week one, she caught two clauses worth escalating to a partner, and ClauseScan's summaries matched her own read of the full text every time.

Hand sketched timeline titled Four weeks, no single moment. Four milestones: Sprint begins week 1, Opens fade week 2, Barely opens week 3, Routine audit highlighted week 5.
Nobody could point to the day it happened. It just wasn't happening anymore by week three.

By week three, Elin's printed sheet still had every flag listed, but she'd stopped annotating most of them, just initialing the routine ones without opening the full text behind the tag. There was no single afternoon where she decided to stop. The summaries had simply been right for so long that opening the file behind them started to feel like double-checking her own shoes were tied.

Hand sketched metaphor scene titled Switch not dial. Left, a green gauge icon labeled Dial, caption many settings. Right, a burgundy box icon labeled Switch, caption two positions only.
It looked like a dial turning down. It was actually a switch that had already flipped.

Deep in week four, ClauseScan flagged an assignment clause in a supplier contract with an unusual carve-out, rights that would transfer to a specific successor entity, not the acquirer itself, a genuinely rare structure. The summary read the same calm way every other flag had for a month: "Standard assignment language, no unusual terms." Elin initialed it and moved to the next file.

The clause was never hidden. It was flagged, correctly, in plain text. It just sounded exactly as ordinary as everything around it, and by week four, ordinary is what she'd stopped opening.

Nobody caught it in the moment, because nothing about the moment looked different from the hundred before it. It surfaced five weeks later during Penwright & Cole's routine quarterly quality audit, a scheduled, unglamorous sample of cleared flags, which found the carve-out and traced it back to Elin's initial without a full open.

Hand sketched labeled parts diagram titled What's inside a flag. Center document icon labeled ClauseScan Flag, with four callouts: one-line summary, novelty marker, confidence tag, full clause link.
The summary and the confidence tag already existed. Nothing on the card said this one was actually rare.

With the redesigned flag, a genuinely unusual clause carries a visible novelty marker, a small, distinct mark meaning "this pattern is rare in our data," separate from and independent of how confident ClauseScan's summary sounds. Run the same sprint forward: the assignment carve-out arrives in week four with that marker on it, stands out from the calm rows around it, and Elin opens the full text the same afternoon, not five weeks later in an audit.

The old flag asked whether the clause was risky. The new one also asks whether it was strange, and stopped assuming those were the same question.

I let every flag read in the same calm voice because differentiating them felt like it would just add noise to a clean screen. It took a routine quarterly audit, not a crisis, to see that the calm voice was exactly what let the one rare clause hide in plain sight.

The five steps, if you want to remember itNot a lecture on burnout. FLIPS is what tells you fatigue is a snap, not a slope.

F
Find the person. Whose morning is this?
Elin Kowalczyk, three years in, known for catching a tax-liability clause nobody else traced through dense cross-references.
Grounds "review fatigue" in one real reviewer instead of a category of user.
L
Locate the habit. What did she stop doing because it worked?
Opening the full clause text behind every flag, a habit that faded because the flags kept turning out routine.
Names the habit as rational, earned by weeks of the tool being right, not as carelessness.
I
Identify the flip. What verb snaps?
Opens every flag's full text, versus barely opens any. A one-week snap, not a six-week slide, is what the data actually showed.
The hardest step, and the direct answer: fatigue is the over-trust flip, not a mood or a slow decline.
P
Pinpoint the old decision. Which choice only made sense before?
Giving every flag the same calm, complete-sounding one-line summary, regardless of how rare the underlying clause pattern actually was.
A real, reversible design choice, not "she should have tried harder."
S
Show the replay. Same trigger, new design.
The novelty marker separates "rare" from "confident." The carve-out gets opened the same afternoon, five weeks earlier than the audit would have caught it.
Ends on something countable: same-day, not five weeks later.
Hand sketched icon list titled The five letters, FLIPS. Five items: F find the person, L locate the habit, I identify the flip, P pinpoint the decision, S show the replay.
Five letters, one method. The one in the middle is the only one that's genuinely hard.

The recap, one line per letter: find is Elin, three years in and trusted with the hard clauses; locate is the habit of opening every flag's full text; identify is the over-trust flip, a one-week snap from opens-everything to opens-almost-nothing; pinpoint is a flag design that sounded equally calm whether the clause was routine or rare; and show is a novelty marker that turns a five-week audit gap into a same-day catch.

And if you want to be sure it really works, try it somewhere elseA different flip family, a port terminal instead of a law firm. Not fatigue this time, a private workaround.

Cascade Bay Terminal uses ManifestGuard, a tool that reads incoming cargo manifests and flags shipments worth a closer customs look. Petra Lindholm is a customs inspector there, and instead of an over-trust flip, her story is a workaround flip: ManifestGuard has no history feature, so when she overrides a flag, that override vanishes the moment she closes the file.

Hand sketched flow diagram titled Petra's private tally. Four boxes: Manifest scanned, Model flags item, Private note kept highlighted, Note never syncs back.
The tool never asked her to keep this record. So she built one it would never see.

Mapped onto FLIPS: find is Petra, six years inspecting cargo, known for catching mislabeled chemical shipments by smell alone before any paperwork confirmed it. Locate is her habit of trusting ManifestGuard's flags outright, since disputing one meant redoing the whole inspection form from scratch. Identify is the workaround flip: instead of trusting the tool as her whole workflow, she starts keeping a private paper tally of every override in a desk drawer, because the tool itself remembers none of it. Pinpoint is the decision never to build a history or notes field into ManifestGuard, reasonable when inspectors handled low volume and rarely needed to look back. Show is that with a real override log built into the tool, Petra's tally becomes the system's own record, visible to the next shift instead of locked in a drawer only she can read.

Overridden flags Cascade Bay could actually trace back to a reason, before and after adding an override log
100% 50 0 12% Before the log 91% After the log
Petra's private tally had the answers the whole time. The system just never asked her for them.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "review fatigue is a snap from opens-everything to opens-almost-nothing, not a slow decline, so design for the snap, not the slope," and stop.
Cost: there's no budget to build a novelty-detection model this quarter. Say so honestly, and start with a much simpler proxy, flagging any clause pattern that's appeared fewer than five times in the firm's own history.
The model gets better, for real: an even more accurate ClauseScan makes the snap happen faster, not slower, since the flip is driven by how often the flag turns out routine, not by how good the underlying model actually is.

Where people run it wrong.
They treat review fatigue as a training problem, telling reviewers to "stay focused," when the actual failure is structural, not a matter of willpower.
They measure success by clearance rate, which looks identical whether a reviewer read the clause or clicked past the tag.
They wait for a dramatic single incident to notice the drift, when the whole point of fatigue is that it rarely announces itself that way.

How to use it live. When someone asks about review fatigue, ask what percentage of flags actually get opened, not cleared. A falling open rate is the tell, and it shows up long before a bad clause ever does.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is review fatigue, in this answer?
Tap to flip
ANSWER
The over-trust flip: checks sometimes, then stops checking at all. It fires on good news, when a flag keeps turning out routine.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Elin Kowalczyk, a junior associate at Penwright & Cole LLP, three years in, known for catching a buried tax-liability clause.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Opening the full clause text behind every flag, instead of trusting the one-line summary alone.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Opens every flag's full text, versus barely opens any at all. The data showed a one-week snap, not a gradual slide.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing every flag's summary in the same calm, complete-sounding voice, regardless of how rare the underlying clause pattern actually was.
6 · THE NUMBER
Fill in the blank: Elin's open rate fell from 95 percent in week 1 to ___ percent by week 4.
Tap to flip
ANSWER
19 percent. Most of that drop happened inside about eight working days, not across the whole sprint.
7 · THE REPLAY
Same week four, redesigned flag. What changes?
Tap to flip
ANSWER
The novelty marker sets the assignment carve-out apart from the calm rows around it. Elin opens it the same afternoon, not five weeks later in an audit.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and which flip family?
Tap to flip
ANSWER
Cascade Bay Terminal's ManifestGuard. The flip there is a workaround flip, a private paper tally built because the tool kept no history of its own.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Elin's full-context open rate dropped from 95 percent to ___ percent between week 1 and week 4 of the sprint.
Show hint
Look at the line chart in Section 1.
Show answer
19 percent. A drop of roughly 76 points, most of it inside about eight working days, which is what makes it a snap and not a slope.
Multiple choice
2. Why is review fatigue better described as a snap than as gradually "checking a bit less carefully"?
  • A. Because reviewers are told to stop checking after a fixed number of days.
  • B. Because the AI model itself changes its behavior after enough correct flags.
  • C. Because the data shows an open rate holding steady, then dropping sharply within about a week, not sliding evenly over months.
  • D. Because law firms require associates to log their open rate weekly.
Show hint
Look at the Identify the flip step and the tells that you picked wrong.
Show answer
C. A gradual slide would be a dial. What the data actually shows, a steady rate followed by a sharp one-week drop, is a flip: two settings, no real middle.
True or false
3. True or false: the fix for review fatigue in this story is to make every flag demand a full open, no exceptions.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Genuinely boilerplate clauses are fine to clear from the summary alone. The fix targets the rare, unusual clause specifically, not every flag equally.
Short answer, apply it yourself
4. Think of something you review often, a spell-checker's suggestions, a spam folder, a code linter. What's a habit you stopped doing because it kept confirming everything was fine?
Show hint
Think about a suggestion tool that was right often enough that double-checking it started to feel pointless.
Show answer
Model answer: Many people stop reading a spell-checker's flagged word before accepting the fix, or stop opening a spam folder's contents before deleting it, once the tool has been right for long enough.
Short answer, where it wouldn't matter
5. Name a kind of clause in Penwright & Cole's system where this exact fatigue problem genuinely doesn't matter.
Show hint
Look at "what I would leave alone" in Section 1.
Show answer
Model answer: A standard, boilerplate governing-law clause with no unusual terms. Clearing it from the summary alone is genuinely fine almost every time, since there's rarely anything underneath worth a second look.
Short answer, the number question
6. If the sprint had run for only one week instead of five, would review fatigue have been a real risk for Elin? Why or why not?
Show hint
Look at how many flags it actually took for the open rate to start dropping.
Show answer
Model answer: Probably not much of one. The flip needed enough repeated, confirmed-routine flags to build the habit of trusting the summary. A single week likely wouldn't supply enough of them to snap the habit yet.
Before you close the answer
Why this works
Tests whether you understand fatigue as a structural, one-way behavior change, not a mood a reviewer could just decide to snap out of.
Follow-up traps
"Couldn't you just tell reviewers to stay vigilant?" Response: the habit forms because checking kept confirming things were fine, which isn't a discipline problem, it's a rational response to a design that never signaled which flags were actually rare.

"Isn't a novelty marker just another number reviewers will eventually stop trusting too?" Response: possibly, over a much longer horizon, which is exactly why detect matters here too, tracking the open rate on marked clauses specifically, not just assuming one fix lasts forever.
If pressed
Penwright & Cole's actual fix also logs how long a flag stayed open before being cleared, since a marked, "rare" clause cleared in under three seconds is a second signal that the marker itself has started being ignored.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more