Explain how review fatigue undermines a human-in-the-loop design.
Penwright & Cole LLP uses ClauseScan, a tool that reads an incoming contract during due diligence and flags clauses worth a lawyer's second look: indemnification, liability caps, assignment rights. Elin Kowalczyk is a junior associate who does the first pass on flagged clauses, and whose habit of actually reading them quietly changed underneath her, weeks before anyone noticed.
- Design for the day the reviewer stops opening the file, not the day they read every word.Why: this is the actual failure mode. A design built only for a careful reader is built for a version of the job that quietly stops existing.
- Mark which flags are genuinely unusual, not just which flags exist.Why: a flat, equally-confident summary on every flag is what makes a genuine outlier invisible inside a long, calm list.
- Track how often a flag actually gets opened, not just how often it gets cleared.Why: a clearance rate looks identical whether someone read the clause or just clicked past the tag.
- Force open a small, disguised sample of routine-looking flags on a fixed schedule.Why: this is the only way to know a reviewer's real open rate, not just their reported one.
- Leave genuinely boilerplate clauses, like a standard governing-law line, safe to skim from the summary alone.Why: not every flag deserves the same suspicion. The fix is for the rare outlier, not for punishing routine work.
How to answer this, stage by stage
Nobody is grading whether you know the phrase "review fatigue." They're grading whether you can name the exact moment it stops being a phrase and starts being a missed clause.
Let's learn
Here is what happens when a tool works fine for months, then fails in a way nobody wrote down as a rule.
ClauseScan reads an incoming contract during due diligence and flags the clauses worth a lawyer's second look, indemnification, liability caps, who a contract's rights transfer to if the company gets sold. Before it, an associate at Penwright & Cole read every page of every contract in a deal's document set by hand, about forty-five minutes a contract, ten contracts a realistic day's work.
Now ClauseScan flags roughly one clause in seven across a document set, and an associate like Elin reviews around 150 flagged clauses a day instead of reading every page of every contract, clearing a deal's whole document set in the time it used to take to read ten contracts start to finish.
Here's the turn: the handful of routine flags she waves through fast were never the danger. The danger is that her habit of opening the full clause behind a flag didn't fade gradually. It snapped, in about a week, from doing it every time to almost never, because the flag had been right, calmly and consistently, for too long to keep questioning.
At its worst, a genuinely unusual clause, one ClauseScan correctly flagged, slides past unread because its one-line summary sounded exactly as calm as the hundred and fifty routine ones around it, and the client signs a deal carrying a risk nobody actually looked at.
What I would leave alone: genuinely boilerplate clauses, a standard governing-law line with no unusual terms, really are fine to clear from the summary alone almost every time. This fix is for the rare outlier, not a blanket demand to reread everything twice.
The lesson: a flag that's always been right gets read like it will always be right. That is exactly the day it needs the closest look, and exactly the day nobody gives it one.
Now here is the same thing as a story
The short version above is what you'd say explaining this to Penwright & Cole's practice group. Read this one for how quietly the habit actually thinned.
Elin Kowalczyk has been a junior associate at Penwright & Cole for three years, and partners hand her the first pass on merger indemnification language specifically because she once caught a clause that would have made a client liable for a seller's pre-existing tax debt, buried in dense cross-reference language nobody else had time to trace.
When a large acquisition's due-diligence sprint began, Elin opened the full clause text behind every flag ClauseScan raised, printing the day's clause summaries onto a single sheet she annotated by hand in the margins as she went. In week one, she caught two clauses worth escalating to a partner, and ClauseScan's summaries matched her own read of the full text every time.
By week three, Elin's printed sheet still had every flag listed, but she'd stopped annotating most of them, just initialing the routine ones without opening the full text behind the tag. There was no single afternoon where she decided to stop. The summaries had simply been right for so long that opening the file behind them started to feel like double-checking her own shoes were tied.
Deep in week four, ClauseScan flagged an assignment clause in a supplier contract with an unusual carve-out, rights that would transfer to a specific successor entity, not the acquirer itself, a genuinely rare structure. The summary read the same calm way every other flag had for a month: "Standard assignment language, no unusual terms." Elin initialed it and moved to the next file.
Nobody caught it in the moment, because nothing about the moment looked different from the hundred before it. It surfaced five weeks later during Penwright & Cole's routine quarterly quality audit, a scheduled, unglamorous sample of cleared flags, which found the carve-out and traced it back to Elin's initial without a full open.
With the redesigned flag, a genuinely unusual clause carries a visible novelty marker, a small, distinct mark meaning "this pattern is rare in our data," separate from and independent of how confident ClauseScan's summary sounds. Run the same sprint forward: the assignment carve-out arrives in week four with that marker on it, stands out from the calm rows around it, and Elin opens the full text the same afternoon, not five weeks later in an audit.
The old flag asked whether the clause was risky. The new one also asks whether it was strange, and stopped assuming those were the same question.
I let every flag read in the same calm voice because differentiating them felt like it would just add noise to a clean screen. It took a routine quarterly audit, not a crisis, to see that the calm voice was exactly what let the one rare clause hide in plain sight.
The five steps, if you want to remember itNot a lecture on burnout. FLIPS is what tells you fatigue is a snap, not a slope.
The recap, one line per letter: find is Elin, three years in and trusted with the hard clauses; locate is the habit of opening every flag's full text; identify is the over-trust flip, a one-week snap from opens-everything to opens-almost-nothing; pinpoint is a flag design that sounded equally calm whether the clause was routine or rare; and show is a novelty marker that turns a five-week audit gap into a same-day catch.
And if you want to be sure it really works, try it somewhere elseA different flip family, a port terminal instead of a law firm. Not fatigue this time, a private workaround.
Cascade Bay Terminal uses ManifestGuard, a tool that reads incoming cargo manifests and flags shipments worth a closer customs look. Petra Lindholm is a customs inspector there, and instead of an over-trust flip, her story is a workaround flip: ManifestGuard has no history feature, so when she overrides a flag, that override vanishes the moment she closes the file.
Mapped onto FLIPS: find is Petra, six years inspecting cargo, known for catching mislabeled chemical shipments by smell alone before any paperwork confirmed it. Locate is her habit of trusting ManifestGuard's flags outright, since disputing one meant redoing the whole inspection form from scratch. Identify is the workaround flip: instead of trusting the tool as her whole workflow, she starts keeping a private paper tally of every override in a desk drawer, because the tool itself remembers none of it. Pinpoint is the decision never to build a history or notes field into ManifestGuard, reasonable when inspectors handled low volume and rarely needed to look back. Show is that with a real override log built into the tool, Petra's tally becomes the system's own record, visible to the next shift instead of locked in a drawer only she can read.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "review fatigue is a snap from opens-everything to opens-almost-nothing, not a slow decline, so design for the snap, not the slope," and stop.
Cost: there's no budget to build a novelty-detection model this quarter. Say so honestly, and start with a much simpler proxy, flagging any clause pattern that's appeared fewer than five times in the firm's own history.
The model gets better, for real: an even more accurate ClauseScan makes the snap happen faster, not slower, since the flip is driven by how often the flag turns out routine, not by how good the underlying model actually is.
Where people run it wrong.
They treat review fatigue as a training problem, telling reviewers to "stay focused," when the actual failure is structural, not a matter of willpower.
They measure success by clearance rate, which looks identical whether a reviewer read the clause or clicked past the tag.
They wait for a dramatic single incident to notice the drift, when the whole point of fatigue is that it rarely announces itself that way.
How to use it live. When someone asks about review fatigue, ask what percentage of flags actually get opened, not cleared. A falling open rate is the tell, and it shows up long before a bad clause ever does.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a novelty marker just another number reviewers will eventually stop trusting too?" Response: possibly, over a much longer horizon, which is exactly why detect matters here too, tracking the open rate on marked clauses specifically, not just assuming one fix lasts forever.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Human-in-the-loop product design
- #1 When should a human be required to approve an AI action rather than merely able to?
- #2 Design the review interface for a human checking 200 AI-generated outputs an hour.
- #4 What is the difference between human-in-the-loop and human-on-the-loop?
- #5 How do you decide which cases get routed to a human?
- #6 Describe a confidence-based routing policy and its failure mode.
- #7 How do you measure whether the human in the loop is adding value?