When should a human be required to approve an AI action rather than merely able to?
GUARD the letter that mailed itself before anyone could stop it
Brookline Loan Servicing uses PathForward, a tool that reads a homeowner's hardship file and recommends approving a repayment plan or denying it and moving the account to the foreclosure track. Delphine Achebe is a loss-mitigation specialist who has worked hardship files for nine years, and who nearly let one bad denial reach a real family.
The direct answer
Require a person's approval, don't just make one available, whenever the action is hard to undo and moves faster than the person it affects can react to it. Here, that means any denial that also queues a foreclosure-track referral or a credit bureau report needs a specialist's sign-off before it fires. Anything reversible, like approving help, can stay merely reviewable.
Do this, in order
Require approval whenever the action is hard to undo and outruns the person it affects.Why: this is the actual line between "required" and "able." Everything else is a refinement of where that line sits.
Split the decision from the next action, so a denial can never auto-queue its own foreclosure referral.Why: merging the two removed the one pause where a person could still catch a bad call before it moved.
Leave low-confidence approvals reviewable, not required, since a wrongful approval is cheap to unwind.Why: mandatory gates cost real time. Spend that cost only where the harm can't be taken back.
Track how often a required gate gets approved without a real look, not just whether it fired.Why: a gate that always says yes in under four seconds isn't a gate anymore, it's a checkbox.
Give the homeowner a visible hold period after any denial, even a gated one, before the next step executes.Why: a required approval still needs a window on the other side, or the person it affects still can't catch a mistake in time.
How to answer this, stage by stage
Nobody is grading whether you can define "human in the loop." They're grading whether you can name the actual line that decides when a checkbox isn't enough.
Stage 1
Scope it to one real decision
Say it like this
"I'll ground this in a mortgage servicer's hardship tool, one that can recommend approving a repayment plan or denying it and starting foreclosure."
Why this works
Keeps the answer from floating into a general policy essay about AI oversight.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's affected. Unequal, where the harm lands hardest. Ability to contest, who can push back and who can't. Reduce, the actual design change. Detect, how you'd know it's working."
Why this works
Signals a method built for power imbalance, not a generic risk checklist.
Stage 3
Name both people
Say it like this
"There's the specialist, who can override the model any time she wants. And there's the homeowner, who finds out about a denial after it's already queued the next step."
Why this works
Makes the power gap concrete instead of abstract, before naming any fix.
Stage 4
Give the actual test for "required"
Say it like this
"Require it when two things are both true: the action is hard to undo, and the person it affects can't realistically catch it and object before the next step fires."
Why this works
This is the direct answer, and it's a test you can apply to a case you've never seen before.
Stage 5
Name the design decision that broke this
Say it like this
"PathForward's denial and its next collections step got merged into one motion, to cut turnaround time. That removed the only pause a person had to catch a mistake."
Why this works
Names a real, reversible product decision, not "add more review" in disguise.
Stage 6
Say what you would not gate
Say it like this
"I wouldn't require approval on auto-approved repayment plans. A wrongful approval costs the company money, but it's cheap to catch and unwind, nobody's harmed by getting help they didn't strictly qualify for."
Why this works
Shows judgment, not blanket caution, which is what separates a real policy from a fear response.
Stage 7
Say how you'd know the gate still works
Say it like this
"I'd watch how long a specialist actually spends on a required approval. If gated cases start clearing in under five seconds, the gate is decoration, not a check."
Why this works
Shows you think past the day the policy ships, to the day it quietly stops meaning anything.
Stage 8
Close on the one line
Say it like this
"Make it required exactly where the person it affects has no real chance to catch a mistake before it moves. Everywhere else, able is enough."
Why this works
Restates the direct answer in one breath, ready for a follow-up.
Let's learn
What happens the moment a machine's decision can't be undone?
PathForward reads a homeowner's hardship file, income, expenses, the reason for the missed payments, and recommends approving a new repayment plan or denying it and moving the account toward foreclosure. Before it, every one of Brookline's roughly 1,400 monthly hardship files went to a specialist team of nine, each file taking about thirty-five minutes to read start to finish, a six-day average turnaround.
Slow, but every single file passed through a person before anything was mailed.
Now PathForward clears about two-thirds of files on its own in under a minute, and only sends the harder or more uncertain ones to Delphine's team, whose average review time has dropped to nine minutes.
Here's the turn: the occasional wrong call was never really the danger. The danger was a design choice that let a denial and the next collections step fire in the same motion, so by the time anyone might catch the mistake, the account was already moving.
The third box used to be a person. Merging steps quietly removed it.
Hardship denials overturned on appeal, gated versus ungated cases, last quarter
Same rough volume of denials, both categories. The gate did not slow down every file, only the ones where being wrong was expensive.
At its worst, a family gets a foreclosure notice mailed before anyone at Brookline even learns their file was misread, and by the time a specialist catches it on appeal, real damage, credit, legal fees, stress, is already done.
One of these two people can stop the letter. The other only finds out it was sent.
The decision I would take back
We merged the denial step and the "queue next collections action" step into one automated motion at launch, purely to cut the six-day turnaround. That made sense when PathForward's accuracy looked strong in testing. It stopped making sense once real volume meant a small error rate turned into real families every month.
What I would leave alone: auto-approved repayment plans don't need a mandatory gate. A wrongful approval costs Brookline money, not the homeowner, and it's easy to catch and correct later. Not every AI decision needs the same brake pedal.
The lesson: the question was never whether PathForward is usually right. It's whether the person on the receiving end has any real window to catch it when it's wrong.
Now here is the same thing as a story
The short version above is what you'd say defending this policy to Brookline's compliance committee. Read this one for how close the near miss actually came.
The loss-mitigation floor at Brookline gets loud around month-end, when the hardship queue fills up ahead of the next payment cycle. Delphine Achebe has worked that floor for nine years, and she is the one who once caught a self-employment income mismatch that would have denied a nurse's repayment plan over a clerical rounding error nobody else had noticed.
Knowledge spark: what's the difference between required and able?
"Able" means a person can step in if they choose to. "Required" means the action is physically blocked until a person signs off. The gap between the two only matters once you ask whether the person affected has time to ask for that step-in before it's too late.
For its first four months, PathForward's volume stayed small enough that Delphine's team saw nearly every flagged file the same day it came in. The system worked exactly as designed: model decides, decision executes, letter mails, all in one motion, because volume was low enough that nothing ever really got missed.
The design didn't change between month one and month five. The volume running through it did.
By month four, hardship volume had nearly tripled after a rate reset hit thousands of adjustable loans at once, and PathForward's denial-and-referral motion kept firing at the same speed as always, just far more often.
The same word, "review," described two very different guarantees, and nobody had picked which one this case needed.
In month five, a new hire on Delphine's team asked her, almost in passing, why a denial letter for a widow's file had already gone out before anyone flagged the income document looked incomplete. Delphine pulled the file. The income document was, in fact, incomplete, missing a pension statement that would have qualified the woman for the plan she'd asked for. The denial, and the foreclosure-track referral behind it, had already mailed two days earlier.
Nobody had decided the widow's file shouldn't get a second look. The system had simply never given anyone the chance to look before the letter was already gone.
Delphine caught it in time to intercept the referral and reverse the denial before the foreclosure track actually advanced, but only because a new hire happened to ask an odd question on a slow afternoon. An audit of the prior six weeks found four more files where the same gap existed, unnoticed, because nothing had gone wrong loudly enough yet.
A file only needs all four rarely. A denial that also starts foreclosure usually has every one of them.
With the redesigned policy, any denial that also triggers a foreclosure-track referral or a credit bureau report is blocked until a specialist actively signs it, a real click, on a real screen, not a default. Run the same month forward: the widow's incomplete file gets caught at the gate itself, before any letter exists to intercept, and the file is resolved in one day instead of nearly reaching a real foreclosure step two days after the fact.
The old design asked whether a person could look. The new one asks whether a person had to.
I let the denial and the referral merge into one motion because splitting them felt like giving back the speed we'd just built. It took a new hire's offhand question to see that speed was never the thing worth protecting on the file that turns out to be wrong.
GUARD, the letters that decide who gets a leverNot a lecture on fairness in the abstract. GUARD is what tells you exactly when "able" stops being enough.
G
Groups. Who's affected.
Delphine, who holds the lever and can override any recommendation. The homeowner, who has no lever at all, only a letter after the fact.
Names both people before naming any fix.
U
Unequal. Where the harm lands.
A wrongful denial's harm, foreclosure risk, credit damage, stress, lands entirely on the homeowner, who often can't respond fast enough to matter before the next step already fired.
Shows the imbalance isn't hypothetical, it's asymmetric by design.
A
Ability to contest. Who never gets to push back.
Merging the denial and the referral into one motion meant the homeowner's only realistic chance to contest, the specialist's sign-off, had already been removed before the letter went out.
The hardest step, and the direct answer: require approval exactly where this window doesn't otherwise exist.
R
Reduce. The actual design change.
Split the decision from the next action. Any denial that also queues a foreclosure referral or credit report is blocked until a specialist actively signs it.
A concrete product decision, not a policy memo or a training session.
D
Detect. How you'd know it's still working.
Track how long a specialist actually spends on a required approval. A gate that clears in under five seconds on every case has quietly become a rubber stamp.
Catches the gate decaying into decoration before an audit has to find it for you.
Two questions, asked every time: can it be undone, and does the person affected have time to ask first.
The recap, one line per letter: groups is the specialist who holds the lever and the homeowner who doesn't, unequal is a wrongful denial's cost landing entirely on the homeowner, ability to contest is the missing pause once decision and action got merged, reduce is splitting them back apart with a real sign-off, and detect is watching how long that sign-off actually takes.
And if you want to be sure it really works, try it somewhere elseSame five letters, a county unemployment office instead of a mortgage servicer. A different harm, the same missing pause.
Marrow County's unemployment office uses a model that flags benefit claims as likely eligible or likely fraudulent based on wage-reporting patterns. Anders Kessler reviews the claims it flags, and until recently, a likely-fraudulent flag froze a claimant's payments the same day, automatically, with a form letter mailed a week later.
Mapped onto GUARD: groups is Anders, who can release a frozen payment any time he wants, against a claimant who has no way to contest a freeze before their rent is due. Unequal is that a false fraud flag can cost someone their apartment before the appeal letter even arrives. Ability to contest is the real gap: the freeze and the mailing both fired the moment the model flagged the claim, with no pause for a person to look first.
Same shape, a different office. Before either system, a person read the file before anything moved.
Days a wrongly frozen claim went unpaid, before and after requiring sign-off on a freeze
Requiring sign-off before the freeze fires did more for this number than any appeals-process speedup ever did.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "require approval exactly where the action is hard to undo and outruns the person it affects, leave the rest merely reviewable," and stop.
Cost: there's no budget to add reviewers to every gate this quarter. Say so honestly, and gate only the subset of denials that also trigger foreclosure or fraud freezes, since that's where irreversible harm actually concentrates.
The model gets better, for real: even a model that's rarely wrong still needs the gate, since the whole point was never accuracy, it was giving the affected person a window that used to not exist at all.
Where people run it wrong.
They make everything "able to review" and call it human-in-the-loop, without asking whether anyone actually has time to use that ability before the harm lands.
They gate everything equally, which burns out the very reviewers meant to catch the cases that matter, on cases that never needed a brake pedal.
They ship the gate and never check whether it's still a real look or a five-second click nobody would notice was automatic.
How to use it live. When someone asks when a human must approve, ask two questions back: can this be undone, and does the person it affects have any real time to object before it moves. If both answers are no, that's your required gate.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "who must approve this" question?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Built to find who can push back and who can't.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Delphine Achebe, a loss-mitigation specialist who has worked hardship files at Brookline Loan Servicing for nine years.
3 · UNEQUAL
Where does the harm of a wrongful denial actually land?
Tap to flip
ANSWER
Entirely on the homeowner, in foreclosure risk and credit damage, not on Brookline, who can absorb a mistake far more easily.
4 · ABILITY TO CONTEST
What removed the homeowner's real chance to catch a mistake?
Tap to flip
ANSWER
Merging the denial decision and the foreclosure referral into one automated motion, so the letter was already mailed before anyone could realistically object.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging the decision step and the next collections action into one motion at launch, to cut turnaround time, which made sense at low volume and stopped making sense at real scale.
6 · THE NUMBER
Fill in the blank: ungated denials were overturned on appeal ___ times last quarter, versus 3 for gated ones.
Tap to flip
ANSWER
31. Roughly ten times as many, on a similar volume of cases.
7 · THE REPLAY
Same month five, redesigned policy. What changes?
Tap to flip
ANSWER
The widow's incomplete file gets caught at the gate itself, resolved in one day, instead of a denial mailing two days before anyone noticed the missing pension statement.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the required gate there?
Tap to flip
ANSWER
Marrow County's unemployment fraud-flag system. The required gate is a caseworker's sign-off before any wage-fraud flag can freeze a claimant's payment.
Check yourself Score: 0 / 0
True or false
1. True or false: an action should require human approval any time the AI might be wrong.
True
False
Show hint
Look at the direct answer's two-part test.
Show answer
False. The AI being wrong isn't the test. The test is whether the action is hard to undo and whether the affected person has real time to catch it first. A required gate on every action would make speed impossible and teach nobody anything.
Multiple choice
2. Why did merging the denial and the foreclosure referral into one motion matter more than PathForward's accuracy?
A. It made the tool slower.
B. It made PathForward's model less accurate over time.
C. It removed the only pause where a person could catch a mistake before it reached the homeowner.
D. It increased Brookline's legal costs directly.
Show hint
Look at the Ability to contest step.
Show answer
C. Accuracy was never the core issue. The issue was that a wrong call and its consequence fired together, with nobody positioned to intervene between them.
Fill in the blank
3. Fill in the blank: an audit of the six weeks before the near miss found ___ more files with the same missing pause.
Show hint
Look at the story's audit paragraph, right after the near miss.
Show answer
Four. Four more files, unnoticed only because nothing had gone wrong loudly enough yet to surface them on its own.
Short answer, apply it yourself
4. Think of an app or service where an automated action of yours was hard to undo. What would have made you feel it should have required your sign-off instead of just letting you review it after?
Show hint
Think of an autopay, an auto-renewal, or an auto-sent message.
Show answer
Model answer: Many people can name an autopay or auto-renewal that charged them before they noticed, precisely because reviewing it after the fact meant the money, or the commitment, was already gone.
Short answer, where it wouldn't matter
5. Name a decision in Brookline's system where "able to review" is genuinely enough, and required approval would just be friction.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Auto-approved repayment plans. A wrongful approval costs Brookline money, not the homeowner, and it's cheap to catch and correct later without any real harm having landed.
Short answer, the number question
6. If gated denials had also started clearing in under five seconds on average, would that change your assessment of whether the fix worked? Why or why not?
Show hint
Look at the Detect step and stage 7 of the walkthrough.
Show answer
Model answer: Yes. A five-second clearance time on a case complex enough to need a gate suggests the sign-off has become a rubber stamp, not a real look, which means the fix has quietly stopped doing its job.
Before you close the answer
Why this works
Tests whether you can name a real, checkable line for "required" instead of a vague comfort level with "human oversight."
Follow-up traps
"Isn't requiring approval on everything just the safest choice?" Response: it burns reviewer attention on cases that don't need it, and a reviewer facing that volume starts clicking through required gates just as fast as they would an optional one.
"What if the specialist just rubber-stamps the required approvals anyway?" Response: that's exactly why detect exists, tracking how long a sign-off actually takes catches a gate quietly becoming a checkbox before an audit has to find it for you.
If pressed
Brookline's actual fix also added a 24-hour hold on any gated denial after sign-off, so even an approved denial has one more window before the letter physically mails, a second, cheaper safeguard behind the first.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.