ConceptIntermediateDesigning for Uncertainty & Trust / Human-in-the-loop product design / #4

What is the difference between human-in-the-loop and human-on-the-loop?

PICK the dollar rule that guarded the wrong thing

Meridian Payments runs Sentry Flag, a tool that scores every transaction for fraud risk in real time. Beatriz Salcedo is a fraud analyst who wears a headset through most of her shift, working a stream of flagged transactions that have already gone through by the time she looks at them.

The direct answer
Human-in-the-loop blocks the action until a person approves it. Human-on-the-loop lets the action complete and gives a person a chance to catch it after. The real test for which one you need isn't how big or risky the action looks, it's whether it can still be undone once a person actually does look. If the harm is already permanent by review time, on-the-loop was never really oversight at all.
Do this, in order
  1. Pick human-in-the-loop for anything that can't be undone once it happens.Why: this is the actual line between the two modes. Everything else is a refinement of where you draw it.
  2. Stop sorting into the two modes by dollar amount alone.Why: a small, irreversible loss can be worse than a large, reversible one. Size was never the right proxy for risk.
  3. Check reversibility per payment type, not per transaction size.Why: two products at the same dollar amount can have completely different real-world reversal windows.
  4. Keep human-on-the-loop for actions with a real, working reversal path.Why: blocking everything burns the very reviewer attention that on-the-loop monitoring is supposed to save.
  5. Re-check the split whenever a new product or transaction type launches.Why: a rule built for one payment type silently stops fitting the moment a new, less reversible one ships under the same threshold.

How to answer this, stage by stage

Nobody is grading whether you can recite two definitions. They're grading whether you can say which one a specific action actually needs, and defend it.

Stage 1
Answer with a position, not a definition
Say it like this
"In-the-loop blocks the action first. On-the-loop lets it happen and reviews it after. I'd pick between them based on whether the action can still be undone, not on how big it is."
Why this works
Commits to an answer immediately instead of drifting into a glossary entry.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my answer up front. Impact, who feels each kind of error. Cost asymmetry, which one is actually more expensive. Kill criteria, what would change my answer."
Why this works
Signals a method built to commit and defend, not hedge with "it depends."
Stage 3
Ground it in one real system
Say it like this
"I'll use a payments company's fraud system. Sentry Flag scores every transaction, and a fraud analyst named Beatriz reviews the ones that already went through."
Why this works
Keeps the answer concrete instead of drifting into a general policy essay.
Stage 4
Name the asymmetry, in real units
Say it like this
"Blocking a legitimate card purchase costs a few minutes of friction, maybe a retried checkout. Letting an irreversible transfer complete and turn out fraudulent costs the entire amount, permanently, no matter how fast Beatriz catches it after."
Why this works
This is the hardest step: showing one error is cheap and the other is expensive, in real terms.
Stage 5
Name the decision that got this wrong before
Say it like this
"Meridian sorted into the two modes by dollar amount, a rule copied from card transactions, which have chargeback rights. Instant transfers don't, and that rule quietly put them in the wrong mode."
Why this works
Shows the mistake was a real, specific design choice, not a vague failure to "be careful."
Stage 6
Say what would change your pick
Say it like this
"If instant transfers ever got a real, working reversal window, a 24-hour recall, on-the-loop would be fine for them too. My pick moves with the actual reversal path, not with a fixed rule."
Why this works
Gives kill criteria, the thing that separates a confident answer from a stubborn one.
Stage 7
Say what wouldn't change
Say it like this
"Routine card purchases stay on-the-loop no matter what. Chargeback rights already give them a real reversal path, so blocking them would just be friction with no matching safety gain."
Why this works
Shows judgment, not a reflex to gate everything the moment one gap gets found.
Stage 8
Close on the one line
Say it like this
"In-the-loop and on-the-loop aren't about how much you trust the model. They're about whether a mistake is still fixable by the time a person actually sees it."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Say we build a fraud system that has to make one decision, over and over, thousands of times an hour: block now, or let it through and look later.

Sentry Flag reads every transaction Meridian processes, card purchases and real-time bank transfers alike, and scores it for fraud risk in about 40 milliseconds. Before it, Meridian held every transaction over 2,000 dollars for manual review, a blanket rule that added a 45-minute average hold to any large purchase, card or transfer, no matter how routine.

Hand sketched metaphor scene titled Turnstile versus camera. Left, a navy box icon labeled Gate, caption blocks first. Right, a blue gauge icon labeled Watch, caption catches after.
One design stops you before you pass. The other only knows you passed.

Now most transactions clear instantly, and only a small tier gets held for a person's approval before completing at all. Everything else, including plenty of large ones, moves through immediately and reaches Beatriz's review stream afterward.

Here's the turn: the split between "block first" and "review after" was never really about the dollar amount. It was about whether the money could still be pulled back once someone looked, and a rule built around size alone quietly stopped asking that question.

Hand sketched timeline titled One quarter, one gap. Four milestones: Instant transfer launches month 1, Ring finds the gap highlighted month 2, Losses climb month 3, Gating added month 4.
The gap didn't open the day instant transfers launched. It opened the day someone noticed the rule didn't ask about them differently.
Flagged fraud losses actually recovered, by payment type, same review mode
100% 50 0 78% Card transactions 2% Instant transfers
Same review mode, same fraud team, same speed of catching it. The difference is whether the money could still move back.

At its worst, a fraud ring finds the gap before anyone else does: thousands of instant transfers, each small enough to clear Meridian's dollar threshold, none of them reversible once the recipient cashes out, all of them technically caught by Beatriz's review, and none of that catching worth anything by the time she sees it.

Hand sketched comparison titled The asymmetry, drawn. Left, a small green scale icon labeled Card fraud, caption 78 percent recovered. Right, a large red gauge icon labeled Instant transfer fraud, caption gone for good.
Two boxes the same size on a dashboard. Nowhere near the same size once you ask what a review can still do about them.
The decision I would take back We sorted transactions into "block first" and "review after" purely by dollar amount, a rule inherited from Meridian's original card-only system, where chargeback rights made a fixed size threshold reasonable. It stopped making sense the day instant transfers launched, a payment type with no reversal window at all, still sorted by the same old rule.

What I would leave alone: routine card purchases stay human-on-the-loop, full stop. Chargeback rights already give them a real path back, so blocking them would just add friction with nothing gained in return.

The lesson: the two modes were never a spectrum of trust in the model. They're a question about whether tomorrow can still fix today's mistake.

Now here is the same thing as a story

The short version above is what you'd say defending this fix to Meridian's risk committee. Read this one for how the gap actually got found.

Beatriz Salcedo has worked fraud review at Meridian for five years, and she is the one who once traced a wave of card disputes back to a single compromised merchant terminal, three weeks before the card networks flagged it themselves.

Knowledge spark: why does chargeback right matter so much here? A chargeback lets a cardholder's bank pull disputed money back from the merchant's bank, even weeks after a purchase clears. It's a built-in undo button most people never think about. Payment types without one don't have that safety net once the money moves.

Meridian's instant transfer product launched quietly, aimed at people splitting rent or paying a contractor same-day, and it inherited the existing dollar-based rule without anyone specifically re-checking whether that rule still made sense for a payment type with no reversal window.

Hand sketched decision tree titled Which one do you actually need. Root: Can this be undone once a person looks. Four branches: yes cheaply leads to Human-on-the-loop, no money's gone leads to Human-in-the-loop, yes but slowly leads to Human-on-the-loop, no harm already landed leads to Human-in-the-loop.
The question was always sitting right here. Nobody had asked it specifically about the new payment type.

A fraud ring found the gap within weeks: thousands of instant transfers, each just under Meridian's block threshold, moved through instantly, reviewed only after the fact, and by the time Beatriz's queue flagged the pattern, every recipient account had already been emptied and closed.

Beatriz caught every single one of those transfers, eventually. Catching them after the money was gone was never the same thing as stopping them.

A routine quarterly risk audit, the same kind Meridian ran every quarter without expecting anything unusual, sampled recovery outcomes across payment types and found the 2-percent recovery rate on instant transfers sitting next to card fraud's 78 percent, same review process, wildly different result.

Hand sketched icon list titled What actually decides it. Four items: a scale icon labeled Can it be undone, a gauge icon labeled How fast harm lands, a question mark box icon labeled How often it's wrong, a box icon labeled Cost of blocking it.
Dollar amount isn't on this list. It was never actually the thing that mattered.

With the redesigned split, instant transfers above a much lower threshold move to human-in-the-loop, blocked until Beatriz or a teammate actively clears them, regardless of how the dollar amount compares to a card transaction. Run the same quarter forward: the ring's pattern gets caught at the third or fourth attempt, blocked before completion, instead of thousands of transfers deep before a routine audit finally connects the dots.

The old rule asked how big the transaction was. The new one asks whether tomorrow can still undo it.

I built the split around dollar amount because it had worked fine for years on a system that was only ever card transactions. It took a routine audit, not a dramatic breach, to see that a new payment type had quietly broken an assumption nobody had actually written down.

PICK, in one screenNot a debate about which mode is "better." PICK is what tells you which one a specific action actually needs.

P
Position. The pick, before the reasoning.
Human-in-the-loop for anything that can't be undone once it happens. Human-on-the-loop for anything that still can.
Commits to an answer immediately, instead of describing both modes and stopping there.
I
Impact. Who feels each kind of error.
A blocked legitimate card purchase costs the customer a few minutes of friction. An irreversible fraudulent transfer that clears costs the full amount, permanently, to whoever it was stolen from.
Names both costs in real units, not as abstractions.
C
Cost asymmetry. The one that's actually expensive.
Friction is cheap and visible, a retried checkout. A completed, irreversible fraud loss is hidden until review and permanent by the time anyone sees it.
The hardest step, and the direct answer: optimize against the hidden, permanent one.
K
Kill criteria. What would flip the pick.
If instant transfers ever gained a real reversal window, on-the-loop would become defensible for them too. The pick tracks the reversal path, not a fixed rule.
Separates a confident, revisable answer from a stubborn one.

The recap, one line per letter: position is in-the-loop for anything irreversible, on-the-loop for anything that isn't, impact is friction against permanent loss, cost asymmetry is that the hidden, unrecoverable loss is the one to design against, and kill criteria is that a real reversal window would change the pick entirely.

And if you want to be sure it really works, try it somewhere elseSame four letters, a telehealth pharmacy instead of a payments company. A different kind of "can't take it back."

Fernglade TeleRx auto-approves routine refills of a patient's existing, stable medication, at the same dose, no pharmacist review before dispensing. Ansel Meurer is a pharmacy technician who reviews a sample of those refills after they've already shipped.

Mapped onto PICK: position is human-on-the-loop for a stable refill at an unchanged dose, human-in-the-loop for any new medication or dosage change. Impact is that reviewing a stable refill after shipping costs almost nothing, since the patient's body already tolerates that exact dose, while reviewing a new or changed dose after it's been taken costs something that can't be undone once swallowed. Cost asymmetry is that a dispensing delay is cheap and visible, while an already-ingested wrong dose is the hidden, permanent one. Kill criteria: if a medication change could be safely reversed by simply not taking the next dose, on-the-loop might be defensible even there.

Hand sketched metaphor scene reused for the pharmacy setting, representing the same gate versus watch distinction applied to a dosage decision instead of a payment.
Swap "transfer" for "dose." The same gate-versus-watch question decides the same way.
Dispensing errors caught before versus after the patient took the dose, by refill type
Refill type, stable to changed, left to right Harm already done when caught Stable refills New dose, new drug
Same review team, same sampling rate. The refill type alone predicts whether catching it after even means anything.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "in-the-loop blocks, on-the-loop watches, and the real test is whether tomorrow can undo today's mistake," and stop.
Cost: there's no engineering time to build a reversal-aware routing system this quarter. Say so honestly, and start by manually reclassifying any payment type with no chargeback right into the blocking tier.
The model gets better, for real: even a much more accurate fraud model doesn't change which mode a payment type needs, since the split was never about model accuracy, it was about whether the harm can be undone.

Where people run it wrong.
They pick the mode by how risky an action feels, rather than by whether it can actually be reversed once flagged.
They copy a threshold from an older product onto a new one without re-checking whether the new product even has a way back.
They treat "we caught it eventually" as equivalent to "we stopped it," when for an irreversible action those are two very different outcomes.

How to use it live. When someone asks for the difference between the two modes, answer with the reversibility question first, then the definitions. It's a stronger answer, and it's the one that actually decides real cases.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an "explain the difference and pick one" question?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Commit to a position, then show the asymmetry behind it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Beatriz Salcedo, a fraud analyst at Meridian Payments, five years in, known for tracing a card-dispute wave to a single compromised terminal.
3 · POSITION
What's the one-line answer to "which mode do I need"?
Tap to flip
ANSWER
In-the-loop for anything that can't be undone once it happens. On-the-loop for anything that still can.
4 · COST ASYMMETRY
Which error is actually the expensive one here?
Tap to flip
ANSWER
A completed, irreversible fraud loss. It's hidden until review and permanent by the time anyone sees it, unlike blocking friction, which is cheap and visible right away.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Sorting transactions into blocking and review-after modes purely by dollar amount, a rule inherited from card-only days that never got re-checked for a payment type with no reversal right.
6 · THE NUMBER
Fill in the blank: card fraud recovered 78 percent of the time. Instant transfer fraud recovered ___ percent.
Tap to flip
ANSWER
2 percent. Same review process both times, entirely different outcome once the money had already moved.
7 · THE REPLAY
Same fraud ring, redesigned split. What changes?
Tap to flip
ANSWER
The pattern gets caught at the third or fourth transfer, blocked before completion, instead of thousands deep before a routine audit connects the dots.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what decides the pick there?
Tap to flip
ANSWER
Fernglade TeleRx's prescription refill system. The pick turns on whether a dose has already been taken, not on how large or small the change looks.

Check yourself Score: 0 / 0

Short answer, apply it yourself
1. In your own words, what's the actual test for whether an action needs human-in-the-loop instead of human-on-the-loop?
Show hint
Look at the direct answer and the Position step.
Show answer
Model answer: Whether the action can still be undone once a person actually looks at it. If the harm is already permanent by review time, on-the-loop isn't real oversight, it's a record of what already happened.
Multiple choice
2. Why did sorting transactions by dollar amount fail for instant transfers?
  • A. Instant transfers are always larger than card transactions.
  • B. Sentry Flag couldn't score instant transfers at all.
  • C. Instant transfers have no chargeback right, so a dollar threshold built around card reversibility didn't apply to them.
  • D. Beatriz didn't have access to the instant transfer queue.
Show hint
Look at the story's knowledge spark on chargeback rights.
Show answer
C. The dollar rule was reasonable for a payment type with a built-in undo button. It silently stopped being reasonable the moment a payment type without one launched under the same rule.
True or false
3. True or false: human-on-the-loop is simply a weaker, less safe version of human-in-the-loop, and should be avoided wherever possible.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. For actions with a real reversal path, like card purchases with chargeback rights, on-the-loop is the right call. Blocking those anyway would just add friction with no matching safety gain.
Fill in the blank
4. Fill in the blank: Sentry Flag scores a transaction for fraud risk in about ___ milliseconds.
Show hint
Look at the opening of Section 1.
Show answer
40 milliseconds. Fast enough that speed was never the actual constraint on the design. Reversibility was.
Short answer, where it wouldn't matter
5. Name a transaction type in Meridian's system where human-on-the-loop is genuinely the right call, no matter how large the amount gets.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Routine card purchases. Chargeback rights already give them a real reversal path, so on-the-loop monitoring is the right fit regardless of dollar amount.
Short answer, the number question
6. If instant transfers had recovered fraud losses at 60 percent instead of 2 percent under the old system, would the same fix still be worth making? Why or why not?
Show hint
Look at the Kill criteria step.
Show answer
Model answer: Probably still worth it, but less urgently. The core test is reversibility, not a specific recovery percentage, so any meaningfully low recovery rate on an irreversible product still argues for moving it to human-in-the-loop.
Before you close the answer
Why this works
Tests whether you understand the two modes as a reversibility question, not just two names for "how much oversight."
Follow-up traps
"Why not just put everything in-the-loop to be safe?" Response: blocking everything burns reviewer attention on cases with a real reversal path anyway, and a reviewer facing that load starts approving fast, which defeats the point of the gate.

"Isn't reversibility itself sometimes hard to know in advance?" Response: it's still more knowable than confidence scores, since it's a property of the payment type and its legal reversal rights, not a guess about how the model feels about a specific case.
If pressed
Meridian's actual fix also added a short delay window, about ninety seconds, on newly launched payment types by default, so any product ships human-in-the-loop until its real reversal properties are confirmed, not assumed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more