CaseAdvancedDesigning for Uncertainty & Trust / Human-in-the-loop product design / #5

How do you decide which cases get routed to a human?

ORDER the queue that ranked the wrong thing first

Union Trust Bank runs WatchLine, a tool that scores every transaction for money-laundering risk and builds the day's alert queue. Folasade Adeyemi is a financial crimes investigator who reviews the alerts that make it through, out of roughly 50,000 scored every month.

The direct answer
Route to a human by asking four questions in order, not by score alone: is this pattern genuinely new, is the underlying action hard to undo, does a fixed legal reporting rule require it regardless, and only then, does capacity allow reviewing it anyway. A pattern nobody has seen before deserves a human's eyes precisely because the model has nothing to compare it to, whether or not it scores as alarming.
Do this, in order
  1. Route any alert flagged as a genuinely new pattern, no matter what its score says.Why: a score reflects how much an alert resembles past caught cases. A pattern built to look like nothing you've seen won't score high, by design.
  2. Route anything tied to an action that's hard to undo once it clears.Why: a wire about to leave the country today is a different kind of urgent than a domestic transfer sitting in a multi-day hold.
  3. Route anything a fixed legal reporting threshold requires, as a floor, not a suggestion.Why: some routing decisions aren't judgment calls at all, they're compliance requirements that don't move with the model's opinion.
  4. Only after those three, sort what's left by score, limited by real investigator capacity.Why: score is still useful, just not first. It's the tiebreaker among alerts that already cleared the more important bars.
  5. Sample auto-cleared alerts months later and check how many turned out real.Why: this is the only way to know whether the ranking is still catching what matters, before a regulator finds out for you.

How to answer this, stage by stage

Nobody is grading whether you can name a scoring model. They're grading whether your ranking survives the one case built specifically to duck it.

Stage 1
Ground it in one queue, one constraint
Say it like this
"I'll use a bank's financial crimes unit, 50,000 alerts a month, and investigator capacity for maybe 2,000 of them. The routing decision is the whole game."
Why this works
Names the real constraint, capacity, before proposing any ranking at all.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what routing is actually trying to protect. Reversibility, which mistake is hardest to undo. Dependency, what has to exist before what. Evidence, what to check cheaply first. Rank, the actual order."
Why this works
Signals a method for ranking by consequence, not a list built from gut feel.
Stage 3
Name the real outcome
Say it like this
"It's not clearing the queue fast. It's catching genuine laundering before funds leave the jurisdiction, without burying real cases under false positives so deep nobody reaches them either."
Why this works
Without this, ranking is just opinion dressed up as a process.
Stage 4
Give the actual rank, not just a list
Say it like this
"First, anything genuinely novel, whatever its score. Second, anything hard to undo. Third, anything a legal floor requires. Only then, score, limited by capacity."
Why this works
This is the direct answer, and it's ordered by what breaks first if skipped, not by convenience.
Stage 5
Name the decision that got this wrong before
Say it like this
"Union Trust originally routed by score percentile alone, top four percent to a human. That worked when the score tracked real risk. It broke the day someone built a pattern specifically to not resemble anything the model had learned."
Why this works
Names a real, specific design choice, not a vague "the model missed it."
Stage 6
Name the dependency, honestly
Say it like this
"You can't detect 'novel' until you've built a library of known patterns to compare against. That had to exist first, which is why this fix took months, not a config change."
Why this works
Shows the ranking wasn't just decided, it was sequenced against a real technical dependency.
Stage 7
Say what wouldn't change
Say it like this
"Routine, low-score, well-understood transaction types don't need this. The fix is for the case built to look ordinary, not for making every alert suspicious."
Why this works
Shows judgment, not a blanket rule that would drown the queue in noise.
Stage 8
Close on the one line
Say it like this
"A score tells you how much something resembles what you've already caught. Routing on score alone guarantees you'll miss whatever was built not to resemble it."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Picture fifty thousand transactions scored every month, and room to actually look closely at maybe two thousand of them.

WatchLine reads every wire transfer and large transaction Union Trust processes and scores it for money-laundering risk. Before it, a fixed-rule engine flagged around 8,000 transactions a month against simple triggers, round-dollar amounts, sanctioned-country lists, and investigators cleared 92 percent of those as false positives, exhausted by volume that mostly went nowhere.

Hand sketched flow diagram titled What unblocks what. Four boxes: Build pattern library, Build novelty scorer, Add override rule, Route the novel ones.
Each box had to exist before the next one could. The fix wasn't a switch, it was a sequence.

Now WatchLine scores all 50,000 alerts and routes only the riskiest slice, about 2,000 a month, to a human, a queue Folasade's team can actually work through carefully instead of skimming.

Here's the turn: the false positives that stayed in the queue were never the real danger. The danger was a genuinely new laundering pattern, built specifically to not resemble anything WatchLine had ever flagged, scoring calm and ordinary precisely because nothing in its training data looked like it either.

Hand sketched comparison titled Reversible or not. Left, a green box icon labeled Reversible, caption swings both ways. Right, a rust box icon labeled Irreversible, caption bolted shut.
Score alone can't tell these two apart. Only asking about the action itself can.
Confirmed laundering cases actually caught, score-only routing versus score-plus-novelty routing
21 10 0 8 Score-only routing 19 Score plus novelty
Same 21 confirmed cases, same investigators. The only thing that changed was what earned a case a place in the queue.

At its worst, a genuinely new laundering structure runs for months underneath a queue that only ever asked "does this look like trouble we've seen before," and the funds are gone, overseas, well before any human ever looks at the case.

The decision I would take back We routed to a human by score percentile alone, top four percent, everything else auto-cleared. That made sense when the score tracked real risk closely in every backtest we ran. It stopped making sense the day someone built a transaction pattern specifically to avoid resembling anything the model had learned to flag.

What I would leave alone: routine, well-understood transaction types, common wire patterns the model has scored thousands of times with a strong track record, don't need this override. The fix targets the case built to look ordinary, not every alert equally.

The lesson: a ranking built entirely on resemblance to the past will always have a blind spot shaped exactly like whatever hasn't happened yet.

Now here is the same thing as a story

The short version above is what you'd say defending this routing change to Union Trust's compliance committee. Read this one for how the gap actually got found.

Folasade Adeyemi has investigated financial crime at Union Trust for seven years, and she is the one who once caught a structuring pattern, several deposits each kept just under a reporting threshold, that a junior analyst had cleared as unrelated coincidence.

Knowledge spark: why would a genuinely risky transaction score low? A risk score is built by learning what past caught cases looked like. A pattern deliberately designed to avoid resembling any of them won't trigger the features the model learned to weight, so it can score as calm as a routine transaction, not because it's actually safe, but because nothing about it matches what the model was ever shown.

Under the old routing rule, WatchLine's top four percent by score reached Folasade's team every month, and for most of a year, that rule caught real cases at a rate the compliance team was proud of. The rule wasn't wrong. It was just answering only one question.

Hand sketched decision tree titled Does this alert route to a human. Root: New alert scored. Four branches: known pattern low score leads to Auto-clear, novel pattern any score leads to Route to human, meets legal reporting floor leads to Route to human, high score known pattern leads to Route if capacity allows.
Score alone only answers the last branch on this tree. It was being asked to answer all four.

A new laundering ring structured its transactions around what the model had already learned to consider safe, small, varied amounts, unusual but plausible-sounding business justifications, timing spread out enough to avoid any pattern WatchLine had seen before. Every one of those transactions scored in the bottom half of the queue, alongside thousands of genuinely routine ones.

WatchLine never missed the pattern by accident. The pattern was built, on purpose, to be the kind of thing WatchLine had no reason to doubt.

Nobody caught it during the run itself. It surfaced when a partner bank in another jurisdiction flagged a linked account during its own routine review and reached out to Union Trust to compare notes, revealing months of transactions that had cleared Folasade's queue without ever reaching it.

Hand sketched timeline titled The rollout, timed. Four milestones: Pattern library ships month 1, Novelty scorer trained month 2, Override rule ships highlighted month 3, First novel case caught month 4.
The fix took three real months to build, in order, because each step genuinely needed the one before it.

With the redesigned routing, an alert now routes to a human if it's flagged as statistically unlike anything in Union Trust's known-pattern library, independent of its score, alongside anything tied to a fast-moving, hard-to-reverse transfer, and anything a fixed legal threshold already requires. Run the same laundering pattern forward: its unfamiliarity, the exact thing that let it score low, is now the reason it routes to a human within days, not the reason it stays hidden for months.

Hand sketched icon list titled What earns a route to a human. Four items: a question mark box icon labeled Pattern never seen before, a gauge icon labeled Action is hard to undo, a scale icon labeled Meets a fixed legal floor, a box icon labeled Score if capacity allows.
Score is still on the list. It's just fourth, not first.

The old queue asked how alarming something looked. The new one also asks how unfamiliar it is, and stopped assuming those were the same question.

I built the routing rule around score percentile because it backtested beautifully against every case we already knew about. It took a call from another bank's compliance team, not our own audit, to see that "backtests well" and "catches what hasn't happened yet" were never the same guarantee.

ORDER, the four questions that decide the queueNot a scoring debate. ORDER is what tells you score should never go first.

O
Outcome. What routing is actually protecting.
Catching real laundering before funds leave the jurisdiction, without burying real cases under noise nobody has time to reach.
Without a stated outcome, any ranking is just opinion.
R
Reversibility. Which mistake is hardest to undo.
A wire that clears internationally today can't be recalled tomorrow. A domestic hold, or an alert that simply waits another day, still can be.
The hardest step, and the direct answer's second bar: rank by what can't be taken back.
D
Dependency. What has to exist before what.
A novelty override needs a known-pattern library to compare against. That library had to ship before "route anything unfamiliar" could mean anything at all.
Shows the fix was sequenced by a real technical dependency, not just decided.
E
Evidence. What to check cheaply first.
A retrospective sample of past confirmed cases against the proposed novelty rule, run before committing months of engineering to build it.
Confirms the fix is worth building before the full cost is spent.
R
Rank. The actual order, defended.
Novel pattern first, hard-to-reverse action second, legal reporting floor third, score fourth, limited by capacity.
A ranking you can argue with, not a vague sense of priority.

The recap, one line per letter: outcome is catching real laundering without drowning in false positives, reversibility is ranking the wire that can't be recalled above the alert that can wait, dependency is the pattern library that had to exist before novelty detection could, evidence is the cheap retrospective check before the full build, and rank is novelty, reversibility, legal floor, then score.

And if you want to be sure it really works, try it somewhere elseSame five letters, an airline's baggage claims desk instead of a bank's crimes unit. A different kind of irreversible.

Falconbridge Airways scores mishandled-baggage claims for likely fraud or error, deciding which claims auto-pay and which route to Rowan Achterberg, a baggage claims agent, for a manual look before any payout goes out.

Mapped onto ORDER: outcome is paying legitimate claims fast without funding serial fraud. Reversibility is that a payout issued as an instant digital gift card is effectively gone the moment it's redeemed, while a claim still pending review can be adjusted freely. Dependency is that detecting a repeat-claimant pattern needs a claimant history database, which had to exist before that rule could route anything at all. Evidence is a cheap backtest checking how many past confirmed fraud cases a repeat-claimant rule would actually have caught. Rank: repeat-claimant pattern first, irreversible payout method second, legal or contractual claim deadlines third, score fourth.

Hand sketched flow diagram reused for the airline setting, representing the same dependency shape: a claimant history has to exist before a repeat-pattern rule can route anything.
Swap "laundering pattern" for "repeat claimant," and the same dependency chain holds up in a different building.
Claim risk score versus payout method, confirmed fraud cases at Falconbridge Airways
Fraud risk score, low to high, left to right Bank transfer payouts Gift card payouts
The riskiest payout method, not the highest score, marks where the real blind spot sits.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "route by novelty and reversibility before score, since score only measures resemblance to what you've already caught," and stop.
Cost: there's no engineering budget to build a novelty scorer this quarter. Say so honestly, and start with a manual flag for any transaction type or claim type nobody has categorized before, a cheap proxy while the real system gets built.
The model gets better, for real: even a much more accurate score doesn't fix this on its own, since accuracy is measured against cases you've already seen, and the blind spot lives specifically in the cases you haven't.

Where people run it wrong.
They treat the model's score as the entire routing decision, instead of one input among several.
They build novelty detection last, as a nice-to-have, when it's actually the piece that catches what everything else is designed to miss.
They wait for an external tip, a partner bank, a regulator, a customer complaint, to reveal the gap, instead of sampling their own auto-cleared cases on a schedule.

How to use it live. When someone asks how to route cases to a human, ask what a case would look like that was built specifically to score low. If your ranking has no answer to that, score is doing too much of the work.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "which cases get routed to a human" question?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Rank by what's hardest to undo, not by convenience.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Folasade Adeyemi, a financial crimes investigator at Union Trust Bank, seven years in, known for catching a structuring pattern a junior analyst missed.
3 · OUTCOME
What is the routing decision actually trying to protect?
Tap to flip
ANSWER
Catching genuine laundering before funds leave the jurisdiction, without burying real cases so deep under false positives that they never get reached either.
4 · THE RANK
What's the actual order, first to last?
Tap to flip
ANSWER
Genuinely novel pattern first, hard-to-reverse action second, fixed legal reporting floor third, score fourth, limited by capacity.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Routing to a human by score percentile alone, which worked while the score tracked real risk and broke once a pattern was built specifically not to resemble anything the model had learned.
6 · THE NUMBER
Fill in the blank: of 21 confirmed cases, score-only routing would have caught only ___.
Tap to flip
ANSWER
8. Adding a novelty-based override raised that to 19 of the same 21 cases.
7 · THE REPLAY
Same laundering pattern, redesigned routing. What changes?
Tap to flip
ANSWER
Its unfamiliarity, the exact thing that let it score low, now routes it to a human within days instead of letting it run for months undiscovered.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what ranks first there?
Tap to flip
ANSWER
Falconbridge Airways' baggage claims desk. A repeat-claimant pattern ranks first, ahead of score, the same way novelty ranked first at the bank.

Check yourself Score: 0 / 0

Multiple choice
1. Why did the laundering pattern in this story score low instead of high?
  • A. WatchLine had a software bug that affected scoring.
  • B. It was built specifically to avoid resembling any pattern the model had learned to flag as risky.
  • C. Folasade's team manually lowered its score by mistake.
  • D. The transaction amounts were too small for WatchLine to score at all.
Show hint
Look at the knowledge spark on why a risky transaction can score low.
Show answer
B. A score measures resemblance to past caught cases. A pattern designed to avoid that resemblance will score calm by construction, not by accident.
True or false
2. True or false: after this fix, score no longer matters at all in deciding which alerts route to a human.
  • True
  • False
Show hint
Look at the Rank step.
Show answer
False. Score still matters, it's just fourth in the order, a tiebreaker among alerts that already cleared the novelty, reversibility, and legal-floor bars.
Fill in the blank
3. Fill in the blank: under the old fixed-rule engine, investigators cleared ___ percent of flagged alerts as false positives.
Show hint
Look at the opening of Section 1.
Show answer
92 percent. Most of the old queue's volume went nowhere, which is exactly the kind of exhaustion WatchLine's scoring was built to reduce.
Short answer, apply it yourself
4. Think of a system you know that ranks items by a risk or relevance score, a spam filter, a fraud check, a content moderation queue. What would a case built specifically to score low, on purpose, look like there?
Show hint
Think about what "designed to look ordinary" means in that specific system.
Show answer
Model answer: In a spam filter, it might be a message written in plain, personal language with no typical spam markers at all, precisely because it was crafted to avoid every signal the filter learned to catch.
Short answer, where it wouldn't matter
5. Name a transaction type at Union Trust where the old score-only routing rule is still fine, unchanged.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Routine, well-understood wire patterns the model has scored thousands of times with a strong track record. The fix targets the case built to look ordinary, not every alert equally.
Short answer, the number question
6. If Union Trust only had capacity to review 200 alerts a month instead of 2,000, would the same four-question rank still make sense? Why or why not?
Show hint
Look at the Rank step and how score fits in as the last, capacity-limited factor.
Show answer
Model answer: Yes, the order itself wouldn't change, but the score-based fourth tier would shrink sharply, since capacity is what decides how deep into the score-ranked leftovers a human actually reaches.
Before you close the answer
Why this works
Tests whether you understand that a risk score is a measure of resemblance to the past, and can name what that guarantees you'll miss.
Follow-up traps
"Couldn't you just lower the score threshold to catch more cases?" Response: a lower threshold still only catches things that resemble past cases more loosely, it doesn't help with a pattern built to resemble nothing at all.

"Won't a novelty rule flood the queue with harmless new-but-normal transactions?" Response: it's paired with the pattern library specifically so "novel" means statistically unlike anything seen, not just new to this customer, which keeps the volume manageable.
If pressed
Union Trust's actual fix also logs which novelty-routed alerts get cleared in under sixty seconds, since a fast clearance on a genuinely unfamiliar pattern is a sign the reviewer is treating the override as noise rather than the signal it's meant to be.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more