CaseAdvancedDesigning for Uncertainty & Trust / Human-in-the-loop product design / #7

How do you measure whether the human in the loop is adding value?

LEADthe number that rings before a stroke patient misses the window

Kestrel Regional Stroke Center runs an AI tool over every incoming CT scan, flagging which ones show signs of a large-vessel occlusion, a blockage that needs a specialist within hours. Ingunn Marsalis is the radiographer who does the second look on every flagged scan before it goes to the on-call neurologist.

The direct answer
Measure the human's value by tracking how often the reviewer's decision genuinely changes the outcome, not by counting reviews completed. Track override rate on the AI's calls, split by scan type, as a trend over time. A rate that falls toward zero without any real improvement in the model underneath it is the first sign the reviewer has stopped adding anything at all.
Do this, in order
  1. Track the reviewer's override rate on the AI's calls, by scan type, as a trend.Why: a falling override rate with no matching gain in model accuracy means the person has quietly stopped catching anything the model would have missed.
  2. Pair that rate with an independent, sampled re-read by a second radiologist.Why: a low override rate could mean the model is excellent, or it could mean nobody is really checking. Only an outside sample tells the two apart.
  3. Track how long the reasoning panel stays open before a decision gets made.Why: a decision made in under two seconds on a case the model itself flagged as uncertain is a decision that was never really reviewed.
  4. Compare cases the reviewer changed against cases later confirmed correct or wrong.Why: value isn't overriding often, it's overriding the right ones. Count how many of those overrides were later proven necessary.
  5. Set a floor: if override rate on any scan type drops under a quarter of its baseline, trigger a mandatory audit that week.Why: a number nobody acts on is decoration. This turns the leading signal into an actual decision.
  6. Leave alone scan types where a low override rate has always held up under audit.Why: some scan categories really are close to unambiguous, and a low override rate there is a fact about the case, not a warning about the person.

How to answer this, stage by stage

The interviewer isn't grading whether you can name a metric. They're grading whether the metric you name would catch a reviewer who has stopped reviewing, before a patient pays for it.

Stage 1
Ground it in one review seat
Say it like this
"I'll ground this in a real seat, a stroke center where an AI tool flags CT scans for a possible large-vessel occlusion, and a radiographer does the second look before it reaches a neurologist."
Why this works
Stops the answer from turning into a generic list of dashboard metrics nobody would actually build.
Stage 2
Reframe what the question is really asking
Say it like this
"This isn't 'is the human present.' It's 'is the human's presence still changing anything.' Those are different questions, and most teams only measure the first one."
Why this works
Separates a review step that exists from a review step that does real work.
Stage 3
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the outcome that actually matters. Early signal, what moves before that outcome does. Abuse, how the metric gets gamed. Decision, what I'd do at each level."
Why this works
Signals a method for finding a leading indicator, not a guess at a dashboard tile.
Stage 4
Give the one measure that answers the question
Say it like this
"Override rate on flagged scans, by scan type, tracked monthly. Ours started around 14 percent and drifted under 2 percent in three months, with no matching jump in the model's own accuracy."
Why this works
This is the direct answer, and the hardest step in LEAD, spoken as one clean number.
Stage 5
Name how that number lies
Say it like this
"A falling override rate can mean genuine trust in a better model, or it can mean a tired reviewer clicking through. I'd pair it with a small, independent, sampled re-read, since that's the only thing that tells those two apart."
Why this works
Shows you know a metric can be hit without the underlying work happening.
Stage 6
Prove it with the six months nobody watched
Say it like this
"The override rate crossed our warning floor in month three. The peer audit that actually found a missed case didn't happen until month six. That's three months this number sat there, unused."
Why this works
Makes the "leading indicator" claim checkable, not asserted.
Stage 7
Say what you'd leave alone
Say it like this
"For scan types where audits confirm the override rate has always been low and correct, I wouldn't treat a low number there as a red flag. Some cases really are close calls without a close call."
Why this works
Shows judgment instead of blanket suspicion of every low number.
Stage 8
Close on the one line
Say it like this
"The human in the loop isn't adding value because they're in the loop. They're adding value on the days the loop actually changes the outcome, and that's the number worth watching."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Picture a reading room with forty flagged CT scans a day, and one number everyone upstream watches: how fast each scan gets a final read.

Kestrel's stroke tool reads an incoming CT scan and flags it as likely showing a large-vessel occlusion, needing a specialist fast, or clear, safe to route normally. Before the tool, every scan waited in a queue for a radiographer to read it cold, slower, but never once skipped.

Hand sketched flow diagram titled How a flagged scan gets reviewed. Four boxes: AI flags scan, Ingunn opens viewer, Reads the reasoning highlighted, Confirms or overrides.
The whole point of the second look lives in that third box. Everything upstream of it is just routing.

Now most clear scans move on in minutes, and Ingunn still opens every flagged one, checking the AI's marked region against what she sees herself, free to override whenever the reasoning doesn't hold up.

Here's the turn: the tool wasn't the problem. Nobody was watching whether Ingunn's second look was still a real second look, only whether flagged scans kept moving. A three-second glance and an eight-minute read look identical on a throughput dashboard.

Override rate on flagged large-vessel-occlusion scans, by month
16% 8 0 4%, warning floor Month 1, 14% Month 3, crosses floor Month 6
Nobody was tracking this line on its own. If they had been, a warning would have fired three months before the peer audit did.

At its worst, an entire scan category gets a rubber-stamped second look for months, and the first anyone learns of it is a peer-review audit finding a delayed diagnosis.

Hand sketched comparison titled Which one rings first. Left, a blue gauge icon labeled Override rate, caption moves in week 3. Right, a red document icon labeled Missed case found, caption found in month 6.
Same underlying problem, two very different clocks. One of them had a five-month head start nobody used.
The decision I would take back Kestrel never built a per-scan-type override-rate trend into the reading room dashboard, only a running count of scans cleared per shift. That made sense at launch, when the point was proving the tool could keep up with volume. It stopped making sense once the tool had been running long enough for a reviewer's habits to quietly change underneath a healthy-looking throughput number.

What I would leave alone: clean scans with no ambiguous shadowing have always had a low override rate, and repeat audits confirm that's correct. A low number there is a fact about the case, not a warning sign about Ingunn.

The lesson: a number that only counts finished reviews can't tell you when the reviewing itself quietly stopped.

Now here is the same thing as a story

The short version above is what you'd say defending this measurement plan to Kestrel's chief of radiology. Read this one for how the number actually crept down.

Ingunn Marsalis has read stroke scans for nine years, and she used to be the radiographer other shifts asked to double-check a borderline case, the one who could see a subtle occlusion nobody else had caught yet.

Knowledge spark: why would an override rate drift down on its own? When a model's flagged region keeps matching what a reviewer already sees, and it's right often enough early on, the reviewer's habit shifts from a full independent read to a quick confirm of the model's marked area. Nobody decides to stop reviewing. It happens a little at a time, and a throughput dashboard has no way to see the difference.

In her first months with the tool, Ingunn overrode about one in seven flagged scans, usually catching a small occlusion the model's marked region had missed by a few millimeters. Over the following months, as the model's highlighted regions got more consistently close to what she'd have marked herself, she found herself overriding less, then rarely, then almost never.

Hand sketched timeline titled Six months to the peer audit. Four milestones: Tool ships override 14 percent, Rate drifting month 2, Below floor month 3 highlighted, Peer audit month 6.
Three months sat between the number crossing its own warning line and anyone outside the reading room finding out why.

A locum radiographer covering her shift one week remarked, half-admiring, "you clear these fast, does the model just nail it every time now?" Ingunn laughed and said something like that. She hadn't actually timed how long she'd spent on each one lately.

Hand sketched quadrant titled Leading vs lagging signals. X axis how soon it appears, Y axis visible on todays dashboard. Override rate placed early and hidden. Scans per day placed early and obvious. Missed case found placed late and obvious. Second-look time placed early and hidden.
The signals worth building live in the hidden-and-early corner, exactly where today's dashboard never looks.

Six months in, a routine peer-review audit sampled twenty recent cleared scans and found one with a genuine small occlusion that should have triggered an urgent referral, the exact kind of case Ingunn used to catch without a second thought.

The patient wasn't missed because the model got worse. She was missed because nobody was watching the one number that would have shown, months earlier, that the second look had quietly stopped being a second look.

With the redesigned dashboard, a per-scan-type override-rate trend sits next to throughput, not instead of it, and any category falling below a quarter of its baseline rate triggers a mandatory sampled re-read that same week. Run the same six months forward: the month-three warning triggers a re-read, catches the drift, and a short refresher on what the model's confident-looking regions can still miss brings the override rate back to a level the audit confirms is genuinely earned.

The old dashboard asked "are scans getting cleared." The new one asks "is the second look still a second look."

I built the throughput dashboard first because it was the easiest thing to show the board was working. It took a peer-review audit to see that "fast" and "still being checked" were never the same claim.

LEAD, in one screenNot a lecture on picking a north star metric. LEAD is what tells you which number rings first.

L
Link. The outcome that actually matters.
Correct, timely referrals for real occlusions, not raw scans cleared per shift.
Separates the real clinical outcome from the number that's easiest to count.
E
Early signal. What moves first.
Override rate on flagged scans, by scan type, tracked over time. It fell below a warning floor three months before a peer audit found a missed case.
The hardest step, and the direct answer to the question.
A
Abuse. How the metric could be gamed.
A low override rate can mean real trust in a better model, or a reviewer clicking through tired shifts. An independent, sampled re-read is the only way to tell them apart.
Shows the metric was chosen with an eye on how people respond to being measured.
D
Decision. What you'd actually do.
Below a quarter of a scan type's baseline override rate, trigger a mandatory sampled re-read that same week, not a note in next quarter's review.
Turns the number into an action, not a dashboard decoration.
Hand sketched icon list titled Four things worth tracking. Four items: a gauge icon labeled Override rate by scan type, a scale icon labeled Confidence vs actual match, a document icon labeled Reasoning panel opened, a question mark box icon labeled Peer audit disagreement.
Four numbers. Only one of them would have rung early on its own, which is why override rate leads the list.

The recap, one line per letter: link is correct, timely referrals rather than raw throughput, early signal is override rate falling below its baseline, abuse is a rubber-stamped rate looking identical to a genuinely earned one without an outside check, and decision is a mandatory sampled re-read the moment the floor is crossed.

Hand sketched labeled parts diagram titled What the review dashboard needs. Center gauge icon labeled Review Dashboard, with four callouts: override trend, calibration drift, open-rate, audit gap.
None of these four require reading a single scan yourself. They just require watching the checking, not just the throughput.

And if you want to be sure it really works, try it somewhere elseSame four letters, a marketplace trust and safety queue instead of a reading room. A completely different kind of checking, the same quiet drift.

Palisade Marketplace uses an AI tool that flags seller listings likely to be counterfeit goods, so a trust and safety analyst doesn't have to open every new listing cold. Ingrid Castellanos reviews flagged listings most mornings and used to remove or request proof on roughly one in five of the AI's flagged items.

Mapped onto LEAD: link is genuine counterfeit removals that hold up on appeal, not listings reviewed per hour; early signal is Ingrid's override rate, which had quietly fallen from one in five to nearly zero as the AI's flagged reasons started reading as more specific and confident, whether or not the underlying detection model had actually improved that much; abuse is that a falling override rate could mean the model genuinely got sharper, or that Ingrid, under a growing review queue, started clearing flagged listings faster than she was reading the evidence behind them.

Hand sketched flow diagram titled Palisades listing review chain. Four boxes: AI flags listing, Ingrid opens case, Reviews the evidence highlighted, Removes or clears.
Swap "scan" for "listing," and the same missing number shows up in a completely different queue.
Counterfeit listings later reinstated on appeal, quarterly sample, before and after tracking override rate
25 12 0 21 Before tracking 3 After tracking
Same quarterly sample size both times. Watching the override rate closed most of the gap without touching the model at all.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "track the reviewer's override rate over time, by category, since a falling one means checking has quietly stopped, weeks before a bad outcome ever surfaces," and stop.
Cost: there's no budget for a full audit team this year. Say so honestly, and start with a small, cheap monthly sample on whichever category has the steepest override-rate drop, since that's where the risk is concentrated.
The model gets better, for real: if the underlying model genuinely improves, override rate should still fall, but slowly and evenly, not in the sudden, sustained way a real behavior change looks like. The shape of the drop, not just its direction, tells you which one happened.

Where people run it wrong.
They build a dashboard that measures throughput and call the human-in-the-loop question answered, since throughput is the easiest number to show a stakeholder.
They treat any low override rate as unambiguous good news, without ever checking whether it reflects trust or fatigue.
They wait for a downstream audit to tell them something's wrong, when the whole point of a leading indicator is to ring before that.

How to use it live. When someone asks how you'd measure the human's value, ask one question first: what number would have moved weeks before the harm became visible to anyone outside the room. Build the measurement around that number, not the one that's easiest to screenshot for a board deck.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "how do you measure this" question?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Find the number that moves before the real outcome does.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ingunn Marsalis, a radiographer at Kestrel Regional Stroke Center who has read stroke scans for nine years.
3 · THE LINK
What's the outcome that actually matters here?
Tap to flip
ANSWER
Correct, timely referrals for real occlusions, not raw scans cleared per shift.
4 · THE EARLY SIGNAL
What number moved first, and how far ahead of the real problem?
Tap to flip
ANSWER
Override rate, which crossed its warning floor in month three, three months before a peer audit found the missed occlusion in month six.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never building a per-scan-type override-rate trend into the dashboard, only tracking raw throughput, which made sense at launch and stopped making sense once habits had time to quietly shift.
6 · THE NUMBER
Fill in the blank: override rate started at 14 percent and fell to about ___ percent by month six.
Tap to flip
ANSWER
2 percent. A seven-fold drop that the throughput dashboard never showed at all.
7 · THE REPLAY
Same six months, redesigned dashboard. What changes?
Tap to flip
ANSWER
The month-three warning triggers a mandatory sampled re-read, catches the drift, and a short refresher brings the override rate back to a level the audit confirms is genuinely earned.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Palisade Marketplace's counterfeit-listing review queue. The early signal is Ingrid Castellanos's override rate falling from one in five toward zero.

Check yourself Score: 0 / 0

Multiple choice
1. Why is override rate a better measure of the human's value here than counting how many scans get reviewed?
  • A. Override rate is easier for a hospital board to understand.
  • B. Reviews completed can happen without any real checking, while override rate reflects whether the review is actually catching something.
  • C. Counting reviews is against hospital policy.
  • D. Override rate is the only number the software can log.
Show hint
Look at the reframe in Stage 2.
Show answer
B. A completed review can be a real second look or a three-second click-through. Only override rate, and what happens to those overrides later, tells the difference.
True or false
2. True or false: a falling override rate always means the AI model has genuinely gotten more accurate.
  • True
  • False
Show hint
Look at the Abuse step.
Show answer
False. A falling override rate can also mean the reviewer is checking less closely. Only an independent, sampled re-read can tell the two apart.
Fill in the blank
3. Fill in the blank: the peer audit sampled twenty recent cleared scans and found ___ genuine occlusion that should have triggered an urgent referral.
Show hint
Look at the story's audit paragraph.
Show answer
One. A single missed case out of twenty, exactly the kind Ingunn used to catch without a second thought before the override rate drifted down.
Short answer, apply it yourself
4. Think of a review step you've seen in any product, a moderator, an editor, an approver. What number would have caught them clicking through faster than they used to?
Show hint
Think about approval queues, content moderation, or expense sign-offs you've watched or used.
Show answer
Model answer: Most people can name an approval or moderation step where the reviewer's override or rejection rate would have quietly dropped well before anyone noticed the reviewing had gotten shallow.
Short answer, where it wouldn't matter
5. Name a scan category at Kestrel where a low override rate isn't actually a warning sign.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Clean scans with no ambiguous shadowing. A low override rate there reflects a genuinely unambiguous case, confirmed by audit, not a checking problem.
Short answer, the number question
6. If the override rate had fallen slowly and evenly over two full years instead of sharply over three months, would the same warning-floor design still make sense? Why or why not?
Show hint
Look at the "swap the trigger" section's note about the shape of the drop.
Show answer
Model answer: Possibly not as a hard alert, since a slow, steady decline might reflect real, gradual model improvement rather than a reviewing habit collapsing. The shape of the drop matters as much as its direction.
Before you close the answer
Why this works
Tests whether you can name a metric that predicts a reviewer disengaging rather than one that only confirms it after a patient is harmed.
Follow-up traps
"Couldn't you just track the model's own confidence score instead of the human's override rate?" Response: confidence alone doesn't show whether a human is still genuinely checking it, and it's the checking, not the score, that catches a real mistake before it reaches a patient.

"Won't a hard override-rate floor just push reviewers to override more so they don't trigger an alert?" Response: pair the floor with the independent sampled re-read, so a reviewer padding overrides for appearance gets caught the same way a rubber-stamping reviewer would.
If pressed
Kestrel's actual fix also logs how long the reasoning panel stayed open before a decision, since a sub-two-second open-to-confirm time is a second, independent signal the override rate alone can't fully capture.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more