CalculationAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #5

How do you set a pass bar when human performance on the same task is 92 percent?

The direct answer
Don't copy the 92 percent onto the model. Split the bar into two numbers: the tool has to catch at least 9 of every 10 real misses the radiologist makes, while flagging no more than about 15 of every 100 scans overall, so the extra reviews stay workable. A missed finding and a wasted five-minute recheck are not the same size mistake, so they don't get the same bar.
Do this, in order
  1. Split the single 92 percent target into two numbers, a catch rate on real misses and a flag rate on clean scans.Why: they pull against each other, and one blended score hides which one you're actually protecting.
  2. Set the catch rate near total, at least 9 of every 10 real misses caught.Why: a missed finding on this task often gets no second chance, so this is the number that can't fail quietly.
  3. Cap the flag rate at what the review desk can actually absorb, around 15 out of every 100 scans.Why: a great catch rate that floods the queue gets ignored, the same way any over-eager tool does.
  4. Move the catch-rate target up or down based on whether a miss has another safety net, not on how rare the finding is.Why: whether a miss gets caught some other way swings the number more than anything else does.
  5. Never grade the tool on raw accuracy measured against 92 percent.Why: only 1 in 10 scans has a real finding here, so a tool that says "normal" every time already scores close to 90 percent while catching nothing.
  6. Recheck the flag-rate ceiling every time staffing on the review desk changes.Why: the same catch-rate target turns unworkable the moment one radiologist covers the desk instead of three.

How to answer this, stage by stage

Nobody's grading you on landing the exact right percentage here. They're grading whether you split the bar on purpose instead of copying one number off the page. Eight moves get you there.

1
Pin down exactly what a flag decides
Say it like this
"Before any numbers, here's what I'm actually pricing. A tool reviews chest CT scans after a radiologist has already read them, and flags any it thinks got missed, for a second look before the report goes final. That's the one decision this bar controls, not accuracy on every task the model might ever touch."
Why this works
Naming exactly what a flag triggers, before any numbers, stops the whole answer from drifting into a vague accuracy debate.
2
Say why 92 percent alone is the wrong target
Say it like this
"First, I want to say why I'm not just aiming to beat 92 percent. That's the radiologist's overall accuracy. A missed cancer and a needlessly flagged clean scan are both 'wrong' inside that one number, but they cost completely different amounts. One number can't hold two different costs."
Why this works
This is the reframe. It tells the interviewer you're not about to hand back a single blended score and call it done.
3
Split the bar into two numbers
Say it like this
"So I'm splitting it. Number one: how many of the radiologist's real misses does the tool catch. Number two: how many clean scans does it flag along the way. Those two trade off against each other on the same threshold, so I have to set them on purpose, not let one fall out of the other."
Why this works
This is the whole method in three sentences. Say it before a single number and the rest of the answer has somewhere to hang.
4
Own the numbers
Say it like this
"Here's my actual bar. Out of 1,000 scans, say 100 have a real finding. The radiologist alone catches 92 of those and misses 8. I want the tool to catch at least 9 of every 10 of those misses, call it 7 of the 8. And I want the total flag rate, real catches plus false alarms, to stay under 15 percent of all scans, so about 150 out of 1,000."
Why this works
Real numbers, not adjectives. A candidate who can't do this arithmetic out loud is guessing with a confident voice.
5
Give the range
Say it like this
"That's not one bar, it's a range. If the finding type has another safety net, say a routine scan already booked in three months, a lower catch rate, around 75 percent, is fine. If there's no other check coming, I'd push the catch rate to 97 percent or higher, even if that means flagging more clean scans than I'd like."
Why this works
A single number implies a confidence the estimator doesn't have. The range shows the bar depends on something real, not a guess.
6
Run the sanity check
Say it like this
"Sanity check: 92 percent sounds like an A grade, but in a hundred real findings that's 8 patients walking out with something missed. And a 15 percent flag rate on a 50-scan overnight shift is about 7 extra rechecks a night, which one radiologist can actually get through."
Why this works
This turns a percentage into something you can picture, a count of people or minutes, so it doesn't sound like a formula recited from memory.
7
Point at the one lever doing the real work
Say it like this
"One more thing before I close. If you pushed on which assumption breaks this the fastest, it isn't how common the finding is. It's whether a miss has a second chance somewhere downstream. That single fact swings the target catch rate more than prevalence or staffing put together."
Why this works
This is the line that separates a good estimator from someone reciting a formula. It names the one lever that actually matters.
8
Say the whole split back in one line
Say it like this
"Pull it together: not 92 percent. Catch at least 9 of every 10 real misses, keep the flag rate under about 15 percent, and move the catch-rate target based on whether a miss gets another chance. Two numbers, chosen by what each mistake costs."
Why this works
One breath, the whole answer, in case that's all the interviewer remembers.
If you remember one thing The catch-rate number and the flag-rate number protect two different things. Catch rate protects the patient. Flag rate protects the review desk. Neither one alone is the bar.

Let's learn

Out of every 100 scans that actually have something wrong on them, a radiologist reading alone catches 92. She misses 8. That's before any tool has touched the work. That's not a bad day. That's what a genuinely skilled person looks like on a task nobody does perfectly.

Now put a tool behind her. It looks at chest CT scans after she has already read them, and flags any scan it thinks she got wrong. A flagged scan gets a second look before the report goes out.

Knowledge spark: what does "92 percent" actually measure here? It's the share of real findings a reader catches on a scan, sometimes called sensitivity. 92 percent means 92 of every 100 real ones get caught. The other 8 slip through, not because anyone was careless, but because some findings are genuinely hard to see.

Now say we set the tool's pass bar the easy way: beat 92 percent. One number, matched to the human.

Knowledge spark: what is "prevalence"? How common the thing you're looking for actually is in the group being scanned. If 1 in 10 scans has a real finding, prevalence is 10 percent. It matters because a tool can score well on plain accuracy just by guessing the common answer, normal, every single time.

Here's the turn. That one number does not tell the tool what to protect. A missed finding leaves a patient with nothing, no treatment starts, and often the next scheduled check isn't for months. A wrongly flagged clean scan costs the radiologist five minutes rereading something that was fine. Grading both mistakes against the same 92 percent treats a life and a coffee break as the same size problem.

At its worst, that mistake compounds. Tune the flag threshold to "match 92 percent accuracy" and it can end up doing almost nothing extra, because "normal" is already the right call on 9 of every 10 scans, and saying nothing scores well on accuracy. The tool passes its own bar and catches almost none of the misses it was built to catch.

Where the numbers land, per 1,000 scans
Real findings (100 total)
92 radiologist
99 of 100 caught
Sent back for a second look (150 total)
143 false alarms
150 of 1,000 scans
Radiologist's own catch (92) AI adds from the misses (7) Still missed (1) False alarms (143)
The AI's real job is the small green slice in row one, seven more real findings caught. The price for that sits in row two: 143 clean scans flagged along the way, out of a 150-flag budget.
A tool graded on accuracy alone can pass its own bar while catching nothing.

The choice I would take back. We'd have set the target as "beat 92 percent overall," because it sounded fair, a bar as good as the best human in the room. I'd take that back. In a population where only 1 in 10 scans has a real finding, a tool that never flags anything already scores 90 percent. The bar has to come from the two real costs, not from one blended number.

What I would leave alone. Not every flag in this system needs this treatment. The tool also flags scans with a technical problem, bad positioning, motion blur, that need a retake. Missing one of those costs a same-day callback. A false flag there costs thirty seconds glancing at a fine image. Both sides are cheap and reversible. A simple bar, beat the scanner's own error rate, is fine there. Save the two-number treatment for mistakes that don't cost the same on both sides.

The lesson. A pass bar is not one number copied from a human's score. It's two numbers, one for each kind of mistake, sized by what that mistake actually costs, not by what sounds fair on a slide.

Now here is the same thing as a story

Skip this part if you already believe two mistakes can cost wildly different amounts. Read on if you don't.

Ingrid Voss runs clinical safety for Cascade Radiology Partners, a chest imaging group that reads overnight for a dozen hospitals. She's the one who signs off on any AI tool before it touches a live scan, and she has turned down more tools than she's approved.

The new one is a second-opinion flagger. It reads every scan behind the radiologist on shift and flags anything it thinks might have been missed.

When the vendor pitched it, their whole deck rested on one line: 94 percent accuracy, beats your radiologists' 92. Ingrid almost signed off on that alone. It was a bigger number. Bigger sounded safer.

She ran it against three months of the group's own scans anyway, before committing, out of habit more than suspicion.

The tool matched the pitch on accuracy. And it caught two of the nineteen real misses hidden in that quarter's data. Two out of nineteen.

The number that sounded like safety was the number that was almost meaningless.

She sat with that for a while. The tool wasn't broken. It had just learned, correctly, that "normal" is the right call on 9 of every 10 scans in this population, so calling almost everything normal was a cheap way to a good accuracy score. It only had to find something on the rare scan where the radiologist's own miss was hiding, and it mostly didn't bother.

A clinical safety lead at a workstation late at night, a monitor showing a scan with one flagged mark, and a notebook with two circled numbers, 90 percent and 15 percent
Before she signs off on anything, she writes down two numbers, not one.

So Ingrid threw out the single number. She sat down with the group's chief radiologist and asked two separate questions instead: of the misses we can find in old cases, how many would this have caught, and of the clean scans, how many would it needlessly send back.

Rerun against the same three months, tuned against those two numbers instead of one blended score: it caught 16 of the 19 misses, and flagged 142 clean scans out of the roughly 900 in the sample, about 16 percent. Worse-looking accuracy number than the vendor's pitch. Far better at the one job that mattered.

She signed off on that version. Not the higher-accuracy one.

The thing I'd want Ingrid to say out loud, if she got asked why in an interview: the vendor's number answered a question nobody was actually asking. The real question was never "is this tool usually right." It was "when it's wrong, which kind of wrong is it, and can the desk survive that."

Five letters, and where each one lands here

This is an estimation question with a cost asymmetry built in, so BOUND fits, not FLIPS. Nobody's habit is fading over months here. It's an argument about two numbers and what each one is allowed to cost.

Number line from 0 to 100 percent marking a 75 percent low bar, a 90 percent proposed bar, the 92 percent human baseline marked separately as a comparison point, and a 97 percent high bar
The range, with the human baseline marked for scale, not as one of the bounds

B, break it down. The bar isn't one score, it's two: a catch rate on the radiologist's real misses, and a flag rate on clean scans, because those two mistakes cost different amounts.
O, own the numbers. Out of 1,000 scans and 100 real findings, the radiologist catches 92 and misses 8. Target: catch at least 7 of those 8. Keep total flags, real catches plus false alarms, under about 150.
U, use a range. 75 percent catch rate when a miss has another safety net downstream. 97 percent or higher when it doesn't. The proposed 90 percent sits between those, closer to the strict end on purpose.
N, nail the sanity check. 92 percent human accuracy still means real patients missed every week at a busy site. A 15 percent flag rate is about 7 extra rechecks a night, which one radiologist can actually clear.
D, direction. Whether a miss has a downstream safety net moves the target catch rate more than prevalence, staffing, or anything else on the list.

What moves the catch-rate target most
Does the miss get a second chance downstream?+8 points
Overnight staffing: one radiologist instead of three−7 points
How common the finding is: 20% instead of 10% of scans+2 points
Minutes a recheck actually costs: 20 instead of 5−2 points
All four are read against the proposed 90 percent catch-rate target. Whether a miss gets caught some other way moves the number more than the other three combined, which is exactly the D step's point.

And if you want to be sure it really works, try it somewhere else

A contract-translation review tool works the same way. A translator drafts the target-language version, then a checking tool flags clauses it thinks got the meaning wrong, for a second translator to look at before the contract goes out.

B, break it down. Same split: how many of the first translator's real clause-level errors does the tool catch, and how many correct clauses does it needlessly flag for a second read.
O, own the numbers. Say a firm processes 200 contracts a month, about 40 load-bearing clauses per contract, 8,000 clauses total. A skilled translator gets the legal meaning right on about 89 percent of those on a first pass, so around 880 clauses a month carry a real error. Target: catch at least 90 percent of those, about 790, while flagging no more than 12 percent of all 8,000 clauses, about 960.
U, use a range. Boilerplate clauses, ones a lawyer reads anyway before signing, can sit at a lower catch rate, maybe 70 percent, since there's a safety net downstream. Liability and indemnity clauses, the ones relied on directly if a dispute ever happens, need a catch rate close to total, 95 percent or higher.
N, nail the sanity check. 89 percent sounds solid for a translator, but across 8,000 clauses a month that's close to 30 real errors a day going into contracts before any check runs. A 12 percent flag rate is about 32 clauses a day sent back, workable for a small review team.
D, direction. Same lever as the radiology case: whether a clause type gets read again downstream by someone else, a lawyer, a counterparty, moves the target catch rate more than the base error rate does.

The lever that generalizes In both the scans and the contracts, the same fact moves the bar most: whether the mistake gets a second chance somewhere downstream. Not how rare the mistake is, and not how good the human's own score already looks.

Swap the trigger and it still runs.
Speed: an interviewer asks how fast the flag has to run to fit inside the radiologist's existing read time. Same split, same two numbers, now bounded by a time budget instead of a workload budget.
Cost: the hospital caps how many rechecks the desk can absorb at 100 a night. Same equation, run backwards, solve for the highest catch rate that flag-rate ceiling allows.
The model got better: a new version catches 95 percent of misses at the same 15 percent flag rate. Same two numbers, just better numbers. The method never changes.

Where people run it wrong.
They set one bar and call it done, so the tool ends up optimizing the wrong half of the mistake without anyone deciding that on purpose.
They chase the catch rate toward 100 percent with no ceiling on the flag rate, and the desk drowns in rechecks until nobody trusts a flag anymore.
They set the bar once at launch and never touch it again, so a staffing change two years later leaves a flag-rate ceiling nobody re-checked.

How to use it live. Say the split out loud before any number: "I'd set two numbers here, not one, a catch rate on real misses and a flag rate on clean scans." That sentence buys you time to actually do the arithmetic, and it tells the interviewer you're not about to hand back a single blended score.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a pass-bar question like this, and why not FLIPS?
Tap to flip
ANSWER
BOUND. There's no person's habit snapping here, just an argument about numbers and cost asymmetry. FLIPS needs a two-setting behavior flip; this question needs arithmetic and a range.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ingrid Voss, clinical safety lead at Cascade Radiology Partners. She signs off on any AI tool before it touches a live scan.
3 · THE NUMBER SHE ALMOST TRUSTED
What number did Ingrid almost sign off on, before she checked it herself?
Tap to flip
ANSWER
94 percent accuracy, the vendor's headline number. It looked like safety. Tested against real cases, the tool caught only two of nineteen real misses.
4 · THE SPLIT
What two numbers replace the single 92 percent bar?
Tap to flip
ANSWER
A catch rate on the radiologist's real misses (at least 90 percent) and a flag rate on clean scans (capped near 15 percent). One protects the patient, one protects the review desk.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Grading the tool against a single 92 percent accuracy target, matched to the radiologist's own score. It sounded fair. It fails because only 1 in 10 scans has a real finding, so a tool that flags nothing already scores close to 90 percent.
6 · THE NUMBER
Fill in the blank: out of 100 real findings, the radiologist alone catches ___ and misses ___.
Tap to flip
ANSWER
92 caught, 8 missed. That's a good radiologist having a normal day, not a failure.
7 · THE REPLAY
Same three months of scans, tuned against the two-number bar instead of one accuracy score. What changes?
Tap to flip
ANSWER
It catches 16 of 19 real misses instead of 2, and flags about 142 clean scans, roughly 16 percent. Worse-looking accuracy. Far better at the actual job.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the split there?
Tap to flip
ANSWER
A contract-translation review tool. Catch rate on clause-level mistranslations the first translator missed, capped flag rate on already-correct clauses sent back for a second read.

Check yourself Score: 0 / 0

Fill in the blank
1. Out of every 1,000 scans in this answer's example, the radiologist alone catches ___ of the 100 real findings and misses ___.
Show hint
92 plus how many equals 100?
Show answer
92 caught, 8 missed. That's the human baseline the whole answer is built around.
Multiple choice
2. Why does this answer refuse to grade the tool against 92 percent accuracy alone?
  • A. Because accuracy is impossible to calculate for AI tools.
  • B. Because in a population where only 1 in 10 scans has a real finding, a tool that flags nothing already scores close to 90 percent, so accuracy alone rewards doing nothing.
  • C. Because 92 percent was invented for this example and isn't a real number.
  • D. Because radiologists always outperform AI tools no matter what bar is set.
Show hint
Think about what "always say normal" scores in a population where most scans are normal.
Show answer
B. A blended accuracy score can be won by ignoring the rare, real findings entirely, which is exactly the opposite of what a flagging tool is for.
True or false
3. True or false: the proposed pass bar asks the tool to catch fewer than half of the radiologist's real misses.
  • True
  • False
Show hint
The proposed catch rate is 9 out of every 10, not 1 out of every 2.
Show answer
False. The proposed bar asks for at least 90 percent of real misses caught, far more than half, because a missed finding often has no second chance.
Short answer, apply it yourself
4. Pick a product you use yourself that flags something for a human to double check, a spam filter, a fraud alert, a spell-checker. Name one pass bar for it that should be two numbers instead of one, and say why the two costs aren't equal.
Show hint
Think about what a missed flag costs versus what a false flag costs.
Show answer
Model answer: "A bank's fraud alert: missing a real fraud costs the customer real money and trust. A false alarm costs them one text message confirming a purchase. The catch-rate number should be very high, the false-alarm number can be much more relaxed, because the two mistakes are nowhere near the same size." Any answer works if it names a real product and two genuinely unequal costs.
Short answer, the number question
5. If prevalence doubled, 200 real findings in 1,000 scans instead of 100, would a flag-rate ceiling of 15 percent of total volume still leave enough room to catch 90 percent of the radiologist's misses? Work through it.
Show hint
Work out the new miss count first, then check it against the same 150-flag budget.
Show answer
Yes, it still fits. At 200 real findings, the radiologist still misses about 8 percent, roughly 16 this time instead of 8. Catching 90 percent of those means catching about 14. The total flag cap stays at 150, since it's defined as 15 percent of overall volume either way, so there's still room: the false-alarm tolerance among the smaller pool of clean scans only needs to move from about 16 percent up to about 17 percent.
Multiple choice
6. What old decision does this answer take back?
  • A. Grading the flagging tool against a single 92 percent accuracy target that matched the radiologist's own score.
  • B. Removing a confidence mark from the radiologist's screen.
  • C. Hiring a second radiologist for the overnight desk.
  • D. Lowering the flag-rate ceiling all the way to zero.
Show hint
Look at "the choice I would take back" in the first section.
Show answer
A. Matching the bar straight to the human's own accuracy sounded fair but rewarded a tool that caught almost nothing, since normal scans dominate the population.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more