How do you set a pass bar when human performance on the same task is 92 percent?
- Split the single 92 percent target into two numbers, a catch rate on real misses and a flag rate on clean scans.Why: they pull against each other, and one blended score hides which one you're actually protecting.
- Set the catch rate near total, at least 9 of every 10 real misses caught.Why: a missed finding on this task often gets no second chance, so this is the number that can't fail quietly.
- Cap the flag rate at what the review desk can actually absorb, around 15 out of every 100 scans.Why: a great catch rate that floods the queue gets ignored, the same way any over-eager tool does.
- Move the catch-rate target up or down based on whether a miss has another safety net, not on how rare the finding is.Why: whether a miss gets caught some other way swings the number more than anything else does.
- Never grade the tool on raw accuracy measured against 92 percent.Why: only 1 in 10 scans has a real finding here, so a tool that says "normal" every time already scores close to 90 percent while catching nothing.
- Recheck the flag-rate ceiling every time staffing on the review desk changes.Why: the same catch-rate target turns unworkable the moment one radiologist covers the desk instead of three.
How to answer this, stage by stage
Nobody's grading you on landing the exact right percentage here. They're grading whether you split the bar on purpose instead of copying one number off the page. Eight moves get you there.
Let's learn
Out of every 100 scans that actually have something wrong on them, a radiologist reading alone catches 92. She misses 8. That's before any tool has touched the work. That's not a bad day. That's what a genuinely skilled person looks like on a task nobody does perfectly.
Now put a tool behind her. It looks at chest CT scans after she has already read them, and flags any scan it thinks she got wrong. A flagged scan gets a second look before the report goes out.
Now say we set the tool's pass bar the easy way: beat 92 percent. One number, matched to the human.
Here's the turn. That one number does not tell the tool what to protect. A missed finding leaves a patient with nothing, no treatment starts, and often the next scheduled check isn't for months. A wrongly flagged clean scan costs the radiologist five minutes rereading something that was fine. Grading both mistakes against the same 92 percent treats a life and a coffee break as the same size problem.
At its worst, that mistake compounds. Tune the flag threshold to "match 92 percent accuracy" and it can end up doing almost nothing extra, because "normal" is already the right call on 9 of every 10 scans, and saying nothing scores well on accuracy. The tool passes its own bar and catches almost none of the misses it was built to catch.
The choice I would take back. We'd have set the target as "beat 92 percent overall," because it sounded fair, a bar as good as the best human in the room. I'd take that back. In a population where only 1 in 10 scans has a real finding, a tool that never flags anything already scores 90 percent. The bar has to come from the two real costs, not from one blended number.
What I would leave alone. Not every flag in this system needs this treatment. The tool also flags scans with a technical problem, bad positioning, motion blur, that need a retake. Missing one of those costs a same-day callback. A false flag there costs thirty seconds glancing at a fine image. Both sides are cheap and reversible. A simple bar, beat the scanner's own error rate, is fine there. Save the two-number treatment for mistakes that don't cost the same on both sides.
The lesson. A pass bar is not one number copied from a human's score. It's two numbers, one for each kind of mistake, sized by what that mistake actually costs, not by what sounds fair on a slide.
Now here is the same thing as a story
Skip this part if you already believe two mistakes can cost wildly different amounts. Read on if you don't.
Ingrid Voss runs clinical safety for Cascade Radiology Partners, a chest imaging group that reads overnight for a dozen hospitals. She's the one who signs off on any AI tool before it touches a live scan, and she has turned down more tools than she's approved.
The new one is a second-opinion flagger. It reads every scan behind the radiologist on shift and flags anything it thinks might have been missed.
When the vendor pitched it, their whole deck rested on one line: 94 percent accuracy, beats your radiologists' 92. Ingrid almost signed off on that alone. It was a bigger number. Bigger sounded safer.
She ran it against three months of the group's own scans anyway, before committing, out of habit more than suspicion.
The tool matched the pitch on accuracy. And it caught two of the nineteen real misses hidden in that quarter's data. Two out of nineteen.
She sat with that for a while. The tool wasn't broken. It had just learned, correctly, that "normal" is the right call on 9 of every 10 scans in this population, so calling almost everything normal was a cheap way to a good accuracy score. It only had to find something on the rare scan where the radiologist's own miss was hiding, and it mostly didn't bother.
So Ingrid threw out the single number. She sat down with the group's chief radiologist and asked two separate questions instead: of the misses we can find in old cases, how many would this have caught, and of the clean scans, how many would it needlessly send back.
Rerun against the same three months, tuned against those two numbers instead of one blended score: it caught 16 of the 19 misses, and flagged 142 clean scans out of the roughly 900 in the sample, about 16 percent. Worse-looking accuracy number than the vendor's pitch. Far better at the one job that mattered.
She signed off on that version. Not the higher-accuracy one.
The thing I'd want Ingrid to say out loud, if she got asked why in an interview: the vendor's number answered a question nobody was actually asking. The real question was never "is this tool usually right." It was "when it's wrong, which kind of wrong is it, and can the desk survive that."
Five letters, and where each one lands here
This is an estimation question with a cost asymmetry built in, so BOUND fits, not FLIPS. Nobody's habit is fading over months here. It's an argument about two numbers and what each one is allowed to cost.
B, break it down. The bar isn't one score, it's two: a catch rate on the radiologist's real misses, and a flag rate on clean scans, because those two mistakes cost different amounts.
O, own the numbers. Out of 1,000 scans and 100 real findings, the radiologist catches 92 and misses 8. Target: catch at least 7 of those 8. Keep total flags, real catches plus false alarms, under about 150.
U, use a range. 75 percent catch rate when a miss has another safety net downstream. 97 percent or higher when it doesn't. The proposed 90 percent sits between those, closer to the strict end on purpose.
N, nail the sanity check. 92 percent human accuracy still means real patients missed every week at a busy site. A 15 percent flag rate is about 7 extra rechecks a night, which one radiologist can actually clear.
D, direction. Whether a miss has a downstream safety net moves the target catch rate more than prevalence, staffing, or anything else on the list.
And if you want to be sure it really works, try it somewhere else
A contract-translation review tool works the same way. A translator drafts the target-language version, then a checking tool flags clauses it thinks got the meaning wrong, for a second translator to look at before the contract goes out.
B, break it down. Same split: how many of the first translator's real clause-level errors does the tool catch, and how many correct clauses does it needlessly flag for a second read.
O, own the numbers. Say a firm processes 200 contracts a month, about 40 load-bearing clauses per contract, 8,000 clauses total. A skilled translator gets the legal meaning right on about 89 percent of those on a first pass, so around 880 clauses a month carry a real error. Target: catch at least 90 percent of those, about 790, while flagging no more than 12 percent of all 8,000 clauses, about 960.
U, use a range. Boilerplate clauses, ones a lawyer reads anyway before signing, can sit at a lower catch rate, maybe 70 percent, since there's a safety net downstream. Liability and indemnity clauses, the ones relied on directly if a dispute ever happens, need a catch rate close to total, 95 percent or higher.
N, nail the sanity check. 89 percent sounds solid for a translator, but across 8,000 clauses a month that's close to 30 real errors a day going into contracts before any check runs. A 12 percent flag rate is about 32 clauses a day sent back, workable for a small review team.
D, direction. Same lever as the radiology case: whether a clause type gets read again downstream by someone else, a lawyer, a counterparty, moves the target catch rate more than the base error rate does.
Swap the trigger and it still runs.
Speed: an interviewer asks how fast the flag has to run to fit inside the radiologist's existing read time. Same split, same two numbers, now bounded by a time budget instead of a workload budget.
Cost: the hospital caps how many rechecks the desk can absorb at 100 a night. Same equation, run backwards, solve for the highest catch rate that flag-rate ceiling allows.
The model got better: a new version catches 95 percent of misses at the same 15 percent flag rate. Same two numbers, just better numbers. The method never changes.
Where people run it wrong.
They set one bar and call it done, so the tool ends up optimizing the wrong half of the mistake without anyone deciding that on purpose.
They chase the catch rate toward 100 percent with no ceiling on the flag rate, and the desk drowns in rechecks until nobody trusts a flag anymore.
They set the bar once at launch and never touch it again, so a staffing change two years later leaves a flag-rate ceiling nobody re-checked.
How to use it live. Say the split out loud before any number: "I'd set two numbers here, not one, a catch rate on real misses and a flag rate on clean scans." That sentence buys you time to actually do the arithmetic, and it tells the interviewer you're not about to hand back a single blended score.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Acceptance criteria for non-deterministic output
- #1 Rewrite this criterion to be testable: the model should not hallucinate.
- #2 How do you express an acceptance criterion as a rate rather than an absolute?
- #3 What is the difference between a threshold criterion and a distributional criterion?
- #4 Write acceptance criteria for an AI feature that extracts fields from an invoice.
- #6 Describe acceptance criteria that account for the severity of different error types.
- #7 How would you write criteria for a feature where the worst case matters more than the average?