What is the product cost of a false positive versus a false negative in a resume-screening feature?
Screenmark reads every application to Strathmere Health Alliance's critical-care nursing openings, across all 14 hospitals, and decides which resumes a recruiter ever sees. Hesper Vasterling owns the call on where its cutoff sits. Leocadio Hazenkirk runs the recruiter team that lives with one kind of mistake every week. Roswitha Tansley found the other kind, sitting untouched in a pile nobody had opened in fourteen months.
- Add a review band below the hard cutoff for scarce, critical-care roles.Why: this is the actual position. Nothing else here fixes the resumes that were never seen.
- Set the band where the audit found the errors, not by lowering the whole cutoff.Why: a blanket lower cutoff floods recruiters on roles that were never the problem.
- Strip the raw employment-gap field from the score, and check for a proxy standing in for it.Why: removing one field does nothing if another field is quietly rebuilding the same signal.
- Leave the hard cutoff alone on roles with a deep applicant pool.Why: on those roles, a missed candidate is replaced by the next one in line, so the extra review buys nothing.
- Set a real kill line: applicant-to-opening ratio, not a feeling.Why: a position with no way to be proven wrong is a habit, not a decision.
- Put a re-check of the decline pile on the calendar, every quarter, going forward.Why: nobody had looked at it in fourteen months. That gap is what let this run silent for so long.
How to answer this, stage by stage
Nobody is grading whether you can define false positive and false negative. They're grading whether you'll actually commit to which one costs more here, and say why, out loud, in dollars.
Let's learn
Before Screenmark, one recruiter read every critical-care resume by hand. About 14 a day, alone, no ranking tool. When a resume had a six-month gap in the dates, that recruiter did the obvious thing: kept reading, looked for the reason, called if the rest of it looked strong.
Screenmark reads about 2,600 applications a year to Strathmere's critical-care postings, ICU, ER, NICU, cardiac stepdown, across all 14 hospitals. It blends credential match, unit-specific experience, keyword overlap with the requisition, and one more signal: how continuous the work history looks. Every resume gets a score from 0 to 100.
At launch, the cutoff sat at 61. Score 61 or above, Screenmark hands the resume to Leocadio's team for a full review, about 22 minutes, phone screen and reference checks included. Score below 61, the resume gets an automatic decline email, and no person ever opens it. About 1,700 of the 2,600 applications a year, 65 percent, never get seen by anyone.
Here's the turn. The interesting mistake was never the 22 minutes Leocadio's team loses on a resume that doesn't pan out. It's what happens to the resumes that never make it to his team at all.
What it costs at its worst: a critical-care seat that stays empty gets covered by a travel nurse, at a real premium Strathmere already tracks, about $4,600 a week over a staff hire. When the strongest applicant for a seat never got seen, Strathmere's own fill-time data shows the seat takes about two extra weeks to close, on average. That's roughly $9,200, once, per seat, every time it happens quietly.
What I would leave alone: Strathmere's general medical-surgical floor postings. Those get about 14 qualified applicants for every opening. A missed one there costs almost nothing, because another equally strong resume is already sitting in the same pile. Building a review band for that pool would just slow recruiters down for no real gain.
The lesson: a cutoff isn't wrong just because it makes mistakes. Every cutoff does. It's wrong when nobody ever checks which side of the line the mistakes are landing on, and for how long.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why fourteen quiet months, not one bad week, is what let this run.
Hesper Vasterling built her name on catching a model's blind spot before it ever reached anyone's inbox. Screenmark's employment-gap penalty was the one that got past her.
She shipped Screenmark eighteen months ago with a simple, defensible design: one score, one cutoff, 61. Above it, a recruiter looks. Below it, an automatic decline, instant, and kind about it in the email. Leocadio's team loved it. Reviews that used to take all morning took an hour. Override rate on the resumes Screenmark did pass stayed low, month after month, and nobody had a reason to second-guess the ones it didn't.
For a while, that felt like enough. Screenmark was fast. It was fair, as far as anyone could tell, because nobody was looking at the pile it was quietly discarding.
Around month eleven, ICU and NICU vacancies at three of the smaller hospitals started running long enough that nursing directors were leaning harder on travel-nurse coverage than anyone liked. Nobody connected it to Screenmark. Staffing is always tight somewhere. It read like weather, not a signal.
Then, in the ordinary course of a fairness review nobody expected to find anything, Roswitha Tansley pulled a stratified sample of 400 resumes from the year's auto-declines and had two senior clinical recruiters rescore them by hand, blind to Screenmark's number, against the real job requirements.
Thirty-four of the 400 should have gone to a phone screen. Not maybe. Both recruiters agreed, independently, before they compared notes.
Twenty-one of those 34, 62 percent, had an employment gap of six months or more. Most weren't unexplained. Caregiving leave. A nurse trained abroad, waiting on state license verification, a process that can run four to six months on its own and has nothing to do with whether she can run a code. Screenmark had learned, from years of past hiring outcomes, that a continuous work history correlated with getting hired. It applied that correlation evenly, to every gap, regardless of what caused it.
Roswitha didn't stop at the finding. She ran the math forward: 8.5 percent of the full 1,700 auto-declines, scaled up, is about 145 wrongly filtered candidates a year, network-wide, for critical-care roles alone. Not all of them would have changed a hiring outcome. But cross-checked against Strathmere's own fill-time data, about 58 of those a year land on seats that measurably took longer to fill because the strongest applicant was never seen. At roughly $9,200 in extra agency coverage per seat, that's about $533,600 a year, quietly, on top of every nurse who simply never got a chance.
Hesper's fix wasn't to lower the cutoff for every role. That would flood Leocadio's team with reviews for postings, like general med-surg, that already have plenty of good applicants. Instead, she added a band: any resume scoring 40 to 60 on a critical-care requisition gets a fast, four-minute triage look, just a license and certification check, not the full 22-minute review. Only below 40 stays fully automatic. About 680 resumes a year fall into that band, costing roughly 45 recruiter-hours a year, a little over a week of one person's time, spread across twelve months.
Run the fix back against the audit sample: 29 of the 34 miscategorized resumes fall inside the new 40 to 60 band. They get seen now, not filtered silently, for about 45 hours a year against roughly $533,600 a year in hidden cost. That's not a close call either, just pointed the other way.
What Hesper would tell her past self, back at launch: a single cutoff is the cheapest thing to build and the easiest thing to demo. It is also silent by design about which side of the line it's getting wrong, and silence is exactly what let this run for fourteen months.
PICK, and the one number an audit can't forgive
Not a way to dress up "false negatives are worse" as an opinion. PICK forces a real commitment, then makes you prove which mistake actually costs more, in the same unit, on both sides.
Three things worth stating directly, since this is where the real judgment sits. The alternative Hesper's team considered, and rejected, was simply lowering the single 61 cutoff for every role, network-wide, to catch more of the same candidates without building a band at all. It lost, because a lower cutoff on general med-surg postings, where 14 qualified applicants already compete for every opening, would have flooded Leocadio's team with hundreds of extra reviews a year for a pool that was never the problem. The AI-specific failure worth naming is a training-data proxy: Screenmark learned that a continuous work history predicted getting hired, from Strathmere's own past hiring outcomes, and applied that pattern to every gap alike, whether it came from caregiving leave, license verification, or nothing worth flagging at all. The guardrail is two-part: strip the raw employment-gap field from the score entirely, and check for a proxy standing in for it, since a feature-importance audit after the fix found "years since first listed job" was quietly reconstructing about 80 percent of the removed signal on its own. And the trade-off is real and stated on purpose: about 45 extra recruiter-hours a year, and a few added days before a band candidate hears back, against an estimated $533,600 a year in hidden agency cost from seats that sat open longer than they should have.
And if you want to be sure it really works, try it somewhere else
Same four letters, a credit union instead of a hospital network, and this time the silent filter wasn't a gap in a resume. It was a thin credit file.
Candorix is Wexbridge Community Credit Union's pre-qualification screener. It reads a loan application and scores it 0 to 100 before anyone at the credit union sees it: income, existing debt, and how long the applicant's credit file runs. Score 58 or above pre-qualifies for a full application. Below it, an automatic decline. Vernice Kettlewick handles the applications that make it through.
Over about 18 months, with no single bad week to point at, Vernice noticed thin-file referrals kept ending the same way: pre-qualified, declined, gone. The credit union's own numbers, once someone finally pulled them, showed a 71 percent auto-decline rate for thin-file applicants against 34 percent for long-file applicants with comparable income and debt. Nothing had crashed. It had just quietly drifted that way and stayed there.
Same rank, different lever: the fix isn't a smarter income model. It's the same shape of band: for thin-file applicants under 24 months of credit history, pull in rent and utility payment history as a substitute signal, and route anyone within 8 points of the cutoff to a human underwriter instead of an automatic decline. Early results suggest about 260 additional applicants a year would clear pre-qualification who were otherwise silently filtered, a false negative that used to cost Wexbridge a member for good and never showed up on any report.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name which miss is cheap and visible, which is hidden and expensive, pick the second one to protect, and say the kill line.
Cost: no budget this quarter for a full audit like Roswitha's. Pull the cheapest slice first, a stratified sample of 100 declines instead of 400, and say plainly the estimate is provisional until the full sample runs.
The model got better, for real: say Screenmark's credential-matching accuracy improves by ten points overall. The pick doesn't change. A better model still needs the band, because the gap penalty was never about accuracy on the parts it was already good at. It was a blind spot on a feature accuracy alone doesn't touch.
Where people run it wrong.
They treat "false negative" as automatically worse, without ever pricing either side in the same unit.
They fix a scarce-role problem by loosening the cutoff everywhere, wasting recruiter or underwriter hours on roles that were never the issue.
They remove one biased feature and declare it fixed, without checking whether a correlated feature is quietly doing the same job.
How to use it live. Ask, before naming a pick: "Which of these two mistakes gets caught by someone, and which one gets caught by no one unless somebody goes looking for it?" Whichever one nobody's watching for is usually the one to protect against.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Doesn't the review band just move the false-negative problem instead of solving it?" Response: no, because the band targets exactly where the audit found the errors concentrated, 29 of 34 miscategorized resumes fall inside it, not spread evenly across the whole score range.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on What changes when the product is probabilistic
- #1 Name three product decisions that change when a feature's output is probabilistic rather than deterministic.
- #2 A traditional feature either works or has a bug. Explain why that framing breaks for an LLM feature.
- #3 What does 'correct' mean for a summarization feature? Give a definition your engineering team could test against.
- #4 QA files a bug that reads: the model gave a wrong answer once. How do you triage it?
- #5 Explain the difference between a defect and an acceptable error rate to a non-technical executive.
- #6 Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?