CaseIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #22

A vendor's benchmark shows 98 percent accuracy. What do you ask next?

BOUND a benchmark number is a claim wearing a decimal point

Say a recruiting firm is choosing an AI resume-screening vendor to pre-rank applicants before a human recruiter looks at them. Halstead Talent is that firm. Marcus Whitfield is Head of Talent Acquisition Product there. ScreenWise is the vendor whose sales deck opens with "98% accuracy."

The direct answer
Ask what "correct" means, on what set of resumes, before asking anything about the number itself. Ninety-eight percent accuracy on a clean, balanced benchmark the vendor built and tuned on is a different fact entirely from ninety-eight percent on your own messy applicant pool, and the gap between those two numbers is usually the whole story.
Do this, in order
  1. Ask what counts as "correct" in their benchmark.Why: accuracy can mean matching a label exactly or just landing in the right general bucket, and those produce very different real numbers.
  2. Ask whose resumes were in the eval set.Why: a benchmark built on clean, well-formatted resumes tells you almost nothing about performance on career changers, employment gaps, or non-native English resumes.
  3. Ask for the accuracy on rejected candidates specifically, not just overall.Why: a heavily imbalanced set of mostly-obvious accepts can hit 98 percent overall while doing badly on the harder reject calls that matter most.
  4. Ask for a range, not a single number.Why: one point estimate with no confidence interval hides how much that 98 could move with a different sample.
  5. Run your own eval on a sample of your real applicant pool before signing.Why: it's the only step that actually tells you what will happen to your candidates, not the vendor's.

How to answer this, stage by stage

Nobody is grading whether you can repeat "correlation isn't causation." They're grading whether you can turn one flashy number back into an equation with parts you can actually question.

Stage 1
Scope it to one real number
Say it like this
"I'll answer this for Halstead Talent, evaluating ScreenWise's claimed 98 percent accuracy on resume screening, before signing a contract for it."
Why this works
Keeps "what do you ask next" from turning into a list of generic due-diligence questions.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, the equation behind the number. Own the numbers, where each assumption comes from. Use a range, not a point estimate. Nail the sanity check. Direction, which assumption swings it most."
Why this works
Tells the interviewer you're about to do arithmetic, not just express skepticism in the abstract.
Stage 3
Break the claim into its parts
Say it like this
"98 percent accuracy equals correct calls, however they define correct, divided by total calls, on whatever set of resumes they used. Change any one of those three things and the number changes without the model changing at all."
Why this works
Shows the number isn't one fact, it's an equation with at least three hidden knobs.
Stage 4
Give the one question, before any reasoning
Say it like this
"The first thing I ask is what 'correct' means and whose resumes were in the eval set, because until I know both, 98 percent tells me nothing I can act on."
Why this works
This is the direct answer, said as one concrete question instead of a vague call for "more rigor."
Stage 5
Own the numbers, out loud
Say it like this
"I'd assume their eval set overlaps maybe sixty percent with our real applicant pool, since off-the-shelf benchmarks rarely include career changers or employment gaps in real proportion, and I'd assume 'correct' means top-one label match, the strictest common definition."
Why this works
Shows you can state an assumption and where it comes from, instead of just gesturing at uncertainty.
Stage 6
Sanity check the adjusted number
Say it like this
"If my adjusted estimate lands around 84 percent, I'd compare that to how often two of our own human recruiters agree on the same resume, roughly 89 percent historically. If the AI's real number is close to or below that, it isn't actually beating a second human opinion."
Why this works
Ties the estimate to something the room already trusts, instead of leaving it floating on its own.
Stage 7
Name the direction, what would move it most
Say it like this
"The single biggest swing factor is whether their eval set actually overlaps our applicant pool. If I only get to ask one follow-up question, that's the one."
Why this works
Shows you know which assumption is load-bearing, not just that assumptions exist.
Stage 8
Close on the one line
Say it like this
"98 percent isn't a fact yet. It's a claim with three hidden knobs, and I want to see all three turned before I let it into a hiring decision."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

Say a recruiting firm screens resumes for warehouse operations roles at a scale of a few thousand applicants a month.

Before ScreenWise, two human recruiters split the pile and moved about forty resumes an hour each, agreeing with each other on roughly eighty-nine percent of pass or reject calls when spot-checked against one another. ScreenWise's sales deck claimed its model matched human judgment ninety-eight percent of the time.

Hand sketched flow diagram titled How a vendor benchmark number gets made. Four boxes: Pick an eval set, Tune on similar data highlighted, Run it once, Report the best number.
Every one of these four steps is a choice. None of them are visible in the headline number alone.

Here's the turn: the ninety-eight isn't wrong, exactly. It's real, on whatever set ScreenWise measured it on. The mistake isn't trusting a false number, it's trusting a true number attached to a question nobody asked.

Hand sketched labeled parts diagram titled What the 98 percent claim doesn't say. A question box icon at the center labeled 98% accuracy, with four callouts around it: what eval set, what counts as correct, tested on whose resumes, what's the range.
A number with none of these four answered is a number you can't act on yet.

At its worst, signing on the headline number alone means quietly changing who gets seen by a human recruiter at all, based on a figure that was never tested on the applicants actually applying.

Building up a real-world accuracy estimate from ScreenWise's 98 percent claim
100% 50% 0 98% Vendor claim -8 eval set -4 definition -2 imbalance 84% Real-world estimate
Three adjustments, none of them exotic, pull the headline number down fourteen points before it ever touches a real applicant.
The number worth sanity-checking against Halstead's own human recruiters agree with each other about 89 percent of the time on the same resume. An 84 percent real-world estimate for ScreenWise means it isn't clearly better than a second human opinion, even before counting the cost of trusting a vendor's own benchmark on faith.

What I would leave alone: ScreenWise's core model architecture and training approach don't need this scrutiny, that's the vendor's business and not something a buyer needs to audit line by line. The scrutiny belongs entirely on the claim, not the method behind it.

The lesson: a benchmark number is the answer to a question. Before trusting it, make sure you know what the question actually was.

Now here is the same thing as a story

The short version above is what you'd say in the vendor selection meeting. Read this one for how the number almost went unquestioned.

Before the ScreenWise pitch, Marcus Whitfield's team screened resumes the slow way: two recruiters, a shared spreadsheet, and enough disagreement between them that a third reviewer settled ties once a week. After the pitch, everything changed for a while.

ScreenWise's deck opened on the number, ninety-eight percent accuracy, in type twice the size of anything else on the slide. The demo ran smoothly. The sales engineer answered every question about integration timelines and data security without hesitation. Nobody in the room, at first, asked what "accuracy" actually meant.

Knowledge spark: why can two vendors both claim "98 percent" and mean different things? Accuracy is just correct calls divided by total calls, but "correct" can be defined loosely (the model's top guess was in the right general category) or strictly (it matched the exact label a human gave). A benchmark that only counts obvious, easy cases as the "total" can look excellent while still failing badly on the hard cases that actually mattered.

It was the newest member of the team, three weeks into the job, who finally asked the question out loud during the vendor Q&A: "Ninety-eight percent correct according to what?" The sales engineer paused, then said the benchmark measured agreement with the vendor's own labeling team on a curated set of five thousand resumes, mostly drawn from tech and finance roles, not warehouse operations.

Hand sketched comparison titled The vendor's eval set vs our real applicant pool. Left, a green document icon labeled Vendor eval set, caption clean balanced resumes. Right, a red document icon labeled Halstead's real pool, caption gaps career changers typos.
Same word, "resumes." Two very different piles of paper.
The number wasn't false. It just answered a question about a different applicant pool than the one about to be screened.

Marcus's team ran their own two-week pilot before signing anything, feeding ScreenWise a thousand real warehouse-role applications and checking its calls against the two human recruiters' own agreed-upon decisions. The model's real accuracy on that set landed at eighty-four percent, fourteen points under the number on the slide.

Eighty-four percent, compared honestly against the recruiters' own eighty-nine percent agreement rate with each other, meant ScreenWise wasn't clearly better than a second human opinion. It was faster. It wasn't proven more accurate, not yet, not on the resumes that actually mattered.

Hand sketched quadrant titled Claimed accuracy vs real world accuracy, by vendor. ScreenWise sits far right on claimed accuracy but only middling on real world accuracy. Vendor B and Vendor C claim less but land closer to their claims.
The vendor with the loudest number wasn't the one whose real number held up best.

BOUND, in one screenNot a warning to "be skeptical of numbers." BOUND is what tells you exactly which three questions to ask first.

B
Break it down. State the equation before touching the number.
98 percent accuracy equals correct calls, by whatever definition, divided by total calls, on whatever eval set was used.
This is the hardest step and the answer to the question: turning one claim into three separate, checkable parts.
O
Own the numbers. State each assumption and where it comes from.
Assume roughly 60 percent overlap between the vendor's eval set and Halstead's real applicant pool, and assume "correct" means the strictest, top-one label match.
Shows you can commit to a working number instead of only raising doubts.
U
Use a range. Low, high, and a point estimate.
A low estimate around 71 percent, a high end at the vendor's claimed 98, with a point estimate near 84 given the adjustments made.
A single number implies more confidence than the situation earns.
N
Nail the sanity check. Compare against something already trusted.
84 percent against the recruiters' own 89 percent inter-rater agreement means the model isn't clearly beating a second human opinion yet.
Anchors the adjusted estimate to a number the room already believes, instead of leaving it floating alone.
D
Direction. Which assumption swings it most.
Eval-set overlap with the real applicant pool moves the estimate more than any other single factor.
Tells you exactly which follow-up question to ask first if you only get one.
Hand sketched timeline titled The range behind one claimed number. Three milestones: Low estimate 71 percent, Vendor claim 98 percent highlighted, Our real traffic 84 percent.
The vendor's number sits at one end of a real range. The point that matters landed closer to the middle.

The recap, one line per letter: break it down is turning "98 percent" into correct calls over total calls on a named set, own the numbers is stating the assumed overlap and definition out loud, use a range is giving 71 to 98 with 84 as the working estimate, nail the sanity check is comparing that 84 to the recruiters' own 89 percent agreement rate, and direction is naming eval-set overlap as the single biggest lever.

Hand sketched icon list titled Questions to ask before trusting a benchmark number. Three items: a question box icon labeled What exactly was tested, a scale icon labeled Correct by whose definition, a funnel icon labeled Whose resumes were excluded.
Ask these three before the second slide of the vendor deck, not after the contract is signed.

And if you want to be sure it really works, try it somewhere elseSame five letters, a seafood plant's AI defect-grading vendor instead of a resume screener. A different hidden knob breaks the second story.

Nordfjord Seafood runs quality grading at a mid-size fish processing plant, considering Gleaner Vision, an AI vendor claiming 95 percent accuracy at spotting defects on filleted fish before packaging. Marit Halvorsen, Head of Quality, is the one deciding whether to trust the number. Mapped onto BOUND: break it down is the same shape, 95 percent equals correct defect calls over total calls, on whatever set of fish images was used; own the numbers assumes the vendor's eval set was shot under studio lighting, not the plant's actual conveyor-line conditions; use a range gives a low estimate near 68 percent and a point estimate near 79 once lighting and belt-speed motion blur are accounted for.

The hidden knob that breaks this story differently isn't the eval set's content, it's the plant's real production speed. Gleaner Vision's benchmark measured accuracy on still images with no motion blur, since that's how their internal test set was captured. Nordfjord's actual conveyor line moves fish past the camera fast enough that motion blur is common, a factor the benchmark never tested for at all, because nobody at the vendor had asked what speed a real plant's belt runs at.

Hand sketched labeled parts diagram titled What the 98 percent claim doesn't say, reused here for Gleaner Vision's 95 percent defect grading claim. Question box center, with what eval set, what counts as correct, tested on whose resumes, and what's the range around it.
Swap "resumes" for "fish images" and the same four gaps in the claim show up again.
Gleaner Vision's defect-detection accuracy, still images versus real conveyor speed
0 swing 16 pts Conveyor motion blur 9 pts Lighting conditions 5 pts Fish species mix 2 pts Camera angle
Conveyor motion blur alone swings the estimate more than the other three assumptions combined, and it's the one Nordfjord's own plant conditions actually produce every shift.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask what counts as correct and whose data it was tested on, then run your own pilot before signing," and stop.
Cost: there's no budget for a two-week pilot before the vendor decision is due. Say so honestly, and shrink the pilot to a few hundred real resumes checked against your own recruiters' agreed calls, rather than skipping verification entirely.
The model gets better, for real: if a follow-up benchmark, run on your own applicant pool, genuinely shows 95 percent, that's the pilot doing its job in reverse, telling you the number has finally earned the trust the sales deck asked for on day one.

Where people run it wrong.
They accept the vendor's own benchmark as sufficient proof, without asking what set it was run on.
They compare the headline number to nothing, no human baseline, no range, so 98 percent sounds impressive purely because it's a big number.
They test on the vendor's easiest cases and skip the ones that actually decide the hiring outcome, the borderline rejects.

How to use it live. The moment someone quotes you a benchmark number, ask back: correct according to what, and tested on whose data? Let the answer to those two questions decide how much weight the number deserves.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an estimation and sizing question like unpacking a vendor's benchmark claim?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, direction. It shows the arithmetic instead of just expressing doubt.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marcus Whitfield, Head of Talent Acquisition Product at Halstead Talent, and the new hire on his team who first asked what "98 percent" actually meant.
3 · BREAK IT DOWN
State the equation behind "98 percent accuracy," out loud.
Tap to flip
ANSWER
Correct calls, by whatever definition of correct was used, divided by total calls, on whatever eval set was tested.
4 · THE FIRST QUESTION
What's the one question this answer says to ask before anything else?
Tap to flip
ANSWER
What does "correct" mean, and whose resumes were in the eval set.
5 · THE SANITY CHECK
What number did Halstead compare the adjusted 84 percent estimate against, and why?
Tap to flip
ANSWER
Their own recruiters' 89 percent inter-rater agreement rate. If the AI's real number is near or below that, it isn't clearly beating a second human opinion.
6 · THE NUMBER
Fill in the blank: after adjusting for eval-set narrowness, strict definition of correct, and class imbalance, ScreenWise's claimed 98 percent became a real-world estimate near ___ percent.
Tap to flip
ANSWER
84 percent. Fourteen points below the number on the sales deck.
7 · DIRECTION
Which single assumption swings the real-world estimate the most, and why?
Tap to flip
ANSWER
How much the vendor's eval set actually overlaps with your real applicant pool. It's the single biggest lever on the final number.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what hidden factor swings the estimate most there?
Tap to flip
ANSWER
Nordfjord Seafood's Gleaner Vision defect-grading vendor. The biggest swing factor there is conveyor-line motion blur, which the vendor's still-image benchmark never accounted for.

Check yourself Score: 0 / 0

Short answer, name the question
1. What is the single most important question to ask before trusting a vendor's "98 percent accuracy" claim?
Show hint
Look at the direct answer and the B step.
Show answer
Model answer: What does "correct" mean, and on whose resumes was it measured? Without both answers, the number can't be acted on.
Multiple choice
2. Why did ScreenWise's real accuracy on Halstead's own applicants come in lower than the sales deck's number?
  • A. The model was retrained and got worse between the pitch and the pilot.
  • B. The vendor's benchmark was measured on a different, cleaner set of resumes than Halstead's real applicant pool.
  • C. Halstead's recruiters made more errors than usual during the pilot.
  • D. The sales engineer had misquoted the number by mistake.
Show hint
Look at the comparison sketch and the waterfall chart.
Show answer
B. The claim was true on the vendor's own eval set; it just wasn't a set that resembled Halstead's real applicants.
True or false
3. True or false: an 84 percent real-world accuracy estimate clearly beats human recruiters' own 89 percent agreement rate with each other.
  • True
  • False
Show hint
Look at the Nail the Sanity Check step.
Show answer
False. 84 is below 89, so the model isn't clearly outperforming a second human opinion, even before counting the risk of trusting an unverified vendor claim.
Fill in the blank
4. Fill in the blank: in the seafood plant story, the assumption that swings the real-world accuracy estimate the most is ___.
Show hint
Look at the tornado bar chart in Section 4.
Show answer
Conveyor motion blur. It swings the estimate by 16 points, more than lighting, species mix, and camera angle combined.
Short answer, apply it yourself
5. Think of a product claim you've seen, "99.9 percent uptime," "95 percent customer satisfaction." What question would you ask to turn it back into an equation with parts?
Show hint
Ask what counts as success, and measured over what population or time period.
Show answer
Model answer: Usually: what counts as the numerator (a success, by what definition), what counts as the denominator (total attempts, over what window), and whose data it was measured on.
Short answer, where it wouldn't matter
6. Name a part of the ScreenWise evaluation where this level of skepticism genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The vendor's internal model architecture and training method. That's their business to manage; the scrutiny belongs on the claim, not the method behind it.
Before you close the answer
Why this works
Tests whether you'll take a headline benchmark number at face value or break it back down into the eval set, the definition, and the class balance that actually produced it.
Follow-up traps
"What if the vendor won't share their eval set details?" Response: that refusal is itself informative, and it makes running your own pilot on real applicants non-negotiable rather than optional.

"Isn't 84 percent still pretty good?" Response: good compared to what; compared to human recruiters' own 89 percent agreement rate, it isn't yet the clear upgrade the sales deck implied.
If pressed
The pilot Halstead ran specifically oversampled borderline reject cases, not just an even random sample, because those are the resumes where a screening error actually costs a qualified candidate an interview, and an evenly sampled pilot would have hidden that weakness the same way the vendor's own benchmark did.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more