Critique a quality report that presents a single accuracy number.
A single accuracy number tells you the model is fine on average. Averages are exactly where the expensive kind of wrong likes to hide.
- Reject the one number, and ask for the miss rate near the decision cutoff first.Why: that's the only score range where a miss actually flips a real hire or no-hire call.
- Ask for the miss rate broken out by rubric dimension and candidate segment next.Why: a rare but severe bias hides longest inside a healthy-looking average, and it does real harm before the average ever moves.
- Check whether the eval set was even graded at that level of detail before asking for the breakdown.Why: you can't report a cut that was never labeled. The real gap usually sits upstream in how the eval set was built, not in the slide.
- Take the free cut first: filter the transcripts already graded by score band, no new labeling needed.Why: it costs nothing and tells you in an afternoon whether there's a decision-band problem worth paying to chase.
- Only commission a deeper, tagged re-grade once the free cut shows a real signal.Why: tagging by dimension and segment is slow and costs real people-days. Spend it once you know there's something to find.
- Track the breakdown every quarter after that, not as a one-time slide.Why: a single good number this quarter says nothing about whether it stays good, and a rising average can hide the same gap for another year.
How to answer this, stage by stage
Nobody's grading whether you know accuracy can be misleading. They're grading whether you can rank what's missing by what it would cost to have missed it, out loud, on the spot.
Let's learn
Yardstick reads a candidate's interview transcript and scores it against a hiring rubric: technical depth, communication clarity, and culture signal, each out of five. The three combine into one fit score, and above a cutoff of 3.0 the tool hands the hiring team a recommend call. Below it, a no-recommend.
Before Yardstick, a hiring coordinator read every transcript by hand and wrote up notes, about 40 minutes a transcript. With the tool doing the first pass, that dropped to about 8 minutes of review per transcript, since the coordinator was checking the tool's call instead of building one from nothing.
Every quarter, a review panel hand-grades 1,000 transcripts and checks whether Yardstick's recommend call matches what the panel itself would have called. For four straight quarters the number held: 92, 93, 91, 92 percent agreement. Steady. The kind of number that makes a slide feel finished.
Then a first-week data analyst sat in on a quarterly review and asked a plain question: which rubric dimension is that 92 percent weakest on? Rilla went to answer it and found she couldn't. The panel had only ever graded each transcript pass or fail against the final call. Nobody had tagged which dimension caused a disagreement, or what kind of candidate it happened to.
Rilla's team took the free step first: they filtered the 1,000 already-graded transcripts by score band, no new grading needed. The overall miss rate held at 8 percent, matching the 92 percent everyone already trusted. But inside the decision band, scores between 2.6 and 3.4, the miss rate was 21 percent. That's the only zone where a miss actually changes a hire.
A deeper, tagged re-grade of 200 transcripts, two people, nine days, found where inside the decision band the 21 percent was concentrated. On the communication-clarity dimension, candidates who weren't native English speakers and who paused or restructured a sentence mid-answer got marked down 34 percent of the time. The model's own written reason read confident and fluent: "hesitant, unclear communication." A human grader reading the same transcript called it competent, just phrased in a second language.
Here's the part that costs the most. This is a confident-wrongness problem, the kind people also call drift: the model hadn't seen enough of this speech pattern to read it correctly, so it guessed with total certainty and got it backwards. Because communication clarity is one of three scores feeding the final call, this pattern was enough on its own to push some of those candidates below the 3.0 cutoff. Nobody saw it. The topline number never moved, because it was never built to notice one dimension going wrong for one kind of person.
What I'd leave alone: Yardstick also shows interviewers a live "quick pulse" score during the call itself, a soft signal next to their own notes, never used to auto-advance or auto-reject anyone. A blended number is genuinely fine there. A person reads every transcript for that score anyway, so there's no hidden decision for a rare miss to quietly control.
The lesson: a report built on one number isn't lying. It's just answering a question nobody actually asked, which is "is the model fine on average," instead of the one that matters, which is "is the model fine where it decides something."
Now here is the same thing as a story
Read this when you want to feel why the free cut mattered, not just know that it did.
Rilla Tazwell has run quality for Yardstick for three years. She can read a quarterly number in about four seconds and tell you whether it's worth a second look.
Most quarters, the review meeting was short. The panel's number came in, someone read it out, 92 percent, and the room moved to the next item on the agenda. Nobody argued with it. It had earned the right not to be argued with, four quarters running.
The first-week analyst who asked the question wasn't trying to make trouble. She'd built exactly this kind of report at her last job, and she'd learned to ask one thing before trusting any of them: what's the number weakest on? She asked it the way you'd ask someone their name. Rilla opened her mouth to answer and realized, mid-sentence, that she didn't actually know.
That was the whole trigger. One plain question, from someone who'd been in the building four days.
Rilla's first instinct was to commission a full re-grade, every transcript from the past year, tagged properly this time, dimension by dimension, segment by segment. Someone had already opened a spreadsheet to scope the headcount by lunch. It would have taken two people something like six weeks and touched roughly 4,000 transcripts.
She asked for an afternoon instead.
The 1,000 transcripts from the current quarter were already graded pass or fail. Nobody had asked to see them any other way. Rilla's team sorted them by score band instead, which cost nothing, no new grading, just a different cut of the same spreadsheet. The overall miss rate matched the 92 percent exactly. The decision band, scores between 2.6 and 3.4, came back at 21 percent.
That number bought the real investigation. A targeted, tagged re-grade of 200 transcripts from inside that band, two people, nine days, not six weeks, found where the 21 percent was concentrated: candidates answering in a second language, marked down on communication clarity, a third of the time, for a pattern that had nothing to do with whether they could do the job.
It was never really about whether the topline number was wrong. It wasn't. Ninety-two percent was true the whole time. What it was never built to say was where the model's confidence and its correctness had quietly come apart, and for whom.
The decision that opened this gap went back to the year the eval pipeline was built, when Yardstick covered one role and a small review panel eyeballed every disagreement personally each week. Someone asked whether it was worth tagging each miss by dimension and segment. The answer, at the time, was no. It would have doubled the grading time for a number small enough to just read by hand. Nobody planned for forty roles and a review meeting too short to argue with anything.
Run the same quarter again with one change: the eval set gets tagged by rubric dimension and candidate segment as it's graded, not after someone asks a question about it. The 21 percent in the decision band shows up on the slide itself, no free filter required, no six-week scramble. Any transcript where communication clarity alone swings the score across the cutoff gets a second human read before the call goes out, instead of an automatic no-recommend. The next quarter, that rule caught 14 transcripts in the decision band from exactly this pattern. Nine of them flipped from no-recommend to recommend once a person read them.
One design trusted a single number to speak for the whole model. The other asks the number what it's standing on before believing it.
What I'd tell myself, back in that first-year meeting: a report that only ever grades right or wrong can never grow into a report that says where or for whom. That has to be decided before the grading starts, not added later when someone finally asks.
O-R-D-E-R, ranked for a slide with one number on it
This isn't a story about a bad slide. It's ORDER, run on a quality report instead of a backlog, ranking what a blended number hides by what it costs to have missed it.
Three things worth stating plainly, since this is where the real judgment sits. The rejected alternative was the full re-grade, all 4,000 transcripts from the past year, tagged properly, before shipping any fix. It lost because it traded speed for completeness nobody needed yet, when a free filter on data already in hand answered the only question that mattered first: is there a signal worth chasing at all. The AI-specific failure worth naming by name is confident wrongness from a pattern the model barely saw in training or in its own eval set, sometimes called drift: it read a second-language speech pattern as a competence problem, and wrote a fluent, certain reason for a call that was backwards. The guardrail is concrete: any transcript where one rubric dimension alone swings the score across the 3.0 cutoff gets a second human read before the call ships, instead of an automatic decision. That guardrail isn't free. Tagging the eval set by dimension and segment going forward slows the grading pipeline and costs real people-hours every quarter, a real quality-versus-speed trade the team accepted on purpose rather than backfilling the whole history at once. And the bar isn't zero misses in the decision band, a scoring model can't promise that. It's a decision-band miss rate under 10 percent on a rolling 200-transcript stratified sample, checked every quarter, with no single rubric dimension responsible for more than a third of the misses inside that band.
And if you want to be sure it really works, try it somewhere else
Same five letters, a claims-triage tool this time, with nothing about hiring or transcripts anywhere in sight.
Farroway Claims runs an AI tool that scores incoming insurance claims and sorts them into an auto-approve fast lane or a route-to-adjuster lane. Vendrick Costigan runs quality for that tool.
O, outcome. Let a claims quality lead decide whether to trust a specific auto-approve or route-to-adjuster call, not produce a slide that looks fine in a board deck.
R, reversibility. A miss near the payout cutoff already sent money out the door or delayed a real claim before anyone noticed. A miss on an obviously simple or obviously complex claim never would have changed the outcome either way.
D, dependency. A routine regulatory audit asked for the miss rate on claims within 500 dollars of the payout cutoff. Nobody could produce it, because the eval set had only ever been graded match or no-match against the adjuster's final call, same gap as Yardstick's, a different industry.
E, evidence. The free cut: filter the already-graded claims by how close they sat to the cutoff. Overall miss rate held at 6 percent. Inside the near-cutoff band, it was 19 percent.
R, rank. A deeper, tagged re-grade found where: water-damage claims from manufactured-home policies, a small volume segment, missed 29 percent of the time on a documentation-completeness dimension. The model read blurrier photos, common from older phones in that segment, as incomplete documentation. Adjusters reading the same photos called them fine.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the near-cutoff cut, don't wait to describe a full segment re-grade before saying anything.
Cost: there's no budget yet for a dedicated re-grading team. Pull the free cut, filter what's already graded by score band or by distance from the cutoff, and say plainly that's as far as it goes until proper tags exist.
The model got better, for real: say overall accuracy climbed from 92 to 96 percent this quarter. That's not proof the decision band or any one segment improved. A rising average can hide the same buried gap for another year, just under a better-looking number.
Where people run it wrong.
They read a stable topline number as proof nothing needs checking, and never ask for a single extra cut.
They demand a full historical re-grade of everything before shipping any fix, instead of the free band filter that would have told them in an afternoon whether it was worth it.
They find the bad segment and retrain the model immediately, without asking whether the real fix is a guardrail, a second human read on near-cutoff segment cases, rather than a blind retrain.
How to use it live. Say one line before agreeing whether a number is good or bad: "What's inside that number?" It buys you a beat to think, and it's the exact question the whole answer is built around.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the decision band is a small share of total volume, does it really deserve to go first?" Response: yes, because size isn't what earns the rank, consequence is. Every transcript inside that band is one where the tool's call is genuinely load-bearing, unlike the other transcripts, where a human would land in the same place regardless of what the model said.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?