CalculationAdvancedQuality, Cost & Token Economics / Eval design for product teams / #5

How do you validate that your judge model agrees with human raters?

A judge model that agrees with humans 87.5 percent of the time sounds validated. Split that number by how hard each call actually was, and the real story is 94 and 68, not one number sitting in between.

The direct answer
Validate a judge model by scoring a held-out sample against real human raters, split by difficulty tier, and report agreement per tier instead of one blended figure. On a weekly 200-call sample at Ferrowick, the judge matched human raters on 94 percent of clear-cut calls but only 68 percent of the borderline ones, a gap the blended 87.5 percent completely hid. Find the rubric line where those misses cluster and route it to a human, because a high average agreement number is not proof the judge is safe everywhere. It is often proof mostly earned on the easy majority.
Do this, in order
  1. Grade the judge against real human raters on a held-out sample, split by tier, never one blended number.Why: a blended 87.5 percent hid a judge that was only right 68 percent of the time on the calls where humans themselves disagreed.
  2. Build the tier split from how humans actually scored the sample, not a guess.Why: a call only counts as borderline when two human raters disagreed with each other first, so the split measures real difficulty, not a convenient bucket.
  3. Check whether the judge's misses cluster on one rubric line or scatter evenly.Why: 11 of 16 borderline misses landed on the empathy and tone item alone, a real blind spot, not noise near a fuzzy line.
  4. Sanity check the number against a fresh human rater on the same hard calls, not just against itself.Why: a third senior rater agreed with the human majority on 78 percent of the borderline calls the judge only got right 68 percent of the time, so the gap is real.
  5. Route the tone and empathy line to a human second read only on calls near the pass cut-off.Why: reviewing every call by hand defeats the point of the judge; routing about 150 borderline-tone calls a week costs roughly 25 analyst hours, a bounded price for the one place a wrong score docks a real bonus.
  6. Re-run the tier split whenever the prompt or model version changes, not on a fixed calendar.Why: agreement drifted quietly the last time a model version shipped without a fresh validation pass.

How to answer this, stage by stage

Nobody's grading whether you can name the word "validation." They're grading whether you can show the arithmetic behind one number and say which part of it you trust least.

1
Scope it to one product before estimating anything in the abstract
Say it like this
"Let's ground this in one product. Scorepoint is a judge model that reads every support call transcript at Ferrowick and scores it against a six item quality rubric. Mikkel Bruvik owns whether that score can be trusted."
Why this works
An abstract validation question turns into a lecture fast. One product turns it into a real number problem.
2
Say your structure out loud before touching a single number
Say it like this
"I'm going to break the estimate into its parts, state where every number came from, give a range by tier instead of one blend, run a sanity check against a fresh human, then say which assumption would move the answer most."
Why this works
Tells the interviewer you have a method, not a vibe, before you've said a single figure.
3
Break the equation down before touching the sample
Say it like this
"Agreement rate is matches divided by calls sampled, and I'd compute that separately for each difficulty tier, not once for the whole sample. A blended number averages away exactly the gap I'm trying to find."
Why this works
Naming the equation first stops you backing into a reassuring number without knowing what it's built from.
4
Own every number and where it came from
Say it like this
"I'll assume a 200 call weekly sample, scored independently by two QA analysts. Where they agree, I call it clear-cut, 150 of the 200. Where they disagree, a senior rater breaks the tie, and I call it borderline, 50 of the 200."
Why this works
A number nobody can trace back to a source is a guess wearing a decimal point.
5
Report the range by tier, not one blended figure
Say it like this
"On this sample, the judge matched the human answer on 141 of 150 clear-cut calls, 94 percent, and on 34 of 50 borderline calls, 68 percent. Blended together that's 175 of 200, 87.5 percent. The honest answer is a range, 68 to 94, not one number in the middle."
Why this works
A single number this high is exactly the trap the question is testing. It was mostly earned on the easy majority.
6
Run the sanity check, twice
Say it like this
"I'd get a third senior rater to blind score the 50 borderline calls fresh. If they agree with the human answer more than the judge does, the gap is real, not just fuzziness anyone would hit. Then I'd pull the judge's 16 borderline misses and see if they land on one rubric line or scatter evenly."
Why this works
A smell test tells you whether 68 percent is bad, or just as good as a human gets on a genuinely hard call.
7
Name the one assumption that would swing the answer most
Say it like this
"The blend assumes the sample's mix, 75 percent clear-cut, 25 percent borderline, matches production. If the calls that actually get coached on skew more borderline, say 40 percent, the real blended agreement drops from 87.5 to about 83.6. That's the assumption I'd stress test first."
Why this works
A good estimator says which number they're least sure of. A bad one lets the reader assume they're all equally solid.
8
Close on the decision, not the arithmetic
Say it like this
"So: validate by tier, never by blend, and route the empathy and tone line to a human near the cut-off. The 87.5 percent was never the risk. The 68 percent hiding inside it was."
Why this works
Ending on the decision, not the last number crunched, is what makes this sound like judgment instead of a spreadsheet read aloud.

Let's learn

Every day, twelve QA analysts at Ferrowick could fully score about 180 calls between them. A full score means listening to the whole call and marking six rubric lines: greeting, resolution, compliance, empathy and tone, de-escalation, and wrap-up. Done right, that's about twelve minutes a call.

Ferrowick's phone lines, running support for a handful of retail and telecom clients, take in about 6,000 calls a day. 180 hand-scored calls is 3 percent of that. Nobody thought that was ideal. It was just what twelve people can do in a shift.

Then Ferrowick built Scorepoint, a judge model that reads a call's transcript and scores all six rubric lines itself, usually within ninety seconds of the call ending. Coverage jumped from 3 percent to 100 percent, overnight. That's the part everyone talked about in the first month. It is not the part that mattered.

Knowledge spark: what's a judge model? A second model whose only job is to read another model's, or a person's, output and grade it against a rubric. It stands in for a human reviewer, at a scale no human review team can match. Whether it's any good is a separate question from whether it's fast.

The real question came a few weeks later, once Scorepoint's score started feeding straight into agent coaching, and eventually into a small slice of quarterly bonus. Once a score touches real money, the question stops being whether Scorepoint saves QA time. It becomes whether Scorepoint's score matches what a human would have said, on this exact call.

The coverage number was never the risk. Whether Scorepoint's score matched a human's, on the calls a human would call close, was.
The choice that mattered Mikkel validated Scorepoint the way most teams do the first time. One 200 call sample, scored against a single QA supervisor, one blended agreement number, 87.5 percent, filed as "validated." That made sense when nothing depended on the score yet. It stopped making sense the day a wrong score could cost an agent part of a bonus.

Here's the arithmetic behind that 87.5. Each week, Ferrowick pulls 200 calls at random and has two QA analysts score them independently, listening to the recording, not reading Scorepoint's output first. When the two analysts agree with each other, that call is labeled clear-cut, 150 of the 200 most weeks. When they disagree, a senior rater breaks the tie, and the call is labeled borderline, 50 of the 200.

Scorepoint's score is then checked against that human answer, tier by tier, instead of once for the whole 200.

The build-up: matches out of calls sampled, by tier
200 100 0 150 sampled Clear-cut 50 sampled Borderline 200 sampled Blended 94% 68% 87.5%
Judge matched the human answerJudge missed it
Clear-cut: 141 of 150 matched, 94 percent. Borderline: 34 of 50 matched, 68 percent. Blended: 175 of 200, 87.5 percent, a figure that only exists once the easy majority gets folded back in.

At its worst, a validated-looking judge is worse than an unvalidated one, because nobody keeps checking a number that's already been signed off. The old manual process, thin as it was, never claimed more coverage than it had. A judge that reports one confident 87.5 percent and gets trusted for it can quietly keep failing on exactly the calls where a human ear would have caught the problem.

What I'd leave alone: the compliance rubric line, whether the agent stated the required identity check, genuinely doesn't need this treatment. It's a factual check, the sentence is either in the transcript or it isn't, and the judge agrees with humans there about 99 percent of the time. Spending audit hours re-checking that line would take time from the one line that's actually shaky.

The lesson: a number can be completely honest and still be the wrong number to trust. 87.5 percent really is what you get from blending the sample. It never said which quarter of the sample was carrying all the risk.

Now here is the same thing as a story

Read the long version below when you want to feel why a number this high still went wrong, not just be told that it did.

Mikkel Bruvik can read a rubric print-out and tell you, before he's finished the first line, which of the six items an agent is about to fail on.

He ran Ferrowick's QA team for four years before Scorepoint existed, back when 180 calls a day was the ceiling and everyone in the room knew it. He never pretended the manual sample covered everything. He just made sure the 180 they picked each day were spread fairly across shifts and agents, so nobody could say they'd been singled out.

Scorepoint launched in the spring. For the first two months, Mikkel ran the full validation himself: a 200 call sample every week, two raters, the tier split done by hand, a report with two numbers on it, not one. Ninety-four and sixty-eight, most weeks, give or take a couple of points. He read both lines every Monday.

By month four, the two numbers had barely moved. He started reading just the blended figure at the top of the report and skimming past the rest. By month six, the tier split was still being computed, quietly, by the script that built the report. Nobody was reading it. It just sat there under a header nobody scrolled to.

It came back on an ordinary Wednesday, not because a call had gone badly, but because a team lead mentioned, almost in passing, that three of her best agents had scored oddly low on empathy that quarter. All three were known for calming down furious customers. None of them had ever sat below the rubric's pass line before.

Mikkel almost said what everyone says: agents have off quarters. Then he remembered the tier split was still sitting under that header nobody scrolled to, and pulled it.

The blended number hadn't moved all quarter. What had moved, quietly, underneath it, was exactly the tier where three good agents were about to lose part of a bonus over a rubric line the judge had never been reliable on.

He pulled the sixteen borderline misses from that week's report and sorted them by which rubric line each one turned on. Eleven of sixteen were the same line: empathy and tone. Scorepoint reads only the transcript. It has no audio, no pace, no pause before a hard sentence. A flat, polite sentence reads the same to it whether the agent behind it sounded warm or sounded like they were reading off a script.

It was never really about whether 87.5 percent was a good number. It was a fine number, blended. What it hid was that on the one rubric line a human ear can catch and a transcript reader can't, the judge was wrong nearly a third of the time, and three real agents were sitting inside that third.

The decision that opened the door went back to month four, the week Mikkel quietly stopped reading the tier split every Monday. It made sense then. The two numbers had been stable for weeks, and a report nobody acts on eventually stops getting read all the way down. Nobody decided to stop watching the empathy line specifically. The whole tier just faded from a report that kept computing it anyway.

Run that Wednesday again with one change: the tier split stays in the coaching report itself, not buried under the blended headline, and any call where Scorepoint's empathy score sits within five points of the rubric's cut-off gets a human second read before it touches a bonus number. The same transcript still reads flat to Scorepoint. But the flag catches it in that Monday's report, two days before the quarterly bonus file locks, not the following quarter, after three agents have already filed a grievance.

What I'd tell myself, back in month four: a report that stops getting read isn't a report that stopped mattering. It's a report that's about to matter again, on a Wednesday nobody picked.

BOUND, the five letters behind that 68 and that 94

This isn't a story question. It's an estimation problem wearing a story's clothes, and BOUND is what turns a single number into an honest range.

BBreak it down. What's the actual equation?
Agreement rate equals matches divided by calls sampled, computed once per tier, not once for the whole set. Blended agreement is just the sum of both tiers' matches over the sum of both tiers' samples.
Say the equation before touching a number, or the number you land on is a guess with a confident face.
OOwn the numbers. Where did each one come from?
200 calls a week, two QA analysts scoring independently. Agree with each other: clear-cut, 150 of 200. Disagree: a senior rater breaks the tie, borderline, 50 of 200. Every figure traces back to a person, not a guess.
This is also where the rejected alternative sits: validating against one rater alone, see below.
UUse a range, not one number.
94 percent on clear-cut, 68 percent on borderline. The honest answer to "does the judge agree with humans" is that range, not the 87.5 percent blend sitting between them.
A single number this high is exactly what makes a bad validation look finished.
NNail the sanity check. Does the number survive being compared to something real?
A fresh senior rater, blind-scoring only the 50 borderline calls, agreed with the human majority on 39 of them, 78 percent. The judge only matched on 34, 68 percent. So the gap is real, not just the ordinary fuzziness of a hard call. Then check whether the misses cluster: 11 of 16 borderline misses shared the same rubric line, empathy and tone, far more than the roughly 3 you'd expect if misses were spread evenly across six rubric lines.
This is the hardest step, and the one most answers skip. A high number with no sanity check is a guess dressed as a result.
DDirection. Which assumption would move the answer most?
The validation sample assumes 75 percent clear-cut, 25 percent borderline, matching a random weekly pull. But the calls that actually get coached on skew harder. If the real mix runs 60/40 instead, the true blended agreement drops from 87.5 to about 83.6, a bigger swing than any other assumption in this estimate.
Naming the shakiest assumption out loud is what a good estimator does that a bad one skips.
What moves the estimate most, if the assumption behind it is wrong
Production tier mix ~4.0 pts Ground truth, 1 vs 2 raters ~2.0 pts Model version since last audit ~1.5 pts
Biggest swingMedium swingSmaller swing
Estimated points of blended agreement moved if each assumption turns out wrong. The production tier mix is the assumption worth stress testing first, because it swings the answer twice as much as the next one.

Three things worth stating directly, since this is where the real judgment sits. The alternative Mikkel's team tried first, and later dropped, was validating Scorepoint against one QA supervisor's rating alone, since it only needed one person's time. It lost because a single rater's own quirks become the ground truth, so a low agreement number can't tell you whether the judge is wrong or that one rater just scores tone stricter than everyone else. A two-rater majority, with a tie-break for the tier split, doesn't have that problem. The AI-specific failure mode worth naming by name is a transcript-only judge missing paralinguistic tone. Scorepoint reads words, not sound, so it cannot hear warmth or coldness in an agent's voice, only what was said. The guardrail is a routing rule: any call where the empathy and tone sub-score sits within five points of the rubric's cut-off goes to a human second read before it reaches a coaching report, not just when someone complains. That guardrail isn't free. It sends roughly 150 calls a week to a human at about ten minutes each, twenty-five analyst hours nobody had budgeted, spent only on the one rubric line proven to be the judge's blind spot, not spread thin across all six. And the bar for letting Scorepoint act alone was never 100 percent agreement. A system reading 6,000 transcripts a day can't promise that. It's a tier-specific bar, cleared separately and checked weekly, not one blended pass mark standing in for six different questions.

And if you want to be sure it really works, try it somewhere else

Same five letters, a veterinary telehealth network instead of a call center, nothing about contact center rubrics anywhere in sight.

Coalspring is a veterinary telehealth network. A vet tech chats with a pet owner, looks over photos the owner uploads, and writes a triage note recommending home care, a same-day visit, or an emergency vet now. Coalspring's judge model, built the same way as Scorepoint, reads that note and scores it against a five item documentation rubric. Kasimir Petitjean runs the validation.

The build-up: on Coalspring's own 150 call weekly sample, split the same way, by whether two supervising vets agreed with each other first, the judge matched the human answer on 120 of 125 clear-cut notes, 96 percent, and on 15 of 25 borderline notes, 60 percent. Blended, that's 135 of 150, 90 percent, a number that again looks fine sitting alone on a dashboard.

The decision Kasimir would take back Trusting the phrase "photo reviewed, consistent with mild irritation" as proof a photo was actually opened, just because the judge scores documentation completeness from the note's own wording.

The sanity check: Kasimir pulled the ten borderline misses and found eight of them shared one thing. The triage note claimed a photo had been reviewed, but the platform's own open-event log showed nobody had actually opened it. The judge scores "photo reviewed" from the sentence, not from any real signal that a photo was opened. A tech who writes that phrase out of habit reads exactly the same to the judge as one who genuinely looked.

Same rank as before: validate by tier, then check where the misses cluster before trusting the blend. The fix is the same shape too: route any note claiming a photo review to a check against the platform's own open-event log, and only let the judge act alone once that log confirms the claim.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: validate by tier, never by blend, and route the one rubric line where misses cluster to a human near the cut-off.
Cost: there's no budget this quarter for both a bigger validation sample and the human-routing rule. The routing rule wins. Twenty-five hours a week on a real blind spot beats a thousand-call sample that just re-measures the same 87.5 percent more precisely.
The model got better, for real: say Scorepoint's overall accuracy climbs to 91 percent next quarter. That's not proof the borderline tier climbed with it. The clear-cut tier could have gotten even easier for the judge while the empathy line stayed exactly as blind as before.

Where people run it wrong.
They read one blended number as proof there's no gap anywhere, and never ask the sample to earn that number tier by tier.
They fix a low borderline score by lowering the pass bar, instead of finding where the judge is actually blind.
They validate once at launch and never again, so a prompt change ships without anyone re-running the split.

How to use it live. Say the real question out loud before answering it: "is one sample-wide percentage actually answering whether this is safe everywhere, or does it need splitting first." That buys a beat to think instead of quoting a number you haven't checked the shape of.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: show the arithmetic, own the assumptions. Built for estimation and sizing questions, not stories about a person's habit.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Mikkel Bruvik, QA analytics lead at Ferrowick, a contact center outsourcer. Ran the QA team by hand for four years before Scorepoint existed.
3 · THE OLD SHORTCUT
What did Mikkel stop checking, because the top-line number always looked fine?
Tap to flip
ANSWER
He stopped reading the weekly tier split and started reading only the blended 87.5 percent figure at the top of the report.
4 · THE HIDDEN GAP
What two numbers does the blended 87.5 percent hide?
Tap to flip
ANSWER
94 percent agreement on clear-cut calls, and only 68 percent on borderline calls, the ones where the two human raters had disagreed with each other.
5 · THE OLD DECISION
What decision would Mikkel take back?
Tap to flip
ANSWER
Validating Scorepoint against one QA supervisor's rating instead of a two-rater majority with a tier split, from the very first validation pass.
6 · THE NUMBER
Fill in the blank: of the 16 borderline misses in that week's report, ___ of them landed on the same rubric line, empathy and tone.
Tap to flip
ANSWER
11. Far more than the roughly 3 you'd expect if misses were spread evenly across all six rubric lines, which is what marks it as a real blind spot.
7 · THE REPLAY
Same bad Wednesday, new design, what changes?
Tap to flip
ANSWER
The tier split stays visible in the coaching report itself, and any call with an empathy score within five points of the cut-off routes to a human. It catches the miss two days before the quarterly bonus file locks, not the following quarter, after a grievance is filed.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the matching blind spot?
Tap to flip
ANSWER
Coalspring, a veterinary telehealth network. Same shape of blind spot: the judge trusts a note's claim that a photo was reviewed instead of checking a real signal that it actually was.

Check yourself Score: 0 / 0

Fill in the blank
1. The judge matched the human answer on ___ percent of clear-cut calls and only ___ percent of borderline calls.
Show hint
Look at the two numbers stacked inside the blended 87.5 percent in the build-up chart.
Show answer
94 percent, then 68 percent. Blended together that's 175 of 200, 87.5 percent, but the range those two numbers describe is the honest answer.
True or false
2. True or false: because the judge's blended agreement rate was 87.5 percent, a fresh human rater scoring the same borderline calls would also land somewhere close to 87.5 percent.
  • True
  • False
Show hint
Check the sanity-check step. It re-scores only the borderline calls, not the whole 200.
Show answer
False. A fresh senior rater, scoring only the 50 borderline calls, agreed with the human majority 78 percent of the time. That's the tier-specific ceiling. The blended 87.5 percent only exists once the easy majority of clear-cut calls gets folded back in.
Multiple choice
3. Why does splitting the validation sample by tier matter more than reporting one blended agreement number?
  • A. Because tier splits make the report look more thorough to auditors.
  • B. Because a blended number is dragged up by the easy majority of clear-cut calls and hides how the judge performs on the hard ones.
  • C. Because the judge model runs faster once the sample is split.
  • D. Because human raters refuse to score without a tier label.
Show hint
Think about what 150 easy calls scoring 94 percent does to a 200-call blend.
Show answer
B. The clear-cut tier is both the bigger group and the easier one, so it pulls the blend up. A high blended number is often mostly earned on the easy majority, exactly the trap this question is testing.
Short answer, name the rejected alternative
4. What alternative did Mikkel's team try first for validating Scorepoint, and why did it lose?
Show hint
Look at the O step in the framework recap, where the numbers' sources get named.
Show answer
Model answer: Validating against one QA supervisor's rating alone, since it only needed one person's time. It lost because a single rater's own quirks become the ground truth, so a low agreement number can't tell you whether the judge is wrong or that one rater just scores tone stricter than everyone else.
Short answer, apply it yourself
5. Pick an AI product you use yourself that gets checked against human judgment somehow. Name one place its agreement number might be hiding a tier-specific gap, and how you'd check.
Show hint
Think of a product that reports one overall accuracy number for cases that are actually a mix of easy and genuinely hard calls.
Show answer
Model answer: An email spam filter might report 99 percent accuracy overall, but that number is mostly obvious spam and obvious real mail. I'd pull a sample of the emails a person would call "maybe spam", promotional mail from a real sender, and check the filter's accuracy on just those, separately from the easy cases.
Multiple choice
6. If the calls that actually get coached on in production run 40 percent borderline instead of the validation sample's 25 percent, what happens to the real blended agreement rate?
  • A. It stays at 87.5 percent, since the tier-specific numbers don't change.
  • B. It drops to about 83.6 percent, since more of the harder, lower-agreement tier is now in the mix.
  • C. It rises, since borderline calls are naturally easier for a transcript-only judge.
  • D. It becomes impossible to estimate without a brand new sample.
Show hint
Redo the blend with 60 percent at 94 percent and 40 percent at 68 percent instead of 75/25.
Show answer
B. 0.6 times 94 plus 0.4 times 68 comes out to about 83.6 percent. The tier numbers don't move, but the mix they're blended in does, and that's the D step: name which assumption swings the answer most.
Before you close the answer
Why this works
Tests whether you'll trust a single validation number or ask what it's built from. Most candidates report the blended figure and stop, exactly the trap the question is testing.
Follow-up traps
"Isn't 87.5 percent already a strong result for an LLM judge?" Response: on its own, maybe. But it's an average of a 94 and a 68, and only the 68 is where a wrong score costs an agent real money.

"Why not just retrain the judge on more empathy examples and move on?" Response: worth trying, but it doesn't ship this week, and until it's validated again, the routing guardrail is what protects the next coaching cycle, not a retrain still in progress.
If pressed
The real production rule was never a single blended pass bar. It was tier-specific: 90 percent or higher on clear-cut before the judge could act alone, and any borderline-tier call within five points of the rubric cut-off routed to a human, since a system scoring 6,000 transcripts a day can't promise zero misses on a probabilistic read.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more