How do you validate that your judge model agrees with human raters?
A judge model that agrees with humans 87.5 percent of the time sounds validated. Split that number by how hard each call actually was, and the real story is 94 and 68, not one number sitting in between.
- Grade the judge against real human raters on a held-out sample, split by tier, never one blended number.Why: a blended 87.5 percent hid a judge that was only right 68 percent of the time on the calls where humans themselves disagreed.
- Build the tier split from how humans actually scored the sample, not a guess.Why: a call only counts as borderline when two human raters disagreed with each other first, so the split measures real difficulty, not a convenient bucket.
- Check whether the judge's misses cluster on one rubric line or scatter evenly.Why: 11 of 16 borderline misses landed on the empathy and tone item alone, a real blind spot, not noise near a fuzzy line.
- Sanity check the number against a fresh human rater on the same hard calls, not just against itself.Why: a third senior rater agreed with the human majority on 78 percent of the borderline calls the judge only got right 68 percent of the time, so the gap is real.
- Route the tone and empathy line to a human second read only on calls near the pass cut-off.Why: reviewing every call by hand defeats the point of the judge; routing about 150 borderline-tone calls a week costs roughly 25 analyst hours, a bounded price for the one place a wrong score docks a real bonus.
- Re-run the tier split whenever the prompt or model version changes, not on a fixed calendar.Why: agreement drifted quietly the last time a model version shipped without a fresh validation pass.
How to answer this, stage by stage
Nobody's grading whether you can name the word "validation." They're grading whether you can show the arithmetic behind one number and say which part of it you trust least.
Let's learn
Every day, twelve QA analysts at Ferrowick could fully score about 180 calls between them. A full score means listening to the whole call and marking six rubric lines: greeting, resolution, compliance, empathy and tone, de-escalation, and wrap-up. Done right, that's about twelve minutes a call.
Ferrowick's phone lines, running support for a handful of retail and telecom clients, take in about 6,000 calls a day. 180 hand-scored calls is 3 percent of that. Nobody thought that was ideal. It was just what twelve people can do in a shift.
Then Ferrowick built Scorepoint, a judge model that reads a call's transcript and scores all six rubric lines itself, usually within ninety seconds of the call ending. Coverage jumped from 3 percent to 100 percent, overnight. That's the part everyone talked about in the first month. It is not the part that mattered.
The real question came a few weeks later, once Scorepoint's score started feeding straight into agent coaching, and eventually into a small slice of quarterly bonus. Once a score touches real money, the question stops being whether Scorepoint saves QA time. It becomes whether Scorepoint's score matches what a human would have said, on this exact call.
Here's the arithmetic behind that 87.5. Each week, Ferrowick pulls 200 calls at random and has two QA analysts score them independently, listening to the recording, not reading Scorepoint's output first. When the two analysts agree with each other, that call is labeled clear-cut, 150 of the 200 most weeks. When they disagree, a senior rater breaks the tie, and the call is labeled borderline, 50 of the 200.
Scorepoint's score is then checked against that human answer, tier by tier, instead of once for the whole 200.
At its worst, a validated-looking judge is worse than an unvalidated one, because nobody keeps checking a number that's already been signed off. The old manual process, thin as it was, never claimed more coverage than it had. A judge that reports one confident 87.5 percent and gets trusted for it can quietly keep failing on exactly the calls where a human ear would have caught the problem.
What I'd leave alone: the compliance rubric line, whether the agent stated the required identity check, genuinely doesn't need this treatment. It's a factual check, the sentence is either in the transcript or it isn't, and the judge agrees with humans there about 99 percent of the time. Spending audit hours re-checking that line would take time from the one line that's actually shaky.
The lesson: a number can be completely honest and still be the wrong number to trust. 87.5 percent really is what you get from blending the sample. It never said which quarter of the sample was carrying all the risk.
Now here is the same thing as a story
Read the long version below when you want to feel why a number this high still went wrong, not just be told that it did.
Mikkel Bruvik can read a rubric print-out and tell you, before he's finished the first line, which of the six items an agent is about to fail on.
He ran Ferrowick's QA team for four years before Scorepoint existed, back when 180 calls a day was the ceiling and everyone in the room knew it. He never pretended the manual sample covered everything. He just made sure the 180 they picked each day were spread fairly across shifts and agents, so nobody could say they'd been singled out.
Scorepoint launched in the spring. For the first two months, Mikkel ran the full validation himself: a 200 call sample every week, two raters, the tier split done by hand, a report with two numbers on it, not one. Ninety-four and sixty-eight, most weeks, give or take a couple of points. He read both lines every Monday.
By month four, the two numbers had barely moved. He started reading just the blended figure at the top of the report and skimming past the rest. By month six, the tier split was still being computed, quietly, by the script that built the report. Nobody was reading it. It just sat there under a header nobody scrolled to.
It came back on an ordinary Wednesday, not because a call had gone badly, but because a team lead mentioned, almost in passing, that three of her best agents had scored oddly low on empathy that quarter. All three were known for calming down furious customers. None of them had ever sat below the rubric's pass line before.
Mikkel almost said what everyone says: agents have off quarters. Then he remembered the tier split was still sitting under that header nobody scrolled to, and pulled it.
He pulled the sixteen borderline misses from that week's report and sorted them by which rubric line each one turned on. Eleven of sixteen were the same line: empathy and tone. Scorepoint reads only the transcript. It has no audio, no pace, no pause before a hard sentence. A flat, polite sentence reads the same to it whether the agent behind it sounded warm or sounded like they were reading off a script.
It was never really about whether 87.5 percent was a good number. It was a fine number, blended. What it hid was that on the one rubric line a human ear can catch and a transcript reader can't, the judge was wrong nearly a third of the time, and three real agents were sitting inside that third.
The decision that opened the door went back to month four, the week Mikkel quietly stopped reading the tier split every Monday. It made sense then. The two numbers had been stable for weeks, and a report nobody acts on eventually stops getting read all the way down. Nobody decided to stop watching the empathy line specifically. The whole tier just faded from a report that kept computing it anyway.
Run that Wednesday again with one change: the tier split stays in the coaching report itself, not buried under the blended headline, and any call where Scorepoint's empathy score sits within five points of the rubric's cut-off gets a human second read before it touches a bonus number. The same transcript still reads flat to Scorepoint. But the flag catches it in that Monday's report, two days before the quarterly bonus file locks, not the following quarter, after three agents have already filed a grievance.
What I'd tell myself, back in month four: a report that stops getting read isn't a report that stopped mattering. It's a report that's about to matter again, on a Wednesday nobody picked.
BOUND, the five letters behind that 68 and that 94
This isn't a story question. It's an estimation problem wearing a story's clothes, and BOUND is what turns a single number into an honest range.
Three things worth stating directly, since this is where the real judgment sits. The alternative Mikkel's team tried first, and later dropped, was validating Scorepoint against one QA supervisor's rating alone, since it only needed one person's time. It lost because a single rater's own quirks become the ground truth, so a low agreement number can't tell you whether the judge is wrong or that one rater just scores tone stricter than everyone else. A two-rater majority, with a tie-break for the tier split, doesn't have that problem. The AI-specific failure mode worth naming by name is a transcript-only judge missing paralinguistic tone. Scorepoint reads words, not sound, so it cannot hear warmth or coldness in an agent's voice, only what was said. The guardrail is a routing rule: any call where the empathy and tone sub-score sits within five points of the rubric's cut-off goes to a human second read before it reaches a coaching report, not just when someone complains. That guardrail isn't free. It sends roughly 150 calls a week to a human at about ten minutes each, twenty-five analyst hours nobody had budgeted, spent only on the one rubric line proven to be the judge's blind spot, not spread thin across all six. And the bar for letting Scorepoint act alone was never 100 percent agreement. A system reading 6,000 transcripts a day can't promise that. It's a tier-specific bar, cleared separately and checked weekly, not one blended pass mark standing in for six different questions.
And if you want to be sure it really works, try it somewhere else
Same five letters, a veterinary telehealth network instead of a call center, nothing about contact center rubrics anywhere in sight.
Coalspring is a veterinary telehealth network. A vet tech chats with a pet owner, looks over photos the owner uploads, and writes a triage note recommending home care, a same-day visit, or an emergency vet now. Coalspring's judge model, built the same way as Scorepoint, reads that note and scores it against a five item documentation rubric. Kasimir Petitjean runs the validation.
The build-up: on Coalspring's own 150 call weekly sample, split the same way, by whether two supervising vets agreed with each other first, the judge matched the human answer on 120 of 125 clear-cut notes, 96 percent, and on 15 of 25 borderline notes, 60 percent. Blended, that's 135 of 150, 90 percent, a number that again looks fine sitting alone on a dashboard.
The sanity check: Kasimir pulled the ten borderline misses and found eight of them shared one thing. The triage note claimed a photo had been reviewed, but the platform's own open-event log showed nobody had actually opened it. The judge scores "photo reviewed" from the sentence, not from any real signal that a photo was opened. A tech who writes that phrase out of habit reads exactly the same to the judge as one who genuinely looked.
Same rank as before: validate by tier, then check where the misses cluster before trusting the blend. The fix is the same shape too: route any note claiming a photo review to a check against the platform's own open-event log, and only let the judge act alone once that log confirms the claim.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: validate by tier, never by blend, and route the one rubric line where misses cluster to a human near the cut-off.
Cost: there's no budget this quarter for both a bigger validation sample and the human-routing rule. The routing rule wins. Twenty-five hours a week on a real blind spot beats a thousand-call sample that just re-measures the same 87.5 percent more precisely.
The model got better, for real: say Scorepoint's overall accuracy climbs to 91 percent next quarter. That's not proof the borderline tier climbed with it. The clear-cut tier could have gotten even easier for the judge while the empathy line stayed exactly as blind as before.
Where people run it wrong.
They read one blended number as proof there's no gap anywhere, and never ask the sample to earn that number tier by tier.
They fix a low borderline score by lowering the pass bar, instead of finding where the judge is actually blind.
They validate once at launch and never again, so a prompt change ships without anyone re-running the split.
How to use it live. Say the real question out loud before answering it: "is one sample-wide percentage actually answering whether this is safe everywhere, or does it need splitting first." That buys a beat to think instead of quoting a number you haven't checked the shape of.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just retrain the judge on more empathy examples and move on?" Response: worth trying, but it doesn't ship this week, and until it's validated again, the routing guardrail is what protects the next coaching cycle, not a retrain still in progress.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Eval design for product teams
- #1 What makes an eval product-relevant rather than research-relevant?
- #2 Design an eval for a feature that drafts email replies.
- #3 How do you decide between automated evals and human review?
- #4 Explain the tradeoffs of LLM-as-judge for a product team.
- #6 Describe a rubric that a non-technical reviewer could apply consistently.
- #7 What is inter-rater reliability and why should a PM care?