What is inter-rater reliability and why should a PM care?
Ask what inter-rater reliability is, and most answers stop at the definition. The real answer is what happens to every other number once that reliability quietly slips.
- Track rater to rater agreement as its own number, measured with raters working apart.Why: it is the reliability of the yardstick, not the model, and a model can hit 90 percent against a yardstick that has already gone soft.
- Set real thresholds for what you do at each band, not just a dashboard tile.Why: 85 percent and up, trust the composite score. 75 to 84, spot check before acting. Under 75, freeze anything that score was used to justify.
- Only count agreement measured before raters have talked to each other.Why: raters comparing notes first inflates the number without the rubric getting one word clearer, which looks like a fix and is not one.
- Run a standing, blind double scored sample every review cycle, not once at launch.Why: reliability erodes quietly as raters rotate, and a one time calibration check goes stale the day a new rater starts.
- When the score drops below the freeze line, fix the rubric and the raters before trusting any number built on it again.Why: a composite score computed against a broken ruler just launders noise into something that looks measured.
- Budget real rater time for this, on purpose.Why: checking agreement costs hours that would otherwise grow the golden set, and a smaller set that means one steady thing beats a bigger one nobody can trust.
How to answer this, stage by stage
Nobody is grading whether you know the textbook definition. They are grading whether you can turn that definition into a threshold you would actually act on.
Let's learn
Rubricast reads a student's essay draft and hands back a suggested letter grade, plus three lines of feedback a teacher can send along or fix first.
Before Rubricast, a teacher graded every essay alone, comparing gut instinct to a rubric taped inside the planner. A full class set of thirty essays took about four hours on a Sunday night.
Rubricast reads the same thirty essays in about ninety seconds and hands back a grade and feedback on each one. The teacher reads it over, agrees, edits, or overrides, and sends it back. Most Sunday nights got their four hours back.
Here is the part that matters. Rubricast's own report card looked fine the entire time, 88 to 91 percent match with the golden label, cycle after cycle. The real trouble was never Rubricast getting worse. It was that the golden label itself had quietly stopped meaning one steady thing.
At its worst, a grading tool teachers stop trusting is worse than no tool at all. The old way, slow as it was, at least meant one adult's judgment stayed consistent essay to essay. Rubricast, once teachers stopped trusting the suggested grade, got overridden constantly, which cost more time correcting it than grading from scratch would have.
What I'd leave alone: the mechanics part of the rubric, spelling and grammar, does not need this. A run on sentence is a run on sentence, and our two raters agree on it almost every time without discussion. Spending review time re-checking agreement there would take time away from the criteria that are actually judgment calls, like whether the evidence in a paragraph is sufficient.
The lesson: a number can be completely honest and still be checked against something shaky. Eighty nine percent match told the team Rubricast was fine. It never told them whether the person behind that ninety percent agreed with themselves, let alone with someone else grading the same page.
Now here is the same thing as a story
Read the short version above when you are being asked this cold. Read this one when you want to feel why a PM has to actually check, not just define.
When Rubricast launched, Feodora Ondrasek pulled the calibration report herself every review cycle. It was a small thing, two trained ex teachers each grading the same 150 essays alone, then a spreadsheet showing how often they landed on the same grade. She read it, noted it, moved on.
For the first two cycles it read 84, then 81 percent. Fine, everyone said. Grading is a judgment call, not arithmetic, some disagreement is normal. Feodora agreed. She had graded essays herself for three years before this job.
By the third cycle, the number was busy competing for her attention with a dozen other things. The top line tile, Rubricast's match rate against the golden label, sat steady at 88 to 91 percent, the same green box it had shown since launch. She started skimming the calibration report instead of reading it. By the fourth cycle, she stopped opening it at all. Nothing in the top line tile ever told her to.
It came back on an ordinary Tuesday, during a routine internal review that had nothing to do with grading accuracy. Someone on the compliance side, checking whether QA processes were being followed on schedule, noticed the calibration report had not been reviewed in over a year and pulled it fresh.
Agreement between the two raters had reached 71 percent. Nobody had set a number for what was too low. Feodora's gut said 71 was too low the moment she saw it.
Half the team wanted to retrain Rubricast that week. Feodora asked for the afternoon instead, to look at where exactly the two raters were splitting, essay by essay, criterion by criterion.
Almost all of the disagreement sat on one line of the rubric: whether the evidence in a paragraph was "sufficient." Thesis clarity, organization, mechanics, the raters barely disagreed on those. On evidence, they were nearly a coin flip apart.
Someone floated a fast fix. Have the two raters score the next batch together, talking each essay through before settling on one number. It felt cheap and it felt fast. The next calibration report came back at 95 percent.
Feodora did not believe it. She pulled in a rater who had never sat in on that conversation, gave them the same essays alone, no discussion. That rater agreed with the group's grade only 74 percent of the time. The rubric line still said "sufficient evidence." Not one word of it had changed.
The decision that let this happen went back to launch, a meeting nobody remembers as important. Someone asked whether the golden set needed two independent raters on every essay, or whether one would do with a spot check later if anything looked odd. The team picked one, because doubling the grading bill for the same 200 essays felt wasteful when the model was brand new and mistakes were expected to be big and obvious. Nobody ever came back to that call as the model got quieter about being wrong.
Run the same Tuesday again with one change: a standing, blind, independently scored sample runs every review cycle from day one, not once at launch. Agreement drops to 80 percent by cycle two, still inside a caution band, not yet a crisis. It gets flagged, the evidence line on the rubric gets rewritten with two worked examples, and by cycle three agreement is back to 87. The 71 percent Tuesday never happens, because nobody let three cycles of quiet slide go unchecked.
One design trusted a single top line number to answer a question it was never built to answer. The other checks the ruler on its own, on a schedule, whether or not anything looks wrong yet.
What I would tell myself, back in that launch meeting: the moment two humans are the definition of "right," ask how you will know if they stop agreeing with each other, before you ever ask how the model is doing against them. Nobody asked. That is on the room, not on the raters.
LEAD, the four checks behind a number you would actually bet on
Not argued from both sides here, there is only one honest position. LEAD is what forces you to find the leading edge instead of admiring the lagging one.
Three things worth stating directly, since this is where the real judgment sits. The alternative the team could have picked instead of fixing the rubric line was scoring every essay with three raters and simply averaging their grades together, without ever separately checking whether those three agreed. It loses because averaging hides disagreement inside a rounded number instead of fixing it, three raters who disagree wildly can still average out to something that looks calm on a dashboard. The AI specific failure mode worth naming by name is ground truth drift, the golden label quietly losing its own reliability while the model's reported accuracy against it holds flat and looks healthy the entire time. The guardrail is a standing, blind, independently scored sample, checked every review cycle, with rater agreement tracked as its own number on its own dashboard tile, never folded into the model's accuracy score. That guardrail is not free. Running it costs roughly ten hours a cycle of two trained raters' time, time that would otherwise go toward growing the golden set with brand new essays instead of re-checking old ones, a real cost accepted on purpose, because a smaller golden set that means one steady thing beats a bigger one nobody can trust. And the bar for "good enough" was never zero disagreement, two careful humans reading the same paragraph will read it slightly differently sometimes, that is fine. It is an audited threshold, checked on a real held out sample of raters working apart, not a promise that people will one day agree perfectly.
And if you want to be sure it really works, try it somewhere else
Same four letters, a claims desk instead of a classroom, nothing about essays anywhere in sight.
Anchorfield Insurance uses a model to draft a suggested severity score on incoming auto claims, low, medium, or high, which decides how fast an adjuster has to look at it. Casimira Halverson runs quality for that model.
Link: the outcome worth protecting is whether the severity label behind every accuracy number means the same thing across every adjuster who has ever double scored a sample claim, not just whichever one happened to score it first.
Early signal: the two adjusters who double score a rotating claims sample used to agree on severity 79 percent of the time. As new adjusters joined during a hiring push, that number slid to 68 percent over eleven weeks, quietly, while the model's own reported accuracy against the golden label stayed near 90 percent the entire stretch. The rate of adjusters overriding the model's suggested severity did not spike until week eleven, when it jumped from 12 percent to 31.
Abuse: the fast fix on the table here was the same one Rubricast tried. Pair the adjusters up for a joint scoring session before the next check. Agreement read 93 percent afterward. A new adjuster tested alone, cold, still only matched the group 70 percent of the time, the real number, unmoved.
Decision: same threshold policy, same three bands. At 68 percent, well under the freeze line, Casimira paused any model tuning that had been justified by the accuracy score and pulled a random, independent sample of 40 claims for three fresh adjusters to score cold, no discussion, before trusting a single number again.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the threshold policy, above the bar trust it, below the bar freeze it, and name the one check that tells you which side you are on.
Cost: there is no budget this quarter for both a routine double scoring sample and a model retrain. The double scoring sample wins every time, a retrain justified by an unverified score is a coin flip dressed up as a fix.
The model got better, for real: say Rubricast's overall match rate improved this cycle. That does not prove the golden label got any more reliable. A model can climb against a ruler that is still bent.
Where people run it wrong.
They treat a rising top line accuracy number as proof there is no problem anywhere, and never check what it was measured against.
They fix a low agreement score by letting raters compare notes, instead of fixing the rubric line causing the disagreement.
They keep the golden label frozen at launch quality forever, instead of re-checking it every cycle as the raters themselves change.
How to use it live. Say the real tension out loud before answering it: "is this asking me to define a term, or to say when I would stop trusting a number built on it." That buys a beat to think instead of racing straight into a textbook definition.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if there's no budget for routine double scoring every cycle?" Response: then say so plainly, the composite score is unverified until that budget exists, which is a different claim than saying you checked a threshold you never actually measured.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Eval design for product teams
- #1 What makes an eval product-relevant rather than research-relevant?
- #2 Design an eval for a feature that drafts email replies.
- #3 How do you decide between automated evals and human review?
- #4 Explain the tradeoffs of LLM-as-judge for a product team.
- #5 How do you validate that your judge model agrees with human raters?
- #6 Describe a rubric that a non-technical reviewer could apply consistently.