ConceptIntermediateEval-Driven Specification / Writing an eval spec / #12
What should the eval spec say about who breaks a tie when scores are ambiguous?
The direct answer
Name one role, not one person, as the tiebreaker, and write it into the eval spec itself. Track two numbers every week: how often a tie happens, and who actually breaks it. If one reviewer keeps ending up as the tiebreaker and keeps saying yes, the eval has stopped measuring the work. It's measuring that one reviewer.
Do this, in order
Write the tiebreaker into the eval spec as a role, not a person.Why: an unwritten "whoever's free" rule quietly turns into "whoever answers fastest," and that's usually the most lenient reviewer.
Track the tie rate every week, next to the pass rate, not instead of it.Why: a pass rate can hold steady for months while more and more cases quietly need a tiebreak to get there.
Log who breaks each tie, alongside the tie rate itself.Why: a tiebreaker doesn't have to cheat on purpose. If ties keep landing with the same person, that person's judgment has become the rubric, whether anyone meant it to or not.
Rotate the tiebreak seat, or route it to a fixed lead, so no reviewer collects most of the close calls.Why: rotation is what actually stops one lenient reviewer from becoming the deciding vote by accident.
Set a real tie rate that trips a rubric review.Why: a tie rate nobody has to answer to is a chart nobody reads, not a defense.
Leave low stakes disagreements, style and tone, off the strict tiebreak process.Why: not every disagreement needs a named owner, and treating them all the same buries the ones that do.
How to answer this, stage by stage
Seven moves. Name the decision before the story, or the answer sounds like a process diagram instead of something you'd actually put in a spec.
1
Scope it to one firm and one product
Say it like this
"Let's make this real. A law firm called Renwick & Hale uses a tool called Casefold to draft the first summary of every brief an associate has to read before a hearing. Elias Thorne runs product for Casefold, and he's the one who has to decide what the eval spec says about a tie."
Why this works
Who breaks a tie stays a vocabulary question until one grader's judgment is actually deciding whether a summary reaches a partner.
2
Say your structure out loud
Say it like this
"I'll use LEAD here. L is the real thing a clear tiebreak rule protects, E is the early signal, how often ties happen and who resolves them, A is how tiebreaking gets gamed, and D is what actually changes at different tie rates."
Why this works
Naming the plan up front tells the interviewer you're about to make a decision, not just describe a process.
3
Reframe what the question is really asking
Say it like this
"This sounds like a process question, who signs off on a close call. It's really asking whether the pass bar stays the same bar every time, or whether it quietly depends on which reviewer happened to answer first."
Why this works
This is where the answer stops being about paperwork and starts being about whether the eval can still be trusted.
4
Give the one decision
Say it like this
"So here's what I'd put in the spec. Name a role as the tiebreaker, not whoever's on duty that day. Track the tie rate and who resolves each one, every week, right next to the pass rate. If one reviewer is breaking most of the ties, that's not neutral anymore, that's the real rubric, and I'd rotate the seat before it gets there."
Why this works
This names a concrete rule you could actually write into a document, not a value like "fairness" with nothing under it.
5
Prove it with a failure
Say it like this
"Here's why that matters. At Renwick & Hale, the weekly tie rate crept from 6 percent to 24 percent over ten weeks, and the pass rate held around 90 percent the whole time, so the dashboard looked fine. Underneath it, one reviewer, Owen Radcliffe, ended up breaking almost 9 out of 10 of those ties, and he passed 90 percent of what he broke, against 55 percent for everyone else. One summary he passed on a tie left out a filing deadline. It reached a partner's binder untouched, and the firm nearly missed the response window."
Why this works
A compressed real failure, with real numbers, does more work than a paragraph of reasoning about what could go wrong.
6
Say how it gets gamed
Say it like this
"Nobody has to cheat here on purpose. A grader with a genuinely borderline case usually just routes it to whoever answers fastest. If that person happens to lean lenient, the close calls start finding their way to them without anyone deciding it should work that way, and the sample pass rate holds steady while the real bar quietly drops."
Why this works
Naming a gaming path that requires no bad intent shows you understand the failure, not just the fix.
7
Say what changes at each tie rate, and close
Say it like this
"Under maybe 8 percent, a tie is just two people reading something differently, note it and move on. Somewhere around 12 to 15 percent, that's not noise anymore, that's the rubric itself going unclear, and it earns a rubric review, not just another tiebreak. Past 20 percent, I'd stop trusting the pass rate as a ship gate until the rubric gets rewritten. If someone asks me whether Casefold's eval still means anything, I want the tie rate and who's breaking them in front of me, not just this week's pass rate."
Why this works
This ties the whole answer to a real action at a real number, so it reads as a plan, not just a definition.
If you only get through two stages
Stages 4 and 6 are the answer. Say the one decision, a named role plus a tracked tie rate, then say how it gets gamed even when nobody means to game it. Everything else here is how you defend that under follow up.
Let's learn
What happens when two graders look at the same AI summary and can't agree whether it's good enough? Say a law firm uses a tool called Casefold. It reads a legal brief, thirty pages of filings and exhibits, and writes the first draft of a one page summary an associate needs before a hearing.
Knowledge spark: what's an eval spec?
A written document that says exactly how a team checks whether an AI tool's output is good enough. What gets scored, what the scale is, what counts as a pass, and who decides when it's not clear.
Before Casefold, an associate read the whole brief by hand and wrote that summary themselves. A thirty page brief took about ninety minutes to read and summarize well.
Now Casefold drafts the same summary in under two minutes. To make sure the drafts hold up, two senior associates score a sample of forty summaries a week against a five point rubric. A score of four or five passes. Anything under that fails and goes back to be rewritten by hand.
Here's the turn. When the two graders' scores land on opposite sides of that line, one says four, one says three, that's not a pass or a fail yet. It's a tie. And the eval spec never said who breaks it. So whoever happens to be free breaks it, case by case, with nobody keeping count.
A pass rate that never moves is not proof every close call got the same fair look. It can be proof the same reviewer is answering all of them.
Knowledge spark: what's an ambiguous score?
When two people grading the same thing land on opposite sides of the pass line. Neither grader is wrong. The case itself is sitting right on the edge.
At its worst, the pass rate held around 90 percent for ten straight weeks, so the weekly QA report looked healthy the whole time. Underneath it, the share of summaries landing as a tie crept from 6 percent to 24 percent, almost one in four. Nearly all of those ties were landing with one reviewer, Owen Radcliffe, who happened to answer fastest on the QA Slack channel and who passed nine out of ten of the ties he broke. One of those passes was a summary that left out a filing deadline. It reached a partner's binder marked clean, and the firm came within two days of missing the response window.
Same tie, same shrug. The deadline still slips off the page.
Share of the weekly sample landing as a tie between the two graders, by week
under the 12 percent floor, ordinary disagreement
over it, the rubric itself is going unclear
The tie rate crossed the 12 percent floor in week six, four weeks before a new associate's offhand question forced anyone to look. By week ten, nearly a quarter of the sample needed a tiebreak.
Confirmed missed-fact errors reaching a partner's binder, before and after ties were tracked
1
A normal quarter, tie rate tracked and rotated
6
The quarter ties went unwatched
One confirmed miss in a normal quarter, the number the manual-only era caught by hand. Six in the quarter the tie rate climbed to 24 percent unwatched. Every one of those six had individually cleared a tiebreak.
The choice I'd take back
When we wrote the eval spec, we left tiebreaking as "whoever's free handles it," because ties looked like a rare edge case, not worth writing a rule for. That was fine when the rubric was brand new and cases were clear cut. It stopped being fine once almost a quarter of the sample needed a tiebreak.
What I'd leave alone. Ties over style, is the summary too long, does a sentence read stiffly, don't need a named tiebreaker. Nothing about the case depends on those, so two graders can just pick one and move on.
The lesson. A tie rate nobody counts isn't proof the rubric is fine. It can be proof one reviewer has quietly become the whole rubric. The only way to know a pass rate still means something is to watch who's deciding the close calls, not just how many pass.
Now here is the same thing as a story
Use this version when you've got the time. The short version is above. This is for when you want to feel why it mattered.
Elias Thorne has run product for Casefold for two years, and he wrote the first version of its eval spec himself, back when the tool was still being tested on one partner's caseload.
For the first several months, the rubric earned its trust the honest way. Every Friday, two senior associates scored a sample of the week's summaries, and on the rare case where they disagreed, whoever was in the office finished the argument out loud, in about two minutes, and moved on.
Then the checking faded, in three small steps, none of them looking wrong at the time. First, because ties were rare enough early on, Elias never wrote down a real rule for who settled them, just a line in the spec saying the graders would sort it out between themselves. Second, as the firm grew, the QA Slack channel filled up with side conversations, and the fastest replier on any given tie started being whoever happened to be at their desk. Third, nobody was logging which reviewer answered, only whether the week's sample passed or failed.
Then Naomi Kessler, a first year associate shadowing QA for the month, asked an offhand question in a Friday review. Why does Owen end up deciding almost every close call?
Elias didn't have an answer. He pulled the raw grading log instead of the weekly pass rate, every score, every tie, and who resolved it, plotted by week. Owen Radcliffe had gone from breaking about a third of the ties in week one to almost nine in ten of them by week ten. And Owen passed 90 percent of what he broke. Everyone else passed about 55 percent.
We did not lose one summary that quarter. We lost the one thing that made a passing score mean the same thing every time.
It would be easy to say Owen did something wrong. He didn't. He was fast, he was available, and every time a close call landed on his desk he made a real call. Nobody decided the tiebreaker should be whoever answers fastest. It just quietly became that, one Slack reply at a time.
So here's the decision Elias would take back.
Two years earlier, in the meeting where the eval spec first got written down, someone asked whether they should name an actual tiebreak owner. It felt like overbuilding a rule for something that would barely ever happen. Elias remembers agreeing they'd deal with it if it ever became a real problem.
I would name the tiebreak owner in the spec from day one instead. Same five point rubric, same weekly sample, but a fixed rotation for who breaks a tie, and a real number on the tie rate itself that someone has to answer to.
Here's the replay. Same ten weeks, tie rate tracked from day one. It crosses the 12 percent floor in week six, and that crossing is the trigger, not Naomi's question in week ten. Elias pulls that week's tied cases, rewrites the two rubric lines graders keep reading differently, and starts a weekly tiebreak rotation from there. The tie rate falls back under 10 percent by week eight. The quarter closes with 1 confirmed missed fact reaching a partner's binder, instead of 6.
One eval spec watches whether this week's summary passed. The other watches whether passing still means the same thing it did in week one.
And the thing Elias would tell himself, back in that first meeting: the pass rate was never the risk. Leaving nobody in charge of the close calls was.
LEAD, run on who gets to say yes when the rubric doesn't
This reads like a process question, who signs off on a form. Underneath it, it's still asking whether the pass bar stays the same bar for every case, or starts depending on who happened to answer. That's LEAD, run on a tiebreak instead of a single metric.
L, link. What a clear tiebreak rule actually protects. Not speed, not fairness to the reviewers themselves, a bar that means the same thing whichever grader was in the room. If a summary needs a tiebreak, its pass or fail should still depend on the work, not on who was free that day.
E, early signal. The earliest thing worth watching isn't the pass rate at all. It's two smaller numbers: how often a tie happens, and who ends up resolving each one. Both drift for weeks before a bad summary ever reaches a partner.
A, abuse. How this gets gamed, usually without anyone meaning to. A close call gets routed to whoever answers fastest, and if that person happens to lean lenient, ties start finding their way there on their own. The pass rate never has to move for this to happen.
D, decision. What actually changes. Under maybe 8 percent, a tie is normal disagreement, note it and move on. Around 12 to 15 percent, that's the rubric going unclear, and it earns a rubric rewrite, not another tiebreak. Past 20 percent, stop trusting the pass rate as a ship gate until the rubric gets fixed.
The check that proves a tiebreak rule is still real
Pull one week at random and ask two things: how many ties were there, and did more than one person break them. If the answer to the second one is close to "just one person," the tiebreak rule isn't neutral anymore, whatever the pass rate says.
And if you want to be sure it really works, try it somewhere else
A used car marketplace called Fairmile Auto Exchange has two inspectors score every vehicle condition report before it goes live: does it list every real issue with the car, yes or no. Same shape of question. A pass rate that holds steady can hide the same thing it hid at Renwick & Hale: one inspector quietly deciding every borderline report, because they're the one who happens to be online when the disagreement comes in.
L. Whether a shopper reading the report actually knows what's wrong with the car, not whether this week's reports cleared review at the same rate as last week.
E. How often the two inspectors' checklists disagree on a report, and which inspector settles it, tracked by week, not the pass rate alone.
A. Nobody has to cheat here either. A tie usually gets routed to whichever inspector is on the overnight shift, since they're awake when the question comes in, and an inspector working alone overnight tends to wave more borderline reports through just to keep the queue moving.
D. Tie rate flat, trust the pass rate. Tie rate climbing and landing mostly with one shift, stop trusting the pass rate and pull a fixed daytime reviewer into the rotation, not just a reminder to be thorough.
Same idea, a car lot at midnight.
Swap the trigger and it still runs
The rubric gets simpler. Doesn't matter. A cleaner five point scale can still produce ties, and an unnamed tiebreaker still drifts toward whoever's fastest.
The QA team gets bigger. Doesn't matter either. More reviewers just means more candidates for one of them to quietly collect most of the close calls.
The model gets genuinely better at the judgment call. Track the tie rate anyway. That's the one case where it should fall on its own, and watching it is what makes "it got better" checkable instead of assumed.
Where people run it wrong
Treating a steady pass rate as proof every close call got the same fair look, with nobody logging who actually broke each tie.
Naming a tiebreaker in theory, "the team will figure it out," instead of naming a role and a rotation in the spec itself.
Waiting for someone to ask an offhand question before pulling the raw log, instead of setting a real tie rate that trips a review on its own.
How to use it live
Say the split first. "Before I answer, I want to separate two questions here, is this one summary a pass or a fail, and is the tiebreak rule itself still fair." That's not stalling. That's naming which question you're actually being asked, and it buys you the room to build the real answer instead of taking the pass rate's word for it.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does it stand for here?
Tap to flip
ANSWER
LEAD. L is the real bar a tiebreak rule protects, E is the early signal, the tie rate and who resolves it, A is how tiebreaking gets gamed, D is what changes at each tie rate.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Elias Thorne, the product manager who has run Casefold, the AI legal brief summarizer, for two years, and who wrote its first eval spec himself.
3 · THE HABIT
What did the team stop doing once the pass rate looked steady?
Tap to flip
ANSWER
Writing down who actually broke each tie. The spec only ever said graders would "sort it out between themselves," so nobody tracked whether that meant ten different people or mostly one.
4 · THE REAL SIGNAL
What number was moving while the pass rate looked flat?
Tap to flip
ANSWER
The tie rate. It crept from 6 percent in week one to 24 percent by week ten, while the overall pass rate held around 90 percent the whole time.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Leaving tiebreaking as an unwritten habit in the original eval spec, because ties looked too rare in the first few months to be worth a real rule.
6 · THE NUMBER
The tie rate went from ______ percent in week one to ______ percent by week ten, crossing a ______ percent floor in week six.
Tap to flip
ANSWER
6 percent, 24 percent, 12 percent. By week ten, Owen Radcliffe was breaking about 9 out of 10 of those ties.
7 · THE REPLAY
Same rubric, tiebreak owner named and tracked from day one, what changes?
Tap to flip
ANSWER
The alert trips in week six instead of going unnoticed until week ten. Elias rewrites the two unclear rubric lines and starts a weekly tiebreak rotation that week. The tie rate falls back under 10 percent by week eight, and the quarter closes with 1 confirmed missed fact reaching a partner instead of 6.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
A used car marketplace's vehicle condition reports. Its early signal is how often the two inspectors disagree, and which inspector settles it, since an overnight inspector working alone tends to wave more borderline reports through.
Check yourself Score: 0 / 0
True or false
1. True or false: if Casefold's weekly pass rate holds steady, every close call is getting a fair, independent look.
True
False
Show hint
Ask what a steady pass rate can hide underneath it.
Show answer
False. A steady pass rate can hide one reviewer quietly breaking almost every tie. Only tracking who resolves each tie shows whether the bar is still the same bar.
Fill in the blank
2. In week one, the tie rate was ______ percent. By week ten, the week Naomi asked her question, it had grown to ______ percent.
Show hint
The numbers sit right under the line chart in Section 1.
Show answer
6 percent. 24 percent. Nearly a fourfold jump, hidden completely behind a pass rate that never moved.
Multiple choice
3. Which of these would actually tell you whether Casefold's eval spec still means the same thing it did in week one?
A. The weekly pass rate.
B. How many summaries Casefold drafts a week.
C. The tie rate, and which reviewer resolves each tie.
D. How fast Casefold drafts a summary.
Show hint
Three of these can look completely normal even after one reviewer has quietly become the whole rubric.
Show answer
C. Only the tie rate and who breaks them shows whether one reviewer has quietly become the real rubric.
Short answer
4. Name a place in this same eval where a tie doesn't need a named tiebreaker at all.
Show hint
Look for disagreements that aren't about missed facts.
Show answer
Model answer: "A tie over style, is the summary too long or does a sentence read stiffly. Nothing about the case depends on that, so two graders can just pick one and move on."
Short answer, apply it yourself
5. Pick a product you use yourself. Name one place it has two people or two systems grade the same thing, where a close call might get quietly settled by whoever's fastest to answer instead of a named rule. What would you check?
Show hint
Look for any place two ratings can disagree with no written rule for who wins.
Show answer
Model answer: "A ride-share app where two different support agents can approve or deny the same refund request. If one agent always seems to get the ambiguous ones and always approves them, I'd track how often each agent resolves a disputed refund, not just the overall approval rate."
Multiple choice
6. If Renwick & Hale had caught the tie rate crossing 12 percent in week six instead of finding out from Naomi's question in week ten, what would most likely have happened to the confirmed missed-fact errors that quarter?
A. No change, it would still be 6.
B. Fewer, because the rubric would have been fixed and the tiebreak rotated before most of the quarter's ties piled up on one reviewer.
C. More, because fixing the rubric mid quarter would have confused the graders.
D. It depends only on how many summaries Casefold drafted that quarter, not on the tie rate.
Show hint
Think about what the tie rate tracking is actually for: catching the drift before it piles up on one reviewer.
Show answer
B. Catching it in week six gives Elias a chance to fix the ambiguous rubric lines and rotate the tiebreak seat before most of the quarter's close calls ever reach Owen alone, exactly what the replay in Section 2 shows.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.