ConceptAdvancedEval-Driven Specification / Writing an eval spec / #9
Describe how you would specify an LLM-as-judge eval and its known weaknesses.
The direct answer
Run every AI-drafted note through an LLM judge scored against a written rubric by default, but never let the judge's score alone approve a note. A separate, deterministic check has to confirm the name, the amount, and the program match the real record first. The judge is fast and steady on tone, but it quietly scores a warm, well-written note higher than a plain one even when a fact inside it is wrong, so the fact check has to sit outside the judge's own score, not folded into it.
How to build the spec, in order
Run the LLM judge on every draft by default, but never let its score alone approve one.Why: speed and consistency are real, but the judge's own score is exactly the thing a wrong-but-warm note can fool.
Check the hard facts, name, amount, and program, against the real gift record before the judge ever scores tone.Why: splitting the two checks means a warm sentence can't buy its way past a wrong number.
Route any note to a donor giving $500 or more this year to a person, no matter what the judge scored it.Why: that's exactly where the hidden mistake costs the most, a lost renewal, not a rewrite.
Name the hidden bias by its real name, verbosity bias, and don't let the rubric reward length or warmth on its own.Why: a bias you haven't named is a bias you can't test for.
Track the real error rate against a fixed kill line every month, not the judge's average score.Why: the average looked fine at 97 in 100 while the 3 that mattered kept slipping through.
Leave the small, first-time gifts on the judge alone, with no extra check.Why: a plain note on a $20 gift costs nothing to get slightly wrong, and checking it steals review time from where it matters.
Six moves for walking the eval spec out loud
This is a tradeoff wearing a spec's clothes, so PICK is the tool, not a checklist of rubric fields.
1
Reframe what the question is actually testing
Say it like this
"Before I get into it, I want to say what this question is really asking. It's not 'can an AI grade another AI's writing.' It's 'where's the line where you stop trusting that grade and put a person on it instead.' That's the spec I'm going to write."
Why this works
Shows the interviewer you see the real shape of the question, not just the surface one about whether LLM judges work.
2
State your position, with the actual mechanism in it
Say it like this
"My position: score every draft with the LLM judge by default. But the judge's score never gets to be the only check. A separate, dumb, deterministic check confirms the name, the amount, and the program match the real gift record first. Only a note that passes that gets the judge's tone score counted as final."
Why this works
PICK rewards commitment. Naming the actual mechanism, not a vague "add oversight," is what makes it a position instead of a hedge.
3
Name who eats each kind of mistake, and how fast
Say it like this
"Split the two kinds of miss. If the judge marks a plain, correct note as weak on tone, someone catches it in the weekly sample and rewrites a line, a couple of minutes, done. If the judge passes a warm note with the wrong program on it, a real donor gets it, and nobody notices until they call, or an audit runs, months later."
Why this works
This is the impact step. It separates a visible, cheap cost from a hidden, expensive one instead of talking about "quality" as one number.
4
Name the hidden bias by its real name
Say it like this
"Here's the asymmetry. The judge likes long, warm-sounding notes. That's not a guess, it's a known failure mode called verbosity bias. So a note that reads beautifully and gets one fact wrong can outscore a plain, correct one. That's the error I'd optimize against, not the one where the judge is too strict."
Why this works
This is the C step. Naming the exact bias by its behavior, not just "it can be wrong," proves you understand LLM judges instead of just distrusting them.
5
Give the spec itself, field by field
Say it like this
"The spec has three parts. One, a deterministic fact layer: name, amount, program, checked against the record, pass or fail, no model involved. Two, the LLM judge, only on notes that passed step one, scoring tone and personalization one to five against a written rubric with a real example of a five and a real example of a one. Three, an escalation rule: any note to a donor giving $500 or more this year goes to a person, regardless of either score."
Why this works
This is the direct answer to "describe how you would specify it," in the actual shape a spec takes, not a description of the idea.
6
Close on what would make you stop trusting it
Say it like this
"I'd stop trusting the judge the moment one of three things shows up: the audited error rate on major-donor notes stays above 1 in 100 for two months running, the judge's score tracks note length better than it tracks the fact check, or it scores its own model family's drafts higher than an equally good draft from a different model. Any one of those, and the judge comes off auto-approval until it's fixed."
Why this works
A pick with no kill criteria sounds stubborn. This is what makes the whole answer sound like a standard instead of a preference.
One more thing before the walkthrough moves on: this spec applies to a slice of the product, not the whole thing. Most candidates hear "specify an eval" and answer with one rubric for every draft everywhere. Say which drafts get the deterministic layer and which don't, and you've answered a harder, better question than the one that was technically asked.
Let's learn
The tool sits inside donor software a nonprofit already uses. A donor gives money, the system pulls their name, their gift, and what program it went to, and drafts a personal thank-you note in under a minute, ready for someone on the donor-relations team to read and send.
One gift, one draft, two different checks before it leaves the building
Knowledge spark: what's an LLM-as-judge eval?
A second AI model reads the first AI's answer and grades it against a written checklist, instead of a person grading every single one. It's fast, it's cheap, and it never gets tired on the four-hundredth note. It's also just another model guessing, which is exactly the part people forget.
Before the tool, a coordinator wrote every note by hand for anyone giving real money, checking the amount and the program against the record each time, about 15 minutes a note. Smaller gifts got a plain mail-merge letter, the same three sentences for everyone, about 2 minutes to send and forgettable to read.
With the tool, a full draft appears in about 30 seconds, already sounding personal. At first, a coordinator still read every line against the record before sending, about 5 minutes. Once the drafts kept coming back right, that dropped to under a minute. Review time across the donor-relations team fell more than 80 percent inside a year.
Here is the turn. The time saved was never the risk. The risk was what a coordinator stopped checking once the drafts kept scoring well. A system that grades its own homework and always hands back a passing grade will, eventually, hand back a passing grade on a wrong answer.
We didn't build an eval for whether the note was true. We built one for whether it sounded true.
At its worst, that looks like this: a note names the wrong program, or the wrong amount, and it reads so smoothly nobody questions it. It goes to a donor who's given for years. They notice. They don't say anything the first time. They just don't renew. Nobody connects the dropped gift to one bad sentence in a thank-you note, because nothing about the note looked wrong.
The choice I would take back
Early on, once the judge's score matched a human rater's judgment 96 percent of the time on a test set of 500 notes, the team let any note scoring above 4 out of 5 auto-send with zero human review. That felt earned, the number was real. I would take it back. A 96 percent match rate on a test set tells you the judge agrees with people most of the time. It says nothing about what the other 4 percent looks like once it's live, at volume, on real donor money.
Cost per incident, in dollars
Caught the same day, cheap
Not caught for months, expensive
Judge wrongly flags a correct, plain note
$2
Judge wrongly passes a warm note with a wrong fact
$180
The cheap bar is barely there on purpose. $2 covers four minutes of staff time to rewrite a line, caught in the weekly sample. $180 covers the time to investigate and fix the record once a donor calls, plus the value of a renewal that doesn't come back, at a $650 average gift and a renewal rate that drops from 65 percent to 40 percent after a factual miss. Ninety times the cost, and it's the one nobody sees coming.
What I would leave alone. The mail-merge tier, gifts under $50, where the note has always been plain and nobody has ever expected personality from it. If the judge lets a slightly generic line through there, it costs nothing. Nobody is checking a $20 acknowledgment against anything. Spending review time there takes it away from the notes where a wrong fact actually lands somewhere.
The lesson. We built the eval to answer "does this sound like a good thank-you note," and it answers that well. We never built a separate eval for "is every fact in it true." Those are two different questions, and only the second one can hurt somebody.
The month Pinehaven stopped trusting the queue
You don't need this to answer the question. Read it if you want to feel why the fact check has to live outside the judge.
Wren Castellano can read a donor's giving history and know, before she's finished the page, exactly which line to open with. She's been the donor-relations coordinator at Pinehaven Land Trust for five years, and every gift that comes in gets a note with her name at the bottom.
Osprey's drafting tool arrived the spring before last, and for months it was just good. A draft landed in about thirty seconds, already reading like Wren had written it herself. She read every line against the gift record, fixed a word, sent it, about five minutes a note. Nine notes on a busy Tuesday used to take her most of a morning. Now it took less than an hour, and she still checked every fact herself.
She never stopped believing in it. She just, slowly, stopped opening the record to check the amount first. Then she stopped checking the program name, because the tool had never once gotten it wrong that she'd caught. By early autumn she was reading a note the way you read a text message, for the shape of it, not the specifics.
Same judge, two very different misses
Then, one Thursday, she was about to send a note thanking a donor named Everett for his gift "to expand the coastal trail." She paused with her hand on the mouse, not because anything about the sentence looked wrong, but because she remembered mentioning the youth education program to Everett at a site visit the month before. She pulled his record. His gift had gone to youth education. The note the tool drafted, and the judge had scored a 4.6, named the wrong program entirely.
She hadn't caught it by checking. She'd caught it by luck, and by knowing Everett personally, which she doesn't for most of the two hundred donors on her list.
We didn't lose one fact. We lost the five seconds Wren used to spend checking every note, and got back a coordinator who now checks all of them again, by hand.
Ottavia Marsh had been the eval lead on the drafting tool since before it shipped. When Wren's near miss reached her, the header number was still fine. Auto-approved notes matched their gift records 96 times in 100 on the last full audit. She almost filed it as an outlier.
Then she pulled the meeting notes from the launch decision. The team had tested the judge's scores against a human rater on 500 notes and found 96 percent agreement, and decided that was strong enough to auto-send anything scoring above 4 out of 5, with nobody checking behind it. Nobody had asked what the other 4 percent actually said. Everett's note, if it had gone out, would have been one of them.
She read a sample of the auto-approved notes that had scored highest, the ones the judge liked best, and found the pattern fast. The judge scored longer, warmer notes higher almost every time, whether or not the specific in them was true. A short, correct note about a $75 gift routinely scored a 3.8. A long, glowing note with one invented detail routinely scored a 4.7. The judge wasn't grading whether the note was right. It was grading whether the note sounded like a good one, and length and warmth were doing most of that work.
So Ottavia split the check in two. Before the judge ever reads a draft, a separate pass confirms the name, the amount, and the program against Pinehaven's own record, plain matching, no model involved. Only a note that passes gets scored for tone at all, and any note to a donor giving $500 or more that year goes to a person no matter what either check says.
The part she'd take back isn't the launch. It's trusting a 96 percent agreement number on a test set to describe what would happen at full volume, on real gifts, month after month. A number that good on a sample of 500 hid exactly the kind of note that would slip through on gift two thousand.
Wren's review time on a verified note is back under a minute. Nine notes on a busy Tuesday take her under fifteen. And the specific mistake she caught by luck, in the two seconds before she clicked send, can't happen again the same way, because nothing with an unmatched fact reaches her inbox scored as done.
PICK, run against the judge itself
This is a tradeoff dressed as a spec, so PICK is the tool. A "how would you measure launch health" question would reach for LEAD instead.
P, position. Default to the LLM judge for tone and personalization. Never let its score alone approve a note. A separate, deterministic fact check runs first, and any note to a $500-plus donor goes to a person regardless of either score.
I, impact. A plain, correct note wrongly marked weak on tone costs a coordinator a couple of minutes on the weekly sample, and nobody outside the team ever sees it. A warm note with a wrong fact reaches a real donor, and the cost, an $18 fix plus a chunk of a renewal that doesn't come back, doesn't show up until someone notices, which can be months.
C, cost asymmetry. The judge has a known weakness called verbosity bias: it scores longer, warmer writing higher, whether or not a fact inside it is true. That means the exact notes most likely to slip past it are the ones that read the best. Spend your caution there, not on the notes the judge is merely too strict about.
K, kill criteria. Stop trusting the judge for auto-approval the moment any one of three things is true: the audited error rate on major-donor notes holds above 1 in 100 for two months running; the judge's score correlates more with note length than with the fact-check result; or it scores its own model family's drafts higher than an equally good draft from a different model. Each one is checkable against real numbers, not a feeling.
Knowledge spark: what is verbosity bias?
A judge model tends to score a longer, more confident-sounding answer higher than a short, plain one, even when the short one is the accurate one. It happens because "reads well" is an easier thing for a language model to spot than "is true." Naming it is the difference between a general worry and a specific check you can run.
Audited error rate on major-donor notes, against the kill line
Fact mismatch found in auto-approved notes, by month
Kill line for $500-plus donor notes, 1 in 100
Four straight months of the fact-check layer made the rate better. It's still 70 percent over the line that matters for major-donor notes. Getting better is not the same question as safe enough, and until this line crosses the kill line, the escalation rule for $500-plus donors stays exactly where it is.
Run PICK again, in a vet's exam room
A veterinary clinic uses a model to draft the after-visit summary a pet owner gets: what happened, what to watch for, and any medication. Same shape of question, a different desk, a different kind of harm.
P. Score every summary with a judge for clarity and tone by default. Never let its score alone clear a summary that names a medication, a dose, or a follow-up date. Those get a vet's eyes before the owner sees them. I. A clunky sentence in a routine wellness note gets caught on the vet's read-back, costs a minute. A wrong dose that reads confidently gets read by an owner at home, at night, with no vet in the room to ask. C. Wording misses are common and cheap, caught constantly, fixed in seconds. A dosage error is rare and, once acted on, can't be taken back. Spend the caution where a mistake is irreversible, not where it's merely embarrassing. K. Pull the judge off auto-clearing any summary the moment an audit finds a dosage or medication-name mismatch above 1 in 1,000, or the judge's clarity score tracks sentence length better than it tracks a match against the vet's own chart notes. Either one, and a person reads every summary with a medication in it, no exceptions.
What I would leave alone, in the exam room
The routine wellness-visit summary with no medication change and no abnormal finding, just "all normal, see you next year." If the judge lets a slightly stiff sentence through there, an owner rereads it once and moves on. Nothing about it can hurt the pet.
Swap the trigger and it still runs
Speed: the judge scores a draft in half a second instead of five. Doesn't move the line. The fact check exists because of what a wrong sentence can do once it's sent, not how fast it arrived.
Cost: running the judge gets ten times more expensive. Also doesn't move the line, since the hidden factual miss is what's actually expensive, not the API call.
The judge gets better: verbosity bias drops until length barely predicts the score at all. Move the line, don't erase it. The deterministic fact check stays, because even a well-calibrated judge is still grading how a sentence reads, and a fact check is grading whether it's true. Those never merge into one question.
Where people run it wrong
Treating the judge's average agreement rate with human raters as proof it's safe at volume, when the rate was measured on 500 notes and the failure shows up on note two thousand.
Making the judge "stricter" as the fix, which only teaches it to reward a different kind of writing, not to check whether a fact is true.
Writing "add a human review step" without saying which notes get it, so the review step ends up reading the easy ones and skipping the ones that actually needed a person.
If you are asked this cold
Say the reframe out loud before you list any rubric fields. "Give me a second, I want to separate what a wrong answer costs from how often the judge is wrong." That's true, it's already stage one of the walkthrough, and it buys you the time to find the real asymmetry instead of a generic list of rubric criteria.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what's its hardest step?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming which judge mistake is cheap and visible, and which is hidden and expensive, then building the spec around the hidden one.
2 · THE PERSON
Who owns the eval spec in this answer, and what did she notice that the launch numbers missed?
Tap to flip
ANSWER
Ottavia Marsh, eval lead on Osprey Civic Tools' drafting product. She noticed the judge's 96 percent agreement rate with human raters didn't mean the passed notes were true, only that they sounded right.
3 · THE HABIT
What did Wren Castellano stop doing once the drafts kept scoring well?
Tap to flip
ANSWER
Checking the dollar amount and the program on each note against the real gift record before sending it. The notes read so smoothly she stopped treating them as guesses.
4 · THE ASYMMETRY
Name the two kinds of judge mistake here and what each one costs.
Tap to flip
ANSWER
A plain, correct note wrongly marked weak on tone: caught in the weekly sample, about $2 to fix. A warm note with a wrong fact: not caught for months, about $180 once you add the fix and the renewal that doesn't come back.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Score every draft with the judge by default, but never let its score alone approve a note; a separate fact check against the gift record has to pass first, and any $500-plus donor note goes to a person regardless of either score.
6 · THE NUMBER
By month four, the audit still found ______ in 100 judge-approved major-donor notes with a fact the record didn't match.
Tap to flip
ANSWER
1.7. Down from 3.4 at month one, but still above the 1-in-100 line the team set, which is why the escalation rule stayed in place.
7 · THE KILL CRITERIA
Name the three things that would make you stop trusting the judge for this feature.
Tap to flip
ANSWER
The audited error rate on major-donor notes stays above 1 in 100 for two months running. The judge's score tracks note length better than the fact check, evidence of verbosity bias. The judge scores its own model family's drafts higher than an equally good draft from a different model, self-preference bias.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where's the reject line there?
Tap to flip
ANSWER
A vet clinic's AI-drafted after-visit summaries. The judge's score alone never clears a summary naming a medication, a dose, or a follow-up date, those always get a vet's eyes first.
Check yourself Score: 0 / 0
Multiple choice
1. Which of these is the hidden, expensive error this eval spec has to guard against?
A. The judge takes a few seconds longer to score a draft than expected.
B. The judge scores a warm, well-written note higher even though it names the wrong program.
C. A coordinator has to log in each morning to review a note.
D. The rubric has more than five scoring criteria.
Show hint
Three of these are annoyances. One is the mistake that reaches a real donor and doesn't get caught.
Show answer
B. A, C, and D are all minor operational friction. B is the one condition that costs the most and is hardest to catch, which is exactly what a cost asymmetry has to be built around.
Fill in the blank
2. Fill in the blank: by month four, the audit of judge-approved major-donor notes still found ______ percent with a fact the gift record didn't match, still above the 1 percent line.
Show hint
It dropped from 3.4 percent at month one, but it didn't cross the kill line.
Show answer
1.7. Better every month, and still 70 percent over the line that matters for donors giving $500 or more, which is why the escalation rule for that group never got lifted.
True or false
3. True or false: once a note passes the deterministic fact check, the judge's tone score alone is enough to send it to any donor, no matter the gift size.
True
False
Show hint
Look at what the escalation rule does for donors giving $500 or more.
Show answer
False. Passing the fact check is necessary, not sufficient. Any note to a donor giving $500 or more that year still goes to a person, regardless of what either the fact check or the judge says.
Multiple choice
4. Why couldn't Osprey just tell the judge to "grade more strictly" instead of building a separate fact check?
A. A stricter judge would cost more to run per note.
B. A stricter judge would still be grading how the note reads, not whether the fact inside it is true, so a warm wrong note could still pass.
C. Coordinators refused to work with a stricter rubric.
D. It was against Pinehaven's policy to change the scoring rules.
Show hint
Ask what a "stricter" version of the same judge would actually be measuring.
Show answer
B. Turning the judge's dial up just teaches it to reward a narrower kind of good-sounding writing. It never gives the judge a way to know a fact is wrong, because it was never checking facts, it was checking tone. The fix has to add a check the judge doesn't do, not tune the one it already does.
Short answer
5. If the audited error rate had stayed near month one's 3.4 percent instead of dropping to 1.7 percent by month four, would routing every $500-plus note to a person still make sense? Walk through it.
Show hint
Redo the $180-per-incident math at roughly double the rate, and compare it to what full auto-approval would have cost instead.
Show answer
Yes, by more, not less. At 1.7 percent the escalation rule was already worth keeping, since each uncaught error runs about $180 against a $2 caught one. At 3.4 percent, the same math roughly doubles the expected cost of skipping the human check. A higher error rate makes the case for the escalation rule stronger, not weaker. It's exactly the kind of number that should tighten a rule, never loosen it.
Short answer, apply it yourself
6. Pick a place in your own work where one AI checks another AI's output. Name one way the checker could be fooled by something that just looks right, and what you'd check outside the checker to catch it.
Show hint
Look for the output the checker never compares against a real source, only against its own sense of what "good" sounds like.
Show answer
Model answer: "Take a tool that uses an LLM to grade AI-written meeting summaries for clarity. A summary that's fluent and well organized can still misstate who agreed to what, and a clarity-only judge would score it high anyway, because misstatement doesn't make a sentence read worse. I'd check the summary's action items against the actual calendar invites and task tracker before letting a clarity score alone close the loop." Any answer works if you can name the fact the checker never actually verifies.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.