ConceptAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #23

How do you write criteria that hold across a model upgrade?

The direct answer
Write every criterion to catch one named mistake, not one blended score, and tag it to the exact failure it was built to catch. When the model changes, don't just watch the same score climb. Pull a fresh batch of its real output, look for a kind of mistake nobody wrote a rule for yet, and check whether the old rules are even testing anything real any more.
Do this, in order
  1. Tie every criterion to one named mistake, and re-test it against fresh real output the moment the model changes.Why: a blended pass rate can climb from 92 to 98 percent while a brand-new kind of mistake goes completely unmeasured underneath it.
  2. Pull a fresh sample of the new model's real output and hunt for mistakes with no rule written for them yet, instead of re-running the old checklist.Why: every existing check was worded around what the old model got wrong, so a mistake the old model never made has zero criteria watching for it.
  3. Recheck any capacity or threshold tied to how often something used to happen, before trusting that it still catches enough.Why: a review trigger sized for 12 percent of output can quietly stop covering anything once the new model triples how often it fires.
  4. Confirm the checks still measure what they claim before blaming coverage.Why: a display or format change can hide the exact evidence a human reviewer would need to catch the new mistake.
  5. Slice the pass rate by the kind of output at risk, not one blended number.Why: the aggregate barely moved while the risky slice went from 4 percent to 61 percent.
  6. Leave the checks that still work alone.Why: the length, caps, and emoji checks still catch what they always caught; rewriting them wastes time the real gap needs.

How to answer this, stage by stage

Six moves. The trap in this question is treating "the score went up" as proof the criteria still work, so two of the six stages are spent forcing that assumption to earn its keep.

1
Pin it to one real rubric and number
Say it like this
"Let me put a number on this. Say there's a marketing platform called Brightloop, and inside it a tool called Hooklines writes the subject line for every email a client sends. Astrid Fenmark runs eval for it, and fourteen months ago she wrote five rules a subject line has to pass before it ships: under sixty characters, no shouting in caps, at most one emoji, and so on."
Why this works
Naming the tool, the person, and the actual rules turns "criteria" from an abstract word into five checkable things, which is what the rest of the answer needs to stand on.
2
Say exactly what "holds across an upgrade" means
Say it like this
"Holding across an upgrade doesn't mean the pass rate stays the same or goes up. It means the rules are still testing for the mistakes the new model actually makes, not just the ones the old one made. Those can be two completely different lists."
Why this works
This is the whole question in one sentence. Skip it and the rest of the answer sounds like generic advice about "keeping evals fresh."
3
State the paradox plainly, then explain it
Say it like this
"Here's the actual shape of it. The five-rule pass rate climbed from 92 to 98 percent the month the new model shipped. That looked like a clean win. But subject lines that state a real number, a price, a percent off, a count left, went from 12 percent of everything Hooklines wrote to 34 percent, because the new model got much better at sounding specific. Of those, the ones inventing a number that wasn't true went from about 4 percent to 61 percent. None of the five rules were built to catch that, so none of them saw it."
Why this works
Naming the mechanism, rules built for old mistakes grading a model with new ones, beats saying "the rubric got stale" and leaving it vague.
4
Give the fix up front
Say it like this
"So here's what I'd do. Add a sixth rule: any subject line that states a real number gets checked against the actual campaign data it's supposed to come from, not skimmed by a person for tone. Re-run that check, and the size of the review team behind it, every time the model changes, instead of trusting a climbing pass rate to mean the same thing it meant before."
Why this works
Matches the direct answer. Naming the fix, not just the flaw, shows you can repair the rubric, not only describe what broke it.
5
Show the timeline and the segment recut
Say it like this
"Here's the timeline. The rubric was written at month one. One rule got added at month nine. The upgrade shipped at month fourteen, and the pass rate hit 98 that same month. Nobody touched the rubric at fourteen, or in the two months after. At month sixteen, a client called Wick and Ember got a subject line reading 'Only 6 left,' on a product with two hundred units in stock, and that's when anyone looked again."
Why this works
Putting real months on both the shipping date and the review date turns "nobody checked" from an accusation into a fact anyone can verify.
6
Rule out a broken check, then run the evidence test
Say it like this
"Before blaming coverage, I'd rule out a broken grader. Astrid's team hand-checked 20 flagged lines against what the automatic rubric said, and agreed on 19. So the scoring itself was fine. The real test: pull 50 real subject lines with a number in them, sent after the upgrade, and check each number against the actual campaign record. The rubric passed 96 percent of them. Only 43 percent of the numbers were true. That's the whole gap, in one comparison. And I wouldn't touch the length or caps checks, they still catch what they always caught."
Why this works
Ending on a number nobody can argue with, then naming what stays untouched, shows judgment instead of a blanket rewrite of every rule.

Let's learn

Hooklines is a feature inside Brightloop, a marketing platform. A client picks an email campaign, and Hooklines writes the subject line, so the client doesn't have to.

Knowledge spark: what's a grounded claim? A grounded claim is a fact in the subject line that's actually true, checked against the real record it came from. "20% off" is grounded if the campaign really is 20% off. It's a made-up claim if the model just guessed a number that sounded convincing.

When Hooklines launched, Astrid wrote five rules a subject line had to pass: a length limit, no shouting in all caps, no more than one emoji, a correctly formatted name tag, and no spammy punctuation. Fourteen months later, after a model upgrade, the pass rate on those five rules climbed from 92 to 98 percent.

Two lines, sixteen months: the rubric's pass rate vs. the real fabricated-claim rate
What the dashboard showed: five-rule pass rate
climbing, 92 to 98, every release wins
Quiet signal: blended fabricated-claim rate, all subject lines
climbing quietly, 0.5 to 7.3 percent
m1m5m9, rule addedm14, upgrade shipsm16
The rubric's own pass rate climbed six points in fourteen months. The real fabricated-claim rate, blended across every subject line Hooklines wrote, climbed almost seven points over the same stretch, starting right where the upgrade shipped. Nothing about that second line ever touched the first.

Here is the turn. Those six extra points on the rubric never had a chance to see the problem. The real story is a kind of subject line the rubric was never built to check: one that states a real number, growing from a small slice of everything Hooklines wrote to more than a third of it, and increasingly making that number up.

The model did not get worse at writing subject lines. It got so good at sounding specific that nobody thought to check if it was telling the truth.

At its worst, this costs Brightloop named clients with campaigns already sent. Wick and Ember, a candle brand with two hundred units of a bestseller in stock, sent a subject line reading "Only 6 left, don't miss it," and spent the next two days answering angry replies from customers who'd tried to buy a candle that was never actually running out.

A hand sketch timeline with four marks: the rubric built at month one with five checks, one rule added at month nine for merge-tag format, the model upgrade shipping at month fourteen with the pass rate at 98 percent, and a red mark at month sixteen where a client named Wick and Ember complains about a fake stock count.
The gap between a climbing pass rate and the mistake it was never built to see

The choice I would take back. Writing the five rules around the old model's actual mistakes, fourteen months ago, was the sensible call. That model almost never wrote a specific number, so nobody had a reason to write a rule for one. I would take back leaving those five rules as the whole rubric, with no standing check tied to "does this claim exist in real data." I would write that sixth rule the day the tool launched, even while it rarely fired, so it would already be there the day it started to matter.

The decision that mattered Grade every subject line that states a fact against the real record it's supposed to come from, not against a rubric built entirely from the old model's habits. Not because the five old rules are wrong. Because a rubric that never learned to ask "is this true" will always agree with itself, right up until a client answers angry replies about candles that were never running out.

What I would leave alone. The length, caps, and emoji rules don't need touching. They still catch what they always caught, at close to the same rate, because those are exactly the mistakes the new model still occasionally makes. Rewriting rules that already work spends time the real gap needs.

The lesson. A rubric is a list of the mistakes you've already seen. The day the model changes, that list does not know it's incomplete, and it keeps handing out a passing grade unless someone goes looking for the mistake it was never written to catch.

The two days Wick and Ember stopped trusting the send button

Read the short version above if you're short on time. This is the long version, for when you want to feel exactly where the sixteen months went.

Astrid Fenmark has run eval for Hooklines since it launched. She can read a week's pass rate and tell, inside a minute, whether a release earned it or got lucky on an easy batch of campaigns.

She wrote the original five rules herself, with two copywriters, the month Hooklines first shipped. Under sixty characters. No shouting in caps. At most one emoji. A correctly formatted name tag. No spammy punctuation. All five came straight from watching that early model's actual mistakes, week after week, and writing a rule for each one.

For the first several months, watching the pass rate climb felt like the model getting sharper. Eighty-eight. Ninety. Ninety-two. Every point felt earned. Somewhere in there, without anyone deciding it on purpose, the review team quietly stopped double-checking every flagged subject line by hand. When only one in eight lines mentioned a number, checking each one against the real campaign was easy. Nobody had a reason to worry about capacity.

Then, at month fourteen, Brightloop rolled out a model upgrade across the whole platform.

Nobody rewrote the rubric. Why would they. It kept scoring in the high nineties.

Two months later, three weeks before a seasonal sale, Ramona Vetch, who ran Wick and Ember out of a small workshop with her brother, noticed a subject line claiming her bestselling candle had six left. She had two hundred in the back room. She thought it was one strange line. Then she pulled the last ten campaigns. Six of the ten had a number that wasn't real.

She paused every scheduled send and started writing her own subject lines by hand, the way she did before she ever used Hooklines.

We did not lose one wrong number in one email. We lost a candle maker's trust in every line the tool would ever write for her again.

I want to say the model got worse at writing subject lines. It did not get worse. It got fluent enough to sound certain about a number it never checked, and none of the five rules had ever been built to notice.

So here is the decision I would take back.

Fourteen months earlier, writing five rules straight from the old model's real mistakes was the sensible call. That model almost never stated a specific number, so nobody wrote a rule asking whether one was true. I would put a sixth rule in from day one: any subject line naming a price, a percent, a count, or a date gets checked against the campaign's real data before it ships, not skimmed by a person for tone. Not because the five old rules were wrong. Because of what it would have shown, two months earlier: the new model stating a real number three times as often as before, and getting more than half of them wrong.

That is the whole difference. One rubric trusts a score that only ever answered for the mistakes it already knew. The other checks that score against the mistake the model just started making.

And the part I would want to tell myself, if I could go back: we wrote a test that only ever answered for the model we had on day one. Nobody decided that on purpose. We just never rebuilt it when the model changed.

What the recut actually showed

Before blaming the rubric's coverage, Astrid's team checked whether the scoring itself was even right. Two people hand-checked 20 flagged subject lines against what the automatic rubric said they should score. They agreed on 19 of 20. The grading was not the problem. That left the coverage.

Share of subject lines versus share of fabricated claims
66%
39%
34%
61%
No specific claim
Generic, no price or count
States a specific claim
A price, percent, count, or date
Share of all subject lines, post-upgrade
Share of fabricated claims found
Claim-bearing subject lines are 34 percent of everything Hooklines writes since the upgrade, up from 12 percent before it. That same 34 percent is behind 61 percent of every fabricated claim the audit found, while the five-rule rubric had zero rules built to check any of them.
Shipping on the rubric's pass rate alone
98 percent on the five-rule rubric, month fourteen
0 of the five rules checked whether a stated price, count, or date in a subject line was actually true
Requiring the grounding check to agree too
Adopted the week Wick and Ember paused its sends
43 percent of 50 real claim-bearing subject lines matched the actual campaign record they were supposed to describe

Three ways a passing rubric can miss a brand-new mistake

Not because anyone cut a corner on purpose. A rubric built entirely from the old model's habits can keep passing, release after release, and still stop meaning anything the day the model starts making a different kind of mistake.

Three hand-sketched panels: a pink card with a question mark labeled no rule for grounding; a gauge labeled review capacity fixed while claim lines tripled; and a document labeled format hid the tell, reviewers never taught to check a filled-in name.
Three separate, checkable ways a rubric can pass and still miss the real mistake
Way 1
No rule was ever written for the new mistake.

All five checks target things the old model got wrong: length, caps, emoji count, tag format, punctuation. None of them ask whether a stated fact is true, because the old model almost never stated one.

How you'd check it: list what each rule was written to catch, and ask which real mistakes in this week's output aren't on that list. If the answer is "several," coverage has a gap.
Way 2
A review trigger sized for the old rate got swamped by the new one.

The team's one manual check, flag any line with a number for a human to verify, covered nearly all of them when only 12 percent of lines had a number. Review headcount never changed. Now 34 percent of lines have one, and the team can only reach about a third of what gets flagged.

How you'd check it: compare how often the trigger fires today against the review capacity it was sized for. Three times the volume against the same headcount is a real gap, not a bad month.
Way 3
A format change hid the evidence a reviewer would need.

The review tool fills in a sample name where the merge tag sits, so a line reads "Sarah, save 20% today" instead of showing the raw tag. Reviewers were trained to check tone, not to ask whether "Sarah" or the "20%" corresponds to anything real, because the old model never produced a line confident enough to need that question.

How you'd check it: ask a reviewer to explain, out loud, what they're checking when they approve a claim-bearing line. If the answer is "it reads naturally," the format is hiding the actual test.

Reading TRACE off a rubric that never learned to doubt a number

This reads like a question about writing good rules, but the real job is diagnosis: work out why a rubric the team trusted for fourteen months quietly stopped predicting real client complaints. GUARD would fit if the harm landed unevenly across a protected group; here the harm is a whole kind of mistake the rubric never learned existed.

T, timeline. The rubric went in at month one, built entirely from the launch model's own mistakes. One rule was added at month nine, a small format fix, not a rebuild. The model upgrade shipped at month fourteen. Nobody touched the rubric at fourteen, or in the two months after. The blended fabricated-claim rate started climbing the same month, from 0.5 to 7.3 percent by month sixteen, while the five-rule pass rate kept climbing from 92 to 98 over the same stretch.
R, recut. Split subject lines by whether they state a specific claim. No-claim lines, 66 percent of volume post-upgrade: pass and stay accurate at close to the old rate. Claim-bearing lines, 34 percent of volume, up from 12 before the upgrade: 61 percent of every fabricated claim found in the audit. A blended rate that moved seven points hid a slice that was failing worse than half the time.
A, assume nothing. Before blaming the rubric's coverage, two people hand-checked 20 flagged lines against what the automatic rubric's own rules said they should score, and agreed on 19. The scoring itself was fine. The gap was real, not a measurement bug.
C, cause candidates. Three, named and separate: the rubric was worded entirely around the old model's failure modes, so no rule exists for a claim being untrue; the one review trigger tied to a number appearing was sized for a rate that has since tripled, so most flagged lines never get a human check; and the review tool's own display fills in a sample name over the merge tag, hiding from reviewers the exact thing they'd need to question.
E, evidence test. Pull 50 real claim-bearing subject lines sent after the upgrade, and check each stated number against the actual campaign record. Rubric pass rate on those same 50: 96 percent. Real match rate: 43 percent. A rubric built for the old model's mistakes can't grade a mistake it was never shown.
Why the evidence test is the hard step Anyone can suspect a rubric has fallen behind a model upgrade. A test earns its place by turning that suspicion into a number: how the rubric's own pass rate compares to a real check of the same subject lines against real data. Do that comparison and you've checked something real. Call a rubric "probably outdated" without it, and you've only said the same worry in a more confident voice.

Same shape, a maintenance reply that promised a technician nobody dispatched

Fallowmere Property Group runs an AI tool that drafts the first reply to a tenant's maintenance request, for a human property manager to lightly edit and send. Declan Whitfield, the operations lead, wrote the rubric ten months before a model upgrade: does the reply name the specific issue, does it give a timeframe, does it avoid promising a technician the office hasn't actually scheduled.

T. The rubric's overall pass rate climbed from 90 to 97 percent across three releases, including the upgrade. Tenant complaints about a broken promise, someone showing up to wait for a technician who never came, held steady around one in thirty replies for months, then started climbing six weeks after the upgrade shipped.
R. Split by whether the reply names a specific date or time. Vague replies, 69 percent of volume: complaint rate barely moved. Specific replies, 31 percent of volume, up from 8 percent before the upgrade: 44 percent of them turned out to promise a visit nobody had dispatched, up from 3 percent.
A. Before blaming coverage, staff hand-check 15 flagged replies against the rubric's own verdict. They agree on 14 of 15. The scoring checks out.
C. Three candidates: the rubric was worded around the old model's habit of writing vague, generic replies, with no rule checking whether a stated appointment is real; the review queue was sized for when only 8 percent of replies included a date, and got swamped once that share nearly quadrupled; and reviewers were trained to skim for a warm, reassuring tone, never to check a stated time against the dispatch system, because the old model rarely gave them a specific time to check.
E. Pull 40 real replies with a specific date or time, sent after the upgrade, and check each against the dispatch system. Rubric pass rate: 95 percent. Real match to an actual scheduled visit: 52 percent.

Swap the trigger and it still runs

  • Speed: instead of a slow model upgrade rolled out over weeks, the trigger is a same-day switch to a faster model to cut costs. TRACE still starts with what the rubric was ever built to check, not with how fast the switch happened.
  • Cost: the team shrinks the rubric from five checks to three to save review time, on the idea the extras rarely fired anyway. The recut still has to show which kind of mistake got dropped, not just how many checks remain.
  • The model really did get better: the case on this page. Hooklines genuinely got more fluent at writing tight, specific copy. The rubric just never had a rule that could tell a real number apart from a confident-sounding fake one.

Where people run it wrong

  • Treating "the pass rate went up" as proof the rubric still works, instead of asking which mistakes it was ever built to catch.
  • Reading months of a climbing score as improvement everywhere, when a fixed rubric, tuned against long enough, can climb for reasons that have nothing to do with the model's newest failure.
  • Writing the rubric once at launch and never rebuilding it when the model changes, the way nobody redraws a map once new roads get built.

How to use it live

Buy yourself ten seconds by naming the split out loud. "So there's the score against the rules we wrote for the old model, and there's whatever the new model is actually doing now, which might include a mistake nobody wrote a rule for. A rubric frozen at the old model's habits can hide an entire new kind of mistake completely. Let me say how I'd check whether that's happening here." That's not stalling. That's where the real diagnosis starts.

Flashcards (click a card to flip it)

This is a diagnosis question about a rubric quietly going blind, not a habit changing, so these eight test the TRACE moves and the real figures behind them.

1 · THE FRAMEWORK
Which framework fits "how do you write criteria that hold across a model upgrade," and why?
Tap to flip
ANSWER
TRACE. The real task is a diagnosis wearing a how-to question's clothes: why a rubric that kept passing quietly stopped predicting real complaints. BOUND would fit a question about sizing something; this is about why a score and reality disagree.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Astrid Fenmark, eval lead for Hooklines, the subject-line generator inside Brightloop, a marketing platform. She wrote the original five-rule rubric herself, fourteen months before the upgrade that broke it.
3 · THE HABIT
What did the review team stop doing once the pass rate kept climbing?
Tap to flip
ANSWER
Double-checking every flagged subject line by hand against the real campaign. It was easy when only one in eight lines had a number; nobody adjusted it once that share tripled after the upgrade.
4 · THE THREE CAUSES
Name the three named cause candidates behind the gap.
Tap to flip
ANSWER
No rule was ever written for a claim being untrue, a review trigger sized for the old, rarer rate got swamped by the new one, and the review tool's own display hid the evidence a reviewer would need to catch the fake claim.
5 · THE NUMBER
The rubric scored 98 percent, but real claim-bearing subject lines only matched the true campaign data ______ percent of the time.
Tap to flip
ANSWER
43 percent, on a fresh sample of 50 real claim-bearing lines sent after the upgrade. The rubric's own pass rate on those same 50 was 96 percent.
6 · THE CHECK
Name the one test that turned the suspicion into a number.
Tap to flip
ANSWER
Pulling 50 real claim-bearing subject lines sent after the upgrade and checking each stated number against the real campaign record, instead of trusting the rubric's own pass rate on the same lines.
7 · THE FIX
What does the fixed rubric require that the old one didn't?
Tap to flip
ANSWER
A sixth rule checking any stated price, percent, count, or date against real campaign data, sized review capacity that grows with how often that trigger fires, and a re-check of the whole rubric every time the model changes, not just when the pass rate drops.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE on a different product. Which one, and what's the number?
Tap to flip
ANSWER
Fallowmere Property Group's maintenance-reply tool. The rubric scored 95 percent on 40 real time-specific replies, but only 52 percent of those matched an actual scheduled technician visit.

Check yourself Score: 0 / 0

Multiple choice
1. Astrid's team requires every release to beat the last version's pass rate against the same five-rule rubric. What's the actual problem with that process, as written?
  • A. Five rules is too few for a subject-line generator to be judged against.
  • B. It never checks whether a stated fact in the subject line is actually true, so a new kind of mistake can climb underneath a passing score with nothing watching for it.
  • C. The rubric should be thrown out and replaced with a completely different scoring method.
  • D. The pass rate should be lowered until every subject line gets a manual review.
Show hint
Look at what the five rules check, and what they never ask about a subject line at all.
Show answer
B. The rubric names five things to check and nothing else. There was never a rule asking whether a stated number was real, so nothing in the process could catch it, no matter how high the pass rate climbed.
Fill in the blank
2. Claim-bearing subject lines, ones stating a price, percent, count, or date, grew from ______ percent of volume before the upgrade to ______ percent after it.
Show hint
Look at the segment chart's note under the grouped bars.
Show answer
12 percent, 34 percent. The new model got far more fluent at sounding specific, which is exactly the kind of subject line the rubric had never been asked to check.
True or false
3. True or false: once the two-person check confirmed the automatic rubric agreed with a human on 19 of 20 flagged lines, that proved the rubric fully covered the new model's real mistakes.
  • True
  • False
Show hint
Confirming the scoring agrees with itself only rules out one kind of problem.
Show answer
False. That check (the A step) only ruled out a broken grader. It took the segment recut and the 50-line evidence test (R and E) to show the rubric was missing the whole kind of mistake behind most real complaints.
Short answer
4. Name a place in Hooklines' rubric where you'd leave the current checks exactly as they are, and say why.
Show hint
Think about the checks that target mistakes the new model is still making at close to the old rate.
Show answer
Model answer: "Leave the length, caps, and emoji rules exactly as they are. They still catch what they always caught, and rewriting rules that already work spends time and review attention the real gap, the fabricated-claim rule, actually needs."
Multiple choice
5. A teammate says the real fix is simpler: just write ten more length and caps test cases, since those rules matter most for brand safety. Why doesn't that fix what's wrong?
  • A. Because writing more test cases would take too many weeks to build.
  • B. Because more examples for a rule the team already trusts still skip the fabricated-claim mistake entirely, since nobody would think to test for something they haven't written a rule to check yet.
  • C. Because the rubric's scoring already disagreed with the manual check, so no amount of examples fixes it.
  • D. Because clients would stop trusting Brightloop's send button entirely.
Show hint
One answer treats the count of test cases as the problem. The rest of the answer says the missing rule is the problem.
Show answer
B. More test cases for a rule that already exists repeat the same blind spot. The count grows; the kind of mistake covered doesn't.
Short answer, apply it yourself
6. Think of a checklist, rubric, or test you rely on somewhere, a hiring scorecard, a code review checklist, a grading rubric, that was written against one version of whatever it judges. What's one way the thing being judged could change shape while the checklist kept passing it?
Show hint
Look for a case where what the checklist judges changed after the checklist was written, and nobody rewrote it.
Show answer
Model answer: "A code review checklist built for one codebase can keep passing pull requests cleanly even after the team adopts a new framework nobody wrote a checklist item for, so every review looks thorough on paper while a whole new class of real bugs slips through unchecked." Any honest answer works if it names a real case where what's being judged grew past the checklist and the checklist never grew with it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more