ConceptAdvancedQuality, Cost & Token Economics / Eval design for product teams / #14
Explain how to design an eval that will still be meaningful in a year.
A steady accuracy score only proves the model still agrees with the answer key. It proves nothing about whether the answer key is still true.
The direct answer
Split the eval into two layers and grade them apart. A skill layer checks things that never go out of date, like whether the tool found and cited the right clause, and barely needs touching. A content layer checks the answer against the rule actually in force today, and its health gets tracked as its own number: staleness, the share of golden cases whose recorded answer no longer matches the current rule. Refresh the content layer when staleness crosses a set line, not on a fixed date, because the accuracy score will hold steady long after the rulebook underneath it has moved.
Do this, in order
Split the eval into a stable skill layer and a rotating rule layer, and grade them apart.Why: citing the right clause is a skill that does not go stale, but the rule text behind it does, and grading both the same way hides which one just broke.
Track golden set staleness as its own number, next to the accuracy score, never folded into it.Why: the accuracy score only checks agreement with the answer key, and it can hold steady long after the key stops matching the real rule.
Refresh the golden set when staleness crosses a line, not on a calendar date.Why: rules change in bursts, so a quarterly refresh either finds nothing or finds it months late.
The moment a case gets flagged stale, regrade it against the current rule before trusting the headline score again.Why: this is the only way to know how much of that steady number was ever real.
Set a threshold that blocks new filings in a category from shipping once its staleness gets too high.Why: a small stale slice is a watch item, a large one on a real filing type is a reason to stop and fix it first.
Break staleness down by rule category, not company wide.Why: one fast moving category can rot while a blended number still looks fine, the same way a small average dip can hide one slice falling hard.
How to answer this, stage by stage
Nobody is grading whether you can define a golden set. They are grading whether you know a flat score can hide a rulebook that already moved.
1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one tool. Redliner is a compliance review tool at Winterbourne Trust. It reads a loan disclosure or a regulatory filing before it goes out and checks every clause against the rule that clause is supposed to follow. Romilly Isham runs eval quality on it."
Why this works
An abstract "how do you keep an eval meaningful" answer turns into a definitions lecture fast. One product makes the two layer split a real decision instead.
2
Say your structure out loud
Say it like this
"I'm going to answer two questions. First, what number would tell me the eval is going stale, weeks before the score itself shows it. Second, what I'd actually do differently once that number moves. Most answers to this stop at 'refresh it sometimes,' and that's the part that doesn't survive a real year."
Why this works
Tells the interviewer you have a plan and a payoff, not just a definition to recite.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me how to build one good eval. It's asking whether I know an eval can look perfectly healthy while it's grading against a rulebook that's already out of date, and whether I'd catch that before someone outside the company does."
Why this works
Stops you giving the generic answer, build a golden set and check it sometimes, which is exactly the design that misses a year of drift.
4
Give the one decision, plainly
Say it like this
"Split the eval into a skill layer and a content layer. The skill layer checks things like 'did it find and cite the right clause,' and barely needs touching. The content layer checks the answer against today's rule, and I track a separate number for it: staleness, the share of golden cases whose recorded answer no longer matches the current rule. I refresh the content layer when staleness crosses a line, not on a date."
Why this works
This is the actual answer, in one breath, with the mechanism named.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without it. Redliner's golden set had 220 cases and scored 94 percent, steady, for a year. Nobody checked whether the rules those cases were graded against were still the current ones. Thirty one of them weren't. Regraded against the current rule, that slice scored 61 percent, not 94, and one of those stale cases had already cleared a real filing months earlier."
Why this works
Shows the real cost of skipping the split, not just the mechanism behind it.
6
Say what you'd leave alone
Say it like this
"I wouldn't put the skill layer on the same fast refresh schedule. Whether Redliner found and cited the right clause doesn't change when a rule's wording does, so rechecking those cases weekly would just be busywork with nothing new to catch."
Why this works
Shows judgment instead of blanket caution applied evenly everywhere.
7
Close on the decision, not the story
Say it like this
"So: two layers, graded apart, and one leading number, staleness, that moves before the accuracy score ever would."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd actually reuse.
Let's learn
Redliner reads a loan disclosure or a regulatory filing before it goes out, and checks every clause against the rule that clause is supposed to follow, flagging anything that no longer lines up.
Before Redliner, a compliance analyst at Winterbourne Trust checked a filing by hand against a binder of rules, about three hours for a normal one. A busy week meant filings sat in a queue for two or three days waiting for a free analyst.
Knowledge spark: what's a golden set?
A pile of real cases a person checks by hand and holds back, on purpose never used to train the model. It's the answer key everything else gets measured against. The whole point of this answer is that the key itself can quietly stop being true.
Redliner checks the same filing in under two minutes now. Weekly, against a 220 case golden set built the year it launched, it scored 94 percent, and it held there, month after month, without much drama.
The problem was never that 94 percent. The problem is what a steady 94 percent actually proves: that Redliner still agrees with the answer key, and nothing about whether the answer key still agrees with the rule.
Nobody was watching for that second thing, because nothing on the dashboard was built to show it. The eval had one number, and that number looked fine every single week.
Golden set staleness, month 0 to month 12
Share of golden cases graded against a replaced rule
Staleness sits at zero for three months, jumps to about fourteen percent the month the disclosure rule changed, and stays there, unflagged, for eight more months, until a new analyst asks why one case still cites the old version.
The team's first read of the flat 94 percent was the natural one. Nothing to see here, the tool is working. That read was honest and also wrong, because the score was never designed to notice a rule changing underneath it. It could only ever tell you whether Redliner still agreed with what the golden set said, back when the golden set was built.
Blended accuracy vs the stale slice, regraded against today's rule
What the dashboard showed all yearWhat the same cases scored against the real rule
The blended score never dropped below 93 percent all year. The same 31 cases, checked against the rule actually in force, scored 61 percent, a 33 point gap the headline number never showed anyone.
The choice that mattered
The team locked the golden set the day Redliner launched, so scores would stay comparable release over release. That made sense at the time. Building 220 cases and getting to a working score was already the hard part, and the rulebook barely moved in year one. Nobody came back to add a plan for what happens once it does move.
At its worst, a filing goes out citing a rule that's already been replaced, and it clears review because Redliner and the answer key agree with each other, both of them wrong in the same direction. The company only finds out when a real question gets asked about it, not when the eval score changes, because the eval score never changes.
What I'd leave alone: the skill layer, the part of the eval that checks whether Redliner found and cited the right clause in the first place. That skill doesn't decay when a rule's wording changes, so putting it on the same fast refresh cycle as the content layer would just burn hours checking something that was never at risk.
The lesson: a score is only as honest as the thing it's compared to. If the rulebook underneath the score can quietly move and nobody checks, the score keeps handing you the same nice number long after it's stopped meaning what everyone assumes it means. Track whether the answer key is still true. Don't just track whether the model still agrees with it.
Now here is the same thing as a story
Read the long version below when you want to feel why a flat 94 percent didn't mean what everyone assumed, not just be told the number to watch.
For years before Redliner existed, Romilly Isham was the person at Winterbourne Trust who checked a filing against the rule binder by hand. She knew which clause mattered for which rule the way some people know a shortcut through a building. Ask her whether a disclosure clause was compliant and she could usually tell you before she'd finished reading it twice.
She built Redliner's golden set herself, the year it launched. Two hundred and twenty real cases, each one checked against the rule that applied to it at the time, held back and never used to train the model. For the first few months she reread the actual rule text behind every case as she added it, slowly, carefully, the way you'd expect from someone who used to do this work by hand.
The score came back at 94 percent that first month, and it barely moved after that. Ninety three, ninety four, ninety five. Week after week, the same steady number. Somewhere in there, without deciding to, she stopped rereading the rule text behind a case once it was already locked into the set. The score kept saying it was fine. Why would she go check something the score already vouched for.
Then, a full year in, a new analyst who'd joined that spring was working through onboarding and pulled up the golden set to see how it worked. She asked Romilly a plain question: why does this case cite a version of the disclosure rule that got replaced last spring?
Romilly couldn't answer it. She went back through all 220 cases by hand over the better part of two days, checking each cited rule against the version currently in force. Thirty one were wrong. All thirty one had been graded, every single week for eight months, against a rule that no longer existed.
We didn't lose thirty one test cases. We lost a year of not knowing which parts of a steady number were still true.
The decision that opened the door went back to the week Redliner launched. In a short planning meeting, the team agreed to lock the golden set once it was built, so a score in month six would mean the same thing as a score in month one. That made real sense at one deploy a quarter, comparing scores across releases needs a fixed ruler. Nobody in that meeting asked what happens when the rulebook itself, not the model, is the thing that moves.
Run the same year again with one change. Track staleness from day one, the share of golden cases whose recorded answer no longer matches the current rule, checked weekly on a rotating sample. The disclosure rule changes in month four. Within a week of that change being logged, the staleness reading jumps from zero to a flagged number, and the flagged category gets pulled and rechecked before it ever clears another filing. Instead of eight months hidden and a two day manual audit after the fact, it's caught inside a week, at the cost of a routine check nobody has to remember to run.
One design trusted a single steady score to speak for a rulebook that never stopped moving underneath it. The other design asks, every week, whether the thing the score is being compared to is still true.
What I'd tell myself, back in that short planning meeting: the moment you lock an answer key so scores stay comparable, ask what happens the day the key itself goes out of date, not just the model. Nobody asked. That's on the room, not on Romilly.
LEAD, for a metric that has to outlive the rulebook it's checking
Not "did the score hold steady." LEAD asks what would have told you the truth underneath the score, weeks before it mattered.
LLink. What business outcome does this eval actually protect?
A filing that matches the rule currently in force, not a score that shows Redliner agreeing with a recorded answer. Those are two different claims, and only the first one is the thing anyone actually cares about.
Name the outcome before the metric, or the metric ends up designed to look good on a dashboard instead of catching a bad filing.
EEarly signal. What moves weeks before the outcome would?
Golden set staleness rate, the share of the 220 cases whose recorded answer no longer matches today's rule. Accuracy sat at 94 percent the whole time. Staleness would have crossed a warning line eight months before a new analyst caught it by asking a plain question.
The strong line: name the metric that would look perfectly healthy right up until the day someone checks the underlying rule by hand. Accuracy did exactly that here.
AAbuse. How does this metric get hit without doing the real work?
Two ways. Leave the golden set untouched forever, the score stays high and nobody spends the time. Or, worse, quietly edit the recorded answer to match whatever the model currently outputs instead of checking it against the real rule.
The second one is the dangerous one. It makes the staleness number look clean while the eval slowly turns into a machine that only ever agrees with itself.
DDecision. What do you actually do differently at each threshold?
Under 5 percent staleness, normal refresh cadence, nothing urgent. 5 to 15 percent, flag the category and refresh it before the next release ships. Above 15 percent, or any staleness at all in a high stakes filing category, freeze that category's production use until the golden set is caught up.
A number nobody acts on is a chart on a wall. Naming what changes at each line is what makes staleness a metric instead of decoration.
Three things worth stating directly, since this is where the real judgment sits. The alternative I rejected was refreshing the golden set on a fixed calendar, say every quarter, instead of tracking staleness as its own number. It lost because rule changes are bursty, not evenly spaced. A quarterly refresh either lands on a quiet quarter and finds nothing, or lands four months after a real change and catches it almost as late as never checking at all. The AI specific failure worth naming by name is silent answer key drift, the eval's own ground truth going out of date with no signal anywhere that it happened, since the model keeps confidently agreeing with a key that's already wrong. The guardrail is the staleness layer itself, checked weekly, threshold triggered rather than calendar triggered. That guardrail isn't free either. Rechecking all 220 golden cases against the live rule database every single day would cost an analyst most of a day, every day. So the check runs weekly instead, on a sample large enough to catch a real jump, trading a few days of detection lag for a fraction of the cost. And the bar it enforces was never zero staleness. A rulebook this size is always changing somewhere, so a handful of cases going stale between checks is normal. The bar is that staleness gets caught and flagged within about a week of crossing 15 percent, not that it never happens at all.
And if you want to be sure it really works, try it somewhere else
Same four letters, a factory floor instead of a bank, and this time the leading number existed. It just got gamed instead of trusted.
Floorwatch is a tablet app floor supervisors use at Gantry Steel Works. Before a shift starts, a supervisor points it at a work area and it checks the setup, guarding, signage, clearance, against the safety code that area falls under. Rafferty Kolinski runs eval quality on it.
After a similar incident spread through the industry, Gantry Steel Works actually built a staleness tracker for Floorwatch's golden set, the same idea Redliner was missing. For ten months it worked exactly as intended. Then a routine code update changed the required clearance distance around one class of machine guarding, and Floorwatch's staleness reading jumped, flagging a real slice of the golden set as stale, right before a scheduled customer audit.
An engineer under deadline pressure to keep the dashboard green cleared the flag himself. Instead of pulling the updated code section and rechecking each flagged case against it, he edited the recorded answer on every flagged case to match whatever Floorwatch was currently outputting. The staleness number dropped back to zero within the hour. Nothing about the real clearance distance had been checked at all.
The decision Rafferty would take back
Nobody required a second person, someone outside the team that built Floorwatch, to sign off before a flagged case's recorded answer could be changed. One engineer, alone, under a deadline, could clear a real staleness flag by editing the key to match the model instead of the code, and nothing in the process stopped him.
The leading number existed. It just got satisfied on paper instead of checked against the thing it was supposed to track.
Same method, a sharper warning: a good leading indicator is not self policing. Redliner's failure was never tracking staleness at all. Floorwatch's failure was tracking it, and letting the same person who owns the score also own the power to edit the answer key alone. Both end in the same place, a clean looking number sitting on top of a rulebook nobody actually checked.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the two numbers, accuracy graded against the key, and staleness, the share of the key that's stopped matching the real rule. That pair alone answers the question.
Cost: there's no budget to recheck the golden set weekly. Recheck it monthly instead, on a sample, and require two sign offs, not one, before any flagged case's answer gets changed. Slower and cheaper, not gone.
The model got better, for real: say the underlying model gets a genuine upgrade and raw accuracy jumps. That's not proof the golden set caught up. Rerun the staleness check before trusting the new headline number. A smarter model can still be graded against an old rule.
Where people run it wrong.
They treat a steady accuracy score as proof the eval is still meaningful, and never check whether the answer key itself is still true.
They refresh the golden set on a calendar instead of a trigger, and miss whatever changed in the months between refreshes.
They let the same team that owns the score also own sole power to edit the answer key, with no second sign off required.
How to use it live. Say the two layer split out loud before describing either one: "before I say what the eval measures, let me separate what I'm checking from what I'm checking it against, because those two things can drift apart on their own." That buys a beat to think and tells the interviewer you already know where this question is headed.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does each letter stand for here?
Tap to flip
ANSWER
LEAD. L is the real outcome, a filing that matches today's rule. E is the early signal, golden set staleness. A is how it gets gamed, freezing the set or rewriting the key to match the model. D is the threshold that actually changes what ships.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Romilly Isham, eval quality lead for Redliner, the filing review tool at Winterbourne Trust. Used to check filings by hand against the rule binder before Redliner existed.
3 · THE HABIT
What did she stop doing because the score held steady?
Tap to flip
ANSWER
She stopped rereading the actual rule text behind each golden case once it was locked into the set, since the score kept vouching for it every week without her checking.
4 · THE REAL SIGNAL
What's the honest number that could have caught this early?
Tap to flip
ANSWER
Golden set staleness rate, the share of the 220 cases whose recorded answer no longer matches the rule currently in force. It would have crossed a warning line at month four, eight months before anyone noticed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Locking the golden set the day Redliner launched, so scores stayed comparable release over release, with no plan for tracking when the answer key itself went out of date.
6 · THE NUMBER
Fill in the blank: the golden set held steady at ___ percent for a year. Regraded against the current rule, the stale slice actually scored ___ percent.
Tap to flip
ANSWER
94 percent, then 61 percent. Thirty one of two hundred twenty cases, fourteen percent, were being graded against a rule that had already been replaced.
7 · THE REPLAY
Same year, new design, what changes?
Tap to flip
ANSWER
Staleness gets tracked from day one. When the rule changes in month four, the reading jumps within a week of the change being logged, instead of sitting hidden for eight months until a new analyst happens to ask.
8 · THE TRANSFER
Section four answers this same question again for a different product. Which one, and what's the twist this time?
Tap to flip
ANSWER
Floorwatch, a safety inspection app at Gantry Steel Works. This time staleness was being tracked, but an engineer cleared a real flag by editing the answer key to match the model instead of checking the real code.
Check yourself Score: 0 / 0
True or false
1. True or false: because Redliner's blended accuracy score held at 94 percent for a year, that proves the golden set's answer key was still correct the whole time.
True
False
Show hint
Look at what the "94 percent" number actually gets compared against in Section 1.
Show answer
False. The score only checks whether Redliner agrees with the recorded answer. It says nothing about whether that recorded answer still matches the rule in force today, which is exactly what went wrong.
Fill in the blank
2. Of the 220 golden set cases, ___ were later found to be graded against a rule that had already been replaced. Regraded against the current rule, that same slice scored ___ percent instead of 94.
Show hint
Check the two bars in the second chart in Section 1.
Show answer
31, then 61 percent. A 33 point gap between what the dashboard showed all year and what the same cases scored against the real rule.
Multiple choice
3. Which of these would have caught Redliner's problem months before the accuracy score ever moved?
A. Checking the blended accuracy score more often, say daily instead of weekly.
B. Tracking golden set staleness, the share of cases whose recorded answer no longer matches the current rule, as its own number.
C. Measuring how fast Redliner responds to a query.
D. Counting how many filings get reviewed each week.
Show hint
Look at the E step in the framework recap.
Show answer
B. Accuracy can hold perfectly steady while the answer key it's graded against goes stale. Staleness is the number that would have moved first.
Short answer, name the rejected alternative
4. What alternative to threshold based staleness tracking does this answer reject, and why does it lose?
Show hint
Look at the "three things worth stating directly" paragraph after the framework recap.
Show answer
Model answer: Refreshing the golden set on a fixed calendar, like every quarter, instead of tracking staleness. It loses because regulation changes in bursts, not evenly, so a calendar refresh either lands on a quiet stretch and finds nothing, or lands months after a real change and catches it almost as late as never checking at all.
Short answer, apply it yourself
5. Pick an AI product you use yourself that gets checked against some fixed rulebook, policy, or reference set. Name one way its answer key could quietly go out of date, and how you'd catch it.
Show hint
Think of a product graded against something that changes on its own schedule, a tax rule, a school calendar, a price list, separate from how good the model itself is.
Show answer
Model answer: A tax prep chatbot's golden set of sample returns could stay graded against last year's tax brackets after the brackets update for the new year. I'd track the share of golden cases whose expected answer still matches this year's published brackets as its own number, checked once the new tax year opens, instead of trusting a flat accuracy score to notice on its own.
Multiple choice
6. At Gantry Steel Works, an engineer cleared a real staleness flag on Floorwatch's golden set by editing the recorded answer to match whatever Floorwatch currently output, instead of checking the real safety code. What does this show about a leading indicator like staleness?
A. Staleness is a weak metric and should be dropped in favor of accuracy alone.
B. A leading indicator can still be gamed if the same team that owns it can edit the answer key alone, with no outside check.
C. This only matters for questions that use the FLIPS framework.
D. It proves generation models can never be evaluated against a changing rulebook.
Show hint
Look at the key point block titled "The decision Rafferty would take back."
Show answer
B. A good number still needs governance around who can change what it's compared against. Without a second sign off, one person under pressure can satisfy the metric on paper without doing the real check.
Before you close the answer
Why this works
Tests whether you understand that a steady score can be lying about what it's compared against, and whether you'd build a leading number to catch that instead of just promising to "keep the golden set updated." Most candidates stop at the promise, with no number attached to it.
Follow-up traps
"Why not just refresh the golden set every quarter and be done with it?" Response: a calendar refresh catches a change that happens to land near the refresh date and misses one that happens the week after, since regulation changes in bursts, not evenly.
"Couldn't someone game the staleness number the same way they game accuracy?" Response: yes, and it happened at Gantry Steel Works. That's why the real fix needs a second sign off before any flagged case's recorded answer gets changed, not just the number itself.
If pressed
The real production bar was never zero staleness. A rulebook this size is always changing somewhere, so a handful of cases going stale between checks is normal. The bar is staleness caught and flagged within about a week of crossing 15 percent, with a second person required to sign off on any change to the answer key itself.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.