ConceptAdvancedEval-Driven Specification / Writing an eval spec / #17
Explain how you would version and archive eval sets over time.
The direct answer
Version every eval set the way you'd version code. Freeze a dated, read-only snapshot every time you score against it, log every edit, a fixed label, an added case, a removed case, as a dated event with a reason, and keep every past snapshot archived for good. Every so often, re-run the current model against an old snapshot on purpose, so you can tell whether a rising score is real progress or a test that quietly got easier.
Do this, in order
Freeze a dated, read-only snapshot of the eval set every time you score against it, and never edit that file again.Why: this is the one habit that makes "the score changed" mean something specific, instead of a question nobody can answer.
Log every edit, a fixed label, an added case, a removed case, as a dated event with a reason, somewhere the whole team can see.Why: a quiet edit and a real model gain look identical on a weekly report; the log is the only thing that tells them apart.
Archive every past snapshot forever, not just the current one.Why: you can't check an old score against a new model if the old test no longer exists anywhere.
Every so often, re-run the current, frozen model against an old snapshot before trusting a rising score.Why: this is the one check that catches drift before it turns into a decision to ship the wrong version.
Put a sign-off step on adding or removing cases, the same weight as changing what a metric means.Why: quiet cleanup is exactly how an eval set drifts without anyone deciding it should.
Leave one-afternoon sanity checks unversioned.Why: a scratch test that never leaves one person's machine and never feeds a ship decision doesn't need the ceremony; save it for the eval set people actually quote.
How to answer this, stage by stage
Eight moves. The trap in this question is treating "keep good records" as the answer, so half the stages are spent proving why a plain version number, not a tidy folder, is the thing that actually saves you.
1
Pin it to one real eval set and number
Say it like this
"Let me put a number on this. Say there's a staffing company called Denman Staffing, and their tool Shortlist reads a stack of resumes and ranks them for a recruiter. Priit Kallas runs eval for it, and the whole team trusts one number: how often Shortlist's ranking matches what a real recruiter would have decided, checked against a 300-case eval set."
Why this works
Naming the company, the tool, and the person turns "version your eval sets" from advice into one real file with a real owner, which the rest of the answer needs to stand on.
2
Say your structure out loud
Say it like this
"I'd walk this like a diagnosis, not a checklist. Timeline, recut, assume nothing, cause candidates, evidence test. TRACE. The question sounds like a how-to, but the real job is working out why a climbing score can stop meaning what everyone assumes it means."
Why this works
Naming the plan up front tells the interviewer you're not about to ramble through a list of best practices.
3
Reframe what "version and archive" is actually for
Say it like this
"Versioning an eval set isn't paperwork. It's the only way to answer one question later: did the model get better, or did the test get easier? Without a dated copy from three months ago sitting somewhere real, you can't even ask that question, let alone answer it."
Why this works
This is the whole question in two sentences. Skip it and the rest sounds like generic advice about keeping tidy files.
4
Give the one decision
Say it like this
"So here's what I'd actually do. Freeze a dated, read-only copy of the eval set every time you score against it. Log every edit, relabeled, added, removed, with a date and a reason. Keep every old snapshot around forever. And every so often, run the current model against an old snapshot on purpose, just to check."
Why this works
Matches the direct answer. Naming the exact mechanism, not a category like "better process," is what a strong candidate says here.
5
Walk the timeline
Say it like this
"Here's the timeline. Priit built a 285-case eval set at launch, month one, score 81 percent. Month two, he fixed 14 mislabeled cases in the same file. By month three, leadership was already quoting the number, whatever it was that week, as a stable fact in the board deck, 'Shortlist matches a real recruiter.' That same month, he added 40 new cases for a warehouse-ops launch, no date logged. Month four, he quietly dropped 25 cases from a line of business Denman stopped serving. By then the score read 94."
Why this works
Real dates turn "the eval set drifted" from a claim into something anyone could check, and show the citation started before the edits even stopped.
6
Show the recut and rule out the model
Say it like this
"Here's the recut. Same frozen model, v4.2, never touched across that whole stretch. Score it against today's 300-case eval set: 94 percent. Score that exact same model against the eval set exactly as it existed three months ago, 285 cases, archived by luck rather than by any policy: 86 percent. Same model, both times. Eight points apart. That rules the model out. The eval set is the only thing left that changed."
Why this works
Holding the model constant is what makes the eval set the only remaining explanation, instead of a suspicion.
7
Name the three ways it drifted, and the one check
Say it like this
"There are three ways an eval set drifts like this without anyone deciding it should. Someone quietly fixes a mislabeled case. Someone adds new cases and nobody notes when. Someone drops old ones as 'no longer relevant.' The one check that proves it: pull an old snapshot, if one exists, and score today's model against it. Priit only had one because an old export happened to sit in someone's downloads folder. That's not a system. That's a coin flip that landed right."
Why this works
Naming three specific, checkable mechanisms beats a vague "the data changed," and shows exactly what the fix removes: luck.
8
Close on the one line
Say it like this
"So that's the answer. Freeze it, log it, archive it, and check it against the model on purpose, not by luck. A score you can't compare to itself six months back isn't a measurement. It's a mood with two decimal places."
Why this works
Ends on the decision and a line the interviewer remembers, not a summary of everything already said.
Let's learn
Shortlist is a tool inside Denman Staffing that reads a stack of resumes for one open job and puts the strongest matches on top, so a recruiter looks at those first.
Knowledge spark: what's an eval set?
A stack of real cases with the right answer already written down next to each one. For Shortlist, one case is a resume, a job, and what a real recruiter actually decided. Score the model against the whole stack, and the percent it gets right is the number everyone watches.
When Denman built Shortlist, Priit put together a test: 285 real cases, each one a resume and a job paired with what a real recruiter had actually decided. He ran the model against that test and got 81 out of 100 right. Nobody argued with the number, because nobody had touched the test.
Here is the turn. That climb is real, on the weekly report. But the test itself got edited three separate times in those months, and not one of those edits got written down anywhere. Some of the 13-point jump is the model. Some of it is a test that quietly stopped being the same test, and right now nobody at Denman can say how much is which.
The model didn't get better by 13 points. Some of that is the model. Some of it is a test that quietly stopped being the same test.
At its worst, this costs Denman a model version that scores higher on a test that's gotten easier, while every recruiter using Shortlist for real hiring gets slightly worse rankings than before, and nobody finds out until candidates who never should have made the shortlist start making it.
The gap between when the test kept changing and when it got called stable
The choice I would take back. Fixing a mislabeled case in the same file, back in month two, felt like basic housekeeping. Adding new cases when a job category launched felt the same way. Dropping old cases for a line of business Denman stopped working in felt like cleanup, not a decision. I would take back treating the eval set like a spreadsheet you tidy. I would freeze a dated copy every time anyone scores against it, and make every edit, a fix, an add, a drop, its own dated, logged event from day one.
The decision that mattered
Freeze a dated snapshot every time you score, log every edit as its own event, archive every snapshot forever, and re-check the current model against an old one on purpose. Not because any single edit was careless. Because a rising score that nobody can compare to itself six months back will always look like progress, right up until someone almost ships the wrong version on the strength of it.
What I would leave alone. A one-afternoon sanity check, run by one engineer on their own machine to see if a prompt tweak is even worth pursuing, doesn't need any of this. It never gets quoted outside that afternoon and never feeds a ship decision. Save the ceremony for the eval set people actually cite.
The lesson. A test is supposed to answer one question: did the thing being measured get better? The day the test itself can change without anyone writing it down, it stops being able to answer that question, and it keeps handing out a climbing score anyway.
The week a new hire asked what test they were even running
Read the short version above if you're short on time. This is the long version, for when you want to feel exactly where those five months went.
Priit Kallas has run eval for Shortlist since before real recruiters ever touched it. He built the first 285-case test himself, sitting with two recruiters for a week, pulling real resumes and real job openings and writing down, case by case, what a recruiter would actually decide.
For nine months, the best part of Priit's Monday was one number in a shared doc: how often Shortlist's ranking matched what a recruiter would have picked, run against that same 285-case test. Eighty-one out of a hundred, that first Monday. Nobody argued with it, because nobody had touched the test.
Then, over the next few months, the number climbed. Eighty-four. Eighty-nine. Ninety-four. Each Monday looked like the one before it, a slightly better score on what everyone assumed was the same test. Leadership started saying it in meetings: "Shortlist matches a real recruiter ninety-four percent of the time now." It became the line in the board deck.
Nobody was hiding anything. Somewhere in those months, Priit had fixed 14 cases where the recorded "right answer" had actually been wrong. He'd added 40 new cases once Denman started placing warehouse-ops roles. He'd dropped 25 old cases for print-production roles, a line of business Denman didn't work in any more. Each change felt small enough to just make, in the same file, on an ordinary Tuesday, without telling anyone or dating anything.
Then a new hire, three weeks into the eval team, looked at the climbing line in the shared doc and asked a question nobody had asked in five months: "Is this the same test as January?"
Priit didn't know. That was the whole problem, said out loud for the first time.
So he went looking. He found one saved copy of the eval set, 285 cases, sitting in an old export in someone's downloads folder, timestamped three months back, more by luck than by any system. He froze the current model, the exact version they were about to sign off on shipping, and ran it against both files.
Against today's 300-case test: 94 percent. Against the 285-case test from three months back: 86 percent. Same model, both runs. Eight points apart. Nothing about the model had changed in that window. The test had.
We didn't nearly ship a better version of Shortlist. We nearly shipped one wearing a better score than it earned.
I want to say the model got worse. It didn't, not really. It's roughly the model it always was. The test underneath it changed shape three separate times, quietly, and the climbing number never once said so.
So here is the decision I would take back. When Priit fixed that first mislabeled case, back in month two, editing it in the same file felt like the sensible, boring thing anyone would do. Nobody would call that reckless. I would take back treating the eval set like a spreadsheet you tidy, instead of a versioned file you freeze. Fix the label in a new dated copy. Add the cases in a new dated copy. Keep every one of them, forever, so the question the new hire asked has an answer that doesn't depend on someone's downloads folder.
That is the whole difference. One eval set gets edited quietly and trusted completely. The other gets frozen, dated, and checked against itself on purpose.
And the part I'd want to tell myself, if I could go back: we built a test to answer one question, is the model getting better, and then we let the test itself change without ever asking whether it could still answer that question at all.
What the frozen model actually proved
Before trusting the eight-point gap, Priit's team checked whether the scoring itself was even right. Two people hand-checked 20 of the disagreements between the model and the recruiter's real call. They agreed with the automatic score on 19 of 20. The grading was not the problem. That left the eval set.
Same frozen model, two versions of the same test
94%
86%
Model v4.2 vs. today's eval set 300 cases, current file
Model v4.2 vs. the archived snapshot 285 cases, three months back
Scored against today's eval set
Scored against the three-months-back snapshot
The model never changed between these two runs. The eight-point gap is the eval set, and only the eval set, since the same frozen version answered both tests.
Trusting the weekly number alone
94 percent on the current eval set, month five
0 of the three edits behind that score were dated, logged, or archived anywhere
Re-running the frozen model against an old snapshot
Done the week the new hire asked the question
8 points of the climb came from the test changing, not the model, once the same model answered both files
Three ways an eval set drifts without anyone deciding it should
Not because anyone cut a corner. Each of these, on its own, looks like ordinary housekeeping. Together, over five months, they changed what the test was measuring.
Three separate, checkable ways the same eval set stopped being the same test
Way 1
A fix nobody dated.
Priit fixed 14 cases where the recorded "right answer" didn't match what had actually happened. The fix itself was correct. Nobody wrote down when he made it, or kept a copy of what the file looked like the day before.
How you'd check it: ask whether anyone can produce the eval set exactly as it looked before the last relabel. If the answer is "let me check my downloads folder," there's a gap.
Way 2
New cases added without a date.
Forty warehouse-ops cases went in the same week Denman launched that job category. Nobody tagged when, and nobody flagged that the blended score now included forty cases the earlier model version was never tested against.
How you'd check it: compare today's case count against what last quarter's report actually said. If the report never mentioned a count at all, that's the gap.
Way 3
Old cases dropped as "no longer relevant."
The 25 hardest cases, all from a line of business Denman stopped serving, came out in one afternoon. They also happened to be the cases the model struggled with most, so a defensible cleanup quietly made the remaining test easier.
How you'd check it: pull whatever got removed and score the current model on it separately, if it still exists anywhere. If it doesn't, that absence is itself the finding.
TRACE, read straight off Shortlist's eval set
This reads like a question about good record-keeping, but the real job is diagnosis: work out why a score everyone trusted quietly stopped being comparable to itself. BOUND would fit a question about sizing something from scratch; here the number already exists, and the job is working out what it's still measuring.
T, timeline. The eval set went in at month one, 285 cases, score 81 percent. It was edited three times across the next few months, a relabel, an add, a removal, none of them dated or logged. Leadership started quoting the score as a stable, comparable fact starting month three, one month before the last of those three edits even happened.
R, recut. Score the same frozen model two ways. Against today's 300-case eval set: 94 percent. Against the archived 285-case snapshot from three months back: 86 percent. Same model, both runs. An eight-point gap that has nothing to do with the model, because the model never moved.
A, assume nothing. Before blaming the eval set, rule out a broken grader. Two people hand-checked 20 disagreements between the model and the real recruiter call, and agreed with the automatic score on 19. The scoring was fine. Holding the model constant across both runs is what rules the model out as the reason the number moved. That leaves the eval set as the only variable left.
C, cause candidates. Three, named and separate: a mislabeled case fixed quietly, with no dated record of what changed; new cases added the same month a job category launched, with nobody noting when; and old cases dropped as "no longer relevant," which happened to be the cases the model found hardest.
E, evidence test. Pull an archived copy of the eval set from before the edits, if one exists, and score today's frozen model against it. Denman had one only because an old export sat in someone's downloads folder. A rebuilt process makes that copy exist on purpose, every time, instead of by luck.
Why the evidence test is the hard step
Anyone can suspect an eval set has drifted. The evidence test turns that suspicion into two numbers on the same model: what it scores on the test today, and what it scores on the test from before the edits. Do that comparison and you've checked something real. Call a rising score "probably fine" without it, and you've only said the same worry in a more confident voice.
Same blind spot, a translation eval set nobody dated either
Nordvale Translations runs a tool that machine-translates safety warnings for medical device manuals into twelve languages, checked against a bilingual-reviewer eval set. Boyan Hristov, the quality lead, built the original 196-segment eval set nine months before a similar drift caught his attention.
T. The reported agreement score climbed from 88 to 97 percent across four months, on a model version that was never touched during that window. The eval set itself was edited three times in the same stretch, none of it dated. R. The same frozen model scored 97 percent against today's 220-segment eval set, and only 90 percent against an archived copy of the eval set as it stood four months ago, 196 segments. Same model, seven points apart. A. Two reviewers hand-checked 15 flagged segments against the automatic score and agreed on all 15. The scoring itself checked out. C. Three candidates, the same shape as before: nine mistranslated reference answers quietly fixed, thirty new segments added for a newly supported language with no date noted, and eighteen old segments dropped for a discontinued device line. E. Pulling the one archived copy that happened to exist, a reviewer's own backup download, and scoring today's frozen model against it, was the only reason anyone could tell real progress apart from a changed test at all.
Swap the trigger and it still runs
Speed: instead of drifting quietly over months, the eval set gets rebuilt in a single rushed afternoon before a launch deadline. TRACE still starts with what the eval set was measuring before that afternoon, not with how fast the rebuild happened.
Cost: the team trims the eval set from 300 cases to 150 to cut review time, keeping the cases that are quick to check and dropping the slow ones, without tracking which is which. The recut still has to show which kind of case got cut, not just how many are left.
The model really did get better: the case on this page, partly. Priit's relabel and his warehouse-ops additions were legitimate improvements. Only the removal of the hardest cases inflated the score past what real progress had actually earned.
Where people run it wrong
Treating a climbing score as proof of real progress, without ever asking whether the eval set is the same test it was last quarter.
Editing the eval set "to fix an obvious mistake" without writing down what changed, because it feels like tidying, not a decision.
Believing an archive exists just because a file happens to sit in someone's downloads folder, instead of a place the whole team can find on purpose.
How to use it live
Buy yourself ten seconds by naming the split out loud. "So there's the score against the test we built back at launch, and there's whatever that test has quietly turned into since. A climbing score can hide a test that's stopped being the same test. Let me say how I'd check whether that's happening here." That's not stalling. That's where the real diagnosis starts.
Flashcards (click a card to flip it)
This is a diagnosis question about a test quietly changing shape, not a habit fading, so these eight test the TRACE moves and the real numbers behind them.
1 · THE FRAMEWORK
Which framework fits "how would you version and archive eval sets over time," and why?
Tap to flip
ANSWER
TRACE. It reads like a how-to question, but the real job is diagnosis: work out why a climbing score can stop meaning what everyone assumes it means. ORDER would fit a question about sequencing work; here the number and reality have quietly split apart.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priit Kallas, eval and product lead for Shortlist, Denman Staffing's resume-ranking tool. He built the original 285-case eval set himself, sitting with two recruiters for a week.
3 · THE HABIT
What did the team stop doing, once the weekly score kept climbing?
Tap to flip
ANSWER
Asking whether the eval set itself had changed. Every fix, add, and removal happened in the same file, undated and unlogged, because each one felt like small housekeeping rather than a decision worth writing down.
4 · THE THREE CAUSES
Name the three named cause candidates behind an eval set drifting like this.
Tap to flip
ANSWER
A mislabeled case quietly fixed, new cases added with nobody noting when, and old cases silently dropped as "no longer relevant."
5 · THE NUMBER
The same frozen model scored 94 percent against today's eval set, but only ______ percent against the eval set as it existed three months back.
Tap to flip
ANSWER
86 percent. Same model, both runs. The eight-point gap comes from the eval set changing, not the model, since the model never moved between the two runs.
6 · THE CHECK
Name the one test that turned the suspicion into a real number.
Tap to flip
ANSWER
Pulling the one archived copy of the old eval set that happened to exist, and scoring today's frozen model against it, instead of trusting the climbing weekly number on its own.
7 · THE FIX
What does "versioning and archiving" actually require, beyond just keeping one saved copy?
Tap to flip
ANSWER
A dated, read-only snapshot every time you score, a logged reason for every single edit, every past snapshot kept forever, and a periodic on-purpose check of the current model against an old snapshot.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE on a different product. Which one, and what's the number?
Tap to flip
ANSWER
Nordvale Translations' safety-warning translation eval set. The frozen model scored 97 percent against today's test and 90 percent against the archived snapshot from four months back.
Check yourself Score: 0 / 0
Fill in the blank
1. Denman's eval set held ______ cases when Priit first built it, and ______ cases by the time the new hire asked whether it was still the same test.
Show hint
Look at the launch count in the timeline, and the count in the "same frozen model, two versions" chart.
Show answer
285 and 300. The count moved twice, up by 40 when warehouse-ops cases were added, then down by 25 when print-production cases were removed, with no record of either change.
Multiple choice
2. Priit fixed 14 mislabeled cases in month two, editing the file in place. Was that decision reckless?
A. Yes, editing a live eval file is always reckless, no matter the reason.
B. No, fixing a genuine mislabel was the sensible call. The problem was making that fix without freezing a dated copy first, not the fix itself.
C. Yes, because it added 3 of the 13 points to the score.
D. No, because relabeling never affects the reported score.
Show hint
Ask whether the fix itself was wrong, or whether it was the lack of a dated record that caused the real problem later.
Show answer
B. The relabel was correct and reasonable. What broke things was doing it invisibly, in the same file, so nobody could later tell that edit apart from the two that followed it.
True or false
3. True or false: once the frozen model scored 94 percent against today's eval set and 86 percent against the eval set from three months back, that proved the model itself had gotten 8 points better.
True
False
Show hint
Look at which single variable was held constant across both scoring runs.
Show answer
False. It proved the opposite. The same model produced both numbers, so the eight-point gap has to come from the eval set changing, not the model.
Short answer
4. Name a place at Denman where this exact versioning discipline would be overkill, and say why.
Show hint
Think about a check that never gets cited outside the moment someone runs it, and never feeds a ship decision.
Show answer
Model answer: "Leave a one-afternoon sanity check alone, the kind one engineer runs on their own machine to see if a prompt tweak is worth pursuing. It never gets quoted outside that afternoon and never feeds a ship decision, so freezing and archiving it would spend ceremony the real eval set actually needs."
Short answer, apply it yourself
5. Think of a test, checklist, or rubric you rely on that gets edited over time, a hiring scorecard, a grading key, a QA checklist. What's one way it could change shape while still looking like the same measurement?
Show hint
Look for a case where what's being judged changed after the checklist was written, and nobody rewrote the checklist to match.
Show answer
Model answer: "A hiring scorecard built around one job's requirements can keep scoring candidates cleanly even after the role's actual day-to-day work shifts, so every candidate looks properly evaluated on paper while the scorecard is quietly measuring a job that no longer exists." Any honest answer works if it names a real case where the thing being measured moved and the measurement never moved with it.
Multiple choice
6. If Denman had only relabeled the 14 cases and never added the 40 or removed the 25, roughly what would today's reported score be closer to?
A. 94 percent, same as now.
B. 84 percent, since the relabel alone only accounted for about 3 of the 13 points, and the add and the removal accounted for the rest.
C. 81 percent, unchanged since launch.
D. There's no way to estimate this from the numbers given.
Show hint
Walk the timeline chart point by point and see how much of the climb happened before the add and the removal.
Show answer
B. The reported score sat at 84 percent right after the relabel, before the 40-case add and the 25-case removal each pushed it higher.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.