CaseIntermediateEval-Driven Specification / Golden datasets and test set ownership / #21

What happens to your golden set when the product's scope expands?

The direct answer
A golden set does not grow with the product on its own. The moment scope expands, check whether the set's own examples still match the real mix of what's being done today. If they do not, pull real examples from the new scope and put them in, and score every retrain against that refreshed set and a live sample of new-scope work side by side, not the old set alone.
Do this, in order
  1. Rebuild the golden set's mix to match today's real routes before trusting any retrain's score.Why: 180 examples were all downtown, single-stop, gas-van routes, while over a third of daily routes are now something else entirely.
  2. Rule out a broken scorer before blaming the set's coverage.Why: if the grading itself disagreed with a manual check, the segment gap would be noise, not proof.
  3. Recut real performance by old scope versus new scope, not one blended number.Why: 94 percent on time downtown, 61 percent on regional routes, both hiding inside one 82 percent average.
  4. Tag every example by zone, vehicle, and stop count, and refresh the set the day scope changes ship, not on a review calendar.Why: regional, electric, and bundled routes launched five months before anyone touched the set.
  5. Give the golden set an owner tied to the launch process, not a standalone eval task nobody remembers to loop in.Why: the team that shipped the scope change never told the team that owns the set.
  6. Leave the downtown segment alone.Why: it already lands 94 percent on time, same as before the expansion, so more review there just slows down what already works.

How to answer this, stage by stage

Eight moves. The timeline itself carries half the argument here, so this one spends two full stages just pinning down the gap between when the roads changed and when anyone touched the test.

1
Pin it to one real product and number
Say it like this
"Let me put a number on this. Say there's a regional delivery company, Cinderpath Delivery, and an AI tool called Pathwise plans the driver's stop order every morning. Zanyar Rahimi, who owns Pathwise's eval, built a golden set of a hundred and eighty route scenarios the month it launched, and every retrain since has had to beat the last one's score against that same set before it ships."
Why this works
Naming the real company, tool, and process gives every later step something concrete to check against, instead of a general worry about old test sets.
2
Say exactly what "scope expanded" means here
Say it like this
"Five months after launch, Cinderpath signed two grocery accounts that needed same-day delivery past the ring road, up to fifty-five miles out. Ops swapped half that fleet to electric vans and let Pathwise bundle up to twenty stops on one run instead of six. That's the scope change. It's not vague. It's a zone, a vehicle, and a stop count, all different from anything the golden set was built from."
Why this works
Turning "scope expanded" into three checkable facts stops the answer from staying abstract.
3
Name exactly why the set stopped representing quality
Say it like this
"Here's the actual problem. Every one of those hundred and eighty examples is a downtown route, six stops or fewer, gas van, no charging stop anywhere in it. A set like that can't grade a route it's never seen the shape of. It's not that the set got worse. It just never got asked to plan the new job."
Why this works
A specific, checkable flaw beats a vague "the set might be a bit stale" every time.
4
Give the fix up front
Say it like this
"So here's what I'd do. Tag every golden-set example by zone, vehicle, and stop count. Check that mix against what's actually running today. Pull in real regional and bundled routes until the set looks like this week, not month one. And score every retrain against both, the golden set and a fresh sample of real new-scope routes, before anyone calls it a win."
Why this works
Matches the direct answer. Naming the fix, not just the flaw, shows you can repair a golden set, not only point at one.
5
Show the timeline: expansion versus last update
Say it like this
"Here's the timeline. The golden set went in at month one. Scope expanded at month five. Nobody touched the set in between, and nobody's touched it since, five retrains later, at month ten. The score climbed from sixty-six to ninety-seven across that whole stretch, on a set that stopped matching the job at month five."
Why this works
Naming the exact gap between "scope changed" and "set updated" turns the drift into a fact on a calendar, not a feeling.
6
Rule out a broken grader before blaming the set
Say it like this
"Before I blame the set's coverage, I'd check whether the score's even being read right. Two dispatchers hand-checked a sample of real routes against what the golden set's rubric would have scored them, and they agreed almost every time. So the grading wasn't the problem. The coverage was."
Why this works
Skipping this is the easy mistake. Blame the examples before ruling out a broken grader and you might fix the wrong thing.
7
Recut real performance, old scope versus new scope
Say it like this
"Split real routes by zone. Downtown, sixty-three percent of daily volume: ninety-four percent land inside their window. Regional, thirty-seven percent of daily volume: sixty-one percent do. And regional routes are barely four percent of the golden set. The set never learned the roads it's now a third of the job."
Why this works
Naming the split, then tying it straight back to how the set was built, is the strongest move TRACE has here.
8
Say what you'd leave alone, then close
Say it like this
"I wouldn't touch how Pathwise plans a downtown route. Ninety-four percent land on time, same as before the expansion, and the golden set already knows that job cold. So: keep the golden set, but stop trusting it for a job it's never seen. A set built for one map will always agree with itself, right up until the road it was never shown."
Why this works
Ending on what stays, not just what changes, shows judgment instead of blanket distrust of the whole set.

Let's learn

Pathwise is a screen a dispatcher already has open before the first van leaves the lot. Every morning it looks at that day's deliveries and works out the order a driver should hit them in.

Knowledge spark: what's a golden set? A small set of example cases a team trusts enough to grade a model against, before letting a new version take over from the one already running. It only stands in for the real job for as long as someone keeps checking that it still looks like the real job.

Before anyone tuned anything, Pathwise's route-match score, how closely its stop order matched what a senior dispatcher would have picked, was rough. It matched only sixty-six out of every hundred of Zanyar's hand-built downtown examples. That felt right. Nobody trusted it yet.

Ten months and five retrains later, that same score sits at ninety-seven. Every retrain beat the one before it. Every retrain replaced what was already live.

Two lines, five retrains over ten months
What the dashboard showed: golden-set match score
climbing, 66 to 97, every retrain wins
Quiet signal: real on-time-window rate, all live routes
drifting down, 90 to 82 percent
m1m3m5, scope expandsm7m9m10
The golden-set score climbed thirty-one points in ten months. The share of real deliveries landing inside their promised window drifted down eight points over the same stretch, starting right around the month scope expanded. An eight-point dip sounds mild. It is hiding a much bigger split.

Here is the turn. Those thirty-one extra points are not the real story. The real story is what never showed up anywhere on that climbing line: how Pathwise plans a route it was never built to test. That split sat underneath the whole ten months, quietly getting worse as regional routes grew from a sliver of the day's work to more than a third of it.

The model did not get worse at planning downtown routes. It just never got asked to plan this one.

At its worst, this costs Cinderpath drivers stranded on roads nobody built a test for. Regional routes, the ones with a far zone, an electric van, or a bundle of stops, only land on time sixty-one percent of the time. And regional routes are no longer rare. They are thirty-seven percent of everything Pathwise plans on a given day.

A hand sketch timeline with five marks: the golden set built with 180 downtown routes, scope expanding at month five to add regional, electric, and bundled routes, two retrain checkmarks continuing to climb, and a red mark at month ten where a van gets stranded on a regional route, labeled the set never heard about this.
The gap between a climbing score and the road it was never shown

The choice I would take back. Building the golden set only from downtown routes, ten months ago, was the sensible call. Regional, electric, and bundled routes did not exist yet. There was nothing else to build it from. I would take back leaving the set frozen there with no trigger tied to scope changes. I would tie a golden-set refresh to every scope launch, the same way a new feature gets a rollout plan.

The decision that mattered Score every retrain against a refreshed set that matches today's real route mix, not a set frozen at launch. Not because the old set is worthless. Because a set that never learned the new roads will always agree with itself, right up until a van is stranded on one of them.

What I would leave alone. The downtown segment does not need touching. It lands ninety-four percent of deliveries on time, same as before the expansion, because that is exactly the job the golden set was built to test. More review there would only slow down the part that already works.

The lesson. A golden set is a snapshot of the world on the day you built it. The day the product's job changes, that snapshot does not know it is out of date, and it keeps grading you as if nothing changed, unless someone tells it otherwise.

A frozen night on a farm road fifty-two miles out

Read the short version above if you are short on time. This is the long version, for when you want to feel exactly where the ten months went.

The dispatch floor at Cinderpath Delivery gets loud around 6am, once the regional runs start rolling out behind the downtown ones. Zanyar Rahimi has run eval and ops scoring for Pathwise for two years. He can look at a week's match score and tell, inside a minute, whether a retrain earned it or got lucky on a slow week.

He built the golden set himself, with two senior dispatchers, the month Pathwise first launched. A hundred and eighty routes pulled from real downtown deliveries, each one checked against what a senior dispatcher would have planned by hand. Fast to build. Easy to trust, because everyone on the team had driven those exact streets.

For the first four months, watching the match score climb felt like the model getting sharper. Sixty-six. Seventy-eight. Eighty-five. Every point felt earned. Somewhere in there, without anyone deciding it on purpose, the team quietly stopped doing something it used to do every few weeks: pull a handful of that week's real routes and check them by hand against what Pathwise had picked. Once the golden score was climbing fast on its own, the manual check started to feel like double work.

Then Cinderpath signed two grocery accounts that needed same-day delivery past the ring road, up to fifty-five miles out. Ops swapped half the exurban fleet to electric vans to meet the contract's terms, and let Pathwise bundle up to twenty stops on one run instead of six, to make the long drive worth it.

Nobody touched the golden set. Why would they. It kept scoring in the nineties.

Three weeks before the holidays, a Pathwise-planned regional route sent a driver fifty-two miles out on a bundled eighteen-stop run, including a charging stop the route had placed wrong. The van sat at four percent battery on a farm road with no signal, for almost three hours. Eleven of the eighteen deliveries on that run were frozen groceries. All eleven had to be written off.

Zanyar's first thought, and everyone's first thought, was that this was a fluke. One bad electric route on one bad day. Then Marisha Tan, the regional account manager, pulled a month of regional deliveries for a client call and found the near miss was not rare at all. In a smaller way, without a stranded van attached to it, it was happening on nearly four regional routes out of every ten.

We did not lose three hours on a farm road. We lost the part of the map the golden set never got shown.

I want to say the model got worse at picking routes. It did not get worse. It just never got asked, not once in a hundred and eighty examples, to plan a route with a charging stop in it, or a bundle past six stops, or a zone past the ring road.

So here is the decision I would take back.

Ten months earlier, building the set only from downtown routes was the sensible call. Regional, electric, and bundled work did not exist yet to pull examples from. I would put a refreshed, tagged sample back alongside it: real regional routes, real bundled routes, checked every time scope changes, not just at launch. Not because the downtown set is wrong. Because of what it would have shown, months earlier: ninety-four percent on time for the job it was built to test, sixty-one percent for the job nobody ever wrote a single example of.

That is the whole difference. One process trusts a score that only ever answered for the roads it already knew. The other checks that score against the roads drivers are actually on tonight.

And the part I would want to tell myself, if I could go back: we built a test that only ever answered for the map we already had. Nobody decided that on purpose. We just never redrew it when the map changed.

What the recut actually showed

Before blaming the golden set's coverage, Zanyar's team checked whether the score was even being read right. Two dispatchers hand-checked a sample of real routes against what the golden set's own rubric said they should score. They agreed on nearly every one. The grading was not the problem. That left the coverage.

Real routes, split by zone
63%
94%
37%
61%
Downtown routes
Under 12 miles, single stop type, gas van
Regional routes
Up to 55 miles, electric van or 7+ stops
Share of daily route volume
On-time delivery-window rate
Downtown routes land on time ninety-four percent of the time, close to what the golden set predicted. Regional routes land on time sixty-one percent of the time, and the golden set has almost none of them: seven of its hundred and eighty examples, about four percent, against thirty-seven percent of real daily routes.
Shipping on the golden score alone
97 on the golden set, ten months of tuning
4 percent of the golden set's own 180 examples covered regional or bundled routes, the same job behind 37 percent of real daily routes
Requiring the real-route check to agree too
Adopted the week the stranded van made the news
0 retrains have shipped since without both checks agreeing

Three ways a golden set stops matching the job

Not because anyone cut a corner on purpose. A set built entirely from the old scope can be met honestly, retrain after retrain, and still stop meaning anything the day the job changes.

Three hand-sketched panels: a tan card labeled frozen at the old map, no charging stops, no bundling; a document labeled old rubric grades the new roads by the wrong checklist; and a card with a question mark labeled two teams, no shared trigger to refresh the set.
Three separate, checkable ways a set can pass and still miss the new job
Way 1
Frozen at the old map. Every example predates the new roads.

All hundred and eighty examples in the golden set are downtown routes: six stops or fewer, a gas van, no charging stop in sight. The set was built before regional, electric, and bundled routes existed to pull examples from, and nobody added any once they did.

How you'd check it: count how many examples in the set are tagged regional, electric, or bundled, and compare that share to the real share of routes now run that way. Four percent against thirty-seven is a real gap.
Way 2
Old rubric, new roads. The scoring criteria never learned the new job.

The match rubric checks whether Pathwise's stop order looks like what a senior dispatcher would pick for a downtown run: order, timing, nothing else. It has no language for where a charging stop should sit in an eighteen-stop bundle. Even the few regional examples in the set get graded on criteria built for a different job.

How you'd check it: list what the rubric actually scores, and check whether it can even represent a charging stop or a bundle order. This one could not.
Way 3
Two teams, no shared trigger. Nobody's job was to tell eval the map changed.

Ops and fleet launched the regional and electric-van expansion. Eval, Zanyar's team, owns the golden set. No process connects the two. The scope change shipped on schedule. The golden set never heard about it.

How you'd check it: ask whether any launch checklist for a scope change includes "update the golden set." At Cinderpath, it did not.

TRACE, for a score that graded the wrong map

This reads like a question about a document, but the real job is diagnosis: work out why a set the team trusted for ten months quietly stopped matching the job. GUARD would fit if the harm were one group losing a fair shot at an appeal; here the harm is a whole kind of route the set never learned existed.

T, timeline. The golden set went in at month one, built entirely from downtown routes. Scope expanded at month five: regional zones, electric vans, and bundled stops all launched together. Nobody touched the set at month five, or at any retrain after it. The real on-time rate started drifting the same month, from ninety to eighty-two percent by month ten, while the golden score kept climbing from sixty-six to ninety-seven over the same stretch.
R, recut. Split real routes by zone. Downtown, sixty-three percent of daily volume: ninety-four percent land on time. Regional, thirty-seven percent of daily volume: sixty-one percent do. A blended eighty-two hid a segment that was already working and a segment nobody was watching.
A, assume nothing. Before blaming the golden set's coverage, two dispatchers hand-checked a sample of real routes against what the set's own rubric said they should score, and agreed almost every time. The scoring itself was fine. The gap was real.
C, cause candidates. Three, named and separate: every one of the hundred and eighty examples predates the regional, electric, and bundled expansion, so the model was never asked to plan that kind of route; the scoring rubric itself has no way to grade a charging stop or a bundle order, built for a job that no longer covers everything Pathwise does; and no process connects the team that launches a scope change to the team that owns the golden set, so the set never heard the map had changed.
E, evidence test. Count how many of the golden set's hundred and eighty examples are regional, electric, or bundled. Seven. About four percent. Real daily volume in that same category: thirty-seven percent. The golden set can't grade a road it was never shown.
Why the evidence test is the hard step Anyone can suspect an old set has fallen behind a bigger product. A test earns its place by turning that suspicion into a number: how much of the set's own mix still matches today's real mix. Do that comparison and you've checked something real. Call a golden set "probably outdated" without it, and you've only said the same worry in a more confident voice.

Same shape, a permit tool that only ever saw houses

Thistlewick County's building department runs PermitLens, a tool that pre-screens permit applications for missing documents before a human reviewer sees them. Ilhan Botros, the model owner, built the acceptance set two years ago: a hundred and thirty applications, all single-family home renovations, scored against a benchmark called PermitCheck. Every quarterly retrain has had to beat the last version's score on that set before it replaces the one screening live applications.

T. PermitCheck's score climbed from seventy-two to ninety-one over six quarterly retrains. The share of commercial and multi-unit applications correctly flagged for missing documents never moved off roughly one in three, the whole time.
R. Split by application type. Residential, sixty-eight percent of volume: PermitLens flags correctly eighty-eight percent of the time. Commercial and multi-unit, thirty-two percent of volume: thirty-four percent.
A. Before blaming coverage, two reviewers blind-recheck thirty flagged applications against PermitLens's own output. They agree on twenty-eight of thirty. The scoring checks out.
C. Three candidates: the acceptance set's hundred and thirty examples are all residential, none commercial; the rubric checks for residential-only documents, a single owner's deed, homeowner's insurance, that a commercial application does not even use, so the few commercial cases in the set get graded on the wrong checklist; and the county's digital-services group added commercial permitting nine months ago without telling the team that owns the acceptance set.
E. Compare the share of commercial and multi-unit applications in the acceptance set. Three of a hundred and thirty, about two percent. Thirty-two percent of real applications now. PermitLens barely tests the application type that is now a third of what it reviews.

Swap the trigger and it still runs

  • Speed: instead of a slow five-month drift, the trigger is a same-week emergency scope change, a partner airline reroutes overflow packages onto Pathwise overnight. TRACE still starts with what the golden set was ever built to represent, not with how fast the change shipped.
  • Cost: the team shrinks the golden set from 180 to 60 examples to cut review time, on the idea a smaller set is close enough. The recut still has to show which routes got cut, not just how many examples remain.
  • The model really did get better: the case on this page. Pathwise genuinely got sharper at downtown routes. The golden set just never had a task that could tell that real gain apart from a fake one on the routes it never tested.

Where people run it wrong

  • Treating "we beat the last version's score" as proof the model works on every kind of route, instead of asking which routes the golden set was ever built from.
  • Reading ten months of a climbing score as the product getting better everywhere, when a fixed set, tuned against long enough, can climb for reasons that have nothing to do with the harder job.
  • Writing the golden set once at launch and never re-deriving it when scope expands, the way nobody circles back to redraw a map once the roads change.

How to use it live

Buy yourself ten seconds by naming the split out loud. "So there's the score against the team's own examples, and there's whatever the product's actually being asked to do today, which might have grown since that set was built. A set frozen at the old scope can hide an entire kind of job completely. Let me say how I'd check whether that's happening here." That's not stalling. That's where the real diagnosis starts.

Flashcards (click a card to flip it)

This is a diagnosis question about a test set outliving its own scope, not a habit changing, so these eight test the TRACE moves and the real numbers behind them.

1 · THE FRAMEWORK
Which framework fits "what happens to your golden set when scope expands," and why?
Tap to flip
ANSWER
TRACE. The real task is a diagnosis wearing a scope question's clothes: why a set that looked fine kept passing while a whole new kind of route slipped through uncovered. GUARD would fit a fairness harm; this is a coverage gap.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zanyar Rahimi, eval and ops lead for Cinderpath Delivery's route planner, Pathwise. He built the 180-route golden set himself, five months before the scope that broke it launched.
3 · THE HABIT
What did the team stop doing once the golden score kept climbing?
Tap to flip
ANSWER
A manual check: pulling that week's real routes and comparing them by hand to what Pathwise picked. It started to feel like double work once the golden score alone looked convincing.
4 · THE THREE WAYS
Name the three ways a golden set stops matching the job after scope expands.
Tap to flip
ANSWER
Frozen at the old map (every example predates the new routes), old rubric on new roads (the scoring criteria can't represent the new job), and two teams with no shared trigger (whoever launches the scope change never tells whoever owns the set).
5 · THE NUMBER
The golden set scored 97, but real regional routes only landed on time ______ of the time.
Tap to flip
ANSWER
61 percent. Thirty-three points under what the golden score alone suggested, and hidden inside a blended 82 percent average.
6 · THE CHECK
Name the one test that turned the suspicion into a number.
Tap to flip
ANSWER
Comparing the share of golden-set examples that are regional, electric, or bundled, about 4 percent of 180, to the real share of daily routes in that same category, 37 percent.
7 · THE FIX
What does the fixed process require that the old one didn't?
Tap to flip
ANSWER
A golden set tagged by zone, vehicle, and stop count, refreshed every time scope changes ship, scored alongside a live sample of real new-scope routes, not a set frozen at launch.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE on a different product. Which one, and what's the number?
Tap to flip
ANSWER
Thistlewick County's PermitLens. PermitCheck scored 91 overall, but caught missing documents on commercial and multi-unit applications only 34 percent of the time, a third of real volume and just 2 percent of the acceptance set.

Check yourself Score: 0 / 0

Multiple choice
1. Zanyar's team requires every retrain to beat the last version's score against the same 180-route golden set. What's the actual problem with that process, as written?
  • A. A hundred and eighty examples is too small a set for a route planner to be judged against.
  • B. It never checks whether the examples still cover the routes Pathwise is actually planning today, so a whole new kind of route can hide inside a passing score.
  • C. The golden set should be thrown out and replaced with a completely different benchmark.
  • D. The score should be lowered until every regional route gets a manual override.
Show hint
Look at what the process checks, and what it never checks, about which kind of route the set represents.
Show answer
B. The process names a score with no check on whether the examples cover the routes the product is actually planning now. That's the real gap, not the count of examples.
Fill in the blank
2. Only about ______ of the golden set's 180 examples were regional, electric, or bundled routes, compared with 37 percent of real daily routes.
Show hint
Look at the evidence-test comparison, and the result strip.
Show answer
4 percent, 7 examples. Nearly every real regional route had no counterpart anywhere in the set.
True or false
3. True or false: once the two dispatchers confirmed Pathwise's real route choices matched what the golden set's rubric said they should, that proved the golden set covered every kind of route well.
  • True
  • False
Show hint
Confirming the scoring agrees with itself only rules out one kind of problem.
Show answer
False. That check (the A step) only ruled out a broken grader. It took the segment recut and the coverage comparison (R and E) to show the set was missing over a third of what Pathwise actually plans.
Short answer
4. Name a place in Cinderpath's process where you'd leave the current golden set exactly as it is, and say why.
Show hint
Think about the segment where the golden set and real routes already agree.
Show answer
Model answer: "Keep it exactly as is for downtown routes, sixty-three percent of daily volume. There, the golden set already reflects the real job, and Pathwise lands ninety-four percent of those on time. A second layer of review there would just slow down the majority of routes that already work."
Multiple choice
5. A teammate says the real fix is simpler: just write fifty more examples for the golden set. Why doesn't that fix what's wrong with it?
  • A. Because writing more examples would take too many weeks.
  • B. Because more examples written the same way, from the same downtown routes the team already knows, would still skip regional, electric, and bundled routes, since nobody on the team would think to write one from scope they haven't built for yet.
  • C. Because the golden set's scoring already disagreed with the manual check, so no amount of examples fixes it.
  • D. Because dispatchers would stop trusting the dashboard entirely.
Show hint
One answer treats the count as the problem. The rest of the answer says the kind of example is the problem.
Show answer
B. More examples from the same old scope repeat the same blind spot. The count grows; the kind of route covered doesn't.
Short answer, apply it yourself
6. Think of a test or checklist you rely on somewhere, a hiring rubric, a certification exam, a code review checklist, that was built for one version of the job. What's one way the job could grow past it while the test kept passing?
Show hint
Look for a case where the job changed shape after the test was written, and nobody rebuilt the test.
Show answer
Model answer: "A hiring rubric built for one role can keep passing candidates cleanly even after the role quietly grows to cover a second skill nobody wrote a rubric question for, so every hire looks like a strong match on paper while the team still has a real gap nobody's measuring." Any honest answer works if it names a real case where the job's scope grew and the test never grew with it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more