CaseAdvancedEval-Driven Specification / Golden datasets and test set ownership / #15
How do you use the golden set during a model migration?
The direct answer
Score the new model against the golden set broken out by category, not by one overall number, and hold the migration until every category holds its own recall. A migration changes how the model actually thinks. A great overall score can hide a brand new blind spot the old pass or fail line was never built to catch.
Do this, in order
Break the golden set out by category before every migration sign off, and hold on the weakest category, not the average.Why: this is the check the whole flip turns on. Skip it and a bad category can hide inside a great overall score.
Treat a migration as its own gate, separate from a routine retrain.Why: swapping the model changes what it's good and bad at. Refreshing the same model on new data doesn't.
Set a real floor per category, for example no category may drop more than a few points against the old model, instead of one pass or fail line.Why: a single threshold can clear easily while one category quietly falls apart underneath it.
Watch real losses by category after go live too, not just the model's own scoring dashboard.Why: the dashboard can hold steady for weeks while one category's real losses climb the whole time.
Leave the category floor off a same architecture retrain.Why: judgment, not blanket caution. The model's failure shape only moves when the architecture does.
How to answer this, stage by stage
Six moves. Most of the weight sits in stage three: this question is really asking what happens the day one great number and six honest ones stop telling the same story. Every stage has the actual words to say.
1
Ground it in one product, one person
Say it like this
"I'll make this concrete. Say a payments processor runs a fraud model that scores every transaction in real time, and a fraud risk lead keeps a golden set, a stack of past transactions with a confirmed right call, to test any new version of that model before it goes live. A migration means swapping the model itself, not just refreshing it, and that's what this question is really about."
Why this works
Nobody can judge how you'd use a golden set against a product and a moment you haven't named yet.
2
Say your structure in one breath
Say it like this
"Five things, fast. Who signs the migration off. What she stopped doing once the score kept coming back better. The two setting switch that snaps. The call I'd take back. And the same migration, replayed with a category floor in place."
Why this works
A route named up front tells the interviewer this isn't being improvised on the spot.
3
Reframe what the question is actually testing
Say it like this
"A golden set doesn't fail during a migration because someone got careless. It fails because a migration changes what the model is good at, and one overall score can rise while it's quietly covering for a category that fell apart."
Why this works
This is what separates a real answer from "just check the golden set," which is the assumption the question is already making.
4
Give the one decision
Say it like this
"Concretely: score the new model against the golden set broken out by category, and hold the migration until every single category holds its own recall, not just the total. An eight or ten point drop in one category is a hold, even if the total is the best score any migration has produced."
Why this works
There's a mechanism and a number in that sentence, not just an instinct to "test thoroughly."
5
Prove it with the compressed failure
Say it like this
"Say a fraud risk lead's golden set covers six fraud patterns. The new model's overall score jumps from 80 to 89 percent, so she signs off the same day. Underneath that number, one category, account takeover, fell from 97 to 79, because the new model reasons about identity well and barely looks at how fast a transaction moves. Three weeks later a fraud ring runs exactly that pattern, and it costs the company 310,000 dollars before the overall score ever moves."
Why this works
Four sentences, and it still lands on the exact moment the total and the truth split apart.
6
Say what you'd measure, what you'd skip, and close
Say it like this
"I'd track every category's own recall against the last model, not just the total. I wouldn't add this floor to a routine retrain on the same model, that's guarding against a problem that only shows up when the model itself changes. So: don't let one good number sign off a migration. Break the golden set out by category, hold on the weakest one, and the score keeps testing the whole job, not just the part the model happens to be good at right now."
Why this works
Shows judgment as well as caution, and ends on the sentence they'll actually repeat back to their own team.
If you remember one thing
A migration doesn't need a bug to hide a regression. It just needs one good number to be easier to trust than six harder ones, and nothing built to check the six before shipping.
Let's learn
Picture a fraud model at a payments processor. Every time a customer pays, it scores the transaction in under a second and decides whether to let it through, hold it, or block it. Before any new version of a model like this ships, it has to beat the old one first.
Knowledge spark: what's a golden set?
A stack of past transactions where someone already confirmed the right call: real fraud, or really fine. New versions of the model get tested against it before anyone trusts them with real money.
For three earlier migrations, the golden set's story was simple. The new model's overall score always came back higher than the old one's, and the six categories underneath it always moved the same direction as the total. So the total became the only number worth reading.
Here's the turn. A migration is not a small update. It changes how the model actually thinks, which means it can get much better at some kinds of fraud and much worse at others, in the same breath. One overall score can rise while it hides that second half completely.
Knowledge spark: what's recall?
Out of every 100 real fraud cases in a category, how many the model actually caught. A model can have great recall overall and terrible recall in one category, and the overall number will never say so on its own.
Recall by fraud category, old model vs. new model, out of 100
Five categories improved. Account takeover, the one category the old model was best at, fell 18 points, from 97 to 79. The overall average across all six still rose from 80 to 89.
The score didn't get better everywhere. It got better everywhere except the one place nobody was still looking.
Account takeover losses, weekly, before and after the migration
Three weeks after go live, account takeover losses hit 310,000 dollars total, against a typical three week run of about 45,000. The golden set's own overall score never moved the whole time.
Twelve hundred transactions make up this golden set, two hundred confirmed cases in each of six categories: card testing, account takeover, synthetic identity, first party fraud, merchant collusion, refund abuse. Every one of those numbers was sitting in the golden set's own breakdown, tagged and ready to read, the whole time.
At its worst, this costs more than the model ever saved. Three hundred and ten thousand dollars moved through one blind spot in three weeks, and the golden set's own overall score never once suggested there was anything left to check.
The decision that mattered
Score the new model against the golden set broken out by category, and hold the migration until every category holds its own recall. Not a bigger golden set. Not a stricter one time review.
What I would leave alone. A routine retrain that keeps the same model, same architecture, just newer data, doesn't need this floor. The model's failure shape doesn't move when the architecture doesn't, so the overall score alone still tells the truth there.
The lesson. A model migration is the one kind of update where a good total number is least trustworthy, because it's the one update big enough to trade a strength for a weakness without anyone deciding to.
Now here is the same thing as a story
Pull this one out when there's room to sit with it, not just tick it off a list.
Yeva Sorokin has read fraud alerts for seven years, the last four of them running fraud risk at Ledgerline Payments, a mid size processor that clears payments for a few thousand online stores. Hand her a flagged transaction and she can usually name which of six fraud patterns it's shaped like before she's finished reading the first line: card testing, account takeover, synthetic identity, first party fraud, merchant collusion, refund abuse.
Ledgerline's fraud model doesn't stay still. It gets replaced every eight or nine months with something meaningfully different, chasing whatever fraud pattern is growing fastest that year. Before any new version goes live, Yeva has to sign off on it, and the golden set is how she does that: twelve hundred real transactions, two hundred per category, each with a confirmed right answer already attached.
The first time she ran a migration, she gave herself a day and a half. She pulled the new model's score for each of the six categories separately, matched it against the old model's, category by category, and only signed off once every single one held its ground. Nothing surprising turned up. The categories all moved the same way the total did.
The second migration, she did the same thing in half a day. Same result: categories and total, moving together.
By the third migration, about eight months before this one, the category checklist had quietly become a second tab on the same dashboard as the total score, folded in during a routine model ops cleanup because running two separate reviews for something that always agreed felt like busywork. Yeva still glanced at the tab. She just stopped opening each category and reading the number underneath it. Everything still matched, so glancing kept being enough.
This time, the fourth migration, the new model's overall score came back at 89, up from 80. That was the best jump any migration had produced. Yeva didn't open the category tab at all. She signed off before lunch.
People are switches, not dials
Three weeks later, an ops analyst on the chargeback team flagged something in the weekly loss report that didn't belong there: account takeover losses were running at nearly seven times their usual rate, and had been climbing every week since the day the new model shipped.
Yeva pulled the golden set apart to find out how that had slipped through. The category was right there in the breakdown the whole time. Account takeover recall had gone from 97 out of 100 under the old model to 79 out of 100 under the new one. Eighteen points, buried inside a total that had gone up nine, because the new model reasoned about identity and device history far better than the old one, and paid much less attention to how fast a string of transactions moved. Every other category had improved. Account takeover alone had fallen off a cliff, and the total was strong enough to hide it completely.
We didn't sign off on a better model. We signed off on a model that traded one strength for another and let the total number keep the secret.
I want to say the model got worse. It didn't, not on the whole. What actually broke was quieter than that. Yeva's only real defense against a trade like this was a habit: opening the category tab and reading every number underneath it. That habit only has two settings. She's either doing it or she isn't. Three migrations of it agreeing with the total, one after another, is exactly what pushes a person from doing it into not doing it, and being right three times running is what keeps anyone from going back to the slower way.
Here's the call I'd take back. Eight months before this migration, during the third migration's own model ops cleanup, someone folded the category by category memo into the same dashboard as the total score, and marked it optional instead of blocking. Nobody meant anything by it. The category memo had agreed with the total for three migrations running, so keeping it as its own required, blocking step looked like duplicate work for a check that always said the same thing as the number sitting right above it.
I'd put the block back. Not a bigger golden set. Not a longer review. Just this: no migration ships until every category holds its own recall, checked on its own, whether the total looks great or not.
Run the same fourth migration again, with that block restored. The category tab isn't optional this time, it's the thing that has to clear before the migration can ship. Account takeover recall shows up at 79 against the old model's 97, an eighteen point drop, well past any floor Yeva would set. The migration holds on day two of the review, the same week, instead of three weeks into real losses. The team blends the old model's speed based scoring back in for anything shaped like account takeover, recall comes back to 95 in the next pass, and the migration ships four days later than planned, with the 310,000 dollars never spent at all, and Yeva not spending her next two weekends walking chargeback logs by hand the way she did after the real one.
If I'm honest, the mistake wasn't trusting three good migrations in a row. Every one of those really was fine. The mistake was writing a checklist that only asked whether the category numbers agreed with the total, and never asked how many times they could agree before agreeing stopped being news.
The switch under the score
The letters matter less than which one breaks first. Here's the same five steps, mapped onto Yeva's queue.
FLIPS, five rows
FFind the person
Who signs the migration off?
Not the model team in the abstract. Whoever holds the pen on the migration and decides when to open the harder tab.
In this answer: Yeva Sorokin, Ledgerline Payments's fraud risk lead, who owns every migration sign off against the twelve hundred transaction golden set.
LLocate the habit
What did she stop doing once the score kept agreeing with itself?
Look for the check that quietly went from routine to optional, not her overall care. Being right three migrations running is what buys the habit its exit.
In this answer: She stopped opening the category tab and reading each number underneath it, once three migrations in a row had every category agreeing with the total.
IIdentify the flip
What two setting switch snaps, with no middle?
"The review got thinner" describes the outcome, not the action. Name the exact two states with nothing between them.
In this answer: Reads every category's own recall before signing off, or reads the one overall score and signs off. Once the category step stopped blocking, there was no smaller version of checking left to fall back on.
PPinpoint the old decision
Which call only made sense before three migrations in a row agreed?
Look for a narrow, defensible call from an early cleanup. "Add a category floor" after the fact doesn't count, that's a new dial.
In this answer: Model ops folded the blocking per category memo into the same dashboard as the total score and marked it optional, because three straight migrations already had it agreeing with the total.
SShow the replay
Same migration, category floor restored. Better ending?
Run the identical trigger through the fixed design and see where it stops. A count or a clock, not an adjective.
In this answer: Account takeover's 79 trips the floor on day two of the review. The migration ships four days late instead, and the 310,000 dollars in losses never happens.
A small move in the score. A hard snap in how she checked it.
"The review got thinner over time" is a diagnosis anyone can offer after the fact. The harder part is naming the exact habit that had to stop first, opening the category tab, and showing there was no smaller version of it left once it did.
And if you want to be sure it really works, try it somewhere else
ClauseBridge translates contracts for law firms, and keeps a golden set of professionally verified translations to certify any new translation model before firms can use it live. Same question, a law firm's contract desk instead of a payments processor, and a flip that isn't over trust this time.
F. Greta Lindholm, senior linguistic QA lead at ClauseBridge, who signs off any new contract translation model before firms can use it. L. For two years she personally read the raw flagged mistranslations in every contract category, NDAs, leases, employment contracts, M&A term sheets, IP licensing, before approving a migration. She handed that reading to a junior reviewer, Teo Halvorsen, once his write ups started matching her own conclusions, migration after migration. I. A different flip. She doesn't skip the review, she hands it down and never takes it back. Personally reads the raw per category examples herself, or reads only Teo's one blended summary and signs, nothing shared in between once the handoff stuck. P. The handoff had no shape to it. Teo wrote one paragraph covering every contract category at once, because nobody had ever asked him to write one memo per category, the way Greta used to keep it herself. After two years of his conclusions matching hers, one paragraph felt like plenty. S. Restore a memo per category, even under Teo's byline, so one category can flag on its own instead of hiding inside one paragraph's overall tone. Same migration: the lease renewal category's own memo would have shown date accuracy at 91 against the old model's 99, and Greta would have held it in the two day migration review instead of a client's paralegal catching a shifted termination date nineteen days after the contract went out.
Days between a migrated model's date error and someone catching it, ClauseBridge
One blended memo, every contract category folded into one paragraph
Old design
19 days
Per category memo restored, lease dates checked on their own
New design
1 day
Old design: caught when a client's own paralegal happened to compare a lease amendment against the English original, nineteen days after it went out. New design: the lease renewal memo flags its own date accuracy drop inside the two day migration review, before the model ever reaches a client's desk.
A second decision worth taking back
Letting a summary stand in for evidence is itself a decision, not a fact about trust. A rule that said the format can change who writes the memo, never how many categories it has to cover on its own, would have kept the per category signal alive underneath the handoff.
Swap the trigger and it still runs
Speed: if ClauseBridge migrated models every six weeks instead of about once a year, a blind spot like the lease category would surface in the very next migration, with far less time for one blended memo to feel safe.
Cost: if a full category by category review cost real money, an outside proofreader per category, someone would have shrunk it to one memo even sooner, and the same date accuracy gap would have opened faster, not slower.
The model got better: this is close to what actually happened at Ledgerline too. The new model wasn't worse. It was better at almost everything, and "better at almost everything" is exactly the result that makes a person stop checking the one thing it isn't.
Where people run it wrong
Blaming the new model for the account takeover losses, when the model never touched anything the golden set didn't already know about, it just never got asked before shipping.
Reaching for a bigger golden set as the fix, when the same twelve hundred transactions, read by category instead of by total, would have caught it for free.
Waiting for the chargeback numbers to reveal the gap, instead of watching each category's own recall on the day the migration ships.
How to use it live
Buy yourself a few seconds by naming the reframe before the fix: "The question isn't whether the total score improved, it did. It's whether a migration can trade a strength for a weakness without the total ever admitting it." Say that, and the rest of the answer is just the mechanism.
Flashcards (click a card to flip it)
Eight fixed slots, pulled straight from the answer above.
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over trust flip. She checks sometimes, then stops checking at all, and it fires because the news keeps being good. A model migration that keeps testing better is a perturbation too, not just a bad one.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Yeva Sorokin, fraud risk lead at Ledgerline Payments, seven years reading fraud alerts. She can name which of six fraud patterns a flagged transaction is shaped like before she finishes the first line.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She stopped opening the golden set's category tab and reading each category's own recall, once three migrations in a row had every category agreeing with the total score.
4 · THE FLIP, HERE
What's the two setting switch in this story?
Tap to flip
ANSWER
Reads every category's own recall before signing off, or reads the one overall score and signs off. No setting in between once the category step stopped blocking.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Model ops folded the blocking, per category migration memo into the same dashboard as the total score, and marked it optional, because three straight migrations already had it agreeing with the total.
6 · THE NUMBER
The overall score rose from ___% to ___%, while account takeover recall fell from ___ to ___.
Tap to flip
ANSWER
80% to 89%. Account takeover recall fell from 97 to 79, an 18 point drop hidden inside a 9 point rise.
7 · THE REPLAY
Same migration, new gate, what changes?
Tap to flip
ANSWER
Account takeover's 79 trips the category floor on day two of the review. The migration ships four days late instead, and the 310,000 dollars in real losses never happens.
8 · CROSS-PRODUCT
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
ClauseBridge, a legal contract translation platform, using the delegation flip: a QA lead hands migration review to a junior reviewer and never takes the per category read back.
Check yourself Score: 0 / 0
Multiple choice
1. What was the flip in Yeva's story, and what were its two settings?
A. She becomes less confident in the fraud model after the migration ships.
B. She reads the golden set broken out by category before signing off, or she reads the one overall score and signs off, with nothing in between.
C. The fraud model got worse at catching account takeover specifically.
D. She asks a colleague to double check the migration before she signs it.
Show hint
Look for something Yeva does with her own attention, not something that happened to the model.
Show answer
B. C describes the trigger's effect on one category, not Yeva's behavior, and the model overall got better, not worse. A describes a feeling, and the story never shows her losing confidence, only stopping a specific check. D describes a fix worth making, but it isn't what actually happened.
Fill in the blank
2. The decision this answer takes back is that the migration checklist's ______ step got folded into the ______ score's dashboard and marked ______ instead of blocking.
Show hint
It's a merged steps reversal: two separate checks became one, and the natural pause where a category could get caught disappeared.
Show answer
Category by category memo, total, optional. Nobody removed the category breakdown, they just stopped requiring anyone to read it before shipping, which had the same effect once the habit of glancing at it faded too.
True or false
3. True or false: the category floor Yeva would add is also needed for a routine quarterly retrain that keeps the same model architecture.
True
False
Show hint
Ask what actually changes about the model's failure shape in a retrain versus a migration.
Show answer
False. A retrain that keeps the same architecture doesn't change what the model is good or bad at, so the overall score alone still tells the truth. The category floor exists to catch a trade a real migration can make that a retrain can't.
Multiple choice
4. Why wasn't "glance at the category tab a bit more carefully" a real option once the habit had faded?
A. Because the golden set only stores one overall number, not a breakdown by category.
B. Because there were only two real states once the checklist stopped blocking: open the tab and read every category, or sign off on the total alone. Nothing forced a quick glance that would have caught a category down 18 points.
C. Because Yeva was not allowed to see the category breakdown before signing off.
D. Because reading the category tab more closely would not have caught the account takeover regression.
Show hint
This is the flip versus dial mistake. A flip has exactly two settings, not a sliding scale of carefulness.
Show answer
B. D is the trap answer. The regression was sitting right in the breakdown the whole time, easy to catch, which is exactly why the missing step mattered more than the model's own accuracy.
Short answer, apply it yourself
5. Think of a dashboard or scorecard you or a team you're on checks before approving something. What's the one overall number on it that could be hiding a bad slice underneath, and what would you check to catch that before it happened?
Show hint
Think of an average customer satisfaction score that could hide one unhappy segment, or a team's overall bug count that could hide one broken feature.
Show answer
Model answer: "Our support team tracks one average response time across every ticket type. It looks fine every week. Nobody ever splits it by ticket type, so a billing dispute could be taking three times as long as a password reset and the average would never say so. Breaking the number out by ticket type, and setting a floor for the slowest one, would catch that before a customer complains about it instead of after." Any real example counts, as long as it names a plausible hidden slice and a concrete way to check it.
Fill in the blank, do the math
6. The chart shows account takeover recall at 97 under the old model, falling to 79 under the new one, while the overall recall across all six categories rose from 80 to 89. How many points did account takeover recall fall, and how many points did the overall score rise?
Show hint
Subtract 79 from 97 for the fall. Subtract 80 from 89 for the rise.
Show answer
18 points down, 9 points up. The rise was exactly half the size of the drop it was hiding, and it's invisible if you only ever check the one overall number, because that number went up the whole time.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.