Who should be able to veto a launch based on golden set results?
- Give the validation lead's sign-off a real vote, tied to every group in the golden set, not the blended average.Why: this is the one change that turns a warning anyone can ignore into a warning that has to be answered before anyone ships.
- Require any override of that veto to be a named, written decision by someone outside the shipping team.Why: without this, "veto power" just moves the same person's judgment call one desk over and calls it a process.
- Rebuild the golden set whenever the real population changes, not only when a version ships.Why: the gate stopped measuring the truth the day the applicant pool changed and nobody touched the test.
- Track each group's numbers in production against the golden set's numbers, every month.Why: a widening gap between the two is the actual warning sign, and it's invisible if the only number anyone checks is the blended one.
- Give the people the score is used on a way to ask for a second look, before months pass, not after.Why: right now nobody outside the score can push back on it at all, because nobody tells them a model set the number first.
- Leave the deterministic checks alone, like whether the score renders as a number one through ten.Why: there's no judgment to get wrong in a count, so don't spend veto power on the part nobody can quietly get wrong.
How to answer this, stage by stage
Seven moves, ending on the one line worth remembering.
Let's learn
The Harrow Score is one number, one through ten, printed at the top of a parole file. Ten means high risk. One means low. A parole board looks at that number before it looks at almost anything else in the room.
For the first three versions, the check was simple. Before any change shipped, a validation lead ran it against a golden set, nine hundred old cases with an outcome a review panel had already settled, and the score had to land within one point of the panel's own number. Version one cleared that bar. So did version two. So did version three, comfortably, at ninety two percent.
Fourteen months ago, the state widened who even qualifies for an early hearing. First-time offenders with low-level charges, people who never used to come up this early, started showing up in the schedule. The applicants changed. The golden set didn't. Then version four shipped, built for a new contract deadline, and it cleared the bar too. Ninety percent, five points above the eighty five percent line. Nobody thought twice.
Split by the same golden set, first-time applicants matched the panel seventy one percent of the time. Returning applicants matched ninety six percent of the time. The gate had one bar, one blended number, and the number that mattered had been sitting inside it the whole time, unread.
At its worst, this doesn't fail loudly. It fails as a longer wait. A wrong high-risk score means a second hearing, added months, for a mistake nobody outside the model has any real way to catch before it happens.
What I would leave alone. Some checks don't need this. Whether the score renders as a whole number, whether the file attaches to the right case ID, that's counting, not judging. No group gets quietly worse off because of a file that failed to attach. Save the real scrutiny for the part where the model is actually guessing at someone's future.
The lesson. I used to think a validation gate was safe once it had a bar and a number to clear. It isn't. A gate only protects the group it was built to measure, and it will keep waving everyone through, right up until someone checks who's actually walking past it.
Now here is the same thing as a story
The short version is above. Read this one when you want to feel why the fix actually matters.
Every Tuesday morning, Thuy Lam pulls the newest model build before anyone else on the team has opened their email. She's run model validation at Harrow Analytics for three years, ever since the company's recidivism-risk score started scoring real cases for the state's Board of Paroles instead of test data. She can tell from the first fifty rows of a run whether it's going to be a quiet Tuesday or a long one.
For the first three versions, it was always quiet. She'd run the new build against the golden set, nine hundred old cases with an outcome a review panel had already agreed on, and split the results eight different ways out of habit, more than anyone actually asked her to. Every slice came back close enough to the last one to sign off by lunch.
She kept splitting the results long after anyone checked her math. Nobody asked her to. It just felt like the honest way to check a number that was going to sit at the top of somebody's file.
Then, eighteen months ago, came the near miss. Version three, two days from shipping, had a labeling error nobody had caught, a batch of cases where a violation flag had been copied one row off from where it belonged. Thuy found it by accident, cross-checking one case that read strange to her, the night before it would have gone out. Nobody had a name for whose job it was to catch that. She'd just happened to look twice.
She didn't stop thinking about it. Not because it shipped wrong, it didn't, she caught it. Because of how close it came to not being caught at all, and because catching it had never been anyone's actual job.
Fourteen months ago, the state changed who qualifies for an early hearing. First-time offenders, low-level charges, people who never used to come up this soon, started filling the Board's schedule. Harrow's applicant mix shifted with it. The golden set, built years earlier from whatever old case files the company could get its hands on, stayed exactly the same nine hundred cases.
Version four was due in six weeks, timed to a new contract with a second state. Thuy ran it the way she always did. Ninety percent, comfortably over the eighty five percent bar. She split it anyway.
First-time applicants: seventy one percent. Returning applicants: ninety six percent.
She wrote it up and sent it to Torsten Bergstrom, the product lead who owned the ship date, two weeks before launch.
Torsten's answer wasn't unreasonable, on its own. The overall number cleared the bar. The contract had a fixed date attached to a fixed dollar figure. He told her they'd tighten the first-time-applicant slice in the next update, and thanked her for catching it. Nobody on that call was wrong about their own job. Thuy's job was to flag it. Torsten's job was to hit the date. Neither job included a way for the flag to actually stop the ship.
Version four shipped on schedule.
Three weeks later, Andre Willett came up for his first parole hearing. No record before this case, a clean file for the fourteen months he'd served. The Harrow Score on his file read a seven. Under Board policy, anything six or above triggers a second review, and the schedule for those runs about four months behind. Andre wasn't denied. He was queued.
Nobody at Harrow read Andre's file next to the golden set the week his score came back. If anyone had, they'd have found eleven similar cases already sitting in it, first-time, low-level, clean disciplinary record inside, and the panel had scored every one of them a three.
Thuy was in the room, three years earlier, when the validation gate was first designed. One gate. One blended pass rate. It was the obvious choice at the time, there was one applicant population, one number described it fine, and building anything more would have been solving a problem they didn't have yet.
Run version four again, but give Thuy's sign-off an actual vote, one that has to clear every named slice of the golden set, not just the blend, and that only a named person outside the shipping deadline can override, on the record. Same six-week timeline. The seventy one percent comes back two weeks before launch, same as before. This time the launch doesn't happen at ninety. It waits, or it ships with a stated exception somebody outside product signed their name to.
Andre's case might still have needed a second look. But it would have needed one because a person decided that on purpose, not because nobody had the power to ask.
GUARD, and the sign-off nobody could enforce
This is a risk question, so the framework is GUARD. A launch veto is a governance decision wearing an engineering-convenience hat, which is exactly why "the overall number cleared the bar" gets said in a launch review without anyone asking whose number it actually was.
And if you want to be sure it really works, try it somewhere else
A state benefits agency runs a model that flags applications for a fraud review before a caseworker approves them. Different program, same five letters, same trap.
G, groups. The fraud-model team that ships changes and decides when a new version goes live, and the caseworkers who inherit whatever gets flagged, and the applicants who never learn a model touched their file first.
U, unequal. The flag rate isn't even. Applicants with irregular income, mostly gig and seasonal workers, get flagged for review at a much higher rate than salaried applicants, and the overall flag rate looks stable because salaried applicants are most of the volume.
A, ability to contest. An applicant can appeal a denial, months later, but nobody tells them a fraud model flagged their file first, so there's nothing to contest about the flag itself, only the final decision it fed into.
R, reduce. A caseworker ombudsperson, outside the model team, gets a real vote on any version whose flag-rate gap by income type crosses a set line, not just the overall flag rate.
D, detect. Track flag rate by income type monthly, and the average time an application sits in review by the same split. A flat overall number can hide one group waiting three times as long.
Swap the trigger and it still runs
- Speed: a new state contract adds a six-week deadline, and the pressure to treat "cleared the bar" as done gets stronger exactly when the golden set is most out of date.
- Cost: rebuilding the golden set with fresh, correctly-labeled cases costs real money and staff time, so it keeps getting deferred to "next quarter" every quarter.
- The model gets better: version five clears ninety six percent blended, and that number looks so good nobody proposes splitting it by group at all, because there's no obvious reason to doubt it.
Where people run it wrong
- Calling it a "veto" when the validation lead's sign-off is still just a comment on a ticket the product lead can close without responding to.
- Checking the overall pass rate and calling it a fairness review, when the whole problem lives inside the one slice the average smooths over.
- Rebuilding the golden set only when a version ships, instead of the moment the real population it's supposed to represent actually changes.
If you're asked this cold
Ask who can stop a launch that clears the bar, by name, and what happens if they say wait. If the honest answer is "nobody, the number already cleared," that's the whole risk, right there, before you say another word.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Golden datasets and test set ownership
- #1 What is a golden dataset and why does the PM usually own it?
- #2 How do you construct a first golden set with no production traffic?
- #3 Describe the composition of a golden set: what proportion should be edge cases?
- #4 How do you keep a golden set representative as your user base changes?
- #5 Explain the risk of a golden set that engineering can see during development.
- #6 What is a holdout set and when would you use one for an AI product?