ConceptAdvancedEval-Driven Specification / Golden datasets and test set ownership / #25

Who should be able to veto a launch based on golden set results?

The direct answer
Give the model validation lead a real vote, not a memo. No launch ships until she signs off, and her sign-off has to check every named group in the golden set on its own, not just the blended number. If she says wait, the only way around her is a documented, named override from someone outside the team that owns the ship date, on the record, not a hallway conversation.
Do this, in order
  1. Give the validation lead's sign-off a real vote, tied to every group in the golden set, not the blended average.Why: this is the one change that turns a warning anyone can ignore into a warning that has to be answered before anyone ships.
  2. Require any override of that veto to be a named, written decision by someone outside the shipping team.Why: without this, "veto power" just moves the same person's judgment call one desk over and calls it a process.
  3. Rebuild the golden set whenever the real population changes, not only when a version ships.Why: the gate stopped measuring the truth the day the applicant pool changed and nobody touched the test.
  4. Track each group's numbers in production against the golden set's numbers, every month.Why: a widening gap between the two is the actual warning sign, and it's invisible if the only number anyone checks is the blended one.
  5. Give the people the score is used on a way to ask for a second look, before months pass, not after.Why: right now nobody outside the score can push back on it at all, because nobody tells them a model set the number first.
  6. Leave the deterministic checks alone, like whether the score renders as a number one through ten.Why: there's no judgment to get wrong in a count, so don't spend veto power on the part nobody can quietly get wrong.

How to answer this, stage by stage

Seven moves, ending on the one line worth remembering.

1
Ground it in one product before you name the acronym
Say it like this
"Say a company builds a score that tells a parole board how likely someone is to break the rules of release. One digit, one through ten, printed at the top of a case file. Before any new version ships, it has to clear a golden set, real past cases with an outcome a review panel already agreed on."
Why this works
Grounds "golden set" and "veto" in one concrete product before you touch a framework, so nothing you say next is a definition.
2
Say your structure out loud
Say it like this
"I'd use GUARD, because this is a risk question about who gets a say, not a rollout plan. Who's in the room, where the harm lands unevenly, who can't push back, what I'd actually build, and how I'd catch it if it slipped through anyway."
Why this works
Two seconds that prove you have a plan before the story starts, and it names the framework without reciting it.
3
Split who's in the room from who isn't
Say it like this
"Three groups sit around a launch like this. The team that owns the ship date, and they decide. The person who validated the model, and she can only recommend. And the people the score gets used on, who never even learn a score exists."
Why this works
This is GUARD's G step, and naming three groups instead of two is what sets up the whole rest of the answer.
4
Turn "it cleared the bar" into a real number
Say it like this
"The blended score was ninety percent, well above the bar. Split the same golden set by applicant type and it's seventy one percent for people up for their first hearing, ninety six percent for everyone else. One number was hiding two very different tests."
Why this works
A number the interviewer can check beats saying a launch "might" have been unsafe.
5
Say who couldn't stop it
Say it like this
"The validation lead flagged it two weeks before launch, in writing. It didn't matter, because her sign-off was a recommendation, not a vote. And the person whose file the score sits on top of never learns the score exists, so he can't ask anyone to look again either."
Why this works
GUARD's hardest, strongest step. This is what separates a real risk answer from a checklist about accuracy.
6
Name the actual build, not a committee
Say it like this
"I'd make her sign-off a gate, not a suggestion, and it has to clear every named slice of the golden set, not the blend. If someone wants to launch anyway, that's a named, written override from someone outside the team with the deadline. Not a hallway yes."
Why this works
A gate someone can point to beats a policy nobody enforces once the deadline gets close.
7
Say how you'd catch a bad call after the fact, and close
Say it like this
"Every month I'd check each group's real numbers against the golden set's numbers, and log every override with a name on it. So: a real vote for the person who validates the model, checked group by group, and a named signature on every time someone overrides her."
Why this works
Ends on the line the interviewer remembers, and it shows you're thinking past launch day, not just about launch day.

Let's learn

The Harrow Score is one number, one through ten, printed at the top of a parole file. Ten means high risk. One means low. A parole board looks at that number before it looks at almost anything else in the room.

Knowledge spark: what a golden set is A stack of real past cases where the right outcome is already agreed on, usually by a review panel. A team runs a new model against it before shipping, to check the model still gets the old cases right.

For the first three versions, the check was simple. Before any change shipped, a validation lead ran it against a golden set, nine hundred old cases with an outcome a review panel had already settled, and the score had to land within one point of the panel's own number. Version one cleared that bar. So did version two. So did version three, comfortably, at ninety two percent.

Fourteen months ago, the state widened who even qualifies for an early hearing. First-time offenders with low-level charges, people who never used to come up this early, started showing up in the schedule. The applicants changed. The golden set didn't. Then version four shipped, built for a new contract deadline, and it cleared the bar too. Ninety percent, five points above the eighty five percent line. Nobody thought twice.

Ninety percent was never a number about every applicant. It was a blend, and the blend hid the one group the new law had just added.

Split by the same golden set, first-time applicants matched the panel seventy one percent of the time. Returning applicants matched ninety six percent of the time. The gate had one bar, one blended number, and the number that mattered had been sitting inside it the whole time, unread.

At its worst, this doesn't fail loudly. It fails as a longer wait. A wrong high-risk score means a second hearing, added months, for a mistake nobody outside the model has any real way to catch before it happens.

The decision I would take back We built one validation gate against one blended pass rate. That was the right call when there was one applicant population and the blend described it fine. It stopped being true the day the state changed who applies, and the gate kept waving every version through anyway, because nobody had told it to look underneath the average.

What I would leave alone. Some checks don't need this. Whether the score renders as a whole number, whether the file attaches to the right case ID, that's counting, not judging. No group gets quietly worse off because of a file that failed to attach. Save the real scrutiny for the part where the model is actually guessing at someone's future.

The lesson. I used to think a validation gate was safe once it had a bar and a number to clear. It isn't. A gate only protects the group it was built to measure, and it will keep waving everyone through, right up until someone checks who's actually walking past it.

Now here is the same thing as a story

The short version is above. Read this one when you want to feel why the fix actually matters.

Every Tuesday morning, Thuy Lam pulls the newest model build before anyone else on the team has opened their email. She's run model validation at Harrow Analytics for three years, ever since the company's recidivism-risk score started scoring real cases for the state's Board of Paroles instead of test data. She can tell from the first fifty rows of a run whether it's going to be a quiet Tuesday or a long one.

For the first three versions, it was always quiet. She'd run the new build against the golden set, nine hundred old cases with an outcome a review panel had already agreed on, and split the results eight different ways out of habit, more than anyone actually asked her to. Every slice came back close enough to the last one to sign off by lunch.

She kept splitting the results long after anyone checked her math. Nobody asked her to. It just felt like the honest way to check a number that was going to sit at the top of somebody's file.

Then, eighteen months ago, came the near miss. Version three, two days from shipping, had a labeling error nobody had caught, a batch of cases where a violation flag had been copied one row off from where it belonged. Thuy found it by accident, cross-checking one case that read strange to her, the night before it would have gone out. Nobody had a name for whose job it was to catch that. She'd just happened to look twice.

She didn't stop thinking about it. Not because it shipped wrong, it didn't, she caught it. Because of how close it came to not being caught at all, and because catching it had never been anyone's actual job.

Fourteen months ago, the state changed who qualifies for an early hearing. First-time offenders, low-level charges, people who never used to come up this soon, started filling the Board's schedule. Harrow's applicant mix shifted with it. The golden set, built years earlier from whatever old case files the company could get its hands on, stayed exactly the same nine hundred cases.

Version four was due in six weeks, timed to a new contract with a second state. Thuy ran it the way she always did. Ninety percent, comfortably over the eighty five percent bar. She split it anyway.

First-time applicants: seventy one percent. Returning applicants: ninety six percent.

She wrote it up and sent it to Torsten Bergstrom, the product lead who owned the ship date, two weeks before launch.

Torsten read the same nine hundred cases as one number, ninety percent. Thuy read them as two different tests, and only one of them was passing.

Torsten's answer wasn't unreasonable, on its own. The overall number cleared the bar. The contract had a fixed date attached to a fixed dollar figure. He told her they'd tighten the first-time-applicant slice in the next update, and thanked her for catching it. Nobody on that call was wrong about their own job. Thuy's job was to flag it. Torsten's job was to hit the date. Neither job included a way for the flag to actually stop the ship.

Two people: a product lead with a hand on a SHIP lever, deciding alone, and a validation lead holding only a written memo, no lever at all
Same file, only one desk has a lever

Version four shipped on schedule.

Three weeks later, Andre Willett came up for his first parole hearing. No record before this case, a clean file for the fourteen months he'd served. The Harrow Score on his file read a seven. Under Board policy, anything six or above triggers a second review, and the schedule for those runs about four months behind. Andre wasn't denied. He was queued.

Nobody at Harrow read Andre's file next to the golden set the week his score came back. If anyone had, they'd have found eleven similar cases already sitting in it, first-time, low-level, clean disciplinary record inside, and the panel had scored every one of them a three.

We didn't lose Andre four months. We lost him to a number nobody had the standing to stop.

Thuy was in the room, three years earlier, when the validation gate was first designed. One gate. One blended pass rate. It was the obvious choice at the time, there was one applicant population, one number described it fine, and building anything more would have been solving a problem they didn't have yet.

Run version four again, but give Thuy's sign-off an actual vote, one that has to clear every named slice of the golden set, not just the blend, and that only a named person outside the shipping deadline can override, on the record. Same six-week timeline. The seventy one percent comes back two weeks before launch, same as before. This time the launch doesn't happen at ninety. It waits, or it ships with a stated exception somebody outside product signed their name to.

Andre's case might still have needed a second look. But it would have needed one because a person decided that on purpose, not because nobody had the power to ask.

GUARD, and the sign-off nobody could enforce

This is a risk question, so the framework is GUARD. A launch veto is a governance decision wearing an engineering-convenience hat, which is exactly why "the overall number cleared the bar" gets said in a launch review without anyone asking whose number it actually was.

G, groups. Three groups sit around this launch, and only one of them votes. Torsten's team owns the ship date and decides go or no-go. Thuy validates the model and can only recommend. Applicants like Andre never learn a score exists at all, let alone that it shipped over a written flag.
U, unequal. The harm doesn't land evenly across applicants. It lands on the group the golden set was never rebuilt to represent, split the real cases by applicant type and the gap shows up in the numbers, not just in one file.
Falsely scored high risk, by applicant type
Golden-set cases the review panel scored low risk, one to three, that version four scored high risk, six or above.
Returning applicants (matches the population the set was built on)
6%
First-time applicants (newly eligible under the state's own rule change)
34%
Almost six times the false-high-risk rate, and the blended score never showed it. It sat at ninety percent, because it was always measured against the same nine hundred cases that happened to skew toward one group.
A, ability to contest. Andre can't appeal a number he doesn't know exists. Nobody tells an applicant a model scored his file, let alone that the model had a documented gap for someone with exactly his case type. The queue he lands in reads like a rule, not a judgment call, so there's nothing in it that invites a second opinion.
A flow from case file to score to second review triggered, with the step where someone should re-check the number missing, straight to a hearing queued four months later
The step that should sit fourth in this chain, and doesn't
R, reduce. Turn the validation lead's sign-off into a real gate, checked against every named group in the golden set, not the blend. A launch that fails on any group needs a named, written override from someone outside the team that owns the ship date.
D, detect. Track each group's real-world numbers against the golden set's numbers every month, so a widening gap gets caught on a calendar, not by a case that happens to look wrong to someone. Log every override, with a name attached, so a pattern of overrides is itself a thing someone can be asked about.
Where this answer would fail If the fix here were an ethics review board, a bias-awareness training, or a values statement, none of it counts. "Her sign-off is a gate, checked by group, overridable only by name" is a build ticket. Somebody can ship it this sprint, and you can check whether they did.

And if you want to be sure it really works, try it somewhere else

A state benefits agency runs a model that flags applications for a fraud review before a caseworker approves them. Different program, same five letters, same trap.

G, groups. The fraud-model team that ships changes and decides when a new version goes live, and the caseworkers who inherit whatever gets flagged, and the applicants who never learn a model touched their file first.
U, unequal. The flag rate isn't even. Applicants with irregular income, mostly gig and seasonal workers, get flagged for review at a much higher rate than salaried applicants, and the overall flag rate looks stable because salaried applicants are most of the volume.
A, ability to contest. An applicant can appeal a denial, months later, but nobody tells them a fraud model flagged their file first, so there's nothing to contest about the flag itself, only the final decision it fed into.
R, reduce. A caseworker ombudsperson, outside the model team, gets a real vote on any version whose flag-rate gap by income type crosses a set line, not just the overall flag rate.
D, detect. Track flag rate by income type monthly, and the average time an application sits in review by the same split. A flat overall number can hide one group waiting three times as long.

Swap the trigger and it still runs

  • Speed: a new state contract adds a six-week deadline, and the pressure to treat "cleared the bar" as done gets stronger exactly when the golden set is most out of date.
  • Cost: rebuilding the golden set with fresh, correctly-labeled cases costs real money and staff time, so it keeps getting deferred to "next quarter" every quarter.
  • The model gets better: version five clears ninety six percent blended, and that number looks so good nobody proposes splitting it by group at all, because there's no obvious reason to doubt it.

Where people run it wrong

  • Calling it a "veto" when the validation lead's sign-off is still just a comment on a ticket the product lead can close without responding to.
  • Checking the overall pass rate and calling it a fairness review, when the whole problem lives inside the one slice the average smooths over.
  • Rebuilding the golden set only when a version ships, instead of the moment the real population it's supposed to represent actually changes.

If you're asked this cold

Ask who can stop a launch that clears the bar, by name, and what happens if they say wait. If the honest answer is "nobody, the number already cleared," that's the whole risk, right there, before you say another word.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about who can veto a launch, and why?
Tap to flip
ANSWER
GUARD, for risk, safety and fairness. The real question is who holds power over a launch decision and who can't push back on it, exactly what GUARD is built to find.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Thuy Lam, model validation lead at Harrow Analytics, three years running golden-set checks before every version ships to the state's Board of Paroles.
3 · THE HABIT
What did Thuy do that nobody asked her to, and why did it matter later?
Tap to flip
ANSWER
She split every validation run by applicant group instead of only checking the blended pass rate. It's the only reason the seventy one percent gap ever got written down at all.
4 · THE GAP
What's the two-number gap that proves the launch passed unevenly, not just imperfectly?
Tap to flip
ANSWER
Ninety percent blended match on the golden set, against seventy one percent for first-time applicants alone and ninety six percent for returning applicants. Same model, same week.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building one validation gate against one blended pass rate. It made sense when there was one applicant population to describe. It stopped being true the day the state changed who applies, and nobody rebuilt the gate.
6 · THE NUMBER
Fill in: first-time applicants were falsely scored high risk ______ percent of the time, against ______ percent for returning applicants.
Tap to flip
ANSWER
Thirty four percent, against six percent. The blended score never showed this split, it sat at ninety percent the whole time.
7 · THE REPLAY
Same six weeks, new design. What changes, and when does it change?
Tap to flip
ANSWER
Thuy's sign-off becomes a real gate, checked by group. The seventy one percent surfaces two weeks before launch, same as before, but this time the launch waits, or ships with a named person's override on record.
8 · TRANSFER
Section four runs GUARD again on a different program. Which one, and what does the reduce step become?
Tap to flip
ANSWER
A state benefits fraud-flagging tool. Reduce: a caseworker ombudsperson outside the model team gets a real vote on any version whose flag-rate gap by income type crosses a set line.

Check yourself Score: 0 / 0

Multiple choice
1. A launch review says: "the model cleared eighty five percent on the golden set, so it's fine to ship." What's the real problem with stopping there?
  • A. Eighty five percent isn't a high enough bar to set in the first place.
  • B. The golden set needs more than nine hundred cases to mean anything.
  • C. An overall number can clear the bar while one named group underneath it is failing badly, and nobody who could catch that has a vote.
  • D. The review panel that scored the golden set should have used two panels instead of one.
Show hint
The problem in this answer was never the size of the bar.
Show answer
C. A, B, and D worry about the bar's height or the set's size. The real risk is that a blended pass rate can hide a group failing underneath it, and the person who found that had no power to stop the ship.
True or false
2. True or false: Torsten was wrong to launch version four, because he ignored Thuy's warning.
  • True
  • False
Show hint
Ask whether he ignored her, or whether the process gave her warning nowhere to land.
Show answer
False. Torsten didn't ignore her, he read the memo and made a reasonable call given his job: the blended number cleared the bar, and the contract had a fixed date. The design flaw was that there was no process where her warning could actually stop a launch, not that he personally disregarded it.
Fill in the blank
3. The golden set showed a ______ percent match rate for first-time applicants, against a ______ percent match rate for returning applicants, blended into one ninety percent overall score.
Show hint
It's the pair of numbers behind stage 4 of the walkthrough.
Show answer
Seventy one percent, and ninety six percent. Same model, same golden set, split by applicant type instead of read as one blended number.
Short answer
4. What old decision does this answer take back, and why did it make sense when Harrow first made it?
Show hint
Look at how the validation gate was originally built, not at who ignored what later.
Show answer
Model answer: "Building one validation gate around one blended pass rate. It made sense at the time because there was one applicant population, so a blended number described it fine. Splitting it into groups back then would have solved a problem the product didn't have yet."
Short answer, apply it yourself
5. Pick a scoring tool you know of, in any field. If the team building it hit its target number tomorrow, who actually has the power to say "not yet" anyway?
Show hint
Look for whoever can only write a memo, not cast a vote.
Show answer
Model answer: "A credit-scoring model at a bank clears its accuracy target every quarter. The model risk officer reviews it, but the business line that owns the revenue target makes the final call. Nobody outside that business line has to sign off before it ships, so the review is a memo, not a vote, same shape as Thuy's."
Short answer
6. If the false-high-risk rate for first-time applicants had been half as bad, seventeen percent instead of thirty four, would giving Thuy a real veto still be the right fix? Why or why not?
Show hint
Ask whether the fix depends on the size of the gap, or on the fact that nobody could act on any gap at all.
Show answer
Model answer: "Yes. Seventeen percent is still nearly three times the returning-applicant rate, and it's still a group carrying more than its share of a real cost, extra months waiting. The actual problem was never the exact size of the gap, it was that nobody who could see it had the power to stop a launch over it. That's true whether the gap is thirty four percent or seventeen."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more