ConceptIntermediateEval-Driven Specification / Golden datasets and test set ownership / #10

What governance do you need around who can edit the golden set?

The direct answer
Put every edit to the golden set behind a log that shows exactly what changed, when, and who changed it, before you build anything else. Then require a second person, never the one whose own work needs that case to pass, to approve the change before it merges. Build the log first: a review policy with nothing to review against is just trust wearing a badge.
The ranking, by what breaks first if skipped
  1. Put every edit to the golden set behind a log: who changed it, when, and what it said before.Why: without this, a bent case doesn't just go unreviewed, it becomes unfindable. There's nothing left to compare it against.
  2. Require a second person to approve the change before it merges, looking at the real before/after.Why: a review policy with no diff to look at is a rubber stamp with extra steps.
  3. Never let the person whose own model or prompt change needs a case to pass also approve the edit to that case.Why: that exact conflict of interest is what let a fifteen-minute deadline fix become a permanent rewrite of the correct answer.
  4. Restrict who can merge into the golden set to a small, named group.Why: a review policy nobody is technically forced through is a suggestion, not a gate.
  5. Spot-check a sample of past edits every quarter, even when nothing looks wrong.Why: cheap, and it's the only control that catches a bent case before a colleague stumbles onto it by accident.

How to answer this, stage by stage

Seven moves. The trap in this question is answering it like a policy memo, naming controls in whatever order a compliance checklist would. An interviewer wants to see you rank them by what actually goes wrong first, and defend the order.

1
Ground it in one real product
Say it like this
"Let me make this real. Say it's Regent, a SaaS platform executive assistants use to run their bosses' calendars, and inside it there's an AI feature called Pace that reads calendars and books meetings on its own. Pace gets tested against a golden set, 240 real scheduling situations with the correct move written down for each. I'll answer against that."
Why this works
A generic answer names controls in the abstract. A grounded one shows you know exactly what each control is standing between and a real mistake.
2
Reframe from "list the rules" to "rank them"
Say it like this
"I can name a handful of controls, but the real question is which one you can't afford to build last. Most of these you can bolt on next sprint. One of them, if it's missing, means you don't just miss a bad edit, you lose the ability to ever prove one happened."
Why this works
Tells the interviewer you're about to use a method, not recite a list you rehearsed the night before.
3
Name the outcome every control is protecting
Say it like this
"Every one of these controls exists to protect one thing: that the golden set stays a real ruler. Nobody, under deadline pressure or otherwise, gets to quietly bend the answer key so their own change looks like it passed."
Why this works
Without stating what's being protected, any ranking that follows is just a preference dressed up as a method.
4
Find the gap that can't be undone
Say it like this
"If there's no log, a rewritten case doesn't get caught late, it becomes invisible. Nobody can point to what it used to say. That's the one gap you can't repair after the fact, because the thing you'd need to repair it, a record of the original, is exactly what's missing."
Why this works
This is the reversibility test, the sharpest move in the ranking, and it's specific instead of "logging is best practice" hand-waving.
5
Show the dependency underneath the review policy
Say it like this
"But 'require a second reviewer' only means something if the reviewer can see an actual before-and-after. Without a log, review is somebody eyeballing a spreadsheet and taking the editor's word for it. So the log isn't just important on its own, it's the thing the review policy is standing on."
Why this works
Shows the controls aren't a flat list, they're built on each other, which is what separates a ranking from an opinion.
6
Prove it with the failure it prevents
Say it like this
"Here's what happens without that log. At Regent, someone under deadline pressure changed golden case 118, the one that says never auto-book over a VIP's existing meeting, to allow it when the overlap is under 15 minutes. No review, no note, just an overwrite. Six weeks later Pace auto-booked a board call straight over a client's one-on-one, because the ruler itself had quietly moved."
Why this works
A concrete, compressed failure carries more weight than any amount of asserting that governance matters.
7
Close with the full ranking, defended
Say it like this
"So, in order: an edit log first, since nothing else means anything without it. A required review with a real diff second. No self-approval for the person whose own work needs the case to pass, third. Then restricted merge access and a quarterly spot-check, both cheap to add once the first three exist."
Why this works
Ends on the literal list the question asked for, with a defended order behind it instead of a memorized sequence.

Let's learn

Every week, an executive assistant used to spend close to five hours untangling a boss's calendar by hand: emailing back and forth across time zones, working out who outranks who when two meetings collide.

Regent is the SaaS platform that EAs run those calendars through, and Pace is the feature inside it that does the untangling on its own. Pace reads every calendar it's given access to, reads the email thread asking for a meeting, and either books it or flags it for a human. To know whether a new version of Pace is any good before it touches a real calendar, Regent runs it against a golden set: 240 real scheduling situations, each with the one correct move already written down.

Knowledge spark: what a golden set is doing here It's the answer key. Before a new version of Pace ships, someone runs it against all 240 situations and checks its answers against the ones already written down. If the new version's answers match, it ships. The golden set only works if nobody involved in writing the new version can also quietly rewrite the answer key.

Before Pace, that manual untangling ate 4.8 hours a week per EA. After Pace shipped, most of that disappeared. Pace now auto-books about 82 percent of meeting requests without a human touching them, and the average EA is down to roughly 35 minutes a week on scheduling, mostly the requests Pace flags rather than books.

EA hours per week spent manually resolving scheduling conflicts, before and after Pace
4.8 hrs 0.6 hrs Before Pace After Pace
This is the number Pace is worth building for. It's also exactly the number that makes the golden set worth protecting, because it's the thing a bent test case can quietly put back at risk.

One of the 240 cases is numbered 118. It reads: a VIP already has a one-on-one on the calendar; a new request from someone senior overlaps it; Pace should never auto-book over that meeting, it should escalate to a human, no matter how small the overlap.

Eighteen months in, an engineer testing a new version of Pace ahead of a big pilot demo kept hitting case 118. The new version wanted to auto-book anyway. Rather than fix the model in time, they opened the golden set's shared editing tool and changed case 118's correct answer: auto-booking is fine now, as long as the overlap is under 15 minutes. No review. No note. The old answer was simply gone, overwritten in place, nothing left to show it had ever said anything else.

The eval turned green. It didn't mean the model had gotten better. It meant the ruler had moved.

Here is the part that matters. Auto-book rate, the number everyone actually watched, didn't drop after that edit. It went up, from 82 to 84 percent, because the loosened rule let more requests through without a human ever seeing them. The number that was supposed to catch a regression looked healthier than ever, for exactly the reason it should have looked worse.

Hand-sketch dependency chain: four boxes, Edit log, Review required, No self-approval, and Periodic audit, connected left to right by arrows, with Edit log circled in amber as the one the rest depend on.
Review required can't mean anything until the first box exists

Six weeks after case 118 was quietly rewritten, Pace auto-booked a board call directly over an existing one-on-one between Regent's biggest customer, Calloway & Vance, and their account lead. The overlap was 13 minutes, comfortably inside the new, secret exception. Both meetings showed up on the same room at the same time. Nobody had escalated it, because the rule that would have caught it no longer said what everyone thought it said.

Pace's auto-book rate over the six weeks after case 118 was rewritten
case 118 edited the overlap 84% 82 week 0 → week 6
This is the line everyone was actually watching. It climbed the whole way. A dashboard was never going to catch this, because the thing that broke wasn't the model's behavior, it was the definition of correct.
The choice I would take back When the golden set was first built, it lived in a shared spreadsheet-style tool anyone on the Pace team could edit directly, because at the time it was one person maintaining it and review felt like ceremony for a document nobody but her touched. Nobody revisited that once the team grew and started shipping model changes every week or two. I'd take that back: put the golden set behind the same kind of gated, logged, reviewed change that the code reading it already goes through, from the day it's created, not the day it first burns someone.

What I would leave alone. Anyone, including EAs on the customer side, should stay free to suggest a new scenario for the golden set at any time; a customer hitting a weird edge case is exactly how good cases get found. The governance is about who can merge a change to an existing answer, not about who can propose a new one.

The lesson. A missing governance control doesn't show up as an error. It shows up as a number that keeps looking fine, because the thing that was supposed to catch the problem is the thing that got quietly changed. You don't find that out from a dashboard. You find it out from someone standing in a doorway that already has two meetings booked in it.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why the log outranks everything else, not just take it on faith.

Every Monday morning, before anyone else at Regent had logged on, Anselm Duckworth pulled up the golden set. He'd been the one who built it, back when Pace was a single spreadsheet tab and "the golden set" meant a folder of transcripts he'd graded himself over a weekend. Three years later it was 240 cases and the thing every model change had to clear before it touched a real calendar.

For most of those three years, the Monday check was Anselm's own ritual, not a rule. He'd scan the diff of whatever had changed since Friday, ask whoever'd touched it why, and move on. It rarely took him more than twenty minutes. The team was small. He knew everyone's handwriting, so to speak.

As Regent grew, model changes stopped shipping every few months and started shipping every week or two. Anselm's Monday scan got shorter. Then it got skipped some weeks, when a launch was close. He told himself the team had gotten careful enough that the check was mostly a formality by then.

It wasn't a formality. It just hadn't been tested.

Six weeks after case 118 quietly changed, Regent's biggest account, Calloway & Vance, called in furious: their CEO and their account lead had both shown up to the same conference room at the same time, because Pace had auto-booked a board call straight over an existing one-on-one. Regent's on-call engineer pulled up the incident, saw Pace had done exactly what case 118 told it to do, and closed it as "model working as specified, escalate the underlying rule." Everyone assumed the rule had always allowed small overlaps. Nobody thought to ask when.

Hand-sketch comparison: on the left, a green box labeled Missing periodic audit, captioned add it any quarter, nothing lost. On the right, a red box labeled Missing edit log, captioned case 118 already got rewritten.
One of these you can fix next quarter. The other one already cost Regent its biggest account's trust.

It took a colleague's remark to unravel it. A product designer, pulling up case 118 for an unrelated design review two days later, said it almost as an aside: "Wait, doesn't this one say never, no matter the overlap? I remember arguing about that exact line with you at launch."

We didn't lose a booking. We lost the one thing that told us Pace was still correct.

Anselm went looking for the original wording of case 118 and found nothing. Not a version, not a comment, not a timestamp. The tool the golden set lived in simply held whatever the current cell said. There was no way to prove what case 118 used to say, who had changed it, or when, only a designer's memory of an argument from launch day against a database that disagreed with her.

Nobody on the team had a number in their head about this. They had a feeling about whether the golden set could still be trusted, and it only had two settings: it's the ruler, or it isn't. One unrecoverable edit flipped it. Getting case 118 fixed the same afternoon didn't flip it back. For the next three weeks, every model change went out with two engineers independently re-checking it by hand against the original design docs, because nobody trusted the golden set enough to run it alone.

A year earlier, when the golden set was still small enough for Anselm to hold in his head, someone had asked whether it needed a real edit log and required review, or whether his Monday scan was governance enough. At the time, with one person maintaining it and every change small, that wasn't a bad call. It just never got revisited once the team grew past the size where one person's memory could stand in for a record.

I'd take that back. With a real log, case 118's rewrite shows up the moment it happens: who changed it, what it said before, and no merge until someone who isn't the engineer under deadline signs off on it. The fifteen-minute deadline fix never reaches a real calendar. Calloway & Vance never books two meetings into the same room.

What I'd tell myself, back in that first review: skipping the log felt like trusting the team. It was actually removing the only thing that could prove the trust was still earned.

Putting the controls in order, and defending it

GUARD would fit if the question were about who gets hurt by a biased model. Nothing here is about bias, it's about ranking candidate controls before one gets skipped, which is ORDER's job.

O, outcome. Every control here is competing to protect one thing: that the golden set stays a real ruler, one nobody, under any pressure, can quietly bend to make their own change look like it passed.
R, reversibility. A missing quarterly audit is fixable any time, you just start sampling. A missing restriction on who can merge is fixable by tightening access next sprint. A missing edit log is different: once time passes with no record, you can't tell what changed, who changed it, or what it used to say. The first sign of the gap is a customer standing in a doubly-booked room.
D, dependency. A "review required" policy can't work on its own. A reviewer needs something to review: the actual before-and-after. That only exists if the edit log exists first. Skip the log, and a review policy is a person eyeballing a spreadsheet and taking the editor's word for it.
E, evidence. Cheap to check: pull the last handful of edits made to the golden set and see if each one has an author, a stated reason, and a visible diff. At Regent, case 118's edit had none of the three, and nobody could say how many others didn't either.
R, rank. Edit log first, since nothing downstream means anything without it. Required review with a real diff second, since that's what turns the log into a gate instead of a record nobody reads. No self-approval third, since that's the specific conflict of interest that broke case 118. Restricted merge access and a quarterly spot-check round it out, both cheap once the first three exist.
The check that keeps this ranking honest Swap the outcome and the order should move. If a wrongly-booked meeting only ever cost someone five minutes and an apology, the edit log could sit lower on this list. It doesn't rank first because logging is rigorous. It ranks first because a real client's leadership walked into the same room twice before anyone could prove why.

Rank it again, somewhere a bent answer key pays out in cash

A grain co-op, Redfield Grain Cooperative, runs an AI model that grades photos of delivered grain for moisture and quality, which sets the price paid to each farmer that day. The model is tested against a golden set of graded sample photos before any update ships.

O. Every control here protects one thing: that a grade the model gives a farmer's delivery is the grade a trained human grader would give, not a grade shaped by whoever last touched the answer key.
R. A missing second reviewer on grading updates is recoverable, add the requirement next season. A missing edit log is not: if a co-op employee quietly loosens the golden grade for a sample that happens to match their own family's usual delivery, and nobody can prove the sample's original grade, that unfair payout can never be traced or clawed back.
D. A conflict-of-interest rule, no employee approving a grade change tied to their own household's deliveries, only works once there's a log showing whose delivery each sample photo belongs to and who touched its grade. Without that link, the rule has nothing to check itself against.
E. Cheap to check: pull the golden set's grading history for the last harvest and see whether any edited samples trace back to an employee's own family account. Redfield found three, all edited by the same person, none reviewed.
R. Same order: edit log first, so every grade change is traceable to a sample and a person. Required review second. The conflict-of-interest rule third, since it can only be enforced once the first two exist. Restricted grading access and a seasonal audit close it out.
Redfield's golden grading samples: confirmed original vs. re-graded from scratch after the audit
156 confirmed original 21 re-graded 3 retired 180 golden grading samples, one harvest
Those 3 samples are gone for good, not because they were wrong, but because nobody could ever prove what they used to say. That's what a missing log actually costs: not just time, but samples you can never trust again.

Swap the trigger and it still runs

  • Regent moves the golden set to a proper version-controlled repo, but still allows direct pushes. The order doesn't move. A log with no required review is a record of the damage, not a way to stop it.
  • A new model provider makes Pace faster to iterate on, so changes ship daily instead of weekly. Same order, and the edit log matters more, not less, since faster shipping means more chances to quietly bend a case under pressure.
  • The team gets smaller instead of bigger, back to two people. Doesn't reorder anything. A two-person team under deadline pressure is exactly the size where "we trust each other" quietly replaces "we can prove it," which is the whole failure.

Where people run it wrong

  • Treating "we haven't had an incident" as proof the golden set is safe, when it only means nobody's looked hard enough yet.
  • Writing the review policy before the edit log exists, so "review required" has nothing real to review.
  • Restricting who can write to the golden set while still letting that same small group approve its own changes, which solves access and leaves the actual conflict of interest untouched.

If you're asked this cold

Say the outcome out loud before naming a single control. "Every one of these exists to protect one thing: that the answer key stays honest even when someone's under deadline pressure to make it agree with them." Then rank from there. Naming the outcome first is what turns a list into an argument.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits "what governance do you need around who can edit the golden set," and why not GUARD?
Tap to flip
ANSWER
ORDER, for ranking candidate governance controls by what's hardest to undo if it's missing. GUARD is for who gets hurt by a biased or unfair model. This question is about ranking controls before one gets skipped, not about naming who's harmed.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Anselm Duckworth, the product manager who built and has owned the golden set for Regent's scheduling assistant, Pace, since it was one spreadsheet tab.
3 · THE HABIT
What habit let a rewritten golden case go unnoticed for six weeks?
Tap to flip
ANSWER
Anselm used to personally scan every change to the golden set each Monday. As the team grew and shipped more often, that scan got shorter, then got skipped in busy weeks, and nothing formal replaced it.
4 · THE DEPENDENCY
Which governance control has to exist before a "review required" policy actually means anything?
Tap to flip
ANSWER
The edit log. A reviewer needs a real before-and-after to check. Without a log recording what a case used to say, review is just someone taking the editor's word for it.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Letting the golden set live in a shared tool with direct, unlogged edits, back when one person maintained it alone and review felt like ceremony. Nobody revisited it once the team grew and started shipping model changes every week or two.
6 · THE NUMBER
After case 118 was quietly rewritten, Pace's auto-book rate went from 82 percent to ___ percent over six weeks, right up to the week of the Calloway & Vance overlap.
Tap to flip
ANSWER
84 percent. The number everyone was watching climbed the whole time, because the loosened rule let more requests through unescalated. It was never going to be the signal that caught this.
7 · THE REPLAY
Same deadline pressure, golden set with a real edit log and required review this time. What changes?
Tap to flip
ANSWER
Case 118's rewrite shows up the moment it happens, with the old answer preserved. It can't merge without a second person's sign-off, and that person isn't the engineer under deadline. The fix never reaches a real calendar, and Calloway & Vance never double-books a room.
8 · THE TRANSFER
Section 4 runs ORDER again on a different product, in a different industry. Which one, and what plays the role of the edit log there?
Tap to flip
ANSWER
Redfield Grain Cooperative's AI grain-grading model. The equivalent is a log tying every grading sample to a specific delivery and a specific editor, so a conflict-of-interest rule (no grading your own family's delivery) has something real to check itself against.

Check yourself Score: 0 / 0

True or false
1. True or false: since Pace's auto-book rate actually went up after case 118 was rewritten, 82 to 84 percent, the golden set didn't really need an edit log to catch this problem. Say why.
  • True
  • False
Show hint
Ask what a loosened auto-book rule does to the one number everyone was actually watching.
Show answer
False. The rate going up is exactly why it couldn't catch the problem. A dashboard metric can't flag a change to the definition of correct, only a log of who changed the golden set and when can do that.
Multiple choice
2. Which two controls does this answer rank above every other governance control, and in what order?
  • A. Restricted merge access, then a quarterly audit
  • B. An edit log, then a required review with a real diff
  • C. A conflict-of-interest rule, then restricted merge access
  • D. A quarterly audit, then an edit log
Show hint
One of these has to exist before the other one can mean anything, per the dependency step.
Show answer
B. The log comes first because nothing downstream works without a record of what changed. Review comes second because it's the log that turns "required review" into an actual check instead of a formality.
Fill in the blank
3. Before Pace, an EA spent about ______ hours a week manually resolving scheduling conflicts. After Pace shipped, that dropped to about ______ hours a week.
Show hint
It's the number that makes the whole product, and by extension the golden set protecting it, worth the trouble.
Show answer
4.8 and 0.6. That's the value the golden set exists to protect. A bent golden case doesn't threaten a metric, it threatens this number going back up without anyone noticing why.
Multiple choice
4. What does the dependency step (D) in this answer's ORDER argue?
  • A. A required review policy can't mean anything until an edit log exists for the reviewer to check against
  • B. Restricted merge access and required review are the same control
  • C. Every engineer should be allowed to edit the golden set freely to keep things fast
  • D. Quarterly audits should replace review entirely, since they're cheaper
Show hint
Ask what a reviewer actually needs in front of them to review anything.
Show answer
A. A reviewer needs a real before-and-after to check. Without a log recording the original answer, "review required" is a person trusting the editor's word.
Short answer, apply it yourself
5. Pick an AI feature you use or are building that gets tested against some kind of answer key or test set. Which control, if missing, would let someone quietly bend that answer key without anyone being able to prove it happened?
Show hint
Look for the control that decides whether a bent answer becomes provable, or just becomes someone's word against a database.
Show answer
Model answer: "A résumé-screening tool's test set of 'correctly ranked' candidate profiles. The control hardest to add back after the fact is a log of who edited a profile's correct ranking and when, so a hiring manager under pressure can't quietly relabel a borderline case to make a new model version look fair."
Short answer, the number question
6. If Regent's golden set had 20 cases instead of 240, would skipping the edit log still be the top-ranked risk? Say what changes and what doesn't.
Show hint
Reversibility is about whether a bent answer can be proven and undone, not about how many other cases exist alongside it.
Show answer
Model answer: "Yes, it would still rank first. A smaller golden set actually makes each case matter more, since one bent case is a bigger share of the whole answer key. What doesn't change: the log is what makes a bent case provable and reversible. Fewer cases don't buy that back, they just mean fewer places to hide the edit."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more