The direct answer
Ship one page next to the golden set itself, not a wiki, not a slide deck. State the one decision it gates, what it covers by language and category and date, who approved each gold answer, what it plainly does not cover yet, and the one rule for adding to it without breaking comparisons across model versions. Leave out a full log of every edit ever made, that belongs to version control, not to a page a stranger reads in five minutes.
Do this, in order
Ship one page next to the data file naming what it gates, what it covers, and what it does not.Why: this is the whole rule, everything below just protects it from rotting or being ignored.
State coverage honestly, by language, category, and the date it was built, not just a total count.Why: a total count hides exactly the kind of gap that let a fast growing slice of a catalog go untested.
Name who approved each gold answer and how, not just that someone signed off.Why: a stranger can't tell a careful label from a rushed one without knowing who checked it and against what.
Write the one rule for adding new examples: append, tag, never edit or delete an existing one.Why: without it, "the same set" quietly stops being the same set between one model version and the next.
Leave out a full diff log of every edit and disagreement.Why: that is version control's job, and a play by play buries the one page nobody has time to read anyway.
Re-check the coverage line every time the product's mix of languages or categories shifts.Why: a page that was honest on day one goes quietly wrong the moment the real world moves and the page doesn't.
How to answer this, stage by stage
Seven moves. Ground it in one real golden set you'd actually be handed, not a lecture on documentation in general.
1
Scope it to one golden set you'd actually inherit
Say it like this
"Say I've just taken over as eval lead for a marketplace's translation tool. Someone built a golden set of real product listings eight months ago, four hundred and twenty of them, and then left for another job. I've got the file. I don't have anything else. Let's start there instead of talking about documentation as a topic."
Why this works
One real file with a number attached beats a general essay on documentation hygiene, and it tells the interviewer you're about to show a decision, not describe a policy.
2
Say your structure out loud
Say it like this
"Here's how I'll walk through it: why an undocumented golden set is actually a trap, not just untidy, the one page I'd write to fix it, what that page leaves out on purpose, and the rule that keeps it working the next time someone inherits it."
Why this works
Two seconds of structure signals a plan before a single specific lands, so the interviewer isn't guessing where the answer is headed.
3
Reframe what documentation is actually for
Say it like this
"The instinct is to think documentation is a courtesy, something nice to leave behind for whoever comes next. It's not. Without it, whoever inherits the set has exactly two moves: trust it blindly, or throw it out and rebuild from nothing. Both are expensive, and neither one is actually a decision, because neither is based on anything real."
Why this works
Shows you understand documentation as risk management, not tidiness, which is the actual reframe the question is fishing for.
4
Give the one page, not a wiki
Say it like this
"So here's what I'd write, one page, saved right next to the data file, not linked from somewhere else it can drift from. Five things: what decision this set gates, a launch or a rollback; what it covers, by language and category and the date it was built; how each gold answer got approved and by whom; what it plainly does not cover yet; and the one rule for adding to it without breaking comparisons across model versions."
Why this works
Naming five fixed things is something an interviewer can picture as an actual document, not a vague promise to "write it down somewhere."
5
Prove it with a failure, in four sentences
Say it like this
"Here's what happens without that page. The set I inherited was built back when the marketplace was almost entirely Spanish listings. Eight months later, a third of the catalog was Portuguese, and only eight percent of the golden set was. The eval score kept coming back at ninety seven percent, and it was really a Spanish only number wearing a marketplace wide one's clothes."
Why this works
A concrete failure with real numbers persuades where a warning never does, and it shows exactly what a document's silence costs.
6
Say what you'd deliberately leave out
Say it like this
"I wouldn't turn that page into an audit trail. Every edit, every argument about a single example, every past version of every label, that belongs in version control, not in the page a stranger reads in five minutes. The moment the page tries to be the full history, nobody reads it, and it stops doing its actual job."
Why this works
Shows judgment instead of overcorrecting into a different kind of unusable document, which is exactly what separates this from "just write everything down."
Say it like this
"So: one page, next to the data, naming what it gates, what it covers, who approved it, what it doesn't cover yet, and the one rule for adding to it. If a stranger can't read that page and decide whether to trust the set inside five minutes, it isn't documentation. It's just another file sitting next to the one nobody trusted either."
Why this works
Interviewers remember the last line most, and this one hands them a concrete test, one page, five minutes, they can run on any golden set they're handed afterward.
If you remember one thing
Stages 4 and 6 are the answer. The five fixed sections that make the page inspectable, and the explicit refusal to let it turn into an audit log. Everything else here is proof it works.
Let's learn
The golden set is a spreadsheet. Four hundred and twenty rows, one per product listing, each with a real seller's Spanish or Portuguese text next to the English a person once approved.
Say a cross border marketplace builds a tool that translates a seller's product listing into a buyer's language, so a couch listed in Spanish shows up in English without anyone retyping a word.
Before a golden set like this existed, checking whether a new version of the translation model was safe to launch meant a person reading a random sample of live listings by hand, about three hours a check, and two reviewers often disagreeing about whether the same listing had actually passed.
Knowledge spark: what provenance means here
The record of where a data point came from. For a golden set, that means which listing it was pulled from, who translated it, and who checked and approved it as correct. Without that record, a gold answer is just text that looks official.
Once the golden set existed, the same check ran in minutes and returned one number, a pass rate. A launch decision that used to take an afternoon dropped to fifteen minutes.
Here's the part that's easy to miss.
The extra speed isn't the real story. The real story is that nobody wrote down what those four hundred and twenty examples actually covered, so the pass rate looked like it meant "the whole marketplace" when it might have meant something much narrower.
What checking a golden set looks like with no documentation attached
At its worst, this costs one of two ways. A bad model version ships because the pass rate covered the wrong slice of the catalog and nobody could tell. Or a perfectly good golden set gets thrown out by someone who can't tell if it's trustworthy, and a team spends three months rebuilding what already existed.
A high pass rate on the wrong slice of a catalog isn't a safe number. It's a confident number that happens to be silent about the part that matters.
An undocumented golden set isn't neutral. It's a coin flip, dressed up as a measurement, waiting for whoever inherits it to guess correctly.
The decision that mattered
Ship one page next to the golden set itself: what it gates, what it covers with real dates, who approved each label, what it doesn't cover yet, and the one rule for adding to it. A stranger should be able to read it in five minutes and know whether to trust the set.
The choice I would take back. The golden set shipped as a spreadsheet and nothing else. No page, no note, just four hundred and twenty rows and the assumption that whoever needed to understand it would always be the person who built it.
What I would leave alone. The actual gold translation text in each row is fine exactly as it is. It doesn't need an essay next to it. The problem was never the examples themselves, it was the total silence around what they added up to.
The lesson. A golden set that only its builder can explain isn't really an asset yet. It's a liability with a due date, and the date is whenever that person leaves the room.
Now here is the same thing as a story
Read this version when you've got three minutes, not thirty seconds, since a message from Brazil makes the point better than a spreadsheet ever could.
Small sellers across Latin America list handmade goods and small batch food on Solera Marketplace, in Spanish or Portuguese. Buyers in the US and Spain read the listings in English, thanks to a translation model most of them never think about.
Zuzana Bartok took over as eval lead eight months into the model's life, after Emeka Osayande, who built the golden set from scratch, left for a job at another company. He explained it to her once, on a fifteen minute call, no notes, just talking through the file while she nodded along. That call was the whole handoff. It felt like enough, because he clearly knew what he was doing, and she trusted that.
For months, that trust held up fine. Before every model launch, Zuzana ran the golden set check, got a number back, and if it cleared ninety five percent she signed off. It always cleared. She never opened the actual spreadsheet to look inside it, she just ran the script Emeka had left behind. Nobody had ever told her she needed to.
The page that didn't exist yet, and what it would have needed to say
Then came a Tuesday message from Paloma Duarte, who runs customer support out of Sao Paulo.
A buyer in Lisbon had bought what the listing called a "cama," a bed, and received a photo of a sofa bed instead, still folded out, with a note asking where the rest of the order was. The original Portuguese listing said "sofa cama." The translation had dropped the sofa and kept only the bed. Paloma flagged it, half joking, half not: "does the golden set even have much Portuguese in it?"
Zuzana didn't know. That was the whole problem in one sentence.
She opened the actual file for the first time in eight months and counted. Three hundred and eighty six Spanish examples. Thirty four Portuguese. Out of four hundred and twenty total, Portuguese was eight percent of the golden set.
Then she checked the live catalog. Solera had expanded hard into Brazil that year. Portuguese listings had grown to thirty one percent of everything sellers posted, and climbing.
Share of the golden set vs. share of the live catalog, by source language
The golden set stayed almost entirely Spanish for eight months while the catalog it was supposed to represent quietly became nearly a third Portuguese. The ninety seven percent pass rate had been true the whole time, just true about the wrong marketplace.
Zuzana pulled a fresh sample of live Portuguese listings and read them by hand, the way people used to before the golden set existed. Roughly one in five had an error that changed the meaning, the sofa cama kind of mistake. Spanish, by comparison, was closer to one in forty.
We didn't ship a slightly worse translation. We shipped one that nobody had actually tested for a third of what it was translating.
She wasn't angry at Emeka, once she'd sat with it. He hadn't done anything careless. He built the golden set when Portuguese was a rounding error, and a fifteen minute call was genuinely enough explanation for a set two people both understood. The problem wasn't his judgment. It was that the golden set had no memory of its own shape, so nobody could see the shape stop matching the world it was meant to stand in for.
The gap the golden set never mentioned, found by a customer instead of a page
So here's what she wrote, that same week. One page. Not a longer spreadsheet, not a wiki with ten linked tabs, one page, saved in the same folder as the data file itself. It named the decision the set gates: whether a new model version launches. It named coverage honestly, by language and category, with the date it was last checked. It named that Portuguese sat at eight percent and was known to be under-represented, in plain words, not buried in a footnote. It named who had approved each gold answer. And it gave one rule for adding more: append new examples in matched batches per language, tag them with a date, never touch an existing row.
Now walk the next launch review forward with that page in the room. A new hire on the eval team, someone who had never met Emeka and never would, opened the page before running anything. Five minutes in, she saw the Portuguese line and asked the question Zuzana had needed eight months and a customer complaint to ask herself. The coverage gap got flagged and fixed before the next launch, not after another sofa turned into a bed.
And the part I'd tell myself, if I could go back: the golden set was never the risk. The silence around it was. A spreadsheet with no page next to it doesn't say "trust me" or "don't." It just sits there, and whoever inherits it has to guess, and guessing is expensive whichever way it goes.
SPARK, applied to the page instead of the product
This question asks for an artifact, not a screen, so SPARK still applies, just pointed at a page instead of an interface. A question asking "how would you measure translation quality" would reach for LEAD instead; deciding what a piece of documentation should actually say is a design problem about what a stranger needs to trust something, and that's exactly SPARK's job.
SPARK, aimed at a page instead of a product
S, situation. Zuzana Bartok, eval lead at Solera Marketplace, inherits a golden set built by someone who has already left, with nothing attached to it but the spreadsheet itself.
P, payoff. Not "a tidier spreadsheet." The habit worth building is that whoever inherits the set can tell, in minutes, whether to trust it, extend it, or flag it, instead of quietly guessing either way.
A, anchor. One page next to the data file: the decision it gates, coverage by language and category with a date, who approved each label, known gaps stated plainly, and the append only rule for adding more.
R, risk. The first time a golden set goes undocumented, whoever inherits it either trusts a stale or skewed set blindly and ships something broken, or distrusts a genuinely good one and rebuilds it from nothing. Both cost real time, and neither is actually informed.
K, keep out. Don't turn the page into a full audit log of every edit and every disagreement between labelers, that's version control's job. Do keep the actual gold translation text exactly as it is, no annotation needed there.
Why the anchor and the risk have to match
Check them against each other: does the one page actually survive the day someone discovers a gap in the set? Only if the page already names the gap before the discovery happens, the way Portuguese being at eight percent should have been sitting in plain text from day one. An anchor that only lists what's covered and stays silent about what isn't hasn't actually reduced the risk, it's just moved the guessing from "is this set any good" to "what did they forget to mention."
What stays off the page, on purpose
And if you want to be sure it really works, try it somewhere else
A farm advisory cooperative keeps a golden set of confirmed crop disease photos to test its diagnosis tool. Different field entirely, same trap, and a different reason the gap goes unnoticed.
Same framework, a different crop, a different season nobody wrote down
S. Sanne de Groot, quality lead for the farm advisory app at Terra Vista Co-op. Field agronomists diagnose crop disease from a photo and a five minute phone call today, no model in the loop yet.
P. The habit worth building: whoever inherits the disease photo golden set can tell in minutes which crops and seasons it actually covers, instead of assuming a high pass rate means the model works everywhere it's asked to.
A. One page next to the photo set: the decision it gates, whether a new diagnosis model ships to the app, coverage by crop and growing season and region, who confirmed each diagnosis by name and license, known gaps stated plainly, built entirely from summer wheat, no winter crop examples yet, and the append only rule for adding more.
R. An undocumented version ships confident on a problem it never saw, a rust outbreak on a crop the set never covered, because the pass rate looked fine on a wheat heavy set, and a farmer growing something else entirely trusts a diagnosis nobody actually tested.
K. Don't log every disagreement between two agronomists arguing over a borderline photo, that's a research question for later, not a line on the page a farmer's trust depends on. Do keep each photo's exact confirmed diagnosis untouched.
Swap the trigger and it still runs
- Speed: even if the eval script ran in one second instead of overnight, speed doesn't tell a stranger what the set covers. A fast wrong answer is still wrong, and the page is still needed.
- Cost: storage gets cheap enough that logging every single edit costs nothing. That still isn't the same job as the one page. A full log answers "what changed." The page answers "can I trust this," and cheap storage doesn't make the second question answer itself.
- The model gets better: the translation model gets meaningfully more accurate across the board. A more accurate model, tested against an under documented set, is still an accuracy number nobody can say applies to the whole catalog. Better accuracy doesn't fix a coverage gap, because accuracy and coverage are two different ways to be wrong.
Where people run it wrong
- Writing the documentation somewhere else, a wiki page, a slide, a doc linked from a ticket, so it rots the moment the data file moves or gets renamed and the page doesn't move with it.
- Turning the page into a full history of every edit and calling that thoroughness, so the one page nobody has time to read replaces the one page everyone would have actually used.
- Treating a high pass rate as proof of coverage, when a high pass rate on a narrow slice just means the slice was easy, not that the slice was the whole picture.
How to use it live
If you're asked this cold, ask the room one question back before answering: "if I handed you this golden set with nobody around who built it, what's the first thing you'd need to know before trusting a single number that came out of it?" That buys a few seconds of real thinking time, and whatever the room answers is usually one of the five things that belong on the page anyway.
Flashcards (click a card to flip it)
1 · THE SITUATION
Who is this answer about, and what does she inherit with no notes attached?
Tap to flip
ANSWER
Zuzana Bartok, eval lead at Solera Marketplace. She inherits a 420 example golden set of translated product listings from Emeka Osayande, who explained it once on a fifteen minute call before he left the company.
2 · THE REFRAME
Why isn't documentation just a courtesy left behind for whoever comes next?
Tap to flip
ANSWER
Without it, whoever inherits a golden set only has two moves: trust it blindly, or throw it out and rebuild from nothing. Both are expensive, and neither is actually an informed decision.
3 · THE ANCHOR
What five things does the one page next to the golden set have to name?
Tap to flip
ANSWER
What decision it gates, what it covers by language, category, and date, who approved each gold answer, what it doesn't cover yet, and the one rule for adding to it safely.
4 · THE RISK
What breaks the first time a golden set ships with no documentation?
Tap to flip
ANSWER
Whoever inherits it either trusts a stale or skewed set blindly and ships something broken, or distrusts a genuinely good set and rebuilds it from nothing. Both cost real time, and neither is based on anything real.
5 · THE PROOF
What actually happened when Zuzana finally opened the spreadsheet?
Tap to flip
ANSWER
She found the golden set was 92% Spanish and 8% Portuguese, while the live catalog had grown to 31% Portuguese. A fresh manual check found roughly 1 in 5 Portuguese listings had a meaning-changing error, versus about 1 in 40 for Spanish.
6 · THE NUMBER
Only ___ of the golden set's 420 examples were Portuguese, even though Portuguese listings had grown to ___ percent of the live catalog.
Tap to flip
ANSWER
34 of 420 (8 percent), and Portuguese had grown to 31 percent of the live catalog. The golden set never moved while the marketplace it was meant to represent quietly did.
7 · THE REPLAY
Same coverage gap, new set card. What changes for the next person who inherits the set?
Tap to flip
ANSWER
A new hire with no history with the set reads the page before running anything, sees the Portuguese line stated plainly, and flags the gap in the first five minutes, instead of eight months and a customer complaint later.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor name?
Tap to flip
ANSWER
Terra Vista Co-op's crop disease golden set. The anchor's one page names that it was built entirely from summer wheat photos, with no confirmed winter crop examples yet, and who confirmed each diagnosis by name and license.
Check yourself Score: 0 / 0
Fill in the blank
1. Once Zuzana pulled a fresh sample and read it by hand, roughly 1 in ___ Portuguese listings had a meaning-changing error, against roughly 1 in ___ for Spanish.
Show hint
Think about how much worse an untested slice of a catalog can look once someone finally checks it by hand.
Show answer
1 in 5, 1 in 40. Portuguese listings were failing roughly eight times more often than Spanish ones, a gap the golden set's 97 percent pass rate never showed because it barely tested Portuguese at all.
True or false
2. True or false: the right fix for an undocumented golden set is to log every edit ever made to it, so anyone can see its full history.
Show hint
Ask what job a full edit log actually does, and who is already doing it.
Show answer
False. A full edit log is version control's job. Turning the one page into a full history buries it under noise nobody reads, and it stops doing what a stranger actually needs from it.
Multiple choice
3. Which of these belongs on the one page that ships next to a golden set?
- A. A full diff log of every edit ever made to every example.
- B. A plain statement that the set is under 10 percent Portuguese and shouldn't be trusted alone for a Portuguese only launch.
- C. The personal notes each labeler kept while reviewing examples.
- D. Nothing. A well built golden set doesn't need documentation.
Show hint
Ask which option a stranger could actually use to decide whether to trust the set, in five minutes.
Show answer
B. A and C are exactly the kind of play by play this answer says to leave out. D is the mistake the whole answer argues against, a good set with no documentation is still a guess waiting to happen.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about what made a fifteen minute call feel like enough of a handoff at the time.
Show answer
Model answer: The golden set shipped as a spreadsheet with no page attached, on the assumption that whoever needed to understand it would always be the person who built it. That made sense while Emeka and Zuzana were both in the same fifteen minute call. It stopped making sense the day only one of them still worked there and the catalog itself had quietly changed shape.
Short answer, apply it yourself
5. Pick a test set, benchmark, or golden set you've used or inherited at work. Did you know what it did not cover? What would you have caught sooner if a one page card had come with it?
Show hint
Look for a gap you only discovered because something broke, not because a document told you first.
Show answer
Model answer: "A support macro test set at a previous job only had examples from our biggest customer segment. Nobody had written that down anywhere. When we launched to a smaller segment with very different phrasing, the model looked fine on our tests and failed constantly in production. A one page card stating the set was built entirely from one segment's tickets would have caught it before launch instead of two weeks after."
Multiple choice
6. What is the deeper risk of an undocumented golden set, according to this answer?
- A. It takes longer for engineers to read the underlying code.
- B. Whoever inherits it either blindly trusts a stale or skewed set and ships something broken, or blindly distrusts a good one and rebuilds it from nothing.
- C. The spreadsheet file eventually gets too large to open quickly.
- D. Engineers complain about inconsistent formatting between rows.
Show hint
Think about the two opposite mistakes a stranger can make with no way to tell which one is right.
Show answer
B. Both blind trust and blind distrust are expensive, and an undocumented set gives a stranger no way to tell which one is warranted. That's the actual risk, not formatting or file size.