Describe how you would build a golden set for a product I name.
- Break the golden set into the product's real categories, and budget typical plus edge-case examples inside each one.Why: a flat random sample can leave a real category almost empty while the total still looks fine.
- Name each category's example count out loud, don't hide a guessed total behind one confident number.Why: an unstated per-category number is a guess wearing a total's clothes.
- Give the total as a range, not one point number.Why: the real total depends on how many categories the product actually ends up serving, and that grows with usage.
- Check the total against how long it would really take a person to hand-label it.Why: a number that implies two days for 190 careful examples is too good to be true, and one that implies a whole quarter is too slow to ship.
- Watch which assumption, more categories or deeper edge-case coverage inside each one, swings the total most.Why: that's the number a sharp interviewer pushes on, and "it depends" isn't an answer.
- Leave the low-stakes, high-volume categories thin, on purpose.Why: padding an already-simple category doesn't catch anything new, it just spends review time a harder category needs more.
How to answer this, stage by stage
Eight moves. The trap is naming one round number and never showing which categories it's actually built from.
Let's learn
Recap is the feature inside Waybridge, a project-management tool, that listens to a meeting and writes the summary and the action items for you.
Before Reza's team built a real eval, they used the same rule every feature team at Waybridge used: twenty hand-labeled examples, picked at random, done in about a day and a half. Back then Recap only summarized two kinds of meeting, so twenty felt like plenty.
Now the team is building one sized on purpose: 190 examples spread across the seven kinds of meeting Recap actually covers today, standups, retros, 1:1s, client calls, incident reviews, kickoffs, and planning reviews.
Daily standup: 10 + 5 = 15
Sprint retro: 15 + 10 = 25
1:1: 12 + 8 = 20
Client status call: 18 + 12 = 30
Incident / postmortem review: 20 + 20 = 40
Kickoff / onboarding call: 15 + 10 = 25
Roadmap / planning review: 20 + 15 = 35
# total (added up, not averaged)
190
Here's the turn. More examples isn't really the fix.
Say it plainly: the old set of twenty had exactly one incident-review example in it, buried among nineteen ordinary standups and retros. The score could look great and still be unable to see the one category where a miss costs the most.
What that costs at its worst: Recap reports a healthy overall number, leadership trusts it, and the one category where a wrong summary actually costs a customer something, an incident review, keeps failing quietly, because the eval was never built with enough of that category to notice.
The choice I would take back. Copying Waybridge's org-wide flat quota, twenty examples, sampled at random, for Recap's eval. That made sense when Recap only summarized two similar meeting types at beta. It stopped making sense the moment Recap grew to seven very different shapes of meeting, and a flat random sample stopped being able to represent any of them well.
What I would leave alone. The daily standup category doesn't need a bigger or more careful golden set. Standups are short, low-stakes, and nearly identical in shape, so 15 is enough there, and adding more would just teach the model to be right on the easiest thing it does.
The lesson. A golden set sized by a flat count only proves you have enough examples in general. It says nothing about whether you have enough of the ones that can actually break.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why one missing category mattered more than a hundred extra examples would have.
Reza Farahani can tell a good meeting summary from a bad one in about four seconds, without reading past the first line. He's the product manager for Recap, and he wrote its very first golden set himself, the night before the beta launch review: twenty examples pulled at random from a shared recordings folder. It took him about a day and a half. Back then Recap only handled two kinds of meeting, daily standups and sprint retros, so twenty examples split roughly in half between two simple formats felt like plenty.
For the better part of a year, that felt right. Recap kept scoring in the mid-nineties against those twenty examples. The team added client calls. Then 1:1s. Then incident reviews, for engineering teams who wanted a written record after something broke. Nobody touched the golden set. Why would they. The number kept looking good.
The habit thinned in three small steps. First, when Recap grew into incident reviews, nobody added examples for that category, there wasn't time before the next release and the score was still fine. Then, when an engineer asked in a planning meeting whether the golden set still matched what Recap actually covered, the answer was "probably," and the meeting moved on. Then, quietest of all, new hires joining the Recap team stopped opening the twenty examples at all. They just watched the dashboard. Ninety-four percent felt like a fact, not a measurement anyone had checked lately.
Then, on a Tuesday, one of Waybridge's client engineering leads ran an incident postmortem call after an outage. Fifty-two minutes, four people, a clear cause. Four minutes in, the call drifted, someone raised an unrelated billing question, then drifted back. Recap wrote up the call and listed three follow-up actions. It missed the fourth: rotate the API key that had been exposed. Nobody caught it that day.
A week later, almost the same outage happened again, same exposed key, still live. The client's lead pulled up the old summary to check what they'd agreed to do, and the fourth action simply wasn't there. It wasn't wrong. It was gone. Recap had followed the conversation right up until the tangent, then treated whatever came after the tangent as the real ending.
Reza pulled the golden set that afternoon. Of the original twenty examples, exactly one was an incident review.
The old decision that put him there was almost a year old, and easy to remember, because it took about ninety seconds. Someone on the platform team had shared a doc: eval quota for new features, twenty hand-labeled examples, sampled at random, ship it. Every team used the same doc. Reza copied it because copying it was fast, and Recap, at the time, only did two things. Nobody in that room was wrong. Twenty random examples covering two simple formats is a reasonable golden set. Twenty random examples covering seven very different ones is a coin flip about which seven you even got.
So he rebuilt it that month: 190 examples, weighted by category, forty of them incident reviews, a third of those built specifically around calls that drift off-topic mid-conversation the way the real one had. Run against the new set, Recap doesn't just report a healthy overall number and stop. It reports: six of the forty postmortem examples get the follow-up action wrong when the call drifts for more than ninety seconds. That's the exact shape of what happened to the client. Caught in an afternoon of testing, instead of found by a customer, twice.
The part I'd go back and tell myself: I used to think a golden set was there to catch mistakes. It's really there to decide, ahead of time, which mistakes you're even able to catch.
BOUND, built from seven categories, not one guess
This is a sizing question wearing a design question's clothes, so BOUND fits and FLIPS doesn't. Nobody's habit snapped on a Tuesday here. A team had to decide how much of each kind of meeting its own eval could afford to skip.
B, break it down. Golden set size equals the sum, across every meeting type Recap covers, of the typical examples for that type plus its edge cases. Seven separate additions, not one guess split evenly.
O, own the numbers. Seven meeting types: standup 15, retro 25, 1:1 20, client status call 30, incident review 40, kickoff call 25, planning review 35. Added up: 190.
U, use a range. At the point estimate, 190, for the seven types Recap covers today. If a team only uses four of them, 110. If a large account runs ten or eleven types, 280. Honest range: 110 to 280.
N, nail the sanity check. 190 examples, at about 30 minutes each to write a gold summary and note the real action items, comes to roughly 95 hours, about two and a half weeks of one reviewer doing nothing else. That's also a sliver of what Recap actually processes, something like 4,000 calls a day company-wide, but a golden set was never meant to sample everything. It's meant to sample what breaks.
D, direction. Adding one more meeting type at the average size moves the total by about 27. Doubling the edge cases inside every category moves it by about 80. Depth beats breadth. That's the number worth watching, not the category count.
And if you want to be sure it really works, try it somewhere else
A mid-size immigration-paperwork translation agency runs a tool called Clarity Check. It doesn't summarize anything, it flags terminology drift, the moments a legal term got translated into something close but wrong, so a human reviewer can catch it before a document gets filed.
Marlena Oduya is the agency's terminology reviewer, and she's the one who ends up sizing its golden set.
B, break it down. Holdout-style golden set size equals the sum, across every document type the agency certifies, of typical examples plus edge cases, added up by document type, not guessed as one number per language pair.
O, own the numbers. Six document types: birth or marriage certificates 12, employer sponsorship letters 18, visa petition letters 18, medical records 22, court transcripts 30, asylum affidavits 35, the highest, because a mistranslated detail there can affect a real asylum claim. Total: 135.
U, use a range. At 135 for six document types. If the agency only certifies four, skipping court transcripts and medical records because those need specialist translators it doesn't keep in-house, 90. If it adds a second high-stakes language pair with different legal terms, 200. Range: 90 to 200.
N, nail the sanity check. At 135 examples, roughly 35 minutes each for a bilingual reviewer to check the translation and write the gold terminology map, that's about 79 hours, close to two weeks of one reviewer's time. Against a monthly grading budget of 50 hours, that's the whole set spread across two months, not one sitting.
D, direction. In principle, doubling edge-case depth would swing Clarity Check's total more than adding a document type, same shape as Recap's answer. But immigration paperwork has a fixed list of document types the agency is legally required to certify. Marlena can't decide asylum affidavits don't count. The lever she can actually pull is which language pair gets full depth first, and that's what she watches.
The old decision Clarity Check would take back is a cousin of Recap's, not a copy. Waybridge had sized by a flat quota because copying the org's standard playbook was fast. The agency's old habit was different: a flat number of examples per language pair, chosen back when it only handled one kind of document, so language pair really was the whole story. Same mistake wearing a different coat: sizing an eval by the axis that was true once, and never checking it again as the real shape of the work changed.
Swap the trigger and it still runs.
Speed: an interviewer wants a number before the meeting ends. Round to 25 per document type, six types, 150 as the working total, revisit the exact split once real case files start arriving.
Cost: the agency can only pay a certified bilingual reviewer for 50 hours of grading a month. At 35 minutes an example, that's about 85 examples a month, so the golden set gets built across two monthly passes, not one sitting.
The model got better: a new legal-terminology glossary cuts down ambiguous terms. That doesn't shrink the golden set. Fewer ambiguous terms just means fewer of the 135 examples turn out to be genuinely hard, which is a different thing from needing fewer of them.
Where people run it wrong.
They size the golden set by the unit that was true at launch, language pair, meeting type, feature, and never check whether that's still the right unit as the product's shape changes.
They let one overall accuracy number stand in for every category, so a badly under-covered high-stakes category hides behind a healthy average.
They set the size once at launch and never revisit it as the stakes grow, more customers, more categories, past what the original number was ever built to answer.
How to use it live. Say the equation before naming a single number: total equals the sum across categories of typical plus edge, never one flat count split evenly. That buys a few real seconds to do the arithmetic instead of guessing out loud.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Golden datasets and test set ownership
- #1 What is a golden dataset and why does the PM usually own it?
- #2 How do you construct a first golden set with no production traffic?
- #3 Describe the composition of a golden set: what proportion should be edge cases?
- #4 How do you keep a golden set representative as your user base changes?
- #5 Explain the risk of a golden set that engineering can see during development.
- #6 What is a holdout set and when would you use one for an AI product?