How do you construct a first golden set with no production traffic?
- Build the golden set from hand-written categories now, instead of waiting for real calls to arrive.Why: waiting means the first pilot customers become the eval set, and a wrong callback number in week one costs more than four days of writing up front.
- Break the total into two real terms: how many categories of hard voicemail exist, and how many examples get written for each one.Why: a single guessed total hides which assumption is actually doing the work.
- Size each category by how much it would hurt if the model got it wrong, not with the same count everywhere.Why: a clean, quiet voicemail needs far fewer examples than one with a spoken dollar amount or a callback number.
- Give the total as a range, 90 to 210, not a single confident number like 135.Why: a golden set built with no traffic is itself an estimate, and one number pretends a precision nobody has yet.
- Check the range against how long it actually takes one person to write and review that many examples.Why: a total that looks fine on paper can still blow a ship date nobody checked it against.
- Name category count as the assumption to watch, and leave room to add a category once real calls start.Why: it's the number that swings the total most, and the one nobody can see yet without a single real call.
How to answer this, stage by stage
Eight moves. The trap is naming one confident total, "about 150 examples," with no visible equation behind it.
Let's learn
Picture this: a feature that has to prove it works before it's ever used for real. That's Ringnote, the voicemail transcription and summary feature built into Cordant, a CRM for small sales and service teams.
Before Ringnote, a rep who missed a call had to listen to the voicemail, often rewinding to catch a number or a name, then type a note before calling back. About 90 seconds a message, once you count the playback and the typing. A rep handling 20 voicemails a day loses half an hour just processing them before dialing a single callback.
Ringnote is built to turn that 90 seconds into about 10. A rep reads two lines instead: "Customer asking about order 48213, wants a callback before 3pm." That's the promise.
Here's the turn. The problem isn't whether Ringnote's transcripts and summaries are accurate. Nobody can say whether they are, because there's no traffic yet to check them against. So the plan on the roadmap was simple: launch, then sample real customer calls after two or three weeks and build the eval set from those.
Clean baseline: 10
Background noise: 15
Strong accent: 15
Spoken numbers: 20
Rambling, long message: 15
Angry or urgent tone: 15
Crosstalk: 15
Silence or hang-up: 10
Non-English or code-switch: 20
# total (added up, not averaged)
135
# range, if the category count itself is wrong
90 to 210
What that costs at its worst: a rep gets a wrong callback number from a bad summary, doesn't double check, and calls the wrong person, or misses a window entirely. Once that happens twice, the rep doesn't just go back to listening to every voicemail. They listen and read the summary, checking one against the other. Ringnote ends up slower than no feature at all, because now there's an extra step on top of the old one.
The choice I would take back. Planning to build the eval set from real customer calls after a soft launch. That made sense in the early planning meeting, when nobody had any idea which categories would actually matter and building a set by hand felt like guessing without data. It stopped making sense the moment beta customers became the ones finding out the hard way.
What I would leave alone. The clean baseline category, a quiet, single caller stating a name and a clear number, doesn't need extra rigor. Every off-the-shelf speech model already gets this right almost every time. Ten examples there just confirm the pipe works end to end. Put the real budget into the categories that are actually hard.
The lesson. You don't need real calls to know what a golden set needs. You need an honest list of the ways a voicemail gets hard, and you can write that list on a whiteboard before a single customer calls.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why a whiteboard list beat waiting for real customers to find the bugs first.
The whiteboard in Marguerite Ashby's office has had the same corner reserved for six months: a list that never quite gets finished. She's the product manager on Ringnote, and she's the one who wrote the original plan, back when the feature was a slide deck and nothing else: launch, then sample real customer calls after two or three weeks and build the eval set from those.
It was a reasonable plan at the time. Nobody on the team had listened to more than a handful of voicemails, all from their own phones, all clean. Building a category list from nothing felt like guessing. Waiting for real data felt like the responsible move.
For months, that plan just sat there, undisturbed, because launch was always a quarter away.
Then leadership pulled the beta forward. A prospect at a trade show had asked for exactly this, and now there were five pilot customers lined up and a date on the calendar: two weeks out.
Foster Lindgren, three weeks into the job as an engineer on the team, asked the question in standup that nobody had a real answer to. "How do we know if it's actually good before we show it to anyone?" The honest answer was: we don't. We were going to find out after beta, from the pilot customers themselves.
Marguerite went back to her office and looked at the whiteboard. The list wasn't actually empty. She'd been adding to it for months, every time someone mentioned a way a voicemail might trip the model up: a warehouse worker calling from the loading dock, a customer reading off a dollar figure, a message left in Spanish because the caller assumed nobody would notice. Nine categories, already there, never written down as a plan.
So she stopped waiting. She and one engineer spent four and a half days writing 135 examples, category by category, the risky ones first: 20 for spoken numbers, 20 for non-English or code-switched messages, 10 for the clean baseline nobody needed to worry about.
Here's the turn. It wasn't about whether Ringnote was accurate. It was that nobody could say whether it was, and the plan to find out involved a real customer discovering it first. Once Marguerite could see that plainly, the fix wasn't more testing. It was writing down what she already knew.
The old decision, from that first slide-deck meeting: sample real traffic after launch, because nobody had the data yet to do anything else. What nobody said out loud in that meeting was that the categories worth testing didn't need real traffic to name. They needed someone willing to write them down before they were proven necessary.
Before the beta ever started, the golden set caught two real problems. The model dropped the decimal in a spoken dollar amount three times out of twenty. It summarized a Spanish-language voicemail as if it had been left in English, twice out of twenty. Both fixed, both re-tested, four days before the first pilot customer ever left a message.
Waiting for traffic would have meant a real customer finding those two bugs, on a call that mattered to them. Instead, a whiteboard list found them first, for the cost of four and a half days of writing, not one unhappy pilot customer.
The part I'd go back and tell myself: I was waiting for the calls to tell me what mattered. I already knew. I just hadn't written it down yet.
BOUND, run with nothing to sample from yet
This is a sizing question with no data behind it, so BOUND fits and FLIPS doesn't. Nobody's habit snapped here. A team just had to decide how big a list to write before anyone could prove the list was right.
B, break it down. Total examples equals the number of hard-voicemail categories, times how many examples get written for each one, added up category by category. Not a single guessed number.
O, own the numbers. Nine categories: clean baseline, background noise, strong accent, spoken numbers, rambling message, angry tone, crosstalk, silence or hang-up, non-English or code-switch. Spoken numbers and code-switching get 20 examples each, the categories that get the review, rambling and crosstalk get 15, and the clean baseline gets 10.
U, use a range. Added up, that's 135 at the point estimate. Six categories instead of nine brings it to 90. Fourteen brings it to 210. That's the honest range: 90 to 210.
N, nail the sanity check. Writing and reviewing one example takes about 15 minutes. At 135, that's 34 hours, just over four days, against the 10 workdays before beta. Even the high end, 210 examples, costs about six and a half days. Both fit, with room to spare at the low end and a real squeeze at the high end.
D, direction. Category count swings the total the most. Moving it from six to fourteen, with examples-per-category held steady, swings the total by 120 examples. Moving examples-per-category from 10 to 20 instead, with category count held steady, only swings it by 90. Category count is the assumption to watch, and it's the one nobody can see clearly until the first real calls start arriving.
And if you want to be sure it really works, try it somewhere else
Brackenfield HVAC builds the dispatch software techs use to route an incoming service call: emergency no-heat, routine maintenance, a warranty claim, a sales inquiry, a billing question, or wrong number and spam. Six categories, fixed by how the business already runs, and the same sizing question shows up wearing a different coat.
B, break it down. Total examples equals six fixed categories, times however many examples each one needs, added up.
O, own the numbers. Silas Braithwaite, Brackenfield's eval lead, sized each category by how much its phrasing actually varies: emergency no-heat at 12, routine maintenance at 10, warranty claim at 18, because customers describe a warranty issue a dozen different ways, sales inquiry at 10, billing question at 12, wrong number or spam at 8.
U, use a range. Added up, that's 70 at the point estimate.
N, nail the sanity check. Each example takes about 20 minutes to write and review, longer than Ringnote's, because a believable warranty claim needs a real product model number and a plausible install date. At 70 examples, that's about 23 hours, just under three days, against a five-day window before a trade-show demo.
D, direction. Here the swing flips. Category count barely moves, the business only has six ways to route a call, so the range from 5 to 7 categories only swings the total by 24 examples. Examples-per-category is the one that swings it, from 8 to 20 per category swings the total by 72 examples, because how wildly a category gets phrased matters more than how many categories exist.
The old plan Brackenfield would take back is a cousin of Marguerite's, not a copy of it. Cordant had assumed real traffic was the only honest source for category names. Brackenfield's team had assumed a flat number, "15 examples per category," would work for a business that only ever sees six kinds of calls, without asking which of those six actually needs the most. Same mistake, different shirt: treating every category as though it needed the same amount of proof.
Swap the trigger and it still runs.
Speed: an interviewer wants a usable answer before the meeting ends. Same nine categories, same equation, examples-per-category stated as a round 10 to 20 instead of researched one by one, tightened once real calls start coming in.
Cost: engineering can only pay for 80 hand-written examples total. Solved backwards: at a fixed 80, the clean baseline and silence categories drop to 5 each first, and spoken numbers and non-English stay at 20, because those are the two that cost a real callback if they're wrong.
The model got better: the speech-to-text vendor ships an update that halves background-noise errors. The category count still has to be estimated the same way, because background noise was never the category swinging the total. A better model changes which examples you expect to pass already, not how many you still need to write.
Where people run it wrong.
They wait for a beta to gather "real" examples, so the first customers get an unmeasured model.
They write the same number of examples per category regardless of how much that category's phrasing actually varies, wasting budget on the easy ones.
They treat the point estimate as fixed and never say what would make it grow, so someone above them finds a new category two days before ship with no time built in for it.
How to use it live. Say the equation out loud before naming a single number: categories, times examples per category, added up, not averaged. Then say plainly that a golden set built with no traffic is itself an estimate, so a range is coming, not a guess dressed up as a fact. That buys thinking time and tells the interviewer a real number is on the way.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Golden datasets and test set ownership
- #1 What is a golden dataset and why does the PM usually own it?
- #3 Describe the composition of a golden set: what proportion should be edge cases?
- #4 How do you keep a golden set representative as your user base changes?
- #5 Explain the risk of a golden set that engineering can see during development.
- #6 What is a holdout set and when would you use one for an AI product?
- #7 How do you handle labelling disagreement inside a golden set?