Describe the composition of a golden set: what proportion should be edge cases?
- Set the edge share at about 30 in 100, chosen on purpose, not sampled at random.Why: a golden set exists to catch trouble, so it should be harder than an average day, not a mirror of one.
- Break the edge share into its real categories before picking one number.Why: overlapping speech, bad audio, a sponsor ad read, and a language switch each fail differently and cost differently, and one percentage hides all of that.
- Move the share up toward 45 in 100 for anything sponsor-billed or legally reviewed, and down toward 15 in 100 for low-stakes hobby use.Why: a missed edge case on a friends-and-family show costs nobody anything; the same miss on a billed ad read costs a client.
- Check the chosen share against what real uploaded episodes actually look like.Why: a number that ignores real traffic is a guess with a decimal point.
- Name which assumption would move the number most: the cost of a miss, not how rare the case is.Why: a rare, expensive miss is worth testing for; a common, cheap one usually isn't.
- Revisit the split whenever the customer mix changes, not just when the model changes.Why: a number frozen at launch stays accurate only for the product that no longer exists.
How to answer this, stage by stage
Nobody is grading whether you land on exactly thirty percent. They are grading whether the split has real arithmetic under it, moves with the stakes, and survives a check against the real world. Seven moves get you there.
Let's learn
Loomcast is a tool that listens to a podcast episode and writes the transcript and the show notes for it, so the host doesn't have to.
Before Loomcast, a solo host spent about two hours after recording just writing show notes by hand: timestamps, a two-line summary, the sponsor read logged for their own records.
Now Loomcast does that first pass itself, usually in about twelve minutes of light editing. On a clean, one-host episode it is close to perfect.
At its worst, a sponsor's ad read gets logged with the wrong timestamp for months without anyone catching it, because the golden set never had a single clip that tested a sponsor read against a busy, overlapping moment. The bill goes out wrong, the client audits their own logs, and the platform spends more time rebuilding trust with that one advertiser than it ever saved any host in show-note time.
The choice I would take back. The golden set's edge share was set once, at 8 in 100, back when every customer was a single hobbyist host. It was never revisited when sponsor-billed shows and legally reviewed interview shows joined the mix.
What I would leave alone. For a hobby show with no ads and no legal review, a thin edge slice is genuinely fine. A wrong line in the show notes there costs nobody anything real.
The lesson. A golden set's edge share isn't a number you set once and trust forever. It's a dial you reset every time the cost of missing an edge case changes, not just when the product does.
Now here is the same thing as a story
Skip this part if you already believe a golden set's edge share needs revisiting every time the stakes change. Read on if you don't.
Mireille Osgood built Loomcast's first golden set three years ago, when the company had nine customers, all solo hosts recording alone in a spare room. She listened to every flagged clip herself back then, checking whether the model's transcript matched what she actually heard.
She set the golden set at 120 clips, ten of them picked as the hard ones. A guest with a cold. A dog barking mid-take. That was basically the whole list of things that ever went wrong.
Loomcast grew. Eighteen months in, it had crossed four hundred shows, including two-guest interview shows, and a dozen networks that sold ads and needed the sponsor's minute logged for billing. For a while, every new feature still got checked against Mireille's original ten hard clips, plus whatever she happened to remember to add.
Then the checking quietly thinned. Someone else took over onboarding the new networks, and the golden set stopped being part of that conversation. It stayed at ten hard clips out of 120, about eight percent, for a full year, even as sponsor-billed shows tripled and interview shows with real crosstalk became half the platform.
The near miss came on a Thursday. A true-crime interview show's producer went to publish an auto-written pull quote for a promo post. The transcript had lifted a line the host said quietly to the producer off mic, layered under the guest's answer during a tense, overlapping moment, and stitched it into the guest's own quote. The producer caught it by reading the clip back before it posted, five minutes before it would have gone out to forty thousand followers.
We didn't almost publish one wrong sentence. We almost put words in a guest's mouth, in public, with her name on them.
Mireille pulled the golden set that week. Of the sixty most-listened-to shows on the platform by then, twenty-two were selling ads, and eleven were the interview format with real overlapping speech. The golden set still had ten hard clips out of 120, and not one of them had two people actually talking over each other in a heated exchange. Every hard clip was calm, alternating turns, nothing like the moment that had almost shipped.
She remembered writing that first list of ten, three years back, at her kitchen table, timing herself: forty minutes to think up every way the transcript could go wrong. A barking dog had felt like the wildest case she could imagine.
She rebuilt the set at 200 clips: 140 typical, and 60 chosen on purpose across four hard categories. She tied the split to what a client actually had on the line. Shows with sponsor billing or legal review got checked against a wider, 45-in-100 version before any new model version could touch them.
Next quarter, the same shape of failure showed up again, this time inside a test run. The wider golden set caught it three days before a similar true-crime show would have published a similar misattributed quote.
The thing I'd tell my kitchen-table self: forty minutes of imagining every way it breaks isn't the same as forty minutes spent asking what happens if we're wrong about the boring parts. I sized the golden set to what I could picture, not to what would actually cost someone something.
The five moves that set the split
This is a sizing question, an arithmetic split between two piles, so BOUND fits. Nobody's trust is flipping here. It's a ratio, and an honest answer for what moves it.
B, break it down. A golden set is two piles added together: typical clips picked to be easy, plus edge clips picked on purpose because they're hard. The edge share is the second pile divided by the whole set.
O, own the numbers. 200 clips total. 140 typical, 70 percent. 60 edge, 30 percent, split across four categories: 20 overlapping speech, 15 rough audio, 15 sponsor ad reads, 10 language switches.
U, use a range. 15 in 100 for a hobby show with no ads and no legal review. 45 in 100 for a sponsor-billed or legally reviewed show. The number moves with what a miss costs, not with how long the product has existed.
N, nail the sanity check. Checked against 500 real uploaded episodes tagged by hand: about 26 in 100 actually had an edge-case feature. 30 sits just above that on purpose, since a golden set exists to catch trouble, not describe an average day.
D, direction. The edge share moves most with the cost of missing a case, not its real-world rarity. A sponsor billing dispute or a misquoted guest costs more than the frequency numbers suggest, so cost, not frequency, sets the dial.
And if you want to be sure it really works, try it somewhere else
A company sells predictive-maintenance alerts for commercial HVAC systems: an AI listens to sensor readings and flags a unit before it fails, so a technician gets sent out before a building loses climate control.
B, break it down. The alert golden set is typical wear signatures, plus edge signatures chosen on purpose: sensor glitches that mimic a real failure, rare multi-part failures, and brand-new equipment models with no history yet.
O, own the numbers. 150 test cases. 105 typical, 70 percent. 45 edge, 30 percent: 20 sensor-glitch false alarms, 15 rare multi-part failures, 10 brand-new unit models.
U, use a range. 15 in 100 is enough for a small residential HVAC company, where a missed edge case means one annoyed homeowner. Push it to 45 in 100 for a hospital or a data center's cooling system, where a missed edge case means a server room overheating or an operating room losing climate control.
N, nail the sanity check. Two years of real service logs show about 20 in 100 callouts started as a signature the model had never been tested on. 30 sits close to that, on the safe side.
D, direction. For this product, the assumption that moves the share most isn't the dollar cost of a miss. It's who's on the other end of the failure. A residential miss costs an afternoon. A hospital miss costs a delayed surgery. Naming who's affected moves the number more than counting dollars does.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule and the split: seventy typical, thirty hard on purpose, checked against 26 in 100 from real uploads. The build-up backs it up if they ask.
Cost: instead of asking what share is right, a client caps the review budget for building the golden set at a fixed number of clips a month. Same rule, solved backward: work out how many of that fixed number should be edge clips, instead of what the total size should be.
The model got better: a new transcription model ships that handles overlapping speech nearly perfectly. The golden set doesn't shrink on its own; re-run the sanity check first, since doing well on ten calm test clips proves nothing about the messy ones.
Where people run it wrong.
They pick a round number, like 20 percent, because it sounds thorough, with nothing behind it.
They set the edge share once at launch and never touch it again as the product's customers and their stakes change.
They check the golden set against itself instead of against something outside it, like real uploaded traffic.
How to use it live. Say the equation before any number: "this is two piles added together, typical clips plus edge clips picked on purpose, and the edge share is the second pile over the whole set." That buys the time to work out what the real split should be, instead of guessing a percentage that sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Golden datasets and test set ownership
- #1 What is a golden dataset and why does the PM usually own it?
- #2 How do you construct a first golden set with no production traffic?
- #4 How do you keep a golden set representative as your user base changes?
- #5 Explain the risk of a golden set that engineering can see during development.
- #6 What is a holdout set and when would you use one for an AI product?
- #7 How do you handle labelling disagreement inside a golden set?