ConceptAdvancedEval-Driven Specification / Golden datasets and test set ownership / #3

Describe the composition of a golden set: what proportion should be edge cases?

The direct answer
Build most of the golden set from easy, typical cases, and set aside about 30 out of every 100 items as edge cases picked on purpose: overlapping speakers, bad audio, non-English segments, a sponsor reading their own ad. Move that share up toward 45 in 100 when a missed edge case costs real money or trust, and down toward 15 in 100 when it barely matters, then check the number against what real uploads actually look like. Reset it whenever the cost of missing an edge case changes, not just when the product does.
Do this, in order
  1. Set the edge share at about 30 in 100, chosen on purpose, not sampled at random.Why: a golden set exists to catch trouble, so it should be harder than an average day, not a mirror of one.
  2. Break the edge share into its real categories before picking one number.Why: overlapping speech, bad audio, a sponsor ad read, and a language switch each fail differently and cost differently, and one percentage hides all of that.
  3. Move the share up toward 45 in 100 for anything sponsor-billed or legally reviewed, and down toward 15 in 100 for low-stakes hobby use.Why: a missed edge case on a friends-and-family show costs nobody anything; the same miss on a billed ad read costs a client.
  4. Check the chosen share against what real uploaded episodes actually look like.Why: a number that ignores real traffic is a guess with a decimal point.
  5. Name which assumption would move the number most: the cost of a miss, not how rare the case is.Why: a rare, expensive miss is worth testing for; a common, cheap one usually isn't.
  6. Revisit the split whenever the customer mix changes, not just when the model changes.Why: a number frozen at launch stays accurate only for the product that no longer exists.

How to answer this, stage by stage

Nobody is grading whether you land on exactly thirty percent. They are grading whether the split has real arithmetic under it, moves with the stakes, and survives a check against the real world. Seven moves get you there.

1
Pin down what one item is, and name the edge categories before touching a percentage
Say it like this
"Before I give you a number, let's agree what we're counting. One item here is a two to five minute clip pulled from a real episode. And 'edge case' isn't one thing. It's at least four things: two people talking over each other, a rough phone line, a guest switching into another language, and a sponsor reading their own ad live."
Why this works
Stops the interviewer from hearing a vague percentage attached to nothing concrete.
2
Say the equation out loud
Say it like this
"The golden set is two piles added together. Easy, typical clips, plus edge clips picked on purpose. The edge share is just the second pile, divided by the whole set."
Why this works
Shows the arithmetic before a single number lands, so what follows reads as a build-up, not a guess with a percent sign on it.
3
Own the numbers for one real golden set
Say it like this
"Say the set has 200 clips. I'd put 140 of them, seventy percent, as clean, typical episodes: one host, a decent mic, English, nobody talking over anyone. The other 60, thirty percent, I'd split across four hard categories on purpose: 20 with overlapping speakers, 15 with rough audio, 15 with a sponsor reading their own ad, and 10 with a guest switching languages mid-sentence."
Why this works
Turns "thirty percent" into a number a reviewer can check, line by line.
4
Give the share a range tied to what a miss actually costs
Say it like this
"That thirty percent isn't fixed. For a friends-and-family show with no ads and nobody reviewing the transcript for legal reasons, I'd let it drop to fifteen in a hundred, since a miss there is just an odd line in the show notes. For a show with sponsor-read ads billed by the minute, or an interview show a lawyer actually reads, I'd push it up to forty-five in a hundred, because a miss there costs real money or a real complaint."
Why this works
Shows the number is tied to what's at stake, not to how long the product has existed.
5
Check the number against what real uploads look like
Say it like this
"We pulled 500 real uploaded episodes and had someone tag which ones actually had overlapping speech, rough audio, a language switch, or a sponsor read. About 26 in 100 did. My 30 percent sits just above that on purpose. A golden set is there to catch trouble, not to describe an average Tuesday."
Why this works
This is the step most estimates skip, the one that turns a plausible-sounding number into one that survives a follow-up question.
6
Name which assumption moves the number most
Say it like this
"If I had to guess which assumption swings this hardest, it isn't how often an edge case shows up. It's what a miss costs. A sponsor's ad read logged with the wrong timestamp turns into a billing dispute and a client we might lose. A guest's off-the-record aside getting quoted by accident is worse than that. Raise the cost of either one and the edge share should climb, even if the real-world frequency never changes."
Why this works
Answers the hardest part of the question directly, and shows judgment instead of just arithmetic.
7
Close on the rule, in one breath
Say it like this
"So: seventy typical, thirty hard on purpose, checked against what real uploads actually look like, and moved up or down depending on what a miss actually costs, not on how the product happens to feel that quarter."
Why this works
Restates the whole decision in one breath, the way a strong close actually sounds.
If you remember one thing Thirty in a hundred is a starting point, not a rule. What actually sets the number is what a miss costs, checked against what real uploads look like.

Let's learn

Loomcast is a tool that listens to a podcast episode and writes the transcript and the show notes for it, so the host doesn't have to.

Knowledge spark: what is a golden set? A small, hand-picked set of examples used to test a model before it ships. Not a random sample. Every item in it was chosen on purpose, usually because it's hard, so a passing score actually means something.

Before Loomcast, a solo host spent about two hours after recording just writing show notes by hand: timestamps, a two-line summary, the sponsor read logged for their own records.

Now Loomcast does that first pass itself, usually in about twelve minutes of light editing. On a clean, one-host episode it is close to perfect.

The mistakes that show up are not spread evenly across every episode. They cluster in a handful of situations nobody tested for on purpose, and the mistake nobody notices for months is the one that reads calm, tidy, and wrong.

At its worst, a sponsor's ad read gets logged with the wrong timestamp for months without anyone catching it, because the golden set never had a single clip that tested a sponsor read against a busy, overlapping moment. The bill goes out wrong, the client audits their own logs, and the platform spends more time rebuilding trust with that one advertiser than it ever saved any host in show-note time.

The decision that mattered Set the edge share by what a miss costs, not by habit. A golden set built for a hobby show and never revisited will keep passing a model that's blind to exactly the clips a sponsor-billed show can't afford to get wrong.

The choice I would take back. The golden set's edge share was set once, at 8 in 100, back when every customer was a single hobbyist host. It was never revisited when sponsor-billed shows and legally reviewed interview shows joined the mix.

What I would leave alone. For a hobby show with no ads and no legal review, a thin edge slice is genuinely fine. A wrong line in the show notes there costs nobody anything real.

The lesson. A golden set's edge share isn't a number you set once and trust forever. It's a dial you reset every time the cost of missing an edge case changes, not just when the product does.

Now here is the same thing as a story

Skip this part if you already believe a golden set's edge share needs revisiting every time the stakes change. Read on if you don't.

Mireille Osgood built Loomcast's first golden set three years ago, when the company had nine customers, all solo hosts recording alone in a spare room. She listened to every flagged clip herself back then, checking whether the model's transcript matched what she actually heard.

Two piles compared. Left, 120 clips, 10 picked as hard, all calm. Right, 200 clips, 60 picked as hard, on purpose, in a darker color.
The golden set's own composition, before and after Mireille rebuilt it

She set the golden set at 120 clips, ten of them picked as the hard ones. A guest with a cold. A dog barking mid-take. That was basically the whole list of things that ever went wrong.

Loomcast grew. Eighteen months in, it had crossed four hundred shows, including two-guest interview shows, and a dozen networks that sold ads and needed the sponsor's minute logged for billing. For a while, every new feature still got checked against Mireille's original ten hard clips, plus whatever she happened to remember to add.

Then the checking quietly thinned. Someone else took over onboarding the new networks, and the golden set stopped being part of that conversation. It stayed at ten hard clips out of 120, about eight percent, for a full year, even as sponsor-billed shows tripled and interview shows with real crosstalk became half the platform.

The near miss came on a Thursday. A true-crime interview show's producer went to publish an auto-written pull quote for a promo post. The transcript had lifted a line the host said quietly to the producer off mic, layered under the guest's answer during a tense, overlapping moment, and stitched it into the guest's own quote. The producer caught it by reading the clip back before it posted, five minutes before it would have gone out to forty thousand followers.

We didn't almost publish one wrong sentence. We almost put words in a guest's mouth, in public, with her name on them.

Mireille pulled the golden set that week. Of the sixty most-listened-to shows on the platform by then, twenty-two were selling ads, and eleven were the interview format with real overlapping speech. The golden set still had ten hard clips out of 120, and not one of them had two people actually talking over each other in a heated exchange. Every hard clip was calm, alternating turns, nothing like the moment that had almost shipped.

She remembered writing that first list of ten, three years back, at her kitchen table, timing herself: forty minutes to think up every way the transcript could go wrong. A barking dog had felt like the wildest case she could imagine.

She rebuilt the set at 200 clips: 140 typical, and 60 chosen on purpose across four hard categories. She tied the split to what a client actually had on the line. Shows with sponsor billing or legal review got checked against a wider, 45-in-100 version before any new model version could touch them.

Next quarter, the same shape of failure showed up again, this time inside a test run. The wider golden set caught it three days before a similar true-crime show would have published a similar misattributed quote.

The thing I'd tell my kitchen-table self: forty minutes of imagining every way it breaks isn't the same as forty minutes spent asking what happens if we're wrong about the boring parts. I sized the golden set to what I could picture, not to what would actually cost someone something.

The five moves that set the split

This is a sizing question, an arithmetic split between two piles, so BOUND fits. Nobody's trust is flipping here. It's a ratio, and an honest answer for what moves it.

B, break it down. A golden set is two piles added together: typical clips picked to be easy, plus edge clips picked on purpose because they're hard. The edge share is the second pile divided by the whole set.
O, own the numbers. 200 clips total. 140 typical, 70 percent. 60 edge, 30 percent, split across four categories: 20 overlapping speech, 15 rough audio, 15 sponsor ad reads, 10 language switches.
U, use a range. 15 in 100 for a hobby show with no ads and no legal review. 45 in 100 for a sponsor-billed or legally reviewed show. The number moves with what a miss costs, not with how long the product has existed.
N, nail the sanity check. Checked against 500 real uploaded episodes tagged by hand: about 26 in 100 actually had an edge-case feature. 30 sits just above that on purpose, since a golden set exists to catch trouble, not describe an average day.
D, direction. The edge share moves most with the cost of missing a case, not its real-world rarity. A sponsor billing dispute or a misquoted guest costs more than the frequency numbers suggest, so cost, not frequency, sets the dial.

The build-up: 200 clips, five pieces
140 typical
20
15
15
10
Typical, clean episodes Overlapping speakers Rough audio Sponsor ad reads Language switches
The two categories most likely to touch real money or a real complaint, sponsor ad reads and overlapping speakers, are only 35 of the 200 clips. Small slices, disproportionate cost if they're missing.
A number line marking the edge share, low to high: 15 in 100 for a hobby show with nothing at stake, 26 in 100 from real uploads as a sanity check, 30 in 100 as the pick for most shows, and 45 in 100 for a sponsor-billed or legally reviewed show.
The edge share, low to high, with the real-upload check marked for scale
What moves the edge share most
Sponsor billing dispute becomes a real cost of a miss+15 pts
Locked to hobby-only customers, nothing at stake−15 pts
Customer mix shifts entirely to sponsor-billed shows+10 pts
Real incidence sampled higher than expected, 26 to 34+4 pts
The two biggest swings both come from the same lever: what a miss actually costs, not how often it happens. The smallest swing, a sample that undercounted real edge cases, barely moves the number at all.

And if you want to be sure it really works, try it somewhere else

A company sells predictive-maintenance alerts for commercial HVAC systems: an AI listens to sensor readings and flags a unit before it fails, so a technician gets sent out before a building loses climate control.

B, break it down. The alert golden set is typical wear signatures, plus edge signatures chosen on purpose: sensor glitches that mimic a real failure, rare multi-part failures, and brand-new equipment models with no history yet.
O, own the numbers. 150 test cases. 105 typical, 70 percent. 45 edge, 30 percent: 20 sensor-glitch false alarms, 15 rare multi-part failures, 10 brand-new unit models.
U, use a range. 15 in 100 is enough for a small residential HVAC company, where a missed edge case means one annoyed homeowner. Push it to 45 in 100 for a hospital or a data center's cooling system, where a missed edge case means a server room overheating or an operating room losing climate control.
N, nail the sanity check. Two years of real service logs show about 20 in 100 callouts started as a signature the model had never been tested on. 30 sits close to that, on the safe side.
D, direction. For this product, the assumption that moves the share most isn't the dollar cost of a miss. It's who's on the other end of the failure. A residential miss costs an afternoon. A hospital miss costs a delayed surgery. Naming who's affected moves the number more than counting dollars does.

Which assumption moves it depends on the product At Loomcast, it's the dollar cost of a miss, a sponsor billing dispute. On the HVAC alerts, it's who gets hurt by a miss, a homeowner versus a hospital. Both are "cost," but they're not counted the same way, and naming which one applies is half the answer.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule and the split: seventy typical, thirty hard on purpose, checked against 26 in 100 from real uploads. The build-up backs it up if they ask.
Cost: instead of asking what share is right, a client caps the review budget for building the golden set at a fixed number of clips a month. Same rule, solved backward: work out how many of that fixed number should be edge clips, instead of what the total size should be.
The model got better: a new transcription model ships that handles overlapping speech nearly perfectly. The golden set doesn't shrink on its own; re-run the sanity check first, since doing well on ten calm test clips proves nothing about the messy ones.

Where people run it wrong.
They pick a round number, like 20 percent, because it sounds thorough, with nothing behind it.
They set the edge share once at launch and never touch it again as the product's customers and their stakes change.
They check the golden set against itself instead of against something outside it, like real uploaded traffic.

How to use it live. Say the equation before any number: "this is two piles added together, typical clips plus edge clips picked on purpose, and the edge share is the second pile over the whole set." That buys the time to work out what the real split should be, instead of guessing a percentage that sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about how a golden set should be split between typical and edge cases, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question about a ratio and its arithmetic, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Mireille Osgood, eval lead at Loomcast, an AI podcast transcription and show-notes tool. She built the company's first golden set three years ago.
3 · THE CHECK THAT THINNED
What kept the golden set's edge share honest at first, and how did it quietly stop?
Tap to flip
ANSWER
Every new feature got checked against Mireille's original ten hard clips, plus whatever she remembered to add. The share stayed near 8 percent for a year even as sponsor-billed shows tripled, because nobody owned revisiting it.
4 · THE SPLIT, IN THIS STORY
What's the two-pile split this answer turns on?
Tap to flip
ANSWER
Typical clips picked to be easy, plus edge clips picked on purpose because they're hard: 140 typical, 60 edge, out of 200.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Setting the edge share once, at 8 in 100, when every customer was a solo hobbyist. It made sense before sponsor billing and legal review existed on the platform.
6 · THE NUMBER
Fill in the blank: real uploaded episodes were sampled and tagged by hand. About ___ in 100 actually had an edge-case feature, which is why 30 in 100 was picked as the golden set's edge share.
Tap to flip
ANSWER
About 26 in 100. The golden set sits a little above that on purpose, since it exists to catch trouble, not describe an average day.
7 · THE REPLAY
Same near miss, new golden set. What changes?
Tap to flip
ANSWER
The wider golden set catches the overlapping-crosstalk failure in a test run the next quarter, three days before a similar show would have published a similar misattributed quote.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's different about which assumption moves the split there?
Tap to flip
ANSWER
A predictive-maintenance alert system for commercial HVAC. There, it's not the dollar cost of a miss that moves the split most, it's who's on the other end of the failure, a homeowner versus a hospital.

Check yourself Score: 0 / 0

Short answer, the number question
1. If the real-world sample had found edge cases in only 10 in 100 uploads instead of 26, should the golden set's edge share still be about 30 in 100? Why or why not?
Show hint
Think about what a golden set is actually for: catching trouble on purpose, versus mirroring an average day.
Show answer
Yes, it should stay close to 30 in 100, maybe trimmed slightly, not dropped to 10. A golden set exists to test the model against trouble on purpose. Dropping it to match a low real-world rate would leave the model almost untested on exactly the cases that cost the most when missed.
Multiple choice
2. Why does BOUND fit a question about how a golden set should be split, rather than FLIPS?
  • A. Because golden sets are always a technical detail, not a product decision.
  • B. Because this is a sizing question, an arithmetic split between two piles, not a person's trust flipping between two settings.
  • C. Because BOUND and FLIPS always get combined on eval questions.
  • D. Because FLIPS only works for questions about call centers.
Show hint
Ask what FLIPS actually needs to work: a habit fading, then a switch that snaps between two settings.
Show answer
B. Mireille's story is about a ratio built wrong and revisited too late, not a habit quietly fading into a flip.
True or false
3. True or false: once a golden set's edge-case share is set, it should stay fixed even as the product's customers change.
  • True
  • False
Show hint
Think about what happened to Loomcast's original 8 percent once sponsor-billed shows joined the platform.
Show answer
False. The whole reversal in this answer is that the share should move with what a miss costs, and that changes every time the customer mix changes, sponsor billing and legal review raise the stakes.
Fill in the blank
4. The golden set in this answer has 200 clips total: ___ typical and ___ picked as edge cases on purpose, split across four categories worth 20, 15, 15, and 10 clips.
Show hint
Check the build-up chart's largest segment and its total.
Show answer
140 typical and 60 edge. 140 + 20 + 15 + 15 + 10 = 200.
Short answer, apply it yourself
5. Pick an AI feature you use yourself. What's one edge case for it, a situation it probably wasn't tested on much, and how costly would it be if the feature got that case wrong?
Show hint
Think of a situation you hit rarely but that would actually matter if the feature got it wrong: an unusual input, a person, a time of day.
Show answer
Model answer: "A voice assistant setting a timer: the edge case is a request said with background noise or a strong accent. Most days it barely matters. But if it's set for something like a medicine reminder, getting it wrong could cost more than an annoyance." Any answer works if it names a real edge case and a real cost.
Multiple choice
6. In this answer, why does the cost of missing an edge case move the golden set's split more than how often that edge case actually happens?
  • A. Because rare things are always more expensive to test.
  • B. Because a cheap, common miss can be ignored, but an expensive miss, like a sponsor billing dispute or a misquoted guest, is worth testing for even if it's rare.
  • C. Because frequency can't be measured, only cost can.
  • D. Because the golden set should always match real-world frequency exactly.
Show hint
Look back at stage 6: which matters more, how often something happens, or what happens when it does.
Show answer
B. A rare, cheap miss barely matters. A rare, expensive miss, like the misattributed quote almost published, is exactly what a golden set exists to catch.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more