How do you prevent the golden set from becoming a proxy for a single loud customer?
- Tag every golden-set example to the account it came from, and cap any one account's share of the set.Why: this is the fix everything else supports. Skip it and the loudest account still gets to write the test.
- Set the cap low enough to catch drift in weeks, not the better part of a year.Why: an eight percent cap flags a runaway account fast. A generous cap just delays the same problem.
- When an account is about to cross the cap, stop taking their examples and force a fresh pull from other accounts and segments.Why: this is what actually restores the broad picture, not just freezes the skew where it happened to stop.
- Watch how well the tool does for the accounts that never complain, not just the golden set's own pass rate.Why: the set's score can hold steady for a year while whole customers it never learned from get a worse answer every time.
- Leave the checks that don't depend on judgment alone.Why: a check that just confirms a merge field filled in correctly doesn't care which account taught the model anything. Re-auditing it wastes review time.
- Don't wait for a different customer's complaint to reveal the skew.Why: by the time it shows up as someone else's bad ticket, the pattern has usually been shipping in real replies for months.
How to answer this, stage by stage
Seven moves. Most of the weight sits in stage three: this question is really asking what happens the day the easiest evidence to act on and the most important evidence to act on stop being the same thing. Every stage has the actual words to say.
Let's learn
Picture a support tool that reads an incoming customer ticket and hands the agent a reply already drafted, ready to send or edit. Before it existed, an agent wrote every reply by hand, and a tricky ticket could eat ten minutes just finding the right words. With a draft sitting there, most tickets close in under two minutes.
For a long stretch, the golden set's own score barely moved. It sat near 97 out of 100, quarter after quarter. Nobody had a reason to look past that number.
Here's the turn. There's an easier way for a new ticket to land in that set than a fair pull across every customer: whichever customer complains the loudest, with the clearest proof, gets folded in first, because that's the path of least resistance. Slowly, without anyone deciding it on purpose, the set stops testing "does this work for everyone" and starts testing "does this work for the one customer who keeps emailing about it."
Four hundred tickets make up this golden set, pulled at the start from across the whole customer base: retailers, clinics, software companies, insurance offices. When one retail account's tickets started filling most of the new slots, nothing in the set was tagged well enough to say so.
At its worst, this costs more than never building a golden set at all. A score that never drops tells the whole team to stop worrying, right as the tool gets worse for every customer who never complains, because they don't know who to complain to, or don't have the time.
What I would leave alone. A check that just confirms the drafted reply pulled in the right order number or the right return window doesn't need a re-audit when one account's share grows. It's graded against a fact, not a judgment call, so one customer's habits can't quietly bend it.
The lesson. A golden set doesn't fail because someone lied to it. It fails because it's a photo of whoever showed up loudest that quarter, and nobody wrote down that a photo is all it ever was.
Now here is the same thing as a story
Pull this one out when there's time to sit with it, not just tick it off a list.
Ligaya Domingo can read a support ticket and tell, before she's finished the first line, whether the reply drafted underneath it is actually right or just sounds right. Four years ago she answered Ticketframe's own support tickets by hand. Now she runs quality for the reply-suggester every agent at every Ticketframe customer sees.
The golden set she built is four hundred tickets wide, pulled from across Ticketframe's whole customer base: software companies, clinics, retailers, insurance offices, anyone who runs their support desk through Ticketframe. For her first two years on the job, she blocked out a full week every quarter. She'd request a handful of fresh, confirmed tickets from a rotating slice of accounts, big and small, and fold them in, so the set kept looking like the whole customer base instead of any one part of it.
Then Marrowstone Outdoor Supply signed on. Marrowstone runs a busy retail support desk, mostly returns and shipping questions, and their support ops lead was sharp: every time the reply-suggester got something wrong, he'd send Ligaya the ticket, the bad draft, and the correct reply, already lined up and ready to use. No chasing an account manager, no redacting, no waiting. Just done.
At first, Ligaya still ran her quarterly pull, but she'd start with whatever Marrowstone had sent that quarter, since it was already sitting in her inbox, and work down toward her target from there. A quarter later, she'd skip the accounts that needed extra coordination, the ones with smaller support teams, since Marrowstone alone covered most of the number she needed anyway. By the third quarter she couldn't have told you the exact day the pull stopped happening on the calendar. She just noticed, one week, that she'd snoozed the reminder for the fourth time, and that time she deleted it instead.
From then on, whenever Marrowstone's support ops lead sent a clean, verified example, she added it. That became the whole practice.
Eleven months in, a customer named Solace Behavioral Health filed a complaint. A patient had asked about a refund on a therapy session still waiting on insurance pre-authorization, and the reply-suggester told the agent to promise a refund in three to five business days. Solace can't promise that. Committing to a refund timeline before the insurer rules is against their own compliance policy, and the agent, trusting the draft, had sent it anyway.
Ligaya pulled the golden set apart to find out how that pattern got in there. It hadn't, not directly. What had happened was quieter. Eleven months of near-total real estate in the golden set had gone to Marrowstone: three percent of the set at the start, forty one percent by the time Solace complained. Marrowstone's tickets are almost all returns and shipping, so the model had learned, very well, that a fast promised timeline is the safe, helpful answer. It just never learned that some customers can't make that promise at all.
I want to say the model got worse. It didn't, not once, not by one setting. What actually happened was quieter than that. Ligaya's only real defense against one account taking over the set was a habit, running the quarterly pull, and a habit only has two settings. She was either doing it or she wasn't. Eleven months of an easier, faster substitute pushed her into not doing it, and being right, fast, and helpful eleven months running is exactly what keeps anyone from going back to the harder way.
Here's the call I'd take back. When the golden set's intake rule got written, in a planning meeting in Ticketframe's first year, the whole rule was two lines: a real ticket, and proof the correct reply was right. Nobody added a limit on how many could come from one account. At the time, no account had ever sent more than two or three examples in a quarter, so a cap felt like solving a problem that didn't exist yet.
I'd put a cap on it. No account gets to be more than eight percent of the golden set. The moment an account is about to cross that line, its examples stop going in automatically, and a fresh, scheduled pull from other accounts and tiers has to happen first.
Run the same eleven months again under that rule. Marrowstone crosses eight percent in month three, not month eleven. The cap flags it, and Ligaya runs a two day pull across twenty other accounts to backfill the set, including a batch from Solace. In that pull, a Solace ticket about a pending insurer decision shows up, and the reply-suggester gets it wrong in the same safe, obvious way it always does with anything it hasn't seen enough of. Ligaya catches it eight months before the real complaint, in a spreadsheet, instead of after a patient got a promise nobody could keep.
If I'm honest with myself, the mistake wasn't trusting Marrowstone's examples. Every one of them was real and correct. The mistake was writing a rule that only asked whether an example was true, and never asked how many true examples from one place a good test can absorb before it stops being a test of anything but that one place.
Where the set actually narrowed, step by step
A golden set doesn't need a bug to stop working. It just needs to keep growing from the one place growth is easiest. Here's the same five letters, mapped onto Ligaya's queue.
"The set drifted toward one customer" is a diagnosis anyone can offer after the fact. The harder part is naming the exact habit that had to stop first, the quarterly pull, and showing there was no smaller version of it left once it did.
And if you want to be sure it really works, try it somewhere else
ParcelDesk reads a permit application for a city planning office and flags anything that fails code, feeding a golden set of confirmed-clean, correctly-flagged applications. Same question, a permit office instead of a helpdesk, and a flip that isn't scope this time.
F. Sunniva Aasgard, senior plan-review analyst, who feeds ParcelDesk's golden set of confirmed-clean applications.
L. For two years she routed every new candidate through a rotation of four junior reviewers, one per specialty, building, electrical, plumbing, zoning, so the set stayed broad across permit types and filers. She stopped, one filer at a time, once deadline pressure meant she started clearing that filer's applications herself.
I. A different flip. She doesn't narrow which accounts get sampled, she takes the task back from the junior rotation for one filer, then never hands it back. Routes new examples through the rotation, or personally fast-tracks and adds them herself, nothing shared in between.
P. City council gave large commercial projects a mandatory ten-day review deadline. Nobody built an exception for how new golden-set examples got sourced when a project like that came in, so the fastest path, Sunniva handling it herself, became the only path for one filer's queue.
S. No filer's applications enter the golden set without also routing through the junior rotation that quarter, deadline or not, even if it means Sunniva does a first pass and a junior reviewer signs off after the fact instead of before. The takeover of the queue gets caught in six weeks instead of about a year, because the rotation log itself shows the imbalance.
Swap the trigger and it still runs
- Speed: if Marrowstone had onboarded onto Ticketframe in a single month instead of ramping for nearly a year, the skew would have shown up in the very first quarterly pull, with far less of the golden set already replaced by their tickets.
- Cost: if running the quarterly pull cost real money each time, say a paid redaction service per account, someone would have started skipping the harder accounts even sooner, and Marrowstone's free, ready-made tickets would have taken over faster, not slower.
- The model got better: this is close to what actually happened. The reply-suggester got very good at Marrowstone's kind of ticket. Nothing about the model was wrong. What was wrong was assuming "very good at one thing" and "very good, period" were the same finding.
Where people run it wrong
- Blaming the model for the Solace complaint, when the model never touched anything outside what the golden set taught it.
- Reaching for a bigger golden set as the fix, when tagging the existing four hundred by account and capping any one account's share would fix it for far less work.
- Waiting for a different customer's complaint to reveal the skew, instead of watching the account-share number directly.
How to use it live
Buy yourself a few seconds by naming the reframe before the fix: "The question isn't whether Marrowstone's tickets were good, they were. It's whether one good, generous customer can quietly become the whole test." Say that, and the rest of the answer is just the mechanism.
Flashcards (click a card to flip it)
Eight fixed slots, pulled straight from the answer above.
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Golden datasets and test set ownership
- #1 What is a golden dataset and why does the PM usually own it?
- #2 How do you construct a first golden set with no production traffic?
- #3 Describe the composition of a golden set: what proportion should be edge cases?
- #4 How do you keep a golden set representative as your user base changes?
- #5 Explain the risk of a golden set that engineering can see during development.
- #6 What is a holdout set and when would you use one for an AI product?