ConceptIntermediateEval-Driven Specification / Golden datasets and test set ownership / #20
Describe the review cadence for a golden dataset.
The direct answer
Give the golden set two triggers, not one. The moment the product ships something that changes what the model actually sees, like a new content type, review the set before the next scheduled date, not after it. Back that with a cheap weekly sample check and a full quarterly manual audit, in that order, because months of a stale set going unwatched are far harder to undo than a single missed example, which gets caught and fixed in a day.
The cadence, ranked by what breaks first if skipped
Trigger a golden-set review the moment the product ships something that changes its shape, before the next scheduled date.Why: a calendar-only cadence is blind to a feature launch, and that blindness is invisible until real content in the new format starts slipping through.
Run a cheap weekly check: pull 20 to 30 real flagged items and see how many don't resemble anything already in the set.Why: it catches slow drift on its own calendar, cheaply, before anyone commits to a full manual relabel that isn't needed yet.
Add real new examples to the golden set every month, pulled from whatever the weekly check flagged.Why: a set that only grows once a quarter can hide a whole season of new phrasing or a whole new content type underneath a score that still looks fine.
Run a full two-person manual re-audit of the entire golden set every quarter.Why: it's the only layer that catches categories nobody's flagged yet and retires examples that no longer look like real usage.
Leave the plain word-blocklist off this cadence entirely, updated the moment a new word gets reported.Why: it's deterministic, it doesn't grade against a golden set, and a calendar slot for it would just be noise.
How to answer this, stage by stage
Six moves, in the order I'd actually say them. This question tempts you to name one number, like "every quarter," and stop. Don't. There are two completely different reasons a golden set goes stale, and most answers only ever plan for one of them.
1
Put a real product and a real owner in the room
Say it like this
"Let me make this concrete. Say it's TinkerNest, an app where kids build shared story worlds, and there's a 600-example golden set one person owns and built by hand. Cadence questions fall apart fast when there's nobody named to actually run the check."
Why this works
An interviewer can't grade "regularly." Naming a product and an owner earns you the room to give a real answer.
2
Split the question into two different clocks
Say it like this
"There are two totally different reasons a golden set stops matching real usage. One is just time passing, slang shifts, kids move on to new games. The other is the product itself changing shape, shipping a content type the set was never built to grade. Those need two different triggers, not one calendar."
Why this works
This reframe is the whole answer. Most candidates only plan for the first clock and never mention the second.
3
Rank the harder failure first
Say it like this
"Everything in this cadence protects one thing: whether the golden set still reflects what kids are actually sending, not what they were sending a year ago. Given that, a single flagged message that slips through and gets caught is a five-minute fix. Months of the set quietly not covering something is the one you can't undo, because nobody can go back and re-check conversations that already happened."
Why this works
This is outcome and reversibility working together. It shows you're ranking by real damage, not by gut feeling.
4
Give the dependency, the one most answers miss
Say it like this
"If the product ships something like a new content type, that has to trigger a review before the next scheduled date, not queue behind it. A quarterly audit only means something if the set already covers what the product looks like right now. Wait for the calendar and you're grading last quarter's app."
Why this works
Shows real ordering logic instead of a wish list. This is the hardest letter in the framework for a reason.
5
Give the cheap check, then the actual ranked cadence
Say it like this
"Before anyone relabels 600 examples by hand, pull twenty or thirty real flagged items from this week and see how many don't look like anything already in the set. That's one afternoon, and it tells you whether the expensive review is even needed yet. So: a product-change trigger first, a cheap weekly check second, a monthly top-up third, and a full two-person audit every quarter, last, because it's the slowest layer."
Why this works
Shows the dependency between a cheap signal and an expensive review, which is exactly what real judgment looks like.
6
Close on the one line
Say it like this
"So: the golden set gets reviewed the moment the product changes shape, checked cheaply every week no matter what, topped up monthly, and fully re-audited every quarter. If you remember one line: a scheduled review only catches what time changed. Something else has to catch what you changed, the same week you changed it."
Why this works
Ends on the actual decision, in a sentence a reader could repeat back cold.
Let's learn
The safety set is a spreadsheet of 600 real examples, each one already marked by a person: let it through, block it, or send it to a human.
Say a kids' app called TinkerNest uses that set to grade an AI moderation model. Kids age seven to eleven build shared "story worlds" together: drawings with captions, and chat inside a shared project. Every message and caption runs past the model before it's visible to anyone else.
Knowledge spark: what a golden set actually is
A stack of real examples with a right answer a person already agreed on, used to grade the model instead of guessing whether it's doing well. TinkerNest's safety set is 600 real flagged messages and captions, each one already marked by a trained reviewer.
For a year, the model's pass rate against the safety set held steady at 97 percent, checked every time the model or its settings changed. Then TinkerNest shipped voice messages inside shared projects. Nothing about the safety set changed on launch day. Eight weeks later, a spot check found the model was only catching real bullying content sent by voice at about 1 in 12, against a rate of about 1 in 90 for typed chat, the number it had held for a year.
Bullying messages wrongly let through, by channel
The safety set's overall pass rate never moved off 97 percent the whole time. It was still grading the same 600 examples it had graded all year, and none of them were written in a voice-transcript format.
Here is the turn. The extra misses were never the real problem. A handful of messages a week, on a platform with thousands of kids, is easy to miss in a dashboard. The real problem is what nobody did. Nobody reviewed the safety set when voice chat shipped, because nothing about the model had changed, and "nothing changed" had quietly become the same thing as "nothing to check."
We didn't have a message that got through. We had eight weeks nobody had checked.
At its worst, this is not an embarrassing miss. A cruel message and a genuinely dangerous one, like a stranger asking a child to meet somewhere, look exactly the same to a model that's never seen a voice transcript before: both score as "unclear," both get waved through the same way. A kids' app that ships a new way to talk to other kids, with an eval that quietly can't see it yet, ends up worse than if it had never shipped the feature at all.
Each later layer only means something once the one before it can be trusted
The choice I would take back
When the safety set's review cadence was first set, in a four-person planning meeting, the question was how to fit "check the checker" into a small team's roadmap without it becoming its own job. The answer: one full audit a quarter, on the same date as the general product review. That made sense when TinkerNest was two content types and nothing changed between quarters. I would take it back, not by making the calendar faster, but by adding a second trigger tied to the product shipping something new, not to a date at all.
What I would leave alone. TinkerNest also runs a plain list of banned words, a simple filter with no model behind it. It doesn't grade against the golden set and it doesn't drift the way the model does. Someone reports a new word, it gets added that day. A calendar slot for it would just be more process around something that already works.
The lesson. A golden set that only gets checked on a date is quietly betting that nothing about the product will change between now and then. Kids don't wait for a quarterly review to start using a new feature differently than anyone expected.
Now here is the same thing as a story
You don't need this to answer the question. Read it slower, when you want to feel why the second trigger matters and not just recite that it exists.
Before TinkerNest, Marika Delacroix spent eight years as a middle-school counselor. She could read a kid's four-word text and know, before the second read, whether "I'm fine" needed a follow-up call. TinkerNest hired her to be the first person whose whole job was watching what its moderation model got wrong.
In her first three months, she built the safety set from nothing: 600 real flagged messages and captions, each one she or a second reviewer had marked by hand. Every Friday at 4, once the week's projects had mostly quieted down, she pulled 25 fresh real examples and checked them against the set herself, not just did the model score them right, but did the set still look like TinkerNest.
For months, that Friday hour was the best part of her week. The pass rate held at 97 percent, quarter after quarter. She'd close her laptop most Fridays having found nothing worth flagging.
So, around month four, she started skipping the odd busy Friday. By month eight, she'd mostly stopped pulling a fresh sample at all, opening the safety set only right before the quarterly review, and trusting the pass rate to say the whole story. By the time voice messages launched inside shared projects, she hadn't pulled a live sample outside the quarterly window in months. The set was still 600 examples. It had been the same 600 for a year.
The single message was never the harder one to undo
Eight weeks in, a parent pressed "flag for a grown-up" on a message inside a shared project called Space Explorers. Marika opened it that afternoon. A voice message, transcribed messily: "ur so dumb lol nobody wants u here go home." A kid who'd been asked to leave the group weeks earlier, back again under a friend's account. The model had scored it "playful." No escalation. Straight through, same as any joke between friends.
We didn't lose one message. We lost eight weeks of not looking.
Marika pulled a real sample of voice-flagged content going back to launch day, the thing she used to do every Friday without being asked. Out of the bullying-flagged voice messages in that window, the model had been catching about 1 in 12. Text chat, the whole year, had held at about 1 in 90. She never had a number in her head, if she's honest. She had a feeling with two settings: green, don't open it; not-green, open everything. Eight straight green quarters, and the feeling never once told her to look, because the pass rate was reading the same 600 examples it always had. It had nothing new to be wrong about.
The decision that made sense at the time got made in that same four-person planning meeting, back when the safety set was three months old. The question on the table was how to check the model without turning it into a second full-time job. The answer: one full audit a quarter, tied to the same date as the general product review. On a small team shipping fast, that was the right call.
I would take that back. Not the audit, not the set. Just the idea that a date was the only thing that could ever start a review. It isn't that quarterly was too slow, either. Monthly wouldn't have caught this: voice chat could have shipped the day after a monthly check and still sat unwatched for weeks. The fix isn't a faster dial. It's a second trigger that isn't a date at all.
Run the same eight weeks with that second trigger in place, and the moment voice chat ships, Marika pulls a live sample within days. She finds the format gap immediately: none of the safety set's bullying examples look like a messy transcript. Fifteen real voice examples go into the set before day ten. The Space Explorers message either gets caught the day it's sent, or worst case, the model's already been retuned on how bullying actually sounds out loud by the time it would have arrived.
One design hands Marika a calendar. The other hands her a tripwire that goes off the week the product changes, not the week the next quarterly invite goes out.
The thing I'd tell myself, looking back: I built a test that could tell me if I'd broken the model. I never built one that noticed when I'd changed the shape of the product out from under it.
ORDER, run as a calendar instead of a scorecard
This is a scheduling and ownership question wearing an eval question's clothes. TRACE would go hunting for a mystery, and there isn't one, the cause is plain once you look. LEAD would go looking for a new early metric, but the metric already exists, the question is only who watches it and when. What's actually being ranked is a set of review triggers by how badly it hurts to skip each one, so the framework is ORDER.
ORDER, run against a calendar instead of a roadmap
O, outcome. Every layer of this cadence protects one thing: the safety set still describes what kids on TinkerNest are actually sending each other this month, not what they sent when it was built. Not "the pass rate held." Whether a real risky message gets caught.
R, reversibility. One flagged voice message that slips through and gets caught by a parent within days is the easy one to undo: add it to the set, done by the weekend. Eight quiet weeks with zero voice examples in the set at all is the hard one: nobody can go back and re-check every voice message that already went out during that window.
Knowledge spark: what "format drift" means here
Not the model getting worse. The world the model sees changing shape underneath it. A safety set built on typed messages has never once seen a messy voice transcript, so it isn't wrong about bullying by voice, it's blind to it. That's a different problem than accuracy, and it needs a different check.
D, dependency. The quarterly full audit only means something if the safety set already reflects the product's current shape, so a real shipped change, like voice messages, has to trigger its own review before the next scheduled date, not queue behind it. That's forced by reality, not preference.
E, evidence. Pulling twenty or thirty real flagged items from the week after a launch and checking how many look nothing like anything already in the set costs one afternoon, and it would have shown the gap in week one instead of week eight.
R, rank. Product-change trigger first. A cheap weekly sample check second, running whether or not anything shipped. A monthly top-up of new real examples third. A full two-person manual audit of the whole set fourth, quarterly, because it's the slowest and most expensive layer and only needs to run once the cheaper layers already trust what they're looking at.
The check that makes ORDER honest
Swap what's actually at stake and the order moves. If a flagged voice message only ever sat in a private queue nobody read for a week, a missed format could go unnoticed for a month before it cost a real kid anything, and the event trigger could slide down the list. It isn't the eight weeks alone that earn the event trigger the top spot under O. It's that the message already reached another kid's chat before anyone but a parent caught it.
Run it where the flaw is a snagged thread, not a cruel sentence
A workwear mill runs a camera over every fabric panel before it's cut, scoring each one for weave slubs and stitch flaws, so a bad panel gets pulled before sixty pairs of trousers get cut around it.
O. Every check here protects one thing: a defective panel gets pulled before it's cut into finished pieces. Not "the camera score passed." Whether the flaw gets caught before the fabric is committed.
R. A panel already cut around a missed defect, sixty pairs of trousers deep into the wrong batch, is the hardest thing here to undo, you can't uncut fabric. Which camera model scores the panels, the mill can swap that any month.
D. The quarterly comparison against real store return rates means nothing until the review confirms the golden set of defect photos still covers the mill's current yarn, like the coarser weave a new supplier started shipping this spring. Reality forces that order, not preference.
E. A quality lead pulls thirty photos flagged this week and checks by eye how many don't resemble anything already in the golden set, one morning's work, and that alone would show a new weave type slipping through weeks before a returns report ever could.
R. Automated regression on every camera or model change first. A weekly cheap sample check by the quality lead second, independent of any change. A monthly top-up of new defect photos third. A quarterly full relabel of the golden set against real returns last.
Swap the trigger and it still runs
The mill installs a faster cutting line, and panels move through the camera twice as fast, blurring fine slubs. The order doesn't move. The event trigger fires on the equipment change, not the calendar.
A new yarn supplier is 20 percent cheaper but weaves slightly coarser. Same order. The event trigger fires the day the supplier switches, not at the next scheduled photo review.
A new camera catches finer defects than the old one ever could. Doesn't move the weekly check up or down. Better is a claim until real photos confirm it, and a missed flaw costs the same either way.
Where people run it wrong
Treating a golden-set score that's stayed green for a year as proof nothing changed, when the yarn, the camera, or the cutting line all moved underneath it.
Only reviewing the golden set when someone changes the model, so a supplier switch can go unrepresented for months with nobody watching.
Handing the weekly sample check to whoever's free that day instead of one named owner, so it quietly stops happening and nobody notices for a quarter.
If you're asked this cold
Say the outcome out loud before you name a single interval. "Every layer of this cadence protects one thing: a risky message reaches a person before it reaches another kid, even from a channel the safety set was never built to expect." Ten seconds, and every interval you name after that has a reason attached to it instead of sounding like a cron schedule.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits "describe the review cadence for a golden dataset," and why not TRACE or LEAD?
Tap to flip
ANSWER
ORDER, for ranking review triggers by what breaks first if skipped. TRACE hunts for a mystery cause, and there isn't one here. LEAD looks for a new early metric, but the number already exists, the question is only when different people look at it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marika Delacroix, TinkerNest's trust-and-safety lead, who built its 600-example safety set from scratch after eight years as a school counselor.
3 · THE HABIT
What did Marika stop doing because the safety set kept passing?
Tap to flip
ANSWER
She stopped pulling a fresh weekly sample of real flagged messages to check by hand. First she skipped busy Fridays, then she only opened the set right before the quarterly review.
4 · THE SWITCH
What's the two-setting switch in this story?
Tap to flip
ANSWER
Marika went from checking real flagged messages herself most weeks to trusting the quarterly calendar completely. No setting in between where she checked just the parts that had actually changed.
5 · THE OLD DECISION
What review-cadence decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Tying the safety set's only trigger to a fixed date, once a quarter, alongside the general product review. It made sense on a four-person team when TinkerNest was two content types and nothing changed between quarters.
6 · THE NUMBER
Miss rate on real bullying messages held around 1 in ___ for text chat. After voice launched, the same safety set caught real bullying voice messages at only about 1 in ___.
Tap to flip
ANSWER
1 in 90 for text; about 1 in 12 for voice, across the eight weeks nobody pulled a fresh sample. The safety set's pass rate stayed at 97 percent the entire time, because it was still grading the same 600 examples it always had.
7 · THE REPLAY
Same eight weeks, cadence fixed: what changes?
Tap to flip
ANSWER
The event trigger fires the week voice chat ships. Marika pulls a live sample within days, finds the format gap, and has fifteen real voice examples added to the set before day ten, instead of eight weeks after.
8 · THE TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what plays the role of the event trigger there?
Tap to flip
ANSWER
A workwear mill's fabric-defect camera. The equivalent is a quality lead pulling this week's flagged photos the day a new yarn supplier or a faster cutting line changes what a normal panel looks like.
Check yourself Score: 0 / 0
True or false
1. True or false: since TinkerNest's moderation model itself never changed when voice chat launched, waiting for the next quarterly safety-set review was the right call. Say why.
True
False
Show hint
Ask what changed even though the model's code didn't.
Show answer
False. The model didn't change, but the product did. Voice messages get transcribed differently than typed chat, and the safety set had zero examples in that format, so nothing about waiting for the calendar was ever going to catch it sooner.
Multiple choice
2. Which is harder to recover from: one bullying message that slips through and gets caught by a parent, or eight quiet weeks with no voice examples in the safety set at all?
A. They're equally easy to fix
B. The eight quiet weeks, because nobody can go back and re-check messages that already went out
C. The single message, because a parent already saw it
D. Neither matters once the safety set passes its quarterly check
Show hint
Think about what you can and can't undo once it's already happened.
Show answer
B. One flagged message gets added to the safety set in a day. Eight weeks of unwatched voice chat means real conversations already happened that nobody can go back and re-check.
Fill in the blank
3. Marika's safety set has ______ real, labeled examples, and it held a 97 percent pass rate for about a year before voice chat launched.
Show hint
It's the number she built from scratch in her first three months on the job.
Show answer
600. That number never changed in the eight weeks after voice chat shipped, which is exactly why the pass rate looked just as healthy as ever while real voice-flagged bullying was getting missed underneath it.
Short answer
4. Would switching from a quarterly safety-set review to a monthly one have caught the voice-chat gap sooner? Say why or why not.
Show hint
Ask what a monthly calendar still doesn't know about.
Show answer
Model answer: "Not reliably. Voice chat could ship the day after a monthly review and still sit unchecked for weeks. A faster calendar is still just a calendar. What actually catches it is a trigger tied to the product shipping something new, not a shorter gap between dates."
Short answer, apply it yourself
5. Pick an AI feature you use or are building. What's a product change to it that should trigger an off-cycle review of its golden set, instead of waiting for the next scheduled one?
Show hint
Look for a change that alters the shape or format of what the model sees, not just its accuracy.
Show answer
Model answer: "A grocery app's photo receipt scanner. The day it starts accepting handwritten farmers-market receipts instead of only printed ones, the golden set built on printed receipts needs a fresh review, not a wait for the next scheduled quarterly check."
Short answer, the number question
6. If the voice-chat miss rate had only slipped from 1 in 90 to 1 in 40, not 1 in 12, would the ranked cadence change? Say what moves and what doesn't.
Show hint
Reversibility is about what happens once it's caught, not about how far it had already slipped.
Show answer
Model answer: "The cadence stays the same. The event trigger still fires the week voice chat ships, owned by the same person, because a smaller drift is just as invisible to a calendar-only review. What changes is urgency: a smaller slip might not need pulling Marika off other work that same day, but it still needs to be caught by a real trigger, not stumbled into by a parent."
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.