ConceptAdvancedEval-Driven Specification / Golden datasets and test set ownership / #13

Describe the tradeoff between a broad golden set and a deep one.

The direct answer
Start broad, with two or three test patients on every trial, so nothing ships completely broken. Then, before any single trial opens for real patient contact, swap that trial's slice of the golden set for a deep one, built from real past screening records. A broad set only proves the tool works in general. It cannot prove the tool won't wave a real patient through a rule combination nobody ever tested.
How to build the golden set, in order
  1. Start broad on every trial, then go deep on whichever trial is about to touch a real patient.Why: a three-patient sample proves nothing's obviously broken. It never proves the tool won't wave through a rule combination it has never once seen.
  2. Build the deep set from real past screening records, not invented cases.Why: the rule combinations that hurt a patient are the ones that actually happen, and a golden set built from guesses misses exactly those, the same way a shallow one does.
  3. Let the broad floor do the one job it's actually good at: catching a trial that's completely broken.Why: a trial with zero matches or way too many is loud and fast to notice. Chasing rare rule combinations with a broad set asks it to do a job it was never built for.
  4. Say the cost gap out loud in real dollars and days, not in feelings.Why: "$8,400 and three weeks" moves a resourcing decision. "Let's be more careful with this one" doesn't.
  5. Recheck how many real rule combinations a trial's golden set actually covers once referrals start flowing, not just at launch.Why: a trial's true complexity only shows itself once real patients start arriving, not when someone first guesses at it.
  6. Leave the broad floor alone on any trial with only a handful of rules that don't interact.Why: three patients already cover almost everything a simple trial can throw at the tool, so going deep there burns hours nobody needed.

Seven moves for picking a golden set

A tradeoff question wants a pick, not a rubric, so PICK does the work here, not a checklist of eval fields.

1
Ground it in one real product before naming the framework
Say it like this
"Let's put this on Enrollwise, the tool a cancer center uses to check a patient's chart against the rules for forty open trials. I'm going to talk about the moment right before one of those trials opens for real patient contact."
Why this works
Stops the answer floating at "AI evals in general" and gives the interviewer something concrete to push on.
2
Name the four moves before making any of them
Say it like this
"I'm going to pick a position first, say who feels each kind of miss and in what units, name which kind is worse, then say what would change my mind. Four moves, in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Reframe what the question is really testing
Say it like this
"This isn't really 'broad or deep.' It's 'do you know a clean, evenly-spread test set can hide the one rule combination that only shows up once a real patient walks in.' So I'm not going to answer with a sample size. I'm going to answer with which mistake I'm protecting against."
Why this works
Shows the interviewer you see past the surface ask to the real judgment being tested.
4
Give the position, with the actual mechanism in it
Say it like this
"My position: build broad first, three patients on every trial, so nothing ships obviously broken. Then, before any trial opens for real patient contact, I replace that trial's slice with a deep set, built from real past screening records, covering the rule combinations that trial actually produces."
Why this works
PICK rewards committing to a real mechanism, not a vague promise to "test more."
5
Name who feels each kind of miss, then prove the asymmetry with a real case
Say it like this
"Here's the split. A trial that only has the broad set and breaks outright, say a wrong lab code that matches nobody, gets caught inside a day because the match count reads zero and someone notices. That costs about four hours of engineering time and touches no patient at all. A missed rule combination is different. We had a trial that needed a five-week gap after a specific class of drug, not the usual three. Our broad set never happened to include a patient recently on that drug, so nobody tested it. A real patient matched at 94 percent, went through three weeks of scans and blood work, and got told no on day twenty-one, after the sponsor's own review caught what our set never did."
Why this works
The real numbers and the real case make the asymmetry checkable, not just asserted.
6
Say what you'd track after launch, not just at launch
Say it like this
"I wouldn't call this done at launch. Every quarter, I'd recheck how many real rule combinations actually showed up in referrals against how many the deep set covers, because a trial's true complexity only shows itself once real patients start flowing through it."
Why this works
Shows you think past launch day, which is where most answers stop.
7
Name the kill criteria and close on the one line
Say it like this
"I'd drop the special deep pass for a trial the moment its own rules stop interacting, once it's down to a handful of criteria that don't combine into anything a golden set can miss. And I'd leave the broad floor exactly where it is on every trial like that, because getting one of those slightly wrong costs a few hours, not three weeks of a patient's time."
Why this works
Shows confidence without stubbornness, and ends on the line the interviewer should walk away remembering.

One more thing before the walkthrough moves on: this pick covers one trial at a time, not the whole golden set forever. Most candidates hear "test the risky one more" and answer by making everything bigger. Say which trial earns the deep pass and which ones stay on the broad floor, and you've shown real judgment instead of reciting a checklist.

Let's learn

Enrollwise reads a patient's chart at Halden Cancer Center, the diagnoses, the labs, the drugs they've already had, and checks it against the rules for forty open trials running through the center at once.

Before Enrollwise, a coordinator like Dominic Kettering read through a stack of new patient charts by hand every week, checking each one against whichever trial's rules seemed to fit, a job that took him most of a Thursday for maybe fifteen real candidates.

A patient record is checked against forty trial rulebooks, a coordinator reviews the match, and screening starts
One chart, forty rulebooks, one coordinator's morning
Knowledge spark: what's a washout period? The stretch of time a trial requires between a patient's last dose of an earlier drug and the day they start a new one. It exists so the old drug's effects don't mix with the new one's. Most drugs need about three weeks. Some need longer.

With Enrollwise checking all forty trials against every new chart overnight, Dominic gets a ranked list by morning, a match score next to each name. Most weeks that's about twelve strong candidates flagged by 8 a.m., work that used to take him most of a day.

To build Enrollwise, the team wrote a golden set: a handful of real test patients per trial, checked by hand, that Enrollwise has to score correctly before anything ships. Three patients a trial, forty trials, a hundred and twenty patients in total. It took about eighty hours to build, and it caught every trial that was flat-out broken before it ever reached Dominic's list.

Here is the turn. A hundred and twenty test patients across forty trials sounds like a lot. Spread three patients over one trial's actual rules, it's thin. The lung trial Dominic works with most, a combination-drug trial with fourteen rules that interact with each other, drug class, timing, organ function, biomarker status, has more real combinations of those rules than three patients could ever touch.

We proved Enrollwise wasn't obviously broken. We never proved it wouldn't wave a real patient through a combination none of the three test patients happened to have.

At its worst, that looks like this: a patient who took a certain drug three weeks ago, which is fine for most of Enrollwise's forty trials, but this one trial needs five weeks for that exact drug class. Enrollwise scores the match at 94 percent. Nobody built a test patient with that timing, so nobody ever saw the tool get it wrong, until an actual patient did.

The choice I would take back We used one golden-set size, three patients, for every trial, because that was the right amount back when only a couple of simple trials were running. I would take that back. The fourteen-rule lung trial needed a deep set built from real screening history, not the same three patients as a trial with two rules that never touch each other.
Cost per event, in dollars
Caught fast, costs the company
Rare, and it costs a patient weeks
Enrollwise flags itself: a broad-only trial breaks outright
$300
Enrollwise waves a patient through an untested rule
$8,400
The cheap bar is barely there on purpose. $300 covers about four hours of engineering time to fix a bad lab-code mapping, found the same day the trial's match count reads zero. $8,400 is what one real screen fail costs in scans, labs, travel, and the sponsor's own review, the shape of what Dominic's March patient went through. Neither number counts the twenty-one days taken from that patient's own shrinking window to find a trial that will actually take them.

What I would leave alone. The broad, three-patient set on the simple trials. Most of Enrollwise's forty trials have two or three rules that don't interact with anything. Three test patients already cover almost every real case those trials will ever see. Building a deep set there would burn hours to fix a problem that doesn't exist.

The lesson. We built the golden set to answer "does Enrollwise work." It answered that well. We never built a separate answer to "does Enrollwise work on the one trial where getting it wrong costs a patient three weeks," and only the second question had a real person's time riding on it.

The five weeks Enrollwise never tested

You don't need this to answer the question. Read it if you want to feel why the deep pass has to happen before a trial goes live, not after.

Dominic Kettering has coordinated cancer trials at Halden for six years. Hand him a chart and he can usually tell inside a minute which of the center's trials it might fit, just from the diagnosis line and the drug list.

Enrollwise arrived eighteen months ago. For most of that time it was just good. Every morning a ranked list waited for him, twelve or so strong matches, and he'd read the top few closely before calling a patient in. It was almost always right about the easy trials, the ones with two or three straightforward rules. He stopped double-checking those from scratch. Why would he. It kept being right.

First he stopped re-reading the full chart on a match above 90 percent. Then he stopped calling the pharmacy to confirm drug history himself, since Enrollwise already had it pulled in. By month ten, if the score cleared 90, he moved straight to booking the first screening visit.

Then, in March, a chart landed with a 94 percent match for the lung combination trial. Everything on the screen looked right, diagnosis, biomarker, no other trial pulling the patient away. Dominic called the patient the same afternoon.

Two kinds of wrong, not the same size: a broad-only trial breaking is caught the same day and costs hours, while a rare rule combination slipping through costs a real patient three weeks of screening
Same tool, two very different kinds of wrong

The patient had finished a course of an EGFR inhibitor exactly three weeks before the referral, well past the three-week gap Enrollwise had learned to expect from Halden's other trials. This trial needed five weeks for that specific drug class, a rule buried two lines down in a protocol most people skim. Enrollwise's golden set had never once included a patient recently on that drug, so the rule had never actually been tested, only assumed.

Screening started the next week. Two scans, three blood draws, two four-hour round trips from the patient's town to Halden. Nineteen days in, the paperwork went to the trial sponsor's own central review, the last check before enrollment. Two days after that, on day twenty-one, it came back: ineligible, washout violation. The patient's cancer hadn't paused for any of it.

We did not lose three weeks of a coordinator's time. We spent three weeks of a patient's.

I want to say the problem is that Enrollwise scored a 94. It's not really that. A rule that was never tested isn't wrong, exactly. It's just unknown, and Enrollwise reported its confidence in a number nobody had ever checked against a real case like this one.

Months before, when the golden set first got built, the team argued for maybe twenty minutes about test-patient count. Five per trial felt slow to build. Three felt fine, since the two trials running at the time barely had any rules that interacted at all. Three became the number for every trial the company would ever add, without anyone deciding that on purpose.

Isobel Krantz, the product manager who owns Enrollwise's release criteria, pulled the sponsor's screen-fail report the week after and found the drug-class rule buried in the protocol. She went back through two years of real referrals and built a deep set for that one trial: sixty patients, real cases, every washout combination the trial had actually produced. Run the same March chart against it now and Enrollwise flags the washout conflict itself, before Dominic ever picks up the phone. The patient's cancer keeps its three weeks.

One design hands Dominic a number he has no way to check. The other hands him the one flag that actually mattered.

And the thing I'd tell myself, if I could go back: we built a golden set to prove the tool wasn't broken. We never asked whether "not broken" and "safe for this one patient" were the same question. They weren't.

PICK, spelled out for a golden set

This is a tradeoff about where to spend eval hours, not a checklist for building any old test set, so PICK is the tool.

P, position. Build broad first, three patients on every trial, as the floor that catches an outright break. Then, before any trial opens for real patient contact, replace that trial's slice with a deep set built from real past screening records.
I, impact. A trial with only the broad set that breaks outright touches no patient and costs Enrollwise's own engineering team about four hours to fix, caught the same day the match count reads zero. A missed rule combination reaches an actual patient: about $8,400 in scans, labs, and review time, and twenty-one days off a shrinking window to find a trial that will actually take them.
C, cost asymmetry. A three-patient sample can only ever test the rule combinations someone happened to think of. The rarer and more interacting a trial's rules are, the more completely a broad set misses them, and the miss doesn't show up on a dashboard. It shows up three weeks later, on a real chart.
K, kill criteria. Drop the deep pass for a trial the moment its own rules stop interacting, once it's down to a handful of criteria nothing else touches. At that point three patients already cover nearly everything the trial can throw at the tool, and a deep set is eighty hours spent proving what the broad one already proved.
Knowledge spark: why not just build a bigger broad set everywhere instead? Because the rule combinations that matter are specific to each trial's own protocol. Six patients spread over forty trials still won't happen to include someone on the one drug this trial cares about. More patients everywhere is still shallow. A deep set means patients chosen because of that trial's own rules, not just more of the same guess.
Interacting rules per trial, against the kill line
Broad floor already covers it
Needs the deep pass
Kill line: 6 interacting rules
0 8 16 6-rule kill line 2 3 4 6 14 Nausea Maintenance Immuno solo Biomarker Lung combo
Most of Enrollwise's forty trials sit well under six interacting rules, where a three-patient broad set already covers nearly every real combination. The lung combination trial sits at fourteen, far past the line, which is exactly where a deep pass earns its eighty extra hours. The biomarker-stratified trial sits right on the line: watched, not yet deepened.

Run PICK again, at a city building department

PermitClear reads a submitted building permit, drawings, equipment lists, occupancy type, and checks it against the code sections for twenty-five permit categories: fences, decks, signs, reroofs, HVAC swaps, and further down the list, fire-suppression retrofits on existing high-rises. Same shape of question, a different kind of harm.

P. Gate any auto-approve recommendation on the fire-suppression retrofit category on a deep golden set built from the department's own past inspection records, covering the real code-section combinations that category produces. Keep the broad, two-application floor on every other category.
I. A false "needs a person to look" flag on a fence permit costs a homeowner a couple of days waiting on a reviewer, felt the same week. A missed code conflict on a fire-suppression retrofit doesn't get caught until a contractor is mid-install, or an inspector finds it after the building's already occupied: tens of thousands in rework, and in the worst case, a real gap in how fast people could get out.
C. Interacting code combinations on a complex retrofit are rare in any small, evenly-spread sample, and a broad set will always be dominated by the twenty simple categories that make up most applications. The miss doesn't cost anyone until the retrofit's already built.
K. Drop the deep pass for a category once it's down to a couple of code sections that don't combine with anything else, the same line Enrollwise uses.

What I would leave alone, at the permit counter The broad floor on fences, decks, and signs. Those permits have a couple of rules each that never touch one another. A reviewer catching a false flag there loses ten minutes, not a rework bill.

Swap the trigger and it still runs

  • Speed: Enrollwise takes thirty seconds to score a chart instead of two. Doesn't move the pick, because the pick is about which rule combinations get tested before a real patient's time gets spent, not how fast the tool answers.
  • Cost: building the deep set for one trial gets three times more expensive. Still doesn't flip it. A deep pass costing even $360,000 is still smaller than a handful of $8,400 screen fails stacking up over a year.
  • The model gets better: if Enrollwise's own recall on interacting rules genuinely closes the gap to the deep set's, per the kill criteria, drop the special pass. That's exactly the evidence that would let one golden set do the job.

Where people run it wrong

  • Building the deep set from invented edge cases instead of a trial's real screening history, testing combinations that sound scary but never actually happen.
  • Making every trial's golden set a little bigger, evenly, instead of concentrating the deep pass on the one trial about to touch a real patient, which burns hours without closing the actual gap.
  • Treating a high pooled match-accuracy number as proof a trial is ready, without ever checking whether the rare rule combinations were tested at all.

If you are asked this cold

Say the reframe out loud before you name a sample size. "Give me a second, I want to separate what each kind of miss actually costs before I pick how deep to test." That's true, it's already stage three of the walkthrough, and it buys you the time to find the real asymmetry instead of reciting a number.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's its hardest step?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a broad set's miss stays hidden until it reaches a real patient, while a broad set breaking outright gets caught the same day.
2 · THE PERSON
Who owns Enrollwise's golden-set decision, and what did she build after the screen fail?
Tap to flip
ANSWER
Isobel Krantz, the product manager who owns Enrollwise's release criteria. After the March screen fail, she pulled two years of real referrals and built a sixty-patient deep golden set for the lung combination trial, covering the washout combinations that trial actually produces.
3 · THE HABIT
What did Dominic Kettering stop doing once Enrollwise kept agreeing with him?
Tap to flip
ANSWER
Re-reading the full chart and calling the pharmacy to confirm drug history himself on any match scoring above 90 percent. By month ten he moved straight to booking the first screening visit.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A broad-only trial breaking outright: caught inside a day, about $300 in engineering time, no patient touched. A missed rule combination inside a trial: a real patient screened for three weeks before a screen fail, about $8,400 in scans, labs, and review time.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Build broad first, three patients a trial, and replace any trial's slice with a deep golden set, built from real screening records, before that trial opens for real patient contact.
6 · THE NUMBER
The lung trial's real washout rule needed a gap of ______ weeks for that specific drug class, not the three weeks Enrollwise had learned from every other trial.
Tap to flip
ANSWER
Five. The three-patient golden set never included anyone recently on that drug class, so the five-week rule was never actually tested, only assumed.
7 · THE KILL CRITERIA
What would make you drop the deep pass and go back to one broad golden set?
Tap to flip
ANSWER
The day a trial's own rules stop interacting, once it's down to a handful of criteria that don't combine into anything a broad set could miss. At that point three patients already cover nearly everything real.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the deep pass sit there?
Tap to flip
ANSWER
PermitClear, a city building department's permit-review tool. The deep pass sits on the fire-suppression retrofit category, not on the twenty simple categories like fences and decks.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these is the actual mechanism behind this answer's pick?
  • A. Build six test patients per trial everywhere instead of three.
  • B. Keep a broad, three-patient floor on every trial, and replace one trial's slice with a deep set built from real screening records before it opens for real patient contact.
  • C. Have a coordinator double-check every single match by hand, regardless of score.
  • D. Raise Enrollwise's overall match-accuracy bar from 90 to 99 percent.
Show hint
Three of these either don't touch the rule combination that mattered, or throw away the time Enrollwise was built to save.
Show answer
B. A spreads the same guesswork thinner, C undoes the whole point of the tool, and D barely touches a rule combination that shows up in maybe one referral out of many. Only B actually targets the miss that matters.
Fill in the blank
2. Fill in the blank: Enrollwise's broad golden set started at ______ test patients per trial across forty trials, about a hundred and twenty patients in total.
Show hint
The same number the flashcards call "the position" size.
Show answer
Three. Three patients per trial, forty trials, a hundred and twenty patients, about eighty hours to build. Enough to catch a trial that's flat-out broken, not enough to test one trial's real rule combinations.
True or false
3. True or false: once Enrollwise clears a high overall match-accuracy score across all forty trials, that alone means it's safe to trust on any one of them.
  • True
  • False
Show hint
Think about what the lung trial's overall accuracy said about the one washout rule that was never tested.
Show answer
False. The lung trial cleared its usual accuracy bar and still waved a real patient through an untested washout rule. Overall accuracy never said anything about that one combination, because the golden set behind it never included a case like it.
Multiple choice
4. Why couldn't the team just make the broad golden set bigger everywhere, instead of building a separate deep set for the lung trial?
  • A. More patients would take too long to review by hand.
  • B. Extra patients spread evenly across forty trials still won't happen to include someone on the one drug the lung trial specifically cares about; the miss is trial-specific, not fixed by a bigger random sample.
  • C. Coordinators refused to work with a larger golden set.
  • D. The trial sponsor required exactly three patients per trial by contract.
Show hint
Ask whether a bigger random sample fixes a miss that's specific to one trial's own rules.
Show answer
B. Six patients spread over forty trials is still a guess about which cases matter. The lung trial's washout rule only gets tested by choosing patients because of that trial's own protocol, which is exactly what a deep set does and a bigger broad set doesn't.
Short answer
5. If the lung trial's real washout rule needed a two-week gap instead of five, and Enrollwise's usual three-week default already covered it, would the deep golden set still be worth building for that specific rule? Walk through it.
Show hint
Think about whether the miss would even happen if the default already covered the real rule.
Show answer
Not for that specific rule. If the trial's real requirement, two weeks, is looser than the three-week default Enrollwise already assumes, a patient clearing the default automatically clears the real rule too, so a broad set would never actually get that one wrong. The deep pass earns its hours on rules stricter than the default, like the real five-week rule, not on ones looser than it. The lung trial still has thirteen other interacting rules though, so the deep set would still matter for those.
Short answer, apply it yourself
6. Pick a place in your own work where one quick check covers many different situations. Name the rare, specific case hiding inside it, and what you'd build to actually test that case.
Show hint
Look for the check that's fast because it only samples a couple of examples, and ask which real situation it's never actually been tried against.
Show answer
Model answer: "Our support macro suggests a reply based on ticket text, tested against three sample tickets per category across sixty categories. Buried in that is refund requests tied to a chargeback already filed with the card issuer, maybe one in two hundred refund tickets, where the standard reply promises a refund the company can't actually send. I'd pull real chargeback-adjacent tickets from the last year and build a dedicated test set just for that combination, instead of trusting the same three generic tickets every other category gets." Any answer works if you can name the rare case the quick check is currently blind to.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more