ConceptAdvancedEval-Driven Specification / Golden datasets and test set ownership / #18

What is the argument for maintaining separate golden sets per customer segment?

The direct answer
Keep one shared golden set as the floor for every market, but split off a segmented golden set for any single market before its ads start spending real money there, built from that market's own flagged and rejected ads. A shared set only proves the copy is decent on average. It cannot prove a line that reads fine in one market won't break a rule that only exists in another.
How to split a golden set by market, in order
  1. Keep the shared golden set as the floor, and split off a segmented set for any market before it starts spending real money.Why: a shared set can't see a rule it was never built against, and that's exactly the gap a segmented set closes.
  2. Build each segmented set from that market's own flagged and rejected ads, not guesses about what might go wrong.Why: the claims that actually hurt a market are the ones that actually happen there, and an invented set misses exactly those.
  3. Let the shared set do the one job it's good at: catching a draft that's badly written or off-brief, in any market.Why: a weak headline is loud and shows up fast on its own. Chasing rare regional claims with the shared set asks it to do a job it was never built for.
  4. Say the cost gap out loud in dollars and days, not in feelings.Why: "$11,400 and a two-week delay" moves a resourcing decision. "Let's be careful with this one" doesn't.
  5. Recheck what a market's segmented set actually catches once real campaigns have run there a while, not just at launch.Why: a market's true divergence from the shared baseline only shows itself once real ads have run, not when someone first guesses at it.
  6. Leave the shared set alone on any market with only a couple of rules that don't interact with anything else.Why: a couple of simple rules is already covered by the shared floor, so building a segmented set there burns hours nobody needed.

The four moves for splitting a golden set

A tradeoff question wants a pick, not a checklist of eval fields, so PICK does the work here.

1
Ground it in one real product before naming the framework
Say it like this
"Let's put this on Fieldcopy, the tool a growth team uses to draft ad copy for three markets: the US, Mexico, and India. I want to talk about the point right before a new market's ads start spending real money."
Why this works
Stops the answer floating at "AI evals in general" and gives the interviewer something concrete to push on.
2
Preview the four moves before making any of them
Say it like this
"I'm going to pick a position first, say who feels each kind of miss and in what units, name which kind is worse, then say what evidence would change my mind. Four moves, in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Say what the question is actually testing
Say it like this
"This isn't really 'one golden set or three.' It's 'do you know that a good ad in one market can be an ad you're not allowed to run at all in another, and that's not just a translation problem.' So I'm not going to answer with a headcount. I'm going to answer with which mistake I'm protecting against."
Why this works
Shows the interviewer you see past the surface ask to the real judgment being tested.
4
Give the position, with the actual mechanism in it
Say it like this
"My position: keep one shared golden set as the floor, so nothing ships obviously bad anywhere. Then, before a market starts spending real money, I replace that market's slice with a segmented set, built from that market's own flagged and rejected ads, graded by someone who actually works in that market."
Why this works
PICK rewards committing to a real mechanism, not a vague promise to "localize more."
5
Name who feels each kind of miss, then prove the asymmetry with a real case
Say it like this
"Here's the split. A weak headline that clears the shared set gets caught fast: click-through runs under benchmark, the ad's paused inside two days, and it costs about $180 in wasted spend, nobody outside the team ever sees it. A missed regional rule is different. Fieldcopy once drafted a Mexico ad that said a camp stove 'cures the mountain cold in seconds.' Nothing in our shared set tested for a claim like that, because nobody writes ads that way in the US. It ran for six days before a compliance review caught it, and the whole Mexico push paused for two weeks while legal checked every line Fieldcopy had ever cleared there, about $11,400 all in."
Why this works
The real numbers and the real case make the asymmetry checkable, not just asserted.
6
Say what you'd track after launch, not just at launch
Say it like this
"I wouldn't call this done at launch. Every quarter I'd recheck how many of a market's real flagged ads its segmented set actually covers, because you don't know how different a market really is until real campaigns have run in it for a while."
Why this works
Shows you think past launch day, which is where most answers stop.
7
Name the kill criteria and close on the one line
Say it like this
"I'd drop a market's segmented set the moment it stops catching anything the shared set wouldn't already catch, usually once that market's down to a couple of rules that don't touch anything else. And I'd leave the shared set exactly where it is for a market like that, because getting it slightly wrong there costs a couple of days, not a paused campaign and a two-week delay."
Why this works
Shows confidence without stubbornness, and ends on the line the interviewer should walk away remembering.

One more thing before the walkthrough moves on: this pick covers one market at a time, not the whole golden set forever. Most candidates hear "test the risky one more" and answer by making everything bigger. Say which market earns the segmented pass and which ones stay on the shared floor, and you've shown real judgment instead of reciting a checklist.

Let's learn

Fieldcopy is a box on a marketer's screen: type in a product and a campaign brief, get back three headlines and two lines of body copy, ready to paste into an ad.

Before Fieldcopy, someone wrote every regional ad by hand, one market at a time: US English one week, Mexico Spanish the next, India English the week after, each one rewritten from scratch to fit that market's own tone, about six hours a week just on copy.

A brief goes into Fieldcopy, which drafts an ad and checks it against one shared golden set, and the same check clears US, Mexico, and India campaigns alike
One rulebook, checking three very different markets
Knowledge spark: what counts as a "claim" in an ad? Any line that promises something specific about what a product does. "Warm, durable, packs small" is description. "Cures the cold" is a claim. Most countries have rules about which claims you can make without proof behind them, and the rules are not the same rules everywhere.

With Fieldcopy drafting from the same brief, a marketer gets three headline options and two lines of body copy back in under a minute, for any market. Most weeks that cuts copy time from six hours down to about forty minutes.

To ship Fieldcopy, the team built a golden set: 150 real ads, graded by hand as "good enough to publish" or not, checked before any draft ships. It took about three weeks to build, almost all of it against US campaigns, since the US was the only market live when the set was built.

Here is the turn. The set answered "is this good ad copy" just fine, for a long stretch. But it only ever asked that question in one market's shape. When Mexico and India launched six months later, the same 150-ad set kept gating drafts in both, since building two more sets sounded like triple the review work for a two-person team.

We proved Fieldcopy could write a decent ad. We never proved it knew which decent ad was against the rules somewhere else.

At its worst, that looks like this: a Mexico ad says a camp stove "cures the mountain cold in seconds," a claim that reads fine to a US eye, because nothing in the shared set was ever graded for a claim shaped like that. Fieldcopy scores it clean. A marketer approves it in under a minute, the way she approves forty ads a week.

The choice I would take back We built one golden set early to save review hours, on the assumption "good ad copy" meant the same thing everywhere it ran. I would take that back. Before any market starts spending real money, I'd split off a segmented set, built from that market's own flagged and rejected ads, sitting on top of the shared floor.
Cost per miss, in dollars
Caught fast, costs the campaign
Rare, and it costs a whole market's push
A weak headline clears the shared set, CTR flags it
$180
A banned claim clears the shared set, runs six days
$11,400
The cheap bar is barely there on purpose. $180 covers about two days of wasted spend on a headline that just underperforms, caught the moment its click-through drops under benchmark. $11,400 is what one regional compliance miss costs in pulled spend, legal review, and a paused launch, the shape of what Mexico's camp stove ad actually cost. Neither number counts the two weeks the region's real push lost while every other live ad got rechecked by hand.

What I would leave alone. The shared set, for anything not yet spending real money in a market: internal drafts, new ad formats being tried out, and generic checks like grammar or a weak call to action that don't change by region. US and English Canada also stay on one set. They're close enough that a second one wouldn't catch anything new.

The lesson. We built the golden set to answer "is this good ad copy." It answered that well. We never built a separate answer to "is this ad copy legal and appropriate in the specific market about to run it," and only the second question could actually stop a campaign.

The claim nobody wrote for the US

You don't need this to answer the question. Read it if you want to feel why the segmented pass has to happen before a market goes live, not after.

The dashboard Esperanza Duval opens every Monday has three tabs, one per market, and for a long stretch of last year, all three said the same thing: green.

She's run growth marketing at Solstice Outfitters for four years, the only person on the team who watched all three markets grow from week one. Every Monday at nine she'd open the tabs, skim the week's drafts, tap approve on whatever cleared the check, and be back on her actual job, growing the channels themselves, by half past nine.

First she stopped reading every headline word for word once Fieldcopy kept getting them right. Then she stopped opening the Mexico and India tabs on their own, just scanning the combined pass rate across all three. By month nine, if the number was green, she moved on without opening a single ad.

Then, on a Thursday call about something else entirely, the Mexico operator said it in passing: "oh, is the cold-weather stove line okay to call it a cure? Our old agency always got twitchy about that word."

Two kinds of wrong, not the same size: a weak headline that slips through gets caught fast and costs a few hundred dollars, while a banned claim that slips through is found days later and costs a market's whole launch
Same tool, two very different ways to be wrong

Esperanza laughed it off for about four seconds. Then she pulled up the live campaign.

The ad had been running six days: "cura el frío de la montaña en segundos," a headline Fieldcopy had drafted, scored clean against the shared set, and she'd approved without a second look, the way she approved forty other ads that same week. Nothing in the 150-ad golden set had ever included a claim shaped like this one, so the claim had never actually been tested, only assumed fine because nothing like it had ever failed before.

We didn't just pull one ad. We paused Mexico's whole push for two weeks while legal checked every line Fieldcopy had ever cleared there.

I want to say the problem is that Fieldcopy scored the ad clean. It's not really that. A claim that was never tested isn't wrong, exactly. It's just unknown, and Fieldcopy reported its confidence in a number nobody had ever checked against a claim like this one.

Months before, when the golden set first got built, the team spent maybe fifteen minutes on the question of one set or three. Building three felt like triple the review work for two people already stretched thin. One set, one bar, felt clean. Nobody sat down and asked whether "good ad copy" meant the same thing in Guadalajara as it did in Ohio.

Esperanza pulled a year of Mexico's own flagged and rejected ads the week after and built a segmented set for that market: sixty real ads, including the kind of claim that had just cost them two weeks. Run the same brief against it now, and Fieldcopy flags the curative-sounding claim itself, before anyone ever taps approve. The ad never goes live. Mexico's push keeps its two weeks.

One design hands Esperanza a green light she has no way to double check. The other hands her the one flag that actually mattered.

And the thing I'd tell myself, if I could go back: we built a golden set to answer "is this a good ad." We never asked whether a good ad in one market could be an ad we're not allowed to run at all in another, and only the second question could actually stop a campaign.

PICK, run against one region

This is a tradeoff about where to spend eval hours, not a checklist for building any old test set, so PICK is the tool.

P, position. Keep one shared golden set as the floor, built to catch a draft that's badly written or off-brief in any market. Then, before any market starts spending real money, replace that market's slice with a segmented set, built from that market's own flagged and rejected ads.
I, impact. A weak headline that clears the shared set costs about $180 in wasted spend, caught inside two days once click-through drops under benchmark, and touches no one outside the team. A missed regional claim reaches real campaigns and real regulators: about $11,400 in pulled spend and legal review, plus two weeks off a market's real push while every other live ad gets rechecked.
C, cost asymmetry. A shared set can only ever test the claims someone happened to think of when they built it. The more a market's own rules differ from the market the set was built in, the more completely the shared set misses them, and the miss doesn't show up on a dashboard. It shows up days later, on a live campaign.
K, kill criteria. Drop the segmented set for a market the moment its own rules settle into a couple that don't interact with anything else. At that point the shared set already covers nearly everything the market can throw at Fieldcopy, and a segmented set is extra hours spent proving what the shared one already proved.
Knowledge spark: why not just build a bigger shared set everywhere instead? Because the claims that matter are specific to each market's own rules. Sixty more ads spread over three markets still won't happen to include the one claim Mexico's rules actually restrict. More ads everywhere is still a guess. A segmented set means ads chosen because of that market's own rules, not just more of the same guess.
Region-specific ad rules, against the kill line
Shared floor already covers it
Needs the segmented pass
Kill line: 3 region rules
0 4 8 3-rule kill line 0 1 2 5 7 US Canada UK Mexico India
US, Canada, and the UK sit under three region-specific rules, where the shared set already covers nearly every real case. Mexico and India sit well past the line, curative-sounding claims and unproven superlative claims respectively, which is exactly where a segmented pass earns its extra review hours. The UK sits closest to the line, watched, not yet segmented.

Run PICK again, at an insurance desk

ClaimSort reads an incoming insurance claim, the policy, the damage report, the photos, and drafts a settlement recommendation across every claim type a mid-size insurer handles: fender-benders, kitchen fires, and further down the list, flood damage on a total loss. Same shape of question, a different kind of harm.

P. Segment ClaimSort's golden set for catastrophic and flood claims before its recommendation touches a real payout, off the shared floor used for routine auto and homeowner claims.
I. A false "needs a person to look" flag on a routine fender-bender costs an adjuster maybe fifteen minutes of double-checking, felt the same day. An underpaid flood claim built on a golden set that never had enough real flood cases doesn't surface until a policyholder appeals or a regulator audits, months later: a corrected payout, and real penalty risk on top of it.
C. Catastrophic claim patterns are rare in any small, evenly-spread sample, and a shared set will always be dominated by the routine claims that make up most of the volume. The miss doesn't cost anyone until the payout's already sent, or worse, already disputed.
K. Drop the segmented pass for a claim type once it settles into a couple of rules that don't interact with anything else, the same line Fieldcopy's team uses.

What I would leave alone, at the claims desk The shared floor on routine auto claims. Those claims follow a couple of rules that never touch one another. An adjuster catching a false flag there loses fifteen minutes, not a corrected payout.

Swap the trigger and it still runs

  • Speed: Fieldcopy answers in five seconds instead of a minute. Doesn't move the pick, because the pick is about which market's rules get tested before real spend goes out, not how fast the draft arrives.
  • Cost: building India's segmented set costs three times more, given the extra language variants. Still doesn't flip it. Even at $18,000, that's smaller than a handful of $11,400 compliance pulls stacking up over a year.
  • The model gets better: if Fieldcopy's own flag recall on regional claims genuinely closes the gap to the segmented set's, per the kill criteria, fold that market back into the shared set. That's exactly the evidence that would let one golden set do the whole job.

Where people run it wrong

  • Building the segmented set from imagined edge cases instead of a market's real flagged and rejected ads, testing claims that sound scary but never actually happen.
  • Splitting every market evenly instead of concentrating the segmented pass on the one about to spend real money, which burns hours without closing the actual gap.
  • Treating a high shared-set pass rate as proof a market is ready, without ever checking whether that market's own rules were tested at all.

If you are asked this cold

Say the reframe out loud before you name a headcount. "Give me a second, I want to separate what each kind of miss actually costs before I say how many golden sets we need." That's true, it's already stage three of the walkthrough, and it buys you the time to find the real asymmetry instead of guessing a number.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's its hardest step?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a missed regional claim stays hidden until real money and real legal review are already spent, while a weak headline gets caught the same week.
2 · THE PERSON
Who owns Fieldcopy's golden-set decision, and what did she build after the Mexico ad?
Tap to flip
ANSWER
Esperanza Duval, the growth marketing PM who owns Fieldcopy's release bar. After the Mexico ad ran six days, she pulled a year of that market's flagged and rejected ads and built a sixty-ad segmented golden set for it.
3 · THE HABIT
What did Esperanza stop doing once Fieldcopy kept clearing every draft?
Tap to flip
ANSWER
Reading every headline word for word, then checking each market's tab on its own, then checking the check itself. By month nine, a green combined pass rate was enough for her to move on without opening a single ad.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A weak headline that clears the shared set: caught in two days by a falling click-through rate, about $180, no one outside the team sees it. A missed regional claim: a live ad running six days before compliance catches it, about $11,400, and a two-week delay to that market's real push.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Keep one shared golden set as the floor for every market, and replace a market's slice with a segmented set, built from that market's own flagged ads, before that market starts spending real money.
6 · THE NUMBER
The Mexico ad ran for ______ days before a compliance review caught the claim Fieldcopy's shared set had never once tested for.
Tap to flip
ANSWER
Six. The shared golden set was built almost entirely from US ads, so a claim shaped like "cures the cold" was never graded good or bad, only assumed fine because nothing like it had ever failed before.
7 · THE KILL CRITERIA
What would make you drop a market's segmented set and go back to one shared golden set?
Tap to flip
ANSWER
The day that market's own rules settle into a couple of rules that don't combine with anything else. Below that line, the shared set already covers nearly everything real.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the segmented pass sit there?
Tap to flip
ANSWER
ClaimSort, an insurance claims-triage tool. The segmented pass sits on catastrophic and flood claims, not on the routine auto and homeowner claims that make up most of the volume.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these is the actual mechanism behind this answer's pick?
  • A. Build one bigger shared golden set with more ads spread across all three markets.
  • B. Keep a shared golden set as the floor for every market, and replace one market's slice with a segmented set built from that market's own flagged ads before it starts spending real money.
  • C. Have a marketer personally re-read every ad by hand before it ships, regardless of score.
  • D. Raise Fieldcopy's overall pass rate from 90 to 99 percent across all markets.
Show hint
Three of these either don't touch the claim that actually broke, or throw away the time Fieldcopy was built to save.
Show answer
B. A spreads the same guesswork thinner, C undoes the whole point of the tool, and D barely touches a claim that shows up in maybe one ad out of hundreds. Only B actually targets the miss that mattered.
True or false
2. True or false: once Fieldcopy clears a high overall pass rate across all three markets, that alone means its copy is safe to run in any one of them.
  • True
  • False
Show hint
Think about what the overall pass rate said about the one claim that was never tested.
Show answer
False. Fieldcopy cleared its usual pass rate and still let a curative-sounding claim run in Mexico for six days. The overall rate never tested that claim, because the golden set behind it never had a case shaped like it.
Fill in the blank
3. Fill in the blank: Fieldcopy's shared golden set started at ______ real ads, graded by hand, built mostly from US campaigns.
Show hint
The same number the flashcards call "the position" size.
Show answer
150. Built over about three weeks, almost all of it against US campaigns, since the US was the only market live at the time. Enough to catch a badly written ad, not enough to test one market's own rules.
Multiple choice
4. Why couldn't the team just make the shared golden set bigger everywhere, instead of building a separate segmented set for Mexico?
  • A. More ads would take too long to review by hand.
  • B. Extra ads spread evenly across three markets still won't happen to include the one claim Mexico's own rules restrict; the miss is market-specific, not fixed by a bigger random sample.
  • C. The regional teams refused to work with a bigger golden set.
  • D. Solstice Outfitters' legal team required exactly 150 ads by policy.
Show hint
Ask whether a bigger random sample fixes a miss that's specific to one market's own rules.
Show answer
B. Sixty more ads spread over three markets is still a guess about what matters. Mexico's rule only gets tested by choosing ads because of that market's own regulations, which is exactly what a segmented set does and a bigger shared set doesn't.
Short answer
5. If the Mexico ad had been caught by the shared set on day one instead of day six, would building a separate segmented set for Mexico still be worth it? Walk through it.
Show hint
Think about whether one lucky catch means the underlying rule got tested, or just that one ad did.
Show answer
Yes, still worth it. Catching that one ad on day one only proves the shared set got lucky once, not that it now knows the rule. Every other untested claim in Mexico's rules would still be sitting there unchecked. The segmented set tests the rule itself, not just one ad written against it, so it's still the fix, even if this specific miss had cost nothing.
Short answer, apply it yourself
6. Pick a place in your own work where one quick check covers several different situations. Name the specific case hiding inside it that check has never actually been tried against, and what you'd build to test it.
Show hint
Look for the check that's fast because it only samples a couple of generic examples, and ask which real situation it's never actually seen.
Show answer
Model answer: "Our support macro drafts a reply from ticket text, tested against ten generic tickets across forty categories. Buried in that is a refund request where the buyer and the card holder are different people, maybe one in three hundred refund tickets, where the standard reply promises to refund a card we don't actually have on file. I'd pull real tickets shaped like that from the past year and build a dedicated test set just for that case, instead of trusting the same ten generic tickets every other category gets." Any answer works if you can name the rare case the quick check has never actually been tried against.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more