Artifact critiqueAdvancedAI Opportunity & Model Strategy / Model selection from a PM lens / #8
Describe a test set you would build to choose between two candidate models.
PICK96 percent looked safer on average, and 71 percent on the slice that actually mattered
Nordwave Stream runs a catalog of about 40,000 titles. Every new title gets tagged for genre, runtime metadata, a short recap, and content warnings, before it can go live. Damon Achterberg leads content operations and has to pick between two candidate models for that tagging job.
The direct answer
Don't build one big random sample and compare overall agreement. Build a stratified test set that deliberately over-samples the rare, high-cost content-warning cases and the titles where two human catalogers once disagreed, then score each model separately on that slice. Pick whichever model wins where being wrong is hidden and expensive, even if it loses slightly on the overall average.
Do this, in order
Score each model on the slice where an error is hidden and expensive, not on overall average agreement.Why: the overall number can look better while quietly hiding a much worse number underneath it.
Build the test set from real titles, over-sampling the rare content-warning cases on purpose.Why: random sampling barely tests the case that matters most, since it's rare by definition.
Include the titles where two human catalogers disagreed.Why: that's where a model's judgment actually gets tested, not just its memory of the easy titles.
Score each error type on its own, and say out loud which one you're optimizing against.Why: over-tagging annoys a viewer for a minute, missing a real warning is a hidden, expensive miss.
Set the kill criterion before you see either model's numbers.Why: separates a real evaluation from picking whichever result looks better after the fact.
Re-run the same test set on every later model version, not just at the first bake-off.Why: a vendor's "improved" update can quietly regress on exactly the slice you built the test set to protect.
How to answer this, stage by stage
Nobody is scoring whether you can name two model vendors. They're scoring whether the test set you describe would actually have caught the mistake that matters, not just the mistake that's easy to measure.
Stage 1
Ground it in the actual decision
Say it like this
"Let's ground this in one real choice. Nordwave needs to pick between two candidate models to tag 40,000 titles: genre, runtime, recap, and content warnings. That's the decision the test set has to settle."
Why this works
Keeps the answer from turning into an abstract lecture on evaluation methodology with no real stakes attached.
Stage 2
State your structure
Say it like this
"I'll run this as PICK. Position, my pick before any reasoning. Impact, who feels each kind of error. Cost asymmetry, which error is cheap and which is hidden and expensive. Kill criteria, what would flip the pick."
Why this works
Signals you can commit to a position instead of hedging with "it depends," which is exactly what this question is testing.
Stage 3
Commit to a position, right away
Say it like this
"My position: I'd pick whichever model catches the most real content warnings, even if its overall tagging accuracy is a couple points lower. A slightly wrong genre tag is a shrug. A missing content warning is the thing that ends up in a viewer complaint."
Why this works
This is the P step and the direct answer to the question, stated before any supporting evidence, not buried at the end.
Stage 4
Name who feels each kind of error
Say it like this
"An over-tagged title annoys one viewer, and the catalog team fixes it in a day once a report comes in. A missing content warning reaches every viewer who watches that title before anyone notices, and there's no report to fix, because nobody knew to look for it."
Why this works
Names both sides of the tradeoff in real units instead of leaving it as an abstract accuracy comparison.
Stage 5
Prove it with the number that flips the pick
Say it like this
"On the full 800-title random sample, Model A scores 96 percent overall, Model B scores 94. But slice out just the 150 titles that needed a rare content warning, and Model A only catches 71 percent of them. Model B catches 95. The 'better' model on average is the worse pick on the slice that actually matters."
Why this works
Compresses the whole case into the one comparison that would have changed the decision.
Stage 6
Close with the kill criterion
Say it like this
"My kill criterion, set before I looked at either model's numbers: if either model catches fewer than 90 percent of the rare-warning slice, it's disqualified regardless of overall score. Model A fails that bar. Model B passes it. That's the whole pick."
Why this works
A kill criterion decided in advance is what separates a real evaluation from rationalizing whichever number came out on top.
Let's learn
What happens the first time two candidate models disagree about whether a documentary needs a self-harm warning?
Nordwave's catalogers used to tag every title by hand: genre, runtime metadata, a short recap, and three content-warning categories, about 15 minutes a title.
Model A took that down to seconds, and its overall agreement with the human-labeled sample was 96 percent. Auto-publish turned on for anything it scored with high confidence.
The step that used to be a pause, checking before publish, had already been folded into one automatic step.
Here's the turn: the 4 percent Model A gets wrong isn't the real problem. The real problem is which 4 percent. Genre slips are cheap. A missed content warning is not the same kind of wrong as a missed runtime number.
An average agreement score treats every mistake as the same size. They are not the same size.
Overall agreement vs. the rare-warning slice, both models
Model AModel B
The model that wins on average is the one that misses more than one in four of the warnings that actually count.
At its worst, that gap sends a documentary about a real mass-casualty event out with no self-harm warning at all, to every viewer who plays it, with no report to catch it because nobody knew to look.
The choice I would take back
We tested both candidate models against one big random sample of last year's catalog and picked whichever one scored higher on overall agreement. Nobody asked which titles were actually in that sample, because the number looked clean and the decision felt done.
What I would leave alone: genre tags and runtime metadata don't need this kind of scrutiny. Any model's occasional slip there costs almost nothing, and building a special test slice for it would be effort spent where nothing is actually at risk.
The lesson: an average is only honest when every mistake underneath it costs the same. The moment one kind of mistake costs more than another, the average is hiding the real decision, not summarizing it.
Now here is the same thing as a story
The short version above is what you'd say in a vendor review. Read this one for what it felt like the week a peer platform's mistake changed what Damon tested before signing off on either model again.
The Nordwave catalog room gets quiet around 4pm, right before the week's new titles go live.
When Model A first shipped, catalogers reviewed every suggested tag before publish, all four fields, for the first six weeks. It kept getting the easy titles right, so review narrowed to just the ones the model flagged low-confidence. By month four, the auto-published titles weren't being spot-checked at all. The flagged ones the model asked about were always fine, so trusting the rest felt earned, not careless.
Content warnings disagree the least often of any field, and cost the most when they do. That combination is exactly what a random sample undercounts.
Then a peer streaming platform, not Nordwave, had a documentary about a real historical tragedy go out with no self-harm warning. A viewer's clip of the missing warning went widely shared before the peer platform pulled it. Damon read about it on a Tuesday morning and couldn't stop thinking about whether Nordwave's own Model A would have made the same call.
Knowledge spark: why does a strong overall score hide a rare, costly miss?
Rare cases contribute almost nothing to an average, by definition. A model that's excellent on the 96 percent of titles that never needed a warning can score high overall while being genuinely bad at the 4 percent that did. The average was never built to protect that slice.
He pulled the original bake-off data and re-sliced it by content-warning category instead of averaging across all fields. Model A's overall 96 percent had been carrying a 71 percent catch rate on exactly the 150 titles where a rare warning was actually needed. Model B, the "worse" candidate at 94 percent overall, caught 95 percent of those same titles.
We didn't pick the model that tags titles best. We picked the model that agrees with us most often, and those turned out to be two different questions.
Nordwave's own near-identical gap had been sitting in the original bake-off data the entire time.
When the original bake-off was designed, nobody suggested slicing the sample by field. "Let's just run both against last year's catalog and compare agreement," someone said in that planning meeting, and it sounded like the obvious, fair way to compare two models.
The real question was never which model agreed with the human labels more often overall. It was which model you could trust on the one field where being wrong doesn't announce itself.
None of these three showed up by accident in the random sample. Each one had to be deliberately built in.
Damon rebuilt the test set: the same 800 titles, but now stratified to guarantee 150 rare-warning cases and 60 cases where two human catalogers had once disagreed. Re-running both models against it surfaced the 71-versus-95 gap immediately. Nordwave switched the content-warning field to Model B, kept Model A for the cheaper genre and runtime fields, and found 43 back-catalog titles Model A had mistagged, fixing all of them before any viewer complaint ever arrived.
Viewer-reported missing-warning flags, month over month, Model A
This was climbing in Nordwave's own numbers before the peer platform's incident ever happened. Nobody had set a bar to notice it.
What I'd tell myself, the morning Damon read about the peer platform's viral clip: we picked a model on a coin that looked fair on average, and never checked which side it was weighted on.
PICK, mapped onto one model choiceNot a script for always picking the model with the lower headline score. PICK is what tells you which score to trust.
P
Position. The pick, before any reasoning.
Model B, for the content-warning field specifically, even though its overall agreement score is two points lower than Model A's.
Interviewers are testing whether you can commit. Stating the pick first, then defending it, is what separates a decision from a hedge.
I
Impact. Who feels each kind of error.
A wrong genre tag costs one viewer a moment of confusion, fixed in a day once reported. A missing content warning reaches every viewer of that title with no report to catch it.
Naming both sides in real units is what makes the tradeoff concrete instead of abstract.
C
Cost asymmetry. Which error is hidden and expensive.
Over-tagging is cheap and visible, a viewer reports it and it's fixed same day. A missed warning is hidden and expensive, it ships to everyone and nobody flags it until real harm happens.
This is the hardest step and the one the whole pick turns on: optimizing against the second kind of error, not the first.
K
Kill criteria. What would flip the pick.
Set before seeing either model's numbers: below 90 percent catch rate on the rare-warning slice disqualifies a model regardless of overall score.
A kill criterion decided in advance is what separates a real pick from rationalizing whichever number looked better afterward.
The two boxes are deliberately unequal in weight, because the two mistakes are not the same size.
The recap, one line per letter: position is Model B for the content-warning field, impact is a shrug for genre versus a silent miss for a real warning, cost asymmetry is optimizing against the hidden and expensive error, and kill criteria is the 90-percent floor on the rare-warning slice that Model A fails and Model B clears.
Four parts, and the two on the right are the ones a plain random sample would never have generated on its own.
And if you want to be sure it really works, try it somewhere elseSame four letters, a translation service instead of a streaming catalog. The stakes change shape, but the seam sits in the same place.
Anniken Solvang runs quality for Glossary Works, choosing between two candidate translation models for medical-device instruction manuals. Mapped onto PICK: position is picking whichever model scores highest on dosage-and-warning sentences specifically, even if its overall fluency score is lower. Impact: a slightly stiff sentence costs a reader a moment of confusion; a mistranslated dosage warning costs a patient real harm, with no built-in way for anyone downstream to catch it. Cost asymmetry: optimize against the hidden, expensive error, not the visible, cheap one. Kill criteria: any model scoring below 98 percent exact-match on the dosage-and-warning sentence slice is disqualified, regardless of its overall BLEU score.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "score on the slice where wrong is hidden and expensive, not on the average," and stop.
Cost: no time to build a full stratified test set before a launch date. Say so honestly, and commit to hand-picking even 50 known rare cases rather than skipping the slice entirely.
The model got better, for real: if a future version of Model A closes the gap and catches 97 percent of the rare-warning slice, that's the moment to re-run the same test set and let the pick change, not to assume the old gap still holds.
Where people run it wrong.
They compare two models on one overall score and stop, without ever asking which mistakes that score is averaging over.
They build the test set from whatever data was easiest to pull, instead of deliberately over-sampling the rare case the decision actually depends on.
They pick the model with the better score and never write down a kill criterion, so any inconvenient result afterward gets explained away instead of acted on.
How to use it live. The moment an interviewer asks how you'd choose between two models, ask yourself: which single mistake, if either model made it, would be the one nobody could undo quietly? Build the test set around finding that mistake, and the rest of the comparison follows.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits choosing between two candidate models with a test set?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It commits to a pick first, then finds which error the whole decision should turn on.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Damon Achterberg, content operations lead at Nordwave Stream, who re-sliced a bake-off after a peer platform's incident and found his own near-identical gap.
3 · THE HABIT
What did catalogers stop doing as Model A kept scoring well?
Tap to flip
ANSWER
They reviewed every tag at first, then only the low-confidence flags, then by month four stopped spot-checking the auto-published titles at all.
4 · THE ASYMMETRY
What's the cost asymmetry this pick turns on?
Tap to flip
ANSWER
Over-tagging is cheap and visible, fixed once reported. Missing a real content warning is hidden, ships to every viewer, and has no built-in way to get caught.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Testing both candidate models on one big random sample and picking whichever scored higher on overall agreement, without checking which titles were actually in that sample.
6 · THE NUMBER
Fill in the blank: on the rare-warning slice, Model A caught ___ percent and Model B caught ___ percent, even though Model A led on overall agreement.
Tap to flip
ANSWER
71 percent and 95 percent. The model that looked worse overall was the one that actually protected the slice that mattered.
7 · THE REPLAY
Same bake-off, stratified test set built from the start. What changes?
Tap to flip
ANSWER
The 71-versus-95 gap surfaces immediately. Nordwave switches the content-warning field to Model B and finds 43 back-catalog mistags before any viewer ever reports one.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product. Which one, and what stays the same?
Tap to flip
ANSWER
Glossary Works' medical-device translation choice. The field changes to dosage warnings, but the seam is identical: pick against the error that's hidden and expensive, not the one that's average and visible.
Check yourself Score: 0 / 0
True or false
1. True or false: this answer recommends picking Model A because it scored higher on the original 800-title random sample.
True
False
Show hint
Look at the direct answer and the grouped bar chart.
Show answer
False. Model A's 96 percent hid a 71 percent catch rate on the rare-warning slice, so the answer picks Model B for that field despite its lower overall score.
Multiple choice
2. Why does a random sample undercount the rare content-warning cases?
A. Random samples always exclude content-warning titles by design.
B. The rare cases make up a small share of the catalog, so a random draw naturally contains few of them.
C. Content-warning titles are removed from the catalog before sampling.
D. Both models refuse to process content-warning titles.
Show hint
Look at the knowledge spark about why a strong overall score can hide a rare miss.
Show answer
B. If only 150 of 40,000 titles need a rare warning, a random sample will barely include enough of them to measure performance on that slice at all.
Fill in the blank
3. Fill in the blank: after switching to Model B for content warnings, Damon's team found ___ back-catalog titles Model A had mistagged, before any viewer complaint arrived.
Show hint
Look at the paragraph right after the test set gets rebuilt.
Show answer
43 titles. All were fixed proactively, the exact outcome the original random-sample bake-off had no way to surface.
Short answer, where it wouldn't matter
4. Name a field in this schema where the choice between the two models barely matters, and say why.
Show hint
Look at "what I would leave alone" and the quadrant diagram.
Show answer
Model answer: Runtime metadata. It's rarely wrong and cheap when it is, so there's no hidden-and-expensive failure mode hiding underneath the overall accuracy number for that field.
Short answer, apply it yourself
5. Think of an app you use that has to make a judgment call sometimes. What's one rare, high-cost mistake it could make that a simple "how often is it right" score would never surface?
Show hint
Look for the mistake that's rare but expensive, not the one that's common but cheap.
Show answer
Model answer: A spam filter: its overall accuracy can look excellent while it rarely, quietly, sends one urgent message from a doctor's office to the spam folder, a mistake a plain accuracy score would never flag as urgent.
Short answer, work the number
6. If the kill criterion had been set at 80 percent instead of 90, would Model A still have been disqualified on the rare-warning slice?
Show hint
Compare Model A's 71 percent against an 80 percent bar instead of a 90 percent one.
Show answer
Model answer: Yes. 71 percent is still below an 80 percent bar, so Model A fails either way, though a much looser bar, say 65 percent, would have let it through and hidden the real gap again.
Before you close the answer
Why this works
Tests whether you'll build a test set around the mistake that actually costs something, instead of trusting whichever overall number happens to be higher.
Follow-up traps
"Isn't over-sampling rare cases just cherry-picking the test to favor one model?" Response: no, because the kill criterion was set before either model's numbers were known, and the slice reflects a real cost difference, not a result chosen after the fact.
"What if Model B is much more expensive to run?" Response: then that's a second, separate tradeoff worth naming, but it doesn't change which model is safe to ship on the field where a miss is hidden and expensive; it might mean using B only for that one field, which is exactly what Nordwave did.
If pressed
The 150-title rare-warning slice itself gets refreshed quarterly with newly added titles from genres prone to needing warnings, so the test set keeps testing against what the catalog is actually adding, not just what it looked like at the original bake-off.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.