ConceptAdvancedAI Opportunity & Model Strategy / Model selection from a PM lens / #17
Explain why a model that is better on average can be worse for your product.
TRACEthe campus average went up, twelve students' goals quietly stopped being exact
Here is what happens when a model gets measurably better and a product gets measurably worse, for exactly the students who could least afford it, while every number on the dashboard says the upgrade worked.
The direct answer
An average score can rise while one small, high-stakes segment quietly gets worse, because the improvement and the regression are hiding inside the same number. Before trusting any "better on average" claim, recut the score by the segment where a wrong answer actually costs something, not just by the overall mean. If you only ever look at the average, you'll ship the regression and call it a win.
Do this, in order
Recut every "improved" score by your highest-stakes segment before trusting it.Why: an average is a blend, and a small segment cratering can hide inside a rising overall mean.
Rule out instrumentation before blaming the model.Why: a scoring rubric change looks exactly like a real regression until you check.
Name three real cause candidates, not just "the model got worse."Why: alignment tuning, training-data shift, and prompt-template mismatch each need a different fix.
Run one evidence test that separates the top candidates.Why: the exact-same-prompt, both-models comparison is the single check that tells you where the fault actually lives.
Build the segment-level check into every future model comparison, not just this one.Why: the next "better on average" model can hide the same problem in a different segment.
How to answer this, stage by stage
The interviewer wants to know if you trust an average, or go looking for the segment it's hiding.
Stage 1
Scope it to one real system
Say it like this
"Let's ground this in one product. GoalScribe drafts IEP goals for a special-education case manager to review and sign. I'll answer for the week the district upgraded to a model that scored better on a general drafting-quality benchmark."
Why this works
Stops "better on average, worse for the product" from staying an abstract paradox with nobody's students in it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as TRACE. Timeline, when things actually shipped and moved. Recut, slicing the average by segment. Assume nothing, ruling out instrumentation first. Cause candidates, naming real hypotheses. Evidence test, the one check that separates them."
Why this works
Signals a repeatable diagnostic method, not a one-off guess about what went wrong.
Stage 3
Reframe: an average is a blend, not a single truth
Say it like this
"An average score doesn't mean everyone got a little better. It means the total went up. A large group improving a little and a small group collapsing can produce the exact same rising average."
Why this works
This is where a strong answer separates from one that just repeats "the benchmark said it was better."
Stage 4
Give the one decision: recut by segment before trusting the average
Say it like this
"I'd never trust an average alone. I'd recut the score by goal-complexity tier, routine goals versus the small number needing exact, legally specific language, and check each one separately."
Why this works
This is the direct answer, and it's the exact check that would have caught the real problem here.
Stage 5
Prove it with the compressed failure
Say it like this
"At Marrow Creek, the campus-wide goal-quality score rose from 78 to 85 after the upgrade. But for the twelve students needing precise, measurable AAC and trial-count language, the same score dropped from 81 to 64, and a compliance audit caught it ten weeks later."
Why this works
Compresses the whole paradox into the one number pair that shows the average was hiding a real regression.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't just a normal product regression is that a model upgrade changes behavior across the board at once, in ways a general benchmark was never built to catch at the segment level. Checking every segment on every model upgrade costs real time. I'd spend it anyway, because the alternative is finding out from an audit instead of a dashboard."
Why this works
Names the load-bearing AI-specific judgment and states the trade-off plainly.
Stage 7
Close on the one line
Say it like this
"So: an average went up, a small high-stakes segment went down, and only recutting the score would have shown both at once."
Why this works
Restates the direct answer in one breath, exactly what a live interview rewards.
Let's learn
GoalScribe listens to a case manager's notes from an IEP meeting and drafts the measurable goals that go into a student's legal education plan.
Before the upgrade, Grace Osei wrote or heavily rewrote most draft goals herself, about 20 minutes per student, across a caseload of around 40 students.
After the district upgraded to a newer model, campus-wide review time dropped, and the general quality score teachers rated drafts on climbed from 78 to 85 out of 100.
Here's the turn: the extra polish wasn't the story. For twelve students on Grace's caseload, the ones needing exact trial counts and specific AAC prompting-level language, the same drafts got quietly, measurably less precise, even as the campus average looked like a clean win.
The cohort that broke wasn't large enough to move the average much. It was exactly specific enough to need it to be right.
At its worst, a model upgrade that looks like a clean win on a dashboard quietly produces legally non-compliant goals for the smallest, most vulnerable slice of a caseload.
The choice I would take back
Marrow Creek's model-comparison process only ever measured one campus-wide average quality score, never broken out by goal-complexity tier. That made sense when nobody expected an upgrade to help most students while hurting a specific few. It stopped making sense the moment a rising average turned out to be hiding a real regression for the twelve students who needed the most precision.
What I would leave alone: the general quality rubric teachers use for routine goals didn't need to change at all. It was measuring the right thing for the vast majority of the caseload. The gap was only ever in the high-specificity slice.
The lesson: an average score answers "did most things get better." It doesn't answer "did the thing I can least afford to get wrong get worse." Ask both, every time.
Goal-quality score, campus average versus the high-specificity cohort
Before upgradeAfter, campus averageAfter, twelve-student cohort
Same upgrade, same week. One number went up seven points. The other went down seventeen.
Now here is the same thing as a story
The short version above is what you'd say in a data review. Read this one for the day the compliance audit actually landed on Grace's desk.
Every Thursday, Grace Osei printed her caseload's draft goals and walked them past each student's specific needs before a single one went into a legal document.
The model upgrade shipped quietly over a weekend, and the following Monday, GoalScribe's drafts read smoother, warmer, more "supportive" in tone, across almost every student on Grace's list.
Knowledge spark: what does "alignment tuning" actually change in a model's writing?
Alignment tuning nudges a model toward responses people rate as more helpful or agreeable in testing. That often means softer, more general, more reassuring language. For most writing, that's an improvement. For a legal document that needs an exact number of trials or a specific prompting level, softer and more general is exactly the wrong direction.
For ten weeks, nothing looked wrong. Teachers liked the drafts. The campus quality score climbed from 78 to 85. Nobody was checking goal-complexity tiers separately, because nobody had ever needed to before.
The real drift started in week two. Nobody looked closely enough to see it until week ten.
A routine compliance audit, the kind scheduled a year in advance, pulled ten random IEPs from students with the most complex needs. Three of them had goals missing an exact trial count or a specific AAC prompting level, both legally required, both present in every one of those same students' goals before the upgrade.
The average told us the drafts got better. It never once mentioned the twelve students for whom "better" meant "less exact."
Specificity score for the high-stakes cohort, ten weeks after the upgrade
The cohort's score crossed the district's own compliance floor three full weeks before the audit that finally noticed.
The team's first instinct was to blame the rubric. It hadn't changed. Then they suspected the new prompt hadn't been updated for the new model. It had, word for word, the same template that worked before.
Three real explanations, each pointing at a different fix. Only one test tells you which is true.
The team ran the exact, unchanged prompt template against both the old and new model, side by side, on a 40-case golden set of the district's highest-specificity goals. The old model held its precise phrasing. The new model, given the identical instructions, still drifted toward general, supportive language on the same cases.
Holding the prompt fixed and changing only the model isolates the one variable that actually moved.
When the upgrade was first approved, someone said, "the new model scores higher across the board on our quality rubric, this is a clear improvement." True, and also incomplete, since nobody had asked the rubric to look at the twelve students who needed something the rubric wasn't built to weigh separately.
The real question was never whether the new model was better. It was better for whom, and the campus average had no way to say.
What I'd tell myself, reading that audit finding: an average is a rumor about everyone. It's never a fact about any one student.
TRACE, run backward from an average that liedNot a lecture on statistics. The one recut that would have caught this before an audit did.
T
Timeline. When did things actually start moving?
The model shipped week 0. The campus average moved immediately. The real specificity drift for the twelve-student cohort started around week 2, invisible until the audit in week 10.
The regression started eight weeks before anyone noticed it.
R
Recut. Slice it by segment.
Recutting by goal-complexity tier showed the high-specificity cohort dropping from 81 to 64 while the campus-wide average rose from 78 to 85 in the same window.
The hardest step: the average alone would never have shown this.
A
Assume nothing. Rule out instrumentation first.
The team checked the scoring rubric and the prompt template before blaming the model. Neither had changed.
A tracking or process change can look exactly like a real regression until you rule it out.
C
Cause candidates. Three named hypotheses.
Alignment tuning toward softer tone, a training-data shift away from rare AAC and trial-count phrasing, or the old prompt template no longer being honored the same way.
Three real explanations, not a shrug of "the new model is different."
E
Evidence test. The one check that separates them.
Running the identical prompt on both models against the same 40-case golden set. Only the new model drifted, which pointed at the model itself, not the prompt template.
The strongest move in the whole framework: one test, one clear answer.
The recap, one line per letter: timeline is a two-week drift hidden for eight more weeks, recut is 85 campus-wide against 64 for the twelve-student cohort, assume nothing is ruling out the rubric and the prompt first, cause candidates is alignment tuning, data shift, or prompt mismatch, evidence test is the same-prompt, both-model comparison that pointed at the model itself.
And if you want to be sure it really works, try it somewhere elseSame five letters, a transit agency's fare-dispute tool instead of a school district. A different cohort at risk this time: paratransit riders needing precise accommodation language.
Dalisay Cruz handles fare-dispute resolution for Corrigan Transit Authority, using a model that drafts response letters to riders disputing a fare or a denied accommodation request. When Corrigan upgraded to a newer, better-scoring model, overall response-quality ratings from a general customer-service rubric rose. But drafts for paratransit riders disputing a specific accessibility accommodation, ones needing exact regulatory citations, quietly got vaguer, replaced with generic apology language that no longer cited the specific rule being invoked. Mapped onto TRACE: timeline is the upgrade shipping and the accommodation-specific drift starting within a week; recut is splitting response quality by dispute type, routine fare disputes versus accommodation disputes; assume nothing is checking whether the response template changed, which it hadn't; cause candidates are the same three, tone-alignment, data shift, or template mismatch; evidence test is running the identical accommodation-dispute template against both models on a fixed set of past disputes and comparing which one drops the regulatory citation.
The same four categories of language, in different words, are exactly what a transit agency's accommodation-dispute golden set would need too.
One test, three possible readings, and each one points to a different fix.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "recut the average by your highest-stakes segment before trusting it, then isolate the cause with a same-prompt test," and stop.
Cost: there's no time to build a full segment-level eval before the next upgrade ships. Say so honestly, and recut just the highest-stakes segment, not every possible slice.
The model really is better everywhere: if a recut confirms every segment improved, that's exactly the outcome the check was built to prove, not just to disprove.
Where people run it wrong.
They trust a rising average without ever asking what it's blending together.
They blame the model immediately, without first ruling out a rubric or prompt-template change.
They test only on easy, high-volume cases, missing the rare segment where the real cost lives.
How to use it live. When an interviewer asks why a better-on-average model can be worse for your product, ask yourself: what's the smallest, highest-stakes slice this average could be quietly hiding? Go recut the number by that slice first.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits explaining why a better-on-average model can be worse for a product?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. It finds which segment an average is quietly hiding.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Grace Osei, a special-education case manager at Marrow Creek Unified School District, managing a caseload of about 40 students.
3 · THE RECUT
What did slicing the score by segment reveal?
Tap to flip
ANSWER
The campus average rose from 78 to 85, while the twelve-student high-specificity cohort's score dropped from 81 to 64 in the same window.
4 · WHY NO MIDDLE READING
Why couldn't the team just assume the rubric was the problem, without testing?
Tap to flip
ANSWER
Because a rubric or prompt change can look exactly like a real model regression on the surface. Only checking each directly rules them in or out.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Measuring only one campus-wide average quality score, never broken out by goal-complexity tier, so the twelve-student regression had nowhere to show up.
6 · THE NUMBER
Fill in the blank: the campus average rose from 78 to ___, while the high-specificity cohort fell from 81 to ___.
Tap to flip
ANSWER
85 campus-wide, 64 for the twelve-student cohort, in the same ten-week window.
7 · THE REPLAY
Same upgrade, segment-level recut already in place. What changes?
Tap to flip
ANSWER
The 81-to-64 drop for the twelve-student cohort shows up in week two or three, not week ten, letting the team fix or flag it before a compliance audit ever needs to.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the equivalent at-risk cohort?
Tap to flip
ANSWER
Corrigan Transit Authority's fare-dispute drafting tool. The equivalent at-risk cohort is paratransit riders disputing a specific accessibility accommodation.
Check yourself Score: 0 / 0
True or false
1. True or false: the campus-wide quality score and the twelve-student cohort's score moved in the same direction after the upgrade.
True
False
Show hint
Look at the grouped bar chart comparing the two scores.
Show answer
False. The campus average rose from 78 to 85 while the twelve-student cohort's score dropped from 81 to 64, in the same window.
Multiple choice
2. What was the first step the team took after noticing the audit's finding, before blaming the model?
A. They immediately rolled back to the old model.
B. They ruled out a rubric or prompt-template change first.
C. They fired the case manager responsible for the flagged goals.
D. They stopped using AI drafting entirely for all students.
Show hint
Look at the "assume nothing" step.
Show answer
B. Checking the rubric and prompt template first rules out instrumentation before attributing the drop to the model itself.
Fill in the blank
3. Fill in the blank: the evidence test ran the identical prompt on both models against a ___-case golden set.
Show hint
Look at the labeled parts diagram and the evidence-test flow.
Show answer
40 cases. Enough to confirm only the new model drifted toward vaguer language on identical instructions.
Short answer, where it wouldn't matter
4. Name a part of Grace's caseload where this exact regression would not have shown up.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Routine goals for students not requiring highly specific measurable language. The general rubric was already measuring the right thing for that majority of the caseload.
Short answer, apply it yourself
5. Think of a tool or app you use that got a "better" update. Is there a specific way you use it where the update might have quietly made things worse?
Show hint
Think of a small, specific use case that's different from how most people use the same tool.
Show answer
Model answer: A spell-checker update that got better at catching common typos overall, but started "correcting" a specific technical term or name that isn't in its dictionary, for the small group of people who actually use that term.
Short answer, work the number
6. If the cohort at risk had been 120 students instead of 12, would the campus average still have risen from 78 to 85?
Show hint
Think about how much weight a larger cohort would carry inside the same overall average.
Show answer
Model answer: Probably not by that much. A larger cohort dropping from 81 to 64 would drag the campus average down too, which is exactly why small, high-stakes segments are the ones most likely to hide inside a rising average unnoticed.
Before you close the answer
Why this works
Tests whether you trust an aggregate metric or go looking for the segment it's hiding. Most candidates stop at the dashboard's headline number.
Follow-up traps
"What if the cohort gap is real but small, does it really matter?" Response: size it against what's actually at stake, a legally required, measurable IEP goal for twelve students is a real compliance risk regardless of how small the group is next to the whole caseload.
"Isn't recutting by segment just slicing until you find a story that fits?" Response: no, because the slice, goal-complexity tier, was chosen for a reason that predates seeing the data, not searched for after the fact to manufacture a finding.
If pressed
The evidence test held the golden set's 40 cases out of any future model's training or fine-tuning data specifically, since a golden set that a vendor could see and optimize against stops being a real test of what actually changed.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.