CaseAdvancedAI Opportunity & Model Strategy / Model selection from a PM lens / #21
How do you handle a model that is best for one customer segment and worst for another?
PICKthey rationed their own attention, and gave it to the wrong accounts
Auberdine Foodservice Supply delivers food and paper goods to restaurants and cafeterias. StockSense is the model that predicts how much each account will order next week, from a chain of forty locations down to a single independent diner. Marisol Etxeberria is the AI PM who has to decide what to do with a model that is excellent for one kind of account and shaky for another.
The direct answer
Don't run one setting for every account. Split the decision by segment: let StockSense run on its own for large chain accounts, where it's already accurate, and require a rep to check its forecast before ordering for small independent accounts, where it's weakest and a bad call costs the whole relationship, not just a shipping fee. Route by where the error actually lands, not by which accounts are easiest to automate.
Do this, in order
Split the model's role by segment, don't apply one blanket setting to every account.Why: an average accuracy number hides that the two segments fail in completely different ways.
Require a human check on the segment where errors are hidden and expensive, not the one that's loudest.Why: a chain's stockout gets a phone call. An independent's just quietly stops reordering.
Measure cost in the same unit for both segments before deciding where to spend review time.Why: dollars lost, not just accuracy percent, is what actually tells you which error to optimize against.
Set a kill criteria for when a segment no longer needs the human check.Why: as independent accounts build order history, the model gets better there too, and the rule should say so.
Watch churn by segment, not just forecast accuracy by segment.Why: churn is the real cost of a bad forecast, and it moves weeks after accuracy already told you something was wrong.
Say plainly that this isn't a fix for the model, it's a fix for how much you trust it, and where.Why: shows a real decision was made, not a promise to "improve accuracy" someday.
How to answer this, stage by stage
Nobody is scoring whether you know that segments exist. They're scoring whether you can name which segment's errors are actually the expensive ones, and commit to a pick before you explain why.
Stage 1
Say your position first, before the reasoning
Say it like this
"My pick: route by segment. StockSense runs alone for large chain accounts. A rep checks its forecast before ordering for small independent accounts. I won't average the two into one setting."
Why this works
Commits to a real position immediately, instead of describing the tradeoff and letting the interviewer wonder what you'd actually do.
Stage 2
Name who feels each kind of error, and in what units
Say it like this
"A chain's under-forecast costs about fifty dollars in an expedited shipment, resolved the same day. An independent's bad forecast costs spoiled food, or worse, a stockout that makes the owner quietly switch suppliers, worth about fourteen thousand dollars a year in lost business."
Why this works
Grounds the tradeoff in real units instead of a vague sense that "small accounts matter too."
Stage 3
Name the cost asymmetry, and optimize against the hidden one
Say it like this
"The chain's error is loud and cheap, someone calls, it gets fixed same day. The independent's error is quiet and expensive, the owner just stops reordering and you find out at renewal. I'm optimizing against the quiet one, since it's the one that actually loses us money without ever showing up as a complaint."
Why this works
This is the heart of PICK: naming which error is hidden and expensive, and choosing to protect against that one specifically.
Stage 4
Give the kill criteria
Say it like this
"Once a specific independent account has enough of its own order history, twelve months or more, and StockSense's accuracy on that account alone clears ninety percent, I'd lift the human-review requirement for that account specifically."
Why this works
Shows the pick isn't permanent stubbornness. There's a real, stated bar that would change it.
Stage 5
Prove it with the compressed evidence, and name the AI-specific reasoning
Say it like this
"As our book of business grew from 800 to 1,400 accounts, reps had less time to double-check every forecast, so they kept manually reviewing the big chains they cared most about and let StockSense run unsupervised on the small independents, exactly the segment where it's weakest because those accounts have the thinnest order history for it to learn from. Independent churn went from 3 to 11 percent a quarter before anyone connected the two."
Why this works
This is the load-bearing, AI-specific judgment: a model is only as good as the order history behind each individual account, and that varies wildly across a customer base.
Stage 6
Say what wouldn't change, then close
Say it like this
"I wouldn't add a human check to mid-size group accounts, since their order history sits in between and a light spot-check catches what little StockSense gets wrong there. For Auberdine, the pick holds: route by segment, protect the quiet expensive error, not the loud cheap one."
Why this works
Closes with real judgment about where the fix doesn't need to reach, and restates the position in one breath.
Let's learn
The label printer at Auberdine's Ohio warehouse prints a pick list before the sun's up, one for every account on the truck that day.
Before StockSense, a rep built each account's weekly order forecast by hand, about 25 minutes per account. With StockSense, that same forecast comes back in under a minute for all 1,400 of Auberdine's accounts, chains and independents alike.
The four letters, held up as one page. Cost asymmetry is the one most segment problems skip.
Here's the turn: StockSense scores 94 percent accuracy on large chain accounts and 71 percent on small independent accounts, and averaged across the whole book, that looks like a fine 85 percent. The average hides that one segment's mistakes are a phone call, and the other segment's mistakes are a relationship that ends without a word.
Cost of a bad forecast, by segment
One error is a same-day fix. The other is worth 280 times as much, and it never generates a ticket.
At its worst, an independent restaurant runs out of a key ingredient on a busy Friday because StockSense underforecast a new menu item, and the owner switches to a competitor supplier that same week, with nobody at Auberdine finding out until the account has already gone quiet for a month.
The mistake wasn't trusting the model too much. It was trusting it least exactly where the mistake costs the most.
The choice I would take back
Auberdine's original launch treated "does an account get an automated StockSense forecast" as one toggle for the whole book of business, on for everyone or off for everyone. That made sense when the book was 800 accounts and mostly chains. It stopped making sense once independents became a third of the book and grew fastest, with the thinnest order history of anyone.
What I would leave alone: I wouldn't add a mandatory human check to mid-size group accounts, since their order history sits between the two extremes and a light spot-check already catches what little goes wrong there.
The lesson: an average accuracy number is a place to hide a segment's real problem, not a place to find one.
Now here is the same thing as a story
The short version above is what you'd say defending a routing decision. Read this one for what it felt like watching two quarters pass with nobody able to name the exact week anything changed.
Her name is Marisol. She has run demand planning at Auberdine for five years, and she can spot a supplier's own excuse before they've finished making it.
When StockSense first launched, reps loved it evenly across every account. Chains got their forecasts in seconds. Independents, which used to take the longest to plan by hand because their orders swing with local events and weather, suddenly took no time at all. For months, everyone reviewed everything lightly, and it held up fine.
Two boxes of deliberately unequal weight. Only one of them ever shows up on a dashboard.
As Auberdine's book grew from 800 accounts to 1,400 over about a year, reps simply had less time. Nobody decided, in any single meeting, to stop double-checking independent forecasts. It just happened, a little at a time: reps kept close tabs on the handful of big chain accounts that made up most of their bonus, since a mistake there was visible and painful to explain. The hundreds of small independent accounts, worth less individually, got whatever attention was left over, which was less and less.
Knowledge spark: why would a model be worse for accounts with less order history?
A forecasting model learns an account's pattern from its own past orders. A chain with three years of steady weekly orders gives the model a clear pattern to follow. A small independent restaurant with six months of orders and a menu that changes with the season gives it much less to work with, so its guesses are shakier, even though the model itself hasn't changed at all.
Nobody could point to the week it changed. There was no single bad order, no one complaint that started it. Just, slowly, less scrutiny landing on exactly the accounts that needed it most.
The second step is where a thin order history quietly becomes a bad prediction, for either segment.
The real question was never whether StockSense was good enough overall. It was whether Auberdine's own reps were spending their limited checking time on the segment where a mistake actually costs the business something, instead of the segment where it just feels urgent in the moment.
The segment with the thinnest history is also the one with the weakest forecasts. That's not a coincidence, it's the mechanism.
When the automated rollout was first designed, someone said, "let's just turn it on for everyone, we'll tune it later," and it sounded reasonable, since at 800 mostly-chain accounts, there wasn't much of a gap to tune.
Independent-account forecast accuracy as order history builds
It's rising, on its own, as order history builds. It just isn't there yet, which is exactly what a fixed kill criteria is for.
Rerun the same two quarters with segment-based routing in place: chains keep running on StockSense alone, independents get a rep's four-minute check before ordering, and independent churn falls back from 11 percent a quarter to 4, close to where it sat before the book of business ever grew.
Three branches, and the busiest one, independents, is the one that finally gets a human step.
What I'd tell myself, watching that churn number climb: an average accuracy score is not a plan. It's a number that looks calm right up until you ask which segment it's actually calm about.
PICK, the decision that ends the whole "which segment do we trust" argumentNot a script for distrusting the model everywhere. PICK is what tells you exactly where to spend a rep's limited attention.
P
Position. Your pick, before the reasoning.
Route by segment. StockSense runs alone for chains. A rep checks its forecast before ordering for independents.
Stating the pick first is what makes this an answer instead of a tour of the tradeoff.
I
Impact. Who feels each kind of error, in what units.
A chain feels a fifty-dollar expedited shipment. An independent feels spoiled food or a churned account worth fourteen thousand dollars a year.
Real units turn a vague sense of "small accounts matter" into a comparison worth acting on.
C
Cost asymmetry. Which error is hidden and expensive.
The chain's error is loud and cheap, resolved the same day. The independent's error is quiet and expensive, discovered only at renewal, if at all.
This is the hardest step: optimizing against the quiet error instead of the loud one takes deliberate attention.
K
Kill criteria. What evidence would flip the pick.
Once a specific independent account has twelve or more months of order history and StockSense clears 90 percent accuracy on that account, the human check lifts for that account alone.
A stated bar is what separates a real decision from stubbornly distrusting a segment forever.
The recap, one line per letter: position is routing by segment instead of one blanket setting, impact is fifty dollars against fourteen thousand, cost asymmetry is choosing to protect against the quiet expensive error over the loud cheap one, and kill criteria is the ninety percent, twelve month bar that would lift the human check.
And if you want to be sure it really works, try it somewhere elseSame four letters, event ticketing instead of foodservice. The segments swap, the quiet expensive error doesn't.
Tomasz Wybicki runs product at Larkspur Ticketing, where FraudGuard flags likely-fraudulent purchases before a ticket ships. Mapped onto PICK: position is routing by buyer type, FraudGuard runs alone on casual, one-ticket buyers, and a human reviews flags on high-volume resellers, where it's worst. Impact: a false flag on a casual buyer costs an annoyed support ticket. A missed fraud case on a reseller account costs thousands in chargebacks across dozens of tickets at once. Cost asymmetry: the casual buyer's error is loud and cheap, a complaint, resolved fast. The reseller's error is quiet until a chargeback wave lands weeks later. Kill criteria: once a specific reseller account has enough clean transaction history, the human review lifts for that account.
Staff had quietly stopped mentioning when they manually overrode FraudGuard's calls on reseller accounts, since it felt like second-guessing a tool leadership had praised. Nobody was tracking how often that override happened, or why, until a chargeback wave made the pattern impossible to ignore.
The same four checks, whether the product ships food or tickets.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "route by segment, protect the error that's hidden and expensive, not the one that's loud," and stop.
Cost: no budget for a human reviewer on the weaker segment. Say so honestly, and start with the highest-value accounts in that segment first, rather than skipping the review everywhere.
The model got better, for real: if a vendor update genuinely improves accuracy on the weak segment, that's exactly when to re-check the kill criteria, not assume the gap has already closed.
Where people run it wrong.
They watch one blended accuracy number and never ask which segment is dragging it down.
They put review effort where mistakes are loudest, not where they're most expensive.
They never write a kill criteria, so a human check meant to be temporary quietly becomes permanent, or gets dropped too early with no evidence behind it.
How to use it live. The moment an interviewer describes a model that's uneven across customer segments, ask yourself: which segment's mistakes are loud and cheap, and which are quiet and expensive? Commit to protecting the quiet one, and the rest of the routing decision follows on its own.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Substitution flip: as their book of business grew, reps rationed their own limited checking time toward the chain accounts they cared most about, leaving the model unsupervised on the segment it handles worst.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Etxeberria, the AI PM at Auberdine Foodservice Supply, who split StockSense's role by customer segment after independent-account churn tripled.
3 · THE HABIT
What did reps stop doing because the book of business grew?
Tap to flip
ANSWER
They stopped double-checking StockSense's forecasts for small independent accounts, keeping their limited attention on the big chain accounts instead.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Reviewing every account's forecast lightly versus reviewing only the chains closely while independents ran fully unsupervised. There was no gradual middle, attention simply drained from one segment to the other.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating automated forecasting as one toggle for the whole book of business instead of a decision that could be sliced by segment from day one.
6 · THE NUMBER
Fill in the blank: independent-account churn rose from ___ percent to ___ percent a quarter as the book of business grew.
Tap to flip
ANSWER
3 percent to 11 percent a quarter.
7 · THE REPLAY
Same growth in accounts, segment-based routing in place. What changes?
Tap to flip
ANSWER
Independents get a rep's four-minute check before ordering, and churn falls back from 11 percent a quarter to 4, close to where it sat before the book of business ever grew.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Larkspur Ticketing's FraudGuard. The flip is concealment: staff quietly stopped mentioning when they overrode fraud flags on reseller accounts, so nobody tracked the pattern until a chargeback wave hit.
Check yourself Score: 0 / 0
True or false
1. True or false: StockSense's blended 85 percent accuracy across the whole book of business was a reliable signal that both customer segments were being served well.
True
False
Show hint
Look at the two accuracy numbers, 94 percent and 71 percent, behind that average.
Show answer
False. The blended number hid a 23-point gap between segments, and the segment doing worst was also the one where a mistake cost the most.
Multiple choice
2. Why did reps end up checking chain forecasts closely while independent forecasts ran unsupervised, the opposite of what the model needed?
A. Independent restaurants asked not to be contacted about their orders.
B. As the book of business grew, reps had less time, and rationed it toward the big chain accounts they cared most about, not the segment the model handled worst.
C. StockSense was programmed to only show alerts for chain accounts.
D. Chain accounts complained more often about forecast accuracy.
Show hint
Look at "the habit thinning" in the story section.
Show answer
B. Limited attention got rationed toward the accounts that felt most urgent to protect, which happened to be exactly backward from where the model actually needed help.
Fill in the blank
3. Fill in the blank: a chain account's forecast error costs about ___ dollars. An independent account's forecast error costs about ___ dollars in average lost annual revenue.
Show hint
Look at the grouped bar chart, "cost of a bad forecast, by segment."
Show answer
50 dollars, 14,000 dollars. A gap of 280 times, and the smaller number is the one that shows up on a dashboard.
Short answer, where it wouldn't matter
4. Name a segment of Auberdine's accounts where adding a mandatory human check would NOT be worth it, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Mid-size group accounts. Their order history sits between chains and independents, and a light spot-check already catches what little goes wrong there.
Short answer, apply it yourself
5. Think of a tool or service you use that treats every customer the same way. Where might it actually be worst for the customers who'd be hurt most by a mistake?
Show hint
Think about which users have the least data behind them, or the least ability to absorb an error.
Show answer
Model answer: A ride-share pricing algorithm trained mostly on dense city trips can misprice a rural route badly, and a rural rider has fewer alternatives to fall back on than a city rider does.
Short answer, work the number
6. If an independent account's forecast accuracy reaches 90 percent at month 15 instead of never quite getting there, would the kill criteria say to lift the human check?
Show hint
Look at both parts of the kill criteria, not just the accuracy number.
Show answer
Model answer: Yes, since the kill criteria only needs twelve or more months of history and 90 percent accuracy on that specific account, both of which would be true by month 15.
Before you close the answer
Why this works
Tests whether you'll look past a single blended accuracy number and find the segment where a mistake is quiet, expensive, and easy to miss until it's already cost you the account.
Follow-up traps
"Isn't adding human review to independent accounts just slower and more expensive?" Response: about four minutes a week per account, far less than the original 25-minute manual forecast, and cheap next to a churned account worth fourteen thousand dollars a year.
"Couldn't you just retrain the model with more independent-account data?" Response: worth doing in parallel, but it takes months to accumulate that history per account, and the human check protects revenue in the meantime, it isn't a substitute for the retraining, it's the bridge to it.
If pressed
The kill criteria is checked per account, not per segment as a whole, since a segment-wide average would let a handful of long-tenured independents drag the whole group's number up while newer independents stayed unreviewed and unprotected.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.