CaseAdvancedModel Fluency & the AI PM Role / Working with ML engineers and researchers / #6

Your ML team wants three months to improve accuracy by two points. How do you evaluate that ask?

PICK · Fernglow's ML team wants a quarter to move a blended caption score by two points, and the return notes keep repeating three words about one category no one tracks on its own

Fernglow writes the caption for you: upload a product photo and a short blurb, and about four seconds later you have a ready to post line and hashtags for Instagram and TikTok Shop. Ilkka Preen runs product for the Pro tier, where captions ship straight to a seller's queue with no human read, gated by a score called Claim Match, the share of captions whose claims are actually backed by the seller's own listing text. Sepideh Reyk leads the ML team, and she wants a quarter to move that blended score from about ninety one percent to about ninety three. Ilkka has to decide whether those two points are worth a quarter of her team's time, and worth it for what, exactly.

The direct answer
Do not sign off on the three months as asked, and do not turn it down either. Ask the ML team to name one real class of caption failures the two points would fix, the kind a seller already feels. Then run a two week check to see if a scoped fix actually moves that class before committing a full quarter to a blended number that might climb for reasons that never touch it.
Do this, in order
  1. Ask what specific class of failing captions the two points would fix, before agreeing to anything.Why: a blended score can climb two points without ever touching the one category driving real complaints.
  2. Run a two week scoped check on the named class before committing the full quarter.Why: it costs a fraction of the full ask and answers the real question fast instead of at the end of a quarter.
  3. Set a kill line with a real number, not a feeling, for what counts as fixed.Why: apparel's care claim rate has to clear a stated cut off, or the quarter has not earned its keep.
  4. If the team can only point to the blended target, send the ask back.Why: ninety one to ninety three with no named class is a target with no test attached to it.
  5. Once the class is confirmed and the check passes, scope the full quarter to that class, not to accuracy in general.Why: a blank check for accuracy tends to spend itself on whatever moves easiest, not on what is actually costing money.
  6. Leave the other five categories alone.Why: none of them show the pattern apparel does, so a scoped fix there would spend budget with nothing real to show for it.

How to answer this, stage by stage

Nobody is grading whether you can define accuracy. They are grading whether you can take a resourcing ask that sounds reasonable on its face and find out, out loud, what it would actually buy.

1
Put a real ask on the table, not a percentage in the air
Say it like this
"Let's ground this. Say there's a company called Fernglow. It writes captions for product photos, and a score called Claim Match decides whether a caption ships with no human read. The ML team wants a quarter to move that score from ninety one to ninety three percent."
Why this works
Keeps the interviewer from grading you on how you feel about the word accuracy instead of a real ask with real numbers in it.
2
Name the framework before you reason
Say it like this
"I'll use PICK here. Position, what I'd actually decide before saying yes or no. Impact, what's lost in each direction. Cost asymmetry, which mistake costs more to make. Kill criteria, the one test that decides it."
Why this works
Two seconds of structure tells the interviewer you have a method, not a mood.
3
Give the position, and refuse the binary
Say it like this
"My position isn't yes and it isn't no. It's: show me what the two points buy before I hand over a quarter. A blended score moving two points can mean everything got a little cleaner, or it can mean one loud category got fixed while the one that's actually costing money never moved."
Why this works
This is the direct answer, said out loud, before a story gets the chance to soften it into it depends.
4
Split the impact both ways, not just the scary one
Say it like this
"If I say yes blind and the two points land somewhere harmless, I've burned a quarter of ML time that could have gone to something sellers would actually notice, like TikTok Shop support. If I say no and the fix really would have caught the apparel problem, I've left forty returns a week on the table for another quarter."
Why this works
Naming what breaks on both sides keeps this from turning into a one sided caution story.
5
Point at the cheap test instead of the expensive guess
Say it like this
"Instead of picking blind, I'd ask for two weeks. Same two engineers, scoped to just the apparel care claim problem, checked against listings we already know have explicit care instructions. That costs about fourteen thousand dollars. Guessing wrong on the full quarter costs about eighty four thousand, and we wouldn't know it was wrong until month three."
Why this works
This is the actual center of PICK: naming which mistake is cheap to make and which one is expensive to walk back.
6
Give the kill line, in a real number
Say it like this
"Before the quarter starts, I want one number: does the apparel material claim rate clear two percent. Right now it's six point two. Below two is where it stops being any different from returns we get for reasons that have nothing to do with the caption, wrong size, a damaged box. That's the line that tells us it's actually fixed."
Why this works
A kill line with no number attached is just an opinion wearing a framework's clothes.
7
Close on the one decision, and say what happens next
Say it like this
"So here's the whole thing. Two weeks, scoped to the class that's actually costing money, checked against a real kill line. If it clears two percent, the full quarter gets approved, scoped to that fix, not to accuracy as a word. If it doesn't, we haven't spent a quarter finding that out the hard way."
Why this works
Restates the direct answer in one breath and hands the interviewer the test that makes it more than a preference.

Let's learn

Before Fernglow's Pro tier existed, every caption went past a person first. A seller read it, caught anything odd, fixed it, posted it. Slower, but nothing shipped that the seller hadn't actually looked at.

Hand sketched labeled parts diagram titled What Fernglow actually does. A gauge icon in the center labeled Claim Match gate, with four labels radiating out: photo plus listing text in, caption drafted for the post, claims checked against the listing, auto post or flagged for review.
One score decides whether a caption ships with no human read at all. That score is blended across six categories, and blended is exactly where this whole question lives.

Now, on the Pro tier, captions that clear Claim Match go straight out. That score checks one thing: does every claim in the caption, the material, the care instructions, the fit, actually show up in the seller's own listing text. Right now it sits at about ninety one percent, blended across six categories. Sepideh's team wants a quarter to move it to about ninety three.

Knowledge spark: what is a golden set? A pile of real examples someone has already checked by hand, with the right answer marked on each one. A model's score only means something if it is tested against a golden set the model never trained on. Fernglow's is six hundred forty seller listings, refreshed every quarter.

Here is the turn. The extra two points were never really about whether the score climbs. Somewhere inside that blended number, one category, apparel, has a much bigger problem. About six percent of its captions state a material or care claim the listing never made. Machine washable. Hypoallergenic. Wrinkle free. The model isn't lying on purpose. It's filling a gap the way any confident guesser fills a gap, and it sounds exactly as sure when it's wrong as when it's right.

What that costs at its worst: about two thousand four hundred apparel captions a week go out on auto post. If roughly a quarter of the ones with a false claim turn into an actual return, that's around forty returns a week, about seven hundred dollars, from captions saying something the product never promised. A blended score that climbs two points can hide that completely, if the fix happens to land somewhere else.

Cost, by the numbers: a scoped check versus a blind quarter
$90,000 $45,000 $0 $14,000 Two week scoped spike $84,000 Blind three month build
Cheap, answer in daysExpensive, answer in month three
The scoped check costs about a sixth of the full ask and tells you fast whether the named class actually moves. Committing the full quarter blind means paying six times more before finding out if it worked.
The two points were never the real ask. The six percent hiding inside one category was.
The choice I would take back Fernglow built one blended Claim Match number across all six categories from the very first version, back when apparel was a small slice of what the tool did. That made sense with one category. It stopped making sense once apparel became half of every auto posted caption and started producing its own kind of mistake nobody else's captions were making, and nobody had built a way to see it on its own.

What I would leave alone: home goods and electronics captions occasionally get a color a little off, calling a bronze finish gold. Buyers can see the actual color in the photo, so it has never once turned into a return. Splitting that category out for its own scoped fix would spend a quarter's attention on a mistake nobody is acting on.

The lesson: a blended score can climb for reasons that have nothing to do with the thing that's actually costing you money. Before you spend a quarter chasing a number, find out which category is really inside it.

Hand sketched metaphor scene titled Chase the dial or name the class. Left panel, a gauge icon labeled The dial, blended score ninety one to ninety three. Right panel, a scale icon labeled The class, apparel care claims named.
One side is a number that moves for almost any reason. The other is a named failure a seller could point to. Only one of them is worth a quarter.

Now here is the same thing as a story

The short version above is what you actually say out loud. Read this one for the Thursday sync that put the real question and the ask on the same desk in the same week.

Ilkka Preen can read a support ticket queue and tell you which three words are about to become a pattern, usually before anyone else notices there is one. Three years running product at Fernglow will do that.

Fernglow launched with one category, apparel, and one blended number for how good the captions were. When home goods, beauty, pet supplies, electronics accessories, and jewelry got added, one at a time over eighteen months, the blended number just got wider. Nobody built six numbers. They kept the one.

For most of that time, the one number was fine to trust. Ninety, ninety one, holding steady, month after month. Ilkka checked the dashboard every Monday out of habit more than worry. It never had anything to say.

Hand sketched horizontal timeline titled Ilkka's desk, four beats. Four milestones: blended score steady at ninety one percent for months running, category view stops with one dashboard covering six categories, the remark, not what it said repeated across three weeks, this milestone emphasized, and the ask reread, same quarter new question.
Nothing about the blended score ever changed. The problem was already sitting inside it, in a category nobody had a reason to look at separately.

Somewhere in there, without anyone deciding it on purpose, apparel quietly became almost half of everything running through the Pro tier's auto post pipeline. The blended score didn't care. It just kept averaging.

Then, in a routine Thursday sync, the support lead mentioned something small. Three weeks running, she said, return notes for apparel kept using almost the same phrase: not what it said. Not a complaint about fit. Not a complaint about shipping. About the caption.

Ilkka pulled the notes herself that afternoon. A merino wrap listed as hand wash only, captioned machine washable. A canvas tote with no water resistance claim in its listing at all, captioned water resistant. Small, specific, and every single one traced back to a caption Fernglow had generated and auto posted, no human read, because it had cleared the blended Claim Match score with room to spare.

Sitting on her desk that same week was the ML team's ask: a quarter, to take that blended score from ninety one to ninety three. She'd been about to approve it the way she approved most roadmap asks that came with a clean number attached. Two points, three months, seemed reasonable on its face.

We didn't need the score to climb. We needed to know which two percent of it was apparel's care claims, and whether the fix would ever touch them.

One option got floated in that same meeting, and it's worth naming because it looked like the safe call: pull apparel out of auto post entirely, force a human read on every caption until the ML team's fix landed. It lost. Thousands of apparel sellers had never filed a single return tied to a caption, and the whole reason the Pro tier exists is that captions ship the moment a seller hits upload. Slowing all of them down to catch a problem that lived in six percent of one category would have fixed the six percent by breaking the other ninety four.

So instead, Ilkka sent the ask back with one question attached: does the plan for these two points name apparel's care claim rate specifically, with its own number, or is it just ninety one to ninety three, blended. Sepideh's answer, to her credit, was honest: the roadmap item as written was aimed at the blended score generally, with no category singled out.

They agreed on two weeks instead of a quarter. Same two engineers, scoped to nothing but apparel's material and care claims, checked against a kill line: does the rate clear two percent, roughly where returns for reasons that have nothing to do with the caption already sit.

Day zero, the rate was six point two percent. By day four, four point one. Day eight, two point nine. Day twelve, one point nine, under the line for the first time. Day fourteen, one point six, and it held.

That's when the quarter got approved, not for accuracy in general, but for exactly this: rolling the same fix past the small test set and building the check into every apparel caption Fernglow ever generates.

What I'd tell myself, looking back at that Thursday sync: a number that's been steady for a year isn't proof nothing's wrong underneath it. It's proof nobody's split it apart recently enough to find out.

PICK, what a quarter has to earn

Not a vote on whether the ML team is right to ask. PICK only works here if you refuse to grade the ask by its headline number and go find out what it actually buys instead.

PPosition. What you'd actually decide.
Do not approve the three months as written, and do not reject it either. Find out first what the two points actually fix, in terms a seller would recognize, before signing off on a quarter.
A blended score moving two points can mean six small things each got a little cleaner, or it can mean one specific, costly thing got fixed. Those are very different quarters, even though the roadmap line reads the same.
Say the position before any number shows up. A position built backward from the story looks like it was reverse engineered from the answer.
IImpact. What's lost each way.
Approve blind, and if the two points land somewhere harmless, you've spent a quarter of ML capacity that could have gone to something sellers would actually notice, and the apparel problem is still there in month four.
Reject outright, and if the fix really would have caught the care claim problem, you've left forty returns a week on the table for another quarter, for no reason but caution.
Naming both losses keeps this from reading as one sided worry dressed up as rigor.
Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a gauge icon labeled Two week spike, about fourteen thousand dollars, answer in days. Right panel, a question mark icon labeled Blind quarter, about eighty four thousand dollars, answer in month three.
One mistake is small, fast, and shows its result inside a sprint. The other is large, slow, and hides its result until the quarter is already spent.
CCost asymmetry. The heart of it.
A two week check, scoped to the one class that's actually costing money, costs about fourteen thousand dollars and gives you the real answer inside a sprint. Committing the full quarter up front, on the blended number alone, costs about eighty four thousand, and you don't find out if it worked until the very end. Start with the small, fast, visible mistake. Spend the large, slow one only once you know it will land somewhere real.
KKill criteria. The one test.
Can the team point to a real, currently failing class of captions, apparel's material and care claims, with its own number, that the fix is built to move. Right now that number is six point two percent, and the line that says it's actually fixed is under two. If they can only offer "ninety one to ninety three, blended," that is not a test, it is a hope with a deadline attached.
The kill line, charted: apparel's care claim rate over the two week check
8% 4% 0% kill line: 2% 6.2% 4.1% 2.9% 1.9%, crossed here 1.6%, held Day 0 Day 4 Day 8 Day 12 Day 14
Above the kill lineCleared the kill line
The rate crossed the two percent kill line by day twelve and held at day fourteen. That result, not the blended score, is what earned the full quarter.
Hand sketched decision tree titled Where the check happens. Root box reading the three month ask lands, branching to four labeled conditions: names a failing case class leading to scope the quarter and approve, only a blended number moves leading to send back and run the spike, spike clears the kill line leading to full quarter scoped build, spike misses the kill line leading to no quarter and rework the plan.
Four branches, one question asked before any calendar time gets promised: does the ask name a class that actually fails today.

The trade worth saying out loud: the fix itself, once scoped, means checking every claim in an apparel caption against the listing text before it ships, which adds a small amount of time and cost per caption. Worth it, because the alternative is a hidden class of returns nobody was watching for.

And if you want to be sure it really works, try it somewhere else

Same four letters, a radiology practice instead of a seller's photo, and this time the thing hiding inside the blended number could cost more than a return.

Ravensbrook Imaging Partners runs a triage tool called Firstpass that flags studies likely to hold a critical finding, so a radiologist reads those first instead of waiting in normal queue order. The vendor's ML team wants a quarter to move Firstpass's blended flagging accuracy from eighty nine percent to ninety one, the same two point ask, a completely different picture at stake.

Where the projected two points would actually come from
2.5 pts 1.25 pts 0 Incidental findings, 1.1 pts Nodule flags, 0.5 pts Pneumothorax, 0.2 pts Other, 0.2 pts 2.0 pts total
Low stakes categoriesThe finding that actually matters
Most of the projected gain sits in incidental findings nobody treats urgently. Only two tenths of a point touches pneumothorax, the finding currently missed on about three point four percent of positive studies.

Position: don't approve or reject the vendor's quarter on the blended number alone, ask which finding class the two points come from. Impact: approve blind and the quarter might fix everything except the one finding that matters; reject outright and a genuine safety relevant gain, if one exists, never ships. Cost asymmetry: two weeks spent breaking the projected gain down by finding type is cheap; a full quarter spent chasing a blended target is expensive if pneumothorax barely moves. Kill criteria: does the plan name pneumothorax sensitivity specifically, with its own target, or is it just eighty nine to ninety one, blended.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: ask what class the number would fix, before anything about cost.
Cost: no budget this quarter for a two week check. Fine, but get the plan to name the class in writing before the full quarter starts, so the whole team is scoped to that class, not just the metric.
The model got better, for real: say the blended score improved on a fresh benchmark, no work needed. Still doesn't tell you whether the one class you actually care about moved. Ask for that number by itself.

Where people run it wrong.
They treat the size of the ask, two points, three months, as if it describes the problem, when it only describes the target.
They assume a blended number that's been steady for months means nothing underneath it is drifting, when steady just means nobody's split it apart recently.
They fix a slow decision by approving the whole thing at once, instead of buying a cheap answer first and committing once they have it.

How to use it live. Before answering yes or no to any ask like this, ask one thing out loud: what specific case, that currently fails, would this fix. If the answer names one, you have something to evaluate. If it doesn't, you have a number looking for a reason.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks you to evaluate a resourcing ask instead of picking a side outright?
Tap to flip
ANSWER
PICK: state a position that refuses the flat yes or no, name what's lost in each direction, find which mistake is cheap versus expensive, then name the one test that would flip your call.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ilkka Preen, who runs product for Fernglow's Pro tier, a caption tool for social media sellers, gated by a score called Claim Match.
3 · THE REAL QUESTION
What's the actual question to ask before saying yes or no to a three month accuracy ask?
Tap to flip
ANSWER
Not "is two points worth three months." It's: what specific, currently failing class of cases would those two points fix, in terms a user would recognize.
4 · THE SPLIT
What's lost in each direction if Ilkka gets this wrong?
Tap to flip
ANSWER
Approve blind and the quarter might move the blended score without ever touching apparel's care claim problem. Reject outright and a real fix, if it exists, doesn't ship, and forty returns a week keep happening.
5 · THE ASYMMETRY
Which mistake is cheap and which one is expensive here?
Tap to flip
ANSWER
A two week scoped check costs about fourteen thousand dollars and gives a fast answer. Committing the full quarter blind costs about eighty four thousand, and you don't find out if it worked until month three.
6 · THE NUMBER
Fill in the blank: blended score sits at ___ percent, the ask would move it to ___. Apparel's claim hallucination rate is ___ percent, and the kill line is ___ percent.
Tap to flip
ANSWER
91.4 percent. 93.4 percent. 6.2 percent. 2 percent. The gap between those last two numbers is the whole argument.
7 · THE KILL TEST
What's the one test that decides whether the three month ask is worth it?
Tap to flip
ANSWER
Can the team name a real class of currently failing cases, with its own number, that the fix is built to move. A blended target with no named class fails this test.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of apparel's care claims there?
Tap to flip
ANSWER
Ravensbrook Imaging Partners' triage tool, Firstpass. Pneumothorax flagging sensitivity, the one finding class that actually matters, against a blended number padded out by low stakes incidental findings.

Check yourself Score: 0 / 0

True or false
1. True or false: the ML team's ask to move the blended Claim Match score from 91 to 93 percent was, on its own, a bad idea.
  • True
  • False
Show hint
Check what Ilkka actually objected to when the ask sat on her desk.
Show answer
False. The number itself wasn't the problem. The problem was approving a quarter of work without knowing whether those two points would ever touch apparel's care claim rate, the thing actually costing money.
Multiple choice
2. Why did Ilkka send the quarter long ask back instead of approving it?
  • A. She didn't trust Sepideh's team's estimate of three months.
  • B. The plan aimed at the blended score generally, with no named class, so there was no way to check if it would fix the actual problem.
  • C. Fernglow didn't have budget for a quarter of ML time.
  • D. The blended score was already high enough not to bother improving.
Show hint
Look at the exact question Ilkka asked Sepideh in the story.
Show answer
B. Sepideh's own answer confirmed the roadmap item as written targeted the blended score generally, with no category singled out.
Fill in the blank
3. Apparel's material and care claim hallucination rate started the two week check at ___ percent and ended at ___ percent, clearing the ___ percent kill line by day ___.
Show hint
Check the numbers on the line chart under the K letter.
Show answer
6.2 percent to 1.6 percent, clearing the 2 percent kill line by day 12. It crossed under the line at day twelve, at one point nine percent, then held at one point six by day fourteen.
Short answer, name the rejected alternative
4. What alternative did the team consider right when the problem surfaced, and why did it lose?
Show hint
Look at the meeting where the ask and the return notes ended up on the same desk.
Show answer
Model answer: Pull apparel out of auto post entirely and require a human read on every caption until the fix landed. It lost because thousands of apparel sellers had never had a single caption related return, and forcing all of them to wait would have broken the entire reason the Pro tier exists, to catch a mistake that lived in only six percent of one category.
Short answer, apply it yourself
5. Think of a tool you use that reports one overall score for something. What's one narrower slice of that score you'd want broken out before trusting it?
Show hint
Look for a tool that blends very different kinds of tasks into one number.
Show answer
Model answer: A food delivery app's overall on time rate blends easy short trips with rare cross town orders. Before trusting it, I'd want the cross town rate on its own, since that's the slice where being late actually changes whether I'd use the app again.
Short answer, work the number
6. If Fernglow's ML team could only test their fix for two days instead of two weeks, would the kill line still be a fair test? Why or why not?
Show hint
Look at how much the rate moved just in the first four days.
Show answer
Probably not on its own. Two days gets you the trend's direction, but the rate moved from 6.2 to 4.1 in the first four days alone. That's not enough time to know if it holds, and holding is the whole point of a kill line, not just a fast dip that creeps back up.
Before you close the answer
Why this works
Tests whether you'll grade a resourcing ask by its own stated size, two points, three months, or go looking for what specific problem it actually fixes. Most candidates either approve because the ask sounds reasonable, or reject because three months sounds long. Neither one is a real evaluation.
Follow-up traps
"What if the two week check comes back inconclusive, not clearly a pass or fail?" Response: extend it by a week at the same scope, not by expanding straight to the full quarter. An inconclusive result still isn't a confirmed class to build against.

"Isn't picking apparel's care claims after the fact just cherry picking a number that made your case?" Response: no, the category was named before the two week check ran, from three weeks of return notes repeating the same phrase, not chosen afterward to fit a result.
If pressed
The scoped fix works by pulling every material or care phrase out of a generated caption and checking it against the seller's own listing text with a lightweight entailment classifier, not the caption model grading its own output, since a model checking its own claims tends to agree with itself. That check adds about one hundred fifty milliseconds and a fraction of a cent per caption.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more