ConceptFoundationalModel Fluency & the AI PM Role / What changes when the product is probabilistic / #9

Give an example of a product where a 95 percent success rate is excellent and one where it is unshippable.

GUARD · a song-recommendation stream vs a mammography triage queue, same headline number

Hummingline queues the next song in a listener's personal stream. Enitan Delgadillo leads its product team. Foretell reads an incoming mammogram at Solandra Health's screening clinics and decides how soon a radiologist looks at it. Priyamvada Vesterling owns what its numbers say. Both ship at a headline 95 percent. Only one of them earns it, and the difference shows up in what happened to Cosima Achterloo's file.

The direct answer
Hummingline's next-song recommendation is excellent at 95 percent: a bad pick costs a four-second skip, spread evenly across every listener, and the person who eats the miss is the same person who can fix it on the spot. Foretell, a mammography triage tool, is unshippable at 95 percent: its missing 5 percent doesn't spread around at all. It lands completely on real cancer patients sorted into a slower reading queue, with no idea it happened and no way to say anything about it. Judge a success rate by who absorbs the miss and whether they can see it, never by the number alone.
Do this, in order
  1. Judge every success rate by who absorbs the miss and whether they can see it, not the number on its own.Why: this is the whole decision the rest of the answer works out.
  2. Audit outcomes by risk tier or confidence band, never one blended number.Why: Foretell's blended 5 percent hid a 31 percent miss rate sitting inside its own borderline calls.
  3. Never let a lower-priority tag quietly mean a slower total wait.Why: Solandra Health's Routine queue crept from 4 days to 21 because nothing capped it independent of the tag.
  4. Raise the bar for any tool whose miss can't be undone by the person it lands on.Why: a patient can't self-correct a triage tag the way a listener self-corrects a bad song.
  5. Give every deprioritized person a visible reason, or at least a real channel to ask.Why: Cosima never knew Foretell existed, so she had nothing to push back against.
  6. Leave Hummingline's 95 percent alone.Why: chasing it higher would trade song discovery for a harm that, at four seconds a miss, doesn't actually exist.

How to answer this, stage by stage

Nobody is grading whether you can say "context matters." They're grading whether you can name two real products at the same number and explain, specifically, why one of them should never have shipped there.

1
Ground it in two named products before touching thresholds
Say it like this
"Let's make this concrete. Hummingline is a song-recommendation stream, Enitan Delgadillo runs its product team. Foretell is a mammography triage tool built into Solandra Health's screening clinics, Priyamvada Vesterling owns its numbers. Both run at 95 percent. That's the whole comparison."
Why this works
A number with no product attached is just an opinion about a percentage.
2
Say the structure out loud
Say it like this
"I'll run this as GUARD. Name who's affected by each tool, find where the miss actually lands, ask who can push back on it, say what I'd change, then say how I'd know it's happening."
Why this works
Two seconds of structure signals a plan, not four scattered observations arriving as they occur to you.
3
Reframe the question before naming either product
Say it like this
"The number by itself tells you nothing. Ninety-five percent excellent and ninety-five percent unshippable look identical on a slide. What's different is who eats the missing five, and whether that person ever finds out."
Why this works
This line is the whole answer. Skip it and the rest sounds like two unrelated anecdotes instead of one argument.
4
Give both examples, plainly
Say it like this
"Hummingline: excellent. A wrong song costs four seconds, spread evenly across every listener, fixed by the same person it happened to. Foretell: unshippable. Its missing five percent lands entirely on real cancer patients, sorted into a slower queue they never see and can't contest."
Why this works
This matches the direct answer, and it's concrete enough that a follow-up question has something to grab onto.
5
Prove it with the compressed story
Say it like this
"Here's what that looked like once. A mammogram got tagged Routine on a borderline call. The Routine queue, quietly under-staffed against Priority, had crept to 21 days by then. She got read on day 21, not day 4. What was an early finding on day 4 needed chemotherapy by day 21."
Why this works
A number with no person attached doesn't land the way this question needs it to.
6
Give the fix, and say what it costs
Say it like this
"I'd cap Routine's wait time regardless of tag, and audit the miss rate by Foretell's own confidence band instead of one blended figure. Raising the threshold that far means more borderline scans bump to Priority, more radiologist hours, more cost per screening. I'd take that trade every time over a queue nobody's watching."
Why this works
Naming the cost is what makes this a real decision instead of a wish that quality and cost were both free.
7
Close on the one line
Say it like this
"So: same 95 percent, opposite verdict, because one miss costs four seconds and fixes itself, and the other costs three weeks a patient never even knows to ask about."
Why this works
Leaves the interviewer with the decision, not just the story.

Let's learn

Here's what happens when the exact same success rate means two completely different things, depending on who has to live with the other five percent.

Hummingline queues the next song in a listener's personal stream: pick a seed, and the model keeps the vibe going, one track at a time. Foretell reads an incoming mammogram and decides how soon a radiologist should look at it, sorted into Priority, read the same day, or Routine, read in normal order. Both call themselves 95 percent accurate.

Hand sketched icon list titled Solandra Health's reading queue, before Foretell. Three numbered rows: one, every mammogram read in the order it arrived. Two, turnaround held near 4 days, no tiers at all. Three, one queue, one radiologist rota, no split.
Before Foretell, there was no tag to trust or distrust. One queue, one line, one pace.

Before Foretell, a radiologist read every mammogram in the order it arrived. No tiers, no tags, turnaround held near four days, week after week. With Foretell, the truly urgent cases get read the same day, and the tool earns its keep there. What matters is the other side: at 95 percent sensitivity for Priority-tagging, 5 in every 100 real cancers get tagged Routine instead, and enter the same queue as everyone else's screening.

Knowledge spark: what does sensitivity actually mean here? Foretell doesn't diagnose anyone. It only decides how soon a person reads the file. Sensitivity is how often it correctly flags a real cancer for a same-day read. 95 percent sensitivity means 5 out of every 100 real cancers get sorted into the slower queue instead.

Solandra Health's clinics screen about 42,000 women a year. At the usual population rate for screening, about 5 cancers turn up per 1,000 women, so roughly 210 real cancers a year. Five percent of 210 is about 10. Ten real women a year, tagged Routine instead of Priority.

Hand sketched quadrant chart titled Same 5 percent, landing two different ways. X axis how visible the miss is, from hidden to obvious. Y axis who absorbs it, from spread across everyone to concentrated on a few. Hummingline miss plotted obvious and spread across everyone. Foretell miss plotted hidden and concentrated on a few.
Hummingline's miss sits where anyone would want a miss to sit: obvious, and spread evenly. Foretell's sits in the worst corner of this chart.

Compare that to Hummingline. A wrong song there costs Tessaly Braelynn about four seconds on her commute: she skips it, or taps thumbs-down and the stream adjusts instantly. Nobody is harmed, because the person who eats the miss is also the person with full power to notice and fix it, right then.

The missing 5 percent isn't spread evenly. It's ten women a year carrying the whole cost the tool was built to prevent.
Foretell's miss rate on Priority-tagging, blended vs by confidence band
35% 17.5% 0% 5% Blended, all calls <1% Confident calls 31% Borderline calls
Blended, every callConfident callsBorderline calls
Foretell already knows which of its own calls are shaky. Nobody audits by that flag. Everybody watches the blended 5.

What it costs at its worst: Cosima Achterloo's mammogram showed a small, genuinely hard-to-call mass, borderline enough that Foretell's own model sat right on the edge of its cutoff. Tagged Routine, her file waited 21 days instead of the 4 the old queue used to promise. What a day-4 read would likely have caught early enough for surgery alone had grown, by day 21, into something that needed chemotherapy too.

Routine queue turnaround, month by month
24d 12d 0d 5-day service commitment Month 1 Month 2 Month 3 Month 4 Month 5, Cosima reads day 21
Routine turnaround, daysWhere Cosima's file landed
Blended sensitivity never moved off 95 percent while this line climbed underneath it for five straight months.
The choice I would take back Solandra Health let Priority and Routine share one radiologist rota, with Routine's wait simply whatever time was left over once Priority was covered. That made sense the day 95 percent looked like a number that made everybody a little safer. It stopped making sense the moment anyone worked out what the missing five percent actually was: not spread out, but concentrated entirely on the roughly ten women a year real enough to need a same-day read and unlucky enough not to get tagged for one.

What I would leave alone: Hummingline's 95 percent doesn't need any of this. A skipped song costs four seconds and self-corrects the instant it happens. Chasing that number higher would mean narrowing the model's willingness to try a new-but-plausible song, trading real discovery for a harm that, at four seconds a miss, was never actually there.

The lesson: a success rate isn't one number, it's a question about who's standing where the miss lands. Ask that question before you ask what the number is.

Now here is the same thing as a story

Read the short version above when you're actually in the interview chair. Read this one when you want to feel exactly what 21 days cost, next to 4.

Priyamvada Vesterling can read Foretell's monthly performance report faster than almost anyone at Solandra Health, and for the first year, she read every part of it, not just the top line.

She joined the product team six months before Foretell launched, back when everyone agreed on one thing: the tool would only work if a Priority tag genuinely meant faster hands on the file, not just a label on a screen. To make sure of it, she pulled the Routine-queue turnaround number by hand every single month for the first year, checked against the sensitivity dashboard everyone else watched.

For a long while, both numbers agreed with each other. Sensitivity held near 95 percent. Routine turnaround held near four, then five, then six days, boring and stable. Priyamvada's monthly pull started feeling like busywork. She let it slide to a quarterly glance. Then to a question she'd ask herself, half out loud, whenever the sensitivity chart moved: is Routine still fine? It always looked fine, whenever she checked.

Nobody can point to the week it stopped being fine. Radiologists reading Priority cases got noticed, praised even, in the weekly huddle, for same-day turnaround on the highest-risk files. Nobody in that huddle ever mentioned Routine, so nobody's day got built around protecting it. Dr. Radu Stanescu, who reads both tiers, did what any reasonable radiologist would: cleared Priority first, every morning, and let Routine wait for whatever hours were left. Four days crept to six. Six to nine. Nine to fourteen. Nobody decided this. It just happened, a little at a time, in the same direction, for five months.

Hand sketched comparison diagram titled Two people, the same tag, very different power. Left panel, a person icon labeled Dr. Stanescu, caption reads the queue, can move a case up. Right panel, a person icon labeled Cosima, caption never sees the tag, can't move anything.
Dr. Stanescu can move a file up the list the moment something looks wrong. Cosima never learns there was a list to move up on.

Cosima Achterloo's mammogram went into that queue in month five. It showed a small, genuinely hard-to-call mass, borderline enough that Foretell's own model sat right on the edge of its cutoff. Tagged Routine. She had no way to know a tag existed, let alone that hers had landed on the wrong side of it. She went home the way she always did after a screening, expecting nothing for a couple of weeks.

She waited twenty-one days for a queue that, a year earlier, would have read her file in four.

Dr. Stanescu read her file on day 21 and called her back that same afternoon. Nobody in that reading room did anything careless. He read every file in the order the system handed it to him. The system just never told him, or her, that the order had quietly stopped being safe.

We didn't lose three weeks to a broken model. We lost them to a queue nobody was watching once the tag told everyone it was fine.
Hand sketched left to right flow diagram titled Where an appeal should sit, and doesn't. Five boxes connected by arrows: mammogram taken, Foretell scores it, tagged Routine, no notice no appeal this box emphasized in red, read in queue order.
The path only runs one way. Nothing in it lets a patient ask the tag a question, or even know one was made.

Cosima's file went into Solandra Health's incident review as a routine delay, one of several that quarter. Nothing in that review asked what Foretell had said about her file first, because nothing in her chart recorded it. She never saw a score. Never saw a queue label. Never had anything to ask a question about, because as far as she knew, a screening just takes a few weeks sometimes.

The decision Priyamvada would take back reaches back to the week Foretell launched, when the team decided Priority and Routine would share one rota, and Routine's wait would just be whatever time was left. That made sense the day 95 percent looked like a number that spread the risk evenly. It stopped making sense the moment anyone understood what the missing five percent actually was.

Run the same five months again, with two changes in place. Every mammogram, Priority or Routine, gets read inside five business days, full stop, independent of the tag. And Foretell's own confidence score gets audited separately every month, not folded into one blended number. Cosima's file still lands on the Routine side of a genuinely close call, because nothing hits 100 percent. But the five-day ceiling means Dr. Stanescu reads it on day 5, not day 21, while the finding is still what a day-4 read would have caught.

One design let a tag quietly decide how long an actual risk could wait. The other lets the wait have its own floor, no matter what the tag says.

What I would tell myself, the day Foretell launched: a tool built to protect the most urgent five percent still needs someone watching what happens to the other ninety-five, because that's exactly where urgency goes to disappear if nobody's looking.

Hand sketched full page metaphor scene titled A miss you feel, and a miss you don't. Left panel, a gauge icon labeled FELT AT ONCE, caption a skipped song, fixed in four seconds. Right panel, a question mark icon labeled NEVER LEARNED, caption a queue tag nobody sees or asks about.
The whole difference between these two products, in one picture. One miss reports itself. The other has to be gone looking for.

GUARD: the same 95 needs a different verdict depending on who's holding the other 5

Not a checklist for spotting risk in general. GUARD forces you to say, specifically, who eats a miss and whether they can do anything about it, before you're allowed an opinion on the number.

GGroups. Who's actually affected.
On Hummingline: every listener, equally. At Solandra Health: Cosima, the patient; Dr. Stanescu, the radiologist who trusts the tag to sort his day; and Priyamvada, who reports the blended number up the chain.
Naming the patient, the operator, and the person reporting the number keeps this from being "the model is risky" in the abstract.
UUnequal. Where the harm lands hardest, and why.
Hummingline's miss spreads evenly, a few seconds per listener, nobody singled out. Foretell's concentrates completely on the roughly ten real cancer patients a year, and disproportionately inside its own borderline-confidence calls: 31 percent there, against under 1 percent on confident calls.
Not "the model is unreliable." The specific reason one small slice carries nearly the entire real cost.
AAbility to contest. Who never gets to push back.
Tessaly sees a bad song the second it plays and fixes it herself, free, in the same breath. Cosima never sees Foretell's output at all. Nothing in her care ever told her a triage layer existed, so she has nothing to push back against, and neither does anyone on her behalf.
The strongest move in GUARD: the gap between a risk that's managed and one that's simply absorbed by whoever it lands on, without them ever knowing.
RReduce. The specific design change.
Raise the sensitivity bar for Priority-tagging, accepting more false positives and more radiologist hours in exchange. Cap Routine's real wait time independent of the tag. Keep Foretell advisory-only, so it can never quietly pull staffing away from the tier it's supposed to protect.
Three concrete changes, not a promise to keep an eye on the queue.
DDetect. How you'd know it's happening.
Audit the miss rate by Foretell's own confidence band, monthly, and track Routine turnaround as its own tracked number, never folded into the blended sensitivity figure. That combination would have shown the pattern by month two, not month five.
Turns "we'd probably notice eventually" into a number a review can actually act on.

And if you want to be sure it really works, try it somewhere else

Same five letters, a county benefits office instead of a screening clinic, and the missing five percent is a family's rent instead of a diagnosis.

Cedarhollow County's Department of Human Services runs Priorly, a tool that reads an incoming application for emergency food assistance and decides whether it qualifies for Expedited processing, three business days by law, or the ordinary Standard queue, thirty days. Marisabel Duquette, one of Cedarhollow's caseworkers, trusts Priorly's tag the way any busy caseworker would: she works Expedited first, because that's the queue the state audits. Priorly runs at 95 percent sensitivity for flagging genuine emergency need. The missing 5 percent doesn't spread evenly across every application. It lands on the small share of genuinely urgent households a borderline call sent to Standard instead, like Solveig Halcott's, filed the week she lost her job.

Hand sketched icon list titled The same three groups, at a county benefits office. Three numbered rows: one, Solveig, never sees Priorly's tag, has no lever. Two, the caseworker, trusts the tag, works the visible queue first. Three, county leadership, watches one blended accuracy number.
Same three seats as Solandra Health's story. A patient's chart became an application file, and a radiologist's rota became a caseload.
The decision Cedarhollow would take back Standard's processing time was left to float on whatever caseworker time was left after Expedited, on the same assumption Solandra Health started with: that 95 percent already meant everyone genuinely urgent had been caught.

Same rank, different lever, mapped straight onto GUARD: the groups are Solveig, who never sees Priorly's tag, and Marisabel, who works the visible queue first because that's the one the state grades her on. The harm concentrates entirely on the households a borderline call sent to Standard, not spread across the whole caseload. Solveig has no way to know a tag decided her case could wait, and no channel to ask about it, it just reads to her as a slow county office. The fix is the same shape: cap Standard's real wait independent of the tag, and audit outcomes by Priorly's own confidence band instead of one countywide accuracy figure. Detecting it means watching Standard's own turnaround climb, the same leading signal that would have caught Solandra Health's queue five months early.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to it: cap the wait independent of the tag, audit by confidence band, never trust one blended number.
Cost: no budget for more caseworker hours this year. Ship the wait cap first, since it costs schedule discipline, not headcount, then add the confidence-band audit once budget allows.
The model got better, for real: say Priorly's sensitivity jumps to 99 percent. Keep the wait cap anyway. A smaller miss still deserves a floor, because the harm on the family who hits it hasn't gotten any smaller.

Where people run it wrong.
They watch one countywide accuracy number and call the rollout a success.
They let the tag decide who gets caseworker attention, instead of only deciding read order.
They treat "we'd probably catch it in the next audit" as a plan, instead of pricing what waiting for that audit actually costs a family.

How to use it live. Ask the coverage question before naming a fix: "does a lower tag ever change how long someone actually waits, or does it only change the order they're seen in?" That question alone tells you whether the harm is contained or compounding.

Three things worth stating directly, since the real judgment sits here. The alternative Solandra Health could have taken instead of the wait ceiling was a quarterly audit of a random sample of Routine-tagged cases. It lost, because an audit only finds a pattern after it's already cost a quarter's worth of women their head start; it detects the harm, it doesn't stop it. The AI-specific failure worth naming is automation complacency on a probabilistic score: once Foretell's blended number stayed near 95 percent, Solandra Health's staffing decisions treated that score as settled fact instead of a call the model itself knew was often close, and nothing in the rollout ever asked Foretell how sure it actually was, case by case. The guardrail is the wait-time floor plus the confidence-band audit, so a close call can never silently turn into an unbounded wait. And the trade-off is real: raising the threshold that catches more Priority cases pushes meaningfully more borderline scans into the Priority queue, more radiologist hours, more cost per screening episode, accepted on purpose, because the alternative is finding out from a callback on day 21 instead of a number that moved on day 40.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks you to name where a rate is fine and where it isn't?
Tap to flip
ANSWER
GUARD: name who's affected, find where the harm concentrates, ask who can push back, name the fix, then say how you'd detect it.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Tessaly Braelynn, a Hummingline listener; Priyamvada Vesterling, who owns Foretell's numbers at Solandra Health; Dr. Radu Stanescu, who reads both tiers; and Cosima Achterloo, whose mammogram got tagged Routine.
3 · THE GROUPS
Who are the groups GUARD asks you to name here?
Tap to flip
ANSWER
Every Hummingline listener, spread evenly. At Solandra Health: the patient (Cosima), the radiologist who trusts the tag (Dr. Stanescu), and the team reporting the blended number (Priyamvada).
4 · WHERE THE HARM CONCENTRATES
Where does the harm land hardest, and why that group specifically?
Tap to flip
ANSWER
Foretell's own borderline-confidence calls: a 31 percent miss rate there, against under 1 percent on confident calls, both hidden inside one blended 5 percent.
5 · THE OLD DECISION
What decision would this answer take back?
Tap to flip
ANSWER
Letting Routine's wait time float on whatever radiologist time was left over after Priority, on the assumption that 95 percent already meant everyone urgent had been caught.
6 · THE NUMBER
Solandra Health screens about ___ women a year. About ___ real cancers occur in that population. At 95 percent sensitivity, about ___ get tagged Routine instead of Priority.
Tap to flip
ANSWER
42,000 women a year. About 210 real cancers. About 10 of those get tagged Routine instead of Priority.
7 · THE FIX, MADE COUNTABLE
Same five months, with the fix in place, what changes?
Tap to flip
ANSWER
Cosima's file still gets tagged Routine, but a five-day wait ceiling means Dr. Stanescu reads it on day 5, not day 21, while it's still what a day-4 read would have caught.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent miss?
Tap to flip
ANSWER
Priorly, Cedarhollow County's benefits-triage tool. The equivalent miss is a genuinely urgent food-assistance application tagged Standard instead of Expedited, like Solveig Halcott's.

Check yourself Score: 0 / 0

Multiple choice
1. Which factor makes Foretell's 5 percent miss worse than Hummingline's 5 percent miss?
  • A. Foretell runs an older model version than Hummingline.
  • B. The harm concentrates entirely on real cancer patients who can't see or contest the tag.
  • C. Foretell processes far more images per day than Hummingline processes songs.
  • D. Radiologists generally distrust new software more than listeners do.
Show hint
Think about who absorbs each miss, and whether that person ever learns it happened.
Show answer
B. Hummingline's miss spreads evenly and self-corrects. Foretell's concentrates on a small group with no way to know or object.
True or false
2. True or false: Foretell's blended 95 percent sensitivity number is, on its own, enough to know the tool's risk is under control.
  • True
  • False
Show hint
Check the bar chart: what happens when the blended number gets split by confidence band?
Show answer
False. Split by confidence band, confident calls miss under 1 percent while borderline calls miss 31 percent. The blended 5 percent hides that entirely.
Fill in the blank
3. Solandra Health's Routine queue turnaround crept from about ___ days to about ___ days over five months, while its blended sensitivity number never moved.
Show hint
Check the line chart in Let's learn.
Show answer
4 days to 21 days. Nobody was tracking Routine turnaround as its own number, so nothing caught the climb until Cosima's case surfaced it.
Short answer, name the reversal
4. What old decision would this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Letting Priority and Routine share one radiologist rota, with Routine's wait simply whatever time was left over. It made sense when 95 percent looked like a number that spread risk evenly across everyone, before anyone worked out that the missing five percent actually concentrated on a small, specific group.
Short answer, apply it yourself
5. Think of a tool you use that ranks or scores you: a loan application, a resume screen, a delivery ETA. Name one situation where a miss there would be invisible to you, the way Cosima's was.
Show hint
Look for a decision made about you that you'd never see a score or a reason for.
Show answer
Model answer: A resume-screening tool that quietly sorts applications into "review soon" and "review eventually." A strong candidate landing in the second pile has no way to know a score put them there, and no channel to ask why, the rejection just arrives late, or never.
Fill in the blank, work the number
6. If Solandra Health's screening volume doubled to 84,000 women a year with the same rates, about how many real cancers a year would get tagged Routine instead of Priority?
Show hint
Work out the cancer count first, then take 5 percent of it.
Show answer
About 20 to 21 women a year. 84,000 at roughly 5 cancers per 1,000 is about 420 real cancers; 5 percent of 420 is about 21, double the original 10.
Before you close the answer
Why this works
Tests whether you'll pick apart a single headline percentage into who bears the miss and whether they can see it, instead of treating 95 percent as one uniform verdict on the whole product.
Follow-up traps
"Isn't raising Foretell's threshold this much overkill, since 95 percent already sounds high?" Response: "High" only describes the model. It says nothing about whether the group eating the miss can undo it, and here they can't.

"Doesn't a wait-time ceiling independent of the tag defeat the whole point of triage?" Response: No, it only bounds the worst case. Priority cases still get read same-day; Routine just stops being allowed to drift past a safe floor.
If pressed
The borderline band isn't a guess. It's the roughly 6 percent of scans where Foretell's own predicted probability lands within 3 points of its Priority cutoff, the exact place a probabilistic model is, by definition, least sure of itself, and exactly where the 31 percent miss rate concentrates.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more