Give an example of a product where a 95 percent success rate is excellent and one where it is unshippable.
Hummingline queues the next song in a listener's personal stream. Enitan Delgadillo leads its product team. Foretell reads an incoming mammogram at Solandra Health's screening clinics and decides how soon a radiologist looks at it. Priyamvada Vesterling owns what its numbers say. Both ship at a headline 95 percent. Only one of them earns it, and the difference shows up in what happened to Cosima Achterloo's file.
- Judge every success rate by who absorbs the miss and whether they can see it, not the number on its own.Why: this is the whole decision the rest of the answer works out.
- Audit outcomes by risk tier or confidence band, never one blended number.Why: Foretell's blended 5 percent hid a 31 percent miss rate sitting inside its own borderline calls.
- Never let a lower-priority tag quietly mean a slower total wait.Why: Solandra Health's Routine queue crept from 4 days to 21 because nothing capped it independent of the tag.
- Raise the bar for any tool whose miss can't be undone by the person it lands on.Why: a patient can't self-correct a triage tag the way a listener self-corrects a bad song.
- Give every deprioritized person a visible reason, or at least a real channel to ask.Why: Cosima never knew Foretell existed, so she had nothing to push back against.
- Leave Hummingline's 95 percent alone.Why: chasing it higher would trade song discovery for a harm that, at four seconds a miss, doesn't actually exist.
How to answer this, stage by stage
Nobody is grading whether you can say "context matters." They're grading whether you can name two real products at the same number and explain, specifically, why one of them should never have shipped there.
Let's learn
Here's what happens when the exact same success rate means two completely different things, depending on who has to live with the other five percent.
Hummingline queues the next song in a listener's personal stream: pick a seed, and the model keeps the vibe going, one track at a time. Foretell reads an incoming mammogram and decides how soon a radiologist should look at it, sorted into Priority, read the same day, or Routine, read in normal order. Both call themselves 95 percent accurate.
Before Foretell, a radiologist read every mammogram in the order it arrived. No tiers, no tags, turnaround held near four days, week after week. With Foretell, the truly urgent cases get read the same day, and the tool earns its keep there. What matters is the other side: at 95 percent sensitivity for Priority-tagging, 5 in every 100 real cancers get tagged Routine instead, and enter the same queue as everyone else's screening.
Solandra Health's clinics screen about 42,000 women a year. At the usual population rate for screening, about 5 cancers turn up per 1,000 women, so roughly 210 real cancers a year. Five percent of 210 is about 10. Ten real women a year, tagged Routine instead of Priority.
Compare that to Hummingline. A wrong song there costs Tessaly Braelynn about four seconds on her commute: she skips it, or taps thumbs-down and the stream adjusts instantly. Nobody is harmed, because the person who eats the miss is also the person with full power to notice and fix it, right then.
What it costs at its worst: Cosima Achterloo's mammogram showed a small, genuinely hard-to-call mass, borderline enough that Foretell's own model sat right on the edge of its cutoff. Tagged Routine, her file waited 21 days instead of the 4 the old queue used to promise. What a day-4 read would likely have caught early enough for surgery alone had grown, by day 21, into something that needed chemotherapy too.
What I would leave alone: Hummingline's 95 percent doesn't need any of this. A skipped song costs four seconds and self-corrects the instant it happens. Chasing that number higher would mean narrowing the model's willingness to try a new-but-plausible song, trading real discovery for a harm that, at four seconds a miss, was never actually there.
The lesson: a success rate isn't one number, it's a question about who's standing where the miss lands. Ask that question before you ask what the number is.
Now here is the same thing as a story
Read the short version above when you're actually in the interview chair. Read this one when you want to feel exactly what 21 days cost, next to 4.
Priyamvada Vesterling can read Foretell's monthly performance report faster than almost anyone at Solandra Health, and for the first year, she read every part of it, not just the top line.
She joined the product team six months before Foretell launched, back when everyone agreed on one thing: the tool would only work if a Priority tag genuinely meant faster hands on the file, not just a label on a screen. To make sure of it, she pulled the Routine-queue turnaround number by hand every single month for the first year, checked against the sensitivity dashboard everyone else watched.
For a long while, both numbers agreed with each other. Sensitivity held near 95 percent. Routine turnaround held near four, then five, then six days, boring and stable. Priyamvada's monthly pull started feeling like busywork. She let it slide to a quarterly glance. Then to a question she'd ask herself, half out loud, whenever the sensitivity chart moved: is Routine still fine? It always looked fine, whenever she checked.
Nobody can point to the week it stopped being fine. Radiologists reading Priority cases got noticed, praised even, in the weekly huddle, for same-day turnaround on the highest-risk files. Nobody in that huddle ever mentioned Routine, so nobody's day got built around protecting it. Dr. Radu Stanescu, who reads both tiers, did what any reasonable radiologist would: cleared Priority first, every morning, and let Routine wait for whatever hours were left. Four days crept to six. Six to nine. Nine to fourteen. Nobody decided this. It just happened, a little at a time, in the same direction, for five months.
Cosima Achterloo's mammogram went into that queue in month five. It showed a small, genuinely hard-to-call mass, borderline enough that Foretell's own model sat right on the edge of its cutoff. Tagged Routine. She had no way to know a tag existed, let alone that hers had landed on the wrong side of it. She went home the way she always did after a screening, expecting nothing for a couple of weeks.
She waited twenty-one days for a queue that, a year earlier, would have read her file in four.
Dr. Stanescu read her file on day 21 and called her back that same afternoon. Nobody in that reading room did anything careless. He read every file in the order the system handed it to him. The system just never told him, or her, that the order had quietly stopped being safe.
Cosima's file went into Solandra Health's incident review as a routine delay, one of several that quarter. Nothing in that review asked what Foretell had said about her file first, because nothing in her chart recorded it. She never saw a score. Never saw a queue label. Never had anything to ask a question about, because as far as she knew, a screening just takes a few weeks sometimes.
The decision Priyamvada would take back reaches back to the week Foretell launched, when the team decided Priority and Routine would share one rota, and Routine's wait would just be whatever time was left. That made sense the day 95 percent looked like a number that spread the risk evenly. It stopped making sense the moment anyone understood what the missing five percent actually was.
Run the same five months again, with two changes in place. Every mammogram, Priority or Routine, gets read inside five business days, full stop, independent of the tag. And Foretell's own confidence score gets audited separately every month, not folded into one blended number. Cosima's file still lands on the Routine side of a genuinely close call, because nothing hits 100 percent. But the five-day ceiling means Dr. Stanescu reads it on day 5, not day 21, while the finding is still what a day-4 read would have caught.
One design let a tag quietly decide how long an actual risk could wait. The other lets the wait have its own floor, no matter what the tag says.
What I would tell myself, the day Foretell launched: a tool built to protect the most urgent five percent still needs someone watching what happens to the other ninety-five, because that's exactly where urgency goes to disappear if nobody's looking.
GUARD: the same 95 needs a different verdict depending on who's holding the other 5
Not a checklist for spotting risk in general. GUARD forces you to say, specifically, who eats a miss and whether they can do anything about it, before you're allowed an opinion on the number.
And if you want to be sure it really works, try it somewhere else
Same five letters, a county benefits office instead of a screening clinic, and the missing five percent is a family's rent instead of a diagnosis.
Cedarhollow County's Department of Human Services runs Priorly, a tool that reads an incoming application for emergency food assistance and decides whether it qualifies for Expedited processing, three business days by law, or the ordinary Standard queue, thirty days. Marisabel Duquette, one of Cedarhollow's caseworkers, trusts Priorly's tag the way any busy caseworker would: she works Expedited first, because that's the queue the state audits. Priorly runs at 95 percent sensitivity for flagging genuine emergency need. The missing 5 percent doesn't spread evenly across every application. It lands on the small share of genuinely urgent households a borderline call sent to Standard instead, like Solveig Halcott's, filed the week she lost her job.
Same rank, different lever, mapped straight onto GUARD: the groups are Solveig, who never sees Priorly's tag, and Marisabel, who works the visible queue first because that's the one the state grades her on. The harm concentrates entirely on the households a borderline call sent to Standard, not spread across the whole caseload. Solveig has no way to know a tag decided her case could wait, and no channel to ask about it, it just reads to her as a slow county office. The fix is the same shape: cap Standard's real wait independent of the tag, and audit outcomes by Priorly's own confidence band instead of one countywide accuracy figure. Detecting it means watching Standard's own turnaround climb, the same leading signal that would have caught Solandra Health's queue five months early.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to it: cap the wait independent of the tag, audit by confidence band, never trust one blended number.
Cost: no budget for more caseworker hours this year. Ship the wait cap first, since it costs schedule discipline, not headcount, then add the confidence-band audit once budget allows.
The model got better, for real: say Priorly's sensitivity jumps to 99 percent. Keep the wait cap anyway. A smaller miss still deserves a floor, because the harm on the family who hits it hasn't gotten any smaller.
Where people run it wrong.
They watch one countywide accuracy number and call the rollout a success.
They let the tag decide who gets caseworker attention, instead of only deciding read order.
They treat "we'd probably catch it in the next audit" as a plan, instead of pricing what waiting for that audit actually costs a family.
How to use it live. Ask the coverage question before naming a fix: "does a lower tag ever change how long someone actually waits, or does it only change the order they're seen in?" That question alone tells you whether the harm is contained or compounding.
Three things worth stating directly, since the real judgment sits here. The alternative Solandra Health could have taken instead of the wait ceiling was a quarterly audit of a random sample of Routine-tagged cases. It lost, because an audit only finds a pattern after it's already cost a quarter's worth of women their head start; it detects the harm, it doesn't stop it. The AI-specific failure worth naming is automation complacency on a probabilistic score: once Foretell's blended number stayed near 95 percent, Solandra Health's staffing decisions treated that score as settled fact instead of a call the model itself knew was often close, and nothing in the rollout ever asked Foretell how sure it actually was, case by case. The guardrail is the wait-time floor plus the confidence-band audit, so a close call can never silently turn into an unbounded wait. And the trade-off is real: raising the threshold that catches more Priority cases pushes meaningfully more borderline scans into the Priority queue, more radiologist hours, more cost per screening episode, accepted on purpose, because the alternative is finding out from a callback on day 21 instead of a number that moved on day 40.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Doesn't a wait-time ceiling independent of the tag defeat the whole point of triage?" Response: No, it only bounds the worst case. Priority cases still get read same-day; Routine just stops being allowed to drift past a safe floor.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on What changes when the product is probabilistic
- #1 Name three product decisions that change when a feature's output is probabilistic rather than deterministic.
- #2 A traditional feature either works or has a bug. Explain why that framing breaks for an LLM feature.
- #3 What does 'correct' mean for a summarization feature? Give a definition your engineering team could test against.
- #4 QA files a bug that reads: the model gave a wrong answer once. How do you triage it?
- #5 Explain the difference between a defect and an acceptable error rate to a non-technical executive.
- #6 Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?