CaseAdvancedResponsible AI & Advanced Practice / AI product case study teardowns / #18

Examine how an AI product changed after a major model upgrade.

LEAD the product is Vantpoint, an AI assistant that flags likely-critical scans for radiologists

Vantpoint scans a radiology image and flags the ones likely to show something critical, so a radiologist reviews those first. Dr. Fennimore Adisa reads scans at a mid-size regional hospital, and her habit of spot-checking Vantpoint's "clear" pile is the whole reason this story has an ending worth telling.

The direct answer
Don't judge the upgrade by its headline sensitivity number alone. Watch radiologists' own spot-check rate on scans Vantpoint marked clear, since that number started falling within three weeks of the upgrade, nine weeks before any missed case showed up in the outcome data.
Do this, in order
  1. Track radiologists' spot-check rate on cleared scans as the real leading signal.Why: it moved weeks before any outcome number would have shown a problem.
  2. Set a decision threshold: below a set spot-check rate, require a mandatory audit of a fixed sample.Why: a metric nobody acts on is just a chart nobody reads.
  3. Watch for the metric being gamed by a rushed, in-name-only spot-check.Why: a compliance quota with no real review defeats the entire point of tracking it.
  4. Keep celebrating the genuine sensitivity gain, don't treat the upgrade as bad news.Why: the model really did get better. The behavior problem is a separate, human-side issue.
  5. Leave manual review alone for scans Vantpoint actively flags as likely-critical.Why: those already get full attention, the risk lives entirely in the cleared pile.
  6. Don't roll back the upgrade over this.Why: the fix is restoring a habit, not reverting a genuinely better model.

How to answer this, stage by stage

Seven moves. The second one reframes the whole question before any numbers show up.

Stage 1
Scope it to one real upgrade
Say it like this
"I'll examine one specific change: Vantpoint's move from its v2 to v3 model, at one hospital, not model upgrades in general."
Why this works
A concrete before-and-after keeps the question answerable instead of abstract.
Stage 2
Reframe: a better model is still a perturbation
Say it like this
"Good news is still a change. The model got better, and that's exactly the kind of change that can cause people to stop checking it entirely."
Why this works
Most candidates only look for problems after bad news. This is the sharper read.
Stage 3
Say your structure out loud
Say it like this
"I'll use LEAD: link, the outcome that matters. Early signal, what moves first. Abuse, how it gets gamed. Decision, what I'd actually do at each level."
Why this works
Signals you're hunting for the leading number, not just repeating the headline metric.
Stage 4
Name the early signal
Say it like this
"Sensitivity went from 82% to 94%, genuinely good news. But radiologists' spot-check rate on cleared scans dropped from 15% to 2% in the first three weeks, and that's the number that actually predicts a missed case."
Why this works
This is the LEAD answer itself, the number that would look healthy right up until the morning it wasn't.
Stage 5
Prove it with the compressed failure
Say it like this
"Week eleven, a genuinely rare finding sits in Vantpoint's cleared pile, missed, since almost nobody was spot-checking cleared scans by then. It's caught two weeks later, at a follow-up visit, not by the hospital's own process."
Why this works
A short, real failure makes the leading indicator's value impossible to argue with.
Stage 6
Name the decision at each threshold
Say it like this
"If the spot-check rate falls under 8%, I'd trigger a mandatory audit of forty cleared scans that week, not just a reminder email."
Why this works
A metric nobody acts on is decoration. This makes it a real trigger.
Stage 7
Close on the direct answer
Say it like this
"The model genuinely improved. What changed for the worse was a human habit, and that's the number worth watching after any upgrade like this."
Why this works
Restates deliverable 0 plainly, closing on the actual judgment being tested.

Let's learn

Vantpoint reads a radiology scan and flags the ones most likely to show something critical, so a radiologist reads those first, before working through the rest of the day's scans in order.

Before the v3 upgrade, Vantpoint caught about 82% of genuinely critical findings, so radiologists like Fennimore kept a habit of spot-checking a random 15% of the "cleared" pile too, just in case.

Knowledge spark: what's sensitivity? How much of the real, genuinely critical findings a tool actually catches, out of all the ones that were really there. High sensitivity means it misses very few. It says nothing about what a person does with the ones it does miss.

With v3, sensitivity jumped to 94%, a real, substantial improvement that showed up clearly in every monthly accuracy report the hospital ran.

Vantpoint's sensitivity, before and after the v3 upgrade
100% 50% 0 82% v2 (before) 94% v3 (after)
A genuinely good number, and exactly the kind of good news that made everyone stop watching the number that actually mattered.

The turn: the accuracy gain was never the risk. The risk was what radiologists quietly stopped doing the moment the headline number looked this good.

At its worst: the hospital reports a record accuracy quarter to its board, while the actual habit that used to catch Vantpoint's rare misses has quietly disappeared, months before anyone connects the two.

The decision I would take back We reported sensitivity as the single headline metric in every upgrade announcement, since it was the clearest, most demoable number to share. That made sense when sensitivity and radiologist behavior moved together. It stopped making sense once a genuinely better model made radiologists trust it enough to quietly stop spot-checking, exactly the behavior the sensitivity number can't see.

What I would leave alone: scans Vantpoint actively flags as likely-critical still get full radiologist attention, no change needed there. The entire risk lives in the pile Vantpoint marks clear.

The upgrade didn't just make the model better. It made radiologists trust the model faster than the model earned that trust on the rare cases.

The lesson: a genuinely better model is still a perturbation. The headline number that improves is rarely the number that predicts what happens next.

Now here is the same thing as a story

Use the short version above under time pressure. Read this one for how close the missed case actually came.

Fennimore has read scans for twelve years and built her own quiet habit years before Vantpoint ever existed: always glance at a handful of the "normal" pile, never just the flagged ones.

Hand sketched flow diagram titled The scan review pipeline. Four boxes: scan taken, Vantpoint flags it, radiologist spot checks clear ones highlighted, report signed.
The spot-check step, third in line, is the one that quietly emptied out after the upgrade.

The weeks after v3 launched felt genuinely great. Vantpoint's flags were sharper, its false alarms dropped, and Fennimore's own caseload of confirmed criticals per week held steady even as her total review time shrank.

Hand sketched metaphor comparison titled Two clocks one rings first. Left panel a gauge icon labeled Sensitivity, caption rings weeks late. Right panel a document icon labeled Spot check rate, caption rings first.
Sensitivity is the clock everyone was watching. It was never going to ring first.

The habit thinned quietly, in three small steps nobody wrote down. First, Fennimore spot-checked ten cleared scans a day instead of fifteen. Then five. Then, some days, none at all, since the flagged pile alone already felt thorough enough to fill her attention.

Hand sketched timeline titled Upgrade to first missed case. Three milestones: v3 ships week 1, spot checks fall highlighted week 3, first missed case week 11.
Eight weeks separate the moment the spot-check habit collapsed from the moment anyone noticed a real problem.

Week eleven, a rare, atypical presentation of a genuinely critical finding sat in Vantpoint's cleared pile. Nobody spot-checked that batch that day. It surfaced two weeks later, at a routine follow-up, not through the hospital's own review process.

Hand sketched quadrant titled Sorting the metrics that could have warned us. Axes how soon it moves from late to early, and how visible on a dashboard from hidden to obvious. Sensitivity sits upper left. Spot check rate sits lower right. Missed case reports sit far upper left. Radiologist survey sits middle.
The metric that would have warned everyone earliest is exactly the one that never made it onto a dashboard.

The review afterward found nothing wrong with Vantpoint's model. What it found was a hospital-wide spot-check rate that had fallen from 15% to 2% within three weeks of the upgrade, a chart nobody had been keeping.

Hand sketched labeled parts diagram titled What the post upgrade dashboard needs. Center gauge icon labeled Dashboard, four callouts: spot check rate, sensitivity trend, missed case log, alert on rate drop.
Three of these four already existed somewhere. Only the alert on a falling spot-check rate had to be built new.

The old dashboard asked the hospital to trust that a rising sensitivity number meant everything downstream was fine. The new one tracks the spot-check rate directly and triggers a mandatory audit the moment it falls too far.

I picked sensitivity as the one number to report every quarter because it was the cleanest, most shareable win. It took one rare finding sitting unspotted for two extra weeks to see that the number I was watching had never been built to catch what radiologists quietly stopped doing.

LEAD, in one screenNot a dashboard of everything. LEAD is what finds the one number that would have rung first.

Hand sketched icon list titled LEAD in one screen. Four items: a gauge icon labeled Link the outcome that actually matters, a document icon labeled Early signal what moves first, a question mark box icon labeled Abuse how it gets gamed, a scale icon labeled Decision what you'd do at each level.
Four letters, and the early signal is the one a headline-metric answer always misses.
L
Link.
The outcome that matters is catching every genuinely critical finding, not the model's own sensitivity score in isolation.
Grounds the metric in patient safety, not in what looks good on a report.
E
Early signal.
Radiologists' own spot-check rate on cleared scans, which fell from 15% to 2% within three weeks, nine weeks before any missed case surfaced.
The hardest, most valuable step: a number that would have looked perfectly fine on any accuracy dashboard.
A
Abuse.
A radiologist could log a spot-check as done without actually reviewing the scan closely, just to satisfy a quota.
Every metric has a way to be hit without doing the underlying work.
D
Decision.
Below an 8% spot-check rate, trigger a mandatory audit of forty cleared scans that same week, not a reminder email.
Turns the metric into an actual action, not a chart nobody acts on.
Radiologist spot-check rate on cleared scans, weeks since the v3 upgrade
16% 8% 0 Wk11: missed case Wk0 Wk6 Wk9
By week nine, the leading signal had already collapsed to a sixth of its starting rate, two weeks before anything showed up in outcome data.

The recap, one line per letter: link is catching every real critical finding, early signal is spot-check rate collapsing weeks ahead of any missed case, abuse is a hollow, rushed spot-check logged just to hit a number, and decision is a mandatory audit triggered automatically below 8%.

And if you want to be sure it really works, try it somewhere elseSame four letters, a waste-management sorting line instead of a hospital.

Sableton runs an AI vision system that sorts recyclables on a conveyor line, flagging contaminated batches for a human sorter to pull before they reach the baler. Bellamy Grennock oversees quality control at a regional materials recovery facility.

Mapped onto LEAD after Sableton's own major model upgrade: link is keeping contamination out of baled recyclable loads, not the model's raw sort-accuracy number. Early signal is how often human sorters manually re-check a batch the model marked clean, which fell sharply the week accuracy jumped, well before any contaminated bale actually got rejected by a buyer downstream. Abuse: a sorter could tap "checked" on a batch without actually pulling it off the line to inspect. Decision: below a set manual re-check rate, the facility automatically routes a larger random sample of "clean" batches to a slow-lane manual inspection station for that shift.

Hand sketched labeled parts diagram titled What the post upgrade dashboard needs, reused for the recycling sorting line example. Center gauge icon labeled Dashboard, four callouts: spot check rate, sensitivity trend, missed case log, alert on rate drop.
Swap "spot-check rate" for "manual re-check rate" and the same dashboard gap holds for a sorting line.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "watch the human spot-check rate, not the accuracy number, it collapsed weeks before any missed case," and stop.
Cost: there's no budget this quarter to build a new dashboard metric. Say so honestly, and start with a simple weekly manual count of spot-checks logged, even without automated tracking yet.
The model gets worse, for contrast: if Vantpoint's next update genuinely regressed instead of improving, the risk flips to substitution or workaround behavior, radiologists rationing trust toward easy cases, a completely different leading signal to watch.

Where people run it wrong.
They treat a model upgrade as purely good news and stop looking for any new risk it might introduce.
They report the headline accuracy metric as the whole story, missing that it can't see a human habit disappearing underneath it.
They build a metric and never attach a real decision to it, so it just sits on a dashboard nobody acts on until it's too late.

How to use it live. When someone describes a product after a big model upgrade, ask yourself: what human behavior might this improvement quietly end. That behavior, not the upgrade's own scorecard, is usually the real story.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "examine how this product changed after a major model upgrade"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It's a metric question about which number reveals the real change first.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dr. Fennimore Adisa, a radiologist of twelve years who had her own long-standing habit of spot-checking cleared scans.
3 · THE HEADLINE NUMBER
What did the model's sensitivity actually do after the upgrade?
Tap to flip
ANSWER
It genuinely improved, from 82% to 94%. The model really did get better.
4 · THE EARLY SIGNAL
What's the actual leading indicator in this story?
Tap to flip
ANSWER
Radiologists' spot-check rate on scans marked clear, which fell from 15% to 2% within three weeks of the upgrade.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Reporting sensitivity as the single headline metric in every upgrade announcement, since it was the clearest number to share.
6 · THE NUMBER
Fill in the blank: the missed case surfaced at week 11, but the spot-check rate started falling by week ___.
Tap to flip
ANSWER
Week 3. An eight-week gap between the leading signal moving and the outcome anyone actually noticed.
7 · THE DECISION RULE
What specific action fires when the spot-check rate falls too low?
Tap to flip
ANSWER
Below 8%, a mandatory audit of forty cleared scans triggers that same week, not just a reminder.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's its equivalent early signal?
Tap to flip
ANSWER
Sableton's recycling sorting line. Its early signal is the human sorters' manual re-check rate on batches the model marked clean.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer treats the v3 model upgrade itself as a mistake that should be reversed.
  • True
  • False
Show hint
Look at the priority list and "what I would leave alone."
Show answer
False. The model genuinely improved. The problem is a human habit that quietly disappeared, and the fix restores the habit, not the old model.
Multiple choice
2. Why does this answer treat the spot-check rate as more useful than the sensitivity number?
  • A. Sensitivity is harder to measure accurately.
  • B. The spot-check rate fell weeks before any missed case appeared in outcome data, while sensitivity looked good the whole time.
  • C. Radiologists don't trust the sensitivity metric.
  • D. Sensitivity only applies to non-critical findings.
Show hint
Look at the "early signal" step and the line chart.
Show answer
B. LEAD's whole job is finding the number that moves before the lagging outcome does.
Fill in the blank
3. Fill in the blank: the spot-check rate fell from 15% to ___% within three weeks of the upgrade.
Show hint
Look at the direct answer and the line chart.
Show answer
2%. A drop to roughly a seventh of its starting rate, nine weeks before the first missed case surfaced.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Reporting sensitivity as the single headline metric, which made sense while it and radiologist behavior moved together, before the upgrade made trust outrun accuracy.
Short answer, apply it yourself
5. Pick a product you use that recently got a big update or improvement. What's one thing you might have quietly stopped double-checking because it got better?
Show hint
Think about an app update, a browser feature, or a tool at work that got noticeably more reliable.
Show answer
Model answer: Many people stop proofreading autocorrect once it feels reliable, right up until it silently changes a word they never meant to send.
Short answer, where it wouldn't matter
6. Name a part of Vantpoint's workflow where this exact risk, over-trust after an upgrade, genuinely wouldn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Scans Vantpoint actively flags as likely-critical. Those already get full radiologist attention regardless of the upgrade.
Before you close the answer
Why this works
Tests whether you treat a model improvement as automatically safe, or whether you can find the human behavior a genuinely better model can still quietly damage, and whether you'd catch it with a real leading metric instead of a headline number.
Follow-up traps
"Isn't tracking spot-check rate just adding more process for no real benefit?" Response: it's one number, tied to one clear trigger, and it would have surfaced this exact risk eight weeks earlier than it actually did.

"What if radiologists just game the spot-check metric by clicking through without really looking?" Response: that's the abuse risk named directly, which is why the decision rule includes an actual mandatory audit of real scans, not just a self-reported count.
If pressed
Vantpoint's v3 upgrade also changed its confidence calibration internally, meaning its own reported confidence scores got quietly more conservative even as raw sensitivity rose, which is a separate detail worth knowing but doesn't change the fact that no radiologist-facing signal ever surfaced that shift.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more