ConceptAdvancedAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #13

What should you ask about a vendor's eval methodology?

LEAD the aggregate number that never moved while one subgroup quietly fell to 71

Vantage Radiology AI sells a tool that flags likely fractures on overnight X-rays. Priya Nataraj buys and manages imaging software for Cascadia Health, a five-hospital network. Renata Vasilenko runs the same kind of job two states over, at a sibling network.

The direct answer
Before you trust any vendor's accuracy number, ask to see their eval set's refresh cadence and a breakdown by patient subgroup and hardware, not just the aggregate score. If they cannot show a held-out set that gets rebuilt after every model update, split by the kinds of cases your own site actually has, their number is a slide, not evidence.
Do this, in order
  1. Demand a subgroup and hardware breakdown, not just the aggregate.Why: an aggregate can look great while one small, high-stakes slice quietly fails underneath it.
  2. Ask how often the eval set gets refreshed, and by whom.Why: a static set gets quietly outrun by a model that's been retuned since, so an old number stops meaning anything.
  3. Ask what a low-confidence case looks like, and whether that gets reported.Why: a vendor with no uncertainty signal gives you no way to know which flagged cases deserve a second look.
  4. Run your own shadow period on your hardest, most representative cases before going live everywhere.Why: your patients and your equipment are not the vendor's benchmark, and the gap only shows up on your own data.
  5. Write a real threshold into the rollout: a mandatory second read when any subgroup falls too far behind.Why: a number with no action tied to it is a dashboard decoration, not a safeguard.

How to answer this, stage by stage

Nobody is scoring you on whether you can say "verify vendor claims." They're scoring whether you can name the specific question that would have caught this before a patient nearly paid for it.

Stage 1
Scope it to one vendor claim
Say it like this
"I'll answer this for one vendor, Vantage Radiology AI, and one published number: 96 percent accuracy on fracture detection."
Why this works
Stops the question from turning into a generic "always audit your vendors" speech.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link to the outcome that actually matters, find the early signal, name how the metric gets gamed, then say what I'd decide at each threshold."
Why this works
Shows a plan before diving in, so the answer doesn't sound like you're improvising.
Stage 3
Reframe what the question is really testing
Say it like this
"The real question isn't whether 96 percent is good. It's whether that one number is the whole truth, or just the part of the truth the vendor chose to measure."
Why this works
Separates a strong candidate from one who just repeats the vendor's number back at the interviewer.
Stage 4
Give the one decision
Say it like this
"Ask for the eval set's refresh cadence and a breakdown by subgroup and hardware. If they can't show both, the 96 percent is a marketing slide, not a reason to sign."
Why this works
This is the direct answer, said plainly, before any story backs it up.
Stage 5
Prove it with the near miss
Say it like this
"Their aggregate number never moved, it stayed at 96 percent the whole time. But on our own older X-ray units, on kids' growth-plate fractures specifically, true accuracy had quietly dropped to 71. A nine-year-old almost went home on it."
Why this works
Turns "ask for a breakdown" from a nice-to-have into the one question that would have caught real harm before it happened.
Stage 6
Name the threshold you'd act on
Say it like this
"If any subgroup's sensitivity falls more than 10 points below the vendor's aggregate number, that subgroup gets a mandatory second human read until the gap closes."
Why this works
A real number needs a real action tied to it, or it's just a chart nobody actually uses.
Stage 7
Close on the one line
Say it like this
"Don't buy the aggregate. Buy the breakdown, the refresh cadence, and a plan for the day one slice of it quietly falls behind."
Why this works
Restates the direct answer in one breath, ready to survive a follow-up.

Let's learn

What happens when a vendor's benchmark number is the only evidence you have?

Say a hospital network buys a tool that reads overnight X-rays and flags the ones that likely show a fracture, so a tired resident knows which film to open first. Before it, every overnight film at Cascadia Health sat in one long queue, read in the order it landed, whether it was a hairline wrist crack or a dislocated shoulder. A priority case took about 40 minutes to reach a radiologist's eyes, because nothing sorted the queue by urgency.

Hand sketched labeled parts diagram titled What a real eval spec holds. A document icon at the center labeled Eval Spec, with four callouts around it: held-out set, refresh cadence, subgroup split, failure disclosure.
This is what a real eval spec has to answer for. A cover slide with one number answers none of it.

The vendor's own materials said the model caught 96 out of every 100 fractures. Priya's team believed it, the same way five other hospital networks already had. With the flag running, a likely fracture reached a radiologist in about 9 minutes instead of 40.

Here's the turn: the 4 missed fractures out of every 100 in that aggregate number were never the real risk. The real risk was what the aggregate was hiding. On the network's older Corvalen-brand X-ray units, and specifically on kids' growth-plate fractures, the model's true accuracy had quietly slipped to 71 out of 100. Nobody knew, because nobody had ever asked the vendor to break the number apart.

True sensitivity on the hardest subgroup, by quarter, against the vendor's frozen aggregate
100% 75 50 vendor's aggregate: 96% Q1 Q4, near miss Q5 93% 71%
The aggregate never moved. A monthly subgroup audit would have shown this falling for a year before the near miss happened.

At its worst: a nine-year-old with a subtle growth-plate fracture, imaged on a Corvalen unit, got flagged "clear." An overnight resident, tired, trusting the flag the way the vendor's number told him he could, sent the child home with an ice pack. The swelling got worse overnight. The family came back the next morning, and a repeat film caught it. No lasting harm. But it was close, and it was close for a reason nobody had gone looking for.

The choice I would take back Priya's team signed the rollout on the vendor's aggregate accuracy number and the fact that other hospital networks had already adopted it. That made sense at the time: the deck looked thorough, and asking for more felt like slowing down a tool that was clearly saving time everywhere else. It stopped making sense the moment one subgroup, on one piece of hardware, needed a very different number than the one on the cover slide.
Knowledge spark: what is a held-out eval set? A held-out set is a pile of cases the model never trained on, kept aside just to test it honestly. If a vendor tests on cases the model already saw during training, the number is inflated, the same way studying the exact questions on an exam beforehand inflates a grade.

What I would leave alone: the flag itself, for the vast majority of studies that look nothing like this subgroup, doesn't need this level of suspicion. A grown adult's clean forearm break on the network's newer units is exactly the case the vendor's aggregate number describes well, and slowing that down with extra review would cost time for no real safety gain.

The lesson: a single accuracy number is never exactly wrong. It's just never the whole answer, and the part it leaves out is usually the part that matters most to the patient least like the vendor's average case.

Now here is the same thing as a story

The short version above is what you'd say defending this to Cascadia's board. Read this one for how close it actually came.

Priya Nataraj can read a radiology vendor's slide deck and tell you within two slides whether the eval section is padding or substance. She's spent eleven years buying and un-buying software for Cascadia Health's five hospitals, and she has a rule: if a benchmark slide doesn't say where the test cases came from, she assumes the answer is wherever made the number look best.

Vantage Radiology AI's deck passed her rule. It named its test set, 40,000 studies, and its overall sensitivity, 96 percent, with a citation to a peer-reviewed paper. Five other hospital networks had already signed. Priya signed too, that spring, for a tool that would flag likely fractures on overnight X-rays so a resident knew which film to open first instead of reading the queue cold.

For the first several months, it did exactly what the deck promised. A likely fracture that used to sit in a 40-minute queue now reached a radiologist's screen in about 9 minutes. Priya's team stopped asking Vantage for updates on the number. It was 96 percent at signing, and nobody had a reason to think it had moved.

Hand sketched flow diagram titled How a stale eval set drifts unnoticed. Five boxes in sequence: Vendor builds set, Model gets updated, Set stays the same, Gap quietly widens (highlighted), Patient feels it.
Vantage kept updating the model. Nobody ever asked whether the test they'd graded it on had kept up.

It had moved. Not for most patients. For one narrow slice: studies from Cascadia's older Corvalen-brand units, still running at two of the five hospitals, and specifically for growth-plate fractures in children, a small share of any hospital's overnight volume. Vantage's own held-out set had only ever included 4 cases like that in every 100, so a model that got a little worse there barely nudged the aggregate at all.

Renata Vasilenko, a product manager at a sibling network two states over, mentioned it almost in passing, over coffee at a vendor conference: "you're still running the version from the original bake-off, right? We pulled our own numbers apart by hardware last spring. Our Corvalen units were ugly." Priya laughed it off. Then she went home and pulled Cascadia's numbers apart the same way.

Hand sketched comparison titled How the metric gets gamed. Left, a document icon labeled 96 percent, caption looks perfect on the cover slide. Right, a person icon labeled Missed patient, caption walks out with a real fracture.
The cover slide and the ten-year-old are both true at the same time. That's what a hidden subgroup does.

They were ugly too. Worse than ugly on the growth-plate subgroup: 71 out of 100, against a vendor number that had never once moved off 96.

Three weeks after that conversation, a nine-year-old came into one of the two hospitals still running Corvalen units, with a wrist that hurt more than it looked. The flag came back clear. The overnight resident, tired, trusting the tool the way its number told him he could, sent the child home with an ice pack.

We did not lose a subgroup average. We nearly lost a family's trust in an emergency room, over a number that was true for almost everyone and false for exactly the wrong nine-year-old.

The child came back the next morning, swelling worse. A repeat film, read cold by a day-shift radiologist who never saw the original flag, caught the fracture. Nothing permanent. But Priya spent that afternoon on the phone with Vantage, and the aggregate number stopped being the only number that mattered.

Hand sketched quadrant titled How much should you trust the number. Axes, sample recency from stale to refreshed, and subgroup detail from aggregate only to broken out. Vantage's slide sits low on both axes. Vantage's real report sits low on recency but high on subgroup detail. A rigorous vendor sits high on both.
The cover slide and the real report were the same vendor. Only one of them was worth signing on.

Here's what I'd take back. When Priya's team reviewed the contract, they treated the vendor's aggregate accuracy and the fact that other networks had already adopted it as enough evidence to sign. That was a reasonable read of a professional deck. It stopped being reasonable the day it turned out the deck's one number was quietly hiding a second one that mattered more to a specific nine-year-old than any average ever could.

I would go back and demand the subgroup breakdown and the refresh cadence before signing, not after a child's swelling forced the question. That's the whole difference. One design trusts a single number because it's the only one on the slide. The other insists there's always a second one hiding under it, and asks to see both before anyone signs anything.

And the part I'd tell my past self: we didn't buy a bad tool. We bought a good tool and never asked it the one question that would have shown us exactly where it was quietly getting worse.

LEAD, in one screenNot "is the number high." LEAD is what tells you whether the number is the whole story or just the flattering half of it.

L
Link. The outcome that actually matters.
Not "is the vendor's number high." Whether a specific patient, on specific hardware, gets caught when it counts, every overnight read, not the network-wide average.
Without this, "measure it" has no target to aim at.
E
Early signal. What moves first.
The subgroup sensitivity gap, tracked quarter over quarter, moved for a year before any patient outcome did. Corvalen-unit growth-plate sensitivity fell from 93 to 71 while the vendor's aggregate sat frozen at 96 the entire time.
This is the hardest step, and the one this whole answer turns on.
A
Abuse. How the metric gets gamed.
A small, high-stakes subgroup can fall apart underneath a healthy-looking aggregate, because it's too small a share of volume to move the average. Vantage's held-out set had only 4 cases like this in every 100.
Not fraud, just a number too coarse to see the thing that mattered.
D
Decision. What you'd actually do.
Subgroup within 10 points of the aggregate, keep going. More than 10 points behind, mandatory second human read on that subgroup until the gap closes. This is the change Cascadia made after the near miss.
A metric nobody acts on is a dashboard decoration, not a safeguard.
Hand sketched icon list titled Ask a vendor these before you sign. Five items: where do test cases come from, how often is the set refreshed, show the subgroup breakdown, what does low confidence look like, who owns the pass bar.
Five questions. Vantage's original deck answered none of them, and nobody noticed until a near miss forced the question.

The recap, one line per letter: link is which patient actually matters, early signal is the subgroup gap widening for months before anyone noticed, abuse is a small subgroup hiding under a healthy average, decision is a real threshold with a real action tied to it.

Hand sketched decision tree titled What Priya does with the vendor's answer. Root, vendor's eval answer. Three branches: refreshed and broken out leads to start a pilot, aggregate number only leads to run your own shadow eval, vendor won't say leads to walk away.
Most of the actual judgment call happens here, not in reading the deck.

And if you want to be sure it really works, try it somewhere elseSame four letters, a wildfire-risk model for home insurance instead of a fracture flag. A different vendor, the same hidden subgroup.

Meline Insurance Group buys a wildfire-risk scoring model from a vendor to help price home insurance premiums. Tobias Ekwueme runs the underwriting technology team. Mapped onto LEAD: link is whether a specific home's actual defensible space and roofing material predict a future claim, not whether the model matches old claims patterns overall. Early signal is the vendor's scoring accuracy specifically on rural and exurban properties, a small share of the vendor's total portfolio, which drifted downward for two years while the aggregate stayed flat. Abuse is the same shape as Cascadia's: a small segment can fall apart underneath a healthy aggregate because it barely moves the average. Decision is a real threshold: any property-density tier more than 8 points below the aggregate triggers a manual underwriter review before binding a policy.

The old decision here isn't an unread deck, it's a different reversal: Meline's fire-safety team reviewed the vendor's validation report once, two years earlier, and never re-asked when the model was updated. That made sense when the review was fresh. It stopped making sense the moment "reviewed once" quietly became "never checked again."

Rural-property claim prediction accuracy, before and after the underwriter review threshold
100% 50 0 91% aggregate 79% rural, before Before review 91% aggregate 90% rural, after After review aggregate accuracy rural segment accuracy
The aggregate barely moved either time. The rural segment is what actually changed, and it's the one number the old process never checked.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask for the subgroup breakdown and the refresh cadence, not just the aggregate," and stop.
Cost: there's no budget to demand a custom audit from every vendor. Say so honestly, and run your own cheap shadow period on your hardest cases instead of paying for the vendor's.
The model gets better, for real: if a later model update genuinely closes the subgroup gap, that's the threshold doing its job, telling you it's safe to relax the mandatory second read, not a reason to stop watching the gap entirely.

Where people run it wrong.
They accept an aggregate number because it's the only number the vendor volunteered.
They treat "the deck looked thorough" as the same thing as "the deck showed the whole picture."
They never write a threshold into the rollout, so a real gap sits in a spreadsheet instead of triggering an actual decision.

How to use it live. The moment someone asks "should we trust this vendor's number," ask back: what does this number look like broken out by the group least like their average case? Let that answer decide, not the confidence of the slide it came from.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you know it's working" and "what should you ask about a metric" questions?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It forces you past the vendor's headline number to the thing that would have moved first.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Nataraj, eleven years buying and managing imaging software for Cascadia Health, a five-hospital network.
3 · THE HABIT
What did Priya's team stop doing once the deck looked thorough?
Tap to flip
ANSWER
Asking Vantage for updates on the accuracy number. It was 96 percent at signing, so nobody had a reason to think it had moved.
4 · THE EARLY SIGNAL
What's the signal that moved months before any patient outcome did?
Tap to flip
ANSWER
Subgroup sensitivity for Corvalen-unit growth-plate fractures, falling from 93 to 71 over five quarters, while the vendor's aggregate stayed frozen at 96.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Signing based on the vendor's aggregate number and other networks' adoption, without demanding a subgroup and hardware breakdown first.
6 · THE NUMBER
Fill in the blank: the vendor's aggregate accuracy stayed at 96 percent the whole time, while true subgroup sensitivity fell to ___ percent by the quarter of the near miss.
Tap to flip
ANSWER
71 percent, a 25-point gap the aggregate number never once showed.
7 · THE REPLAY
Same near miss, but the subgroup threshold and mandatory second read are already in place. What changes?
Tap to flip
ANSWER
The falling subgroup number trips the 10-point threshold months earlier, a human reads every Corvalen-unit growth-plate flag, and the child's fracture gets caught the same night, not the next morning.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, using the same framework. Which product, and what's the reversal?
Tap to flip
ANSWER
Meline Insurance Group's wildfire-risk pricing model. The reversal is reviewing the vendor's validation report once, two years earlier, and never re-checking it after model updates.

Check yourself Score: 0 / 0

True or false
1. True or false: the near miss happened because Vantage's model got worse across the board.
  • True
  • False
Show hint
Look at what happened to the aggregate number the whole time.
Show answer
False. The aggregate stayed at 96 percent throughout. Only one small subgroup, on one type of hardware, quietly got worse.
Multiple choice
2. Why did the subgroup failure never show up in the vendor's own aggregate number?
  • A. The vendor deliberately hid the number.
  • B. The subgroup was too small a share of total volume to move the average.
  • C. Growth-plate fractures are always easy to detect.
  • D. Cascadia's radiologists stopped reading the flagged studies.
Show hint
Look at the Abuse step in the LEAD recap.
Show answer
B. A small, high-stakes subgroup can fall apart underneath a healthy-looking average, simply because it's too small a slice to move it.
Fill in the blank
3. Fill in the blank: true subgroup sensitivity for Corvalen-unit growth-plate fractures fell from 93 percent in quarter one to ___ percent by the quarter of the near miss.
Show hint
Look at the line chart in Section 1.
Show answer
74 percent. It kept falling to 71 percent the following quarter, before the fix went in.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "The choice I would take back."
Show answer
Model answer: Signing based on the vendor's aggregate number and other networks' adoption. It made sense because the deck looked thorough and other hospitals had already trusted it.
Short answer, where it wouldn't matter
5. Name a place in Cascadia's rollout where this extra scrutiny genuinely would not matter.
Show hint
Look at "What I would leave alone."
Show answer
Model answer: A clean adult forearm break on the network's newer X-ray units. That's exactly the case the vendor's aggregate number describes well, so extra review there costs time for no real safety gain.
Short answer, apply it yourself
6. Think of a product you use that reports one overall quality number. What group of users is least like its average case, and how would you find out if the number is quietly worse for them?
Show hint
Look for a group that's a small share of total usage, since that's exactly what an aggregate number is bad at showing.
Show answer
Model answer: Ask the product for a breakdown by that specific group instead of trusting the headline number, the same move as demanding a subgroup report from a vendor.
Before you close the answer
Why this works
Tests whether you'll accept a vendor's one headline number, or go looking for the subgroup it's built to hide, and whether you can name a real threshold that turns a number into an action.
Follow-up traps
"Isn't demanding a subgroup breakdown for every vendor unrealistic?" Response: no, run your own cheap shadow period on your hardest cases when a vendor won't provide the breakdown, rather than skipping the check entirely.

"What if the subgroup is too small to ever get a reliable number?" Response: that's exactly when a mandatory human read matters most, since a small subgroup with a bad number and a small subgroup with a good number look identical without one.
If pressed
Cascadia's fix wasn't a bigger model. It was retraining specifically with more Corvalen-unit and pediatric growth-plate cases, since more clear-sky, adult-forearm data would have done nothing to close this particular gap.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more