What should you ask about a vendor's eval methodology?
Vantage Radiology AI sells a tool that flags likely fractures on overnight X-rays. Priya Nataraj buys and manages imaging software for Cascadia Health, a five-hospital network. Renata Vasilenko runs the same kind of job two states over, at a sibling network.
- Demand a subgroup and hardware breakdown, not just the aggregate.Why: an aggregate can look great while one small, high-stakes slice quietly fails underneath it.
- Ask how often the eval set gets refreshed, and by whom.Why: a static set gets quietly outrun by a model that's been retuned since, so an old number stops meaning anything.
- Ask what a low-confidence case looks like, and whether that gets reported.Why: a vendor with no uncertainty signal gives you no way to know which flagged cases deserve a second look.
- Run your own shadow period on your hardest, most representative cases before going live everywhere.Why: your patients and your equipment are not the vendor's benchmark, and the gap only shows up on your own data.
- Write a real threshold into the rollout: a mandatory second read when any subgroup falls too far behind.Why: a number with no action tied to it is a dashboard decoration, not a safeguard.
How to answer this, stage by stage
Nobody is scoring you on whether you can say "verify vendor claims." They're scoring whether you can name the specific question that would have caught this before a patient nearly paid for it.
Let's learn
What happens when a vendor's benchmark number is the only evidence you have?
Say a hospital network buys a tool that reads overnight X-rays and flags the ones that likely show a fracture, so a tired resident knows which film to open first. Before it, every overnight film at Cascadia Health sat in one long queue, read in the order it landed, whether it was a hairline wrist crack or a dislocated shoulder. A priority case took about 40 minutes to reach a radiologist's eyes, because nothing sorted the queue by urgency.
The vendor's own materials said the model caught 96 out of every 100 fractures. Priya's team believed it, the same way five other hospital networks already had. With the flag running, a likely fracture reached a radiologist in about 9 minutes instead of 40.
Here's the turn: the 4 missed fractures out of every 100 in that aggregate number were never the real risk. The real risk was what the aggregate was hiding. On the network's older Corvalen-brand X-ray units, and specifically on kids' growth-plate fractures, the model's true accuracy had quietly slipped to 71 out of 100. Nobody knew, because nobody had ever asked the vendor to break the number apart.
At its worst: a nine-year-old with a subtle growth-plate fracture, imaged on a Corvalen unit, got flagged "clear." An overnight resident, tired, trusting the flag the way the vendor's number told him he could, sent the child home with an ice pack. The swelling got worse overnight. The family came back the next morning, and a repeat film caught it. No lasting harm. But it was close, and it was close for a reason nobody had gone looking for.
What I would leave alone: the flag itself, for the vast majority of studies that look nothing like this subgroup, doesn't need this level of suspicion. A grown adult's clean forearm break on the network's newer units is exactly the case the vendor's aggregate number describes well, and slowing that down with extra review would cost time for no real safety gain.
The lesson: a single accuracy number is never exactly wrong. It's just never the whole answer, and the part it leaves out is usually the part that matters most to the patient least like the vendor's average case.
Now here is the same thing as a story
The short version above is what you'd say defending this to Cascadia's board. Read this one for how close it actually came.
Priya Nataraj can read a radiology vendor's slide deck and tell you within two slides whether the eval section is padding or substance. She's spent eleven years buying and un-buying software for Cascadia Health's five hospitals, and she has a rule: if a benchmark slide doesn't say where the test cases came from, she assumes the answer is wherever made the number look best.
Vantage Radiology AI's deck passed her rule. It named its test set, 40,000 studies, and its overall sensitivity, 96 percent, with a citation to a peer-reviewed paper. Five other hospital networks had already signed. Priya signed too, that spring, for a tool that would flag likely fractures on overnight X-rays so a resident knew which film to open first instead of reading the queue cold.
For the first several months, it did exactly what the deck promised. A likely fracture that used to sit in a 40-minute queue now reached a radiologist's screen in about 9 minutes. Priya's team stopped asking Vantage for updates on the number. It was 96 percent at signing, and nobody had a reason to think it had moved.
It had moved. Not for most patients. For one narrow slice: studies from Cascadia's older Corvalen-brand units, still running at two of the five hospitals, and specifically for growth-plate fractures in children, a small share of any hospital's overnight volume. Vantage's own held-out set had only ever included 4 cases like that in every 100, so a model that got a little worse there barely nudged the aggregate at all.
Renata Vasilenko, a product manager at a sibling network two states over, mentioned it almost in passing, over coffee at a vendor conference: "you're still running the version from the original bake-off, right? We pulled our own numbers apart by hardware last spring. Our Corvalen units were ugly." Priya laughed it off. Then she went home and pulled Cascadia's numbers apart the same way.
They were ugly too. Worse than ugly on the growth-plate subgroup: 71 out of 100, against a vendor number that had never once moved off 96.
Three weeks after that conversation, a nine-year-old came into one of the two hospitals still running Corvalen units, with a wrist that hurt more than it looked. The flag came back clear. The overnight resident, tired, trusting the tool the way its number told him he could, sent the child home with an ice pack.
The child came back the next morning, swelling worse. A repeat film, read cold by a day-shift radiologist who never saw the original flag, caught the fracture. Nothing permanent. But Priya spent that afternoon on the phone with Vantage, and the aggregate number stopped being the only number that mattered.
Here's what I'd take back. When Priya's team reviewed the contract, they treated the vendor's aggregate accuracy and the fact that other networks had already adopted it as enough evidence to sign. That was a reasonable read of a professional deck. It stopped being reasonable the day it turned out the deck's one number was quietly hiding a second one that mattered more to a specific nine-year-old than any average ever could.
I would go back and demand the subgroup breakdown and the refresh cadence before signing, not after a child's swelling forced the question. That's the whole difference. One design trusts a single number because it's the only one on the slide. The other insists there's always a second one hiding under it, and asks to see both before anyone signs anything.
And the part I'd tell my past self: we didn't buy a bad tool. We bought a good tool and never asked it the one question that would have shown us exactly where it was quietly getting worse.
LEAD, in one screenNot "is the number high." LEAD is what tells you whether the number is the whole story or just the flattering half of it.
The recap, one line per letter: link is which patient actually matters, early signal is the subgroup gap widening for months before anyone noticed, abuse is a small subgroup hiding under a healthy average, decision is a real threshold with a real action tied to it.
And if you want to be sure it really works, try it somewhere elseSame four letters, a wildfire-risk model for home insurance instead of a fracture flag. A different vendor, the same hidden subgroup.
Meline Insurance Group buys a wildfire-risk scoring model from a vendor to help price home insurance premiums. Tobias Ekwueme runs the underwriting technology team. Mapped onto LEAD: link is whether a specific home's actual defensible space and roofing material predict a future claim, not whether the model matches old claims patterns overall. Early signal is the vendor's scoring accuracy specifically on rural and exurban properties, a small share of the vendor's total portfolio, which drifted downward for two years while the aggregate stayed flat. Abuse is the same shape as Cascadia's: a small segment can fall apart underneath a healthy aggregate because it barely moves the average. Decision is a real threshold: any property-density tier more than 8 points below the aggregate triggers a manual underwriter review before binding a policy.
The old decision here isn't an unread deck, it's a different reversal: Meline's fire-safety team reviewed the vendor's validation report once, two years earlier, and never re-asked when the model was updated. That made sense when the review was fresh. It stopped making sense the moment "reviewed once" quietly became "never checked again."
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask for the subgroup breakdown and the refresh cadence, not just the aggregate," and stop.
Cost: there's no budget to demand a custom audit from every vendor. Say so honestly, and run your own cheap shadow period on your hardest cases instead of paying for the vendor's.
The model gets better, for real: if a later model update genuinely closes the subgroup gap, that's the threshold doing its job, telling you it's safe to relax the mandatory second read, not a reason to stop watching the gap entirely.
Where people run it wrong.
They accept an aggregate number because it's the only number the vendor volunteered.
They treat "the deck looked thorough" as the same thing as "the deck showed the whole picture."
They never write a threshold into the rollout, so a real gap sits in a spreadsheet instead of triggering an actual decision.
How to use it live. The moment someone asks "should we trust this vendor's number," ask back: what does this number look like broken out by the group least like their average case? Let that answer decide, not the confidence of the slide it came from.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the subgroup is too small to ever get a reliable number?" Response: that's exactly when a mandatory human read matters most, since a small subgroup with a bad number and a small subgroup with a good number look identical without one.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Evaluating AI vendors as a buyer
- #1 List the ten questions you would ask every AI vendor before a pilot.
- #2 How do you evaluate a vendor's quality claims without running your own eval?
- #3 Design the pilot you would run to evaluate two competing AI vendors.
- #4 What contractual terms matter specifically for AI vendors and not for other software?
- #5 How do you assess a vendor's model dependency and what happens if their provider changes terms?
- #6 Describe the data handling questions you would put to a vendor on behalf of your security team.