How do you test a vendor's product on your hardest cases rather than their demo cases?
Veridetect sells a camera-based visual defect inspection system to parts manufacturers. Doreen Vasquez is the quality engineering manager at Ashgrove Machining, a precision fastener maker. Baptiste Onyema runs quality at a sibling plant using the same vendor.
- Build a test set from your own historical hardest cases before go-live, not the vendor's demo set.Why: a vendor's demo shows you their best cases, not the ones that have actually cost you money before.
- Test rare defect categories separately, not folded into one aggregate score.Why: a rare, hard defect can fail badly while barely moving an overall accuracy number built mostly on common cases.
- Require a confidence signal per defect category, not just pass or fail.Why: without it, nobody can tell which category the system is actually unsure about.
- Get sign-off on the hard-case results before launch, not after an incident.Why: a gap found before go-live is a fix. The same gap found after a customer return is a recall.
- Re-run the hard-case archive test on a schedule, not just once.Why: the vendor's model changes over time, and a test that only ran once at signing goes stale.
How to answer this, stage by stage
Nobody is scoring you on whether you can say "test thoroughly." They're scoring whether you can name the specific gap a demo-only test would have missed.
Let's learn
Here is what happens when a vendor's demo is flawless and your worst case is not on it.
Say a parts manufacturer buys a camera system that inspects every fastener leaving the line for visible defects. Before it, Ashgrove's inspectors checked every part by hand, about 45 seconds a piece, catching a historically hard defect, a hairline stress crack near the fastener head, in about 91 percent of real cases. Veridetect's camera checks a part in about 3 seconds. Its own demo, built on common defects like surface scratches and missing holes, showed 99.2 percent accuracy.
Ashgrove went live plant-wide on the strength of that 99.2 percent number. Nobody had tested Veridetect specifically against Ashgrove's own archive of past hairline stress-crack cases, because the demo's aggregate score looked strong enough to cover everything.
Here's the turn: the 0.8 percent Veridetect missed on its own demo was never the real problem. The real problem was that the demo barely contained the one defect category that had actually cost Ashgrove money before. When Doreen's team finally tested the archive directly, catch rate on hairline stress cracks specifically was 64 percent, not 99.2.
At its worst: a batch of parts with an undetected hairline crack ships to a customer, and the defect is caught only by luck, in the customer's own incoming inspection, after the parts have already left the building.
What I would leave alone: Veridetect's handling of the common defects, scratches, missing holes, dimensional flaws, doesn't need this level of scrutiny. Those categories are well represented in any reasonable demo, and re-testing them constantly would waste time on a part of the system that was never actually in question.
The lesson: an aggregate accuracy number tells you how a vendor does on the cases it chose to show you. It tells you nothing about the one case you actually needed it for, until you go find that case yourself.
Now here is the same thing as a story
The short version above is what you'd say defending this test to Ashgrove's plant manager. Read this one for how the gap actually surfaced.
The label printer on Ashgrove's inspection line prints one tag for every part the vendor's system flags. It has printed the same four words for two years: hold for review.
Doreen Vasquez has run quality engineering at Ashgrove for nine years. Before Veridetect, she could tell a real hairline crack from a tooling shadow at a glance, the kind of judgment call that takes a decade to build and about a second to make.
Veridetect went live two years ago. For most of that time, it did exactly what the demo promised: fast, consistent, catching the common defects Ashgrove's inspectors used to catch by hand. Doreen's team stopped running a separate manual check on hairline cracks specifically, because the vendor's number covered "defects" as one category, and 99.2 percent felt like more than enough.
Baptiste Onyema runs quality at a sibling Ashgrove plant, same vendor, same equipment line. Eight months ago, his plant shipped a batch of about 3,400 fasteners to a customer. Forty of them carried a hairline stress crack near the fastener head. Veridetect had cleared all forty. The customer's own incoming inspection caught it, not Baptiste's line.
No injury. No field failure. A recall, a paused contract review, and a very uncomfortable phone call. Baptiste told Doreen about it at a supplier quality meeting, the way you tell someone about a near miss you still think about at 2am.
Doreen went back to Ashgrove and pulled the plant's own archive: two hundred real, historical images of hairline stress cracks, collected over years of manual inspection records. She ran Veridetect against that archive directly, for the first time since the system went live.
Sixty-four percent. Not fraud, not negligence on the vendor's part, just a rare category thin enough in the original demo that nobody, on either side, had ever actually confirmed it worked.
Here's what Doreen's team had actually been doing, without ever writing it down: after the first year, a few inspectors had quietly started pulling a random handful of Veridetect-cleared parts each week, specifically from the fastener-head area, for an off-the-books manual re-check. Nobody had asked them to. Nobody above them knew it was happening. It was never in a procedure document, never budgeted, never mentioned in a status report.
Here's what I'd take back. Ashgrove accepted Veridetect's demo pass rate as the whole story before going live, plant-wide, on every defect category at once. That was a reasonable read of a genuinely strong number. It stopped being reasonable the moment the demo turned out to barely include the one defect Ashgrove had actual scar tissue over.
I would go back and require the archive test, category by category, before go-live, not two years and a sibling plant's recall later. And the part I'd tell myself: we didn't get a bad vendor. We got a good one, tested on the wrong questions, and the gap sat there quietly enough that the fix ended up being an inspector's private habit instead of a decision anyone made on purpose.
The five steps, if you want to remember itNot "test more." FLIPS is what tells you exactly which case a demo-only test will never show you.
And if you want to be sure it really works, try it somewhere elseA different flip family, an airline's baggage-damage claims tool instead of a fastener line. This time the tool never even sees the hard cases at all.
Skyfare Regional uses an AI tool to assess photos of damaged luggage and recommend a settlement amount. Greta Sundqvist manages claims operations there. Mapped onto FLIPS with a different family, the substitution flip: F is Greta's claims agents, competent at spotting real structural damage in a glance. L is the habit of trusting the AI tool's recommendation on straightforward scuff-and-dent claims. I is the flip, but this time it's substitution, not workaround: agents learned early on that the tool struggled with complex claims, broken wheels plus frame damage, high dollar value, so they quietly began routing every complex claim to a senior human assessor by hand, while letting the tool handle only the easy claims it was good at.
The result: the tool's measured accuracy stayed excellent on the dashboard, because in practice it only ever saw the cases it was already good at. The real gap on complex claims was never tested, never fixed, and never visible to management, because the input mix had quietly shifted around it.
The old decision here isn't a missed archive test at go-live, it's a related reversal: Skyfare's original vendor evaluation used only the vendor's own sample claim photos, all simple and common, and never built a test set from Skyfare's own historically hardest structural-damage claims. That made sense when the sample photos looked convincing and nobody had reason to doubt them. It stopped making sense once agents, on their own, discovered the tool's real limits and quietly worked around them, leaving management's own dashboard blind to a gap the agents had already found by hand.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "test on your own hardest historical cases before go-live, not the vendor's demo set," and stop.
Cost: there's no budget for a full custom evaluation on every vendor. Say so honestly, and prioritize the archive test for the rare, highest-cost defect categories first, not every category equally.
The model gets better, for real: if a vendor's next release genuinely closes the gap on your hardest category, that's the re-test schedule doing its job, telling you the standing check can relax slightly, not that it can stop.
Where people run it wrong.
They accept a vendor's aggregate demo score as proof it works on the case that actually matters to them.
They test once at signing and never again, missing drift as the vendor's model changes.
They let a private workaround quietly cover a real gap instead of surfacing it as a decision for someone to actually make.
How to use it live. The moment someone says "the vendor's numbers look great," ask back: great on whose hardest case, yours or theirs? Let that answer decide whether to sign, not the number on the slide.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if you don't have historical data for a brand-new hard case?" Response: that's exactly when a per-category confidence signal matters most, since without history, the only way to know the system is uncertain is if it says so itself.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Evaluating AI vendors as a buyer
- #1 List the ten questions you would ask every AI vendor before a pilot.
- #2 How do you evaluate a vendor's quality claims without running your own eval?
- #3 Design the pilot you would run to evaluate two competing AI vendors.
- #4 What contractual terms matter specifically for AI vendors and not for other software?
- #5 How do you assess a vendor's model dependency and what happens if their provider changes terms?
- #6 Describe the data handling questions you would put to a vendor on behalf of your security team.