CaseIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #18

How do you test a vendor's product on your hardest cases rather than their demo cases?

FLIPS the workaround flip, a shadow queue nobody was told to build

Veridetect sells a camera-based visual defect inspection system to parts manufacturers. Doreen Vasquez is the quality engineering manager at Ashgrove Machining, a precision fastener maker. Baptiste Onyema runs quality at a sibling plant using the same vendor.

The direct answer
Before going live, pull your own archive of past real defects, especially the rare, hard ones, and run the vendor's system against that archive yourself. Do not accept the vendor's demo pass rate as proof it works on the case that has actually hurt you before. An aggregate number built on common defects will look excellent while hiding a specific, costly blind spot.
Do this, in order
  1. Build a test set from your own historical hardest cases before go-live, not the vendor's demo set.Why: a vendor's demo shows you their best cases, not the ones that have actually cost you money before.
  2. Test rare defect categories separately, not folded into one aggregate score.Why: a rare, hard defect can fail badly while barely moving an overall accuracy number built mostly on common cases.
  3. Require a confidence signal per defect category, not just pass or fail.Why: without it, nobody can tell which category the system is actually unsure about.
  4. Get sign-off on the hard-case results before launch, not after an incident.Why: a gap found before go-live is a fix. The same gap found after a customer return is a recall.
  5. Re-run the hard-case archive test on a schedule, not just once.Why: the vendor's model changes over time, and a test that only ran once at signing goes stale.

How to answer this, stage by stage

Nobody is scoring you on whether you can say "test thoroughly." They're scoring whether you can name the specific gap a demo-only test would have missed.

Stage 1
Scope it to one vendor, one defect category
Say it like this
"I'll answer this for Ashgrove Machining's rollout of Veridetect, and one specific defect: a hairline stress crack near the fastener head."
Why this works
Keeps "test on your hardest cases" from turning into a generic QA lecture.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Shows a real method for finding where a demo-only test actually breaks, not just a gut instinct to "test more."
Stage 3
Reframe what the question is really testing
Say it like this
"The real question isn't whether Veridetect is accurate. It's whether the vendor's demo cases and your worst historical cases are even the same test."
Why this works
Separates a strong candidate from one who just repeats "high accuracy" back at the interviewer.
Stage 4
Give the one decision
Say it like this
"Build your own archive of past real defects, run it against the vendor's system yourself, and require sign-off on it before go-live, category by category, not as one blended score."
Why this works
This is the direct answer, said plainly, before any story backs it up.
Stage 5
Prove it with the near miss
Say it like this
"Veridetect's own demo showed 99.2 percent accuracy. When we finally ran it against our own archive of hairline stress-crack cases, the real number was 64 percent. We only found out after a sibling plant shipped a batch with the same defect to a customer."
Why this works
Turns "test on your hardest cases" from advice into a specific, avoidable number.
Stage 6
Say what you'd measure going forward
Say it like this
"Re-run the hard-case archive quarterly, not just at signing, and track catch rate by defect category, not one blended number."
Why this works
Shows the fix isn't a one-time test, it's a standing check.
Stage 7
Close on the one line
Say it like this
"A demo shows you the vendor's best day. Your archive shows you your worst one. Test against the second, before you sign anything."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

Here is what happens when a vendor's demo is flawless and your worst case is not on it.

Say a parts manufacturer buys a camera system that inspects every fastener leaving the line for visible defects. Before it, Ashgrove's inspectors checked every part by hand, about 45 seconds a piece, catching a historically hard defect, a hairline stress crack near the fastener head, in about 91 percent of real cases. Veridetect's camera checks a part in about 3 seconds. Its own demo, built on common defects like surface scratches and missing holes, showed 99.2 percent accuracy.

Hand sketched labeled parts diagram titled The five letters. A circle at the center labeled FLIPS, with five callouts around it: F find the person, L locate the habit, I identify the flip, P pinpoint the decision, S show the replay.
The whole method, in one picture. The hardest letter is the one this story turns on.

Ashgrove went live plant-wide on the strength of that 99.2 percent number. Nobody had tested Veridetect specifically against Ashgrove's own archive of past hairline stress-crack cases, because the demo's aggregate score looked strong enough to cover everything.

Here's the turn: the 0.8 percent Veridetect missed on its own demo was never the real problem. The real problem was that the demo barely contained the one defect category that had actually cost Ashgrove money before. When Doreen's team finally tested the archive directly, catch rate on hairline stress cracks specifically was 64 percent, not 99.2.

Veridetect's accuracy, aggregate versus the one hardest defect category
100% 50 0 99.2% Aggregate, demo set 64% Hairline crack, own archive
The demo number and the number that actually matters to Ashgrove are two different tests wearing the same accuracy label.

At its worst: a batch of parts with an undetected hairline crack ships to a customer, and the defect is caught only by luck, in the customer's own incoming inspection, after the parts have already left the building.

The choice I would take back Ashgrove accepted Veridetect's own demo pass rate as sufficient validation before going live, and never built a required test against Ashgrove's own archive of past hard cases first. That made sense at the time: 99.2 percent looked like an unambiguous win, and asking for more testing felt like slowing down an obvious upgrade. It stopped making sense the moment the one defect category that actually mattered turned out to barely be represented in that number at all.
Knowledge spark: why would a vision model do worse on a rare defect? A model learns mostly from what it sees most often. If hairline stress cracks are a small share of the training and demo examples, the model has less practice recognizing them, even while it gets very good at the common defects that make up most of its exposure.

What I would leave alone: Veridetect's handling of the common defects, scratches, missing holes, dimensional flaws, doesn't need this level of scrutiny. Those categories are well represented in any reasonable demo, and re-testing them constantly would waste time on a part of the system that was never actually in question.

The lesson: an aggregate accuracy number tells you how a vendor does on the cases it chose to show you. It tells you nothing about the one case you actually needed it for, until you go find that case yourself.

Now here is the same thing as a story

The short version above is what you'd say defending this test to Ashgrove's plant manager. Read this one for how the gap actually surfaced.

The label printer on Ashgrove's inspection line prints one tag for every part the vendor's system flags. It has printed the same four words for two years: hold for review.

Doreen Vasquez has run quality engineering at Ashgrove for nine years. Before Veridetect, she could tell a real hairline crack from a tooling shadow at a glance, the kind of judgment call that takes a decade to build and about a second to make.

Veridetect went live two years ago. For most of that time, it did exactly what the demo promised: fast, consistent, catching the common defects Ashgrove's inspectors used to catch by hand. Doreen's team stopped running a separate manual check on hairline cracks specifically, because the vendor's number covered "defects" as one category, and 99.2 percent felt like more than enough.

Hand sketched comparison titled Small move, big snap. Left, a gauge icon labeled Old and new value, caption a slow crawl, 91 to 64 percent. Right, a scale icon labeled Trust: fine then not, caption no middle setting.
The number crept. The trust didn't crawl with it. It held, then it snapped.

Baptiste Onyema runs quality at a sibling Ashgrove plant, same vendor, same equipment line. Eight months ago, his plant shipped a batch of about 3,400 fasteners to a customer. Forty of them carried a hairline stress crack near the fastener head. Veridetect had cleared all forty. The customer's own incoming inspection caught it, not Baptiste's line.

No injury. No field failure. A recall, a paused contract review, and a very uncomfortable phone call. Baptiste told Doreen about it at a supplier quality meeting, the way you tell someone about a near miss you still think about at 2am.

We did not have a worse camera. We had the exact same camera, tested against the exact same demo, and neither plant had ever once shown it the one defect that had actually cost us money before.

Doreen went back to Ashgrove and pulled the plant's own archive: two hundred real, historical images of hairline stress cracks, collected over years of manual inspection records. She ran Veridetect against that archive directly, for the first time since the system went live.

Hand sketched quadrant titled Defect types, by how often and how well caught. Axes, how often seen from rare to common, and how well Veridetect catches it from poorly to well. Surface scratches and missing holes sit common and caught well. Hairline stress crack sits rare and caught poorly.
The demo lived entirely in the top right. The defect that actually hurt Ashgrove lived in the bottom left, alone.

Sixty-four percent. Not fraud, not negligence on the vendor's part, just a rare category thin enough in the original demo that nobody, on either side, had ever actually confirmed it worked.

Here's what Doreen's team had actually been doing, without ever writing it down: after the first year, a few inspectors had quietly started pulling a random handful of Veridetect-cleared parts each week, specifically from the fastener-head area, for an off-the-books manual re-check. Nobody had asked them to. Nobody above them knew it was happening. It was never in a procedure document, never budgeted, never mentioned in a status report.

Hand sketched flow diagram titled How the hidden workaround formed. Five boxes in sequence: Veridetect clears part, No confidence shown, Doreen doubts hard case highlighted, Private re-check begins, Nobody told, made official.
A private safety net, built by hand, because the official system gave nobody a way to say which parts it was actually unsure about.

Here's what I'd take back. Ashgrove accepted Veridetect's demo pass rate as the whole story before going live, plant-wide, on every defect category at once. That was a reasonable read of a genuinely strong number. It stopped being reasonable the moment the demo turned out to barely include the one defect Ashgrove had actual scar tissue over.

I would go back and require the archive test, category by category, before go-live, not two years and a sibling plant's recall later. And the part I'd tell myself: we didn't get a bad vendor. We got a good one, tested on the wrong questions, and the gap sat there quietly enough that the fix ended up being an inspector's private habit instead of a decision anyone made on purpose.

The five steps, if you want to remember itNot "test more." FLIPS is what tells you exactly which case a demo-only test will never show you.

F
Find the person.
Doreen Vasquez, nine years running quality engineering at Ashgrove, able to spot a real hairline crack from a tooling shadow at a glance.
A specific person with real competence, not a generic "the QA team."
L
Locate the habit.
Her team stopped running a separate manual check on hairline cracks specifically, since the vendor's blended accuracy number felt like more than enough coverage.
A rational habit, formed because the tool genuinely worked on everything they could see.
I
Identify the flip.
Trusting Veridetect's clearance across every defect type, on one setting, versus quietly building a private, off-the-books re-check queue for the hardest category, on the other. No official middle setting existed.
The hardest step, and the one the whole story turns on.
P
Pinpoint the old decision.
Accepting the vendor's demo pass rate as sufficient before go-live, instead of requiring a test against Ashgrove's own hard-case archive first.
A specific, one-time evaluation decision, not an ongoing dial anyone turned up.
S
Show the replay.
With the archive test required before go-live, the 64 percent gap surfaces during evaluation, not after a customer return. Veridetect gets retuned for that category, and catch rate rises to 93 percent before a single part ships.
A countable ending: fixed before launch, not after a recall.
Hand sketched metaphor scene titled Switch, not dial. Left, a gauge icon labeled Assumed, caption partial trust, many settings. Right, a box icon labeled Actual, caption official process or secret workaround.
There was never a middle setting on offer. Only a decision nobody made, and a workaround that filled the gap instead.
Hand sketched icon list titled What a real hard-case test should include. Four items: your own historical archive not theirs, the rare defect categories especially, a confidence score per category, sign-off before go-live not after.
None of these four appeared in Ashgrove's original rollout checklist. All four are exactly where the 64 percent was hiding.

And if you want to be sure it really works, try it somewhere elseA different flip family, an airline's baggage-damage claims tool instead of a fastener line. This time the tool never even sees the hard cases at all.

Skyfare Regional uses an AI tool to assess photos of damaged luggage and recommend a settlement amount. Greta Sundqvist manages claims operations there. Mapped onto FLIPS with a different family, the substitution flip: F is Greta's claims agents, competent at spotting real structural damage in a glance. L is the habit of trusting the AI tool's recommendation on straightforward scuff-and-dent claims. I is the flip, but this time it's substitution, not workaround: agents learned early on that the tool struggled with complex claims, broken wheels plus frame damage, high dollar value, so they quietly began routing every complex claim to a senior human assessor by hand, while letting the tool handle only the easy claims it was good at.

The result: the tool's measured accuracy stayed excellent on the dashboard, because in practice it only ever saw the cases it was already good at. The real gap on complex claims was never tested, never fixed, and never visible to management, because the input mix had quietly shifted around it.

Hand sketched quadrant reused for Skyfare Regional, showing common versus rare claim types plotted against how well the tool handles each, with simple scuff claims common and well handled, and complex structural damage claims rare and poorly handled.
A different airline, a different defect, the same empty bottom-left corner nobody tested.

The old decision here isn't a missed archive test at go-live, it's a related reversal: Skyfare's original vendor evaluation used only the vendor's own sample claim photos, all simple and common, and never built a test set from Skyfare's own historically hardest structural-damage claims. That made sense when the sample photos looked convincing and nobody had reason to doubt them. It stopped making sense once agents, on their own, discovered the tool's real limits and quietly worked around them, leaving management's own dashboard blind to a gap the agents had already found by hand.

Monthly count of hairline-crack parts cleared in error, before and after the archive test
6 3 0 Baptiste's incident Month 1: 3 Month 4: 6 Month 6: 0
The archive test and retuning didn't just stop the bleeding, they reversed a trend that had been climbing for months before anyone noticed.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "test on your own hardest historical cases before go-live, not the vendor's demo set," and stop.
Cost: there's no budget for a full custom evaluation on every vendor. Say so honestly, and prioritize the archive test for the rare, highest-cost defect categories first, not every category equally.
The model gets better, for real: if a vendor's next release genuinely closes the gap on your hardest category, that's the re-test schedule doing its job, telling you the standing check can relax slightly, not that it can stop.

Where people run it wrong.
They accept a vendor's aggregate demo score as proof it works on the case that actually matters to them.
They test once at signing and never again, missing drift as the vendor's model changes.
They let a private workaround quietly cover a real gap instead of surfacing it as a decision for someone to actually make.

How to use it live. The moment someone says "the vendor's numbers look great," ask back: great on whose hardest case, yours or theirs? Let that answer decide whether to sign, not the number on the slide.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip: the tool becomes one step inside a longer manual chain the team invents, and management can't see it, because the tool gave no way to flag which cases it was unsure about.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Doreen Vasquez, nine years running quality engineering at Ashgrove Machining, able to spot a real hairline crack from a tooling shadow at a glance.
3 · THE HABIT
What did Doreen's team stop doing because the vendor's number looked strong?
Tap to flip
ANSWER
Running a separate manual check on hairline stress cracks specifically, since Veridetect's blended 99.2 percent accuracy felt like more than enough.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Trusting Veridetect's clearance across all defect types on one setting, versus a private, off-the-books re-check queue for the hardest category on the other. No official middle setting existed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Accepting Veridetect's demo pass rate as sufficient before go-live, instead of requiring a test against Ashgrove's own hard-case archive first.
6 · THE NUMBER
Fill in the blank: Veridetect's demo showed 99.2 percent aggregate accuracy, but on Ashgrove's own archive of hairline stress-crack cases, the real accuracy was ___ percent.
Tap to flip
ANSWER
64 percent, a gap the blended aggregate number never once revealed.
7 · THE REPLAY
Same rollout, but the archive test runs before go-live. What changes?
Tap to flip
ANSWER
The 64 percent gap surfaces during evaluation. Veridetect gets retuned for that category, raising catch rate to 93 percent before a single part ships, not after a sibling plant's recall.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Skyfare Regional's baggage-damage claims tool, using the substitution flip: agents quietly reserved the tool for easy claims and routed hard ones to humans, hiding the real gap from the dashboard.

Check yourself Score: 0 / 0

Short answer, name the flip
1. What was the flip in Doreen's story, and what were its two settings?
Show hint
Look at the I step in the FLIPS recap.
Show answer
Model answer: Trusting Veridetect's clearance across every defect type, versus a private, off-the-books re-check queue for the hardest category. There was no official middle setting.
Fill in the blank
2. Fill in the blank: Veridetect's demo showed 99.2 percent aggregate accuracy, but against Ashgrove's own archive of hairline stress-crack cases, the real accuracy was ___ percent.
Show hint
Look at the bar chart in Section 1.
Show answer
64 percent. A 35-point gap the blended aggregate score never showed.
Multiple choice
3. Why couldn't Doreen's inspectors have just "checked a little more carefully" instead of building a full workaround queue?
  • A. They were told to build a formal review process.
  • B. Veridetect asked them to double-check its work.
  • C. There was no official partial-trust option, only full trust or a private full re-check.
  • D. Checking more carefully would have cost extra money the plant didn't have.
Show hint
Look at "the flip has a middle setting" in the flip taxonomy's tells.
Show answer
C. The system offered pass or fail with no confidence signal, so the only real options were full trust or a fully separate, unofficial re-check.
True or false
4. True or false: the near miss happened because Veridetect's underlying camera hardware was faulty.
  • True
  • False
Show hint
Look at "why would a vision model do worse on a rare defect."
Show answer
False. The hardware worked fine. The model simply had far less training and demo exposure to the rare hairline-crack category.
Short answer, where it wouldn't matter
5. Name a defect category at Ashgrove where this level of hard-case testing genuinely wouldn't matter.
Show hint
Look at "What I would leave alone."
Show answer
Model answer: Common defects like surface scratches and missing holes. They're well represented in any reasonable demo, so retesting them constantly wastes time on a part of the system that was never in question.
Short answer, apply it yourself
6. Pick a tool you use that reports one overall quality score. What's the rare, hard case it's least likely to have been tested on, and how would you check?
Show hint
Ask what case is rare enough to barely affect an aggregate score, but costly enough to matter if it fails.
Show answer
Model answer: Build or find your own small archive of that hard case and test the tool against it directly, the same move as Ashgrove's stress-crack archive.
Before you close the answer
Why this works
Tests whether you'll accept a vendor's aggregate demo number, or go build your own test from the case that has actually hurt you before, and whether you can name the specific gap a blended score is built to hide.
Follow-up traps
"Isn't building a custom archive test for every vendor unrealistic?" Response: prioritize it for the rare, highest-cost categories first, the ones a demo is least likely to represent, rather than testing everything equally.

"What if you don't have historical data for a brand-new hard case?" Response: that's exactly when a per-category confidence signal matters most, since without history, the only way to know the system is uncertain is if it says so itself.
If pressed
Ashgrove's actual fix wasn't a bigger vendor model. It was a targeted threshold adjustment on the hairline-crack category specifically, since retraining on more common-defect data would have done nothing to close a gap that was never about volume.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more