CaseAdvancedAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #3
Design the pilot you would run to evaluate two competing AI vendors.
SPARK the dockworker who noticed the demo reel only ever showed the easy containers
Greymarsh Terminal handles inbound shipping containers for a mid-size port. Iris Halvorsen runs container inspection there. Portalis and Quaydeck both sell AI that reads camera footage of a container and flags likely shipping damage before a truck driver takes possession.
The direct answer
Run both vendors side by side, in shadow mode, on the exact same live container stream for several weeks, scored against a blind human-adjudicated answer, never each vendor's own reported number. Neither system gets to auto-release or auto-bill a single container during that window. The whole design exists to stop either vendor from grading its own homework, which a same-day bake-off using a sales demo can never do.
Do this, in order
Test both vendors on the same live, unfiltered container stream, not a curated demo reel.Why: a reel picked by the vendor's own sales team will always favor the vendor.
Score both against a blind human-adjudicated answer, not their own self-reported accuracy.Why: letting a vendor grade its own homework is the fastest way to buy the wrong tool with confidence.
Keep a human in the loop for every release decision during the whole pilot.Why: a wrong call that never reaches a human is a wrong call nobody in the pilot ever sees.
Break results out by shift and lighting condition, not just a single daily average.Why: a vendor that's strong in daylight and weak at night looks fine in an average that hides the weak half entirely.
Decide in advance what "good enough to switch" actually means, in a number.Why: without a stated bar, the loudest sales rep tends to win the decision instead of the data.
How to answer this, stage by stage
Nobody is scoring whether you can name two vendors. They're scoring whether your pilot design stops either vendor from controlling its own evidence.
Stage 1
Scope it to one real bake-off
Say it like this
"I'll design this for a port terminal comparing two container-damage inspection vendors before committing to either one."
Why this works
Keeps "design a pilot" from turning into an abstract essay on vendor selection.
Stage 2
Say the structure out loud
Say it like this
"I'll use SPARK. Situation, the job today with neither vendor. Payoff, the habit I want to build. Anchor, the one design decision everything hangs on. Risk, what breaks if I'm wrong. Keep out, what I deliberately won't build yet."
Why this works
Shows a repeatable design method instead of a list of nice-to-haves.
Stage 3
Ground it in the job today
Say it like this
"Right now, inspectors visually check every container by hand, about four minutes each, catching most damage but missing some, with no AI in the loop at all."
Why this works
The anchor only means something once the interviewer can picture the manual process it's replacing.
Stage 4
Give the anchor
Say it like this
"Both vendors run in shadow mode on the same containers for six weeks, scored blind against a human's independent call, and neither one is allowed to auto-release anything during that window."
Why this works
This is the direct answer, concrete enough that a panel could ask exactly how the blind scoring works.
Stage 5
Prove it with the near miss
Say it like this
"Before this design existed, one vendor's sales rep ran a live side-by-side using a reel of forty containers he'd hand-picked. It looked great. A dockworker on the floor pointed out every one of them was a clean, easy call, nothing like the messy ones that actually cause disputes."
Why this works
Turns "design a fair evaluation" from theory into a specific afternoon that nearly went the wrong way.
Stage 6
Close on the one line
Say it like this
"Never let a vendor pick which containers you judge them on. Run both on the same real stream, score them blind, and keep a human deciding until the numbers earn otherwise."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Let's learn
Here is what happens when a team lets a fast, exciting demo stand in for a real evaluation.
Before either vendor, Greymarsh's inspectors checked every inbound container by eye, about four minutes each, roughly two hundred a day across the team, catching around eight of every ten billable damage cases before a truck drove off with it. An AI read on the same camera footage could flag likely damage in under a second, at every gate, all day.
This is the job neither vendor has touched yet. The anchor only makes sense next to this.
Here's the turn: speed was never the risk in this design. The risk was that Greymarsh's original evaluation plan merged two separate questions, does this even work, and which of the two is better, into a single same-day bake-off, using whatever containers each vendor's own rep chose to show. That felt efficient with peak season two weeks out. It also meant each vendor got to pick its own best evidence.
Damage caught against a blind human answer, after four pilot weeks
Portalis caught more damage, but the second chart below shows why that isn't the whole story.
At its worst, this doesn't just risk picking a slightly worse vendor. It risks signing a multi-year contract based on a reel that was never a fair test, then discovering the gap only once real containers, and real disputed claims, are already on the line.
The choice I would take back
The team folded "does either tool work" and "which one is better" into one same-day side-by-side, using each vendor's own chosen sample of containers. That made sense with peak season pressing and no time to spare. It stopped making sense the moment a dockworker noticed the sample never included a single hard call, only the ones a demo reel is built to make look easy.
What I would leave alone: the terminal's basic gate-camera hardware doesn't need re-evaluating alongside the vendors. Both companies work with the same existing camera feed, so the hardware itself was never the variable in question.
The lesson: a fast bake-off feels efficient right up until you realize you let the two companies being judged also choose the evidence they're judged on. The fix isn't more time. It's who controls the sample.
Now here is the same thing as a story
The short version above is what you'd say defending the pilot design to the terminal's operations committee. Read this one for how close the reel came to deciding it instead.
Iris Halvorsen had run container inspection at Greymarsh for nine years, sharp enough to spot a hairline dent through a truck's side mirror at thirty feet. When both vendors came calling with pilot offers, she scheduled what felt like the obvious next step: a live side-by-side demo, one afternoon, both systems watching the same gate.
The demo day never tested either side of this picture. It just showed easy containers where nobody was ever wrong.
Knowledge spark: what does "shadow mode" mean?
Shadow mode means an AI system watches and makes its call in the background, but a human still makes the real decision. Nobody acts on the AI's answer yet, so you can compare it against reality without any risk to an actual container release.
Portalis's sales rep arrived with a reel: forty containers, hand-picked from a customer reference site, each with an obvious dent or a torn seal. Quaydeck's rep, seeing the bar had been set, brought a similar reel of his own. Both systems called every container correctly. Iris nearly signed a letter of intent with Portalis that same week, since it had called its forty a beat faster.
A dockworker named Nils Aabo, restacking pallets nearby, glanced at the monitor mid-demo and said something that stuck with Iris: "These are all the easy ones. Where's the container from Tuesday, the one with the scuff that took three of us arguing to call?" Nobody had an answer.
Both vendors had passed a test neither of them could have failed. The reel wasn't proof of anything except that a demo, picked by the person selling it, will always look perfect.
Iris pulled the letter of intent and redesigned the evaluation from scratch: six weeks, both vendors reading the same live, unfiltered gate stream, scored against her own team's independent call on every container, released to blind review so nobody knew which vendor's flag they were checking.
This is what replaced the reel. Every part of it exists to take control away from whoever's selling.
The near miss happened in week five of the redesigned pilot, not week one of the sales demo. That's the whole point of shadow mode.
SPARK, in one screenNot a longer bake-off. SPARK is what stops a vendor demo from quietly writing its own grade.
S
Situation. The job today, with neither vendor.
Inspectors visually check every container by hand, catching roughly eight of ten billable damage cases before release.
Grounds the whole design in a real workflow instead of an abstract comparison.
P
Payoff. The habit worth building.
Inspectors stop re-checking every container and start spending their time only on the ones either system escalates as unclear.
This is what the pilot is actually for, not just a vendor score.
A
Anchor. The one decision everything hangs on.
Both vendors run in shadow mode on the same live stream, scored blind against an independent human call, for six weeks, with no auto-release.
This is the hardest step, and the whole pilot design turns on getting it right.
R
Risk. What breaks if this is wrong.
If inspectors know which containers are being watched closely, they may unconsciously check those more carefully, so the comparison has to stay blind to them too.
Naming this in advance is what keeps the anchor from quietly failing in a way nobody notices until later.
K
Keep out. What's deliberately not built yet.
Neither vendor gets auto-release or auto-billing authority during the pilot, no matter how confident either system reports itself to be.
Shows judgment instead of a wish list of every feature both vendors demoed.
None of these are missing by accident. Each one is a decision the pilot deliberately postpones.
The recap, one line per letter: situation is the manual inspection process both vendors are meant to change, payoff is inspectors focusing only on escalated cases, anchor is the blind, parallel, shadow-mode design, risk is inspectors behaving differently if they know which containers are being watched, and keep out is withholding auto-release authority from both vendors for the entire pilot.
Portalis caught more damage. It also called more false alarms. Neither number alone would have told Iris that.
And if you want to be sure it really works, try it somewhere elseSame five letters, a recycling plant's contamination-sorting cameras instead of a shipping port. A reporting decision breaks the second story, not a sales reel.
Gullwash Recycling piloted Clearbin and Binscope, two vendors selling AI cameras that spot contaminants on a sorting line before they reach a bale. Mapped onto SPARK: situation is line workers manually pulling contaminants off a moving belt today, payoff is workers focusing only on the items either camera flags as unclear, anchor is the same shadow-mode design, both cameras watching the same belt for six weeks, scored blind against a worker's own call, risk is lighting changing sharply between day and night shifts in ways a vendor's daytime demo would never reveal, and keep out is neither camera getting authority to auto-reject a bale during the pilot.
The old decision here isn't a sales reel, it's a reporting one: the plant's original test setup had each vendor's dashboard report a single end-of-day accuracy number, with no breakdown by shift. That made sense, since daily summaries were how every other machine on the line got tracked. It stopped making sense once the night shift's dim lighting turned out to be exactly where both systems struggled most, hidden completely inside a daily average that looked fine.
This is what the shift breakdown made possible. A single daily average could never have produced this rule.
Contamination missed per shift, day versus night, before and after the reporting broke out by shift
The daily average told Gullwash almost nothing. The night shift was carrying nearly all of the real risk, invisibly, until the report was broken apart.
Swap the trigger and it still runs.
Speed: an interviewer caps you at a minute. Say "shadow mode, same stream, blind scoring, no auto-release," and stop.
Cost: there's no budget to run both vendors for six full weeks. Say so honestly, and shrink the window rather than the sample, testing fewer weeks on the same real, unfiltered stream instead of a shorter test on a curated one.
The vendor gets better, for real: if a vendor's next model version genuinely narrows the gap with a competitor, that's the blind scoring doing its job, telling you the decision is closer than the last quarter's numbers suggested.
Where people run it wrong.
They let each vendor supply its own demo sample instead of using one shared, unfiltered stream.
They score against a vendor's own reported accuracy instead of a blind, independent human call.
They report a single average number instead of breaking results out by the conditions that actually vary, like shift, lighting, or container type.
How to use it live. If an interviewer pushes for more detail, name the one design choice that would have caught the near miss: whoever controls the sample controls the result, so never let the vendor choose it.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "design a pilot" or "design a feature" question?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It runs forward, the way real design decisions actually happen.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Iris Halvorsen, who has run container inspection at Greymarsh Terminal for nine years, and redesigned the pilot after a dockworker's remark.
3 · THE ANCHOR
What's the one design decision the whole pilot hangs on?
Tap to flip
ANSWER
Both vendors run in shadow mode on the same live stream, scored blind against an independent human call, with no auto-release for either.
4 · THE RISK
What breaks first if the anchor is designed wrong?
Tap to flip
ANSWER
Inspectors, knowing which containers are under close comparison, unconsciously check those more carefully, which quietly biases the whole result.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging "does it work" and "which is better" into one same-day bake-off using each vendor's own curated demo sample.
6 · THE NUMBER
Fill in the blank: on the blind ground truth, Portalis caught 88 percent of damage, and Quaydeck caught ___ percent.
Tap to flip
ANSWER
79 percent, but with fewer false alarms, which is why the decision needed both numbers, not just the catch rate.
7 · THE REPLAY
Same demo day, same two vendors, but the shadow-mode design already exists. What changes?
Tap to flip
ANSWER
There is no single demo day at all. Both vendors run quietly for six weeks on real containers, and the letter of intent waits for the blind scoring, not a reel.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Gullwash Recycling's contamination-sorting camera pilot. The reversal is a reporting choice, a single daily average that hid the night shift's much higher miss rate.
Check yourself Score: 0 / 0
True or false
1. True or false: the near miss happened because one vendor's AI model was clearly worse than the other's.
True
False
Show hint
Look at what Nils pointed out during the demo.
Show answer
False. Both vendors passed a demo reel of hand-picked easy containers. The reel never tested either model against a real hard case.
Multiple choice
2. Why does the anchor require blind scoring against a human call, instead of trusting each vendor's own reported accuracy?
A. Vendors are legally required to under-report their accuracy.
B. A vendor grading its own homework will always look better than it is.
C. Human inspectors are always more accurate than any AI model.
D. It's cheaper than asking the vendor for a number.
Show hint
Look at the anchor step and the direct answer.
Show answer
B. Self-reported numbers can't be checked against anything independent, which is exactly what let the demo reel look perfect.
Fill in the blank
3. Fill in the blank: Portalis caught 88 percent of damage against the blind answer, while Quaydeck caught ___ percent.
Show hint
Look at the grouped bar chart in Section 1.
Show answer
79 percent. Lower catch rate, but also fewer false alarms, which the quadrant diagram shows matters just as much.
Short answer, where it wouldn't matter
4. Name a part of this pilot design where extra scrutiny of the two vendors genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The gate camera hardware itself. Both vendors read the same existing camera feed, so the hardware was never the variable being tested.
Short answer, apply it yourself
5. Think of two competing tools you've seen compared at work. Was the comparison run on a shared, unfiltered sample, or did each side supply its own?
Show hint
Ask who chose the test cases, and whether either side could have picked easier ones for themselves.
Show answer
Model answer: Most informal bake-offs let each vendor bring its own examples, which quietly favors whichever vendor curates the better reel, not the better product.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Merging the "does it work" question and the "which is better" question into one same-day demo. It made sense with peak season two weeks away.
Before you close the answer
Why this works
Tests whether you'll design an evaluation that a vendor can't quietly control, and whether you know the difference between a demo and a real test.
Follow-up traps
"Isn't six weeks too long when you need a decision fast?" Response: shrink the window before you shrink the sample; a shorter pilot on the same real, unfiltered stream still beats a longer one on a curated reel.
"What if one vendor refuses to run in shadow mode without auto-release authority?" Response: that refusal is itself useful information about how much control you'd actually have after signing.
If pressed
Greymarsh's blind scoring used a rotating third inspector who never saw which flag came from which vendor, specifically to stop even an unconscious house preference from creeping into the human ground truth itself.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.