Artifact critiqueAdvancedAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #7

What does a good vendor eval report look like and what should make you suspicious?

AUDIT the one meeting where nobody could say which model version made the number

Silverbrook Animal Alliance runs six shelters and about nine hundred adoptions a year. Selin Marroquin runs day to day operations there. Amberdale Match is the vendor pitching an AI tool that scores how well an adopter and an animal fit. Jono Reyes sits on Silverbrook's volunteer board and spent a decade in a data team before that.

The direct answer
A good vendor eval report names its own eval set, pins the exact model version behind the number, and shows you the cases it got wrong, not only the ones it got right. Be suspicious the moment a report hands you one confident number and nothing to check it against. That number tells you what the vendor wants you to believe, not what your own cases will actually get.
Do this, in order
  1. Ask for the eval set and the model version before trusting any number.Why: without both, the number can't be checked or reproduced, so it isn't evidence yet.
  2. Run the vendor's claim on a slice of your own real cases before signing.Why: a report is the vendor's best day. Your own cases are every day.
  3. Read what the report leaves out as closely as what it includes.Why: no failure cases and no confidence range usually means those numbers exist and look bad.
  4. Ask who paid for the eval and who scored it.Why: an eval the vendor ran and graded itself has a built-in reason to look good.
  5. Set a re-test schedule tied to model version changes, not contract renewal.Why: a vendor can update its model quietly, and last year's report says nothing about today's version.
  6. Skip this level of scrutiny for small, no-cost, easy-to-reverse tools.Why: not every vendor decision carries a year-long lock-in, so treating them all the same wastes real hours.

How to answer this, stage by stage

Nobody is scoring whether you can list five red flags from memory. They're scoring whether you can name the one question that actually breaks a bad report open.

Stage 1
Scope it to one real report
Say it like this
"I'll answer this for Silverbrook Animal Alliance, looking at one vendor's eval report ahead of a twelve month contract, not vendor evaluation in general."
Why this works
Keeps the question from turning into a generic due-diligence checklist.
Stage 2
Name your structure out loud
Say it like this
"I'll use AUDIT. Ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself."
Why this works
Shows the interviewer a repeatable method, not just a gut feeling about vendors.
Stage 3
Reframe what the question is really testing
Say it like this
"This isn't really asking if I can spot a bad vendor. It's asking whether I'll trust a confident number or go find out what's actually behind it."
Why this works
Separates a real answer from a list of generic red flags anyone could recite.
Stage 4
Give the one check that matters most
Say it like this
"The single question that breaks a bad report open is: which exact model version made this number, and can I run it again myself on my own cases?"
Why this works
This is the direct answer, said the way you'd actually say it out loud.
Stage 5
Prove it with the near miss
Say it like this
"Amberdale Match's report claimed ninety six percent. When I ran their live tool on our own forty held-out cases myself, it got seventy one."
Why this works
Turns "be careful with vendor claims" into one specific, checkable afternoon.
Stage 6
Say what you'd still leave alone
Say it like this
"For a free volunteer-scheduling tool we can drop in a week, I wouldn't run any of this. Nothing is locked in, and nothing costs us if it's wrong."
Why this works
Shows judgment instead of turning every vendor conversation into a full audit.
Stage 7
Close on the one line
Say it like this
"A good report shows its own work. A suspicious one shows you one number and asks you to trust it. Test the number before the contract, not after."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

What happens the first time nobody checks a vendor's own number before signing a year-long contract on it?

Silverbrook Animal Alliance runs six shelters and about nine hundred adoptions a year. Before any AI tool existed, a volunteer read each adopter's paper questionnaire and matched them to an animal by hand, about forty minutes of reading and judgment per case. Roughly seven times in ten, that match held. Three times in ten, the animal came back within ninety days.

Hand sketched flow diagram titled Signing off on a vendor report, before AUDIT. Five boxes in sequence: Vendor sends report, Skim the headline number, Nod along, Sign the contract highlighted, Find the gap in production.
This is the whole path a report used to travel. Nothing on it asked the report to prove itself.

Amberdale Match pitched a tool that reads the same questionnaire and scores compatibility in under a second, backed by a one-page claim of ninety six percent accuracy, a number that would cut Silverbrook's return rate by more than half.

The real question was never whether that tool works. It was whether Selin could tell the difference between a number Amberdale proved and a number Amberdale just said.
Amberdale Match's claimed accuracy versus what Silverbrook's own cases actually got
100% 50% 0 96% Vendor's report 71% Selin's own replication
Twenty five points of gap, on the exact same forty cases. The report was not lying about its own test. It just wasn't Silverbrook's test.

At its worst, trusting an unverified number doesn't just risk one bad match. It risks locking a nonprofit's whole shelter network into a year-long, eighteen-thousand-dollar contract built on a number nobody outside the vendor ever checked. Running a real replication check costs about two weeks nobody had budgeted for. Against a year-long contract that size, that trade is an easy one to accept.

The choice I would take back Silverbrook's vendor review only ever asked for a one-page summary of results. That was fine when every tool it tried was free or month to month. It stopped being fine the moment a vendor's single page was about to justify a year-long contract nobody could walk away from cheaply.

What I would leave alone: the free volunteer shift-scheduling assistant Silverbrook already uses gets none of this scrutiny. Nobody's adoption outcome depends on its output, so demanding a full replication test there would burn real hours for no real protection.

The lesson: a good number and a good report are not the same thing. A report earns your trust by showing its own work. A number just asks for it.

Now here is the same thing as a story

The short version above is what you'd say defending this call to Silverbrook's board. Read this one for how close the mistake actually came.

Selin Marroquin could read a shelter intake form and know, just from the tone of the handwritten notes, which volunteer should make the follow-up call. She had run Silverbrook's operations for six years and vetted a dozen small tools the same easy way: ask for a one-page summary, skim it, decide.

Hand sketched comparison titled A trustworthy report versus a suspicious one. Left, a green document icon labeled Trustworthy, caption named eval set, version pin, failures shown. Right, a red-orange question mark box labeled Suspicious, caption one number, no set named, no version.
Amberdale's one pager was entirely the right side of this picture, and it read as fine for months.

For most of those tools, that was plenty. They were free, or month to month, and swapping one out cost an afternoon, not a contract.

Amberdale Match's demo went well. The compatibility score matched what shelter staff already believed about three sample cases. The vendor's one-pager claimed ninety six percent accuracy, tested against "a broad sample of real placements."

Knowledge spark: what's a version pin, and why does it matter? An AI model gets updated all the time, sometimes with no announcement. A version pin just means writing down exactly which version made a given number, the way a recipe card names its own ingredients. Without one, a great score from March tells you nothing about the model actually running in July.

Jono Reyes had joined Silverbrook's board four months earlier, after a decade in a data team at an insurance company. Sitting in on the pitch, he asked one question: "Which version of their model made that ninety six percent number, and whose cases were in the test?" Selin didn't know. Neither, it turned out, did Amberdale's own sales rep on the call.

Hand sketched quadrant titled Vendor claims, mapped, axes Verifiable and Favors the vendor. Reproducible benchmark sits high on verifiable, low on favors vendor. Cherry picked demo and marketing case study sit low on verifiable, high on favors vendor. Reference call sits in the middle.
Amberdale's one pager sat exactly where the cherry picked demo sits on this map.

Selin had considered simply asking Amberdale for another live demo instead of a harder replication test. She dropped that idea. A demo is scripted to the three cases a vendor already knows will look good. It could never tell her anything about the other thirty seven.

She asked Amberdale Match directly for their eval set and the version pin behind the number. Amberdale stalled for twelve days, then sent a spreadsheet with forty case descriptions, no names, no dates, no model version listed anywhere. Selin ran those same forty cases through Amberdale's live tool herself. It matched its own report on twelve of them, and disagreed sharply on the rest. Real accuracy against Silverbrook's own cases: seventy one percent, not ninety six.

The old decision that let this get so close to a signature wasn't the demo, and it wasn't Amberdale's number. It was Silverbrook's habit of treating one page as due diligence, a habit built for tools that cost nothing to walk away from.

Selin pulled the vendor packet off the board's agenda and sent Amberdale five questions: the eval set, the version, the failure cases, the confidence range, and who ran the test. She gave them ten business days to answer all five before Silverbrook would even reopen the conversation.

Six months later, a second vendor pitched a similar tool in the same board room, with the same kind of confident one-pager. This time Selin ran AUDIT before the meeting ever happened, asking for the eval set and version in her very first reply. The vendor answered honestly within nine days. Their real accuracy on Silverbrook's own cases came out to eighty four percent, close enough to their claim to trust. Silverbrook signed that one, and caught the gap two weeks into evaluation instead of three months into a contract.

AUDIT, in five checksNot a lecture on being suspicious of every vendor. AUDIT is the specific test a report has to pass before its number earns your trust.

Hand sketched icon list titled AUDIT, the five checks. Ask who paid for it. Uncover the eval set. Demand the version pin. Isolate what's missing. Test it yourself.
Five checks, in the order that actually breaks a bad report open.
A
Ask who paid for it.
Amberdale ran and scored its own eval, with nobody outside the company checking the work.
An eval nobody but the vendor touched carries the vendor's own reason to look good.
U
Uncover the eval set.
Amberdale's one-pager never named where its test cases actually came from.
A benchmark score means nothing if you can't see or recheck the cases behind it.
D
Demand the version pin.
Amberdale's own sales rep couldn't say which model version made the ninety six percent number, or when.
This is the step that actually separates a real report from a snapshot with no expiration date.
Hand sketched decision tree titled Do you trust this number. Root, vendor hands you a benchmark score. Three branches: set named and reproducible leads to verify once then trust, no eval set named leads to ask again and hold, won't share version leads to walk away.
Amberdale sat on the middle branch for twelve days before finally answering at all.
I
Isolate what's missing.
No failure cases, no confidence range, one number and nothing standing behind it.
What a report leaves out is usually more honest than what it puts on the front page.
T
Test it yourself.
Selin ran Amberdale's own forty cases through their live tool and got seventy one percent, not ninety six.
This is the only step that catches a gap a well-written report can hide.
Hand sketched labeled parts diagram titled What's inside a real eval report. A document icon at center labeled Eval Report, with four callouts: named eval set, version pin, failure cases, confidence range.
Amberdale's one pager had none of these four. That absence was the whole warning.

The recap, one line per letter: ask is checking who ran the eval and who it flatters, uncover is finding the actual cases behind the number, demand is pinning the exact model version, isolate is noticing what the report never shows, and test is running it yourself before any signature.

And if you want to be sure it really works, try it somewhere elseSame five checks, a used vehicle marketplace instead of an animal shelter network. A different old decision breaks the second report.

Wrenchbay runs an online marketplace for used trucks and trailers. Devraj Okafor, a product manager there, was evaluating Panelscan, a vendor pitching AI damage assessment from uploaded photos, meant to replace an inspector's forty-minute walkaround. Mapped onto AUDIT: ask is checking that Panelscan's own claims team ran and scored the eval with no outside review. Uncover is finding that Panelscan's report only listed "a large internal photo set," with no named source. Demand is asking which exact model version produced the ninety one percent damage-detection rate, since Panelscan had shipped two model updates that same quarter. Isolate is noticing the report showed only caught damage, never missed damage. Test is running twenty of Wrenchbay's own listing photos through Panelscan's live tool before signing anything.

The old decision here isn't a one-page-summary habit, it's a different reversal: Wrenchbay's procurement process merged the technical vendor review and the contract negotiation into a single meeting, run by the same category manager under a deadline to close the quarter. That merge made sense when Wrenchbay only ever bought off-the-shelf software with no real model behind it. It stopped making sense once one meeting had to both judge an AI vendor's claims and negotiate its price, with no natural pause between the two.

Hand sketched timeline titled The report's real lifecycle, Contract signed emphasized. Report written, caption vendor picks its best run. Report arrives, caption headline number only. Contract signed, caption nobody re-ran it. Gap surfaces, caption three months later.
Devraj split his single meeting back into two, so the middle step got a real pause again.

Devraj split the meeting in two: a technical AUDIT review first, contract terms only after Panelscan passed it. On the first pass, Panelscan's real detection rate against Wrenchbay's own twenty photos came out to seventy eight percent, good enough to proceed, but with a required quarterly retest written into the contract instead of a one-time check.

Back at Silverbrook, the same pattern held over the following year: every time Amberdale pushed a new model version, Selin re-ran the same forty held-out cases before trusting the new number.

Amberdale Match's accuracy on Silverbrook's own cases, across four silent model updates
100% 50% 0 retest floor, 80% Jan Mar May Jul 89% 71% 78%
A silent update in May dropped real accuracy below the floor for two months before anyone outside Amberdale would have known, if Selin hadn't kept checking.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "check who ran the eval, on what data, on which version, then verify it yourself," and stop.
Cost: there's no time to run a full replication before a decision is due. Say so honestly, and check the smallest, riskiest slice of cases instead of skipping the check entirely.
The model gets better, for real: if Amberdale's next version genuinely closes the gap and matches its own claim, that's still worth re-verifying once, since a good version now says nothing about the version running next quarter.

Where people run it wrong.
They treat a confident one-pager as due diligence, because reading it feels like the work of checking.
They accept "a broad internal test" as an answer instead of asking for the actual eval set.
They check the number once at signing and never again, even after the vendor ships a new model version.

How to use it live. The moment a vendor hands you a single number, ask back: which version made it, on whose data, and can I run it again myself? Let the answer to that decide how much you trust the rest of the pitch.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an artifact-critique question like judging a vendor's eval report?
Tap to flip
ANSWER
AUDIT: ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Selin Marroquin, operations lead at Silverbrook Animal Alliance, who could read an intake form and know who should make the call.
3 · THE QUESTION THAT MATTERED
What single question exposed the report?
Tap to flip
ANSWER
"Which version of their model made that number, and whose cases were in the test?" Nobody in the room, including the vendor, could answer it.
4 · THE GAP
What was the actual gap this story found?
Tap to flip
ANSWER
The vendor's report claimed 96 percent accuracy. Selin's own replication on the same forty cases got 71 percent.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating a one-page vendor summary as real due diligence, a habit that made sense for free trial tools and stopped making sense before a year-long contract.
6 · THE NUMBER
Fill in the blank: the vendor's report claimed 96 percent accuracy, but Selin's own replication on the same forty cases got only ___ percent.
Tap to flip
ANSWER
71 percent, a twenty five point gap on the exact same cases the vendor's own report claimed to cover.
7 · THE REPLAY
Same kind of confident one-pager, six months later. What changes?
Tap to flip
ANSWER
Selin runs AUDIT before the meeting happens instead of after. The gap surfaces two weeks into evaluation instead of three months into a signed contract.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different company. Which one, and what's different about its old decision?
Tap to flip
ANSWER
Wrenchbay, a used-vehicle marketplace, evaluating Panelscan. Its reversal is merging technical review and contract talks into one meeting, not a one-page-summary habit.

Check yourself Score: 0 / 0

Multiple choice
1. Why isn't a single headline number in a vendor's report enough to trust on its own?
  • A. Vendors are legally required to round their numbers up.
  • B. Reports always use the wrong font for the data.
  • C. It can't be reproduced or checked without a named eval set and a model version.
  • D. Interviewers dislike any answer with a percentage in it.
Show hint
Look at the direct answer and the U and D steps.
Show answer
C. Without the eval set and the version, the number can't be checked again, so it isn't evidence yet.
Fill in the blank
2. Fill in the blank: Selin's own replication on Silverbrook's forty held-out cases got ___ percent accuracy, against the vendor's claimed 96 percent.
Show hint
Look at the bar chart in Section 1.
Show answer
71 percent. A twenty five point gap on the exact same cases the vendor's report claimed to cover.
True or false
3. True or false: Amberdale Match lied about its own ninety six percent number.
  • True
  • False
Show hint
Look at the chart note after the first bar chart.
Show answer
False. The number was real on the vendor's own test. It just wasn't Silverbrook's test, which is exactly why replication matters.
Short answer, where it wouldn't matter
4. Name a tool at Silverbrook where this level of scrutiny genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The free volunteer shift-scheduling assistant. Nobody's adoption outcome depends on it, so a full replication test there would waste real hours.
Short answer, apply it yourself
5. Think of a vendor report or claim you've seen accepted at face value somewhere you've worked. What single question would have tested it?
Show hint
Ask whether anyone outside the party making the claim ever checked it on real, independent cases.
Show answer
Model answer: Usually the missing question is "who checked this besides the person who wrote it, and on what."
Short answer, the number question
6. If Selin's replication had come out to 90 percent instead of 71, would running AUDIT still have been worth it? Why or why not?
Show hint
Look at the priority list's reasoning about a year-long contract.
Show answer
Model answer: Yes. Even a passing result is only trustworthy because it was checked. An untested 90 percent is still just a claim, not evidence.
Before you close the answer
Why this works
Tests whether you treat a vendor's number as a fact or as a claim, and whether you know the one question that actually breaks a report open instead of a generic list of red flags.
Follow-up traps
"Isn't this just being suspicious of every vendor by default?" Response: no, AUDIT is a check you run once per report, not a permanent stance; a vendor that survives it earns real trust afterward.

"What if the vendor simply refuses to share their eval set?" Response: that refusal is itself the answer, per the decision tree; either walk away or price the unknown risk into the contract terms.
If pressed
Silverbrook now requires the replication slice to be held out entirely from the vendor beforehand, cases Amberdale never saw before quoting a number, so there's no way to tune the pitch to the test.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more