What does a good vendor eval report look like and what should make you suspicious?
Silverbrook Animal Alliance runs six shelters and about nine hundred adoptions a year. Selin Marroquin runs day to day operations there. Amberdale Match is the vendor pitching an AI tool that scores how well an adopter and an animal fit. Jono Reyes sits on Silverbrook's volunteer board and spent a decade in a data team before that.
- Ask for the eval set and the model version before trusting any number.Why: without both, the number can't be checked or reproduced, so it isn't evidence yet.
- Run the vendor's claim on a slice of your own real cases before signing.Why: a report is the vendor's best day. Your own cases are every day.
- Read what the report leaves out as closely as what it includes.Why: no failure cases and no confidence range usually means those numbers exist and look bad.
- Ask who paid for the eval and who scored it.Why: an eval the vendor ran and graded itself has a built-in reason to look good.
- Set a re-test schedule tied to model version changes, not contract renewal.Why: a vendor can update its model quietly, and last year's report says nothing about today's version.
- Skip this level of scrutiny for small, no-cost, easy-to-reverse tools.Why: not every vendor decision carries a year-long lock-in, so treating them all the same wastes real hours.
How to answer this, stage by stage
Nobody is scoring whether you can list five red flags from memory. They're scoring whether you can name the one question that actually breaks a bad report open.
Let's learn
What happens the first time nobody checks a vendor's own number before signing a year-long contract on it?
Silverbrook Animal Alliance runs six shelters and about nine hundred adoptions a year. Before any AI tool existed, a volunteer read each adopter's paper questionnaire and matched them to an animal by hand, about forty minutes of reading and judgment per case. Roughly seven times in ten, that match held. Three times in ten, the animal came back within ninety days.
Amberdale Match pitched a tool that reads the same questionnaire and scores compatibility in under a second, backed by a one-page claim of ninety six percent accuracy, a number that would cut Silverbrook's return rate by more than half.
At its worst, trusting an unverified number doesn't just risk one bad match. It risks locking a nonprofit's whole shelter network into a year-long, eighteen-thousand-dollar contract built on a number nobody outside the vendor ever checked. Running a real replication check costs about two weeks nobody had budgeted for. Against a year-long contract that size, that trade is an easy one to accept.
What I would leave alone: the free volunteer shift-scheduling assistant Silverbrook already uses gets none of this scrutiny. Nobody's adoption outcome depends on its output, so demanding a full replication test there would burn real hours for no real protection.
The lesson: a good number and a good report are not the same thing. A report earns your trust by showing its own work. A number just asks for it.
Now here is the same thing as a story
The short version above is what you'd say defending this call to Silverbrook's board. Read this one for how close the mistake actually came.
Selin Marroquin could read a shelter intake form and know, just from the tone of the handwritten notes, which volunteer should make the follow-up call. She had run Silverbrook's operations for six years and vetted a dozen small tools the same easy way: ask for a one-page summary, skim it, decide.
For most of those tools, that was plenty. They were free, or month to month, and swapping one out cost an afternoon, not a contract.
Amberdale Match's demo went well. The compatibility score matched what shelter staff already believed about three sample cases. The vendor's one-pager claimed ninety six percent accuracy, tested against "a broad sample of real placements."
Jono Reyes had joined Silverbrook's board four months earlier, after a decade in a data team at an insurance company. Sitting in on the pitch, he asked one question: "Which version of their model made that ninety six percent number, and whose cases were in the test?" Selin didn't know. Neither, it turned out, did Amberdale's own sales rep on the call.
Selin had considered simply asking Amberdale for another live demo instead of a harder replication test. She dropped that idea. A demo is scripted to the three cases a vendor already knows will look good. It could never tell her anything about the other thirty seven.
She asked Amberdale Match directly for their eval set and the version pin behind the number. Amberdale stalled for twelve days, then sent a spreadsheet with forty case descriptions, no names, no dates, no model version listed anywhere. Selin ran those same forty cases through Amberdale's live tool herself. It matched its own report on twelve of them, and disagreed sharply on the rest. Real accuracy against Silverbrook's own cases: seventy one percent, not ninety six.
The old decision that let this get so close to a signature wasn't the demo, and it wasn't Amberdale's number. It was Silverbrook's habit of treating one page as due diligence, a habit built for tools that cost nothing to walk away from.
Selin pulled the vendor packet off the board's agenda and sent Amberdale five questions: the eval set, the version, the failure cases, the confidence range, and who ran the test. She gave them ten business days to answer all five before Silverbrook would even reopen the conversation.
Six months later, a second vendor pitched a similar tool in the same board room, with the same kind of confident one-pager. This time Selin ran AUDIT before the meeting ever happened, asking for the eval set and version in her very first reply. The vendor answered honestly within nine days. Their real accuracy on Silverbrook's own cases came out to eighty four percent, close enough to their claim to trust. Silverbrook signed that one, and caught the gap two weeks into evaluation instead of three months into a contract.
AUDIT, in five checksNot a lecture on being suspicious of every vendor. AUDIT is the specific test a report has to pass before its number earns your trust.
The recap, one line per letter: ask is checking who ran the eval and who it flatters, uncover is finding the actual cases behind the number, demand is pinning the exact model version, isolate is noticing what the report never shows, and test is running it yourself before any signature.
And if you want to be sure it really works, try it somewhere elseSame five checks, a used vehicle marketplace instead of an animal shelter network. A different old decision breaks the second report.
Wrenchbay runs an online marketplace for used trucks and trailers. Devraj Okafor, a product manager there, was evaluating Panelscan, a vendor pitching AI damage assessment from uploaded photos, meant to replace an inspector's forty-minute walkaround. Mapped onto AUDIT: ask is checking that Panelscan's own claims team ran and scored the eval with no outside review. Uncover is finding that Panelscan's report only listed "a large internal photo set," with no named source. Demand is asking which exact model version produced the ninety one percent damage-detection rate, since Panelscan had shipped two model updates that same quarter. Isolate is noticing the report showed only caught damage, never missed damage. Test is running twenty of Wrenchbay's own listing photos through Panelscan's live tool before signing anything.
The old decision here isn't a one-page-summary habit, it's a different reversal: Wrenchbay's procurement process merged the technical vendor review and the contract negotiation into a single meeting, run by the same category manager under a deadline to close the quarter. That merge made sense when Wrenchbay only ever bought off-the-shelf software with no real model behind it. It stopped making sense once one meeting had to both judge an AI vendor's claims and negotiate its price, with no natural pause between the two.
Devraj split the meeting in two: a technical AUDIT review first, contract terms only after Panelscan passed it. On the first pass, Panelscan's real detection rate against Wrenchbay's own twenty photos came out to seventy eight percent, good enough to proceed, but with a required quarterly retest written into the contract instead of a one-time check.
Back at Silverbrook, the same pattern held over the following year: every time Amberdale pushed a new model version, Selin re-ran the same forty held-out cases before trusting the new number.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "check who ran the eval, on what data, on which version, then verify it yourself," and stop.
Cost: there's no time to run a full replication before a decision is due. Say so honestly, and check the smallest, riskiest slice of cases instead of skipping the check entirely.
The model gets better, for real: if Amberdale's next version genuinely closes the gap and matches its own claim, that's still worth re-verifying once, since a good version now says nothing about the version running next quarter.
Where people run it wrong.
They treat a confident one-pager as due diligence, because reading it feels like the work of checking.
They accept "a broad internal test" as an answer instead of asking for the actual eval set.
They check the number once at signing and never again, even after the vendor ships a new model version.
How to use it live. The moment a vendor hands you a single number, ask back: which version made it, on whose data, and can I run it again myself? Let the answer to that decide how much you trust the rest of the pitch.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the vendor simply refuses to share their eval set?" Response: that refusal is itself the answer, per the decision tree; either walk away or price the unknown risk into the contract terms.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Evaluating AI vendors as a buyer
- #1 List the ten questions you would ask every AI vendor before a pilot.
- #2 How do you evaluate a vendor's quality claims without running your own eval?
- #3 Design the pilot you would run to evaluate two competing AI vendors.
- #4 What contractual terms matter specifically for AI vendors and not for other software?
- #5 How do you assess a vendor's model dependency and what happens if their provider changes terms?
- #6 Describe the data handling questions you would put to a vendor on behalf of your security team.