CaseIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #2
How do you evaluate a vendor's quality claims without running your own eval?
LEAD the coffee-break question that found the gap five weeks before anyone else did
Lissett Group is a mid-size law firm handling vendor contracts for a string of manufacturing clients. Priya Sundaram runs legal operations there. Clarisage sells an AI tool that reads contracts and flags risky clauses, and publishes a 94 percent clause-extraction score on its own benchmark.
The direct answer
Don't try to rebuild the vendor's whole eval. Instead, pull twenty of your own real documents, run them through the tool, and spot-check the ones it marks "high confidence, no issues." Watch how often you overturn that confident call. That gap, between what they claim and what survives your own check, is the number that will tell you the truth weeks before a published benchmark ever could.
Do this, in order
Test the claim on a small sample of your own documents before trusting it at scale.Why: a benchmark score describes someone else's documents, not yours.
Spot-check the outputs marked "high confidence," not the ones it flags as uncertain.Why: uncertain flags already get a second look; confident wrong answers are the ones that slip straight through.
Ask for the vendor's own eval set and test conditions in writing.Why: a score with no visible test set could mean anything, including a set stacked with easy documents.
Track the disagreement rate weekly, not just at the end of the pilot.Why: it moves for weeks before a bad document actually gets through, giving you time to act early.
Set a threshold in advance for when you'd pause and demand retraining.Why: without one, a rising gap just feels like bad luck until it's already cost you a real mistake.
How to answer this, stage by stage
Nobody is scoring whether you can recite what a good eval set looks like. They're scoring whether you know what to check when you don't have time to build one.
Stage 1
Scope it to one real claim
Say it like this
"I'll answer this for a real case: a law firm's legal ops team deciding whether to trust an AI contract-review vendor's published accuracy score."
Why this works
Keeps a broad evaluation question from turning into a lecture on statistics.
Stage 2
Say the method out loud
Say it like this
"I'll use LEAD. Link the claim to what actually matters here, find the early signal, name how it gets gamed, then say what I'd do at each threshold."
Why this works
Shows the interviewer you have a repeatable move for "trust but verify," not just a gut feeling.
Stage 3
Reframe what the claim should be tested against
Say it like this
"A 94 percent score on their benchmark doesn't answer the real question, which is how it does on our contracts, with our defined terms and our clause structures."
Why this works
This is where a strong answer separates from someone who just repeats the vendor's number back.
Stage 4
Give the one decision
Say it like this
"Run twenty of our own contracts through it, then spot-check every 'high confidence' result by hand. If the tool and my team disagree less than one time in twenty, I trust it more. If it's one in five, something is wrong before we ever find out from a real mistake."
Why this works
This is the direct answer, concrete enough that a panel could push on the exact number.
Stage 5
Prove it with the near miss
Say it like this
"Three NDAs went to a client marked 'high confidence, no issues.' One had a missing indemnification cap. Priya caught it at eight at night doing a spot-check nobody had asked her to do."
Why this works
Turns an abstract trust question into a specific afternoon that could have gone the other way.
Stage 6
Close on the one line
Say it like this
"You don't need their whole eval. You need your own twenty documents and the honesty to check the ones marked confident, not just the ones marked unsure."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Let's learn
Here is what happens when a team accepts a vendor's published number as the whole story.
Before Clarisage, Priya and two paralegals split contract review by hand, about forty-five minutes each, roughly ninety hours a month across a caseload of a hundred and twenty vendor contracts. Clarisage read every contract in seconds and flagged clauses to check, which cut the team's hands-on time to a fraction of that, on paper.
Lissett Group picked the left side for the first five weeks, without deciding to.
Here's the turn: the extra speed was never the risk. The risk was that Lissett's leadership accepted Clarisage's published 94 percent score as the bar for going live, and never asked what documents that score was actually tested on. It felt reasonable: the firm had never bought an AI tool before, and a benchmark number felt like the same kind of proof a case citation would be.
Weekly disagreement rate: how often Priya's team overturned a "high confidence" call
The gap climbed for a month before it reached a client's desk. Nobody was watching this number, because nobody had agreed it was the number that mattered.
At its worst, trusting a published score without checking it doesn't just waste review time. It risks a client receiving a contract with a real hole in it, stamped "no issues" by a tool everyone assumed had been checked.
The choice I would take back
Leadership accepted Clarisage's own 94 percent benchmark as the bar for launch, instead of asking what kind of documents it had been tested on. That made sense for a firm buying its first AI tool with no template yet for skepticism. It stopped making sense the moment real contracts, full of the firm's own defined terms, turned out to look nothing like the benchmark's clean sample set.
What I would leave alone: Clarisage's basic clause-flagging on plain, boilerplate sections doesn't need this scrutiny. Standard governing-law and notice clauses are the same in almost every contract, and the tool catches those reliably regardless of which benchmark it was trained on.
The lesson: a number a vendor built to sell you the tool is not the number that protects you from it. The one that matters is smaller, cheaper to get, and it's the one nobody thinks to ask for until something almost goes out the door wrong.
Now here is the same thing as a story
The short version above is what you'd say defending the pilot to the partners. Read this one for how close the missed clause actually came.
Priya Sundaram had run legal ops at Lissett Group for six years, long enough to read a vendor contract and know which clause a client would ask about first. When Clarisage went live, its "high confidence, no issues" tag agreed with her own read often enough that within a month she'd stopped opening every flagged contract herself, trusting the green tag to mean what it said.
Nobody at Lissett Group had actually measured the right side until five weeks in.
Knowledge spark: why would a benchmark score not carry over to real use?
A benchmark is a fixed set of test documents a vendor picked to grade itself against. If those documents are cleaner or simpler than what a real customer actually has, the score can be true and still tell you almost nothing about your own files.
Five weeks in, over coffee with a legal ops peer at another firm, Priya mentioned the tool's accuracy score in passing. The peer set down her cup: "Wait, you just believed their number? Did you test it on your own stuff first?" Priya hadn't. That night, she pulled twenty recent contracts and ran them back through Clarisage herself.
Three NDAs from the same vendor family had been marked "high confidence, no issues" and were sitting in a folder ready to go out to a manufacturing client. All three were missing a standard indemnification cap the firm always included, a clause Clarisage's training documents had apparently never needed to flag, because the benchmark it was built on didn't use the firm's specific templates.
The tool wasn't lying. It had just never been asked whether it understood Lissett Group's own definition of "no issues," and nobody had checked before letting three contracts sit one folder away from a client's inbox.
Priya pulled the three NDAs the same night, fixed the missing clause by hand, and built a standing rule the next morning: every "high confidence" tag on a new document type gets a human spot-check for the first month, whether it feels necessary or not.
This is what Priya started asking Clarisage for, in writing, after that night.
The colleague's remark didn't create the gap. It just made someone finally look for it.
LEAD, in one screenNot a replacement for a real eval. LEAD is what tells you the fastest, cheapest way to catch a bad claim before it costs you.
L
Link. The outcome that actually matters.
Whether Lissett Group's own contracts get flagged correctly, not how Clarisage did on a benchmark set built for every customer at once.
Without naming this, a vendor's own score becomes the accidental north star.
E
Early signal. What moves before the outcome does.
The weekly disagreement rate between Clarisage's confident calls and a human spot-check, which climbed for a month before any real document actually slipped through.
This is the hardest step, and it's the entire reason LEAD exists instead of just waiting for a mistake.
A
Abuse. How this metric gets gamed.
A benchmark built on clean, simple documents will always score high; a vendor never has to say so unless you ask what's actually in the test set.
Every metric can be hit without doing the real work; a published score is the easiest one to game unintentionally.
D
Decision. What you'd actually do at each threshold.
Below one disagreement in twenty, trust the flow with light spot-checks. Above one in five, pause new document types and demand retraining before going further.
A metric nobody acts on is a dashboard decoration, not a decision.
Four moves, none of them requiring you to build a full evaluation pipeline first.
The recap, one line per letter: link is tying trust to Lissett's own contracts, not the vendor's benchmark, early signal is the weekly disagreement rate that rose for a month before the near miss, abuse is a benchmark quietly built on easy documents, and decision is a stated threshold for when confidence stops being enough.
The published score sits in the worst corner. An afternoon of your own testing moves you to the best one.
And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary chain's diagnostic-imaging vendor instead of a law firm. A calibration setting breaks the second story, not a benchmark.
Cutlass Veterinary Group piloted Vetraya, an AI tool reading X-rays for fractures, published at 91 percent sensitivity on a public veterinary imaging set. Mapped onto LEAD: link is trusting Vetraya's read on Cutlass's own patients and their own aging X-ray machine, not the vendor's benchmark equipment, early signal is the rate at which a vet overturns a "clear" read on a small sample of recent films, abuse is that public imaging benchmarks skew toward common breeds and obvious fractures, and decision is Owen Kessler's team agreeing in advance: above one overturned "clear" read in fifteen, pause and recalibrate before reading another film.
The old decision here isn't an accepted benchmark, it's a calibration one: the clinic used Vetraya's default confidence threshold straight out of the box, tuned on the vendor's own newer imaging equipment, without adjusting it for the contrast levels their older X-ray unit actually produced. That made sense in the vendor's demo, which used its own machine. It stopped making sense once Cutlass's equipment quietly shifted what "confident" should even mean on their films.
Contracts and films needing correction after being marked clean, per month
Both drops came from the same move: a small sample of real documents or real films, checked by hand, before trusting the confident tag.
Swap the trigger and it still runs.
Speed: an interviewer caps you at a minute. Say "twenty of your own documents, spot-check the confident ones, watch the gap," and stop.
Cost: there's no time or budget for even a small sample before go-live next week. Say so honestly, and add a mandatory human check on every confident result for the first month instead of skipping the check entirely.
The vendor gets better, for real: if a later version genuinely closes the gap between claimed and measured accuracy, that's the early signal doing its job, telling you it's safe to loosen the spot-check rate.
Where people run it wrong.
They treat a published benchmark score as proof instead of a claim still waiting to be tested.
They spot-check only the outputs the tool already flagged as uncertain, missing the confident wrong answers entirely.
They wait for a real mistake to reach a client instead of watching the disagreement rate that was rising for weeks first.
How to use it live. If an interviewer asks how you'd evaluate a claim with zero time to build anything, say "twenty real documents and an honest spot-check of the confident ones" and give the actual disagreement threshold you'd set. A number beats a philosophy every time in this kind of question.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "how do you measure or trust this" question?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It finds the number that moves before the real outcome does.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Sundaram, legal operations manager at Lissett Group for six years, who caught the missed clause with an unplanned late-night spot-check.
3 · THE HABIT
What did Priya stop doing because the green tag kept agreeing with her own read?
Tap to flip
ANSWER
Opening every flagged contract herself. Within a month she trusted the "high confidence, no issues" tag without a second look.
4 · THE EARLY SIGNAL
What number would have told the real story weeks before the near miss?
Tap to flip
ANSWER
The weekly disagreement rate between Clarisage's confident calls and a human spot-check, which climbed from 8% to 27% over a month.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Accepting Clarisage's published 94 percent benchmark score as the launch bar, without asking what documents it was actually tested on.
6 · THE NUMBER
Fill in the blank: on Lissett Group's own twenty-document sample, the measured accuracy was ___ percent, against a published 94.
Tap to flip
ANSWER
61 percent, measured by meaning, not just by whether a clause got extracted at all.
7 · THE REPLAY
Same three NDAs, same missing clause, but the spot-check rule already exists. What changes?
Tap to flip
ANSWER
A junior associate catches the missing indemnification cap during the mandatory first-month spot-check, days before the contracts would have shipped, not the night before.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Cutlass Veterinary Group's Vetraya pilot. The reversal is a calibration default, using the vendor's out-of-box confidence threshold on older imaging equipment it was never tuned for.
Check yourself Score: 0 / 0
Multiple choice
1. Why is spot-checking "high confidence" outputs more important than checking the uncertain ones?
A. Confident outputs are always wrong more often.
B. Uncertain outputs don't need to be checked at all.
C. Uncertain flags already get a second look; confident wrong answers slip straight through.
D. Vendors only lie about confident results.
Show hint
Look at priority bullet 2 and the E step.
Show answer
C. A tool flagging its own uncertainty already invites a human check. A confident wrong answer is the one nobody re-examines.
True or false
2. True or false: Clarisage's published 94 percent score was a fabricated or dishonest number.
True
False
Show hint
Look at the knowledge spark about benchmarks.
Show answer
False. The score was likely true on the vendor's own benchmark. It just didn't describe how the tool would do on Lissett Group's specific contracts.
Fill in the blank
3. Fill in the blank: the weekly disagreement rate rose from 8 percent in week one to ___ percent in the week before the near miss.
Show hint
Look at the line chart in Section 1.
Show answer
27 percent. It had been climbing for a full month before anyone was watching it as a signal.
Short answer, where it wouldn't matter
4. Name a part of Clarisage's output where this extra scrutiny genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Plain boilerplate clauses like governing law or notice provisions. They're nearly identical across contracts, so the tool catches them reliably regardless of the training benchmark.
Short answer, apply it yourself
5. Think of a tool at your own job that came with a published accuracy or satisfaction number. What would a twenty-item version of your own eval look like?
Show hint
Pick a small, real sample from your own work and imagine spot-checking its most confident results.
Show answer
Model answer: Pull a small batch of your own real cases, run them through the tool, and check the results it seemed most sure about first, since those are the ones nobody else is double-checking.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Accepting the vendor's own published benchmark as the launch bar. It made sense for a firm with no prior template for testing an AI vendor's claims.
Before you close the answer
Why this works
Tests whether you'll accept a vendor's own number at face value or find a cheap, fast way to check it against your own reality before it costs you.
Follow-up traps
"Isn't twenty documents too small a sample to mean anything?" Response: it's small on purpose, it's meant to catch a large, obvious gap fast, not replace a full statistical eval; a large gap on twenty documents is still a real signal.
"What if the vendor won't share their eval set at all?" Response: that refusal is itself useful information, and it's exactly why testing your own sample matters more than trusting theirs.
If pressed
Lissett Group's fix wasn't just more spot-checks, it was retraining Clarisage on twenty of the firm's own past contracts with the missing clause pattern labeled, which closed most of the gap within three weeks.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.