How do you eval a feature whose quality depends on the user's own documents?
Varro's shared benchmark barely moved for a year. The first time anyone checked its flags against one customer's real vendor contracts, precision fell twenty nine points, and on one clause type, fifty six.
- Replace the one shared benchmark score with a rotating, per customer live sample, reviewed against that customer's real documents.Why: no shared ground truth exists when quality depends on documents that differ customer to customer.
- Size each customer's sample past the point where a couple of contracts could flip the number, before reporting anything.Why: Foscari Freight's first honest number needed ninety six flags across forty one contracts. Five contracts would have told nobody anything.
- Report precision and recall as a range tied to sample size, never a single blended percent.Why: a number like sixty one percent, stated flat, claims a confidence the sample doesn't actually support.
- Treat a gap between the shared benchmark and a customer's live sample as the finding itself, and go find the clause driving it.Why: this is what actually found the fuel pricing problem, not a vague sense that something felt off.
- Keep the shared benchmark only as a minimum floor check between releases, never as the real quality number for a paying customer.Why: the benchmark and a live customer's paperwork are two different questions wearing the same percentage sign.
- Re-run a customer's live sample whenever their document mix changes meaningfully, not on a fixed calendar alone.Why: Foscari's trailer leasing contracts were new to their upload mix the same quarter recall on that clause fell to eighteen percent.
How to answer this, stage by stage
Nobody's testing whether you know the words precision and recall. They're testing whether you'll trust a number that looks stable, or ask whose documents it was actually measured against.
Let's learn
What does it mean for an AI tool to be "right," when every customer feeds it a completely different pile of paperwork?
Before Varro, Foscari Freight's two person procurement and legal team read every new or renewed vendor contract by hand before anyone signed it. About forty five minutes a contract, done carefully when there was time and done in a rush during renewal season when there wasn't. Looking back at contracts that later caused a problem, the team figures they caught maybe four risky clauses out of five, missing the rest under deadline pressure.
Then Basford built Varro. Upload a contract, get flags back in under a minute, each one naming the clause and the reason it's risky. Company wide, checked against Basford's shared benchmark, Varro's flags matched a real risk clause ninety percent of the time, and had for over a year.
The turn is not that Varro got worse. Company wide, on the clause types the benchmark did cover, Varro's recall stayed close to eighty four percent all year. The turn is that ninety percent was always a benchmark number, not a Foscari Freight number, and those had quietly become two different questions.
Here's the arithmetic behind the numbers that finally got checked. This quarter, Foscari Freight, a trucking company, uploaded forty one vendor contracts: fuel supply agreements, trailer leasing agreements, roadside maintenance contracts. Varro raised ninety six flags across all forty one. A compliance reviewer checked every one of those flags against her own read of the contract.
The thirty two price escalation flags were nearly all on Foscari's fuel supply contracts. Fuel pricing in trucking commonly uses an index linked formula, the price moves with a published fuel index inside agreed bounds, which is normal and low risk. Varro kept reading that formula as the unbounded, vendor sole discretion clause its benchmark had been built to catch, because the benchmark had never once included a fuel contract.
Recall told the other half of the story. Eleven of Foscari's forty one contracts were trailer leases, and every one of them carried a one sided early termination clause: the vendor could end the lease with no notice, Foscari needed ninety days. Varro caught that clause in two of the eleven, eighteen percent, against roughly eighty four percent recall company wide on the clause types the benchmark did cover.
At its worst, a procurement team that trusts a confident sounding flag report on faith could let a genuinely dangerous termination clause ride for the next three years of a fuel relationship, and the miss would never show up in the company wide number anyone was watching, because that number was never measuring Foscari's contracts to begin with.
What I'd leave alone: Varro's precision on NDA and confidentiality only contracts doesn't need this treatment. Those contracts are heavily standardized across every industry Basford serves, and precision there holds in the mid nineties no matter which customer uploads them. Building a special live sample process for a customer that only ever uploads NDAs would spend audit time where there's no real gap to find.
The lesson: a benchmark score can be completely honest and still be the answer to a different question than the one that matters. Ninety percent told Basford Varro was fine. It never told them ninety percent right on the contracts the benchmark had already met, and thirty four percent right on the one clause type unique to a customer whose paperwork looked nothing like it.
Now here is the same thing as a story
Read the long version below when you want to feel why a number this stable still went wrong, not just be told that it did.
Zevi Dobrescu can read a fifty page vendor contract and find the one clause that actually matters before he's finished the second page. He spent three years pricing supply chain risk before he joined Basford, and he built Varro's original benchmark himself, sixty contracts, hand labeled, the week the product first shipped.
For the first year, Zevi did two things every week: he read the benchmark score off the release dashboard, and he pulled a handful of live customer flags at random, across whichever customers had signed up recently, and checked them by eye. Nothing formal. Just a feel for whether the real world still looked like the benchmark.
The benchmark score sat in the high eighties and low nineties, month after month. Somewhere around month nine, Zevi's spot checks thinned from a handful a week to one or two. The company was onboarding customers faster than one person could eyeball, and the number on the dashboard kept being fine, so there was always somewhere more urgent to look.
By month sixteen, Zevi wasn't spot checking live flags at all. He read the weekly benchmark number, saw ninety, and moved on. It had been steady for so long it had stopped feeling like something that needed checking.
It came back on an ordinary Thursday, not from an incident, but from a customer success rep mentioning, almost in passing, that a procurement lead at Foscari Freight had asked, half joking, why Varro never seemed to flag anything on their fuel contracts. Either they'd gotten lucky with clean vendors, she said, or the tool wasn't really looking.
Zevi's first instinct was the sensible one. Foscari might just have unusually clean paperwork. He'd seen customers like that before. But the comment sat with him, and instead of writing it off he did what the dashboard hadn't made him do in months: he pulled Foscari's own contracts and had someone actually check.
Forty one contracts, ninety six flags, checked by hand. Fifty nine held up. Sixty one percent, against a benchmark that had read ninety for a year. Narrowed to the price escalation flags on Foscari's fuel contracts specifically, it fell to thirty four. And eleven trailer leases, each with a one sided termination clause, that Varro had caught on only two.
The decision that opened the door went back to the week Zevi built that first sixty contract benchmark. It was a fine decision then. Basford had a dozen customers, mostly manufacturers, and public template contracts plus a few donated agreements covered what those twelve customers actually uploaded. Nobody decided, on purpose, that the benchmark would still be the whole quality story two years and two hundred and sixty customers later. It just kept being the number on the dashboard, and a number that keeps being fine stops looking like a number anyone chose.
Run that Thursday again with one change: a rotating live sample, a slice of customers checked against their own real contracts every quarter, weighted toward whichever customers had changed their document mix most since their last check. Foscari Freight, adding trailer leases to their upload mix for the first time, gets pulled into that sample automatically in week three of the quarter, not month sixteen after a customer has to ask why. Same forty one contracts, same reviewer, about seven hours of work, and the gap gets caught before three years of fuel contracts ride on a flag report nobody had actually checked.
One design let a two year old number speak for every customer forever after. The other asks each customer's paperwork the question that's actually theirs to answer.
What I'd tell myself, back the week we called sixty contracts enough: a benchmark this stable going quietly stale is the outcome that never announces itself. It just sits at ninety percent, forever, right up until someone whose contracts it never saw finally asks why.
BOUND: the arithmetic behind a benchmark that went quietly stale
Not a story question wearing a framework's clothes. This is an estimation problem, and BOUND is what turns "the benchmark says ninety" into an honest, checkable range.
Three things worth stating directly, since this is where the real judgment sits. The alternative Zevi's team tried first, and later dropped, was simply growing the shared benchmark, sixty contracts up to two hundred, hoping more examples would generalize better across customers. It lost because a bigger fixed benchmark is still a fixed benchmark: even two hundred hand picked contracts didn't include a single trailer lease or a fuel index pricing clause, because nobody picking those two hundred knew to go looking for either. The number barely moved. The AI specific failure worth naming by name is benchmark drift: a fixed eval set that stops representing the real, growing population of documents flowing through the feature, while its own score stays perfectly stable and gives no sign anything changed. The guardrail is the rotating live sample itself, weighted toward whichever customers changed their document mix most, plus one hard rule: any gap between the benchmark and a live sample gets investigated, not shrugged off as noise. That guardrail isn't free. It costs roughly 160 reviewer hours a quarter, spent on the customers most likely to be unlike the benchmark, against 1,950 hours for reviewing everyone, a bounded price for the one place a miss actually rides on a three year contract. And the bar was never zero misses. A feature reading thousands of contracts a quarter can't promise that. It's a range, checked per customer, tightened over time as the benchmark absorbs whatever a gap turns out to be, not one blended percentage standing in for two hundred and sixty different piles of paperwork.
And if you want to be sure it really works, try it somewhere else
Same five letters, patient lab results instead of vendor contracts, nothing about procurement anywhere in sight.
Amberlight is a patient portal from Purbeck Health, a telehealth network. Patients upload their own lab result PDFs, and Amberlight reads them and flags anything worth a follow up call: an out of range value, a concerning trend across visits. Ovidia Stancik runs eval on it.
The build-up: Amberlight's shared benchmark, fifty lab report templates from major national lab chains, holds flag precision at ninety two percent, stable for most of a year. Purbeck signed a new fertility clinic partnership this quarter. On a live sample of thirty of those patients, seventy flags, a nurse reviewer confirmed only thirty nine, fifty five percent.
The sanity check: Ovidia pulled the thirty one misses and found most shared one cause. The fertility clinic's lab used a differently ordered reference range column than the national chains in the benchmark, and Amberlight kept reading a normal hormone level against the wrong column, flagging it as abnormal. Recall told the same story from the other side: a real abnormal thyroid marker in that lab's own layout got caught only forty five percent of the time, against ninety percent company wide.
Same rank as before: sample each new document source's live output before trusting the shared benchmark, then chase the gap. The fix is the same shape too: route any flag from a lab format outside the benchmark to a nurse second read, until that lab's own live sample clears its own bar.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: never trust one shared benchmark for documents that differ by customer, sample each customer live, and chase the gap when it shows up.
Cost: there's no budget this quarter for both a bigger benchmark and a live sampling process. The sampling process wins. A hundred and sixty hours spent on the customers most likely to be different beats a bigger benchmark that still can't see every customer's own paperwork.
The model got better, for real: say Varro's overall company wide precision climbs to ninety three percent next quarter. That's not proof Foscari's own number moved with it. The customers the benchmark already resembles could have gotten even easier, while the fuel contract gap stayed exactly where it was.
Where people run it wrong.
They read one stable benchmark score as proof nothing anywhere is broken, and never check it against a real customer's own documents.
They fix a low live sample number by tweaking the model's confidence threshold, instead of asking whether the benchmark ever saw this kind of document at all.
They validate once at launch and never again, so the customer base grows past what the benchmark represents and nobody notices until someone asks.
How to use it live. Say the real question out loud before answering it: "is this one shared score actually measuring what this specific customer's documents look like, or does it need its own check." That buys a beat to think instead of quoting a number you haven't asked whose documents it came from.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just make the benchmark bigger instead of building a whole sampling process?" Response: already tried, sixty grew to two hundred contracts and the number barely moved, because a hand picked set still can't guess every customer's specific vendor wording in advance.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Eval design for product teams
- #1 What makes an eval product-relevant rather than research-relevant?
- #2 Design an eval for a feature that drafts email replies.
- #3 How do you decide between automated evals and human review?
- #4 Explain the tradeoffs of LLM-as-judge for a product team.
- #5 How do you validate that your judge model agrees with human raters?
- #6 Describe a rubric that a non-technical reviewer could apply consistently.