What metric captures a model's willingness to say it does not know?
A flat accuracy score can sit still for months while the one number that would have warned you drifts the whole time. The job is building that number on purpose, not admiring the one that already looks good.
- Track abstention rate on a golden set built only from questions the model should not answer with confidence.Why: overall accuracy blends thousands of easy, well supported answers with the handful of hard ones that actually cost money, so it can hold flat while the real risk drifts for weeks.
- Pair it with accuracy on the questions the model does answer.Why: without that second number, a model that hedges on everything scores a perfect abstention rate while never actually helping anyone.
- Gate every release on both numbers, not on a slower quarterly check.Why: a small prompt tweak shipped between quarterly reviews is exactly when a drift gets three or four weeks to run before anyone looks at it.
- Set real thresholds, not a feeling.Why: a number nobody has to clear before shipping is a chart on a dashboard, not a metric anyone acts on.
- Leave the easy, document grounded questions tuned toward being direct.Why: hedging on a question with one clear right answer sitting in the contract costs real usefulness for no safety gained.
- Treat a drop in accuracy on answered questions as a different, more urgent bug than a drop in abstention rate.Why: one means the model is getting confident answers wrong, the other means it is over answering, and they need different fixes.
How to answer this, stage by stage
Nobody's grading whether you can name a metric. They're grading whether the metric you name would have caught anything before it turned into an eighty five thousand dollar problem.
Let's learn
Marlein Vantry has run answer quality for Caselight for two years. Caselight is an AI research tool for in house counsel. It answers questions about contracts, regulations, and precedent, in plain paragraphs, no citation menu to dig through. Before it, counsel at a client company spent about three hours a day chasing down answers by hand. After it, that dropped to about forty minutes, a real two hours and twenty minutes back, every single day.
For nine months, the number Marlein read out at the Monday review barely moved. Overall accuracy: 92 percent, give or take a tenth of a point, graded monthly against 1,000 real questions by a senior counsel panel. Clients ran something like fourteen thousand questions a month through Caselight: lease clauses, non compete enforceability, a new state privacy rule, whether a boilerplate indemnity clause meant what it looked like it meant. When Caselight answered in its flat, confident voice, it was right almost every time. When it hedged, it was usually genuinely unclear too. Counsel learned fast that the confident voice meant something, and stopped double checking the confident ones. Why would they. It kept being right.
In week fourteen, the team shipped a prompt change to fix a real complaint: users kept saying Caselight hedged too much, that "this may depend on jurisdiction" on a routine question was just noise. The fix worked, on paper. Complaints about unhelpful hedging fell from about 120 a week to 45 within a month. Nobody re ran the hard question review before shipping it, because that review only ran once a quarter, and the number that gated the release, overall accuracy, hadn't moved.
What that fix actually did was teach Caselight to sound sure more often, including on questions nobody should sound sure about. Correct abstention on the hard set slid from 82 percent to 58 percent over the following ten weeks. Restated the other way: the rate of confidently wrong answers on that same hard slice climbed from about 18 percent to about 42 percent.
The extra confidently wrong answers on the hard set, on their own, are not really the problem. It's 150 questions out of fourteen thousand a month, barely a dent in overall accuracy. The real problem is what in house counsel did next: nothing. They kept trusting every confident answer exactly the way they always had, because nobody told them the confident tone had quietly stopped meaning what it used to mean.
The lesson: overall accuracy is a blended number, built mostly from thousands of easy questions with one obvious right answer. It can hold still for months while the one number that would have warned you, correct abstention on the hard slice, quietly drifts the entire time underneath it. Logging a number and gating a release on it are two different decisions. Caselight only ever made the first one.
Now here is the same thing as a story
Read this when you want to feel why the gap between the two numbers mattered, not just know that it did.
Ask Marlein Vantry which questions Caselight should never answer with a flat yes, and she can list them without checking a doc: anything where two appellate districts disagree, anything about a state Caselight has thin coverage in, anything where the real answer depends on a fact the contract simply never states. Two years running answer quality there will do that to a person.
The habit at client companies thinned in three quiet steps nobody wrote down anywhere. First, routine lease renewals stopped getting a second look, because Caselight's confident answers on those had been right for months. Then most contract questions stopped getting one. By month seven, the rule inside most legal teams using Caselight had become simple, unspoken, and completely reasonable given what they'd seen: if it sounds sure, it's sure.
The trigger, when it came, was one wording change nobody thought was risky. In week fourteen, the team shipped a prompt update meant to cut down on a real complaint: too much hedging on routine questions. It worked. Complaints fell from about 120 a week to 45 within a month. Everyone was pleased. The quarterly hard question review, the one graded by senior counsel against genuinely unclear cases, was not due again for another eight weeks, and the number that gated the release, overall accuracy, hadn't moved.
Five weeks later, a contracts lead asked Caselight a plain question: would a non solicit clause in a departing employee's agreement survive termination, in a state where a recent appellate ruling had actually split the districts on exactly that point. Caselight answered in the same flat, confident voice it always used for a clean question. "Yes, the clause survives termination." No hedge. No mention of the split.
The company relied on that answer inside a live negotiation. Three weeks later, outside counsel on the other side of the deal caught the split, and the whole agreement had to be reopened. Eighty five thousand dollars in renegotiation costs and outside counsel fees. Three weeks of the deal team's time, gone. Then came the overcorrection: legal ops mandated a second lawyer review every single Caselight answer used in any negotiation, starting that week, no exceptions. The tool that used to hand counsel back two hours and twenty minutes a day gave almost none of it back anymore. Every answer needed a human read regardless, so the software cost and the review time both got paid, on every deal, not just the risky ones.
Marlein's first instinct was to ask whether the model itself had gotten worse. The quarterly hard question review happened to be scheduled for the following week anyway, so she pulled it forward. The senior counsel panel re graded all 150 questions in the hard set. Caselight had been getting a smaller share of them right for ten straight weeks. Not because anyone told it to. Because the same wording change that quieted the hedging complaints had also taught it to sound sure more often, including on the ones nobody should sound sure about.
Here's the part that actually explains it. That number, correct abstention on the hard set, had been logged every week since Caselight launched. It just was never the number that gated a release. Overall accuracy was, and overall accuracy barely blinked, because 150 hard questions are a rounding error next to fourteen thousand easy ones a month.
So here is the decision I would take back. We built the hard question check at launch. We decided, quietly, that it only needed senior counsel grading once a quarter, because that grading took real lawyer time and doing it every release felt like overkill. I would make correct abstention on the hard set a release gate from day one, checked weekly on a rolling sample, not necessarily senior counsel graded every single week, but graded often enough that a ten week slide gets caught inside two.
Run the same fourteen weeks again with that gate in place. Correct abstention on the hard set still starts sliding in week ten, same as before. But now someone has to clear it before week eleven's release ships. The release that caused the slide gets held. The eighty five thousand dollar negotiation never happens, because the wrong answer never went out the door.
One design trusted a number because it happened to already look healthy. The other watches the number built specifically to catch what the healthy one hides.
What I'd tell myself, back at launch: the number you gate a release on is a decision, not a default. We picked the flattering one because it was already there, cheap to compute, and it had never once caught a real problem. That last part was never proof it was the right number. It was only proof we had never tested it against something like this.
LEAD, run on Caselight's own numbers
Four letters instead of five, built for a question that asks you to name a number, not diagnose a drop.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was gating abstention on how many documents the retrieval step pulled back for a question, treating a thin document count as the signal for uncertainty. It lost because a genuinely contested question can return five documents that simply disagree with each other, so a full retrieval count looks exactly like a confident, well supported answer even when the underlying law is split. The AI specific failure mode worth naming by name is a particular kind of confident hallucination: the model treating a genuine legal split as settled because its training and its retrieved documents lean one way, and stating that lean as fact. The guardrail is the hand graded hard question golden set itself, refreshed by senior counsel every quarter so it keeps testing genuinely unsettled questions instead of ones the model has since learned cold. That guardrail isn't free. Building and re grading a golden set of genuinely hard questions costs real senior lawyer hours every quarter, and gating a release on it will occasionally hold back a prompt tweak that was actually fine, a small delivery speed cost accepted only because a confidently wrong answer on this slice costs far more than a slower release cycle. And the bar isn't zero wrong guesses, no probabilistic model can promise that. It's an abstention rate above 85 percent on a rolling 100 question sample of the hard set, paired with accuracy above 95 percent on whatever it does answer, checked before every release, not once a quarter.
And if you want to be sure it really works, try it somewhere else
Same LEAD, a claims coverage tool for insurance adjusters this time, nothing about lawyers or contracts anywhere in it.
Coverwell is an AI tool that answers coverage questions for claims adjusters at a regional insurer: does this policy cover this specific loss, given the exact wording of this endorsement. Rhian Onyema leads claims quality there.
L, link. Fewer confidently wrong coverage calls that get paid out or denied and then have to be reversed, since a claim flagged for a human underwriter is recoverable, and a confidently wrong call that already went out often costs a clawback, or worse, a regulator's attention.
E, early signal. Abstention rate on a golden set of about 120 genuinely ambiguous scenarios: conflicting endorsements, disputed cause of loss wording, a claim missing its inspection report. Checked weekly, apart from overall coverage call accuracy, which sat near 95 percent the whole time and moved less than half a point.
A, abuse. Tell Coverwell to kick every ambiguous seeming claim to a human and abstention rate looks perfect, while claims that used to clear in a day now sit in a queue for a week. This number only means something next to accuracy on the calls it makes on its own.
D, decision. Abstention rate on the hard set above 80 percent, accuracy on the calls it makes alone above 96 percent, both gating every model or prompt update, not just the quarterly one.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the one number that gates a release, don't spend the time justifying why overall accuracy isn't enough.
Cost: there's no senior underwriter graded golden set funded yet at a smaller shop. Don't skip the gate, hand grade a stratified sample of forty or fifty scenarios yourself, weekly, until a real one exists.
The model got better, for real: say overall coverage call accuracy improved a full point that quarter. That's not the same claim as the hard slice being calibrated. A model can improve on average while one narrow, expensive slice drifts the entire time underneath the average.
Where people run it wrong.
They watch the number that already looks good, because it's the one that was already on the dashboard.
They read a flat overall number as proof nothing needs fixing, and skip building the harder golden set entirely.
They build the abstention rate gate and forget to pair it with accuracy on the calls made alone, so a model that hedges on everything sails through looking perfectly safe.
How to use it live. Say the two numbers are usually not the same number, out loud, before naming either one: "the number that says it's fine and the number that would have warned me early are usually not the same number." That buys you a beat to pick the real one instead of naming whichever metric comes to mind first.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the model is right to sound confident more often than your golden set assumes?" Response: that's why the golden set is graded by senior counsel against genuinely split precedent and thin coverage, not a guess, and it gets re reviewed. A model earning real confidence on a question the panel called ambiguous is a reason to update the golden set, not evidence the metric was wrong.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?