ConceptAdvancedQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #16

What metric captures a model's willingness to say it does not know?

A flat accuracy score can sit still for months while the one number that would have warned you drifts the whole time. The job is building that number on purpose, not admiring the one that already looks good.

The metric
Track abstention rate on a golden set built only from questions the model should not answer with confidence: split precedent, thin document coverage, facts the contract never states. Check it every week, on its own, never folded into overall accuracy, because overall accuracy can sit flat for months while this number quietly slides. Pair it with accuracy on the questions the model does answer, or a model tuned to hedge on everything will look perfect on this metric alone while doing none of its job.
In this order
  1. Track abstention rate on a golden set built only from questions the model should not answer with confidence.Why: overall accuracy blends thousands of easy, well supported answers with the handful of hard ones that actually cost money, so it can hold flat while the real risk drifts for weeks.
  2. Pair it with accuracy on the questions the model does answer.Why: without that second number, a model that hedges on everything scores a perfect abstention rate while never actually helping anyone.
  3. Gate every release on both numbers, not on a slower quarterly check.Why: a small prompt tweak shipped between quarterly reviews is exactly when a drift gets three or four weeks to run before anyone looks at it.
  4. Set real thresholds, not a feeling.Why: a number nobody has to clear before shipping is a chart on a dashboard, not a metric anyone acts on.
  5. Leave the easy, document grounded questions tuned toward being direct.Why: hedging on a question with one clear right answer sitting in the contract costs real usefulness for no safety gained.
  6. Treat a drop in accuracy on answered questions as a different, more urgent bug than a drop in abstention rate.Why: one means the model is getting confident answers wrong, the other means it is over answering, and they need different fixes.

How to answer this, stage by stage

Nobody's grading whether you can name a metric. They're grading whether the metric you name would have caught anything before it turned into an eighty five thousand dollar problem.

1
Ground the metric in one real tool before naming any number
Say it like this
"Let's make this concrete. Caselight is an AI research tool for in house counsel. It answers questions about contracts, regulations, and precedent in plain paragraphs. Marlein Vantry owns answer quality there."
Why this works
A metric question answered in the abstract turns into a list of buzzwords. One real product makes it a design decision you can defend.
2
Say what the question is actually testing
Say it like this
"This isn't really asking me to name a metric. It's asking whether I know the difference between a model that's wrong and a model that's confidently wrong on something nobody should be confident about, and whether I'd build a number that catches the second one on purpose."
Why this works
Naming the real question up front stops you from reaching for the first metric that comes to mind, which is what most candidates do.
3
Give the direct answer, cold, before any story
Say it like this
"I'd track abstention rate on a golden set built only from questions the model shouldn't answer with confidence: split precedent, thin document coverage, facts the contract doesn't state. Weekly, on its own, never folded into overall accuracy. And I'd pair it with accuracy on whatever it does answer, because a model that hedges on everything would look perfect on abstention rate alone."
Why this works
A reader who stops here already knows exactly what you'd build. Everything after this is proof.
4
Show why the obvious metric hides the real problem
Say it like this
"Overall accuracy at Caselight sat near 92 percent for nine months, an easy, consistent number. It barely moved after a bad incident, 91.8. If that's the only number you're watching, you'd have said nothing was wrong, right up until a lawyer used a confidently wrong answer in a live negotiation."
Why this works
This answers "why not just use accuracy" before the interviewer has to ask it.
5
Name the gaming risk before anyone else does
Say it like this
"Here's the trap with abstention rate on its own. Tell the model to hedge on everything and it hits 100 percent, and now it never actually answers a question a lawyer needed answered today. So this number only means something sitting next to a second one: accuracy on what it did answer."
Why this works
Naming your own metric's failure mode before being asked for it is what makes the rest of the answer sound like judgment, not a slogan.
6
State the thresholds, not a vibe
Say it like this
"Concretely, I'd want abstention rate on the hard set above 85 percent, and accuracy on the questions it does answer above 95 percent, both measured on a rolling weekly sample, both gating every release, not just the quarterly one. Drop below either bar and the release doesn't ship."
Why this works
A calibrated bar beats "the model must always know when it doesn't know," which no probabilistic system can promise.
7
Close with what changes, and prove it with a number
Say it like this
"Same drift, new design. The abstention rate on the hard set still starts sliding around week ten. But now it's a release gate, not a background log, so the release that caused it gets blocked in week eleven, before the incident that cost eighty five thousand dollars ever happens."
Why this works
Closing on a countable, prevented cost is what turns this from a definition into a decision.

Let's learn

Abstention rate on the hard question golden set, week 1 to week 20
90% 50% week 19: the incident Wk 1 Wk 10 Wk 20
Correct abstention rate on the hard set
The number holds near 82 percent for nine weeks, then slides for ten straight weeks, down to 58 percent by week 20. Nobody watched it, because it wasn't the number that gated a release.

Marlein Vantry has run answer quality for Caselight for two years. Caselight is an AI research tool for in house counsel. It answers questions about contracts, regulations, and precedent, in plain paragraphs, no citation menu to dig through. Before it, counsel at a client company spent about three hours a day chasing down answers by hand. After it, that dropped to about forty minutes, a real two hours and twenty minutes back, every single day.

Knowledge spark: what does abstain mean here? Saying "I don't have enough to answer this for sure" instead of guessing. A doctor who orders more tests before naming a diagnosis is abstaining. A doctor who guesses anyway is not. For an AI tool, abstaining means flagging a question as unclear instead of answering it in a confident voice.

For nine months, the number Marlein read out at the Monday review barely moved. Overall accuracy: 92 percent, give or take a tenth of a point, graded monthly against 1,000 real questions by a senior counsel panel. Clients ran something like fourteen thousand questions a month through Caselight: lease clauses, non compete enforceability, a new state privacy rule, whether a boilerplate indemnity clause meant what it looked like it meant. When Caselight answered in its flat, confident voice, it was right almost every time. When it hedged, it was usually genuinely unclear too. Counsel learned fast that the confident voice meant something, and stopped double checking the confident ones. Why would they. It kept being right.

The decision that mattered Caselight also logged a second number every week from day one: correct abstention rate on a golden set of 150 questions built only to be hard, split precedent, thin coverage, facts the contract never states. That number just was not the one that gated a release. Overall accuracy was. The hard set number sat several clicks deep, opened only at the quarterly review.

In week fourteen, the team shipped a prompt change to fix a real complaint: users kept saying Caselight hedged too much, that "this may depend on jurisdiction" on a routine question was just noise. The fix worked, on paper. Complaints about unhelpful hedging fell from about 120 a week to 45 within a month. Nobody re ran the hard question review before shipping it, because that review only ran once a quarter, and the number that gated the release, overall accuracy, hadn't moved.

What that fix actually did was teach Caselight to sound sure more often, including on questions nobody should sound sure about. Correct abstention on the hard set slid from 82 percent to 58 percent over the following ten weeks. Restated the other way: the rate of confidently wrong answers on that same hard slice climbed from about 18 percent to about 42 percent.

Two numbers, the same six weeks around the incident
100% 0% Weeks 1 to 9 Weeks 16 to 20 92.1% 18% 91.8% 42%
Overall accuracyConfidently wrong on the hard set
Overall accuracy moves three tenths of a point. The rate of confidently wrong answers on the hard set more than doubles, from about 18 percent to about 42 percent, in the same stretch.
We didn't hand her a worse model. We handed her the same confident voice, aimed at a question it had no business being sure about.

The extra confidently wrong answers on the hard set, on their own, are not really the problem. It's 150 questions out of fourteen thousand a month, barely a dent in overall accuracy. The real problem is what in house counsel did next: nothing. They kept trusting every confident answer exactly the way they always had, because nobody told them the confident tone had quietly stopped meaning what it used to mean.

What I would leave alone The bulk of Caselight's traffic is document grounded: a statute's plain text, a clause that's simply sitting in the uploaded contract. Tuning those toward being direct and decisive is right. Hedging on a question with one clean answer costs real usefulness for no safety gained. Only the hard slice needed the tighter gate, not the whole product.

The lesson: overall accuracy is a blended number, built mostly from thousands of easy questions with one obvious right answer. It can hold still for months while the one number that would have warned you, correct abstention on the hard slice, quietly drifts the entire time underneath it. Logging a number and gating a release on it are two different decisions. Caselight only ever made the first one.

Now here is the same thing as a story

Read this when you want to feel why the gap between the two numbers mattered, not just know that it did.

Ask Marlein Vantry which questions Caselight should never answer with a flat yes, and she can list them without checking a doc: anything where two appellate districts disagree, anything about a state Caselight has thin coverage in, anything where the real answer depends on a fact the contract simply never states. Two years running answer quality there will do that to a person.

The habit at client companies thinned in three quiet steps nobody wrote down anywhere. First, routine lease renewals stopped getting a second look, because Caselight's confident answers on those had been right for months. Then most contract questions stopped getting one. By month seven, the rule inside most legal teams using Caselight had become simple, unspoken, and completely reasonable given what they'd seen: if it sounds sure, it's sure.

The trigger, when it came, was one wording change nobody thought was risky. In week fourteen, the team shipped a prompt update meant to cut down on a real complaint: too much hedging on routine questions. It worked. Complaints fell from about 120 a week to 45 within a month. Everyone was pleased. The quarterly hard question review, the one graded by senior counsel against genuinely unclear cases, was not due again for another eight weeks, and the number that gated the release, overall accuracy, hadn't moved.

Five weeks later, a contracts lead asked Caselight a plain question: would a non solicit clause in a departing employee's agreement survive termination, in a state where a recent appellate ruling had actually split the districts on exactly that point. Caselight answered in the same flat, confident voice it always used for a clean question. "Yes, the clause survives termination." No hedge. No mention of the split.

We didn't lose eighty five thousand dollars to a wrong answer. We lost it to a right habit nobody updated.

The company relied on that answer inside a live negotiation. Three weeks later, outside counsel on the other side of the deal caught the split, and the whole agreement had to be reopened. Eighty five thousand dollars in renegotiation costs and outside counsel fees. Three weeks of the deal team's time, gone. Then came the overcorrection: legal ops mandated a second lawyer review every single Caselight answer used in any negotiation, starting that week, no exceptions. The tool that used to hand counsel back two hours and twenty minutes a day gave almost none of it back anymore. Every answer needed a human read regardless, so the software cost and the review time both got paid, on every deal, not just the risky ones.

Marlein's first instinct was to ask whether the model itself had gotten worse. The quarterly hard question review happened to be scheduled for the following week anyway, so she pulled it forward. The senior counsel panel re graded all 150 questions in the hard set. Caselight had been getting a smaller share of them right for ten straight weeks. Not because anyone told it to. Because the same wording change that quieted the hedging complaints had also taught it to sound sure more often, including on the ones nobody should sound sure about.

Here's the part that actually explains it. That number, correct abstention on the hard set, had been logged every week since Caselight launched. It just was never the number that gated a release. Overall accuracy was, and overall accuracy barely blinked, because 150 hard questions are a rounding error next to fourteen thousand easy ones a month.

So here is the decision I would take back. We built the hard question check at launch. We decided, quietly, that it only needed senior counsel grading once a quarter, because that grading took real lawyer time and doing it every release felt like overkill. I would make correct abstention on the hard set a release gate from day one, checked weekly on a rolling sample, not necessarily senior counsel graded every single week, but graded often enough that a ten week slide gets caught inside two.

Run the same fourteen weeks again with that gate in place. Correct abstention on the hard set still starts sliding in week ten, same as before. But now someone has to clear it before week eleven's release ships. The release that caused the slide gets held. The eighty five thousand dollar negotiation never happens, because the wrong answer never went out the door.

One design trusted a number because it happened to already look healthy. The other watches the number built specifically to catch what the healthy one hides.

What I'd tell myself, back at launch: the number you gate a release on is a decision, not a default. We picked the flattering one because it was already there, cheap to compute, and it had never once caught a real problem. That last part was never proof it was the right number. It was only proof we had never tested it against something like this.

LEAD, run on Caselight's own numbers

Four letters instead of five, built for a question that asks you to name a number, not diagnose a drop.

LLink. The business outcome that actually matters, not the model's own score.
Fewer confidently wrong answers reaching a lawyer who then acts on them without double checking, since a flagged "not sure" is recoverable, counsel just digs further, and a confidently wrong answer folded into a live negotiation often isn't.
This step stops you naming a number before you've said what it's actually protecting.
EEarly signal. The thing that moves weeks before the outcome does.
Correct abstention rate on a golden set built only from questions the model shouldn't answer with confidence, checked weekly, on its own, never blended into overall accuracy. This is the answer to the actual question: overall accuracy at 92 percent told nobody anything for six straight weeks while this number slid from 82 to 58.
Accuracy at 92 percent tells you nothing. A hard set abstention rate falling from 82 to 58 tells you the whole story, three weeks before anyone screenshots the proof.
AAbuse. How this metric gets gamed, by the team or the model.
Tell the model to hedge on every single question and abstention rate hits 100 percent, while the tool answers nothing a lawyer actually needed today. The metric alone rewards a model that has stopped doing its job.
Naming your own metric's failure mode before being asked for it is what makes the rest of the answer sound like judgment instead of a slogan.
DDecision. What you'd actually do differently at each threshold.
Abstention rate on the hard set above 85 percent and accuracy on answered questions above 95 percent, both gating every release. Either one dropping below its bar blocks the ship, not just a note in a review deck.
A metric nobody acts on is a dashboard decoration. This is the part that makes it a real gate.
Hand sketched comparison titled A metric that passes while the product fails. Left panel a gauge icon labeled abstention rate, caption one hundred percent, never wrong. Right panel a person icon labeled the lawyer, caption gets no answer every time.
Push a model to hedge on everything and abstention rate hits a perfect number while the product stops answering anyone at all.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was gating abstention on how many documents the retrieval step pulled back for a question, treating a thin document count as the signal for uncertainty. It lost because a genuinely contested question can return five documents that simply disagree with each other, so a full retrieval count looks exactly like a confident, well supported answer even when the underlying law is split. The AI specific failure mode worth naming by name is a particular kind of confident hallucination: the model treating a genuine legal split as settled because its training and its retrieved documents lean one way, and stating that lean as fact. The guardrail is the hand graded hard question golden set itself, refreshed by senior counsel every quarter so it keeps testing genuinely unsettled questions instead of ones the model has since learned cold. That guardrail isn't free. Building and re grading a golden set of genuinely hard questions costs real senior lawyer hours every quarter, and gating a release on it will occasionally hold back a prompt tweak that was actually fine, a small delivery speed cost accepted only because a confidently wrong answer on this slice costs far more than a slower release cycle. And the bar isn't zero wrong guesses, no probabilistic model can promise that. It's an abstention rate above 85 percent on a rolling 100 question sample of the hard set, paired with accuracy above 95 percent on whatever it does answer, checked before every release, not once a quarter.

And if you want to be sure it really works, try it somewhere else

Same LEAD, a claims coverage tool for insurance adjusters this time, nothing about lawyers or contracts anywhere in it.

Coverwell is an AI tool that answers coverage questions for claims adjusters at a regional insurer: does this policy cover this specific loss, given the exact wording of this endorsement. Rhian Onyema leads claims quality there.

L, link. Fewer confidently wrong coverage calls that get paid out or denied and then have to be reversed, since a claim flagged for a human underwriter is recoverable, and a confidently wrong call that already went out often costs a clawback, or worse, a regulator's attention.
E, early signal. Abstention rate on a golden set of about 120 genuinely ambiguous scenarios: conflicting endorsements, disputed cause of loss wording, a claim missing its inspection report. Checked weekly, apart from overall coverage call accuracy, which sat near 95 percent the whole time and moved less than half a point.
A, abuse. Tell Coverwell to kick every ambiguous seeming claim to a human and abstention rate looks perfect, while claims that used to clear in a day now sit in a queue for a week. This number only means something next to accuracy on the calls it makes on its own.
D, decision. Abstention rate on the hard set above 80 percent, accuracy on the calls it makes alone above 96 percent, both gating every model or prompt update, not just the quarterly one.

Hand sketched quadrant titled Which number would have warned you first. X axis how early it moves before the incident, from late to early. Y axis how far it actually moves, from barely to a lot. Overall accuracy sits low on both axes. Hedging complaints sits high on movement but late on timing. Hard set abstention rate sits high on both, early and large.
Three numbers, three different shapes. Only the one in the top right corner would have rung before the incident did.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the one number that gates a release, don't spend the time justifying why overall accuracy isn't enough.
Cost: there's no senior underwriter graded golden set funded yet at a smaller shop. Don't skip the gate, hand grade a stratified sample of forty or fifty scenarios yourself, weekly, until a real one exists.
The model got better, for real: say overall coverage call accuracy improved a full point that quarter. That's not the same claim as the hard slice being calibrated. A model can improve on average while one narrow, expensive slice drifts the entire time underneath the average.

Where people run it wrong.
They watch the number that already looks good, because it's the one that was already on the dashboard.
They read a flat overall number as proof nothing needs fixing, and skip building the harder golden set entirely.
They build the abstention rate gate and forget to pair it with accuracy on the calls made alone, so a model that hedges on everything sails through looking perfectly safe.

How to use it live. Say the two numbers are usually not the same number, out loud, before naming either one: "the number that says it's fine and the number that would have warned me early are usually not the same number." That buys you a beat to pick the real one instead of naming whichever metric comes to mind first.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a question asking what metric to use?
Tap to flip
ANSWER
LEAD: find the outcome that actually matters, then the number that moves before it does, name how that number gets gamed, then say what you'd do differently at each threshold.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marlein Vantry, who leads answer quality for Caselight, an AI research tool that answers contract, regulation, and precedent questions for in house counsel.
3 · THE HABIT
What habit had quietly formed before the drift?
Tap to flip
ANSWER
In house counsel stopped double checking Caselight's confident sounding answers, because for nine good months, a confident tone reliably meant well supported. Nobody told them that stopped being true.
4 · TWO NUMBERS
What are the two numbers in this story, and which one actually moved first?
Tap to flip
ANSWER
Overall accuracy, which stayed near 92 percent, and the hard set abstention rate, which drifted from about 82 percent to 58 percent over ten weeks. Only the second one moved before the incident.
5 · THE OLD DECISION
What old decision does this answer take back?
Tap to flip
ANSWER
Caselight logged the hard set abstention rate every week from day one, but only overall accuracy gated a release. The hard set number sat unread except at the quarterly review.
6 · THE NUMBER
Fill in the blank: the hard set abstention rate held near ___ percent for months, then drifted to ___ percent by week 20, while overall accuracy barely moved.
Tap to flip
ANSWER
82 percent, then 58 percent. Overall accuracy only moved from about 92.1 to 91.8 percent over the same stretch, which is why nobody caught it by watching that number.
7 · THE REPLAY
Same drift, new design, what changes?
Tap to flip
ANSWER
The hard set abstention rate becomes the release gate, not a background log. The same slide that starts in week ten blocks a release in week eleven, before the incident that cost eighty five thousand dollars ever happens.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the parallel?
Tap to flip
ANSWER
Coverwell, a claims coverage tool for insurance adjusters, run by Rhian Onyema. Same LEAD steps: an abstention rate on a golden set of genuinely ambiguous coverage scenarios, paired with accuracy on the calls it makes alone.

Check yourself Score: 0 / 0

True or false
1. True or false: because overall accuracy at Caselight stayed near 92 percent through the whole drift, the slide in the hard set abstention rate wasn't a real problem.
  • True
  • False
Show hint
Check the grouped bar chart. Did the two numbers move by the same amount?
Show answer
False. The hard set is a tiny slice of total volume, so overall accuracy barely moved even while the confidently wrong rate on that slice climbed from about 18 percent to 42 percent, and that slice is exactly where a wrong answer costs real money.
Multiple choice
2. Why would a model tuned to hedge on every single question score perfectly on abstention rate while still being a bad answer to the interview question?
  • A. Abstention rate only counts questions from the easy general golden set.
  • B. Abstention rate alone can't tell an honestly uncertain model from one that never commits to anything, so it needs to be paired with accuracy on what it does answer.
  • C. Abstention rate is only measured once a quarter, so gaming it wouldn't matter.
  • D. Lawyers would immediately notice a hedging model and stop using it, so the metric doesn't matter either way.
Show hint
Look at the A step in the LEAD recap, and the hand sketch of the metric passing while the lawyer walks away with no answer.
Show answer
B. A model that abstains on everything hits 100 percent on this metric alone while doing zero of its job. Pairing it with accuracy on answered questions is what stops that gaming path.
Fill in the blank
3. The hard set abstention rate held near ___ percent for months, then drifted down to ___ percent by week 20, while overall accuracy moved only from about 92.1 to ___ percent.
Show hint
Look at the line chart right under "Let's learn."
Show answer
82 percent, 58 percent, 91.8 percent. The gap between how much each number moved is the whole argument for why abstention rate has to be its own gate.
Multiple choice, name the rejected alternative
4. What alternative did Caselight's team reject when deciding how to gate abstention, and why did it lose?
  • A. Gating abstention on how many documents the retrieval step returned. Rejected because a genuinely contested question can return several documents that simply disagree with each other, so a full document count doesn't catch a real split.
  • B. Replacing the model with a fixed rules engine that never generates free text.
  • C. Having every single answer read by a lawyer before it reaches a client, forever.
  • D. Removing the hard question golden set entirely and trusting overall accuracy alone.
Show hint
Look at the closing paragraph of the LEAD recap section for the named rejected alternative.
Show answer
A. Document count looks like confidence even when the documents disagree, which is exactly the case where the model should be hedging instead of answering.
Short answer, apply it yourself
5. Pick an AI tool you use yourself. Name a question it should probably decline to answer with confidence, and say what you'd check to know if it's getting worse at knowing that.
Show hint
Think of a tool giving advice in a field where being confidently wrong costs more than saying "not sure."
Show answer
Model answer: A tax prep AI answering whether a specific home office deduction is safe to claim. I'd build a small set of genuinely gray cases, mixed personal and business use, a home audited before, and check every month how often it correctly says "this depends on facts a form can't capture, talk to a preparer" instead of just answering yes.
Short answer, reason about the number
6. If Caselight's accuracy on answered questions dropped to 80 percent while the hard set abstention rate stayed healthy at 85 percent, would that be the same kind of problem as the one in this story? Why or why not?
Show hint
Compare what each number is actually measuring, not just whether either one looks bad.
Show answer
Model answer: No. A healthy abstention rate means the model is still correctly declining the hard questions. A falling accuracy on answered questions means it's getting confident answers wrong even on the ones it chose to answer, which is a different bug, closer to a data or hallucination problem, and it needs a different fix, not a calibration fix.
Before you close the answer
Why this works
Tests whether you'd reach for the number that already looks healthy, or go build the harder, adversarial one nobody's incentivized to build. Most candidates say "track hallucination rate" and stop there, without naming a real metric or its trap.
Follow-up traps
"Isn't building a golden set of genuinely ambiguous questions expensive? You're asking for real lawyer time." Response: yes, and that expense is exactly why it gets skipped, which is the whole failure in this story. Start with a small stratified sample, even 50 questions, and grow it. A small honest golden set beats a large easy one every time.

"What if the model is right to sound confident more often than your golden set assumes?" Response: that's why the golden set is graded by senior counsel against genuinely split precedent and thin coverage, not a guess, and it gets re reviewed. A model earning real confidence on a question the panel called ambiguous is a reason to update the golden set, not evidence the metric was wrong.
If pressed
The actual thresholds used at Caselight: abstention rate on the hard set above 85 percent and accuracy on answered questions above 95 percent, both on a rolling weekly sample of at least 100 questions, both gating every release, not just the quarterly one.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more