CaseAdvancedDesigning for Uncertainty & Trust / Designing for failure and graceful degradation / #12

How do you design for the case where the model is confidently wrong?

LEAD the number that would have rung weeks before the settlement did

The artifact here is a single line of text: a risk label on a contract clause. Claritas is a contract review tool that reads vendor agreements and labels each clause's risk. Petra Aldrich is a contracts manager who trusted one of those labels, on one particular Tuesday, more than it deserved.

The direct answer
Stop showing one overall confidence number and start showing each clause category's actual track record: how often a human reviewer has overturned this exact kind of label before. Track that overturn rate by category as your real metric, not overall accuracy, and flag any category where it's climbing, even while the aggregate score still looks perfect.
Do this, in order
  1. Replace the single confidence score with a per-category track record.Why: an aggregate number that looks great can still be hiding one narrow category that's confidently wrong most of the time.
  2. Track the human overturn rate by clause category, not overall accuracy.Why: overall accuracy stays flat for months while one category quietly drifts, and that drift is the actual leading signal.
  3. Flag any category whose overturn rate is climbing, before it crosses a real incident.Why: waiting for a dispute to notice a bad category means finding out the expensive way, every time.
  4. Watch for reviewers rubber-stamping agreement to save time, and audit a real sample.Why: an overturn rate can look calm simply because nobody is actually checking anymore, which is its own kind of confidently wrong.
  5. Require a second look on any category above the flagged threshold, regardless of the model's stated confidence.Why: the model's own confidence is exactly the number that's already wrong for that category.
  6. Leave high-volume, well-calibrated categories, like standard payment terms, exactly as fast as they are today.Why: this fix targets the rare categories quietly going wrong, not the ones already working.

How to answer this, stage by stage

Nobody is grading whether you know that models can be wrong. They're grading whether you can name the number that would have told you before the wrongness cost anyone money.

Stage 1
Ground the artifact in one real label
Say it like this
"I'll use Claritas, a contract review tool, and one label it gave a contracts manager named Petra: 'Low risk, standard terms' on an indemnification clause that wasn't standard at all."
Why this works
Turns an abstract question about confident wrongness into one concrete, arguable label.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link to the real business outcome, early signal that moves first, abuse, how the metric gets gamed, and decision, what you'd actually do at each threshold."
Why this works
Shows you're answering a metric question with a metric method, not a design wish list.
Stage 3
Name the early signal
Say it like this
"The metric that would have rung first isn't overall accuracy, it's the human overturn rate for this one clause category, which was already climbing for months while the aggregate score sat at 96 percent."
Why this works
This is LEAD's whole point: naming the number that moves before the damage, not after.
Stage 4
Give the one design decision
Say it like this
"Show each category's own track record on the label, not one confidence number for everything, and flag any category whose overturn rate crosses a threshold, whatever the model's own confidence claims."
Why this works
This is the direct answer, said as a concrete interface and metric decision, not a general call for caution.
Stage 5
Name how the metric gets gamed
Say it like this
"A calm overturn rate can just mean reviewers stopped pushing back and started agreeing to move faster. I'd audit a real random sample against outside counsel, not just trust the in-app disagreement count."
Why this works
Shows you'd distrust your own metric enough to check it, which is what a metric question is actually testing.
Stage 6
Close on the one line
Say it like this
"A model that's confidently wrong isn't a calibration bug you patch once. It's a category you haven't found yet. Design to find it early, not to make the aggregate score look nicer."
Why this works
Restates the direct answer in one breath, and separates "confident" from "correct" cleanly.

Let's learn

The artifact here is a small green label reading "Low risk, standard terms." What should a reader do with that label, and what would tell you it's about to be wrong?

Before Claritas, Petra's team read every vendor contract clause by hand, flagging anything unusual for outside counsel, a slow process that took a paralegal roughly two hours per contract but rarely missed a genuinely dangerous clause.

Hand sketched flow diagram titled How a clause gets its risk label. Five boxes: Clause parsed, Risk labeled, Confidence shown highlighted, Human review, Filed or flagged.
The third box is where the whole problem lives. What gets shown there decides whether anyone looks twice.

Now Claritas reads a full contract and labels every clause's risk in under a minute, with an overall accuracy rate the team monitors closely: 96 percent agreement with human reviewers, month over month, steady as a rock.

Human overturn rate for indemnification-clause risk labels, month over month
40% 20 0 20% concern line Month 1, 4% Month 5, dispute
The line crossed the concern threshold in month four, a full month before the dispute. The overall accuracy dashboard never once flinched.

Here's the turn: the model's overall accuracy was never the problem. One narrow category, indemnification clauses from newer, smaller vendors, was being labeled "Low risk" wrong more than a third of the time, and that number was climbing for months while the number everyone was actually watching stayed perfect.

Overall label agreement versus one category's real agreement, same month
100% 50 0 Overall, 96% Indemnification, 61%
A 35-point gap, sitting inside a headline number that looked fine every single month.

At its worst, a contracts manager signs off on a clause the tool called routine, and finds out only after a dispute how far from routine it really was.

Hand sketched two-panel metaphor scene titled Two clocks. Left, a blue gauge icon labeled Aggregate score, caption looks fine for months. Right, a red gauge icon labeled Category signal, caption rings weeks earlier.
Both clocks were ticking the whole time. Only one of them was ever going to ring before the dispute did.
The decision I would take back Claritas's team chose to surface one overall confidence score per contract, since a single number was simpler to explain to customers than a breakdown by clause category. That made sense when most categories performed similarly. It stopped making sense the moment one narrow, newer category started drifting badly while every other category kept the average looking healthy.

What I would leave alone: high-volume, well-understood categories, like standard payment or confidentiality terms, don't need this scrutiny. Their overturn rate has stayed low and flat for years, and adding friction there protects against nothing.

The lesson: confidence and correctness are two different numbers, and a model can be perfectly sure of itself in exactly the category where it's most wrong.

Now here is the same thing as a story

For eight months, Claritas had been reliable enough that a routine vendor contract went from Petra's desk to signed in under a day.

Petra Aldrich manages vendor contracts for a mid-size manufacturer, and she'd learned to trust the "Low risk, standard terms" label the way she'd trust a paralegal she'd worked with for years, because for eight straight months, it had never once steered her wrong.

Knowledge spark: why would a model be confidently wrong in just one category? A model's overall accuracy is an average across everything it sees, and most categories are common enough that it's seen thousands of examples. A newer, rarer clause type, like an indemnification clause from a smaller vendor's own template, might appear in a tiny fraction of its training data. The model can still sound just as sure about it, because nothing in its confidence score tracks how rare or unfamiliar a specific case actually was.

A vendor contract came in from a smaller, newer supplier, using its own indemnification language rather than the manufacturer's standard template. Claritas read it and labeled the clause "Low risk, standard terms," the same phrasing it used on hundreds of contracts before. The clause was, in fact, an uncapped liability clause, one with no ceiling on how much the vendor could be forced to pay if something went wrong, dressed in ordinary-sounding legal phrasing.

Hand sketched timeline titled Petra's contract, month by month. Five milestones: Contract signed labeled Low risk, Dispute surfaces months later 180,000 dollars highlighted, Audit finds the pattern same clause other contracts, Category flag ships track record shown, Replay caught before signature.
Five months from a signed label to a real number on a settlement, and a pattern nobody had gone looking for yet.

Petra signed off on schedule. Five months later, a dispute with that same vendor surfaced, and because the liability clause had no cap, the manufacturer ended up covering a 180,000 dollar settlement that a properly flagged, capped clause would have limited many times over.

Claritas never lied about being confident. It just never said that confidence, for this one category, had stopped meaning anything.

Legal ops ran a routine quarterly audit two months after the dispute, pulling a random sample of "Low risk" labeled contracts for a second look. Digging into the pattern, they found four more newer-vendor contracts carrying the same uncapped liability language, all quietly labeled "Low risk, standard terms," none of them yet in a dispute, but all of them one bad vendor relationship away from becoming one.

Hand sketched two-panel metaphor scene titled How it gets gamed. Left, a green document icon labeled On paper, caption overturn rate looks low. Right, an orange person icon labeled In reality, caption reviewer just agrees, to save time.
A calm overturn number and a genuinely calibrated model can look identical from a dashboard. Only an outside audit tells them apart.

With the redesign, that same contract's label now reads: "Low risk (indemnification, newer vendor: flagged for extra review, human overturn rate 35% this category)" instead of a flat "Low risk, standard terms." Run the same month forward: Petra sees the flag, routes it to outside counsel as a precaution, and counsel catches the missing liability cap before anyone signs anything.

Hand sketched labeled parts diagram titled What the redesigned label shows. A document icon at the center labeled Risk Label, with four callouts around it: category track record, not just the overall score, extra scrutiny flag, this month's overturn rate.
None of these four existed on the old label. All four came out of one audit two months too late for the first contract.

The old label asked Petra to trust a single number that had quietly stopped applying to this one category. The new one tells her exactly which category still needs her eyes.

I approved the single confidence score because it was simpler to build and simpler to explain to customers. It took a 180,000 dollar settlement, and four more contracts an audit found only by accident, to see that simple had been hiding exactly the category where it mattered most.

LEAD, in one screenNot a lecture on calibration. LEAD is what tells you which number would have rung first.

L
Link. The business outcome that actually matters.
Catching a dangerous, uncapped clause before signature, not the model's own accuracy score.
Grounds the metric in real cost, a settlement, not an abstract quality number.
E
Early signal. The number that moves first.
The human overturn rate for indemnification clauses from newer vendors, climbing from 4 percent to 35 percent over five months while overall accuracy stayed at 96 percent.
This is the hardest step and the whole answer: naming the leading number the aggregate score was hiding.
A
Abuse. How this metric gets gamed.
A calm overturn rate can simply mean reviewers started rubber-stamping agreement to move faster, not that the model actually got more accurate.
Forces a real outside audit instead of trusting the in-app disagreement count on its own.
D
Decision. What you'd actually do at each threshold.
Any category crossing a 20 percent overturn rate gets a mandatory second look and a visible scrutiny flag, no matter what the model's own confidence claims.
Turns the metric into an action, not a number nobody actually uses.
Hand sketched quadrant titled Which clause categories need a scrutiny flag. Axes: how often this clause appears rare to common, human overturn rate low to high. Payment terms and confidentiality sit common and low overturn. Termination sits in the middle. Indemnification clauses from new vendors sit rare and high overturn.
Only the rare, high-overturn corner needs a scrutiny flag. Everywhere else, the aggregate score is telling the truth.

The recap, one line per letter: link is catching a dangerous clause before signature, early signal is the category-level overturn rate climbing while the aggregate held steady, abuse is reviewers quietly rubber-stamping to save time, and decision is a mandatory second look once a category crosses a real threshold.

And if you want to be sure it really works, try it somewhere elseSame four letters, a consumer allergen-checking app instead of a contract tool. The confidently wrong label here reads "Safe for you," not "Low risk."

PlateSafe is an app that scans a restaurant menu and tells a user with food allergies which dishes are safe. Tomasz Adebayo has a tree-nut allergy and relies on it to eat out. Mapped onto LEAD: link is avoiding a real allergic reaction, not the app's overall accuracy; early signal is the human-reported correction rate for dishes containing pesto or marzipan-adjacent ingredients, a rare cross-contamination source that climbed for weeks before Tomasz had a reaction, while PlateSafe's overall "Safe" label accuracy sat at 98 percent the entire time.

The abuse risk here is structurally the same: PlateSafe could hit a low correction rate simply because most users never report a near-miss at all, only actual reactions, so a rising rate of quiet, unreported near-misses would stay invisible unless the team went looking for it directly, the same audit discipline that caught Claritas's pattern. The decision: flag any ingredient category whose reported correction rate crosses a threshold with a visible "less common ingredient, double-check with the kitchen" note, instead of a flat green "Safe for you."

Hand sketched two-panel metaphor scene titled Two clocks, reused here for PlateSafe. Left, a blue gauge icon labeled Aggregate score, caption looks fine for months. Right, a red gauge icon labeled Category signal, caption rings weeks earlier.
Same two clocks, a different kind of danger behind the slower one.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "track the overturn rate by category, not overall accuracy, and flag whichever category is climbing," and stop.
Cost: there's no time to build category-level tracking for every clause type this quarter. Say so honestly, and start with whichever category carries the highest dollar risk if it's wrong, here, liability and indemnification.
The model gets better, for real: if Claritas's overall accuracy genuinely improves next quarter, that's still not a reason to stop tracking by category, a hidden bad category can persist quietly inside a rising average exactly as easily as inside a flat one.

Where people run it wrong.
They watch one aggregate accuracy number and call the model healthy, without ever slicing it by the kind of case involved.
They treat a low disagreement rate as proof of good calibration, without checking whether anyone is still actually disagreeing.
They wait for an incident to reveal a bad category, instead of auditing quietly, on a schedule, before anything goes wrong.

How to use it live. When someone asks how to design for a confidently wrong model, don't reach for "add a confidence score." Ask which number would have moved first, weeks before anyone noticed, and design around watching that one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a metric or "how do you design for X" measurement question?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Find the number that moves weeks before the outcome does.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Petra Aldrich, a contracts manager who trusted a "Low risk, standard terms" label on an uncapped liability clause after eight months of Claritas being reliable.
3 · THE EARLY SIGNAL
What number would have moved first?
Tap to flip
ANSWER
The human overturn rate for indemnification clauses from newer vendors, which climbed from 4 to 35 percent over five months while overall accuracy stayed at 96 percent.
4 · HOW IT GETS GAMED
How could the overturn rate look healthy without the model actually improving?
Tap to flip
ANSWER
Reviewers could simply agree with the model more often to save time, making the disagreement count drop without any real gain in calibration.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Showing one overall confidence score per contract instead of a per-category breakdown, reasonable while categories performed similarly, wrong once one category started drifting badly.
6 · THE NUMBER
Fill in the blank: the dispute cost the manufacturer ___ dollars, thanks to an uncapped liability clause.
Tap to flip
ANSWER
180,000 dollars, on a clause the tool had confidently labeled "Low risk, standard terms."
7 · THE REPLAY
Same contract, redesigned label. What changes?
Tap to flip
ANSWER
The label shows a scrutiny flag with the 35 percent category overturn rate, Petra routes it to outside counsel, and the missing liability cap is caught before signature.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
PlateSafe, a consumer allergen-checking app. The early signal is the correction rate for dishes with pesto or marzipan-adjacent ingredients, climbing before an actual reaction happened.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: overall label agreement stayed at 96 percent, but agreement on indemnification clauses from newer vendors was only ___ percent that same month.
Show hint
Look at the bar chart comparing overall versus category accuracy.
Show answer
61 percent. A 35-point gap the aggregate number never revealed.
Multiple choice
2. Why is overall accuracy the wrong metric to watch for "confidently wrong" cases, according to this answer?
  • A. Overall accuracy is too expensive to calculate regularly.
  • B. A rare, narrow category can drift badly while the aggregate stays flat, hiding the exact place trouble is building.
  • C. Overall accuracy only applies to large enterprise customers.
  • D. It requires the model to be retrained every month.
Show hint
Look at the Early Signal step.
Show answer
B. An average across everything can look perfectly healthy while one specific slice is confidently wrong most of the time.
True or false
3. True or false: this answer recommends adding the same extra scrutiny to every clause category, including high-volume ones like standard payment terms.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Well-understood, high-volume categories with a long track record of low overturn rates don't need the extra flag.
Short answer, apply it yourself
4. Think of a tool that gave you one overall score or rating that turned out to hide a much worse number for your specific situation. What was the gap?
Show hint
Ask whether an average rating or score ever masked a narrower, worse-performing case that applied to you specifically.
Show answer
Model answer: Most people can recall an average rating that didn't reflect their specific case, the same gap between Claritas's 96 percent and its 61 percent category.
Short answer, why no middle setting
5. Why wouldn't simply "checking the confidence score more carefully" have caught this problem?
Show hint
Look at the knowledge spark about confidence versus correctness.
Show answer
Model answer: The model's confidence score never tracked how rare or unfamiliar a case was, so it sounded just as sure about the rare, wrong category as about the common, correct ones. There was no confidence signal to catch more carefully.
Short answer, the number question
6. If the concern threshold were set at 40 percent overturn instead of 20 percent, would the indemnification category still have been caught before the dispute? Why or why not?
Show hint
Look at the line chart's monthly values.
Show answer
Model answer: No. The overturn rate only reached 35 percent by the month of the dispute itself, so a 40 percent threshold would have missed it entirely, which is exactly why the threshold has to be set below the danger line, not at it.
Before you close the answer
Why this works
Tests whether you can separate a model's confidence from its actual correctness, and whether you'd design around a leading, category-level signal instead of an aggregate score that looks fine until it doesn't.
Follow-up traps
"Won't showing a track-record number on every label just overwhelm the user with statistics?" Response: no, most categories stay quietly in the background at their normal, low overturn rate. The flag only appears when a category actually crosses the concern threshold, which should be rare.

"What if a category is new and simply doesn't have enough history yet to have a track record?" Response: a category with too little history for a reliable overturn rate should be flagged for exactly that reason, insufficient track record, rather than shown a default "Low risk" label it hasn't earned yet.
If pressed
Claritas's redesigned tracking runs the category-level overturn rate as a rolling 90-day window rather than an all-time average, so a category that was fine for years but has started drifting in the last quarter gets flagged well before an all-time average would ever notice the shift.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more