ConceptIntermediateDesigning for Uncertainty & Trust / Designing for failure and graceful degradation / #21

What failure telemetry does design need in order to improve the experience?

LEAD the number that would have rung weeks before the state audit did

The City of Alder Falls uses an AI tool that pre-screens building permit applications, flagging each one as likely compliant or needing a closer look. Wilhelmina Otieno is the permits clerk who reviews the flagged ones, and who first noticed something felt off almost a year before anyone could prove it.

The direct answer
Track the override rate, how often a person changes the AI's recommendation, split by category, over time. A rate that used to sit around 12 percent and quietly falls toward zero is design's earliest warning that people have stopped checking, weeks before any wrongly approved case is ever discovered downstream.
Do this, in order
  1. Track override rate on the AI's recommendation, by category, as a trend over time.Why: a falling override rate is the earliest sign that checking has quietly stopped, long before a wrong decision surfaces anywhere else.
  2. Track whether the reviewer actually opens the AI's stated reasoning before deciding.Why: an override rate can look fine even when nobody's really reading the reasoning behind it, so open rate catches the gap the override number alone would miss.
  3. Track confidence calibration, whether the AI's stated certainty still matches how often it's actually right.Why: a model can drift quietly, staying just as confident-sounding while its real accuracy slips underneath that confidence.
  4. Pair the override rate with an independent, sampled ground-truth audit, not just the reviewer's own record.Why: a low override rate could mean genuine trust, or it could mean reviewers rubber-stamping to hit a throughput target. Only an outside sample tells the two apart.
  5. Set a floor: if override rate on any category falls below roughly a quarter of its historical average, trigger a mandatory review.Why: a number nobody acts on is decoration. This turns the leading signal into an actual decision.
  6. Leave alone any category where the override rate has always been low and audits confirm it should be.Why: some application types really are simple and safe to approve quickly. Not every low override rate is a red flag.

How to answer this, stage by stage

Nobody is grading whether you can name a metric. They're grading whether the metric you name would have rung before the actual problem did.

Stage 1
Anchor it in one review queue
Say it like this
"I'll ground this in a real case, a city permits office using an AI tool that flags applications as likely compliant or needing a closer look."
Why this works
Keeps the answer from turning into a list of generic dashboard metrics.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the outcome that actually matters. Early signal, what moves before that outcome does. Abuse, how the metric could be gamed. Decision, what I'd actually do at each level."
Why this works
Signals a method for finding a leading indicator, not a guess at a dashboard.
Stage 3
Name the outcome that actually matters
Say it like this
"The real outcome isn't approvals per week, it's correct approvals, permits that don't later get flagged as non-compliant on a state audit."
Why this works
Separates the business outcome from the vanity metric that's easiest to count.
Stage 4
Find the signal that moves first
Say it like this
"Override rate, how often a reviewer changes the AI's call, on that same category. It used to sit around 12 percent and fell under 1 percent months before anyone found a problem."
Why this works
This is the direct answer to the question, and LEAD's hardest step.
Stage 5
Say how that signal could be gamed
Say it like this
"A falling override rate could mean real trust, or it could mean reviewers rubber-stamping to hit a quota. I'd pair it with a small, independent, sampled audit to tell the two apart."
Why this works
Shows the metric isn't naive, it accounts for how people actually respond to being measured.
Stage 6
Say what you'd do at each threshold
Say it like this
"If a category's override rate falls below roughly a quarter of its historical level, that triggers a mandatory sample review that week, not a note for next quarter's planning."
Why this works
Turns a dashboard number into an actual decision, not decoration nobody acts on.
Stage 7
Prove it with the six months nobody watched
Say it like this
"The override rate crossed the warning floor in month four. The state audit that actually found the problem didn't happen until month seven. That's three months this signal had, sitting unused."
Why this works
Makes the "leading" claim checkable instead of asserted.
Stage 8
Close on the one line
Say it like this
"Design doesn't need to know a decision is wrong. It needs to know the moment people stopped checking whether it was."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Picture a review queue with three hundred permit applications a week, and only one number anybody upstream watches: how many got approved on time.

Alder Falls' permit-screening tool reads a building permit application and flags it likely compliant, ready for a fast approval, or needing a closer human look for something like a zoning setback or a load-bearing change. Before it, every application waited in a queue for a person to check it by hand, a slower process that never once skipped a check.

Hand sketched flow diagram titled How a permit gets reviewed. Four boxes: AI flags file, Reviewer sees score, Opens reasoning highlighted, Reviewer decides.
The whole design depends on that third box actually happening, every single time.

Now most likely-compliant applications move through in under a day. Wilhelmina still reviews every flagged application, and can override the AI's call whenever the reasoning doesn't hold up.

Here's the turn: the tool wasn't the problem. The problem is that nobody was watching whether Wilhelmina's checking was actually still happening, only whether applications kept moving. A review that opens a reasoning panel and clicks approve without reading it looks identical, on a throughput dashboard, to a review that read every word.

Override rate on AI likely-compliant recommendations, by month
15% 7 0 3%, warning floor Month 1, 12% Month 4, crosses floor Month 6
Nobody was tracking this line. If they had been, the warning would have fired three months before the audit did.

At its worst, a whole category of applications gets waved through on autopilot for months, and the first anyone learns of it is a state auditor's letter.

Hand sketched comparison titled Two clocks, two speeds. Left, a blue gauge icon labeled Leading Signal, caption rings in month 4. Right, a red gauge icon labeled Lagging Signal, caption rings in month 7.
Same underlying problem. One clock had a three month head start nobody used.
The decision I would take back Alder Falls never built a per-category override-rate trend into the review dashboard, only a running tally of applications processed. That made sense at launch, when the point was proving the tool could handle volume at all. It stopped making sense once the tool had been running long enough for a reviewer's habits to quietly change underneath a healthy-looking throughput number.

What I would leave alone: simple applications, like a fence permit with no zoning question at all, have always had a low override rate and audits confirm that's correct. A low number there isn't a warning sign, it's just an easy category.

The lesson: a number that only counts what happened can't tell you when people stopped actually checking it.

Now here is the same thing as a story

The short version above is what you'd say defending this telemetry plan to Alder Falls' city council. Read this one for how the number actually crept down.

Wilhelmina Otieno has reviewed building permits for the city for eleven years, and she used to be the person who caught the setback variance nobody else noticed, the kind of thing that saves a homeowner from tearing out a wall two years later.

Knowledge spark: why would an override rate ever drift down on its own? When a model's reasoning sounds consistently confident, and it's right often enough early on, a reviewer's habit shifts from reading closely to skimming, then to trusting the label outright. Nobody decides to stop checking. It happens a little at a time, and a throughput dashboard has no way to see the difference.

In her first months with the tool, Wilhelmina overrode about one in eight of its likely-compliant recommendations, usually catching a zoning detail the model's summary had glossed over. Over the following months, as the model's write-ups got more confident and specific-sounding, she found herself overriding less, then rarely, then almost never.

Hand sketched metaphor scene titled How the metric gets gamed. Left, a green gauge icon labeled Override Rate, caption looks perfectly fine. Right, a red person icon labeled Real Checking, caption quietly stopped.
The rate looked healthy from a distance. What it was actually measuring had already changed.

A colleague reviewing a batch alongside her one afternoon remarked, half-joking, "you haven't overridden one of these in weeks, has it gotten that good?" Wilhelmina laughed it off. She hadn't timed how long she'd actually spent reading each one.

Hand sketched timeline titled Six months to the audit. Five milestones: Tool ships month 1, Override 12 percent month 1, Rate drifting month 3, Falls below 3 percent month 4 highlighted, State audit month 7.
Three months sat between the number crossing its own warning line and anyone outside the queue finding out why.

Seven months in, a routine state compliance audit sampled twenty recent likely-compliant approvals and found four with a genuine setback issue that should have triggered a closer look, the exact kind of case Wilhelmina used to catch without a second thought.

The permits weren't wrong because the model got worse. They were wrong because nobody was watching the one number that would have shown, months earlier, that checking had quietly stopped.

With the redesigned dashboard, a per-category override-rate trend sits next to throughput, not instead of it, and any category falling below a quarter of its historical rate triggers a mandatory sampled review that same week. Run the same six months forward: the fourth-month warning triggers a review, catches the drift, and a short retraining session on what the model's confident write-ups can still miss brings the override rate back to a level the sampled audit confirms is genuinely earned.

The old dashboard asked "is work getting done." The new one asks "is the checking still real."

I built the throughput dashboard first because it was the easiest thing to prove the tool was working. It took a state auditor's letter to see that "fast" and "still being checked" were never the same claim.

LEAD, in one screenNot a lecture on picking a north star. LEAD is what tells you which number rings first.

L
Link. The outcome that actually matters.
Correct approvals, permits that hold up later, not raw approvals-per-week.
Separates the real business outcome from the number that's easiest to count.
E
Early signal. What moves first.
Override rate on the AI's likely-compliant recommendations, by category, tracked over time. It fell below a warning floor three months before the state audit found anything.
The hardest step, and the direct answer to the question.
A
Abuse. How the metric could be gamed.
A low override rate can mean real trust, or reviewers rubber-stamping to hit a throughput target. An independent, sampled audit is the only way to tell them apart.
Shows the metric was chosen with an eye on how people respond to being measured.
D
Decision. What you'd actually do.
Below a quarter of a category's historical override rate, trigger a mandatory sampled review that same week, not a note in a quarterly report.
Turns the number into an action, not a dashboard decoration.
Hand sketched icon list titled What failure telemetry needs. Four items: a gauge icon labeled Override rate by category, a scale icon labeled Confidence vs actual match, a document icon labeled Reasoning panel opened, a question mark box icon labeled Downstream audit result.
Four numbers. Only one of them would have rung early on its own, which is why override rate leads the list.

The recap, one line per letter: link is correct, durable approvals rather than raw throughput, early signal is override rate falling below its historical level, abuse is a rubber-stamped rate looking identical to a genuinely earned one without an outside check, and decision is a mandatory sampled review the moment the floor is crossed.

Hand sketched labeled parts diagram titled What the failure dashboard needs. Center gauge icon labeled Failure Dashboard, with four callouts: override rate trend, calibration drift, reasoning open rate, audit disagreement.
None of these four require reading a single permit application yourself. They just require watching the checking, not just the output.

And if you want to be sure it really works, try it somewhere elseSame four letters, a public library's book-tagging assistant instead of a permit office. A completely different kind of override, the same quiet drift.

Larkspur Public Library Network uses an AI assistant that suggests subject tags and shelf categories for newly catalogued books, so librarians don't have to tag every arrival from scratch. Beatrix Solheim catalogues new arrivals most mornings and used to correct the AI's suggested tags on roughly one book in six.

Mapped onto LEAD: link is patrons actually finding the right book under the right subject heading, not just how many books get catalogued per day; early signal is Beatrix's tag-override rate, which had quietly fallen from one in six to nearly zero as the AI's suggestions started sounding more specific and confident, whether or not the underlying tagging model had actually improved that much; abuse is that a falling override rate could mean the model genuinely got better, or that Beatrix, under a cataloguing backlog, started accepting suggestions faster than she was reading them.

Hand sketched flow diagram titled Larkspur's book-tagging chain. Four boxes: AI tags book, Librarian sees tags, Override tracked highlighted, Drift caught early.
Swap "permit" for "book," and the same missing number shows up in a completely different building.
Books later found under the wrong subject heading, before and after tracking override rate
40 20 0 34 Before tracking 4 After tracking
Same quarterly sample size, both times. Watching the override rate is what closed most of the gap.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "track override rate over time, by category, since a falling one means checking has quietly stopped, weeks before a wrong decision ever surfaces," and stop.
Cost: there's no budget for a full independent audit team this year. Say so honestly, and start with a small, cheap monthly sample on the category with the steepest override-rate drop, since that's where the risk is concentrated.
The model gets better, for real: if the underlying tagging or screening model genuinely improves, override rate should still fall, but slowly and evenly, not in the sudden, sustained way a real behavior change looks like. The shape of the drop, not just its direction, tells you which one happened.

Where people run it wrong.
They build a dashboard that measures throughput and call it done, since throughput is the easiest number to show a stakeholder.
They treat a low override rate as unambiguous good news, without ever checking whether it reflects trust or fatigue.
They wait for a downstream audit to tell them something's wrong, when the whole point of a leading indicator is to ring before that.

How to use it live. When someone asks what failure telemetry design needs, ask one question first: what number would have moved weeks before the problem became visible to anyone outside the room? Build the dashboard around that number, not the one that's easiest to screenshot.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "what telemetry does design need" question?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Find the number that would have moved before the real outcome did.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Wilhelmina Otieno, a permits clerk for the City of Alder Falls who has reviewed building applications for eleven years.
3 · THE LINK
What's the outcome that actually matters here?
Tap to flip
ANSWER
Correct, durable approvals, permits that don't later get flagged as non-compliant, not raw applications processed per week.
4 · THE EARLY SIGNAL
What number moved first, and how far ahead of the real problem?
Tap to flip
ANSWER
Override rate, which crossed its warning floor in month four, three months before a state audit found the actual non-compliant permits in month seven.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never building a per-category override-rate trend into the dashboard, only tracking raw throughput, which made sense at launch and stopped making sense once habits had time to quietly shift.
6 · THE NUMBER
Fill in the blank: the override rate started at 12 percent and fell under ___ percent by month six.
Tap to flip
ANSWER
1 percent. A twelve-fold drop that the throughput dashboard never showed at all.
7 · THE REPLAY
Same six months, redesigned dashboard. What changes?
Tap to flip
ANSWER
The month-four warning triggers a mandatory sampled review, catches the drift, and a short retraining brings the override rate back to a level the audit confirms is genuinely earned.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Larkspur Public Library Network's book-tagging assistant. The early signal is the librarian's tag-override rate falling from one in six toward zero.

Check yourself Score: 0 / 0

Multiple choice
1. Why is override rate a better early-warning signal here than the model's overall approval accuracy?
  • A. Override rate is easier for a city council to understand.
  • B. Override rate reflects whether a human is still genuinely checking, which starts drifting weeks before a wrong approval is ever discovered.
  • C. Approval accuracy can't be measured at all in this system.
  • D. Override rate is required by state law.
Show hint
Look at the Early Signal step.
Show answer
B. A wrong approval only becomes visible once something downstream catches it. The checking behavior that would have caught it drifts away much earlier, and override rate is what shows that drift.
True or false
2. True or false: a falling override rate always means the AI model has genuinely gotten more accurate.
  • True
  • False
Show hint
Look at the Abuse step.
Show answer
False. A falling override rate can also mean reviewers are checking less closely. Only an independent, sampled audit can tell the two apart.
Fill in the blank
3. Fill in the blank: the state audit sampled twenty recent approvals and found ___ with a genuine setback issue that should have triggered a closer look.
Show hint
Look at the story's audit paragraph.
Show answer
Four. A fifth of the sample, exactly the kind of case Wilhelmina used to catch without a second thought before the override rate drifted down.
Short answer, apply it yourself
4. Think of a product where you started clicking "accept" or "approve" faster over time. What number would have caught that shift before you noticed it yourself?
Show hint
Think about auto-suggestions you stopped double-checking, in email, code, or spell-check.
Show answer
Model answer: Most people can name a suggestion feature, autocomplete, spell-check, a code assistant, where their own override or edit rate quietly dropped well before they consciously noticed they'd stopped checking.
Short answer, where it wouldn't matter
5. Name a category in Alder Falls' system where a low override rate isn't actually a warning sign.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Simple, low-risk applications like a basic fence permit with no zoning question. A low override rate there reflects a genuinely easy category, confirmed by audit, not a checking problem.
Short answer, the number question
6. If the override rate had fallen slowly and evenly over two full years instead of sharply over four months, would the same warning-floor design still make sense? Why or why not?
Show hint
Look at the "swap the trigger" section's note about the shape of the drop.
Show answer
Model answer: Possibly not as a hard alert, since a slow, steady decline might reflect real, gradual model improvement rather than a behavior shift. The shape of the drop matters as much as its direction.
Before you close the answer
Why this works
Tests whether you can name a metric that predicts a failure rather than one that only describes it after the fact.
Follow-up traps
"Couldn't you just track the model's confidence score directly instead of the human's override rate?" Response: confidence alone doesn't show whether a human is still genuinely checking it, and it's the checking, not the score, that catches a real mistake before it ships.

"Isn't setting a hard override-rate floor just going to make reviewers override more to avoid triggering an alert?" Response: pair the floor with the independent sampled audit, so a reviewer padding their override count for appearance gets caught the same way a rubber-stamping reviewer would.
If pressed
Alder Falls' actual fix also logs how long the reasoning panel stayed open before a decision, since a sub-second open-to-approve time is a second, independent signal that the override rate alone can't fully capture.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more