CalculationAdvancedShipping & Model Lifecycle / Incident management for AI products / #21

How would you measure whether your incident response is improving?

The direct answer
Track self-detection rate: the share of routing incidents your own monitoring catches before a merchant calls in, not incident count or average time-to-resolve. Pair it with a 14-day recurrence rate for the same root cause, since detection rate alone can be gamed. Treat 80 percent weekly detection as the floor, and freeze new routing-model rollouts the moment it drops under 50.
Do this, in order
  1. Track self-detection rate as the number that actually answers the question.Why: it moves weeks before incident count or time-to-resolve show anything is wrong.
  2. Pair it with a 14-day same-root-cause recurrence rate.Why: self-detection rate alone can be gamed by loosening what counts as "caught." Recurrence catches a quick patch that never fixed the real thing.
  3. Build the detector independent of the routing model's own confidence.Why: a model that's confidently wrong on a stale traffic feed will never flag itself. A GPS-deviation check that never asks the model how sure it is will.
  4. Attach a real action to each threshold, not a flat "zero incidents" rule.Why: a routing engine recomputing thousands of stops a day in live traffic will never hit zero surprises. A flat rule just gets quietly redefined instead of met.
  5. Score narrow-delivery-window accounts on a tighter bar than the average.Why: the average can sit fine while the one segment a merchant would actually leave over is exactly where detection is failing.
  6. Put both numbers in front of whoever owns the merchant relationship, not just engineering.Why: a number nobody outside the ops desk ever sees is a chart on a wall, not a decision.

How to answer this, stage by stage

Seven moves. The hardest one is naming the two metrics you're refusing to use before you name the one you'd use instead.

1
Scope it to one company and one real incident
Say it like this
"Let's ground this. Thrace Dispatch routes about fourteen thousand stops a day for six hundred merchant accounts, through Waybound, the engine that plans and re-plans every driver's route in real time. Damaris Verrall runs incident response there. She's the one this question is really asking."
Why this works
Grounds the answer in real numbers before any metric talk starts, so it never turns into a definitions lecture.
2
Name the two metrics you're rejecting, before you name your own
Say it like this
"Before I answer, I want to say what I'm not using. Time-to-resolve and monthly incident count are the two numbers most ops teams reach for, and I'd throw both out as the answer to 'is it improving.' You can make time-to-resolve look great by fixing the fast, obvious incidents quickly, while a slower, real problem never even becomes a ticket."
Why this works
Naming a rejected alternative up front is what tells the interviewer you made a call, not that you reached for the first metric on the dashboard.
3
Name the real outcome underneath the question
Say it like this
"Here's what actually matters: whether a merchant who goes through a rough week still trusts Thrace enough to keep routing volume through it. Not 'fewer incidents happen.' Whether the ones that do happen cost you the account."
Why this works
This is the L step, said plainly, before any number. Skip it and the metric you name next has nothing real to aim at.
4
Give the early signal, the actual answer
Say it like this
"So the number I'd watch every week is self-detection rate: of every routing incident that happens, what share does Waybound's own monitoring catch before a merchant calls in? That number moves weeks before incident count or time-to-resolve tell you anything's wrong."
Why this works
This matches deliverable zero exactly. A metric answer with no mechanism behind it is just a nicer-sounding guess.
5
Say how someone would hit that number without fixing anything
Say it like this
"And I'd say straight away how that number lies. You can raise self-detection rate on paper by loosening what counts as 'detected,' counting a low-confidence flag as a catch even when nobody confirmed a real incident happened. Or you hide a repeat failure by filing it under a new incident ID with a slightly different description, so it never shows up as the same problem twice."
Why this works
Naming the abuse case before anyone else does is what separates someone who understands metrics from someone who just picked one off a list.
6
Attach a real action to every threshold
Say it like this
"So I'd pair it with a second number: how often the exact same root cause comes back within fourteen days. Above 80 percent self-detection most weeks, normal cadence. Between 50 and 80, the narrow-window merchants get their own tighter, stratified audit. Under 50, or two repeats of the same root cause in a month, any pending routing-model update gets frozen until an independent check clears it, not the model's own confidence."
Why this works
A metric nobody acts on is decoration on a dashboard. This line proves the number actually changes what the team does.
7
Close on the number that would have rung early
Say it like this
"That's the whole point of watching self-detection instead of incident count. A rate sliding from 74 percent to 31 over five months would have told Damaris the real story two months before a merchant cut its volume by 40 percent. Incident count alone told her the opposite story the entire time."
Why this works
Closes on a number someone could go check, not a promise to "watch things more closely" next quarter.

Let's learn

What happens when the number a team is proud of and the number a merchant actually feels stop moving in the same direction?

Say a last-mile delivery company runs an AI engine that plans, and constantly re-plans, a driver's route as traffic and delays change through the day. Thrace Dispatch runs one, called Waybound, across about fourteen thousand stops a day for six hundred merchant accounts.

Underneath Waybound sits an anomaly monitor: a rule that watches every stop and flags anything that drifts too far off plan. At a smaller scale, tuned tight, that monitor caught close to three out of every four real routing incidents before a merchant ever called in. Call that self-detection rate: 74 percent.

Knowledge spark: what's an anomaly-detection threshold? A rule that decides when a computer should tap someone on the shoulder. Drift past the line, get flagged. Stay under it, nothing happens, even if something underneath is actually wrong.

Then the company grew. Fourteen thousand stops a day meant that same tight threshold was throwing off close to three thousand alerts daily, almost all of them nothing: a driver grabbing coffee, a red light, five ordinary minutes of traffic. The desk couldn't keep up, and they stopped trusting the alerts at all.

So the team widened it. Instead of flagging any stop running more than twelve minutes late, Waybound would only flag one once it drifted past forty. Alert volume dropped 90 percent overnight. Everyone breathed out.

Over the next five months, nothing dramatic happened. Monthly incident count drifted down. Average time to resolve got faster too. Every chart in the weekly ops review looked like a team getting better at its job.

Self-detection rate was doing the opposite the entire time. It just wasn't on anyone's chart.
Two hand sketched gauge panels side by side under the title Which clock rings first. Left panel, incident count and time to resolve, caption looks fine right up until the merchant leaves. Right panel, self-detection rate, caption moves first, weeks before anyone else notices.
Two numbers went into the same weekly report. Only one of them was telling the truth about what was coming.

Here's what incident count never showed. Most of Thrace's real routing problems don't blow one stop by forty-five minutes. They blow twenty or twenty-five stops at once, each one twenty to thirty minutes late, in the same cluster, because Waybound re-routed a whole zone around a road closure that had already reopened, or because a traffic feed sat stale for an hour and nobody knew it. No single stop in a run like that ever crosses the forty-minute bar on its own. The whole cluster sails past the monitor clean, every time.

Self-detection rate, before and after the threshold widened
80% 40% 0 50% floor before month 1 month 2 month 3 month 4 month 5
Self-detection rate crossed under the 50 percent floor by month three, a full month before any merchant felt it. Nothing on the incident-count chart that same month said a thing was wrong.

Cold Harbor Grocers found out a different way. Cold Harbor ships frozen goods on narrow half-hour windows, and across three separate weeks that spring, one of its delivery clusters ran twenty-five to thirty-five minutes behind, batch after batch. Nobody at Thrace called Cold Harbor about it. Cold Harbor's own customers called Cold Harbor.

The choice I'd take back Widening the anomaly threshold from twelve minutes to forty, as a permanent default, without ever re-checking whether the wider setting still caught what actually mattered to narrow-window merchants once the company scaled up. It was the right call for the noise problem it was fixing. Nobody ever went back and asked what it stopped catching.

At its worst: Cold Harbor didn't wait for a renewal conversation. The quarter after that third bad week, it quietly moved 40 percent of its volume to a second carrier. No complaint filed, no meeting requested. Thrace's own dashboard, that same quarter, said incidents were down.

Share of Cold Harbor's volume routed through Thrace, before and after
100% 50% 100% 60% before next quarter
A 40 percent cut, made quietly, with no ticket and no call. This is the outcome self-detection rate was trying to protect. Incident count never touched it.

What I'd leave alone. A single paperback arriving forty minutes late to someone who wasn't watching the clock. Chase a threshold tight enough to catch that, and you rebuild the exact noise problem that got the threshold widened in the first place, and the desk stops trusting alarms all over again.

The lesson. A number that goes down for two quarters straight isn't proof anything got better. It's only proof of what the team decided to count.

Now here is the same thing as a story

Use this version when you want to feel why a new hire's question landed harder than any chart in the room could.

Damaris Verrall doesn't guess. Five years on the incident desk at Thrace Dispatch taught her that a merchant's tone in an email tells you almost as much as the numbers do, and by the time Zola Mabaso joined the team that spring as an operations analyst, Damaris had a reputation for being the person who could explain any chart in the building cold, no prep, on the spot.

She built that reputation on real numbers. Waybound's rollout, her second year on the job, gave the desk something to be proud of: live rerouting, real traffic, drivers who stopped calling dispatch every time a road closed because the app just handled it. The anomaly monitor, tuned tight at twelve minutes, caught close to three out of four real incidents before anyone outside the company noticed. Damaris pulled that self-detection number herself every week for two years. She used to print it out.

Then the volume came. Fourteen thousand stops a day meant the twelve-minute rule was throwing three thousand alerts daily, and reviewing three thousand alerts a day isn't a job, it's a flood. Widening the threshold to forty minutes was the obvious fix, agreed in one Thursday meeting, barely debated. Alert volume dropped 90 percent. The desk could breathe again.

Damaris stopped printing the self-detection number sometime that year. Not on purpose. Incident count and resolve time both looked so good, so steadily, that it stopped occurring to her to ask what else had moved.

Nothing broke, for months. The good numbers kept being good.

Then came the quarterly review prep, four months after the threshold changed, a Tuesday afternoon with sandwiches going stale on the conference table. Zola, still new enough to ask the questions everyone else had stopped asking, was building the merchant-health slide and noticed something that didn't fit: Cold Harbor Grocers, a five-year account, had cut its shipped volume through Thrace by nearly 40 percent the previous quarter. She asked Damaris, half out loud, half to herself: "if incidents are down and resolve time is down, why did they leave?"

Damaris didn't have an answer. Not because she hadn't looked. Because the dashboard she'd been trusting had nothing in it that could have told her.

We didn't lose forty percent of an account. We lost the ability to notice we were losing it.

She pulled the raw incident logs herself that night, not the dashboard, the actual logs, and rebuilt what happened to Cold Harbor's account over the quarter. Three separate weeks, one cluster each time, twenty to twenty-five stops running twenty-five to thirty-five minutes behind their promised windows. Every one of those stops sat comfortably under the forty-minute bar. Not one alert fired. Cold Harbor's own customers had been first to notice, every single time, and Cold Harbor had been quietly routing around Thrace for a full quarter before anyone at Thrace had the numbers in one place.

The meeting where the threshold moved from twelve minutes to forty happened on a Thursday, eighteen months earlier. Someone on the desk said the twelve-minute rule was "basically noise at this point," and nobody in the room disagreed, because on the incidents anyone could remember, it mostly was. Nobody asked what a real incident would look like at the new scale, spread thin across two dozen stops instead of piled onto one.

The replay, run the same five months forward with self-detection rate on the wall instead of just incident count: by month three, the rate had already crossed under 50 percent, a full month before Cold Harbor's first bad week. A stratified audit on narrow-window accounts, the kind the new design triggers the moment detection drops under 50, would have caught the cluster pattern on its first real occurrence, not its ninth.

What I'd tell myself, back in that Thursday meeting: "basically noise" is a sentence about last month's incidents, never next year's. I signed off on a default that fit the traffic I already had, and never once asked what traffic I didn't have yet.

LEAD, the four questions Damaris should have asked before the QBR

This is a metric question wearing an operations review's clothes, so the framework is LEAD, not FLIPS. Nothing in this story is a single moment that snapped. It's a number that quietly stopped telling the truth, which is exactly the shape LEAD is built for.

L
Link. The real business outcome, not the model's own score.
Not "fewer incidents." Whether the incidents that happen still cost you the account.
Here, it's whether a merchant who goes through a rough patch still trusts Thrace enough to keep routing volume through it.
E
Early signal. The number that moves before the outcome does.
This is the answer to the question, and the reason LEAD exists at all.
Self-detection rate: the share of a week's real routing incidents Waybound's own monitoring catches before a merchant reports one.
A
Abuse. How the number gets hit without the real problem going away.
Every metric has a cheap way to be satisfied.
Count a low-confidence flag as "detected" without confirming a real incident happened, or relabel a repeat root cause under a fresh incident ID so it never counts as a recurrence.
D
Decision. What you'd actually do at each threshold.
A metric nobody acts on is a dashboard decoration.
Above 80 percent, normal cadence. 50 to 80, a stratified audit on narrow-window accounts. Under 50, or two repeats of one root cause in a month, freeze routing-model rollouts until an independent check clears it.
Two hand sketched panels under the title How the leading signal gets gamed. Left panel, a gauge, caption self-detection rate creeping back up. Right panel, a person, caption same routing gap, filed under a new incident ID.
A rate can climb back up on the dashboard while the actual gap it was built to catch just gets renamed.
The check that makes LEAD honest Try swapping the outcome in the L step. Change it from "the merchant still trusts us" to "the model never makes a routing mistake." The number changes completely, and it stops being something you could ever hit, because a live routing model recomputing thousands of stops a day in real traffic will make mistakes. If your metric survives any outcome you plug into it, you never really tied it to one.

And if you want to be sure it really works, try it somewhere else

Kessington County runs Corvalt, an AI tool that pre-screens building-permit applications for missing documents and code violations before a human reviewer signs off, so applicants aren't rejected for a missing signature after a three-week wait.

L
Whether applicants and contractors keep trusting the county's electronic review enough to keep filing online, instead of demanding an in-person review from day one.
E
Self-detection rate: the share of wrongly approved applications Corvalt's own weekly audit sample catches before a field inspector finds the same gap on-site.
A
Audit only the easy checklist items, page count, a signature present, and never the substantive code-compliance calls, a setback distance, a load rating, the ones an inspector would actually catch first.
D
Above the floor, normal cadence. Below it, stratify the audit toward whichever code section keeps slipping through. Two repeats of the same code section in 30 days moves that section to mandatory human sign-off for 60 days, no matter what Corvalt's own flag says.
Hand sketched flow diagram titled Corvalt's audit trail, permit by permit, five connected boxes: filed, pre-screened, audit sample highlighted, inspector visit, repeat section.
Same shape, a different desk. The audit sample is still the step that catches a gap before an inspector's visit does.
Same shape, different cost At Thrace the missing check was a threshold nobody re-checked after scaling up. At Kessington it's an audit sample that only ever looked at the easy half of the checklist. Different mechanism, same real finding: a leading metric only leads if the thing it's sampling is the thing that actually breaks.

Swap the trigger and it still runs.
Speed: an interviewer caps you at 90 seconds. Skip straight to the number and the pairing, self-detection rate plus recurrence, and leave the story out entirely.
Cost: the incident desk gets cut in half this quarter. Don't drop self-detection sampling to zero to save headcount. Narrow it to the highest-risk segment, narrow-window merchants, and keep full coverage there.
The model got better, for real: suppose Waybound's next version genuinely cuts real incidents in half. Self-detection rate is still the right number to watch, because it tells you whether the incidents that do still happen are the ones you're catching, not just that there are fewer of them.

Where people run it wrong.
They watch incident count falling without ever asking whether the definition of "incident" quietly got narrower underneath them.
They treat a widened anomaly threshold as a permanent setting instead of something that needs re-checking every time volume or the merchant mix changes.
They average self-detection rate across every merchant instead of scoring the narrow-window, highest-stakes accounts on their own tighter bar.

How to use it live. Say the outcome before the number: "the real question isn't whether incidents went down, it's whether the merchants who had one still trust you, so let me say what number would tell you that first." That's not stalling. It's the L step, and it buys you the room to land on self-detection rate instead of blurting out "MTTR" and hoping nobody asks why that's the wrong answer.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and why not FLIPS?
Tap to flip
ANSWER
LEAD, for a metric question. FLIPS finds the moment someone's behavior snaps after a change. Here nothing snapped, a number just quietly stopped telling the truth, which is what LEAD is built to catch.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Damaris Verrall, incident response lead at Thrace Dispatch, five years on the desk, who used to print the self-detection number every week and pin it up.
3 · THE HABIT
What did she stop doing because the other numbers looked so good?
Tap to flip
ANSWER
Checking self-detection rate every week. Once incident count and resolve time both looked great after the threshold widened, she stopped asking what else had moved.
4 · THE L STEP
What's the real outcome underneath "is incident response improving"?
Tap to flip
ANSWER
Whether a merchant who goes through a rough patch still trusts you enough to keep routing volume through you. Not whether incident count went down.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Widening the anomaly threshold from 12 minutes to 40 as a permanent default when volume scaled up, and never re-checking whether it still caught what mattered to narrow-window merchants.
6 · THE NUMBER
Fill in the blank: self-detection rate fell from ___ percent to ___ percent over five months, while Cold Harbor cut its volume by ___ percent.
Tap to flip
ANSWER
74 percent to 31 percent. Cold Harbor cut its volume by 40 percent the following quarter.
7 · THE REPLAY
Same five months, new design, what changes?
Tap to flip
ANSWER
Self-detection rate crosses under 50 percent by month three, a full month before Cold Harbor's first bad week. A stratified audit on narrow-window accounts catches the cluster pattern on its first real occurrence.
8 · CROSS PRODUCT TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
Corvalt, a building-permit pre-screening tool at Kessington County. Its early signal is the share of wrongly approved applications the county's own audit sample catches before a field inspector does.

Check yourself Score: 0 / 0

True or false
1. True or false: once monthly incident count and average time-to-resolve both improved for five straight months, that was solid evidence Thrace's incident response was actually getting better.
  • True
  • False
Show hint
Ask what kind of incident a 40-minute single-stop threshold would never catch.
Show answer
False. Both numbers can improve while a widened anomaly threshold quietly stops catching a whole category of real incidents, ones spread thin across a cluster of stops, before merchants notice them.
Multiple choice
2. Which of these would raise Thrace's self-detection rate on paper without making incident response actually better?
  • A. Hiring two more people to the incident desk.
  • B. Counting a low-confidence Waybound flag as "detected" even when nobody confirmed a real incident happened.
  • C. Lowering the anomaly threshold back to 12 minutes for every merchant.
  • D. Auditing narrow-window merchants on a tighter bar than the average.
Show hint
Three of these change what the team actually does. One just changes what gets written down.
Show answer
B. That's the abuse case named in the LEAD recap: a flag counted as a catch with nobody checking whether a real incident happened moves the number without moving the truth.
Fill in the blank
3. The anomaly threshold widened from ___ minutes to ___ minutes when Thrace scaled up, and alert volume dropped by ___ percent overnight.
Show hint
Both numbers are in "Let's learn," right after the anomaly-monitor spark.
Show answer
12; 40; 90. That 90 percent drop in noise is exactly why the wider default felt like an obvious win at the time.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the Thursday meeting, not a dial someone could just turn back up.
Show answer
Model answer: Widening the anomaly threshold from 12 to 40 minutes as a permanent default, without ever re-checking it as volume scaled. It made sense at the time because the tight threshold was genuinely mostly noise on the incidents anyone remembered, and cutting alert volume 90 percent let the desk breathe again.
Short answer, apply it yourself
5. Pick a tool you use at your own job that reports a health number of some kind. Name one thing it could quietly stop catching if that number's definition never got re-checked as usage grew.
Show hint
Look for a rule tuned once, early, that nobody has revisited since.
Show answer
Model answer: A helpdesk tool flags a ticket as "urgent" if it sits unanswered for two hours, a rule set when the team handled forty tickets a day. At four hundred tickets a day, a genuinely urgent one could sit inside a busy queue for ninety minutes and never cross that two-hour line, while the dashboard still shows "urgent tickets: 0" and looks perfectly healthy.
Short answer, the number question
6. If Cold Harbor's cluster had run just under thirty minutes late instead of twenty-five to thirty-five, would the 40-minute threshold still have missed it? Would a tighter, stratified audit on narrow-window accounts still have caught it?
Show hint
Compare the single-stop threshold's bar against a stratified audit's actual job.
Show answer
Yes, and yes. Anything under 40 minutes per stop clears the anomaly threshold regardless of how many stops are affected, so the flat rule misses it either way. A stratified audit doesn't wait for a single stop to cross a bar. It samples narrow-window accounts on a fixed schedule and checks the whole cluster's pattern, so a batch of twenty-eight-minute misses across two dozen stops still gets caught.
Before you close the answer
Why this works
Tests whether you'll reach for the metric that already sits on the dashboard, incident count, time-to-resolve, or build the one that actually predicts merchant churn, and whether you can say out loud how your own metric gets gamed before the interviewer finds the hole first.
Follow-up traps
"Isn't self-detection rate just another number the same broken monitor produces? How is it different from incident count?" Response: it's produced by a check that's independent of Waybound's own confidence, comparing a driver's actual GPS position against the plan, so it can catch exactly the case where the model is confidently wrong and would never flag itself.

"Why not just tighten the anomaly threshold back to 12 minutes for everyone?" Response: that recreates the exact noise problem that got it widened in the first place, three thousand alerts a day at current volume. The fix is a stratified, segment-aware audit on the accounts that actually can't absorb a miss, not one global number applied to every merchant regardless of what a delay costs them.
If pressed
The GPS-deviation detector doesn't compare a driver's live position to Waybound's current plan. It compares it to the plan Waybound was confident in about ninety seconds earlier, before its own last recompute, because checking against the very latest plan means the model marking its own homework in real time, the exact trap that let the stale-feed problem hide in the first place. Running that check inline, on every recompute, would add real delay to a routing engine already working on a tight loop, so it runs a short lag behind instead, a small, accepted delay in exchange for not slowing the live routing decision itself.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more