ConceptIntermediateShipping & Model Lifecycle / Incident management for AI products / #1

What counts as an incident for an AI feature but not for a normal one?

The direct answer
A normal feature has an incident when it stops doing its job: it's down, it errors, it's slow. An AI feature can do its job exactly as built, answer fast, sound sure, never throw an error, and still be having an incident, because the job is stating a fact about someone's account, and the fact is wrong. Track how often the assistant states an unsupported claim on an action a customer can't undo, sampled and checked against the real record, and treat a sustained rise on that number as the incident, paged like an outage, even while the uptime graph reads green.
Do this, in order
  1. Track how often the assistant states something about an account that isn't true, as its own signal, separate from uptime and error rate.Why: a chatbot can be running perfectly and still be wrong, and uptime can't tell the difference.
  2. Check that signal against the real account record on a routine sample, not on a delay tied to complaints.Why: most customers who get bad information from a bot assume they misunderstood, not that the bot lied, so complaints under-count the real damage for weeks.
  3. Split high-stakes categories from everyday ones and watch each on its own line.Why: a small, dangerous category can drown inside one big healthy average and never move it.
  4. Set a real bar on that line and page the on-call owner the moment it's crossed, same severity as a system outage.Why: a number nobody acts on is decoration, not an incident rule.
  5. Check claims after the chat ends instead of grading every live reply before it sends.Why: grading every message before it goes out roughly doubles the cost and the wait on every single chat; checking after accepts a slower catch instead of slowing down and pricing up every real conversation.
  6. Watch for the assistant learning to hedge its way under the bar instead of getting more accurate.Why: a model that says "not sure, let me check" to everything passes the safety number and quietly breaks the product underneath it.

How to answer this, stage by stage

Nobody is grading whether you can say "monitor the model." They are grading whether you can name the specific, checkable thing that makes an AI mistake different from a normal outage. Six moves get you there.

1
Scope it to one real product and one real company
Say it like this
"Let's ground this. Pellworth Mobile runs a chat assistant called Switchboard, about twenty-two thousand conversations a day, covering billing, roaming, and porting your number in from another carrier. Idara Merrow owns Switchboard's incident policy on the platform team."
Why this works
Grounds the answer in a real daily volume before any definition gets argued in the abstract.
2
Say plainly why "it's still running" isn't proof nothing's wrong
Say it like this
"Here's the reframe I'd open with. For a normal feature, broken means down, or slow, or throwing an error. An AI feature can pass every one of those checks and still be broken, because its job isn't 'respond,' it's 'respond correctly about your account.' Those are two different failure modes, and only one of them shows up on an uptime graph."
Why this works
Names the actual distinction the question is asking about, before any metric gets proposed.
3
Name the signal itself
Say it like this
"So here's what I'd track as the incident signal: how often Switchboard states something specific about a customer's account, a port status, a credit amount, a plan change, with no hedge, that doesn't match the real record. Call it the unsupported-claim rate. It's not model accuracy in general. It's specifically the confident, checkable claims."
Why this works
This is the answer to the question, stated as one countable thing instead of a general worry.
4
Own the real numbers
Say it like this
"Switchboard checks about four percent of daily chats automatically against the account record. In the quiet weeks that number sat around zero point four percent. Over about five weeks it climbed to three point one percent, all inside the porting and SIM-swap categories. Uptime the entire time held at ninety-nine point nine eight percent. The system never once looked broken."
Why this works
Gives the leading number a real path, and shows the lagging number staying healthy on purpose, so the gap between the two is visible, not just claimed.
5
Set the bar and the actual decision at it
Say it like this
"The rule I'd write: if the unsupported-claim rate on an account action a customer can't undo holds above one point five percent for two days running, that's an incident. Paged to the on-call model owner, same as a P1 outage, and Switchboard's answers in that category freeze to a person until it's fixed. Not a ticket in a backlog. A page."
Why this works
A signal with no bar and no owner is a chart nobody reads. This turns it into an actual rule someone acts on.
6
Name what gets gamed, then close on one line
Say it like this
"One thing I'd watch for: the easiest way to shrink that number isn't making the model more accurate, it's having it hedge everything, 'not fully sure, please check your account,' on questions it used to answer fine. That would pass the safety number and quietly wreck the product it's supposed to protect. So here's the actual answer: an AI incident is a sustained rise in confident, wrong, unfixable claims, tracked as its own signal and paged like an outage, not whatever the uptime page happens to say that day."
Why this works
Closes on a definition someone could go check against a real dashboard, not a promise to "watch quality closely."
If you remember one thing An AI incident isn't defined by whether the system is up. It's defined by whether it just stated something false, with confidence, on an action nobody can take back. Uptime and that number can move in completely opposite directions at the same time.

Let's learn

Say a phone company builds a chat assistant that answers billing questions, roaming questions, and questions about bringing your number in from another carrier. Before it existed, a person on the support line pulled up your account and read the real record before saying anything final, like "yes, your port is done." That took a while: on a bad evening, the wait to reach a person ran past eleven minutes.

Then Pellworth Mobile built Switchboard. It answers in seconds instead. About twenty-two thousand chats a day go through it, checking the same account systems a person would, just faster.

Knowledge spark: what's an unsupported claim? A specific, checkable thing the assistant says about your account, stated as fact, that the real record doesn't back up. "Your data balance is 4 gigabytes" when the account shows 2. "Your port is complete" when it's still pending. Not a wrong opinion, a wrong fact.

Here's the part that matters. A few extra wrong answers are not the real problem. The real problem is what kind of thing gets it wrong. Getting your data balance off for a minute costs you a refresh. Telling you your number port is finished when it isn't costs you your phone service, because the sensible next move is cancelling the carrier you're leaving.

Hand-sketched comparison of two gauges. Left, the uptime dashboard, staying green the whole time, nothing crashed. Right, the unsupported claim rate, climbing for weeks before one ticket shows up.
Two dials on the same product. One of them was moving the entire time. Nobody was watching it.
The system was never down. It just started being confidently wrong, and confidently wrong looks exactly like confidently right on an uptime graph.
The decision that mattered Stop waiting for a complaint to trigger a look at Switchboard's transcripts. Sample continuously, check the specific claims against the real account record, and treat a sustained rise on unfixable categories as an incident, not a coaching note for later.

At its worst, Switchboard tells a run of customers their port is complete while the losing carrier's system still shows it pending. They cancel their old line, the way anyone would. For a few days, until the real port finishes, they have no working phone at all. No calls out, no calls in.

The choice I would take back. Switchboard's transcripts only ever got a human read when a customer filed a complaint about one, roughly three in every thousand chats. That made sense at launch, when Switchboard only handled billing and data balance, and a wrong answer meant a customer tapped refresh. It stopped making sense once porting and SIM swaps became categories where a wrong answer doesn't just annoy someone, it strands them.

What I would leave alone. Not every category needs this kind of watching. If Switchboard misquotes this month's data balance by a gigabyte, the customer taps refresh and it corrects off the same account system in real time. Nobody loses anything a second look doesn't fix in five seconds. That category can stay on the old occasional spot-check; there's no reason to page anyone over it.

The lesson. A number that looks healthy in total can be hiding exactly the mistake that matters, because the categories where a model is quietly wrong are rarely where most of the volume sits. If the harmful mistake is rare on purpose, tucked inside a category most people never touch, watching the average will never find it.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how a number can climb for five weeks while every other light on the board stays green.

Idara Merrow has run the platform team behind Switchboard for two years, the group that owns what happens once the model gets something wrong, not just whether it answers fast enough. She wrote the current complaint-triggered review process herself, back when Switchboard only handled two things: "what's my bill" and "how much data do I have left."

Switchboard grew. Roaming status. Discount eligibility. And, eighteen months in, porting: telling a customer whether their number had finished moving over from another carrier. For most of that year the complaint queue stayed thin, three or four a week, mostly people annoyed about tone, not facts. Idara read it Monday mornings with her coffee and moved on.

When the complaint count first crept past five in a week, Idara's instinct wasn't to build something new. She proposed tightening the existing process instead: drop the count that triggered a deeper look from five down to two. She dropped that idea within a day. A queue that only fills up when someone notices they were lied to doesn't get more honest by watching it more closely, it just means reading the same few complaints about tone more often, while the one category that could actually strand someone stays exactly as invisible as before.

She used to read every complaint line by line. Then she skimmed for anything with the word "wrong" in it. Then, most weeks, she just checked the count stayed under five and closed the tab.

There wasn't a Tuesday where it happened. Nobody can point to the day it started. Sometime over about five weeks, in a stretch nobody was watching closely, Switchboard started telling a small slice of porting customers their number transfer was complete a little early, before the losing carrier's system had actually released it.

Customers did the sensible thing. Told their port was done, they cancelled the carrier they were leaving. Most real ports finished within a day of that message anyway, so most people never noticed a gap. A few didn't. For a handful of people, the real port dragged an extra two, three, four days behind what Switchboard had told them, and during that stretch they had no working line at all.

We didn't lose a few wrong chat messages. We handed a stranger's actual phone line to a guess.

It was never really about how many chats got a wrong answer. Switchboard answered every one of its twenty-two thousand daily chats exactly the way it was built to: instantly, in full sentences, with no error. The failure wasn't in whether it answered. It was in whether the specific thing it said was true, and nothing in the pipeline was built to ask that question on its own.

The complaint-triggered review process went back to a planning meeting eighteen months earlier, when Switchboard only handled billing and data balance. Someone asked how they'd catch a bad answer. The answer in the room was, "if it's actually a problem, someone will complain, and we'll see it in the queue." Nobody was wrong to say it. At the time, a wrong answer meant a customer refreshed a number. There was nothing porting-shaped in the product yet for anyone to worry about.

Run the same five weeks through the sampled unsupported-claim rate instead. It reads zero point four percent in week one, same as always. By week three it's crossed one point five percent on the porting category, held for two days straight, and the rule fires automatically, before Idara has opened her laptop that morning. Porting answers freeze to a human handoff by lunch. Six customers got an early port confirmation before the fix shipped two days later. Not sixty-three.

What I'd tell myself, back in that first planning meeting: "someone will complain if it's a real problem" is only true for mistakes people can tell happened to them. A wrong port confirmation doesn't look like a bug from the other side of the chat window. It looks like good news, right up until your phone stops working, and by then most people are calling their old carrier, not us.

LEAD: the four things worth tracking instead of uptime

This is a metric question, what actually tells you the AI feature is failing, not a system-health question, so LEAD fits and a normal incident checklist doesn't.

L, link. The real outcome underneath a "port complete" message: whether a customer who trusts it still has a working phone line a week later, not the model's own confidence. Every category Switchboard answers ties back to some outcome like this; porting's is unusually unforgiving because the customer's next move, cancelling the old line, can't be undone.
E, early signal. The unsupported-claim rate: how often Switchboard states a specific, checkable fact, no hedge, that the real account record disagrees with, sampled continuously and checked automatically. This is the number that moved for five straight weeks while uptime sat at 99.98 percent the entire time.
A, abuse. The rate drops if the model just hedges more, "not sure, please check," on questions it used to answer confidently and correctly. That would look like an improvement on the dashboard and be a worse product underneath it.
D, decision. Below 1.5 percent, sustained: nothing changes, that's normal model noise. At 1.5 percent held for two days on a category a customer can't undo: page the on-call model owner, freeze that category to a human handoff, same severity ladder as a system outage. Below that line nobody gets woken up over a stray wrong data-balance guess; above it, someone does, on purpose.

One more thing worth saying: 1.5 percent isn't a number pulled from nowhere. It's roughly four times the rate the same check found on Switchboard's lowest-stakes category, data balance questions, where the model runs the exact same setup and trips over the exact same kind of confusing phrasing. If porting's rate matched that quiet-day baseline, nobody would call it an incident. It's sitting four times above its own noise floor that makes the bar real instead of made up.

Hand-sketched comparison. Left, a document showing the metric on paper, unsupported claim rate reads low because the model just hedges more. Right, a person, what the customer actually got, told the port finished, cancels the old line.
A metric can be satisfied on paper while the thing it stands in for is still going wrong. That's why the abuse check is its own line, not an afterthought.
The leading signal: unsupported claims on porting, sampled weekly
1.5% 0% 3.5% Wk 1 Wk 2 Wk 3 Wk 4 Wk 5 Wk 6 bar crossed
Under the old process, nothing reads this chart until week six. Under the rule Idara wrote after, the page fires the moment week three's second day crosses the line.
The lagging outcome: support tickets mentioning a porting problem, by week
Week 13 tickets
Week 24 tickets
Week 3, the week the leading signal crossed 1.5%4 tickets
Week 45 tickets
Week 56 tickets
Week 6, when the old process finally noticed41 tickets
Three whole weeks separate the leading signal crossing its bar and the lagging signal finally saying something loud enough to notice. That gap is the actual cost of waiting for complaints.

And if you want to be sure it really works, try it somewhere else

Ardmoor Terminal runs Dockline, a chat assistant dispatchers and truckers use to check whether a container has cleared customs and can leave the gate. Roan Selkirk runs dispatch operations there.

L, link. Whether a container that leaves the gate is actually allowed to, not whether Dockline sounds confident about it. Moving an uncleared container off the terminal is a real compliance violation, with fines attached to the trucking company and the terminal both.
E, early signal. The share of "cleared for release" answers Dockline gives that don't match the customs system's own record, sampled continuously, not caught by asking whether the chat widget loaded or answered fast.
A, abuse. During a customs-system integration change, the rate climbed from a baseline of 0.6 percent to 2.8 percent over three weeks. Nobody was gaming it on purpose; the new integration was quietly serving Dockline a cached status a few hours stale, and stale looked exactly like current on the screen.
D, decision. Under the old setup, with no separate signal for this, 11 containers left the gate before their customs release was actually confirmed, each one a real fine. Running the same 1.5-percent, two-day rule Idara wrote for Switchboard against Dockline's numbers, the bar gets crossed in the first week of the integration change; the projected count under that rule is 2 containers, not 11, because clearance answers would have frozen to a live check against the customs system itself instead of trusting a cached number.

Same shape, different stakes At Pellworth a missed signal costs a customer their phone line for a few days. At Ardmoor it costs a trucking company a fine and the terminal a compliance mark. The definition doesn't change: whatever the assistant states as fact about an action nobody can take back gets its own watched number, separate from whether the system is up.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the reframe: an AI incident is a sustained rise in confident, wrong, unfixable claims, not whatever uptime happens to say.
Cost: engineering can only afford to check a small slice of daily chats. Don't drop the high-stakes categories out of the sample to save budget; shrink the sample size on the low-stakes categories instead, since a wrong data-balance answer costs a refresh either way.
The model got better: a new Switchboard version turns out to be more accurate on porting than the one it replaced. That doesn't make the signal pointless, it's still the only way anyone would know that for sure instead of assuming it from a quiet complaint queue.

Where people run it wrong.
They watch uptime and error rate and call that "monitoring the AI feature," when those checks only catch the system stopping, not the system confidently lying while it keeps running.
They wait for complaints to define the problem, when most people hurt by a wrong AI answer assume the mistake was theirs, not the bot's, and never file one.
They set one blended number across every category the assistant answers, so a small, dangerous category disappears inside a big, healthy average.

How to use it live. Say the reframe before naming a single number: "the incident isn't whether it's up, it's whether it just stated something false about an action the person can't undo." That buys the room to ask what this specific product's unfixable categories actually are, instead of reciting a generic uptime checklist.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits "what counts as an incident for an AI feature," and why not FLIPS?
Tap to flip
ANSWER
LEAD. This is a metric question, what's the actual signal that something's wrong, not a story about a person's habit flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Idara Merrow, who has run the platform team behind Pellworth Mobile's Switchboard chat assistant for two years, and owns its incident policy.
3 · THE HABIT
What did Idara stop doing to Switchboard's complaint queue because it kept looking fine?
Tap to flip
ANSWER
She went from reading every complaint line by line, to skimming for the word "wrong," to just checking the weekly count stayed under five.
4 · THE TWO SIGNALS
What's the difference between what uptime shows and what the unsupported-claim rate shows?
Tap to flip
ANSWER
Uptime held at 99.98 percent the whole five weeks, perfectly healthy. The unsupported-claim rate climbed from 0.4 to 3.1 percent in the same stretch, the only number that actually moved before anyone noticed.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Only reviewing chat transcripts when a customer complained. It made sense when Switchboard only handled billing and data balance, where a wrong answer cost a refresh. It stopped making sense once porting made a wrong answer cost someone their phone line.
6 · THE NUMBER
Fill in the blank: the incident rule fires when the unsupported-claim rate on an unfixable category holds above ___ percent for ___ days running.
Tap to flip
ANSWER
1.5 percent, for 2 days running. That's roughly four times the rate the same check finds on Switchboard's lowest-stakes category on a quiet week.
7 · THE REPLAY
Same five weeks, new rule in place. What changes?
Tap to flip
ANSWER
The rule fires automatically in week three, the moment the rate holds above 1.5 percent for two days. Porting freezes to a human handoff by lunch. Six customers get a wrong port confirmation before the fix ships two days later, instead of sixty-three.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the unfixable category?
Tap to flip
ANSWER
Ardmoor Terminal's Dockline assistant. Its unfixable category is telling a trucker a container is cleared for release when customs hasn't actually released it, a mistake that can't be undone once the container leaves the gate.

Check yourself Score: 0 / 0

Multiple choice
1. Why couldn't Pellworth Mobile's uptime and error-rate monitoring catch Switchboard's early port confirmations?
  • A. Because the monitoring system itself was down during those five weeks.
  • B. Because Switchboard was answering every chat correctly by the system's own definition of working, fast, no errors, full sentences, while the specific facts it stated were wrong.
  • C. Because uptime monitoring only works for features that don't use a model at all.
  • D. Because the porting category wasn't included in Switchboard's launch plan.
Show hint
Think about what "the system is up" actually measures, versus what it doesn't measure.
Show answer
B. Uptime measures whether the system responds. It says nothing about whether what it says is true, and that gap is exactly where an AI-specific incident can hide.
True or false
2. True or false: reviewing chat transcripts only when a customer files a complaint is a reasonable way to catch a rare, high-stakes AI mistake.
  • True
  • False
Show hint
Think about what a customer who got a wrong but good-sounding answer actually assumes happened.
Show answer
False. A customer told their port is complete has no reason to think the bot lied, it sounds like good news. Most people in that spot blame their old carrier or their own timing, not Switchboard, so the complaint queue stays quiet exactly when the real damage is building.
Fill in the blank
3. The unsupported-claim rate on porting climbed from ___ percent in week one to ___ percent by week five, while uptime held at ___ percent the entire stretch.
Show hint
Check the framework recap's E step and the leading-signal chart.
Show answer
0.4 percent; 3.1 percent; 99.98 percent. The gap between those two numbers, one climbing, one staying perfectly flat, is the whole argument for why uptime isn't a substitute for a truthfulness signal.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the planning meeting from eighteen months before the porting category existed.
Show answer
Model answer: Relying on customer complaints to trigger a review of Switchboard's transcripts. It made sense at launch, when the assistant only handled billing and data balance, and a wrong answer cost a customer a few seconds to refresh. It stopped making sense once porting turned a wrong answer into a lost phone line.
Short answer
5. Name a category inside Switchboard where a wrong answer would NOT count as an incident under this rule, and say why.
Show hint
Look at "what I would leave alone" in Section 1.
Show answer
Model answer: The data balance category. If Switchboard misstates how much data is left, the customer taps refresh and sees the real number in seconds, off the same account system. Nothing about that mistake is unfixable, so it stays on the old occasional spot-check instead of the paged incident rule.
Short answer, apply it yourself
6. Think of an AI feature you use yourself. What's a category inside it where a confident, wrong answer would be worse than the app going down for an hour?
Show hint
Look for something the app states as fact that you'd act on right away, and can't take back once you do.
Show answer
Model answer: A navigation app confidently rerouting you onto a road it says is open, when the road's actually closed for construction. The app going down for an hour just means you drive without it. The app confidently sending you down a closed road means you're stuck, possibly somewhere without a safe place to turn around.
Before you close the answer
Why this works
Tests whether you define "incident" by what a normal engineering team already tracks, uptime, errors, latency, or notice that an AI feature's real failure mode is confident, plausible wrongness that none of those dashboards can see. Most candidates reach for "monitor the model closely" and stop; the strong answer names one specific, checkable signal and a real bar on it.
Follow-up traps
"Isn't this just a fancy accuracy number, not a real incident definition?" Response: it becomes an incident definition the moment it has a stated bar, a named owner, and an action tied to crossing it, which is exactly what turns a chart nobody reads into a page someone answers.

"Won't sampling four percent of chats miss the exact bad conversation that matters?" Response: yes, and that's an accepted trade, not an oversight. A single missed chat in the sample is expected noise; the rule is built on a sustained rate crossing the bar over two days, not on catching any one wrong answer as it happens.
If pressed
Of the 63 customers who got an early port confirmation before the old process caught the problem in week six, 9 went more than 24 hours with no working phone line at all, the number that actually justified treating this as a page-worthy incident rather than a quality backlog item.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more