How does the severity distribution of errors matter more than the error rate?
Two systems can share the exact same error rate and carry nowhere near the same risk. What decides that is where the mistakes land, not how many there are.
- Split the miss rate by severity tier, and give the rarest, worst tier its own bar.Why: one blended number lets a miss on the worst tier hide inside thousands of correct calls on the mild one.
- Audit the top severity tier in full, not by random sample.Why: at low volume, a random sample almost never picks up enough of the rarest tier to know if it's actually healthy.
- Set a tighter pass bar for that tier than the rest of the queue tolerates.Why: a miss that never reaches a person costs far more on the worst tier than on the mild one, so the same bar can't cover both.
- Re-pull the severity split whenever the blended number goes quiet for a long stretch.Why: calm for months is exactly when a person stops opening the file that would have caught this.
- Close the routing or sampling gap that let the miss through, not the whole model.Why: the check that mattered showed the model's overall judgment hadn't moved, the sampling design had the blind spot.
- Accept that some borderline posts wait longer for a person to clear them.Why: tightening the net around risky language holds back a few real jokes too, a cost worth paying only on the tier where a miss is this expensive.
How to answer this, stage by stage
Nobody's grading whether you know severity matters. They're grading whether you'd trust a flat number, or ask what's actually inside it before you call a system healthy.
Let's learn
Here is what happens when a number stays perfectly calm while something underneath it gets worse.
Fenmarsh is a short video app. Every clip and caption that gets posted runs through a flagging model the moment it goes up. The model sorts anything it's unsure about into three severity tiers: mild rule-breaks like spam or a rude joke, moderate problems like bullying, and severe cases like a real threat of violence. Ilsabet Kalman runs trust and safety there.
Before the flagging model, a small team of moderators only looked at posts someone had reported by hand, working through one shared queue in the order the reports came in. They cleared about 22,000 reports a week, roughly 4 minutes each, and on a busy weekend the queue backed up more than 30 hours. A mild joke report and a genuine threat report sat in the exact same line, waiting their turn.
Now the flagging model checks all 3.4 million uploads a day on its own, clears almost everything instantly, and sends about 40,000 posts a week to a person, tagged by tier. Once a week, a senior reviewer hand-grades a random sample of 2,000 of the model's decisions, spread across all three tiers, and reports back one number: how often that sample agrees with what the model did. For fourteen months straight, that number held between 92.4 and 93.1 percent.
Then, in week 51, one post got scored wrong in the one direction that actually mattered. Someone posted a specific, named threat of violence against a real person, worded in the same short, punchy style as a common joke format the model had learned to wave through. The model tagged it mild, not severe, and cleared it instantly. It stayed up for 36 hours, earning ad views the whole time, until a friend of the target reported it outside the app, straight to the police.
That same week, the blended number didn't dip. It rose, to a new high of 93.1 percent, because the model also caught an unusually clean batch of spam and copycat mild posts.
Once Ilsabet went back and hand-graded the full population of severe-tier decisions from that week, not just the slice the random sample happened to touch, the real picture came out: mild-tier posts missed at about 3 percent that week, moderate at about 5 percent, and the severe tier, only 80 posts total, missed at nearly 9 percent. One in eleven.
What it cost, at its worst: police called Fenmarsh 36 hours after the post went up, asking why a threat naming a real address had been live and monetized the whole time. Nobody inside Fenmarsh had flagged anything wrong, because the one number everyone watched had just had its best week ever.
What I would leave alone: the mild tier's number is genuinely fine as a blended random sample. A missed spam post or an overturned rude joke costs almost nothing, and it usually gets caught on the next report anyway. Auditing that tier in full every week, the way the severe tier now needs, would spend real reviewer time on a slice where being wrong barely matters.
The lesson: a number can be completely true and still be the wrong number to trust. Ninety-three percent correct told Ilsabet the model was fine. It never had a way to tell her that the seven percent it missed had wildly different costs sitting inside it, all averaged into one line. The average has no idea which mistake was cheap and which one had a name and a street address.
Now here is the same thing as a story
Read this version when you want to feel why a calm number can still be the wrong number to trust, not just be told that it is.
Ilsabet Kalman can look at a spike in Fenmarsh's report queue and know, inside a couple of minutes, whether it's a coordinated pile-on or an ordinary Tuesday. Three years running trust and safety reviews there will do that to a person.
For most of that time, Monday mornings were the calmest meeting on her calendar. She'd open the week's audit: 2,000 of the flagging model's decisions, hand-graded by a senior reviewer, checked against what a person would have called. The number came back somewhere between 92 and 93 percent, same as the month before, and the one before that. She'd read it out, note it, and move on before her coffee went cold.
When Fenmarsh first split flagged posts into three severity tiers, eighteen months back, Ilsabet pulled a second report alongside the blended one for a while: a rough breakdown by tier, mild, moderate, severe. For the first two months she opened both every Monday. Severe-tier volume was so small back then, a handful of posts a week, that the tier number never told her anything the blended one hadn't already said. She started skipping it some weeks, when she remembered. By month nine, she'd stopped opening it at all. The blended number had never once disagreed with it in a way that mattered, so pulling a second file felt like checking a lock she'd already checked.
Nothing about that showed up in a number. It showed up on a Thursday, in a message from a stranger: a screenshot of a post naming a real person, a real street, and exactly what would happen to them, still live, still earning ad views, thirty-six hours after it posted.
Ilsabet's first instinct, and half the team's, was to treat this as proof the model had gotten worse and start scoping a retrain. Someone had a plan drafted by lunch. Ilsabet asked for the afternoon instead, to pull the full population of severe-tier decisions from that week, all eighty of them, not just whatever slice had landed in the random sample.
The afternoon wasn't spent proving the model wrong. It was spent finding out the model's judgment on the other severe posts that week was still fine, and that the miss was sitting somewhere the Monday audit had never had the numbers to reach.
It was never really about whether the blended number moved. It hadn't. What moved was something that number was never built to measure: whether a random sample of 2,000, drawn the same way every week regardless of what each tier actually cost to get wrong, could ever really tell her the severe tier was safe. It couldn't, and it never had.
The decision that opened the door went back to the week tiers were added. Someone asked, in passing, whether the audit needed its own severe-tier sample. The answer was no, the tier barely produced ten posts a week back then, and adding a second dashboard for ten items felt like process for its own sake. Nobody revisited it as that tier's stakes grew, only its volume.
Run the same Thursday again with one change: every severe-tier decision gets hand-graded in full each week, not sampled, about eighty posts, twenty minutes of a senior reviewer's time. The same post, the same wording, still gets misread by the model as a mild joke. But it's sitting in a full-review queue a person checks that same day, not a random slice that may or may not have picked it up. Caught in about six hours. No stranger's screenshot required.
One design trusted one sample size to answer two very different questions. The other asks each tier the question its own stakes actually deserve. Those aren't the same audit wearing different clothes. One has a hole the rarest, worst case falls straight through. The other doesn't.
What I'd tell myself, back in that meeting when tiers were first added: the moment a metric starts covering cases with wildly different costs, ask whether one sample can still answer for all of them, or whether it only ever could because the expensive case was too rare to notice yet. Nobody asked. That's on the room, not on the model.
The five moves, for a rate that never blinked
This isn't a diagnosis of a bug. It's FLIPS run on a trust metric, with the sampling method as the thing that quietly broke.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was retraining the flagging model's severity classifier from scratch, the plan half the team had drafted by lunch. It lost because a full audit of the model's other severe-tier calls that same week showed its judgment hadn't measurably worsened; the gap was in how the audit sampled, not in what the model knew, and a months-long retrain would have spent real engineering time on a model that wasn't the problem while leaving the sampling hole open for the next borderline case. The AI-specific failure mode worth naming by name is a confident misroute driven by surface phrasing: the threat's caption matched a common joke template closely enough that the model scored it mild with real confidence, mistaking the shape of the sentence for its content. The guardrail is concrete and two-part: audit the severe tier in full every week instead of by random sample, and route any post matching a short list of named-target or explicit-threat phrases to a person regardless of the model's own tier confidence. That guardrail isn't free. Tightening the auto-clear threshold around threat-adjacent language means roughly 300 more borderline posts a week, mostly real dark humor, wait for a person to clear them instead of posting instantly, a real friction cost accepted only on that narrow slice, not rolled out to the whole app. And the bar that decides whether the severe tier is safe enough isn't zero misses, a system scoring 3.4 million uploads a day can't promise zero on a probabilistic call. It's an audited miss rate under 2 percent on a full weekly review, checked against roughly 5 percent tolerated on the mild tier's random-sample audit, tight enough that the review team is usually the one who catches the rare miss, not the police.
And if you want to be sure it really works, try it somewhere else
Same five letters, an insurance claims desk instead of a social app, with nothing about content moderation anywhere in sight, and a different flip family doing the work.
Ashwick Mutual runs an AI tool that scores incoming auto-repair claims from "routine" up to "escalate now." Endymion Haskin is the senior fraud investigator who oversees how flagged claims get handled.
F, find the person. Endymion, five years running claims review, who could once spot a staged-accident pattern from three lines of an intake form.
L, locate the habit. Once the flagged-claim tool's blended accuracy held near 90 percent for a year, Endymion delegated every flagged claim, whatever tier the model gave it, to two junior adjusters, and stopped personally reviewing anything below the model's own top escalation tier.
I, identify the flip (delegation flip). Junior adjusters handle every flagged claim, or Endymion takes the rarest, highest tier back himself. Three claims from the same repair shop in one month, each scored "routine" by the model, its lowest tier, got cleared by a junior adjuster in minutes apiece, $118,000 paid out before anyone connected the three.
P, pinpoint the old decision. A year earlier, when severity tiers were added to the fraud model's output, Ashwick kept the same "junior adjusters clear anything the model flags" rule for every tier, since a separate rule for the rarest tier felt like process for a case that came up a handful of times a year.
S, show the replay. Endymion adds one standing rule: any repair shop appearing on two or more flagged claims inside 30 days gets pulled to him automatically, whatever tier the model gave it. Same three claims, same month: the second one trips the rule. All three get reviewed together within a day, before the third payment goes out, instead of surfacing six weeks later in a routine audit.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the one check, pull the full population for the rarest tier, don't trust the random sample's silence.
Cost: there's no dedicated audit team funded yet for the rare tier. Don't skip the check, a senior person hand-grades the small full population themselves until one exists.
The model got better, for real: say the blended accuracy improved that quarter. That's not proof the rarest, priciest tier improved with it. A model can get better on average while one narrow slice stays exactly as blind as before.
Where people run it wrong.
They read a flat blended number as proof there's no gap anywhere, and never check whether the rare tier ever got a real sample size.
They promise a full retrain the moment something severe slips through, before checking whether the model's judgment actually moved.
They fix the one case by hand and never change the sampling or the routing, so the next rare case sails through the exact same gap.
How to use it live. Say the real question out loud before answering it: "is the model wrong, or did the check never really look." That buys you a beat to think instead of guessing out loud in front of the interviewer.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the severe tier grows past the point where a full review is realistic?" Response: then it stops being a full review and becomes its own stratified sample, oversampled well past its share of traffic, with a sample size chosen to actually detect a miss rate that low, not proportional sampling dressed up as coverage.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?