ConceptAdvancedQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #20

How does the severity distribution of errors matter more than the error rate?

Two systems can share the exact same error rate and carry nowhere near the same risk. What decides that is where the mistakes land, not how many there are.

The direct answer
Track the miss rate by severity tier, not as one blended number, because a single rare miss on the worst tier can hide inside thousands of correct calls on the mild tier and never move the topline at all. Give the smallest, highest-stakes tier its own pass bar and check it in full, not by random sample, since it is exactly the tier a big denominator is built to hide. A flat error rate only proves the average is fine. It says nothing about whether that average is built out of pieces with wildly different costs.
Do this, in order
  1. Split the miss rate by severity tier, and give the rarest, worst tier its own bar.Why: one blended number lets a miss on the worst tier hide inside thousands of correct calls on the mild one.
  2. Audit the top severity tier in full, not by random sample.Why: at low volume, a random sample almost never picks up enough of the rarest tier to know if it's actually healthy.
  3. Set a tighter pass bar for that tier than the rest of the queue tolerates.Why: a miss that never reaches a person costs far more on the worst tier than on the mild one, so the same bar can't cover both.
  4. Re-pull the severity split whenever the blended number goes quiet for a long stretch.Why: calm for months is exactly when a person stops opening the file that would have caught this.
  5. Close the routing or sampling gap that let the miss through, not the whole model.Why: the check that mattered showed the model's overall judgment hadn't moved, the sampling design had the blind spot.
  6. Accept that some borderline posts wait longer for a person to clear them.Why: tightening the net around risky language holds back a few real jokes too, a cost worth paying only on the tier where a miss is this expensive.

How to answer this, stage by stage

Nobody's grading whether you know severity matters. They're grading whether you'd trust a flat number, or ask what's actually inside it before you call a system healthy.

1
Scope it to one product before saying anything about error rates in general
Say it like this
"Let's make this concrete. Fenmarsh is a short video app. Its flagging model scores every upload and sorts anything risky into three severity tiers. Ilsabet Kalman runs trust and safety there."
Why this works
A metrics question answered in the abstract turns into a lecture. One product, one person, makes it a decision you can defend.
2
Name what the question is actually testing
Say it like this
"This isn't really asking if error rate matters. It's asking whether I'd trust one blended number, or ask what's actually inside it before I call a system healthy."
Why this works
Naming the real question stops you from giving the generic "measure everything" answer most candidates default to.
3
Give the direct answer, cold, before any story
Say it like this
"A flat error rate only tells you the average is fine. It says nothing about whether that average is built out of pieces with wildly different costs. So the first thing I'd do is split the miss rate by how bad each miss actually is, not trust the one blended number."
Why this works
A reader who stops here already knows the whole answer. Everything after this is proof.
4
Explain why a blended sample structurally can't catch this
Say it like this
"If the worst tier is a tiny slice of total volume, a random weekly sample of even a couple thousand items barely touches it. You're not measuring whether it's healthy. Most weeks, you're not really measuring it at all."
Why this works
This is the move that turns "trust matters more than rate" into something a sampling method can actually fail at.
5
Run the one check that separates a real regression from a sampling blind spot
Say it like this
"I'd go back and pull every decision the model made on the worst tier that week, not just whatever fell into the random sample, and hand-grade all of it. If the full set still looks fine outside the one miss, the sampling was the problem, not the model."
Why this works
Turns a debate about whether the model got worse into one number you can actually check.
6
Name the trade-off out loud instead of pretending the fix is free
Say it like this
"Tightening the net on risky language means some real dark jokes wait for a person instead of posting instantly. I'd take that friction on purpose, only on the slice where a miss is this expensive, not across the whole app."
Why this works
Shows a real decision, not a wish that speed, cost, and safety could all be free at once.
7
Close with the specific fix and the number that proves it worked
Say it like this
"The fix isn't retraining the whole flagging model, the full audit says its judgment on other severe posts didn't change. It's checking that tier in full every week and giving it its own bar. Rerun the same week with that in place, and the miss gets caught in about six hours, not found by an outside report two days later."
Why this works
Closing on a narrow, provable fix is what makes the answer sound like a decision, not a promise.

Let's learn

Here is what happens when a number stays perfectly calm while something underneath it gets worse.

Fenmarsh is a short video app. Every clip and caption that gets posted runs through a flagging model the moment it goes up. The model sorts anything it's unsure about into three severity tiers: mild rule-breaks like spam or a rude joke, moderate problems like bullying, and severe cases like a real threat of violence. Ilsabet Kalman runs trust and safety there.

Knowledge spark: what's a severity tier? A label the model puts on a flagged post for how bad it would be if the post stayed up. It isn't how sure the model is. It's how much it would cost everyone if this one, specifically, turned out to be wrong.

Before the flagging model, a small team of moderators only looked at posts someone had reported by hand, working through one shared queue in the order the reports came in. They cleared about 22,000 reports a week, roughly 4 minutes each, and on a busy weekend the queue backed up more than 30 hours. A mild joke report and a genuine threat report sat in the exact same line, waiting their turn.

Now the flagging model checks all 3.4 million uploads a day on its own, clears almost everything instantly, and sends about 40,000 posts a week to a person, tagged by tier. Once a week, a senior reviewer hand-grades a random sample of 2,000 of the model's decisions, spread across all three tiers, and reports back one number: how often that sample agrees with what the model did. For fourteen months straight, that number held between 92.4 and 93.1 percent.

Blended weekly agreement rate, week 1 to week 60
95% 90% week 51: the missed threat Wk 1 Wk 30 Wk 60
Blended agreement rate, all tiers, random sample
The number holds between 92.4 and 93.1 percent for over a year. Week 51, the week a real threat got waved through, is the high point on the whole chart, not a dip.

Then, in week 51, one post got scored wrong in the one direction that actually mattered. Someone posted a specific, named threat of violence against a real person, worded in the same short, punchy style as a common joke format the model had learned to wave through. The model tagged it mild, not severe, and cleared it instantly. It stayed up for 36 hours, earning ad views the whole time, until a friend of the target reported it outside the app, straight to the police.

That same week, the blended number didn't dip. It rose, to a new high of 93.1 percent, because the model also caught an unusually clean batch of spam and copycat mild posts.

The model did not get meaningfully worse that week. One severe miss got buried under thousands of easy, correct mild ones.

Once Ilsabet went back and hand-graded the full population of severe-tier decisions from that week, not just the slice the random sample happened to touch, the real picture came out: mild-tier posts missed at about 3 percent that week, moderate at about 5 percent, and the severe tier, only 80 posts total, missed at nearly 9 percent. One in eleven.

Miss rate by severity tier, week 51, full population
10% 0% 3% 5% 8.75% Mild, 36,000 Moderate, 3,900 Severe, 80
Mild and moderate tiersSevere tier
The severe tier's real miss rate ran nearly three times higher than the mild tier's. Because it's only 80 of the 40,000 posts sent for review that week, it moved the blended number almost nothing.
Hand sketched comparison titled same blended number very different unit. Left panel mild tier thirty six thousand posts a week a few dozen wrong. Right panel severe tier eighty posts a week one real threat missed.
Same audit, same week, two completely different units. One tier can afford to be wrong a little. The other can't afford to be wrong at all, and it's the one the sample barely touches.

What it cost, at its worst: police called Fenmarsh 36 hours after the post went up, asking why a threat naming a real address had been live and monetized the whole time. Nobody inside Fenmarsh had flagged anything wrong, because the one number everyone watched had just had its best week ever.

The decision that mattered Eighteen months earlier, when severity tiers were first added to the flagging model's output, the team folded the new tier field straight into the same one weekly random-sample audit that already existed, instead of building a separate check for the rarest tier. That was fine when the severe tier barely produced any volume. It stopped being fine the day a tier existed where being wrong could cost someone their safety, and a random sample of 2,000 could go whole weeks without really touching it.

What I would leave alone: the mild tier's number is genuinely fine as a blended random sample. A missed spam post or an overturned rude joke costs almost nothing, and it usually gets caught on the next report anyway. Auditing that tier in full every week, the way the severe tier now needs, would spend real reviewer time on a slice where being wrong barely matters.

The lesson: a number can be completely true and still be the wrong number to trust. Ninety-three percent correct told Ilsabet the model was fine. It never had a way to tell her that the seven percent it missed had wildly different costs sitting inside it, all averaged into one line. The average has no idea which mistake was cheap and which one had a name and a street address.

Now here is the same thing as a story

Read this version when you want to feel why a calm number can still be the wrong number to trust, not just be told that it is.

Ilsabet Kalman can look at a spike in Fenmarsh's report queue and know, inside a couple of minutes, whether it's a coordinated pile-on or an ordinary Tuesday. Three years running trust and safety reviews there will do that to a person.

For most of that time, Monday mornings were the calmest meeting on her calendar. She'd open the week's audit: 2,000 of the flagging model's decisions, hand-graded by a senior reviewer, checked against what a person would have called. The number came back somewhere between 92 and 93 percent, same as the month before, and the one before that. She'd read it out, note it, and move on before her coffee went cold.

When Fenmarsh first split flagged posts into three severity tiers, eighteen months back, Ilsabet pulled a second report alongside the blended one for a while: a rough breakdown by tier, mild, moderate, severe. For the first two months she opened both every Monday. Severe-tier volume was so small back then, a handful of posts a week, that the tier number never told her anything the blended one hadn't already said. She started skipping it some weeks, when she remembered. By month nine, she'd stopped opening it at all. The blended number had never once disagreed with it in a way that mattered, so pulling a second file felt like checking a lock she'd already checked.

Hand sketched comparison titled a habit that thins then just stops. Left panel the habit thinning checks the tier split weekly then some weeks then rarely. Right panel the flip stops opening the tier file at all calls the blended number enough.
Not a dial turned down. A file she used to open, and then, one Monday, simply didn't anymore.

Nothing about that showed up in a number. It showed up on a Thursday, in a message from a stranger: a screenshot of a post naming a real person, a real street, and exactly what would happen to them, still live, still earning ad views, thirty-six hours after it posted.

Ilsabet's first instinct, and half the team's, was to treat this as proof the model had gotten worse and start scoping a retrain. Someone had a plan drafted by lunch. Ilsabet asked for the afternoon instead, to pull the full population of severe-tier decisions from that week, all eighty of them, not just whatever slice had landed in the random sample.

The afternoon wasn't spent proving the model wrong. It was spent finding out the model's judgment on the other severe posts that week was still fine, and that the miss was sitting somewhere the Monday audit had never had the numbers to reach.

We didn't need a smarter model. We needed the smallest tier checked in full, not by chance.

It was never really about whether the blended number moved. It hadn't. What moved was something that number was never built to measure: whether a random sample of 2,000, drawn the same way every week regardless of what each tier actually cost to get wrong, could ever really tell her the severe tier was safe. It couldn't, and it never had.

The decision that opened the door went back to the week tiers were added. Someone asked, in passing, whether the audit needed its own severe-tier sample. The answer was no, the tier barely produced ten posts a week back then, and adding a second dashboard for ten items felt like process for its own sake. Nobody revisited it as that tier's stakes grew, only its volume.

Run the same Thursday again with one change: every severe-tier decision gets hand-graded in full each week, not sampled, about eighty posts, twenty minutes of a senior reviewer's time. The same post, the same wording, still gets misread by the model as a mild joke. But it's sitting in a full-review queue a person checks that same day, not a random slice that may or may not have picked it up. Caught in about six hours. No stranger's screenshot required.

One design trusted one sample size to answer two very different questions. The other asks each tier the question its own stakes actually deserve. Those aren't the same audit wearing different clothes. One has a hole the rarest, worst case falls straight through. The other doesn't.

What I'd tell myself, back in that meeting when tiers were first added: the moment a metric starts covering cases with wildly different costs, ask whether one sample can still answer for all of them, or whether it only ever could because the expensive case was too rare to notice yet. Nobody asked. That's on the room, not on the model.

The five moves, for a rate that never blinked

This isn't a diagnosis of a bug. It's FLIPS run on a trust metric, with the sampling method as the thing that quietly broke.

FFind the person. Whose morning is this?
Ilsabet Kalman, trust and safety lead at Fenmarsh, three years running the weekly review of the flagging model's agreement number. One product, one person, keeps this from turning into a lecture on metrics in the abstract.
Not "the trust and safety team." A specific person's Monday morning.
LLocate the habit. What did she stop doing because it worked?
Pulling a severity-tier breakdown alongside the blended weekly number. Weekly for two months, then some weeks, when she remembered, then not at all for the five months before the incident, because the tier file never once said anything the blended number hadn't already said.
This is the habit, the not-checking, the story is really about. Not the model itself.
IIdentify the flip. What verb snaps? (over-trust flip)
Opens the tier breakdown and checks it, or trusts the one blended number alone and never opens it. Two settings, no in-between: once the blended number has held steady long enough to feel proven, she stops pulling the second file entirely, not just less often.
This is the flip that fires on good news. A calm number for over a year is what made checking feel pointless.
PPinpoint the old decision. Which choice only made sense before?
Eighteen months earlier, folding the brand-new severity tier field into the same one weekly random-sample metric that already existed, instead of building a second, tier-aware audit, because the severe tier barely had any volume yet to justify one.
Reasonable when tiers were new and tiny. Wrong the day tier stakes mattered more than tier count.
SShow the replay. Same bad week, better ending?
Audit the severe tier in full every week, not by random sample, about 80 items, roughly 20 minutes of a senior reviewer's time. Same post, same wording, same week: caught in about six hours during the full tier review, held before the public ever saw it, instead of surfacing 36 hours later through an outside report to police.
A modest fix. Twenty minutes a week, aimed at the one slice that actually needed it.
Hand sketched numbered list titled five letters one flipped switch. F find the person whose morning is this. L locate the habit what did she stop opening. I identify the flip what verb snaps, shown in red. P pinpoint the old decision which choice made sense before. S show the replay same bad week better ending.
The five moves, in order. The I row is the hard one. Everything before it is setup, everything after it is proof.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was retraining the flagging model's severity classifier from scratch, the plan half the team had drafted by lunch. It lost because a full audit of the model's other severe-tier calls that same week showed its judgment hadn't measurably worsened; the gap was in how the audit sampled, not in what the model knew, and a months-long retrain would have spent real engineering time on a model that wasn't the problem while leaving the sampling hole open for the next borderline case. The AI-specific failure mode worth naming by name is a confident misroute driven by surface phrasing: the threat's caption matched a common joke template closely enough that the model scored it mild with real confidence, mistaking the shape of the sentence for its content. The guardrail is concrete and two-part: audit the severe tier in full every week instead of by random sample, and route any post matching a short list of named-target or explicit-threat phrases to a person regardless of the model's own tier confidence. That guardrail isn't free. Tightening the auto-clear threshold around threat-adjacent language means roughly 300 more borderline posts a week, mostly real dark humor, wait for a person to clear them instead of posting instantly, a real friction cost accepted only on that narrow slice, not rolled out to the whole app. And the bar that decides whether the severe tier is safe enough isn't zero misses, a system scoring 3.4 million uploads a day can't promise zero on a probabilistic call. It's an audited miss rate under 2 percent on a full weekly review, checked against roughly 5 percent tolerated on the mild tier's random-sample audit, tight enough that the review team is usually the one who catches the rare miss, not the police.

And if you want to be sure it really works, try it somewhere else

Same five letters, an insurance claims desk instead of a social app, with nothing about content moderation anywhere in sight, and a different flip family doing the work.

Ashwick Mutual runs an AI tool that scores incoming auto-repair claims from "routine" up to "escalate now." Endymion Haskin is the senior fraud investigator who oversees how flagged claims get handled.

F, find the person. Endymion, five years running claims review, who could once spot a staged-accident pattern from three lines of an intake form.
L, locate the habit. Once the flagged-claim tool's blended accuracy held near 90 percent for a year, Endymion delegated every flagged claim, whatever tier the model gave it, to two junior adjusters, and stopped personally reviewing anything below the model's own top escalation tier.
I, identify the flip (delegation flip). Junior adjusters handle every flagged claim, or Endymion takes the rarest, highest tier back himself. Three claims from the same repair shop in one month, each scored "routine" by the model, its lowest tier, got cleared by a junior adjuster in minutes apiece, $118,000 paid out before anyone connected the three.
P, pinpoint the old decision. A year earlier, when severity tiers were added to the fraud model's output, Ashwick kept the same "junior adjusters clear anything the model flags" rule for every tier, since a separate rule for the rarest tier felt like process for a case that came up a handful of times a year.
S, show the replay. Endymion adds one standing rule: any repair shop appearing on two or more flagged claims inside 30 days gets pulled to him automatically, whatever tier the model gave it. Same three claims, same month: the second one trips the rule. All three get reviewed together within a day, before the third payment goes out, instead of surfacing six weeks later in a routine audit.

Three claims from one repair shop, before the pattern was caught
Claim A, $38,000 Claim B, $41,000 Claim C, $39,000 Total exposure: $118,000
Claim AClaim BClaim C, the one that surfaced the pattern
Each claim alone looked routine to the model. Stacked together, from the same shop, inside one month, they were a ring. The blended accuracy number never had a reason to look at all three at once.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the one check, pull the full population for the rarest tier, don't trust the random sample's silence.
Cost: there's no dedicated audit team funded yet for the rare tier. Don't skip the check, a senior person hand-grades the small full population themselves until one exists.
The model got better, for real: say the blended accuracy improved that quarter. That's not proof the rarest, priciest tier improved with it. A model can get better on average while one narrow slice stays exactly as blind as before.

Where people run it wrong.
They read a flat blended number as proof there's no gap anywhere, and never check whether the rare tier ever got a real sample size.
They promise a full retrain the moment something severe slips through, before checking whether the model's judgment actually moved.
They fix the one case by hand and never change the sampling or the routing, so the next rare case sails through the exact same gap.

How to use it live. Say the real question out loud before answering it: "is the model wrong, or did the check never really look." That buys you a beat to think instead of guessing out loud in front of the interviewer.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: checks something sometimes, then stops checking it at all once a number holds calm long enough to feel proven. Fires on good news, like a steady low error rate, not on bad news.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ilsabet Kalman, trust and safety lead at Fenmarsh, a short video app whose flagging model sorts flagged posts into three severity tiers.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Pulling a separate severity-tier breakdown alongside the blended weekly agreement number. Once the tier split kept matching the blended number for months, she stopped opening it at all.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Opens the severity-tier file and checks it, or trusts the one blended number alone. Once the blended number holds steady long enough, she stops opening the tier file at all, no in-between.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Folding the brand-new severity tier field into the same one weekly random-sample metric instead of building a separate, tier-aware audit, made when the severe tier barely had any volume yet.
6 · THE NUMBER
Fill in the blank: the blended weekly agreement number held between ___ and ___ percent for fourteen months, including the week of the missed threat.
Tap to flip
ANSWER
92.4 to 93.1 percent. The week of the missed threat actually hit the high point, 93.1, the same week a real threat got waved through.
7 · THE REPLAY
Same bad week, new design, what changes?
Tap to flip
ANSWER
The severe tier gets audited in full every week, not by random sample. The same threat post gets caught in about six hours during that full review, instead of surfacing 36 hours later through an outside report to police.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Ashwick Mutual, an insurance company, and its claims fraud-flagging tool. This one runs on the delegation flip: a senior investigator reclaims review of the rarest severity tier after junior adjusters clear a ring of claims the blended number never caught.

Check yourself Score: 0 / 0

True or false
1. True or false: since Fenmarsh's blended weekly agreement number never dropped, and even hit a new high the week of the incident, the flagging model's overall judgment must have gotten better that week.
  • True
  • False
Show hint
Check what happened to the mix of mild versus severe posts sampled that week, not just the topline number.
Show answer
False. A clean batch of easy, correct mild-tier catches pushed the blended number up and buried the one severe miss inside it. A number going up on average says nothing about the tail.
Multiple choice
2. Why couldn't a bigger random sample alone have fixed the blind spot?
  • A. A bigger sample would have cost too much reviewer time to be worth it.
  • B. Random samples can never contain severity tiers at all.
  • C. Even a much bigger random sample, drawn in the same proportion as the traffic, would still barely touch the severe tier, because it's such a tiny share of total volume.
  • D. The severe tier isn't something a person can grade by hand.
Show hint
Compare the severe tier's weekly volume to the total review queue in the tier chart.
Show answer
C. The fix isn't a bigger sample of everything, it's oversampling the rare, high-stakes tier on purpose, since proportional sampling will always undercount it.
Fill in the blank
3. The severe tier made up about ___ posts a week out of roughly 40,000 sent for human review, which is why a random sample of 2,000 could go whole weeks without really checking it.
Show hint
Look at the tier miss rate chart's third bar and its label.
Show answer
80 posts a week. About 0.2 percent of the review queue, easy for a random sample to miss almost entirely.
Short answer, name the rejected alternative
4. What alternative fix did this answer reject, and why did it lose?
Show hint
Look at what half the team wanted to do within a day of the incident, in the framework recap.
Show answer
Model answer: Retraining the flagging model's severity classifier from scratch. It lost because a full audit of the model's other severe-tier calls that same week showed its judgment hadn't measurably worsened. The real gap was in how the audit sampled, not in what the model knew.
Short answer, apply it yourself
5. Pick an AI product you use yourself. Name one severity split hiding inside a single error rate you've seen reported for it, and say which slice you'd actually want tracked on its own.
Show hint
Think of a product where getting it wrong sometimes costs nothing and sometimes costs a lot, then ask if its published accuracy treats those the same.
Show answer
Model answer: A spam filter that's right 99 percent of the time. Most of that 1 percent is a harmless newsletter caught by mistake. But if the same 1 percent sometimes buries a fraud alert from your bank, that's a completely different cost hiding in the same number. I'd want the fraud-adjacent slice tracked on its own, not folded into the blended 99 percent.
Multiple choice
6. Why did the fix set a tighter pass bar for the severe tier, under 2 percent, instead of using the same under 5 percent bar the mild tier tolerates?
  • A. Because a system scoring millions of posts a day can promise zero misses if the team just tries harder.
  • B. Because a miss on the severe tier costs far more than a miss on the mild tier, so the same tolerance would quietly accept a much bigger real-world risk.
  • C. Because the mild tier's bar was a typo and should also be under 2 percent.
  • D. Because the severe tier gets checked by a different, less accurate model.
Show hint
Think about what a miss actually costs on each tier, not how often each tier is wrong.
Show answer
B. The pass bar has to match the cost of being wrong on that tier, not stay uniform for convenience. A probabilistic system can't promise zero, but it can promise a tighter bar where a miss costs more.
Before you close the answer
Why this works
Tests whether you'll trust a calm topline number or ask what's actually inside it. Most candidates stop at "the error rate looks fine" without asking what that rate is an average of.
Follow-up traps
"Isn't auditing every severe-tier post just a bigger, more expensive version of the same sample?" Response: no, it's a different question entirely. A sample estimates a rate. A full review of a tier small enough to check in full, 80 posts, 20 minutes, answers the actual question: was this specific week clean.

"What if the severe tier grows past the point where a full review is realistic?" Response: then it stops being a full review and becomes its own stratified sample, oversampled well past its share of traffic, with a sample size chosen to actually detect a miss rate that low, not proportional sampling dressed up as coverage.
If pressed
The actual bar used at Fenmarsh: the severe tier needs an audited miss rate under 2 percent on a full weekly review, not zero, since a probabilistic system scoring 3.4 million uploads a day can't promise zero. The mild tier keeps its looser bar, under 5 percent, on the same random-sample audit it always used.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more