CaseAdvancedQuality, Cost & Token Economics / Eval design for product teams / #21

Explain how to eval safety behaviour without a huge annotation budget.

A small labeling budget forces a choice: spread it evenly across every category, or point it at the rare one where a miss can't be taken back. Only one of those protects the streamer who can't ask for help mid-crisis.

The direct answer
Don't spread a small annotation budget evenly across every violation category the way traffic volume tempts you to. Weight it toward the rare, high stakes behavior the model is least sure about, self-harm signaling and grooming-adjacent chat, by routing the model's lowest confidence outputs on those categories to a human reviewer first. Back that with a small, scripted set of adversarial cases for the behavior too rare to collect naturally, because the streamer having a crisis on camera has no way to flag the model's mistake themselves.
Do this, in order
  1. Weight the review budget by severity, not by how often a category shows up in traffic.Why: a split that matches volume gives the rarest, highest stakes categories almost nothing, exactly where the model is least sure.
  2. Route the model's lowest confidence outputs on the rare, high stakes categories to a human first.Why: active learning spends scarce hours where the model is most likely wrong, instead of re-confirming what it already gets right.
  3. Build a small, scripted set of adversarial cases for behavior too rare to collect naturally.Why: real self-harm or grooming-adjacent examples take months to pile up. A written scenario gets calibration signal now.
  4. Watch a proxy in production for the person who can't self-report.Why: track how often a moderator manually escalates something the model scored "no action," by category, since the streamer mid-crisis can't file that complaint themselves.
  5. Re-run the stratified sample and the adversarial set on every model swap, not just at launch.Why: a new model shifts most on the classes it has the least practice with, and that's exactly the rare, severe ones.
  6. Leave the routine, high volume, low stakes categories on a lighter spot-check cadence.Why: budget spent re-checking what the model already gets right is budget stolen from the ten minutes that actually matter.

How to answer this, stage by stage

Nobody is grading whether you know the word "active learning." They are grading whether a small budget, in your hands, ends up protecting the person who can't push back, or just the categories that were already easy.

1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one product. Emberwatch is the real time moderation engine inside Duskrail, a livestream platform. It watches every second of video and chat and can mute a message, flag a clip, or cut a stream outright. Anzu Callahan runs trust and safety on it."
Why this works
An abstract "how do you eval safety on a budget" answer turns into a shopping list of tools fast. One product makes it a real decision with real stakes.
2
Say your structure out loud before diving in
Say it like this
"I'll answer this in two parts. First, how I'd spend a small labeling budget so it protects the person who can't push back. Then, what I'd watch once it's live, since a fixed budget can't review everything by itself."
Why this works
Tells the interviewer you have a plan and a payoff, not just a list of budget-saving tricks.
3
Reframe the question before answering it
Say it like this
"This isn't really asking how to save money on labeling. It's asking who a tight budget quietly stops protecting, because wherever the budget is thinnest is wherever the model is least checked."
Why this works
Stops you from giving a tooling answer, "use active learning," instead of naming who the budget decision actually affects.
4
Give the one decision, plainly
Say it like this
"I would not split the review budget the way most teams default to, proportional to how often each category shows up in traffic. I'd weight it by severity and by how unsure the model already is, routing its lowest confidence outputs on the rare, dangerous categories to a human first, and I'd fill the gap with a small, scripted adversarial set for cases too rare to occur naturally yet."
Why this works
This is the actual answer to the question, in one breath.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without that. Emberwatch launched with two hundred human-reviewed clips a week, split to match traffic, so self-harm-adjacent and grooming-adjacent chat, together under one percent of volume, got about eight clips a week. A streamer showing real signs of a crisis got scored low confidence, and the model's default under low confidence was no action. The stream stayed live for fourteen minutes, seen by up to thirty two hundred people, before a moderator happened to click into it and cut it by hand."
Why this works
Shows the real cost of the wrong split, not just the mechanism behind it.
6
Say what you would measure going forward
Say it like this
"I'd track how often a moderator manually escalates something the model scored 'no action,' broken out by category, since the streamer can't file that complaint themselves. And I'd re-run the stratified review and the adversarial set on every model swap, not only at launch."
Why this works
Shows you're thinking past launch day, and names the one signal that stands in for a complaint the subject can't file.
7
Say what you'd leave alone
Say it like this
"I wouldn't put the same review hours on routine, high volume categories, a copyrighted song playing in the background, a common profanity flag. Those stay on a light spot check. The scarce hours go where a miss can't be taken back."
Why this works
Shows judgment instead of blanket caution spread evenly over everything.
8
Close on the decision, not the story
Say it like this
"So: a small labeling budget isn't an excuse to check everything a little. It's a reason to check the rare, dangerous cases a lot, and to check them again every time the model changes."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd actually reuse.

Let's learn

A single dashboard tile, red when a stream needs a look, sits in front of every moderator on shift at Duskrail. That tile is Emberwatch, the real time moderation engine behind the platform. It watches video and chat frame by frame, and it can mute a message, flag a clip, or cut a stream on its own.

Before Emberwatch, forty moderators split six thousand concurrent streams between them, glancing into each one every few minutes. They caught the loud stuff fast, nudity, slurs, anything a viewer would screenshot and report. Anything quieter, a person talking themselves toward something dangerous in a calm voice, only got caught if a moderator happened to be watching that exact stream at that exact minute.

Knowledge spark: what is active learning? Instead of picking clips to review at random, you show a human the clips the model is least sure about first. A small labeling budget goes a lot further when it is spent exactly where the model is guessing.

Now Emberwatch watches every second of every stream at once. The easy categories, the loud ones, get caught in under two seconds, and viewers stopped seeing the worst of the obvious stuff before it even finished loading.

We didn't need to review more clips. We needed to stop reviewing the ones the model was already right about.

The extra speed was not the problem. The problem was where the team spent its labeling budget once the easy categories started scoring well. Two hundred clips a week went to a paid review vendor, split the same way traffic was split. Ninety six percent of that budget went to nudity, hate speech, and spam, the categories the model was already good at. The categories where getting it wrong costs someone their safety, self-harm signaling and grooming-adjacent chat, got what was left over. About eight clips a week.

Hand sketched comparison. Left, a person labeled Moderator on duty, caption can pull the stream the second it looks wrong. Right, a person labeled Streamer mid crisis, caption can't see the model's score, can't appeal it, can't ask for help right now.
One of these two people can stop a bad fourteen minutes before it starts. The other one is the one it happens to.
Rare, high stakes categories' share of the weekly 200-clip review budget
4% 35% Split by traffic volume (old) Split by severity & uncertainty (new)
Proportional-to-traffic splitSeverity-weighted split
Under the old split, self-harm-adjacent and grooming-adjacent chat, together under one percent of traffic, got about eight of the two hundred weekly clips. Doubling the total budget without changing the rule would still leave that category near sixteen clips, a marginal fix. Changing the rule instead moved it to about seventy.
The choice that mattered When Emberwatch first launched, the team split the two hundred clip review budget the same way the platform's traffic split, no category dominating badly, so nobody wrote a special rule for the rare ones. That made sense at a small size. It stopped making sense once the platform tripled in traffic and a half a percent category became thousands of real streams a month, checked almost never.

What I'd leave alone: the routine stuff, a copyrighted song caught by a filter, a common swear word, doesn't need the same review hours. A miss there costs a warning that should have been a mute. A miss on the rare category costs a real person a real fourteen minutes.

The lesson: a model that scores ninety nine percent right company wide can still be almost blind on the one category where being wrong actually hurts someone, if that is exactly the category nobody had the budget to check. A small budget is not a reason to check less carefully. It is a reason to point the checking somewhere on purpose.

Now here is the same thing as a story

Read the short version above for the two minute answer. Read this one when you want to feel why splitting the budget by traffic felt like the fair call at the time.

Before Emberwatch existed, Anzu Callahan built Duskrail's very first moderation rota by hand: forty moderators, six thousand streams, a spreadsheet of who was watching what and when. For three years, nothing serious slipped past that rota without Anzu's team catching it within the hour.

Emberwatch launched, and the good months were good. The loud stuff, nudity, slurs, dropped to under two seconds from flag to cut, company wide. Anzu's vendor reviewed two hundred clips a week, split the way the platform's traffic was split, and every week the numbers came back clean.

The habit thinned in three small beats. First, the weekly split stopped getting a second look before approval, since it had passed every week for months. Then, when a rare category came back with zero flagged clips two weeks running, the vendor quietly sent it fewer clips, to spend the hours where flags were actually turning up. Then, when Duskrail's traffic tripled in a year, nobody asked whether two hundred clips a week, split the same old way, was still enough for the categories it had always been thin on.

The trigger was a near miss. A moderator caught a borderline stream by pure luck, scrolling past it between two other tasks, and flagged it up the chain herself before the model ever did. She asked Anzu afterward: "What would have happened if I hadn't clicked into that one?" Anzu didn't have a real answer.

We didn't take a review hour away from anyone. We just never asked what eight clips a week were actually covering.

She never got the chance to answer that question properly before the real one happened. A streamer showing real signs of a crisis scored low confidence, twice in a row, and the model's own default under low confidence was to do nothing, because it had almost no confident examples of that category in either direction. The stream ran for fourteen minutes, up to thirty two hundred people watching at once, before a moderator clicked in by chance and cut it by hand.

The decision that opened the door went back to Emberwatch's first planning meeting, eighteen months earlier. The team agreed the two hundred clip review budget would split the same way the platform's traffic split, since at that size no single rare category was worth a special rule. Nobody wrote down when that rule should be revisited. It just never was, until the scale had multiplied and the rule hadn't moved with it.

Run the same week again with one change. The review budget now splits by how unsure the model is, not by traffic, so the rare high stakes categories get about seventy clips a week instead of eight, and a small scripted adversarial set stands in for the real examples still too rare to collect. On the next near identical case, the model flags it as unsure within two seconds, a moderator is paged automatically, and the stream is cut at the forty second mark, seen by about ninety people instead of thirty two hundred.

One design let the loudest, easiest categories decide where the labeling hours went. The other design let the rarest, most dangerous category decide, on purpose, before it ever had to prove itself in a real fourteen minutes.

What I'd tell myself, back in that first planning meeting: a budget split that makes sense at your current size is a decision with an expiry date, not a rule. Nobody put a date on it. That's on the room, not on Anzu.

GUARD, run against a budget too small to spend evenly by accident

Not a checklist for a review board. Five questions that build toward the one that actually decides everything: who never gets to push back, and whether the labeling budget was ever pointed at them.

GGroups. Who is actually affected by this model, on both ends?
Anzu's trust and safety team, who can pull a stream or mute a message the second the model flags it. And the streamer having a crisis on camera, who has no idea a model, or a labeling budget, is deciding anything about them right now.
Name both, or the answer stays about labeling logistics instead of who it's actually deciding for.
UUnequal. Where does the harm land unevenly, and why that group specifically?
It lands hardest on exactly the categories a traffic-matched budget was always going to starve: the rare, quiet, high stakes ones. A calm-voiced streamer heading toward real danger doesn't trip the loud filters the budget was built to catch.
A company wide accuracy number can look excellent while the harm concentrates entirely in the half a percent of traffic nobody was checking.
AAbility to contest. Who never gets a chance to push back?
A streamer mid-crisis can't see the model's confidence score, can't appeal a "no action" decision, and often can't even ask for help clearly in that moment. The only lever belongs to whichever moderator happens to be watching that exact stream.
This step decides everything else. If the subject could push back the moment it happened, you'd just fix it live. They can't, so the fix has to happen before the stream ever airs.
Hand sketched flow diagram titled Where an appeal step should sit and doesn't. Four boxes in a row: Model unsure, No action taken, No appeal step circled in red, Noticed 14 min later.
There is no box between "no action taken" and "noticed, fourteen minutes later." That gap is the whole problem.
RReduce. What's the actual design change, not a policy?
Weight the review budget by severity and by the model's own uncertainty, not by traffic share. Route its lowest confidence outputs on the rare, dangerous categories to a human first, and fill the gap with a small scripted adversarial set for cases too rare to occur naturally yet.
A memo about "taking safety seriously" doesn't change where the review hours actually go. A different split does.
DDetect. How would you know this is happening in production, before someone outside the company tells you?
Track how often a moderator manually escalates something the model scored "no action," broken out by category, since the subject can't self-report. A spike in one rare category, even while the company wide rate holds steady, is the real signal.
The subject can't tell you directly. This proxy has to be watched by category, not averaged into one comfortable number.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was a flat, random one percent review sample across all traffic, simple, unbiased, easy to explain to an auditor. It lost because most randomly picked clips still land in the easy, high volume categories the model already handles well, so a flat sample spends most of its scarce hours re-confirming what's already known instead of finding what isn't. The AI specific failure worth naming by name is silent miscalibration on a rare class: without enough labeled examples in a category, a model's confidence threshold quietly defaults toward "no action" on anything unclear, not because it checked and found nothing wrong, but because it was never taught what wrong looks like there. The guardrail is the severity-weighted active learning routing plus the scripted adversarial set, since together they hand the model real signal on a class it would otherwise learn nothing about until a real incident supplied it. And the trade being accepted on purpose is real: pointing the small budget at rare categories means the routine, high volume ones get checked less often, so a slow drift in something ordinary, a copyright filter getting a little too trigger happy, would now take longer to notice. That's a real cost, traded against a subject who has no way to tell you when the miss goes the other way.

The five, in one line each:
G: name the moderator and the streamer, both.
U: the harm concentrates in whichever category the budget starved.
A: a streamer mid crisis can't appeal a "no action" score, so the fix has to happen before the stream airs.
R: weight the review budget by severity and uncertainty, not traffic share.
D: watch the manual override rate by category, since that's the closest thing to a complaint you'll get.

Same five letters, a different budget problem entirely

Not every scarce labeling budget is scarce because of money. Sometimes it's scarce because the only qualified reviewer is a specialist you can't clone.

Cloverhitch runs a livestock auction marketplace. PastureGuard, its screening tool, checks listing videos for welfare red flags, visible lameness, overcrowding, distress, before a listing goes live. Emryk Vantassel runs quality on it.

PastureGuard's review budget isn't measured in dollars. It's measured in a contracted vet's hours: about three hundred listings a month out of roughly forty thousand. Split proportional to how often each issue type shows up, most of that vet's time went to routine paperwork mismatches, a weight recorded wrong, a breed misfiled, since those are common and easy to confirm. Visible lameness and overcrowding, together under one percent of listings, got a sliver of that time, even though those are exactly the categories where a wrong call means an injured animal goes to auction with real bids on it before anyone with real judgment sees the footage.

The decision Emryk would take back PastureGuard's vet hours split the same way the platform's issue types split in the traffic, a rule that made sense when the marketplace was smaller and no single welfare category needed protecting on its own. Nobody revisited it once volume tripled and the rare categories stayed exactly as rare, and exactly as unchecked.
PastureGuard's catch rate on its lowest confidence welfare flags, week by week
90% 0% wk 3: vet hours re-weighted Wk 1 Wk 3 Wk 6
Catch rate on lowest confidence welfare flags
The catch rate on PastureGuard's own lowest confidence welfare flags sat near 22 percent for the first two weeks, while the company wide accuracy number looked fine the whole time. Once the vet's hours were re-weighted toward those flags instead of routine paperwork checks, the rate climbed to 81 percent by week six.

Same method, different shape: a streamer mid-crisis and an injured animal on an auction block look nothing alike, but both are a model's call being made about someone, or something, with no lever to push back, before anyone outside the company would ever know.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the line: find who can't push back, then point the smallest slice of the budget at exactly where that risk sits, not where the traffic sits.
Cost: there's no budget this quarter for any extra vet hours at all. Shrink the sample, a hundred listings a month instead of three hundred, but keep the same weighting rule. The size can shrink; which categories get first pick of it can't go back to matching traffic.
The model got better, for real: say the newer model is measurably more accurate on average. That's exactly when this matters most, since a stronger model finds smoother, less obviously wrong ways to reach the same rare failure, and a model that's confidently wrong on a rare case is harder to catch than one that's clumsily wrong on a common one.

Where people run it wrong.
They split the review budget to match how often a category shows up in traffic, instead of how likely the model is to be wrong on it.
They treat a scripted adversarial set as a one time launch check instead of refreshing it on every model swap.
They report one company wide accuracy number, which hides a rare category running badly below it the whole time.

How to use it live. Say the two things out loud before answering: "how big is the review budget, and who does the smallest slice of it currently protect." That single question buys a beat of thinking time, and it turns the rest of the answer into naming what protects the one with no lever.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
GUARD: name who can't push back. On a tight labeling budget, that means spending the scarce hours where that person is least protected, not where review is easiest.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Anzu Callahan, trust and safety lead on Emberwatch at Duskrail, a livestream platform. Built the platform's original forty person moderation rota by hand.
3 · THE HABIT
What did the team stop doing once the weekly review split felt reliable?
Tap to flip
ANSWER
They stopped questioning whether the traffic-matched budget split still made sense as the platform scaled, and let the launch rule stand in for every size that came after it.
4 · THE TWO SETTINGS
What's the two setting switch this answer turns on?
Tap to flip
ANSWER
A review budget split proportional to traffic volume, versus the same budget weighted by category severity and by how unsure the model already is.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Splitting the two hundred clip weekly review budget to match traffic share, with no rule for when that split should be revisited as the platform scaled.
6 · THE NUMBER
Fill in the blank: the crisis stream stayed live for ___ minutes, seen by up to ___ people, before a moderator caught it by chance.
Tap to flip
ANSWER
14 minutes, up to 3,200 people. The rare, high stakes categories were getting about 8 of the 200 weekly review clips, the signal nobody had reweighted as traffic tripled.
7 · THE REPLAY
Same bad week, new design, what changes?
Tap to flip
ANSWER
The rare, high stakes categories now get about 70 of the 200 weekly clips, plus a scripted adversarial set. The next near identical case gets flagged in 2 seconds and cut at the 40 second mark, seen by about 90 people instead of 3,200.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and who or what is the subject who can't contest, this time?
Tap to flip
ANSWER
PastureGuard, a listing screening tool at Cloverhitch, a livestock auction marketplace. The subject is the animal shown in the listing, which has no way to flag its own distress.

Check yourself Score: 0 / 0

True or false
1. True or false: raising the total review budget from 200 to 400 clips a week, while keeping the same proportional-to-traffic split, would have caught the fourteen minute incident.
  • True
  • False
Show hint
Look at the bar chart in Section 1. Doubling the total budget doubles the rare category's share too, but only from a very small number.
Show answer
False. Under the old split, doubling the budget would move the rare categories from about 8 clips a week to about 16, still a sliver. The problem was never the total size of the budget. It was the rule deciding where each clip went.
Multiple choice
2. What does weighting the review budget by severity and model uncertainty buy you that a bigger budget alone doesn't?
  • A. It makes the model flag content faster.
  • B. It spends the scarce hours exactly where the model is most likely wrong, instead of re-confirming what it already gets right.
  • C. It removes the need for any human review at all.
  • D. It guarantees the model never misses a case again.
Show hint
Think about what a random or traffic-matched sample mostly ends up re-checking.
Show answer
B. A flat or traffic-matched sample mostly lands on the easy, high volume categories the model already handles well. Uncertainty routing sends the humans toward the calls the model is least sure about instead.
Fill in the blank
3. Fill in the blank: the crisis stream stayed live for ___ minutes, seen by up to ___ people at once, before a moderator caught it by chance.
Show hint
The number is in Section 2's story and repeated on flashcard 6.
Show answer
14 minutes, up to 3,200 people. No streamer mid-crisis could flag it themselves, so it took a moderator clicking into that exact stream by chance to catch what a properly weighted review budget would have caught before it ever aired.
Short answer, name the rejected alternative
4. What review split did Emberwatch's team choose at launch, and why did that same choice stop making sense eighteen months later?
Show hint
Look at the block-key box titled "The choice that mattered" in Section 1.
Show answer
Model answer: They split the 200 clip weekly budget the same way the platform's traffic split, since no rare category was worth a special rule at launch size. Once traffic tripled, that same half a percent category became thousands of real streams a month, and the split never got revisited.
Multiple choice
5. PastureGuard's company wide accuracy looked fine while its catch rate on its own lowest confidence welfare flags sat near 22 percent for weeks. What does that combination tell you?
  • A. The model is working fine, since the company wide number is what matters.
  • B. The vet hours were being spent on the easy, common issue types, leaving the rare welfare categories under-checked, with the animal on the receiving end unable to flag it itself.
  • C. PastureGuard should stop reviewing any listings at all to save money.
  • D. A steady company wide number proves there's no rare category running badly underneath it.
Show hint
Look at what the line chart shows before week 3, compared to what a single blended accuracy number would have shown.
Show answer
B. A steady blended number hid a badly under-covered category. And the animal in the listing has no way to contest a missed flag, which is exactly the "ability to contest" gap GUARD asks you to name.
Short answer, apply it yourself
6. Pick an AI product you use yourself where a small review or labeling budget probably gets spent on the easy, common cases instead of the rare, high stakes ones. What's one thing you'd want that team to check first?
Show hint
Think of a product where a rare miss costs someone real harm and they never learn it happened, not just an easy miss you'd notice yourself.
Show answer
Model answer: A resume screening tool likely spends its review hours confirming it correctly filters obvious spam applications, since those are common and easy to check. I'd want the team to check its rare false negatives instead: qualified candidates from an unusual career path who got filtered out and never learn why, since they have no way to appeal a rejection they don't know happened.
Before you close the answer
Why this works
Tests whether you treat a small labeling budget as an excuse to check everything a little worse, or as a forcing function to check the rare, dangerous cases on purpose. Most candidates answer "hire more labelers" and stop there.
Follow-up traps
"Why not just hire more labelers instead of reweighting?" Response: more labelers help, but a bigger budget spent the same proportional way still starves the rare category, just a little less badly. The split matters more than the size on its own.

"Isn't routing by model uncertainty just training the model to hide its own mistakes?" Response: no, it's the opposite. Uncertainty routing sends the model's least confident calls to a human first, which is the fastest way to find exactly where the model doesn't know yet, not a way to bury it.
If pressed
The score used to rank clips for review wasn't the model's raw confidence output, which is often miscalibrated on exactly the rare classes that matter most. It was that raw score passed through a small held-out calibration set first, so the review queue was ordered by a number actually trustworthy enough to prioritize on.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more