CaseAdvancedShipping & Model Lifecycle / Incident management for AI products / #3

Your model starts producing offensive output. Describe the first hour.

The direct answer
Cut the model's publish path everywhere on the platform in the first five minutes, not just on the account that got caught, because the failure lives in the model, not in one creator. Spend the rest of the hour pulling down every reply that matches the same pattern and finding the people it reached, since none of them knew a machine was answering and none of them has a lever to ask it to stop. Don't turn auto-publish back on for a language until a red-teamed eval set built for that language clears a real bar, not a feeling that the fix probably worked.
Do this, in order
  1. Kill the model's auto-publish path platform-wide, in the first five minutes, not just for the flagged account.Why: the failure is in the model every account shares, so leaving anyone else live just moves the harm to a different community while you're still staring at this one.
  2. Pull down the flagged reply and every other reply posted in the last 24 hours that matches the same pattern.Why: the people who received them never knew a machine wrote them and have no way to ask for one to come down themselves.
  3. Reach out directly to everyone who received one of those replies, before they have to complain.Why: a follower mocked in a creator's name has no appeal button and no address to write to, so the company has to go find her.
  4. Log the exact model version, the prompt template, and real examples before anything gets rolled back.Why: without a snapshot, tomorrow's fix is a guess about what changed, not a diagnosis of what actually broke.
  5. Set the reopen bar as a recall number checked per language on a red-teamed set, not one blended global number.Why: a 99 percent overall score is exactly what let this gap hide for as long as it did.
  6. Put the safety classifier's flag rate, broken out by detected input language, on a live dashboard going forward.Why: that's what turns the next version of this into an internal alert instead of a stranger's screenshot.

How to answer this, stage by stage

Nobody is grading whether you can say "we'd take it seriously." They are grading whether you can name who gets hurt, who can't complain about it, and what you'd actually flip off in the first sixty minutes. Seven moves, timed against the hour itself.

T+0
Scope it to one real product and one real alert
Say it like this
"Let's ground this. Nettleframe is a social app. About 2,300 creators use its auto-reply tool, ReplyForge, which posts roughly 61,000 replies a day in each creator's own voice. Selim Bragoli runs trust and safety for it. At 2:14 on a Tuesday, an on-call alert lands: a reply posted under a wellness creator's name, to a fan's comment written in Tagalog, reads as cruel and dismissive of something hard she'd just shared."
Why this works
Grounds the question in a real scale and a real clock before any response gets described.
T+2
Name both people before touching a single dashboard
Say it like this
"Before I open anything, I want two names in my head, not one. Mira Sondergaard, the creator, who never wrote this reply and found out about it the same way everyone else did. And Jinky Ramos, the fan, who left a real comment about something hard and got mocked back, in Mira's name, by a machine she didn't know was answering."
Why this works
This is the G step. Naming the operator and the subject in the same breath stops the answer from staying about the account and starts it being about the person on the other end.
T+5
Kill the publish path everywhere, not just for Mira
Say it like this
"First real move, by five past two: flip ReplyForge's auto-publish path to draft-only, platform-wide. All 2,300 accounts, not just Mira's. We did talk about just suspending her account and reviewing it, that's the smaller, faster-looking move. I'd reject it, because the model behind this reply is the same model behind all 61,000 replies a day. Leaving anyone else live just means the next one lands on a different creator's followers while we're still looking at this one."
Why this works
This is the R step, and naming the rejected alternative out loud shows a real judgment call, not the only option that occurred to you.
T+15
Say plainly why this landed harder on some people than others
Say it like this
"Here's the part I'd say even though it's uncomfortable. This probably wasn't the first offensive reply ReplyForge ever produced, it's the first one anyone noticed. The safety classifier gating every reply was tuned and tested mostly on English text, 99.2 percent recall on a 5,000-example set. Once we built a matching set in Tagalog after the fact, recall was 71 percent. That gap doesn't spread evenly. It sits on exactly the followers writing in a language our own safety net was never really checked against."
Why this works
This is the U step. A number with a comparison, not just a single recall figure, is what makes "lands unevenly" a fact instead of a feeling.
T+25
Pull back what's already out, and go find the person it reached
Say it like this
"Next twenty minutes: pull the flagged reply, and sweep the last 24 hours for anything matching the same pattern, non-English input, a high mockery score. That sweep turns up six more. Jinky never knew this was a machine talking. She has no button to appeal to and no email to write. So support reaches out to her directly, today, before she has to find us."
Why this works
This is the A step in action. Naming that she can't contest it is the strongest line in the whole answer, and it turns straight into a concrete task instead of staying a sad observation.
T+45
Log it before anything gets reset
Say it like this
"Before we roll the model back to the last version we trust, someone captures the exact version number, the prompt template, and ten more examples like this one. Otherwise tomorrow this is a story everyone half-remembers instead of a bug we can actually point at."
Why this works
Shows you think past the hour, toward the fix, without letting the logging step delay the kill switch that already happened at T+5.
T+55
Set the real bar to turn it back on, then close on one line
Say it like this
"Auto-publish doesn't come back for a language until the classifier clears 98 percent recall on a red-teamed set built for that language, not 100, because zero misses isn't a real bar for a probabilistic classifier, and not one blended global number, because that's exactly what hid this gap. Going forward, the classifier's flag rate by detected language sits on a live dashboard, so the next gap like this shows up there before it shows up as a screenshot. So in one sentence: cut the switch everywhere in five minutes, spend the hour pulling down what's already out and finding who it reached, and don't turn it back on language by language until a real eval set says so."
Why this works
This is the D step plus the close. A stated threshold and a live metric are things an interviewer can actually check you understand, not a promise to "monitor closely."
If you remember one thing A creator-level suspension fixes the account that got caught. A platform-wide kill switch fixes the hour. The model, not the account, is what's broken, and the fix has to match where the failure actually lives.

Let's learn

Say a social app builds a helper that writes and posts replies to fan comments by itself, in a creator's own voice, so a creator getting thousands of comments a day doesn't have to answer them one by one.

Before a tool like this, a creator with a big following might get 2,000 comments under one post. Answering even a tenth of them by hand took a couple of hours. Most comments got no answer at all.

Knowledge spark: what's a red-teamed eval set? A batch of test examples built on purpose to find a weakness, not pulled from everyday use. Real traffic tells you how the model does on what usually happens. A red-teamed set tells you how it does on the case you're worried about, whether or not that case has shown up yet.

Now the tool posts about 61,000 replies a day across every creator who's turned it on. Each one goes out in seconds. Nobody waits, nobody reviews it first.

The extra replies are not the problem. The problem shows up the day one of those replies is genuinely cruel, posted under a real person's name, to someone who never got a chance to check it before it went out.

The reply did not fail a spot check. It failed a person who never got to run one.
The decision that mattered Cut the model's publish path everywhere the moment it's caught, not just on the account it happened to. The failure lives in the model every creator shares, so a fix scoped to one account leaves the same gap open for the next 2,299.

At its worst, a machine mocks someone for something hard they just shared, wearing a creator's face while it does it. The follower has no way to know it wasn't really her. The screenshot spreads to a whole community faster than any takedown can follow it, and the company finds out from a stranger's post instead of its own dashboard.

The choice I would take back. When the safety classifier that gates every reply was first signed off, the bar was one blended recall number across all languages combined, 99.2 percent, measured on a 5,000-example set that was almost entirely English. That felt thorough. It also meant a language with real, serious gaps could sit under that one number for months, invisible, because the number itself never got broken apart.

What I would leave alone. Nettleframe also lets creators type and send their own replies by hand, no auto-publish involved. That path doesn't change. This incident is about what a machine posts on its own; a blunt reply a real person typed and chose to send is a different problem with a different owner.

The lesson. A recall number that looks strong in aggregate can be hiding a language, a dialect, or a group it barely works for at all. If you never split the number apart, you'll only find the gap the same way this one got found, after someone screenshots it.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how close one blended number came to never getting questioned.

Selim Bragoli has run trust and safety product for Nettleframe for three years. He was the one who pushed to launch ReplyForge with a safety classifier in front of it at all, back when the first draft of the plan was to ship without one and add it "if it turns out to be a problem."

ReplyForge's first eight months were genuinely good. Creators who'd been drowning in comments started answering almost everyone. Engagement on comment sections climbed. The safety classifier caught the obvious stuff, insults, slurs, spam, at a rate everyone was proud of: 99.2 percent recall on the launch eval set.

Nobody had asked, in that launch review, how the model did on comments that weren't in English.

At 2:14 on a Tuesday, Selim's phone lit up with a screenshot forwarded from the on-call channel. A fan named Jinky Ramos had left a comment on wellness creator Mira Sondergaard's latest post, in Tagalog, sharing that she'd just gotten a hard health update and didn't know how to feel about it. ReplyForge answered, in Mira's voice, with something that read as flippant and mocking, the kind of reply that trivializes exactly the thing someone just worked up the nerve to say out loud.

Jinky didn't know it was a machine. Mira didn't know it had happened until three of her other Tagalog-speaking followers tagged her, confused and upset, within about forty minutes.

It was never really about the one reply. It was about how long a gap that size could sit there, quietly, inside a number nobody had split apart.

Selim's first move wasn't to suspend Mira's account and start a slow review. Someone on the call suggested exactly that, contain it to the account that got caught, keep the rest of the platform running while it gets sorted out. Selim said no. The model behind that reply was the same model behind every other account still posting live. Containing it to Mira's account meant accepting that the next bad reply, in some other language, under some other creator's name, was still a live possibility for as long as the review took.

Two figures side by side. Left, Selim, the operator, labeled as holding the switch, able to turn auto-publish off for every account. Right, Jinky, the subject, labeled as having left a real comment and been mocked back by a machine she never knew existed.
Selim has a lever. Jinky doesn't even know there was a machine to argue with.

By 2:19, ReplyForge was in draft-only mode, everywhere, for all 2,300 accounts. By 2:39, a sweep of the last 24 hours had turned up six more replies matching the same pattern, non-English input, a high mockery score from the classifier's own logs, none of them caught before publish. By 2:59, support had reached out directly to every follower who'd received one, including Jinky, rather than waiting for any of them to file a complaint that might never come.

What stayed with Selim wasn't the six replies themselves. It was how ordinary the original launch review had felt. Nobody had been careless. A 99.2 percent recall number is a genuinely good number. It just happened to be an average of a language the classifier handled almost perfectly and a language it barely handled at all, and averaging the two together is exactly what let the second one hide.

The redesigned safety bar now grades the classifier per language, not once, blended. Every language on the platform gets its own red-teamed eval set, and auto-publish only turns on for a language once that set clears 98 percent, checked weekly, not the day the fix ships and never again. Run the same launch review under that design from the start, and the Tagalog gap gets found and closed months before any real comment could have hit it.

The thing Selim would tell his past self, back at that first launch review: a number that looks strong right up until you ask what it's an average of usually is.

GUARD, letter by letter: who gets the lever and who doesn't

This is a risk and safety question about a live harm, not a sizing problem or a metric to track, so GUARD fits and BOUND or LEAD don't.

G, groups. Two of them, always named together. The operator: Mira Sondergaard, the creator, whose voice and name the reply was posted under, and who trusted ReplyForge to speak for her. The subject: Jinky Ramos, the fan, who left a real comment and received a machine's reply with no idea it wasn't Mira herself.
U, unequal. The safety classifier's recall was 99.2 percent on a 5,000-example English eval set, and 71 percent on a 300-example Tagalog set built after the fact, with an 84 percent figure on a 400-example Spanish set landing in between. The harm from that gap concentrates on exactly the followers writing in a language the safety net was weakest against.
A, ability to contest. Jinky never knew a machine had replied. She has no appeal button, no email address for "this wasn't really her," and no reason to think the offensive reply came from anywhere but Mira herself. The company has to find her; she has no lever to find the company.
R, reduce. Flip ReplyForge's auto-publish path to draft-only, platform-wide, for all 2,300 accounts, not a review scoped to the one flagged account. The rejected alternative, suspending only Mira's account while the rest keeps running, was considered and turned down because the failure is model-level, shared by every account, not something specific to hers.
D, detect. Two things going forward: a red-teamed eval set per language, with a 98 percent recall bar checked weekly before that language's auto-publish path reopens, and the classifier's own flag rate broken out by detected input language on a live dashboard, so a gap like this shows up as an internal number trending the wrong way, before it shows up as a screenshot.

A left to right flow diagram: comment posted, reply drafted, safety score, auto published, then a gap marked in red labeled no appeal step, showing there is nowhere in the pipeline for the person who received the reply to contest it.
Every step in the pipeline has an owner. The step where the subject gets to say "that wasn't right" was never built.

One trade-off worth saying plainly: draft-only mode is slower. A reply that used to post in seconds now sits in a review queue, and someone, a moderator or the creator, has to spend real time on it. That's a real cost in speed and in people's hours, accepted on purpose while the per-language fix ships, because the alternative is an unmonitored gap staying open for however long "review it later" actually takes.

Recall by language, on the classifier gating every reply
English (5,000-example eval set)99.2%
Spanish (400-example red-teamed set)84%
Tagalog (300-example red-teamed set)71%
One blended number, 99.2 percent, was the launch bar. It was never the real story for two of the platform's own languages.
Where the first hour actually went
Kill switch flipped everywhere0 to 5 min
Sweep and pull matching replies5 to 25 min
Reach the people it reached25 to 45 min
Log it, set the reopen bar45 to 60 min
The switch is the fast part. Reaching the people already affected takes longer than turning the model off ever does.

And if you want to be sure it really works, try it somewhere else

Vantree Health runs Concierge, an AI chat that gathers symptoms from patients before a video visit, across about 40 clinics, roughly 9,000 intake chats a day. Zita Buchholz is the product manager who owns it.

G, groups. The operator: the clinic staff and physicians who read Concierge's intake summary and trust it to have captured what a patient said. The subject: the patient typing into the chat, often describing something serious in their own words before ever talking to a person.
U, unequal. Concierge's tone-reading layer, the part that decides how urgently to flag a message, was tuned mostly on clinical phrasing. On a 2,000-example eval set written the way a doctor's note reads, it flagged self-harm-adjacent language correctly 97 percent of the time. On a 250-example set built afterward using casual, texting-style phrasing, the kind a teenager or a stressed adult actually types, that dropped to 68 percent. The harm concentrates on patients who don't write like a clinical chart.
A, ability to contest. A patient who types something serious in casual language and gets a flat, low-urgency response has no way to know Concierge misjudged it. They don't know a person never saw the message. They may just conclude nobody was worried, and not bring it up again.
R, reduce. Zita's team didn't retrain the tone layer first and call it done, the fast-looking fix. They added a second, keyword-based safety net running in parallel: any message hitting a fixed list of self-harm and crisis terms routes to a live nurse queue automatically, regardless of what the tone layer scored it. Two independent checks, not one smarter version of the same check.
D, detect. Every message where the tone layer says "low urgency" and the keyword net says "route it anyway" gets logged as a disagreement. That disagreement rate is now a weekly number the team watches, because a rising rate means the tone layer's blind spot is showing up more often, not less.

Same shape, different stakes At Nettleframe, an unnoticed reply is cruel and public. At Vantree, an unnoticed miss is a person in crisis who never got flagged. The GUARD shape doesn't change: name both people, find where the harm concentrates, and build a second, independent check rather than trust one smarter version of the first one.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the kill switch: cut the publish path everywhere in the first five minutes, then spend the rest of the hour finding who was already reached.
Cost: the trust and safety team only has a few hours of engineering time this week. Don't shrink the per-language eval sets to fit the budget, keep draft-only mode on longer instead; that's a real, statable trade-off, not a silent one.
The model got better: the next version of the classifier clears 98 percent on every language's red-teamed set on the first try. That doesn't make the per-language bar pointless, it's still the only way anyone would know that for certain, instead of assuming it from one blended number.

Where people run it wrong.
They suspend the one account that got caught and call the incident closed, when the failure lives in the model every account shares.
They grade a safety classifier on one blended number and never ask what groups that average is hiding.
They wait for the person who was harmed to complain, instead of going to find them, when that person often has no idea a machine was ever involved.

How to use it live. Say the two people before you say a single fix: "there's the operator who trusted the tool, and there's the person on the receiving end who never opted into a machine speaking for someone else." That buys you the room to ask who, specifically, this product's version of that second person is, instead of reciting a generic incident-response checklist.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits "your model starts producing offensive output, describe the first hour," and why not BOUND or LEAD?
Tap to flip
ANSWER
GUARD. This is a risk and safety question about a live harm reaching a real person, not a sizing problem or a metric to track over time.
2 · THE TWO PEOPLE
Who are the operator and the subject in this story?
Tap to flip
ANSWER
Mira Sondergaard, the creator, is the operator whose voice the reply posted under. Jinky Ramos, the fan, is the subject who received the offensive reply with no idea it came from a machine.
3 · THE UNEVEN PART
Why did this harm land specifically on Tagalog-speaking followers rather than everyone equally?
Tap to flip
ANSWER
The safety classifier gating every reply was tuned mostly on English text. Recall was 99.2 percent in English but only 71 percent on a Tagalog red-teamed set, so the gap concentrated on exactly the followers writing in that language.
4 · WHO CAN'T PUSH BACK
Why couldn't Jinky just report the reply and get it sorted out herself?
Tap to flip
ANSWER
She never knew a machine had answered. She had no reason to think the reply wasn't really Mira, no appeal button, and no address to write to, so the company had to go find her rather than wait for a complaint.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Grading the safety classifier's launch readiness on one blended recall number, 99.2 percent, across all languages combined. It made sense because the number looked strong and thorough, right up until someone asked what it was an average of.
6 · THE NUMBER
Fill in the blank: recall was ___% on the English eval set and ___% on the Tagalog red-teamed set.
Tap to flip
ANSWER
99.2% in English. 71% in Tagalog.
7 · THE REPLAY
Same incident, redesigned safety bar. What changes?
Tap to flip
ANSWER
Every language gets its own red-teamed eval set and a 98 percent recall bar, checked weekly, before auto-publish reopens for it. Run under that design from the start, the Tagalog gap gets found and closed months before any real comment could have hit it.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD on a different product in a different domain. Which product, and what's its version of the safety gap?
Tap to flip
ANSWER
Vantree Health's Concierge intake chat. Its tone-reading layer missed self-harm-adjacent language written in casual, texting-style phrasing far more often than clinically phrased language, so the team added a second, independent keyword-based safety net.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Selim reject suspending only Mira's account and reviewing it on its own?
  • A. Suspending one account would have upset Mira unfairly.
  • B. The failure lived in the model every account shared, so leaving the rest live left the same gap open for the next creator.
  • C. Suspending accounts required a legal review that would take too long.
  • D. Nettleframe's policy required a platform-wide response for any complaint.
Show hint
Think about where the actual defect lived, in one account's settings or in the shared model.
Show answer
B. The classifier gap was language-based, not account-based, so 2,299 other accounts were still exposed to the exact same risk the moment the review started.
True or false
2. True or false: once ReplyForge's auto-publish path was switched off platform-wide, the incident was effectively over.
  • True
  • False
Show hint
Think about what had already been posted before the switch flipped.
Show answer
False. Turning the switch off stops new harm, but six more matching replies were already live and had to be pulled, and the people who received them, including Jinky, still had no idea a machine had spoken for Mira until support reached out directly.
Fill in the blank
3. The classifier's launch bar was one blended recall number of ___ percent across all languages. Auto-publish now only reopens for a language once its own red-teamed set clears ___ percent.
Show hint
Check the GUARD recap's U and D steps.
Show answer
99.2 percent; 98 percent. The new bar is lower than the old blended number on purpose, because it's now measured on the harder, per-language set, not the easier average.
Short answer
4. Name a place in Nettleframe's own product where this exact kind of incident would NOT need the same platform-wide kill switch.
Show hint
Look at the "what I would leave alone" paragraph in Let's learn.
Show answer
Model answer: Replies a creator types and sends herself, with no auto-publish model involved. A blunt or badly worded reply from a real person is a different problem, handled by the creator's own judgment, not a model-level safety gap.
Short answer, apply it yourself
5. Think of an AI tool you use or have seen in the world. Who's a group of people it produces output for, or about, who would have no real way to contest a mistake it made?
Show hint
Look for someone who receives the output secondhand, without ever seeing the tool itself.
Show answer
Model answer: An AI tool that drafts landlord messages to tenants about lease issues. The tenant receiving the message has no way to know it was AI-drafted, no way to check if it correctly represented the landlord's actual position, and no address to contest a tone or a claim that was wrong.
Short answer, the number question
6. If Nettleframe had only suspended Mira's account and left the other 2,299 live, roughly how many replies would still have gone out platform-wide, unmonitored, in that same first hour, given the 61,000-a-day rate?
Show hint
Divide the daily rate by 24 to get a rough hourly figure, then note that one account out of 2,300 barely changes it.
Show answer
Roughly 2,500. 61,000 divided by 24 is about 2,540 replies an hour platform-wide. Removing Mira's single account from that pool changes the number by almost nothing, which is exactly why a single-account fix doesn't address a model-level gap.
Before you close the answer
Why this works
Tests whether you reach for the account-level fix that looks fast and contained, or the model-level fix that actually matches where the failure lives. Most candidates stop at "suspend the account and investigate." The strong answer also names who received the harm and has no way to contest it, then builds a concrete step around that person, not just around the model.
Follow-up traps
"Isn't a platform-wide shutdown a huge overreaction to one bad reply?" Response: the shutdown lasts as long as draft-only mode takes to catch up, hours, not a launch reversal, and the alternative, leaving 2,299 other accounts exposed to the same model, is the bigger risk, not the smaller one.

"How do you know six more matching replies is the full extent of it?" Response: it's the full extent of what the pattern sweep found in 24 hours using the classifier's own logs; the per-language eval work that follows is specifically there to find what a pattern sweep alone can't.
If pressed
The disagreement rate isn't unique to Vantree's second product. Nettleframe added the same idea after this incident: a lightweight keyword net runs alongside the main classifier, and any case where the two disagree gets logged and reviewed weekly, a second independent check rather than one smarter version of the first.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more