Your model starts producing offensive output. Describe the first hour.
- Kill the model's auto-publish path platform-wide, in the first five minutes, not just for the flagged account.Why: the failure is in the model every account shares, so leaving anyone else live just moves the harm to a different community while you're still staring at this one.
- Pull down the flagged reply and every other reply posted in the last 24 hours that matches the same pattern.Why: the people who received them never knew a machine wrote them and have no way to ask for one to come down themselves.
- Reach out directly to everyone who received one of those replies, before they have to complain.Why: a follower mocked in a creator's name has no appeal button and no address to write to, so the company has to go find her.
- Log the exact model version, the prompt template, and real examples before anything gets rolled back.Why: without a snapshot, tomorrow's fix is a guess about what changed, not a diagnosis of what actually broke.
- Set the reopen bar as a recall number checked per language on a red-teamed set, not one blended global number.Why: a 99 percent overall score is exactly what let this gap hide for as long as it did.
- Put the safety classifier's flag rate, broken out by detected input language, on a live dashboard going forward.Why: that's what turns the next version of this into an internal alert instead of a stranger's screenshot.
How to answer this, stage by stage
Nobody is grading whether you can say "we'd take it seriously." They are grading whether you can name who gets hurt, who can't complain about it, and what you'd actually flip off in the first sixty minutes. Seven moves, timed against the hour itself.
Let's learn
Say a social app builds a helper that writes and posts replies to fan comments by itself, in a creator's own voice, so a creator getting thousands of comments a day doesn't have to answer them one by one.
Before a tool like this, a creator with a big following might get 2,000 comments under one post. Answering even a tenth of them by hand took a couple of hours. Most comments got no answer at all.
Now the tool posts about 61,000 replies a day across every creator who's turned it on. Each one goes out in seconds. Nobody waits, nobody reviews it first.
The extra replies are not the problem. The problem shows up the day one of those replies is genuinely cruel, posted under a real person's name, to someone who never got a chance to check it before it went out.
At its worst, a machine mocks someone for something hard they just shared, wearing a creator's face while it does it. The follower has no way to know it wasn't really her. The screenshot spreads to a whole community faster than any takedown can follow it, and the company finds out from a stranger's post instead of its own dashboard.
The choice I would take back. When the safety classifier that gates every reply was first signed off, the bar was one blended recall number across all languages combined, 99.2 percent, measured on a 5,000-example set that was almost entirely English. That felt thorough. It also meant a language with real, serious gaps could sit under that one number for months, invisible, because the number itself never got broken apart.
What I would leave alone. Nettleframe also lets creators type and send their own replies by hand, no auto-publish involved. That path doesn't change. This incident is about what a machine posts on its own; a blunt reply a real person typed and chose to send is a different problem with a different owner.
The lesson. A recall number that looks strong in aggregate can be hiding a language, a dialect, or a group it barely works for at all. If you never split the number apart, you'll only find the gap the same way this one got found, after someone screenshots it.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how close one blended number came to never getting questioned.
Selim Bragoli has run trust and safety product for Nettleframe for three years. He was the one who pushed to launch ReplyForge with a safety classifier in front of it at all, back when the first draft of the plan was to ship without one and add it "if it turns out to be a problem."
ReplyForge's first eight months were genuinely good. Creators who'd been drowning in comments started answering almost everyone. Engagement on comment sections climbed. The safety classifier caught the obvious stuff, insults, slurs, spam, at a rate everyone was proud of: 99.2 percent recall on the launch eval set.
Nobody had asked, in that launch review, how the model did on comments that weren't in English.
At 2:14 on a Tuesday, Selim's phone lit up with a screenshot forwarded from the on-call channel. A fan named Jinky Ramos had left a comment on wellness creator Mira Sondergaard's latest post, in Tagalog, sharing that she'd just gotten a hard health update and didn't know how to feel about it. ReplyForge answered, in Mira's voice, with something that read as flippant and mocking, the kind of reply that trivializes exactly the thing someone just worked up the nerve to say out loud.
Jinky didn't know it was a machine. Mira didn't know it had happened until three of her other Tagalog-speaking followers tagged her, confused and upset, within about forty minutes.
Selim's first move wasn't to suspend Mira's account and start a slow review. Someone on the call suggested exactly that, contain it to the account that got caught, keep the rest of the platform running while it gets sorted out. Selim said no. The model behind that reply was the same model behind every other account still posting live. Containing it to Mira's account meant accepting that the next bad reply, in some other language, under some other creator's name, was still a live possibility for as long as the review took.
By 2:19, ReplyForge was in draft-only mode, everywhere, for all 2,300 accounts. By 2:39, a sweep of the last 24 hours had turned up six more replies matching the same pattern, non-English input, a high mockery score from the classifier's own logs, none of them caught before publish. By 2:59, support had reached out directly to every follower who'd received one, including Jinky, rather than waiting for any of them to file a complaint that might never come.
What stayed with Selim wasn't the six replies themselves. It was how ordinary the original launch review had felt. Nobody had been careless. A 99.2 percent recall number is a genuinely good number. It just happened to be an average of a language the classifier handled almost perfectly and a language it barely handled at all, and averaging the two together is exactly what let the second one hide.
The redesigned safety bar now grades the classifier per language, not once, blended. Every language on the platform gets its own red-teamed eval set, and auto-publish only turns on for a language once that set clears 98 percent, checked weekly, not the day the fix ships and never again. Run the same launch review under that design from the start, and the Tagalog gap gets found and closed months before any real comment could have hit it.
The thing Selim would tell his past self, back at that first launch review: a number that looks strong right up until you ask what it's an average of usually is.
GUARD, letter by letter: who gets the lever and who doesn't
This is a risk and safety question about a live harm, not a sizing problem or a metric to track, so GUARD fits and BOUND or LEAD don't.
G, groups. Two of them, always named together. The operator: Mira Sondergaard, the creator, whose voice and name the reply was posted under, and who trusted ReplyForge to speak for her. The subject: Jinky Ramos, the fan, who left a real comment and received a machine's reply with no idea it wasn't Mira herself.
U, unequal. The safety classifier's recall was 99.2 percent on a 5,000-example English eval set, and 71 percent on a 300-example Tagalog set built after the fact, with an 84 percent figure on a 400-example Spanish set landing in between. The harm from that gap concentrates on exactly the followers writing in a language the safety net was weakest against.
A, ability to contest. Jinky never knew a machine had replied. She has no appeal button, no email address for "this wasn't really her," and no reason to think the offensive reply came from anywhere but Mira herself. The company has to find her; she has no lever to find the company.
R, reduce. Flip ReplyForge's auto-publish path to draft-only, platform-wide, for all 2,300 accounts, not a review scoped to the one flagged account. The rejected alternative, suspending only Mira's account while the rest keeps running, was considered and turned down because the failure is model-level, shared by every account, not something specific to hers.
D, detect. Two things going forward: a red-teamed eval set per language, with a 98 percent recall bar checked weekly before that language's auto-publish path reopens, and the classifier's own flag rate broken out by detected input language on a live dashboard, so a gap like this shows up as an internal number trending the wrong way, before it shows up as a screenshot.
One trade-off worth saying plainly: draft-only mode is slower. A reply that used to post in seconds now sits in a review queue, and someone, a moderator or the creator, has to spend real time on it. That's a real cost in speed and in people's hours, accepted on purpose while the per-language fix ships, because the alternative is an unmonitored gap staying open for however long "review it later" actually takes.
And if you want to be sure it really works, try it somewhere else
Vantree Health runs Concierge, an AI chat that gathers symptoms from patients before a video visit, across about 40 clinics, roughly 9,000 intake chats a day. Zita Buchholz is the product manager who owns it.
G, groups. The operator: the clinic staff and physicians who read Concierge's intake summary and trust it to have captured what a patient said. The subject: the patient typing into the chat, often describing something serious in their own words before ever talking to a person.
U, unequal. Concierge's tone-reading layer, the part that decides how urgently to flag a message, was tuned mostly on clinical phrasing. On a 2,000-example eval set written the way a doctor's note reads, it flagged self-harm-adjacent language correctly 97 percent of the time. On a 250-example set built afterward using casual, texting-style phrasing, the kind a teenager or a stressed adult actually types, that dropped to 68 percent. The harm concentrates on patients who don't write like a clinical chart.
A, ability to contest. A patient who types something serious in casual language and gets a flat, low-urgency response has no way to know Concierge misjudged it. They don't know a person never saw the message. They may just conclude nobody was worried, and not bring it up again.
R, reduce. Zita's team didn't retrain the tone layer first and call it done, the fast-looking fix. They added a second, keyword-based safety net running in parallel: any message hitting a fixed list of self-harm and crisis terms routes to a live nurse queue automatically, regardless of what the tone layer scored it. Two independent checks, not one smarter version of the same check.
D, detect. Every message where the tone layer says "low urgency" and the keyword net says "route it anyway" gets logged as a disagreement. That disagreement rate is now a weekly number the team watches, because a rising rate means the tone layer's blind spot is showing up more often, not less.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the kill switch: cut the publish path everywhere in the first five minutes, then spend the rest of the hour finding who was already reached.
Cost: the trust and safety team only has a few hours of engineering time this week. Don't shrink the per-language eval sets to fit the budget, keep draft-only mode on longer instead; that's a real, statable trade-off, not a silent one.
The model got better: the next version of the classifier clears 98 percent on every language's red-teamed set on the first try. That doesn't make the per-language bar pointless, it's still the only way anyone would know that for certain, instead of assuming it from one blended number.
Where people run it wrong.
They suspend the one account that got caught and call the incident closed, when the failure lives in the model every account shares.
They grade a safety classifier on one blended number and never ask what groups that average is hiding.
They wait for the person who was harmed to complain, instead of going to find them, when that person often has no idea a machine was ever involved.
How to use it live. Say the two people before you say a single fix: "there's the operator who trusted the tool, and there's the person on the receiving end who never opted into a machine speaking for someone else." That buys you the room to ask who, specifically, this product's version of that second person is, instead of reciting a generic incident-response checklist.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"How do you know six more matching replies is the full extent of it?" Response: it's the full extent of what the pattern sweep found in 24 hours using the classifier's own logs; the per-language eval work that follows is specifically there to find what a pattern sweep alone can't.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Incident management for AI products
- #1 What counts as an incident for an AI feature but not for a normal one?
- #2 Write the severity definitions for AI quality incidents.
- #4 How do you triage an incident where the code is fine and the model is the problem?
- #5 What is the AI equivalent of a rollback, and when is it not available?
- #6 Describe the on-call runbook entry for a sudden quality drop.
- #7 How do you decide whether to disable a feature or degrade it during an incident?