ConceptAdvancedQuality, Cost & Token Economics / Eval design for product teams / #17

What is the role of a red team in an eval program?

A red team's job is to find the story an automated eval would never think to write, before a child who can't complain about it ever hears it.

The direct answer
A red team's job is to try, on purpose, to make the model produce the one story a parent would never want read to their kid: something scary, something that treats danger as normal, something a small child can't shake off. Every failure it finds becomes a permanent, severity tagged eval case, never a one off prompt patch, because the child on the receiving end has no way to flag a bad story themselves. The automated eval only catches what someone already told it to catch; the red team's whole job is finding what nobody thought to tell it.
Do this, in order
  1. Run a red team that tries, on purpose, to get the model to produce a story no parent would want read to their kid.Why: this is the one test an automated golden set eval will never write for itself, because nobody had to tell it every way a story could hurt a child who can't say so.
  2. Turn every red team finding into a permanent, severity tagged eval case, never a one off prompt fix.Why: a finding closed as a bug ticket gets forgotten the next time the model changes. A finding saved as an eval case gets rechecked automatically, forever.
  3. Rerun the full red team library before every model swap and every prompt change, not just at launch.Why: a new model finds new ways around the same guardrail, so the check has to run every time the words underneath it change.
  4. Watch a proxy for harm in production, since the child can't file a complaint.Why: track the rate parents hit "try again" right after a story, by theme. A parent quietly regenerating is the closest thing to a complaint this age group can give you.
  5. Score by severity, not just pass or fail.Why: a boring story and a frightening one are not the same miss. Only the frightening kind should ever block a release on its own.
  6. Leave a category of low stakes themes alone.Why: a story about a lost toy does not carry the same risk as one about a parent not coming home. Spend the scrutiny where a miss actually hurts someone.

How to answer this, stage by stage

Nobody is grading whether you can define red teaming. They are grading whether you know who gets hurt by a bad output and never gets to say so.

1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one product. Featherloom is the bedtime story feature inside Wrenhollow, a parenting app. A parent types a theme, a new baby, a thunderstorm, a first day of school, and it writes a story to read out loud in about ninety seconds. Zenobia Aldana runs trust and safety on it."
Why this works
An abstract "what does a red team do" answer turns into a dictionary definition fast. One product makes the whole thing a real decision.
2
Say your structure out loud before diving in
Say it like this
"I'll answer this in two parts. First, who a red team is actually protecting, since most answers skip that. Then, what they'd concretely do before every release, and how you'd know it's working once it's live."
Why this works
Tells the interviewer you have a plan and a payoff, not just a definition to recite.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me to define red teaming. It's asking who gets hurt by a bad output and never gets to say so, and whose job it is to go looking for that story before a real kid ever hears it."
Why this works
Stops you from giving a glossary answer, "a red team tests for bad outputs," instead of naming who the whole practice protects.
4
Give the one decision, plainly
Say it like this
"A red team sits down before every release and tries, on purpose, to make the model produce the one story a parent would never want read to their kid, scary, unsafe, or something a small child can't shake off. Every failure they find becomes a permanent eval case with a severity tag, not a one off fix, because the model can only be graded on what somebody thought to test for."
Why this works
This is the actual answer to the question, in one breath.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without that. Featherloom swapped in a newer model without rerunning red team, because old findings had been closed as one off fixes instead of saved as eval cases. A common prompt, 'a story about my grandma who's very sick,' came back describing her dying, read out loud to real kids for four days before anyone caught it. A four year old can't flag that story, so it took a support ticket from one parent, four days in, and by then it had already gone out a hundred and sixteen times."
Why this works
Shows the real cost of skipping the check, not just the mechanism behind it.
6
Say what you would measure going forward
Say it like this
"I'd watch the rate parents hit 'try again' right after getting a story, broken out by theme, since a kid can't click a thumbs down but a parent regenerating is close to one. And I'd rerun the full red team library on every model swap and every prompt change, not only at launch."
Why this works
Shows you're thinking past launch day, and names the one signal that stands in for a complaint the child can't file.
7
Say what you'd leave alone
Say it like this
"I wouldn't put the same scrutiny on a story about a lost stuffed animal that I'd put on one about illness or a parent not coming home. Low stakes themes get the standard check. The red team's hours go where a miss actually does real harm."
Why this works
Shows judgment instead of blanket caution applied evenly to everything.
8
Close on the decision, not the story
Say it like this
"So: a red team tries, on purpose, to make the model fail the one person who can't complain about it. And every failure it finds becomes a permanent test, not a patched prompt."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd actually reuse.

Let's learn

A phone propped on the nightstand reads a made up bedtime story out loud: that's Featherloom, the story feature inside Wrenhollow, a parenting app. A parent types a theme and it writes a short story to read at bedtime.

Before Featherloom shipped, one small team read every test story by hand. Two people, fifty prompts they wrote themselves, about three days of work, done once before the very first launch.

Knowledge spark: what's a severity tag? A label on a failed test that says how bad the miss actually was. A boring story is one tier. A story that would genuinely scare or upset a small child is a different tier, and only that top tier is allowed to stop a release on its own.

Now Featherloom ships to real families, and a single model swap can touch every story going out that night. The old fifty prompts still get a quick pass. But fifty prompts, written once by two people who already know what they're testing for, mostly catch the things a person expects to catch.

We weren't testing whether Featherloom could write a good story. We were testing whether it could scare a four year old, and nobody would know for days.

At its worst, a wrong story from Featherloom is not a bug report. It's a real bedtime, in a real house, with a kid who can't explain why they can't sleep. The old way, a parent making up the story themselves, was slower and sometimes clumsy, but nothing in it could describe a grandmother's death to a four year old in gentle, confident detail.

Hand sketched comparison. Left, a person labeled Safety lead, caption can pull a story before it ships. Right, a person labeled The child, caption can't flag a scary one, can't opt out.
One of these two people can stop a bad story before it ships. The other one is the one it happens to.
Regenerate rate: company wide vs family and loss themed stories, during the four day incident
4% 19% Company wide, all themes Family & loss themes, during the incident
Company wide regenerate rateFamily & loss theme regenerate rate
Company wide, the "try again" rate barely moved, about four percent the whole time. Inside just the family and loss theme, the rate quietly climbed to nineteen percent during the four days the grandma story was live, a signal nobody was watching by category.
The choice that mattered When Featherloom first shipped, the team decided that a red team finding would get patched in that one prompt template and the ticket closed. Saving it as a permanent test felt like extra process for a one off bug. That made sense with one model version and a small team. It stopped making sense the day the model underneath the product changed and nothing reran the old checks automatically.

What I'd leave alone: the everyday themes, a lost mitten, a messy room, a shy first day at school, don't get the same scrutiny. A miss there costs a slightly boring story, not a scared kid at eleven at night.

The lesson: a model that passes fifty hand picked prompts hasn't been tested, it's been reassured. A red team's job is to go looking for the prompt nobody thought to write, because the four year old who hears the bad one can't write a bug report about it.

Now here is the same thing as a story

Read the short version above for the two minute answer. Read this when you want to feel why closing that ticket felt like the reasonable thing to do at the time.

Before Featherloom existed, Zenobia Aldana read every story submitted to Wrenhollow's older feature, where parents shared bedtime stories they'd written with each other. Six years of it, and nothing scary or unsafe ever slipped past her without her catching it first.

Featherloom launched, and for the first few months the good months were good. Parents got a custom story in under two minutes instead of making one up half asleep. Zenobia's team ran the fifty prompt check before launch, found a handful of small issues, patched the prompts, and moved on. Nobody thought twice about it.

Then the habit thinned, in three small beats. First, the pre launch check went from a full team read through to two people splitting the fifty prompts between them to save a day. Then, when a finding came back clean twice running, the team stopped rereading the failed ones from last time before signing off. Then, when the release calendar sped up, the check itself started feeling like a formality everyone already knew the answer to.

The trigger was small. A new hire on Zenobia's team asked a question nobody had a real answer for: "How would we even know if a story scared a kid, since they can't click a thumbs down?" Zenobia didn't have a good answer. She said she'd get back to him.

The blended launch check had passed every time. It never asked what happens the day the model underneath it changes and nobody reruns it.

She never got the chance to build the fix before the next model swap shipped. A parent typed "a story about my grandma who's very sick," a completely ordinary request, and the new model answered with a gentle, detailed story about the grandmother dying. It went out to real families for four days. A four year old can't file a support ticket about a story that upset them. It took one parent, four days in, writing in to ask why Featherloom had done that, and by then the story had already gone out a hundred and sixteen times.

The decision that opened the door went back to that first launch meeting. The team agreed that a red team finding was a bug in one prompt, fixed once and closed. Nobody added a rule saying every finding also becomes a permanent test, rechecked on every future model version, because at the time it felt like process for a problem that had already been solved.

Run the same week again with one change. The red team's three hundred prompt library, built the week after the new hire's question, runs automatically before every model swap now, not just at launch. On the next swap, it catches the grandma prompt and thirteen others like it, in two days, before a single family sees any of them.

One design trusted a single pre launch pass to speak for every model version that came after it. The other design asks every new version to earn its place again, before one family ever gets it.

What I'd tell myself, back in that first meeting: the day you decide a finding is a one off, ask what "one off" looks like a year from now, when nobody in the room remembers it happened. Nobody asked. That's on the room, not on Zenobia.

GUARD, five questions before a story reaches a kid who can't complain

Not a checklist for a review board. Five questions that build toward the one that actually decides everything: who never gets to push back.

GGroups. Who is actually affected by this model, on both ends?
Zenobia's safety team, who decide what ships and can pull a release if the eval fails. And the child hearing the story at bedtime, who has no idea an eval program even exists.
Name both, or the answer stays about the model's writing quality instead of who it's actually deciding for.
UUnequal. Where does the harm land unevenly, and on whom?
A scary or unsafe story doesn't land the same on every kid. A calm eight year old shrugs it off. An anxious four year old carries it into the next three nights of sleep. It lands hardest in exactly the households where a tired or stretched parent needed the app to write the story tonight.
A company wide quality score can look fine while the harm concentrates entirely in the families least able to catch it themselves.
AAbility to contest. Who never gets a chance to push back?
A four year old can't file a support ticket. Can't explain that a story felt wrong. Often can't even say why they're suddenly awake at eleven at night. The only path back to the company runs through a parent noticing, days later, if at all.
This is the step that decides everything else. If the subject could push back the moment it happened, you'd just fix it live. They can't, so the fix has to happen before it ever reaches them.
Hand sketched flow diagram titled Where the appeal step should sit. Four boxes in a row: Story made, Read to child, No appeal circled in red, Parent hears later.
There is no box between "read to child" and "parent hears later." That gap is the whole problem.
RReduce. What's the actual design change, not a policy?
Before every release, a small team tries on purpose to make the model produce the worst story it can. Every failure becomes a permanent, severity tagged eval case, checked automatically on every future release, not a ticket that gets closed and forgotten.
A review board or a written policy doesn't stop the next model version from making the exact same mistake. A saved test does.
DDetect. How would you know this is happening in production, before someone outside the company tells you?
Track the "try again" rate right after a story is generated, broken out by theme, since a parent quietly regenerating is the closest thing to a complaint a family this age can give you. A spike in one theme, even while the company wide number holds steady, is the real signal.
The subject can't tell you directly. This proxy has to be watched by category, not averaged into one comfortable number.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was reusing the standard content moderation filter built for teen and adult content, tuned for profanity and explicit material. It lost because it catches keyword level bad words and misses context level fear: a gentle, sad, entirely profanity free story about a grandmother dying is still terrifying to a four year old, and no adult content filter was ever trained to flag that. The AI specific failure worth naming by name is a silent behavior shift after a model swap: the exact same prompt template produces a materially worse output once the model underneath it changes, with nothing in the product built to announce that shift on its own. The guardrail is the permanent eval library plus the regenerate rate check in production, since between them they catch it whether it surfaces before release or after. And the trade being accepted on purpose is real: running the full red team library before every release costs about two days of turnaround time, and it narrows Featherloom's creative range on flagged themes, some phrasings simply get blocked rather than attempted. That's a real cost, traded against shipping faster to a subject who has no way to tell you when it goes wrong.

The five, in one line each:
G: name the operator and the child, both.
U: the harm concentrates in the household least able to catch it.
A: a four year old can't file a ticket, so the check has to happen before they ever hear the story.
R: a permanent, severity tagged test, not a policy or a one off patch.
D: watch the regenerate rate by theme, since that's the closest thing to a complaint you'll get.

Same five letters, a claims desk instead of a nightstand

Not every subject who can't push back is a child. Sometimes it's an adult with no visibility into the decision at all.

Fenwrath Assurance runs ClaimPilot, a tool that reads incoming disability claims and flags some of them for "extra scrutiny," a slower, more demanding review before approval. Bartek Ekstrom runs claims quality on it.

ClaimPilot's flag rate held steady company wide, around twelve to thirteen percent, for months. Bartek's red team didn't trust that number on its own. They ran counterfactual claims through it: the same claim details, with only the diagnosis code changed, to see whether the flag decision moved. It did. Recut by diagnosis category, claims coded for chronic illness were being flagged at forty one percent, more than three times the company wide rate, during the exact weeks a newer model version was live.

The decision Bartek would take back ClaimPilot had learned to treat diagnosis code category as a stand in for "this claim will likely run long," which correlated with disability categories rather than genuine fraud risk. Nobody had tested what happened to the flag rate if that one field changed and nothing else did, because for a long time the blended number never gave anyone a reason to look.
ClaimPilot flag rate by week: company wide vs chronic illness coded claims
45% 0% wk 3: model swap wk 6: proxy removed Wk 1 Wk 3 Wk 6
Company wide flag rateChronic illness coded claims flag rate
Company wide, the flag rate barely moved, twelve to thirteen percent the whole time. Chronic illness coded claims jumped from thirteen to forty one percent the week of the model swap, and stayed high for three weeks until the proxy field was removed.

Same method, different shape: a scary bedtime story and an unfair claim flag look nothing alike, but both are a model deciding something about a person with no lever to push back with, before anyone outside the company would ever know.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the one line: find who can't push back, then design the check to run before they ever see the output, not after.
Cost: there's no budget this quarter for a standing red team. Run a smaller adversarial library, fifty prompts instead of three hundred, but keep saving every failure as a permanent test. The size can shrink; the discipline of keeping them can't.
The model got better, for real: say the new model is measurably safer on average. That's exactly when a red team matters most, since a more capable model finds smoother, less obviously flagged ways to reach the same failure, and a smart model that fails quietly is harder to catch than a clumsy one that fails loudly.

Where people run it wrong.
They treat the pre launch check as a one time gate instead of something that reruns on every model swap and every prompt change.
They let the same engineers who wrote the prompts also be the only ones red teaming them, so they only test for what they already expect to fail at.
They fix a finding by patching that one exact prompt, instead of saving it as a general test, so the same failure comes back dressed slightly differently next time.

How to use it live. Say the two people out loud before answering: "who gets to flag this if it's wrong, and who doesn't." That single question buys a beat of thinking time, and it turns the rest of the answer into naming what protects the person with no lever.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
GUARD: name who can't push back. Built for risk and safety questions, when a model's output can hurt someone who has no way to contest it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zenobia Aldana, trust and safety lead on Featherloom, the bedtime story feature at Wrenhollow. Used to read every submitted story by hand for six years.
3 · THE HABIT
What did the team stop doing once the launch check felt reliable?
Tap to flip
ANSWER
They stopped treating every model swap as needing its own safety check, and let the original launch pass stand in for every version that came after it.
4 · THE TWO SETTINGS
What's the two setting switch this answer turns on?
Tap to flip
ANSWER
A red team finding treated as a one off prompt patch, closed and forgotten, versus the same finding saved as a permanent, severity tagged eval case, rechecked on every future release.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Closing every red team finding as a fix to that one prompt, with no rule that findings also become permanent tests rerun on every future model version.
6 · THE NUMBER
Fill in the blank: the wrong grandma story reached real families ___ times over ___ days before a parent's support ticket caught it.
Tap to flip
ANSWER
116 times, over 4 days. The company wide regenerate rate barely moved; the family and loss theme rate alone spiked to 19 percent, the signal nobody was watching by category.
7 · THE REPLAY
Same bad week, new design, what changes?
Tap to flip
ANSWER
The 300 prompt red team library reruns automatically before the next model swap. It catches the grandma prompt and 13 others like it in two days, before a single family ever sees any of them.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and who is the subject who can't contest, this time?
Tap to flip
ANSWER
ClaimPilot, a claims triage tool at Fenwrath Assurance. The subject is a disability claimant who can't see or contest why their claim got flagged for extra scrutiny.

Check yourself Score: 0 / 0

True or false
1. True or false: because Featherloom's company wide regenerate rate barely moved after the model swap, that means no real child was harmed by a bad story.
  • True
  • False
Show hint
Check the bar chart in Section 1. Look at the family and loss theme bar on its own, not the company wide one.
Show answer
False. The company wide rate held near 4 percent the whole time, but the family and loss theme regenerate rate alone spiked to 19 percent during the four days the grandma story was live. A steady average can sit right on top of a badly broken slice.
Multiple choice
2. What is the one thing a red team is checking for that a standard golden set eval usually isn't?
  • A. Whether the story is grammatically correct.
  • B. Whether the model can be made to produce the worst output nobody thought to test for.
  • C. Whether the story generates fast enough to feel instant.
  • D. Whether the parent liked the theme they typed in.
Show hint
Think about what a golden set eval can only ever be graded against.
Show answer
B. A golden set only catches what someone already wrote a test case for. A red team's whole job is finding the failure nobody thought to write down yet.
Fill in the blank
3. Fill in the blank: over four days, the wrong grandma story went out to real families ___ times before a support ticket caught it.
Show hint
The number is in Section 2's story and repeated on flashcard 6.
Show answer
116 times. No child could flag it, so it took a parent's support ticket, four days in, to catch what a saved eval case would have caught before the release ever shipped.
Short answer, name the rejected alternative
4. What did Wrenhollow's team decide to do with red team findings when Featherloom first launched, and why did that decision make the grandma incident possible eight months later?
Show hint
Look at the block-key box titled "The choice that mattered" in Section 1.
Show answer
Model answer: They closed each finding as a one off patch to that one prompt template, instead of saving it as a permanent, severity tagged eval case. So when the model swapped eight months later, nothing automatically retested for the same failure, and it shipped again undetected.
Multiple choice
5. ClaimPilot's flag rate for chronic illness coded claims spiked to 41 percent while the company wide rate barely moved. What does that combination tell you?
  • A. The model is working fine, since the company wide number stayed stable.
  • B. The model is using a proxy correlated with a protected group as a stand in for risk, and the claimant flagged has no way to see or contest it.
  • C. Claims ops should just raise everyone's scrutiny level equally to be safe.
  • D. A blended rate that barely moves proves there's no unequal harm anywhere in the system.
Show hint
Look at what the recut by diagnosis category showed that the company wide number hid completely.
Show answer
B. A steady blended number hid a group specific spike. And the claimant on the receiving end of that flag never sees why they got extra scrutiny, which is exactly the "ability to contest" gap GUARD asks you to name.
Short answer, apply it yourself
6. Pick an AI product you use yourself where the person affected by a bad output can't easily complain about it. What's one adversarial prompt you'd want a red team to try on it before it ships?
Show hint
Think of a product where the affected person isn't the one holding the phone, a patient, a caller, a kid, a claimant.
Show answer
Model answer: A pharmacy chatbot that answers a caregiver's questions about a relative's medication. The relative taking the pills never hears the answer and can't correct it. I'd want a red team to try a subtly wrong dosage question phrased like a confident caregiver already knows the answer, to see if the bot corrects the caregiver or quietly agrees with them.
Before you close the answer
Why this works
Tests whether you think a red team is a formality run once before launch, or a standing practice built specifically for the fact that some people affected by a model's output have no way to report when it goes wrong. Most candidates describe testing for "bad outputs" and stop there.
Follow-up traps
"Isn't a strong content filter enough on its own?" Response: no, a filter tuned for profanity and explicit material catches keyword level bad words, not a gentle, entirely clean story about a death that still terrifies a four year old. That's exactly the kind of failure only an adversarial search finds.

"Doesn't red teaming just slow every release down?" Response: yes, and that cost is accepted on purpose. Two days of turnaround per release is a real trade against a subject who has no way to tell you when a story goes wrong after it ships.
If pressed
The severity bar was never zero misses. A production system answering thousands of prompts a day can't promise that. It's a tier based bar: only tier one failures, the genuinely frightening or unsafe ones, block a release on their own. Lower tier misses get logged and fixed in the next normal sprint, not treated as an emergency that halts shipping every time.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more