ConceptIntermediateResponsible AI & Advanced Practice / Responsible AI as a product requirement / #2

What safety requirements belong in every AI PRD regardless of feature?

ORDER the product is Lumenreel, a video streaming service, and its AI-written "previously on" episode recaps

Lumenreel streams licensed and original shows. Its AI recap tool writes the "previously on" summary that plays before a new episode. Amara Osei is the only person at the company whose job is to review every team's AI feature PRD before it ships, and she keeps a shared tablet on her desk that three different reviewers pass around during launch meetings.

The direct answer
Five things belong in every AI PRD no matter what the feature does: a named harm, a held-out eval set that tests for it, a named owner who gets paged when it fails, a kill switch that actually works, and a disclosure that the output is AI-made. Rank them by what's hardest to add after launch, not by how the team feels about them, and the named harm and its eval set come first, because you cannot build either one after the fact without redoing the work you skipped.
Do this, in order
  1. Name the specific harm this feature could cause before writing anything else.Why: every other requirement below depends on this one existing first.
  2. Build a held-out eval set that actually tests for that named harm.Why: without it, "we thought about safety" is a sentence, not a check.
  3. Name one owner who gets paged when the eval fails.Why: a failing check with nobody responsible for it is a check nobody actually watches.
  4. Build a kill switch that can pull the feature back in under a day.Why: this one is comparatively cheap to bolt on later, so it can wait behind the first three if a launch is truly rushed.
  5. Disclose to users when they're looking at AI-made output.Why: cheapest of the five to retrofit, but skipping it quietly erodes trust the other four can't repair alone.

How to answer this, stage by stage

The interviewer already knows "add safety requirements" is the right instinct. What they're testing is whether you can rank them under time pressure instead of listing five equally-weighted bullet points.

Stage 1
Pick one shipped feature to stand on
Say it like this
"I'll ground this in Lumenreel's AI recap tool, since a real feature keeps the five requirements from turning into a values speech."
Why this works
A concrete feature gives every abstract requirement a place to actually attach.
Stage 2
State your method up front
Say it like this
"I'll use ORDER: the outcome all five protect, which is hardest to undo, what depends on what, what's cheap to learn first, then the actual rank."
Why this works
Signals a ranking is coming, not a flat list, before you've said a single requirement.
Stage 3
Reframe the question
Say it like this
"This isn't really 'what should be in the PRD.' It's 'which of these can I still add after launch, and which one is gone for good if I skip it now.'"
Why this works
This is the actual insight the question is testing for, stated in one breath.
Stage 4
Name the five, in order
Say it like this
"Named harm, eval set, named owner, kill switch, disclosure. The first two you basically can't retrofit. The last three you can, in a pinch."
Why this works
This is the direct answer to the question, said plainly, before any story.
Stage 5
Show what skipping the top one cost
Say it like this
"At Lumenreel, the recap tool shipped with no named harm written down anywhere, and it hallucinated a plot detail about a fictional character that had to be publicly corrected."
Why this works
Compresses the whole story into a defensible, four-sentence proof.
Stage 6
Close on the rank
Say it like this
"So: name the harm, build the eval set, name an owner, then a kill switch, then a disclosure label. In that order, because that's the order you lose them in."
Why this works
A ranked close feels finished. A flat list of five things trails off.

Let's learn

Say a company ships eleven different AI features across eleven different teams in a year. Each team writes its own PRD, and each one has its own idea of what "think about safety" means.

For a while, that worked out fine at Lumenreel. Amara would spot-check maybe every third PRD that crossed her desk, mostly the ones a team had already flagged as "sensitive" themselves. The ones that passed her spot-check looked genuinely careful. So she trusted the teams that didn't ask for a review to have made the same call correctly.

Knowledge spark: what's a kill switch, here? A way to pull an AI feature's output back to a safe fallback, fast, without a full re-deploy. For a recap tool, that might mean falling back to a short, plain, human-written summary instead of the AI-generated one, the moment something's flagged.

The recap tool was one of the eleven. It had shipped fourteen months earlier with no rollback plan and no named owner, because nobody on that team had marked it as a feature that "needed" a safety section, and Amara hadn't pulled it for a spot-check.

Eleven AI features, one requirement at a time
11 6 0 Harm 4 Eval set 3 Owner 3 Kill switch 3 Disclosure 6
The recap tool had none of the five. Disclosure, the cheapest one, was the only category most teams bothered with at all.

At its worst: the recap for a popular drama's season finale invented a plot detail that never aired, describing a side character as having cheated on their partner in an episode where nothing of the kind happened. Viewers screenshotted it as a "leaked spoiler" before the studio's own social team caught it and had to post a correction.

The decision I would take back We let each team's PRD template ask "if applicable, describe safety considerations," leaving the applicable part up to the team itself. That made sense back when only one or two teams were shipping anything AI-touched and Amara could eyeball each one personally. It stopped making sense once eleven teams were shipping in parallel and "if applicable" quietly became "if we felt like it."

What I would leave alone: a feature like an AI-suggested watchlist ordering, where the worst case is a slightly worse row of recommendations, doesn't need the same five-item gate. Ranking the same five requirements as equally mandatory for every feature, regardless of what's actually at stake, would slow the harmless ones down for no real gain.

The recap tool didn't fail because nobody cared about safety. It failed because "safety considerations, if applicable" let eleven different teams each decide for themselves what applicable meant.

The lesson: a safety section that's optional by default isn't a safety requirement. It's a suggestion wearing a requirement's name.

Hand sketched icon list titled Five things every AI PRD needs. Five items: a scale icon labeled name the specific harm, a gauge icon labeled a held-out eval set for it, a box icon labeled a kill switch that works fast, a person icon labeled a named owner who gets paged, a document icon labeled a label for AI-made output.
Five things, and the recap tool's original PRD had a line about none of them.

Now here is the same thing as a story

The short version above is what you'd say defending this ranking on the spot. Read this one for how Amara actually found the gap.

Amara is the only person at Lumenreel whose full-time job is reviewing AI feature launches, across every team, not just her own. She's done it for three years, and for most of that time a spot-check felt like plenty.

She'd pull roughly every third PRD, mostly the ones a team had already flagged as touching something sensitive. The ones she checked were, almost without exception, thoughtful. So over time she stopped worrying much about the ones nobody flagged. If a team didn't think their feature needed a safety section, she trusted that judgment.

The recap tool's team never flagged it. Recaps felt low-stakes, a convenience feature, not a safety-adjacent one. It shipped fourteen months ago with a one-line "safety considerations: N/A" in a PRD nobody outside that team ever read closely.

Hand sketched timeline titled A PRD's safety lifecycle. Five milestones: draft, harm named. Red-team pass, 2 days. Eval built, week 1. Launch, gated. Post-launch audit, week 6, emphasized.
This is the lifecycle Lumenreel has now. The recap tool, fourteen months ago, skipped straight from draft to launch.

Then came the near miss. A season finale recap, generated fresh for that week's new episode, described a fan-favorite side character as having cheated on their partner. It never happened in the show. The model had blended two separate subplots from earlier episodes into one invented detail, confidently, the way these models do when nothing is stopping them.

It went out to a small percentage of viewers before someone on the studio's social team, prepping the same week's promotional post, noticed the recap didn't match the actual episode and escalated it within the hour. Lumenreel pulled the recap and posted a correction before it spread much further. Genuinely lucky timing, nothing more.

Hand sketched decision tree titled What skipping it actually costs. Root: which requirement got skipped. Four branches: kill switch missing leads to add it in a day, eval set missing leads to redo discovery three weeks, owner missing leads to nobody gets paged, harm never named leads to no eval set is even possible.
Two of these four are cheap to fix after the fact. Two of them are not, and the recap tool had skipped both of the expensive ones.

When Amara pulled the recap team's PRD afterward, there was no named harm anywhere in it, so there had never been anything to build an eval set against. Building one from scratch, after the fact, meant going back through eight months of past recaps by hand to find the pattern, three full weeks of work that could have been two days if the harm had been named on day one.

We did not almost lose one recap. We almost let an invented plot detail travel further than a quiet correction could catch, and the only reason it didn't was a studio employee's unrelated Tuesday task.

Hand sketched quadrant titled Ranking the five by what breaks first. Axes how hard to retrofit from easy later to nearly impossible later, and how bad if missing from minor to severe. Name the harm and eval set sit top right, severe and nearly impossible to retrofit. Kill switch and named owner sit lower left of them. Disclosure label sits lowest.
The two hardest to retrofit are also the two that matter most. That's not a coincidence, that's why the rank exists.

The team rewrote the PRD template that month: every AI feature, no exceptions, states a named harm before a single design mock gets drawn. Ranked in the order they're hardest to add back: named harm, eval set, named owner, kill switch, disclosure.

Hand sketched flow diagram titled What has to come first. Five boxes: name the harm highlighted, build the eval set, set the threshold, assign an owner, ship with a kill switch.
Everything else in the pipeline depends on the first box. Skip it, and the rest can't really happen at all.

I used to think "safety considerations, if applicable" was a reasonable trust to place in eleven capable teams. It took watching one of them, in complete good faith, decide a recap tool didn't need it, to see that the word "applicable" was doing all the actual work, and nobody had ever defined it.

ORDER, the five rankedNot a checklist of five equal boxes. ORDER is what forces you to say which one you lose for good if you skip it.

O
Outcome. What all five protect.
Viewer trust and the studio partnerships behind every licensed show, both of which a publicly corrected "spoiler" quietly damages.
Without a named outcome, ranking five requirements is just opinion.
R
Reversibility. Which one you lose for good.
Named harm and its eval set. Skip them at launch and rebuilding them later means redoing eight months of discovery by hand, three weeks instead of two days.
The hardest step, and the one that actually produces a ranking instead of a list.
D
Dependency. What unblocks what.
You cannot build a meaningful eval set until the harm is named. The order isn't a preference, it's forced by what the next step needs.
Some of the rank isn't judgment at all, it's just reality.
E
Evidence. What's cheap to learn first.
A one-day red-team pass, a handful of people trying to break the recap tool on purpose, would have surfaced this exact hallucination pattern before launch.
Cheap to learn early, expensive to learn from a viewer's screenshot instead.
R
Rank. State the order, defend the top.
Named harm, eval set, named owner, kill switch, disclosure. The top pick wins because everything else is retrofittable and this one isn't.
Closes the framework the same way it opened: with a defensible top pick, not a shrug.
Cost of building it upfront versus after the fact
0 2 days, upfront 21 days, after the fact
Naming the harm on day one is ten times cheaper than reconstructing it after a launch already needed a public correction.

The recap, one line per letter: outcome is viewer and studio-partner trust, reversibility is why named harm and eval set sit at the top, dependency is the eval set needing the named harm first, evidence is the cheap red-team pass that would have caught this, and rank closes with the actual five in order.

And if you want to be sure it really works, try it somewhere elseSame five letters, a gig-work marketplace instead of a streaming service. A completely different harm, and the hardest-to-retrofit item flips.

Sparefolk connects homeowners with freelance handypeople, and its AI tool drafts each job listing's description from a worker's photos and a short note about their experience. Ines Bouchard, who owns that feature, ran the same five requirements against it. The outcome: a homeowner's trust that the listing describes a worker's real qualifications. The named harm here isn't a hallucinated plot detail, it's the AI tool inventing a certification a worker never actually holds, since the model fills gaps in a thin note with plausible-sounding specifics. The hardest one to retrofit turned out to be disclosure, not the eval set: workers had already seen their AI-drafted listings go live for eight months with no label saying the description was AI-written, so freelancers had built their whole profile's voice around text they didn't write. Adding a disclosure label after that much time meant reopening a trust conversation with every worker on the platform at once, harder to undo than it looked on paper.

Hand sketched labeled parts diagram titled What's in a safety section. Center document icon labeled Safety Section, with five callouts: named harm, eval set, named owner, kill switch, disclosure.
Five parts, and which one is hardest to fix later depends on the feature, not on the list itself.

Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "named harm and eval set first, because you can't rebuild those after launch," and stop.
Cost: there's no time to build all five before a hard launch date. Say so honestly, and ship with named harm and a named owner at minimum, since those two are the ones you genuinely can't add back later.
The model gets better, for real: if the recap model's overall accuracy improves, the named-harm eval set still matters just as much, since a rare, better-hidden hallucination is exactly what a coarse eval set misses first.

Where people run it wrong.
They write all five as an equally-weighted checklist, which tells a rushed team nothing about what to cut under pressure.
They rank by how scary a requirement sounds instead of by what's actually hardest to add back later.
They assume "if applicable" is a fine default, when in practice it just means whichever team is busiest decides for itself.

How to use it live. When someone asks what belongs in every AI PRD, ask yourself one question before answering: if we skip this one and launch anyway, can we still add it next month for cheap. Whatever answers no goes first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what belongs in every AI PRD, ranked"?
Tap to flip
ANSWER
ORDER: name the outcome, rank by reversibility, map the dependencies, find cheap evidence, then state the rank.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Amara Osei, the only reviewer of every AI feature PRD across Lumenreel's eleven teams.
3 · THE OLD HABIT
What did Amara stop doing because the spot-checks kept looking fine?
Tap to flip
ANSWER
Questioning a team's own call that their feature "didn't need" a safety section. She trusted that judgment instead of checking every PRD.
4 · THE FIVE
What are the five universal requirements, in rank order?
Tap to flip
ANSWER
Named harm, eval set, named owner, kill switch, disclosure. Ranked by which one you can't add back after launch.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing "safety considerations, if applicable" into the PRD template and letting each team decide what applicable meant.
6 · THE NUMBER
Fill in the blank: rebuilding the missing eval set after launch took ___ days instead of the 2 it would have taken upfront.
Tap to flip
ANSWER
21 days. Roughly ten times slower than doing it before the feature shipped.
7 · THE REPLAY
Same recap tool, redesigned PRD template. What changes?
Tap to flip
ANSWER
The team names the hallucination risk on day one, builds a small eval set in two days, and the season-finale recap gets caught before it ever reaches a viewer.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and which requirement flips to hardest-to-retrofit there?
Tap to flip
ANSWER
Sparefolk's AI listing generator. There, disclosure turns out to be the hardest to retrofit, since workers had already built trust around undisclosed AI-written text.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer treats all five requirements as equally urgent for every feature, with no ranking between them.
  • True
  • False
Show hint
Look at the quadrant diagram sorting the five by how hard they are to retrofit.
Show answer
False. Named harm and its eval set rank first because they're nearly impossible to retrofit, while a kill switch and disclosure can wait if a launch is genuinely rushed.
Multiple choice
2. Why do "named harm" and "eval set" rank above "kill switch" in this answer?
  • A. They're cheaper to build than a kill switch.
  • B. Skipping them means redoing weeks of discovery later, while a kill switch can be added in about a day.
  • C. Legal requires them by law.
  • D. They were mentioned first in the original PRD template.
Show hint
Look at the "cost, upfront versus after the fact" bar chart.
Show answer
B. That's ORDER's reversibility step: rank by what breaks first and is hardest to undo, not by cost or convention.
Fill in the blank
3. Fill in the blank: of Lumenreel's eleven shipped AI features, only ___ had a named owner who'd get paged if the feature's eval failed.
Show hint
Look at the bar chart of eleven features by requirement.
Show answer
3. Disclosure was the most common of the five at 6 of 11, but named owner, eval set, and kill switch all sat at just 3.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Writing "safety considerations, if applicable" into the template. It made sense when only one or two teams shipped AI features and Amara could review each personally.
Short answer, where it wouldn't matter
5. Name a kind of AI feature where the full five-item gate genuinely isn't worth the overhead.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An AI-suggested watchlist ordering, where the worst case is a slightly worse row of recommendations, not a real harm worth the same five-item gate.
Short answer, apply it yourself
6. Pick an AI feature you've used. If its team had to cut two of the five requirements to hit a deadline, which two do you think they'd actually be safe cutting, and why?
Show hint
Ask which two could still be added back next month without redoing earlier work.
Show answer
Model answer: Most people land on kill switch and disclosure, the same two this answer ranks last, since both can be bolted on after launch without redoing any discovery work.
Before you close the answer
Why this works
Tests whether you can turn "think about safety" into a ranked, defensible list under time pressure, instead of a flat five-item checklist that gives a rushed team no guidance on what to cut.
Follow-up traps
"Isn't a kill switch more important than a disclosure label?" Response: probably, but both are retrofittable, so the rank between them matters less than the gap between them and named harm, which isn't retrofittable at all.

"What if the team genuinely doesn't know what harm a new feature could cause yet?" Response: that uncertainty is itself the reason to spend a day red-teaming before launch, not a reason to skip naming a harm entirely.
If pressed
Lumenreel's real fix ties the "named owner" field in the PRD template to an on-call rotation automatically, so naming an owner isn't just a text field, it actually pages a real person within minutes of a failed eval.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more