Critique a launch plan with no safety evaluation.
Brightledger is a personal finance app. Spend Coach AI reviews a user's transaction history and recommends specific spending cuts, like "you could save 40 dollars a month by canceling this subscription." Marcus Doyle, a senior growth PM, wrote the nine-page launch plan. Talia Renwick, who leads responsible-AI review, was asked to sign off on it two days before the leadership readout.
- Add a safety evaluation section with its own hard gate, before any rollout percentage is approved.Why: right now nothing in this plan can actually block a launch on safety grounds, only on app-store ratings or crash rate.
- Name the harm categories this feature can actually cause, specifically.Why: "test for safety" means nothing until someone writes down what going wrong looks like here.
- Build a golden set that includes financially vulnerable scenarios, not just typical ones.Why: the highest-harm case, someone in real financial distress getting bad advice, is exactly the one a convenience sample would miss.
- Pin every test result to the exact model version and date it came from.Why: a passing result with no version pin can't be reproduced or trusted six weeks from now.
- Ask who set the launch date and whether that date created pressure to skip this section.Why: this plan wasn't rushed by accident. Something drove the timeline, and it's worth naming out loud.
- Keep the existing rollout, marketing, and support sections largely as they are.Why: those sections are genuinely well built. The problem isn't the plan's quality, it's its missing page.
How to answer this, stage by stage
Six moves. A critique question rewards precision over volume, so keep each one tight.
Let's learn
Say a personal finance app builds a feature that reads a user's transaction history and tells them exactly where to cut spending. "You could save 40 dollars a month by canceling this subscription." Specific, personal, and delivered with total confidence.
Before this feature, users got a generic budgeting dashboard: categories, totals, a pie chart. It worked, in the sense that nothing about it could actively steer someone wrong. It also didn't do much.
The new launch plan for this feature runs nine pages. A five-stage rollout schedule, from 1% of users up to everyone. A four-week marketing calendar. A support runbook covering six common questions. A rollback trigger tied to app-store rating and crash rate.
Here is the turn. None of that tells you whether the actual advice is any good, or who it might hurt. A rollback trigger watching app-store ratings will catch complaints about a confusing button. It will not catch a recommendation that tells someone living paycheck to paycheck to cancel something that turns out to be a medical alert subscription, because that user might not leave a bad review. They might just quietly stop trusting the app, or worse, follow advice that makes a bad week worse.
At its worst: this ships, a user near a zero balance gets told to cut something they actually needed, and the first anyone hears about it is a support ticket routed as "billing confusion," not "the AI gave harmful advice," because nothing in the plan ever defined that second category.
What I would leave alone: the rollout schedule and rollback trigger are genuinely good, and shouldn't be rebuilt. A staged rollout with a rating-based rollback is the right tool for catching bugs and confusing UX. It was never going to be the right tool for catching bad financial advice, and it doesn't need to try.
The lesson: a launch plan's length is not evidence of its safety. Nine careful pages about rollout and marketing can sit right next to zero pages about harm, and nothing about the document's size will tell you that on a skim.
Now here is the same thing as a story
The short version above is the critique. Read this one for how Talia actually caught it.
Talia Renwick has led responsible-AI review at Brightledger for a year and a half. She reads launch plans the way an editor reads a manuscript: not for whether it's well written, but for what it quietly assumes.
Marcus Doyle, the growth PM who wrote the plan, had done real work. The rollout schedule was sensible. The marketing calendar was tight. The support runbook covered every question his team could think of. He sent it to Talia two days before a leadership readout tied to a Q3 metric: 300,000 users onboarded to the premium tier through Spend Coach recommendations.
Talia read all nine pages twice before she noticed what wasn't there. Not a thin safety section. No safety section. No mention of what kinds of advice could go wrong, no test set of real transaction histories, no mention of which model version had been checked against anything.
She pulled the launch plan template Brightledger had used for its last eight feature releases and checked how many had included a documented safety evaluation section at all. The pattern surprised her more than any single plan had.
We didn't wake up one day and decide safety evaluation didn't matter. We let it erode one launch at a time, each one a little faster than the last, until the section quietly stopped being required at all.
Talia asked Marcus directly what had happened to the safety section from the template. He hadn't removed it on purpose. Nobody had explicitly owned it in over a year, so each launch plan since had quietly copied the previous one's shortened version, a little thinner each time.
Talia mapped Spend Coach against Brightledger's other recent launches on the two axes that actually mattered: how much harm a wrong output could cause, and how much safety review it had actually received.
The fix wasn't rebuilding the plan. It was adding one required section, with its own owner and its own hard gate, and testing it against real vulnerable-user scenarios before the leadership readout, not after.
What I would tell myself, a year before this: a missing section doesn't announce itself. It just looks like eight good pages, and nobody double-checks a document that reads as finished.
AUDIT, the check Talia actually ranNot "does this look thorough." AUDIT forces you to ask what a thorough-looking document still leaves out.
The recap, one line per letter: ask is a growth metric with no safety stakeholder attached, uncover is a nonexistent eval set, demand is a missing version pin, isolate is a plan thorough everywhere except the one section that matters most, and test is running it yourself against vulnerable cases before signing anything.
And if you want to be sure it really works, try it somewhere elseSame five letters, a used-car trade-in tool instead of a finance app. A completely different field.
Larkfield Auto Exchange runs TradeValue AI, a tool that gives sellers an instant trade-in offer based on their car's condition photos and mileage. A launch plan for a new version, expanding it to older, higher-mileage vehicles, lands on a reviewer's desk with the same structure: rollout, marketing, support, no safety section.
Mapped onto AUDIT: ask is whether the expansion date is tied to a dealer-network growth target that created pressure to skip harder testing on older vehicles. Uncover is that there's no golden set of older, higher-mileage vehicle photos, only the newer-car set the original model was built on. Demand is no version pin on which model handled the older-vehicle expansion. Isolate is that the plan is thorough on rollout and dealer communications and silent on whether the pricing model actually works for the exact vehicles it's being expanded to cover. Test is running the tool against fifty real older, high-mileage listings before the expansion ships, checking specifically for systematic underpricing that could cost sellers real money.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "thorough everywhere except the one section that decides whether it should ship, and that section doesn't exist," and stop.
Cost: there's no time before the readout to build a full vulnerable-user golden set. Say so honestly, and start with ten manually reviewed real cases, even a small set, rather than approving with zero.
The model gets better, for real: if Spend Coach's overall recommendation accuracy improves in testing, that's still not a reason to skip the safety section. A more accurate model can still be confidently, specifically wrong for the one user who can least afford it.
Where people run it wrong.
They mistake a long, detailed launch plan for a safe one, when length has nothing to do with what's covered.
They let a rollback trigger built for bugs and ratings stand in for a safety gate, when it was never built to catch that kind of harm.
They treat a missing safety section as an oversight to note, rather than a hard blocker to the launch date itself.
How to use it live. If you're handed a document like this in an interview, don't critique it page by page. Ask what question this document was actually trying to answer, and check whether it answers that question at all. Here, the real question was never asked.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Wouldn't adding this section just delay the launch and hurt the Q3 metric?" Response: possibly a short delay, but the alternative is shipping financial advice to real people with zero evidence it's safe for the people it could hurt most.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Responsible AI as a product requirement
- #1 How do you turn a responsible AI principle into a testable product requirement?
- #2 What safety requirements belong in every AI PRD regardless of feature?
- #3 Describe how you would assess a feature for potential harm before building it.
- #4 Explain the difference between a safety issue and a quality issue.
- #5 How would you handle a feature that works well overall but poorly for one demographic?
- #6 What is a content policy and who should own it in a product organization?