Artifact critiqueIntermediateResponsible AI & Advanced Practice / Responsible AI as a product requirement / #12

Critique a launch plan with no safety evaluation.

AUDIT the artifact is a nine-page launch plan for Spend Coach AI, Brightledger's personal-finance nudging feature

Brightledger is a personal finance app. Spend Coach AI reviews a user's transaction history and recommends specific spending cuts, like "you could save 40 dollars a month by canceling this subscription." Marcus Doyle, a senior growth PM, wrote the nine-page launch plan. Talia Renwick, who leads responsible-AI review, was asked to sign off on it two days before the leadership readout.

The direct answer
This plan isn't missing a section. It's missing the one section that determines whether the other eight are worth reading: no named harm categories, no golden set of real spending scenarios, no test against financially vulnerable users, no version pin on which model was tested, and no bar that blocks launch on safety grounds. A rollout schedule and a rollback trigger tied to app-store ratings are not a safety evaluation. Don't approve this plan until that section exists and has its own hard gate, separate from marketing and support readiness.
Do this, in order
  1. Add a safety evaluation section with its own hard gate, before any rollout percentage is approved.Why: right now nothing in this plan can actually block a launch on safety grounds, only on app-store ratings or crash rate.
  2. Name the harm categories this feature can actually cause, specifically.Why: "test for safety" means nothing until someone writes down what going wrong looks like here.
  3. Build a golden set that includes financially vulnerable scenarios, not just typical ones.Why: the highest-harm case, someone in real financial distress getting bad advice, is exactly the one a convenience sample would miss.
  4. Pin every test result to the exact model version and date it came from.Why: a passing result with no version pin can't be reproduced or trusted six weeks from now.
  5. Ask who set the launch date and whether that date created pressure to skip this section.Why: this plan wasn't rushed by accident. Something drove the timeline, and it's worth naming out loud.
  6. Keep the existing rollout, marketing, and support sections largely as they are.Why: those sections are genuinely well built. The problem isn't the plan's quality, it's its missing page.

How to answer this, stage by stage

Six moves. A critique question rewards precision over volume, so keep each one tight.

Stage 1
Name what the artifact actually is
Say it like this
"This is a nine-page launch plan for an AI feature that gives specific spending advice. I'll critique it as that specific document, not as a launch plan in the abstract."
Why this works
Grounds the critique in the real document instead of a general lecture on launch readiness.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who set the timeline, uncover what eval set exists, demand the version pin, isolate what's missing, and say how I'd test it myself before signing off."
Why this works
Signals a repeatable critique method, not a one-off gut reaction to the document.
Stage 3
Name what's actually missing
Say it like this
"The rollout schedule, marketing calendar, and support runbook are all thorough. There's no safety evaluation section at all: no harm categories, no golden set, no vulnerable-user testing."
Why this works
Names the specific gap instead of a vague "this needs more rigor."
Stage 4
Ask who set the timeline, and why
Say it like this
"This plan is tied to a Q3 board metric. That's worth asking about directly, since a deadline tied to a leadership readout is exactly the kind of pressure that quietly drops the section with no internal champion."
Why this works
Names the incentive behind the gap instead of treating it as a random oversight.
Stage 5
Say what you'd test yourself, before signing off
Say it like this
"Before I'd approve this, I'd run the model myself against ten real transaction histories from users near a zero balance, and check whether any recommendation tells them to cut something they actually need."
Why this works
Moves from criticism to a concrete action you'd take before the plan goes forward.
Stage 6
Close on the one line
Say it like this
"So: good plan, missing the one page that decides whether it should exist. I wouldn't sign off until that page has its own hard gate."
Why this works
Leaves the interviewer with a clear verdict, not just a list of concerns.

Let's learn

Say a personal finance app builds a feature that reads a user's transaction history and tells them exactly where to cut spending. "You could save 40 dollars a month by canceling this subscription." Specific, personal, and delivered with total confidence.

Before this feature, users got a generic budgeting dashboard: categories, totals, a pie chart. It worked, in the sense that nothing about it could actively steer someone wrong. It also didn't do much.

Knowledge spark: what's a launch plan supposed to prove? Not just that a feature works. That it's ready for the specific ways it can fail, and that someone has checked those ways on purpose, before real users find them first.

The new launch plan for this feature runs nine pages. A five-stage rollout schedule, from 1% of users up to everyone. A four-week marketing calendar. A support runbook covering six common questions. A rollback trigger tied to app-store rating and crash rate.

Here is the turn. None of that tells you whether the actual advice is any good, or who it might hurt. A rollback trigger watching app-store ratings will catch complaints about a confusing button. It will not catch a recommendation that tells someone living paycheck to paycheck to cancel something that turns out to be a medical alert subscription, because that user might not leave a bad review. They might just quietly stop trusting the app, or worse, follow advice that makes a bad week worse.

Launch plan section coverage, items documented out of the standard checklist
6 3 0 Rollout Marketing Support Safety eval 5/5 4/4 6/6 0/6
Every other section is fully documented. The safety evaluation section doesn't exist at any level, not partially, not thinly. It's absent.

At its worst: this ships, a user near a zero balance gets told to cut something they actually needed, and the first anyone hears about it is a support ticket routed as "billing confusion," not "the AI gave harmful advice," because nothing in the plan ever defined that second category.

The decision I would take back We let the launch checklist template stay the same as it was for low-stakes features, marketing calendar, support runbook, rollback trigger, and never added a required safety section for features that give direct financial advice. That template made sense for a theme picker or a bill reminder. It stopped making sense the moment a feature started telling people what to do with their money.

What I would leave alone: the rollout schedule and rollback trigger are genuinely good, and shouldn't be rebuilt. A staged rollout with a rating-based rollback is the right tool for catching bugs and confusing UX. It was never going to be the right tool for catching bad financial advice, and it doesn't need to try.

A launch plan can be thorough and still be missing the one page that decides whether the feature should exist at all.

The lesson: a launch plan's length is not evidence of its safety. Nine careful pages about rollout and marketing can sit right next to zero pages about harm, and nothing about the document's size will tell you that on a skim.

Now here is the same thing as a story

The short version above is the critique. Read this one for how Talia actually caught it.

Talia Renwick has led responsible-AI review at Brightledger for a year and a half. She reads launch plans the way an editor reads a manuscript: not for whether it's well written, but for what it quietly assumes.

Hand sketched icon list titled What the launch plan actually has. Four items: a gauge icon labeled rollout schedule five stages, a document icon labeled marketing calendar four weeks, a box icon labeled support runbook six scenarios, a red question mark box labeled safety evaluation not present.
Three items, fully built. The fourth doesn't have a thin version. It just isn't there.

Marcus Doyle, the growth PM who wrote the plan, had done real work. The rollout schedule was sensible. The marketing calendar was tight. The support runbook covered every question his team could think of. He sent it to Talia two days before a leadership readout tied to a Q3 metric: 300,000 users onboarded to the premium tier through Spend Coach recommendations.

Talia read all nine pages twice before she noticed what wasn't there. Not a thin safety section. No safety section. No mention of what kinds of advice could go wrong, no test set of real transaction histories, no mention of which model version had been checked against anything.

Hand sketched comparison diagram titled What's in the plan vs what's missing. Left panel, a document icon labeled Fully written, caption rollout marketing support. Right panel, a red question mark box labeled Not written, caption any safety evaluation at all.
Both of these were true on the same nine pages, at the same time.

She pulled the launch plan template Brightledger had used for its last eight feature releases and checked how many had included a documented safety evaluation section at all. The pattern surprised her more than any single plan had.

Share of Brightledger's last eight launches with a documented safety evaluation
100% 50% 0 Spend Coach
Each launch had slightly less safety documentation than the one before it, until this one had none at all. Nobody decided this all at once.

We didn't wake up one day and decide safety evaluation didn't matter. We let it erode one launch at a time, each one a little faster than the last, until the section quietly stopped being required at all.

Hand sketched flow diagram titled Where the plan skips a step. Five boxes: build feature, write plan, safety check highlighted, leadership signs, ship.
The third box in this pipeline is the one Marcus's plan quietly skipped straight past.

Talia asked Marcus directly what had happened to the safety section from the template. He hadn't removed it on purpose. Nobody had explicitly owned it in over a year, so each launch plan since had quietly copied the previous one's shortened version, a little thinner each time.

Hand sketched labeled parts diagram titled What this launch plan needs added. Center document icon labeled Launch Plan, with four callouts: vulnerable set, named harms, version pin, hard block bar.
Four additions. None of them touch the rollout, marketing, or support sections, which stay exactly as they are.

Talia mapped Spend Coach against Brightledger's other recent launches on the two axes that actually mattered: how much harm a wrong output could cause, and how much safety review it had actually received.

Hand sketched quadrant titled This launch, next to recent ones. Axes safety review received and harm if advice is wrong. Spend Coach sits top left, severe harm and no review. Theme picker sits bottom right, small harm and thorough review.
Spend Coach sits exactly where nothing should: the highest possible harm, the least review of anything on the list.

The fix wasn't rebuilding the plan. It was adding one required section, with its own owner and its own hard gate, and testing it against real vulnerable-user scenarios before the leadership readout, not after.

What I would tell myself, a year before this: a missing section doesn't announce itself. It just looks like eight good pages, and nobody double-checks a document that reads as finished.

AUDIT, the check Talia actually ranNot "does this look thorough." AUDIT forces you to ask what a thorough-looking document still leaves out.

A
Ask who paid for it. Whose timeline is this.
The plan is tied to a Q3 board metric on premium-tier onboarding, owned by growth, with no safety stakeholder named anywhere in it.
Names the pressure behind the gap instead of treating it as a random accident.
U
Uncover the eval set. There isn't one.
No golden set of transaction scenarios, no named harm categories, nothing distinguishing a typical user from a financially vulnerable one.
A missing eval set is itself a finding, not just an absence.
D
Demand the version pin. There isn't one either.
The plan never names which model version was tested, on what date, or by whom.
Without this, even a claimed test result couldn't be trusted or reproduced later.
I
Isolate what's missing. The whole finding.
Rollout, marketing, and support are fully documented. Safety evaluation is entirely absent, not thin, not partial.
What a thorough-looking document leaves out is usually the most important thing about it.
T
Test it yourself, before it reaches a signature.
Run the model against ten real, financially vulnerable transaction histories before approving the plan, not after launch.
Turns a critique into a concrete action taken before the risk ships.

The recap, one line per letter: ask is a growth metric with no safety stakeholder attached, uncover is a nonexistent eval set, demand is a missing version pin, isolate is a plan thorough everywhere except the one section that matters most, and test is running it yourself against vulnerable cases before signing anything.

And if you want to be sure it really works, try it somewhere elseSame five letters, a used-car trade-in tool instead of a finance app. A completely different field.

Larkfield Auto Exchange runs TradeValue AI, a tool that gives sellers an instant trade-in offer based on their car's condition photos and mileage. A launch plan for a new version, expanding it to older, higher-mileage vehicles, lands on a reviewer's desk with the same structure: rollout, marketing, support, no safety section.

Mapped onto AUDIT: ask is whether the expansion date is tied to a dealer-network growth target that created pressure to skip harder testing on older vehicles. Uncover is that there's no golden set of older, higher-mileage vehicle photos, only the newer-car set the original model was built on. Demand is no version pin on which model handled the older-vehicle expansion. Isolate is that the plan is thorough on rollout and dealer communications and silent on whether the pricing model actually works for the exact vehicles it's being expanded to cover. Test is running the tool against fifty real older, high-mileage listings before the expansion ships, checking specifically for systematic underpricing that could cost sellers real money.

The same labeled parts pattern reused for a second product: a vulnerable-scenario set, named harms, a version pin, and a hard block bar, this time for a car trade-in pricing tool.
Same four additions. A pricing tool needs them for exactly the same reason a spending-advice tool does.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "thorough everywhere except the one section that decides whether it should ship, and that section doesn't exist," and stop.
Cost: there's no time before the readout to build a full vulnerable-user golden set. Say so honestly, and start with ten manually reviewed real cases, even a small set, rather than approving with zero.
The model gets better, for real: if Spend Coach's overall recommendation accuracy improves in testing, that's still not a reason to skip the safety section. A more accurate model can still be confidently, specifically wrong for the one user who can least afford it.

Where people run it wrong.
They mistake a long, detailed launch plan for a safe one, when length has nothing to do with what's covered.
They let a rollback trigger built for bugs and ratings stand in for a safety gate, when it was never built to catch that kind of harm.
They treat a missing safety section as an oversight to note, rather than a hard blocker to the launch date itself.

How to use it live. If you're handed a document like this in an interview, don't critique it page by page. Ask what question this document was actually trying to answer, and check whether it answers that question at all. Here, the real question was never asked.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "critique a launch plan with no safety evaluation"?
Tap to flip
ANSWER
AUDIT: ask, uncover, demand, isolate, test. Isolate is the step that names the actual finding.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Talia Renwick, who leads responsible-AI review, and Marcus Doyle, the growth PM who wrote the launch plan.
3 · THE HABIT
What did each launch plan quietly do to the safety section over the last eight releases?
Tap to flip
ANSWER
Each one copied the previous plan's shortened version, a little thinner each time, until it disappeared completely.
4 · THE FINDING
What's the actual gap this critique identifies?
Tap to flip
ANSWER
No safety evaluation section at all: no harm categories, no golden set, no vulnerable-user testing, no version pin.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Keeping the same launch checklist template for a financial-advice feature that low-stakes features used, with no required safety section and no named owner for it.
6 · THE NUMBER
Fill in the blank: the safety evaluation section scored ___ out of 6 required items, versus 5 of 5, 4 of 4, and 6 of 6 for the other three sections.
Tap to flip
ANSWER
0 out of 6. Every other section was fully documented; this one didn't exist at any level.
7 · THE FIX
Same launch plan, with the safety section added. What changes before the readout?
Tap to flip
ANSWER
The model gets tested against real financially vulnerable transaction histories first, with a hard gate that can block launch on safety grounds, separate from the rollback trigger.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's missing there?
Tap to flip
ANSWER
Larkfield Auto Exchange's TradeValue AI, expanding to older vehicles. It's missing a golden set of older, high-mileage listings and any check for systematic underpricing.

Check yourself Score: 0 / 0

True or false
1. True or false: the rollback trigger tied to app-store rating and crash rate would likely catch a case where the AI recommends cutting an essential expense.
  • True
  • False
Show hint
Think about what actually generates an app-store review or a crash log.
Show answer
False. A user quietly stops trusting or gets hurt by bad advice; that rarely shows up as a crash or a low star rating specifically about the recommendation itself.
Multiple choice
2. Why does this critique treat the missing safety section as the central finding, rather than one item on a list of gaps?
  • A. Because the other three sections are also incomplete.
  • B. Because it's the one section that determines whether the feature should ship at all, not just how smoothly it ships.
  • C. Because leadership specifically asked for a safety section by name.
  • D. Because the marketing calendar depends on the safety section's results.
Show hint
Look at the direct answer and the Isolate step.
Show answer
B. Rollout, marketing, and support all assume the feature is safe to ship. Nothing in the plan actually checks that assumption.
Fill in the blank
3. Fill in the blank: over the last eight launches, the share with a documented safety evaluation fell from 100% down to ___%, this launch's share.
Show hint
Look at the declining line chart.
Show answer
0%. The decline happened gradually across eight quarters, not all at once, which is why nobody caught it sooner.
Short answer, where it wouldn't matter
4. Name a Brightledger feature where this level of safety-section scrutiny genuinely would not be necessary.
Show hint
Look at the quadrant diagram's lowest-harm item.
Show answer
Model answer: The app's theme picker. A wrong recommendation there costs nothing more than a mismatched color scheme.
Short answer, apply it yourself
5. Pick a document or plan you've reviewed yourself. What section was thoroughly written that quietly assumed something never actually got checked?
Show hint
Think about a plan that read as complete but skipped validating its own core assumption.
Show answer
Model answer: A common one: a detailed marketing launch plan that assumes a feature is stable and ready, with no actual QA sign-off attached anywhere in the document.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Using the same launch checklist template regardless of feature risk. It made sense while most features were low-stakes, like a theme picker or a bill reminder.
Before you close the answer
Why this works
Tests whether you can spot a structural gap inside a document that reads as complete, and whether you'll name the incentive behind that gap instead of treating it as a random oversight.
Follow-up traps
"Isn't a rollback trigger enough of a safety net on its own?" Response: no, since it only catches harm that shows up as a crash or a low rating, and the harm here is quiet, personal, and unlikely to generate either.

"Wouldn't adding this section just delay the launch and hurt the Q3 metric?" Response: possibly a short delay, but the alternative is shipping financial advice to real people with zero evidence it's safe for the people it could hurt most.
If pressed
Brightledger's eventual safety section required the golden set to be refreshed every quarter with real, anonymized transaction histories from users near a zero balance specifically, not just a one-time set built at launch.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more