ConceptFoundationalEval-Driven Specification / Writing a PRD for an AI feature / #1
What sections does an AI PRD need that a standard PRD does not?
The direct answer
Write the guardrail section first: the exact list of things the model must never say. A model swap you can redo next month. A listing description with steering language in it, already sent to outside listing sites, you cannot call back. Only once that section exists do an eval spec, a failure behavior section, and a staged rollout plan even make sense to write. A PRD that opens with a launch date instead of that list is a demo plan wearing a PRD's cover page.
Write the sections in this order
Write the guardrail section first, the exact list of things the model must never say.Why: it is the one thing here you cannot undo once a bad description reaches a site you don't control.
Build the eval golden set and rubric next, before anyone argues about which model or prompt is "good."Why: you cannot set a rollout threshold, or judge any candidate, against a "good" that was never written down.
Write the failure behavior section: what the model does when it isn't sure.Why: a wrong description that sounds confident is worse than one that visibly asks a person to check it.
Write the data and consent section, covering agent photos and past listings.Why: that has to be settled before a single photo trains, prompts, or few-shots anything.
Write the staged rollout and kill switch section, sized against the eval pass rate.Why: this is where you buy the right to be wrong small, before you're wrong at full scale.
Write the cost and latency budget last.Why: it's the easiest thing in the whole document to change after launch, so it doesn't earn a place near the top.
How to answer this, stage by stage
Six moves, in the order I'd actually say them. This question tempts you to recite a table of contents. Don't. Rank it, and defend the top line.
1
Reframe the question before you answer it
Say it like this
"Before I list sections, here's what I think this question is actually asking. It's not 'what extra headings go in the doc.' It's 'which extra section can't wait,' because some of them protect you from a mistake you can undo, and one or two protect you from a mistake you can't."
Why this works
Shows the interviewer you're about to rank, not recite a checklist you memorized.
2
Name the outcome every extra section is protecting
Say it like this
"Everything extra in an AI PRD is trying to protect one thing: an agent publishes a description they trust enough to use as is, and nothing ever goes out that a compliance person catches after the fact. Not 'the model sounds fluent.' That."
Why this works
Without a stated outcome, ranking the sections is just opinion wearing a plan's clothes.
3
Say which missing section is hardest to undo, and lead with it
Say it like this
"If I only get to add one section a standard PRD doesn't have, it's the guardrail section: the exact list of things the model must never say. A caption model, I can swap again next month. A listing that told a buyer a house is 'perfect for a young family,' already sitting on three outside listing sites, I can't quietly take that back."
Why this works
This is the actual answer to the question, not a hedge dressed up as a list.
4
Order the rest by what depends on what
Say it like this
"After guardrails, the eval spec has to exist before I can write the rollout section, because I can't set a go or no go number for adding more agents if I never defined what 'good' looks like in the first place. And data and consent has to exist before either of those, because I don't get to build a golden set from listing photos I don't have the right to use yet."
Why this works
This is what dependency actually tests: telling apart what reality forces from what you get to choose.
5
Say what you'd check cheaply before writing the full document
Say it like this
"Before I write twenty pages, I run sixty real past listings through whatever prompt already exists and read the output myself. That's an afternoon of work, and it told us one in twelve leaned on language a fair housing reviewer would flag. That's worth more than a week of arguing about section order in a doc nobody's tested against a real listing yet."
Why this works
Cheap evidence beats a confident guess about what the PRD needs to cover.
6
Give the order, and close on the one line
Say it like this
"So the order is: guardrails, eval spec, failure behavior, data and consent, staged rollout, then cost and latency last, because that's the one I can change on a Tuesday without anyone outside the company ever knowing I changed it. If you only remember one line: write the section you can't take back first, and write the section you can fix anytime last."
Why this works
Ends the answer on the actual decision, not a summary of everything that was said.
Let's learn
Say a real estate company builds a tool that writes listing descriptions. An agent uploads photos and types three or four rough notes, "3bd 2ba, updated kitchen, big backyard," and the tool writes the polished paragraph that goes on the listing.
A normal PRD for a normal feature at this company covers six sections: the problem, the users, the requirements, the metrics, the timeline, and what's out of scope. Halima Bassey, a product lead, has written that PRD eleven times in four years, for a saved search alert, a commission dashboard, a kitchen renovation calculator. All eleven shipped on time. None of them ever needed a seventh section.
She writes the same six sections for this one too. Under requirements, one paragraph says the model reads the photos and the notes and drafts a description. Nothing in the doc says what happens if it's wrong.
Knowledge spark: what a guardrail rule actually is
A plain sentence telling the model something it must never do. Not "be careful with neighborhood language." "Never describe who a home is a good fit for, and never guess at a school, a commute, or a neighborhood's character." One rule, one sentence, checkable by a person reading the output.
Here is the turn. The extra mistakes an AI feature makes are not the real problem. Every feature ships with some mistakes. The real problem is that this one publishes its mistakes somewhere a normal bug fix can't reach.
The wrong word was never the real risk. Not being able to call it back once a stranger has already read it is.
Say the team skips the guardrail section and ships fast. The model writes fluent, warm copy, and about one in twelve descriptions leans on a phrase like "perfect for a young family" or "ideal for empty nesters," because that's the kind of sentence that reads well and the model was never told not to write it. Some of those phrases sound harmless. Some of them describe who a house is for, which is exactly the kind of language that gets a real brokerage a real fine. And by the time anyone notices, the description is already synced to two or three outside listing sites, cached, and read by people who never saw it on the company's own page at all.
None of the other sections are worth writing until this one exists
Before any of this ships, a real eval spec gets written: two hundred real past listings, across houses, condos, land, and multi family, each one scored by a fair housing trained reviewer and a copy editor for what the description should actually say. Then a staged rollout: internal review only for two weeks, five percent of agents who opt in for three weeks, twenty five percent for two weeks, then everyone.
Knowledge spark: what a golden set is
A stack of real examples with a right answer already agreed on by a person. Not a guess at what "good" means. A specific set you can run any new model or prompt against and get a real score back.
Guardrail flag rate, sample test to full rollout
The rate doesn't fall because the model gets smarter mid rollout. It falls because each stage is small enough to catch what the guardrail rules missed and fix the prompt before more agents see it.
The choice I would take back
Eight months ago, the engineer who first built this prompt told the model to invent neighborhood color, lines like "perfect for a young family" or "ideal for professionals," to make the copy read less robotic in demos. It worked. Everyone in the room liked it. Nobody was asking "what could this actually say once it ships" at the time, because nothing had shipped yet. I would take that default back and constrain the model to facts the agent actually gave it. No invented amenities, no guessing who the house is for.
What I would leave alone. GreenLatch also has a small internal tool that suggests a price range to the agent. It's never published, never shown to a buyer. That one doesn't need a guardrail section or a staged rollout. If it's wrong, the agent just doesn't use the number. Nothing leaves the building before a person looks at it anyway.
The lesson. We wrote AI PRDs like software PRDs because writing was the part that felt familiar. A normal PRD assumes any mistake can be patched next release. An AI PRD has to assume some mistakes get read by a stranger before you even find out about them.
The line in review that stopped the meeting
You don't need this to answer the question. Read it slower, when you want to feel why the order matters and not just recite it.
Halima Bassey has written product requirements at GreenLatch Realty for four years. Eleven launches, all on the same six-section template the company has used since before she got there. Nobody has ever pushed back on that template. Why would they. It has never once let anyone down.
In March her director hands her the brief for Draft Desk, the tool that turns an agent's rough notes into a full listing description. Halima writes the PRD the way she's written the last eleven: problem, users, requirements, metrics, timeline, out of scope. Under requirements, one paragraph. The model reads the photos, reads the notes, drafts the copy. She schedules the review for the following Tuesday.
For those two weeks, Draft Desk is the best thing anyone in the office has seen. A designer runs three bullet points through it at a stand up, just for fun, and the room actually laughs at how good the sentence sounds. Her director starts using the word "launch" in messages that used to say "prototype."
Then comes the review meeting. A backend engineer, reading the requirements section for the third time, asks one question. "What happens when it's wrong?"
Halima goes to answer and realizes she doesn't have one. The template has a section for what the feature does. It has no section for what the feature does when it fails, because none of her last eleven features ever failed in public. A saved search alert that's wrong just alerts you about the wrong house. You close it. This one writes a sentence, and the sentence leaves the building.
Her director's first instinct sounds reasonable. "We're three weeks from the fall listing push. Let's ship the happy path and patch anything that comes up." It's the kind of line that sounds fine in a status update.
Halima says no, and the reason matters more than the no itself. She grabs a marker and draws two doors on the whiteboard instead of arguing about it.
She drew this instead of debating the deadline
Which model runs behind Draft Desk, she can change that door as many times as she wants. The launch date, same door. But a description that's already synced to an outside listing site is a door bolted shut. Getting a wrong sentence off somebody else's server, once it's cached and re-published, isn't a patch. It's a much slower, much more expensive kind of fix, and some of it never fully comes back.
They didn't ship a badly worded sentence. They shipped a sentence nobody could unsend.
So the plan runs in the other order. The guardrail rules go up on the wall that afternoon, six sentences, each one a thing the model must never say. The eval spec starts the next morning, two hundred real listings, scored by hand. The staged rollout follows the eval, not the calendar: internal only, then five percent, then twenty five, then everyone, each stage waiting for the flag rate to clear its own bar before the next one opens.
By week eight, Draft Desk is live for all three hundred and forty agents. The flag rate caught in review before publish has dropped from eight percent in that first sample test to under one percent. And the number that actually matters: zero guardrail flagged descriptions ever reached a live listing. Every one of them got caught inside the company, before it left.
The thing Halima would tell herself, looking back: she wrote eleven PRDs assuming every mistake could wait for the next release. Draft Desk doesn't wait. It publishes.
ORDER, run against a table of contents instead of a calendar
This is a prioritization question in a documentation costume. FLIPS looks for the moment someone's behavior snaps after a product changes. Nothing has shipped yet here, nothing has changed on anyone's desk. The whole question is which sections earn a place, and in what order you'd write them, so the framework is ORDER.
ORDER, worked against a document instead of a deadline
O, outcome. Every extra section in this PRD competes to protect one thing: an agent publishes a description they trust enough to use as is, and nothing goes out that a person would need to walk back later. Not "the model sounds fluent." A buyer never reads model fluency.
R, reversibility. Skipping the guardrail section is the hardest mistake here to undo, because a bad phrase reaches outside listing sites and gets cached before anyone inside the company sees it. Skipping the cost and latency section is the easiest, you can always retune that next sprint. So the hard to undo gap gets written first, and the easy to undo one waits.
D, dependency. The rollout section can't be written before the eval spec exists, because a rollout threshold means nothing without a defined "good" to measure against. The eval spec can't be built before the data and consent section is settled, because you can't build a golden set out of agent photos you don't yet have the right to use. Neither of those is a judgment call. Reality forces both.
Knowledge spark: what a shadow test is
Running a model on real material the normal way, scoring what it produces, and never showing that output to anyone outside the team. The customer sees nothing new the whole time. It's a rehearsal with real material, not a performance.
E, evidence. Sixty real past listings, run through the existing prompt and read by a person, for the cost of one afternoon, tells you more about what the PRD actually needs to cover than a week of debating section order in the abstract.
R, rank. Guardrail rules first, because they're the hardest to undo. Eval spec and data consent second, because they're forced by dependency, not preference. Failure behavior third. Staged rollout fourth, sized against the eval pass rate. Cost and latency last, because it's the one section you can always rewrite without anyone outside the building noticing.
The check that makes ORDER honest
Swap what's public and watch the order move. If Draft Desk's output only ever sat in an internal draft folder, and an agent had to copy and paste it out by hand, a wrong phrase would cost a click to fix and the guardrail section could sit fourth on this list instead of first. It isn't the wording itself that earns it the top spot. It's that a stranger already reads it before anyone inside the company gets a second look. Change what's actually irreversible, and the order changes with it, which is exactly what should happen.
Run it somewhere a government file replaces a search index
A farm co-op builds a tool that turns a field worker's voice memo into a pesticide application log, filed with the state agriculture department. Same shape of question, a different kind of already-submitted record.
O. Every application gets logged accurately, the same day, in a form the state accepts. Not "the transcript reads cleanly." Whether the filed record is right.
R. A filed government report is the hardest thing here to undo. You can't quietly edit a submission after the fact, you need a formal correction, and by then an inspector may already have the wrong version. Which transcription model drafts it, you can swap that again next season.
D. The guardrail rules, never let the model guess a chemical name or an amount it didn't hear clearly, have to exist before the eval golden set, since the golden set needs real examples of unclear audio to test against.
E. Shadow test the model against thirty already-approved past logs, read by a licensed applicator, before it ever drafts one that actually gets filed.
R. Guardrails first, eval spec by crop and chemical type next, shadow test, then a staged rollout by farm. Already filed logs never get quietly re-drafted. Ever.
Swap the trigger and it still runs
The launch date moves up to two weeks instead of two months. The order doesn't change, only the slack. Guardrails and the eval spec still go first, you just run the sixty sample test and the rubric in parallel instead of one after the other.
The AI vendor doubles the per-description price right before launch. Same order. The eval spec still goes first, now to prove a cheaper model clears the same bar, not because the schedule got tighter.
The new model turns out to make fewer mistakes than the old one, in early testing. Doesn't move the guardrail section. Better is still a guess until the eval set says so, and a wrong phrase at scale costs the same either way.
Where people run it wrong
Reading a handful of outputs and calling the prompt "basically fine," instead of running the sixty sample test.
Writing the rollout plan and the launch date before the eval spec exists, so nobody can actually say what "ready" means.
Treating every listing the same, when a wrong phrase on a house in a residential neighborhood carries far more risk than one on an empty commercial lot.
If you're asked this cold
Say the outcome out loud before you name a single section. "Everything extra in this PRD is protecting one thing: agents publishing something they trust." Ten seconds, and it gives you a shelf to hang every section on, instead of reciting a list.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits "what sections does an AI PRD need," and why not FLIPS?
Tap to flip
ANSWER
ORDER, for prioritizing which extra section matters most. FLIPS finds the moment a person's habit snaps after something changes. Nothing has shipped yet here, the whole question is what order you'd write the sections in before anything does.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Halima Bassey, a product lead who has written eleven PRDs at GreenLatch Realty on the standard six-section template, all shipped on time.
3 · WHAT NO ONE HAD TO ASK BEFORE
Before Draft Desk, what question never came up in one of Halima's PRD reviews?
Tap to flip
ANSWER
What happens when it's wrong. Every earlier feature could be patched next release. None of them published something to a stranger before a person inside the company could check it.
4 · THE SECTION YOU CAN'T TAKE BACK
Which section is hardest to undo if it's missing, and why?
Tap to flip
ANSWER
The guardrail section. A caption model can be swapped again next month. A description with steering language, already sitting on outside listing sites, cannot be quietly called back.
5 · THE OLD DECISION
What choice from Draft Desk's early build would you take back, and why did it make sense then?
Tap to flip
ANSWER
The original prompt told the model to invent neighborhood color, like "perfect for a young family," to sound less robotic in demos. It worked great on stage. Nobody was asking what it could actually say once it shipped, because nothing had shipped yet.
6 · THE NUMBER
The sixty sample test found about ______ descriptions leaning on language a fair housing reviewer would flag.
Tap to flip
ANSWER
About 1 in 12, roughly 8 percent. That number is what got the guardrail section written before anything else.
7 · THE REPLAY
Same launch, sections written in the right order: what does week eight look like?
Tap to flip
ANSWER
Draft Desk is live for all 340 agents, the flag rate caught before publish has dropped from 8 percent to under 1 percent, and zero guardrail-flagged descriptions ever reached a live listing.
8 · THE TRANSFER
Section 4 runs ORDER on a completely different product. Which one, and what plays the role the guardrail section plays here?
Tap to flip
ANSWER
A farm co-op's tool that drafts pesticide application logs for a state filing. The equivalent is the rule against the model guessing a chemical name or amount it didn't clearly hear, since a filed government report can't be quietly re-drafted.
Check yourself Score: 0 / 0
True or false
1. True or false: the cost and latency budget section should be written before the guardrail section in this PRD. Say why.
True
False
Show hint
Ask which of the two is easiest to change after launch without anyone outside the company noticing.
Show answer
False. Cost and latency is the easiest section to change after launch, so it goes last. The guardrail section goes first because it protects the one mistake nobody can quietly undo.
Fill in the blank
2. The sixty sample test found about ______ descriptions using language a fair housing reviewer would flag.
Show hint
It's the fraction that made Halima write the guardrail section before anything else.
Show answer
1 in 12, about 8 percent. That's the number the whole reversibility argument leans on, small enough to feel manageable, common enough that ignoring it would have reached real listings within days.
Multiple choice
3. Which of these has to exist before Halima can write the rollout section?
A. The cost and latency budget
B. The eval golden set and rubric
C. The out of scope list
D. The launch announcement email
Show hint
You can't set a go or no go number for the next rollout stage without a defined "good" to measure against.
Show answer
B. A rollout threshold is meaningless without an eval spec to measure it against. That's the dependency the D in ORDER is testing.
Short answer
4. Name a feature at GreenLatch where this same extra-sections treatment genuinely would not matter, and say why.
Show hint
Look for whatever a person always checks before it goes anywhere.
Show answer
Model answer: "The internal price-suggestion tool. It's never published, never shown to a buyer. If it's wrong, the agent just doesn't use the number. Nothing leaves the building before a person looks at it anyway, so a guardrail section and a staged rollout wouldn't buy anything here."
Short answer, apply it yourself
5. Pick an AI feature you use or are building. If it shipped something wrong today, what's the one thing about it you could never quietly undo?
Show hint
Look for whatever's already out in the world and hard to call back, not whatever's merely embarrassing.
Show answer
Model answer: "A support tool that drafts email replies to customers. If a draft with the wrong refund amount went out signed as final, I couldn't unsend it, the customer already has it in their inbox. Which model drafts the reply, I can change that any time. The sent email, I can't."
Short answer, the number question
6. If the sixty sample test had found 1 in 40 flagged descriptions instead of 1 in 12, would the order of sections change? Say what moves and what doesn't.
Show hint
Reversibility is about what happens on the day it's wrong, not about how often it's wrong right now.
Show answer
Model answer: "The order stays the same. The guardrail section still goes first, because a lower error rate doesn't make a published mistake any easier to undo, it just makes it rarer. What changes is pace: with fewer flags, the golden set might need fewer examples to build confidence, and the rollout stages could move faster, but nothing skips ahead of guardrails and the eval spec."
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.