ConceptAdvancedEval-Driven Specification / Writing a PRD for an AI feature / #22

Describe how you would version a PRD as model behaviour changes underneath it.

The direct answer
Tag every claim in the PRD to the exact model or prompt version it was checked against. The moment that version changes, flag the claim as unverified automatically, so nobody can keep treating it as true from memory. Re-check only the claims the new version actually touches, not the whole document.
Do this, in order
  1. Tag every claim in the PRD to the exact model or prompt version it was checked against, and flag it unverified the moment that version changes.Why: this is the fix. It would have caught the drift on day eleven instead of day nineteen, before a number ever reached a supplier.
  2. Wire the version-change flag into the release pipeline, not a reviewer's memory.Why: a habit that depends on someone remembering to reopen the PRD is exactly what let eleven quiet days turn into more.
  3. Split which claims need a fresh look by whether they depend on the model's own wording, not a blanket "review everything."Why: date math and other fixed rules don't drift when the model does, and reviewing them anyway burns time that should go to the claims that actually can.
  4. Leave the PRD itself as the one place people go to know what the tool does.Why: swapping it for a live dashboard nobody reads loses the readable contract; the problem was never the document, it was the missing version tag.
  5. Watch how many claims are currently unverified, not just whether escalation tickets are quiet.Why: eleven quiet days looked fine and weren't; a rising unverified count would have shown the risk before any ticket arrived.
  6. Don't respond to a drift incident by rewriting the whole PRD from scratch on every release.Why: a full rewrite every time the model changes is the heavy version of the same mistake. Only the claims tied to what actually changed need a fresh look.

How to answer this, stage by stage

Seven moves. This question sounds like it wants a process. It's really about the day a document stops matching the thing it describes. Each stage has the words you'd actually say.

1
Scope it to one concrete document
Say it like this
"Let me make this concrete. Say a procurement team uses a tool that reads every contract, flags the ones coming up for renewal, and drafts a negotiation brief with a suggested counter-offer range. That's the product this whole question sits on top of."
Why this works
"A PRD" has no shape until it's pinned to one real document describing one real product.
2
Say your structure out loud
Say it like this
"I'll walk through five things. Who actually treats this PRD as the truth. What they stop double-checking once it's signed off. What changes the day the model underneath it moves but the document doesn't. Which old decision I'd take back. And what the same bad day looks like once the PRD is versioned properly."
Why this works
Two seconds of structure stop you rambling through an abstract word and show the interviewer you already have a plan for it.
3
Reframe what "done" actually means
Say it like this
"A PRD getting signed off doesn't mean the product is finished changing. It means someone wrote down what the model did on one particular day. If nobody ties that sentence to which model wrote it, the sentence keeps getting read as true long after it's stopped being true."
Why this works
This is the actual insight the question is fishing for. Going stale is the PRD's normal state, not a failure state, unless something fights it.
4
Give the one decision
Say it like this
"Here's what I'd actually do. Tag every behavior claim in the PRD to the exact model or prompt version it was checked against. The moment that version changes underneath it, flag the claim as unverified, and don't let anyone treat it as still-true until someone's actually looked again."
Why this works
This is the direct answer, said out loud. It's a specific mechanism, not a promise to "keep the PRD updated," which is the thing everyone already says and nobody does.
5
Run the compressed failure
Say it like this
"Say Chidera signs off every escalation ticket against a PRD that says the tool never states a counter-offer as final, only a range. Engineering ships a wording change to sound more confident. Eleven days later, a buyer's ticket asks why the brief said '$42,000, final' instead of a range. Chidera, going off memory of the PRD, says that's not something the tool does, and closes it. The buyer forwards the number to the supplier that afternoon, and it costs the company about $5,400 they didn't need to give away."
Why this works
Four sentences, and it lands exactly on the day a static PRD stopped matching the model behind it.
6
Say what you'd measure, and what you'd leave alone
Say it like this
"I'd track how many claims in the PRD are currently unverified against the live model, not just whether escalation tickets are quiet. And I wouldn't re-check every claim on every release. A claim like 'flags renewals ninety days out' is date math, not generated wording. It can't drift when the model changes, so it doesn't need a fresh look every time something else does."
Why this works
Shows judgment past launch day, and stops the fix from turning into re-reading the whole PRD every release, which nobody will actually keep doing.
7
Close it in one breath
Say it like this
"So: a PRD stops being true the moment the model behind it changes and nobody wrote that moment down. Tag every claim to the version it was checked against, flag it the second that version moves, and the PRD stays a live contract instead of a photo everyone forgets was ever taken on a particular day."
Why this works
Restates the decision and why, in one breath. That's the line an interviewer remembers.
If you remember one thing Stages 3 and 4 carry this answer. A PRD isn't wrong the day it's written, it's wrong the moment the model moves and nothing says so. Tag the claim to the version, and let the version change do the flagging instead of a person's memory.

Let's learn

Here's what happens when a tool changes on a quiet Monday and the document describing it does not.

Say a procurement team uses a tool that reads every contract coming up for renewal and writes a short brief telling the buyer what range to offer the supplier.

Knowledge spark: what's a PRD? Short for product requirements document. The write-up that says what a product is supposed to do, so engineers, reviewers, and anyone answering a question about it are all working from the same page.

Before this tool existed, a buyer spent about three hours before each renewal digging through three years of invoices and drafting their own opening ask by hand. With the tool, that dropped to about ten minutes. It reads the contract, checks the price history, and drafts a brief with a suggested range.

Here's the part that matters. It isn't that the tool got something wrong. It's that nobody had a rule for what happens the day the tool's own wording changes and the sign-off document doesn't move with it.

Share of negotiation briefs stating a number as final, days since a routine wording update
0% 10% 20% 30% 40% Day 0 0%, wording ships Day 8 9% Day 11 13%, ticket closed Day 19 31%, loss found
A claim tagged to a version would have shown "unverified since v2.4" the moment the day-11 ticket opened. Nothing was built to say that, so the drift ran eight more days before finance found it the hard way.
Knowledge spark: what's a model or prompt version? Every time the AI behind a product gets swapped for a newer one, or the instructions it's given get rewritten, that's a new version. Nothing about the screen changes. The way it answers can change completely.

Twelve claims went into the PRD the week the tool shipped. Claim four says the tool never states a number as final, only a range, so the buyer keeps room to move in the real conversation. When the underlying model got a routine wording update, meant to make the tool sound more confident, claim four didn't change on paper. Nothing tagged that sentence to a version, so nothing told anyone it might now be wrong.

Two identical three by four grids of twelve squares, each one a claim in the PRD. Left grid, all twelve outlined in green, checked against version two point three. Right grid, eleven still green, one outlined in red, measured against version two point four eleven days later.
Twelve claims, same size grid, side by side. The broken count barely moved.
The PRD was still right about last month's model. Nobody had told it a new one had moved in.

At its worst, this costs more than never writing a brief at all. A document everyone trusts and nobody checks doesn't just fail to help. It tells everyone downstream to stop looking exactly when they should be looking hardest.

The decision that mattered Tag every claim in the PRD to the model or prompt version it was checked against, and flag it the moment that version changes. Not a bigger PRD. Not a stricter sign-off meeting.

What I would leave alone. The claim that Renewio flags a contract renewing within ninety days never needed a second look. It's pulled straight from the contract's own date, not from anything the model writes, so it can't drift just because the model's wording changes.

The lesson. A PRD is not a photo of a product. It's a claim about a moving target, and if nobody writes down which version of the target was checked, the claim quietly stops being true and there's no way to tell when.

Now here is the same thing as a story

Use this version when you have room to let it land, not just list it.

The laptop on Chidera Okonjo's desk has one browser tab that's been pinned since the week the tool launched: the PRD.

Chidera has run trust and enablement at Windermere Components for four years. Hand her any escalation ticket about one of the company's tools and she can tell you inside a minute whether it's a real bug or someone misreading what the tool was built to do.

Windermere buys parts from more than two hundred suppliers. Until eighteen months ago, a buyer prepping for a renewal did all the digging by hand. Then the procurement team got Renewio. It reads a contract, flags it ninety days before it renews, and drafts a brief: the supplier's price history, where Windermere has leverage, and a counter-offer range. A range, on purpose, so the buyer keeps room to move once they're actually talking to the supplier.

Twelve claims went into Renewio's PRD the week it shipped. Claim four says the tool never states a number as final. Chidera signed off on all twelve, and for the first six weeks, every time an escalation ticket came in, even the routine ones, she pulled the live chat transcript and checked it against all twelve claims herself, line by line. It always matched.

So she stopped pulling transcripts. She started answering from memory of what the PRD said, the way you answer a question about a rule you've read so many times you stop thinking you need the page.

Then, on a Monday nobody flagged as different, engineering shipped a small wording update to Renewio's prompt. The goal was reasonable: buyers weren't acting on the suggested ranges enough, so the update made the tool sound more confident. Claim four still read the same on paper. Nobody had tagged that sentence to a version, so nobody had a reason to go check it against the new one.

Eleven days later, a newer buyer named Soraya opened a ticket. Renewio's brief for the ArcMetal Supply renewal read, "We're offering $42,000, final." Not a range. She asked Chidera if that was normal.

Chidera didn't pull the transcript. She didn't think she needed to. Claim four. She told Soraya that wasn't something Renewio did, and closed the ticket.

That afternoon, Soraya forwarded the line to ArcMetal's rep, the way the PRD had taught everyone the tool was supposed to work: trustworthy enough to act on without a second read. ArcMetal held the number. Windermere's own price data said the top of a fair range was $37,000. The deal closed at $42,000, with no room left to talk it back down, because Windermere had already called it final.

We didn't lose five thousand four hundred dollars. We taught Soraya to stop reading her own negotiation brief before she hit send.

I want to say the model got worse. It didn't, not really. It got more confident, which was the whole point of the update. The real problem was that Chidera never had a running number for how much of the PRD still matched the live tool. She had a habit, and the habit had exactly two settings: pull the transcript and check, or trust the page. Six quiet weeks flipped it from the first to the second, and nothing was built to flip it back.

Left, a dial with many marks labelled how often to re-check the PRD against the live model, captioned what we assumed she'd do. Right, a two position switch labelled checks every ticket, or trusts the PRD completely, captioned no middle setting.
People are switches, not dials

So here's the decision I'd take back. When Renewio's PRD was first written, someone on the review floated tagging each claim to the model and prompt version it was tested against, so a version bump would show which claims needed a fresh look. The team talked about it for ten minutes and shelved it. There was only one version of the model at the time. Writing a version number next to every sentence read like paperwork for a problem that didn't exist yet. That was a fair call the week it was made. It stopped being fair the day engineering shipped a change to claim four and the PRD had no way of saying so.

I'd build that tag. Not a bigger PRD, not a stricter sign-off meeting, just a version tied to every claim and a flag when the version underneath it moves. Run the same eleven days through that design. The ticket Soraya opens shows "claim four, unverified since v2.4, 11 days" before Chidera reads a single word of the transcript. She pulls it in eight minutes, sees the drift, and tells Soraya to hold. The $42,000 line never reaches ArcMetal. The negotiation runs at the real range, and Windermere closes closer to $36,000.

That's the whole difference. One design hands the PRD a finish line and calls the job done. The other hands it a pulse, something that keeps checking itself against the model whether or not anyone remembers to ask.

And the part I'd tell myself, if I could go back: we asked whether the PRD was accurate the week we wrote it. We never asked how anyone would find out the week it stopped being.

FLIPS, for a PRD that outlives its own model version

This question sounds like it wants a process. It's really a Perturbation question: a document built to describe a product keeps insisting the product hasn't moved. FLIPS runs straight down the line, F to S.

Five stacked rows, F L I P S, each a letter in a coloured box, a step name, and a question. The I row is outlined in red.
FLIPS, in five rows
FFind the person
Who reads the PRD as ground truth?
Not "the team." One person, one document, one calendar habit.
In this answer: Chidera Okonjo, the trust and enablement lead at Windermere Components who signs off every Renewio escalation ticket against its PRD.
LLocate the habit
What did they stop re-checking once it was "done"?
The habit is the PRD working. What a clean sign-off covers for once nobody's job is to look past it anymore.
In this answer: She stopped pulling the live chat transcript for every ticket and checking it against the PRD's twelve claims, after six weeks where it always matched.
IIdentify the flip
What two-setting switch snaps, with no middle?
Not "the PRD got out of date." A specific behavior, exactly two settings, and no drift back once the habit stops.
In this answer: Pulls the transcript and checks it against the PRD, or answers straight from memory of the page, trusting it hasn't drifted from the model behind it. No setting in between once she stopped pulling transcripts.
PPinpoint the old decision
Which choice only made sense before the model moved?
Small, specific, reasonable at the time. Never "add more review."
In this answer: Never tagging a PRD claim to the model or prompt version it was checked against, because at launch there was only one version and it read like paperwork for a problem that didn't exist yet.
SShow the replay
Same bad day, versioned PRD. Better ending?
Run the same trigger through the fixed design. End on something you can count.
In this answer: The escalation ticket flags "claim four, unverified since v2.4, 11 days" the moment it opens. Chidera checks the transcript in eight minutes, and the $42,000 line never reaches the supplier.
Two panels. Left, a gently rising line labelled the model's own wording, from hedged range to reads like a done deal. Right, a line that starts flat and high, labelled checks every ticket, snaps straight down, and runs flat at trusts memory only.
A small move in the model. A hard snap in who was watching it.
Why I is the hard step Anyone can say the PRD "went stale." The hard part is naming the exact habit that had to stop for that staleness to go unnoticed, and proving there's no setting between "checks the transcript" and "trusts the page." "Reads it a little less carefully" is a dial. "Stopped checking for good the day it went six weeks without being wrong" is a switch. If your flip has a middle, keep looking.

And if you want to be sure it really works, try it somewhere else

An HVAC dispatch tool tells field technicians whether to repair or replace a broken unit. Same question, a completely different product, and a different flip.

F. Corbin Ashworth, dispatch supervisor at Fairwind Mechanical Services, who trains every new technician on Unit Advisor's onboarding PRD before their first week on a truck.
L. New technicians read that PRD once, during onboarding, and never open it again, since nothing about the job ever seems to call for it.
I. A different flip. Technicians don't check the tool's advice less carefully. They stop opening the one document that would tell them the advice changed. No middle setting: the PRD gets read during onboarding, or it doesn't get read again, ever, once the first week ends.
P. The onboarding promise: new hires were told the PRD was the one thing they needed to read to understand Unit Advisor, once, and that promise gave their trust nothing to check back against later.
S. Tag the repair-versus-replace rule to the pricing module's version, and force a fresh look the moment the vendor bumps it. A drift that took two hundred and forty days to surface through an angry customer review gets caught in six, before the next onboarding class ever reads the old numbers.

How long the drift ran before anyone caught it, HVAC dispatch
No version tag, PRD read once at onboarding
Old design
240 days
PRD claim tagged to the pricing-module version
New design
6 days
Old design: caught when a customer's public review complained that every technician was pushing replacement, 240 days after the pricing module changed. New design: caught by a forced re-check the moment the vendor bumped the module, six days later, before the next onboarding class ever read the old numbers.
A second decision worth taking back Telling new hires a document was the one-time thing to read is itself a decision, not a fact about training. A line that said "check this again the day the tool's numbers move" would have given them a reason to open it a second time.

Swap the trigger and it still runs

  • Speed: if the wording update had shipped as a single hotfix instead of a routine release, there'd have been no testing window at all, and the same eleven days would have run faster with even less warning.
  • Cost: if pulling a transcript cost real analyst time per ticket, the six weeks of checking everything would have ended even sooner, and the drift would have had longer to run before anyone looked.
  • The model got better: close to what actually happened here. The wording change was meant to make Renewio sound more useful, not worse, which is exactly the kind of change that slips past a review built to catch problems, not improvements.

Where people run it wrong

  • Blaming the buyer for forwarding a number the tool itself printed.
  • Fixing it by writing a longer, stricter PRD instead of tagging the one that already exists to a version.
  • Waiting for a customer complaint instead of watching how many claims have gone unverified since the last model change.

How to use it live

Say the reframe out loud before naming a fix. "So the real question isn't whether the PRD was written well, it's whether anyone would know the moment it stopped being true." That costs five seconds, and it's where the real answer to "how do you version it" actually starts.

Flashcards (click a card to flip it)

Eight fixed slots, pulled straight from the answer above.

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Checks sometimes, then stops checking at all. It fires on a habit that kept working, not on something visibly breaking.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Chidera Okonjo, trust and enablement lead at Windermere Components, four years in. She signs off every Renewio escalation ticket against its PRD.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She stopped pulling the live chat transcript for every escalation ticket and checking it against the PRD's twelve claims, after six weeks where it always matched.
4 · THE FLIP, HERE
What's the two-setting switch in this story?
Tap to flip
ANSWER
Pulls the transcript and checks it against the PRD, or answers straight from memory of the page. No setting in between once she stopped pulling transcripts.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never tagging a PRD claim to the model or prompt version it was checked against, because at launch there was only one version and it read like paperwork for a problem that didn't exist yet.
6 · THE NUMBER
The PRD's twelve claims held ___ broken the whole eleven days. The line that broke cost about $___.
Tap to flip
ANSWER
1 out of 12 broke, claim four. It cost Windermere about $5,400, the gap between the real $37,000 range top and the $42,000 the tool called final.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The ticket flags "claim four, unverified since v2.4, 11 days" the moment it opens. Chidera checks the transcript in eight minutes, and the $42,000 line never reaches the supplier.
8 · CROSS-PRODUCT
Section 4 answers this same question for a different product, with a different flip family. Which product, which family?
Tap to flip
ANSWER
Unit Advisor, an HVAC repair-or-replace dispatch tool, using the abandonment flip: new technicians read the onboarding PRD once and never reopen it, so a pricing-module change goes unnoticed for months.

Check yourself Score: 0 / 0

Multiple choice
1. What was the flip in Chidera's story, and what were its two settings?
  • A. She reads escalation tickets a little less carefully than she used to.
  • B. She pulls the live transcript and checks it against the PRD, or she trusts the page from memory, with nothing in between.
  • C. Renewio's negotiation briefs went from suggesting ranges to stating final numbers.
  • D. She asks a coworker to double-check any ticket she's unsure about.
Show hint
A flip is a verb the person does, not a change in the model, and it has exactly two settings.
Show answer
B. C describes the model's own change, which is the trigger, not Chidera's behavior. A is a dial, there was no "checks it a bit less" setting she landed on. D is a fix, not what happened in the story.
True or false
2. True or false: Windermere should re-check every claim in Renewio's PRD, including the ninety-day renewal flag, every time the model's wording changes.
  • True
  • False
Show hint
Look at the "what I would leave alone" paragraph. What decides the ninety-day flag?
Show answer
False. The ninety-day flag is date math pulled from the contract, not generated wording. It can't drift when the model's phrasing changes, so re-checking it on every release burns review time on a claim that was never at risk.
Fill in the blank
3. The decision this answer takes back is having no automatic ______ between a PRD claim and the model ______ it was checked against, so the only thing standing between the two was ______.
Show hint
It's the reversal category called "absent state": something was never built to keep a record of what had actually been checked.
Show answer
Tag, version, Chidera's memory. Nobody tied claim four to the version it was verified against, so there was no way to tell the PRD was describing an old model until a buyer had already acted on the new one.
Multiple choice
4. Which claim in Renewio's PRD did NOT need a fresh look when the model's wording changed, because it doesn't depend on generated language?
  • A. Never state a counter-offer as final.
  • B. Flags a contract renewing within ninety days.
  • C. Drafts a range grounded in three years of price history.
  • D. Every claim needed the same re-check at the same intensity.
Show hint
Look for the claim that's decided by a date on a contract, not by anything the model writes.
Show answer
B. D fails the "what I would leave alone" discipline. A good answer names somewhere the change genuinely doesn't matter instead of treating every claim as equally at risk.
Short answer, apply it yourself
5. Pick a document at your own job, or in your own life, that you treat as still true because it was true once. What would "still accurate" stop proving if the thing it describes quietly changed underneath it?
Show hint
Think of an onboarding doc, a recipe, a setup guide, an insurance policy summary.
Show answer
Model answer: "My team's onboarding doc says 'ping the on-call channel for deploys.' We switched tools eight months ago and nobody updated it. It still reads as a calm, confident sentence. New hires follow it and get nothing back, because the channel it names doesn't exist anymore. The doc never got worse. It just stopped being tied to anything that could tell it the tool had changed." Any honest answer works if it names a real document and a real way it could silently stop matching reality.
Fill in the blank, do the math
6. The unmatched-briefs chart shows the final-sounding share at 13% on day 11, the day Soraya's ticket was closed, and 31% on day 19, the day finance found the loss. If the day-11 ticket had been caught instead of dismissed, about how many fewer days did the drift need to run, compared with what actually happened?
Show hint
Subtract day 11 from day 19.
Show answer
8 days. Catching it on day 11 instead of day 19 would have stopped the drift eight days earlier, before the $42,000 line ever reached ArcMetal.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more