ConceptIntermediateEval-Driven Specification / Writing a PRD for an AI feature / #14

Describe how the PRD should handle the eval spec: inline, linked or separate?

The direct answer
Write the eval spec inline, in its own labeled section of the PRD, for a feature owned by one pod. Only move it to a separate linked document once three or more features are drawing on the exact same criteria, and even then, name an owner and a version date right inside the PRD's own text. A link nobody owns is how a wrong number quietly becomes the real spec.
How to place the eval spec, in order
  1. Write it inline, in its own labeled section of the PRD, for a single-pod feature.Why: a reader can't skim past a number that lives in the same document they're already signing off on.
  2. Name an owner and a version date on any number in the spec, inline or linked.Why: an unowned number is the one that gets edited by a stranger and nobody notices.
  3. Save linking for criteria genuinely shared by three or more features, not for a shorter PRD.Why: that's the actual kill line, below it a link buys you nothing but risk.
  4. Treat an overlong inline PRD as a same-day trim, not a reason to hide the spec behind a link.Why: the two costs aren't equal, and the cheap one is the one people worry about.
  5. Re-read the linked doc's current numbers before every sign-off, if it is linked at all.Why: this is the one habit that would have caught the drift before engineering built to it.
  6. Watch for a second PM's name showing up in a shared doc's edit history.Why: that's the earliest sign the doc has outgrown being shared informally.

The four questions that settle where the eval spec lives

This is a tradeoff between two real documents, not a style question, so PICK does the work: commit to one, then show why the other one costs more than it looks like it does.

1
Ground it in one real PRD, and say your structure up front
Say it like this
"Let's put this on a real one. Reeltide is a video-hosting platform, and Chapter Marks is the feature that watches an upload and writes chapter markers with a one-line description at each timestamp. I'm going to pick a position on where its eval spec lives, then say who feels each kind of miss, then which kind actually costs more, then what would change my mind."
Why this works
Gives the interviewer one real PRD to push on, and tells them the shape of the answer before you start it.
2
Say what the question is really testing
Say it like this
"This isn't really 'where does the paperwork live.' It's 'who's allowed to quietly change what counts as correct, and would you even find out.' A PRD is the one document everyone in the room actually reads and signs off on. Anything outside it is a document some people read."
Why this works
Shows you see past the surface ask to the judgment it's actually testing.
3
Give the position, with the actual mechanism in it
Say it like this
"My position: for a feature this size, one pod, one PRD, the eval spec lives inline, in its own labeled section, right next to the requirements it's grading. I only pull it into a separate linked doc once three or more features genuinely need the same criteria, and even then that doc gets a named owner and a version date, stated in the PRD, not just a bare link."
Why this works
PICK rewards a real mechanism, not a vague promise to keep docs in sync.
4
Name who feels each kind of miss, then prove the asymmetry with a real number
Say it like this
"Here's the split. If a PRD runs long because the eval spec sits inline, a reviewer flags it the same day, in the same meeting, and it gets trimmed. That cost about half a day once, on Chapter Marks, when the doc grew from four pages to nine. But last year, our chapter-timing tolerance lived in a shared doc used by three other features. Someone else's PM loosened it from three seconds to eight for a different feature and never touched our PRD. Engineering built to eight. We'd promised three. Auditing all four features that shared that doc took ten engineering days, and we had to walk back a claim on the marketing page."
Why this works
Real numbers make the asymmetry checkable instead of asserted.
5
Say what you'd leave alone
Say it like this
"I wouldn't fight this hard over a two-person internal prototype that never ships past a beta test. Nobody but the writer and one engineer will ever open that PRD, so a link can't quietly drift, there's no second author to drift it. Save the rule for anything a stranger might one day edit."
Why this works
Shows judgment instead of applying the same rule everywhere just because you can.
6
Name the kill criteria
Say it like this
"I'd move the eval spec out of the PRD the day three or more shipped features genuinely need the exact same number, because duplicating it three times inline is now the drift risk instead of the fix. Until then, inline wins, because nobody in that meeting can quietly rewrite a number that's sitting in the document they're all looking at."
Why this works
Shows confidence without stubbornness, one testable line either side of the pick would flip on.
7
Close on the one line worth remembering
Say it like this
"So: inline by default, for a feature this size, because a bloated PRD gets caught in the room. A stale link gets caught by finance, or a customer, weeks later. I'd rather trim a page than explain a number nobody owned."
Why this works
Restates the position and the reason in one breath, the line an interviewer remembers after the call ends.

One more thing before this moves to the long version: most candidates hear "inline, linked, or separate" and answer it as a documentation-tidiness question, whichever keeps the PRD shortest. Say who actually gets hurt by each choice, in real hours or real dollars, and you've shown the judgment the question is actually testing.

Let's learn

The feature is a strip of small markers under Reeltide's video player, one dot per scene change, with a short label at each, put there by nobody the viewer will ever meet.

Before Chapter Marks, a viewer hunting for one moment in a 40-minute upload dragged the scrub bar back and forth, guessing, for about 100 seconds on average. A creator writing chapter titles by hand for that same video spent close to 18 minutes doing it after the upload finished.

Knowledge spark: what's an eval spec? The written rule for what counts as the AI feature working. A number, or a check, that says pass or fail. Without one, "it works" just means "it looked fine to whoever tested it that day."

Now Chapter Marks writes the whole chapter list during the upload itself. A viewer taps the list and lands on the right second in about 4 seconds. That's the win, and it's real.

Here is the turn. Every now and then a marker lands a few seconds off the real scene cut, and that alone is not the real problem, most viewers never notice five seconds on a 40-minute video. The real problem is that two different documents disagreed about how many seconds counted as "off," and nobody in the room who signed off on Chapter Marks knew that.

We didn't lose five seconds of accuracy. We lost the one number that said what "accurate" meant.

At its worst, that looks like this. The PRD for Chapter Marks promised markers within 3 seconds of the real scene cut, the number the marketing page turned into "precise, second-accurate chapters." The actual eval criteria lived in a shared doc used by three other features too. Someone else's PM, working on a different feature entirely, loosened that shared tolerance to 8 seconds and never touched the Chapter Marks PRD, because it wasn't their document to touch. Engineering built exactly to what the shared doc said. It just no longer said what the PRD promised.

The choice I would take back We put the eval spec in a shared linked doc so every feature's PRD could stay short and consistent. I'd take that back for a feature this size. Write it inline, in its own section, so the only way to change what "accurate" means is to edit the document everyone in the sign-off meeting is already reading.
Cost, in engineering days, of each kind of miss
0.5 day 10 days 4 features re-audited Inline PRD runs long, trimmed in review Linked doc drifts, caught after ship
The green bar is small on purpose: a long PRD gets flagged by a reviewer in the same meeting and trimmed the same day. The red bar is what the actual drift cost: when the shared tolerance doc quietly moved from 3 seconds to 8, every one of the 4 features that linked to it had to be re-audited by hand, 2.5 engineering days each, 10 days total, plus a walked-back line on the marketing page.

What I would leave alone. A two-person internal prototype that never ships past its own team. Nobody outside that pair will ever open the PRD, so there's no second author who could quietly edit a linked number even if one existed.

The lesson. We built the shared doc to keep every PRD short and consistent. That was the right call when two features existed. Nobody ever revisited it once there were four, and "consistent" quietly turned into "nobody's specifically watching this."

Now here is the same thing as a story

The short version is above. Read this one when you want to feel why it matters, not just know the rule.

Ingmar can smell a PRD that's about to grow legs before anyone else in the room. Four years on Reeltide's platform team, and he's the one people hand a draft to before it goes to review, just to see what he cuts.

When the quality pod formed, eighteen months back, someone suggested pulling every feature's eval criteria into one shared doc, so PRDs would stay short and every team would grade things the same way. It was a good idea. At the time, two features existed, and Ingmar wrote both of them himself.

For months, that felt like a win. PRDs stayed five, six pages. Reviewers read them in ten minutes flat and said so, out loud, in standups. Ingmar linked to the shared eval doc, glanced at the summary line at the top to confirm nothing had obviously changed, and moved on.

Then he stopped glancing at the summary line. He'd click the link, see the doc load, and close the tab before it finished scrolling into view. Then he stopped clicking at all. The tag in the PRD said "current," and four months running, it had been.

Then, in a Tuesday sync, Reeltide's support lead mentioned it the way you'd mention traffic. A creator's ticket had come in: "Chapters are landing in the wrong scene entirely on my top video." Not a complaint about a whole feature failing. Just one line, read off a list of eleven other tickets.

Two unequal boxes: a small plain grey box labeled a PRD runs long, caught by a reviewer the same day, next to a large jagged red-orange box labeled engineering builds against a stale linked doc, nobody notices for weeks.
Same PRD decision, two very different costs

Ingmar opened the shared eval doc for the first time in months. The chapter-timing tolerance no longer read 3 seconds. It read 8. Someone else's PM had widened it four months earlier, for a completely different feature, one that genuinely needed more slack. The edit was reasonable on its own. It had just landed, silently, on three other features that never asked for it.

The chapters weren't broken. The promise was, and nobody had signed their name to it.

It was never really about the five seconds. Chapter Marks had never had a number in the room. It had a document, and the document only had two states: linked and current, or something to worry about later.

Auditing the four features that shared that one doc took ten engineering days. Every one of them had shipped, every one had "passed" whatever the shared doc said that week, and every one of them was quietly grading against a promise the PRD itself no longer matched.

Eighteen months before any of this, in the room where the quality pod first proposed the shared doc, someone did ask whether each PRD should just keep its own copy of its own numbers. It felt like duplicate work for two features that would obviously stay in sync, since the same three people wrote both. They built the shared doc, linked both PRDs to it, and moved on to shipping the next thing.

Ingmar moved Chapter Marks' eval spec inline the month after the audit. Three seconds, in the PRD itself, owned by him, dated. Run the same drift again, and the other PM's edit never touches this document at all, because it isn't their document to edit.

One design lets a stranger's edit quietly rewrite your promise. The other makes that literally impossible, because the words live in the file with your name on it.

And the thing I'd tell myself, back when we built that shared doc: I wanted every PRD I touched to stay five pages. I never asked what happens the day someone I've never met edits the one page that decides whether "accurate" still means what it meant when I wrote it.

PICK, the four letters mapped onto Chapter Marks

This is a tradeoff about where one number lives, not a rule for organizing documentation in general, so PICK is the tool.

P, position. Write the eval spec inline, in its own labeled section of the PRD, for a feature owned by one pod. Move it to a separate, named, versioned linked doc only once three or more features genuinely share the same criteria.
I, impact. A reviewer feels a bloated PRD immediately, in the same meeting, and can ask for a trim on the spot. Engineering, and eventually a customer, feels a stale linked doc weeks later, after they've already built and shipped against a number the PRD never actually re-confirmed.
C, cost asymmetry. A long inline PRD is loud and cheap: someone complains, it gets trimmed, done in half a day. A stale linked doc is quiet and expensive: it passes every review that only checks the PRD's own text, and the real cost, ten engineering days and a walked-back claim, shows up only after the drift has already shipped.
K, kill criteria. Move the spec out of the PRD the day three or more shipped features need the exact same number, because duplicating it inline three times over is now the bigger drift risk. Below that count, inline wins every time.
Knowledge spark: why not just ask engineering to double-check the linked doc every sprint? Because "double-check it" isn't a design, it's a hope. Nothing in the process forces the check to happen, and the whole point of the drift here is that nobody noticed the doc had changed for four months.
Features sharing one eval doc, against the kill line
Below the kill line
Where the drift actually happened
Kill line: 3 features
1 feature 3 features 4 features Kill line: 3 features Chapter Marks, inline, alone The shared doc that drifted
At 1 feature, Chapter Marks sits well below the kill line, so it stays inline. The shared timing-tolerance doc was already past that line, at 4 features, when the drift happened, which is exactly the case where a named, versioned, linked doc would have earned its keep, if anyone had actually treated it that way.

Try it on a claims photo, not a video at all

Ambercross Insurance runs an AI feature that reads photos of car damage and estimates repair cost, so an adjuster can approve small claims same day. Its eval spec has to say how close that estimate must land to what a licensed in-person assessor would say, and the same question comes up: does that number live in the PRD, or somewhere it points to.

P. Here the position flips. Keep the tolerance, within $150 of a licensed assessor's number, in one owned, versioned, linked "AI Assessment Standards" doc, not duplicated inline across every claims PRD.
I. An engineer on any one of Ambercross's six claims products feels a duplicated inline copy the moment it silently diverges from the other five, since nothing forces six separate PRDs to agree. Compliance feels a bare, unowned link the day a regulator asks which number is the real one and nobody can say for certain.
C. Here, six duplicated inline copies are the quiet, expensive risk, each one can drift on its own with nobody checking all six against each other. One owned, versioned, linked doc, audited on a fixed schedule, is the loud, cheap option: a missed audit shows up immediately, because it's the one thing compliance already checks every quarter.
K. Fold it back into each PRD inline the day this tolerance governs only one or two products, or the day compliance drops the quarterly cross-check, the same evidence that would have kept Chapter Marks' number in a shared doc too, if Reeltide had ever built the equivalent check.

What I would leave alone, on Ambercross's claims estimator The photo-quality guidance, blur, glare, distance from the vehicle. It's advice, not a pass or fail number, and every product's PRD can restate it in its own words without any of them disagreeing about what "accurate" means.
Quarterly compliance-audit hours, one shared doc across six products
Read the central standards doc
2 hrs
Check 6 products against it, 3 hrs each
18 hrs
Reconcile and file findings
4 hrs
24 hours a quarter, about three working days, and it's a planned, bounded cost, because there's exactly one document to check. The alternative, six inline copies quietly drifting on their own schedules, has no equivalent bound, which is why the position flips here and stays inline back at Reeltide.

Swap the trigger and it still runs

  • Speed: Chapter Marks starts generating chapters in 5 seconds instead of upload-time. Doesn't move the pick, because the position is about who can edit the number, not how fast the feature runs.
  • Cost: keeping the eval spec inline adds two pages to every PRD it touches. Still doesn't flip it, two pages a reviewer reads once is cheaper than ten engineering days spent auditing a drift nobody saw.
  • The model gets better: if Chapter Marks' scene-detection accuracy improves enough that a marker landing 8 seconds off basically never happens, the tolerance number itself matters less, but the question of who owns it and where it lives doesn't go away, it just gets asked less often.

Where people run it wrong

  • Picking "inline" or "linked" as a house style for every PRD forever, instead of re-checking it against how many features actually share the number today.
  • Linking to a shared doc with no named owner and no version date, so "current" is just whatever the tag says, not something anyone actually re-verified.
  • Treating a short PRD as automatically the safer choice, when the thing that got shortened was the one number everyone downstream is building against.

If you are asked this cold

Say the reframe out loud before picking a side. "Give me a second, I want to know how many other features would need this exact same number before I say where it lives." That's true, it's already stage two of the walkthrough, and it buys you time to find the real threshold instead of guessing at a house style.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's the hardest step to nail?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a stale linked doc is quiet and expensive while a bloated inline PRD is loud and cheap, instead of assuming both cost about the same.
2 · THE PERSON
Who owns Chapter Marks' PRD, and what's he good at?
Tap to flip
ANSWER
Ingmar Sonderby, four years on Reeltide's platform team, the PM other people hand a draft to just to see what he cuts.
3 · THE HABIT
What did Ingmar stop doing once the shared doc kept coming back "current"?
Tap to flip
ANSWER
Glancing at the shared doc's summary line before sign-off. Then clicking the link at all. Four months of it saying "current" was enough to stop checking.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A bloated inline PRD: caught by a reviewer same day, about half a day to trim. A stale linked eval doc: cost 10 engineering days to audit the 4 features that shared it, plus a walked-back marketing claim.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Write the eval spec inline, in its own PRD section, for a feature this size, and only link it out once three or more features genuinely share the same number, with a named owner attached.
6 · THE NUMBER
Auditing the features that shared the drifted eval doc took ______ engineering days.
Tap to flip
ANSWER
10 engineering days, 2.5 per feature across the 4 features that linked to the same shared tolerance doc when it quietly moved from 3 seconds to 8.
7 · THE KILL CRITERIA
What evidence would flip this position back the other way?
Tap to flip
ANSWER
Three or more shipped features genuinely needing the exact same eval number. Below that, duplicating it inline is safer than sharing it; at or above it, a named, versioned, linked doc starts earning its keep.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the position land there?
Tap to flip
ANSWER
Ambercross Insurance's photo-based claims estimator. The position flips to linked, one owned, versioned "AI Assessment Standards" doc, because six products genuinely share that number, well past the kill line.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these is the actual mechanism behind this answer's pick?
  • A. Always write every PRD's eval spec inline, no matter how many features share the number.
  • B. Write it inline for a single-pod feature, and only move it to a named, versioned, linked doc once three or more features genuinely share the same criteria.
  • C. Always link to a shared doc, so every PRD stays as short as possible.
  • D. Skip writing an eval spec at all until after the feature ships.
Show hint
Two of these are "always" rules. A real pick has a threshold that can flip it.
Show answer
B. A and C both apply the same rule forever, which isn't a pick, it's a house style. D isn't a pick at all. Only B names the actual threshold, three or more shared features, that decides which way it goes.
True or false
2. True or false: "it depends on the team" is a strong way to answer this question in an interview.
  • True
  • False
Show hint
PICK's first letter is position, stated before any reasoning, not a hedge.
Show answer
False. "It depends" isn't wrong, exactly, but it fails the question, because it never commits to a position an interviewer can push on. Say inline, name the size of feature it's for, and give the number that would change your mind.
Fill in the blank
3. Fill in the blank: the choice this answer takes back is that Chapter Marks' eval criteria lived in a ______ document, not inside the PRD that made the promise.
Show hint
The word describes a document more than one PRD points to.
Show answer
Shared, linked. It made sense when only two features existed and Ingmar wrote both. It stopped making sense once a stranger to this PRD could edit the one number it depended on.
Multiple choice
4. In which of these cases would this exact pick matter least?
  • A. A feature whose eval number is quoted directly on the public marketing page.
  • B. A feature whose eval spec is shared with three other live products.
  • C. A two-person internal prototype that will never ship past its own team.
  • D. A feature already burned once by a silent tolerance change.
Show hint
Ask where a second author could actually reach the number and change it without anyone noticing.
Show answer
C. With only two people who will ever touch that PRD, there's no stranger who could quietly drift the number either way. A and B are exactly where the pick matters most, and D is the scar this whole answer is built around.
Short answer, apply it yourself
5. Pick a PRD or spec you've written or read yourself. Name one number in it that would be dangerous to leave in a doc nobody owns.
Show hint
Look for the number a reader would assume is still true just because the document is still linked and still loads.
Show answer
Model answer: "A support-response SLA spec that linked out to a shared 'response time standards' doc used by three teams. The four-hour target for our tier got quietly changed to six hours for a different team's workload, and our own PRD never said so. I'd have caught it in a minute if the number had just been typed into our own document." Any answer works if it names a specific number and who could have changed it unnoticed.
Short answer
6. If the shared tolerance had drifted from 3 seconds to only 3.5 seconds instead of 8, would the same position, and the same urgency, still hold? Walk through it.
Show hint
Separate "was the drift itself huge" from "was the fact that nobody could see it happening a problem."
Show answer
The position holds, the urgency drops. A drift to 3.5 seconds is small enough that most viewers would never notice, so the ten-day audit and the walked-back claim probably wouldn't have happened. But the actual danger was never the size of the drift, it was that nobody could see it happening at all. The same silent edit could just as easily have been the 8-second version next time. The fix, put the number where a stranger can't quietly change it, is worth doing regardless of how big this particular drift turned out to be.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more