How do you specify the data dependencies of a feature in a PRD?
- Break the data-dependency section into the feature's real categories, and size a minimum number for each, typical cases and edge cases both.Why: a single vague line can leave the riskiest category almost empty while the section still reads as complete.
- Name every category's number out loud in the PRD.Why: an unstated per-category count is a guess wearing a total's clothes.
- Give the total as a range, not one fixed number.Why: the real total depends on how many categories the feature ends up covering, and that grows with scope.
- Check the total against how long it would really take to source and clean that much data.Why: a number that implies two days for 165 careful examples is too good to be true, and one that implies a whole quarter is too slow to ship.
- Name which assumption, category count or coverage depth inside one, swings the total most.Why: that's the assumption a sharp reviewer pushes on, and "we need more data" isn't an answer.
- Leave the low-risk, well-covered categories thin, on purpose.Why: padding an already-safe category doesn't reduce risk, it just spends review time a riskier one needs more.
How to answer this, stage by stage
Seven moves. The trap is writing one confident sentence about "training data" and never showing which categories it's actually built from.
Let's learn
Curbwrite is the feature inside Fieldpost, a real-estate agent platform, that turns an agent's photos and a few rough notes into a full listing description.
Before Noelia's team wrote a real data-dependency section, the PRD covered it in one line, written in about ten minutes during a two-week pilot that only covered single-family homes: "training and eval data: real listing descriptions, pulled from Fieldpost's own archive." Five hundred examples, pulled at random. That felt like plenty, because single-family was the only category that existed yet.
Single-family resale: 30 + 15 = 45
Condo / co-op: 18 + 7 = 25
Multi-family (2-4 unit): 12 + 8 = 20
Land / vacant lot: 10 + 10 = 20
New construction: 14 + 6 = 20
Luxury ($1.5M+): 15 + 20 = 35
# total (added up, not averaged)
165
Here's the turn. More data isn't really the fix.
Say it plainly: of the old 500 archive examples, the great majority were single-family, because that's what's most common in the archive to begin with. A random pull doesn't fix that skew, it repeats it. A category can have almost no real coverage and the total can still say 500.
What that costs at its worst: Curbwrite reports a healthy overall "sounds natural" score, the team trusts it, and the one category most likely to draw a fair-housing complaint, luxury, keeps producing risky phrasing quietly, because the data behind the eval was never built to catch it.
The choice I would take back. Writing "training and eval data: 500 listings from the archive" as the whole section, instead of breaking it down by property type with a minimum edge-case count in each. That made sense during the two-week pilot, when single-family was the only category to describe. It stopped making sense the moment Curbwrite grew to six very different ones.
What I would leave alone. The single-family category doesn't need special sizing. Archive coverage there is already deep, the descriptions are the most boilerplate, and the fair-housing risk is the lowest of any category, so 45 is enough, and padding it further wouldn't teach the model anything it doesn't already do well.
The lesson. A data-dependency line sized as one flat pull only proves the feature has data in general. It says nothing about whether it has enough of the kind that turns into a complaint.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why one thin category mattered more than four hundred extra single-family examples would have.
Noelia Frayne can read six lines of a listing description and tell you, before she finishes the paragraph, whether it will pass Fieldpost's compliance check. She's the product manager for Curbwrite, and she wrote its first PRD herself, during a two-week pilot covering single-family homes only. The data-dependency section was one line: pull 500 examples from Fieldpost's own archive. It took about ten minutes to write, and with one property type live, it was hard to argue with.
For the better part of a year, that felt right. Curbwrite kept scoring in the high nineties on the internal "sounds natural" review. The team added condos. Then multi-family. Then land, new construction, and luxury, for the high-end agents who wanted a first draft before their own polish pass. Nobody touched the data-dependency line. Why would they. The score kept looking good.
The habit thinned in three small steps. First, when Curbwrite grew into luxury listings, nobody added examples for that category specifically, there wasn't time before the next release and the score was still fine. Then, when a new hire on the compliance team asked whether the eval data still matched what Curbwrite actually covered, the answer was "probably, it's pulled from the whole archive," and the meeting moved on. Then, quietest of all, the Curbwrite team stopped opening the underlying examples at all. They just watched the dashboard. Ninety-six percent felt like a fact, not a measurement anyone had checked lately.
Then, on a Tuesday, Fieldpost's quarterly fair-housing compliance audit pulled twenty live luxury listings, the way it does every quarter, and read them line by line. One, drafted by Curbwrite six days earlier for a house two blocks from a well-rated elementary school, read: "A short walk to one of the area's best schools, perfect for a growing family." Familial-status steering, in a sentence nobody had written on purpose.
Noelia pulled the 500 archive examples that afternoon. Of the 500, four hundred and something were single-family. Luxury, the category the flagged listing belonged to, had eleven.
The old decision that put her there was almost a year old, and easy to remember, because it took about ten minutes. Someone on the platform team had shared a note: for a new AI feature's data-dependency section, pull five hundred examples from the archive and call it done. Every team used roughly the same approach. Noelia wrote it that way because it was fast, and Curbwrite, at the time, only drafted one kind of listing. Nobody in that room was wrong. Five hundred random single-family examples is a reasonable base for a feature that only writes about single-family homes. Five hundred random examples pulled from an archive that's mostly single-family is a coin flip about whether luxury shows up at all.
So she rebuilt it that month: 165 examples, weighted by category, thirty-five of them luxury, twenty of those built specifically around the kind of phrasing that nudges toward a protected class without naming it. Run against the new set, Curbwrite doesn't just report a healthy overall score and stop. It reports: five of the thirty-five luxury examples still draw a steering flag before the fix, the exact shape of what reached that live listing. Caught in an afternoon of testing, instead of found by a fair-housing audit, twice.
The part I'd go back and tell myself: I used to think a data-dependency section was there to say how much data a feature needs. It's really there to decide, ahead of time, which categories you're even able to check.
The five letters, run against Curbwrite's one line
This is a sizing question wearing a design question's clothes, so BOUND fits and FLIPS doesn't. Nobody's habit snapped on one Tuesday, it thinned over months. What changed was how much of each property type that one line could actually account for.
B, break it down. Data-dependency size equals the sum, across every property type Curbwrite covers, of the typical examples for that type plus its edge cases. Six separate additions, not one number split evenly.
O, own the numbers. Six property types: single-family 45, condo 25, multi-family 20, land 20, new construction 20, luxury 35. Added up: 165.
U, use a range. At the point estimate, 165, for the six types Curbwrite covers today. If launch scope narrows to the four an agent touches daily, 125. If it expands to short-term rental copy and teardown listings, past 220. Honest range: 125 to 220.
N, nail the sanity check. 165 examples, at about 25 minutes each to source a real listing, check it against the MLS, and flag any steering language, comes to roughly 69 hours, a bit under two weeks of one reviewer doing nothing else. That's also a sliver of what Fieldpost actually processes, something like 3,000 new listings a month platform-wide, but a data-dependency set was never meant to sample everything. It's meant to sample what breaks.
D, direction. Adding one more property type at the average size moves the total by about 28. Doubling the edge cases inside every category moves it by about 66, more than twice as much. Depth beats breadth. That's the number worth watching, not the category count.
And if you want to be sure it really works, try it somewhere else
A mid-size municipal sanitation authority runs a tool called Noticeline. It doesn't draft anything customer-facing, it turns a field inspector's photos and a few scribbled notes into the formal violation notice mailed to a property owner after illegal dumping, an overflowing bin, or hazardous material left on public land.
Thaddeus Kolb is DraySide Sanitation's compliance program manager, and he's the one who ends up sizing Noticeline's data-dependency section.
B, break it down. Notice-drafting data-dependency size equals the sum, across every violation type Noticeline covers, of typical examples plus edge cases, added up by violation type, not guessed as one number for "code violations."
O, own the numbers. Five violation types: illegal dumping 20, overflow or blocked bins 10, hazardous material dumping 15, construction debris dumping 15, repeat-offender escalation notices 10. Total: 70.
U, use a range. At 70 for five violation types. If the agency only issues notices for the three most common, dumping, overflow, debris, and holds off on hazardous material and repeat-offender language until legal signs off, 45. If it adds two more categories, graffiti-adjacent dumping and abandoned-vehicle debris, past 95. Range: 45 to 95.
N, nail the sanity check. At 70 examples, roughly 20 minutes each for a code-compliance reviewer to check the cited ordinance and the wording, that's about 23 hours, under one working week for one reviewer. Against the roughly 400 notices the agency issues a month, that's a small slice, built to test what breaks, not to sample every notice ever sent.
D, direction. In principle, doubling the repeat-offender edge cases would swing Noticeline's total the way deepening edge cases swung Curbwrite's. But a hazardous-material notice and a repeat-offender notice carry more legal weight than an overflow-bin notice, a wrong citation there gets disputed and can void the fine. So the lever Thaddeus actually pulls first is depth in those two highest-stakes categories, not category count in general. Same shape as Curbwrite's answer, a different category doing the work.
The old decision Noticeline would take back is a cousin of Curbwrite's, not a copy. Fieldpost had sized by a flat archive pull because copying the fastest option was easy. DraySide's old habit was different: one shared boilerplate template for every violation type, with a fill-in-the-blank paragraph for whichever ordinance applied, so nobody ever separately sized how many real examples backed each type. Same mistake wearing a different coat: sizing an eval by the unit that was true once, one generic notice, and never checking it again as the real number of violation types grew.
Swap the trigger and it still runs.
Speed: an interviewer wants a number before the meeting ends. Round to 15 per violation type, five types, 75 as the working total, revisit the exact split once real case files start arriving.
Cost: the agency can only pay a compliance reviewer for 20 hours of grading a month. At 20 minutes an example, that's about 60 examples a month, so the data set gets built across two monthly passes, not one sitting.
The model got better: a cleaner ordinance-citation lookup cuts down ambiguous wording. That doesn't shrink the data set. Fewer ambiguous cases just means fewer of the 70 examples turn out to be genuinely hard, which is a different thing from needing fewer of them.
Where people run it wrong.
They size the data-dependency section off the whole feature at launch, one line, one number, and never revisit it as the feature grows into new categories.
They let one overall "looks right" score stand in for every category, so a badly under-covered high-risk category hides behind a healthy average.
They size once and never re-check as the stakes grow, more customers, more categories, past what the original number was ever built to answer.
How to use it live. Say the equation before naming a single number: total equals the sum across categories of typical plus edge, never one flat sentence covering everything at once. That buys a few real seconds to do the arithmetic instead of guessing out loud.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Writing a PRD for an AI feature
- #1 What sections does an AI PRD need that a standard PRD does not?
- #2 Write the problem statement section for an AI meeting-summary feature.
- #3 How do you specify expected behaviour when the output is generated text?
- #4 Describe how to document the failure modes section of an AI PRD.
- #5 What belongs in the scope section about what the model will explicitly not do?
- #7 Write the quality bar section for a customer-facing classification feature.