ConceptIntermediateEval-Driven Specification / Writing a PRD for an AI feature / #5
What belongs in the scope section about what the model will explicitly not do?
The direct answer
List it in three tiers, in this order. First, anything that could cause real harm if the model just did it quietly, like applying an edit to a live contract with nobody confirming that exact contract. Second, the limits on what the model can actually judge, like whether a clause holds up in a given state's courts. Third, what you're simply not building yet, marked as later, not never. Every single item needs a one-line reason next to it, or it isn't a boundary, it's a guess wearing a boundary's clothes.
The ranking, by what breaks first if skipped
Write the will-not-do list in three tiers: safety exclusions, then capability limits, then deferrals, each with a stated reason.Why: skip the ordering and the list reads as one flat pile, and the item that actually protects someone gets no more weight than the item that's just not built yet.
Name the model's real job before writing a single exclusion.Why: dependency. You can't rank boundaries around a job nobody's named, so whoever builds next decides the job by default, usually toward whatever's fastest to ship.
Check today's list for the cheap tell: does every item carry a real reason, or does it just say "keep it simple"?Why: evidence. A reason-free item is already drifting toward whatever's easiest to build, and this catches it before a user hits the gap.
Watch for the moment a new feature quietly erases a boundary the list assumed.Why: reversibility. This is exactly where an unstated limit turns into a real, unfixable cost, usually through a shortcut nobody thought to check against the list.
Keep the deferral tier loose, and let it move as the product matures.Why: these are reversible by nature. Spending ranking energy protecting a "later" item is wasted effort that belongs on the safety tier instead.
Revisit the whole list whenever a new feature changes what "review" or "confirm" even means.Why: a list written once at launch quietly keeps protecting whatever the product looked like then, not what it does now.
How to answer this, stage by stage
Seven moves. The trap in this question is answering it as a features-we-skipped list, when it's actually asking whether you can tell a boundary that protects someone from a boundary that's just a todo item.
1
Pick one real product to answer against
Say it like this
"Let me make this real. Say a legal tech company called Marlstone Legal Technologies builds a tool named Draftline. It reads a vendor contract and suggests the edits a lawyer would mark up by hand, before anyone at the client company opens the document. I'll answer against that."
Why this works
Grounds an abstract scope question in one real product, so the ranking that follows isn't hypothetical.
2
Name your method before you use it
Say it like this
"I'd use ORDER for this. Rank the categories of 'will not do' by what breaks first if that boundary is missing, not just list everything the model happens to avoid."
Why this works
Signals a method up front, so the answer doesn't sound like a brainstorm someone's making up on the spot.
3
Say what the question is actually testing
Say it like this
"This isn't really asking me to list features I chose not to build. It's asking whether I know the difference between a boundary that protects someone and a boundary that's just a todo item for later. Those two belong in different places on the page."
Why this works
Separates the real judgment call from a surface reading that treats the list as a wish list.
4
Give the ranked answer straight
Say it like this
"Three tiers, in this order. First, anything that could cause real harm if the model just did it quietly, like applying an edit to a live contract with nobody confirming that exact contract. Second, the limits on what the model can actually judge, like whether a clause holds up in a given state's courts. Third, what you're simply not building yet, like non English contracts. Every single item needs a one-line reason next to it, or it doesn't belong on the list."
Why this works
This is deliverable 0, said out loud, in the order that actually matters.
5
Show what has to be true before you can write it
Say it like this
"None of this ranking works until you've named the model's actual job. Draftline proposes edits, a person applies them. Once that's settled, the will-not-do list is just the fence around that one sentence. Skip naming the job, and engineering names it for you, usually by building whatever's fastest to ship."
Why this works
This is the dependency step. It shows the order isn't arbitrary, one decision genuinely has to happen before the next one can.
6
Back it with the one incident that proves the order
Say it like this
"Here's what it cost when this went unranked. Marlstone let a batch feature ship that could auto-apply Draftline's suggestions to a whole stack of contracts, nobody opening them individually. Four months in, a contracts specialist was auto-applying eight in ten batches. One of them quietly dropped the liability cap on a six-hundred-thousand-dollar vendor deal. It went out, got signed by both sides, and by the time anyone reopened that file, it wasn't a draft anymore."
Why this works
A cheap ranking rule plus a real cost is what makes the order defensible instead of just tidy.
7
Close on the rule, not the list
Say it like this
"So: safety exclusions first, because those are the ones you can't take back once a person's relied on the gap. Capability limits second, because those tell a reviewer where their own judgment still has to do the work. Nice-to-have deferrals last, marked as later, not never. If an item on that list doesn't have a reason next to it, it isn't a boundary. It's a guess wearing a boundary's clothes."
Why this works
Ends on the literal ranking the question asked for, defended instead of just listed.
Let's learn
Here's what happens when a product plan says everything a tool will do, and stays quiet about what it won't: the missing half gets written anyway, by whoever builds next, usually toward whichever choice ships fastest.
Say a legal tech company builds a tool that reads a vendor contract and suggests edits, the way a lawyer would mark up a paper draft with a pen.
Before it, a three-person contracts team at a manufacturing company read every vendor agreement by hand. About 40 minutes each, roughly 15 a week.
The tool cut that to about 90 seconds a contract. It read the draft, suggested the changes a lawyer would likely make, and left every suggestion sitting there for a person to accept or reject, one at a time.
Then the company added a second setting: apply every suggested edit to a whole batch of contracts at once, nobody opening them individually. Nothing in the plan said the tool couldn't. Nobody had ever written down that it shouldn't.
Knowledge spark: what's an indemnification cap?
A line in a contract that limits how much one side has to pay the other if something goes wrong. Without it, the amount owed has no ceiling. A vendor contract with the cap in place might limit a payout to what the deal is worth in a year. Without it, a single bad shipment could cost far more than the contract was ever for.
Naming Draftline's actual job comes first. Ranking the exclusions and handing them to engineering only makes sense after that.
Four months in, the contracts team was auto-applying about eight in ten batches without a second look. One of those held a vendor NDA where the tool matched the wrong template and quietly dropped the cap on how much the company could owe if the deal went wrong. It went out, got signed by both sides, and by the time anyone opened that file again, it wasn't a draft anymore. It was a contract.
We didn't just miss a clause. We lost the only chance to catch it before a signature made it permanent.
Share of vendor NDA batches auto-applied with nobody opening them, by week
Batches auto-applied with no individual review
By week ten, four in five batches went out with nobody opening them. The one that mattered was in that four.
The choice I would take back
Marlstone's launch plan for Draftline covered "a person is always in the loop" in one line, with no reason attached and no mention of what a batch feature would do to it. I would write it as its own rule: no edit reaches a live contract without someone confirming that exact contract, batch mode included. Engineering can still own the batching itself, the queue, the send button. That part really is theirs.
What I would leave alone. Draftline guessing which clauses need a second look in the first place is fine to leave loose. If it flags a boilerplate line that didn't need attention, a reviewer skips it in two seconds. That kind of miss costs nothing. Only the exclusions that skip the person entirely are worth a written rule.
The lesson. A plan that only says what a tool will do is half a plan. The other half is what happens the day someone builds a feature the first half never imagined.
Now here is the same thing as a story
The short version is above. Keep reading if you want to feel why the boundary needed its own line, not an assumption riding on the word "review."
Dario Renfrew could read a vendor NDA once and tell you within a page whether Somerhale Manufacturing was carrying more risk than it should. Nine years in contracts does that to a person.
Draftline arrived in February. Every Tuesday morning, Dario opened its queue of suggested edits and worked through them one at a time: accept, tweak a phrase, reject the one that didn't fit. By ten, he'd cleared a stack that used to eat his whole morning.
The suggestions kept being right. So he stopped reading every line of each redline and started skimming the summary instead. Then he stopped skimming most weeks, since the queue always cleared clean. By June, opening Draftline's queue was mostly a formality.
Then Somerhale rolled out batch auto-apply for routine paperwork, mutual NDAs mostly, the kind that rarely needed a human touch anyway. Dario turned it on for his slowest week of the quarter, the way anyone would.
He kept using it. Forty contracts a batch, sometimes two batches a week, and none of them ever came back wrong. So auto-apply stopped being the exception and became how Tuesdays worked.
The mistake, when it came, wasn't loud. Draftline matched one NDA to a template from a different kind of deal, one with no liability cap at all, and suggested removing the cap language. Auto-apply accepted it. Nobody opened that file. It went back to the vendor. It came back signed.
One of these you can fix whenever you get to it. The other already came back with a signature on it.
It was never about the clause being wrong. It was about the moment nobody could take it back anymore.
Six weeks later, the vendor's lawyer called to confirm the missing cap was intentional. It wasn't. Somerhale was now carrying no ceiling on what it could owe under a contract worth about six hundred thousand dollars a year, and there was no version of that Tuesday where opening the file three weeks earlier wouldn't have caught it.
A year earlier, in the meeting where the team scoped Draftline for launch, someone had asked whether "never applies without review" needed its own line in the plan, or whether it was already covered by "a person is always in the loop." The second answer felt obviously true. Nobody wrote down what "in the loop" would still mean once a batch button existed.
I would go back and write it as its own rule: no edit reaches a live contract without someone confirming that specific contract, batch mode included. With that rule, Dario still has to open the file that mattered. Under a minute, not three weeks late, because Draftline's redline was right on 39 of the 40 contracts in that batch. The fortieth gets caught before a signature, not after one.
What I'd tell myself, back in that meeting: "review" isn't one word with one meaning. The day somebody builds a faster way around it, you find out which meaning you actually meant.
The ranking underneath "will not do," spelled out
GUARD would fit if the question were about who gets hurt when Draftline is wrong. This sits earlier than that: which boundary has to be named first, and what order it gets written down in. That's ORDER's job.
O, outcome. Every item on Draftline's will-not-do list protects one thing: that a person's trust in "this was reviewed" is never wrong. Not that the tool is cautious. That the specific claim it makes about itself is true.
R, reversibility. If the exclusion never gets written down, engineering ships whatever's fastest, batch auto-apply included, and the gap between "a person reviewed this" and "nobody did" only shows up after a contract's already signed. You can't undo the six weeks a six-hundred-thousand-dollar deal sat with no liability cap.
D, dependency. None of this ranks until Draftline's actual job is named: it proposes, a person applies. Skip that, and there's nothing to rank exclusions around, so whoever builds next decides the job by default.
E, evidence. Cheap to check: does every item on today's list carry a real reason, or does it just say "keep it simple"? Marlstone's launch plan had one line for the whole boundary, no reason attached, and no mention of what a batch feature would do to it.
R, rank. Safety exclusions first: no edit applies to a live contract without per-contract confirmation. Capability limits second: no read on whether a clause holds up in a given jurisdiction. Deferrals last, marked as later: no non-English contracts yet. Writing the whole list at once, unranked, doesn't save time. It just moves the decision to whoever's fastest to ship.
The check that keeps this ranking honest
Swap the outcome and the order should move. If a wrong Draftline edit only ever cost someone a quick fix and an apology, the safety tier could sit lower, and a lighter rule might genuinely be enough. It ranks first here because a real company carried a real six-figure exposure for six weeks on the strength of an edit nobody reviewed.
Rank it again, somewhere the wrong call reaches a pet, not a contract
Overmere Veterinary Group, a small chain of clinics, built Roundnote: a tool that turns a vet's spoken notes during an exam into a discharge summary and a take-home dosage sheet for the pet's owner.
O. Every version of Roundnote's will-not-do list protects one thing: that a dosage on a take-home sheet reflects what the vet actually decided, not what the model guessed from a similar case.
R. A missing language on the handout is fixable any week, someone translates it. A dosage sheet already emailed to an owner, based on an edit the vet never saw, is not. The pet's already home.
D. The clinic's "quick-send" shortcut only makes sense once someone's decided what counts as a dosage the vet actually confirmed, a clinical call, not a formatting one.
E. Cheap to check: pull the last month of quick-send batches and ask how many dosages the vet actually opened before they went out. Overmere's first quarter of quick-send: about one in forty.
R. Same order: the lead veterinarian owns the rule that every dosage gets a one-tap confirmation, even inside quick-send. The engineering team owns the batching and the send queue underneath it.
Discharge notes a vet confirmed before they reached an owner, out of 40 sent through quick-send that month
Same tool, same quick-send shortcut. The only thing that changed was whether the boundary had its own written line.
Swap the trigger and it still runs
Somerhale doubles its contract volume overnight. The order doesn't move. The safety tier matters more, not less, since more volume means more chances for a batch to slip past review.
Draftline's underlying model gets noticeably more accurate. Doesn't reorder either. A more accurate model still needs someone to confirm the one time it's wrong, or nobody can tell that time from the other 39.
Marlstone starts selling Draftline to outside law firms instead of in-house teams. Doesn't reorder. The boundary protects the same thing regardless of who's on the other side of it.
Where people run it wrong
Treating "the model rarely gets it wrong" as a reason to skip stating the boundary, when rare-and-irreversible is exactly the combination that needs a written rule.
Writing the will-not-do list once at launch and never revisiting it when a feature like batch mode changes what "review" means.
Letting engineering write the whole list because they'll build it anyway, as if deciding the boundary and building around it were the same job.
How to use it live
Say the outcome out loud before naming any exclusion. "Every item on this list protects one thing, that review means what we say it means." Then rank from there. Naming the outcome first turns a list of features into an argument you can defend.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits "what belongs in the will-not-do list," and why not GUARD?
Tap to flip
ANSWER
ORDER, for ranking which boundary has to be named first. GUARD is for naming who can't push back once a system causes harm. This question is about ranking judgment calls before anyone's been harmed, not a power imbalance.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dario Renfrew, senior contracts specialist at Somerhale Manufacturing, the company using Draftline, an AI contract-redlining tool built by Marlstone Legal Technologies.
3 · THE HABIT
What habit let a missing boundary go unnoticed for months?
Tap to flip
ANSWER
Dario used to open Draftline's queue and check every suggested edit by hand. As the suggestions kept being right, he stopped opening most batches individually, and auto-apply quietly became the default.
4 · THE DEPENDENCY
What has to be decided before a will-not-do list can even be written?
Tap to flip
ANSWER
The model's actual job. Draftline proposes edits, a person applies them. Every item on the list is a boundary around that one sentence, and you can't rank boundaries around a job nobody's named.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Writing "a person is always in the loop" as one line covering all review, instead of its own explicit rule. It made sense before a batch feature existed. Nobody revisited it once one did.
6 · THE NUMBER
The auto-applied share of Draftline batches climbed from ___ percent to ___ percent over ten weeks.
Tap to flip
ANSWER
5 percent to 82 percent. That climb is the boundary quietly disappearing, in a number you can actually watch happen.
7 · THE REPLAY
Same batch, same wrong template match, the rule written this time. What changes?
Tap to flip
ANSWER
Auto-apply still runs on 39 of the 40 contracts. The fortieth, the one with the mismatched liability clause, needs a one-tap confirmation first, and gets caught before a signature instead of six weeks after one.
8 · THE TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what plays the role of Draftline's safety tier there?
Tap to flip
ANSWER
Overmere Veterinary Group's Roundnote. The equivalent is a rule that every take-home dosage gets a one-tap vet confirmation, even inside the quick-send shortcut.
Check yourself Score: 0 / 0
Fill in the blank
1. Before Draftline, Somerhale's contracts team spent about ______ minutes redlining a vendor contract by hand. Draftline cut that to about ______ seconds.
Show hint
It's the number that makes the whole tool worth building in the first place.
Show answer
40 minutes, and 90 seconds. That's the value a will-not-do list exists to protect, not slow down.
Multiple choice
2. Which three tiers does this answer rank the will-not-do list into, in order?
A. Nice-to-have deferrals, then safety exclusions, then capability limits
B. Safety exclusions, then capability limits, then nice-to-have deferrals
C. Capability limits, then safety exclusions, then deferrals
D. One flat list, with no order between the items
Show hint
Ask which category you can't take back once someone's relied on the gap.
Show answer
B. Safety exclusions come first because they're the ones a person could rely on and get burned by. Capability limits come second. Deferrals, the things you're simply not building yet, come last.
True or false
3. True or false: since Draftline rarely suggested a wrong edit, the batch auto-apply feature didn't need its own stated boundary. Say why.
True
False
Show hint
Think about what happens when something rare also can't be undone.
Show answer
False. Rare and irreversible together is exactly the combination that needs a written rule. A mistake that happens once in 40 but can't be undone once it ships is worse than one that happens often but costs nothing.
Multiple choice
4. What does the dependency step argue in this answer's ORDER?
A. Engineering should decide the model's job, since they build it anyway
B. The model's actual job has to be named before any exclusion can be ranked around it
C. The will-not-do list should be written before the model's job is decided
D. Safety and capability limits can be ranked in either order, it doesn't matter
Show hint
Ask what has to exist before you can rank anything at all.
Show answer
B. Without a named job, "propose, not apply," there's nothing for the exclusions to be boundaries around, so whoever builds next decides the job by default.
Short answer, apply it yourself
5. Pick an AI feature you use or are building. What's one thing it's quietly deciding not to do that's never actually written down anywhere?
Show hint
Look for the boundary a user would only discover by hitting it.
Show answer
Model answer: "A meeting-notes tool that never flags when it mis-hears a name. Nothing states that limit, so a reader trusts every name in the summary equally, even the ones the model guessed at."
Short answer, the number question
6. If Draftline's auto-applied share had climbed to only 20 percent instead of 82, would the safety exclusion still need to rank first? Say what changes and what doesn't.
Show hint
Reversibility is about whether a wrong edit can be undone, not about how often the feature gets used.
Show answer
Model answer: "What changes: the odds of hitting the gap in any given month, so it might take longer to surface. What doesn't change: one auto-applied contract with a dropped cap is still a signed, unrecoverable contract, so the exclusion still ranks first regardless of the percentage."
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.