ConceptIntermediateEval-Driven Specification / Writing a PRD for an AI feature / #3
How do you specify expected behaviour when the output is generated text?
The direct answer
Replace "expected behaviour" with a weekly blind condition-match rate on the descriptions themselves. Every week, pull a random 75 of the new listings the generator wrote, and have a person check every claim about the item's condition, material, and function against the real intake photos and the checklist the warehouse team filled out. That percent match is the spec. At 94 percent or higher, the generator keeps publishing straight to the site. Between 86 and 94, an editor reads the sample before anything goes live. Under 86, that category goes back to a human writer until two clean weeks pass.
Do this, in order
Rewrite "expected behaviour" as a weekly blind condition-match rate with a real sample size and a cutoff.Why: a line nobody can fail a build on lets an overstated listing publish for months before anyone checks it.
Build the checking step, a person comparing each claim to the real intake record, before the generator publishes a single listing live.Why: measuring this after launch means the first bad batch has already been rented out.
Score condition and function claims on a tighter cutoff than styling language.Why: the blended average can sit fine while the handful of big, expensive returns, the ones with real damage, are exactly where the generator gets it wrong.
Score the claim against the record, never against how nice the sentence reads.Why: a warm, confident description of a stained sofa still reads beautifully right up until someone unwraps it.
Watch for the generator learning to make no checkable claim at all.Why: a description with nothing specific in it passes the rate and still leaves the renter with nothing to trust.
Post the weekly number where merchandising and renter support both see it, not one engineer's notebook.Why: a number nobody outside the team looks at won't stop the next overstated listing from going live.
How to answer this, stage by stage
Six moves. Say the sample size and the two cutoffs out loud, by name, or the number you give next just sounds like a nicer sentence.
1
Say plainly why the line can't fail a build
Say it like this
"'Expected behaviour' sounds like exactly the right bar for a writing feature. But nobody can point to the day we'd pull it off the site because of that line, and there's no sample or number behind it. I want to say that plainly before I fix it, because that's most of the job right here."
Why this works
Naming the flaw first proves you're not about to bolt a percent sign onto the same vague sentence and call it done.
2
Put one real product and one real listing under it
Say it like this
"Say this is Roost, a furniture rental company. When a table or a sofa comes back from a renter, it goes through intake, gets a condition checklist, and the generator writes the new listing straight off that checklist and the photos. Get a condition claim wrong and the next renter opens a box expecting 'excellent condition' and finds a water ring someone already wrote down."
Why this works
A requirement about "the assistant" stays empty until you can say who acts on a wrong sentence and what it costs them.
3
Split what "expected" is protecting from how the sentence sounds
Say it like this
"'Expected behaviour' is really two things wearing one phrase. One is: does the sentence read well. The other is: is every fact in it true. A description can be warm, specific, and completely wrong about the stain on the arm. That gap is the whole danger sitting inside this requirement."
Why this works
This is the reframe, the L step. Skip it and the rate you write next only grades the prose, not the truth.
4
Name the sample, the check, and the bands
Say it like this
"Here's the number. Every week we pull 75 of the roughly 600 new listings the generator wrote that week. A person checks every condition, material, and function claim in the text against the real intake photos and checklist, blind to what the generator said. Match 94 percent or higher, it keeps publishing straight to the site. Between 86 and 94, an editor reads the sample before it goes live. Under 86, that category goes back to a human writer until two clean weeks."
Why this works
This is the E step, the actual answer to the question. Everything before it was clearing ground for this one line.
5
Say out loud how the number gets gamed
Say it like this
"The cheap way to keep this score high is to write descriptions that never make a checkable claim at all, 'ready to rent,' 'great value,' nothing specific enough to fail. I'd flag that risk myself, because a description with no facts in it passes the rate and still leaves the renter with nothing real to trust."
Why this works
Naming the abuse case yourself tells the interviewer you understand metrics, not that you picked one off a list.
6
Close on what changes at each number
Say it like this
"Above 94, we trust the machine and move fast. Between 86 and 94, a person reads before it ships. Below 86, we stop and go back to a person writing it. That's the whole point of the number. It tells us, on any given Tuesday, whether to keep trusting the generator or pull it back, instead of waiting for a renter to find the stain first."
Why this works
A metric nobody acts on is a dashboard decoration. This line is what makes it a real spec, not a wish.
If you only get through two stages
Stages 3 and 4 are the answer. Say what "expected behaviour" was actually hiding, then say the sample size and the two cutoffs. Everything else on this list is how you defend that rate under follow-up.
Let's learn
Say a furniture rental company writes the public listing for every item the moment it's ready to rent again, using AI instead of a person.
Before the generator, one in-house writer worked from the warehouse's condition checklist and photos, about six minutes a listing, and could get through around 45 a day.
Now the generator drafts a listing the second intake closes: under a second, no backlog, all 600 or so items Roost turns over in a week, cleared without anyone breaking a sweat.
Here's the part that matters. The sentences didn't get worse. They read warmer and more specific than the old ones ever did. What changed is that the generator started describing returned, used furniture the way it would describe brand new stock, calling a scuffed dresser "excellent condition" while the checklist sitting right next to it, filled in by hand at intake, said otherwise.
The listings didn't get less accurate right away. They got more confident about things nobody had checked.
A five-star listing and a stain that's already on file, in the same afternoon
What that costs at its worst: Odette Charles was moving into a studio and rented a dining table listed "excellent condition, no visible wear." The intake note on file said otherwise: a water ring across the top and one leg that wobbled on a flat floor. The generator had pulled "excellent condition" from the general furniture description template instead of reading the checklist fields that would have caught it. Odette's table arrived exactly as the checklist described and nothing like the listing. Roost paid for a swap delivery and a partial refund, and Odette cancelled her plan the same week, with a review that read "the listing was not the table I got."
Needs the harder band
Claims that can cost a return
Condition language: "excellent," "no visible wear," "like new"
Material claims: "solid oak" versus veneer, real fabric content
Function claims: "drawers slide smoothly," "no wobble"
Dimensions on large pieces, where a wrong number means a stairwell problem
These are the claims a wrong word actually costs a swap, a refund, or a cancelled plan. They get the tighter bar.
95 percent is already plenty
Styling and taste language
"Would look great in a small apartment"
General color and pairing suggestions
Room-fit suggestions ("pairs well with a light rug")
Tone and phrasing choices that don't claim a fact
A wrong guess here costs a scroll past, not a return delivery. Leave these loose.
Knowledge spark: why a checklist and a description aren't automatically the same thing
The condition checklist is filled in by a person at intake, box by box: scratches, stains, missing parts, how the drawers move. The description is written afterward, from that checklist plus the photos. Nothing stops a text generator from reading the photos and reaching for a generic, positive word instead of the specific box someone already checked.
The leading edge: weekly condition-match rate, sampled blind
94 percent or higher, keep publishing live
86 to 94, an editor reads the sample first
under 86, category goes back to a human writer
The rate crossed under 94 in week 3 and under 86 in week 6. Nobody was watching it, because nobody had built anything to watch.
The lagging outcome: condition disputes filed by renters, per month
4 a month
Months 1 to 2, average
31 a month
Month tied to Odette's dispute
Disputes lag the real drift by weeks, because a renter has to receive the item, notice the gap, and file a complaint. The condition-match rate above had already crossed both thresholds before a single one of these disputes reached a support queue.
The choice I'd take back
I wrote "generated descriptions must accurately and appealingly represent the item's true condition" into the PRD as the whole quality bar, and it sounded like the responsible, catch-all thing to require, so nobody in the room argued with it. Nobody could turn it into a number either, so no one built a checking step. I'd take that back and build the weekly random 75, checked blind against the real intake record, before the generator published a single listing live.
What I'd leave alone. The styling and pairing language stays loose, and that's correct, not a shortcut. Guessing wrong on "would suit a small apartment" costs one more scroll, never a swap delivery. Rates are for the claims with real money riding on them. Turning every sentence into a sampled, scored fact, including the harmless ones, just adds a checking job nobody needs.
The lesson. If you can't say the number that would make you pull a feature back to a human, you haven't written a real requirement. You've written a hope wearing a rule's clothes. Every acceptance line for something a model has to judge needs a sample, a denominator, and a cutoff sitting right next to it, or it isn't a requirement yet, it's a wish.
The Monday check nobody noticed stopping
You don't need this to answer the question. It's here so "600 listings a week" stops being an abstraction and starts being a Tuesday.
Every Monday morning, for the first five months, Saoirse Byrne pulled twelve fresh listings and read them next to the real intake photos, just to see how the generator was doing. She's been a product manager at Roost for three years and wrote the PRD for the listing writer herself. She knows the catalog well enough to tell a veneer dresser from a solid one in a thumbnail photo.
For most of those five months, that Monday check was the best fifteen minutes of her week. The listings sounded right. Specific, warm, and, as far as she could tell, true. So the twelve became six, and then whichever ones she had time for, the way a habit thins when nothing bad ever seems to come of thinning it.
Nothing about the generator changed in that stretch. What changed was the catalog underneath it: more suppliers, more returned stock, and a growing share of listings for used furniture instead of new. The generator kept reaching for the same confident, positive language it always had. On new stock that was true. On returned stock, more and more often, it wasn't.
There was no single bad Tuesday. It built the way a slow leak does. The first sign anyone outside the team saw was a message from Iris Novak, one of the warehouse condition checkers, in a shared channel: "why does the listing say excellent condition on the oak table when I marked it moderate wear a week ago?" Nobody had an answer ready.
The stain wasn't missed by the checklist. It was missed by the sentence that got written after the checklist was already right.
Two weeks after Iris's question, Odette Charles rented that same style of table, listed the same way: "excellent condition, no visible wear." Her copy shipped with the water ring and the wobble both intact, both of them sitting in the intake record the whole time. She filed a dispute, Roost swapped the table and refunded part of the month, and she cancelled her plan before the replacement even arrived.
Saoirse went back to the sign-off meeting in her head, the one from a year earlier where the PRD's quality line got approved with a single sentence under it: generated descriptions must accurately and appealingly represent the item's true condition. Everyone nodded. It sounded stricter than any number anyone could think of on the spot. Nobody turned it into a figure in that room, so it shipped unnumbered, and the meeting moved on to launch timing.
Here's the replay. If a random weekly sample of 75 listings, checked blind by a person against the real intake record, had existed from week one, the rate would have shown 97, then 95, then 93 by week 3, already under the line where an editor should be reading every sampled listing. By week 6 it would have shown 84, under the line where that category comes off automatic publishing entirely. Both of those land before Odette's table ever ships, weeks before Iris had to ask her question in a channel at all.
With that rate running, the oak table style gets pulled into the editor-review band the moment the score first slips under 94. A person reads the batch against the real checklist. The table with the water ring never goes out described as flawless.
What Saoirse would tell herself, back in that sign-off meeting: she let a sentence with no way to fail through, because it sounded stricter than a percentage. It wasn't stricter. It was just untestable, and untestable always loses, quietly, to a number someone is actually checking.
One line became four questions: LEAD, applied here
This is a spec question, not a behavior question, so the framework is LEAD, not FLIPS or GUARD. FLIPS needs a person whose habit already snaps between two settings, and nothing here snaps, the requirement was broken from the day it was written. GUARD asks who can't push back against a decision already made. The whole question here is which number would have caught the wrong sentence first.
L, link. The real outcome, not how polished the sentence sounds or how fast it gets published. Here, it's whether a listing's stated condition is actually true, so a renter's item arrives the way it was described.
E, early signal. The number that moves before trust or cost breaks. Here, the percent of a random weekly sample of 75 generated listings a blind, independent check confirms is actually true, against the real intake photos and checklist.
A, abuse. How the number gets hit without the real problem going away. Write descriptions that never make a checkable condition claim at all, so there's nothing left to fail, and the score climbs while the renter still learns nothing useful.
D, decision. What changes at each score. At 94 or higher, the generator keeps publishing live. Between 86 and 94, an editor reads the sample within a day. Under 86, that category goes back to a human writer until two clean weeks pass.
The check that keeps the rate honest
Try stripping every condition word out of a batch of listings on purpose and see what the score does. If it climbs while the listings say less, the rate is rewarding silence, not accuracy, and the sample needs to score "made no checkable claim" as its own kind of failure.
And if you want to be sure it really works, try it somewhere else
A veterinary clinic's AI writes an after-visit summary email to pet owners, covering the exam findings, the diagnosis, and what to do next. "The summary should be accurate and reassuring" is exactly as untestable as the furniture version.
L. The owner follows the plan correctly, so a sick pet doesn't miss a dose or a needed follow-up visit.
E. Each week, pull 40 summaries at random. A vet tech who wasn't in the room checks every medical claim, the diagnosis wording, the dose, the follow-up urgency, against the real chart notes.
A. Only sampling routine wellness visits, the easy ones, and skipping urgent or multi-medication cases, where a wrong dose line actually hurts something.
D. Above the bar, summaries send straight to the owner's inbox. In the middle band, a vet tech reads it before it sends. Below it, that visit type goes back to a person writing the summary until the model's fixed.
Same rate, a different desk
Swap the trigger and it still runs
The catalog gets bigger. Doesn't matter. Sampling 75 listings a week takes the same afternoon whether the warehouse holds 500 items or 50,000.
The generator gets slower. If a new photo pipeline triples the time to draft a listing, the weekly sample still stays at 75. You just check it on the same schedule.
The generator gets better than planned. If it starts reading the condition checklist more carefully than a person skimming photos would, rerun the same weekly check against the wider set it now covers. Only what's being sampled changes.
Where people run it wrong
Sampling only the categories the generator has always nailed, so the score looks great on the population that was never the risk.
Treating the sample as a one-time launch check instead of a weekly habit, so a new supplier or a new photo angle slips in unwatched for months.
Averaging every listing into one blended score instead of scoring the big, hard-to-return items separately, so a bad run hides inside a good overall number.
How to say it if you're asked this cold
Buy yourself the time to build the number properly. "Before I give you a rate, let me say what 'accurate' is actually protecting here." That's not stalling. It's the L step, said out loud, and it gives you somewhere honest to stand while the real numerator and denominator take shape in your head.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and why not FLIPS or GUARD?
Tap to flip
ANSWER
LEAD, for a spec question about generated text. FLIPS needs a habit that already snaps between two settings; nothing snaps here, the requirement was broken from day one. GUARD asks who can't push back on a decision already made. LEAD finds the number that would have caught the wrong sentence before real cost showed up.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Saoirse Byrne, a product manager at Roost, a furniture rental company, who wrote the listing writer's PRD and used to read a Monday sample of listings against the real intake photos herself.
3 · THE HABIT
What did Saoirse stop doing once the listings kept sounding right?
Tap to flip
ANSWER
Reading a weekly sample of listings against the real intake photos. Twelve a week became six, then whichever ones she had time for, then mostly nothing.
4 · THE RATE
What's the actual rate in this story, in plain words?
Tap to flip
ANSWER
Out of a random 75 of the roughly 600 new listings the generator writes each week, the percent where every condition, material, and function claim in the text matches the real intake checklist and photos, checked blind by a person.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing "generated descriptions must accurately and appealingly represent the item's true condition" into the PRD as the whole quality bar, with no sample, denominator, or cutoff attached, because it sounded like the responsible catch-all and nobody could turn it into a number on the spot.
6 · THE NUMBER
The rewritten rate samples ______ listings a week, and needs ______ percent matched to keep a category publishing live.
Tap to flip
ANSWER
Seventy-five listings. Ninety-four percent. Between 86 and 94, an editor reads the sample first. Under 86, that category goes back to a human writer.
7 · THE REPLAY
Same drift, new design, what changes?
Tap to flip
ANSWER
The rate drops under 94 percent in week 3 and under 86 in week 6, both weeks before Odette's table ever ships. The oak table style gets pulled into editor review the moment it first slips under 94, and the table with the water ring never goes out described as flawless.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
A veterinary clinic's AI after-visit summary email. Its early signal is the percent of a weekly sample of 40 summaries a vet tech confirms matches the real chart notes on diagnosis, dose, and follow-up urgency.
Check yourself Score: 0 / 0
True or false
1. True or false: this gets fixed by telling the generator to sound less certain when it isn't sure about an item's condition.
True
False
Show hint
Ask whether a softer tone changes what a person would find when they check the claim against the real intake record.
Show answer
False. A hedge changes the tone, not whether the water ring is real. It could even pass an "accurate-sounding" read by never quite committing to a claim, while still leaving the renter to find out the hard way. There's still no sample checking the real facts against the record.
Multiple choice
2. Which of these would count as a genuine pass under the rewritten rate?
A. The listing includes three positive adjectives about the item.
B. The renter didn't leave a bad review.
C. A person, working from the real intake checklist and photos, confirms every condition claim in the text is true.
D. The listing published within a second of intake closing.
Show hint
Three of these describe how the listing looked or landed. Only one describes a person checking it against the real item.
Show answer
C. Nice adjectives, a quiet renter, and a fast publish are all facts about the listing's surface, not proof the stated condition was true. The only real pass is a person checking it against the real record.
Fill in the blank
3. The rewritten rate samples ______ listings a week, needs ______ percent matched to keep a category publishing live, and pulls it back to a human writer below ______ percent.
Show hint
The first number is in the walkthrough's stage 4. The other two are in the direct answer at the top.
Show answer
75. 94. 86. Between 86 and 94, an editor reads the sample before the category keeps publishing unsupervised.
Short answer
4. Name a place in this same catalog where leaving "accurate and appealing" loose is actually fine, not a rate.
Show hint
Look for the kind of line where a wrong guess costs a scroll past, not a return delivery.
Show answer
Model answer: "Styling and pairing language, like 'would look great in a small apartment' or general color suggestions. A wrong guess there costs one more scroll, not a swap delivery, so it doesn't need a sampling harness."
Short answer, apply it yourself
5. Pick a product you use yourself. Name one place it makes an "accurate"-sounding promise you've never actually seen tested. What would the weekly-sample version of that check look like?
Show hint
Look for a claim with no number attached, like "verified," "in stock," or "matches your size."
Show answer
Model answer: "A secondhand clothing app that writes AI descriptions saying 'true to size, no flaws.' The weekly check: pull 50 items marked no flaws, and have a person compare the listing text against the real photos the seller uploaded. If the confirmed-accurate rate drifts, I'd know the model was gaming size and condition claims months before returns spiked."
Multiple choice
6. Which of these would make the weekly rate look good without actually protecting a renter?
A. Raising the sample from 75 listings to 150.
B. Writing descriptions that never make a specific condition claim, so there's nothing to check.
C. Having two people independently check each sampled listing.
D. Posting the weekly score where merchandising can see it.
Show hint
The dangerous move removes the very thing the rate is supposed to be checking.
Show answer
B. A vague, claim-free description can score close to perfect because there's nothing false left in it to catch, while the renter still learns nothing real about the item's condition. That's exactly the abuse the rate has to be designed against.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.