Artifact critiqueFoundationalEval-Driven Specification / Writing a PRD for an AI feature / #16

Critique this requirement: the assistant should give helpful and accurate responses.

The direct answer
Replace it with a weekly blind fact-check rate. Every week, pull 100 factual chat answers at random and have a person, working only from the real order record, the real product spec sheet, and the current return policy, check whether every material fact in the bot's answer is actually right. That percent matched is the criterion. At 96 percent or higher, the bot keeps answering that category on its own. Between 88 and 96, it still answers, but a person spot-checks the sample within a day. Under 88, that category comes off live answering until two clean weeks pass.
Do this, in order
  1. Rewrite "helpful and accurate" as a weekly blind fact-check rate with a real sample size and a cutoff.Why: a line nobody can fail a build on lets a wrong answer ship for months before anyone checks it.
  2. Build the checking harness, a person re-deriving the true answer while blind to the bot's own answer, before the bot answers a single question live.Why: measuring this after launch means the first quarter's cost is already spent.
  3. Score dimension, material, and return-window questions on a tighter cutoff than color and styling questions.Why: the blended average can sit fine while the handful of large-furniture answers, the expensive ones, are exactly where it's wrong.
  4. Score the answer's own facts, never the tone or the thumbs-up.Why: a warm, apologetic wrong answer earns a five-star rating from a customer who hasn't found out yet that it was wrong.
  5. Track the return report too, but act on the weekly rate, not on the report.Why: a quarterly finance report is months behind the drift it's describing.
  6. Post the weekly number where support and product both see it, not one engineer's notebook.Why: a number nobody outside the bot team looks at won't stop the next wrong answer from shipping.

How to answer this, stage by stage

Seven moves. Say the numerator and the denominator out loud, by name, or the number you give next just sounds like a guess wearing a percent sign.

1
Say plainly why the line can't fail a build
Say it like this
"Both of those words sound like exactly the right thing to require. But neither one has a sample behind it, a number, or a day where we'd actually stop shipping because of it. Before I rewrite it, I want to say that plainly, because that's most of the job right here."
Why this works
Naming the flaw first proves you're not about to bolt a percent sign onto the same vague sentence and call it fixed.
2
Put one real product and one real question under it
Say it like this
"Say this is Driftwood Home Goods, a furniture site, and their chat widget answers questions like 'will this bed frame's box fit up my stairwell' using the product's spec sheet and the order record. Get that one wrong and a customer pays real money for a return she didn't need to make."
Why this works
A criterion about "the assistant" is empty until someone would actually act on what it says.
3
Split "helpful" from "accurate," they're not one promise
Say it like this
"'Accurate' means the fact in the answer matches the real order or the real product. 'Helpful' means the customer didn't need to call a person afterward. A wrong answer can still feel helpful in the moment, warm, fast, confident. That gap between the two is the whole danger in this requirement."
Why this works
This is the reframe. Say it in plain words, or the rate you write next only measures tone, not truth.
4
Name the numerator and the denominator
Say it like this
"Here are the two words that decide everything. The denominator is every factual chat answer the bot gives in a week, roughly 1,900 of them. The numerator is how many of those, checked by a person working from the real order system and the real spec sheet, turn out to actually be right."
Why this works
Get either of these wrong, especially the denominator, and the rate can look great while missing exactly the answers it exists to catch.
5
Pick the sample and set the bands
Say it like this
"We don't check all 1,900 by hand. We pull 100 at random every week, and a person re-derives the true answer blind to what the bot said. At 96 percent matched or higher, the bot keeps answering that category live. Between 88 and 96, it still answers, but a person spot-checks the sample within a day. Under 88, that category comes off live answering until two clean weeks pass."
Why this works
A metric nobody acts on is a dashboard decoration. This line is what turns it into a real criterion.
6
Say out loud how the number gets gamed
Say it like this
"The cheap way to keep this score high is to only sample order-status questions, the easy yes-or-no ones, and skip the dimension and material questions, which are exactly where a wrong answer costs a return freight bill. I'd name that risk myself before anyone else finds it the hard way."
Why this works
Naming the abuse case yourself is what tells an interviewer you understand metrics, not just picked one off a list.
7
Close on why a good thumbs-up score would have hidden this
Say it like this
"CSAT sat at 4.6 out of 5 the entire time this was drifting, because a warm, confident wrong answer still gets a thumbs-up before the customer finds out it was wrong. The fact-check rate would have crossed 96 percent by week three and 88 percent by week six, months before a return report ever caught it."
Why this works
This ties the rate back to what's actually at stake, and it's the line worth leaving the room with.
If you only get through two stages Stages 4 and 5 are the answer. Say what the numerator and the denominator actually are, then say the sample size and the bands. Everything else on this list is how you defend that rate under follow-up.

Let's learn

Say a furniture site builds a chat widget that answers questions about an order, a return, or a product, right on the page, instead of a customer waiting for a person.

Before the bot, a support team of eleven handled every one of the roughly 2,400 chats a month by hand, about nine minutes each, with a six-minute wait at the busiest hour.

Now the bot answers most of it itself: first reply in under ten seconds, about 78 percent of chats closed with no human needed at all, and a thumbs-up rate sitting at 4.6 out of 5 for eight straight months.

Here's the part that matters. That steady 4.6 was never the problem, and it was never really the win either. It's what happens to a wrong answer that feels warm and certain while it's being said, before anyone finds out it was wrong.

A score can hold at 4.6 out of 5 for eight months while the exact answer that's going to cost real money hasn't been checked once.
Two panels side by side. Left, a form with a thumbs-up and the number 4.6 out of 5, labeled warm answer, hasn't found out yet. Right, a person shrugging, labeled the box won't fit, same chat, three weeks later.
A five-star chat and a customer who hasn't found out yet, in the same moment

What that costs at its worst: a queen bed frame ships in one long carton, not flat panels. A customer named Priscilla, moving into a third-floor walk-up, asked the widget if the box would fit up her 34-inch stairwell. The bot answered from the "assembled width" field on the product page, 30 inches, and said yes. The real shipping carton was 38 inches. It didn't fit. Her movers had to haul it back down and come again. Driftwood paid $210 in return freight and a $75 goodwill credit for the second trip, and Priscilla's review was titled "the chatbot lied to me about the box."

Needs the harder band

Questions that can cost real money

  • Carton dimensions versus assembled dimensions for large furniture
  • Return window and final-sale exclusions
  • Material composition, for allergy or care questions
  • Delivery-date guarantees tied to a refund
These are the answers a wrong call actually costs a return, a refund, or a customer's trust. They get the tighter bar.
95 percent is already plenty

Low-stakes, reversible questions

  • What colors a rug comes in
  • Whether two pieces "match" in a styled room photo
  • Store pickup hours
  • General pairing or styling suggestions
A wrong guess here costs one more message, not a return freight bill. Leave these loose.
Knowledge spark: why "box size" and "assembled size" aren't the same number A dresser assembled is maybe 30 inches wide. Shipped, its parts get packed into a carton built to protect corners and hardware in transit, which is often bigger than the finished piece, sometimes by a foot. A spec sheet that only lists the assembled size is missing the one number a stairwell actually cares about.
The leading edge: weekly fact-check match rate, sampled blind
97% 96% 93% 91% 89% 86% 84% 82% Wk 1 Wk 2 Wk 3 Wk 4 Wk 5 Wk 6 Wk 7 Wk 8, finance flags it
96 percent or higher, keep answering live
88 to 96, a person spot-checks the sample
under 88, category comes off live answering
The rate crossed under 96 in week 3 and under 88 in week 6. Nobody was watching it, because nobody had built anything to watch.
The lagging outcome: returns coded "item not as described" after a chat
40 a month
Months 1 to 2, average
260 a month
Month 8, when finance flagged it
The return report lags the real drift by months, because a customer has to receive the item, discover the mistake, and file a return. The fact-check rate above had already crossed both thresholds before a single one of these returns showed up on a spreadsheet.
The choice I'd take back I wrote "the assistant should give helpful and accurate responses" as the quality line in the PRD, and it sounded like the responsible, catch-all thing to require, so nobody in the room argued with it. Nobody could turn it into a number either, so no one built a sampling harness. I'd take that back and build the weekly random 100, checked blind, before the bot answered a single question live.

What I'd leave alone. The styling and color questions stay loose, and that's correct, not a shortcut. Guessing wrong on "does this look nice with a grey sofa" costs one more message back and forth. It never costs a return freight bill or a customer's trust. Rates are for the answers with real money riding on them. Turning every chat reply into a sampled, scored rate, including the harmless ones, just adds a checking job nobody needs.

The lesson. If you can't say the number that would make you pull a feature off live answering, you haven't written a real requirement. You've written a hope wearing a rule's clothes. Every acceptance line for something a model has to judge needs a sample, a denominator, and a cutoff sitting right next to it, or it isn't a requirement yet, it's a wish.

The 4.6 that never moved

You don't need this to answer the question. It's here so "1,900 answers a week" stops being an abstraction and starts being a Tuesday.

Mai Nguyen has been a product manager at Driftwood Home Goods for four years, the last two owning support tooling. She shipped the return-flow redesign that cut phone wait times from eleven minutes to under three, and she knows the furniture catalog well enough to spot a mislabeled fabric swatch from across the room.

The chat widget launched under her PRD eight months before this story starts. For most of that time it was the thing she was proudest of. A customer typed a question about a dresser's drawer material or a couch's delivery window and got an answer in seconds instead of a six-minute wait. Early on, Mai read a stack of transcripts every Friday, just to see how the bot was doing. It kept sounding right. Confident, friendly, specific. So the Friday read became a monthly skim, and then it mostly stopped, the way a habit thins when nothing bad ever seems to come of thinning it.

Nothing about the bot changed in that time. What changed was the catalog underneath it: new suppliers, updated carton sizes, a couple of products where the "assembled size" field on the page quietly stopped matching the real shipping box. The bot kept answering from whatever field the question seemed to match, the way it always had. It just started being wrong slightly more often, on exactly the questions where being wrong costs the most.

There was no single bad Tuesday. It built up the way weather does. The number that finally said something was three rows down in the quarterly finance report, "returns coded item not as described, chat-assisted orders," which had gone from about 40 a month to 260 over two quarters. Nobody had been watching that line specifically. It went to a different team, filed under freight cost, with no line back to the widget at all.

One of the 260 that final month belonged to Priscilla, moving into a third-floor apartment with a narrow, turning stairwell. She'd asked the exact question the widget exists to answer: would the bed frame's box make it up. The bot found the width field on the product page, 30 inches, well under her 34-inch stairwell, and said yes, confidently, the way it always did. It never checked the field against the real carton, because nobody had ever asked it to, and nobody was checking whether it had.

We didn't ship her a wrong number. We shipped her a confident one, and confidence is the thing she had no way to check.

The box arrived at 38 inches. Her movers got it halfway up before they had to turn around and carry it back down, then come back a second day. Driftwood paid the return freight and the second trip as goodwill. Priscilla left a review with the title still sitting on the product page months later: "the chatbot lied to me about the box."

Mai went back to the sign-off meeting in her head, the one from eight months earlier where the PRD's quality section got approved with one line under it: the assistant should give helpful and accurate responses. Everyone nodded. It sounded stricter than any number anyone could think of on the spot. Nobody turned it into a figure in that room, so it shipped unnumbered, and the meeting moved on to launch timing.

Here's the replay. If a random weekly sample of 100 factual answers, checked blind by a person against the real spec sheet, had existed from week one, the rate would have shown 97, then 96, then 93 by week 3, already under the line where a person should be spot-checking every sampled answer. By week 6 it would have shown 86, under the line where that category comes off live answering entirely. Both of those would have landed weeks before the finance report even existed, let alone got read.

With that rate running, Priscilla's stairwell question gets pulled into the harder band before her order ever ships. A person checks the carton size against the real product record. The box that doesn't fit never leaves the warehouse.

What Mai would tell herself, back in that sign-off meeting: she let a sentence with no way to fail through, because it sounded stricter than a percentage. It wasn't stricter. It was just untestable, and untestable always loses, quietly, to a number someone is actually checking.

Four letters, one number: LEAD in this answer

This is a spec question, not a behavior question, so the framework is LEAD, not FLIPS or GUARD. FLIPS needs a person whose habit already snaps in two settings; nothing here snaps, the requirement was broken from the day it was written. GUARD asks who can't push back against a decision already made. The whole question here is which number would have caught the bad answer first.

L, link. The real outcome, not the model's own confidence or the customer's thumbs-up. Here, it's whether a chat answer's stated fact is actually true, so a customer doesn't make a costly decision, a purchase, a return, on something wrong.
E, early signal. The number that moves before trust or cost breaks. Here, the percent of a random weekly sample of 100 factual chat answers a blind, independent check confirms is actually right, against the real order, spec sheet, and policy.
A, abuse. How the number gets hit without the real problem going away. Sample only the easy, high-confidence categories like order status, and skip the harder judgment calls, dimensions, materials, return windows, exactly where the cost lives.
D, decision. What changes at each score. At 96 or higher, the bot keeps answering that category live. Between 88 and 96, a person spot-checks the sample within a day. Under 88, that category comes off live answering until two clean weeks pass.
The check that keeps the rate honest Try shrinking the denominator on purpose. Sample only order-status questions, the ones the bot has always nailed. If the score barely moves, the sample was already close to the truth. If it jumps, the denominator was hiding the real problem, and that's exactly the shortcut a busy team reaches for without meaning to.

And if you want to be sure it really works, try it somewhere else

A field-service app for HVAC technicians has an assistant that answers "will this part fit this unit" questions from a parts catalog. Same shape of question. "The assistant should give helpful and accurate part guidance" is just as untestable as the furniture version.

L. A technician orders the right part on the first try, so a truck doesn't roll twice for the same repair.
E. Each week, pull 100 part recommendations at random. A parts specialist who hasn't seen the assistant's reasoning checks the part number against the real unit's manual and confirms fit.
A. Only sampling common residential units, the easy matches, and skipping older or commercial models, where the harder compatibility calls actually live.
D. Above the bar, the assistant's recommendation ships straight to the order. In the middle band, a specialist confirms before the part ships. Below it, that unit family goes back to manual lookup until the catalog data is fixed.
Two panels side by side. Left, a plain box labeled ordered on the bot's word, part number never checked by hand. Right, a gauge labeled the weekly sample check, catches the wrong part before the truck rolls.
Same rate, a different desk

Swap the trigger and it still runs

  • The catalog gets bigger. Doesn't matter. Sampling 100 answers a week takes the same afternoon whether the catalog has 500 products or 50,000.
  • The bot gets slower. If a new inventory system triples response time, the weekly sample still stays at 100. You just check it on the same schedule.
  • The bot gets better than planned. If it starts reading updated supplier spec sheets reliably, rerun the same weekly check against the wider set of products it now covers. The process doesn't change. Only what's being sampled does.

Where people run it wrong

  • Sampling only the categories the assistant has always been strong on, so the score looks great on the population that was never the risk.
  • Treating the sample as a one-time launch check instead of a weekly habit, so a catalog update slips in unwatched for months.
  • Averaging every answer into one blended score instead of scoring the financially binding categories separately, so a bad run hides inside a good overall number.

How to say it if you're asked this cold

Buy yourself the time to build the number properly. "Before I give you a rate, let me say what 'accurate' is actually protecting here." That's not stalling. It's the L step, said out loud, and it gives you somewhere honest to stand while the real numerator and denominator take shape in your head.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and why not FLIPS or GUARD?
Tap to flip
ANSWER
LEAD, for a spec/metric question. FLIPS needs a habit that already snaps between two settings; nothing snaps here, the requirement was broken from day one. GUARD asks who can't push back on a decision already made. LEAD finds the number that would have caught the bad answer before real cost showed up.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Mai Nguyen, a product manager at Driftwood Home Goods who wrote the chat widget's PRD and had already shipped the team's return-flow redesign before it launched.
3 · THE HABIT
What did the team stop doing once the thumbs-up rate looked fine?
Tap to flip
ANSWER
Reading chat transcripts by hand. A weekly Friday read became a monthly skim, then mostly stopped, because the bot kept sounding right.
4 · THE RATE
What's the actual rate in this story, in plain words?
Tap to flip
ANSWER
Out of a random 100 of the roughly 1,900 factual chat answers the bot gives each week, the percent a blind, independent check confirms is actually right against the real order, spec sheet, and policy.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing "the assistant should give helpful and accurate responses" into the PRD as the quality bar, with no sample, denominator, or cutoff attached, because it sounded like the responsible catch-all and nobody could turn it into a number on the spot.
6 · THE NUMBER
The rewritten rate samples ______ answers a week, and needs ______ percent matched to keep a category answering live.
Tap to flip
ANSWER
One hundred answers. Ninety-six percent. Between 88 and 96, a person spot-checks the sample within a day. Under 88, that category comes off live answering.
7 · THE REPLAY
Same drift, new design, what changes?
Tap to flip
ANSWER
The rate drops under 96 percent in week 3 and under 88 in week 6, both weeks before finance's report surfaces anything in week 8. Priscilla's stairwell question gets the harder band before her order ships, and the box that doesn't fit never leaves the warehouse.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
An HVAC field-service assistant recommending replacement parts. Its early signal is the percent of a random weekly sample of 100 part recommendations a parts specialist confirms actually fits the real unit.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these would count as a genuine pass under Mai's rewritten rate?
  • A. The bot's answer includes a link to the returns policy page.
  • B. The customer didn't send a follow-up question.
  • C. A person, working from the real order record and product spec sheet, confirms the bot's stated carton width matches the real one.
  • D. The chat ended with a thumbs-up rating.
Show hint
Three of these describe how the chat ended. Only one describes a person checking the real item.
Show answer
C. A link, a quiet customer, and a thumbs-up are all facts about the conversation, not proof the stated fact was right. The only real pass is a person checking it against the real record.
Fill in the blank
2. The rewritten rate samples ______ answers a week, needs ______ percent matched to keep a category answering live, and pulls that category off live answering entirely below ______ percent.
Show hint
The first number is in the walkthrough's stage 5. The other two are in the direct answer at the top.
Show answer
100. 96. 88. Between 88 and 96, a person spot-checks the sample before the category keeps running unsupervised.
True or false
3. True or false: this gets fixed by telling the bot to sound less certain when it isn't sure.
  • True
  • False
Show hint
Ask whether a softer tone changes what a person would find if they checked the fact by hand.
Show answer
False. A hedge changes the tone, not whether the box actually fits. It could even pass an "accurate" check by never committing to a claim, while still leaving the customer to find out the hard way. There's still no sample checking the real facts.
Short answer
4. Name a place in this same chat widget where leaving "helpful and accurate" loose is actually fine, not a rate.
Show hint
Look for the question where a wrong guess costs a follow-up message, not a return.
Show answer
Model answer: "Styling and color questions, like what colors a rug comes in or whether two pieces match in a room photo. A wrong guess there costs one more message back and forth, not a return freight bill, so it doesn't need a sampling harness."
Short answer, apply it yourself
5. Pick a product you use yourself. Name one place it makes a "helpful and accurate"-sounding promise you've never actually seen tested. What would the weekly-sample version of that check look like?
Show hint
Look for a claim with no number attached, like "verified," "in stock," or "compatible."
Show answer
Model answer: "A grocery delivery app that marks items 'in stock.' The weekly check: pull 100 orders where an item was marked in stock, and have a person compare that against what the store actually had on the shelf that day. If the confirmed-in-stock rate drifts, I'd know the inventory feed was going stale months before a wave of cancelled orders told me."
Multiple choice
6. Which of these would make the weekly rate look good without actually protecting a real customer?
  • A. Raising the sample from 100 answers to 300.
  • B. Only sampling order-status questions, the easy yes-or-no ones, and skipping dimension and material questions.
  • C. Having two people independently check each sampled answer.
  • D. Publishing the weekly score where the whole support team can see it.
Show hint
The dangerous move hides the hardest, most expensive questions inside an average full of easy ones.
Show answer
B. The blended score can sit at 96 percent while the riskiest cases, box sizes and material questions on large furniture, never get checked at all. That's exactly the abuse the rate has to be designed against.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more