ConceptAdvancedEval-Driven Specification / Writing a PRD for an AI feature / #23
What does the success criteria section look like when the metric is subjective quality?
The direct answer
Write two numbers into the success criteria, not one. The pass rate, and the rate at which two reviewers, scoring the same image on their own, land on the same verdict. Score every sampled image twice, blind, before either number gets reported. If the two reviewers agree on fewer than about eight close calls in ten, the rubric isn't a bar yet, it's opinion with a percentage sign on it, and no pass rate built on top of it means anything.
Do this, in order
Write the bar as two numbers: the pass rate and the agreement rate.Why: a pass rate that nobody checks against a second reviewer is one person's taste wearing a number.
Score every sampled image twice, independently, before reporting either number.Why: split the week's sample across reviewers to save time, like this team did, and you can never even measure whether two people see the same thing.
Track the agreement number from the first pre-launch batch, not the week before ship.Why: a bad bar shows up in week two if someone is watching. It only shows up in week eight if nobody is.
Write the rubric in checkable lines, not one soft phrase like "looks professional."Why: a vague line lets two different tastes both feel justified passing 9 in 10 images, which is exactly how a loose rubric games itself.
Set a real agreement floor that halts a launch and forces a rewrite, not a number on a slide nobody reopens.Why: a threshold nobody has to answer to is decoration, not a gate.
Leave the obvious defects, a missing product, a floating shadow, off this whole process.Why: two people rarely disagree about a broken render, so agreement tracking only earns its keep on the real judgment calls.
How to answer this, stage by stage
Six moves. Name the two numbers before the story, or the answer sounds like a definition instead of something you'd actually put in a PRD.
1
Scope it to one tool and one review pair
Say it like this
"Let's make this concrete. Say we build Backlot, a tool that takes a seller's plain product photo and drops it into a staged scene, a backdrop, a couple of props, real-looking light. Yusuke Handa runs product for Backlot, and he's the one who has to decide what the success criteria section actually says."
Why this works
A subjective-quality question stays abstract until one person's document has to survive someone reading it literally.
2
Name the plan before the reasoning
Say it like this
"I'll use LEAD here. L is the real thing a subjective bar is supposed to protect, E is the early signal, whether reviewers actually agree on the rubric, A is how a subjective bar gets gamed, and D is what changes at different agreement levels."
Why this works
Stating the plan up front tells the interviewer you're about to make a decision, not describe a feeling.
3
Say what the question is actually checking
Say it like this
"This sounds like a formatting question, what goes in one section of a document. It's really asking whether 'good enough to ship' means the same thing to everyone allowed to say yes, or whether it secretly means whatever the person grading that day happened to think."
Why this works
This is where the answer stops being about a document template and starts being about whether the metric can be trusted at all.
4
State the decision straight
Say it like this
"So here's what goes in that section. Two numbers, not one. Pass rate needs to clear 90 percent, and inter-rater agreement, two reviewers scoring the same image blind, needs to clear a real floor before I trust that 90 percent at all. If agreement is low, I don't ship on the pass rate. I rewrite the rubric first."
Why this works
This names an actual line you could paste into a PRD, not a value like "quality" with nothing under it.
5
Make the hidden number real
Say it like this
"Here's why that matters. At Backlot, the weekly pass rate sat right around 90 percent for eight straight weeks, so the review looked healthy every time. Underneath it, when we finally had two reviewers score the same 50 images instead of splitting them, they agreed on the verdict about 56 percent of the time by week eight. One image they disagreed on had a reflection that didn't match the light we'd added. It shipped. A seller flagged it as fake within a day."
Why this works
A real number that contradicts the number everyone was watching does more work than a paragraph explaining the risk.
6
Name the gaming path, set the thresholds, and close
Say it like this
"Nobody has to cheat for this to happen. The rubric said 'looks professional and realistic,' one soft line, and two different reviewers can both feel completely right passing 9 in 10 images under a line that loose, even while applying two different bars. Under about a 65 percent agreement floor, that's not a healthy metric, it's noise with a percentage sign on it. Above 75 percent, I trust the pass rate on its own. Between the two, I keep both numbers on the same dashboard and don't ship past a week where agreement is dropping, whatever the pass rate says."
Why this works
This closes on a real threshold and a real action, so the answer reads as a plan you'd actually run, not a definition you memorized.
If you only get through two stages
Stages 4 and 6 are the answer. Say the two numbers, then say why a soft rubric line lets both of them look fine at once. Everything else here is how you defend that when someone pushes back.
Let's learn
What happens when "good enough" depends on which person happens to be looking at it? Say we build Backlot, a tool that takes a seller's plain product photo, a lamp on a kitchen counter, and drops it into a staged scene: a matching backdrop, a couple of props, warm light that actually falls the right way.
Knowledge spark: what's a success criteria section?
The part of a plan that says exactly what "good enough to ship" means, in numbers, so nobody has to guess later whether it worked.
Before Backlot, a small seller paid a freelance photographer about 75 dollars a shot and waited a week for a batch back, or shot it themselves on a kitchen counter and hoped it looked fine. Backlot turns the same photo into a staged version in under thirty seconds, included in the plan they already pay for.
Here's the turn. A script can check that the product itself rendered cleanly, no missing handle, no cropped edge. Whether the whole staged scene actually looks professional enough to publish is a judgment call. Only a person can make that call. So the team wrote one line into the rubric, looks professional and realistic, and one number into the success criteria: ship when 90 percent of a weekly sample of 50 images score 4 or 5 out of 5. Nobody wrote down whether two people looking at the very same photo would land on the same score.
A pass rate that never moves is not proof the bar held. It can be proof nobody ever checked whether two people were reading the same bar.
Knowledge spark: what's inter-rater agreement?
How often two people scoring the same thing, on their own, land on the same verdict. Not whether they're right. Whether they're reading off the same bar.
At its worst, the pass rate held right around 90 percent for eight straight weeks, so the weekly review looked healthy every single time. Underneath it, once two reviewers finally scored the same 50 images instead of splitting the pile between them, they agreed on the verdict barely 56 percent of the time. The number that was supposed to prove the bar was solid had never actually tested whether it was one bar.
Passed on paper. Redone by hand a week later.
Inter-rater agreement between two reviewers scoring the same weekly sample, by week
75% and up, agreement backs up the pass rate
65 to 75%, the rubric is going shaky
under 65%, the rubric isn't one bar anymore
Agreement crossed the 65 percent floor in week five, three weeks before a new reviewer's onboarding exercise forced anyone to look. By week eight it had drifted to 56 percent, while the pass rate everyone was watching never moved off 90.
Sellers asking to redo a staged photo by hand, first 30 days live
24
Agreement tracked, rubric fixed at week five
187
What actually happened, gap found in week eight
Twenty four redo requests in the version where agreement got caught and fixed early. A hundred and eighty seven in the version it didn't. Every one of those had individually cleared a 4 or 5 score from at least one reviewer.
The choice I'd take back
We wrote "ship at 90 percent, scored by our review team," as if our review team scored with one shared brain. We never had two reviewers score the same image and compare notes before we shipped. That was fine while the model mostly handled plain products nobody could disagree about. It stopped being fine the day glass, jewelry, and patterned fabric entered the sample, and reflection and texture turned into real judgment calls.
What I'd leave alone. A render with a missing product, a floating shadow, or a cropped edge doesn't need two people's agreement. Everyone calls that a fail on sight. Keep that one as a single reviewer and a checklist. Agreement only earns its keep on the calls two reasonable people could actually split on.
The lesson. A pass rate only tells you the number moved where you wanted. It never tells you whether the two people deciding it were pointing at the same bar. For a subjective metric, the real success criteria isn't the score. It's whether two people scoring blind would hand you the same one.
Now here is the same thing as a story
Use this version when you've got the time to sit with it. The short version is above. This is for when you want to feel why it mattered.
Yusuke Handa has run product for Backlot for two years, since before the staging model could reliably handle anything shinier than a ceramic mug.
For the first several months, the rubric earned its trust honestly. Every Friday, two stylists scored a sample of the week's staged photos, and on the rare image where their scores landed on opposite sides of the line, whoever was in the room argued it out loud for two minutes and moved on.
Then the checking faded, in three small steps, none of them looking wrong at the time. First, to save review hours as the weekly sample grew, the team started splitting the 50 images between two reviewers instead of doubling up on each one, so a single image only ever got one opinion. Second, the rubric still said one line, looks professional and realistic, because writing anything longer felt like overbuilding a document for something a trained eye would just know. Third, nobody was saving which reviewer scored which image, only the week's overall pass rate.
Ottilie Baranski had reviewed staged photos since the rubric's first draft, and she was good at it, fast, and fair by her own lights. She passed images that had the right mood even when a shadow fell a little wrong. Nobody told her that was generous. Nobody was checking.
Then Dashiell Okonkwo joined the review rotation in week eight, and Yusuke gave him an onboarding exercise: blind re-score last week's already-passed batch before joining the live rotation.
Dashiell's scores came back, and 22 of the 50 images he flagged as an obvious fail, mismatched reflections, shadows falling the wrong way for the added light, had already passed the week before. Yusuke pulled the actual overlap number. Fifty-six percent. Barely better than a coin flip.
We didn't lose a single image that quarter to a bad render. We lost the one thing that made a passing score mean the same thing twice.
It would be easy to say Ottilie was too soft. She wasn't. She was fast, she was consistent with herself, and every score she gave was a real judgment, carefully made. Nobody decided the bar should be whatever Ottilie's eye called good enough. It just quietly became that, one Friday at a time, because nothing was ever checked against a second opinion.
So here's the decision Yusuke would take back.
Two years earlier, in the meeting where the rubric first got written down, someone asked whether "looks professional and realistic" was specific enough. It felt like overbuilding a document for something a trained eye would just know. Yusuke remembers agreeing they'd tighten it if it ever became a real problem.
I'd write the agreement check into the success criteria from day one instead. Same rubric, same weekly sample, but every image gets scored twice, blind, and the agreement rate rides right next to the pass rate on the same weekly report.
Here's the replay. Same eight weeks, agreement tracked from week one. It crosses the 65 percent floor by week five instead of surfacing by accident in week eight. That crossing is the trigger, not Dashiell's onboarding exercise. Yusuke pulls that week's disagreements and adds three checkable lines to the rubric: lighting direction has to match the backdrop, reflections have to match what's actually in frame, texture has to read as real up close. Agreement climbs to 81 percent by week eight. Backlot still ships on schedule. In the first month live, sellers ask to redo a staged photo by hand about 24 times instead of the 187 it actually took before anyone rewrote the rubric.
One success criteria section watches whether this week's sample passed. The other watches whether passing still means the same thing it did in week one.
And the thing Yusuke would tell himself, back in that first meeting: the pass rate was never the risk. Writing a rubric loose enough that nobody would ever be caught disagreeing with it was.
LEAD, applied to a bar two people have to agree on
This reads like a formatting question, what goes in one section of a document. Underneath it, it's still asking whether "good enough" means the same thing to everyone allowed to say yes. That's LEAD, run on a rubric instead of a single score.
L, link. What a subjective success bar actually protects. Not a percentage, a bar two different reviewers would land on the same way. If "good enough" depends on which person is scoring, the number is decoration.
E, early signal. The earliest thing worth watching isn't the pass rate. It's whether two reviewers scoring the same image, blind, agree with each other. That number can sit low for weeks while the pass rate everyone's watching looks perfectly fine.
A, abuse. How this gets gamed, usually without anyone meaning to. Write the rubric as one soft line, and almost any honest reviewer can pass 9 in 10 images under it and feel right doing it, because a vague bar bends to whoever's reading it. The pass rate always looks healthy. It's just not measuring one bar anymore.
D, decision. What actually changes. Above 75 percent, agreement looks fine, trust the pass rate. Between 65 and 75, that's the rubric going unclear, and it earns a rewrite before the next sample. Under 65, stop trusting the pass rate as a ship gate until the rubric is specific enough that two people land in the same place.
The check that proves the bar is real
Pull one week's sample and have two reviewers score it blind, separately, the same way. If they land within a point of each other on most images, the rubric is a bar. If they don't, it's a suggestion with a percentage sign on it.
And if you want to be sure it really works, try it somewhere else
An AI tool called Aftercare drafts the discharge note a vet hands a pet owner at checkout, in plain language, so they know what to watch for at home. A vet at Denhollow Animal Hospitals reads each draft before it goes out and scores it on one line: clear enough for a worried owner to follow. Same shape of problem. A pass rate that holds steady can hide the same thing it hid at Backlot, one vet's idea of clear quietly becoming the whole bar, because nobody's checking whether a second vet would call the same note clear.
L. Whether a pet owner actually understands what to watch for at home, not whether this week's notes cleared review at the same rate as last week.
E. Whether two vets, reading the same draft note on their own, agree it's clear enough, tracked by week, not the pass rate by itself.
A. Nobody has to cut a corner here either. "Clear enough" bends to whichever vet is reading it, and a vet fluent in the jargon can read a note as obviously clear that a first-time pet owner would find frightening or confusing.
D. Agreement steady and high, trust the pass rate. Agreement dropping, especially on notes about serious conditions, stop trusting the pass rate and get two vets reading the draft, not a reminder to write clearly.
Same note. Two calls.
Swap the trigger and it still runs
The model gets faster. Doesn't matter. A quicker draft can still be judged by one soft rubric line, and the agreement problem doesn't care how fast the draft arrived.
The review team gets bigger. Doesn't help on its own. More reviewers just means more candidates for the bar to quietly become whichever one of them reads the most images.
The model gets genuinely better at the judgment call. Track agreement anyway. That's the one case it should climb on its own, and watching it is what makes "it got better" a fact instead of a feeling.
Where people run it wrong
Treating a steady pass rate as proof the bar held, with nobody ever checking two scores against each other.
Writing the rubric as one mood word, professional, clear, on-brand, instead of lines specific enough that two people would read them the same way.
Waiting for an onboarding exercise or a customer complaint to reveal the gap, instead of scoring double from the first pre-launch week.
How to use it live
Say the split first. "Before I answer, I want to separate two questions, is this one image good enough, and is the rubric itself something two people would actually agree on." That's not stalling. It's naming which question you're really being asked, and it buys you room to build the real answer instead of taking the pass rate's word for it.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does each letter stand for here?
Tap to flip
ANSWER
LEAD. L is the real bar a subjective metric protects, E is the early signal, whether reviewers actually agree, A is how a loose rubric gets gamed, D is what changes at each agreement level.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Yusuke Handa, who has run product for Backlot, an AI tool that stages product photos for e-commerce sellers, for two years, and who wrote its rubric himself.
3 · THE HABIT
What did the team never start doing, even while the pass rate looked fine?
Tap to flip
ANSWER
Scoring the same image with two reviewers and comparing notes. They split the weekly sample between reviewers to save time, so no image ever got a second opinion.
4 · THE REAL SIGNAL
What number was actually low while the pass rate looked healthy?
Tap to flip
ANSWER
Inter-rater agreement. Two reviewers scoring the same 50 images agreed on the verdict only about 56 percent of the time by week eight, while the pass rate held near 90 percent the whole time.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing the rubric as one soft line, "looks professional and realistic," and never scoring an image twice to check whether that line meant the same thing to two different reviewers.
6 · THE NUMBER
Agreement was about ______ percent by week eight. Above ______ percent, the answer says you can trust the pass rate on its own.
Tap to flip
ANSWER
56 percent. Above 75 percent. Below 65, the number stops being a healthy metric and starts being noise with a percentage sign on it.
7 · THE REPLAY
Same rubric, agreement tracked from day one, what changes?
Tap to flip
ANSWER
The floor gets crossed and caught in week five instead of surfacing by luck in week eight. Yusuke adds three checkable rubric lines, agreement climbs to 81 percent by week eight, and sellers redoing a photo by hand drops from 187 in the first month to about 24.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
Aftercare, an AI tool that drafts vet discharge notes. Its early signal is whether two vets, reading the same draft on their own, agree it's clear enough for a worried pet owner.
Check yourself Score: 0 / 0
Fill in the blank
1. The weekly pass rate held around ______ percent for eight straight weeks, while two reviewers scoring the same images actually agreed only about ______ percent of the time by week eight.
Show hint
The numbers sit in the paragraph right after the turn in the first section.
Show answer
90 percent. 56 percent. A pass rate that never moved was hiding an agreement number that was never healthy.
True or false
2. True or false: because Backlot's pass rate stayed near 90 percent every week, the rubric was clearly working.
True
False
Show hint
Ask what a steady pass rate can hide underneath it.
Show answer
False. A steady pass rate only shows the number moved where the team wanted. It doesn't show whether two reviewers were reading the rubric the same way.
Multiple choice
3. Which decision would you actually take back, based on this answer?
A. Hiring a second reviewer to work through the sample faster.
B. Writing the rubric as one soft line and never scoring an image twice to check it.
C. Letting Backlot generate staged photos at all.
D. Paying reviewers by the image instead of by the hour.
Show hint
Look for the decision named right after the highlight line in the story section, not a new dial.
Show answer
B. The reversal is writing one vague rubric line and never checking whether two people agreed on it, not a staffing or pricing change.
Short answer
4. Name a place in this same rubric where a disagreement doesn't need agreement tracking at all.
Show hint
Look for the kind of mistake nobody would argue about.
Show answer
Model answer: "A render with a missing product, a floating shadow, or a cropped edge. Everyone calls that a fail on sight, so it doesn't need two reviewers comparing notes, just one reviewer and a checklist."
Short answer, apply it yourself
5. Pick a product you use yourself where a person judges something subjective, is this review helpful, is this photo a good fit, is this joke funny. What would you check to know if the bar means the same thing to two different reviewers?
Show hint
Look for anywhere a person, not a script, decides pass or fail.
Show answer
Model answer: "A food delivery app where support agents decide whether a photo of a damaged order qualifies for a refund. If one agent's approval rate runs much higher than another's, I'd have two agents blind-score the same batch of photos and check how often they actually agree, not just compare their approval rates."
Multiple choice
6. If Backlot had caught the agreement drop in week five instead of week eight, what would most likely have happened to the number of sellers redoing a staged photo by hand that first month?
A. No change, it would still be 187.
B. Fewer, because the rubric would have been rewritten and agreement fixed before most of the month's images shipped under the loose bar.
C. More, because rewriting the rubric mid quarter would have confused the reviewers.
D. It depends only on how many photos Backlot generated that month, not on agreement.
Show hint
Think about what catching agreement earlier actually buys you: time to fix the rubric before it ships broadly.
Show answer
B. Catching it in week five gives Yusuke time to add checkable rubric lines and rebuild trust in the pass rate before most of the month's images ever ship under the loose bar, exactly what the replay in Section 2 shows.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.