CaseIntermediateQuality, Cost & Token Economics / Eval design for product teams / #2
Design an eval for a feature that drafts email replies.
A quality rubric can only tell you an email reads well. It cannot tell you whether it is true. Those are two different questions, and most evals for a drafting feature only ever answer the first one.
The direct answer
Build the eval on a golden set of real historical email threads with the reply the person actually sent as the answer key, split into routine, sensitive, and judgment-call buckets. Grade the sensitive and judgment buckets against a claims-and-commitments checklist, not a fluency-and-tone score, because a wrong price or an invented promise reads exactly as smooth as a true one. Ship once those two buckets clear a set pass rate across two straight weekly audits, not one clean run, and leave matching each sender's personal voice for a later version.
Do this, in order
Build the golden set from real threads, with the actual sent reply as the answer key, sliced by routine, sensitive, and judgment-call.Why: without a real reference answer, there is nothing honest to grade a draft against, and a fake rubric fills the gap instead.
Grade the sensitive and judgment tiers against a claims-and-commitments checklist, not a writing-quality score.Why: a wrong fact or an invented promise reads just as polished as a correct one, so fluency scoring can't see it.
Automate the routine tier's check, and spend human review only on the sensitive and judgment tiers.Why: full human review on every reply is not possible at real volume, so the expensive check has to go where a miss actually costs something.
Set the ship bar as a pass rate held across two straight weekly audits, not one perfect run.Why: a model that drafts probabilistically will never hit zero misses, so the real question is whether the miss rate is low enough that a person catches what's left.
Refresh the golden set every week with fresh sent replies.Why: prices and features change, and a golden set that never updates starts grading against answers that are themselves out of date.
Leave tone-matching to each sender's personal voice out of version one.Why: it's a polish problem, not a trust-breaking one, and building it now spends review budget the fact-checking layer needs more.
How to answer this, stage by stage
Nobody's grading whether you can name a rubric. They're grading whether the rubric you name could actually catch the mistake that gets someone fired.
1
Scope it to one real team before designing anything abstract
Say it like this
"Let's ground this in one real team. Northwind sells dispatch software to trucking companies. Six SDRs answer about a hundred fifty inbound prospect replies a day between them. Right now every one of those gets typed by hand."
Why this works
An eval design question answered in the abstract turns into a list of metric names. One team makes it a real decision with real volume behind it.
2
Say your structure out loud
Say it like this
"I'll walk through this as SPARK: where the person starts today, what habit I want the eval to build, the one design decision it hangs on, what happens the first time that decision is wrong, and what I'm leaving out on purpose."
Why this works
Tells the interviewer you have a plan for the next four minutes instead of thinking out loud with no landing point.
3
Name the failure before naming the eval
Say it like this
"Before I design the eval, I want to say plainly what it has to catch. Not 'does the draft sound professional.' It has to catch a draft that sounds completely professional and is still wrong."
Why this works
Stops you from defaulting to a generic writing-quality rubric, the trap almost every candidate falls into on this question.
4
Give the one decision the eval hangs on
Say it like this
"The eval runs on a golden set: four hundred real threads, each one paired with the reply the person actually sent, split into routine, sensitive, and judgment-call buckets. On the sensitive and judgment buckets, drafts get graded against a claims-and-commitments checklist, not a writing score."
Why this works
This is the concrete, defensible answer to the question. Everything else in the answer exists to protect this one decision.
5
Prove it against a real near miss
Say it like this
"Here's why that has to be the anchor. One of our SDRs, Jarrah, almost sent a draft telling a prospect our starter plan supports single sign-on. It doesn't, that's an enterprise-only feature. The draft read beautifully. Our old quality check would have passed it without blinking. She caught it herself, on her last read, by luck."
Why this works
A real near miss makes the risk concrete instead of hypothetical, and shows you can name a specific failure, not just gesture at one.
6
State the ship bar as a threshold, not a promise
Say it like this
"I wouldn't ship this off one clean run. I'd ship once the sensitive and judgment tiers clear ninety percent on the checklist across two straight weekly audits, because a model that drafts probabilistically is never going to hit zero. The real question is whether what's left gets caught before it goes out."
Why this works
Shows you think in eval-set thresholds, not deterministic promises. A model output is never "always right," and a strong answer says so out loud.
7
Close by naming what you're deliberately not doing yet
Say it like this
"One thing I'd skip in version one: grading whether the draft sounds like that specific SDR's own voice. That's a polish problem. A false SSO claim is a trust problem. I'd rather spend the review budget catching the second one."
Why this works
Naming what you're leaving out on purpose is what separates a designed eval from a wish list that tries to grade everything at once.
Let's learn
Every day, six SDRs at Northwind type out about a hundred and fifty replies to prospects, by hand, and Jarrah Bishara does it faster and cleaner than anyone else on the team.
Northwind sells dispatch software to trucking companies, the kind of tool that shows a fleet manager where every truck is and when it'll arrive. A prospect writes back with a question about price, or how the tool connects to the routing software they already run, and someone on the SDR team has to answer fast, in a way that sounds like a person, not a form letter.
Jarrah personally answers about 35 of those a day. Each one takes her close to five minutes: reread the thread, check the price sheet if money comes up, remember what the last rep already promised, then write. That's just under three hours of her day, gone before lunch.
Knowledge spark: what's a golden set?
A pile of real past examples, each one paired with the answer a person actually gave, held back and never used to train the model. It's the answer key an eval gets checked against, instead of someone's opinion about whether an output looks fine.
Then Northwind ships Outrigger, a feature that drafts the reply for the SDR to check and send. Review and edit takes about 40 seconds. Reply time drops to well under an hour a day. The first month's numbers look great: on the weekly five-star check a manager runs by hand, SDRs rate 92 percent of their own edited drafts "good or great."
Ninety-two percent good was never telling anyone whether Outrigger was safe. It was telling them Outrigger writes nice sentences. Whether those sentences were true was a completely different question, and nobody had built a way to check it yet.
Old fluency rubric vs. claims-checked pass rate, by reply tier
Old rubric: reads wellNew eval: claims and commitments checked
The old rubric sits high and flat across all three tiers, since it only ever measured whether the writing sounded good. The claims-checked pass rate agrees on routine replies, then drops hard on the two tiers where a wrong fact actually costs something.
At its worst, a fast, fluent reply that states something untrue costs Northwind more than the slow, careful version ever did. A prospect who's told the wrong plan supports single sign-on, and finds out later it doesn't, doesn't just lose trust in one email. They stop trusting the sales team's word on everything else in the deal.
The choice that mattered
When Outrigger first shipped, quality was checked with a five-star "does this read well" score, sampled by a manager on about twenty drafts a week. That made sense for a tool that mostly saved typing time. It stopped making sense the day Outrigger started making claims on Northwind's behalf, and nobody went back to ask whether a fluency sample was still the right check for that new job.
What I'd leave alone: the routine tier, replies like confirming a call time or sending a spec sheet, doesn't need a claims checklist. Those are almost impossible to get factually wrong, and spending review time matching every one against a source document would take time away from the tier where a miss actually costs something.
The lesson: a draft can be completely well written and still be wrong in a way that costs real money. Ninety-two percent "reads great" told the team nothing about whether Outrigger's facts were straight. It only ever measured the one thing that was never actually at risk.
Now here is the same thing as a story
Read the long version below when you want to feel why a rubric that only checks fluency is such an easy trap to fall into, not just be told it is one.
Cass Ridler runs sales operations at Northwind, four years in, the person other reps come to when a deal's numbers don't add up. She built the weekly quality sample herself: twenty drafts pulled at random every Monday, read cover to cover, scored one to five on "does this sound like us."
Outrigger shipped in the spring. For the first few months, Cass's Monday sample came back steady, 4.6, then 4.7. She'd read all twenty, nod, close the tab, and get back to pipeline reviews.
She started reading twelve of the twenty. Then five. By month four she was skimming the average score and moving on without opening a single draft.
Then Jarrah forwarded her one email. Subject line: "did you see this." Nothing dramatic, just a screenshot of a draft that told a prospect Northwind's starter plan came with single sign-on. It doesn't, that feature sits on the enterprise plan only. Jarrah caught it on her own last read, right before she hit send.
Cass didn't retrain anything that afternoon. She pulled the full week of sensitive and judgment tier replies, forty-three of them, not the usual random twenty, and read every single one against the actual pricing page.
The rubric hadn't missed one lucky email. It had never once been built to look for this kind of mistake.
The real cost wasn't one wrong sentence sitting in a draft folder. It was that Cass no longer knew whether her Monday number had ever meant anything at all.
The decision that opened the door went back to the week Outrigger's QA process got set up. Someone asked, in passing, whether pricing and contract replies needed their own review instead of riding along in the random sample. The answer was no, twenty a week already covers the mix. Nobody came back to that question as the sensitive tier's stakes kept growing while the sample size stayed the same twenty.
Run the same near miss again with one change: every sensitive and judgment tier reply gets checked in full each week against a claims-and-commitments checklist, cross-referenced against the live pricing page, not sampled at random. The same false SSO claim still gets drafted by the model on day one. But it's caught in Tuesday's audit, a day later, six minutes of a reviewer's time, not fourteen months of a wrong assumption nobody revisited.
One design trusted one weekly number to answer for every kind of reply Northwind sends. The other asks the risky replies a harder question than the easy ones.
What I'd tell my own past self, back in that rollout meeting: the moment a check has to cover cases with wildly different costs, ask whether one sample size can really answer for all of them, or whether it only looked like enough because nothing expensive had gone wrong yet. Nobody asked. That's on the room, not on Jarrah, and not on the model.
Five moves, so the eval doesn't just grade nice writing
This is SPARK, run on an eval instead of a feature. The eval is the thing being designed, and it has to survive being wrong just like any other design.
SSituation. How does the job get done today, without you?
Jarrah, five minutes a reply, checking the thread, the price sheet, and what was already promised, entirely by hand, about 35 times a day.
Ground the eval in a real workflow, or the whole design floats free of what actually breaks.
Four things she cross-checks in her head, every single time, before the tool ever existed.
PPayoff. What habit do you want this to build?
Not "the team trusts Outrigger." The habit is Cass grading drafts against what the sender actually kept and sent, the real reply, instead of a rubric that only asks whether the writing is nice.
A habit stated as a feeling can't be checked. A habit stated as a comparison against a real answer can.
AAnchor. The one decision everything else hangs on.
A golden set: 400 real threads with the reply the person actually sent as the answer key, sliced into routine, sensitive, and judgment-call buckets, graded differently by tier.
This is the actual answer to the question. Every other letter exists to protect this one.
The reference answer sits in the same folder as the tier it belongs to, not off in a separate spreadsheet nobody opens.
RRisk. What breaks the first time the anchor is wrong?
A fluency-and-tone score passes a wrong fact every time, because a false claim reads exactly as smooth as a true one. That's what almost went out under Jarrah's name.
Naming the risk before it happens is what makes the anchor a design decision instead of a guess that got lucky twice.
The same sentence answers two different questions completely differently, and only one of the two questions was being asked.
KKeep out. What you deliberately won't build yet.
Grading whether a draft sounds like each SDR's own personal voice. That's a polish problem, and chasing it now spends review budget the fact-checking layer needs more.
Saying what you're not building is what makes this a decision, not a wish list dressed up as a spec.
Three things worth saying straight out, since this is where the real judgment sits. The alternative that got floated first was scoring every draft with a generic AI-quality prompt, an LLM asked to rate clarity, tone, and professionalism, with no source document attached. It lost because a score like that can only ever confirm the email is well written. It has no way to know whether "yes, we support SSO on the starter plan" is true, since truth was never part of what it was asked to grade. The AI-specific failure worth naming by name is a model stating a false capability or an unauthorized promise in the exact same confident, fluent voice it uses for a fact that happens to be correct, so nothing in the tone gives it away. The guardrail is the claims-and-commitments checklist, cross-checked against Northwind's live pricing and feature sheet, required on every sensitive and judgment tier draft before send, not just sampled after the fact. That guardrail isn't free. Full human review on those two tiers runs about six minutes a draft, and covering all forty-odd sensitive and judgment replies a day is close to four hours of someone's week, time that has to come from somewhere. It's worth spending there and nowhere else: the routine tier, roughly a hundred and ten replies a day, runs on an automated check instead, matching each draft against the nearest real historical reply for anything that looks like a changed fact, since a miss on a scheduling confirmation costs almost nothing. And the ship bar was never zero misses. Outrigger drafts probabilistically, and it will occasionally get a fact wrong the same way a tired human rep would. The bar is a ninety percent clean rate on the checklist held across two straight weekly audits on the tiers that matter, checked and re-checked, not asserted once and left alone.
And if you want to be sure it really works, try it somewhere else
Same five moves, a home-services company instead of a software sales team, nothing about trucking anywhere in sight.
Perigee Home Services answers customer quote requests by email after an HVAC inspection. Merilee Fielder runs client operations there, and owns the same question Cass does: is the drafted reply actually good, or does it just read well.
Situation: today, office staff read a technician's raw inspection notes, half-abbreviated, plus the customer's original question, and hand-type a quote follow-up. It takes about seven minutes a reply, and Perigee answers around 60 of these a day.
Payoff: the habit worth building isn't "customers feel reassured." It's grading a draft against what the office actually sent last time a similar job came through, sliced into routine appointment confirmations, sensitive replies like a warranty dispute or a price already quoted by phone, and judgment-call replies like a multi-system estimate with financing terms attached.
Anchor: a golden set of real past quote emails, each paired with the technician's notes and the reply the office actually sent, sliced the same three ways.
Risk: a draft that sounds warm and reassuring while quoting the wrong unit price, or promising a warranty term Perigee doesn't actually offer on that model.
Keep out: matching each office rep's own writing style, same call as Northwind, for the same reason.
Perigee: sensitive-tier claims-checklist error rate, six weeks after adopting the same design
Sensitive-tier error rate, checked against the claims checklist
Fourteen percent of sensitive-tier drafts had a wrong price or a promise Perigee doesn't offer, in week one of checking. By week six, weekly audits and the same tiered design had it down to three percent.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, a golden set with the real sent reply as the answer key, sliced by risk.
Cost: there's no budget this quarter for both full human review and a friendlier-sounding draft voice. The claims check wins every time, a warmer voice sitting on a broken fact is just decoration.
The model got better, for real: say Outrigger's overall accuracy improves this quarter. That's not proof the judgment tier improved with it. A model can get better on average while the one tier that costs the most stays exactly as blind as before.
Where people run it wrong.
They read one high aggregate "reads great" score as proof nothing's hiding underneath it.
They fix a caught mistake by adding one more example to a prompt, instead of building the tiered check that would have caught the next one too.
They wait for a customer complaint to reveal a wrong fact, instead of catching it in a weekly golden-set audit first.
How to use it live. Say the real tension out loud before answering: "is this asking me to build a writing-quality rubric, or to catch a wrong sentence hiding inside a polished one." That buys a beat to think instead of guessing out loud in front of the interviewer.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
SPARK: design against the failure before you build. Built for design questions, here applied to designing the eval itself, not the drafting feature.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Cass Ridler, sales operations lead at Northwind, a B2B software company that sells dispatch software to trucking companies. Owns proving Outrigger's drafts are actually good, not just polished.
3 · THE HABIT
What habit should the eval design build in Cass?
Tap to flip
ANSWER
Grade drafts against what the sender actually kept and sent, using the real sent reply as the answer key, instead of a rubric that only checks whether the writing sounds good.
4 · THE ANCHOR
What's the one design decision the eval hangs on?
Tap to flip
ANSWER
A golden set of 400 real historical threads, each paired with the reply that actually got sent, sliced into routine, sensitive, and judgment-call tiers, graded differently by tier.
5 · THE RISK
What breaks the first time the anchor is wrong?
Tap to flip
ANSWER
A fluency-and-tone score passes a draft that gets a fact or a promise wrong, because a false claim reads just as smoothly as a true one. That's what almost went out under Jarrah's name.
6 · THE NUMBER
Fill in the blank: on the sensitive tier, the old fluency rubric scored drafts at about 94 percent good. The claims-and-commitments checklist found the real pass rate was only ___ percent.
Tap to flip
ANSWER
84 percent. About one in six sensitive-tier drafts carried a wrong fact or an unauthorized promise the rubric had no way to catch.
7 · THE KEEP-OUT
What did version one deliberately not try to eval?
Tap to flip
ANSWER
Whether a draft matches each SDR's own personal writing voice. That's a polish problem, not a trust-breaking one, and building it in v1 would spend review budget the fact-checking layer needed more.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the same tension?
Tap to flip
ANSWER
Perigee Home Services, an HVAC company drafting quote follow-up emails. Same tension: a draft can read warm and reassuring while quoting the wrong price or promising a warranty Perigee doesn't offer.
Check yourself Score: 0 / 0
True or false
1. True or false: because Outrigger's weekly quality score stayed around 92 to 94 percent the whole time, the golden set audit found nothing wrong with its sensitive-tier drafts.
True
False
Show hint
Check the grouped bar chart in Section 1 for the sensitive and judgment tiers, not just the routine one.
Show answer
False. The claims-checked pass rate on the sensitive tier came in at 84 percent, well below the 94 percent the old fluency rubric reported. About one in six sensitive-tier drafts had a wrong fact or an unauthorized promise the rubric never saw.
Fill in the blank
2. The golden set holds ___ real historical threads, split into routine, sensitive, and judgment-call tiers.
Show hint
It's named directly in the Anchor step of the framework recap, and shown as the center label on the second hand-sketch.
Show answer
400. Each one paired with the reply the sender actually sent, which is what makes it a real answer key instead of an invented rubric.
Multiple choice
3. Why does the eval grade sensitive and judgment tier drafts against a claims-and-commitments checklist instead of a fluency score?
A. Because fluency scoring takes longer to run than a checklist.
B. Because a wrong fact or an invented promise reads exactly as smooth as a true one, so fluency scoring can't tell them apart.
C. Because the SDRs already edit every draft for tone before sending it.
D. Because Northwind's engineers built the checklist first and the rubric got added later.
Show hint
Think about what a fluency score is actually capable of measuring, and what it was never asked to look at.
Show answer
B. A confidently written sentence and a confidently written false sentence look identical to a rubric that only checks tone and clarity. Catching the false one needs a source of truth to check against, which is what the checklist provides.
Short answer, name the rejected alternative
4. What alternative did Cass consider for grading Outrigger's drafts, and why did it lose?
Show hint
Look at the paragraph right after the Keep-out step in the framework recap, the one naming what got floated first.
Show answer
Model answer: Scoring every draft with a generic AI-quality prompt, an LLM asked to rate clarity, tone, and professionalism, with no source document attached. It lost because a score like that can only confirm the email is well written. It has no way to know whether a stated fact or promise is actually true, since truth was never part of what it was asked to grade.
Short answer, apply it yourself
5. Pick an AI product you use yourself. Name one place its confident writing might be hiding a wrong fact, and how you'd check.
Show hint
Think of a product that writes in the same confident tone whether the underlying fact is simple or genuinely uncertain.
Show answer
Model answer: A grocery delivery app's auto-generated substitution note ("swapped for a similar item") sounds equally sure whether it swapped one brand of pasta for another or swapped a dairy item for a non-dairy one on an account flagged for a dairy allergy. I'd check by pulling a sample of substitutions on flagged-allergy accounts and grading them by hand against the account's own notes, instead of trusting the app's confident phrasing.
Short answer, reason about the threshold
6. This week's sensitive-tier audit comes back at 84 percent on the checklist. Last week's came back at 91 percent. Has Outrigger cleared the ship bar for the sensitive tier?
Show hint
Re-read the ship bar stated in stage 6 of the walkthrough and the closing paragraph of the framework recap. It isn't a single good week.
Show answer
No. The bar requires the sensitive tier to clear ninety percent across two straight weekly audits. This week's 84 percent breaks the streak, so the count resets, even though last week cleared the bar on its own.
Before you close the answer
Why this works
Tests whether you'll design an eval around real stakes or default to a generic quality rubric because it's easier to build. Most candidates stop at "we'd have an LLM judge rate the drafts one to five."
Follow-up traps
"Why not just have a human read every single draft before it sends?" Response: volume. A hundred fifty replies a day team-wide makes full human review impossible, which is exactly why the routine tier gets an automated check and only the two risky tiers get a person.
"Isn't a 400-thread golden set going to go stale as pricing changes?" Response: yes, which is why it gets refreshed weekly with fresh sent replies instead of built once and left alone. The audit cadence is part of the design, not an afterthought bolted on later.
If pressed
The automated routine-tier check works by matching a draft's stated facts against the nearest historical sent reply on the same topic, not against a static rulebook, so it catches a new kind of wrong claim without anyone having to write a fresh rule every time pricing changes.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.