Artifact critiqueAdvancedEval-Driven Specification / Writing an eval spec / #4

Write a scoring rubric for the quality of a generated customer support reply.

The direct answer
Score each reply on five named things, not one: whether it got the policy right, whether it actually moved the problem forward, whether it answered everything asked, its tone, and whether it should have gone to a person instead. Write what a low score and a high score look like on each one, and let policy accuracy fail a reply outright no matter how good the rest reads. A single 1 to 5 quality number cannot do that job. It can only tell you whether the reply sounded good.
Do this, in order
  1. Score five named things, not one flat quality number.Why: a flat score can average a dangerous mistake away behind two good ones, and that's exactly what let a wrong promise through.
  2. Make policy accuracy a gate, not just one more number in the average.Why: without a gate, a warm, confident, wrong promise still scores high enough to pass.
  3. Write a real low example and a real high example for every dimension, not just a label.Why: a labeled scale with nothing anchoring it drifts back into gut feel inside a month.
  4. Keep the dimension count to five, not fifteen.Why: too many sub-checks slow review down until only a sliver of replies ever get read, and reviewers start grading the checklist instead of the reply.
  5. Leave exact brand-voice phrasing out of the tone score for now.Why: chasing every regional wording variant on day one delays shipping the rubric that would have actually caught the missed promise.
  6. Calibrate every reviewer against the same sample replies before anyone scores live traffic.Why: without a shared reference, five people reading the same five labels still land in five different places.

How to answer this, stage by stage

Six moves. Ground it in one reply type that's already auto-sending, before touching the rubric itself.

1
Scope the rubric to one real reply, not "support quality" in general
Say it like this
"Say I'm the PM for Wingman, the reply drafter inside Silvertail Airlines' live chat. I'm not going to score everything Wingman writes the same way. I'll build this rubric against baggage-delay replies first, since that's the one category we're already letting go out with no human glance at all."
Why this works
A rubric built for one real category is something an interviewer can picture. A rubric built for "quality" in general never gets specific enough to argue with.
2
Say your structure out loud
Say it like this
"Here's how I'll walk through it: what scoring looked like before any of this, what a score actually has to settle, the five things I'd grade instead of one, what breaks if the rubric's too loose or too strict, and what I'd leave out of it for now."
Why this works
Two seconds of structure tells the interviewer you have a plan, so they're following your answer instead of guessing where it's headed.
3
Reframe what the score actually has to settle
Say it like this
"Most people hear 'write a quality rubric' and go looking for the right five-point scale. The real gap is that a single number can't tell you whether the reply lied. It can only tell you whether it read nicely, and a reply can do both at once."
Why this works
This is the actual insight being tested. Skip it and you've written a scale, not a rubric.
4
Give the anchor, all five, and name the gate
Say it like this
"So the rubric scores five things: policy accuracy, whether the reply actually moves the problem forward, whether it answers everything the customer asked, tone, and whether it should've gone to a person instead. Policy accuracy is different from the other four. Score it a one or a two, and the whole reply fails, no matter what the other four say."
Why this works
Naming all five and marking one as a gate is something an interviewer can picture as an actual spec line, not a promise to "be more careful."
5
Prove it with the reply that shouldn't have shipped
Say it like this
"Here's what happens without the gate. A reply offers a customer two hundred dollars for a six-hour bag delay, when the real rule needs twenty-four hours. It reads warm and confident, so it scores a five. It goes out automatically for eleven days, thirty-four times, before a customer's own bank dispute catches what the rubric never did."
Why this works
A specific near miss, with a real count, does more work than "this could go wrong" ever will.
6
Close on the one line
Say it like this
"So: five named things scored instead of one, policy accuracy able to fail a reply by itself, and brand-voice nuance left for later on purpose. The number I can point to is how many wrong promises actually reach a customer. Under the old score, thirty-four. Under the gate, zero, because the same thirty-four get pulled out before they ever send."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.

Let's learn

Here is what happens when a support reply gets judged on one number, and nobody ever writes down what that number is actually supposed to check.

Say we build Wingman, a tool that drafts the first reply inside Silvertail Airlines' live chat, so a support agent can read it, send it as is, fix it, or throw it out and type their own.

A hand-sketch of a desk with two reviewers holding printed chat transcripts, one scrawling a big 5 on a sticky note, the other a 3, with thought bubbles reading sounds warm and reads fine to me
What scoring a reply looked like, before this rubric existed

Before anyone scored anything on paper, every Friday four reviewers split a sample of two hundred replies from that week and gave each one a number, one to five, for quality. Nothing more specific than that. Two straight weeks, baggage-delay replies averaged 4.3. That crossed the line the team had already agreed on: hold 4.2 or higher for two weeks running, and that kind of reply stops going through an agent at all. Wingman's draft goes straight to the customer.

Knowledge spark: what is auto-send? A setting that lets Wingman's draft go straight to the customer with no agent glance first. Most replies still get read by a person before they're sent. Auto-send turns that off for one kind of reply, once its quality score has stayed high for long enough.

On a normal week that saved real time. Baggage-delay chats ran about ninety a day, and skipping the agent glance freed up minutes for the calls only a person could take, the kind that spike hardest during a storm delay.

Then one reply inside that well-scoring batch offered a customer two hundred dollars for a bag delayed six hours. Silvertail only owes that two hundred dollars once a bag has been missing twenty-four hours. Under six hours, an agent is supposed to say the bag is being traced and offer to cover receipts if the delay passes a day. The reviewer who read it gave it a five. It sounded warm. It sounded confident. It sounded exactly like the tone Silvertail wanted.

Here's the part that's easy to miss. A dollar figure that wasn't owed was never really the problem, because that was one reply. The real problem was that "quality" had never meant anything sharper than does this sound like something we'd send. A reply that broke policy in its second sentence could still be the best-reading reply in the whole sample.

The reply that shouldn't have shipped, scored two ways
5 0 5 old score 1 policy acc. 4 resolution 4 complete. 5 tone 3 escalation
the old rubric: one flat number policy accuracy, the gate the other four dimensions
The old rubric gave this exact reply a flat 5. Scored on five dimensions instead, four of them are genuinely good, resolution, completeness, and tone all read fine. Policy accuracy is the one that fails, a 1, because the reply promised something the real policy didn't owe. A flat number averages that 1 away behind three good scores. A gate can't.
We didn't score whether the promise was true. We scored whether it sounded true.

At its worst, this costs real money and a real fight. The wrong reply kept going out for eleven days after auto-send switched on, thirty-four times out of the nine hundred and ninety baggage-delay replies sent in that window, every one to a customer whose bag hadn't actually been gone a full day. Silvertail chose to honor every promise it had already made rather than call thirty-four customers back to take money away, an unplanned $6,800, on top of a specialist's whole afternoon spent walking one customer's bank through why a payment she was owed had never landed.

The decision that mattered Score five named things, not one: policy accuracy, resolution progress, completeness, tone, and escalation safety, each with what a low score and a high score actually look like. Make policy accuracy a gate. Score it low, and the reply fails outright, no matter what the other four say.

The choice I would take back. When this scoring was first built, reviewers pushed back on anything more than one number, said a five-part scorecard would triple how long the Friday sample took, and back then Wingman only touched a slice of chats, so a slower review felt like the safer trade. That held up fine while a flat five just meant "read this one again if you're not sure." It stopped holding up the day a flat five became the only thing standing between a bad promise and a live customer.

What I would leave alone. A reply to something like "what's my seat number" doesn't need this. There's one right answer, it's already sitting in the reservation system, and the reply either matches it or it doesn't. No policy judgment call lives in there for a rubric to catch. Only score the kind of reply where a person had to decide what to promise.

The lesson. Splitting a score into five parts doesn't fix reviewers disagreeing with each other. It fixes what any one reviewer's number was ever measuring. If a wrong promise and a right one can land on the same score, we never wrote a rubric. We ran a taste test.

Now here is the same thing as a story

Read this one when you've got a few minutes, because eleven quiet days say more than any rubric ever could.

Saanvi Rajguru owns Wingman's quality bar at Silvertail Airlines, inherited from the contractor team that had built the scoring process the year before she took the role.

She trained all four reviewers herself, in one shared room, reading transcripts together and arguing out loud about what a five actually meant next to a three. For months she sat in on every Friday session.

Then storm season hit and her calendar filled with two other launches. She stopped sitting in on Fridays. Then she stopped reading the transcripts behind any score above a three, trusting the number on its own. Reviewers, scoring faster as volume tripled through October, stopped writing why behind a score and just typed the digit.

On the Tuesday that mattered, baggage-delay hit its second straight week at 4.3. Auto-send flipped on for that intent automatically, the way the team had agreed it would, months earlier, back when 4.2 sounded like a number worth trusting.

Eleven days later, a complaints-line email landed with a screenshot attached: a customer's own bank dispute form. Her bag had been six hours late. Wingman had told her two hundred dollars was coming. It never did, because the offer had never been real to begin with. She hadn't asked for a refund. She'd asked her bank why a payment Silvertail promised her, in writing, hadn't shown up.

Saanvi pulled the Friday scores for that whole window. Every one of them read 4.1 or higher. Nothing on the dashboard had ever looked wrong.

She wasn't grading whether Silvertail could keep the promise. She was grading whether the sentence sounded like one worth keeping.

She found the reviewer who'd scored the original reply a five. It hadn't been carelessness. Read on its own, it was a good reply: warm, clear, fast. Nothing in the one number he'd been asked to give had ever asked him to check the twenty-four-hour rule against the timestamp sitting right above it in the same chat.

Saanvi pulled all eleven days of transcripts. Thirty-four of the nine hundred and ninety baggage-delay replies sent automatically in that window had made the same offer to a customer whose bag hadn't been gone a full day. Nine hundred fifty-six had been fine.

Close hand-sketch of an index-card rubric with five rows, policy accuracy, resolution progress, completeness, tone, escalation safety, with a red gate mark next to policy accuracy
The anchor: one rubric card, five named rows, one of them a gate

Here's the rubric she wrote after. Five things, not one: policy accuracy, resolution progress, completeness, tone, and escalation safety, each with a written line for what a low score looks like and what a high one does. Policy accuracy stands apart from the other four. Score it a one or a two, on its own, and the whole reply fails, no matter how the rest reads.

Two panels: left, a wrong money promise going straight out to a customer unblocked; right, the same reply stopped by a drawn gate with a person glancing at it before it sends
The day it was wrong, with and without the gate in place

Run the same eleven days forward with the rubric in place. The reply reads exactly the same, warm, clear, fast. But policy accuracy checks the twenty-four-hour rule against the chat's own timestamp before anything sends, and this reply fails that one check by itself. It doesn't auto-send. It queues for the next agent, who glances at it, fixes one sentence, and sends the real answer in under a minute. The other nine hundred fifty-six replies in that window never touch a human at all, because they never trip the gate.

Thirty-four wrong promises become zero. Nine hundred fifty-six replies still move at the same speed they always did.

And the thing I'd tell myself, back when the reviewers first asked for one simple number: I answered how fast can we score them. I never asked what happens the day the fast number is wrong in a way that reads as right.

SPARK, sized for a rubric instead of a screen

This question asks for a scoring rubric, not an interface, but it's still one concrete design decision about the exact moment a reply either earns trust or breaks it, so SPARK still fits. A question asking how to measure Wingman's overall quality across the whole product would reach for LEAD instead.

S, situation. Any reply drafter, before a written rubric exists: whoever's sampling that week reads a reply and decides, alone, whether it sounds like something worth sending.
P, payoff. Not "fewer wrong promises." The habit worth building: any reviewer, any week, lands close to the same score on the same reply, one they could actually defend if someone official asked why auto-send flipped on.
A, anchor. Five named things, not one number: policy accuracy, resolution progress, completeness, tone, and escalation safety, each scored with a written low and a written high. Policy accuracy gates the other four: score it low, and the reply fails, whatever the rest says.
R, risk. Leave the five dimensions as loose labels with no written low and high, and reviewers drift right back to one shared gut feeling wearing five separate numbers. Make every dimension as detailed as policy accuracy, checklist-precise down to exact wording, and review slows so far that only a sliver of replies ever get read, while reviewers start grading whether a reply matches the checklist's words instead of whether it's actually right.
K, keep out. No attempt yet to score exact brand-voice phrasing for every market or channel Silvertail runs chat in. Tone only checks whether a reply reads calm and professional for the moment, not whether it uses a style guide's precise preferred wording.
Why the anchor survives the risk Check it against that Tuesday. Does the gate still catch a reply that reads warm and confident? Yes, because policy accuracy checks the promise against the real policy, not against how nice the sentence sounds. A rubric that only graded tone and clarity would have waved the same reply through a second time.

And if you want to be sure it really works, try it somewhere else

A home-insurance claims chat is a different business entirely, wearing a different kind of promise, but the same gap between "it read well" and "it was true" shows up there too.

A hand-sketch of a desk with a claims-chat dashboard, a similar rubric card, and a small water-damage house icon, for a home-insurance claims team
Same framework, a different desk, a different kind of promise

S. Tamsin Halloran leads claims intake at Cresthaven Mutual, where a chat tool drafts the first reply to anyone filing a water-damage or fire claim. Today, without a written rubric, a reply gets judged on whether it sounds reassuring, with nobody checking whether the settlement figure it quotes matches the actual policy.
P. The habit worth building: whoever reviews a claims-chat log can tell in one read whether the number it quoted was ever real, not just polite.
A. Same shape, different gate. Coverage accuracy, checked against the real policy's sublimit table for that kind of damage, fails the reply outright if the quoted figure doesn't match, the same way policy accuracy did for Wingman.
R. Loosen the gate to a vibe check, does this number feel about right, and it never actually fires, the same failure as before, wearing new clothes. Tighten it to require a full manual policy lookup on every single reply, even a plain "we've received your claim" with no figure in it at all, and response time collapses during the one week claims volume actually spikes, a real storm.
K. No attempt yet to score empathetic phrasing differently for a fire claim versus a flood claim. Just whether the number quoted was ever true.

It took a homeowner calling to ask where her fifteen-thousand-dollar repair estimate had gone, when her policy's water-damage sublimit was five thousand, before anyone checked whether the reviewers' warm scores had ever looked at a number at all.

Swap the trigger and it still runs

  • Speed: even if reviewers scored every reply the second it drafted instead of once a week, a fast wrong score is still wrong. Speed doesn't fix what the number was ever checking.
  • Cost: if the Friday sample got free tomorrow, run by a second model instead of four people, that still wouldn't decide what "quality" means. A free judge giving one flat number just misses the same promise faster.
  • The model gets better: if Wingman's accuracy climbed to 99 in 100, the one reply left over would still need a gate to catch it. A rare miss dressed in confident language is exactly the kind a flat score was already missing.

Where people run it wrong

  • Writing five named dimensions, then averaging them into one number anyway, so a bad policy score still gets buried under four good ones.
  • Naming the dimensions but never writing what a low score and a high score actually look like on each, so reviewers drift back to one shared gut feeling within a month.
  • Trying to score brand voice, translation, and regional tone in the very first version, so the rubric ships months late while ungated replies keep going out the whole time.

How to use it live

If you're asked this cold, ask what the worst possible reply would look like for this product, the one that would actually hurt someone. Then ask whether today's scoring could ever give that reply a low score. That question finds the missing gate faster than listing five dimensions from a blank page.

Flashcards (click a card to flip it)

1 · THE SITUATION
Before this rubric existed, how did a Wingman reply get judged as good or bad?
Tap to flip
ANSWER
Whoever was on that week's Friday sample read it and gave it one number, one to five, based on whether it sounded like something Silvertail would send. Nobody checked it against the actual policy.
2 · THE PAYOFF
What's the real habit a five-part rubric is trying to build?
Tap to flip
ANSWER
Any reviewer, any week, lands close to the same score on the same reply, one they could defend if someone official asked why auto-send switched on for that kind of reply.
3 · THE ANCHOR
Name the five things this rubric scores, and which one can fail a reply by itself.
Tap to flip
ANSWER
Policy accuracy, resolution progress, completeness, tone, and escalation safety. Policy accuracy gates the reply: score it a one or a two, and the whole thing fails, whatever the other four say.
4 · THE RISK
What went wrong the one time this rubric didn't exist yet?
Tap to flip
ANSWER
A reply offered a customer $200 for a six-hour bag delay, when Silvertail's rule needs 24 hours. It read warm and confident, scored a 5, and auto-sent to real customers for 11 days before anyone caught it.
5 · THE PROOF
How did Saanvi actually find out the wrong promise had gone out?
Tap to flip
ANSWER
Not from her own dashboard, which showed a steady 4.1-plus the whole time. A customer emailed Silvertail's complaints line with a screenshot of her own bank dispute, asking where the promised $200 had gone.
6 · THE NUMBER
___ of the ___ baggage-delay replies auto-sent over those 11 days repeated the same wrong $200 promise.
Tap to flip
ANSWER
34 of 990. Baggage-delay chats ran about 90 a day for 11 days; 34 of those went to customers whose bags hadn't actually been gone a full 24 hours.
7 · THE REPLAY
Same 11 days, rubric in place. What changes?
Tap to flip
ANSWER
The risky reply still reads warm and confident, but policy accuracy checks the 24-hour rule against the chat's own timestamp and fails it. It queues for a human instead of auto-sending. The other 956 replies never touch a person, because they never trip the gate. 34 wrong promises become zero.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor add or change?
Tap to flip
ANSWER
Cresthaven Mutual's home-insurance claims chat. Its anchor keeps the same shape, a gating dimension, but the gate checks a quoted settlement figure against the policy's real sublimit table instead of a delay-based promise.

Check yourself Score: 0 / 0

Fill in the blank
1. The rubric that would have caught the wrong promise scores each reply on ___ named things, and only ___ can fail a reply outright, whatever the other four say.
Show hint
Count the labeled rows on the rubric card, and look for the one marked with a gate.
Show answer
Five; policy accuracy. The other four (resolution progress, completeness, tone, escalation safety) get scored the usual way, but a low policy-accuracy score fails the reply on its own, no averaging.
Multiple choice
2. Which rubric design matches the anchor this answer argues for?
  • A. One 1-to-5 quality score, no dimensions.
  • B. Five named dimensions, averaged together, with no dimension able to fail the reply alone.
  • C. Five named dimensions, each with a written low and high, and policy accuracy able to fail the reply by itself.
  • D. Twelve dimensions covering every possible wording choice, scored on every reply.
Show hint
The anchor needs both named dimensions and a gate that can override the average.
Show answer
C. A has no dimensions at all. B has dimensions but no gate, so a bad promise still gets averaged away. D is the too-rigid version this answer warns against.
Short answer
3. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about why a five-part scorecard sounded like a bad trade back when Wingman only touched a slice of chats.
Show answer
Model answer: Early on, reviewers pushed back on scoring more than one number, since a five-part scorecard would triple how long the Friday sample took, and Wingman only touched a slice of chats at the time, so a slower review felt like the safer trade. That held up while a flat five just meant "look again if unsure." It stopped holding up the day a flat five became the only thing standing between a bad promise and a live customer.
True or false
4. True or false: this rubric should also score whether a reply matches the exact preferred wording for every regional market and channel, starting with day one.
  • True
  • False
Show hint
Think about what the keep-out step in this answer names on purpose.
Show answer
False. Matching exact brand-voice phrasing per market and channel is the keep-out. The tone dimension only checks whether a reply reads calm and professional, not whether it uses a style guide's precise words. Chasing every wording variant on day one would have delayed the rubric that actually caught the missed promise.
Short answer, apply it yourself
5. Pick something you track yourself, at work or anywhere. If it looked good on the surface, would you actually catch it if the substance underneath was wrong? What would need to change about how you check it?
Show hint
Look for whether your own check ever verifies the underlying fact, or only whether the output looks fine.
Show answer
Model answer: "I track how many support tickets get marked 'resolved' each week. Right now a ticket counts as resolved once an agent clicks the button, but nobody checks whether the customer's actual problem got fixed. A real check would need to look at whether the customer writes back again within a week, not just whether the button got clicked."
Multiple choice
6. Based on this answer's own numbers, if the policy-accuracy gate had fired on every single baggage-delay reply instead of only the ones with a real mistake, about how many of the 990 replies sent over those 11 days would have needed a human glance?
  • A. 990, all of them.
  • B. 34, only the ones that actually broke the 24-hour rule.
  • C. 0, the gate never actually fires.
  • D. About 500, roughly half.
Show hint
The gate checks each reply against the real 24-hour rule. It doesn't slow down every reply, only the ones that actually fail that check.
Show answer
B. Only 34 of the 990 replies actually broke the 24-hour rule. The other 956 pass policy accuracy cleanly and keep auto-sending at the same speed, which is what makes a targeted gate cheap instead of a blanket slowdown.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more