CaseAdvancedEval-Driven Specification / Writing an eval spec / #16

How do you handle evals for subjective qualities like tone or brand voice?

The direct answer
Score tone with paired good and bad example lines pulled from real situations, not a one to ten "sounds like us" rating. Mark the exact words in each pair that put a reply on the brand's side or off it. A rating scale just moves the guesswork from the model to whoever is grading it.
Do this, in order
  1. Build the tone check from paired good and bad example lines, with the specific words flagged, instead of a "sounds like us" score.Why: a score is still a guess dressed up as a number. A marked pair is something two people can actually agree on.
  2. Base every pair on a real complaint type that actually comes up, not an abstract tone rule.Why: a rule like "be warm" doesn't tell anyone which words to change. A real pair does.
  3. Keep a separate pair type for genuinely serious complaints.Why: the same casual phrasing that works for a noisy air conditioner reads as flippant on a safety complaint.
  4. Check that a passing reply names something real from the actual review, not just that it sounds friendly.Why: a generic line can sound warm and still never mention the guest, the room, or the actual problem.
  5. Score tone drift on the lowest star drafts first, and leave every other stylistic dimension for later.Why: that's where a generic sounding reply costs the most, right where a guest is already unhappy.
  6. Re-run the check against the same batch after any prompt change, and read every flagged case.Why: a check nobody rereads is just a rule that used to matter.

How to answer this, stage by stage

Eight moves. Ground it in one drafted reply before touching the check itself.

1
Scope it to one flagged reply, not brand voice in general
Say it like this
"Say I'm the PM for Porchlight, the tool that drafts guest review replies for Harborcrest Hotels. I'm not going to talk about brand voice in the abstract. I'll build the check against the reply Porchlight drafted for a three star review about a noisy air conditioner in room 214, because that's the one a manager almost posted as is."
Why this works
A check built for one real reply is something an interviewer can picture. A check built for "tone" in general never gets specific enough to argue with.
2
Say your structure out loud
Say it like this
"Here's how I'll walk through it: what today's reply actually got wrong, the habit I want this check to protect, the check itself, what breaks if I set it too tight or too loose, and what I'm leaving out of it for now."
Why this works
Two seconds of structure tells the interviewer you have a plan, so they're following your answer instead of guessing where it's going.
3
Reframe what a tone eval actually has to catch
Say it like this
"Most people hear 'eval for tone' and go looking for a sentiment score, some one to ten rating for how warm a reply sounds. The real gap is that a reply can get every fact right and still read like it came from a hotel with no face behind it. A score can't tell you which words did that. A marked example can."
Why this works
This is the actual insight being tested. Skip it and you've built a rating scale, not a check.
4
Give the anchor: paired examples with the words flagged
Say it like this
"So the check is a stack of paired lines. For the AC complaint, the line Porchlight drafted said, 'we sincerely apologize for any inconvenience experienced during your stay,' and I'd mark 'sincerely apologize,' 'inconvenience experienced,' and 'your stay' as the words that flag it. The line that passes says, 'that AC in two fourteen, we heard you, we're swapping it before you're back in June,' and I'd mark the room number, the plan, and the date as what makes it pass."
Why this works
Naming the actual flagged words is something an interviewer can picture as a real rubric line, not a promise to "write more personally."
5
Say the payoff out loud
Say it like this
"The point of this isn't a nicer sounding reply for its own sake. It's that a manager can skim a passing draft and trust it actually sounds like Harborcrest, and we catch a generic drift in the regular check, before a guest's own comment says it for us."
Why this works
Naming the habit you're protecting, not just the symptom you're fixing, is what keeps a check from turning into busywork.
6
Name the risk, in both directions
Say it like this
"Set the check too tight, and a real safety complaint gets a casual, swap it out reply, because a banned word list can't tell a noisy AC apart from bedbugs. Set it too loose, and 'we're so grateful you shared this with us' passes because it sounds warm, even though it never mentions the room, the guest, or the actual problem."
Why this works
Naming both failure directions shows the check is a dial you can set wrong two different ways, not just one.
7
Prove it with the near miss
Say it like this
"Here's what almost happened. A guest on her third stay wrote that the neighboring unit's AC kept her up all night. Porchlight's draft never mentioned that she'd stayed three times, or which room. Riverside's general manager had her cursor on post before she reread it and realized it sounded like it came from a chain she'd never worked for."
Why this works
A specific near miss, with a real detail like the third stay, does more work than "this could go wrong" ever will.
8
Close on the one line
Say it like this
"So: build the check from paired good and bad lines with the specific words marked, keep a separate pair type for serious complaints, and check that a passing reply names something real. The number I can point to: before the check, twenty two of a hundred and thirty low star drafts got rewritten from scratch. After it, three."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.

Let's learn

Here is what happens when a reply gets every fact right, and still doesn't sound like anyone who worked there wrote it.

Say we build Porchlight, a tool inside Harborcrest Hotels, a small chain of eighteen boutique properties, that reads a guest's review and drafts a reply in the hotel's own voice. A front desk manager reads the draft, tweaks it, and posts it instead of writing one from nothing.

A hand-sketch of a guest experience manager at a desk with a small stack of printed reply drafts, marking one with a pencil checkmark and another with a pencil question mark, a coffee mug and a small plant nearby, no rubric card in sight
Checking a drafted reply, before this check existed
Knowledge spark: what's a tone check like this? A stack of real example lines, sorted into two piles: ones that sound like the brand, and ones that don't. Each line has the exact words marked that put it in one pile or the other. Not a score. A set of comparisons anyone can check a new draft against.

Before Porchlight, a manager wrote every reply by hand. A busy weekend brought in about fifteen reviews at one property. Writing a real, personal reply to each one took about ten minutes. That's two and a half hours, on top of running a front desk.

With Porchlight, a draft appears the moment a review comes in. Reading it, changing a word or two, and hitting post took about ninety seconds. Fifteen reviews went from two and a half hours down to about twenty minutes.

Here is the important part. The minutes saved were never the story. The real question was whether the draft still sounded like Harborcrest, and getting the facts right turned out to be a completely different thing from getting that right. A manager reading fast, deciding whether to tweak a word or hit post, stops reading the way she used to read when she wrote every line herself. She's proofing now, not composing. And a generic sentence can pass a fast proofing check that would never have survived her writing it from scratch.

It wasn't wrong. It just didn't sound like anyone who worked there wrote it.

At its worst, that cost more than a clunky reply. One quarter, twenty two of a hundred and thirty low star drafts, about one in six, got rewritten entirely by a manager because they read as boilerplate. And Harborcrest's whole pitch to a guest is that it isn't a big anonymous chain. A reply that sounds like one, sitting right under a review on the same page as that pitch, argues against the hotel's own brand.

Low star drafts rewritten from scratch, before and after the check
25 0 22 before the check of 130 drafts 3 after the check of 130 drafts
fully rewritten, no check fully rewritten, with the check
Same chain, same rough volume of low star reviews each quarter, about 130. The check didn't change how many reviews came in. It changed how many replies needed a manager to start over from a blank line.

At its worst, this let a specific kind of mistake through with nobody watching for it. A guest on her third stay at the Riverside property wrote that the unit next to hers had rattled all night, again. Porchlight drafted a reply that promised a fix and never once said which room, or that this was her third visit.

The decision that mattered Score tone with paired good and bad example lines for the situations that actually come up, each one marked with the specific words that put it on the brand's side or off it.

The choice I would take back. When Porchlight first shipped, the instructions we gave it said to reply politely and address the guest's issue, on the idea that a well written model would sound personable without anyone telling it how. That held up fine for the first few weeks, when volume was low enough that someone read every draft before it posted. It stopped holding up once the chain leaned on Porchlight across all eighteen properties, and nobody had ever built a specific check for whether a reply sounded like Harborcrest instead of just any hotel.

What I would leave alone. A five star review with a cheerful "thanks so much, glad you enjoyed your stay" reply costs nothing, even if it's a little generic. Almost nobody rereads a happy reply. Spending check effort there takes it away from the drafts that actually need it.

The lesson. Sounding right isn't a bonus feature sitting on top of being correct. It's a separate thing, and it needs its own specific check, built from real examples, not a hope that good instructions will make it happen on their own.

Now here is the same thing as a story

Read this one when you've got a few minutes. The short version is above. This is for when you want to feel why it mattered.

Innes Whitlock could read a guest review and tell you, before finishing the first sentence, whether it was a real complaint or someone just having a bad Tuesday. She'd run Harborcrest's guest experience team for four years, long before anyone built Porchlight, back when eighteen properties meant eighteen managers writing every reply themselves, some good at it, some not.

Porchlight arrived in the spring. For months, it was the best thing that had happened to those front desks. A review came in, a draft appeared, a manager read it, changed a word, and posted it. What used to eat an hour of a Saturday morning now took fifteen minutes. Innes read a sample of the drafts every week at first, and they sounded like Harborcrest: specific, a little wry, never like a form.

She stopped reading the sample every week. Then every month. The ones she checked always looked fine.

Then came the third stay review.

A guest wrote that this was her third time at the Riverside property, and that the unit next to hers had rattled all night, again. Porchlight drafted a reply: "We sincerely apologize for any inconvenience experienced during your stay and have noted your feedback for review by our maintenance team."

Nothing in that line was wrong. It addressed the noise. It promised a fix. Riverside's general manager had her cursor on the post button when something made her stop and reread it. It never said room two fourteen. It never said third stay. It could have gone out under any hotel's name in the country.

Close hand-sketch of a two column card, one side labeled sounds like anyone with the corporate apology line and its flagged words, the other side labeled sounds like us with the specific reply and its flagged words
The anchor: one pair, the words that flip it

She deleted the draft and wrote her own line instead. But she also sent Innes a message: "this could've gone out. is that what we want going out?"

We didn't lose a guest over the AC. We almost lost her over the reply.

I want to say the problem was that Porchlight got the facts wrong. It didn't. Every fact in that draft was true. But Innes had never given anyone, or anything, a specific way to check for the other thing, the thing that actually made a Harborcrest reply sound like Harborcrest. She had a feeling for it. Porchlight didn't.

So here is the decision she made differently.

When the team first wrote Porchlight's instructions, they told it to reply politely and address the issue. That was the whole guidance on tone. It made sense then, when volume was low and every draft got a human read before it went anywhere. Nobody had imagined leaning on it for every property inside a year.

Innes built a check instead. Not a score. A stack of real pairs: a bad line, and next to it, the line a real manager would actually send, with the words marked that made the difference. For the AC complaint, marked words like "sincerely apologize" and "your stay" put a line on the wrong side. A room number, a plan, and a date put it on the right side.

Two panels: left labeled too rigid, showing banned words crossing out a straight apology a real safety complaint needs, with a note that the anchor's own serious complaint pair still lets it through; right labeled too loose, showing a generic grateful sounding line passing the check without naming anything real
Checked against its own risk, both directions

Run the same quarter forward with the check in place. Low star drafts get checked against the pairs before anyone reads them by feel. The following quarter, out of a hundred and thirty low star drafts, only three needed a full rewrite. The rest passed, or got flagged and fixed in under a minute, because the check named exactly which words were the problem instead of asking someone to feel it out.

One design hands a manager a feeling. The other hands her a reason.

And the thing I'd want to tell myself, back when we wrote Porchlight's first instructions: I asked it to be polite. I never told it what polite was supposed to sound like, coming from us.

SPARK, sized to one guest reply

This question asks how to handle evals for something subjective, not literally design a feature, but it's still one concrete decision about the exact moment a reply sounds like Harborcrest or doesn't, so SPARK still fits. A question asking how to measure Porchlight's overall reply quality across the whole chain would reach for LEAD instead.

S, situation. Before a check like this exists, whoever reviews Porchlight's drafts reads them by feel and either catches a generic one or doesn't, depending on how closely they happen to be paying attention that week.
P, payoff. Not "nicer sounding replies." The habit worth building: a manager can skim a passing draft and trust it actually sounds like Harborcrest, and the team catches drift in a regular check, before a guest's own comment says it first.
A, anchor. Score tone with paired good and bad example lines for situations that actually come up, each one marked with the specific words that put it on the brand's side or off it. Keep a separate pair type for genuinely serious complaints.
R, risk. Set the check too tight, and a real safety complaint gets a casual reply because a banned word list can't tell a noisy AC from bedbugs. Set it too loose, and a generic but warm sounding line passes without ever naming the room, the guest, or the actual problem.
K, keep out. No attempt yet to score every stylistic dimension, humor, emoji use, formality by star rating, on every kind of review. Low star and neutral drafts first, since that's where a generic reply costs the most.
Why the anchor survives the risk Check it against the near miss. Does the check still catch a warm sounding line that names nothing real? Yes, because a pass requires a real detail, not just friendly words. Does it still let a genuinely serious complaint use straighter language? Yes, because serious complaints get their own pair type instead of one universal banned word list.

And if you want to be sure it really works, try it somewhere else

A food rescue nonprofit is a different kind of operation entirely, but the same gap between technically fine and actually right shows up in a donor's thank you note.

S. Ottoline Prescott runs operations for Second Table, a food rescue nonprofit that picks up surplus food from grocers and restaurants and gets it to community pantries the same day. Today, without a check, a coordinator drafts a thank you note from a template and either personalizes it by hand or lets it go out generic, depending on how busy the week is.
P. The habit worth building: a donor reads a thank you note and believes it, because it names something real that happened with her actual gift, not because the note used the word grateful enough times.
A. Same shape, a different pair. Bad: "thank you for your generous contribution, which helps further our mission." Good: "because of your gift this week, three hundred forty pounds of bread that would have hit a dumpster fed about two hundred families at the Elm Street pantry." The marked words: a real number, a real place, a real result.
R. Tighten the check to require a precise pound count on every note, and a ten dollar monthly donor gets a note that sounds like the org is bragging about her ten dollars, which reads as showy instead of grateful. Loosen it to "sounds warm," and "we're so grateful for your support" passes again, saying nothing real about what her gift actually did.
K. No attempt yet to personalize every note by donor history, holiday, or gift size. One check, does the note name something real that happened, for every note tied to an actual pickup, is enough to start.

A hand-sketch two column card for Second Table: left labeled generic gratitude with a form-letter thank you line, right labeled names something real with a line naming a specific pound count, place, and result
Same anchor, a different desk, a donor note instead of a review reply

It took a longtime ten dollar monthly donor writing in to ask why her note thanked her for three hundred forty pounds rescued that week, when her own gift couldn't have covered a tenth of that, before anyone noticed the check needed a second guard rail: match the specific claim to the size of the actual gift, not just to whether a pickup happened that week.

Swap the trigger and it still runs

  • Speed: even if Porchlight drafted every reply instantly, which it already does, a fast wrong sounding reply is still wrong. Speed never fixes what the check is testing.
  • Cost: if running the check cost nothing and needed no manager time at all, that still wouldn't tell you which pair type applies. A check still needs to know a safety complaint isn't an AC complaint.
  • The model gets better: if Porchlight's facts became near perfect, never missing a room number or a detail, it could still default to safe, generic phrasing, because that's often the safest sounding text. Better facts don't fix a tone problem. They're different things.

Where people run it wrong

  • Writing one banned word list and calling it done, so it blocks straight talk in the one case, a real safety complaint, where straight talk was the right call.
  • Writing the good example without marking which words make it good, so two managers argue about the same line for months.
  • Trying to score every stylistic dimension, humor, emoji use, formality by star rating, in the first version, so the check ships months late while generic drafts keep going out the whole time.

How to use it live

If you're asked this cold, ask what a technically correct but wrong reply would actually sound like for this product. Write down the specific words that make it wrong. Then ask whether today's check could catch it. That question finds the missing rubric faster than listing every tone rule from a blank page.

Flashcards (click a card to flip it)

1 · THE SITUATION
What's the situation, before this check exists?
Tap to flip
ANSWER
A drafted reply gets every fact right and still reads like a generic hotel chain form letter, and nobody has a specific way to check for that. Whoever reviews it just reads by feel.
2 · THE PAYOFF
What's the real habit this check is trying to build?
Tap to flip
ANSWER
A manager can skim a passing draft and trust it actually sounds like Harborcrest, and the team catches drift in a regular check, before a guest's own comment says it first.
3 · THE ANCHOR
What does the tone check actually check?
Tap to flip
ANSWER
Paired good and bad example lines for real complaint types, each one marked with the specific words that tip it onto the brand's side or off it. Serious complaints get their own separate pair type.
4 · THE RISK
What breaks if the check goes too far either way?
Tap to flip
ANSWER
Too rigid, and a genuinely serious complaint gets a tone deaf, over casual reply because a banned word list can't tell it apart from a small complaint. Too loose, and a generic but warm sounding line passes without naming anything real.
5 · THE PROOF
What almost went out the door?
Tap to flip
ANSWER
Riverside's general manager nearly posted Porchlight's boilerplate reply to a third time guest's AC complaint, catching it seconds before hitting post because it never mentioned her stay or the room.
6 · THE NUMBER
___ of ___ low star drafts needed a full rewrite before the check; ___ after.
Tap to flip
ANSWER
22 of 130; 3. Same rough volume of low star reviews each quarter, about one in six needed a full rewrite before the check existed.
7 · THE REPLAY
Same quarter, check in place. What changes?
Tap to flip
ANSWER
Low star drafts get checked against the pairs before anyone reads them by feel. Only 3 of 130 need a full rewrite. The rest pass, or get flagged and fixed in under a minute.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor check?
Tap to flip
ANSWER
Second Table, a food rescue nonprofit's donor thank you notes. Its anchor keeps the same shape, a marked good and bad pair, but checks whether the note names a real weight, place, and result instead of a real room and date.

Check yourself Score: 0 / 0

True or false
1. True or false: the reply Porchlight drafted for the room 214 complaint contained a factual error.
  • True
  • False
Show hint
Look at what the story says was wrong with the draft. Was it a fact problem, or something else?
Show answer
False. It got every fact right: it addressed the noise and promised a fix. The problem was that it never mentioned the room number or that the guest was on her third stay, so it sounded like it could have come from any hotel chain.
Multiple choice
2. Which design matches the anchor this answer argues for?
  • A. A single one to ten score for how much a draft "sounds like us."
  • B. Paired good and bad example lines for real complaint types, with the specific words marked, and a separate pair type for serious complaints.
  • C. A banned word list applied the same way to every kind of review, with no exceptions.
  • D. A generic tone rule like "always sound warm and friendly," left for each manager to apply as they see fit.
Show hint
The anchor needs marked examples and a way to handle a genuinely serious complaint differently.
Show answer
B. A is the vague scale this answer argues against. C is the too rigid version that would block a straight apology a safety complaint actually needs. D never gets specific enough for anyone to check against.
Fill in the blank
3. Before the check, ___ of the ___ low star drafts each quarter got rewritten entirely because they read as boilerplate. After the check, that number dropped to ___.
Show hint
This is the number the whole argument leans on. It shows up twice: once in the story, once in the chart.
Show answer
22; 130; 3. About one in six low star drafts needed a full rewrite before the check existed. After it, the check caught the same kind of generic line before it ever reached a manager's screen.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about why "reply politely and address the issue" sounded like enough guidance back when Porchlight only served a handful of drafts a week.
Show answer
Model answer: The team's original instructions to Porchlight said to reply politely and address the issue, with no specific tone check. That made sense when volume was low and someone read every draft before it posted. It stopped making sense once the chain leaned on Porchlight across all eighteen properties, and nobody had ever built a check for whether a reply sounded like Harborcrest instead of just any hotel.
Short answer, apply it yourself
5. Pick something you or your team writes at volume, emails, support replies, social captions. If it were being drafted for you, would you actually catch it drifting generic before someone else pointed it out? What specific check would catch it?
Show hint
Look for whether your own read of it checks for something specific, or just whether it feels roughly right.
Show answer
Model answer: "I approve social captions for my shop before they post, and I mostly just skim for typos. A real check would be: does the caption name the actual item, color, or price, not just 'new arrival, check it out.' If it can't name one real detail, it fails, no matter how catchy it sounds."
Multiple choice
6. Based on this answer's own numbers, if the fully rewritten rate had stayed at about one in six instead of dropping after the check, roughly how many of the next quarter's 130 low star drafts would managers have needed to rewrite from scratch?
  • A. About 22.
  • B. About 3.
  • C. All 130.
  • D. About 65, roughly half.
Show hint
One in six of 130 is close to the same number the story already gives you for the quarter before the check existed.
Show answer
A. One in six of 130 is about 22, which matches the actual pre-check quarter. The check is what brought that down to 3, not a change in how many low star reviews came in.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more