How do you handle evals for subjective qualities like tone or brand voice?
- Build the tone check from paired good and bad example lines, with the specific words flagged, instead of a "sounds like us" score.Why: a score is still a guess dressed up as a number. A marked pair is something two people can actually agree on.
- Base every pair on a real complaint type that actually comes up, not an abstract tone rule.Why: a rule like "be warm" doesn't tell anyone which words to change. A real pair does.
- Keep a separate pair type for genuinely serious complaints.Why: the same casual phrasing that works for a noisy air conditioner reads as flippant on a safety complaint.
- Check that a passing reply names something real from the actual review, not just that it sounds friendly.Why: a generic line can sound warm and still never mention the guest, the room, or the actual problem.
- Score tone drift on the lowest star drafts first, and leave every other stylistic dimension for later.Why: that's where a generic sounding reply costs the most, right where a guest is already unhappy.
- Re-run the check against the same batch after any prompt change, and read every flagged case.Why: a check nobody rereads is just a rule that used to matter.
How to answer this, stage by stage
Eight moves. Ground it in one drafted reply before touching the check itself.
Let's learn
Here is what happens when a reply gets every fact right, and still doesn't sound like anyone who worked there wrote it.
Say we build Porchlight, a tool inside Harborcrest Hotels, a small chain of eighteen boutique properties, that reads a guest's review and drafts a reply in the hotel's own voice. A front desk manager reads the draft, tweaks it, and posts it instead of writing one from nothing.
Before Porchlight, a manager wrote every reply by hand. A busy weekend brought in about fifteen reviews at one property. Writing a real, personal reply to each one took about ten minutes. That's two and a half hours, on top of running a front desk.
With Porchlight, a draft appears the moment a review comes in. Reading it, changing a word or two, and hitting post took about ninety seconds. Fifteen reviews went from two and a half hours down to about twenty minutes.
Here is the important part. The minutes saved were never the story. The real question was whether the draft still sounded like Harborcrest, and getting the facts right turned out to be a completely different thing from getting that right. A manager reading fast, deciding whether to tweak a word or hit post, stops reading the way she used to read when she wrote every line herself. She's proofing now, not composing. And a generic sentence can pass a fast proofing check that would never have survived her writing it from scratch.
At its worst, that cost more than a clunky reply. One quarter, twenty two of a hundred and thirty low star drafts, about one in six, got rewritten entirely by a manager because they read as boilerplate. And Harborcrest's whole pitch to a guest is that it isn't a big anonymous chain. A reply that sounds like one, sitting right under a review on the same page as that pitch, argues against the hotel's own brand.
At its worst, this let a specific kind of mistake through with nobody watching for it. A guest on her third stay at the Riverside property wrote that the unit next to hers had rattled all night, again. Porchlight drafted a reply that promised a fix and never once said which room, or that this was her third visit.
The choice I would take back. When Porchlight first shipped, the instructions we gave it said to reply politely and address the guest's issue, on the idea that a well written model would sound personable without anyone telling it how. That held up fine for the first few weeks, when volume was low enough that someone read every draft before it posted. It stopped holding up once the chain leaned on Porchlight across all eighteen properties, and nobody had ever built a specific check for whether a reply sounded like Harborcrest instead of just any hotel.
What I would leave alone. A five star review with a cheerful "thanks so much, glad you enjoyed your stay" reply costs nothing, even if it's a little generic. Almost nobody rereads a happy reply. Spending check effort there takes it away from the drafts that actually need it.
The lesson. Sounding right isn't a bonus feature sitting on top of being correct. It's a separate thing, and it needs its own specific check, built from real examples, not a hope that good instructions will make it happen on their own.
Now here is the same thing as a story
Read this one when you've got a few minutes. The short version is above. This is for when you want to feel why it mattered.
Innes Whitlock could read a guest review and tell you, before finishing the first sentence, whether it was a real complaint or someone just having a bad Tuesday. She'd run Harborcrest's guest experience team for four years, long before anyone built Porchlight, back when eighteen properties meant eighteen managers writing every reply themselves, some good at it, some not.
Porchlight arrived in the spring. For months, it was the best thing that had happened to those front desks. A review came in, a draft appeared, a manager read it, changed a word, and posted it. What used to eat an hour of a Saturday morning now took fifteen minutes. Innes read a sample of the drafts every week at first, and they sounded like Harborcrest: specific, a little wry, never like a form.
She stopped reading the sample every week. Then every month. The ones she checked always looked fine.
Then came the third stay review.
A guest wrote that this was her third time at the Riverside property, and that the unit next to hers had rattled all night, again. Porchlight drafted a reply: "We sincerely apologize for any inconvenience experienced during your stay and have noted your feedback for review by our maintenance team."
Nothing in that line was wrong. It addressed the noise. It promised a fix. Riverside's general manager had her cursor on the post button when something made her stop and reread it. It never said room two fourteen. It never said third stay. It could have gone out under any hotel's name in the country.
She deleted the draft and wrote her own line instead. But she also sent Innes a message: "this could've gone out. is that what we want going out?"
I want to say the problem was that Porchlight got the facts wrong. It didn't. Every fact in that draft was true. But Innes had never given anyone, or anything, a specific way to check for the other thing, the thing that actually made a Harborcrest reply sound like Harborcrest. She had a feeling for it. Porchlight didn't.
So here is the decision she made differently.
When the team first wrote Porchlight's instructions, they told it to reply politely and address the issue. That was the whole guidance on tone. It made sense then, when volume was low and every draft got a human read before it went anywhere. Nobody had imagined leaning on it for every property inside a year.
Innes built a check instead. Not a score. A stack of real pairs: a bad line, and next to it, the line a real manager would actually send, with the words marked that made the difference. For the AC complaint, marked words like "sincerely apologize" and "your stay" put a line on the wrong side. A room number, a plan, and a date put it on the right side.
Run the same quarter forward with the check in place. Low star drafts get checked against the pairs before anyone reads them by feel. The following quarter, out of a hundred and thirty low star drafts, only three needed a full rewrite. The rest passed, or got flagged and fixed in under a minute, because the check named exactly which words were the problem instead of asking someone to feel it out.
One design hands a manager a feeling. The other hands her a reason.
And the thing I'd want to tell myself, back when we wrote Porchlight's first instructions: I asked it to be polite. I never told it what polite was supposed to sound like, coming from us.
SPARK, sized to one guest reply
This question asks how to handle evals for something subjective, not literally design a feature, but it's still one concrete decision about the exact moment a reply sounds like Harborcrest or doesn't, so SPARK still fits. A question asking how to measure Porchlight's overall reply quality across the whole chain would reach for LEAD instead.
And if you want to be sure it really works, try it somewhere else
A food rescue nonprofit is a different kind of operation entirely, but the same gap between technically fine and actually right shows up in a donor's thank you note.
S. Ottoline Prescott runs operations for Second Table, a food rescue nonprofit that picks up surplus food from grocers and restaurants and gets it to community pantries the same day. Today, without a check, a coordinator drafts a thank you note from a template and either personalizes it by hand or lets it go out generic, depending on how busy the week is.
P. The habit worth building: a donor reads a thank you note and believes it, because it names something real that happened with her actual gift, not because the note used the word grateful enough times.
A. Same shape, a different pair. Bad: "thank you for your generous contribution, which helps further our mission." Good: "because of your gift this week, three hundred forty pounds of bread that would have hit a dumpster fed about two hundred families at the Elm Street pantry." The marked words: a real number, a real place, a real result.
R. Tighten the check to require a precise pound count on every note, and a ten dollar monthly donor gets a note that sounds like the org is bragging about her ten dollars, which reads as showy instead of grateful. Loosen it to "sounds warm," and "we're so grateful for your support" passes again, saying nothing real about what her gift actually did.
K. No attempt yet to personalize every note by donor history, holiday, or gift size. One check, does the note name something real that happened, for every note tied to an actual pickup, is enough to start.
It took a longtime ten dollar monthly donor writing in to ask why her note thanked her for three hundred forty pounds rescued that week, when her own gift couldn't have covered a tenth of that, before anyone noticed the check needed a second guard rail: match the specific claim to the size of the actual gift, not just to whether a pickup happened that week.
Swap the trigger and it still runs
- Speed: even if Porchlight drafted every reply instantly, which it already does, a fast wrong sounding reply is still wrong. Speed never fixes what the check is testing.
- Cost: if running the check cost nothing and needed no manager time at all, that still wouldn't tell you which pair type applies. A check still needs to know a safety complaint isn't an AC complaint.
- The model gets better: if Porchlight's facts became near perfect, never missing a room number or a detail, it could still default to safe, generic phrasing, because that's often the safest sounding text. Better facts don't fix a tone problem. They're different things.
Where people run it wrong
- Writing one banned word list and calling it done, so it blocks straight talk in the one case, a real safety complaint, where straight talk was the right call.
- Writing the good example without marking which words make it good, so two managers argue about the same line for months.
- Trying to score every stylistic dimension, humor, emoji use, formality by star rating, in the first version, so the check ships months late while generic drafts keep going out the whole time.
How to use it live
If you're asked this cold, ask what a technically correct but wrong reply would actually sound like for this product. Write down the specific words that make it wrong. Then ask whether today's check could catch it. That question finds the missing rubric faster than listing every tone rule from a blank page.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Writing an eval spec
- #1 What is an eval spec and who is its audience?
- #2 List the components of a complete eval spec.
- #3 How do you define a task-level success criterion for a summarization feature?
- #4 Write a scoring rubric for the quality of a generated customer support reply.
- #5 Describe the difference between an eval spec and a test plan.
- #6 How many examples belong in a first eval set and how do you choose them?