ConceptIntermediateEval-Driven Specification / Writing an eval spec / #5

Describe the difference between an eval spec and a test plan.

The direct answer
A test plan checks that the output runs and comes out the right shape: right length, right fields filled in, the required text there. An eval spec checks whether the output is actually good: whether the tone fits what's happening, whether it could read like an accusation. Write the eval spec's own quality criteria from scratch, and never let a pass on the test-plan checklist stand in for one.
Do this, in order
  1. Write the eval spec's own quality criteria, separate from the test-plan checklist.Why: this is the decision the whole flip turns on. Skip it and a hundred percent pass rate keeps meaning nothing about tone or harm.
  2. Score the things a checklist can't: tone, proportion, whether the copy could read like an accusation.Why: these are exactly the dimensions a pass/fail test plan is built to miss.
  3. Keep the test-plan checks running too.Why: length, fields, disclosures, and leaked data are still worth catching mechanically, and it's cheap to keep catching them.
  4. Route anything touching a regulated or high-stakes screen through a person, until the quality rubric has real data behind it.Why: the cost of getting it wrong sits highest on exactly those screens.
  5. Watch for long clean streaks on the eval spec.Why: a run of all-green weeks is exactly when nobody's reading closely, and that's when the miss slips through.

How to answer this, stage by stage

Eight moves. Most of the weight sits in stage four: this question is really asking whether you understand that a checklist can pass every single time and still never once ask if the thing is good. Every stage has the actual words to say.

1
Ground it in one product, one person
Say it like this
"Let me make this concrete. Tidemark is a fintech app, and it's got a tool that writes the copy shown across signup: welcome screens, the identity-check prompts, the error messages. Rahel Tesfaye runs product content there, and eighteen months in, she finds out a line that passed every single check in her eval spec had been telling new customers their account might get frozen, over a totally routine ID check."
Why this works
Nobody can judge whether a test plan and an eval spec are different things without a real line of copy doing real damage to a real person.
2
Say your structure in one breath
Say it like this
"Five things, fast. What she stopped doing once the checklist kept coming back green. The switch with no middle setting. The old call I'd take back. What I'd measure differently. And the same kind of copy, replayed with a real quality rubric in place."
Why this works
A named route up front tells the interviewer you have a plan, not a story you're inventing live.
3
Reframe what the question is actually testing
Say it like this
"This isn't really 'define an eval spec for me.' It's asking whether I get that a checklist built to test 'does it run' can pass a hundred percent of the time and still never once ask 'is this good.' Those are two different questions, and a test plan structurally can't ask the second one."
Why this works
That gap, between reciting the definitions and seeing why the gap exists, is the trap this question sets.
4
Give the one decision
Say it like this
"Concretely: a test plan checks that the output runs and comes out the right shape, right length, right fields filled, required text present. An eval spec checks whether the output is actually good, whether the tone fits what's happening, whether it could read like an accusation. I'd keep the test-plan checks, they're cheap and they still catch real bugs, but I'd never let a pass on them stand in for the eval spec. I'd write the eval spec's own quality criteria from scratch."
Why this works
There's an actual mechanism in that sentence, not just "be more careful."
5
Prove it with the compressed failure
Say it like this
"Say the eval spec started as a copy of the QA team's test-plan template: right length, right fields, disclosure text present, clean toxicity filter. It ran clean for twenty straight weeks. Then a new variant passed every one of those checks and told a customer their account might get frozen over a routine ID check. It ran live for eleven days, about 3,400 people saw it, and completion on that screen fell from 71% to 38%, before a support lead's offhand comment caught it."
Why this works
Four sentences, and it still lands on the exact question the checklist never had a field for.
6
Show the fix, concretely
Say it like this
"I'd add a real quality rubric, scored by two people, on anything touching identity or account limits: does this imply wrongdoing with no cause, is the tone right for what's actually happening, would a stressed reader feel accused. That runs alongside the test-plan checks, not instead of them."
Why this works
Shows this is additive, not a rejection of the checks that were already working fine.
7
Say what you'd leave alone
Say it like this
"A tooltip that just says a field auto-fills from your linked bank account doesn't need any of this. Nobody's anxious reading that line, and there's no accusation it could make by accident. That one stays a pass/fail check for good."
Why this works
Shows judgment, not a blanket new process bolted onto every screen in the app.
8
Close on the one line
Say it like this
"So here's the whole thing in one breath. A test plan asks if the output works. An eval spec asks if it's good. If you only ever answer the first question, a hundred percent pass rate can still be actively hurting the person reading it."
Why this works
Ends on the exact sentence an interviewer can repeat back to their own team.
If you remember one thing A test plan can pass every single check and still be actively hurting the person reading it, because nothing in a pass/fail checklist was ever built to ask if something is good. That question needs its own criteria, on purpose, or nobody's asking it at all.

Let's learn

Every year, before any of this existed, two writers on Rahel Tesfaye's team wrote every line of copy shown during Tidemark's signup by hand. A new screen took about two weeks: a draft, a compliance read, a legal read, a round of changes. In a good year they touched maybe twenty screens total.

Then Tidemark built a tool that drafts that copy instead: the welcome screen, the identity-check prompts, the error messages, the little nudges that keep someone finishing signup. Last quarter alone it produced 140 new variants, across screens, customer groups, and the three languages Tidemark supports.

Before any of that copy ships, it has to pass an eval spec Rahel and a QA engineer built eighteen months ago, when the tool first launched. It checks the character limit. It checks that the customer's name and the dollar amount fill in correctly. It checks that the required legal disclosure text is there, word for word. It checks for leaked personal data. It checks that a toxicity filter comes back clean. Six checks, every one pass or fail.

Knowledge spark: what's a KYC check? Short for "know your customer." A routine step where a bank or app reverifies who you are, usually triggered by something ordinary, like a first big transfer. It's a normal rule that applies to everyone. It isn't an accusation.

For the first two months, Rahel read every new variant herself before it shipped, even the ones that already passed. Then she started reading only the first of each new type. Then, once the eval spec had come back green for twenty straight weeks, she stopped reading personally at all. A pass was the ship signal, on its own.

One of those variants was built for a specific moment: a customer's first transfer over $2,000, which triggers Tidemark's routine identity reverification. It read: "We've noticed unusual activity on your account. Verify your identity now or your account may be frozen." It passed every check. Right length. Right amount filled in. Disclosure text present. No leaked data. Clean toxicity filter.

Extra support tickets mentioning "frozen" or "flagged," cumulative, over 11 days
0 15 30 45 remark, caught day 1 day 3 day 6 day 9 day 11
Forty seven extra tickets, over eleven days, all tracing back to one line that passed every check in a spec built to test whether copy runs, not whether it's kind.

It ran for eleven days. About 3,400 people saw it. Completion on that screen fell from its usual 71% down to 38%, and the ticket count above kept climbing, a little more each day, until a support lead noticed the tag count on a Friday and messaged Rahel almost as a joke: was Tidemark trying to scare people during a normal ID check?

The extra tickets were not really the problem. The real problem is that nothing in the eval spec had ever asked whether that line sounds like an accusation. A hundred percent pass rate never once meant the copy was good. It only ever meant the copy ran.

We didn't lose eleven days of review. We handed new customers a false accusation, and called it a pass.

At its worst, this is what a regulator reads when they ask why a routine ID check sounds like a fraud accusation to the person who got it. Nobody at Tidemark meant to accuse anyone of anything. That's not what shows up in a support chat log, though, and it's not what a customer who just wanted to move their own money feels reading it.

The decision that mattered Write the eval spec's own quality criteria, scored by real people against real questions: does this imply wrongdoing, is the tone right for the stakes. Not a stricter pass threshold. Not a bigger compliance meeting.

What I would leave alone. A tooltip explaining that a field auto-fills from a linked bank account doesn't need any of this. Nobody's anxious reading that line, and there's no accusation it could make by accident. That one can stay a pass/fail check for good.

The lesson. A checklist built to test whether something runs will always miss whether it's good, because that was never the question it was built to ask. You have to build a second, different question on purpose, or nobody's asking it at all.

Now here is the same thing as a story

Pull this one out when there's room to sit with it, not just tick it off a list.

Every Monday morning, Rahel opened the eval-spec dashboard before she opened anything else, and for a long time, all it ever showed her was green.

She'd written Tidemark's identity-check copy by hand for three years before the generator existed, and she had a good ear for the one thing that mattered most there: whether a line sounded like a routine ask, or like an accusation. Ask her to read a paragraph of KYC copy and she could tell you in five seconds which one it was.

The generator launched eighteen months ago. Rahel and a QA engineer built its eval spec in one afternoon meeting, working from the QA team's existing test-plan template for a different feature, because building something new from scratch felt like a week nobody had. Six checks. Right length, right fields, disclosure text present, no leaked data, clean toxicity filter, reading level under a target grade. All pass or fail.

For the first two months, Rahel read every new variant herself before it shipped. Not because she distrusted the checklist. Because the checklist was new, and she wasn't yet. Every Monday she'd open the week's new copy and read it top to bottom, and every week it read fine.

Then she started reading only the first of each new type, the rest she trusted to look the same. Then, once the eval spec had run clean for twenty straight weeks, not one flagged issue, she stopped opening the dashboard's copy tab at all. She just watched the pass count. Green meant done.

Left, a dial with many marks labelled how carefully to read a new variant, somewhere between skimming and word by word, captioned what we assumed she had. Right, a switch with two positions, trusts the checklist completely or reads it herself first, pushed to trusts the checklist completely, captioned what she actually had.
People are switches, not dials

The trigger was small. A support team lead, doing a normal end-of-week look at ticket tags, saw a cluster mentioning "frozen" and "flagged" that hadn't been there a month ago. He messaged Rahel almost as a joke: were they trying to scare people during a routine ID check, or was that on purpose?

Rahel pulled up the eval-spec dashboard first. Fully green, same as every week. Then, for the first time in months, she opened the actual copy and read it herself. "We've noticed unusual activity on your account. Verify your identity now or your account may be frozen." Written for a customer's first transfer over $2,000, a completely routine check. It had passed every one of the six checks. It had been live for eleven days. About 3,400 people had read it.

She didn't just fix that one line and move on. She and a compliance lead spent two days reading every new variant touching identity or account limits from the last twenty weeks, about sixty four of them, and found eleven more with the same pattern: technically clean, functionally fine, and written in a way that would make a stressed reader feel accused.

We didn't take away Rahel's eleven days. We took away the one question the checklist was never built to ask.

I want to say the problem is that someone wrote a careless line. They didn't. "Verify now or your account may be frozen" is a completely reasonable thing to write if the only questions in front of you are about length and fields and disclosure text, and nothing tells you tone is also your job. That's not really the story either. Rahel never had a number in her head for how carefully to read a new variant. She had a feeling, and it only had two settings: this checklist is proof, or this checklist is a start. Twenty clean weeks flipped it for good, and one bad line flipped it back.

Here's the call I'd take back. In that first afternoon meeting, eighteen months ago, when we opened the QA team's test-plan template and started filling it in, nobody in the room was doing anything careless. There wasn't a real quality template yet, and everyone had used a test plan before, so it was the fastest thing on hand. Nobody ever came back and asked whether a checklist built to test "does it run" could also carry the weight of "is it kind."

I'd add a real quality rubric. Not instead of the six checks, alongside them. Two people score anything touching identity or account limits on three questions: does this imply wrongdoing with no cause, is the tone right for the actual stakes, would a stressed reader feel accused. That runs before ship, same as the six checks, just asking a different kind of question.

Run the same trigger again, three months after the fix. A new market needs its own KYC copy, at a lower transfer threshold. The first draft reads a little urgent, some pressure language that isn't quite an accusation but is close. It passes all six of the old checks, clean as anything. The quality rubric catches the tone inside a day. It ships rewritten. No real customer ever sees the first draft.

If I'm honest, nobody made a bad call in that first afternoon meeting. The bad call was mine: three years writing that copy by hand, and I still reached for the QA team's template instead of asking what a checklist for "good" would even need to look like.

The five moves behind a checklist that couldn't see tone

The letters matter less than which one breaks first. Here's the same five steps, mapped onto Rahel's eval spec.

Five stacked rows, F L I P S, each a hand lettered capital in a circle, a step name, a short question, and the answer in this story. The I row is outlined in red-orange.
FLIPS, five rows
FFind the person
Whose morning is this?
Not "content ops" in the abstract. Whoever actually owns the eval spec the copy generator has to pass.
In this answer: Rahel Tesfaye, senior product content lead at Tidemark, three years writing the app's identity-check copy by hand before the generator existed.
LLocate the habit
What did she stop doing because it kept working?
Look for the check that quietly went from routine to skipped, not her overall care. A long clean streak is what buys a habit like this its opening.
In this answer: She went from reading every new variant herself, to reading only the first of each type, to not reading any of them, once the eval spec ran clean for twenty straight weeks.
IIdentify the flip
What two setting switch snaps, with no middle?
"She got less careful about reading copy" describes the outcome, not the action. Name the exact two states with nothing between them.
In this answer: Trusts a checklist pass as proof it's ready to ship and never reads a variant herself, or treats a pass as necessary but not enough and reads anything touching a regulated step first. Nothing in between once she found the line that passed everything and still read like an accusation.
PPinpoint the old decision
Which call only made sense before there was real volume and real stakes riding on it?
Look for a narrow, defensible call from the eval spec's earliest days. "Read copy more carefully" after the fact doesn't count, that's a new dial.
In this answer: Copying the QA team's test-plan template for the eval spec at launch, because nobody had built a real quality rubric yet and the room needed something fast.
SShow the replay
Same kind of copy, quality rubric restored. Better ending?
Run the identical trigger through the fixed design and see where it stops. A count or a clock, not an adjective.
In this answer: A new KYC variant reads a little urgent. The quality rubric catches the tone inside a day, not eleven, and no real customer ever sees the first draft.
Two panels. Left, a gently rising line labelled how long the checklist stayed green, from 2 weeks clean to 20 weeks clean. Right, a line labelled how often she reads a new variant herself, that starts flat and high at reads every one, then jumps straight down and runs flat and low at reads none, with the drop marked the snap.
A small move in the streak. A hard snap in what she did about it.

"She got less careful about reading copy" is a diagnosis anyone can offer after the fact. The harder part is naming the exact habit that had to break first, opening the copy tab on a Monday, and showing there was no smaller version of it left once it did.

And if you want to be sure it really works, try it somewhere else

Saltmarsh Family Pharmacy runs a tool that drafts the counseling scripts pharmacists read aloud, or print, at pickup, across its 40 stores. Same question, a pharmacy counter instead of a signup screen, and a flip that fires on a joke, not a bad number.

F. Jocelyn Panganiban, pharmacy operations lead at Saltmarsh, who set the standard for what a safe counseling script sounds like across every store.
L. Whenever a pharmacist felt a draft script sounded off, too clinical, too rushed for an older patient, they'd tag it in the shared queue: "AI draft, needs a look." That tagging habit is what fed Jocelyn's team the pattern of what needed fixing.
I. A different flip. Pharmacists don't stop using the tool, and they don't check less. They stop saying they used it. Flags every off script for review, or quietly tweaks a word and lets it go with no flag at all. Nothing in between once a regional supervisor joked, in front of the team, that a pharmacist had "just let the robot write it" after a patient complained.
P. The flag was visible by design, named next to whoever raised it, so Jocelyn's team could ask follow-up questions. That was free when using the draft was normal. It became a cost the moment it was a mark against you.
S. Make the flag anonymous: it records that a script was AI-drafted and needs a look, never who flagged it. A new script for a complex, high-dose medication gets flagged again within days, not silently patched by an embarrassed pharmacist and left in rotation.

Days before a problem script got caught, Saltmarsh
Named flag, after the huddle joke
Old design
19 days
Anonymous flag restored
New design
under 1 day
quietly used, unflagged then found by luck
Old design: 14 days where pharmacists used the draft and tweaked it silently, plus 5 more before someone happened to notice, 19 days total. New design: an anonymous flag routes the same script to review the same day it's drafted.
A second decision worth taking back Making tool use visible by name is itself a decision, not a neutral fact about how a queue works. A queue that flagged the draft without flagging the person would have caught the collapse-risk script in a day, not three weeks.

Swap the trigger and it still runs

  • Speed: if Tidemark only shipped new copy twice a year instead of continuously, that line would still slip through eventually, it would just take longer to reach twenty clean weeks and lull anyone watching.
  • Cost: if a real quality rubric took a full day of two people's time per screen, teams would ration it toward whatever felt riskiest that week, not whatever the pattern actually was.
  • The model got better: this is really what happened at Saltmarsh. The scripts weren't getting worse. The tool kept meeting new drug combinations nobody had written a script for yet, and "nothing looks broken" is exactly what makes anyone stop double-checking.

Where people run it wrong

  • Blaming the person who wrote the first checklist, when nobody told them a quality rubric would still be load-bearing eighteen months later.
  • Adding a bigger, stricter compliance review board, when the real gap is that nothing anyone reviews ever asks about tone at all.
  • Waiting for a customer complaint or a regulator's letter to reveal the gap, instead of checking right now how much of what ships only ever gets checked by a pass/fail list.

How to use it live

Buy yourself a few seconds by naming the reframe before the fix: "The question isn't whether we have an eval spec, we do. It's whether anything in it can tell the difference between copy that runs and copy that's actually good." Say that, and the rest of the answer is just the mechanism.

Flashcards (click a card to flip it)

Eight fixed slots, pulled straight from the answer above.

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Rahel checks sometimes, then stops checking at all, once a long clean streak makes the checklist feel like proof instead of a starting point.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Rahel Tesfaye, senior product content lead at Tidemark, who wrote the app's identity-check copy by hand for three years before the generator existed.
3 · THE HABIT
What did she stop doing because it kept working?
Tap to flip
ANSWER
She stopped personally reading new copy variants. First she read only the first of each new type, then nothing at all, once the eval spec ran clean for twenty straight weeks.
4 · THE FLIP, HERE
What's the two setting switch in this story?
Tap to flip
ANSWER
Trusts a checklist pass as proof it's ready to ship and never reads it herself, or treats a pass as necessary but not enough and reads it first. Nothing in between once she found the line that passed everything and still read like an accusation.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Copying the QA team's test-plan template for the eval spec at launch, because nobody had built a real quality rubric yet and the room needed something fast.
6 · THE NUMBER
The bad KYC line ran live for ___ days, seen by about ___ people, before a colleague's remark caught it.
Tap to flip
ANSWER
11 days, about 3,400 people. Completion on that screen fell from 71% to 38% while it was live, and every one of the six pass/fail checks stayed green the whole time.
7 · THE REPLAY
Same kind of copy, quality rubric restored, what changes?
Tap to flip
ANSWER
A new KYC variant reads a little urgent. The quality rubric catches the tone inside a day, not eleven, and no real customer ever sees the first draft.
8 · CROSS-PRODUCT
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Saltmarsh Family Pharmacy's patient counseling-script generator, using the concealment flip: pharmacists stop flagging AI-drafted scripts once one becomes a joke at a regional huddle.

Check yourself Score: 0 / 0

Short answer
1. In your own words, what's the real difference between a test plan and an eval spec?
Show hint
Think about what question each one is even capable of asking, not just what it happens to check.
Show answer
Model answer: "A test plan checks whether the output runs and comes out the right shape, right length, right fields, required text present. An eval spec checks whether the output is actually good, whether the tone fits the moment, whether it could read like an accusation. A test plan can pass a hundred percent of the time and still never once answer the second question, because it was never built to ask it."
Multiple choice
2. What old decision does this answer take back, and why did it make sense when it was made?
  • A. Tidemark should have hired an outside compliance vendor to write the copy instead of building a generator.
  • B. Copying the QA team's test-plan template for the eval spec at launch, because nobody had built a real quality rubric yet and the room needed something fast.
  • C. The generator's underlying model got worse over the eighteen months.
  • D. Add a stricter compliance review meeting before any copy can ship.
Show hint
The right answer names a specific, small choice from the eval spec's earliest days, not a new process bolted on afterward.
Show answer
B. D is the trap answer, "add more review" is a new dial, not an old choice taken back. C never happened, the model didn't get worse, the checklist just never measured tone. A is a plausible-sounding fix but it isn't the actual decision this story reverses.
True or false
3. True or false: Rahel could have fixed this by simply reading new variants a little more carefully from now on, without ever changing what the eval spec actually measures.
  • True
  • False
Show hint
Ask what happens to the eleven other variants that already shipped during the same clean streak.
Show answer
False. Eleven more variants with the same pattern had already shipped during those twenty weeks, and reading more carefully going forward doesn't catch what's already live, and doesn't scale once volume grows again. The checklist still wouldn't ask about tone on the next variant either. Only adding a real quality criterion catches it going forward.
Fill in the blank, do the math
4. The ticket chart shows 47 extra tickets by day 11, building from 2 on day 1. If the pattern had only run for about half as long before anyone caught it, roughly how many extra tickets would you expect, based on the chart's own numbers?
Show hint
Half of eleven days is close to day 6. Look at where the line already sits by then.
Show answer
About 19 tickets. That's roughly where the line sits at day 6, the halfway point. The count also wasn't growing at a steady pace, it climbed faster in the second half, which is exactly why eleven days of silence cost so much more than six would have.
Short answer, apply it yourself
5. Think of a checklist, test suite, or benchmark you rely on at your own work. What question can it never answer, no matter how many times it passes?
Show hint
Ask what the checklist was originally built to test, not what people have started using it for since.
Show answer
Model answer: "Our support macros get checked against a script-validity test: does every variable resolve, does the ticket close correctly, is the reply under our length limit. It's never once asked whether the reply actually sounds like it cares. A macro can pass every check we have and still read as cold to the person receiving it." Any real example counts, as long as it names a specific dimension the checklist structurally cannot see, not just "we should test more."
Multiple choice
6. Which new copy variant would still be fine to ship on nothing but the old pass/fail checklist, no quality rubric needed?
  • A. A tooltip explaining that a field auto-fills from a linked bank account, with no coverage or accusation risk in it.
  • B. A message telling a customer their transfer is under review for a large-amount check.
  • C. A message asking a customer to reverify their identity after a routine trigger.
  • D. Any message shown on a KYC or account-limit screen, no exceptions.
Show hint
Ask whether the line could ever be misread as an accusation, or whether it's just describing something mechanical.
Show answer
A. An auto-fill tooltip is a purely mechanical fact with no tone risk in it. B and C both sit on exactly the regulated, high-stakes screens where tone can read as an accusation, which is what the quality rubric exists for. D is too broad, treating every KYC-adjacent line as equally sensitive misses the real distinction between "explains a fact" and "implies something about the customer."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more