Artifact critiqueIntermediateQuality, Cost & Token Economics / Eval design for product teams / #19

Critique an eval that uses the same examples used to write the prompt.

Score it on rows nobody used while writing the prompt, or the number on the dashboard is measuring memory, not skill.

The direct answer
An eval built from the same examples used to write the prompt cannot fail. It is not testing whether the prompt works, it is testing whether the prompt remembers the rows it was shown while someone wrote it. Build the eval from a fresh set nobody used or looked at while the prompt was being built, keep that set sealed, and only trust the score it gives you.
Do this, in order
  1. Score the prompt on a fresh, held out set nobody used or looked at while writing it.Why: an eval built from the same rows as the prompt cannot fail, it can only remember.
  2. Split any eval score by whether an example's category was in the writing set.Why: a flat, high blended number can hide near total failure on anything outside that shape.
  3. Check who is grading, not just what they're grading.Why: a grader who wrote the original examples scores from memory, not from the words on the screen.
  4. Run the fresh set score as the real evidence test before touching the prompt again.Why: it is the one number that tells contamination and a genuine prompt problem apart.
  5. Keep the held out set sealed and rebuild it on a schedule, not on demand.Why: the moment someone peeks at it while tuning the prompt, it has already stopped being held out.

How to answer this, stage by stage

Nobody is grading whether you know what an eval is. They are grading whether a 94 percent score would make you suspicious or make you relax.

1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one tool. Bulletforge is Ferngrove's resume bullet rewriter. A job seeker pastes in one flat line, like 'handled customer complaints,' and it hands back a stronger version with a number in it. Zorana Vaduva owns the eval that decides whether that rewrite is any good."
Why this works
An abstract "what's wrong with this eval" answer turns into a definition of the word eval fast. One product turns it into a real decision.
2
Say your structure out loud before diving in
Say it like this
"I'm going to name the actual flaw first, then show you the one check that proves it, then say what I'd build instead. The flaw is simple: if the eval and the prompt were built from the same folder of examples, the eval can't tell you anything true."
Why this works
Tells the interviewer you have a plan and a payoff, not just a definition to recite.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me to spot a technical bug. It's asking whether I know the difference between a prompt that works and a prompt that has just memorized the answer key it was written next to."
Why this works
Stops you from listing eval best practices instead of critiquing the actual thing you were shown.
4
Give the one decision, plainly
Say it like this
"Score it on a fresh set of examples that nobody touched while writing the prompt. Ones the writer never saw, never edited, never used as a hint. If the score holds up there, the prompt actually generalizes. If it doesn't, the first number was never real."
Why this works
This is the actual answer to the question, in one breath, with the fix included.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without that. Zorana's team wrote Bulletforge's prompt using 24 example bullets, mostly from office jobs like sales and marketing. The eval, graded on those same 24 rows, read 94 percent and stayed there for eight weeks straight. The first time someone ran it against 150 fresh bullets from warehouse workers and home health aides, jobs nowhere in the original 24, it scored 61."
Why this works
Shows the real cost of the flaw, not just the mechanism behind it.
6
Say what you would measure going forward
Say it like this
"I'd track two numbers side by side forever, not one. The score on the held out set, and the share of users who keep the rewrite without editing it, split by job category. The moment those two numbers stop agreeing, something's contaminated again."
Why this works
Shows you're thinking past this one bad score, into the thing that catches the next one early.
7
Say what you'd leave alone
Say it like this
"I wouldn't rebuild the eval every week. Once the held out set is genuinely untouched, re-running it after every small prompt change is enough. Rebuilding the set itself that often just burns the one thing that makes it worth anything: nobody's looked at it yet."
Why this works
Shows judgment instead of blanket caution applied at the same cost everywhere.
8
Close on the decision, not the story
Say it like this
"So: never grade a prompt on the examples that taught it the answer. Hold out a fresh set, seal it, and let that be the only number you actually trust."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd reuse on the next eval you're handed.

Let's learn

Bulletforge is a tool on Ferngrove's career coaching site. A job seeker pastes in one flat resume line, and it hands back a sharper version with a real number in it, in under ten seconds.

Knowledge spark: what's an eval? A fixed test you run a prompt against every time it changes, so you get one number back instead of a feeling. It's only honest if the test never leaks into the thing it's testing.

Before Bulletforge, most job seekers rewrote bullets by hand. It took about 25 minutes a bullet to find a real number to put in it, and most people gave up after fixing three or four, leaving the rest flat and weak.

Zorana Vaduva's team wrote Bulletforge's prompt in a three day sprint, using 24 hand picked resume bullets, mostly from office jobs: sales, marketing, and a few in software and admin. Those same 24 rows became the eval. Checked every Monday, it read 94 percent, and it kept reading 94 percent, week after week, for eight weeks straight.

Weekly eval score vs. the share of rewrites users kept unedited, week 1 to week 8
100% 40% wk 8: fresh audit scores 61 Wk 1 Wk 4 Wk 8
Eval score, on the original 24 rowsReal users who kept the rewrite unedited
The eval line barely moves, it's the same 24 rows every single week. The real keep rate slides from 84 percent down to 58 percent as more users outside the original 24 job types show up, and nobody was watching that number at all.
Knowledge spark: what's a held out set? A group of examples nobody used while building or tuning something, kept aside just to grade it later. The moment someone looks at it while the thing is still being built, it has stopped being held out.

The eval never once said anything was wrong, because it never got to ask a new question. It was the same 24 rows Zorana's team had already looked at, line by line, while writing the prompt. Passing them wasn't proof the prompt worked. It was proof the prompt could recite what it had already seen.

The eval wasn't lying. It just never got asked a question it didn't already know the answer to.

Zorana pulled 150 fresh resume bullets that had actually come in from real users, chosen on purpose from job categories the original 24 never touched: warehouse work, home health aide, HVAC technician, freight dispatcher. She graded the same rubric against those, by hand, with someone who hadn't written a single one of the original 24.

The evidence test: the original 24 examples vs. 150 fresh, never seen bullets
94% 61% The original 24 examples 150 fresh, never seen bullets
Score used to write the promptScore on rows nobody used to write it
A 33 point gap between the two scores, on the exact same rubric. That gap is the proof: the prompt never learned resumes in general, it learned those 24 rows.
The choice that mattered Three days from launch, with no labeled data yet, the team reused the same 24 row file for two jobs at once: writing the prompt and grading it. Somebody said they'd build a real eval set later. Later never came, because the number on the dashboard never gave anyone a reason to go looking.

At its worst, a tool that hands out confident, generic bullets to warehouse workers and home health aides, while reading 94 percent on a dashboard nobody questions, is worse than no tool at all. The old way, rewriting a bullet by hand, was slow, but a person knew when their own sentence sounded weak. A flat, confident rewrite from Bulletforge looked exactly like a good one.

What I'd leave alone: the rubric itself. Checking for a number, an action word, and a real result is a fine bar and doesn't need to change. The examples feeding the eval were the problem, not the questions the eval was asking.

The lesson: a score that never moves isn't reassuring, it's a sign the test stopped asking anything new eight weeks ago. Ninety four percent told Zorana's team Bulletforge was fine. It never told them fine meant fine on the same 24 rows they'd already memorized by hand.

Now here is the same thing as a story

Read the long version below when you want to feel why a 94 percent score can be a lie nobody meant to tell, not just be told to build a fresh set.

Before Ferngrove, Zorana Vaduva spent three years building eval sets for a hiring assessment startup. She could smell a leaky eval before she'd read the second row of a spreadsheet, which is exactly the skill she brought with her when Ferngrove hired her as their first data scientist.

Bulletforge launched, and for two good months it did what it was built to do. Job seekers pasted in a flat line and got back something sharper in seconds. The Monday morning dashboard read 90, then 91, then 94, and held there. Zorana glanced at it each week and moved on to the next thing on her list, the way you do with a number that's never once given you a reason to stop.

The habit thinned out in three small steps. First she stopped opening the underlying spreadsheet to spot check a few rows by hand, since the number was fine anyway. Then she stopped sampling real user output at all, trusting the weekly score to stand in for it. Then, most weeks, she only half read the number before closing the tab.

The trigger was small. A new data analyst, three weeks into the job, asked over lunch whether the eval file and the prompt's example file were the same csv. Zorana said no, obviously not, and then actually went and checked.

They were the same file. Not similar. The same 24 rows, opened for two different reasons, on the same afternoon back in week one.

She spent that afternoon pulling 150 real bullets that had come in from actual users over the past two months, on purpose choosing job categories nobody had thought about while writing the prompt: warehouse work, home health aide, dispatch, HVAC. Someone who hadn't written a single one of the original 24 graded them blind, using the same rubric the eval always used.

Sixty one percent. Not 94. For two months, a real slice of Bulletforge's users, mostly in lower paid, non office jobs, had been getting rewrites that sounded confident and were often generic or a little off, while the dashboard read fine the entire time.

The meeting where the old decision got made was nothing dramatic. Three days before launch, someone said they already had 24 good examples, so why not eval against those for now, they'd build a real set once things calmed down. It made sense that week. Nobody had labeled data yet, and a fake number felt better than no number at all. Nobody ever circled back, because the number never once told them to.

Run the same eight weeks again with one change: the held out set gets built in week one, sealed the same day, never opened by anyone touching the prompt. The 61 percent shows up on day one instead of week eight. Zorana fixes the prompt for warehouse and home health aide language before a single one of those users ever sees a weak rewrite that reads like a good one.

One design let a single spreadsheet play two roles and trusted the number that came out of it. The other design keeps the exam locked away from whoever is teaching the class.

What I'd tell myself, back on the afternoon that 24 row file got reused: if the same rows can answer two different questions, you don't actually have two files. You have one file, quietly lying to whichever question it answers second.

TRACE, applied to a score that was too flat to be real

Not five guesses about what's wrong with the prompt. TRACE rules candidates out on purpose, until one check actually separates a leaky eval from a genuine failure.

TTimeline. When exactly did the eval stop meaning anything, and what shipped right before that?
The eval and the prompt were built the same week, from the same 24 row file. The score never moved for eight weeks after that, which should have been the tell, not the reassurance.
A number that never moves at all is often not measuring anything live, it's just asking the same question it already knows the answer to.
RRecut. Split the score by whatever line divides the data.
Blended, the eval read 94 percent. Split by whether a bullet's job category showed up anywhere in the original 24 examples, the categories that were represented scored fine, and the ones that weren't, warehouse, home health aide, dispatch, HVAC, dropped far below that every time anyone actually checked by hand.
A flat, healthy blended score can be hiding a category that was never really being tested at all.
AAssume nothing. Rule out the ruler before you blame the prompt's writing.
First check: was the rubric itself broken, confusing instructions, an inconsistent grader. It wasn't. Outside graders applied it the same way Zorana's team did. The ruler was fine. What it was allowed to measure, the same 24 fixed rows, forever, was the actual problem.
A broken ruler and a rigged test look identical from the dashboard. Check which one you actually have before you trust either number.
CCause candidates. Name the short list, not everything possible.
Three named suspects: the prompt genuinely doesn't generalize past the shape of its own 24 examples, the model is quietly reciting patterns close to those examples rather than reasoning fresh, or the graders scoring the eval are the same people who wrote the 24 rows, and score generously from memory.
Naming who or what owns each candidate is what makes the next step possible instead of a guess.
Hand sketched panel titled Bulletforge eval three suspects for the 94 percent. Three boxes side by side. Prompt never saw this job, marked confirmed, warehouse and home aide roles missing from the 24 examples. Model reciting examples, marked ruled out, wording checked against the 24 rows. Graders remembered answers, marked contributing, same two people wrote and graded it.
The prompt itself never had a chance to be right on a job it had never seen written down.
EEvidence test. The one check that tells the causes apart.
Build a fresh set, chosen specifically to include categories missing from the original 24, and grade it with someone who never wrote or looked at those 24 rows. Score: 61 percent, a 33 point gap from the reported 94. The gap sits almost entirely in the categories the prompt never saw an example of, which points straight at the prompt not generalizing, not at reciting or grader memory being the main driver.
The cheapest, strongest check in the whole method, and the one most teams skip because the fix already feels obvious before they've run it.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was tightening the prompt further using more examples pulled from the same file the eval already leaned on, the instinct half the team had when the real world numbers first looked soft. It lost because every new example just tightens the loop between what the prompt is taught and what it's graded against, making the contamination worse, not better, like studying an exam using an answer key that's already leaked. The AI specific failure worth naming by name is eval contamination, an overlap between the examples used to write a prompt and the examples used to score it, so the score can't fail even when the prompt would. The guardrail is two part: the held out set lives in a different file with a different owner, access logged, and a scheduled blind audit runs on a rotating fresh sample every quarter, not only when someone gets suspicious. That guardrail isn't free. Building a genuinely fresh, stratified set and paying an outside grader to score it by hand costs a few hours and a real person's time, against nearly nothing for reusing what's already labeled, a real trade accepted on purpose, because a slow, honest number beats a fast, fake one. And the bar it enforces was never a promise that every rewrite is perfect. It's a probability bar, checked against the held out set: the prompt passes when it clears 80 percent on a rotating fresh sample, not a claim that the top result is always right.

And if you want to be sure it really works, try it somewhere else

Same five letters, a different industry, and this time the prompt is innocent. The graders are the actual cause.

Pawscript is a tablet tool at Thistlemere Animal Clinic. After a visit, it drafts the after visit care instructions a pet owner takes home, feeding, medicine, what to watch for. Godric Muvunyi runs quality on it.

Pawscript's prompt was written using 20 example visit write ups, all common, routine cases: ear infections, vaccinations, dental cleanings. The same two vet techs who wrote those 20 examples also graded the weekly eval, since they were the ones who knew what a good discharge note looked like. It read 96 percent for months.

Godric ran the evidence test anyway. He took the exact same 40 drafts, half from routine visits like the original 20, half from rarer cases the original 20 never covered, limping and joint pain, chronic kidney diet changes, and had a vet tech from a different clinic, who had never seen the original 20, grade every one blind. The two techs who wrote the originals gave the drafts a 96. The outside tech, grading the identical drafts blind, gave them a 71. The gap showed up almost evenly across routine and rare cases alike, not concentrated in the unfamiliar ones.

The decision Godric would take back The same two techs who wrote Pawscript's original 20 examples were also the only two people who ever graded the eval, every single week. Nobody outside that pair had checked Pawscript's real quality since the day it launched.
Hand sketched panel titled Pawscript eval three suspects for the 96 percent. Three boxes side by side. Prompt did not generalize, marked ruled out, blind score on new cases held near 92 percent. Model invented details, marked ruled out, matched the real visit record every time. Graders scored from memory, marked confirmed, same two techs wrote and graded every week.
The drafts were mostly fine. The two people scoring them were remembering their own examples, not reading the screen.

Same method, opposite result: because the gap between the internal graders and the outside grader showed up evenly across routine and rare cases, that ruled out a prompt that only worked on familiar visits. It pointed straight at who was holding the pen, not what the prompt had or hadn't seen written down before. Bring in a grader who never touched the original examples before assuming the prompt needs work.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the one test, regrade the same output with someone who never wrote or graded it before, and see if the score holds.
Cost: there's no budget to build a whole new held out set from scratch this quarter. At minimum, swap who grades the existing eval to someone who never wrote it, that alone can surface rater contamination for free.
The model got better, for real: say the underlying model gets upgraded to a newer version. That's not proof the eval is clean. A better model can still be graded generously by contaminated raters, or can still overfit to a stale example set. Rerun the fresh set test before trusting the number again.

Where people run it wrong.
They read a stable, unmoving score as proof nothing is wrong, instead of asking why it never changes.
They fix a weak real world number by feeding the prompt more examples pulled from the same file the eval already leans on, tightening the loop instead of loosening it.
They check what the eval is testing and never check who is doing the grading.

How to use it live. Say the two part check out loud before guessing at a fix: "before I answer that, can I ask whether the eval set was built before or after the prompt, and by the same person who wrote it?" That buys a beat to think instead of guessing which cause is real in front of the interviewer.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
TRACE: rule out, then narrow. Built for diagnosis and critique questions, when a number looks too good and you need to find exactly which cause is inflating it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zorana Vaduva, data scientist who owns evals for Bulletforge, Ferngrove's resume bullet rewriter. Spent three years building eval sets before joining Ferngrove.
3 · THE HABIT
What did Zorana stop doing because the eval score looked fine?
Tap to flip
ANSWER
She stopped opening the underlying spreadsheet to spot check rows by hand, then stopped sampling real user output at all, and just watched the weekly number.
4 · THE TWO SUSPECTS
What two stage level suspects is this answer choosing between?
Tap to flip
ANSWER
A prompt that never learned to handle job categories outside its 24 examples, versus graders scoring generously because they wrote the examples themselves.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Reusing the same 24 row file for both writing the prompt and grading it, decided three days before launch with a plan to build a real eval set later that never happened.
6 · THE NUMBER
Fill in the blank: the eval score on the original 24 examples held at ___ percent for eight weeks, while a fresh, never seen set of 150 real bullets scored only ___ percent.
Tap to flip
ANSWER
94 percent, then 61 percent. The 33 point gap between those two numbers is what proved the eval had been grading memory, not skill.
7 · THE REPLAY
Same eight weeks, new design, what changes?
Tap to flip
ANSWER
The held out set gets built and sealed in week one instead of week eight. The 61 percent shows up on day one, and the prompt gets fixed for warehouse and home health aide language before any of those users see a weak rewrite that reads like a good one.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and which cause turns out to be the real one this time?
Tap to flip
ANSWER
Pawscript, at Thistlemere Animal Clinic. This time the prompt is fine. The confirmed cause is rater contamination, the same two vet techs who wrote the examples also graded them every week.

Check yourself Score: 0 / 0

True or false
1. True or false: because Bulletforge's eval score stayed near 94 percent for eight weeks straight, that's good evidence the prompt was solid and unlikely to be hiding a problem.
  • True
  • False
Show hint
Think about what those 24 rows were also used for, back in week one.
Show answer
False. A score that can't move because it's the same 24 rows every week isn't evidence of anything. It never got asked a question it didn't already know the answer to.
Multiple choice
2. What's the one check that actually tells you whether an eval score is inflated by contamination between the prompt's examples and the eval set?
  • A. Add more examples to the eval, pulled from the same file used to write the prompt.
  • B. Regrade against a fresh, held out set that nobody used while writing the prompt, ideally graded by someone who didn't write the original examples.
  • C. Lower the model's temperature until the output looks more consistent.
  • D. Rewrite the rubric's instructions to be more detailed.
Show hint
Think about what Zorana did the same afternoon the new analyst asked their question.
Show answer
B. A fresh set, graded by someone with no memory of the original examples, is the only way to tell whether the prompt really works or just remembers.
Fill in the blank
3. Zorana's team built Bulletforge's prompt using ___ example resume bullets, and the eval, graded on those same rows, read 94 percent for ___ weeks straight before anyone checked it against fresh examples.
Show hint
Both numbers are named in the direct answer's story and repeated on the evidence test chart.
Show answer
24 examples, eight weeks. The same 24 rows fed both the prompt and the eval for two full months before a fresh set was ever built.
Short answer, name the rejected alternative
4. What did half the team want to do first when the real world keep rate first looked soft, and why would it have made the problem worse, not better?
Show hint
Look at what the rejected alternative in the TRACE recap says the instinct was, and what it assumes by default.
Show answer
Model answer: Add more examples to the prompt, pulled from the same file the eval already leaned on. It would have made things worse because every new example tightens the loop between what the prompt is taught and what it's graded against, which is exactly the contamination the eval already had.
Short answer, apply it yourself
5. Pick an AI feature you've used yourself. What's one sign its eval might have been built from the same handful of examples used to write its prompt, and how would you check?
Show hint
Think about a feature that works well on the obvious, common case but gets noticeably worse the moment your situation is a little unusual.
Show answer
Model answer: A recipe app's "convert to metric" feature might work perfectly on common ingredients like flour and sugar, but quietly fumble odd units like a "knob" of butter or a "handful" of herbs, because those were never in whatever small example set someone used to write the conversion prompt. I'd check by pulling a fresh batch of real recipes with unusual units, ones the writer never saw, and grading those apart from the common cases.
Multiple choice
6. Suppose Bulletforge's fresh, held out set had scored 90 percent instead of 61. Would that still be strong proof the original eval was contaminated?
  • A. Yes, any gap at all between the two scores proves contamination.
  • B. No, a small gap close to normal scoring noise wouldn't be strong evidence. The size of the gap is what makes it convincing.
  • C. Yes, because 90 is still a little below 94.
  • D. No, because contamination can only be shown with a line chart, not a bar chart.
Show hint
Compare a 4 point gap to the actual 33 point gap the evidence test found.
Show answer
B. A 4 point gap could just be normal noise between two runs. The 33 point gap Zorana actually found is what makes the case, because it's far too large to explain any other way.
Before you close the answer
Why this works
Tests whether you'll trust a score that never moves, or go looking for what it was actually allowed to test. Most candidates critique the rubric's wording and miss the actual leak underneath it.
Follow-up traps
"Isn't reusing your best examples in the eval normal early on, everyone does that?" Response: it's a fine sanity check in week one, when there's no labeled data yet. The mistake is trusting that same number for eight weeks instead of building a fresh, sealed set the moment the prompt starts shipping to real users.

"What if the fresh set also overlaps a little with the model's general training data?" Response: that's a different, harder problem, general pretraining exposure, not prompt example leakage. It's why the evidence test compares scores across job categories the prompt's own examples never touched, not just whether the words are technically new.
If pressed
The held out set gets rebuilt from a rotating pool every quarter, and the current file is access logged, so Zorana herself can see if anyone opened it while the prompt was mid tuning. That log is what keeps the held out set honest over time, not just on the day it was first built.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more