Critique criteria that specify a benchmark score without specifying the benchmark's relevance.
The direct answer
Don't accept "score at least 90 on the benchmark" as a whole criterion. A number with no stated reason is a hope, not a bar. Rewrite it to say what a 90 is supposed to predict, and check that the benchmark's own examples actually cover the reports that matter, not just the easy majority.
Do this, in order
Rewrite the criterion so a score always states what it's supposed to predict, and prove that before the number counts.Why: "score at least 90," with no stated reason, lets a team ship on a number nobody has checked means anything.
Rule out a broken benchmark before you trust or blame the score.Why: if the reference examples are inconsistent or two graders disagree, a plateau or a gap is noise, not proof either way.
Recut the score by the kind of task before trusting one blended number.Why: a 94 average can sit on top of a 96 for the easy kind and a much thinner real result for the kind that actually decides an outcome.
Compare what the benchmark tests against what real work actually looks like.Why: a benchmark built mostly from the common, easy case can't tell you anything about the rare case that matters most.
Name the exact way the score could be lying, not a vague worry about "overfitting."Why: a narrow benchmark, leaked examples, and a reward that favors the wrong thing are three different bugs with three different fixes.
Leave the easy majority alone.Why: where the benchmark and real behavior already agree, adding review weight there just slows down the reports that already work.
How to answer this, stage by stage
Seven moves. This is a diagnosis question wearing a critique's clothes, so most of the work is proving the number is measuring the wrong thing before you're allowed to rewrite it.
1
Pin the criterion to one real product and one number
Say it like this
"Let me put a number on this. Say we've got a tool, Fathom, that reads a two-hundred to three-hundred page market report and drafts a two-page brief for a consulting team going into a client meeting. A year ago, the eval lead wrote the acceptance bar as one line: the draft has to score at least 90 out of 100 on Brief-Match, a benchmark that checks it against a brief a senior consultant wrote by hand."
Why this works
A vague "critique this criterion" question invites a vague answer. Naming the actual number and the actual benchmark gives every later step something to point at.
2
Say what's actually missing from the line
Say it like this
"Here's what's wrong with it, plainly. It says what score to hit. It never says why 90, or what a 90 is supposed to prove once the brief is in front of a client. A benchmark score stands in for something real. This criterion never named the real thing it's standing in for."
Why this works
Turns "this criterion looks fine on paper" into a specific, defensible complaint an interviewer can't wave away.
3
Give the fix up front
Say it like this
"So here's what I'd do. I'd rewrite the line. Not 'score at least 90,' but 'score at least 90 on a benchmark whose own examples match the mix of reports the team actually pulls, checked every time that mix shifts.' If the benchmark is mostly the easy, common report, a 90 proves nothing about the rare one that actually changes a recommendation."
Why this works
Matches deliverable zero. Saying the actual rewrite, not just "add more testing," shows you can fix a spec, not only critique one.
4
Rule out a broken benchmark before you call it irrelevant
Say it like this
"Before I say the benchmark is measuring the wrong thing, I'd check whether it's even measuring consistently. I'd have two senior consultants re-score the same thirty drafts blind and see if they land near Brief-Match's own numbers. If they don't, the score is just noisy, and that's a different fix entirely."
Why this works
Skipping this is the easy mistake. Calling a number irrelevant before checking it's even reliable means fixing the wrong problem.
5
Show the case where a high score and a bad outcome sat side by side
Say it like this
"Here's the case that made it real. Six weeks after the score first crossed 90 and the team retired its weekly manual check, a partner presented a market-entry recommendation built on a Fathom brief that scored 94. The report was about packaging materials. One paragraph, buried on page 240, said a new rule was about to cut demand by a third. The brief never mentioned it. The client's own head of procurement caught it live."
Why this works
One real failure, with a score and a page number attached, is worth more than a general warning about "irrelevant benchmarks."
6
Name the three ways a line like this goes wrong, and the one check that tells them apart
Say it like this
"Three things could be going on. One, Brief-Match's 120 reference reports are ninety-one percent the easy, steady kind, so it barely tests the rare report where the real finding hides. Two, forty of those same reports were also used to write Fathom's own prompt examples, so the model's basically seen the answers. Three, the score rewards matching a lot of the reference brief's words, not catching the one fact that actually matters. The check that separates these: compare the mix of report types inside Brief-Match's own 120 examples against the mix in a real quarter's traffic. Ours came back seventy-eight percent steady, twenty-two percent shock. Brief-Match's own set was ninety-one and nine."
Why this works
Naming three checkable causes, then one test that actually separates them, is the strongest move in the framework. A vague "the benchmark isn't representative" doesn't survive a follow-up question.
7
Say what you'd leave alone, then close
Say it like this
"I wouldn't touch the bar on the easy, steady reports. There, the benchmark and real behavior already agree, associates barely touch those drafts, twelve percent get a rewrite. So: keep the bar, but make it earn the number. A score with no stated reason and no check on its own coverage isn't a bar. It's a hope with two decimal places."
Why this works
Ending on what stays, not just what changes, shows judgment instead of blanket distrust of every benchmark in the document.
If you remember one thing
Stage 4 and stage 6 are what's being graded. Rule out your own measurement first. Then compare what the benchmark actually tests against what real work actually is. A score with no stated target will always look fine, right up until the report that actually mattered.
Let's learn
Every quarter, an engagement team at Thornwell Partners used to read a market report cover to cover before it ever reached a client. A packaging report, a payments report, a hundred and fifty to three hundred pages. Seven hours a report, one associate, start to finish, before a two-page brief went out the door.
Fathom changed that. Feed it the report, and it drafts the two-page brief in about three minutes.
A year ago, Ishaan Bhatt, who owns the eval process for Fathom, wrote the line that decides when a new version is allowed to ship: the draft has to score at least 90 out of 100 on Brief-Match, a benchmark that checks a draft against a brief a senior consultant wrote by hand for the same report. For eleven weeks after that, the team tuned Fathom's prompt against Brief-Match. The score climbed steadily: 68 the first week, past 90 by week nine, 95 by week eleven.
Two lines, eleven weeks of tuning
What the dashboard showed: Brief-Match score
Quiet signal: share of briefs an associate substantially rewrites before a client meeting
wk 1wk 2wk 3wk 4wk 5wk 6wk 7wk 8wk 9wk 10wk 11
Brief-Match climbed twenty-seven points in eleven weeks. The share of briefs an associate still substantially rewrote before a client meeting never left a four-point band. A score that measures the job should move that number. This one didn't.
Here is the turn. Ninety-five out of a hundred sounds like the number worked. It didn't. The problem was never that the score was too low. The problem is what a 90 was ever supposed to prove, and nobody had written that part down.
We never wrote down what a 90 was supposed to prove. We just wrote down 90.
Knowledge spark: what makes a benchmark relevant
A benchmark is relevant to a real job only if the mix of examples inside it looks like the mix of real work it stands in for. A benchmark built mostly from the easy case will always look healthy, because it's mostly testing the part that was never in danger.
The cost showed up six weeks after the score first crossed 90, once the team had retired the manual weekly check that used to catch a bad brief before a partner ever saw it. A partner presented a market-entry recommendation to a client, built on a brief that scored 94. The report covered a packaging-materials market. One paragraph, on page 240, said a new rule was about to cut addressable demand by a third. The brief never surfaced it. The client's own head of procurement, who already knew about the rule, caught it live, in the room.
The gap between the mark that looked like success and the mark that wasn't
The choice I would take back. A year ago, writing "score at least 90" as the whole bar was the sensible call. Brief-Match was new, the score was 68, and there was plenty of obvious room to improve before anyone needed to ask what the number meant. I would take it back and add one line the day the criterion was written: the benchmark's own examples have to match the real mix of report types, checked and re-checked, not assumed once and left alone.
What I would write instead
Score at least 90 on Brief-Match, and Brief-Match's own 120 examples have to match this quarter's real report mix within a few points, checked every time the traffic shifts. A score with no coverage check isn't a bar. It's a claim nobody tested.
What I would leave alone. The easy majority doesn't need touching. On the steady, common kind of report, seventy-eight percent of everything the team pulls, Fathom scores 96 on Brief-Match and associates rewrite only about 12 percent of those drafts. The benchmark and real life agree there. Adding a second layer of review to the reports that already work would just slow down the four in five reports Fathom already handles well.
The lesson. A score that climbs is not proof of anything by itself. I wrote a bar with a number on it and no reason attached. A number with no reason is just a hope wearing a decimal point.
The six weeks after nobody was watching
Ishaan Bhatt can read an eval dashboard and tell you, before the second chart even loads, whether a climbing score is a real fix or a lucky sprint. He built Brief-Match himself, a year before any of this, working with two senior consultants to pull 120 archived reports from Thornwell's library and write a reference brief for each one. Then he wrote the bar: 90 or better, or the update doesn't ship.
For most of that year, the number felt like the right thing to watch. It climbed slowly, from the high 60s into the 80s, one deliberate sprint at a time, and every point felt earned. Somewhere in there, without anyone deciding it on purpose, the team stopped running its old weekly habit: a manual spot check where a senior associate read three fresh briefs against their source reports before Friday. Once Brief-Match was closing in on 90, that check started to feel like double work. Ishaan signed off on retiring it the same week the score first cleared the bar.
Six weeks after that, a partner walked into a client meeting with a market-entry recommendation built on a Fathom brief that had scored 94. The client's own head of procurement stopped the meeting to ask where the brief's "steady demand" line had come from, because her own team already knew about a rule change that was about to shrink that exact market. The partner had nothing to point to. Thornwell sent a correction two days later. The client kept the account, but asked, plainly, whether every brief they'd been handed that quarter needed a second look.
It did. Two associates spent the better part of a week, about seventy hours between them, rereading every Fathom brief from the previous quarter against its source report. Those were the exact hours Fathom was bought to remove.
What the mix actually was
Before blaming the benchmark's scope, Ishaan's team checked whether Brief-Match was even measuring consistently. Two senior consultants blind-rescored thirty drafts and landed within a few points of Brief-Match's own numbers on all but two of them. The scoring itself wasn't broken. That left the coverage.
Score by report type, next to what associates actually did about it
96%
12%
92%
74%
Steady-state reports 78% of last quarter's 340 reports
Shock-type reports 22% of last quarter's 340 reports
Brief-Match score
Share an associate substantially rewrites
Brief-Match barely tells the two kinds of report apart, 96 against 92. Real behavior tells them apart completely: a 12 percent rewrite rate on the common report, a 74 percent rewrite rate on the rare one where the finding is a reversal. A blended average of 94 hides both numbers.
The 74 percent rewrite rate on shock-type reports is also why the failure was rare rather than routine. Most of the time, an associate still caught the miss and fixed it before a client saw it. The partner's report was one of the roughly one in four shock-type briefs where nobody did, the exact odds a criterion built on a blended 94 will never show you.
What Brief-Match tests, next to what the quarter actually was
91%
9%
Brief-Match's own 120 reference examples
78%
22%
This quarter's 340 real reports
Steady-state reports
Shock-type reports
This is the evidence test: compare what the benchmark actually tests against what real traffic actually is. Brief-Match under-tests shock-type reports by more than half, 9 percent against a real 22 percent, the exact kind of report where a wrong take is worth the most and gets caught the least.
Three ways a 90 stops meaning much
Not because anyone cut a corner on purpose. A criterion that names a score and nothing else can be met honestly, sprint after sprint, and still stop meaning anything.
Three separate, checkable ways to hit a score without meeting its point
Way 1
Narrow slice. The benchmark barely tests the report that matters most.
Brief-Match's 120 reference reports were pulled from whatever was easiest to find good hand-written briefs for, which turned out to be the steady, common kind. Only 11 of the 120 were the shock type, where the real finding sits in one buried paragraph.
How you'd check it: pull the report-type mix inside the benchmark's own example set and compare it to real traffic. Ninety-one against seventy-eight is a real gap.
Way 2
Leaked examples. A third of the test also trained the answer.
Forty of Brief-Match's 120 reference reports had also been pulled, months earlier, into Fathom's prompt library as few-shot examples for the model. On those forty specifically, the model wasn't summarizing a new report. It was closer to reciting one it had already seen graded.
How you'd check it: cross-reference the benchmark's example list against the prompt library's example list. Score those forty separately from the other eighty and see if the gap closes.
Way 3
Wrong reward. The score checks for shared words, not the decisive fact.
Brief-Match scores overlap with the reference brief's wording and structure. It never checks whether the one sentence carrying the report's real finding made it into the draft at all. A brief can echo ninety percent of a two-hundred-page report's language and still miss the one paragraph that changes a client's answer.
How you'd check it: for a sample of high-scoring drafts, check whether the source report's single most decision-relevant sentence appears anywhere in the summary, independent of the overlap score.
Shipping on the blended score alone
94 on Brief-Match, last quarter's average across both report types
26 percent of shock-type reports where a wrong take slipped past unnoticed
Requiring the coverage check too
A criterion that also checks the benchmark's mix against real traffic
0 prompt updates have shipped since without that check passing
TRACE, run on a line that read fine on paper
This is a diagnosis question wearing an artifact critique's coat: the real question is why a criterion that reads cleanly stopped predicting real quality. GUARD would fit if the harm here were about fairness between groups; it's about a number nobody checked against the work it was supposed to stand in for.
T, timeline. The score climbed from 68 to 95 over eleven weeks. The share of briefs an associate still substantially rewrote before a client meeting stayed inside a four-point band the entire time. The two lines should have moved together. Only one of them had a dashboard, and only that one got trusted enough to retire a manual check.
R, recut. Split by report type. Steady-state: Brief-Match score 96, rewrite rate 12 percent. Shock-type: Brief-Match score 92, rewrite rate 74 percent. A blended 94 hid a real result on the report type that almost never mattered and a real result on the one that always did.
A, assume nothing. Before blaming the benchmark's coverage, Ishaan's team checked its own scoring: two senior consultants blind-rescored thirty drafts and landed close to Brief-Match's own numbers on all but two. The scoring itself was fine. The gap was real.
C, cause candidates. Three, named and separate: the reference set is ninety-one percent the easy report type, forty of its 120 examples also trained the model's own prompt, and the score rewards word overlap rather than whether the report's one decisive fact made it in.
E, evidence test. Compare the report-type mix inside Brief-Match's own 120 examples against the mix in this quarter's real 340 reports. Ninety-one against nine, versus seventy-eight against twenty-two. The benchmark under-tests the exact report type where the miss happened by more than half.
Why E is the hard step
Anyone can suspect a criterion is too narrow. A test earns its place by putting a number on the mismatch between what gets tested and what actually happens, not by restating the worry in a more confident tone. Compare a benchmark's own coverage to real traffic, and you've checked something. Call it "probably not representative" without that comparison, and you've only guessed louder.
Same shape, a QA line that grades its own calls
VoxLine runs the customer service line for a mid-size insurer's claims department. Its AI drafts a post-call summary for the supervisor who reviews it, checked against a benchmark called Call-Match: does the draft cover what a supervisor's own summary would cover for the same call. Ronan Gallagher, VoxLine's QA analytics lead, wrote the launch criterion two years ago: the draft has to score at least 88 on Call-Match before a model update ships.
T. Call-Match's score climbed from 79 to 93 over eight weeks of prompt tuning. The rate of calls a supervisor overrides or escalates after reading the draft summary never moved off roughly one in six, the whole time. R. Split by call type. Routine billing calls: Call-Match score 95, override rate 4 percent. Disputed-charge calls: Call-Match score 90, override rate 41 percent. Nearly the same score, a ten-times difference in what a supervisor actually did next. A. Before blaming coverage, Ronan has two senior supervisors re-score twenty-five drafts blind against Call-Match's own numbers. They land close together. The scoring itself checks out. C. Three candidates: Call-Match's 90 reference calls are eighty-five percent routine billing, the prompt was tuned on wording that recurs across those same routine calls, and the rubric's "notes next action" check is satisfied by any next action, not necessarily the correct one for a dispute. E. Compare Call-Match's call-type mix, eighty-five percent routine, against this quarter's real call mix, sixty-five percent routine and thirty-five percent disputed. The benchmark under-tests the calls a regulator would actually ask to review.
Swap the trigger and it still runs
Speed: instead of eleven weeks of steady tuning, the trigger is a rushed model swap approved in two days before a client renewal. TRACE still starts with what the number was ever supposed to predict, not with how fast the swap got approved.
Cost: the team shrinks Brief-Match's example set to save review time, on the idea a smaller set is close enough. The recut still has to show what kind of report got cut, not just how many.
The model really did get better: the case on this page. Fathom's tone and phrasing genuinely improved. The benchmark just never checked whether the one paragraph that mattered made it in, so a real gain and a fake one showed up as the same number.
Where people run it wrong
Treating "we hit the bar" as proof of readiness, instead of asking what the bar was ever built to predict.
Reading a climbing score as the product improving, when a fixed benchmark, tuned against long enough, can climb for reasons that have nothing to do with the real job.
Writing a criterion once and never rechecking whether the mix of real work still looks anything like the mix the benchmark tests.
If you are asked this cold
Buy yourself ten seconds by naming the gap out loud. "So there's what the benchmark tests, and there's the actual mix of work it's supposed to stand in for. A score can hold steady even while those two drift apart. Let me say how I'd check whether they still match." That's not stalling. That's where the real critique starts.
Flashcards (click a card to flip it)
This is a diagnosis question about a written criterion, not a flip story, so these eight test the TRACE moves and the real numbers instead of a habit changing.
1 · THE FRAMEWORK
Which framework fits critiquing a criterion that names a score but not the benchmark's relevance, and why?
Tap to flip
ANSWER
TRACE. The real task is a diagnosis wearing a critique's clothes: working out why a rule that reads fine on paper stopped predicting real quality. GUARD would fit a fairness harm; this is a coverage gap.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ishaan Bhatt, eval and PM lead for Fathom at Thornwell Partners. He built the Brief-Match benchmark and wrote its 90-point bar himself, a year before the gap surfaced.
3 · RULING IT OUT
What did Ishaan's team check about Brief-Match itself before blaming its coverage?
Tap to flip
ANSWER
Whether the scoring was even consistent. Two senior consultants blind-rescored thirty drafts and landed close to Brief-Match's own numbers on all but two, so the scoring itself wasn't the problem.
4 · THE THREE WAYS
Name the three ways a criterion like "score at least 90" can be met honestly and still miss the point.
Tap to flip
ANSWER
A benchmark built mostly from the easy, common report type, the same examples doing double duty as the model's own training examples, and a score that rewards matching words instead of catching the one fact that matters.
5 · THE NUMBER
Brief-Match's own 120 examples were ______ percent the easy, steady report type, against a real quarter's traffic that was only 78 percent that type.
Tap to flip
ANSWER
91 percent. Its own set was only 9 percent the rarer, harder report type that a real quarter ran at 22 percent.
6 · THE CHECK
Name the one test that turned a suspicion into a number.
Tap to flip
ANSWER
Compare the report-type mix inside Brief-Match's own 120 examples against the report-type mix in a real quarter's traffic. Ninety-one against nine, versus seventy-eight against twenty-two, is the whole gap in two numbers.
7 · THE FIX
What does the rewritten criterion actually say that the old one didn't?
Tap to flip
ANSWER
It states what the score is supposed to predict, and adds a second condition: Brief-Match's own examples have to match the real report mix, checked on a schedule, not assumed once and forgotten.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE on a different product with a similar gap. Which one, and what's the number?
Tap to flip
ANSWER
VoxLine's call-summary benchmark: 85 percent routine calls inside the benchmark's own examples, against 65 percent routine calls in real traffic, so the disputed calls that actually risk a regulator complaint were barely tested.
Check yourself Score: 0 / 0
Multiple choice
1. Ishaan's criterion says Fathom's brief must score at least 90 on Brief-Match. What's the actual problem with that line, as written?
A. Ninety is too high a bar for a summarization tool to hit reliably.
B. It never states what a 90 is supposed to predict, or whether Brief-Match's own examples match the reports that matter most.
C. Brief-Match should be replaced with a completely different scoring method.
D. The bar should be lowered until associates stop rewriting the briefs.
Show hint
Look at what the line says, and what it never says, about why 90 or what it proves.
Show answer
B. The line names a score with no stated reason and no check on whether the benchmark's own coverage matches real work. That's the actual gap, not the number itself.
Fill in the blank
2. Brief-Match's own reference set was ______ percent steady-state reports, versus 78 percent in a real quarter's traffic.
Show hint
Look at the top segment of the first bar in the composition chart.
Show answer
91 percent. Against 78 percent in real traffic, meaning the benchmark's rarer, harder report type was only 9 percent of its examples, versus a real 22 percent.
True or false
3. True or false: once the two senior consultants confirmed they agreed with Brief-Match's scoring, that proved the 90-point bar was measuring the right thing.
True
False
Show hint
Confirming the scoring agrees with itself only rules out one kind of problem.
Show answer
False. That check (the A step) only ruled out a broken measurement. It took the composition comparison (the E step) to show the benchmark was testing the wrong mix of reports.
Short answer
4. Name a place in Fathom's process where you'd leave the current Brief-Match bar exactly as it is, and say why.
Show hint
Think about the report type where the benchmark and real behavior already agree.
Show answer
Model answer: "Keep it exactly as is for steady-state reports, seventy-eight percent of real traffic. There, Brief-Match's 96 and the real 12 percent rewrite rate already agree. A second review layer there would just slow down the majority of reports that already work."
Short answer, apply it yourself
5. Think of a pass or fail bar you rely on somewhere, a code review checklist, a test suite, a certification score. What's one way that bar could hit its number while missing the real thing it's supposed to stand in for?
Show hint
Look for a case where the measured thing is common and easy, while the thing that actually matters is rare and hard.
Show answer
Model answer: "A support team's first-contact resolution rate can climb because agents quietly reclassify a hard, unresolved ticket as resolved, hitting the number without actually fixing anything for the customer." Any honest answer works if it names a real case where the measured thing and the real thing quietly came apart.
Multiple choice
6. A teammate says the real fix is simpler: just raise Brief-Match's bar from 90 to 95. Why doesn't that fix what's wrong with the criterion?
A. Because 95 is not achievable given the current model.
B. Because a higher bar on the same narrow, skewed benchmark still says nothing about the 22 percent of reports the benchmark barely tests.
C. Because associates would stop reading the dashboard entirely.
D. Because Brief-Match's scoring already disagrees with itself, so no bar on it means anything.
Show hint
One of these treats the number as the problem. The rest of the answer says the coverage is the problem.
Show answer
B. Raising the number is a dial turned up on the same narrow benchmark. It does nothing about the fact that the benchmark barely tests the report type where the real failure happened.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.