CaseIntermediateEval-Driven Specification / Writing an eval spec / #3
How do you define a task-level success criterion for a summarization feature?
The direct answer
Write the criterion around two checks, not one. First, can the editor decide from the summary alone, with the source paper closed. Second, does the summary keep the one fact that would actually change that decision, the caveat, the limitation, the qualifier, checked against the full paper on a real sample every week. A summary that hits every required field and still drops that one fact should fail, no matter how clean it reads.
Do this, in order
Write the criterion as two checks: can the editor act on it, and does it keep the fact that would change their mind.Why: a summary that only proves it covered the topics has proven nothing about whether it caught what mattered.
Audit that second check on a real sample every week, against the full paper, not against the rubric.Why: the field-check can hold steady for months while this number drifts underneath it, unwatched.
Treat the "must include these fields" rubric as the floor, not the finish line.Why: hitting every field is necessary and proves nothing about whether the one fact that mattered got through.
Set a real floor that stops editors from acting on the summary alone below it.Why: a threshold nobody has to answer to is a number on a slide, not a gate.
Leave the purely descriptive fields, funding source, journal section, off this whole audit.Why: nobody disagrees about facts like that, so spend the checking hours on the papers with a real judgment call in them.
How to answer this, stage by stage
Seven moves. Name the two checks before the story, or the answer sounds like a definition instead of something you'd actually write into a spec.
1
Scope it to one tool and one editor
Say it like this
"Let's make this concrete. Say we build Precis, a tool that reads a submitted manuscript and drafts a plain abstract for the handling editor, so they can decide whether to send it to peer review without reading all forty pages first. Thao Bui runs product for Precis, and this success criteria section is hers to write."
Why this works
A quality question stays abstract until one specific document has to survive someone reading it literally.
2
Name the plan before the reasoning
Say it like this
"I'll use LEAD here. L is the real outcome a good criterion protects, E is the early signal, A is how a summarization criterion gets gamed, and D is what changes at different accuracy levels."
Why this works
Stating the plan up front tells the interviewer you're about to make a decision, not describe a feeling about summaries.
3
Say what the question is actually checking
Say it like this
"This sounds like a question about summary quality, does it read well, does it cover the topics. It's really asking whether the editor can trust the summary enough to never open the source, or whether they're quietly still doing the reading the tool was supposed to save them from."
Why this works
This is where the answer stops being about writing style and starts being about whether the summary can carry a real decision.
4
State the decision straight
Say it like this
"So here's what goes in that section. Two checks. Can the editor make the send-to-review call from the summary alone. And does the summary keep the one fact that would flip that call, the caveat, the limitation, checked against the full paper on a sample every week. Hit every required field and still drop that one fact, and it fails, no matter how clean it reads."
Why this works
This names an actual line you could paste into a spec, not a value like "accurate" with nothing under it.
5
Prove it with the near miss
Say it like this
"Here's why that matters. Precis's field-check, does it name the method, the sample size, the result, held at 97 percent for two straight months. Underneath it, the number that actually tracks whether the one deciding fact survived fell from 89 percent to 61 percent over eight weeks. A coastal-erosion paper's summary said the new model 'outperformed the baseline across the test sites.' True. It just left out that the model failed at two of the twelve sites, the detail sitting on page twenty-six. The handling editor almost fast-tracked it before he opened the full paper anyway, out of nothing but old habit."
Why this works
A real number that contradicts the number everyone was watching does more work than a paragraph explaining the risk.
6
Set the thresholds and name the gaming path
Say it like this
"Above about 85 percent decision-fact capture, I'd trust the summary and let the editor skip the source. Between 70 and 85, I'd flag those weeks and have a second reader check the summary's caveat before the editor decides. Under 70, I wouldn't let editors act on the summary alone at all, because a wrong summary said with confidence is worse than no summary. And nobody has to cut a corner for this to happen. A model rewarded for hitting every field learns the safest way to do that on a complicated paper is to summarize the average and quietly drop the exception."
Why this works
This closes on real thresholds and a real action, so the answer reads as a spec you'd actually ship, not a definition you memorized.
7
Close on the one line
Say it like this
"So the one thing I'd actually do: build the success criteria around whether the summary keeps the fact that changes the decision, not whether it covers the required fields. A rubric that only checks presence will always look healthy, right up until the week it isn't."
Why this works
It restates the decision in one breath, so it's the last thing the interviewer hears, not one option among several.
If you only get through two stages
Stages 4 and 6 are the answer. Say the two checks, then say why a field-only rubric lets a summary pass while missing the one thing that mattered. Everything else here is how you defend that when someone pushes back.
Let's learn
Before we built anything, Rutger read every manuscript that landed on his desk, start to finish, before he decided whether it earned a reviewer's time. Forty pages, most mornings, before his coffee went cold.
Knowledge spark: what's a desk decision?
An editor's early call on a new paper: send it out for review, or say no before it ever reaches a reviewer. Made straight from the manuscript, no outside help.
It took him close to fifty minutes a paper. On a heavy Tuesday, with nine new submissions waiting, that's most of his morning gone before he answers a single email.
Then we built Precis. It reads the manuscript and drafts a plain abstract, the research question, the method, the sample size, the headline result, in about ninety seconds. Rutger could read that instead of the paper and make the same call in under five minutes.
For a long stretch, that trade worked. The check that ran on every abstract, does it name the required fields, held steady, month after month, around 97 percent. Nobody had a reason to look twice.
Knowledge spark: what's a decision-fact?
The one detail in a paper that would actually change whether an editor sends it to review. Not any true sentence. A caveat, a small sample, a result that only held for part of the data.
Here's the turn. Hitting every field is not the same as keeping the fact that would actually change Rutger's mind. And the number that tracks that, whether the abstract kept the paper's one real caveat, was falling the entire time nobody was watching it. It dropped from 89 percent to 61 percent over eight weeks, while the field-check stayed at 97 the whole way down.
Decision-fact capture, the audited abstracts by week, against the field-check that never moved
85% and up, trust the summary alone
70 to 85%, the rubric is drifting
under 70%, a confident wrong summary
Decision-fact capture crossed the 70 percent floor in week six, two weeks before a manuscript nearly got fast-tracked on page twenty-six's missing caveat. The field-check that everyone was actually watching sat at 97 percent every single week of the drop.
A summary can hit every line in the rubric and still hand the editor a decision built on the one sentence the paper actually needed him to read.
At its worst, that gap sends a flawed paper to reviewers as if it were clean, or tells an editor a paper is weaker than it is, because the one thing that would have argued for it never made it into the page he actually reads. Either way, the journal is trusting a summary that passed a checklist and missed the point.
Papers sent to the wrong review track in a two-month window, before and after the decision-fact check shipped
7
Before, field-check only, gap found by luck
1
After, decision-fact tracked from week one
Seven papers went to the wrong track, fast-tracked when they needed full review, or the reverse, in the two months the drift went unwatched. One did in the two months after decision-fact capture rode on the same weekly report as the field-check.
Every field ticked. The one fact that mattered kept walking.
The choice I'd take back
We wrote the success criteria around whether the abstract covered the required fields, method, sample size, result, because that's the part a script can check without a person reading anything. We never wrote in a check for whether the model kept the paper's own caveat, because that felt like something only a careful reader would catch. That was fine while most submissions were straightforward. It stopped being fine once more submissions were multi-site studies, results that only held for one subgroup, and the field-check kept passing anyway.
What I'd leave alone. A note about who funded the study, or which section of the journal a paper belongs in, doesn't need this kind of check. Nobody argues about facts like that. Save the checking for the parts where the summary has to make a judgment call about what mattered most in forty pages.
The lesson. A rubric that only checks whether something got mentioned will always look healthy, right up until the week it isn't. If a criterion can be satisfied by hitting a template, it was never really testing whether the summary did its job.
Now here is the same thing as a story
Use this version when you've got the time to sit with it. The short version is above. This is for when you want to feel why it mattered.
Rutger Van Dijk has handled submissions for the Journal of Coastal Hydrology for eleven years. Ask him and he'll tell you the real skill was never reading fast. It was knowing, by about page six, whether a paper had a real question in it.
For years, that meant every submission got the same forty minutes. He'd read the introduction slowly, skim the method, and slow down again for the results and discussion, because that's where a paper usually admits what it didn't manage to do.
Then Precis arrived. Thao Bui had built it to draft a plain abstract from every submission the moment it landed, the question, the method, the sample size, the headline result, so editors like Rutger could triage without reading each paper cover to cover.
For the first several months, it earned his trust honestly. He'd read the Precis summary, skim the discussion section himself just to check the tone matched, and move on. Fifteen minutes, not fifty.
Then the checking faded, in three small steps, none of them looking wrong at the time. First, the backlog grew, and skimming the discussion section on top of the summary started to feel like doing the job twice. Second, the field-check that ran automatically on every summary, does it name the method, the sample size, the result, never once failed, so there was no obvious reason to keep double-checking behind it. Third, nobody was tracking whether the summary's one caveat matched the paper's real one, only whether the summary existed at all.
Rutger stopped skimming the discussion section sometime that spring. He couldn't tell you the week. It just stopped feeling necessary.
Then a manuscript came in modeling coastal erosion across twelve monitoring sites, testing a new prediction method against the standard one everyone in the field already used.
Precis's summary said the new model "outperformed the baseline method across the test sites," and named the method, the sample size, twelve sites, forty months of data, and the headline result. Every required field, present and correct. Rutger read it in under two minutes and reached for the rapid-review track, the fast lane for papers whose findings were clean enough not to need the usual three reviewers.
Then, out of nothing but habit from years before Precis existed, he clicked through to page twenty-six anyway, not because anything looked wrong.
The new model had failed at two of the twelve sites. Badly. Not a footnote, a paragraph, the authors' own honest account of where their method broke down. It was true that the model outperformed the baseline overall. It was also true that a sixth of the sites told a different story, and that difference was exactly the kind of thing three reviewers needed to weigh in on, not wave through on the fast track.
We didn't lose the paper's headline finding. We nearly lost the one paragraph that made the finding worth arguing about.
It would be easy to say Precis got it wrong. It didn't, not by the rubric it was scored against. Every field was there. The summary was accurate, sentence by sentence. Nobody had ever written down that "accurate" and "keeps the fact that changes the decision" were two different bars, so the tool had never been asked to clear the second one.
So here's the decision Thao would take back.
Two years earlier, in the meeting where the success criteria first got written down, the team scored abstracts on whether a script could confirm the required fields were present, because that was the check a script could actually run without a person reading anything. Someone asked whether they should also check for the paper's stated caveats. It felt like something a good editor would just catch anyway. Thao remembers agreeing they'd add it if it ever became a real problem.
I'd write the decision-fact check in from day one instead. Same weekly sample, but every abstract gets compared against the paper's own stated limitation, by a person, and that agreement number rides on the same weekly report as the field-check.
Here's the replay. Same eight weeks, decision-fact capture tracked from week one. It crosses the 70 percent floor in week six, the same week it actually did, but this time somebody is looking, instead of the flag coming from Rutger's old habit on page twenty-six. Thao pulls the failing cases and finds the pattern, multi-site studies, where the model learned to summarize the average and drop the outlier. She adds one line to the prompt: always state where the result didn't hold, not just where it did. Decision-fact capture climbs back to 88 percent by week eight. In the two months after, papers sent to the wrong track drop from seven to one.
One success criteria section checks whether the abstract said the required things. The other checks whether it kept the one thing that would have changed what Rutger did next.
And the thing Thao would tell herself, back in that first meeting: the risk was never a summary that got a fact wrong. It was a summary that got every fact right and still, on the page that mattered, said nothing at all.
LEAD, run on a summary instead of a score
This reads like a question about writing style, does the summary sound right. Underneath it, it's still asking whether the editor can act on the summary and be right. That's LEAD, run on a paragraph instead of a single number.
L, link. What a task-level success criterion actually protects. Not whether the summary reads well. Whether the editor can decide from it alone and be right. If Rutger still has to open the manuscript to trust the call, the summary hasn't done its job, no matter how clean the sentences are.
E, early signal. The earliest thing worth watching isn't whether the summary covers the required fields, method, sample size, result. It's whether the summary kept the one fact that would change the editor's decision, checked against the full paper on a real sample every week. That number can fall for weeks while the field-check stays perfect.
A, abuse. How this gets gamed, usually without anyone cutting a corner on purpose. A model rewarded for hitting every required field learns to write one clean, template-shaped sentence per field, method here, sample size here, result here, and the safest way to do that on a complicated paper is to summarize the average and quietly drop the exception. The field-check always looks healthy. It was never built to notice a missing exception.
D, decision. What actually changes. Above about 85 percent decision-fact capture, trust the summary, let the editor skip the source. Between 70 and 85, that's the rubric drifting, flag those weeks for a second reader before shipping the next batch. Under 70, stop letting editors act on the summary alone, because a confident wrong summary is worse than no summary.
The check that proves the criterion is real
Pull one week's sample. Have a person read the full paper and the summary side by side, and ask one question: is there anything in the paper that would have changed the editor's call, and is it in the summary. If it usually is, the criterion is real. If it usually isn't, the rubric is checking for politeness, not judgment.
And if you want to be sure it really works, try it somewhere else
An AI tool called Riser reads a technician's full inspection report after every elevator service call and drafts a two-line summary for the building's maintenance manager, so they don't have to read six pages of readings to know whether anything needs attention. Iveta Sokolova manages maintenance for a block of towers at Corrado Lift Services' biggest client and reads a Riser summary after almost every visit instead of the full report. Same shape of problem. A summary that checks every required field, cables, brakes, doors, tested and logged, can still drop the one reading that was actually creeping toward the line.
L. Whether Iveta can decide, from the summary alone, whether a technician needs to come back before the next scheduled visit. Not whether the report got summarized.
E. Whether the summary kept any reading that was outside its normal range, even a small one, checked against the technician's full report on a sample of visits each week, not just whether the summary mentions cables, brakes, and doors.
A. Nobody has to skip a step. A model rewarded for covering every required system, cables, brakes, doors, learns to write one clean line per system, and the easiest clean line for a reading that's drifting but not yet failing is "within range," true enough to pass, vague enough to bury the trend.
D. Readings-kept score high and steady, trust the summary. Score dropping, especially on the older towers' hydraulic systems, stop trusting the two-line version and read the full report until it's fixed.
Every stop logged. The drifting reading is a separate question.
Swap the trigger and it still runs
The model gets faster. Doesn't matter. A quicker draft can still hide behind a field-check that never asks about the exception.
The review budget shrinks. Doesn't help. Fewer human spot-checks just means the gap goes unnoticed longer, not that it's any less real.
The model gets genuinely better at catching exceptions. Track decision-fact capture anyway. That's the one case it should climb on its own, and watching it is what turns "it got better" into a fact instead of a guess.
Where people run it wrong
Treating a steady field-check as proof the summary is trustworthy, with nobody checking it against the actual source.
Writing the rubric as "cover the required points" without ever naming what the one decision-changing fact usually looks like for this kind of document.
Waiting for a near miss, a fast-tracked paper, a missed elevator reading, to reveal the gap, instead of auditing decision-fact capture from week one.
How to use it live
Say the split first. "Before I answer, I want to separate two questions: does the summary cover the right topics, and does it keep the one fact that would actually change what the reader does next." That's not stalling. It's naming which question is really being asked, and it buys you the room to build the real answer instead of taking the field-check's word for it.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does each letter stand for here?
Tap to flip
ANSWER
LEAD. L is the outcome a criterion protects, an editor who can decide without the source. E is the early signal, decision-fact capture. A is how it gets gamed, hitting every field while dropping the exception. D is what changes at each capture level.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Thao Bui, who built the success criteria for Precis, a tool that drafts abstracts for a journal's handling editors, and who wrote its rubric around required fields only.
3 · THE HABIT
What did Rutger stop doing, once the field-check never failed?
Tap to flip
ANSWER
Skimming the discussion section himself after reading the Precis summary, the fifteen minutes of double-checking that used to sit between the summary and his decision.
4 · THE REAL SIGNAL
What number was falling while the field-check looked perfect?
Tap to flip
ANSWER
Decision-fact capture, whether the summary kept the paper's real caveat. It dropped from 89 percent to 61 percent over eight weeks while the field-check held at 97.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Scoring abstracts only on whether a script could confirm the required fields were present, and never checking whether the summary also kept the paper's stated limitation.
6 · THE NUMBER
Decision-fact capture was about ______ percent by week eight. Above ______ percent, the answer says you can trust the summary on its own.
Tap to flip
ANSWER
61 percent. Above 85 percent. Below 70, a wrong summary said with confidence is worse than no summary at all.
7 · THE REPLAY
Same rubric, decision-fact capture tracked from day one, what changes?
Tap to flip
ANSWER
The floor gets crossed in week six, same as before, but this time someone is watching for it instead of catching it by habit on page twenty-six. Thao adds a prompt line about stating where the result didn't hold, capture climbs to 88 percent by week eight, and papers sent to the wrong review track drop from seven to one over two months.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
Riser, which summarizes elevator inspection reports. Its early signal is whether the summary kept any reading that was drifting out of range, checked against the technician's full report.
Check yourself Score: 0 / 0
Multiple choice
1. Which of these would make the strongest task-level success criterion for Precis's abstracts?
A. The abstract uses natural, professional-sounding language.
B. The abstract stays under 200 words.
C. The abstract covers every required field, and is separately checked against the full paper for whether it kept the one fact that would change the editor's decision.
D. The abstract gets drafted in under two minutes.
Show hint
Look for the option with two separate checks, not one soft quality word.
Show answer
C. A single field-count or a style judgment can both look fine while hiding a dropped caveat. Only checking against the source catches that.
Fill in the blank
2. The field-check held around ______ percent for two straight months, while decision-fact capture actually fell from ______ percent to ______ percent over eight weeks.
Show hint
The numbers sit in the paragraph right after the turn in the first section.
Show answer
97 percent. 89 percent to 61 percent. A field-check that never moved was hiding a decision-fact number that had been falling the whole time.
True or false
3. True or false: because Precis's field-check stayed near 97 percent, its abstracts were reliably safe for the editor to act on without reading the source.
True
False
Show hint
Ask what a steady field-check can hide underneath it.
Show answer
False. A steady field-check only shows the required topics got mentioned. It doesn't show whether the one fact that would change the decision survived.
Short answer
4. Name a place in this same rubric where a decision-fact check wouldn't matter.
Show hint
Look for a field nobody would ever argue about.
Show answer
Model answer: "A field like the funding source, or which section of the journal a paper belongs in. Nobody disagrees about facts like that, so a script can check it once and move on. Save the auditing for the parts where the summary has to judge what mattered most."
Short answer, apply it yourself
5. Pick a product you use yourself where a summary or a score stands in for something longer, a review summary, a credit score, a meeting recap. What would you check to know if it kept the one fact that would actually change what you do next?
Show hint
Look for anywhere you act on the short version and never open the long one.
Show answer
Model answer: "A meeting-notes tool that turns a call into three bullets. I'd check whether the bullets kept the one disagreement or open question from the call, not just the decisions everyone already agreed on, because that's usually the fact that changes what I do next."
Multiple choice
6. If Thao had caught the decision-fact drop in week two instead of week six, what would most likely have happened to papers sent to the wrong review track over the next two months?
A. No change, still seven.
B. Fewer, because the prompt fix would have been in place before most of that period's papers were summarized under the drifting rubric.
C. More, because changing the prompt mid-quarter would have confused the editors.
D. It depends only on how many papers came in, not on decision-fact capture.
Show hint
Think about what catching the drop earlier actually buys you: time to fix the prompt before it ships broadly.
Show answer
B. Catching it earlier gives Thao time to add the missing-exception fix and rebuild capture before most of the period's papers ever ship under the drifting rubric, the same logic the replay in Section 2 shows for week six.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.