ConceptIntermediateEval-Driven Specification / Writing an eval spec / #22

What is the minimum eval you would accept before a first internal release?

The direct answer
Before a first internal release, size the eval by the real resume shapes the parser will actually meet, not by whatever resumes are already sitting in a folder. Test a floor of resumes in each real format, roughly ninety total across six shapes for a typical ATS mix, hold an aggregate accuracy bar plus a separate floor per shape, and check that total against how long a person actually takes to grade it by hand. Push the number down toward forty five for a closed, low-stakes alpha, and up toward a hundred fifty or more once real hiring decisions start riding on it.
Do this, in order
  1. Size the eval by real resume-format coverage, not by what's already sitting in a folder.Why: a number with no categories behind it can't tell you what it never tested.
  2. Break the applicant pool into its real shape categories before picking a total.Why: single-column, two-column, DOCX, scanned, non-English, and odd layout each fail differently, and one percentage hides all of that.
  3. Set a floor per category, not just an aggregate bar.Why: an average can look great while one shape is badly broken underneath it.
  4. Move the whole number with what the release actually risks, not with the deadline.Why: a closed alpha and a release that feeds real candidate records don't deserve the same minimum.
  5. Check the minimum against how long a person actually takes to grade it by hand.Why: a number nobody can review before launch is a number that gets skipped under pressure.
  6. Name which assumption would break the number first: a missing format, not reviewer time.Why: that's the one worth stress-testing before anyone trusts the number.

How to answer this, stage by stage

Nobody is grading whether you land on exactly ninety resumes. They are grading whether the number has real coverage under it, moves with the stakes, and survives a check against how long it takes a human to actually run. Seven moves get you there.

1
Scope the release, and name the real resume shapes before naming a number
Say it like this
"Before I give you a number, let's agree what 'minimum' means here. We're not asking whether the parser is perfect, we're asking what has to be true before recruiters even see it. And a resume isn't one shape. There's a clean single-column PDF, a two-column one with a sidebar, a Word doc, a scanned image, one in another language, and the odd layout someone exports from LinkedIn."
Why this works
Stops the interviewer from hearing a bare number attached to nothing concrete.
2
Say the equation out loud
Say it like this
"The minimum eval is two things added together. A floor of test resumes in every real shape the parser will meet, plus a bar it has to clear in each one. Skip either half and the number stops meaning anything."
Why this works
Shows the arithmetic before a single number lands, so what follows reads as a build-up, not a guess.
3
Own the numbers for one real applicant pool
Say it like this
"Say the pool splits roughly like this: forty five percent clean single-column PDFs, twenty percent two-column, fifteen percent Word docs, ten percent scanned, six percent non-English, four percent odd layout. I'd test thirty, twenty, fifteen, fifteen, five, and five against those six shapes, ninety resumes total, checked on the five fields a recruiter actually uses: name, most recent title, most recent employer, dates, and phone."
Why this works
Turns "ninety resumes" into a number a reviewer can check, line by line.
4
Set the bar, aggregate and per shape
Say it like this
"I wouldn't just average it. Ninety percent of fields right across the whole set, but no single shape allowed under eighty percent. Otherwise one badly broken format hides behind a good overall score."
Why this works
This is the part that actually catches trouble later. An average can look fine while one shape is quietly failing every time.
5
Stretch the range to the release's real stakes
Say it like this
"That ninety isn't fixed. If this is a closed alpha, three recruiters kicking the tires, nobody's hiring decision riding on it yet, I'd drop it to about forty five. If this release feeds real candidate records recruiters act on, reject, shortlist, I'd push it toward a hundred fifty or higher before I'd sign off."
Why this works
Shows the number is tied to what's actually at stake, not to how close the deadline is.
6
Sanity check it against how long a person takes to grade it
Say it like this
"Ninety resumes, checking five fields against the real PDF, that's about two minutes each for someone who knows what to look for. Three hours. One working day. Recruiters used to fill these fields by hand at ninety seconds a resume, so grading the eval is barely slower than the job it's replacing. That tells me the number is sized for a person, not a research team."
Why this works
This is the step most estimates skip, the one that turns a plausible-sounding number into one that survives a follow-up question.
7
Name which assumption moves it most, then close
Say it like this
"If I had to bet on what breaks this number, it isn't reviewer time. You can always spread grading over two days. It's finding out there's a seventh shape nobody counted, one that was zero in the eval and isn't zero in the real world. So: ninety resumes, six real shapes, ninety percent overall, eighty percent floor per shape, checked in about three hours, moved up or down by what's actually riding on the release, not by the calendar."
Why this works
Answers the hardest part of the question directly and closes in one breath, the way a strong answer actually sounds.
If you remember one thing Ninety resumes across six shapes is a starting point, not a rule. What actually sets the number is what a miss costs, checked against how long a human takes to grade it.

Let's learn

Resume Fill is a feature inside Slatework, an applicant tracking system, that reads an uploaded resume and fills in a candidate's name, most recent title, most recent employer, dates, and phone number, so a recruiter doesn't type them in by hand.

Knowledge spark: what is an eval? A structured test you run on a model before you trust it with something real. Not a vibe check. A set of real examples, each with a right answer already known, scored on purpose before anyone else sees the output.

Before Resume Fill, a recruiter typed those fields by hand from every resume, about ninety seconds a candidate. A team working through two hundred applications a week spent about five hours a week on nothing but that typing.

Now Resume Fill does it in about three seconds, and on a clean, standard resume it is close to perfect.

The mistakes that show up are not spread evenly across every resume. They cluster in the shapes nobody tested on purpose, and the shape nobody notices for weeks is the one that reads calm, tidy, and wrong.

At its worst, a two-column resume gets its most recent job silently swapped with an older one pulled from a sidebar summary, the wrong years of experience get logged, and a strong candidate quietly drops out of a shortlist search, because the recruiter trusts the filled fields and never opens the source PDF again to check.

The decision that mattered Set the eval's coverage by the real shapes a parser will meet, not by convenience. An eval built entirely from whatever resumes are already in one folder can only prove the parser handles the resumes already in that folder.

The choice I would take back. The pre-launch eval was twenty resumes pulled from the team's own shared hiring folder, because they were fast to grab. It was never checked against what shape those twenty actually were.

What I would leave alone. A low-stakes field like a candidate's portfolio link doesn't need the same fifteen-per-shape floor as the fields that decide who gets shortlisted. A wrong or missing link there costs nobody a callback.

The lesson. A minimum eval isn't the smallest set that looks clean. It's the smallest set that has a real chance of showing you what you don't already know. Any eval built entirely from resumes you already have can only prove the parser handles the resumes you already have.

Now here is the same thing as a story

Skip this part if you already believe a first-release eval should be sized by shape coverage, not convenience. Read on if you don't.

Kaveri Nadira has run resume tooling at Slatework for two years. When something in the pipeline looked off, she was the one who pulled the actual PDF and read it herself, not just the ticket describing it.

The Thursday before Resume Fill's pilot launch, engineering had it parsing cleanly. Kaveri needed one thing before three pilot recruiters got access on Monday: proof it worked. Someone on the team grabbed twenty resumes from the shared hiring folder, ran the parser against them in a twenty-minute meeting, and got nineteen out of twenty fields right. Nobody asked what shape those twenty resumes happened to be. It didn't come up. They were all clean, single-column, English PDFs, because that was what the team's own past hires had sent in.

Two panels compared. Left, one uniform stack of 20 resumes, all the same shape, labeled the Thursday test. Right, six stacks of different heights, 30, 20, 15, 15, 5, and 5, each labeled by resume format, labeled the real minimum, 90 resumes total.
What the pre-launch test actually covered, next to what the real applicant pool looks like

For six weeks, the pilot went well. Every morning around nine, a batch of new applications would land, and by five past, Resume Fill had already filled in the name, title, employer, and dates for every one of them. At first the three recruiters checked every filled resume against the PDF before saving. After two weeks of it always matching, they checked maybe one in five. By week four, they didn't check at all. They just saved and moved to the next candidate.

Then a hiring manager on a global sales search asked Kaveri a small, ordinary question. "Why does this candidate's most recent job say 2019? Her LinkedIn says she left that job last year." No error message. No dashboard alert. Just a question Kaveri couldn't answer without opening the file.

The resume was a two-column layout, a photo and a skills sidebar on the left, the real work history on the right. On that shape, the parser kept grabbing a line from the sidebar summary and reading it as the most recent job, instead of the actual latest entry in the timeline column. Kaveri pulled three weeks of applications from that hiring manager's roles, which drew a lot of candidates using that same European-style template. On the two-column resumes, most recent employer was wrong close to forty percent of the time, and nothing on the screen ever said so.

We didn't get an eval that was twenty percent too small. We got an eval that had never once looked at a two-column resume.

Kaveri remembered the Thursday meeting clearly, twenty resumes, nineteen right, done in twenty minutes. It had felt thorough at the time. Nobody in that room had asked what those twenty resumes actually were, because asking felt like slowing down a release that was already running on time.

She rebuilt the eval that week: ninety resumes across six real shapes, weighted by how often each one actually showed up, with a ninety percent bar overall and an eighty percent floor in every shape. Run against the model as it stood, the two-column category scored sixty two percent, well under the floor. A fix to stop the parser reading the sidebar as the timeline took two days. Rerun, the category cleared eighty eight percent before the wider recruiter group got access. The next near miss, a similar sales search three weeks later, got caught inside the eval, three days before it would have reached a hiring manager the same way.

The thing I'd tell my Thursday-afternoon self: twenty resumes proving the parser works is not the same claim as twenty resumes proving nothing's broken. I had the second one. I told the room I had the first.

BOUND, sized for a Friday launch

This is a sizing question about coverage and a floor, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. A minimum eval is a floor of test resumes in every real shape the parser will meet, plus a bar each shape has to clear. Skip either half and the number stops meaning anything.
O, own the numbers. Ninety resumes total, weighted to the applicant pool's real mix: thirty single-column PDF, twenty two-column, fifteen DOCX, fifteen scanned, five non-English, five odd layout. Checked on five fields: name, most recent title, most recent employer, dates, phone.
U, use a range. About forty five for a closed alpha with no hiring decision riding on it. About a hundred fifty or more once a release feeds real candidate records recruiters act on.
N, nail the sanity check. Grading ninety resumes by hand, five fields each, at about two minutes a resume, is three hours, one working day, barely slower than the ninety-seconds-a-resume job it replaced.
D, direction. The number moves most on a missing shape, not on reviewer time. A category that was never in the eval and is real in the world is worth more than any amount of extra grading time on the shapes already covered.

The build-up: 90 resumes, six shapes
30
20
15
15
5
5
Single-column PDF Two-column DOCX Scanned / OCR Non-English Odd layout
Two-column, the shape that actually broke, is only 20 of the 90 resumes. A modest slice of the eval, an outsized share of the real cost.
What moves the minimum most
Release moves from a closed alpha to real hiring decisions+90
A resume shape shows up that wasn't one of the six+18
Locked to a closed alpha only, nothing riding on it−45
Reviewer grading time cut from a day to an afternoon+0
The two biggest swings both come from what's riding on the release or what shape got missed, not from how fast a person can grade it. Cutting the review window barely moves the number at all, it just means grading spreads over two days instead of one.

And if you want to be sure it really works, try it somewhere else

A water utility is about to let an AI read a customer's photo of their meter and log the usage number, instead of sending a technician out to read it in person.

B, break it down. The minimum eval is coverage across real photo conditions: an analog dial meter, a digital LCD meter, a meter photographed at an angle, one in low light, one behind a fogged glass cover, and one with glare across the display.
O, own the numbers. A hundred twenty test photos. Forty analog dial, thirty digital LCD, twenty angled, fifteen low light, ten fogged, five glare-heavy, weighted by how the real photo submissions actually break down. Bar: within one digit of the true reading on ninety five percent, no shape under eighty five percent.
U, use a range. Forty test photos is enough for an internal pilot among utility staff reading their own meters, low stakes, easy to catch by hand. Push to two hundred or more once the reading feeds a customer's actual bill.
N, nail the sanity check. Grading a hundred twenty photos by hand, about a minute each against the true reading on file, is two hours. A technician's truck-roll to read one meter in person already costs about twenty minutes plus drive time, so even two hours of hand-grading is far cheaper than the process it's replacing.
D, direction. For this product, the number doesn't move most on how many photo conditions exist. It moves on who pays for a wrong reading. A reading that changes a customer's actual bill needs a tighter floor than a reading that's just informational, with a person reconciling it before it ever reaches an invoice.

Which assumption moves it depends on the product At Slatework, it's a missing resume shape, a category that was never in the eval and is real in the world. On the meter-reading tool, it's who's exposed to a wrong reading, a customer's bill versus an internal estimate someone double-checks. Both are "what a miss costs," but they're counted differently, and naming which one applies is half the answer.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the number and the bar: ninety resumes, six shapes, ninety percent overall with an eighty percent floor per shape. The build-up backs it up if they ask.
Cost: instead of asking what the total should be, a manager caps how many resumes a reviewer can grade in a day, say sixty. Same rule, solved backward: allocate that sixty across the six shapes by risk and how often each one shows up, instead of picking a total first.
The model got better: a new version ships that reads two-column resumes almost perfectly. The minimum doesn't shrink on its own. Rerun the per-shape floor check first, since acing the shapes it always handled proves nothing about the one that used to fail.

Where people run it wrong.
They pick a round number, like fifty, because it sounds thorough, with no shape categories behind it.
They set the minimum once before launch and never revisit it as new resume formats or new markets show up.
They grade the eval against the parser's own confident output instead of against the real source file, so a fluent wrong answer passes.

How to use it live. Say the equation before any number: "the minimum eval isn't a count, it's a coverage floor across every shape the parser will actually meet, sized to what a miss costs." That buys the time to work out the real split instead of guessing a number that sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about the minimum eval before a first release, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question about coverage and a floor, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Kaveri Nadira, product manager at Slatework, an applicant tracking system. She owns Resume Fill, the feature that auto-fills a candidate's fields from an uploaded resume.
3 · WHAT THE TEST ACTUALLY PROVED
What did the Thursday test actually prove, and what did the room treat it as proving?
Tap to flip
ANSWER
It proved the parser could read twenty standard, single-column resumes already sitting in one folder. The room treated a 19-out-of-20 score as proof it could read any resume.
4 · THE BUILD-UP, IN THIS STORY
What's the six-shape build-up this answer turns on?
Tap to flip
ANSWER
30 single-column PDF, 20 two-column, 15 DOCX, 15 scanned, 5 non-English, 5 odd layout. Ninety resumes total, weighted by how often each shape really shows up.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Testing on twenty resumes grabbed from a shared folder instead of the real shape categories. It made sense because it was fast, and the team's own past hires all happened to look the same.
6 · THE NUMBER
Fill in the blank: the rebuilt eval's aggregate bar was set at ___ percent field accuracy, with no single shape allowed under ___ percent.
Tap to flip
ANSWER
90 percent aggregate, 80 percent per-shape floor. The two-column shape scored 62 percent against that floor before the fix, and 88 percent after.
7 · THE REPLAY
Same near miss, new eval. What changes?
Tap to flip
ANSWER
The rebuilt eval catches a similar two-column mix-up inside a test run three weeks later, three days before it would have reached a hiring manager the same way the first one did.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's different about which assumption moves the minimum there?
Tap to flip
ANSWER
A water utility's meter-photo reading tool. There, the minimum moves most on who's affected by a wrong reading, a customer's bill versus an internal estimate, not on a missing photo category.

Check yourself Score: 0 / 0

Short answer, the number question
1. If Slatework's applicant pool doubled in size overnight but the same six resume shapes still cover it, should the total minimum eval size double too? Why or why not?
Show hint
Think about what the ninety resumes are actually there to test: whether each shape works, not how many resumes arrive each week.
Show answer
No, not really. It should stay close to ninety, maybe nudged up slightly if the mix shifts toward a riskier shape. Doubling the applicant count doesn't create a new shape the parser has to handle, it just means each shape shows up more often. The eval size should track how many real shapes exist and what a miss costs, not the raw volume of applications.
Multiple choice
2. Why doesn't testing on whatever resumes are already sitting in a shared folder work as a minimum eval?
  • A. Because a shared folder can't be read by an eval script.
  • B. Because resumes in a shared folder are usually too old to matter.
  • C. Because a folder of convenience is rarely built to cover every real shape, so a high score on it only proves the parser handles whatever's already there.
  • D. Because a shared folder always has fewer than twenty resumes in it.
Show hint
Ask what a high score on the Thursday test actually proved, and what it never touched.
Show answer
C. The twenty resumes in the folder were all one shape, single-column PDF, so a good score there said nothing about two-column, scanned, or non-English resumes.
True or false
3. True or false: a low-stakes field like a candidate's portfolio link needs the same fifteen-per-shape floor as the fields that decide who gets shortlisted.
  • True
  • False
Show hint
Think about what actually changes for a candidate if a portfolio link is missed, versus if their most recent job title is wrong.
Show answer
False. A missed or wrong portfolio link costs nobody a callback. The same rigor only matters for fields, like most recent title or employer, that change who gets shortlisted or rejected.
Fill in the blank
4. The rebuilt eval has 90 resumes total: ___ typical single-column PDFs and 60 spread across the five other shapes, checked at a ___ percent aggregate bar with an 80 percent floor per shape.
Show hint
Check the build-up chart's largest segment and the aggregate bar named in the O step.
Show answer
30 single-column PDFs, and a 90 percent aggregate bar. 30 + 20 + 15 + 15 + 5 + 5 = 90.
Short answer, apply it yourself
5. Pick an AI feature you use yourself. Before its first release, what's one real-world "shape" of input it probably wasn't tested enough on?
Show hint
Think of a version of the input that's less common but still real: an accent, a file type, an unusual layout, a second language.
Show answer
Model answer: "A receipt-scanning app probably tested crisp, flat receipts from big chains first. A crumpled, handwritten farmers-market receipt is a real shape it might never have been tested on before release." Any answer works if it names a real input shape and why it's plausibly under-tested.
Multiple choice
6. What old decision does this answer actually take back?
  • A. Adding a person to manually review every parsed resume forever.
  • B. Testing the parser on twenty resumes from one shared folder instead of a structured, six-shape minimum.
  • C. Removing a confidence score from the recruiter's screen.
  • D. Lowering the release bar so the pilot could ship faster.
Show hint
A dial turned up or down doesn't count. Look for the actual decision made in the Thursday meeting.
Show answer
B. The Thursday meeting chose convenience, twenty resumes already on hand, over structure, and never checked what shape those twenty actually were.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more