What is the minimum eval you would accept before a first internal release?
- Size the eval by real resume-format coverage, not by what's already sitting in a folder.Why: a number with no categories behind it can't tell you what it never tested.
- Break the applicant pool into its real shape categories before picking a total.Why: single-column, two-column, DOCX, scanned, non-English, and odd layout each fail differently, and one percentage hides all of that.
- Set a floor per category, not just an aggregate bar.Why: an average can look great while one shape is badly broken underneath it.
- Move the whole number with what the release actually risks, not with the deadline.Why: a closed alpha and a release that feeds real candidate records don't deserve the same minimum.
- Check the minimum against how long a person actually takes to grade it by hand.Why: a number nobody can review before launch is a number that gets skipped under pressure.
- Name which assumption would break the number first: a missing format, not reviewer time.Why: that's the one worth stress-testing before anyone trusts the number.
How to answer this, stage by stage
Nobody is grading whether you land on exactly ninety resumes. They are grading whether the number has real coverage under it, moves with the stakes, and survives a check against how long it takes a human to actually run. Seven moves get you there.
Let's learn
Resume Fill is a feature inside Slatework, an applicant tracking system, that reads an uploaded resume and fills in a candidate's name, most recent title, most recent employer, dates, and phone number, so a recruiter doesn't type them in by hand.
Before Resume Fill, a recruiter typed those fields by hand from every resume, about ninety seconds a candidate. A team working through two hundred applications a week spent about five hours a week on nothing but that typing.
Now Resume Fill does it in about three seconds, and on a clean, standard resume it is close to perfect.
At its worst, a two-column resume gets its most recent job silently swapped with an older one pulled from a sidebar summary, the wrong years of experience get logged, and a strong candidate quietly drops out of a shortlist search, because the recruiter trusts the filled fields and never opens the source PDF again to check.
The choice I would take back. The pre-launch eval was twenty resumes pulled from the team's own shared hiring folder, because they were fast to grab. It was never checked against what shape those twenty actually were.
What I would leave alone. A low-stakes field like a candidate's portfolio link doesn't need the same fifteen-per-shape floor as the fields that decide who gets shortlisted. A wrong or missing link there costs nobody a callback.
The lesson. A minimum eval isn't the smallest set that looks clean. It's the smallest set that has a real chance of showing you what you don't already know. Any eval built entirely from resumes you already have can only prove the parser handles the resumes you already have.
Now here is the same thing as a story
Skip this part if you already believe a first-release eval should be sized by shape coverage, not convenience. Read on if you don't.
Kaveri Nadira has run resume tooling at Slatework for two years. When something in the pipeline looked off, she was the one who pulled the actual PDF and read it herself, not just the ticket describing it.
The Thursday before Resume Fill's pilot launch, engineering had it parsing cleanly. Kaveri needed one thing before three pilot recruiters got access on Monday: proof it worked. Someone on the team grabbed twenty resumes from the shared hiring folder, ran the parser against them in a twenty-minute meeting, and got nineteen out of twenty fields right. Nobody asked what shape those twenty resumes happened to be. It didn't come up. They were all clean, single-column, English PDFs, because that was what the team's own past hires had sent in.
For six weeks, the pilot went well. Every morning around nine, a batch of new applications would land, and by five past, Resume Fill had already filled in the name, title, employer, and dates for every one of them. At first the three recruiters checked every filled resume against the PDF before saving. After two weeks of it always matching, they checked maybe one in five. By week four, they didn't check at all. They just saved and moved to the next candidate.
Then a hiring manager on a global sales search asked Kaveri a small, ordinary question. "Why does this candidate's most recent job say 2019? Her LinkedIn says she left that job last year." No error message. No dashboard alert. Just a question Kaveri couldn't answer without opening the file.
The resume was a two-column layout, a photo and a skills sidebar on the left, the real work history on the right. On that shape, the parser kept grabbing a line from the sidebar summary and reading it as the most recent job, instead of the actual latest entry in the timeline column. Kaveri pulled three weeks of applications from that hiring manager's roles, which drew a lot of candidates using that same European-style template. On the two-column resumes, most recent employer was wrong close to forty percent of the time, and nothing on the screen ever said so.
Kaveri remembered the Thursday meeting clearly, twenty resumes, nineteen right, done in twenty minutes. It had felt thorough at the time. Nobody in that room had asked what those twenty resumes actually were, because asking felt like slowing down a release that was already running on time.
She rebuilt the eval that week: ninety resumes across six real shapes, weighted by how often each one actually showed up, with a ninety percent bar overall and an eighty percent floor in every shape. Run against the model as it stood, the two-column category scored sixty two percent, well under the floor. A fix to stop the parser reading the sidebar as the timeline took two days. Rerun, the category cleared eighty eight percent before the wider recruiter group got access. The next near miss, a similar sales search three weeks later, got caught inside the eval, three days before it would have reached a hiring manager the same way.
The thing I'd tell my Thursday-afternoon self: twenty resumes proving the parser works is not the same claim as twenty resumes proving nothing's broken. I had the second one. I told the room I had the first.
BOUND, sized for a Friday launch
This is a sizing question about coverage and a floor, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. A minimum eval is a floor of test resumes in every real shape the parser will meet, plus a bar each shape has to clear. Skip either half and the number stops meaning anything.
O, own the numbers. Ninety resumes total, weighted to the applicant pool's real mix: thirty single-column PDF, twenty two-column, fifteen DOCX, fifteen scanned, five non-English, five odd layout. Checked on five fields: name, most recent title, most recent employer, dates, phone.
U, use a range. About forty five for a closed alpha with no hiring decision riding on it. About a hundred fifty or more once a release feeds real candidate records recruiters act on.
N, nail the sanity check. Grading ninety resumes by hand, five fields each, at about two minutes a resume, is three hours, one working day, barely slower than the ninety-seconds-a-resume job it replaced.
D, direction. The number moves most on a missing shape, not on reviewer time. A category that was never in the eval and is real in the world is worth more than any amount of extra grading time on the shapes already covered.
And if you want to be sure it really works, try it somewhere else
A water utility is about to let an AI read a customer's photo of their meter and log the usage number, instead of sending a technician out to read it in person.
B, break it down. The minimum eval is coverage across real photo conditions: an analog dial meter, a digital LCD meter, a meter photographed at an angle, one in low light, one behind a fogged glass cover, and one with glare across the display.
O, own the numbers. A hundred twenty test photos. Forty analog dial, thirty digital LCD, twenty angled, fifteen low light, ten fogged, five glare-heavy, weighted by how the real photo submissions actually break down. Bar: within one digit of the true reading on ninety five percent, no shape under eighty five percent.
U, use a range. Forty test photos is enough for an internal pilot among utility staff reading their own meters, low stakes, easy to catch by hand. Push to two hundred or more once the reading feeds a customer's actual bill.
N, nail the sanity check. Grading a hundred twenty photos by hand, about a minute each against the true reading on file, is two hours. A technician's truck-roll to read one meter in person already costs about twenty minutes plus drive time, so even two hours of hand-grading is far cheaper than the process it's replacing.
D, direction. For this product, the number doesn't move most on how many photo conditions exist. It moves on who pays for a wrong reading. A reading that changes a customer's actual bill needs a tighter floor than a reading that's just informational, with a person reconciling it before it ever reaches an invoice.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the number and the bar: ninety resumes, six shapes, ninety percent overall with an eighty percent floor per shape. The build-up backs it up if they ask.
Cost: instead of asking what the total should be, a manager caps how many resumes a reviewer can grade in a day, say sixty. Same rule, solved backward: allocate that sixty across the six shapes by risk and how often each one shows up, instead of picking a total first.
The model got better: a new version ships that reads two-column resumes almost perfectly. The minimum doesn't shrink on its own. Rerun the per-shape floor check first, since acing the shapes it always handled proves nothing about the one that used to fail.
Where people run it wrong.
They pick a round number, like fifty, because it sounds thorough, with no shape categories behind it.
They set the minimum once before launch and never revisit it as new resume formats or new markets show up.
They grade the eval against the parser's own confident output instead of against the real source file, so a fluent wrong answer passes.
How to use it live. Say the equation before any number: "the minimum eval isn't a count, it's a coverage floor across every shape the parser will actually meet, sized to what a miss costs." That buys the time to work out the real split instead of guessing a number that sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Writing an eval spec
- #1 What is an eval spec and who is its audience?
- #2 List the components of a complete eval spec.
- #3 How do you define a task-level success criterion for a summarization feature?
- #4 Write a scoring rubric for the quality of a generated customer support reply.
- #5 Describe the difference between an eval spec and a test plan.
- #6 How many examples belong in a first eval set and how do you choose them?