ConceptAdvancedEval-Driven Specification / Writing an eval spec / #8

What is the role of a human baseline in an eval spec?

The direct answer
A human baseline is the reference an eval spec is actually built around, not a footnote added after the model ships. Before anyone writes a target like "90 percent agreement," score a real person doing the exact same job, on the exact same set of cases, judged by the exact same rule, including the hard cases, not just the easy ones sitting in the drawer. That score is the baseline. Every pass line in the spec gets written as a gap from that number, never as a percentage floating on its own with nothing to be judged against.
Do this, in order
  1. Build the human baseline first, before any target goes in the spec.Why: a target with nothing to be judged against is a hope, not a requirement.
  2. Score the baseline on the same population the model will actually screen, hard cases included.Why: leaving out the hard cases makes the baseline look stronger than it is, and hides exactly where the model needs the most watching.
  3. Have the baseline scored blind, by someone who never sees the model's own call.Why: a grader who's already seen the model's answer tends to agree with it, which quietly erases the whole comparison.
  4. Write the pass bands off the gap to baseline, not off the raw score.Why: 91 percent means something completely different next to a baseline of 84 than it does next to a baseline of 98.
  5. Re-check the baseline whenever the mix of real cases shifts.Why: a baseline measured on last year's cases stops being the right yardstick the moment the population it was built on changes.
  6. Post the baseline next to the model's score everywhere the score gets shown.Why: a hidden baseline lets a number get reported as good news without anyone able to check it.

How to answer this, stage by stage

Seven moves. Say the baseline's number out loud before the model's, or the score you give next is a guess dressed up as a fact.

1
Say what a bare score can't tell you
Say it like this
"'91 percent agreement' sounds like a real result, but on its own it doesn't say if that's good. Ninety-one percent compared to what? Before I answer the question, I want to say that plainly, because spotting that gap is most of the job."
Why this works
Naming the empty number first proves you understand why a baseline matters, instead of just reciting the term back.
2
Put one real product and one real person under it
Say it like this
"Say this is a credit union, and it's built a tool that pre-screens small consumer loan applications before a loan officer ever opens the file. Devika Rao runs the eval spec for it."
Why this works
A baseline for "the model" in general is empty. It only means something once you can say whose job it's standing in for.
3
Name the real outcome the score has to serve
Say it like this
"What we actually care about isn't 'the model matched a label in a spreadsheet.' It's whether the loan decision that comes out the other end is the right one, the kind that holds up if a regulator or an applicant ever asks us to explain it."
Why this works
This is the L step. Skip it, and the target you write next is just a number nobody can defend.
4
Build the baseline before the model's number means anything
Say it like this
"Before I score the model at all, I get real underwriters to score the exact same batch of files, using the exact same rule. Say two senior officers agree with the confirmed outcome 84 percent of the time on the batch the model actually sees. That's the baseline. Now the model's score has something to stand next to."
Why this works
This is the E step, and it's the whole answer. A model score with no baseline is a fact about the model. A model score with a baseline is a fact about whether to trust it.
5
Guard the baseline sample so it can't be gamed
Say it like this
"I'd ask exactly how that baseline sample got picked. If it's built from the easiest, already-obvious files because those are fast to check, the baseline comes back looking stronger than it really is, and it hides the very cases the model has the hardest time with."
Why this works
This is the A step. Naming the cheat yourself tells an interviewer you understand the metric, not just the word.
6
Write the decision bands off the gap, not the raw score
Say it like this
"Below the baseline by a real margin, the model stays advisory only, a human reads every file. Within a few points either side, the model's call gets a second person's sign-off on anything it flags. Clearly above the baseline, and holding for weeks, the smallest, lowest-risk loans can fast-track on the model's call alone."
Why this works
This is the D step. A baseline nobody acts on is trivia. This is what turns it into a real decision.
7
Close on the question that catches teams without one
Say it like this
"If an examiner ever asks me how the model's number compares to what a real underwriter gets on these same files, I want an answer ready, not a scramble. That question is the whole reason to build the baseline before anyone needs it."
Why this works
This is the line worth leaving the room with. It shows you're treating the spec as something that has to survive being questioned, not just something that has to ship.
If you only get through two stages Stages 4 and 5 are the answer. Say what the baseline is, scored on the real cases, and say how you'd guard it from being built out of the easy ones. Everything else on this list is how you defend that baseline under follow-up.

Let's learn

Say a credit union builds a tool that reads a stack of small consumer loan applications and tells a loan officer which ones look safe to fast-track.

Before the tool, four loan officers read every one of the roughly 90 applications that came in each morning, in full, about 18 minutes each. That's 27 hours of reading landing on four people who also had customers standing at the counter. By Thursday, the queue was already two days behind.

Now the model reads every file in under a minute and recommends fast-track, hold, or flag for a closer look. In testing, it agreed with the officer's final decision 91 percent of the time.

Here's the part that matters. The eval spec also listed a human baseline: 98 percent. Someone had pulled 50 already-closed, obviously clean files, no self-employed income, no thin-file cases, and had one officer re-score them in an afternoon. Next to that number, 91 percent looked like the model falling seven points short of a person. So the model was kept advisory only, never trusted to skip a review, and every officer had to add a short note explaining why they agreed or disagreed with its flag.

The model wasn't seven points behind a person. It was being measured against a number nobody had actually earned.
Two panels side by side. Left, a small stack of clean files labeled easy fifty, clean files scored in an afternoon, 98 percent agreement. Right, a larger reddish block labeled hard cases, self-employed, thin file, never sampled.
The baseline that shipped in the spec never met the files the model actually screens

What that costs at its worst: the note took about four minutes, even on a file the officer would have approved in thirty seconds anyway. Review time went from 18 minutes a file to 22. The queue that used to clear by Tuesday started running into Thursday. The tool didn't get the credit union caught up. It made the wait longer than before anyone built it.

Needs the closer read either way

Where officers disagree most

  • Self-employed income with no W-2 to check against
  • A thin file, under two years of credit history on record
  • Debt-to-income sitting right at the cutoff, no clean paycheck to confirm it
Officers and the model both struggle here. This is where a real baseline earns its keep.
The baseline says this is fine

Where nobody really disagrees

  • W-2 income, verified in one phone call
  • Credit score above 740, no flags on the file
  • An amount under the credit union's smallest consumer-loan tier
Officers barely disagree with each other or the model here. Fast-tracking these is where the tool actually pays for itself.
Knowledge spark: what actually counts as "the officer got it right" Not whether the officer's call matched the model's flag. It's whether the officer's call matched the outcome that held up after any appeal, checked by someone who never saw the model's stated reason. That's the only thing allowed to count in the baseline.
The leading edge: hard-case share of the weekly queue
12% 17% 22% 28% 34% Month 1 Month 2 Month 3 Month 4 Month 5, exam
under a fifth of the queue is hard cases
climbing past a fifth
over a quarter, baseline never re-checked
The hard-case share crossed 25 percent in month 4. The baseline in the spec was still the same 50 easy files from January.
The lagging outcome: declines overturned on appeal, hard-case files only
2 a month
February, launch month
13 a month
June, the month of the exam
The hard-case share had already crossed 25 percent two months before this number got noticed, because nobody was watching the segment, only the one aggregate score.
The choice I'd take back The 98 percent baseline went into the spec almost as an afterthought, built from a stack of files nobody would ever argue about. I'd take that back and insist the baseline get scored on the same mix of files the model actually screens, hard cases included, before anyone treated 91 percent as good or bad.

What I'd leave alone. The credit union's hard-block rules stay absolute, and that's correct. An application tied to an active bankruptcy, or a Social Security number that doesn't match, gets declined by policy, full stop. There's no judgment call there, so there's nothing to baseline. A baseline is only for the calls a careful person could honestly go either way on.

The lesson. If your spec has a number the model has to hit, it needs a matching number for what a real, careful person gets on the exact same files, or the model's score is just a figure with nothing to be judged against.

The number that went unquestioned for five months

You don't need this to answer the question. It's here so "91 percent" stops being an abstraction and starts being a Tuesday.

Devika Rao has run the eval spec for Briarcliff Credit Union's lending tools for three years. Ask her what's being scored on any given model and she can name the metric, the sample, and who signed off on it inside a minute.

The pre-screening tool went live in February. For the first few months, it was mostly good news. Officers still read every file end to end, and the tool just sat there flagging its own guess, approve, hold, or a closer look, matching what the officer decided nine times out of ten. Nobody had to change how they worked. It felt like a second pair of eyes that happened to be fast.

Around month three, the credit union opened a new small-business lending line, and the mix of applications moving through the queue began to change. More self-employed income. More thin-file borrowers with no long credit history to lean on. Nothing about the tool changed. The files it was being asked to judge did.

Nobody decided to stop trusting the sign-off number. It just stopped meaning as much. The 98 percent baseline in the spec had been measured back in January, on fifty files that closed clean months before the small-business push. It never got touched again. Nobody's job was to touch it.

Five months in, a state examiner sat down with Devika for the credit union's routine fair-lending review. Halfway through, she asked one plain question: "You cite a 90 percent target and a 98 percent human baseline here. Show me the file set that 98 percent was scored on."

Devika pulled it up. Fifty files. Every one closed by month two. Not one self-employed applicant. Not one thin file. None of it looked like the roughly one in three applications now moving through the queue.

She didn't get defensive about it. She went and built the number that should have existed from day one: a random weekly sample of sixty files, matched to whatever the real queue actually looked like that week, scored blind by two senior officers who never saw each other's call or the model's flag, checked against the outcome that held up after any appeal.

On that sample, the two officers agreed with the confirmed outcome 84 percent of the time. Loan calls on thin-file, self-employed cases are genuinely hard. Careful people disagree with each other on them too.

The model, scored on that same representative batch: 87 percent.

We hadn't been asking the model to catch up to a person. We'd been asking it to catch up to a number nobody had actually earned.

Here's what that cost while nobody was checking. Because the team believed the model was seven points behind a real underwriter, they never let it carry any weight on its own, every file still needed a full read plus the new disagreement note. Meanwhile the segment where the model actually needed watching, the hard cases specifically, got no extra scrutiny at all, because everyone was staring at one aggregate number instead of the segment that mattered. Declines on those hard files that got overturned on appeal climbed from 2 a month in February to 13 a month by June.

Devika went back to the sign-off meeting in her head, the one from January. The spec had one line under quality: "90 percent target, human baseline 98 percent." Nobody in the room asked where the 98 came from. It sounded high enough to be safe, and a high enough number rarely gets questioned in a room that's eager to ship.

Here's the replay. If the representative baseline, 84 percent, had been in the spec from day one, the model's 87 would have read as what it actually was: already ahead, on the exact cases that mattered most. The credit union could have let it carry the easy fast-tracks on its own by month two, freed officers to spend their fuller reviews on the hard segment specifically, and caught the rising overturned-decline count in March instead of finding it by exam in July.

What I'd tell myself, back in that January meeting: I let a number into the spec because it looked reassuring, not because anyone had checked what it was standing on. A baseline you don't interrogate is just a guess wearing a percentage.

The four letters that make a score defensible

This is a concept question about writing an eval spec, but the shape underneath it is the same as any metric question: a number means nothing until you know what it's supposed to move ahead of. That's LEAD, not GUARD. GUARD asks who can't push back against a decision already made. Here, nobody had even decided what the right decision looks like yet.

L, link. The real thing a loan decision has to be right about: whether the call holds up, for the credit union and for the applicant, not whether the model matched a label in a spreadsheet.
E, early signal. The human baseline itself. A real underwriter, scored on the same files, established before the model's number gets treated as good or bad.
A, abuse. Build the baseline sample from the easiest files, the ones nobody would argue about, and it comes back inflated, hiding exactly the cases the model needs the closest watch on.
D, decision. Below baseline by a real margin, the model stays advisory only. Within a few points either side, a second person signs off on anything it flags. Clearly and consistently above baseline, the smallest, lowest-risk loans can fast-track on the model's call alone.
The check that keeps a baseline honest Try scoring the baseline on only the easiest slice, on purpose. If the number barely moves, the first sample was close to representative. If it jumps, the easy slice was hiding the real difficulty, and that's the corner a busy team cuts without meaning to.

And if you want to be sure it really works, try it somewhere else

A veterinary teleconsult service screens photos of pet skin lesions to decide which pets need a same-day vet visit and which can wait for a routine appointment. Same shape of question. "The model should already be as good as a vet" is just as ungrounded as a bare percentage, until someone measures what a vet actually gets.

L. A pet with something serious gets seen the same day, and a pet with something ordinary doesn't get an unnecessary rushed visit.
E. Two vets independently score the exact same batch of real lesion photos, blind to each other's call and to the model's flag, before the model's own score means anything.
A. Score the baseline on the clearest, most obvious photos, sunlit and in focus, and skip the blurry phone shots most owners actually send, and the baseline comes back higher than any real batch of photos could match.
D. Below the baseline, the tool only ever suggests, a vet reviews every photo. Close to it, a second vet checks anything flagged same-day. Clearly above it, the tool can triage the routine cases on its own, with weekly spot checks.
A labeled diagram centered on a document icon marked baseline first, with four labels radiating around it: two vets score it, blind to each other, real lesion photos, before the AI score counts.
Same idea, a different desk

Swap the trigger and it still runs

  • It gets slower. Doesn't matter. Scoring sixty files a week takes the same afternoon whether the model answers in a second or ten.
  • It gets more applications. If the small-business push doubles the queue, the weekly sample still stays at sixty. You just check it as often.
  • It gets better than planned. If the model starts handling self-employed income cleanly, rerun the same weekly check against the wider set it now covers. The process doesn't change. Only what's inside the sample does.

Where people run it wrong

  • Building the baseline once at launch and never touching it again, while the real population of cases keeps shifting underneath it.
  • Letting the same person who built the model also pick which files go into the baseline sample.
  • Reporting one aggregate score instead of scoring the hard segment separately, so a real problem hides inside a healthy-looking average.

How to say it if you're asked this cold

Buy yourself the time to build the number properly. "Before I tell you if that score is good, let me say what it's actually being compared to." That's not stalling, that's the E step, said out loud, and it gives you somewhere honest to stand while the real baseline takes shape in your head.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what does it stand for here?
Tap to flip
ANSWER
LEAD. L is the real outcome the score has to serve, E is the human baseline itself, established before the score means anything, A is how that baseline gets gamed, D is what changes at each gap between the model and the baseline.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Devika Rao, who has run the eval spec for Briarcliff Credit Union's lending tools for three years and can name the metric, the sample, and the sign-off for any of them inside a minute.
3 · THE HABIT
What did the team stop doing once the tool's flag kept matching the officer's call?
Tap to flip
ANSWER
Questioning where the baseline number in the spec had actually come from. It sat unchanged in the document for five months while the real mix of applications shifted underneath it.
4 · THE BASELINE
In plain words, what is a human baseline?
Tap to flip
ANSWER
How well a real, careful person does on the exact same set of cases the model is being scored on, measured before anyone decides if the model's own score is good.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting a 98 percent baseline into the spec because it looked reassuring, without anyone checking that it was built from fifty already-easy files instead of the real mix the model would actually screen.
6 · THE NUMBER
On the real, representative sample, officers agreed with the confirmed outcome ______ percent of the time, and the model scored ______ percent, meaning the model was already ______ the baseline.
Tap to flip
ANSWER
84 percent. 87 percent. Above the baseline, not seven points behind an unearned 98.
7 · THE REPLAY
Same eval spec, real baseline from day one, what changes?
Tap to flip
ANSWER
The model's 87 percent reads as already ahead of the 84 percent baseline. The credit union lets it carry the easy fast-tracks by month two, and the rising overturned-decline count on hard files gets caught in March instead of surfacing in a July exam.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
A veterinary teleconsult tool screening pet skin-lesion photos. Its early signal is two vets independently scoring the same real photo batch, blind to each other and to the model, before the model's score counts for anything.

Check yourself Score: 0 / 0

Fill in the blank
1. The rewritten baseline was scored on a random weekly sample of ______ files, by ______ senior officers working blind to each other, and came back at ______ percent agreement with the confirmed outcome.
Show hint
The numbers are in walkthrough stage 4 and in the story's replay.
Show answer
60 files. Two officers. 84 percent. That's the number the model's 87 percent should have been compared to from the start, not the 98 percent built from fifty easy files.
Multiple choice
2. Which of these would count as a real human baseline for this tool?
  • A. The vendor's stated accuracy for similar screening tools.
  • B. The model's own confidence score, averaged across a week.
  • C. Two officers independently scoring the same real files the model screens, checked against the confirmed outcome.
  • D. How often officers agreed with the model's flag, tracked over time.
Show hint
Three of these are facts about the model or about someone else's tool. Only one is a real person doing the same job on the same cases.
Show answer
C. Agreement with the model's own flag isn't independent, it just measures whether officers copy the model, not whether either of them is right. A real baseline needs a person judged against the real outcome, not against the model.
True or false
3. True or false: this problem gets fixed by telling the model to be more careful with self-employed applicants.
  • True
  • False
Show hint
Ask how anyone would know whether the model actually got more careful.
Show answer
False. An instruction to the model can't be checked without a baseline. You'd still have no way to know if "more careful" closed the gap, or what gap you were even closing it against.
Short answer
4. Name a place in this same spec where an absolute rule is fine, and doesn't need a human baseline at all.
Show hint
Look for a rule that's an exact match, not a judgment call.
Show answer
Model answer: "The hard-block rules, an active bankruptcy on file, or a Social Security number that doesn't match. Those are exact matches with no ambiguity in them, so there's no human judgment to baseline against. Rates and baselines are for the calls a careful person could honestly go either way on."
Short answer, apply it yourself
5. Pick a product you use yourself. Name one place it reports a score with no baseline attached. What would the human baseline look like?
Show hint
Look for a headline number with nothing to compare it to, like "95 percent accurate" or "highly rated."
Show answer
Model answer: "A resume-screening tool that claims '95 percent match accuracy.' The baseline: have two experienced recruiters independently score the same batch of real resumes against the same job, blind to the model's call, and see how often they agree with each other and with who actually got hired."
Multiple choice
6. If the baseline sample had been built only from the easiest 20 files instead of a representative 60, what would most likely happen to the reported baseline number?
  • A. It would drop, because easy files are harder to score consistently.
  • B. It would rise, making the model look worse than it actually is by comparison.
  • C. It would stay exactly the same either way.
  • D. It would become impossible to compute at all.
Show hint
Easy files are the ones people rarely disagree about.
Show answer
B. Humans agree with each other and with the outcome almost automatically on easy files, so a baseline built only from those comes back inflated. That's exactly the abuse the A step warns about.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more