ConceptAdvancedResponsible AI & Advanced Practice / Building an AI PM portfolio / #8
What role should an eval report play in a portfolio?
LEAD Petra Nováková built an eval report for a second-read screening tool modeled on Cascade Diagnostic Imaging, a radiology practice
Petra Nováková spent five years as a radiology technologist before starting to build AI product work. Her portfolio project is a second-read screening tool modeled on Cascade Diagnostic Imaging, flagging scans that likely warrant a closer look before a radiologist signs off. Marcus Teague hires for a health-tech company's AI product team, and he reads eval reports the way most people read the fine print, closely, and rarely believing the headline number alone.
The direct answer
An eval report's job is to be evidence, not the headline. It should support one specific judgment call you made, name the exact eval set and model version behind it, and show the cases it got wrong. A number with no eval set, no version, and no failure case attached predicts nothing about how you'll actually perform. A rigorous one predicts it weeks before anyone checks your work directly.
Do this, in order
Name the exact eval set and model version behind every number in the report.Why: a score with no named set and no version pin is a claim, not evidence.
Show the cases the model got wrong, not just the ones it got right.Why: a report with zero visible failures is the clearest sign nothing was actually tested rigorously.
Tie the report to one specific decision, not a general capability claim.Why: "94% accurate" proves nothing on its own; "this threshold catches these cases and misses these" proves judgment.
Place the eval report as supporting evidence, behind the demo and the judgment write-up, not as the headline.Why: most reviewers need to see the artifact work before they're ready to read the rigor behind it.
Watch for the ways an eval report gets gamed, and avoid them in your own.Why: a cherry-picked or leaked eval set produces an impressive number that predicts nothing real.
Keep it short enough that a reviewer will actually read it.Why: a rigorous report nobody finishes reading does the same job as no report at all.
How to answer this, stage by stage
This question isn't asking whether you know what an eval report is. It's asking whether you understand what it's actually for, and what it's not.
Stage 1
Scope it to one real report
Say it like this
"I'll answer this for one eval report inside a candidate's portfolio, sitting next to a working demo and a judgment write-up."
Why this works
Grounds an abstract question about "role" in a concrete artifact sitting on a real page.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the outcome that actually matters. Early signal, what predicts it before anyone checks directly. Abuse, how it gets gamed. Decision, what a reviewer does at each threshold."
Why this works
Treats an eval report as a signal to be evaluated, not just a section to be filled in.
Stage 3
Name the outcome that actually matters
Say it like this
"The real outcome isn't 'does the model score well.' It's whether this candidate will reason carefully about model quality once they're hired."
Why this works
Reframes the eval report away from being about the model, and back onto what a hiring decision actually needs to predict.
Stage 4
Give the early signal
Say it like this
"A named eval set, a pinned model version, and visible failure cases predict good on-the-job judgment weeks before any reference check would. A bare accuracy number predicts nothing at all."
Why this works
This is the direct answer to the question: the report's rigor, not its score, is the actual signal.
Stage 5
Name how it gets gamed
Say it like this
"Petra's first draft tested the model on the same twenty scans she'd used to tune its threshold. The number looked great and meant almost nothing."
Why this works
Shows a real, specific way the signal breaks, not just an abstract warning.
Stage 6
Say what you'd do at each threshold
Say it like this
"If a report names its eval set, version, and failure cases, I weight it heavily. If it's just a number with none of that, I treat it as decoration and ask for the underlying set directly."
Why this works
A signal nobody acts on differently is just a dashboard number, not a real decision input.
Stage 7
Close on where the report sits
Say it like this
"The eval report supports a judgment call, it isn't the headline. Lead with the demo, back it with the report, and let the rigor speak for itself once someone's already engaged."
Why this works
Restates the direct answer and closes on where the artifact belongs, not just what it should contain.
Let's learn
Picture a hiring panel deciding, weeks before any real evidence exists, whether a candidate reasons carefully about a model's quality.
Before Petra reworked her eval report, it read: "93% sensitivity on flagged scans." One sentence, no set named, no version, no failure case. It told Marcus almost nothing he could act on.
Knowledge spark: what's a version pin?
A record of exactly which version of a model, on which date, produced a given result. Models get quietly updated. A number with no version pin can't be reproduced by anyone, including the person who first reported it.
After the rework, the same underlying project carried a report naming a 60-scan public radiology eval set, the exact model checkpoint and date used, and four cases where the model missed something a radiologist would have caught, each with a short note on why.
Reviewer trust in the eval report, draft by draft
Trust rose steadily as the eval set, version, and failure cases got added, well before Marcus ever verified any of it directly.
At its worst: a headline number gets quoted straight into an interview debrief, someone later asks what set it came from, nobody can answer, and the whole claim quietly gets discounted, along with everything else the candidate said.
The early signal that actually predicts good judgment
A named eval set, a pinned model version, and visible failure cases predict careful on-the-job reasoning weeks before any reference check or first project could confirm it directly. A bare accuracy number, no matter how high, predicts nothing on its own.
What I would leave alone: the headline number itself is fine to keep in the report. The problem was never having a number. It was having only a number, with nothing behind it a reviewer could check.
A number with no eval set behind it isn't evidence. It's a guess wearing a lab coat.
The lesson: an eval report was never supposed to prove the model is good. It was supposed to prove the person building it knows exactly how they'd find out if it wasn't.
Now here is the same thing as a story
The short version above is what you'd say defending your own eval report in an interview. Read this one for how Petra actually rebuilt hers.
Petra could tell, from years of prepping scans, exactly which images tended to trip up even an experienced radiologist on a busy afternoon.
Her first eval report felt, to her, thorough: a clean paragraph, one confident percentage, written the way she'd seen quality metrics presented in her old clinic's monthly reports.
One version of the report asks to be the whole argument. The other version asks to support one, smaller, provable claim.
A more experienced friend, reading a draft, asked a plain question: "which twenty scans is that ninety-three percent even from?" Petra realized, mid-answer, that it was the same twenty she'd used to pick the model's threshold in the first place.
One clock only rings ninety days into a new job. The other rings today, in the report itself, if you know to look at it.
She rebuilt the whole thing over a weekend: sixty scans from a public dataset she hadn't touched while building the model, the exact checkpoint version and date labeled clearly, and four real misses written up honestly instead of hidden.
Four parts, and the second one, the version pin, is the one candidates most often skip entirely.
Four things separate a real eval report from a confident-sounding sentence wearing the same name.
She also moved the report itself, placing it behind the working demo and a short judgment write-up, instead of leading with it.
The eval report earns the most trust once a reviewer has already touched something real, not before.
Most eval artifacts fail on one axis or the other. A well-framed report is rare enough that it stands out on its own.
The old report asked Marcus to trust a number pulled from nowhere. The new one shows him exactly which sixty scans, which model version, and which four misses, and lets him decide for himself how much that's worth.
I wrote my first eval report the way I'd seen quality numbers reported for years, as a clean, confident summary with nothing messy attached. Hearing my friend ask which twenty scans it came from is what showed me the summary was never the evidence. The twenty scans were.
LEAD, for the eval report itselfNot proof the model is good. Proof of how carefully you'd find out if it wasn't.
L
Link. The outcome that actually matters.
Not whether the model scores well, but whether this candidate reasons carefully about model quality on the job.
Reframes the whole question away from the model and onto the candidate's judgment.
E
Early signal. What predicts it, before anyone checks directly.
A named eval set, a pinned version, and visible failure cases predict good judgment weeks before any reference check could confirm it.
The hardest step, and the direct answer to the whole question.
A
Abuse. How it gets gamed.
Testing on the same data used to tune the threshold, or quietly curating only the easy cases, both produce an impressive number that predicts nothing real.
Every real signal has a way to be hit without actually doing the underlying work.
D
Decision. What a reviewer does at each threshold.
A rigorous report gets weighted heavily. A bare number gets treated as decoration, and the reviewer asks for the real eval set directly.
A signal nobody acts on differently is just a number on a page, not a real decision input.
Interview-to-offer conversion, portfolios with and without a real eval report
Same underlying skill, radically different conversion, once the eval report actually became checkable instead of asserted.
The recap, one line per letter: link is the candidate's real judgment, not the model's score; early signal is the eval set, version, and failure cases predicting that judgment early; abuse is testing on the same data used to tune the model; and decision is weighting a rigorous report heavily while treating a bare number as decoration.
And if you want to be sure it really works, try it somewhere elseSame four letters, a pharmacy drug-interaction checker instead of a radiology tool. A different candidate, and this time the abuse case is far more dangerous if missed.
Nkechi Obi built an eval report for a drug-interaction flagging tool aimed at community pharmacies, checking new prescriptions against a patient's existing medication list. Applied to LEAD: link is whether Nkechi reasons carefully about a model that, if wrong, could miss a genuinely dangerous interaction. Early signal: the same three elements, a named eval set of real interaction cases, a version pin, and visible misses, this time weighted even more heavily given the stakes. Abuse: the most dangerous version of gaming this report is quietly excluding rare but severe interactions from the eval set because they're hard to source data for, producing a high score that hides exactly the failure mode that matters most. Decision: Nkechi's report explicitly calls out which interaction severity tiers were and weren't covered by her eval set, rather than letting a single blended accuracy number imply even coverage across all of them.
The same four parts matter even more here, since the cost of an unnoticed gap in the eval set is measured in patient safety, not just interview outcomes.
Swap the trigger and it still runs.
Speed: an interviewer wants your answer in one breath. Say "evidence, not headline," and stop.
Cost: you don't have access to a large real eval set. Use a small, honestly labeled public or synthetic one and say so plainly, rather than skipping the report entirely.
The model gets better, for real: even if the underlying model's accuracy improves on its own, the eval report's job doesn't change, it still needs to show the set, the version, and what's still missed.
Where people run it wrong.
They lead with the eval report before a reviewer has any reason to care about the number yet.
They report a single blended score that hides uneven performance across different case types.
They test on the same data used to tune the model, producing a number that looks rigorous but predicts nothing.
How to use it live. When asked what role an eval report should play, answer with what it isn't first: not the headline, not proof the model is good. Then name what it is: evidence, in support of one specific, named judgment call.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "what role should an eval report play in a portfolio"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. The early signal step names what actually predicts good judgment: a named set, a version pin, and shown failures.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Petra Nováková, a former radiology technologist building a second-read tool for Cascade Diagnostic Imaging, and Marcus Teague, the hiring manager reviewing her eval report.
3 · THE HABIT
What habit did Petra have to drop, once she understood the real problem?
Tap to flip
ANSWER
Reporting a single confident percentage with no named set behind it, the way she'd seen quality numbers presented for years.
4 · THE EARLY SIGNAL
What actually predicts good on-the-job judgment, weeks before a reference check could confirm it?
Tap to flip
ANSWER
A named eval set, a pinned model version, and visibly shown failure cases, not the size of the headline accuracy number.
5 · THE ABUSE CASE
What decision would you take back from Petra's first eval report?
Tap to flip
ANSWER
Testing the model on the same twenty scans she'd used to tune its threshold, which made the number look strong while meaning almost nothing.
6 · THE NUMBER
Fill in the blank: portfolios with a real eval report converted interviews to offers at about ___ percent, versus 8 percent for a bare accuracy claim alone.
Tap to flip
ANSWER
35 percent. The gap between 8 and 35 is what a checkable eval report actually buys a candidate.
7 · THE REPLAY
Same interview, redesigned eval report. What does Marcus actually do differently?
Tap to flip
ANSWER
He weights the report heavily instead of treating it as decoration, since it now names its own set, version, and failure cases, all checkable in minutes.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the higher-stakes abuse case there?
Tap to flip
ANSWER
Nkechi Obi's drug-interaction checker. There, quietly excluding rare, severe interactions from the eval set produces a high score that hides exactly the failure mode that matters most.
Check yourself Score: 0 / 0
Short answer, name the role
1. In one sentence, what role does this answer say an eval report should play in a portfolio?
Show hint
Look at the direct answer at the top.
Show answer
Model answer: It should be evidence supporting one specific judgment call, not the headline of the portfolio, and it needs a named eval set, a version pin, and shown failure cases.
Multiple choice
2. Why did Petra's first eval report actually predict very little about her judgment, despite its impressive-sounding number?
A. Radiology data is always too small to eval anything meaningfully.
B. It was tested on the same scans used to tune the model's threshold, so the number didn't reflect real, unseen performance.
C. Ninety-three percent is objectively too low a number to matter.
D. Hiring managers never read eval reports closely anyway.
Show hint
Look at the "abuse" step and the story's mock-review moment.
Show answer
B. Testing on the same data used to tune the model is a classic way an eval score gets inflated without reflecting real performance.
True or false
3. True or false: this answer recommends leading a portfolio page with the eval report, ahead of the working demo.
True
False
Show hint
Look at the flow diagram showing where the eval report sits in the reading order.
Show answer
False. The eval report sits behind the demo and the judgment write-up, since it earns the most trust once a reviewer already cares about the underlying claim.
Fill in the blank
4. Fill in the blank: interview-to-offer conversion rose from 8 percent to about ___ percent once the eval report named its set, version, and failure cases.
Show hint
Look at the bar chart in the framework recap section.
Show answer
35 percent. Same underlying skill, radically different outcome, once the report became checkable rather than merely asserted.
Short answer, where it wouldn't matter
5. Name a situation where a very short, informal eval note would be genuinely fine, without full rigor.
Show hint
Think about the stakes of the underlying decision the model supports.
Show answer
Model answer: A low-stakes internal tool with no safety or fairness consequences, where a quick, honestly labeled sanity check matters more than a fully rigorous report.
Short answer, apply it yourself
6. Think of a claim you've made about something you built or tested. What eval set, version, or failure case is missing that would make it checkable instead of just asserted?
Show hint
Look for a number in your own work with no named source behind it.
Show answer
Model answer: Most people can name at least one confident claim they've made that, on inspection, has no named test set or version attached to it at all.
Before you close the answer
Why this works
Tests whether you understand that an eval report's value comes from being checkable, not impressive, and whether you'd know how to spot a report that's gamed instead of just trusting the headline number.
Follow-up traps
"Isn't a smaller, honest number worse than a bigger, vague one?" Response: no, a smaller number with a named set and version pin is more useful than a bigger one with nothing behind it, since only the first one is actually checkable.
"What if the eval set is genuinely too small to be statistically meaningful?" Response: say so plainly in the report itself; naming that limitation is still more honest and more useful than hiding it behind a single confident percentage.
If pressed
Petra's final eval report used a specific checkpoint of an open second-read vision model, dated to the exact day she ran the sixty-scan test, so anyone could re-run the identical comparison later and get the same four flagged misses.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.