CaseIntermediateResponsible AI & Advanced Practice / Building an AI PM portfolio / #18

How would you use a portfolio project to open a conversation with a hiring manager?

AUDIT the portfolio piece is Framecheck, an AI tool that scores real estate listing photos and ranks which to feature first

Larkspur Realty Group is a fictional real estate brokerage. Nikolai Bauer built Framecheck there: it scores a listing's photos for lighting, clutter, and framing, then tells an agent which one to feature first. Isolde Van Praet, a hiring manager, is the first person Nikolai practices his opening line on.

The direct answer
Before you lead with any claim, audit it the way a skeptical hiring manager would: who verified it, what eval set it's based on, which model version produced it, what's missing, and could you reproduce it live. Open with whichever claim survives all five checks, not the flashiest number. An audited claim opens a real conversation. An unaudited one invites the first hard question to shut it down.
Do this, in order
  1. List every claim you could open with before picking one.Why: you can't audit a claim you haven't named yet.
  2. Ask who verified each one, you alone, or someone else with real data.Why: a self-graded number is the first thing a sharp interviewer probes.
  3. Check whether each claim came from a real eval set or a convenient sample you happened to have.Why: a number with no real set behind it is a guess dressed up as evidence.
  4. Try to reproduce each claim on the spot before the interview, not during it.Why: if you can't reproduce it calmly at your desk, you can't defend it under pressure either.
  5. Lead with the claim that survives every check, even if it's less flashy than the others.Why: a modest, bulletproof number opens more doors than an impressive, shaky one.
  6. Keep the shakier claims in reserve, not deleted, and name their real limits if asked.Why: an honest "here's what I haven't verified yet" is still stronger than pretending a weak claim is solid.

How to answer this, stage by stage

Nobody is grading whether your opening line sounds impressive. They're grading whether it survives the very first question back.

Stage 1
Scope it to one real project, three real candidate claims
Say it like this
"I'll answer this for Framecheck, a listing-photo scorer, and the three different numbers I could have opened with."
Why this works
Grounds an abstract "how do you open a conversation" question in a real, specific choice.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT on my own claims first. Ask who verified it, uncover the eval set, demand the version pin, isolate what's missing, then test it myself before I ever say it out loud."
Why this works
Signals you're applying the same scrutiny to your own portfolio that a skeptical reviewer would.
Stage 3
Ask who verified each claim
Say it like this
"My 91% accuracy number, I graded that myself, on 50 photos I picked. My click-through number came from Larkspur's own marketing team, off a real live test."
Why this works
Separates a claim only you vouch for from one someone else already checked.
Stage 4
Uncover the eval set behind each
Say it like this
"That 91% came from a set of 50 photos I chose myself. The click-through number came from every real listing photo Larkspur published over six weeks."
Why this works
A convenient small sample and a real, uncherry-picked population are not the same kind of evidence.
Stage 5
Try to reproduce each one
Say it like this
"I could pull Larkspur's click-through numbers back up right now, live, off their own dashboard. I couldn't re-time my old 'two minutes' claim, since it was just me and a stopwatch, once, months ago."
Why this works
A claim that survives a live reproduction check is the one that survives a live follow-up question too.
Stage 6
Pick the survivor, and open with it
Say it like this
"So I open with the click-through number, not the accuracy figure, because it's the one that holds up if you push on it right now."
Why this works
Ends on a clear, defensible choice, not just a list of numbers you happen to have.

Let's learn

A listing-photo scorer is a tool that looks at a set of real estate photos and tells an agent which one is actually going to make someone stop scrolling.

Nikolai had three numbers he could open a conversation with. First: "our model hits 91% top-3 accuracy picking the best photo." Second: "listings using our top-scored photo got 31% more clicks in a live test." Third: "cut an agent's photo-picking time from 12 minutes to under 2."

Knowledge spark: what's a version pin? The exact model name and the date it produced a specific result. Without one, a number can't be checked again later, because the thing that made it might already be different.

Only one of those three numbers came from an independent source. The other two lived entirely inside Nikolai's own head and his own stopwatch.

AUDIT checks passed, by candidate opening claim (out of 5)
5 2.5 0 1 of 5 91% accuracy 2 of 5 Time saved 5 of 5 Click-through test
Only the click-through claim survives every check. The other two rely entirely on Nikolai's own word.

Here's the turn: the extra impressiveness of the 91% figure was never the point. What mattered was whether the number could survive being poked at, and the flashiest claim was the one most likely to fall apart under one direct question.

The most impressive number and the most defensible number were not the same number. Only one of them could survive being asked "how do you know that?"

At its worst: Nikolai opens with the 91% figure, an interviewer asks one plain question about his eval set, and the rest of the conversation becomes about defending a shaky number instead of discussing the actual product decisions behind Framecheck.

The decision that mattered Audit every candidate opening claim the same way a skeptical hiring manager would, and lead with whichever one survives, even if it's the least flashy of the three.

What I would leave alone: the 91% and time-saved numbers don't need to be deleted. They're fine to mention later, clearly labeled as self-tested, once the conversation has real trust behind it.

The lesson: the strongest opening line isn't the one that sounds best. It's the one that's already been tested against the exact question that's coming next.

Now here is the same thing as a story

The short version above is what you'd say explaining your opening line out loud. Read this one for how Nikolai actually picked his.

Nikolai Bauer spent three years staging homes for open houses before he ever trained a model, and he can tell within a glance which photo of a listing is going to stop someone mid-scroll.

He'd practiced his opening line a dozen times before his first real interview, always reaching for the 91% accuracy figure, since it was the biggest number he had.

Hand sketched comparison diagram titled Self-graded versus independently verified. Left panel, a document icon labeled Self-timed, caption just Nikolai twice. Right panel, a scale icon labeled Real A B test, caption Larkspur's own data.
Two numbers, both true. Only one of them was ever checked by someone other than Nikolai.

A friend doing a mock interview asked the obvious question: "How'd you get to 91%?" Nikolai explained he'd hand-picked 50 photos from listings he liked, run the model, and counted how often it agreed with his own judgment.

Hand sketched decision tree titled Reading 91 percent two ways. Root, the 91 percent claim, branching to his own curated set leads to unproven, independent eval set leads to trustworthy.
The exact same number reads two completely different ways, depending entirely on who chose the 50 photos.

His friend didn't say it was wrong. She just said, plainly, "that's your opinion of your own model, dressed up as a statistic." Nikolai sat with that for a while.

Hand sketched flow diagram titled Testing three claims before the interview. Five boxes: list candidates, ask who checked, check the set highlighted, try to redo it, pick the survivor.
Checking the eval set behind each claim is the step that actually separated a real number from a convenient one.

Going back through his own materials, he found the click-through number he'd almost forgotten about: Larkspur's marketing team had run a real, live test across six weeks of published listings, comparing click rates on listings that used Framecheck's top pick versus ones that didn't. Thirty-one percent more clicks, tracked entirely outside his own hands.

Hand sketched labeled parts diagram titled The claim that survived. Center gauge icon labeled 31 percent more clicks, with four callouts: real A B test, Larkspur's analytics, reproducible today, not self graded.
All four of these were true about the click-through number. None of them were true about the 91% figure.

Rebuilding his opening line around the click-through result took Nikolai about twenty minutes, mostly digging up the original marketing report to make sure he could still cite it accurately. Practicing it against the same hard question took another mock interview to get comfortable.

Hand sketched quadrant titled Sorting three candidate openers. Axes how verifiable and how impressive it sounds. The 91 percent claim sits flashy and shaky. Click through test sits flashy and solid. Time saved claim sits modest and shaky.
The click-through claim was the rare one that landed in the corner that actually matters: impressive, and solid.

Leading with the 91% number had felt natural, since it was the biggest, roundest figure Nikolai had. It stopped feeling like the right choice the moment one honest friend asked where it actually came from.

I opened with my biggest number because it was the most impressive one I had. It took one plain question from a friend to see that impressive and defensible were never the same test.

AUDIT, turned on your own opening lineNot a critique of someone else's report. AUDIT is what tells you which of your own claims is safe to lead with.

A
Ask who verified it.
The 91% figure: Nikolai alone. The click-through number: Larkspur's own marketing team.
A self-graded claim and an independently checked one are not the same kind of evidence.
U
Uncover the eval set.
50 hand-picked photos versus every real listing photo published over six weeks.
A convenient sample and a real population tell very different stories.
D
Demand the version pin.
Nikolai couldn't name which exact model version produced the 91% number months later. He could pull the click-through report by date.
A number nobody can trace back to a specific version can't really be re-checked.
I
Isolate what's missing.
The 91% figure had no failure case attached. The click-through number came with its own real methodology, right in the report.
A headline number with nothing behind it is decoration, not evidence.
T
Test it yourself.
Nikolai could reproduce the click-through result live, off Larkspur's dashboard. He couldn't re-run his own stopwatch test the same way.
The hardest check, and the one that actually decides which claim survives a live follow-up.
Hand sketched icon list titled AUDIT applied to your own claim. Items: who verified it, what eval set, which model version, what is missing, can you reproduce it live.
Run every candidate opening line through these five questions before you ever say one out loud.
Interviewer engagement over the first eight minutes, by opening choice
10 5 0 Min 0 Min 2 Min 4 Min 6 Min 8 Audited opener Unaudited opener
Both openers start strong. Around minute four, the unaudited one meets a real follow-up question and falls apart. The audited one keeps climbing.

The recap, one line per letter: ask who verified it separates self-graded from independently checked, uncover the eval set separates a curated sample from a real population, demand the version pin tests whether a number can be traced back, isolate what's missing checks for a stated limit, and test it yourself is the final, hardest check.

And if you want to be sure it really works, try it somewhere elseSame five letters, a fisheries co-op instead of a real estate brokerage. A different building, and this time the flashiest number is the one an interviewer already suspects.

Brackwater Fisheries Collective is a fictional fishing co-op. Halvard Espen built an AI tool there that predicts the best days to bring boats out based on weather and catch history. Junko Arai reviews his opening claim.

Mapped onto AUDIT: Halvard had two candidate openers, "our model beat the co-op's own forecast by 22% on trip yield" and "fishers using our recommendation saved an average of 40 minutes of planning time each morning." Asking who verified each: the yield number came from the co-op's own catch-logging system, tracked independently of Halvard. The time-saved number came only from three fishers he personally interviewed. Uncovering the eval set: the yield comparison ran across an entire fishing season's real trips, while the time-saved figure came from three conversations. Demanding the version pin: the yield number was tied to a specific model version locked before the season started, fully traceable. Isolating what's missing: the time-saved claim had no stated sample size in Halvard's notes until he went looking for it. Testing it himself: he could pull the yield comparison straight from the co-op's own records, live, in front of anyone. He led with the yield number, and kept the time-saved story as a supporting anecdote instead.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "here's the one number I could pull up and defend live, right now," and open with that.
Cost: there's no time to build a full independent test before tomorrow's interview. Say plainly which numbers are self-tested and which aren't, rather than let one blur into the other.
The model gets better, for real: if Framecheck's accuracy improves later, that's still worth mentioning, but it doesn't change which claim you open with, since the audit is about verification, not about how good the number is.

Where people run it wrong.
They lead with the biggest number instead of the most defensible one, and get stopped by the very first follow-up question.
They assume any claim they can state confidently is safe to open with, without ever checking who actually verified it.
They delete the shakier claims entirely instead of keeping them in reserve, clearly labeled, for later in the conversation.

How to use it live. When someone asks how to open a conversation with a hiring manager, ask yourself one question first: which of my claims could I reproduce right now, on the spot, if asked. Open with that one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you open a conversation with a hiring manager using your portfolio"?
Tap to flip
ANSWER
AUDIT, turned on your own claims: ask who verified it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Nikolai Bauer, a former home stager who built Framecheck, a listing-photo scorer, at Larkspur Realty Group.
3 · THE HABIT
Which number did Nikolai default to opening with, and why?
Tap to flip
ANSWER
His 91% accuracy figure, simply because it was the biggest, roundest number he had, without checking who else could vouch for it.
4 · THE MECHANISM
What made the click-through claim survive every audit check?
Tap to flip
ANSWER
It came from Larkspur's own marketing team, off a real six-week test across every published listing, not from Nikolai's own judgment.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Practicing his opening line around the 91% figure a dozen times before ever asking who else had checked it.
6 · THE NUMBER
Fill in the blank: the click-through claim passed ___ of 5 audit checks, versus 1 of 5 for the accuracy claim.
Tap to flip
ANSWER
5 of 5. The time-saved claim passed 2 of 5, landing in the middle.
7 · THE REPLAY
Same interview, new opening line. What changes over the first eight minutes?
Tap to flip
ANSWER
Engagement climbs steadily instead of spiking early and collapsing once a real follow-up question hits the unverified number.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what does the candidate lead with?
Tap to flip
ANSWER
Halvard Espen's fishing-trip prediction tool at Brackwater Fisheries Collective. He leads with the co-op-verified yield number, not his own time-saved anecdote.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at what Nikolai practiced saying before his mock interview.
Show answer
Model answer: Leading with the 91% figure because it was his biggest number. It made sense since he'd never had a reason to doubt his own eval set until someone else asked about it.
Multiple choice
2. Why does this answer recommend opening with the click-through number instead of the 91% accuracy figure?
  • A. Because 31% is a bigger number than 91%.
  • B. Because the click-through number was independently verified, drawn from a real population, and reproducible live, while the accuracy figure was self-graded on a hand-picked sample.
  • C. Because accuracy numbers are never appropriate to mention in an interview.
  • D. Because Larkspur required him to use their own metric by policy.
Show hint
Look at the bar chart of audit checks passed.
Show answer
B. Bigger isn't the same as more defensible. The click-through claim survives scrutiny because of who checked it and how.
True or false
3. True or false: this answer says the 91% accuracy figure should be deleted from Nikolai's portfolio entirely.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. It stays in the portfolio, just not as the opening line, and clearly labeled as self-tested if it comes up later.
Short answer, where it wouldn't matter
4. Name a moment in the conversation where a self-graded, unverified number is still fine to mention.
Show hint
Think about timing, not just content.
Show answer
Model answer: Later in the conversation, once real trust is established, clearly labeled as self-tested rather than presented as independently verified.
Short answer, apply it yourself
5. Think of a project you'd open an interview conversation with. List two numbers you could lead with, then ask yourself: which one could you reproduce live, right now, if asked?
Show hint
Try to actually name the two numbers, not just the general idea of "a good result."
Show answer
Model answer: Most people find one number they personally calculated once and one that an outside system or person already tracks. The second one is almost always the safer opener.
Short answer, the number question
6. If Nikolai's click-through test had only run for one week instead of six, would it still be the stronger opener? Why or why not?
Show hint
Think about which of the five audit checks depends on sample size.
Show answer
Model answer: Likely still stronger than the 91% figure, since it would still be independently verified and reproducible, but he'd need to state the shorter window honestly rather than imply six weeks of data.
Before you close the answer
Why this works
Tests whether you can apply the same skepticism to your own claims that you'd apply to someone else's report, and whether you understand that an opening line's real job is to survive the first question, not just sound impressive.
Follow-up traps
"Isn't the 91% number still worth mentioning at some point?" Response: yes, later, clearly labeled as self-tested, once the conversation has enough trust built up that a shakier number won't derail it.

"What if you genuinely have no independently verified number at all?" Response: then open with the claim you can reproduce most confidently on the spot, and say plainly that it's self-tested, rather than let a reviewer assume it's more than it is.
If pressed
Larkspur's real click-through test ran as a genuine randomized split, half of new listings got Framecheck's top-photo recommendation and half didn't, over the same six-week window, which is exactly the kind of design that makes a 31% difference mean something instead of just reflecting which listings happened to be better already.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more