ConceptFoundationalResponsible AI & Advanced Practice / Building an AI PM portfolio / #2

Describe the three artifacts that make the strongest AI PM portfolio.

SPARK Adaeze Nwafor is building a portfolio around an injury-risk feature for Pacefinder, a running-coaching app she has used herself for three years

Adaeze Nwafor spent six years as a physical therapist before starting to teach herself AI product work at night. Her sample project is a feature for Pacefinder, a running-coaching app: a model that flags when a runner's training load looks like it's heading toward an injury. Grant Ferriera is Head of Product at a different fitness-tech company, and he's the one deciding whether to bring her in for a first call.

The direct answer
Build exactly three things: one small feature you actually shipped and can demo live, one page where you name a real judgment call you made and defend it, and one short log showing you tested the model against real cases instead of trusting it on sight. A resume claims competence. Those three things let a stranger check it themselves.
Do this, in order
  1. Build and ship one small feature someone can actually click and run.Why: it's the only artifact that proves you can finish something, not just describe it.
  2. Write one page naming a real trade-off you chose, and why.Why: this is the only artifact that shows judgment, not just execution.
  3. Keep one short eval log, real cases, not a single accuracy number.Why: it's the only artifact that proves you actually checked the model instead of trusting it.
  4. Cut anything that's polish without proof, like a slide deck with no working link.Why: a fourth artifact that just looks nice dilutes the three that actually carry weight.
  5. Make sure the build survives being wrong at least once, visibly.Why: a demo that only works when the model is right hasn't proven anything about your judgment.
  6. Keep the scope small enough to finish in a few weekends.Why: three finished small things beat one half-built ambitious one.

How to answer this, stage by stage

Nobody's checking whether you can list ten things a portfolio could contain. They're checking whether you can commit to the three that actually matter and say why the rest don't make the cut.

Stage 1
Scope it to one real candidate
Say it like this
"I'll answer this for someone building a portfolio from scratch, aiming at an AI PM role, with maybe six weekends of real time to spend."
Why this works
Grounds "what makes a strong portfolio" in a real, limited amount of time and effort.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, what they're doing today. Payoff, the habit I want a reviewer to build. Anchor, the actual artifacts. Risk, what breaks. Keep out, what I won't build."
Why this works
Signals you're designing the portfolio on purpose, not listing things that felt good to add.
Stage 3
Name today's situation
Say it like this
"Today, most candidates have a resume and a cover letter. Both ask the reader to just trust the claims inside them."
Why this works
Shows what's actually missing before jumping to the fix.
Stage 4
Give the anchor: the three artifacts
Say it like this
"One small shipped build you can click. One judgment write-up naming a trade-off. One eval log against real cases. That's the whole portfolio."
Why this works
This is the concrete answer to the actual question, stated plainly.
Stage 5
Prove it survives being wrong
Say it like this
"Adaeze's demo flags a runner's risk, and when it's wrong, it says so and shows why, instead of just being quietly wrong and hoping nobody notices."
Why this works
Shows the anchor was designed to survive its own failure, not just to look good in a happy-path demo.
Stage 6
Say what you'd leave out, and close
Say it like this
"I wouldn't build a fourth artifact, like a slide deck, day one. Three finished small things beat four half-finished impressive ones. The build, the judgment page, the eval log: that's the portfolio."
Why this works
Restates the anchor and shows judgment about scope, not just ambition.

Let's learn

Picture a portfolio built by someone with no AI job title yet, trying to prove they can think like an AI product manager anyway.

Before Adaeze built anything, her application was a resume that said "improved patient outcomes through data-informed care plans," and a cover letter that said she was "excited about applying AI to health and fitness." Grant reads roughly fifteen of these a week, and none of that told him anything he could check.

Knowledge spark: what's a judgment write-up? A short page, not a case study, where a candidate names one real decision they made building something, like choosing to flag a runner early even at the cost of some false alarms, and says why. It's the part of a portfolio that shows thinking, not just output.

After she built the three artifacts, the same resume line stayed, but underneath it sat a live demo of the injury-risk flag, a one-page write-up titled "why I chose to flag early, not late," and a short log of twenty real training-load cases she'd tested the model against, including the four it got wrong.

What a reviewer can verify, artifact by artifact
100% 50% 0 Resume alone + shipped build + judgment page + eval log 5% 60% 80% 95%
Each artifact adds something a reviewer can actually check for themselves, not just take on faith.

At its worst: a candidate spends a full month building a slick, animated slide deck about their "AI philosophy," and a reviewer closes it in under a minute, because a philosophy isn't something you can click.

The anchor: the three artifacts One small shipped build a reviewer can run themselves. One page naming a real trade-off and why it was chosen. One eval log showing the model tested against real cases, including the ones it got wrong. Nothing else in the portfolio needs to exist for a first call to happen.

What I would leave alone: a short resume and a two-line summary at the top still matter. They're what gets someone to click through to the three artifacts in the first place. The fix isn't deleting them, it's making sure they're not the only thing there.

A resume tells a reviewer what you say you did. Three built artifacts let them check it in the time it takes to drink a coffee.

The lesson: a portfolio isn't a longer resume with pictures. It's the difference between asking to be believed and giving someone something to test.

Now here is the same thing as a story

The short version above is what you'd say defending your own portfolio choices out loud. Read this one for how Adaeze actually built hers.

For six years, Adaeze could watch a runner's gait for ten seconds and tell you which knee was about to become a problem, long before an X-ray would show anything.

She spent her first month of self-taught AI work reading everything, and building nothing, convinced she needed to understand transformers before she was allowed to touch a real project.

Hand sketched comparison diagram titled Two kinds of proof. Left panel, a document icon labeled TELLS, caption trust my word. Right panel, a gauge icon labeled SHOWS, caption check it yourself.
A resume tells. A build shows. Adaeze had spent a month getting better at telling.

A friend who'd just been hired as an AI PM looked at her draft resume and asked one plain question: "What's the thing I could click on right now?" There wasn't one.

That was the whole trigger. No catastrophe, just one honest question with no answer behind it.

Hand sketched timeline titled Six weekends to three artifacts. Four milestones: Weekends 1 and 2, build the flag feature, highlighted. Weekend 3, write the judgment call. Weekend 4, build the eval log. Weekends 5 and 6, polish, ship, share.
Six weekends, spent on three finished things instead of one endless one.

She picked the smallest real problem she actually understood: a running app flags injury risk with one confidence number, and nobody explains what to do with it. She built a version that explains the flag and lets you dismiss it with a reason.

Hand sketched labeled parts diagram titled The anchor, close up. Center gauge icon labeled Injury-Risk Flag, with four callouts: the score, the reason why, a way to contest it, what if it's wrong.
Four small parts, and the third one, a way to contest it, is the part most candidates skip entirely.

The build wasn't going to be right every time, so she made sure it survived that on purpose.

Hand sketched comparison diagram titled The day the model is wrong. Left panel, a question mark box icon labeled Silent bad call, caption runner keeps going. Right panel, a person icon labeled Flagged, explained, caption runner checks with coach.
A wrong call that says nothing sends a runner straight into the next hard workout. A wrong call that explains itself sends them to their coach instead.

Then she wrote the judgment page: why she chose to flag early and eat some false alarms, rather than flag late and risk missing a real injury. And the eval log: twenty real training weeks pulled from public running forums, with the four the model got wrong written up honestly, not hidden.

Hand sketched decision tree titled What goes in, what waits. Root: Building the portfolio. Four branches: proves real judgment leads to Ship it. Just looks polished leads to Cut it. Needs a full team leads to Not day one. Fits one weekend leads to Build it.
Every idea she had got sorted through the same four questions before it earned a place in the six weekends.
Hand sketched icon list titled The three artifacts. Three items: a box icon labeled One small shipped build, a document icon labeled One judgment write-up, a gauge icon labeled One eval log, real cases.
Three items. Everything else she'd built or planned to build got measured against whether it belonged on this exact list.

The old portfolio asked Grant to believe a paragraph. The new one hands him a running demo, a real trade-off, and twenty test cases, and asks him to spend ten minutes instead of trusting a sentence.

I spent a whole month reading about model architecture because it felt like the responsible thing to do before building anything real. Watching a friend ask "what can I click" is what showed me the responsible thing was actually to just build the smallest real version and let it be checked.

SPARK, for a portfolio of exactly three thingsNot a longer list. The one anchor decision underneath all three.

S
Situation. What's happening today, without this.
Most candidates have a resume and a cover letter, both asking to be trusted rather than checked.
Names the real gap before proposing the fix.
P
Payoff. The habit you want built.
You want a reviewer to stop reading claims and start clicking proof, in under ten minutes.
The habit, not the hours saved, is the actual thing being designed for.
A
Anchor. The three artifacts.
One small shipped build. One judgment write-up. One eval log against real cases. This is the direct answer to the question.
The hardest step, and the one everything else in the answer hangs on.
R
Risk. What breaks the first time you're wrong.
The demo has to survive the model being wrong, visibly, not just look good in a rehearsed happy path.
Designs the anchor to hold up under exactly the question a good interviewer will ask.
K
Keep out. What you won't build.
No fourth artifact, like a polished slide deck, before the first three are actually finished.
Shows judgment about scope, which is exactly what the three artifacts are meant to prove in the first place.
Time a candidate spent, by artifact, across six weekends
24h 12h 0 Shipped build 24h Judgment page 6h Eval log 10h Slide deck 0h
The slide deck she almost built would have taken about as long as the eval log, for something no reviewer could click.

The recap, one line per letter: situation is a resume nobody can check, payoff is a reviewer clicking instead of trusting, anchor is the three artifacts, risk is the build surviving a wrong call in public, and keep out is refusing a fourth artifact before the first three are real.

And if you want to be sure it really works, try it somewhere elseSame five letters, an elder-care check-in app instead of a running app. A different candidate, and this time the risk is far more serious than a missed workout.

Baris Aydin is building a portfolio project around a daily check-in call assistant for an elder-care service, flagging when an older adult's voice patterns suggest they should get a nurse follow-up instead of a routine callback. Applied to SPARK: situation is that most portfolios in this space describe "AI for elder care" in the abstract, with no working example. Payoff is getting a reviewer to trust that Baris can be careful with a vulnerable population, not just skilled with a model. Anchor is the same three artifacts: a small shipped call-flagging demo, a one-page write-up on choosing to over-flag rather than under-flag, and an eval log of forty real call transcripts including six flagged wrongly. Risk is what happens when the model wrongly clears someone who needed a follow-up, which is why the demo shows a human nurse callback as the fallback, not a dead end. Keep out is a full voice-AI pipeline with real-time transcription, which Baris deliberately left for "not day one," using recorded sample calls instead.

Hand sketched labeled parts diagram reused here for a different anchor: a central icon labeled the flagging feature, with callouts for the score, the reason why, a way to contest it, and what happens if it's wrong.
The same four-part anatomy holds up in elder-care too. Only the stakes behind "what if it's wrong" changed.

Swap the trigger and it still runs.
Speed: an interviewer gives you two minutes to describe your portfolio. Say "one build, one judgment page, one eval log," and stop.
Cost: you can't afford real user data to test against. Use public data or recorded samples honestly labeled as such, rather than skipping the eval log entirely.
The bar gets higher, for real: if more candidates start shipping real builds, the three artifacts still hold, they just need to be a little more polished to stand out, not replaced with something else.

Where people run it wrong.
They build one impressive thing and call it a portfolio, skipping the judgment write-up that shows why they built it that way.
They write extensive philosophy pages about AI ethics with nothing shipped underneath them.
They hide the cases their model got wrong instead of including them in the eval log.

How to use it live. When asked what your portfolio should contain, count on your fingers as you answer: one build, one judgment call, one eval log. Naming exactly three, out loud, is more convincing than describing ten possible pieces.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "describe the three artifacts that make the strongest AI PM portfolio"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. The anchor step names the three artifacts directly.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Adaeze Nwafor, a former physical therapist building a portfolio around an injury-risk flag for Pacefinder, and Grant Ferriera, the Head of Product reviewing it.
3 · THE HABIT
What habit did Adaeze have to stop, before she could build anything real?
Tap to flip
ANSWER
She stopped reading about model architecture as a way of delaying building, and started with the smallest real feature she actually understood.
4 · THE ANCHOR
Name the three artifacts, in order.
Tap to flip
ANSWER
One small shipped build a reviewer can run, one page naming a real trade-off, and one eval log against real cases, including the ones the model got wrong.
5 · THE RISK
What decision made sure the build survived being wrong?
Tap to flip
ANSWER
Designing the flag to explain itself and offer a way to contest it, instead of quietly giving a wrong answer with no next step.
6 · THE NUMBER
Fill in the blank: adding the eval log took a resume claim from about 5 percent checkable to about ___ percent checkable.
Tap to flip
ANSWER
95 percent. The shipped build alone got it to 60 percent, and the judgment write-up added another 20 points.
7 · THE REPLAY
Same interview, redesigned portfolio. What does Grant actually do differently?
Tap to flip
ANSWER
He clicks the demo, reads the judgment page, and checks two of the four wrong cases in the eval log, all in about ten minutes, instead of reading a paragraph and moving on.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's different about the risk?
Tap to flip
ANSWER
Baris Aydin's elder-care check-in flagging tool. There, a wrong call risks missing a genuinely vulnerable person, which is why the fallback is a human nurse callback, not a dead end.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer says a strong AI PM portfolio should include as many polished write-ups as possible.
  • True
  • False
Show hint
Look at the K, keep out, step.
Show answer
False. It says exactly three artifacts, and explicitly warns against a fourth one, like a slide deck, that just looks polished without proving anything.
Multiple choice
2. Which of the three artifacts is most often skipped by candidates, according to this answer?
  • A. The shipped build
  • B. The resume
  • C. The eval log showing where the model was wrong
  • D. The cover letter
Show hint
Look at "where people run it wrong."
Show answer
C. Candidates often hide the cases their model got wrong instead of including them honestly in an eval log.
Fill in the blank
3. Fill in the blank: in the six-weekend build, the shipped feature took about ___ hours, more than the judgment page and eval log combined.
Show hint
Look at the bar chart of hours spent by artifact.
Show answer
24 hours. The judgment write-up took about 6 hours and the eval log about 10, together still less than the build alone.
Short answer, name the anchor
4. Name the three artifacts this answer says make up the anchor of a strong portfolio.
Show hint
Look at the key-point block titled "the anchor."
Show answer
Model answer: One small shipped build you can run, one page naming a real trade-off and why, and one eval log tested against real cases, including the ones the model got wrong.
Short answer, where it wouldn't matter
5. Name a situation where a candidate might reasonably skip building a fourth polished artifact even under pressure to "stand out more."
Show hint
Think about what happens to attention when there are more than three things to review.
Show answer
Model answer: Almost always. A fourth artifact rarely adds proof a reviewer will actually reach, and it risks diluting attention away from the three that carry the real weight.
Short answer, apply it yourself
6. Think of a skill you want to prove you have. What's one small thing you could build or do this weekend that would let someone check it, instead of just taking your word for it?
Show hint
Think about the smallest real version of the skill, not the most impressive one.
Show answer
Model answer: Most people can find a version small enough to finish in a weekend once they stop trying to prove the whole skill at once and pick one real, narrow case instead.
Before you close the answer
Why this works
Tests whether you understand that a portfolio's job is to be checkable, not impressive, and whether you can commit to a small, finished set of proof instead of an ever-expanding list of things that sound good.
Follow-up traps
"What if I don't have three weekends to spare?" Response: build the eval log first, even a rough one, since it's the cheapest of the three and it's the one candidates skip most, so it stands out fastest.

"Isn't a slide deck useful for walking someone through your thinking?" Response: only after the three artifacts exist. A deck describing a build that isn't there yet is exactly the "tells, doesn't show" problem this answer is built to avoid.
If pressed
Adaeze's eval log used a threshold she could defend by name: flagging anything above a 60 percent risk score, chosen because it caught 90 percent of real injuries in her twenty test cases while only over-flagging on six of them.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more