Artifact critiqueIntermediateResponsible AI & Advanced Practice / Building an AI PM portfolio / #5
Critique a portfolio built entirely from case study write-ups with no build.
PICK Callum Whitfield's portfolio has six write-ups about claims-triage work at Anchor and Vale Insurance, and not one working link. Noor Haddad is reviewing it.
Callum Whitfield spent three years on the claims operations team at Anchor and Vale Insurance before deciding to move into AI product work. His portfolio holds six detailed case-study write-ups about triage and fraud-flagging ideas from that job, no code, no demo, nothing clickable. Noor Haddad reviews portfolios for a hiring committee and reads Callum's as part of a first pass.
The direct answer
This portfolio is weaker than it needs to be: strong writing, zero proof. Cut it to two case studies and spend the freed-up time building one small, rough, real thing, even something far smaller than what the write-ups describe. A shipped build that breaks in front of a reviewer is worth more than a sixth polished story about work nobody can check.
Do this, in order
Build one small, real thing before writing a seventh case study.Why: a portfolio with zero shipped builds can't prove anything a resume line couldn't already claim.
Cut the six write-ups down to two, and let the build carry the rest.Why: past a reviewer's third write-up, attention is already gone, so the fourth through sixth aren't earning their space.
Pick the smallest real problem from the six ideas, not the most impressive one.Why: a small finished build beats an ambitious unbuilt one every time a reviewer actually opens the page.
Make sure the build can fail visibly, in front of the reviewer.Why: a build's errors are cheap and caught immediately; a write-up's errors are invisible until it's too late.
Keep the remaining write-ups honest about what was actually shipped versus proposed.Why: blurring "I designed this" with "I built this" is exactly the gap a build is meant to close.
Only skip the build entirely if the candidate has years of independently verifiable shipped work already.Why: for someone with that track record, a toy demo can look condescending rather than additive.
How to answer this, stage by stage
This isn't a writing critique. It's a test of whether you understand why a stack of well-argued essays isn't the same thing as evidence.
Stage 1
Scope it to the actual portfolio in front of you
Say it like this
"I'll critique this specific case: six write-ups about claims-triage ideas, no build, no demo, no repo."
Why this works
Grounds the critique in the real artifact instead of a general lecture about portfolios.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my stance on write-ups versus builds. Impact, who's hurt by each. Cost asymmetry, which error is worse. Kill criteria, when this portfolio would actually be fine as-is."
Why this works
Signals a structured critique, not just a gut reaction to the writing quality.
Stage 3
State your position, before any reasoning
Say it like this
"My position: six write-ups and zero builds is a weaker portfolio than two write-ups and one small real build, full stop."
Why this works
Commits to a clear stance immediately, instead of hedging into "it depends on the writing quality."
Stage 4
Name who's hurt by each option
Say it like this
"Noor can't verify anything in a write-up in the time she has. Callum, meanwhile, spent real hours polishing prose that a reviewer skims past by write-up three."
Why this works
Shows the cost lands on both sides, not just as a one-sided complaint about the reviewer.
Stage 5
Name the cost asymmetry
Say it like this
"A build's errors are visible and cheap: it crashes, you see it, you fix it. A write-up's errors are hidden and expensive: an overstated claim about judgment gets absorbed by good prose and never gets caught until a bad hire, or never."
Why this works
This is the heart of PICK: naming which error is the one worth optimizing against.
Stage 6
Prove it with a failure
Say it like this
"Write-up four claims Callum 'would reduce false positives by prioritizing claim age.' There's no way to check that. If it had been built, even roughly, we'd know within a minute whether the idea actually held up."
Why this works
Shows the exact cost of the missing build, not just the abstract principle.
Stage 7
Give kill criteria, and close
Say it like this
"If Callum had ten years of independently verifiable shipped AI work already, a toy build here would look condescending, not additive. He doesn't yet, so the fix stands: cut to two write-ups, build one small real thing."
Why this works
Shows the critique has a real limit, and states plainly why it applies in this specific case.
Let's learn
Picture a portfolio that reads well, argues well, and proves nothing at all.
Before anyone pointed it out, Callum's page opened with six case studies, each about 800 words, each describing a triage or fraud-flagging idea he'd proposed at Anchor and Vale. Every one ended with a version of "this would have reduced review time significantly."
Knowledge spark: what's the difference between a case study and a build?
A case study describes what you would do, or say you did, in prose. A build is something that exists and runs, right now, that a stranger can open and test for themselves. One is a claim. The other is evidence.
Noor opened write-up one, read it fully, skimmed write-up two, and by write-up three had already formed her opinion: smart writer, no way to know if any of it actually worked.
How far Noor's attention actually reaches, write-up by write-up
Three of Callum's six write-ups get zero attention. The time spent polishing them was, functionally, wasted.
At its worst: Noor recommends passing on Callum for a role he was probably qualified for, not because his ideas were bad, but because nothing on the page let her tell good ideas from merely well-written ones.
The cost asymmetry that matters here
A build's errors are cheap and visible: it breaks, you see it, you fix it in front of everyone. A write-up's errors are hidden and expensive: an inflated claim about judgment gets absorbed by confident prose, and it isn't caught until a bad hire, or it's never caught at all. Optimize against the hidden one.
What I would leave alone: Callum's writing itself is genuinely good, and that's worth keeping. Two well-chosen write-ups, trimmed to the ones that pair with the actual build, still carry real value. The problem was never his prose. It was that prose was asked to do a job only a build can do.
A write-up's mistake hides inside a well-written paragraph. A build's mistake happens on screen, in front of the person deciding whether to hire you.
The lesson: six stories about judgment aren't six times as convincing as one. Past the third one, they're not read at all, and the one thing that would have been read, a working build, was the one thing missing.
Now here is the same thing as a story
The short version above is what you'd say critiquing this portfolio out loud. Read this one for how the critique actually landed.
Callum could spot a fraudulent claim pattern in a stack of paperwork faster than most of his old team, three years of pattern-matching nobody had to teach him twice.
He spent five careful weekends turning his best ideas from that job into six polished write-ups, proud of each one, certain the writing itself would carry the argument.
One kind of mistake happens where everyone can see it. The other kind hides behind good writing.
He submitted the portfolio to eleven roles over six weeks and heard back from exactly one, a polite pass with no real feedback attached.
Attention runs out well before the sixth write-up. Callum had no way to see that from the outside.
Noor, doing a portfolio-review favor for a mutual contact, gave him the plainest feedback he'd gotten: "I read your first two case studies. I have no idea if any of these ideas actually work."
Every path through the old portfolio ends the same way: nothing left to click.
He picked the smallest of his six ideas, a rule that flagged claims for review based on claim age and adjuster notes, and rebuilt a rough version over one real weekend using a public dataset of synthetic insurance claims.
Four gaps, all in the same write-up, and all four disappear the moment something real gets built instead.
The six write-ups all sit in the same corner: bold claims, nothing behind them. One small build sits somewhere none of them do.
Cut down to two write-ups plus the one working build, the new portfolio lets Noor click the tool, run three sample claims through it, and see it flag one wrong, with Callum's note on why.
One small build did what six write-ups couldn't: gave a reviewer something to actually test.
The old portfolio asked Noor to take six claims of judgment on faith. The new one hands her one small tool, lets her break it, and shows her exactly what Callum did when it broke.
I believed good writing could stand in for evidence, because writing was the skill I trusted myself to prove. Hearing Noor say she had no idea if any of it worked is what showed me that trust was never going to transfer through a paragraph, no matter how well it was written.
PICK, on write-ups versus a real buildNot a style question. A question about which kind of mistake you can actually see.
P
Position. Stated before any reasoning.
Six write-ups and zero builds is a weaker portfolio than two write-ups and one small, real build.
The hardest step, and the direct answer to the whole critique.
I
Impact. Who's hurt by each option.
Noor can't verify anything in prose alone within her review window. Callum spent real hours on writing that goes unread past the third page.
Shows the cost lands on both sides of the exchange, not just one.
C
Cost asymmetry. Which error is worse.
A build's mistake is visible and cheap, caught the moment it happens. A write-up's mistake is hidden and expensive, absorbed by confident prose until it's too late to matter.
The heart of the critique: optimize against the hidden, expensive error.
K
Kill criteria. What would change this.
A candidate with ten years of independently verifiable shipped AI work doesn't need a toy build; it would look condescending relative to their real body of work.
Names the real limit of the critique, rather than treating "always build something" as an absolute rule.
Callback rate before and after the rebuild, same eleven applications' worth of roles
Fewer write-ups, one small real build, and roughly five times the callback rate on a similar batch of applications.
The recap, one line per letter: position is that a build beats a write-up, stated up front; impact is the cost landing on the unread pages as much as the unconvinced reviewer; cost asymmetry is the hidden expense of an unverifiable claim against the cheap visibility of a build's own mistakes; and kill criteria is a candidate with years of independently verifiable work who genuinely doesn't need a toy demo.
And if you want to be sure it really works, try it somewhere elseSame four letters, a waste-management routing portfolio instead of insurance claims. A different candidate, and this time even one build wasn't quite enough on its own.
Beatrix Almeida had the opposite problem: three shipped builds in her portfolio, each a small route-optimization tool for waste-collection trucks, and zero write-ups explaining the judgment behind any of them. Applied to PICK: position is that three unexplained builds are still weaker than one build paired with one clear judgment write-up. Impact: a reviewer can click Beatrix's tools, but has no way to know which trade-off she chose on purpose versus which one she just happened to land on. Cost asymmetry: an unexplained build's hidden cost is that a reviewer assumes the easy, safe choice was made, when actually a harder, more interesting one was, and that judgment goes uncredited. Kill criteria: if a build is simple enough that its logic is fully self-evident on sight, a write-up adds little, but none of Beatrix's three were that simple.
The same four gaps show up in reverse here: three real builds, but no named trade-off, no failure case explained, nothing in Beatrix's own words.
Swap the trigger and it still runs.
Speed: an interviewer wants your critique in one sentence. Say "strong writing, zero proof, fix the ratio," and stop.
Cost: the candidate genuinely has no time to build anything before an application deadline. Submit the two strongest write-ups honestly labeled as proposals, not as case studies implying they were built.
The candidate's experience grows, for real: even after years of real shipped work, at least one build-plus-write-up pairing stays valuable, since it's still the fastest way for an unfamiliar reviewer to trust a specific, current claim.
Where people run it wrong.
They assume more write-ups signal more depth, when past the third one a reviewer has usually already stopped reading.
They let a case study's language blur "I proposed this" with "I built this," which is exactly the gap a build is meant to close.
They build something impressive with no explanation of the judgment behind it, leaving a reviewer to guess at what was actually a hard decision.
How to use it live. When critiquing any write-up-only portfolio, count the shipped builds first, before reading a single case study. If that number is zero, say so immediately and specifically, then read the writing for what it argues, not for what it proves.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits critiquing a portfolio built entirely from case-study write-ups with no build?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Position states up front that a build beats a write-up.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Callum Whitfield, whose portfolio held six claims-triage case studies from Anchor and Vale Insurance with no build, and Noor Haddad, the reviewer critiquing it.
3 · THE HABIT
What habit did Callum have to drop, once the critique landed?
Tap to flip
ANSWER
Trusting good writing alone to carry the weight of proof, instead of building even a small, rough version of one idea.
4 · THE ASYMMETRY
Which is worse: a build's visible mistake, or a write-up's hidden one?
Tap to flip
ANSWER
The write-up's hidden mistake. It gets absorbed by confident prose and isn't caught until a bad hire, or never at all, while a build's mistake is caught immediately.
5 · THE OLD DECISION
What decision would you take back, if redoing this portfolio from scratch?
Tap to flip
ANSWER
Spending five weekends polishing six write-ups instead of spending one weekend building a small, rough, real version of the best idea.
6 · THE NUMBER
Fill in the blank: three of Callum's six write-ups got about ___ percent of a reviewer's attention.
Tap to flip
ANSWER
0 percent. Write-up 3 only got about 20 percent read before the reviewer stopped entirely.
7 · THE REPLAY
Same eleven-application batch, redesigned portfolio. What changes?
Tap to flip
ANSWER
Callback rate goes from 1 out of 11 to 5 out of 12 on a similar batch, once the portfolio drops to two write-ups plus one small, working build.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different candidate. Who, and what was their opposite problem?
Tap to flip
ANSWER
Beatrix Almeida, whose waste-routing portfolio had three shipped builds and zero write-ups, leaving her real judgment calls uncredited and unexplained.
Check yourself Score: 0 / 0
Multiple choice
1. According to this critique, why is a write-up's overstated claim more dangerous than a build's visible bug?
A. Write-ups take longer to produce than builds.
B. Builds are always technically more impressive.
C. The write-up's mistake is hidden by confident prose and often goes uncaught until it's too late.
D. Reviewers never read write-ups at all.
Show hint
Look at the cost asymmetry key-point block.
Show answer
C. A build's error is visible and cheap; a write-up's error is hidden and expensive precisely because good writing absorbs it.
True or false
2. True or false: this critique says every candidate must always include at least one shipped build, no exceptions.
True
False
Show hint
Look at the K, kill criteria, step.
Show answer
False. A candidate with years of independently verifiable shipped work already doesn't need a toy build, which could even look condescending.
Fill in the blank
3. Fill in the blank: after the rebuild, Callum's callback rate went from 1 out of 11 to ___ out of a similar batch of 12 applications.
Show hint
Look at the callback rate chart in the framework recap section.
Show answer
5. Roughly five times the callback rate, from cutting four write-ups and adding one small, working build.
Short answer, name the position
4. State this critique's position on write-ups versus builds, in one sentence.
Show hint
Look at the direct answer and the P step.
Show answer
Model answer: Six polished write-ups with no build is a weaker portfolio than two write-ups paired with one small, real, working build.
Short answer, where it wouldn't matter
5. Name a candidate profile where this critique would apply less strongly.
Show hint
Think about who already has an independently verifiable track record.
Show answer
Model answer: A candidate with a decade of shipped, publicly attributable AI product work, where a small toy demo would add little and might even undersell their real experience.
Short answer, apply it yourself
6. Look at something you've written to describe your own work, like a resume bullet or a project summary. Is there a claim in it that a small, real build could prove instead of just asserting?
Show hint
Look for a phrase like "would improve" or "was designed to," which usually signals a proposal rather than a proven result.
Show answer
Model answer: Most people find at least one claim phrased as a proposal in disguise, something they could turn into a small working example instead of leaving as a described idea.
Before you close the answer
Why this works
Tests whether you can separate the quality of an argument from the existence of evidence for it, and whether you'd know how to fix a real portfolio instead of just describing the problem in the abstract.
Follow-up traps
"What if the write-ups describe work that genuinely can't be rebuilt due to confidentiality?" Response: rebuild a smaller, analogous version on public or synthetic data instead of skipping the build entirely; the point is proof, not exact replication.
"Isn't cutting four write-ups just as risky as having zero builds?" Response: no, since the two that remain actually get read and remembered, while all six competed for attention nobody had past the third one anyway.
If pressed
Callum's rebuilt tool used a synthetic dataset of 500 labeled claims, correctly flagging 71 of 80 fraud cases in a held-out test, a number that exists only because something was actually built and tested, not just described.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.