Walk me through your strongest portfolio piece as if I were hiring you.
Norwood Telecom Group is a fictional regional carrier. Idris Falana built Capacity Draft there: it drafts a recommendation for which cell towers need a capacity upgrade before the next traffic season. Marguerite Okonjo is interviewing him, and asked him to walk through his strongest piece as if she were deciding whether to hire him right now.
- Show the reasoning behind each recommendation, not just the final number.Why: a flat number gives nobody a way to spot-check where it might be wrong.
- Flag any region where recent traffic diverges from the model's training data.Why: that's exactly where a confident-looking number is most likely to be quietly wrong.
- Let the senior spot-check only the flagged cases, not every recommendation.Why: that's what actually restores real delegation instead of a review in disguise.
- Name the alternative I considered and rejected: adding a second human reviewer layer.Why: it would have fixed the symptom, doubled the review cost, and never fixed the underlying trust problem.
- Keep the junior planner's authority over the unflagged, routine recommendations.Why: the goal was restoring delegation, not quietly replacing it with disguised oversight.
- Measure success by how much gets spot-checked, not just by accuracy.Why: accuracy alone hid the exact problem that broke trust the first time.
How to answer this, stage by stage
Nobody's grading whether the project sounds impressive. They're grading whether you can walk them through it like you actually lived it.
Let's learn
Capacity Draft is a tool that looks at a cell tower's traffic history and recommends whether it needs a capacity upgrade before the next busy season hits.
For six years, Marcus Webb reviewed every single upgrade recommendation at Norwood personally. Once Capacity Draft launched, his junior teammate Elena Cho started approving the routine recommendations herself, using the tool's draft. For four months, that worked well.
Then a fast-growing suburban region got a recommendation that looked just as polished as every other one Elena had approved that month. It was wrong. The model had barely seen a region grow that quickly before, and nothing in its output said so.
Here's the turn: the missed forecast itself cost Norwood one delayed upgrade. What it actually cost was Marcus's trust in delegating at all, and within a month, two people were doing the review work that used to take one.
At its worst: Elena, once confident and trusted with routine calls, starts feeling second-guessed on everything, even the recommendations she's clearly getting right, and the two of them end up in a tense conversation about whether she can be trusted at all.
What I would leave alone: the core forecasting model itself didn't need retraining. It was accurate on the vast majority of steady, well-represented regions. The fix was entirely about surfacing which regions it was actually uncertain about.
The lesson: a delegation only survives the first real mistake if the person delegating can see exactly where to look, instead of having to re-check everything to find out.
Now here is the same thing as a story
The short version above is what I'd say defending this project cold, in an interview. Read this one for how it actually played out.
Marcus Webb spent six years at Norwood before Capacity Draft existed, and he could look at a tower's traffic graph and know, almost by feel, whether it needed help before the numbers technically said so.
For the first four months after launch, the good months, Elena approved dozens of routine upgrade recommendations a week using Capacity Draft's drafts, and Marcus spot-checked fewer and fewer of them, since every one he checked looked exactly like his own work used to.
The trigger was one region, a fast-growing suburb where traffic had climbed faster than almost anywhere else in Norwood's footprint. The model, trained mostly on slower-growing areas, gave a confident-looking recommendation that turned out to be badly wrong, and nothing about the output had ever said "this one's different."
Within a week, Marcus was re-reviewing every single recommendation Elena produced, not just the new ones, going back over weeks of already-approved drafts he'd previously trusted without a second look.
The real cost was never the one missed forecast. It was that Elena, who'd been doing the job well for four months, suddenly had every decision she'd ever made quietly re-litigated, and neither of them had a fast way to tell which of her calls actually deserved a second look.
Shipping one clean number back at launch had made complete sense: it was simple, it matched exactly what Marcus used to write down himself, and nobody had asked for anything more. It stopped making sense the moment that one number needed to earn back trust it had just lost, and gave nobody anywhere to look.
With the redesign, every recommendation now shows the regional drivers behind it and a flag whenever recent traffic breaks from the pattern the model trained on. Marcus spot-checks the flagged one in ten. Elena keeps full authority on the other nine, the same authority she'd earned back in month one.
The old design asked Marcus to trust a number with nothing behind it. The new one shows him exactly which numbers are worth a second look, and lets the rest of the delegation actually hold.
I shipped one flat number because it was simple and it matched what I used to write down myself. It took watching Marcus quietly take back four months of delegated trust, over one region, to see that simple and trustworthy were never the same thing.
The five steps, walked through liveNot a slide. FLIPS is the one thing you can actually run in your head mid-interview, under real pressure.
The recap, one line per letter: find the person is Marcus's six years of real competence, locate the habit is trusting Elena's tool-assisted drafts without a second look, identify the flip is delegation reversing all at once, pinpoint the old decision is the flat number with no reasoning, and show the replay is spot-checking one in ten instead of all ten.
And if you want to be sure it really works, try it somewhere elseSame five steps, an insurance claims desk instead of a telecom network. A different flip family entirely: this time, nobody takes anything back. They just stop telling anyone.
Coppergate Insurance Group is a fictional insurer. Wilhelmina Brandt built a tool there that drafts claim-denial justification letters for adjusters, checked off as "AI-assisted" in a shared team log. Salim Otieno reviews her case study.
Mapped onto FLIPS with a different family, the concealment flip: the person is a claims adjuster, competent for years at writing denial letters that held up under appeal. The habit was checking the "AI-assisted" box openly in team reviews, since quality was good and there was no cost to saying so. The flip, concealment: once a few AI-drafted letters came back legally shaky and got quietly corrected without escalation, adjusters stopped checking the box at all, not because they stopped using the tool, but because admitting it had become a liability. The old decision: making disclosure a per-person, visible checkbox, fine while quality was good, costly the moment it wasn't. The replay: switching to an anonymous, team-wide dashboard showing the overall AI-assist rate, with no individual attribution, restored honest reporting without any one adjuster bearing the social cost alone.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "one flat number broke a delegation. Showing the reasoning behind it fixed it," and stop.
Cost: there's no budget to build a full reasoning trace for every recommendation. Start with the divergence flag alone; it catches the exact case that broke trust here, for a fraction of the build.
The model gets better, for real: if Capacity Draft's accuracy improves next quarter, that's still not a reason to remove the divergence flag. A better average model can still be wrong in exactly the one region it's never seen before.
Where people run it wrong.
They add a second human reviewer layer to fix a trust problem, which doubles the review cost without fixing why trust broke in the first place.
They treat the missed forecast as the whole story, instead of the delegation collapse it actually caused.
They measure success by overall accuracy alone, the exact number that hid the regional gap in the first place.
How to use it live. When someone asks you to walk through your strongest piece, ask yourself one question first: what's the moment someone's behavior changed, not just a number. Start there, not with the architecture.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't Marcus have just trusted the tool a little less across the board?" Response: no, that's exactly the dial he didn't have. With no per-recommendation signal, "trust it a little less" meant re-checking everything, which is what actually happened.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Building an AI PM portfolio
- #1 What does a hiring manager actually open first in an AI PM portfolio?
- #2 Describe the three artifacts that make the strongest AI PM portfolio.
- #3 How do you present a shipped AI project when you cannot share the internal metrics?
- #4 What does a portfolio project need to prove that a resume line cannot?
- #5 Critique a portfolio built entirely from case study write-ups with no build.
- #6 How do you build a credible AI PM portfolio with no AI job experience?