InterviewAdvancedResponsible AI & Advanced Practice / Building an AI PM portfolio / #20

Walk me through your strongest portfolio piece as if I were hiring you.

FLIPS the portfolio piece is Capacity Draft, an AI copilot that drafts network capacity upgrade recommendations for Norwood Telecom Group

Norwood Telecom Group is a fictional regional carrier. Idris Falana built Capacity Draft there: it drafts a recommendation for which cell towers need a capacity upgrade before the next traffic season. Marguerite Okonjo is interviewing him, and asked him to walk through his strongest piece as if she were deciding whether to hire him right now.

The direct answer
My strongest piece is Capacity Draft, a tool that let a senior planner delegate first-pass upgrade recommendations to a junior teammate. It broke when the tool's flat, one-number output gave the senior no way to spot-check where the junior's judgment might be shaky, so he quietly took all the review work back himself. The fix: show the regional reasoning behind each number, and flag any region where recent traffic breaks from what the model has actually seen before.
Do this, in order
  1. Show the reasoning behind each recommendation, not just the final number.Why: a flat number gives nobody a way to spot-check where it might be wrong.
  2. Flag any region where recent traffic diverges from the model's training data.Why: that's exactly where a confident-looking number is most likely to be quietly wrong.
  3. Let the senior spot-check only the flagged cases, not every recommendation.Why: that's what actually restores real delegation instead of a review in disguise.
  4. Name the alternative I considered and rejected: adding a second human reviewer layer.Why: it would have fixed the symptom, doubled the review cost, and never fixed the underlying trust problem.
  5. Keep the junior planner's authority over the unflagged, routine recommendations.Why: the goal was restoring delegation, not quietly replacing it with disguised oversight.
  6. Measure success by how much gets spot-checked, not just by accuracy.Why: accuracy alone hid the exact problem that broke trust the first time.

How to answer this, stage by stage

Nobody's grading whether the project sounds impressive. They're grading whether you can walk them through it like you actually lived it.

Stage 1
Scope it to one real piece
Say it like this
"My strongest piece is Capacity Draft, a copilot that drafts network capacity upgrade recommendations at a telecom carrier."
Why this works
Names the project by name in the first breath, no warm-up needed.
Stage 2
Say your structure out loud
Say it like this
"I'll walk you through it with FLIPS: find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals you have a method for telling this story, not just a memorized pitch.
Stage 3
Find the person
Say it like this
"Marcus Webb, a senior capacity planner, reviewed every single upgrade recommendation for six years before the tool existed."
Why this works
Names a real competence before anything goes wrong, so the loss later actually means something.
Stage 4
Locate the habit
Say it like this
"Once Elena, his junior teammate, started using the tool, Marcus stopped reviewing every draft line by line. He trusted her, and the output looked as solid as his own used to."
Why this works
Shows the habit forming because the tool was genuinely working, not out of laziness.
Stage 5
Identify the flip
Say it like this
"Delegation flip: he'd handed the review down. Once one bad forecast slipped through, he took every single review back, all at once, no middle ground."
Why this works
Names a real behavior with two settings, not a vague loss of confidence.
Stage 6
Pinpoint the old decision
Say it like this
"We'd shipped one flat capacity number with no reasoning behind it. That was fine when it was right. It gave Marcus nothing to spot-check once he needed to trust it less."
Why this works
A specific, reversible decision, not a vague "the model wasn't good enough."
Stage 7
Show the replay
Say it like this
"With the reasoning and the divergence flag, Marcus now spot-checks about one recommendation in ten instead of all of them, and Elena keeps her authority on the other nine."
Why this works
Ends on something countable: one in ten, not "much better now."
Stage 8
Close on the one line
Say it like this
"So that's the piece: a delegation that broke because we shipped one number instead of the reasoning behind it, and the fix that let the delegation actually hold."
Why this works
Restates the whole decision in one breath, ready for whatever question comes next.

Let's learn

Capacity Draft is a tool that looks at a cell tower's traffic history and recommends whether it needs a capacity upgrade before the next busy season hits.

For six years, Marcus Webb reviewed every single upgrade recommendation at Norwood personally. Once Capacity Draft launched, his junior teammate Elena Cho started approving the routine recommendations herself, using the tool's draft. For four months, that worked well.

Knowledge spark: what's a distribution shift? When the real-world data a model sees starts looking different from the data it was trained on. A model trained mostly on steady, predictable regions can quietly struggle in one that's growing fast or shifting seasonally, even while it still sounds confident.

Then a fast-growing suburban region got a recommendation that looked just as polished as every other one Elena had approved that month. It was wrong. The model had barely seen a region grow that quickly before, and nothing in its output said so.

Recommendations Marcus personally re-reviewed each week
100% 50% 0 Week 1 Week 2 Week 3 Week 4 5% 15% 60% 98%
Marcus didn't ease back into reviewing more. He went from spot-checking almost nothing to re-checking almost everything, inside a single month.

Here's the turn: the missed forecast itself cost Norwood one delayed upgrade. What it actually cost was Marcus's trust in delegating at all, and within a month, two people were doing the review work that used to take one.

We didn't lose one bad forecast. We lost a working delegation, all at once, over a single region nobody had flagged as different.

At its worst: Elena, once confident and trusted with routine calls, starts feeling second-guessed on everything, even the recommendations she's clearly getting right, and the two of them end up in a tense conversation about whether she can be trusted at all.

The decision I would take back We shipped Capacity Draft with one flat recommendation number, no regional reasoning, no flag for when a region's recent traffic pattern diverged from what the model had trained on. That was fine while the number was reliably right. It stopped being fine the moment it needed to earn trust back.

What I would leave alone: the core forecasting model itself didn't need retraining. It was accurate on the vast majority of steady, well-represented regions. The fix was entirely about surfacing which regions it was actually uncertain about.

The lesson: a delegation only survives the first real mistake if the person delegating can see exactly where to look, instead of having to re-check everything to find out.

Now here is the same thing as a story

The short version above is what I'd say defending this project cold, in an interview. Read this one for how it actually played out.

Marcus Webb spent six years at Norwood before Capacity Draft existed, and he could look at a tower's traffic graph and know, almost by feel, whether it needed help before the numbers technically said so.

Hand sketched timeline titled Before the trigger. Four milestones: Marcus alone reviews everything, good months Elena drafts approved, habit thins spot checks only, the trigger a missed forecast highlighted.
The habit didn't break in one step. It thinned across months, and the trigger itself was small: one region, one missed call.

For the first four months after launch, the good months, Elena approved dozens of routine upgrade recommendations a week using Capacity Draft's drafts, and Marcus spot-checked fewer and fewer of them, since every one he checked looked exactly like his own work used to.

Hand sketched comparison diagram titled Small move, big snap. Left panel, a gauge icon labeled Trust fades, caption slow over weeks. Right panel, a red box icon labeled Review snaps, caption to all at once.
Trust thinned slowly, over weeks. The review habit didn't thin the same way. It snapped, all at once, the week after the missed forecast.

The trigger was one region, a fast-growing suburb where traffic had climbed faster than almost anywhere else in Norwood's footprint. The model, trained mostly on slower-growing areas, gave a confident-looking recommendation that turned out to be badly wrong, and nothing about the output had ever said "this one's different."

Hand sketched labeled parts diagram titled What the flat number hid. Center gauge icon labeled One capacity number, with four callouts: regional drivers, confidence flag, training distribution, seasonal pattern.
All four of these existed inside the model already. None of them ever reached the screen Marcus and Elena actually looked at.

Within a week, Marcus was re-reviewing every single recommendation Elena produced, not just the new ones, going back over weeks of already-approved drafts he'd previously trusted without a second look.

Hand sketched icon list titled What Marcus's day looks like now. Items: re-reviews every draft, a hard talk with Elena, three extra hours weekly, delegation effectively paused.
Four real costs, and not one of them shows up on a dashboard tracking model accuracy.

The real cost was never the one missed forecast. It was that Elena, who'd been doing the job well for four months, suddenly had every decision she'd ever made quietly re-litigated, and neither of them had a fast way to tell which of her calls actually deserved a second look.

Shipping one clean number back at launch had made complete sense: it was simple, it matched exactly what Marcus used to write down himself, and nobody had asked for anything more. It stopped making sense the moment that one number needed to earn back trust it had just lost, and gave nobody anywhere to look.

Hand sketched metaphor scene titled Switch not dial. Left, a gauge icon labeled Assumed, caption adjustable trust. Right, a red box icon labeled Actual, caption on or off.
We built the whole delegation assuming trust could dial down gradually if something went wrong. It didn't. It switched off entirely.

With the redesign, every recommendation now shows the regional drivers behind it and a flag whenever recent traffic breaks from the pattern the model trained on. Marcus spot-checks the flagged one in ten. Elena keeps full authority on the other nine, the same authority she'd earned back in month one.

The old design asked Marcus to trust a number with nothing behind it. The new one shows him exactly which numbers are worth a second look, and lets the rest of the delegation actually hold.

I shipped one flat number because it was simple and it matched what I used to write down myself. It took watching Marcus quietly take back four months of delegated trust, over one region, to see that simple and trustworthy were never the same thing.

The five steps, walked through liveNot a slide. FLIPS is the one thing you can actually run in your head mid-interview, under real pressure.

F
Find the person.
Marcus Webb, six years reviewing every upgrade recommendation himself, before the tool existed.
A named competence, so the loss later has real weight.
L
Locate the habit.
He stopped reviewing every draft line by line, once Elena's tool-assisted work matched his own for four straight months.
The habit formed because the tool was genuinely working, not carelessness.
I
Identify the flip. Delegation.
Handed the review down to Elena. One missed forecast, and he took every review back, all at once.
The hardest step: a real verb with exactly two settings, not a mood.
P
Pinpoint the old decision.
Shipping one flat number with no regional reasoning or divergence flag behind it.
Reversible, specific, and genuinely sensible when it was made.
S
Show the replay.
Marcus now spot-checks one flagged recommendation in ten. Elena keeps authority on the other nine.
Ends on a countable ratio, not a vague "trust restored."
Hand sketched decision tree titled Reading a confident draft two ways. Root, a confident AI draft, branching to matches real judgment leads to approved fast, quietly diverges leads to nobody notices yet.
This is the tree hiding inside every confident-looking output, and it's exactly what a flag is built to surface.
Recommendation accuracy, by region growth rate
100% 50% 0 95% Steady regions 89% Moderate growth 61% Fast growth
Fast-growth regions sat 34 points below steady ones, and that gap was invisible until someone actually recut the number by region.

The recap, one line per letter: find the person is Marcus's six years of real competence, locate the habit is trusting Elena's tool-assisted drafts without a second look, identify the flip is delegation reversing all at once, pinpoint the old decision is the flat number with no reasoning, and show the replay is spot-checking one in ten instead of all ten.

And if you want to be sure it really works, try it somewhere elseSame five steps, an insurance claims desk instead of a telecom network. A different flip family entirely: this time, nobody takes anything back. They just stop telling anyone.

Coppergate Insurance Group is a fictional insurer. Wilhelmina Brandt built a tool there that drafts claim-denial justification letters for adjusters, checked off as "AI-assisted" in a shared team log. Salim Otieno reviews her case study.

Mapped onto FLIPS with a different family, the concealment flip: the person is a claims adjuster, competent for years at writing denial letters that held up under appeal. The habit was checking the "AI-assisted" box openly in team reviews, since quality was good and there was no cost to saying so. The flip, concealment: once a few AI-drafted letters came back legally shaky and got quietly corrected without escalation, adjusters stopped checking the box at all, not because they stopped using the tool, but because admitting it had become a liability. The old decision: making disclosure a per-person, visible checkbox, fine while quality was good, costly the moment it wasn't. The replay: switching to an anonymous, team-wide dashboard showing the overall AI-assist rate, with no individual attribution, restored honest reporting without any one adjuster bearing the social cost alone.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "one flat number broke a delegation. Showing the reasoning behind it fixed it," and stop.
Cost: there's no budget to build a full reasoning trace for every recommendation. Start with the divergence flag alone; it catches the exact case that broke trust here, for a fraction of the build.
The model gets better, for real: if Capacity Draft's accuracy improves next quarter, that's still not a reason to remove the divergence flag. A better average model can still be wrong in exactly the one region it's never seen before.

Where people run it wrong.
They add a second human reviewer layer to fix a trust problem, which doubles the review cost without fixing why trust broke in the first place.
They treat the missed forecast as the whole story, instead of the delegation collapse it actually caused.
They measure success by overall accuracy alone, the exact number that hid the regional gap in the first place.

How to use it live. When someone asks you to walk through your strongest piece, ask yourself one question first: what's the moment someone's behavior changed, not just a number. Start there, not with the architecture.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Delegation flip: handed the task down, then took it back. A senior person reclaims work once a junior's tool-assisted output loses trust.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Marcus Webb, a senior capacity planner with six years of real experience, and Elena Cho, his junior teammate.
3 · THE HABIT
What did Marcus stop doing because it worked?
Tap to flip
ANSWER
Reviewing every recommendation line by line. Elena's tool-assisted work matched his own for four straight months.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Fully delegated review versus fully reclaimed review. Marcus went from spot-checking almost nothing to re-checking almost everything, with no middle setting.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping one flat capacity number with no regional reasoning or divergence flag, fine while it was reliably right, useless the moment it needed to earn trust back.
6 · THE NUMBER
Fill in the blank: fast-growth regions had ___ percent accuracy, versus 95 percent in steady regions.
Tap to flip
ANSWER
61 percent. A 34-point gap that stayed invisible until someone actually recut accuracy by region.
7 · THE REPLAY
Same bad day, new design. What changes?
Tap to flip
ANSWER
Marcus spot-checks the one flagged recommendation in ten. Elena keeps full authority on the other nine, restoring the delegation instead of replacing it.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Wilhelmina Brandt's claim-denial letter tool at Coppergate Insurance Group, running the concealment flip: adjusters quietly stopped disclosing AI use once quality dipped.

Check yourself Score: 0 / 0

True or false
1. True or false: Marcus gradually increased how much he reviewed, over several months, as trust slowly returned.
  • True
  • False
Show hint
Look at the line chart of weekly re-review rate.
Show answer
False. He went from re-reviewing almost nothing to almost everything within a single month, a snap, not a gradual dial.
Multiple choice
2. Why couldn't Marcus have just "checked a bit more carefully" instead of reclaiming every review?
  • A. Because he didn't trust Elena as a person.
  • B. Because nothing in the tool's flat output told him which recommendations were actually worth a closer look, so "checking more carefully" meant checking everything.
  • C. Because company policy required a full re-audit after any single error.
  • D. Because Elena asked him to review everything herself.
Show hint
Look at "what the flat number hid."
Show answer
B. With no signal for which recommendations were uncertain, a middle ground like "check a bit more" wasn't actually available to him.
Fill in the blank
3. Fill in the blank: with the redesign, Marcus now spot-checks about one recommendation in ___.
Show hint
Look at the replay, stage 7 of the walkthrough.
Show answer
Ten. The flagged one gets a second look. The other nine stay fully in Elena's hands.
Short answer, where it wouldn't matter
4. Name a part of Capacity Draft that this whole redesign left completely untouched.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The core forecasting model itself. It stayed accurate on steady regions the whole time; only the missing reasoning and divergence flag needed to change.
Short answer, apply it yourself
5. Pick a tool you use yourself. What's one habit it built in you that you'd stop doing entirely, not just less of, if it got a little worse?
Show hint
Look for a habit with two settings, not a gradual "I'd be more careful."
Show answer
Model answer: Most people can name a tool where one bad result would make them stop trusting it entirely, rather than trust it "a little less," the same all-or-nothing shape as Marcus's flip.
Short answer, the number question
6. If fast-growth regions had run at 80 percent accuracy instead of 61, would the same divergence-flag fix still be the right call? Why or why not?
Show hint
Think about whether the fix depends on the exact gap, or on the fact that a gap existed and was invisible.
Show answer
Model answer: Likely still yes, since the core problem was never the size of the gap, it was that no one could see which regions had one until they went looking.
Before you close the answer
Why this works
Tests whether you can narrate your own strongest project as a real story with a real behavior change in it, under live pressure, instead of reciting an architecture summary that never actually answers "walk me through it."
Follow-up traps
"Why not just add a second reviewer to catch what the junior misses?" Response: I considered that and rejected it. It doubles the review cost and never actually tells anyone where to look, which is the real problem.

"Couldn't Marcus have just trusted the tool a little less across the board?" Response: no, that's exactly the dial he didn't have. With no per-recommendation signal, "trust it a little less" meant re-checking everything, which is what actually happened.
If pressed
The divergence flag itself works by comparing a region's most recent traffic growth rate against the range of growth rates the model actually saw during training, and flags anything outside the 90th percentile of that training range, a concrete, checkable rule, not a vague confidence score.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more