InterviewAdvancedShipping & Model Lifecycle / Incident management for AI products / #22

Run me through your response to an incident I describe.

The direct answer
Run the same five step method on whatever incident you hand me, not a memorized story. Name the one person whose morning breaks, the habit that had quietly stopped needing attention, and the exact behavior that snaps with no middle setting. Then name the one past decision that only made sense before this happened, and replay the same day with it undone, ending on a real number.
Do this, in order
  1. Run the five step method live, not a memorized story.Why: whatever incident you get, find the person, the habit, the flip, the old decision, the replay gives you a shape to fill in on the spot.
  2. Name one real person and the habit that was quietly working, before you say a word about the model.Why: an incident with nobody in it is a systems diagram, not an answer an interviewer can picture.
  3. Find the exact behavior that snaps, with no middle setting, not just a number that moved.Why: this is the one hard step. Everything else falls out once you have it right.
  4. Name the one past decision you would take back, told as a memory of a real meeting.Why: "add more review" or "retrain the model" is a new dial, not a decision, and interviewers can tell the difference.
  5. Replay the same day with that decision undone, and end on a count, a clock, or a cost.Why: "it would work much better" convinces nobody. A number someone could go check does.
  6. Prove it is a method, not a one time story, by running the same five letters on a second incident.Why: this is the actual thing being tested, that you can do this again on a scenario you have never rehearsed.

How to answer this, stage by stage

This question has no fixed content to prepare, because the interviewer supplies the incident on the spot. What they are grading is whether you have a repeatable shape to pour it into. Seven moves get you there, using one real incident as the worked example.

1
Repeat the incident back in one sentence before you say anything else
Say it like this
"So the picture is, a phone caller and a text booking end up holding the same table at the same time, and neither channel knew about the other. That the right incident?"
Why this works
Buys two seconds of real thinking time and proves you were listening, not just waiting for your turn to talk.
2
Say your structure out loud, before any detail
Say it like this
"I run any incident the same way. Find the person, find the habit, find what snaps, find the old decision, then replay it with the fix in place."
Why this works
Tells the interviewer you have a method, not a story you happen to remember from somewhere.
3
Find the person whose morning breaks
Say it like this
"Rosalie runs reservations at Marrow. Four years in, she can seat a walk in party by eye before the host stand even checks the board."
Why this works
A specific person with a real, stated skill is what makes the rest of the answer land as a real incident instead of a hypothetical.
4
Locate the habit, then name the exact flip
Say it like this
"She used to trust the shared seating board completely and glance at it once an hour. The flip is this: she trusts the board fully and logs nothing herself, or she re-logs every single phone hold the second she takes it. There's no version where she just checks a few more."
Why this works
This is the line an interviewer is actually listening for, two settings and no dial in between.
5
Pinpoint the old decision, and say why it made sense at the time
Say it like this
"When Hearth first launched, phone and text each got a fast, independent check against the board, no shared lock between them, because that lock would've added a beat of delay to every booking, for a race nobody expected to matter yet. That was the right call, back when it was true."
Why this works
Pointing at a decision instead of a person is what a genuinely AI shaped answer requires, since the failure came from how two channels checked each other, not from anyone's judgment.
6
Show the replay, and stop at a number
Say it like this
"Same Friday, but now both channels write to one shared hold check before they confirm. The second request just gets offered the next table automatically, mid call, no delay. Six clashes in six weeks becomes zero."
Why this works
An interviewer remembers a countable ending. "Much smoother" is forgettable. "Six becomes zero" is not.
7
Prove it is a method by offering to run it again, live
Say it like this
"If you want to hand me a second one right now, a dispatcher, a claims bot, whatever it is, I'll run the same five steps on it cold."
Why this works
This is the actual thing being graded, not whether you memorized one good story, but whether you can do this on command.

Let's learn

Say Marrow is a restaurant that takes bookings by phone and by text, both handled by Hearth, the reservation and waitlist assistant a hospitality tech company called Skiffmoor built for restaurants like it. A caller says what they want out loud. A texter types it. Hearth turns either one into a held table.

Before Hearth, one person answered the phone and kept a paper book, and a text booking simply did not exist as an option. About one table a night got double booked anyway, a phone call and a walk in colliding, and the fix was always a comped round of drinks.

With Hearth running both channels, Marrow takes about 340 bookings a week, and for a long stretch, a table clash between the phone channel and the text channel happened about once every six weeks. Rare enough that nobody built anything special around it.

Then a local food writer's piece sent Friday call volume up sharply. People were hanging up during the pause Hearth makes a caller sit through while it checks for a clash, and Skiffmoor's team had a real number showing it. So they shortened that pause, from about four seconds down to under one.

Here is the turn. The shorter pause is not the problem by itself. The problem is what happens the moment a phone hold and a text hold reach for the same table within the same few seconds. Confirming a hold isn't a simple database check. It's Hearth's own model deciding, in the moment, how sure it is the table is free. Each channel runs that check on its own local view, sees a table open, and confirms it confidently, because neither one waited long enough to hear from the other. Table clashes went from about one every six weeks to six in six weeks, twice in a single Friday service.

Table clashes per six week window, before and after the pause got cut
6 3 0 Before the cut 1 After the cut 6, in six weeks
Same six week window, same restaurant, same two booking channels. The only thing that changed was how long each channel waited before trusting its own read of the board.

At its worst, this costs more than a comped table. A guest who gets told their confirmed table is gone stops trusting the confirmation at all. They start calling back to double check every booking by hand, and Marrow's phone volume lands right back where the shorter pause was trying to fix it in the first place.

Two channels that can both say yes at the same time are not two channels. They are a coin flip wearing a nice interface.

The choice I would take back is not the shorter pause itself. It is that the two channels each checked the board against their own read of it, never against each other. Shortening the pause just shortened how long each channel had to notice a race it was never built to see.

What I would leave alone: a text booking placed two days out never touched this problem, because nobody is mid call waiting on it. The real time check only ever needed to matter for holds happening within the same few seconds of each other.

The lesson: a pause is not a safety check. It just delays the moment two systems each confidently do the wrong thing at once.

Hand sketched two panel diagram titled Switch not dial. Left panel, what we assumed, a dial we could trust a little less as the caller pause got shorter. Right panel, what was true, a switch, and it had already flipped the moment two channels could both confirm the same table.
The team treated the caller pause like a dial they could trim a little at a time. It was never a dial. Two channels racing for one table is a switch, and it flips the first time they both reach for it together.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how ordinary the Friday that broke it actually looked.

Rosalie Oyelaran runs the reservation desk at Marrow. Four years in, she can look at a walk in party of five and know, before she even checks the board, which table they're getting and how long it'll hold. She built that eye long before Hearth ever existed, back when the book was paper and the only channel was a landline.

For most of two years, Hearth was the best part of her shift. Calls came in, texts came in, and both landed on one shared board she could read at a glance. She used to keep her own paper tally her first year on the desk, cross checking every hold by hand. Once Hearth proved itself, she let that habit go. She'd glance at the board once an hour, see it matched what she remembered, and trust it.

Then the local food column ran. Friday calls nearly doubled overnight, and a few weeks later, Skiffmoor shortened the pause Hearth makes a caller sit through while it checks for a clash, from about four seconds to under one. Nobody at Marrow was told. It wasn't the kind of change anyone thought to announce.

The first clash happened on a quiet Tuesday, and it barely registered. A text booking for a four top and a phone hold for the same table, both confirmed within a few seconds of each other. Rosalie moved the walk in party who'd have taken it anyway, apologized, comped a round, and didn't think much more about it.

Nobody decided to stop trusting the board that Tuesday. It just cost nothing yet.

Three weeks later, on a Friday, it happened twice in one service, forty minutes apart. Two different tables, two different pairs of guests standing at the host stand with a phone confirmation in hand and no table to seat them at. That was the night Rosalie stopped trusting the board the way she had for two years.

She didn't ask Skiffmoor to fix anything, not yet. She built her own system instead. Every phone hold she took, she typed straight into a note on her own phone, texted to herself, the second she took it, before Hearth's board even updated. Then she'd glance at that note against the board every fifteen minutes instead of once an hour, cross checking her own memory against the system that used to be trustworthy on its own.

Hand sketched two panel comparison titled Small move, big snap. Left panel, the number, a caller wait time that quietly dropped from four seconds to under one second. Right panel, the behavior, flat for two years then jumping with no warning, a reservations manager going from trusting the shared board completely to re logging every phone hold herself.
The pause dropping from four seconds to under one looked like a small tuning change. What it did to Rosalie's Friday nights was not small at all.

Here is what that cost, and it was never really the twenty minutes she spent typing notes to herself. It was the twenty minutes at the top of every Friday shift she now spent cross checking her own memory against a system that was supposed to be the one keeping track, time she used to spend walking the floor before doors opened, checking that the corner booth's wobbly chair had finally been fixed.

The old decision, told as a memory of a meeting: back when Hearth first launched both channels, the build meeting for it was simple. Give phone and text each a fast, independent check against the board, because a shared lock between them would add a beat of delay to every booking, on both channels, for a race that essentially never happened. Nobody in that room was wrong about what the data showed them then. Months later, when Friday call volume nearly doubled, a different, smaller meeting shortened the caller pause to stop real hangups, and nobody in that second room had a reason to revisit the first decision at all.

Knowledge spark: why does a shorter pause cause this at all? Hearth checks a hold against the board the moment a caller or a texter confirms. That check is a quick guess about whether the table is really free, and each channel makes its own guess without asking the other channel first. A longer pause used to leave enough time, almost by accident, for one channel's guess to catch up with the other's. A shorter pause removed that accident.
The decision that mattered Each booking channel checked the shared board against its own read of it, and never against what the other channel was doing in the same instant. That was invisible when calls were slow enough that races almost never happened. Shortening the caller pause did not create the gap. It just removed the last bit of accidental cover it had.

The replay, run the same way, with the fix in place: instead of a longer pause, both channels write to one shared hold ledger the instant a caller or texter reaches for a table. Each one checks that ledger, not just its own guess, before confirming. Whichever request lands first gets the table. The second one gets offered the next best table automatically, mid call, with no added wait for the caller. Same Friday, same near doubled call volume, same two guests reaching for table 12 within eight seconds of each other. This time, one gets table 12, the other gets table 9, and neither of them ever knows a race happened. Rosalie doesn't type a single note to herself. She's done confirming the night's floor plan by 6:10, same as any other night, instead of standing at the host stand until doors open, checking her phone against the board.

What I would tell my past self, in that meeting about the caller pause: cutting a wait time and closing a gap are not the same fix, and we only ever measured the one we could see on a dashboard.

The five questions, run on Marrow, letter by letter

This is a live incident question dressed as a behavioral one. Something changed, a person's routine flipped in response, and the fix is a specific decision taken back. FLIPS fits, run forward from the incident itself back to the meeting that caused it.

Hand sketched vertical list titled The five letters, run live on any incident. F find the person, whose morning breaks first. L locate the habit, what they stopped doing because it worked. I identify the flip, shown in a different color, which verb snaps with no middle. P pinpoint the old decision, what made sense before. S show the replay, same day, better ending.
This is the shape I hold in my head walking into the question, before I know a single detail of the incident I'm about to get.
F
Find the person. Whose morning breaks first when this incident lands?
Not "the front desk team." A name, a desk, a real, stated skill.
Rosalie Oyelaran, reservations manager at Marrow, four years on the desk, able to seat a walk in party by eye.
L
Locate the habit. What did they stop doing because it kept working?
The habit is what the tool actually shipped, not a footnote to the story.
Trusting the shared board completely, and checking it once an hour instead of keeping her own tally the way she did in year one.
I
Identify the flip. What verb snaps, with exactly two settings and no middle?
The hardest step, and the one most answers skip past on the way to a fix.
Trusts the board fully and logs nothing herself, or re-logs every single phone hold the instant she takes it. And this is specifically AI shaped: each channel's own confidence in its read of the board was never checked against the other channel, so a race was invisible to both of them at once.
P
Pinpoint the old decision. Which choice only made sense before the incident?
A specific, reversible, once reasonable call, not a character flaw.
Building phone and text as two independent channels, each checking only its own read of the board, made when Hearth first launched, because a shared cross channel lock would have added a beat of delay to every booking, for a race that essentially never happened yet.
S
Show the replay. Same bad Friday, fixed design, better ending?
Ends in a number someone could go check, not a promise that it feels smoother.
A shared hold ledger both channels check before confirming turns six clashes in six weeks into zero, with no added wait for the caller.

Two things worth saying out loud, since this is exactly where an AI PM interview earns its name. First, the fix on the table is a rejected alternative most candidates reach for by reflex: put the pause back to four seconds. That was considered, and ruled out, because it would undo a real, measured drop in callers hanging up, trading one measured cost for another instead of removing the actual race. Second, the real bar for the fix isn't "never let two channels see the same table." Hearth's table match score, party size against table type, already only auto assigns above a calibrated bar tuned against a same day eval set, and anything below that bar still routes to Rosalie's own queue instead of guessing. The shared ledger is the same idea applied to timing instead of matching, a deterministic check the model's own confidence has to clear before it acts alone.

And if you want to be sure it really works, try it somewhere else

Same five questions, a completely different incident, a different flip family, and a different industry, so the method proves itself instead of just repeating a story you happened to prepare.

Coalter Mechanical dispatches HVAC technicians across a metro area, using Dovecote, a scheduling assistant that scores which tech fits which job and routes the easy calls without a human touching them.

F. Etta Verlaine, senior dispatcher, sixteen years routing technicians, able to tell a five minute filter fix from a four hour compressor job just from how a customer describes the noise on the phone.
L. Once Dovecote's match scores proved reliable, Etta let a newer scheduler run the whole emergency board alone during her lunch hour, three months running, because Dovecote's picks kept being good ones.
I. A different flip than Marrow's. This is a delegation flip, not a workaround. Dovecote started misrouting big commercial refrigerant jobs, which need a licensed tech, to techs who were great at routine residential calls but not licensed for that work. Three of those went out wrong in two weeks. Etta didn't spot check the commercial jobs after that. She took the entire emergency board back, residential included, even though residential was never the problem.
P. Dovecote launched with one blended confidence bar covering every job type, because commercial jobs were a small enough slice of daily volume that a second, separate bar felt like extra engineering for a handful of calls a week. Reasonable, at the time, given how thin that slice of the data actually was.
S. Give commercial jobs their own, stricter auto dispatch bar, calibrated against a commercial only eval set, so anything below it routes straight to Etta instead of auto assigning. Replay the same two weeks: the junior scheduler still clears about ninety five percent of the board alone, and Etta reviews only the three or so commercial calls a day that actually need her, not all forty.

Same shape, different cost At Marrow, the missing check was one channel talking to the other before confirming. At Coalter, it was one confidence bar covering two job types that needed different bars. Different flip family, different fix, same real finding: the decision worth taking back is almost always the one that quietly let a system trust itself without checking something outside itself.

Swap the trigger and it still runs.
Speed: an interviewer cuts you off after ninety seconds. Skip straight to naming the decision that let two channels both say yes, and what replaces it.
Cost: engineering says the shared ledger can't ship for six weeks. Don't leave both channels racing in the meantime. Ship the cheap partial version first, the busier channel checks the last few seconds of the other channel's holds before confirming, even before the fix is wired both ways.
The model got better, for real: suppose Hearth's parsing of what a caller actually wants got measurably more accurate. That still isn't the same claim as two channels never racing each other. A better read of "the corner booth for four" doesn't stop it from being confirmed twice within the same eight seconds.

Where people run it wrong.
They point the answer at the person, "she should have double checked," instead of at the decision that let two channels each trust their own read of the board.
They propose "add a review step" or "tighten the confidence bar," which is a new dial turned up, not a decision taken back.
They stop at "make the model more accurate," which does nothing for a race between two systems that both already thought they were right.

How to use it live. Repeat the incident back in one plain sentence before you answer anything. It buys you real thinking time without looking like you're stalling, and it's the same first move every single time, no matter what incident lands on the table.

Flashcards (click a card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip. Rosalie doesn't check more or check less. She builds a private process, re-logging every phone hold herself, around a tool that kept no shared memory of what was already in flight.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Rosalie Oyelaran, reservations manager at Marrow, four years on the desk, able to seat a walk in party by eye before the board even updates.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Keeping her own tally of holds by hand. Once the shared board proved reliable, she dropped it and started trusting the board completely, checking it once an hour.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Trusts the shared board fully and logs nothing herself, or re-logs every single phone hold the second she takes it. No version where she just checks a few more of them.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Building phone and text as two independent channels that each check only their own read of the board, made when Hearth first launched. It made sense because a shared lock would have slowed every booking, for a race that essentially never happened yet.
6 · THE NUMBER
Fill in the blank: table clashes went from about one every ___ weeks to ___ clashes in six weeks after the pause got cut.
Tap to flip
ANSWER
Six weeks; six clashes. The rate went from rare enough to ignore to almost weekly.
7 · THE REPLAY
Same bad Friday, new design, what changes?
Tap to flip
ANSWER
A shared hold ledger both channels check before confirming turns six clashes in six weeks into zero, with no added wait for the caller, and Rosalie stops re-logging holds entirely.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Dovecote, the HVAC dispatch assistant at Coalter Mechanical. Delegation flip: a senior dispatcher takes the emergency board back from a junior scheduler once the model starts misrouting licensed commercial jobs.

Check yourself Score: 0 / 0

True or false
1. True or false: the right fix here is to make Hearth's caller pause longer again, back to about four seconds.
  • True
  • False
Show hint
Check what the shorter pause was already fixing before it caused this.
Show answer
False. Lengthening the pause would undo a real, measured drop in callers hanging up. The fix has to remove the race between the two channels, not reintroduce the wait.
Multiple choice
2. Why is this a workaround flip and not a verification flip?
  • A. Because Rosalie checks a sample of bookings instead of checking every single one.
  • B. Because she builds her own private log around the tool, instead of changing how much of the tool's own output she checks.
  • C. Because the model got better, so she stopped checking anything at all.
  • D. Because she stopped opening the shared board altogether.
Show hint
Look at what she actually does with her hands after the Friday it happens twice.
Show answer
B. A verification flip changes how much of the tool's own output gets checked. Here, she invents a whole second system, her own private log, because the tool itself has no shared memory of what's already in flight.
Fill in the blank
3. The caller hold pause dropped from about ___ seconds to under ___ second, and over the next six weeks, table clashes went from about one every six weeks to ___ in six weeks.
Show hint
Check the chart in "Let's learn" and stage 5 of the walkthrough.
Show answer
four; one; six. A pause cut for a real reason turned a once every six weeks problem into an almost weekly one, without anyone changing the model itself.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory, not a dial anyone could just turn back up.
Show answer
Model answer: Building phone and text as two independent channels that each check only their own read of the board, back when Hearth first launched. It made sense because a shared cross channel check would have slowed every booking, for a race that essentially never happened at that volume.
Short answer, apply it yourself
5. Think of a tool you use with more than one way to feed it input, a form and a chat box, say. What's a race between those two paths that could quietly produce two conflicting results?
Show hint
Look for two paths that can both write to the same record within the same few seconds.
Show answer
Model answer: A shared calendar tool with a web form and an email-to-book address. If both create a hold on the same slot within a few seconds of each other, whichever one syncs last could silently overwrite the other's booking instead of flagging the clash to anyone.
Short answer, the number question
6. If Marrow's Friday call volume had only gone up by ten percent instead of nearly doubling, would the old four second pause likely have stayed long enough to keep catching these clashes? Show the reasoning.
Show hint
Think about why the team cut the pause in the first place, and how big a jump it took to force that call.
Show answer
Probably yes, for a while. The whole reason the pause got cut was to survive a near doubling of calls without losing callers to hangups. A milder ten percent bump likely never forces that same tradeoff, so the pause might have stayed at four seconds and the race might never have surfaced at all. The size of the jump is what forced the decision, not volume by itself.
Before you close the answer
Why this works
Tests whether you can run a real method on an incident you have never heard before, live, instead of retelling one you rehearsed. The second, unrelated incident in Section 4 is the actual proof.
Follow-up traps
"What if you can't think of a real, named person on the spot?" Response: invent one plausible role with one stated line of competence. The interviewer is grading the structure of the answer, not whether the name is a real employee.

"Isn't putting the pause back to four seconds the obvious, cheap fix?" Response: no, because that undoes a real, measured drop in call abandonment. The fix has to remove the race between the two channels, not reintroduce the wait that caused a different, already solved problem.
If pressed
The shared hold ledger uses a short lived key with a compare and swap write and a ninety second expiry, so whichever channel's confirm reaches it first wins the table, the loser is told to retry against the next candidate table with no added latency to the caller, and an abandoned hold never locks a table forever.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more