Run me through your response to an incident I describe.
- Run the five step method live, not a memorized story.Why: whatever incident you get, find the person, the habit, the flip, the old decision, the replay gives you a shape to fill in on the spot.
- Name one real person and the habit that was quietly working, before you say a word about the model.Why: an incident with nobody in it is a systems diagram, not an answer an interviewer can picture.
- Find the exact behavior that snaps, with no middle setting, not just a number that moved.Why: this is the one hard step. Everything else falls out once you have it right.
- Name the one past decision you would take back, told as a memory of a real meeting.Why: "add more review" or "retrain the model" is a new dial, not a decision, and interviewers can tell the difference.
- Replay the same day with that decision undone, and end on a count, a clock, or a cost.Why: "it would work much better" convinces nobody. A number someone could go check does.
- Prove it is a method, not a one time story, by running the same five letters on a second incident.Why: this is the actual thing being tested, that you can do this again on a scenario you have never rehearsed.
How to answer this, stage by stage
This question has no fixed content to prepare, because the interviewer supplies the incident on the spot. What they are grading is whether you have a repeatable shape to pour it into. Seven moves get you there, using one real incident as the worked example.
Let's learn
Say Marrow is a restaurant that takes bookings by phone and by text, both handled by Hearth, the reservation and waitlist assistant a hospitality tech company called Skiffmoor built for restaurants like it. A caller says what they want out loud. A texter types it. Hearth turns either one into a held table.
Before Hearth, one person answered the phone and kept a paper book, and a text booking simply did not exist as an option. About one table a night got double booked anyway, a phone call and a walk in colliding, and the fix was always a comped round of drinks.
With Hearth running both channels, Marrow takes about 340 bookings a week, and for a long stretch, a table clash between the phone channel and the text channel happened about once every six weeks. Rare enough that nobody built anything special around it.
Then a local food writer's piece sent Friday call volume up sharply. People were hanging up during the pause Hearth makes a caller sit through while it checks for a clash, and Skiffmoor's team had a real number showing it. So they shortened that pause, from about four seconds down to under one.
Here is the turn. The shorter pause is not the problem by itself. The problem is what happens the moment a phone hold and a text hold reach for the same table within the same few seconds. Confirming a hold isn't a simple database check. It's Hearth's own model deciding, in the moment, how sure it is the table is free. Each channel runs that check on its own local view, sees a table open, and confirms it confidently, because neither one waited long enough to hear from the other. Table clashes went from about one every six weeks to six in six weeks, twice in a single Friday service.
At its worst, this costs more than a comped table. A guest who gets told their confirmed table is gone stops trusting the confirmation at all. They start calling back to double check every booking by hand, and Marrow's phone volume lands right back where the shorter pause was trying to fix it in the first place.
The choice I would take back is not the shorter pause itself. It is that the two channels each checked the board against their own read of it, never against each other. Shortening the pause just shortened how long each channel had to notice a race it was never built to see.
What I would leave alone: a text booking placed two days out never touched this problem, because nobody is mid call waiting on it. The real time check only ever needed to matter for holds happening within the same few seconds of each other.
The lesson: a pause is not a safety check. It just delays the moment two systems each confidently do the wrong thing at once.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how ordinary the Friday that broke it actually looked.
Rosalie Oyelaran runs the reservation desk at Marrow. Four years in, she can look at a walk in party of five and know, before she even checks the board, which table they're getting and how long it'll hold. She built that eye long before Hearth ever existed, back when the book was paper and the only channel was a landline.
For most of two years, Hearth was the best part of her shift. Calls came in, texts came in, and both landed on one shared board she could read at a glance. She used to keep her own paper tally her first year on the desk, cross checking every hold by hand. Once Hearth proved itself, she let that habit go. She'd glance at the board once an hour, see it matched what she remembered, and trust it.
Then the local food column ran. Friday calls nearly doubled overnight, and a few weeks later, Skiffmoor shortened the pause Hearth makes a caller sit through while it checks for a clash, from about four seconds to under one. Nobody at Marrow was told. It wasn't the kind of change anyone thought to announce.
The first clash happened on a quiet Tuesday, and it barely registered. A text booking for a four top and a phone hold for the same table, both confirmed within a few seconds of each other. Rosalie moved the walk in party who'd have taken it anyway, apologized, comped a round, and didn't think much more about it.
Three weeks later, on a Friday, it happened twice in one service, forty minutes apart. Two different tables, two different pairs of guests standing at the host stand with a phone confirmation in hand and no table to seat them at. That was the night Rosalie stopped trusting the board the way she had for two years.
She didn't ask Skiffmoor to fix anything, not yet. She built her own system instead. Every phone hold she took, she typed straight into a note on her own phone, texted to herself, the second she took it, before Hearth's board even updated. Then she'd glance at that note against the board every fifteen minutes instead of once an hour, cross checking her own memory against the system that used to be trustworthy on its own.
Here is what that cost, and it was never really the twenty minutes she spent typing notes to herself. It was the twenty minutes at the top of every Friday shift she now spent cross checking her own memory against a system that was supposed to be the one keeping track, time she used to spend walking the floor before doors opened, checking that the corner booth's wobbly chair had finally been fixed.
The old decision, told as a memory of a meeting: back when Hearth first launched both channels, the build meeting for it was simple. Give phone and text each a fast, independent check against the board, because a shared lock between them would add a beat of delay to every booking, on both channels, for a race that essentially never happened. Nobody in that room was wrong about what the data showed them then. Months later, when Friday call volume nearly doubled, a different, smaller meeting shortened the caller pause to stop real hangups, and nobody in that second room had a reason to revisit the first decision at all.
The replay, run the same way, with the fix in place: instead of a longer pause, both channels write to one shared hold ledger the instant a caller or texter reaches for a table. Each one checks that ledger, not just its own guess, before confirming. Whichever request lands first gets the table. The second one gets offered the next best table automatically, mid call, with no added wait for the caller. Same Friday, same near doubled call volume, same two guests reaching for table 12 within eight seconds of each other. This time, one gets table 12, the other gets table 9, and neither of them ever knows a race happened. Rosalie doesn't type a single note to herself. She's done confirming the night's floor plan by 6:10, same as any other night, instead of standing at the host stand until doors open, checking her phone against the board.
What I would tell my past self, in that meeting about the caller pause: cutting a wait time and closing a gap are not the same fix, and we only ever measured the one we could see on a dashboard.
The five questions, run on Marrow, letter by letter
This is a live incident question dressed as a behavioral one. Something changed, a person's routine flipped in response, and the fix is a specific decision taken back. FLIPS fits, run forward from the incident itself back to the meeting that caused it.
Two things worth saying out loud, since this is exactly where an AI PM interview earns its name. First, the fix on the table is a rejected alternative most candidates reach for by reflex: put the pause back to four seconds. That was considered, and ruled out, because it would undo a real, measured drop in callers hanging up, trading one measured cost for another instead of removing the actual race. Second, the real bar for the fix isn't "never let two channels see the same table." Hearth's table match score, party size against table type, already only auto assigns above a calibrated bar tuned against a same day eval set, and anything below that bar still routes to Rosalie's own queue instead of guessing. The shared ledger is the same idea applied to timing instead of matching, a deterministic check the model's own confidence has to clear before it acts alone.
And if you want to be sure it really works, try it somewhere else
Same five questions, a completely different incident, a different flip family, and a different industry, so the method proves itself instead of just repeating a story you happened to prepare.
Coalter Mechanical dispatches HVAC technicians across a metro area, using Dovecote, a scheduling assistant that scores which tech fits which job and routes the easy calls without a human touching them.
F. Etta Verlaine, senior dispatcher, sixteen years routing technicians, able to tell a five minute filter fix from a four hour compressor job just from how a customer describes the noise on the phone.
L. Once Dovecote's match scores proved reliable, Etta let a newer scheduler run the whole emergency board alone during her lunch hour, three months running, because Dovecote's picks kept being good ones.
I. A different flip than Marrow's. This is a delegation flip, not a workaround. Dovecote started misrouting big commercial refrigerant jobs, which need a licensed tech, to techs who were great at routine residential calls but not licensed for that work. Three of those went out wrong in two weeks. Etta didn't spot check the commercial jobs after that. She took the entire emergency board back, residential included, even though residential was never the problem.
P. Dovecote launched with one blended confidence bar covering every job type, because commercial jobs were a small enough slice of daily volume that a second, separate bar felt like extra engineering for a handful of calls a week. Reasonable, at the time, given how thin that slice of the data actually was.
S. Give commercial jobs their own, stricter auto dispatch bar, calibrated against a commercial only eval set, so anything below it routes straight to Etta instead of auto assigning. Replay the same two weeks: the junior scheduler still clears about ninety five percent of the board alone, and Etta reviews only the three or so commercial calls a day that actually need her, not all forty.
Swap the trigger and it still runs.
Speed: an interviewer cuts you off after ninety seconds. Skip straight to naming the decision that let two channels both say yes, and what replaces it.
Cost: engineering says the shared ledger can't ship for six weeks. Don't leave both channels racing in the meantime. Ship the cheap partial version first, the busier channel checks the last few seconds of the other channel's holds before confirming, even before the fix is wired both ways.
The model got better, for real: suppose Hearth's parsing of what a caller actually wants got measurably more accurate. That still isn't the same claim as two channels never racing each other. A better read of "the corner booth for four" doesn't stop it from being confirmed twice within the same eight seconds.
Where people run it wrong.
They point the answer at the person, "she should have double checked," instead of at the decision that let two channels each trust their own read of the board.
They propose "add a review step" or "tighten the confidence bar," which is a new dial turned up, not a decision taken back.
They stop at "make the model more accurate," which does nothing for a race between two systems that both already thought they were right.
How to use it live. Repeat the incident back in one plain sentence before you answer anything. It buys you real thinking time without looking like you're stalling, and it's the same first move every single time, no matter what incident lands on the table.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't putting the pause back to four seconds the obvious, cheap fix?" Response: no, because that undoes a real, measured drop in call abandonment. The fix has to remove the race between the two channels, not reintroduce the wait that caused a different, already solved problem.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Incident management for AI products
- #1 What counts as an incident for an AI feature but not for a normal one?
- #2 Write the severity definitions for AI quality incidents.
- #3 Your model starts producing offensive output. Describe the first hour.
- #4 How do you triage an incident where the code is fine and the model is the problem?
- #5 What is the AI equivalent of a rollback, and when is it not available?
- #6 Describe the on-call runbook entry for a sudden quality drop.