CaseAdvancedResponsible AI & Advanced Practice / Agent product management specifics / #7
Design the interruption and takeover experience for a running agent.
FLIPS the product is Marrow Route, a shipment-rerouting agent for Harrow Logistics
Harrow Logistics moves refrigerated freight across five states. Marrow Route is the agent that reroutes and rebooks shipments during storms and port delays. Casey Lindqvist supervises the night shift and used to keep the live action feed open on a second monitor every single hour.
The direct answer
Don't design the interruption experience assuming someone is watching. Design it assuming they've stopped, because a long streak of correct decisions is exactly what makes a person stop watching. Force a brief, explicit acknowledgment checkpoint on a fixed schedule, not a passive feed anyone can ignore, and keep a one-tap full takeover always available for the moment that checkpoint catches something wrong.
Do this, in order
Force an active checkpoint on a fixed schedule, never a passive feed someone merely could watch.Why: a live feed only protects you if someone's eyes are actually on it, and a good streak is exactly when they won't be.
Keep a one-tap full takeover control available at every moment, not just at checkpoints.Why: the rare near-miss that isn't caught by the schedule still needs a fast way to grab the wheel.
Show a plain-language summary of consequential actions at each checkpoint, not a raw action log.Why: a tired supervisor skims a log. They read three sentences.
Tune the checkpoint cadence to the disruption's pace, not a single fixed number.Why: during an active storm, ten actions can happen in two minutes; a fixed time-only cadence would miss that entirely.
Don't force a checkpoint on routine, low-consequence actions with no real streak to watch.Why: interrupting every single reroute regardless of stakes trains people to click through without reading, which recreates the same problem.
How to answer this, stage by stage
Seven stages. The interviewer is watching for whether your design assumes continuous human attention, because that's the assumption that always breaks.
Stage 1
Scope it to one real agent
Say it like this
"I'll answer this for Marrow Route, an agent that reroutes and rebooks refrigerated freight shipments during storms and port delays."
Why this works
Keeps "design the interruption experience" from turning into a UI pattern library recitation.
Stage 2
Name the trap in the question itself
Say it like this
"Most interruption designs assume someone is watching closely enough to know when to interrupt. The real failure mode is that a good streak makes people stop watching entirely, so the design has to work even when nobody's looking."
Why this works
Reframes the question before answering it, which is what separates a strong candidate from a list of button designs.
Stage 3
Find the person and the habit
Say it like this
"Casey watched every single reroute live for the first few weeks. After a long run with zero mistakes, she stopped watching the feed at all and only checked a daily summary."
Why this works
Grounds the abstract "over-trust" idea in one person's actual, sensible behavior.
Stage 4
Name the flip precisely
Say it like this
"There's no half-watching a live feed. Either her attention is on it right now, or it isn't. That's the flip: watches every action, or watches none, with nothing in between."
Why this works
Shows this is a real two-setting switch, not a gradual dial, which is the hardest part of any FLIPS answer.
Stage 5
Give the actual design
Say it like this
"A forced checkpoint every ten actions or fifteen minutes, whichever comes first, with a plain-language summary Casey has to actively acknowledge, plus a one-tap takeover button always visible."
Why this works
This is the concrete decision the whole answer turns on, stated plainly enough to defend.
Stage 6
Name the cost you're accepting
Say it like this
"This adds a few minutes of pause during an active disruption, when speed matters most. I'd rather pay a few minutes than let a bad reroute run unnoticed for hours."
Why this works
States the latency trade-off out loud instead of pretending the checkpoint is free.
Stage 7
Close on what you'd leave alone
Say it like this
"Low-stakes reroutes, ones that are easy to reverse, don't need this checkpoint at all. Forcing it on everything would just train people to click through without reading."
Why this works
Shows judgment about where the fix genuinely matters, not blanket caution applied everywhere.
Let's learn
Marrow Route watches weather and port data and reroutes or rebooks freight shipments in real time, sometimes touching dozens of shipments in a single disruption event.
Before Marrow Route, a dispatcher rerouted shipments by hand during a storm, one call at a time, catching problems immediately because they were the one making every decision.
Knowledge spark: what's a "forced checkpoint"?
A pause the system makes for you, not one you have to remember to take. Instead of a feed you could check any time, the agent stops and waits for an actual acknowledgment before it keeps going, on a schedule the design controls.
With Marrow Route live, a storm night that used to take a dispatcher three hours to work through by hand now resolves in under ten minutes of agent time. That's the number in the ops review. It's also not the number that should have worried anyone.
The turn. The extra speed isn't the problem. What changed is Casey's own behavior. A single-turn routing suggestion tool never has this problem, because a person reads every suggestion before acting on it by definition. An agent that acts on its own removes that built-in read, and whether anyone's watching becomes a design choice instead of an automatic fact.
Share of Marrow Route's actions actively watched live, by week
Nobody flipped a switch to stop watching. The line just kept drifting down, week after week, as the streak of correct reroutes kept getting longer.
The decision I would take back
We built the monitoring experience as a passive live feed on a second monitor, with no forced acknowledgment step, because it felt lighter-weight and less annoying to check when Casey wanted to. That made sense in week one, when she was genuinely watching it constantly anyway. It stopped making sense once a long run of correct reroutes made watching feel unnecessary, and a feed nobody's required to look at stops functioning as oversight the moment nobody's actually looking.
What I would leave alone: low-stakes, easily reversible reroutes, a minor delay adjustment on a non-perishable shipment, genuinely don't need a forced checkpoint. Adding one there just trains people to acknowledge without reading, which recreates the exact problem the checkpoint exists to prevent.
We didn't lose a habit of catching mistakes. We lost a habit of looking, and those turned out to be the same thing.
The lesson: a monitoring design that depends on someone choosing to keep watching will eventually stop being watched, and it will stop exactly when things have been going well, which is the worst possible time for it to stop.
Now here is the same thing as a story
The short version above is what you'd say defending this checkpoint design to Harrow's operations leadership. Read this one for how the gap actually got found.
Casey Lindqvist has supervised Harrow's night shift for five years. She can tell which storm cell will actually disrupt a route before the weather service updates its own forecast.
Nothing dramatic happened at any single point on this timeline. That's exactly what made it easy to miss.
Marrow Route launched in the spring, and for the first six weeks Casey kept the live feed open every hour of her shift, watching every reroute as it happened. It never once got a decision wrong that mattered. By week fourteen, she'd settled into checking only the next-morning digest, since the live feed had stopped showing her anything she hadn't already seen play out correctly a hundred times before.
The exact same number of reroutes happened both weeks. Only one thing changed, and it wasn't on any dashboard.
Then came a night in October when a coastal storm forced Marrow Route to reroute forty shipments in under twenty minutes. One of them, a truck of refrigerated seafood, got rerouted onto a route with a documented three-hour weigh-station delay the agent's own data had access to but weighed as acceptable against a slightly shorter total distance.
The monitoring design was built for a dial that slowly needed less attention. What actually happened was a switch that turned off entirely.
Nobody caught it until the next morning's digest, by which point the seafood had sat past its safe window and the shipment was a total loss. Casey hadn't been careless. She'd done exactly what six weeks of a perfect track record taught her to do: trust it.
Harrow's team rebuilt the monitoring experience around a forced checkpoint instead of a passive feed: every ten actions, or every fifteen minutes during an active event, Marrow Route pauses and shows a plain-language summary of what it just did, and Casey has to actively acknowledge it before it continues.
The third box didn't exist before October. It's the only thing that changed, and it's the whole fix.
Replayed under the new design: the same coastal storm forces the same forty reroutes. At action twenty-three, the checkpoint pauses and shows Casey a plain summary including the seafood truck's new route and its flagged delay. She catches it in the eight minutes it takes to reach that checkpoint, reroutes that one truck again herself, and the shipment arrives on time.
Minutes until a bad reroute is caught, old design vs. new design
The old design didn't catch the bad reroute faster because nobody watching the feed was any less capable. It caught it slower because nobody was watching at all.
I built the monitoring feed as passive because in week one, that lighter design felt like the right amount of friction for someone watching constantly anyway. It took losing a full truck of seafood to see that the same lightness becomes the exact gap a long, genuine streak of success would eventually walk right through.
FLIPS, the five lettersThe I step, over-trust, is the flip family that fires when things are going well, which is exactly why it's easy to miss.
F
Find the person.
Casey Lindqvist, Harrow's night-shift supervisor, who could once spot which storm cell would actually cause trouble before the forecast updated.
Grounds the flip in one real person's actual, reasonable behavior.
L
Locate the habit.
She stopped watching the live feed hour by hour, then day by day, until only the next-morning digest remained.
Names the habit as a rational response to six good weeks, not a failure of attention.
I
Identify the flip.
Watches every action live, or watches none at all. There's no half-watching a live feed; either her attention is on it or it isn't.
The over-trust family: the trigger here is good news, a long correct streak, which is why this flip is so easy to miss designing for.
P
Pinpoint the old decision.
We built the monitoring feed as passive, with no forced acknowledgment step, because it felt lighter-weight when Casey was watching constantly anyway.
A removed-affordance reversal: taking out the friction of a forced pause because, at launch, nobody needed it.
S
Show the replay.
The same storm, the same forty reroutes, but the bad one gets caught at the eight-minute checkpoint instead of the next morning's digest.
Ends in something countable: 8 minutes instead of 360, and a shipment saved instead of lost.
The recap in one image: five letters, and the I step is the one this whole design problem lives inside.
The recap, one line per letter: find the person is Casey watching every reroute at launch, locate the habit is her watching thinning from hourly to daily, identify the flip is watching everything versus watching nothing with no in-between, pinpoint the old decision is the passive feed with no forced pause, and show the replay is catching the bad reroute at minute eight instead of minute 360.
And if you want to be sure it really works, try it somewhere elseSame five letters, a pharmacy refill agent instead of a freight router. A completely different flip family this time: not over-trust, but a workaround, someone building their own private check because the takeover view gave them nothing to trust.
Coldkeep Pharmacy runs a refill-authorization agent that approves routine prescription refills on its own. Benedek Farkas is the pharmacist who reviews anything the agent flags as unusual.
Find the person: Benedek, five years at Coldkeep, knows which patients' refill patterns are worth a second look before the agent even flags them. Locate the habit: he used to trust the agent's takeover view completely when he needed to step in. Identify the flip, a different family this time, the workaround flip: he doesn't stop watching, he starts keeping his own paper log of every takeover, because the agent's takeover screen shows only its current recommendation with no history of what it already decided or reversed on that same patient. Pinpoint the old decision: the team never built a run-history or comparison view into the takeover experience, assuming a single current recommendation was enough context for anyone stepping in. Show the replay: with a redesigned takeover view that shows the agent's last three decisions for that patient side by side, Benedek stops needing his private paper log entirely, since the actual product now gives him the comparison he was building by hand.
A dosage change flagged for review needs a hard stop even though it's rare. A routine refill needs almost nothing, even though it happens constantly.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "design for the moment nobody's watching, not the moment someone is, with a forced checkpoint and a one-tap takeover," and stop.
Cost: there's no budget this quarter for a full checkpoint system. Say so honestly, and start with a single mandatory acknowledgment at the halfway point of any multi-action event, since even one forced pause beats a purely passive feed.
The model gets better, for real: if Marrow Route's routing accuracy improves further, that's still not a reason to remove the checkpoint, since a longer streak of good decisions is exactly what makes the next miss harder for a person to catch on their own.
Where people run it wrong.
They build a visible "stop" button and call the interruption experience done, without asking whether anyone will actually be watching closely enough to press it.
They design the checkpoint cadence around a fixed clock instead of the pace of actions, missing a fast-moving disruption entirely.
They force the same heavy checkpoint on every action regardless of stakes, which trains people to acknowledge without reading.
How to use it live. If you get stuck designing an interruption experience, ask yourself one question: does this design still work on the day the agent has been right for six straight weeks? If your honest answer is "only if someone happens to still be watching," you haven't designed the interruption yet, you've designed a feed.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "design the interruption and takeover experience for a running agent"?
Tap to flip
ANSWER
FLIPS, using the over-trust flip family: the flip fires when a long streak of correct decisions makes a person stop watching entirely.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Casey Lindqvist, who supervises Harrow Logistics' night shift and could once spot a storm cell's real risk before the forecast updated.
3 · THE HABIT
What did Casey stop doing because the agent kept working?
Tap to flip
ANSWER
Watching the live action feed. She went from checking it hourly to checking only the next-morning digest, over about fourteen weeks.
4 · THE FLIP
What's the two-setting switch here?
Tap to flip
ANSWER
Watching every action live, or watching none at all. There's no half-watching a live feed in real time.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building the monitoring feed as passive, with no forced acknowledgment step, since it felt like the right amount of friction while Casey was watching constantly anyway.
6 · THE NUMBER
Fill in the blank: under the old design, the bad reroute was caught after ___ minutes, versus 8 minutes under the new forced-checkpoint design.
Tap to flip
ANSWER
360 minutes. That's the next morning's digest, by which point the refrigerated shipment was already a total loss.
7 · THE REPLAY
Same storm night, redesigned checkpoint. What changes?
Tap to flip
ANSWER
Casey catches the bad reroute at the eight-minute checkpoint, reroutes the seafood truck herself, and the shipment arrives on time instead of spoiling.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Coldkeep Pharmacy's refill-authorization agent, using the workaround flip: a pharmacist builds his own private paper log because the takeover view showed no history.
Check yourself Score: 0 / 0
Short answer, recall the flip
1. What was the flip in Casey's story, and what were its two settings?
Show hint
Look at the "identify the flip" step.
Show answer
Model answer: Watching every action live versus watching none at all. There's no partial setting for a live feed, since it only protects you if someone's eyes are on it in that moment.
Multiple choice
2. Why couldn't Casey have just "watched the feed a bit more carefully" instead of stopping entirely?
A. She was told not to watch it anymore.
B. A live feed only works as oversight if someone's attention is actually on it right now; there's no partial version of that.
C. The feed was removed from her dashboard.
D. She didn't trust the agent enough to keep checking.
Show hint
Think about why a flip needs exactly two settings, no middle.
Show answer
B. "Watching more carefully sometimes" isn't a real state for continuous, real-time monitoring. Either it's happening or it isn't.
True or false
3. True or false: this answer recommends forcing a checkpoint on every single reroute Marrow Route makes, regardless of how minor it is.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Low-stakes, easily reversible reroutes don't get the forced checkpoint, since applying it everywhere would just train people to click through unread.
Fill in the blank
4. Fill in the blank: the redesigned checkpoint triggers every ___ actions, or every 15 minutes during an active event, whichever comes first.
Show hint
Look at the redesigned interrupt flow.
Show answer
Ten. That cadence caught the bad reroute at action twenty-three during the replay, well before all forty actions completed.
Short answer, where it wouldn't matter
5. Name a place in Marrow Route's design where this forced-checkpoint fix genuinely wouldn't matter.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A minor delay adjustment on a non-perishable shipment. It's low-stakes and easy to reverse, so a forced checkpoint there adds friction without much real protection.
Short answer, apply it yourself
6. Pick a product you use yourself that runs on its own for a while. What's one habit of watching it closely that you'd probably stop doing if it kept working well for months?
Show hint
Think about an automated feature you checked carefully at first but stopped double-checking once it never seemed to go wrong.
Show answer
Model answer: A common one: checking that an automatic bill payment or a scheduled backup actually ran correctly, which most people stop verifying after enough uneventful months.
Before you close the answer
Why this works
Tests whether you'll design an interruption experience that assumes continuous human vigilance, the assumption that always breaks, or one that forces attention on a schedule the product controls instead of hoping someone's still watching.
Follow-up traps
"Won't a forced checkpoint every ten actions get annoying during a big storm with hundreds of reroutes?" Response: the cadence scales with the event's pace, using whichever threshold hits first, actions or minutes, so a slow night barely triggers it while a fast one catches problems quickly without demanding constant attention.
"Isn't a visible stop button enough on its own?" Response: no, a stop button only helps if someone's already watching closely enough to know to press it, which is exactly the assumption a long correct streak breaks.
If pressed
Harrow's actual checkpoint summary highlights any action whose outcome falls outside the agent's own predicted range, like a route with an unusually long delay, rather than listing every action with equal weight, so the three-sentence summary stays genuinely short even during a busy event.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.