Design the recovery path for a user whose agent took a wrong action.
Ticketwell is a support-desk platform whose resolve-agent can read a ticket and act on it directly: issue a refund, print a return label, close the case, no human click required. Greta Vandenberg runs support operations for Voltframe Audio, a headphone maker that sells through resellers, and one ordinary Tuesday morning her agent did something she never asked it to.
- Give every high-value agent action its own reversible card with a single "reverse this" button.Why: a recovery path buried in a scrolling ticket thread is not a recovery path, it's a scavenger hunt.
- Show what's still reversible versus already in motion, honestly.Why: a refund not yet settled and a wire already sent are different problems, and pretending otherwise wastes the first ten minutes.
- Scope the undo to exactly what the agent did, never a blanket account rollback.Why: reversing everything from the last hour would undo good actions along with the one bad one.
- Add a pause before any action past a dollar or unit-count threshold, not before every action.Why: friction on every single-unit refund would slow the 99 percent of cases the agent already gets right.
- Notify the affected outside party automatically the moment a reversal fires.Why: a reseller who quietly gets their refund clawed back without a word trusts the platform less than one who never got the wrong refund at all.
- Leave small, single-unit refunds fully autonomous, no pause added.Why: those were never the tickets that went wrong, and slowing them down fixes nothing.
How to answer this, stage by stage
Nobody is grading whether you can describe an undo button. They're grading whether your undo button actually matches the size of what went wrong.
Let's learn
What actually goes wrong the day an autonomous agent's action turns out to be a mistake? Not the mistake itself. It's that nobody, not the agent, not the person who deployed it, can say in the moment what's still fixable and what's already gone.
Before the resolve-agent, a Voltframe support rep read every reseller ticket by hand, checked the order history, and issued a refund that matched exactly what was wrong, usually in about twenty minutes per ticket.
Now the resolve-agent reads the same ticket and issues the refund in under a second, no rep required, for the overwhelming majority of cases.
Here's the turn: the extra speed was never the problem. The problem showed up the first time the agent misread scope, not accuracy, treating a partial complaint as a whole-order return, and there was no card, no flag, and no button, just a line in a ticket thread that looked exactly like every correctly-handled ticket next to it.
At its worst, a reseller relationship built over years takes a real hit, and finance spends a full day untangling a refund and a set of return labels nobody meant to send, while every other correctly-handled ticket that week sits right next to it looking identical.
What I would leave alone: single-unit, low-dollar refunds don't need a pause at all. Those were never the tickets that went wrong, and slowing them down protects against nothing.
The lesson: speed was never the risky part of automating a refund. Scope was. A fast wrong answer and a slow wrong answer cost the same amount of money, the fast one just leaves less time to notice.
Now here is the same thing as a story
The short version above is what you'd say defending this design to Voltframe's finance team. Read this one for the actual Tuesday morning.
Greta Vandenberg has run support operations at Voltframe Audio for five years, long enough to know which resellers call the moment something's wrong and which ones quietly stop reordering instead.
Ticketwell's resolve-agent had been live for four months, handling the flood of "my headphones arrived with a scratch" tickets that used to eat her team's whole morning. It was, by every measure she tracked, working.
A reseller ticket came in at 9:02am: "2 of the 40 units in our last order arrived with cracked cases, please advise." At 9:03am, the agent issued a full refund and a return label for all 40 units. At 9:04am, the shipping confirmation email went out to the reseller, who was, understandably, delighted and confused in equal measure.
Nobody noticed until the next morning, when finance flagged an unusual refund during a routine reconciliation, more than a full day after it happened. By then the reseller had already started boxing up all 40 units to ship back, confused about whether to send the two broken ones or all of them.
With the redesigned system, that same 9:03am refund becomes its own action card the moment it fires, past a 500-dollar threshold: "Refunded 3,800 dollars for 40 units, ticket described 2 units as damaged. Reverse this?" Greta sees it at 9:06am, three minutes after it happened, not the next day. One click cancels the unused return labels automatically and reissues a correct 190-dollar refund, and a second message goes to the reseller explaining the correction before they've boxed up a single working unit.
The old system asked Greta to trust that a fast resolution was a correct one. The new one shows her exactly what the agent did and lets her undo just that, in minutes, not a full news cycle later.
I signed off on instant execution because a pause felt like exactly the kind of friction this product was built to remove. It took one pallet-sized refund to see that removing the pause removed the only place a mistake could still be caught small.
SPARK, in one screenNot a lecture on undo buttons. SPARK is what tells you which single decision the whole recovery path actually depends on.
The recap, one line per letter: situation is a rep matching the refund to the actual problem by hand, payoff is teaching Greta to trust a scoped reverse instead of losing faith in the whole agent, anchor is the action card with its single reverse button, risk is the return labels already shipped, and keep out is holding back full account-wide rollback.
And if you want to be sure it really works, try it somewhere elseSame five letters, a logistics dispatch agent instead of a support desk. The thing already in motion is a truck, not a wire transfer.
RouteHollow Logistics runs a dispatch agent that reassigns delivery routes in real time when a driver calls in sick or traffic shifts. Nadim Calloway is a dispatcher there. Mapped onto SPARK: situation is a dispatcher manually swapping routes by radio, checking which driver has room and which deliveries are time-sensitive; payoff is the habit to build, trusting a scoped route-reversal instead of grabbing the radio and re-routing everything by hand the moment something looks off.
The anchor is structurally the same idea: every route reassignment above a certain lateness-risk score becomes its own card, naming the delivery, the new driver, and the delay it introduces, with one button to send it back to the original driver. The risk RouteHollow designed against: a truck carrying a same-day medical supply delivery had already left the depot on the new, slower route by the time anyone noticed the reassignment was wrong, so the card has to say plainly "already en route, 12 minutes behind schedule" instead of implying a click undoes a truck that's already on the highway.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "card every high-value action with a scoped reverse button, and show what's already in motion," and stop.
Cost: there's no engineering time to build a full action-card system this quarter. Say so honestly, and start by cording just the single highest-dollar action type first, refunds over 500 dollars, before expanding to the rest.
The model gets better, for real: if the agent's scope-reading genuinely improves and this exact mistake gets rare, that's still not a reason to remove the card, a rarer mistake is exactly the one a team stops watching for.
Where people run it wrong.
They build a single "undo my agent" button that reverts everything recent, which undoes good actions right along with the bad one.
They treat every autonomous action as equally risky and add a pause everywhere, which slows down the 99 percent the agent already gets right.
They wait for a customer complaint or a finance audit to catch the mistake, instead of asking upfront what a wrong action even looks like on a dashboard.
How to use it live. When someone asks you to design a recovery path for a wrong agent action, ask yourself one question first: if this action is wrong, can the person recovering from it tell, in one glance, what's still reversible? Design around the answer being no by default.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the agent's mistake is small enough to stay under the threshold every time?" Response: that's what the detection dashboard from the general blast-radius design is for, watching for a pattern of small, correct-looking actions that add up, not just single large ones.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Designing for failure and graceful degradation
- #1 What should happen in the UI when the model returns nothing usable?
- #2 Design the fallback experience for an AI feature when the provider is down.
- #3 Explain the difference between failing loudly and failing silently, and which you prefer.
- #4 How do you design a feature that degrades to a non-AI version rather than breaking?
- #5 Describe three failure modes to design for before launch.
- #6 What error message would you write for a model timeout, and what would you avoid saying?