CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #19

How would you handle uncertainty in an agent that has already acted?

ORDER rank the response by what a rider can't get back, not by what's easiest to check first

Solent Rail runs regional trains along the south coast. Anaya Fitzgerald has led Customer Recovery Operations there for six years. Waypoint is the agent that rebooks a passenger automatically the moment her train is cancelled, without waiting for anyone to approve it first.

The direct answer
Rank the response to an already-executed action by what's hardest to undo, not by what's easiest to check first. Tell the affected rider immediately, in plain words, with a real person to call, before you touch the model's threshold or write the incident report. A missed connection can't be undone once the replacement service has already left without her.
Do this, in order
  1. Notify the affected rider immediately, in plain language, with a real person to call.Why: a missed connection can't be undone once it's already missed, and every silent minute makes it worse.
  2. Give every autonomous action a working undo or human-override path, not just a record of it.Why: a log entry helps you learn later, it does nothing for the rider standing on the wrong platform right now.
  3. Log the action's confidence and reasoning the moment it happens, before any human reviews it.Why: you can't audit a decision you can't reconstruct, and memory fades fast on a busy dispatch floor.
  4. Audit similar past actions at the same confidence level, going back at least a month.Why: one caught mistake is rarely the first one, it's just the first one anyone happened to notice.
  5. Only then adjust the autonomy threshold, once the audit shows where it actually breaks.Why: tightening the threshold before you know the real failure mode just moves the risk somewhere else unseen.
  6. Leave low-stakes, easily-undone actions, like a same-train seat change, fully autonomous.Why: adding a pause to something reversible in seconds slows the product down for no real protection.

How to answer this, stage by stage

The interviewer isn't grading whether you'd add a review step. They're grading whether you know which step actually can't wait.

Stage 1
Scope it to one agent, one incident
Say it like this
"I'll answer this for Waypoint, the agent Solent Rail lets rebook passengers automatically the moment a train is cancelled, using a real incident at Tarnwick Junction."
Why this works
Keeps "an agent that already acted" from staying a hypothetical with no stakes attached.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what we're actually protecting. Reversibility, which response is hardest to undo. Dependency, what unblocks what. Evidence, what's cheap to learn first. Rank, the actual order."
Why this works
Tells the interviewer this is a method, not a list of things that sound responsible.
Stage 3
Name the real outcome, not the system's uptime
Say it like this
"The outcome isn't 'Waypoint rebooked the passenger fast.' It's whether she actually reaches somewhere real, safely, the same day it happens."
Why this works
Without this, "rank the response steps" is just a list with no way to argue for the order.
Stage 4
Give the rank, before any reasoning
Say it like this
"Notify the rider and offer a live line first. Log and audit second. Adjust the confidence threshold third. Rewrite the wider policy last. That's the order I'd defend."
Why this works
This is the direct answer, said as a concrete sequence instead of a pile of good intentions.
Stage 5
Prove the rank with the near miss
Say it like this
"Waypoint once rebooked forty riders onto a bus route that had been suspended that morning. It took nearly two hours before anyone even knew there was a problem, because the message it sent never said why, and there was no one to call."
Why this works
Turns "notify first" into a specific, believable near miss instead of a generic safety slogan.
Stage 6
Close on the one line
Say it like this
"Rank by what a person can't get back, not by what's easiest to check first. Telling her outranks fixing the model, every time, because the model can wait a week and she can't wait an hour."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

Waypoint is an agent a regional rail operator lets act on its own. The moment a train gets cancelled, it rebooks every connecting passenger onto the next real service, without waiting for a person to click approve.

Before Waypoint, a dispatcher called ahead to confirm a connecting service before rebooking anyone, about twelve minutes per group. On a bad night, a team could only get through maybe twenty groups before running out of time, so plenty of riders got no help at all.

Hand sketched flow diagram titled What happens after Waypoint acts. Five boxes: Train cancelled, Agent rebooks instantly, Passenger notified, Human confirms new service highlighted, Recovery if wrong.
The whole question lives in that fourth box. Solent Rail had quietly skipped it for partner services.

Now Waypoint moves hundreds of riders in seconds during a major disruption. No queue, no wait, no dispatcher touching a keyboard for the routine cases.

Here's the turn: the occasional wrong rebooking was never really the problem. The real problem only shows up after the agent has already acted, because a rider who trusts a rebooking message has no way to know, in that moment, whether the plan just written for her is actually still true.

Minutes between an autonomous rebooking and a rider getting a real explanation, by month
12 min 6 0 5 min concern line Month 1 Month 4 Month 5
The gap crossed five minutes back in month two, quietly, as Waypoint's scope grew to include partner bus routes. Nobody was watching that particular number yet.

At its worst, that silent gap strands a rider at a junction near midnight, holding a message that says her journey is "updated," with no way to reach anyone and no idea the replacement service it named stopped running six hours earlier.

Hand sketched comparison titled Reversible or not. Left, a teal circle icon labeled Send the message, caption minutes, easy to undo. Right, an orange box icon labeled Change the threshold, caption weeks, not undone fast.
Solent Rail's incident review kept reaching for the right-hand box first. The rank runs the other way.
The decision I would take back Waypoint's team merged "the agent acts" and "the agent tells you" into one silent step, on the assumption that an automatic push notice counted as a real explanation. That made sense when Waypoint only rebooked riders onto Solent Rail's own trains, a low-stakes, fully reversible move. It stopped making sense the day Waypoint's scope grew to include a partner operator's bus routes it couldn't directly verify were still running.

What I would leave alone: rebooking a rider automatically within Solent Rail's own network, seat to seat, train to train, doesn't need this scrutiny. That's reversible in the time it takes to print a new ticket.

The lesson: an agent that has already acted owes a different kind of honesty than one that's still deciding. It can't ask permission anymore, so it has to make the undo fast and the explanation faster.

Now here is the same thing as a story

The short version above is what you'd say defending this rank to Solent Rail's safety board. Read this one for how close the near miss actually came.

The dispatch floor at Solent Rail gets loud whenever a storm rolls in off the channel, and the night Tarnwick Junction lost its signal feed was no exception. Anaya Fitzgerald has run Customer Recovery Operations there for six years, and can read a service disruption off the board before the first phone call even comes in.

For the first several months after Waypoint launched, it handled routine cancellations cleanly, rebooking riders onto the next train on Solent Rail's own line within seconds, and Anaya's team barely had to touch a keyboard.

Knowledge spark: why would a rebooking agent even touch a partner's bus route? When a train line is fully down, the only real alternative is often a replacement bus run by a separate transport company. The rail operator's own system has to reach out to that partner's schedule feed to know which runs are actually operating, and that feed isn't always refreshed as often as the rail operator assumes.

As Waypoint's scope grew to include those partner bus routes, not just Solent Rail's own trains, the team quietly stopped double-checking those rebookings too, since months of them had gone fine.

That Tuesday night, a signal fault cancelled the 9:56pm service through Tarnwick Junction. Waypoint rebooked all forty connecting passengers onto replacement bus route 217 in about two seconds. Route 217 had been suspended that morning for road works, a fact sitting in a feed Waypoint's confidence check hadn't refreshed since 6am. Every rider got the same push notice: "Your journey has been updated." Nothing about why, nothing about what to do if it looked wrong.

Hand sketched timeline titled Hour by hour, the night Waypoint acted alone. Five milestones: Train cancelled 9 56pm, 40 riders rebooked 9 58pm, No one told why 10 05pm highlighted, Platform found empty 11 42pm, Recovery coach sent 12 20am.
Nearly two hours between the silent rebooking and anyone finding out it was wrong.

At 11:42pm, a station attendant doing a final round noticed platform three at Tarnwick was empty when forty riders should have been waiting for a bus. She called it in.

We did not lose forty riders to a bad model. We lost them the moment the message that rebooked them stopped being an offer and started being a decision with nobody's number on it.

Solent Rail dispatched a recovery coach, and every rider made it home, the last one at 1:05am, an hour and forty minutes after they should have been on their way. Nobody was hurt. It could easily have gone differently on a colder night.

Hand sketched icon list titled Signs an agent already made a bad call. Three items: a question box icon labeled No one can say why it acted, a gauge icon labeled Confidence stayed high on stale data, a funnel icon labeled Complaints reach the wrong team first.
Anaya's team found all three signs sitting in Waypoint's own logs once they went looking.

Anaya's team pulled Waypoint's logs after the fact and found this wasn't the first stale-feed rebooking, just the first one anyone happened to catch. With the redesigned response, Waypoint still acts in two seconds, but now sends a real explanation alongside the rebooking, and it holds any partner-service rebooking behind a live status ping before finalizing it.

Hand sketched labeled parts diagram titled What's in a good post action message. A document icon at the center labeled Disclosure, with four callouts around it: what happened, why it happened, how to undo it, who to call now.
All four were sitting in Waypoint's own logs already. None of them had ever reached a passenger's phone.

Run the same Tuesday night forward: the stale bus route fails its status ping automatically, Waypoint holds the forty passengers on their original delayed train instead of a phantom bus, and pages Anaya's team directly. They are moved by 10:20pm, not 1:05am.

The old design let the agent's silence stand in for confidence. The new one makes the agent say what it doesn't actually know yet.

I thought the fast rebooking was the whole product. It took forty people standing on an empty platform to see that the message telling them what happened mattered just as much as the decision itself.

ORDER, in one screenNot a generic incident checklist. ORDER is what tells you which step on that checklist actually goes first.

O
Outcome. What we're actually protecting.
Not "Waypoint's average rebooking speed," but whether the rider actually reaches somewhere real, safely, the same day.
Without this, ranking the response steps is just opinion.
R
Reversibility. Which response is hardest to undo.
Telling a rider immediately and offering a live line is fast and cheap, and impossible to backdate once the moment's passed. Changing the autonomy threshold is slow but can wait.
This is the hardest step, and the one the whole rank turns on.
D
Dependency. What unblocks what.
You can't reasonably change the confidence threshold before an audit tells you which kind of action was actually the risky one.
Some of this order is forced by reality, not just judgment.
E
Evidence. What's cheap to learn first.
Pull thirty days of past rebookings at the same confidence band and see how many touched a service that, in hindsight, wasn't actually running.
Buys real information before committing to a threshold change.
R
Rank. State the order and defend the top pick.
Notify and offer a live line first, log and audit second, adjust the threshold third, rewrite the wider policy last.
Gives the interviewer a real, arguable sequence instead of a vague "act responsibly."
Hand sketched quadrant titled Response steps ranked. Axes how hard to undo from easy to hard, and how urgent from can wait to must happen now. Notify passenger and offer human override sit low on hard to undo and high on urgent. Audit past cases sits in the middle. Change threshold sits high on hard to undo and low on urgent.
The two most urgent steps are also the two easiest to undo. That's not a coincidence, that's the whole rank.

The recap, one line per letter: outcome is naming what a rider actually needs, not a system-uptime number, reversibility puts telling the rider above touching the model, dependency is the audit gating the threshold change, evidence is thirty days of past cases pulled cheaply before committing to anything bigger, and rank is the order itself, notify first, threshold last.

And if you want to be sure it really works, try it somewhere elseSame five letters, a municipal power utility instead of a rail operator. A different dependency breaks the second story.

Loadwatch is an agent a municipal electric utility uses to cycle down non-critical home loads, like pool pumps and water heaters, during a heat-wave peak, to prevent a wider blackout. Idris Bramwell, the utility's Grid Operations Lead, handles it after Loadwatch cycles off equipment for eleven thousand homes without asking anyone first. Mapped onto ORDER: outcome is keeping the actual neighborhood grid up, not "peak load reduced by X percent"; reversibility puts restoring power and notifying affected homes above adjusting Loadwatch's own trigger threshold, since a family with a medically necessary device losing power for even twenty minutes is far harder to make right than tuning a setting next week.

The dependency here runs in reverse from Solent Rail's case: at Loadwatch, the safety-critical accounts list has to be checked before any cycling happens at all, not audited afterward, because some homes on that grid depend on powered medical equipment and can never be included in an automatic cutoff, no matter how confident the model is about the wider peak. The evidence step there was matching Loadwatch's cutoff list against the utility's own registered medical-equipment accounts, a check that already existed for a different purpose and had simply never been wired into the automated agent.

Hand sketched labeled parts diagram titled What's in a good post action message, reused here for Loadwatch's outage notice. Center document icon labeled Disclosure, with what happened, why it happened, how to undo it, and who to call now around it.
Same four parts. At a power utility, "how to undo it" means restoring the circuit directly, not offering an apology.
Hours to reverse each response, once an agent has already acted
400 hours 0 Send a real explanation, about 3 minutes Offer a human override, 0.5 hours Audit similar past cases, 40 hours Change the autonomy threshold, 400 hours
Sending the explanation costs almost nothing and can't wait. Changing the threshold is nearly a hundred times slower, and it can.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "notify the affected person and offer a real fallback first, log and audit second, only then touch the model," and stop.
Cost: there's no budget for a live status-ping check on every partner rebooking right now. Say so honestly, and start with just the audit, since finding out how often this has already happened costs almost nothing and tells you how urgent the rest actually is.
The model gets better, for real: if Waypoint's confidence calibration genuinely improves and stale-feed rebookings stop happening for six months, that's still not a reason to remove the live status ping, a good track record earns a wider autonomy scope, not a shortcut past checking.

Where people run it wrong.
They treat the incident report and the threshold change as the urgent work, because those feel like "fixing the real problem."
They let an automatic notification count as disclosure, without ever asking whether the person receiving it can actually act on it.
They wait for a complaint to learn an agent's already-taken action was wrong, instead of auditing past actions on their own schedule.

How to use it live. The moment someone describes an agent that's already acted, ask yourself one question before ranking anything: if this turns out to be wrong, what does the affected person lose for every minute nobody tells them? Put that step first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "what would you do first" question about an agent that already acted?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. It ranks the response by what's hardest to undo, not by what's easiest to check first.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Anaya Fitzgerald, Head of Customer Recovery Operations at Solent Rail, who has run that desk for six years.
3 · THE OUTCOME
What's the real outcome this response is protecting?
Tap to flip
ANSWER
Whether the rider actually reaches somewhere real, safely, the same day, not a generic measure of rebooking speed.
4 · THE RANK
State the actual response order, top to bottom.
Tap to flip
ANSWER
Notify the rider and offer a live line, log and audit, adjust the confidence threshold, rewrite the wider policy.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging "the agent acts" and "the agent tells you" into one silent step, treating an automatic push notice as a real explanation.
6 · THE NUMBER
Fill in the blank: it took almost ___ hours between the silent rebooking and Solent Rail sending a recovery coach.
Tap to flip
ANSWER
About 2 hours (9:58pm rebooking to a coach arriving after 11:42pm). The gap between acting and explaining was the whole near miss.
7 · THE REPLAY
Same Tuesday night, redesigned response. What changes?
Tap to flip
ANSWER
The stale bus route fails its status ping automatically, riders stay on their original train, and Anaya's team moves them by 10:20pm instead of 1:05am.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and how does the dependency run differently there?
Tap to flip
ANSWER
Loadwatch, a municipal utility's demand-response agent. There, the safety-critical accounts check has to happen before any cycling, not audited after, since some homes can never be cut off automatically.

Check yourself Score: 0 / 0

True or false
1. True or false: Waypoint's forty rebooked riders were told, at the time, exactly why their journey had changed.
  • True
  • False
Show hint
Look at the push notice they actually received.
Show answer
False. They got a generic "your journey has been updated" message, with no explanation and no way to reach anyone.
Multiple choice
2. Why does this answer rank notifying the rider above adjusting the confidence threshold?
  • A. Notifying the rider is required by transport regulation.
  • B. A missed connection can't be undone once it's already missed, while a threshold change can happen next week.
  • C. Notifications are cheaper to build than a threshold change.
  • D. Riders always complain about threshold changes.
Show hint
Look at the Reversibility step.
Show answer
B. Rank by what's hardest to undo. The rider's lost time can't be recovered later; the threshold can be tuned whenever the audit finishes.
Fill in the blank
3. Fill in the blank: at Loadwatch, auditing similar past actions takes about 40 hours, while changing the autonomy threshold takes about ___ hours.
Show hint
Look at the horizontal bar chart.
Show answer
400 hours. Roughly ten times longer than the audit, which is exactly why the audit has to come first.
Short answer, apply it yourself
4. Think of an automated system you've used that acted on your behalf without asking first, like an auto-payment or an auto-renewal. What would you have wanted it to tell you the moment it acted?
Show hint
Ask what you'd have needed to know to undo it fast, if it was wrong.
Show answer
Model answer: Most people want to know what changed, why, and who to contact right now, the same three things Waypoint's redesigned message added.
Short answer, name the reversal
5. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Merging the act and the notification into one silent step. It made sense when Waypoint only touched Solent Rail's own trains, a low-stakes, reversible move, and stopped making sense once it reached a partner's routes.
Short answer, where it wouldn't matter
6. Name a case where this same "notify immediately" scrutiny genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Rebooking a rider automatically within Solent Rail's own network. That's reversible in the time it takes to print a new ticket, so the added pause would only slow things down.
Before you close the answer
Why this works
Tests whether you treat an agent's own action as the finish line, or recognize that uncertainty handling has to keep working after the action, when the person affected has already started acting on it.
Follow-up traps
"Isn't a live status ping just adding latency back into a system built to remove it?" Response: only for the small share of actions crossing into a partner's own service; leave Solent Rail's own network fully instant, since that's the reversible case.

"What if the audit finds the mistake rate is actually tiny?" Response: a tiny rate is still worth knowing before touching the threshold, since the real argument was never "the rate is high," it was "nobody was watching it at all."
If pressed
The redesigned status ping specifically times out after four seconds and defaults to holding the rider on her original plan rather than guessing, since a slow confirmation is itself a signal the receiving service's own systems may be unreliable that day.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more