ConceptAdvancedResponsible AI & Advanced Practice / Agent product management specifics / #16

What is the right latency expectation for an agent, and how do you set it with users?

PICK the product is SwiftPath, Falkenridge Air's disruption rebooking agent

Interviewer's question: "What is the right latency expectation for an agent, and how do you set it with users?" Falkenridge Air's SwiftPath agent rebooks passengers during weather cancellations and mechanical delays. Nadira Vosloo is the product lead. Osric Dumaine is a passenger stranded on a cancelled connection.

The direct answer
Show a fast, honestly-labeled interim result within a few seconds for every disrupted passenger, then keep working underneath and swap in a better itinerary if one appears, telling the passenger when it changes. Never make someone wait in silence for the "real" answer, and never let the fast answer pretend to be final.
Do this, in order
  1. Show an honestly-labeled interim hold within seconds, never a silent wait.Why: a passenger decides whether to trust the system in the first five seconds, long before any answer is complete.
  2. Confirm interim seat inventory against the live reservation system, not the agent's own memory.Why: a fast answer built on stale inventory is a hallucinated hold, not a real one.
  3. Keep optimizing in the background and notify the passenger if a better itinerary appears.Why: the first hold should never be treated as the final word, or the product quietly ships a worse routing.
  4. Cap the background optimization pass so a mass-disruption event doesn't run unlimited lookups per passenger.Why: the full search costs real money in interline API calls, and a storm night can multiply that cost by thousands of passengers at once.
  5. Watch how many passengers never see an upgrade after the interim hold.Why: that number tells you whether "temporary" is actually landing, or people are quietly accepting the fast answer as final.

How to answer this, stage by stage

Nobody is grading whether you know agents can feel slow. They are grading whether you can name which kind of wait actually costs you the passenger.

Stage 1
Ground it in one moment
Say it like this
"I'll answer this for SwiftPath, Falkenridge Air's rebooking agent, during a weather cancellation, since that's when latency actually matters."
Why this works
A latency question answered in the abstract turns into a generic UX opinion. A real disruption event keeps the stakes concrete.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, impact, cost asymmetry, kill criteria."
Why this works
Names a method before picking a side, so the answer doesn't turn into a loose list of pros and cons.
Stage 3
Take the position
Say it like this
"Show a fast interim hold within a few seconds, clearly marked temporary, then keep optimizing underneath."
Why this works
The P step, and the direct answer. States the pick before any reasoning, which is what "it depends" answers fail to do.
Stage 4
Show the impact on both sides
Say it like this
"The passenger standing at the gate feels every extra second as anxiety. The airline feels a bad interim hold as rework later, a worse seat, a second agent call."
Why this works
The I step. Names who feels each kind of error, in real units, not just "user experience."
Stage 5
Find the cost asymmetry
Say it like this
"A wrong interim hold costs us a few seconds of correction later. A silent wait costs us the passenger, who rebooks themselves the moment it feels broken."
Why this works
The C step, the heart of PICK. Says which error is cheap and visible and which is hidden and expensive, and optimizes against the second one.
Stage 6
Give the kill criteria
Say it like this
"If most passengers accept the interim hold and never see an upgrade notice, that tells me 'temporary' isn't landing, and I'd slow the first answer down instead."
Why this works
The K step. Shows the position is a confident choice, not a stubborn one, by naming what would flip it.
Stage 7
Close on the actual tradeoff
Say it like this
"Fast and honest beats slow and complete, because the passenger standing there decides whether to trust us in the first five seconds, not the last ninety."
Why this works
Restates the direct answer in the shape of the actual question asked.

Let's learn

SwiftPath watches for flight disruptions and rebooks affected passengers automatically, without a phone agent working the case by hand.

Before SwiftPath, a passenger calling about a cancelled flight during a storm night waited an average of twenty two minutes on hold before a human agent could even start searching for a new seat.

Hand sketched timeline titled One passenger's 90 seconds. Four milestones: Second 0 hold shown emphasized, Second 3 inventory confirmed, Second 45 still optimizing, Second 90 upgrade or lock in.
Ninety seconds sounds slow next to a spinner. It sounds fast next to twenty two minutes on hold.

Now, the moment a flight is marked cancelled, SwiftPath shows the passenger a held seat on the next available flight within about three seconds, clearly marked as temporary, while it keeps searching every interline option underneath for up to ninety seconds.

Knowledge spark: what does "interline" mean here? A rebooking on a different airline entirely, one Falkenridge Air doesn't operate. Searching those options takes longer, since it means checking another airline's live seat inventory, not just Falkenridge's own.

Here is the turn: three seconds versus ninety seconds is not really the argument. The real argument is whether the wait, however long it runs, is silent or honest. A passenger who sees an instant, clearly-labeled hold relaxes. A passenger staring at a spinner for even ten seconds starts calling the airline's help line on a second phone, undoing the whole point of automating this.

How long passengers waited for any answer, before and after SwiftPath
22 min 11 min 0 22 min hold Before SwiftPath 3 sec hold, 90 sec final After SwiftPath
The bar on the right is not actually invisible, it is just that fast. The whole design bet is that the gap between them is what a passenger will forgive.

At its worst: a passenger's interim hold gets treated as final because nothing on screen said otherwise, they board a routing that cost them a four hour layover, and never learn a better one existed forty seconds after they stopped looking.

Hand sketched metaphor scene titled The asymmetry, drawn. Left, a circle icon labeled seen and fixed, caption a hold that upgrades in view. Right, a scale icon labeled hidden and costly, caption a silent wait, a lost booking.
Both panels are labeled as costs. Only one of them can quietly cost you the passenger forever.
The decision that mattered Show a fast interim hold within a few seconds, marked clearly as temporary, and keep optimizing underneath instead of making the passenger wait in silence for a "complete" answer.

What I would leave alone: a routine same-day rebooking with no disruption, a passenger asking to move to an earlier flight on a normal Tuesday, doesn't need any of this. The answer there is already fast and already final, and adding a "temporary" label would just create doubt where none belongs.

The lesson: latency is not one number to hit. It's a decision about which kind of wait a passenger will forgive, and silence is never the kind they forgive.

Now here is the same thing as a story

The short version above is what you'd say defending this design in a product review. Read this one for how the gap actually got found.

Osric Dumaine's connection got cancelled at 9:40pm on a night Falkenridge Air grounded eleven flights for weather. He had a hotel booked forty minutes from the wrong airport and a nine year old asleep on his shoulder.

He opened the app expecting to wait. Instead, three seconds after the cancellation posted, SwiftPath showed him a seat on the 6:05am flight the next morning, marked "Temporary hold, still checking for a better option."

Hand sketched flow diagram titled SwiftPath, second by second. Four steps: disruption hits, instant hold emphasized, keep optimizing, notify upgrade.
The second step is the one the whole design bet rides on.

He didn't love the 6:05am flight. But he had a seat, in three seconds, on the worst night of Falkenridge's month. He put the phone down and went to find his son a blanket.

Eighty seconds later, quietly, SwiftPath found a 4:50am departure on a partner airline with a shorter layover. It swapped the hold and sent a notification. Osric didn't see it until he woke his son at 4am for the shuttle, and by then the better seat was simply the one waiting for him.

Osric never felt like he waited at all. He felt like the system had been quietly working the whole time he wasn't looking.

Nadira Vosloo's team almost shipped a different version. In an early design, SwiftPath said nothing for the first ninety seconds, so the very first thing a passenger ever saw was the fully optimized answer, no interim hold, no "still checking." It tested beautifully in calm conditions, where ninety seconds barely registered.

Then came a real storm night with four thousand passengers rebooking inside twenty minutes. The optimization queue backed up. Some passengers waited past four minutes for any answer at all, staring at a spinner, with no idea whether the app had frozen or was simply thinking.

Hand sketched decision tree titled Which answer to show now. Root: passenger needs rebooking. Three branches: inventory confirmed live leads to show instant hold, no live confirmation leads to queue and disclose wait, better fare found later leads to upgrade and notify.
The middle branch is the one the silent design never built.

Support calls spiked, not because the rebooking was wrong, but because nobody could tell if anything was happening. Several passengers called Falkenridge's help line and got rebooked manually by a human agent, moments before SwiftPath's own optimized answer would have landed, wasting both the passenger's patience and the agent's time on a case the system was already solving.

The decision Nadira's team took back: silence was never neutral. Waiting for a "complete" answer with nothing shown in between reads, under real load, exactly like a broken product. They rebuilt it to show the instant hold first, every time, live-inventory-confirmed, and let the optimization keep running visibly underneath.

Replayed on the same storm night with the new design: every one of those four thousand passengers sees a seat within three seconds. The ones whose hold later improves get a notification instead of silence. Support calls on disruption nights dropped by more than half, not because SwiftPath got smarter, but because nobody was left wondering if it was even working.

I once believed the "real" answer was the honest one and anything shown before it was a lesser version pretending to be finished. It took one bad storm night to see that a fast, clearly-labeled interim answer is not a lesser version of honesty. Silence dressed up as thoroughness is the actual dishonesty.

PICK, in one screenNot a UX debate about spinners. PICK is what tells you which kind of wait a passenger will forgive.

P
Position. Say the pick first.
Show a fast interim hold within a few seconds, honestly labeled, and keep optimizing underneath.
The direct answer, stated before any reasoning, which is what separates a pick from "it depends."
Hand sketched icon list titled What every interim hold needs. Four items: inventory confirmed in the last 30 seconds, a clear temporary label, background optimizing still running, a path to notify the upgrade.
Skip the first item and a fast answer becomes a confidently wrong one.
I
Impact. Who feels each error.
The passenger feels every extra second of silence as rising doubt. The airline feels a bad interim hold as rework, a worse seat, a repeat support call.
Names both sides in real units, not a vague "user experience" tradeoff.
C
Cost asymmetry. The heart of it.
A wrong interim hold costs a few seconds of correction, in public, with the passenger watching it improve. A silent wait costs the whole booking, quietly, the moment the passenger stops trusting and calls someone else.
The hardest step. Names which error is cheap and visible and which is hidden and expensive, and optimizes against the second one.
K
Kill criteria. What would flip it.
If most passengers keep the interim hold without ever seeing an upgrade, "temporary" isn't landing as a real signal, and the position should shift toward a slower, more clearly final first answer.
Shows the pick is a confident choice, not a stubborn one, by naming the exact evidence that would change it.
Share of rebooked passengers who never saw an upgrade notice, by week since launch
60% 30% 0% kill line: 50% Week 1 Week 4 Week 8 Week 12
The line sits well under the kill line the whole way. If it ever climbed toward it, that would be the signal to slow the first answer down instead.

The recap, one line per letter: position is a fast honest hold, not a silent wait, impact is the passenger's anxiety against the airline's rework cost, cost asymmetry is a visible correction beating a hidden lost booking, and kill criteria is watching how many passengers never see an upgrade at all.

And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary telehealth line instead of an airline gate. A different cost asymmetry, the same shape of decision.

Thornvale Veterinary Telehealth runs an agent that answers a pet owner's call at 11pm, gives immediate triage guidance, and can either resolve the call or queue a full callback from an on-duty vet who reviews the pet's history first.

Mapped onto PICK: position is the same shape, show fast interim guidance immediately ("keep them calm, don't give food or water, here's what to watch for") rather than making a frightened owner wait in silence for a vet to become free. Impact: the owner feels every silent minute as rising panic over a possibly sick animal; the clinic feels a rushed, incomplete answer as a real risk of missing something serious. Cost asymmetry, and here it flips: an interim answer that's too confident about a symptom can be genuinely dangerous, not just annoying, so the interim guidance has to be strictly limited to safe, reversible advice, never a diagnosis, while the full vet callback is what actually cost asymmetry protects. Kill criteria: if pet owners start treating the interim guidance as a full diagnosis and skip the vet callback altogether, that is evidence the interim answer needs to state its limits more forcefully, not just faster.

Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "fast honest interim beats slow silent complete, because silence is what a passenger can't forgive," and stop.
Cost: if interline API calls get too expensive to run on every passenger, cap the optimization pass by time or by fare-class tier instead of removing the interim hold, since the interim hold is the part that protects trust.
The model gets better, for real: if SwiftPath's full search gets fast enough to run in three seconds flat, the interim/final split still matters during a true mass-disruption event, when queueing thousands of full searches at once would recreate the exact same silent wait a faster median hides.

Where people run it wrong.
They treat "make it faster" as the whole answer, and never ask what the wait actually communicates while it's happening.
They ship a spinner with no interim result because it feels more honest than showing something that might change, and then watch trust collapse under real load anyway.
They let the fast answer look identical to the final one, so nobody notices when it quietly upgrades, and nobody notices when it doesn't.

How to use it live. When someone asks you about agent latency, ask yourself one thing out loud: if this exact wait went on twenty times longer than expected tonight, would the user know something was happening, or would they assume it broke.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question about the right latency for an agent?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Built for tradeoff questions where you have to commit to a side.
2 · THE PEOPLE
Who are the two people this answer centers?
Tap to flip
ANSWER
Nadira Vosloo, SwiftPath's product lead, and Osric Dumaine, a passenger rebooked during a storm-night mass cancellation.
3 · THE POSITION
What is the actual pick in this answer, in one line?
Tap to flip
ANSWER
A fast, honestly-labeled interim hold within seconds, with the optimization continuing visibly underneath, beats a silent wait for a "complete" answer.
4 · THE ASYMMETRY
Which error is cheap and which is hidden and expensive here?
Tap to flip
ANSWER
A wrong interim hold is cheap and visible, it corrects in view. A silent wait is hidden and expensive, it quietly costs the whole booking.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Showing nothing for the first ninety seconds so the first thing a passenger ever saw was the fully optimized answer, with no interim result in between.
6 · THE NUMBER
Fill in the blank: before SwiftPath, a disrupted passenger waited on hold for an average of ___ minutes.
Tap to flip
ANSWER
Twenty two. Against that number, even a ninety second full optimization reads as fast, as long as the wait isn't silent.
7 · THE KILL CRITERIA
What evidence would flip this pick toward a slower, more final first answer?
Tap to flip
ANSWER
Most passengers keeping the interim hold and never seeing an upgrade notice, which would mean "temporary" isn't landing as a real signal.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and how does the cost asymmetry flip there?
Tap to flip
ANSWER
Thornvale Veterinary Telehealth. There, an interim answer that's too confident can be actually dangerous, so it has to stay limited to safe, reversible advice.

Check yourself Score: 0 / 0

True or false
1. True or false: the right fix for SwiftPath's early design was simply to make the full optimized search run faster.
  • True
  • False
Show hint
Look at "the decision that mattered."
Show answer
False. The fix was adding an honest interim hold, not just speeding up the final answer. Even a fast final answer can fail under a real load spike, silence is the actual problem.
Multiple choice
2. Why is a silent wait worse than a wrong interim hold, according to this answer?
  • A. A wrong interim hold never actually happens in practice.
  • B. A wrong interim hold corrects in view and costs a few seconds. A silent wait is hidden, and it costs the whole booking the moment trust breaks.
  • C. Silent waits are more expensive in server costs.
  • D. Passengers always prefer to wait for the most accurate answer.
Show hint
Look at the cost asymmetry stage.
Show answer
B. A visible, correctable error is cheap. A hidden one that breaks trust is expensive, because the passenger has already rebooked themselves elsewhere by the time anyone notices.
Fill in the blank
3. Fill in the blank: SwiftPath shows a passenger a temporary hold within about ___ seconds of a cancellation.
Show hint
Look at the bar chart comparing before and after.
Show answer
Three. That is the number that has to be confirmed against live inventory, not the agent's own memory, or the fast hold becomes a hallucinated one.
Short answer, where it wouldn't matter
4. Name a rebooking situation in this story where the interim-versus-final distinction wouldn't matter.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine, non-disrupted rebooking, like a passenger moving to an earlier flight on a normal day. That answer is already fast and already final.
Short answer, apply it yourself
5. Think of a tool you use that shows a loading spinner with no other information. What would an honest interim result look like there instead?
Show hint
Think about a search tool, a food delivery app, or a customer support chat.
Show answer
Model answer: A delivery app that shows "still finding the fastest driver" with a rough estimate, instead of a blank spinner, lets you trust the wait instead of wondering if the app froze.
The number question
6. If the share of passengers who never see an upgrade climbed from 30 percent toward the 50 percent kill line, what should change about the design?
Show hint
Look at the kill-criteria line chart.
Show answer
Model answer: The "temporary" framing isn't landing as a real signal to passengers, so the design should shift toward a slightly slower but more clearly final first answer, rather than continuing to treat the interim hold as good enough on its own.
Before you close the answer
Why this works
Tests whether you can name the real cost asymmetry in a latency tradeoff instead of just saying "make it fast." Most candidates stop at "reduce latency" without saying what the wait needs to communicate.
Follow-up traps
"Isn't showing an interim result just as confusing as a spinner?" Response: no, because it's confirmed against live inventory and clearly labeled temporary, so it's a real, honest answer, not a placeholder pretending to be one.

"What if passengers get annoyed by the upgrade notification changing their seat?" Response: that's a much smaller cost than the alternative, a passenger who silently keeps a worse seat forever because nobody told them a better one existed.
If pressed
The background optimization pass is capped at ninety seconds and a fixed number of interline lookups per passenger specifically during a mass-disruption event, so a storm night affecting thousands of passengers at once can't run unlimited searches and blow through cost or queue up new silent waits of its own.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more