CaseFoundationalDesigning for Uncertainty & Trust / Designing for failure and graceful degradation / #1

What should happen in the UI when the model returns nothing usable?

SPARK design the anchor for the moment the model has nothing to say

Halcyon Cove Hotels runs Aravel, an AI concierge tool its front-desk staff open on a shared tablet. Priya Nakashima has worked the evening desk at the Cove for six years. Here is what Aravel did the night it could not find a safe dinner match, and what should have happened instead.

The direct answer
Never render a blank screen or a spinner that just times out. When the model can't produce a usable result, show what it tried, name what's uncertain, and give one concrete next step, usually a person to call. For anything touching safety, make that handoff the default, not a link buried behind a "why" tap.
Do this, in order
  1. Never show a blank screen for "nothing usable." Show a named not-sure state with a next step.Why: a blank screen reads as "no options exist," not "the model failed."
  2. For anything safety-related, make the escalate-to-a-person path the default, not a second tap away.Why: the misses that can actually hurt someone are exactly the ones staff are least likely to go looking for on their own.
  3. Tell the person handling it what the model actually attempted, not just that it came up empty.Why: without that, staff either trust the blank screen too much or waste minutes guessing what broke.
  4. Log every "nothing usable" event by request category, not only as one failure rate.Why: a single number hides which category is quietly the dangerous one.
  5. Leave the fast, confident path alone for high-volume, low-stakes requests.Why: a fallback state on wifi passwords and checkout times is friction with nothing to protect.
  6. Route repeat "nothing usable" categories back to the product team weekly, not just to the front desk.Why: a gap the night shift quietly works around every week never gets fixed if it only ever lives in their heads.

How to answer this, stage by stage

Nobody is grading whether your empty state looks polished. They're grading whether it survives the one night a guest's safety was riding on it.

Stage 1
Scope it to one screen, one moment
Say it like this
"I'll answer this for Aravel, a concierge tool at Halcyon Cove Hotels, for the exact moment it returns nothing usable to a front-desk agent standing in front of a guest."
Why this works
Keeps the question from staying a general UX opinion about error states.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK: situation, payoff, anchor, risk, keep out. Anchor is the actual design decision, and it's the answer to this question."
Why this works
Signals a method, not a list of screen ideas pulled from memory.
Stage 3
Ground it in what happens today, without Aravel
Say it like this
"Before Aravel, an agent pulled out a laminated binder, called a partner restaurant directly, and waited a few minutes for a real, specific answer."
Why this works
Shows the fallback path already existed once. The redesign is bringing it back, not inventing it.
Stage 4
Give the anchor, the one decision
Say it like this
"Instead of a blank card, show: 'Couldn't confirm a safe match. Here are three partners to call directly, or ring the on-call concierge lead at extension 402.' Always a name, always a number."
Why this works
This is the direct answer, stated as an actual on-screen sentence instead of a design philosophy.
Stage 5
Prove the anchor survives its own risk
Say it like this
"The first time this happens on a genuinely dangerous request, an allergy, a medical need, the screen never leaves the agent guessing. It always names a person to call, whether or not the model ever figures out why it struggled."
Why this works
Answers the real question underneath: what happens the one time the stakes are highest and the model has nothing.
Stage 6
Close on the one line
Say it like this
"An empty result and a wrong result are not the same problem. An empty result should always hand the person a name and a number, never a blank screen."
Why this works
Restates the direct answer in one breath, ready for a follow-up push.

Let's learn

What should a screen say when it has nothing to say? Aravel is the case that answers it. It's an AI concierge tool a boutique hotel chain gives its front-desk staff, so they can answer guest questions fast without a phone call.

Before Aravel, an agent kept a laminated binder of partner restaurants and called ahead to confirm a dietary fit, a real answer, but it took Priya about six minutes per guest on a busy night.

Hand sketched flow diagram titled Today, without Aravel. Five boxes: Guest asks at desk, Check partner binder, Call restaurant direct highlighted, Confirm dietary fit, Tell guest confident.
Aravel replaced the first two boxes with something instant. It never replaced what happens when the third box has no answer at all.

Now Aravel answers most requests in about 20 seconds, checkout times, wifi passwords, local transit, no call needed.

Here's the turn: the speed was never the problem. The problem showed up on the rare, compound request, mixing an allergy with a specific cuisine, where Aravel sometimes has nothing usable to return. The screen just goes blank, or spins and times out, and staff, trained to trust the instant answer, read the blank screen as "there is nothing available" instead of "the tool couldn't finish the job."

Share of requests where Aravel returns nothing usable, by category
10% 5 0 Checkout, 0.5% Wifi, 0.2% Transit, 2% Dietary, 8%
The failure rate on dietary and allergy requests is sixteen times the failure rate on checkout questions. The old design gave every one of those categories the exact same blank screen.

At its worst, someone with a real allergy gets told, in effect, there is nowhere safe to eat tonight, when three partner restaurants down the street could have handled it easily. Nobody lied to that guest. A blank screen just did the job a lie usually does.

Hand sketched comparison titled The night it was wrong. Left, a grey box icon labeled Old screen, caption blank then a spinner. Right, a blue document icon labeled New screen, caption names the gap, one next step.
Same guest, same allergy, same night. Only what the screen was willing to admit changed.
The decision I would take back Aravel's screen was built with exactly two states: a result card, or nothing. There was no third state for "I'm not sure, ask a person." That made sense while the model rarely failed to answer at all. It stopped making sense once compound dietary requests grew common enough, about one in twelve, that the blank state started showing up on the requests that mattered most.

What I would leave alone: checkout times, wifi passwords, pool hours. These almost never come back empty, and even when one does, nothing bad happens if a guest waits thirty seconds for a human answer. They don't need a hard escalate path built around them.

The lesson: a blank screen is not neutral. It reads as an answer, "no," even though nothing was ever decided. If a screen has nothing to say, it has to say that, out loud, with somewhere to go next.

Now here is the same thing as a story

The short version above is what you'd say defending this redesign to Halcyon Cove's guest experience team. Read this one for how close the near miss actually came.

Priya can calm down an angry guest in under a minute flat, a skill six years on the evening desk will teach anyone. Aravel arrived about a year ago, and for months it made her job faster without changing how she thought about it at all.

A guest would ask about a good dinner spot nearby, Priya would tap it into the tablet, and an answer would come back in seconds, matching what she'd have told them anyway. Slowly, she stopped double-checking Aravel's answers against her own knowledge of the neighborhood, since it kept being right.

Knowledge spark: why would an AI tool return nothing at all? Sometimes a model can't find a confident match for what it was asked, especially a request with several conditions stacked together, an allergy plus a cuisine plus a price range. Rather than guess and risk a wrong answer, it may return an empty or low-confidence result. That's often the safer choice for the model to make. The mistake is in what the screen does with that emptiness next.

On a busy Friday night, a guest named Mr. Okonkwo-Hart stopped at the desk. He had a serious shellfish allergy, wanted something close by, and was hungry after a long travel day. Priya typed the request into Aravel. The screen loaded, spun for a few seconds, and then simply cleared, no card, no message, nothing.

Hand sketched timeline titled The night the screen went blank. Five milestones: Guest asks 7 48pm, Aravel returns nothing 7 49pm highlighted, Priya says none available 7 50pm, Shift lead overhears 7 52pm, Guest seated safely 8 05pm.
Two minutes between a blank screen and a guest being told there was nothing. A near miss, not a headline, only because someone happened to be nearby.
The screen never said "no." It just never said anything, and Priya filled in the silence herself.

Priya, seeing nothing on the screen and trusting Aravel's usual speed, told the guest she didn't have anything close by that could safely handle his allergy tonight. He thanked her and turned toward the elevator, planning to skip dinner.

The shift lead, walking past the desk, overheard the tail end of it and asked what Aravel had actually returned. Priya turned the tablet around: a blank card, nothing else. The shift lead knew of three partner restaurants within six blocks that handle shellfish allergies as a matter of routine, called one directly, and the guest was seated within fifteen minutes.

Hand sketched labeled parts diagram titled The anchor, close up. A document icon at the center labeled New empty state, with four callouts around it: what it tried, what is unsure, one next step, escalate now.
The old screen showed none of these four things. It just showed nothing, and let the person in front of it guess.

With the redesigned screen, Aravel now says: "Couldn't confirm a safe shellfish-free match nearby. Three partners worth calling directly: Marchetti's, Bellhaven Kitchen, The Tide Table. Or ring the on-call concierge lead, ext. 402." Run the same night forward: Priya calls Bellhaven Kitchen herself at 7:50pm, and Mr. Okonkwo-Hart is seated by 8:02pm, three minutes sooner than the version where a stranger happened to be walking by.

The old screen asked Priya to read a blank space and guess what it meant. The new one just tells her what it tried and who to call.

We built the empty state to be quiet, on purpose, so it wouldn't look broken. It took a shift lead's lucky timing to see that quiet and blank read as exactly the same thing to the person holding the tablet.

SPARK, in one screenNot a lecture on making an empty state feel gentle. SPARK is what forces you to design against the one night it actually matters.

S
Situation. How this gets handled today, without the tool.
An agent pulls a laminated binder and calls a partner restaurant directly to confirm a dietary fit, slower, but always a real answer.
Grounds the design in a fallback that already exists, rather than one invented from scratch.
P
Payoff. The habit this should build.
Reading "not sure, here's who to call" as a real answer worth acting on, instead of reading a blank screen as "there is nothing."
Names the actual behavior change, not just a feeling of reassurance.
A
Anchor. The one decision everything hangs on.
Never a blank card. Always: what it tried, what's unsure, and one concrete next step, a name and a number, especially for anything safety-related.
This is the hardest step and the direct answer: a concrete, arguable design decision, not a mood.
R
Risk. What breaks the first time it's wrong.
The first time the model has nothing usable on a genuinely dangerous request. The anchor survives it because the escalate path never depends on the model figuring out why it struggled.
Proves the anchor was built against its own failure, not just described in a deck.
K
Keep out. What we won't build, day one.
No auto-booking the fallback restaurant, no AI-written explanation of the medical risk. Those are judgment calls for a person, not a screen.
Shows restraint, instead of a longer feature list dressed up as a fix.
Hand sketched quadrant titled Where a hard escalate path earns its place. Axes how often it fails from rare to common, and how much a miss can hurt from small to serious. Wifi password and late checkout sit low on both axes. Transit timing sits in the middle. Dietary match sits high on how much a miss can hurt.
Only the top of this chart ever needs a hard, always-on escalate path. Everywhere else, the fast answer alone is fine.

The recap, one line per letter: situation is the laminated binder and the phone call, payoff is teaching staff that "not sure, call this person" is a real answer, anchor is the redesigned empty state and its always-visible next step, risk is the night the model has nothing on a dangerous request, and keep out is holding back auto-booking and AI-written medical explanations.

Hand sketched icon list titled What we left for later. Three items: a box icon labeled No auto booking the fallback restaurant, a document icon labeled No AI written medical explanation, a gauge icon labeled No change to the simple fast answers.
Each of these is a real feature someone will eventually ask for. None of them belongs in the first version.

And if you want to be sure it really works, try it somewhere elseSame five letters, a veterinary triage assistant instead of a hotel concierge. A different anchor, aimed at a different empty screen.

Fernbrook Animal Clinic runs an intake assistant that helps front-desk staff triage phone calls about a sick pet before a vet is free to call back. Deja Whitfield answers those calls. Mapped onto SPARK: situation is a receptionist today, working from a printed symptom checklist and her own judgment about which calls need a vet immediately; payoff is the habit to build, treating "I can't tell from this description" as a real answer that gets a call escalated, instead of a guess dressed up as reassurance.

The anchor here aims at a different failure: when the assistant can't confidently rank how urgent a call is, given vague symptoms over the phone, it never shows a default "routine" label just to fill the screen. It shows "couldn't rank this call. Escalate to the on-call vet now," with the specific symptoms it couldn't reconcile listed underneath. The risk the clinic's team designed against was a receptionist seeing a blank urgency field, assuming that meant routine, and telling a worried owner to wait until morning for a pet that needed same-day care.

Hand sketched labeled parts diagram titled The anchor, close up, reused here for the veterinary intake assistant. Center document icon labeled New empty state, with what it tried, what is unsure, one next step, and escalate now around it, relabeled for a call triage screen.
Swap "safe dinner match" for "how urgent is this call," and the same four parts still do the same job.
Same-day vet visits missed on the first call, with and without a forced escalate label
20% 10 0 label ships Month 1, 14% Month 5, 3%
The drop tracks the same month the forced escalate label shipped, not any change in the underlying triage model itself.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "never a blank screen, always what it tried and one next step, hard-coded for anything safety-related," and stop.
Cost: there's no time to redesign every empty state across the product this quarter. Say so honestly, and start with the categories where a miss can actually hurt someone, not the ones that are simply annoying.
The model gets better, for real: if Aravel's match rate improves and "nothing usable" becomes rare, that's still not a reason to remove the named next step, a rarer miss on a dangerous category is exactly the miss this design exists to catch.

Where people run it wrong.
They treat "no result" and "wrong result" as the same design problem, when a blank screen needs an entirely different fix than a mistaken one.
They bury the escalate path behind a "why didn't this work" link nobody taps under time pressure.
They wait for a complaint to notice the gap, when the person most at risk is the one least likely to know enough to complain.

How to use it live. When someone asks what should happen when the model has nothing, ask yourself one question first: what does the person in front of the screen do in the next ten seconds? Design the empty state to give them an actual next move, not a blank space to interpret on their own.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "what should the UI do when the model fails" question?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Ground the anchor in what already happens without the tool, then prove it survives being wrong.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Nakashima, a front-desk agent at Halcyon Cove Hotels for six years, who could calm an angry guest in under a minute.
3 · THE SITUATION
How does this get handled today, without Aravel?
Tap to flip
ANSWER
An agent pulls out a laminated binder of partner restaurants and calls one directly to confirm a dietary fit, about six minutes per guest.
4 · THE ANCHOR
What's the one design decision this answer hangs on?
Tap to flip
ANSWER
Never a blank card. Always show what it tried, what's unsure, and one next step, a name and a number, hard-coded for safety-related requests.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building Aravel's screen with only two states, a result card or nothing, with no third state for "I'm not sure, ask a person."
6 · THE NUMBER
Fill in the blank: dietary and allergy requests return nothing usable about ___ percent of the time, versus 0.5 percent for checkout questions.
Tap to flip
ANSWER
About 8 percent. Sixteen times the checkout rate, and the exact category where a blank screen can genuinely hurt someone.
7 · THE REPLAY
Same night, redesigned screen. What changes?
Tap to flip
ANSWER
Priya sees three named partners and an on-call number, calls one herself, and the guest is seated by 8:02pm instead of relying on a shift lead happening to overhear.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the anchor there?
Tap to flip
ANSWER
Fernbrook Animal Clinic's phone intake assistant. The anchor forces a visible "couldn't rank this, escalate now" label instead of a blank urgency field defaulting to routine.

Check yourself Score: 0 / 0

Multiple choice
1. Why does the redesigned Aravel screen always name a person to call, instead of just trying to explain why the match failed?
  • A. Guests prefer talking to a human over reading text.
  • B. The escalate path has to work even when the model can't say why it struggled, especially on safety-related requests.
  • C. It reduces how many requests Aravel has to process.
  • D. Hotel policy requires a phone call for every guest request.
Show hint
Look at the Risk step.
Show answer
B. The whole point of the anchor is that it never depends on the model explaining itself correctly, it just always hands off.
True or false
2. True or false: the redesigned empty state adds a hard escalate path to the wifi password and checkout time requests too.
  • True
  • False
Show hint
Look at "what I would leave alone" and the quadrant diagram.
Show answer
False. Those requests rarely fail and cost little when they do, so the redesign leaves them exactly as they were.
Fill in the blank
3. Fill in the blank: Aravel returned nothing at 7:49pm, and by ___ the shift lead had gotten the guest seated at a partner restaurant.
Show hint
Look at the timeline diagram.
Show answer
8:05pm. The redesigned screen would have gotten him seated by about 8:02pm, without needing a shift lead to overhear anything.
Short answer, name the reversal
4. What old design decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Building the screen with only two states, a result or nothing. It made sense while the model rarely failed to answer at all, so a third "not sure" state felt like unneeded complexity.
Short answer, apply it yourself
5. Think of a time an app you used showed you a blank screen or a spinner that never resolved. Did you read that as "nothing exists," even if that wasn't true?
Show hint
Ask what you did next: did you try again, give up, or look for another way?
Show answer
Model answer: Most people can recall treating a blank or stuck screen as a final "no," the same mistake Priya made, and the exact gap this answer's anchor is built to close.
Short answer, where it wouldn't matter
6. Name a request type in Aravel where this hard escalate design genuinely doesn't need to apply.
Show hint
Look at the bottom-left of the quadrant diagram.
Show answer
Model answer: Wifi password lookups. They almost never fail, and even when one does, nothing bad happens if a guest waits a moment for a human answer.
Before you close the answer
Why this works
Tests whether you design an empty state around what a real person does in the next ten seconds, or just make a blank screen look tidier and call the problem solved.
Follow-up traps
"Isn't naming three specific restaurants a lot of manual upkeep for the product team?" Response: yes, and that's an honest cost. The list only needs to cover the small number of categories that are actually safety-related, not every category Aravel handles.

"What if the agent ignores the escalate prompt anyway, the way Priya almost did?" Response: that's exactly why the redesign puts a name and a number directly in the message, not behind a link, so acting on it takes no extra decision.
If pressed
Aravel logs every "nothing usable" event with the specific attributes the model couldn't reconcile, not just a failure flag, so the product team can see a request like "shellfish allergy plus Italian cuisine" clustering before a guest ever gets hurt by it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more