CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #15

How do you show uncertainty in a voice interface with no screen?

PICK the technician whose hands are inside the machine the assistant is talking about

Arctic Fields Refrigeration services industrial chillers for food-processing plants. Rui Tanaka is a field technician. Wrenchmate is the voice assistant built into his headset, reading live sensor data and answering spoken questions while both his hands stay on the equipment.

The direct answer
Don't hedge every answer the same amount. Answer routine, well-heard readings plainly and fast. But the moment a reading is tied to a live safety action, like opening a valve, or the moment the assistant's own confidence is genuinely low, it stops and reads back what it heard before saying anything else. The extra half-second of confirmation only shows up exactly where a mistake would actually cost something.
Do this, in order
  1. Gate the read-back-and-confirm step on stakes, not on every single answer.Why: confirming every routine reading trains technicians to tune the confirm step out; confirming only high-stakes ones keeps it meaningful.
  2. Hedge with a plain spoken phrase, never a spoken percentage.Why: a number like "eighty-seven percent confident" needs a beat of mental math nobody has spare with their hands inside a compressor.
  3. Always let the technician say "wait" or "repeat" and have it actually re-measure, not just replay the same words.Why: a repeated answer that's wrong twice is worse than useless, it's false reassurance.
  4. Track how often technicians say yes to a confirm without pausing first.Why: a reflexive "yes" means the confirm step has become noise, the exact failure it was built to prevent.
  5. Leave routine, low-stakes queries exactly as fast as they are now.Why: adding caution everywhere would slow down the dozens of harmless checks a technician runs every shift for no real safety gain.

How to answer this, stage by stage

The interviewer already knows a screen can show a number quietly. What they want to know is whether you'll notice that a voice has to say its own doubt out loud, and that saying it too often is its own kind of failure.

Move 1
Anchor it to one technician's shift
Say it like this
"I'll answer this for Rui, a field technician at Arctic Fields Refrigeration, using a headset assistant called Wrenchmate while his hands are inside a running chiller."
Why this works
Keeps "no screen" concrete instead of a design-theory abstraction.
Move 2
State your framework
Say it like this
"I'll use PICK. Position first, then impact, then the cost asymmetry, then what would flip my mind."
Why this works
Signals you're about to commit to something, not explore options out loud.
Move 3
Commit to a position before explaining
Say it like this
"My position: confirm before answering only when the reading is tied to a live safety action, or the assistant's own confidence is genuinely low. Everywhere else, answer straight."
Why this works
This is the direct answer, said plainly, before the reasoning tries to earn it.
Move 4
Name who feels each kind of error
Say it like this
"If it confirms everything, Rui eats a few extra seconds dozens of times a shift, and starts saying yes without listening. If it confirms nothing, one misheard reading near a live valve could actually hurt him."
Why this works
Names both costs in the same breath, which is what makes the tradeoff real instead of hypothetical.
Move 5
Show the near miss that proves the asymmetry
Say it like this
"Compressor noise made Wrenchmate mishear 'unit four' as 'unit fourteen.' It answered confidently, for the wrong unit, right as Rui reached for a service valve."
Why this works
Turns "voice interfaces can mishear things" into one specific, checkable near miss.
Move 6
Say what would change your mind
Say it like this
"If technicians start saying yes to the confirm step without pausing, that's my kill criteria. It means I'd need an active repeat-back instead of a passive one."
Why this works
Shows the position isn't stubborn, it has a stated condition that would overturn it.
Move 7
Land the close
Say it like this
"A voice has to say its own doubt out loud. Say it too rarely, and a real mistake slips through unheard. Say it every time, and nobody's actually listening anymore."
Why this works
Restates the position in one breath, ready for a live follow-up.

Let's learn

Picture doing your job with both hands inside a running machine, and the only screen nearby is one you can't look at anyway.

Wrenchmate reads live sensor data off industrial chillers and answers a technician's spoken questions, like "what's the suction pressure on unit four." Before it existed, a technician either read a printed gauge by hand or radioed a base tech for a second opinion, taking about 5 minutes per check. With Wrenchmate, most answers land in about 30 seconds.

Hand sketched metaphor scene titled A screen shows doubt silently, a voice has to say it. Left, a gauge icon labeled On a screen, caption a number glanced at. Right, a person icon labeled In your ear, caption has to say the doubt out loud.
A screen can hedge by just looking a little uncertain. A voice has no "a little." It either says the doubt, or it doesn't.

Here's the turn: when Wrenchmate first launched, it read back every single measurement out loud for confirmation, a safety-first choice made when usage was rare. As technicians started using it dozens of times a shift for routine checks, the team cut that automatic read-back to speed things up, for every kind of query, not just the routine ones.

Time and risk cost of each design, per shift
200 min 100 0 4 min Always confirm, per shift 180 min One near-miss investigation
The cheap error and the expensive one aren't close. Always confirming costs minutes. One missed confirm on the wrong reading costs a stand-down.

At its worst, an unqualified voice answer near a live safety action doesn't just waste a technician's time. It can send them to open a line believing it's been checked, when it hasn't.

The decision I would take back Cutting the automatic read-back applied to every kind of query at once, high-stakes and routine alike, because at launch the team hadn't yet separated the two categories. That made sense when usage was low and every query felt equally worth double-checking. It stopped making sense once routine checks made up the vast majority of use, and the confirm step started feeling like noise on queries where a mishearing genuinely didn't matter.

What I would leave alone: routine, low-stakes queries, like checking a temperature reading nowhere near a live action, don't need any of this. Adding a confirm loop there would just slow down dozens of harmless checks a shift for no real safety gain.

The lesson: a voice interface doesn't get to stay quiet about its own doubt the way a screen can. It has to actually say something, and saying the same cautious thing every single time teaches people to stop hearing it at all.

Now here is the same thing as a story

The short version above is what you'd say out loud in an interview. Read this one for how close the actual near miss came.

His name is Rui Tanaka. He's serviced industrial refrigeration units for nine years, mostly by feel and by ear, long before Wrenchmate existed. He could tell a failing compressor by the pitch of its hum before any sensor confirmed it.

The good months with Wrenchmate were loud, busy ones. By mid-morning most days, Rui had cleared four or five chiller checks with a quick spoken question and a quick spoken answer, hands never leaving the machine, headset doing the reading for him.

Knowledge spark: why would a voice assistant mishear a plant floor at all? Industrial compressors run loud, often past 80 decibels, close to the volume of a passing truck. Speech recognition software listens for patterns in sound waves, and heavy background noise can bend a short word, like "four," into a similar-sounding one, like "fourteen," especially if the two units aren't equally common in its training data.

On a Tuesday, three separate chillers had already needed attention by 10 a.m., and the compressor room was running at full volume. Rui asked, "What's the suction pressure on unit four?" reaching for the service valve as he spoke, expecting the usual quick, plain number.

Hand sketched comparison titled The two costs are not the same size. Left panel, a box icon labeled Extra confirm, caption costs about 5 seconds mildly annoying. Right panel, a question box icon labeled Misheard reading near a valve, caption costs a near miss a safety stand-down.
Rui's shift ran on the left panel, dozens of times a day. The right panel only had to happen once.

Wrenchmate answered, flat and certain: "Reading normal, 42 psi." It had heard "unit fourteen," a chiller three rooms over that had already been serviced that morning, not the live unit Rui's hand was on.

Rui didn't open the valve because of the number Wrenchmate gave him. He opened it because of the habit underneath the number: nine years of trusting his own hand on the gauge first, out of nothing more than old caution, before he trusted anything spoken to him.

His hand found the physical gauge out of ninety habit, not because Wrenchmate had asked him to check. The physical reading didn't match the number he'd just heard. He stopped, re-asked the question slower, and this time Wrenchmate correctly heard "unit four" and gave the real, still-pressurized reading. Nothing happened to him. But the near miss reached his supervisor's desk within the hour.

Hand sketched flow diagram titled The voice loop with the read-back restored. Five boxes: Rui asks, Wrenchmate hears it, checks its confidence highlighted, reads back if unsure, gives the answer.
The middle steps used to be skipped for every query. Now they only fire when it actually matters.

The redesign didn't bring back the old always-confirm habit. It made the confirm step conditional: any query naming a unit right before a safety action, like opening a valve, now triggers an automatic read-back, whether or not Wrenchmate feels confident. Everything else stays as fast as it was.

Hand sketched decision tree titled When does Wrenchmate confirm before answering. Root: is this tied to a live safety action. Branches: opening a valve or line leads to always confirm first, routine check leads to answer directly, confidence is low anyway leads to hedge offer to repeat.
One question decides whether the extra half-second happens at all.

PICK, spelled outNot a lecture on speech recognition. PICK is what decides how often the caution is worth the interruption.

P
Position. The pick, stated first.
Confirm before answering only when a reading is tied to a live safety action, or genuine low confidence. Everywhere else, answer straight.
This is the hardest step and the direct answer: committing to when caution earns its place, instead of always or never.
I
Impact. Who feels each kind of error.
Always confirming costs Rui a few seconds, dozens of times a shift. Never confirming risks a misheard reading near a live valve.
Names both costs in the units that actually matter to him: seconds versus safety.
C
Cost asymmetry. Which error is the real one.
4 minutes a shift, lost to always confirming, against roughly 180 minutes and a mandatory stand-down for one missed high-stakes mishearing.
Makes the tradeoff a real decision instead of a coin flip between two equal options.
K
Kill criteria. What would flip the pick.
If technicians start saying yes to the confirm step without pausing to actually listen, the passive confirm has become noise, and it needs to become an active repeat-back instead.
Shows the position is testable, not just confident.
Hand sketched icon list titled Ways to show doubt with no screen at all. Four items: hedge with a plain word, ask a yes-or-no check, offer to repeat the reading, pause half a beat first.
None of these need a screen. All four fit inside a sentence someone can hear while their hands stay busy.
Misheard-command rate vs. background noise level, across job sites
15% 7.5 0 5% concern line 50 dB 70 dB 75-80 dB
Rui's compressor room runs at 75-80 decibels most mornings, right where the mishear rate starts climbing past what a quiet office would ever see.

The recap, one line per letter: position is stakes-gated confirmation, not always-on or always-off, impact is seconds for Rui versus a real safety risk, cost asymmetry is 4 minutes against 180, and kill criteria is the moment a passive confirm stops meaning anything.

And if you want to be sure it really works, try it somewhere elseSame four letters, a grocery fulfillment warehouse instead of a refrigeration plant. This time the voice assistant is picking orders, not reading pressure gauges.

Northfield Grocery Fulfillment uses a voice-pick headset system so a picker named Grace Okonkwo can hear which item to grab next and confirm the count out loud, hands free to actually carry the bins. Mapped onto PICK: position is confirming out loud only when an item is one of a small set that's easy to misgrab, like two nearly identical spice jars sold in different sizes, and answering straight everywhere else. Impact is Grace losing a couple of seconds on a routine pick versus a wrong item reaching a customer's order and triggering a refund and a re-pick. Cost asymmetry: a routine confirm costs about 2 seconds; a wrong-item shipment costs roughly 25 minutes once you count the refund, the re-pick, and the apology email. Kill criteria: if pickers start confirming by habit without actually checking the label, the system needs to require reading the item's exact printed size back, not just saying yes.

Hand sketched labeled parts diagram titled What's inside a good spoken uncertain answer, reused here for Northfield's voice-pick system. Center icon document labeled Spoken Response, with four callouts: the hedge word, the number if any, offer to repeat, what wait does.
Same four parts as Wrenchmate's spoken answer. A pressure reading and a spice jar size turn out to need the same shape of honesty.

The old decision at Northfield wasn't about read-backs at all, it was a merged-steps problem of its own kind: the system used to announce the item name and the expected count in one breath, with no pause for the picker to actually look at the shelf before confirming. That was fine when the catalog had few lookalike items. It stopped being fine once two nearly identical spice jars, in different sizes, sat on the same shelf.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "confirm only where a mistake would actually cost something, answer plainly everywhere else," and stop.
Cost: there's no budget to add confidence scoring to the voice model right now. Say so, and start with the cheap fix, a hard-coded list of high-stakes actions that always trigger a confirm, no machine confidence needed at all.
The model gets better, for real: if Wrenchmate's speech recognition improves and the mishear rate drops well under the 5% concern line even in loud rooms, the confirm rule can loosen, the same threshold running in reverse.

Where people run it wrong.
They add a confirm step to every single voice answer, and the technician stops actually listening to it.
They remove all confirmation to speed things up, without separating routine queries from safety-tied ones.
They use a spoken percentage as their hedge, which takes longer to process than the task in front of someone's hands.

How to use it live. When someone asks how you'd show uncertainty with no screen, ask back: which of these answers, if wrong, would actually cost someone something? Let that list, not a blanket rule, decide where the voice slows down.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an A-or-B tradeoff question like this one, and what does each letter do?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It forces a committed answer first, then proves it with the unequal cost of each error.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Rui Tanaka, a field technician at Arctic Fields Refrigeration, who has serviced industrial chillers for nine years, mostly by feel and by ear.
3 · THE POSITION
What's the committed pick this answer lands on?
Tap to flip
ANSWER
Confirm with a read-back only when a reading is tied to a live safety action or the model's confidence is genuinely low. Answer plainly everywhere else.
4 · THE ASYMMETRY
Which error does this answer optimize against, and why?
Tap to flip
ANSWER
The rare, expensive one: a misheard reading near a live safety action. It costs roughly 180 minutes and a safety stand-down, against 4 minutes for always confirming.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Cutting the automatic read-back for every kind of query at once, instead of only for the routine ones, when usage grew past the rare, careful early days.
6 · THE NUMBER
Fill in the blank: Rui's compressor room runs at about ___ decibels, right where the misheard-command rate starts climbing past 5%.
Tap to flip
ANSWER
75 to 80 decibels. About the volume of a passing truck, loud enough to bend a short word into a similar-sounding one.
7 · THE REPLAY
Same noisy morning, new design with stakes-gated confirmation. What changes for Rui?
Tap to flip
ANSWER
Because the query names a unit right before a valve action, Wrenchmate automatically reads back "unit fourteen, is that right?" catching the mishearing before any pressure number is stated.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the equivalent high-stakes trigger there?
Tap to flip
ANSWER
Northfield Grocery Fulfillment's voice-pick headset system. The high-stakes trigger is picking one of a small set of easily confused, nearly identical items, like two spice jars in different sizes.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: always confirming every reading costs about ___ minutes per shift, while one missed high-stakes mishearing costs about 180 minutes.
Show hint
Look at the grouped bar chart in Section 1.
Show answer
4 minutes. A 45-times difference in cost, which is what makes stakes-gating the confirm step the real decision, not a coin flip.
Multiple choice
2. Why does this answer reject having Wrenchmate speak a numeric confidence percentage out loud?
  • A. Percentages are considered rude in professional speech.
  • B. A spoken number needs a moment of mental math someone with their hands inside a machine doesn't have to spare.
  • C. The headset hardware can't play back numbers accurately.
  • D. Regulations forbid stating confidence levels in industrial settings.
Show hint
Look at the rejected alternative in the priority list.
Show answer
B. A plain hedge phrase can be processed instantly by ear; a percentage requires a beat of thought nobody has to give mid-task.
True or false
3. True or false: at Northfield Grocery Fulfillment, the fix that mattered most was adding a spoken confidence score to every pick.
  • True
  • False
Show hint
Look at Section 4's old decision.
Show answer
False. The fix was pausing between naming the item and confirming the count, so the picker had a real moment to look at the shelf first.
Short answer, apply it yourself
4. Think of a voice assistant you use, in a car or at home. Does it ever double-check what it heard before acting? When would you want it to, and when would that just be annoying?
Show hint
Think about which of its actions would actually cost you something if it misheard you.
Show answer
Model answer: Most voice assistants confirm nothing or confirm everything. Almost none gate it by how much a mistake would actually cost.
Short answer, why no middle setting
5. Why wasn't "just make Wrenchmate speak a little more carefully" a real fix for the near miss?
Show hint
Look at the disqualified fixes under "where people run it wrong."
Show answer
Model answer: Speaking "more carefully" doesn't change what was actually heard. Only an explicit read-back catches a mishearing before it becomes an answer.
Short answer, where it wouldn't matter
6. Name a kind of query on Rui's shift where this stakes-gated confirmation genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine temperature check with no live safety action attached. Confirming that would just slow Rui down for no real safety gain.
Before you close the answer
Why this works
Tests whether you'll design one uniform level of caution for a voice interface, or notice that the same caution applied everywhere trains people to stop noticing it at all.
Follow-up traps
"What if the technician is in a hurry and just wants to skip the confirm step?" Response: for high-stakes queries, don't allow a skip; the whole point of gating by stakes is that this is the one moment where speed isn't the priority.

"Couldn't you just lower the mishearing rate instead of adding a confirm step?" Response: you should do both; better recognition reduces how often the confirm step fires, but it never gets you to zero in a genuinely loud room, so the confirm step stays as the backstop.
If pressed
The high-stakes trigger list isn't just "any query near a valve," it's built from Arctic Fields' own incident log of near misses, so it grows every time a new kind of misheard command turns out to matter, instead of staying frozen at launch.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more