ConceptIntermediateDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #3

What is the risk of showing a percentage confidence that users cannot interpret?

GUARD a number nobody explains is a decision about who never gets to ask

Larkwood Advisory is a wealth management firm with two tiers: a premium tier with a dedicated human advisor, and a cheaper self-serve tier with none. Isabel Marchetti is a Client Wealth Associate on the premium side. ClearPath is the AI tool that attaches a confidence percentage to every portfolio rebalancing suggestion, for clients in both tiers.

The direct answer
The real risk isn't confusion, it's that a percentage nobody can interpret gets treated as a decision anyway, and the harm lands hardest on whoever has no one to ask what it means. A client with a dedicated advisor calls and gets it explained. A client with no advisor reads a dropping number alone and acts on a guess.
Do this, in order
  1. Never show the raw percentage without a plain-language line bundled in the same view.Why: a number with no explanation attached gets read as a verdict, not a range, by whoever has no other way to check it.
  2. Give every tier, not just the premium one, a real way to ask what a number means.Why: the self-serve tier is exactly the group that never gets a lever, and that's a design decision, not an accident of pricing.
  3. Require a hard pause on any large single-session drop in the score before a client can act on it.Why: a drop from 78 to 61 in one sitting is exactly the moment someone acts on a guess instead of a fact.
  4. Track same-day liquidations after a score drop, by tier, every month.Why: this is the early number that would have shown the harm building before one bad morning made the news.
  5. Never let a "we kept the premium tier's model simple to explain" excuse cover for the self-serve tier having nothing at all.Why: pricing a person out of an explanation is still a design choice, even when nobody meant it as one.

How to answer this, stage by stage

Handle this one plainly. Nobody's grading whether you can define "confidence interval." They're grading whether you noticed that some people get to ask a question and some people don't.

Stage 1
Scope it to one real product with two tiers
Say it like this
"I'll answer this for Larkwood Advisory's ClearPath, which shows the exact same confidence percentage to a premium client with an advisor and a self-serve client with none."
Why this works
Keeps the question from becoming an abstract debate about numeracy in general.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's affected. Unequal, where the harm actually lands. Ability to contest, who never gets a lever. Reduce, the design change. Detect, how I'd know it's happening."
Why this works
Shows the interviewer you're going straight to who gets hurt, not a general list of confusion risks.
Stage 3
Name both groups
Say it like this
"Isabel's premium clients see the same number, but they can call her and ask what it means. The self-serve tier sees the exact same number with no one to call at all."
Why this works
This is the move that turns "confusing UI" into a real fairness question.
Stage 4
Give the direct risk, before any reasoning
Say it like this
"The risk isn't that people get confused. It's that a number nobody can interpret still gets acted on, and it's the tier with no one to ask that pays for it."
Why this works
This is the direct answer, said plainly, before the story does any convincing.
Stage 5
Prove it with the one catastrophe
Say it like this
"A retired teacher on the self-serve tier watched her score drop from 78 to 61 overnight. She read that as a 39 percent chance of losing money and sold her entire portfolio at 7:19 that morning. The market recovered by Friday. Nobody at Larkwood was there for her to ask first."
Why this works
Turns "some people can't interpret this" into a specific, checkable loss instead of a vague worry.
Stage 6
Say the design change and the detection signal
Say it like this
"Any drop that big in one sitting triggers a mandatory plain-language explainer before she can act. And every month, I'd track same-day sell-offs after a score drop, split by tier, because that gap is the warning sign, not the news story after."
Why this works
Turns a moral concern into an actual product decision and a real number to watch.
Stage 7
Close on the one line
Say it like this
"A number nobody can read isn't dangerous by itself. It's dangerous the moment it reaches someone with nobody to ask, and no way to check it."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Here is what a confidence percentage actually risks once it reaches someone who has no way to check it.

Before ClearPath, Larkwood's advisors rebalanced client portfolios by hand, on a quarterly review call, walking each client through the reasoning in plain terms. That took about forty-five minutes per client. ClearPath now flags a suggested rebalance and a confidence score for every account overnight, cutting advisor prep time to about ten minutes per call.

Hand sketched metaphor scene titled One of them can call someone, one cannot. Left, a blue person icon labeled Premium tier, caption advisor explains the number. Right, a red box icon labeled Self-serve tier, caption just the bare percentage.
Same product, same number. Only one of the two groups has a person to call about it.

Here's the turn: the confidence score itself was never wrong. It correctly reflected more market uncertainty that quarter. The problem was that the self-serve tier read that honest uncertainty as a verdict, with nobody there to say otherwise.

Clients who sold a major position within one week of a large confidence-score drop
15% 7.5 0 14% Self-serve tier 1% Premium tier
Same number, same size drop. Fourteen times as many self-serve clients sold, because the only difference between the two groups was whether anyone was there to explain it.
Hand sketched flow diagram titled Where the appeal should sit and does not. Four boxes: Score drops on the app, Client feels alarmed, No one to ask why highlighted, Client sells that day.
The third box is empty on purpose in this diagram. That empty box is exactly where the harm happens.

At its worst, an unexplained confidence score doesn't just confuse someone for a moment. It replaces a real financial decision with a guess made under panic, and the person who made that guess has no way to know it was ever a guess.

The decision I would take back Larkwood priced live advisor access as a premium add-on, to keep the self-serve tier affordable for smaller accounts. That made sense as a way to serve more clients at a lower price. It stopped making sense the moment ClearPath's confidence score became the only thing standing between a self-serve client and a decision she had no way to check.

What I would leave alone: the premium tier's quarterly advisor call itself doesn't need a redesign. Isabel's clients already have a real person to ask, which is exactly the lever the self-serve tier is missing.

The lesson: a number is only as safe as the ability of the person reading it to ask what it means, and pricing that ability out of a tier is still a design decision, whether anyone meant it as one or not.

Now here is the same thing as a story

The short version above is what you'd say defending this fix to Larkwood's compliance team. Read this one for how the morning actually unfolded.

The client was a retired schoolteacher, sixty-eight, who had moved her retirement savings into Larkwood's self-serve tier two years earlier because the fees were lower and her account was modest. She checked the ClearPath app most mornings with her coffee, the way she'd check the weather.

That Tuesday, the app showed her portfolio's rebalance confidence at 61 percent, down from 78 percent the week before. Nothing else on the screen explained why, or what "confidence" was actually measuring.

Knowledge spark: what does a "rebalance confidence" score actually measure? It's usually the model's own estimate of how likely a suggested change is to improve returns over a set time window, based on current market volatility. A drop often means the market got harder to read that week, not that anything about the client's own account got worse.

She read 61 percent the way she'd read a weather forecast: a little over a coin flip's chance of something bad. She had no advisor to call, only a support chat with an average wait of four hours. By 7:19 that morning, before the chat queue had even opened for the day, she'd sold her entire portfolio through the app herself.

Hand sketched timeline titled The morning a score drop became a real loss. Four milestones: Score shows 61 percent 7:04 am, Client reads it as risk 7:06 am, She sells everything highlighted 7:19 am, Market recovers by Friday.
Fifteen minutes between seeing the number and locking in the loss. No step in between where anyone could have stopped her.
She didn't misread the market. She misread a number that was never built to be read alone, by someone with nobody to ask.

The market recovered by that Friday. Her account did not, because the loss was already locked in. Isabel's own premium clients had seen the identical confidence drop that same week. Every one of them called first. Not one of them sold.

Hand sketched labeled parts diagram titled What's on the confidence disclosure screen. A document icon at the center labeled ClearPath Screen, with four callouts around it: the raw percentage, plain language range, talk to someone button, what changed since last time.
Four things that could be on this screen. Before this, the self-serve tier only ever had the first one.

GUARD, in one screenNot a lecture on financial literacy. GUARD is what tells you whose number this actually is.

G
Groups. Who is affected.
Premium-tier clients with a dedicated advisor, and self-serve-tier clients with none, both looking at the exact same confidence percentage.
Naming both groups is what turns a UX complaint into a real fairness question.
U
Unequal. Where the harm actually lands.
Self-serve clients sold within a week of a big score drop fourteen times as often as premium clients seeing the identical drop.
Shows the harm isn't evenly spread, it lands specifically on the tier priced out of an explanation.
A
Ability to contest. Who never gets a lever.
The self-serve client has no advisor, and a support queue that opens after most people have already acted on the number alone.
This is the hardest step and the real risk the question is asking about.
R
Reduce. The design change.
A mandatory plain-language explainer and a forced pause on any large single-session score drop, before a self-serve client can act on it.
A real product decision, not a training module or a disclaimer nobody reads.
D
Detect. How you'd know it's happening.
Same-day liquidations after a confidence-score drop, tracked monthly, split by tier, well before a client complaint or a regulator ever asks.
Catches the pattern from the data, instead of waiting for the next bad morning.
Hand sketched icon list titled Signs a percentage is doing harm. Three items: a gauge icon labeled sell offs spike right after a score drop, a question box icon labeled support chats ask what the number means, a person icon labeled only one tier ever asks at all.
Any one of these three shows up in the data long before a client ever files a complaint.

The recap, one line per letter: groups is the premium and self-serve tiers seeing the same number, unequal is the fourteen-times gap in same-week sell-offs, ability to contest is the self-serve tier having no advisor and a queue that opens too late, reduce is a mandatory pause and plain-language line on big drops, and detect is watching same-day liquidations by tier every month.

And if you want to be sure it really works, try it somewhere elseSame five letters, a staffing agency's hiring-risk score instead of a wealth platform. A different subject has no lever at all this time.

Vantage Staffing Group uses HireSignal, an AI tool that scores job candidates on predicted attrition risk for hiring managers reviewing applications. Callum Bryce is the HR screening manager who reads those scores each week. Mapped onto GUARD: groups here are Callum, who sees the score and decides who advances, and the applicant, who never sees it at all; unequal is that the harm lands specifically on applicants whose resumes show employment gaps, since the model reads a gap as risk regardless of its cause.

The ability-to-contest gap is sharper here than at Larkwood. A self-serve wealth client at least sees the number that hurt her. A rejected job applicant never even learns a risk score existed, let alone what pushed it up. One candidate, a single parent who'd taken eighteen months off for caregiving, was scored high-risk and filtered out before a human ever read her resume, with no path to explain the gap at all.

Hand sketched quadrant titled Who gets hurt worst by a bare number. X axis access to someone who explains it, from none to a dedicated advisor. Y axis cost of a wrong read, from low to high. Retired teacher self serve and rejected applicant sit low on access, high on cost. Premium client and hiring manager sit high on access.
The retired teacher and the rejected applicant land in the same corner, on two completely different products.
Share of high-attrition-risk flags later found to have a caregiving or medical gap behind them
40% 20 0 Month 1 Month 6
Nobody was tracking this rate until an audit asked for it. By month six, over a third of high-risk flags traced back to a gap that had nothing to do with reliability.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the risk lands on whoever has nobody to ask what the number means, name that group and give them a lever," and stop.
Cost: there's no budget to add live advisor access to every tier this year. Say so honestly, and start with the cheap fix, a mandatory plain-language explainer and pause on big drops, which costs far less than a support line.
The model gets better, for real: if ClearPath's forecasts genuinely get more accurate, the confidence scores will still swing with real market volatility sometimes, and the tier with nobody to ask still needs the same pause, better accuracy doesn't remove the need for someone to explain a drop.

Where people run it wrong.
They treat "add a tooltip" as solving the problem, when the real gap is having nobody real to ask.
They assume the harm is spread evenly across users, instead of checking which tier or group actually carries it.
They wait for a complaint or a lawsuit to find the pattern, instead of tracking the leading number every month.

How to use it live. When someone asks about the risk of a confusing number, ask yourself: who reading this has someone to call, and who doesn't? Design for the second group first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a risk or fairness question like "what's the risk of an uninterpretable percentage"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It names who can push back and who can't, on the exact same number.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Isabel Marchetti, Client Wealth Associate at Larkwood Advisory, on the premium tier where clients have a real advisor to call.
3 · THE GROUPS
Who are the two groups this answer names?
Tap to flip
ANSWER
Premium-tier clients with a dedicated advisor, and self-serve-tier clients with none, both seeing the exact same confidence percentage.
4 · ABILITY TO CONTEST
Who never gets a lever in this story, and why?
Tap to flip
ANSWER
The self-serve tier client. She has no advisor, and the support chat queue opens hours after most people would already have acted on the number alone.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pricing live advisor access as a premium add-on, which left the self-serve tier with only the bare percentage and no way to ask what it meant.
6 · THE NUMBER
Fill in the blank: self-serve clients sold a major position within a week of a big score drop ___ times as often as premium clients.
Tap to flip
ANSWER
Fourteen times as often (14 percent versus 1 percent), seeing the exact same size drop.
7 · THE FIX
What's the actual design change this answer recommends?
Tap to flip
ANSWER
A mandatory plain-language explainer and a forced pause on any large single-session confidence-score drop, before a self-serve client can act on it.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and who has even less of a lever there?
Tap to flip
ANSWER
Vantage Staffing Group's HireSignal attrition-risk score. A rejected job applicant never even learns the score existed, let alone what pushed it up, unlike the wealth client who at least sees the number that hurt her.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the retired teacher's confidence score dropped from 78 percent to ___ percent before she sold her entire portfolio.
Show hint
Look at the timeline diagram.
Show answer
61 percent. She read that drop as roughly a 39 percent chance of losing money, with nobody there to correct that reading.
Multiple choice
2. What was the actual risk this answer identifies, according to the direct answer?
  • A. That the confidence score was mathematically wrong.
  • B. That a number nobody can interpret still gets acted on, and the harm lands on whoever has nobody to ask.
  • C. That percentages are always a bad way to show data.
  • D. That the model needed more training data.
Show hint
Reread the direct answer at the top of the page.
Show answer
B. The score itself was honest. The harm came from who had, and didn't have, someone to ask about it.
True or false
3. True or false: premium-tier clients at Larkwood saw a different, more detailed confidence score than self-serve clients did.
  • True
  • False
Show hint
Look at the Groups step.
Show answer
False. Both tiers saw the exact same percentage. The only difference was whether a real person was available to explain it.
Short answer, apply it yourself
4. Think of a score or rating an app has shown you. If you'd gotten a surprising result, did you have a real person to ask about it, or only a help page?
Show hint
Compare a product with live support to one with only a FAQ.
Show answer
Model answer: Most consumer apps offer only a help page, which is exactly the "no lever" position this answer says to design around, not assume away.
Short answer, name who can't push back
5. In the Section 4 story, why is the rejected job applicant's situation worse than the self-serve wealth client's?
Show hint
Look at what each person actually gets to see.
Show answer
Model answer: The wealth client at least saw the number that hurt her. The rejected applicant never learns a risk score existed at all, so there's nothing to even question.
Short answer, where it wouldn't matter
6. Name a part of Larkwood's process where this same design concern genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The premium tier's quarterly advisor call. Isabel's clients already have a real person to ask, so that part of the product needs no redesign.
Before you close the answer
Why this works
Tests whether you'll treat "users can't interpret this number" as a UX inconvenience or notice that the harm concentrates on whoever has no way to ask a real person what it means.
Follow-up traps
"Isn't a plain-language explainer just as easy to misread as the number?" Response: a well-tested plain-language line, paired with a mandatory pause before acting, gives someone a chance to slow down, which a bare number racing past on a screen never does.

"Couldn't you just remove the score entirely for the self-serve tier?" Response: removing it hides real, honest uncertainty from someone who's entitled to know it exists; the fix is making sure they can act on it safely, not keeping them in the dark.
If pressed
Larkwood's actual fix added a forty-eight-hour cooling-off step for any self-serve action taken within one hour of a confidence-score drop larger than fifteen points, long enough for a same-day support callback to reach the client before the trade could finalize.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more