InterviewAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #22

Show me how you would redesign a confident-sounding AI answer to be honest instead.

FLIPS over-trust: the flip that fires when the tone gets more confident, not when the model gets worse

Solum Finance is a budgeting app. Amadou Kessler is a freelance graphic designer who has used it for eight months. Here is the affordability answer Solum gave him the day he asked about a used car, and the redesign that would have caught what it left out.

The direct answer
Redesign the answer to state a graduated likelihood, name the specific things driving the uncertainty, and show what happens in the person's worst realistic month, not just their average one. Never let confident-sounding copy claim more certainty than the number underneath it actually has.
Do this, in order
  1. State a graduated likelihood, never a flat yes.Why: "likely affordable" leaves room for a caveat. "Yes" doesn't, and it never should for a decision this size.
  2. Name the specific factors driving uncertainty, not a generic disclaimer.Why: "your income varies" is forgettable. "Your income varied 30 percent over 3 months" is something a person can actually act on.
  3. Show the worst realistic month, not just the average one.Why: averages hide exactly the month that breaks a budget, which is the only month that ever actually matters.
  4. Never let the copywriting layer round the model's real confidence up to certainty.Why: a confident tone is a writing choice, and a writing choice can quietly out-run what the number underneath actually supports.
  5. Offer a next step, not just a caveat with nowhere to go.Why: "see the math" or "talk to an advisor" turns a warning into something a person can actually do something with.
  6. Leave small, low-stakes purchases as a quick, confident yes.Why: being wrong about a 12 dollar subscription costs almost nothing, so it doesn't need this treatment.

How to answer this, stage by stage

Nobody is grading whether you can spot a bad sentence. They're grading whether you can rewrite it into something specific enough to actually act on.

Stage 1
Scope it to one product, one confident answer
Say it like this
"I'll use Solum Finance's affordability answer, the one it gave a user named Amadou the day he asked whether he could afford a used car."
Why this works
Gives you a real sentence to rewrite instead of a vague idea of "confident AI answers."
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Shows you have a method for finding what's actually wrong, not just a gut feeling that the tone is off.
Stage 3
Name the person and the habit the confident tone built
Say it like this
"Amadou used to check every big purchase against his own spreadsheet. Once Solum's answers started sounding certain, he stopped checking at all."
Why this works
Puts a real cost behind "the tone was too confident," instead of leaving it as a style complaint.
Stage 4
Identify the flip: what changed was the tone, not the number
Say it like this
"This is an over-trust flip. The model's actual confidence number never changed. The copywriting did, and Amadou couldn't tell the difference from the outside."
Why this works
Names exactly why this is an AI-specific problem, not a generic UX complaint about wording.
Stage 5
Name the old decision behind it
Say it like this
"Solum told users at onboarding to trust the answer because it's built on their real data. That was fair when the copy still hedged. Nobody revisited it once the copy stopped hedging."
Why this works
Finds the actual decision that broke, instead of just blaming "bad copywriting" in the abstract.
Stage 6
Show the redesign, side by side
Say it like this
"Before: 'Yes, this fits your budget! You'll have 340 dollars left over each month.' After: 'Likely affordable, with two things to watch: your income varied over 30 percent in the last 3 months, and your insurance renews in 6 weeks at about 180 dollars more. Your average month leaves 340 dollars. Your lowest month this year would have left you about 95 dollars short. Want to see the math, or talk to an advisor?'"
Why this works
This is the actual deliverable the question is asking for, not a description of what a redesign might look like.
Stage 7
Close on the replay
Say it like this
"Run the same afternoon forward with the new answer, and Amadou waits two pay cycles before signing. He never sees an overdraft at all."
Why this works
Ends on something countable, not just a claim that the new version "feels more honest."

Let's learn

Solum is a budgeting app with a feature that answers one question: can I actually afford this purchase, right now?

Before Solum, freelancers like Amadou tracked irregular income by hand in a spreadsheet, checking their own numbers before any big purchase, a slow, careful habit that took real time but caught real problems.

Hand sketched icon list titled The five letters. Five items with icons: F find the person, L locate the habit, I identify the flip, P pinpoint the old decision, S show the replay.
The whole method, in five rows. The hard one is the third.

Now Solum answers instantly, pulling in real transaction history, no spreadsheet required.

Here's the turn: the extra convenience was never the issue. The real problem showed up once Solum's answers started sounding more certain than the underlying number actually was, and users stopped doing their own math, because a confident tone reads as a confident fact, whether or not it is one.

Share of affordability answers users double-checked against their own numbers, by month
80% 40 0 20% concern line Month 1 Month 4 Month 5, car purchase
The line crossed the concern threshold in month four, the same month the copy rewrite fully shipped. Nobody was watching a checking rate, only whether the answers "sounded right."

At its worst, someone makes a large purchase believing "Yes, this fits" is a settled fact, then finds out three weeks later, at the worst possible moment, that it never was.

Hand sketched comparison titled Small move, big snap. Left panel, a blue gauge icon labeled Gentle rise, caption checking less bit by bit. Right panel, a red box icon labeled Flat then jump, caption no checking all at once.
The checking rate didn't drift down slowly forever. It fell off a shelf the month the copy changed.
The decision I would take back Solum's onboarding told users its affordability answers were "powered by your real transaction history, so you can trust the answer," language written when the actual answer copy still hedged. That made sense at the time. It stopped making sense once a later copywriting pass, aiming to sound more helpful, rewrote every answer to sound flatly certain, regardless of what the model's own internal confidence number actually said.

What I would leave alone: small, low-stakes purchases don't need this treatment. A quick, confident yes on a 12 dollar subscription is fine, since being wrong about it costs almost nothing.

The lesson: a confident tone is a writing choice, not a fact about the world. The moment the writing sounds surer than the model actually is, it's lying by tone, even if every number behind it is technically accurate.

Now here is the same thing as a story

The short version above is what you'd say defending this redesign to Solum's product council. Read this one for how the trust actually built up.

Every other Sunday, Amadou Kessler used to sit down with a spreadsheet and his last two months of invoices, checking what he could actually afford before he let himself buy anything big. He's a freelance graphic designer, and for years before Solum, he could tell within about fifty dollars what a slow month would do to his plans.

Solum's early affordability answers were careful. "Likely affordable, but your last invoice hasn't cleared yet," the kind of hedge Amadou's own spreadsheet would have caught too. He kept checking his own numbers for the first few months anyway, out of habit, and Solum kept agreeing with him.

Knowledge spark: why would a model's confidence number and its wording ever disagree? A model can produce a real, calibrated number, say, seventy percent sure. What a user actually reads is whatever sentence a separate copywriting or prompt layer wrote to describe that number. If nobody constrains that sentence to reflect the number underneath, the words can drift toward sounding more certain than the number ever claimed, purely because confident writing tests better with users.

After a while he checked before every big purchase, then only the ones over a few hundred dollars, then, once Solum's answers started sounding flatly certain, a copywriting update rewrote its hedges into plain yeses, he stopped checking at all, since the tone alone read as more sure than any spreadsheet he could build.

Hand sketched metaphor scene titled Switch, not dial. Left, a blue gauge icon labeled Dial, caption checking by degrees. Right, a red box icon labeled Switch, caption check all or none.
Solum's team assumed Amadou had a dial he could turn down slowly. He only ever had a switch.

A new engineer on Solum's product team, running a routine audit of the affordability copy, asked in a meeting, "wait, why does this always say 'Yes, this fits' no matter what the underlying number actually is? Did the model get more confident, or did we just make the copy sound more confident?" Nobody on the team had a clean answer.

Hand sketched timeline titled Before, the thinning, the trigger, the replay. Five milestones: Checked every big buy month 1, Checked only the big ones month 3, Copy turns confident month 5, A new hire asks why month 8 highlighted, Redesigned answer ships month 9.
The audit that followed the new hire's question is what turned up Amadou's case, three months after it happened.

Three months into the confident-tone version, that audit found it: Amadou had asked Solum whether he could afford a 28,000 dollar used car. Solum said "Yes, this fits your budget! You'll have 340 dollars left over each month after this purchase," with no hedge at all. Amadou signed the loan that afternoon.

Solum did not lie about the average month. It just never mentioned the month that actually mattered.

Three weeks later, an insurance renewal came in 180 dollars higher than expected, the same month one of his two biggest clients paid nine days late. His account went 95 dollars into overdraft, a gap Solum's own transaction history could have shown months in advance, if the answer had ever mentioned it.

What Solum's answer implied, versus what actually happened
+400 0 -150 +340 What the answer implied -95 His actual lowest month
A 435 dollar gap between the confident sentence and the real number, sitting in data Solum already had.

With the redesigned answer, restored hedges and named specifics, Solum tells Amadou: "Likely affordable, with two things to watch: your income varied by more than 30 percent over the last three months, and your car insurance renews in six weeks at an estimated 180 dollar increase. In your average month you'd have about 340 dollars left over, but in your lowest month this year, you'd have been about 95 dollars short. Want to see the math, or talk to a real advisor?" Run the same afternoon forward: Amadou waits two pay cycles before signing, and never sees an overdraft at all.

The old copy asked Amadou to trust a tone. The new one shows him the actual month that would have broken his budget.

I signed off on the confident rewrite because it tested well, people liked how sure it sounded. It took a new hire's plain question, and a stranger's overdraft, to see that sure-sounding and sure were never the same thing.

The five steps, if you want to remember it

F
Find the person. Whose morning is this?
Amadou Kessler, a freelance designer who used to check his own numbers by hand before every big purchase.
Grounds the whole flip in one specific person's real habit, not "users" in general.
L
Locate the habit. What did he stop doing?
Cross-checking Solum's affordability answers against his own spreadsheet, since Solum agreed with him for months.
Names the exact thing the product's convenience quietly replaced.
I
Identify the flip. Over-trust, not under-trust.
Checking every answer himself, versus checking none at all, flipped once the copy sounded more confident, not once the model's real accuracy changed.
This is the hardest step: the trigger here is an improvement in tone, which is exactly what makes over-trust easy to miss.
P
Pinpoint the old decision.
Telling users to trust the answer because it's built on real data, a promise made while the copy still hedged, never revisited once the copy stopped hedging.
Finds the specific, reasonable-at-the-time choice that quietly stopped being true.
S
Show the replay.
With the redesigned answer, Amadou sees the specific risk and the worst realistic month, and waits two pay cycles instead of signing that afternoon.
Ends on something countable: no overdraft, not just "a better experience."

The recap, one line per letter: find the person is Amadou and his old spreadsheet habit, locate the habit is the cross-checking he quietly stopped doing, identify the flip is over-trust triggered by a more confident tone, pinpoint the old decision is an onboarding promise nobody revisited, and show the replay is the redesigned answer catching the risk three weeks before it would have mattered.

And if you want to be sure it really works, try it somewhere elseA different flip family, a home-insurance claims assistant instead of a budgeting app. Confidence gets gamed from the other direction this time.

Covent Assurance runs an AI claims assistant that gives renters a fast read on whether a claim is likely to be approved. Noor Falkenberg filed a stolen-bicycle claim through it. This story runs on a different family: the pre-editing flip, not over-trust.

Over months, claimants like Noor learned which phrasing got a fast, confident "likely approved" read from the assistant, and started pre-editing their own claim descriptions to match it, stripping out the messy, true details, whether the lock was cut or simply left unlocked, a small gap in the timeline, that made a case genuinely harder to assess, in exchange for a smoother, more confident-sounding answer. The old decision behind it: Covent's early claims copy explicitly told users that certain word choices "help us process your claim faster," a well-meant onboarding tip that quietly taught claimants to write toward the model's comfort zone instead of writing what actually happened.

Hand sketched flow diagram titled How a claim gets edited toward the model. Five boxes: Claimant describes it plainly, Assistant hesitates asks more, Claimant learns the phrasing highlighted, Description gets edited, Approval feels instant.
The habit forms in the third box. By the fifth box, the model is confidently approving a story that isn't quite the real one.
Hand sketched decision tree titled How confident should the copy sound. Root: what does the number actually support. Three branches: confidence high stable leads to say it plainly no hedge, confidence moderate leads to name the specific risk, confidence low or volatile leads to show the worst realistic case.
The same tree that fixed Solum's affordability answer also tells Covent's assistant when a claim needs a real human look instead of an instant, confident yes.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the specific risk factors, show the worst realistic case, and never let the copy sound surer than the model's real number," and stop.
Cost: there's no time to rewrite every answer template right now. Say so honestly, and start with the highest-stakes decisions, a car, a claim, a loan, since that's where a confident-sounding mistake costs the most.
The model gets better, for real: if Solum's underlying affordability model genuinely gets more accurate, that's still not a reason to let the copy sound more certain than its actual confidence number, the redesign constrains the words to the number, whatever the number turns out to be.

Where people run it wrong.
They test confident-sounding copy for how much users like it, without ever checking whether it matches the model's real number.
They add one disclaimer at the bottom of the screen and call the honesty problem solved, when users have already learned to skip past it.
They assume a wrong answer will generate a complaint, when most people who get burned quietly blame themselves instead.

How to use it live. When someone hands you a confident-sounding AI answer and asks you to make it honest, ask yourself one question before rewriting a single word: what does the number underneath this sentence actually support? Rewrite until the words say exactly that, no more.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust: checks sometimes, then stops checking at all. It fires here because the trigger is an improvement, the copy sounding more confident, not the model getting worse.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Amadou Kessler, a freelance graphic designer who used to check his own irregular income by hand before every big purchase.
3 · THE HABIT
What did he stop doing because it worked?
Tap to flip
ANSWER
Cross-checking Solum's affordability answers against his own spreadsheet, since Solum had agreed with him every time for months.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Checking every affordability answer himself, versus checking none at all. It flipped once the copy started sounding certain, not once the model's real accuracy changed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling users at onboarding to trust the answer because it's powered by their real data, a promise made when the copy still hedged, never revisited once the copy stopped hedging.
6 · THE NUMBER
Fill in the blank: Solum's confident answer implied Amadou would have 340 dollars left over. In his actual lowest month, he was really about ___ dollars short.
Tap to flip
ANSWER
About 95 dollars short. The gap between the implied number and the real one is the whole story.
7 · THE REPLAY
Same purchase, redesigned answer. What changes?
Tap to flip
ANSWER
Amadou sees the specific risk and the worst realistic month, waits two pay cycles, and never sees an overdraft at all.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Covent Assurance's claims assistant, and the pre-editing flip: claimants learn to rewrite their own claim descriptions toward the words the model responds well to.

Check yourself Score: 0 / 0

Multiple choice
1. Why does this count as an over-trust flip rather than the model simply getting less accurate?
  • A. The model's actual internal confidence number changed for the worse.
  • B. The copywriting made every answer sound certain, regardless of what the model's real confidence number said.
  • C. Amadou stopped using the app for several months.
  • D. Solum removed the affordability feature entirely.
Show hint
Look at the Identify the flip step.
Show answer
B. The number underneath never changed. Only the sentence describing it did, which is exactly why over-trust is easy to miss.
True or false
2. True or false: Solum's underlying model became less accurate right before Amadou's car purchase.
  • True
  • False
Show hint
Look at the knowledge spark about confidence numbers versus wording.
Show answer
False. The model's confidence number didn't change. A copywriting pass changed how confident the wording sounded, on top of the same underlying number.
Fill in the blank
3. Fill in the blank: Solum's confident answer implied Amadou would have 340 dollars left over each month, but in his real lowest month he was about ___ dollars short.
Show hint
Look at the implied-versus-actual bar chart.
Show answer
95 dollars. A 435 dollar gap between the confident sentence and the real number, sitting in data Solum already had.
Short answer, apply it yourself
4. Think of a confident-sounding answer you got from an AI tool recently. Was its confidence coming from the actual data behind it, or just from how it was worded?
Show hint
Ask whether the tool ever showed you a real number or reasoning, or just a certain-sounding sentence.
Show answer
Model answer: Most people can name at least one case where the tone was doing more work than the underlying data, exactly the gap this answer redesigns around.
Short answer, why no middle setting
5. Why couldn't Amadou have just "checked a little more carefully" instead of stopping entirely?
Show hint
Look at the small move, big snap diagram.
Show answer
Model answer: A flatly confident tone reads as a settled fact, and there's no natural halfway point between trusting a stated fact and doubting it. Once the tone crossed into certainty, checking stopped feeling necessary at all.
Short answer, where it wouldn't matter
6. Name a place in Solum where this same confident-tone scrutiny genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Small, low-stakes purchases, like a 12 dollar subscription. Being wrong about one of those costs almost nothing, so a quick, confident yes is fine there.
Before you close the answer
Why this works
Tests whether you can actually rewrite dishonest-sounding confidence into something specific and honest, not just critique the old copy in the abstract.
Follow-up traps
"Won't a hedged answer just make users trust the app less overall?" Response: no, the redesigned answer is more specific, not more vague. Naming the exact risk and the worst month builds more real trust than a flat yes that occasionally turns out wrong.

"Couldn't you just add a disclaimer at the bottom instead of rewriting the main answer?" Response: a disclaimer users have already learned to skip past doesn't do the job. The caveat has to live inside the sentence carrying the actual verdict.
If pressed
Solum's redesign specifically constrains the copy-generation layer to never round the model's own internal confidence score up by more than one graduated band, so a later wording pass can't out-run what the number underneath actually supports.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more