ConceptFoundationalDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #1

What does a user need to see to trust an AI recommendation?

SPARK the product is Renlio, a marketplace where people resell secondhand clothes and furniture, and its Suggested Price button

Renlio is a resale marketplace. Its Suggested Price button looks at recent sold listings and hands a seller one number to list at. Dara Whitfield resells secondhand coats and furniture on Renlio most evenings after her day job, at her kitchen table, phone in hand.

The direct answer
Show the evidence the recommendation is built on, not just the number it produces. For a price, that means the actual comparable sold listings and how much data backed the guess. A number alone forces a seller to trust it completely or ignore it completely. Evidence lets trust be partial, and partial trust is the only kind that survives being wrong sometimes.
Do this, in order
  1. Show the real comparable listings behind the number, not just the number.Why: a bare number gives a seller nothing to check it against.
  2. Show how much data backed the guess, not a fixed tone of confidence.Why: a guess from six listings and a guess from four hundred should never look equally sure.
  3. Let the seller change one input, like condition, and watch the price move.Why: a number you can poke at is a number you can actually judge.
  4. Say plainly when there isn't enough data, instead of guessing anyway with a straight face.Why: a confident wrong answer costs more trust than an honest shrug.
  5. Keep the whole thing behind one tap, not bolted onto the main screen.Why: most sellers don't want to dig, and the fast path should stay fast for them.
  6. Watch how often sellers tap "why this price," as an early warning sign.Why: that number moves weeks before anyone complains out loud.

How to answer this, stage by stage

Nobody's grading whether you know AI can be wrong. They're grading whether you can name the exact thing on the screen that lets a normal person tell a good guess from a shaky one.

Stage 1
Scope it to one real screen
Say it like this
"I'll answer this for Renlio's Suggested Price button, the number a seller sees the moment they photograph an item to list."
Why this works
Turns a broad question about "trust" into one concrete screen you can actually design.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, payoff, anchor, risk, keep out. The anchor is the actual answer to this question."
Why this works
Tells the interviewer you're building forward, not just reacting to a hypothetical.
Stage 3
Reframe the question
Say it like this
"This isn't really 'how do I make people believe the AI.' It's 'what does a seller need in front of them to tell a good guess from a shaky one, without me having to explain the model to them.'"
Why this works
Separates a real design answer from a vague pitch about "building confidence."
Stage 4
Give the one decision
Say it like this
"Instead of 'Suggested price: $28,' I'd show '$28, based on 3 sold matches like yours' with a small tag saying how sure the match is."
Why this works
This is a real screen, not a category of screen. A reader can picture exactly what changed.
Stage 5
Prove it with a failure
Say it like this
"Say the price is wrong on a rare item. With a bare number, the seller either trusts it and loses money, or stops trusting the whole tool. With the evidence shown, she can see there were only six thin matches, so she checks that one herself and trusts the other ninety-nine like she always did."
Why this works
Shows the design surviving the exact moment it's supposed to protect against.
Stage 6
Say what you'd measure
Say it like this
"I'd watch how often sellers tap 'why this price' by category. A rising tap rate on one category is a warning sign, weeks before anyone complains."
Why this works
Shows you're thinking past launch day, not just past the pitch.
Stage 7
Close on the one line
Say it like this
"Show the receipt, not just the price. A number asks for total trust. A receipt lets trust be partial, and partial trust is the kind that survives a bad Tuesday."
Why this works
Restates the direct answer in one breath, ready for whatever gets pushed on next.

Let's learn

Renlio's Suggested Price button looks at what nearly identical items sold for recently and hands a seller one number to list at.

Before the button existed, a seller dug through eight or ten sold listings by hand, about fifteen minutes of searching, to guess a fair price. With the button, that same guess arrives in under two seconds.

Knowledge spark: what makes a "comparable" listing? A sold item close enough in brand, condition, and style that its price is a fair stand-in for yours. Two comparables from a deep pool of matches is a much stronger guess than two from a pool of six.

Renlio's own numbers say three out of every four sellers list at exactly the suggested price without changing it at all.

Comparable sold listings found, by category
340 170 0 T-shirt 340 Sneakers 210 Sweater 9 Coat 6
Every one of these gets shown as one confident-looking number. Only two of the four actually have the data to earn that confidence.

The turn: the odd time the suggested price is wrong isn't really the problem. Renlio's price lands in a fair range on most listings. The real problem is a seller can't tell, from looking at the number alone, whether this one is a well-backed guess or a shaky one. That leaves exactly two choices: trust it completely, or check everything yourself and get none of the time back.

The number was never the thing that broke. What broke was a seller having no way to tell a well-backed guess from a shaky one, so she had to pick one rule and apply it to every item, forever.

At its worst: a lightly-worn coat gets priced the same as a heavily-worn one, because the model lumped "coat" into one bucket with too little data underneath. It sits unsold for a month, or sells for far less than it should have, and the seller decides the whole tool can't be trusted, going back to pricing everything by hand, including the ninety percent of items it would have nailed.

The decision I would take back We designed the price screen to show one clean number, no listings, no confidence tag, because a single number tested as simpler and more confident-looking in early demos. That was fine while the model only ever saw items with plenty of close matches. It stops being fine the moment the model has to guess thin, and the seller has no way to tell those two situations apart.

What I would leave alone: for high-volume categories, plain t-shirts, common paperback books, the single clean number is fine exactly as it is. There's so much sold-listing data behind it that adding a receipt would be clutter with no real benefit.

Sellers who changed the suggested price, by category, over 6 months
40% 20% 0% Month 1 Month 6 Common: ~4% Thin-data: 35%
Common categories barely moved. Thin-data categories climbed steadily for six months before anyone in the room noticed the pattern.

The lesson: a screen's confidence should match the model's actual confidence, not a fixed coat of polish painted on every answer whether it's earned that day or not.

Now here is the same thing as a story

The short version above is what you'd say defending this design in a review. Read this one for how Dara actually found the gap.

Dara Whitfield has resold secondhand coats and furniture on Renlio for eight months, most evenings after her day job, at the kitchen table, phone in one hand and a mug going cold in the other.

Before the Suggested Price button, she'd search sold listings herself, comparing brand and condition, and she was good at it. She could eyeball a coat's wear and land within a few dollars of a fair price.

Hand sketched flow diagram titled Pricing an item before the button existed. Four boxes: photograph item, search sold listings highlighted, guess a price, list it.
Before the button, the middle step, searching sold listings by hand, was the one that ate her evening.

For months after the button arrived, it was the best part of listing. She'd photograph an item, tap the button, and a price appeared in two seconds flat. She listed it and moved to the next item.

She stopped double-checking the number around month three. It kept being close enough, and she had a stack of items to get through.

Then came a Tuesday night in her resale group chat, nothing to do with a bad sale at all. A friend mentioned, almost in passing, "wait, you just list at whatever number it gives you? I always check three sold listings first before I trust it."

Dara went quiet for a second. She realized she hadn't checked a single suggested price in five months.

So she pulled up her last six listings and actually looked. A wool coat with heavy pilling had been priced the exact same as a nearly-new one she'd sold the month before. Both had listed at the same suggested number. The pilled one had sat for three weeks before finally selling for twelve dollars less than it should have.

Hand sketched comparison diagram titled The day the price is wrong. Left panel, a question mark box icon labeled Bare number, caption seller trusts it blind. Right panel, a document icon labeled Number plus proof, caption seller can check it.
Same wrong number, two very different Tuesdays, depending on whether Dara had anything to check it against.
She hadn't lost trust in the button. She'd lost the ability to tell which items it was actually sure about.

Here's the decision I'd take back. We built the button to show one clean number because it tested well in early demos, back when the model only ever saw coats with dozens of close matches. Nobody planned for the rare item with six thin matches to look exactly as confident on screen as the common one with three hundred.

Hand sketched labeled parts diagram titled The anchor, close up. Center document icon labeled Suggested Price, with four callouts: 3 sold matches, how sure, adjust condition, why this price.
The actual screen: not a bare number, but a price with its own receipt attached, one tap away.

I'd put the receipt back. Not the full pricing algorithm, just the three matches it actually used and a plain "how sure" tag: strong match, or thin match, worded exactly like that.

Hand sketched metaphor scene titled A guess versus a receipt. Left, a gauge icon labeled One Number, caption trust it whole. Right, a document icon labeled The Receipt, caption check the parts.
One design hands Dara a number to swallow whole. The other hands her a receipt she can actually read.

Replay the same pilled coat under the new screen: the tag reads "thin match, 6 listings." Dara taps once, sees the comparables are all lightly-worn items, adjusts the condition slider herself, and lists at a fair price instead of an average one. It sells in six days instead of sitting for three weeks.

The old screen handed her one number and asked her to trust the whole thing or none of it. The new one hands her a receipt, and lets her trust ninety percent of it completely and check the other ten percent for two seconds.

I built the clean version because it looked more confident in a demo. It took a friend's offhand question in a group chat to see that looking confident and being trustworthy aren't the same screen.

SPARK, in one screenNot a pitch. SPARK is what forces a design decision to survive its own worst day before it ships.

S
Situation. The job today, without you.
Dara searches sold listings by hand, about fifteen minutes per item, to guess a fair price.
Grounds the design in a real task someone already does, not an abstract need.
P
Payoff. The habit you want to build.
Stop researching every price by hand. Trust the fast path for most items, and know exactly which ones to double-check.
Names the thing she'll stop doing, which is the actual product.
A
Anchor. The one decision everything hangs on.
Show the price with its three real comparable matches and a plain "how sure" tag, one tap away, not buried and not forced on the main screen.
The hardest step, and the actual answer to the question being asked.
R
Risk. What breaks the first time it's wrong.
A thin-data guess on a rare item, like the pilled coat, needs to look visibly less certain than a guess backed by hundreds of matches.
Names the exact failure the anchor has to survive.
K
Keep out. What stays off day one.
No full algorithm explainer, no price-history graph, no seller-reputation weighting shown. Just the matches and the confidence tag.
Shows judgment, not a wish list of every feature that could theoretically help.
Hand sketched icon list titled What Renlio skipped on day one. Three items: full pricing algorithm explained, full price history graph, seller reputation weighting shown.
Every one of these would add real detail. None of them is what stopped the pilled coat from being mispriced.

The recap, one line per letter: situation is Dara pricing by hand, payoff is trusting the fast path without researching everything, anchor is the receipt behind the number, risk is a thin-data guess looking as confident as a deep one, keep out is everything that isn't the receipt.

And if you want to be sure it really works, try it somewhere elseSame five letters, a freelance translator instead of a reseller. A different screen, and this time the anchor is an ambiguous phrase, not a price.

Ingrid Bascomb translates legal documents for a freelance-work marketplace. A translation assistant suggests a phrase for each sentence, and she used to look up every ambiguous term herself in two reference dictionaries before typing anything.

Mapped onto SPARK: situation is Farida cross-checking every ambiguous phrase by hand, about a minute each, across a forty-page contract. Payoff is trusting the suggestion for plain sentences and knowing exactly which phrases still need her eye. Anchor is different from Dara's: instead of showing a confidence tag on a number, the assistant highlights the one word in a sentence that has more than one common legal meaning, and shows both readings side by side. Risk: the day a highlighted word is actually unambiguous in context, and the tool flagged it anyway. Keep out: no full grammar explainer, no dictionary browser bolted onto the screen, just the flagged word and its two readings.

Hand sketched quadrant titled Which trust signals actually earn trust. Axes specific from vague to exact, and checkable from must take our word to seller can check. Bare price sits low on both. Confidence tag sits middle. Disclaimer text sits low on both. Sold matches shown sits high on both.
A disclaimer and a confidence tag both sound reassuring. Only one of them actually gives the reader something to check.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show the receipt, not just the answer," and stop.
Cost: there's no engineering time this quarter to build the confidence tag. Start with the three comparable listings alone; even without a formal confidence label, seeing thin data is often obvious on sight.
The model gets better, for real: if Renlio's pricing model gets more accurate overall, that's still not a reason to hide the receipt. A better average model can still be badly wrong on the one rare item a seller happens to be holding.

Where people run it wrong.
They treat "explain the AI" as a wall of text about how the model works, which nobody reads.
They show a confidence score with no way to check it, which is just a second number to blindly trust.
They build the full explanation experience for every item, including the ones with so much data that a bare number was already fine.

How to use it live. When someone asks what a user needs to see to trust a recommendation, ask yourself: could this exact reader tell, from the screen alone, when the recommendation is standing on solid ground versus thin ice? If not, name the smallest thing you'd add so they could.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what does a user need to see to trust an AI recommendation?"
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Anchor is the concrete design decision that answers the question.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dara Whitfield, who resells coats and furniture on Renlio most evenings, and used to research every price by hand.
3 · THE HABIT
What did Dara stop doing because the button worked?
Tap to flip
ANSWER
She stopped double-checking the suggested price against sold listings herself, around month three.
4 · THE ANCHOR
What's the one concrete design decision this answer commits to?
Tap to flip
ANSWER
Show the price with its real comparable sold matches and a plain "how sure" tag, one tap away, instead of a bare confident number.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Showing one clean number with no listings and no confidence tag, since it tested well in early demos before thin-data items were common.
6 · THE NUMBER
Fill in the blank: the pilled coat had only ___ comparable sold listings behind its price.
Tap to flip
ANSWER
6. A plain t-shirt, by contrast, had 340 comparable sold listings behind the exact same kind of number.
7 · THE REPLAY
Same pilled coat, redesigned screen. What changes?
Tap to flip
ANSWER
Dara sees a "thin match" tag, checks the 6 comparables herself, adjusts for condition, and it sells in 6 days instead of sitting 3 weeks underpriced.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the anchor there?
Tap to flip
ANSWER
Farida's legal translation assistant. There, the anchor is highlighting the one ambiguous word and showing both readings, not a confidence tag on a number.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: about three out of every four sellers list at exactly the ___ price without changing it.
Show hint
Look at the first paragraph under "Let's learn."
Show answer
Suggested. That's exactly why a wrong guess on a thin-data item matters so much: most sellers never second-guess it.
Multiple choice
2. Why does this answer recommend showing comparable listings instead of just a confidence percentage?
  • A. Because percentages are illegal to show to consumers.
  • B. Because a percentage alone still can't be checked; the listings give the seller something to actually verify.
  • C. Because percentages are always inaccurate.
  • D. Because listings load faster than a percentage.
Show hint
Look at the quadrant diagram comparing trust signals.
Show answer
B. A checkable signal, like real listings, does more work than a confident-sounding number with nothing behind it.
True or false
3. True or false: this answer says every item should show its full comparable-listings breakdown by default, on the main screen.
  • True
  • False
Show hint
Look at priority list item 5, and the "keep out" step.
Show answer
False. It stays one tap away, so the fast path stays fast for sellers who don't want to dig into every listing.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Showing one clean number with no evidence behind it. It made sense while the model mostly saw items with deep, reliable comparable data.
Short answer, where it wouldn't matter
5. Name a category where the bare number, with no receipt, is genuinely fine as-is.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A high-volume category like plain t-shirts, where hundreds of comparable sold listings back every guess.
Short answer, apply it yourself
6. Pick an AI recommendation you've used yourself. What's one piece of evidence you wish it had shown you, instead of just a confident-sounding answer?
Show hint
Think of a time a recommendation felt like a black box, and what would have let you check it.
Show answer
Model answer: Many people point to a streaming or shopping recommendation and say they wish it showed which of their own past choices it was actually based on.
Before you close the answer
Why this works
Tests whether you can name a specific, buildable interface decision instead of a vague value like "transparency," and whether that decision actually survives the moment the recommendation is wrong.
Follow-up traps
"Isn't showing the comparable listings just moving the trust problem one level down, since the listings themselves could be wrong?" Response: yes, partly, but a wrong listing is something a seller can actually spot and dispute, unlike a wrong number floating with no source at all.

"Won't showing 'thin match, 6 listings' just make sellers distrust the tool on rare items entirely?" Response: that's the point. Distrust aimed at exactly the right ten percent of items is far cheaper than blanket distrust of everything, which is what a bare wrong number eventually causes.
If pressed
Renlio's real confidence tag isn't a raw model score shown to users; it's bucketed into three plain labels, strong, fair, and thin match, since a raw percentage tested as more confusing than helpful to non-technical sellers.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more