What does a user need to see to trust an AI recommendation?
Renlio is a resale marketplace. Its Suggested Price button looks at recent sold listings and hands a seller one number to list at. Dara Whitfield resells secondhand coats and furniture on Renlio most evenings after her day job, at her kitchen table, phone in hand.
- Show the real comparable listings behind the number, not just the number.Why: a bare number gives a seller nothing to check it against.
- Show how much data backed the guess, not a fixed tone of confidence.Why: a guess from six listings and a guess from four hundred should never look equally sure.
- Let the seller change one input, like condition, and watch the price move.Why: a number you can poke at is a number you can actually judge.
- Say plainly when there isn't enough data, instead of guessing anyway with a straight face.Why: a confident wrong answer costs more trust than an honest shrug.
- Keep the whole thing behind one tap, not bolted onto the main screen.Why: most sellers don't want to dig, and the fast path should stay fast for them.
- Watch how often sellers tap "why this price," as an early warning sign.Why: that number moves weeks before anyone complains out loud.
How to answer this, stage by stage
Nobody's grading whether you know AI can be wrong. They're grading whether you can name the exact thing on the screen that lets a normal person tell a good guess from a shaky one.
Let's learn
Renlio's Suggested Price button looks at what nearly identical items sold for recently and hands a seller one number to list at.
Before the button existed, a seller dug through eight or ten sold listings by hand, about fifteen minutes of searching, to guess a fair price. With the button, that same guess arrives in under two seconds.
Renlio's own numbers say three out of every four sellers list at exactly the suggested price without changing it at all.
The turn: the odd time the suggested price is wrong isn't really the problem. Renlio's price lands in a fair range on most listings. The real problem is a seller can't tell, from looking at the number alone, whether this one is a well-backed guess or a shaky one. That leaves exactly two choices: trust it completely, or check everything yourself and get none of the time back.
At its worst: a lightly-worn coat gets priced the same as a heavily-worn one, because the model lumped "coat" into one bucket with too little data underneath. It sits unsold for a month, or sells for far less than it should have, and the seller decides the whole tool can't be trusted, going back to pricing everything by hand, including the ninety percent of items it would have nailed.
What I would leave alone: for high-volume categories, plain t-shirts, common paperback books, the single clean number is fine exactly as it is. There's so much sold-listing data behind it that adding a receipt would be clutter with no real benefit.
The lesson: a screen's confidence should match the model's actual confidence, not a fixed coat of polish painted on every answer whether it's earned that day or not.
Now here is the same thing as a story
The short version above is what you'd say defending this design in a review. Read this one for how Dara actually found the gap.
Dara Whitfield has resold secondhand coats and furniture on Renlio for eight months, most evenings after her day job, at the kitchen table, phone in one hand and a mug going cold in the other.
Before the Suggested Price button, she'd search sold listings herself, comparing brand and condition, and she was good at it. She could eyeball a coat's wear and land within a few dollars of a fair price.
For months after the button arrived, it was the best part of listing. She'd photograph an item, tap the button, and a price appeared in two seconds flat. She listed it and moved to the next item.
She stopped double-checking the number around month three. It kept being close enough, and she had a stack of items to get through.
Then came a Tuesday night in her resale group chat, nothing to do with a bad sale at all. A friend mentioned, almost in passing, "wait, you just list at whatever number it gives you? I always check three sold listings first before I trust it."
Dara went quiet for a second. She realized she hadn't checked a single suggested price in five months.
So she pulled up her last six listings and actually looked. A wool coat with heavy pilling had been priced the exact same as a nearly-new one she'd sold the month before. Both had listed at the same suggested number. The pilled one had sat for three weeks before finally selling for twelve dollars less than it should have.
Here's the decision I'd take back. We built the button to show one clean number because it tested well in early demos, back when the model only ever saw coats with dozens of close matches. Nobody planned for the rare item with six thin matches to look exactly as confident on screen as the common one with three hundred.
I'd put the receipt back. Not the full pricing algorithm, just the three matches it actually used and a plain "how sure" tag: strong match, or thin match, worded exactly like that.
Replay the same pilled coat under the new screen: the tag reads "thin match, 6 listings." Dara taps once, sees the comparables are all lightly-worn items, adjusts the condition slider herself, and lists at a fair price instead of an average one. It sells in six days instead of sitting for three weeks.
The old screen handed her one number and asked her to trust the whole thing or none of it. The new one hands her a receipt, and lets her trust ninety percent of it completely and check the other ten percent for two seconds.
I built the clean version because it looked more confident in a demo. It took a friend's offhand question in a group chat to see that looking confident and being trustworthy aren't the same screen.
SPARK, in one screenNot a pitch. SPARK is what forces a design decision to survive its own worst day before it ships.
The recap, one line per letter: situation is Dara pricing by hand, payoff is trusting the fast path without researching everything, anchor is the receipt behind the number, risk is a thin-data guess looking as confident as a deep one, keep out is everything that isn't the receipt.
And if you want to be sure it really works, try it somewhere elseSame five letters, a freelance translator instead of a reseller. A different screen, and this time the anchor is an ambiguous phrase, not a price.
Ingrid Bascomb translates legal documents for a freelance-work marketplace. A translation assistant suggests a phrase for each sentence, and she used to look up every ambiguous term herself in two reference dictionaries before typing anything.
Mapped onto SPARK: situation is Farida cross-checking every ambiguous phrase by hand, about a minute each, across a forty-page contract. Payoff is trusting the suggestion for plain sentences and knowing exactly which phrases still need her eye. Anchor is different from Dara's: instead of showing a confidence tag on a number, the assistant highlights the one word in a sentence that has more than one common legal meaning, and shows both readings side by side. Risk: the day a highlighted word is actually unambiguous in context, and the tool flagged it anyway. Keep out: no full grammar explainer, no dictionary browser bolted onto the screen, just the flagged word and its two readings.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show the receipt, not just the answer," and stop.
Cost: there's no engineering time this quarter to build the confidence tag. Start with the three comparable listings alone; even without a formal confidence label, seeing thin data is often obvious on sight.
The model gets better, for real: if Renlio's pricing model gets more accurate overall, that's still not a reason to hide the receipt. A better average model can still be badly wrong on the one rare item a seller happens to be holding.
Where people run it wrong.
They treat "explain the AI" as a wall of text about how the model works, which nobody reads.
They show a confidence score with no way to check it, which is just a second number to blindly trust.
They build the full explanation experience for every item, including the ones with so much data that a bare number was already fine.
How to use it live. When someone asks what a user needs to see to trust a recommendation, ask yourself: could this exact reader tell, from the screen alone, when the recommendation is standing on solid ground versus thin ice? If not, name the smallest thing you'd add so they could.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Won't showing 'thin match, 6 listings' just make sellers distrust the tool on rare items entirely?" Response: that's the point. Distrust aimed at exactly the right ten percent of items is far cheaper than blanket distrust of everything, which is what a bare wrong number eventually causes.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Trust, transparency and explainability in UX
- #2 Explain the difference between explainability and transparency in a product context.
- #3 How do citations change user behaviour, and what happens when they are wrong?
- #4 Design the disclosure that tells a user they are talking to an AI.
- #5 When does showing the model's reasoning help, and when does it reduce trust?
- #6 Critique a design that surfaces a chain of thought to end users.
- #7 How much should you tell users about which model powers a feature?