CaseAdvancedDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #11

How do you build trust with a user who has had a bad experience with AI elsewhere?

ORDER the product is Ridgeplan Assist, an AI feature in Ridgeplan that drafts client status updates for Onyxleaf Creative

Onyxleaf Creative is a mid-size marketing agency. Grace Adeyemi runs operations there and has been with the agency for six months. Before Onyxleaf, an AI status-update tool at her last job auto-sent a client an update claiming a milestone was finished. It wasn't.

The direct answer
Never let the AI publish anything client-facing on its own. Show the exact draft, with the specific ticket or update it pulled from, and require one deliberate click before anything reaches a client. Let a wary user earn trust on low-stakes internal drafts first, and only ask them to trust it in front of a client once they've watched it be right, and correctable, many times over.
Do this, in order
  1. Never auto-send anything client-facing. Always require a review click.Why: an unreviewed send is nearly impossible to take back once a client has read it.
  2. Start a wary user on internal-only drafts, not client-facing ones.Why: low-stakes practice is what lets someone learn the tool's shape before anything is riding on it.
  3. Show which specific ticket or update each line of a draft came from.Why: a spot-check that takes ten seconds beats a rewrite that takes ninety minutes.
  4. Track a user's own override rate over time as real evidence.Why: a falling override rate is proof the tool is earning trust, not just a claim that it should.
  5. Keep the review-required rule permanent for client-facing sends, even after trust builds.Why: the moment a tool "graduates" to auto-send is exactly the moment a rare mistake reaches a client unseen.

How to answer this, stage by stage

This isn't a question about being extra nice to one nervous user. It's a question about which trust-building move to do first, and which ones can wait.

Stage 1
Scope it to one real person
Say it like this
"I'll answer this for Grace, an operations manager at a creative agency, who was burned once already by an AI tool that sent a client a false update."
Why this works
Turns "a wary user" into a specific person with a specific reason to be wary.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what we're actually trying to earn. Reversibility, which mistake is hardest to undo. Dependency, what has to happen before what. Evidence, what proves it cheaply. Rank, the actual order."
Why this works
Signals this is a ranking problem, not a list of nice ideas in no particular order.
Stage 3
Reframe the question
Say it like this
"This isn't really 'how do I make Grace comfortable.' It's 'what's the one design decision that, if I skip it, guarantees the same failure happens to her again.'"
Why this works
Moves from a feelings question to a design-priority question, which is what ORDER is built to answer.
Stage 4
Rank by what's hardest to undo
Say it like this
"An unreviewed message that already reached a client can't be unsent. A slow rollout, or a plainer interface, can always be sped up or prettied up later. So the review-before-send rule goes first, above everything else."
Why this works
This is ORDER's actual argument: rank by reversibility, not by what's easiest to build.
Stage 5
Name the dependency
Say it like this
"Grace won't trust the tool on anything client-facing until she's watched it be right, and easy to correct, somewhere low-stakes first. Internal trust has to come before external trust. There's no shortcut around that order."
Why this works
Shows which step genuinely can't happen before the one in front of it.
Stage 6
Prove it with a near miss
Say it like this
"A new hire almost let Ridgeplan Assist auto-send an unreviewed update straight to a major client, because auto-send was the default setting. Grace caught it by chance, glancing at the wrong screen at the right moment."
Why this works
Shows the reversibility risk isn't theoretical, it nearly repeated the exact failure Grace already lived through.
Stage 7
Say what the evidence would look like
Say it like this
"I'd watch her override rate on internal drafts drop over the first month. Forty percent in week one, down near eight percent by week four, is the cheap evidence that trust is building for real, before anything client-facing is ever on the table."
Why this works
Names a real number you'd track, not just a hope that trust improves.
Stage 8
Close on the one line
Say it like this
"Review before send comes first because it's the one thing that can't be undone if you skip it. Everything else, speed, polish, confidence, can be added later. Trust can't be un-lost twice."
Why this works
Restates the ranking logic in one breath, ready for whatever gets pushed on next.

Let's learn

Grace had already been burned once by an AI update that lied to a client. She was not going to let it happen twice, and the interesting part is what that meant she actually did.

Ridgeplan Assist reads a project's tickets, comments, and recent activity, then drafts a plain-language status update a person can send to a client.

Knowledge spark: what's an override rate? How often a person edits or rewrites an AI draft before using it. A high override rate means the draft usually needs real fixing. A falling override rate over time is a sign the tool is getting more of the small things right on its own.

Before Ridgeplan, Grace wrote every client status update herself, reading across four project boards, which took about 90 minutes each Friday. Onyxleaf's other account leads had switched to Ridgeplan Assist within weeks. Grace kept writing hers by hand, quietly, for six months.

Time to produce one client status update, by method
90 min 45 min 0 By hand: 90 min Reviewed draft: 12 min
Grace never got the 78 minutes back, because she never trusted the tool enough to try it on anything that mattered.

The turn: Grace's caution was never really about the model being unreliable. Every account lead using Ridgeplan Assist had fine results. The real risk was that nobody had built a version of the tool that let someone like her try it without betting a client relationship on the first attempt.

The decision I would take back Ridgeplan Assist shipped with auto-send turned on by default for any update marked "routine," to save a click for teams moving fast. That made sense for internal updates nobody was worried about. It stopped making sense the moment a routine-marked update was headed to a client, and a person who'd never opted into that setting had no idea it was on.

What I would leave alone: auto-send for genuinely internal-only summaries, like a team's own daily standup notes, is fine to leave as-is. Nobody outside the team ever sees those, and a small mistake there costs nothing worse than a confused Slack message.

Grace didn't need a smarter model. She needed proof, on something small, before anyone asked her to risk something big.

The lesson: a person who's already been burned once isn't asking you to convince them the model is good. They're asking you to prove they'll get a chance to catch it if it isn't, before it costs them anything real again.

Hand sketched comparison diagram titled Grace's Friday, before and after. Left panel, a document icon labeled Before, caption 90 minutes four boards by hand. Right panel, a gauge icon labeled After, caption 12 minutes one reviewed draft.
The 78 minutes were always available. Grace just needed a reason to believe the draft wouldn't need a full rewrite.

Now here is the same thing as a story

The short version above is what you'd say defending this rollout order to a skeptical product lead. Read this one for how Grace actually came around.

Grace could once read a client's tone in a single line of feedback and know, before anyone else on the account, whether a relationship was drifting sideways. That instinct is what made the old fabricated update sting so badly. She hadn't caught it. The tool had, by design, never given her the chance to.

At Onyxleaf, she watched her teammates use Ridgeplan Assist for months without incident, drafting client updates that always looked clean. She was glad for them, and she kept writing hers by hand anyway. Nobody made her switch. Nobody really noticed she hadn't.

Hand sketched timeline titled Grace's timeline. Four milestones: old job fabricated update lost a client, joins Onyxleaf avoids Assist entirely, new hire near miss almost auto sent unreviewed highlighted, review required ships override rate falls to 8 percent.
The near miss didn't happen to Grace directly. It happened next to her, which turned out to matter just as much.

Then a new hire, three weeks into the job and unaware of any of Grace's history, opened a client thread with Ridgeplan Assist's draft already queued to send automatically, since the update had been auto-marked "routine." Grace happened to glance at his screen while walking past, saw the client's name in the send field, and said, quietly, "wait, did you read that first?"

He hadn't. Neither had the system asked him to.

Hand sketched decision tree titled Reversible or not. Root: Ridgeplan Assist finishes a draft. One branch, auto send is on, leads to sent to client hard to undo. Other branch, review required, leads to Grace approves easy to undo.
Onyxleaf had been living on the left branch without knowing it. The fix was making the right branch the only one that exists.

What that near miss cost, even though nothing actually went wrong: a full afternoon of Grace walking every account lead through their auto-send settings by hand, and a much louder version of the fear she'd already been carrying quietly for six months, that this could happen to anyone, not just to her.

I shipped the "routine" auto-send default because it saved active teams a genuinely annoying extra click, and testing with confident, fast-moving users never surfaced a problem. It took watching it nearly repeat, almost verbatim, the exact failure that had already cost someone a client once, to see that a saved click is never worth an unreviewable send.

Hand sketched flow diagram titled What has to unlock what. Four boxes: preview always shown highlighted, trust on internal drafts, sources visible, client facing allowed.
Client-facing trust sits at the end of this chain, not the front of it. Nothing unlocks it faster than skipping a step.

With review-required as the only setting, Grace started small: her own team's internal weekly notes, nothing a client would ever see. Her override rate started high and fell fast as she learned exactly which lines Ridgeplan Assist got right without help.

Hand sketched icon list titled What earns default trust, in order. Four items: always show the draft first, start on internal only summaries, show which ticket each line came from, track her override rate as evidence.
Grace didn't need to be convinced. She needed four weeks of evidence she could see for herself.

Six weeks in, she used Ridgeplan Assist on a client update for the first time, still reviewed it fully, still caught and fixed one line, and sent it herself. The old process asked her to trust the tool completely, on day one, with something she couldn't take back. The new one let her build that trust in a place where nothing was riding on it yet.

ORDER, the trust queueNot a comfort checklist. ORDER is what forces you to rank which trust-building move actually has to come first.

O
Outcome. What we're earning.
Grace using Ridgeplan Assist on client-facing work without a second AI trust failure.
Without a named outcome, any ranking is just opinion.
R
Reversibility. The hard step.
A message already sent to a client can't be unsent. A slow rollout can always be sped up later.
This is the actual argument for putting review-before-send first.
D
Dependency. What unlocks what.
Trust on internal drafts has to come before trust on client-facing ones. There's no shortcut past that order.
Names what genuinely can't happen out of sequence.
E
Evidence. What proves it cheaply.
Grace's own override rate on internal drafts, watched weekly, before any client-facing use is even discussed.
Cheap, real, and specific to the one person you're trying to convince.
R
Rank. The actual order.
Review before send, then internal-first practice, then visible sourcing, then tracked evidence, then, only then, client-facing use.
States the order plainly and defends the top pick.
Grace's override rate on internal drafts, over four weeks
40% 20% 0 40% 24% 14% 8% Week 1 Week 2 Week 3 Week 4
This is what "trust building" looks like as a real number, not a feeling. It's also exactly the evidence step of ORDER.

The recap, one line per letter: outcome is real client-facing use without a repeat failure, reversibility is why review-before-send outranks everything else, dependency is internal trust before external trust, evidence is the falling override rate, and rank states the order plainly.

And if you want to be sure it really works, try it somewhere elseSame five letters, a smart-home security app instead of a status-update tool. A different industry, and this time the old failure was a false alarm, not a fabricated update.

Kittery Robotics makes a smart-home security system with an AI monitoring assistant. Dov Ashkenazi just switched to it after his previous smart-home AI, at a different company, wrongly flagged him as an intruder in his own home and called the police.

Mapped onto ORDER: outcome is Dov trusting the system enough to leave it fully armed while he's away. Reversibility is that a false police dispatch can't be undone once officers arrive, while a quiet notification can always be escalated later if needed. Dependency is that Dov needs to see the system correctly explain its own uncertain moments before he'll trust it on a genuine alarm. Evidence is his own rate of dismissing false alerts, tracked over his first month, falling as the system's explanations prove reliable. Rank: always show a clear reason before any alarm escalates, start him on notification-only mode, let him watch its explanations for a few weeks, track his dismissal rate, and only then offer full auto-escalation as something he opts into himself.

Hand sketched quadrant titled Sorting trust building moves, cost vs speed. Axes cost to build from cheap to expensive, how fast it rebuilds confidence from slow to fast. Always preview mode sits top left, cheap and fast. Explain each alert sits middle. Full audit log sits bottom right, expensive and slow. Manual arm override sits lower middle.
The cheap, fast-acting moves are exactly where to start with someone who's already been burned once.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "review before send comes first, because it's the one mistake you can't undo," and stop.
Cost: there's no budget for a full override-tracking dashboard this quarter. Say so, and start with the review-required rule alone, since it costs almost nothing and carries most of the actual risk reduction.
The model gets better, for real: even if Ridgeplan Assist's draft quality improves next quarter, that's not a reason to relax the review-before-send rule. A better model can still be wrong exactly once, in front of exactly the wrong client.

Where people run it wrong.
They try to win back trust with reassurance and messaging, instead of an actual design decision the wary person can watch work.
They let a good override rate "graduate" a user straight to auto-send, quietly reintroducing the exact risk they were trying to remove.
They treat every user's caution as identical, instead of starting the specifically wary one on lower-stakes practice first.

How to use it live. When asked how to build trust with a burned user, ask yourself first: which single mistake, if it happened again, could never be taken back? Rank your fix for that one first, and let everything else follow behind it.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you build trust with a user burned by AI elsewhere"?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. The reversibility step is what puts review-before-send at the top.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Grace Adeyemi, an operations manager at Onyxleaf Creative, previously burned by a fabricated AI status update at her last job.
3 · THE HABIT
What did Grace do instead of using Ridgeplan Assist, for six months?
Tap to flip
ANSWER
She wrote every client status update by hand, across four project boards, taking about 90 minutes each Friday.
4 · THE REVERSIBILITY ARGUMENT
Why does review-before-send outrank every other trust-building move?
Tap to flip
ANSWER
A message already sent to a client can't be unsent. Every other improvement, speed or polish, can still be added later.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Defaulting "routine"-marked updates to auto-send, to save active teams a click, without telling users the setting existed.
6 · THE NUMBER
Fill in the blank: Grace's override rate on internal drafts fell from 40 percent in week one to ___ percent by week four.
Tap to flip
ANSWER
8 percent. That decline is the evidence step of ORDER, made countable.
7 · THE REPLAY
Same near miss, redesigned defaults. What changes?
Tap to flip
ANSWER
There's no auto-send setting left to trigger at all. Every client-facing draft, routine or not, waits for a deliberate review click.
8 · CROSS PRODUCT TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what plays the role of "review before send"?
Tap to flip
ANSWER
Kittery Robotics' smart-home security app. There, "always show a clear reason before any alarm escalates" plays that role.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Grace's old status-update tool auto-sent a client an update claiming a ___ was finished, when it wasn't.
Show hint
Look at the lede.
Show answer
A milestone. That single fabricated claim, sent without review, cost her the client relationship at her previous job.
Multiple choice
2. Why does this answer rank "never auto-send client-facing updates" above "show which ticket each line came from"?
  • A. Because sourcing is technically harder to build.
  • B. Because an unreviewed send that reaches a client can't be undone, while a missing source line can be added later without harm.
  • C. Because clients don't care about sourcing.
  • D. Because auto-send only matters for large agencies.
Show hint
Look at ORDER's reversibility step.
Show answer
B. ORDER ranks by what's hardest to undo, not by what's easiest or hardest to build.
True or false
3. True or false: this answer says a wary user's override rate should eventually let auto-send turn back on for client-facing updates.
  • True
  • False
Show hint
Look at the last priority-list item.
Show answer
False. Review-required for client-facing sends stays permanent, regardless of how low anyone's override rate gets.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Defaulting routine-marked updates to auto-send, to save fast-moving teams a click. It made sense for internal use and stopped making sense the moment a routine tag reached a client.
Short answer, apply it yourself
5. Think of a tool you distrust because of a past bad experience, with this tool or a different one. What's the one low-stakes way you'd want to test it again before trusting it with something that matters?
Show hint
Think about what Grace did: internal-only practice before anything client-facing.
Show answer
Model answer: Most people land on some version of "let me watch it be right, and correctable, on something small, before I let it touch something I can't take back."
Before you close the answer
Why this works
Tests whether you can turn "make a nervous user comfortable" into an actual ranked design decision, instead of a reassurance strategy with nothing concrete behind it.
Follow-up traps
"Isn't requiring review forever just admitting the model can't be trusted?" Response: no, it's admitting that the cost of one unreviewed mistake reaching a client is far higher than the small, permanent cost of a review click.

"What if the wary user just never tries the internal-only version either?" Response: then the real problem isn't the design, it's that nobody's shown them a genuinely low-stakes reason to try it at all, which is a separate, earlier fix.
If pressed
Onyxleaf's real fix logs every override alongside the specific line a user changed, not just a yes-or-no edited flag, so the product team can see which kinds of lines Ridgeplan Assist still gets wrong most often, not just how often people fix something.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more