Design the first-run experience for a user who has never used an AI feature.
- Show the reasons behind every result, never a bare confidence score.Why: a score asks for total trust and breaks completely on the first miss. Reasons let trust be partial, so one wrong result doesn't end the relationship.
- Make the first wrong result something the user can explain to someone else in one sentence.Why: the real risk isn't a wrong answer, it's the user going quiet about the tool afterwards. An explainable miss stays visible. An unexplainable one goes underground.
- Give them a private way to disagree before you ask them to trust it publicly.Why: the first real decision shouldn't be "back this in front of my boss." Let them override it safely once or twice before it carries any social cost.
- Track overrides and rejections automatically, without waiting for the user to report them.Why: if trust breaks quietly, the only warning sign is behaviour you can observe. Nobody files a ticket saying "I've stopped believing this."
- Hold accuracy at "good enough" and spend the effort on the explanation instead.Why: a slightly imperfect result costs almost nothing if they're skimming anyway. An unexplained miss costs the whole relationship.
- Keep the very first real use low-stakes if the product can arrange it.Why: trust calibrates on a track record. One high-stakes first call means a single miss teaches the lesson before any track record exists to balance it.
How to answer this, stage by stage
Eight moves. Each one has the words you'd actually say. Read them aloud once, then say them your own way.
Let's learn
Say we build a tool that reads a stack of resumes and ranks them. It goes into a small dental clinic. The office manager will be the first person there to touch any AI tool at all.
Before we had it, she read all 40 resumes by hand, one evening a year, after the clinic closed. It took her four hours. She was good at it. She could tell in ten seconds if someone had never worked a front desk.
Now she presses a button and gets a ranked list in ten seconds. That is not the interesting part.
This is her first time trusting a machine with a decision like this. She has no feel yet for when to believe it and when not to. That feel has to come from somewhere. If we don't build it in, she builds her own, and we don't get to choose what she learns.
So the real design question is not "how do we make a good ranked list." It is what happens the first time the list is wrong.
What I would leave alone. The ranking itself does not need to be perfect on day one. A slightly imperfect order is fine, because she is going to skim the top ten either way, not blindly hire whoever is first. Chasing ranking accuracy before we have earned her trust in the reasons is solving the wrong problem first.
The long version, for when you want to feel why it matters
You don't need this to answer the question. Read it when you want to understand why the answer is right, not just what it is.
Nadia has run the front desk and the books at a six-chair dental clinic for eleven years. Hiring is not her whole job. It is two weeks a year, usually in January, when someone leaves and she has to find the next one. She is good at it in a way nobody ever taught her. She can read a resume and know, almost without thinking, whether someone has actually sat at a front desk with a phone ringing and a patient waiting, or whether they just think they'd be good at it.
Every January until this year, she did it the same way. She'd wait until the clinic closed, make a cup of tea, and go through the whole stack at the kitchen table. Forty resumes, one evening, four hours. By the end she'd have six names on a sticky note and she trusted every one of them, because she'd read every word herself.
This January there is a new tool built into the practice software. It reads the resumes and hands her a ranked list before she has even opened her laptop, the one with the cracked hinge she keeps meaning to replace.
The list is good. Genuinely good. The top three are exactly who she would have picked. She hires the first name on it, a woman named Beth, and Beth is wonderful. In March, Nadia mentions it to Dr. Halvorsen, almost proudly. "The new tool actually found her. Ranked her first."
Then, in June, a second front desk role opens. She runs it again. It ranks a young man named Colton at the top. He interviews well. She hires him.
Colton does not work out. Not dramatically. He is polite and slow to learn the insurance codes and twice sends a patient's file to the wrong fax number. Nothing that gets him fired outright. Just enough that by August, Dr. Halvorsen asks her, gently, what happened with the hiring this time.
And here is the thing Nadia does. She does not say the tool ranked him first. She says she liked his interview. That is not even false. She did like his interview. She just leaves out the part where the ranking is what put him in front of her.
She keeps using the tool. Nobody told her to stop. But from that August on, she reads every resume herself again, the whole stack, the way she used to, and only opens the ranked list afterwards, alone, to see if it agrees with her. If it does, she feels a small private relief. If it doesn't, she never mentions that either.
We lost the one moment in March where she said, out loud, to the person she answers to, that a machine had done something worth trusting. That sentence does not come back once it stops getting said.
I want to say the problem was that Colton was a bad ranking. But Beth was a good one, and it didn't matter. One good pick and one bad pick, and Nadia landed exactly where a fair person would land after that: private use, public silence.
So here is the decision, made back when this was still a screen in a design file. We were going to show her a single number next to each name. Ninety-four percent for Beth. Eighty-nine for Colton. Clean, confident, and it tests well in a demo.
Take that number off before it ever reaches her. In its place, three or four lines under each name. For Colton: matched two years of reception experience, comfortable with the phone system. Missing: any mention of insurance billing, which is half the job.
Now walk August forward with that version instead. Colton is struggling with insurance codes. Nadia opens the tool's old note on him, out of curiosity more than anything. It already told her, in June, that billing was missing. It was not that the tool was wrong about Colton. It told her exactly where the gap was. She hired him anyway for the parts he was strong on, which was a real and reasonable choice.
Now when Dr. Halvorsen asks what happened, she has an answer that costs her nothing to say. "The tool flagged that he had no billing background. I hired him anyway because the rest was strong, and that's the part that didn't hold up."
A score asks someone to trust a verdict. A reason lets them share one. We built the version that asks the harder thing first, on the one day she had no practice trusting us at all.
SPARK, in one screen
This is a design question, so the framework is SPARK. (A "what if the error rate doubled" question would use FLIPS instead. Different question shapes get different tools.)
Run SPARK on a different product
A small accounting firm rolls out an AI tool that drafts the first pass of a client's tax return from their uploaded documents. Same framework, completely different risk.
S. Marcus, a solo bookkeeper with about ninety small-business clients, all due the same week every April.
P. The habit you want: he stops re-entering every number by hand and starts reviewing the draft instead.
A. Flag which documents are incomplete or unusual for that client at intake, before the draft is generated, so a messy client visibly asks for more attention from the start.
R. If one bad draft slips through on a messy client and it costs him a call from the tax office, he doesn't review harder. He starts running only his simplest clients through the tool and quietly types the messy ones from scratch. Full trust at the top of the client list, none at the bottom.
K. Don't try to handle the messy clients automatically on day one. Flagging them is enough, and it's the part that keeps him from writing them off the tool entirely.
Swap the trigger and it still runs
- Speed: the tool starts taking eight seconds instead of being instant. His habit doesn't change, because eight seconds still beats typing.
- Cost: the firm starts charging per return processed. He saves it for higher-fee clients and hand-does the small ones. That's the substitution flip wearing the scope flip's clothes.
- The model gets better: draft accuracy jumps. He stops skimming the "everything matched" cases, and the one time a document is subtly wrong with nothing flagged, it ships inside a filed return.
Where people run it wrong
- Adding a big red warning banner on every draft. Everyone tunes out a banner that fires every time.
- Putting the missing-item flags in a separate report the person has to remember to open. If it isn't on the same screen as the draft, it doesn't exist.
- Writing the flags in the tool's language instead of the person's. "Schema mismatch on Form 4562" tells a bookkeeper nothing. "No mileage log attached" does.
If you're asked this cold
Buy yourself ten seconds by naming the person out loud first. "Let me put someone in this. Say it's a solo bookkeeper doing ninety returns a year." Naming the person is not stalling. It's stage 1 of the answer, and it buys you the time to find the real flip instead of reaching for the first one you think of.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Onboarding users to probabilistic products
- #2 How do you teach users what the AI can and cannot do without a manual?
- #3 Explain the role of example prompts in onboarding and their downside.
- #4 What is the risk of an onboarding that oversells capability?
- #5 How would you set expectations about errors during onboarding?
- #6 Describe progressive disclosure for a complex AI feature.
- #7 Critique an onboarding that starts with a blank chat box.