CaseIntermediateDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #1

Design the first-run experience for a user who has never used an AI feature.

The direct answer
Show the reasons behind each result, not a confidence score. A first run is teaching someone how much to trust the next hundred results, and a score is a verdict they can only swallow or hide.
Do this, in order
  1. Show the reasons behind every result, never a bare confidence score.Why: a score asks for total trust and breaks completely on the first miss. Reasons let trust be partial, so one wrong result doesn't end the relationship.
  2. Make the first wrong result something the user can explain to someone else in one sentence.Why: the real risk isn't a wrong answer, it's the user going quiet about the tool afterwards. An explainable miss stays visible. An unexplainable one goes underground.
  3. Give them a private way to disagree before you ask them to trust it publicly.Why: the first real decision shouldn't be "back this in front of my boss." Let them override it safely once or twice before it carries any social cost.
  4. Track overrides and rejections automatically, without waiting for the user to report them.Why: if trust breaks quietly, the only warning sign is behaviour you can observe. Nobody files a ticket saying "I've stopped believing this."
  5. Hold accuracy at "good enough" and spend the effort on the explanation instead.Why: a slightly imperfect result costs almost nothing if they're skimming anyway. An unexplained miss costs the whole relationship.
  6. Keep the very first real use low-stakes if the product can arrange it.Why: trust calibrates on a track record. One high-stakes first call means a single miss teaches the lesson before any track record exists to balance it.

How to answer this, stage by stage

Eight moves. Each one has the words you'd actually say. Read them aloud once, then say them your own way.

1
Scope it to one real product, in your first sentence
Say it like this
"Let me make this concrete so I'm not designing in the abstract. Say it's a resume-ranking tool built into practice software for small dental clinics, and the office manager using it has never touched an AI product before."
Why this works
An interviewer can't grade "it depends." Naming the product and the person in ten seconds buys you permission for the whole rest of the answer, and it stops you drifting into generic advice.
2
Say your structure out loud before you use it
Say it like this
"I'll cover what a first run actually has to accomplish, the one design decision I'd get right, what I'd watch after launch, and what I'd deliberately not fix on day one."
Why this works
Two seconds of structure stops you rambling and tells the interviewer you have a plan. It also lets them steer you if they only care about one part.
3
Reframe what the question is really about
Say it like this
"The instinct is to design a tour, some tooltips, maybe sample data. But a first run isn't for showing off the output. It's teaching this person how much to trust the next hundred outputs. Whatever they learn on day one, they apply for months."
Why this works
Most candidates answer with an onboarding checklist. This is the line that separates you from them, and it makes everything you say afterwards land harder.
4
Give one concrete decision, not a principle
Say it like this
"So the one thing I'd get right is: show reasons, not a confidence score. Not 'Top match, 94 percent.' Instead: 'Matched: front desk experience, insurance billing. Missing: weekend availability.'"
Why this works
A specific decision can be argued with, which is exactly what an interviewer wants to do next. "Build trust gradually" can't be argued with, because it doesn't mean anything yet.
5
Prove it with a failure, in four sentences
Say it like this
"Here's why the score version fails. She hires the top-ranked candidate, and three months in it isn't working out. Her boss asks what happened, and with a score all she can say is 'the tool rated him 94.' That sounds like she outsourced the decision, so next time she just doesn't mention the tool. She keeps using it privately, and never credits it again."
Why this works
A concrete failure persuades where a principle doesn't. It also proves you thought past the happy path, which is most of what a design question is testing.
6
Name a signal that moves before anyone complains
Say it like this
"I'd watch two things. Whether people mention the tool to colleagues, and whether the override rate is climbing. Both of those drift for weeks before anyone files a complaint or cancels."
Why this works
Shows you think past launch day. Most candidates stop at the design and never say how they'd know whether it worked.
7
Say what you would deliberately not fix
Say it like this
"What I wouldn't chase on day one is ranking accuracy. She's skimming the top ten either way, so a slightly imperfect order costs almost nothing. An unexplained miss costs everything."
Why this works
Shows judgment instead of blanket caution. A candidate who wants to fix everything has prioritised nothing, and interviewers notice.
8
Close on the one line, so it's the last thing they hear
Say it like this
"So: reasons over scores. The first run is teaching trust, and a score is a verdict you either swallow or hide. A reason is something you can disagree with out loud."
Why this works
Interviewers remember your first and last lines most clearly. Don't let a strong answer trail off on a tangent.
If you remember one thing Stages 3 and 4 are the answer. Everything else is packaging. If you're short on time, reframe what a first run is for, then name one concrete decision, and stop.

Let's learn

Say we build a tool that reads a stack of resumes and ranks them. It goes into a small dental clinic. The office manager will be the first person there to touch any AI tool at all.

Before we had it, she read all 40 resumes by hand, one evening a year, after the clinic closed. It took her four hours. She was good at it. She could tell in ten seconds if someone had never worked a front desk.

Now she presses a button and gets a ranked list in ten seconds. That is not the interesting part.

Two panels comparing March with zero wrong of 40 picks, and June with one wrong of 40 picks
The number barely moved
Wrong picks out of 40, by hiring round
0
March
1
June
Same measure, two points in time, drawn at true scale on purpose. The bar barely moves. What she does next is the actual story.

This is her first time trusting a machine with a decision like this. She has no feel yet for when to believe it and when not to. That feel has to come from somewhere. If we don't build it in, she builds her own, and we don't get to choose what she learns.

So the real design question is not "how do we make a good ranked list." It is what happens the first time the list is wrong.

We didn't lose a feature. We lost the one thing that made it useful: her trusting it enough to skip the four hours.
Knowledge spark: score vs. reasons A confidence score is one number standing in for the model's whole judgement. 94 percent, 89 percent, done. Reasons are the actual matched and missing points behind that number. A score asks someone to trust a verdict. Reasons let them share one.

What I would leave alone. The ranking itself does not need to be perfect on day one. A slightly imperfect order is fine, because she is going to skim the top ten either way, not blindly hire whoever is first. Chasing ranking accuracy before we have earned her trust in the reasons is solving the wrong problem first.

The long version, for when you want to feel why it matters

You don't need this to answer the question. Read it when you want to understand why the answer is right, not just what it is.

Nadia has run the front desk and the books at a six-chair dental clinic for eleven years. Hiring is not her whole job. It is two weeks a year, usually in January, when someone leaves and she has to find the next one. She is good at it in a way nobody ever taught her. She can read a resume and know, almost without thinking, whether someone has actually sat at a front desk with a phone ringing and a patient waiting, or whether they just think they'd be good at it.

Every January until this year, she did it the same way. She'd wait until the clinic closed, make a cup of tea, and go through the whole stack at the kitchen table. Forty resumes, one evening, four hours. By the end she'd have six names on a sticky note and she trusted every one of them, because she'd read every word herself.

Three boxes: by hand, trusts the tool, checks it all again
The same evening, three ways

This January there is a new tool built into the practice software. It reads the resumes and hands her a ranked list before she has even opened her laptop, the one with the cracked hinge she keeps meaning to replace.

The list is good. Genuinely good. The top three are exactly who she would have picked. She hires the first name on it, a woman named Beth, and Beth is wonderful. In March, Nadia mentions it to Dr. Halvorsen, almost proudly. "The new tool actually found her. Ranked her first."

Then, in June, a second front desk role opens. She runs it again. It ranks a young man named Colton at the top. He interviews well. She hires him.

Colton does not work out. Not dramatically. He is polite and slow to learn the insurance codes and twice sends a patient's file to the wrong fax number. Nothing that gets him fired outright. Just enough that by August, Dr. Halvorsen asks her, gently, what happened with the hiring this time.

And here is the thing Nadia does. She does not say the tool ranked him first. She says she liked his interview. That is not even false. She did like his interview. She just leaves out the part where the ranking is what put him in front of her.

She keeps using the tool. Nobody told her to stop. But from that August on, she reads every resume herself again, the whole stack, the way she used to, and only opens the ranked list afterwards, alone, to see if it agrees with her. If it does, she feels a small private relief. If it doesn't, she never mentions that either.

We did not lose her four hours back. We lost something quieter than that.

We lost the one moment in March where she said, out loud, to the person she answers to, that a machine had done something worth trusting. That sentence does not come back once it stops getting said.

I want to say the problem was that Colton was a bad ranking. But Beth was a good one, and it didn't matter. One good pick and one bad pick, and Nadia landed exactly where a fair person would land after that: private use, public silence.

A dial with many settings versus a switch with two settings
Switch, not dial

So here is the decision, made back when this was still a screen in a design file. We were going to show her a single number next to each name. Ninety-four percent for Beth. Eighty-nine for Colton. Clean, confident, and it tests well in a demo.

Take that number off before it ever reaches her. In its place, three or four lines under each name. For Colton: matched two years of reception experience, comfortable with the phone system. Missing: any mention of insurance billing, which is half the job.

Now walk August forward with that version instead. Colton is struggling with insurance codes. Nadia opens the tool's old note on him, out of curiosity more than anything. It already told her, in June, that billing was missing. It was not that the tool was wrong about Colton. It told her exactly where the gap was. She hired him anyway for the parts he was strong on, which was a real and reasonable choice.

Now when Dr. Halvorsen asks what happened, she has an answer that costs her nothing to say. "The tool flagged that he had no billing background. I hired him anyway because the rest was strong, and that's the part that didn't hold up."

A score asks someone to trust a verdict. A reason lets them share one. We built the version that asks the harder thing first, on the one day she had no practice trusting us at all.

SPARK, in one screen

This is a design question, so the framework is SPARK. (A "what if the error rate doubled" question would use FLIPS instead. Different question shapes get different tools.)

SPARK: situation, payoff, anchor, risk, keep out
SPARK, for design questions
S, situation. Nadia, office manager at a six-chair dental clinic, eleven years in. Hires once a year. Reads 40 resumes by hand over one evening, about four hours.
P, payoff. The habit you want to build: she stops reading all forty herself and skims the tool's top ten instead. That habit is the product.
A, anchor. Show matched and missing reasons under each name. Never a bare confidence score. This is the decision everything else hangs on.
R, risk. The first time a pick goes wrong, she stops mentioning the tool to her boss and keeps using it privately. Credited, then never credited. No middle setting.
K, keep out. Don't chase ranking accuracy on day one. She's skimming the top ten either way, so a slightly imperfect order is safe to leave alone.
Why the anchor and the risk have to match Check them against each other: does showing reasons actually protect against her going quiet? Yes, because a reason is something she can repeat to her boss without sounding like she outsourced the decision. If your anchor doesn't visibly defuse your risk, you have two unrelated halves instead of an answer.

Run SPARK on a different product

A small accounting firm rolls out an AI tool that drafts the first pass of a client's tax return from their uploaded documents. Same framework, completely different risk.

Marcus before and after one bad call
Same framework, different risk

S. Marcus, a solo bookkeeper with about ninety small-business clients, all due the same week every April.
P. The habit you want: he stops re-entering every number by hand and starts reviewing the draft instead.
A. Flag which documents are incomplete or unusual for that client at intake, before the draft is generated, so a messy client visibly asks for more attention from the start.
R. If one bad draft slips through on a messy client and it costs him a call from the tax office, he doesn't review harder. He starts running only his simplest clients through the tool and quietly types the messy ones from scratch. Full trust at the top of the client list, none at the bottom.
K. Don't try to handle the messy clients automatically on day one. Flagging them is enough, and it's the part that keeps him from writing them off the tool entirely.

Swap the trigger and it still runs

  • Speed: the tool starts taking eight seconds instead of being instant. His habit doesn't change, because eight seconds still beats typing.
  • Cost: the firm starts charging per return processed. He saves it for higher-fee clients and hand-does the small ones. That's the substitution flip wearing the scope flip's clothes.
  • The model gets better: draft accuracy jumps. He stops skimming the "everything matched" cases, and the one time a document is subtly wrong with nothing flagged, it ships inside a filed return.

Where people run it wrong

  • Adding a big red warning banner on every draft. Everyone tunes out a banner that fires every time.
  • Putting the missing-item flags in a separate report the person has to remember to open. If it isn't on the same screen as the draft, it doesn't exist.
  • Writing the flags in the tool's language instead of the person's. "Schema mismatch on Form 4562" tells a bookkeeper nothing. "No mileage log attached" does.

If you're asked this cold

Buy yourself ten seconds by naming the person out loud first. "Let me put someone in this. Say it's a solo bookkeeper doing ninety returns a year." Naming the person is not stalling. It's stage 1 of the answer, and it buys you the time to find the real flip instead of reaching for the first one you think of.

Flashcards (click a card to flip it)

1 · THE OPENING MOVE
What's the first thing you say when asked to design an experience?
Tap to flip
ANSWER
Scope it to one concrete product and user, in a single sentence, so the answer can't drift into generic advice.
2 · THE REFRAME
What is a first-run experience actually for?
Tap to flip
ANSWER
Teaching the user how much to trust the next hundred results. Not showing off the output. Whatever they learn on day one, they apply for months.
3 · THE DECISION
What's the one concrete design decision in this answer?
Tap to flip
ANSWER
Show reasons (matched and missing), never a bare confidence score.
4 · THE RISK
SPARK's R: what's the two-setting switch this design protects against?
Tap to flip
ANSWER
Crediting the tool out loud, or never mentioning it at all. Nothing in between. She keeps using it either way.
5 · THE PROOF
How do you prove a design decision in an interview?
Tap to flip
ANSWER
Show the concrete failure that happens without it, in about four sentences. A principle doesn't persuade; a specific bad Tuesday does.
6 · THE SIGNAL
What would you measure, and why those?
Tap to flip
ANSWER
Whether users mention the tool to colleagues, and whether override rate climbs. Both drift for weeks before anyone complains or churns.
7 · THE RESTRAINT
What do you say you would NOT fix, and why does that help you?
Tap to flip
ANSWER
Ranking accuracy on day one. Naming something you'd leave alone shows judgment. A candidate who fixes everything has prioritised nothing.
8 · THE CLOSE
How should the answer end?
Tap to flip
ANSWER
One line restating the decision and why. "Reasons over scores. A score is a verdict you swallow or hide. A reason is something you can disagree with out loud."

Check yourself Score: 0 / 0

Multiple choice
1. You're asked this question cold. What's your first move?
  • A. List the onboarding screens you'd build.
  • B. Name one concrete product and user, so the answer can't stay abstract.
  • C. Ask the interviewer which product they mean.
  • D. Explain that it depends on the user segment.
Show hint
Which one gives the interviewer something they can actually grade?
Show answer
B. Scoping it yourself is stage 1. Asking the interviewer to pick (option C) hands back the decision and reads as stalling. "It depends" is the weakest possible opening.
True or false
2. True or false: Nadia stopped trusting the tool completely after Colton didn't work out.
  • True
  • False
Show hint
She kept using it. Look at what actually changed.
Show answer
False. She used it every round after that. What changed was that she stopped telling anyone, and started reading every resume herself first as a private check. Trust didn't disappear. Attribution did.
Fill in the blank
3. A first run isn't for showing off the output. It's for teaching the user ______.
Show hint
It's about the next hundred results, not this one.
Show answer
How much to trust the next hundred outputs. This is stage 3, the reframe, and it's the line that separates a good answer from an onboarding checklist.
Multiple choice
4. Which of these is a real answer to "what's the one design decision," and which is a principle wearing an answer's clothes?
  • A. "I'd make sure we build trust gradually with the user."
  • B. "Instead of 'Top match, 94 percent,' show 'Matched: front desk experience. Missing: insurance billing.'"
  • C. "I'd focus on transparency and explainability."
  • D. "I'd run user research before deciding."
Show hint
Which one could an interviewer actually disagree with?
Show answer
B. A and C sound sophisticated but can't be argued with, which means they say nothing. D defers the question. Only B gives the interviewer something specific to push on, and that's the point.
Short answer
5. Name something you would deliberately NOT fix in this product, and say why that answer helps you.
Show hint
Think about a mistake that's cheap and obvious right away, versus one that's hidden and expensive later.
Show answer
Model answer: "I wouldn't chase ranking accuracy on day one. She's skimming the top ten either way, so a slightly imperfect order costs almost nothing. An unexplained miss costs everything." Naming a non-priority shows judgment. A candidate who wants to fix everything has prioritised nothing.
Short answer, apply it yourself
6. Pick a product you use yourself. Run stages 3 and 4 on it out loud: what is its first run really teaching, and what one decision would you make?
Show hint
Think of something you've stopped double-checking because it's always been right. Maps, spell check, a payment app that remembers your card.
Show answer
Model answer: "Take GPS directions. The first run is teaching you whether you still need to glance at street signs. Most people stop within a week. The one decision I'd make: when it's routing you an unusual way, say why in four words ('avoiding a closure ahead'), because without that, the first time it sends you somewhere strange you stop trusting it on unfamiliar roads, which is exactly where it was most useful."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more