ConceptIntermediateDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #3
Explain the role of example prompts in onboarding and their downside.
PICK the product is PaceForge, an AI training-plan chatbot for age-group triathletes
Zanele Buthelezi coaches a squad of thirty age-group triathletes. Between in-person sessions, her athletes use PaceForge, a chatbot that answers training questions on their phone. The first screen shows five example questions to tap.
The direct answer
Example prompts earn their place because a blank box is scary and five taps get someone talking. The downside is that most people never leave those five. Ship them, but rotate one out every few weeks and always include one genuinely odd example, so the five read as a starting point instead of the whole menu.
Do this, in order
Keep example prompts, but design them to show range, not just the five safest asks.Why: a narrow set teaches people the tool's edges are narrower than they really are.
Rotate at least one example regularly and include one deliberately unusual one.Why: a fixed set of five becomes a ceiling nobody questions after week one.
Add one plain line: "these are a start, ask anything training-related."Why: without it, five examples silently read as five categories.
Watch the share of questions that resemble the shown examples versus genuinely new ones.Why: that ratio is the real signal of whether anchoring is happening.
Don't remove example prompts entirely to fix this.Why: a blank box has its own cost, most people never type a first message at all.
How to answer this, stage by stage
Six stages, because the question is narrower than a full design case. Say each one out loud, in order.
Stage 1
Scope it to one real product
Say it like this
"I'll answer this for PaceForge, a training-plan chatbot age-group triathletes use between coaching sessions."
Why this works
Turns "explain the role of example prompts" from a definition into a real design question.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my pick up front. Impact, who feels each error. Cost asymmetry, which one is actually expensive. Kill criteria, what would change my mind."
Why this works
Signals a structured tradeoff answer, not a one-sided opinion.
Stage 3
Take your position, before any reasoning
Say it like this
"Keep the example prompts. A blank box is the bigger risk on day one. But design the five to show range, not just the safest, most obvious asks."
Why this works
Interviewers are testing whether you can commit, not whether you can list pros and cons forever.
Stage 4
Name the cost asymmetry
Say it like this
"Someone typing a question that doesn't match the five examples is a cheap, visible, one-off event. Someone quietly never asking anything outside the five, for months, is hidden and it caps the entire product."
Why this works
This is the heart of PICK: naming which error actually costs you something.
Stage 5
Prove it with a number
Say it like this
"After sixty days, 74 percent of all questions were close copies of the original five examples, and 41 percent of athletes said, in a survey, that they didn't know PaceForge could help with anything else."
Why this works
A real number turns "anchoring is a risk" into "anchoring is already happening."
Stage 6
Close on the kill criteria
Say it like this
"If, after rotating examples and adding an edge case, the share of novel questions still doesn't move, I'd stop blaming the examples and look at whether the product can actually answer those questions well."
Why this works
Shows what evidence would change your mind, not just stubborn confidence.
Let's learn
What happens the first time a new athlete opens PaceForge and sees a blank message box with nothing in it?
Most people freeze, or close the app. That's why PaceForge, like most chat products, shows five example prompts on the first screen: "What should I eat before a long run?" "How do I taper before a race?" and three more like them. Tap one, and you're talking to the tool within two seconds.
For the first few weeks, this worked exactly as intended. New athletes tapped an example, got a good answer, and kept using the app. Usage climbed fast.
Where questions went, 60 days after onboarding, out of 1,000 athletes
Only 6 percent of questions touched topics PaceForge actually handles well but never showed on day one, like altitude taper plans or injury-flare adjustments.
Knowledge spark: what's anchoring?
A person's first look at a small sample of something shapes what they believe the whole thing can do, even after they've used it for months. Five example prompts aren't just a starting point, they quietly become the definition of "what this app is for."
At its worst: a 41 percent survey of athletes said they genuinely didn't know PaceForge could help with things like altitude adjustment or coming back from an injury flare. Some of them asked Zanele those exact questions by text message instead, the same questions the chatbot could have answered in seconds.
The decision I would take back
We removed the full "browse everything it can help with" menu early on and replaced it with five example prompts, since the full menu looked cluttered and slowed the first five seconds. That was fine when PaceForge did a handful of things well. It stopped being fine once it could genuinely help with a much wider range of questions than five bullets could ever show.
What I would leave alone: for a brand-new athlete's very first message, a blank box really is worse than five examples. The fix isn't removing the examples, it's making sure they don't quietly become a ceiling.
The five examples were never the whole product. They were the only five things anyone believed the product was for.
The lesson: onboarding examples do two jobs at once, whether you plan for it or not. They get someone started, and they teach a boundary. If you only design for the first job, the second one happens anyway, just badly.
Now here is the same thing as a story
The short version above is the tradeoff, argued. This one is how Zanele actually noticed it.
Zanele can read a runner's stride off a phone video before the athlete even feels what's wrong. She's coached age-group triathletes for eleven years, and she signed her whole squad up for PaceForge the month it launched, mostly so she'd stop getting the same taper questions by text at 10pm.
A menu with five dishes on it reads as "this is what the kitchen makes," even when the kitchen makes forty things.
It worked. Athletes tapped one of the five prompts, got a good answer, and came back the next day. For a while, Zanele's phone got quieter.
Nobody decided to stop trying new questions. It just never occurred to most athletes that there was anything else to ask.
Then one of the other coaches on her team made an offhand remark during a staff meeting: "Have you noticed everyone asks it the exact same five things? It's like they don't know it does anything else."
Zanele pulled the chat logs that week. Seventy-four percent of every question, across thirty athletes and sixty days, was a close copy of one of the original five examples.
The top right corner is the real loss: questions PaceForge answers well, that almost nobody thought to ask.
We did not build a narrow product. We built a product that looked narrow, because the only five doors we ever showed anyone were the same five doors, month after month.
One error costs a person nothing. The other one costs the whole product its own range, silently.
The fix wasn't removing the five examples. Zanele's squad still needed a fast way to start a conversation. The fix was making sure the five stopped acting like a ceiling.
The second callout, one weird edge case, did more work than the other three combined.
Now the five examples rotate, one swaps out every few weeks, and one has always been something a little unusual, an altitude question, an injury-flare question, a mental-fatigue question. A single line under the five reads: "these are a start, ask anything training-related." Novel questions climbed from 6 percent to 22 percent of all traffic within a month.
I built the narrow five because it was fast to ship and easy to test. It took a colleague's offhand remark in a staff meeting, not a support ticket or a churn number, to notice that the product's actual range and the range athletes believed it had were two very different things.
PICK, in one screenNot a preference. PICK is what forces the tradeoff to name which error you're actually protecting against.
P
Position. The pick, before any reasoning.
Keep example prompts. A blank box costs more on day one than a narrow first impression.
States the commitment first, before the argument for it.
I
Impact. Who feels each error.
A blank box costs a new athlete a stalled first session. A narrow example set costs an experienced athlete months of never asking a question the tool could answer.
Names both sides in real terms, not just "pros and cons."
C
Cost asymmetry. Which one is actually expensive.
The blank-box cost is visible and short-lived. The narrow-example cost is invisible and it compounds for months, 74 percent of traffic stuck on five topics.
The heart of PICK: says plainly which error you're optimizing against.
K
Kill criteria. What would change the pick.
If rotating examples and naming "ask anything" doesn't move the novel-question share, the problem isn't the examples, it's whether the product genuinely handles those questions well.
Separates a confident answer from a stubborn one.
Novel-question share, before and after the rotating examples
The fixed five never moved on their own. Rotation, plus the "ask anything" line, is what actually changed the number.
The recap, one line per letter: position is keep the examples, impact is a stalled first message versus months of unseen range, cost asymmetry is the narrow-example cost since it's invisible and it compounds, and kill criteria is watching whether rotation actually moves the novel-question share.
And if you want to be sure it really works, try it somewhere elseSame four letters, a washing-machine repair chat instead of a training app. This time the five examples hide a completely different capability.
Grzegorz Wilk repairs household appliances and uses FixLine, an AI chat assistant, to help diagnose faults before he drives out with parts. FixLine's first screen shows five example questions: a leaking door seal, a drum that won't spin, an error code, and two more like them.
Mapped onto PICK: position is keep the five examples, since a technician standing in someone's kitchen needs a fast answer, not a blank box. Impact is a technician typing a rare fault by hand, which costs thirty extra seconds, versus a technician who quietly stops trusting FixLine for anything beyond the five shown faults, which costs a wasted truck trip when the real issue turns out to be something FixLine could have flagged. Cost asymmetry favors the second one, it's the one nobody notices happening. Kill criteria: if truck trips for "no fault found" don't drop after adding a wider set of examples, the real problem is diagnostic accuracy, not the prompt list.
The second branch is where FixLine quietly earns or loses trust, a real fault the five examples never mentioned.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "keep the examples, rotate them, name the edge," and stop.
Cost: there's no budget to build a rotation system this quarter. Say so honestly, and start by manually swapping one example each month, since even a manual fix beats a static one.
The model gets better, for real: if PaceForge's answer quality improves across the board, that's still not a reason to leave the same five examples in place forever, a better model with a narrower perceived range is still a narrower product.
Where people run it wrong.
They treat example prompts as a one-time onboarding decision instead of a setting that needs revisiting.
They assume low usage of a feature means nobody wants it, when it might mean nobody knows it exists.
They remove example prompts entirely to fix anchoring, trading one real cost for a worse one.
How to use it live. When someone asks about example prompts, ask yourself one question first: what percentage of real queries still look like the original examples, months later. Say that number, real or estimated, before you say anything about design philosophy.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "explain the role of example prompts and their downside"?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Cost asymmetry is what turns "there's a downside" into a real design decision.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zanele Buthelezi, an eleven-year triathlon coach whose squad uses PaceForge between in-person sessions.
3 · THE HABIT
What did athletes stop doing without realizing it?
Tap to flip
ANSWER
They stopped asking anything that didn't resemble one of the five example prompts, treating those five as the whole product instead of a starting point.
4 · THE ASYMMETRY
What's the two-sided cost this answer is built on?
Tap to flip
ANSWER
Cheap: someone types a question that doesn't match the examples, costs thirty seconds. Hidden: someone never tries, for months, and the whole product looks narrower than it is.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Removing the full capability menu in favor of five fixed examples, since it made the screen faster while PaceForge only did a few things, and stopped making sense as it grew.
6 · THE NUMBER
Fill in the blank: after sixty days, ___ percent of all questions were close copies of the original five examples.
Same squad, rotating examples added. What changes?
Tap to flip
ANSWER
Novel-question share climbs from 6 percent to 22 percent within a month, instead of staying flat forever under the fixed five.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the real risk there?
Tap to flip
ANSWER
FixLine, an appliance-repair chat assistant. There, the risk is a wasted truck trip when a real fault falls outside the five shown examples.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: sixty days in, ___ percent of PaceForge questions were close copies of the original five example prompts.
Show hint
Look at the grouped bar chart of where questions went.
Show answer
74 percent. Only 6 percent were genuinely novel topics PaceForge already handled well.
Multiple choice
2. According to this answer, what's the actual fix for anchoring on example prompts?
A. Remove all example prompts and start every chat with a blank box.
B. Rotate the examples, include a deliberately unusual one, and say plainly that they're just a start.
C. Add a tenth example prompt so there are more choices.
D. Replace text examples with a video tutorial.
Show hint
Look at the priority list and the P step.
Show answer
B. The line chart shows novel-question share only moved once examples started rotating and the "ask anything" line was added.
True or false
3. True or false: this answer recommends removing example prompts entirely because they cause anchoring.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. A blank box has its own real cost. The fix keeps the examples but stops them from acting like a ceiling.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Removing the full capability menu for five fixed examples. It made sense while PaceForge only did a handful of things well, and stopped making sense as its range grew.
Short answer, apply it yourself
5. Pick an AI product you use yourself. What are its example prompts, and is there something you only found out later it could do that those examples never hinted at?
Show hint
Think about the suggestion chips or sample questions you saw the first time you opened a chat assistant.
Show answer
Model answer: Most people can name a capability, drafting an email in a specific tone, analyzing a spreadsheet, that they only discovered by accident, well after the first five suggestions had already shaped what they thought the tool was for.
Before you close the answer
Why this works
Tests whether you can hold two true things at once, that examples help and that they quietly narrow perceived capability, and pick a real design response instead of arguing only one side.
Follow-up traps
"Why not just add more example prompts instead of rotating them?" Response: more fixed examples still become a fixed ceiling eventually, just a slightly higher one; rotation is what keeps the set from calcifying at all.
"Isn't 74 percent close-copy usage actually a good sign that the examples are working?" Response: only if the goal was narrow usage. The 41 percent of athletes who didn't know PaceForge could help with other things says the examples were working too well, at the cost of everything else.
If pressed
PaceForge's real fix logs which example, if any, a session started from, so the team can tell a truly novel first message apart from a slightly reworded copy of an example, which the raw 74 percent number alone couldn't distinguish.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.