CaseAdvancedDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #21
What onboarding content would you build for a feature with a narrow competence area?
GUARD the product is Keyhole, a home-services booking marketplace, and Nook is the AI assistant built into its booking chat
Keyhole connects homeowners with local contractors for repairs. Nook is the chat assistant embedded in every booking, but it only really knows three things: rescheduling a visit, cancellation windows, and arrival times. Halyna Petrenko booked a plumber through Keyhole after a pipe under her sink started leaking.
The direct answer
Build onboarding that draws a visible line, not a written one. The first message Nook ever sends should name exactly what it can help with, in plain terms, and anything outside that line should look and feel different the moment it comes up, not just get a hedge tacked onto the end of a confident-sounding answer. A homeowner should never have to guess whether they've wandered past what Nook actually knows.
Do this, in order
Show a plain "I can help with" list the moment someone starts chatting with Nook.Why: a boundary nobody sees at the start doesn't exist when someone actually crosses it later.
Make out-of-scope replies look visually different from in-scope ones, not just worded differently.Why: in a stressful moment, people react to tone and shape before they read the words carefully.
Route out-of-scope questions to a human immediately, with no attempted answer first.Why: a confident guess on a damage claim costs far more than the extra minutes a human reply takes.
Track a boundary-miss rate: how often Nook engages with an out-of-scope topic before handing off.Why: this is the number that catches the failure in production, before it becomes a support complaint.
Give the contractor side the same visible boundary, not just the homeowner.Why: a contractor disputing a damage claim deserves the same clear line, since Nook can't arbitrate that either.
How to answer this, stage by stage
Nobody's grading whether you can list three onboarding screens. They're grading whether you can name who gets hurt when a narrow tool sounds like a broad one.
Stage 1
Scope it to one real assistant
Say it like this
"I'll answer this for Nook, Keyhole's booking assistant, which only really covers rescheduling, cancellation, and arrival windows."
Why this works
Grounds "narrow competence" in one specific, real boundary instead of a vague notion of limitations.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's affected. Unequal, where the harm lands hardest. Ability to contest, who can't push back. Reduce, the actual design change. Detect, how I'd know in production."
Why this works
Signals a method built for risk questions, not a generic list of onboarding tips.
Stage 3
Name who's affected, and how unevenly
Say it like this
"Anyone can ask Nook a question outside its competence, but it's the homeowner in the middle of a stressful repair, not someone calmly browsing, who's least likely to double-check a confident-sounding wrong answer."
Why this works
Names the specific person who bears the cost, not a generic "users" statement.
Ask who can't push back
Say it like this
"Right now, nothing tells Halyna she's asked Nook something it isn't actually built to answer. Same calm tone, same confident reply, whether it's a reschedule or a damage claim."
Why this works
This is GUARD's hardest step, and the one that turns a vague worry into a specific design gap.
Stage 5
Give the actual design change
Say it like this
"Nook's first message names exactly what it can help with. And the moment someone asks about a damage claim or a price dispute, the reply looks different, a different color, a clear 'connecting you to a person,' not a normal chat bubble with a hedge at the bottom."
Why this works
A concrete onboarding content decision, specific enough to build tomorrow.
Stage 6
Say how you'd detect it in production
Say it like this
"I'd track a boundary-miss rate: how often Nook engages with an out-of-scope topic, damage claims, price disputes, before handing off, instead of waiting for a support complaint to tell us."
Why this works
Shows you'd catch the failure yourself, not rely on the person it hurt to report it.
Stage 7
Name the rejected alternative, and close
Say it like this
"I considered letting Nook attempt an answer with a hedge at the end, 'I'm not fully sure, you may want to double-check.' I rejected it, because in a stressful moment people react to the confident part first and the hedge barely registers."
Why this works
Proves this was a real choice, and shows exactly why the softer option fails the person it's meant to protect.
Let's learn
How would you handle a tool that's genuinely good at three things and dangerously confident about everything else?
Nook is the AI assistant built into Keyhole's booking chat. It handles rescheduling, cancellation windows, and confirming arrival times.
Knowledge spark: what's a narrow competence area?
A part of a problem an AI system genuinely handles well, surrounded by a much larger area where it wasn't trained to help at all. The danger isn't the narrow part. It's that the tool often sounds equally confident everywhere.
Before onboarding addressed this at all, Nook answered every question in the same friendly, certain tone, whether it was "can I move Tuesday to Thursday" or "the plumber flooded my kitchen, what do I do."
Confident wrong-scope answers given, before and after the redesign
Both out-of-scope topics used to get a confident direct answer more often than not. After the redesign, nearly all of them route straight to a person.
At its worst: Halyna, mid-crisis with water pooling under her sink, asks Nook whether Keyhole will cover the water damage. Nook answers in the same calm, certain voice it uses for a reschedule, and she believes it, because nothing about the reply told her she'd asked the one kind of question Nook was never built to handle.
The decision I would take back
We gave Nook one undifferentiated chat interface, the same tone and shape for every reply, in-scope or not. That made sense when Nook only handled simple scheduling and rarely got asked anything else. It stopped making sense once homeowners started bringing real problems to the one chat box they already had open.
What I would leave alone: for the three things Nook actually does well, rescheduling, cancellations, arrival windows, the confident, instant tone is exactly right. Slowing those down with extra caveats would just make the tool worse at what it's genuinely good at.
Nook was never wrong about the reschedule. It was wrong about sounding exactly the same when Halyna asked something it had no business answering.
The lesson: a narrow competence area isn't dangerous because the tool is bad. It's dangerous because nothing marks where "good" ends, and the person least equipped to notice the edge is usually the one standing right on it.
Now here is the same thing as a story
The short version above is what you'd say defending this onboarding decision to Keyhole's trust and safety team. Read this one for how the gap actually reached Halyna.
For eight months, Nook did exactly what it was built for. Rescheduling took ten seconds. Cancellation windows were always right there in the first reply.
Keyhole decided what Nook would and wouldn't attempt to answer. Halyna never got a say in that, and had no way to see where the line even was.
Then a pipe under her kitchen sink split at ten at night. Water was pooling across the floor. She opened the same Keyhole chat she'd used to book the plumber in the first place and asked Nook whether the marketplace would cover the damage.
Under the old design, a damage claim question got answered directly, with no escalation offered at all. The new design routes it to a person and shows the handoff happening.
Nook answered in its usual calm, certain tone: something reassuring about most water damage being covered under standard service terms. It sounded exactly like every other answer Nook had ever given her. She didn't call a specialist that night. She didn't take extra photos right away. She trusted the tone, because the tone had never once been wrong before.
This is the entire onboarding content. Not a manual, not a settings page. Four lines, visible before the first real question ever gets typed.
It took two days, and a much harder conversation with actual support staff, to sort out what Nook's reassuring answer had gotten wrong. The damage claim itself resolved fine in the end. What didn't resolve as easily was Halyna's sense that the chat box she'd trusted for eight months had quietly let her down exactly when it mattered most.
The redesigned reply doesn't just say something different. It looks different, the instant the topic crosses the line.
The redesigned Nook now shows a plain list of what it covers the moment anyone opens the chat, and the instant a damage claim or price dispute comes up mid-conversation, the reply changes shape entirely: a different color, a short "connecting you to a person now," and an actual human joining within minutes.
The boundary shows up twice now: once at the very start, and again, visibly, the moment it actually matters.
The old design treated every Nook reply as interchangeable, one voice for everything. The new one treats "in scope" and "out of scope" as two visibly different kinds of moment, because they are.
We built the single undifferentiated chat voice because it felt simpler and more consistent, one Nook, one personality, one tone. It took watching Halyna trust a wrong answer at ten at night to see that consistency in tone and consistency in trustworthiness were never actually the same thing.
GUARD, and who never got a sayNot a policy document. A visible line, drawn before anyone has to guess where it is.
G
Groups. Who's affected.
Keyhole's team, who decided what Nook would attempt to answer, and Halyna, who had no way to see that decision.
Names the operator and the subject on opposite sides of the same design choice.
U
Unequal. Where harm lands.
A homeowner mid-crisis, not calmly browsing, is the one least likely to double-check a confident wrong answer.
Shows the harm isn't spread evenly; it concentrates on exactly the person with the least room to double-check.
A
Ability to contest.
Nothing told Halyna she'd asked Nook something outside its competence. Same tone, same confidence, every time.
The hardest step, and the one that turns a vague worry into an actual design gap worth naming.
R
Reduce. The design change.
A visible "I can help with" list up front, and a visually distinct, instant human handoff for anything outside it.
A real product decision, not a training document or a policy nobody reads.
D
Detect. How you'd know.
A boundary-miss rate: how often Nook engages with an out-of-scope topic before handing off, tracked before a complaint ever arrives.
Catches the failure in production instead of waiting for the person it hurt to report it.
Homeowner escalation complaints, weeks before and after the redesign
Complaints don't drop to zero instantly, but the line bends hard the same week the visible boundary ships, and keeps falling from there.
The recap, one line per letter: groups is Keyhole's team and Halyna, unequal is the harm concentrating on homeowners mid-crisis, ability to contest is the missing signal that told her nothing had changed, reduce is the visible boundary and instant human handoff, and detect is the boundary-miss rate tracked before any complaint arrives.
And if you want to be sure it really works, try it somewhere elseSame five letters, a payroll platform instead of a home-services marketplace. A different domain, and the narrow competence line moves to pay questions instead of home repairs.
Payspine is a payroll platform used by mid-size employers. Its onboarding chatbot answers questions about when pay arrives and current tax withholding, but it was never built to handle compensation disputes or benefits enrollment. Corentin Osei, an HR generalist, rolled the chatbot out to new hires without a clear line drawn around what it actually covers.
Mapped onto GUARD: groups is the HR team that built the chatbot, and new hires who ask it real questions about their pay. Unequal is that a new hire disputing a paycheck error, often someone with the least standing to push back at a new job, is exactly who's least likely to escalate confidently if the bot answers smoothly instead of routing them to HR. Ability to contest is that nothing in the chatbot's tone signals when it's stepped outside pay-timing questions into a dispute it can't actually resolve. Reduce is a visible "I can help with" list on first use, plus an immediate, visually distinct handoff to HR for anything about disputes or benefits. Detect is tracking how often the bot engages with dispute-shaped language before routing away.
Two branches the bot should answer. Two it should never try to, and routing is the whole design decision.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "draw the boundary visibly, and make crossing it look different, not just read differently," and stop.
Cost: there's no engineering time this quarter for a custom visual treatment. Say so, and start with a plain, differently-labeled message type, cheaper, still visually distinct from a normal reply.
The model gets better, for real: even as Nook's in-scope answers get more accurate, the boundary still matters, since a broader, more confident-sounding tool is if anything more likely to be trusted past its actual edge.
Where people run it wrong.
They document the tool's limits in a help article nobody reads before the moment they actually need it.
They let the assistant attempt an answer with a hedge at the end, which barely registers against a confident first impression.
They wait for a support complaint to reveal the gap instead of tracking a boundary-miss rate directly.
How to use it live. When someone asks what onboarding content a narrow-competence feature needs, ask yourself first: what does it look like, right now, the moment someone asks it something it can't actually answer? If the answer is "exactly the same as everything else," that's the gap to close.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "what onboarding content for a feature with a narrow competence area"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It's a risk question, since the danger is a user trusting the tool past its actual competence.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Halyna Petrenko, a homeowner who asked Nook about water damage coverage during a plumbing emergency.
3 · THE UNEQUAL HARM
Where does the harm land hardest, and why?
Tap to flip
ANSWER
On homeowners mid-crisis, not calmly browsing. They're the least likely to double-check a confident-sounding wrong answer in the moment.
4 · THE ABILITY TO CONTEST
What's missing that would let someone push back?
Tap to flip
ANSWER
A visible signal that a question has crossed outside Nook's competence. Without it, every reply looks equally trustworthy, in scope or not.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Giving Nook one undifferentiated chat voice for every reply, in scope or not. It made sense while Nook was rarely asked anything outside scheduling.
6 · THE NUMBER
Fill in the blank: before the redesign, damage claim questions got a confident direct answer ___ percent of the time.
Tap to flip
ANSWER
68 percent. After the redesign, that dropped to about 3 percent, with the rest routed straight to a person.
7 · THE REJECTED OPTION
What alternative design was considered and rejected?
Tap to flip
ANSWER
Letting Nook attempt an answer with a hedge at the end. Rejected because people react to the confident part first, and the hedge barely registers in a stressful moment.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and where does its competence boundary sit?
Tap to flip
ANSWER
Payspine, a payroll platform. Its bot should answer pay-timing and withholding questions, but must route compensation disputes and benefits enrollment straight to HR.
Check yourself Score: 0 / 0
Multiple choice
1. Why does this answer reject letting Nook attempt an answer with a hedge at the end?
A. Hedged answers take longer to generate.
B. In a stressful moment, people react to the confident part of an answer first, and a hedge at the end barely registers.
C. Hedges are against Keyhole's brand voice guidelines.
D. Hedged answers are harder to track in analytics.
Show hint
Look at stage 7 of the walkthrough.
Show answer
B. A hedge tacked onto a confident-sounding answer doesn't change how the answer actually lands with someone under stress.
True or false
2. True or false: this answer recommends adding a hedge or disclaimer to Nook's in-scope answers, like rescheduling.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. The confident, instant tone stays exactly as-is for the three things Nook actually does well. Only out-of-scope replies change shape.
Fill in the blank
3. Fill in the blank: the metric this answer uses to catch the problem in production is called the ___ rate.
Show hint
Look at the "detect" step in the GUARD recap.
Show answer
Boundary-miss. It tracks how often Nook engages with an out-of-scope topic before handing off, catching the gap before a complaint does.
Short answer, name who can't push back
4. Who never gets to push back in the original design, and what past decision caused that?
Show hint
Look at the "ability to contest" step.
Show answer
Model answer: Halyna, and anyone else who asked Nook an out-of-scope question, had no signal they'd crossed a line. The decision was giving Nook one undifferentiated tone for every kind of reply.
Short answer, where it wouldn't matter
5. Name a part of Nook's job where this whole redesign genuinely doesn't need to change anything.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Rescheduling, cancellation windows, and arrival times. Nook is genuinely good at these, and the confident, instant tone is exactly right there.
Short answer, apply it yourself
6. Pick a product you use yourself. Where's the edge of what its AI feature actually knows, and how would you know if you'd crossed it?
Show hint
Think of a chatbot, a support assistant, or a smart-reply feature that answers some questions confidently and others not so much.
Show answer
Model answer: Many banking chatbots answer balance and transaction questions well but shouldn't be trusted on fraud disputes. Most give no visible signal when a question has crossed that line, the exact gap this answer is about.
Before you close the answer
Why this works
Tests whether you understand that a narrow-competence feature's real danger is an even, confident tone across its whole surface, and whether you can design onboarding content that makes the edge of that competence visible in the moment, not just documented somewhere.
Follow-up traps
"Won't a visible handoff message just feel like a worse experience than a direct answer?" Response: it feels slower in the moment, but it's the honest version, and the alternative is a confident wrong answer that costs far more trust once it's discovered.
"What if the boundary itself is unclear, not every question is obviously in or out of scope?" Response: route borderline cases to a human by default rather than guessing; the cost of an unnecessary handoff is much smaller than the cost of a wrong confident answer.
If pressed
Keyhole's real boundary-miss detector runs a lightweight topic classifier against the conversation before Nook replies, so an out-of-scope topic gets caught and routed before Nook ever generates an answer, not after the fact.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.