CaseIntermediateDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #5

How would you set expectations about errors during onboarding?

FLIPS the product is StrokeLine, a poolside AI kiosk that scores swimmers' stroke technique

Say we build a kiosk. A swimmer finishes a lap, StrokeLine reviews the recorded footage, and within seconds it scores their stroke and flags anything that looks off.

The direct answer
Tell people, on day one, the specific way it fails, not just that it sometimes fails. Say plainly: this tool is built against common technique patterns, so a genuinely unusual but effective stroke can get flagged as a problem it isn't. Then build a standing rule that catches that exact failure, instead of a general warning nobody remembers by week three.
Do this, in order
  1. Name the specific failure shape during onboarding, not a vague "it isn't perfect."Why: a vague warning gives nobody anything to actually watch for.
  2. Build a standing rule that catches that exact failure automatically.Why: a warning fades from memory in weeks; a rule doesn't.
  3. Never claim a report is "validated" without saying validated against what.Why: an unqualified claim gives trust nothing to calibrate against later.
  4. Watch how often the person double-checks the tool's own flags, over time.Why: a falling check rate is the real leading signal, weeks before anything visibly breaks.
  5. Don't set error expectations by adding more review meetings.Why: a scheduled review catches problems on its own clock, not the day they start.

How to answer this, stage by stage

Seven stages. Say each one out loud, in order, and you've got the whole answer.

Stage 1
Scope it to one real product
Say it like this
"I'll answer this for StrokeLine, a poolside kiosk that scores a swimmer's stroke technique from recorded footage, and for Bekah, the coach who reads its reports every day."
Why this works
Turns an abstract "how do you set expectations" question into one real report a real person reads.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a structured answer instead of a list of onboarding best practices.
Stage 3
Reframe the question
Say it like this
"Setting error expectations isn't a disclaimer you say once. It's the difference between someone knowing how the tool fails and someone treating a good streak as proof it never will."
Why this works
Separates a real answer from a generic "communicate clearly" line.
Stage 4
Give the one decision
Say it like this
"I'd tell every coach on day one: StrokeLine is built against common technique patterns. A stroke that's genuinely different, not wrong, can still get flagged. And I'd build a rule: any swimmer flagged three times gets an automatic film pull, no coach has to remember to ask for one."
Why this works
Concrete, and it matches the direct answer word for word.
Stage 5
Prove it with a failure
Say it like this
"Bekah stopped pulling film herself after StrokeLine agreed with her for months. A quarterly audit found two swimmers flagged for a 'flaw' that was really just an unusual, effective stroke, one of which she'd already started correcting in practice."
Why this works
Shows the real cost of unset expectations, not a hypothetical one.
Stage 6
Say what you'd measure
Say it like this
"I'd track what share of flagged swimmers get an actual film check each week. If that number is falling toward zero, the tool looks fine and the safety net is already gone."
Why this works
Shows you'd catch the drift weeks before a scheduled audit would.
Stage 7
Close on the one line
Say it like this
"Name the specific way it fails, and build a rule that catches that failure automatically. That's the whole answer."
Why this works
Restates the decision plainly, ready for any follow-up.

Let's learn

Say we build a kiosk. A swimmer finishes a lap, StrokeLine reviews the footage, and within seconds it scores their stroke and flags anything that looks off, elbow position, kick width, breathing rhythm.

Before StrokeLine, Bekah reviewed film for every swimmer on her roster, every week, about 30 swimmers, roughly nine hours of footage. She caught real issues, but it took her whole Sunday.

With StrokeLine, review takes minutes instead of hours, and for months, it agreed with what Bekah already saw. So she started pulling film herself less often, from nine swimmers a week down to two.

The extra speed was never the problem. The problem was what happened the one time StrokeLine's flag was wrong, and nobody was still checking.

Share of flagged swimmers Bekah actually double-checked on film, by week
100% 50% 0 10% by week 8 Week 1: 90% audit, week 10
The number that mattered was falling weeks before the audit ever ran. Nobody was watching it, because nobody knew it was the number to watch.
Knowledge spark: what's a common technique pattern? A model like StrokeLine learns from thousands of recorded strokes to know what "typical" looks like. A stroke that's different from that pattern, but still fast and safe, can still get flagged, because the model was never taught to tell "different" apart from "wrong."

At its worst: a program-director audit pulled ten swimmers' film at random and compared it against a month of StrokeLine reports. Two swimmers had been flagged for a "flaw" that was, on closer look, just an unusual stroke that happened to work. One of them had already had three weeks of practice time spent "correcting" a problem that never existed.

We did not lose three weeks of practice time to a slow model. We lost it to a report that never said the difference between "wrong" and "different" was one it couldn't always tell.
The decision I would take back We told coaches, on day one, that StrokeLine's flags were "validated against elite technique models," a strong, unqualified claim, without also saying how it treats a stroke that's genuinely unusual but effective. That felt like the confident, professional thing to say at launch. It stopped making sense the moment we had our first swimmer whose stroke simply didn't look like the training data.

What I would leave alone: for swimmers with textbook technique, the vast majority, StrokeLine's flags are reliably right. The fix isn't softening every report, it's naming the one edge where it isn't.

The lesson: an onboarding that says "it's not perfect" and stops there hasn't set an expectation at all. An expectation has to name the actual shape of the failure, or it's just a disclaimer nobody remembers by week three.

Now here is the same thing as a story

The short version above is what you'd say in an interview. This one is how it actually happened, week by week.

Her name is Bekah Lindqvist. She has coached competitive swimmers for eleven years, and before StrokeLine, she knew every stroke on her roster the way you know a friend's handwriting.

Hand sketched icon list titled FLIPS, the five letters. Five items: F find the person, L locate the habit, I identify the flip, P pinpoint the old decision, S show the replay.
Five letters, and the third one, identify the flip, is the one worth slowing down on.

StrokeLine arrived the spring before last. For months, it was a genuine gift: Bekah could scan a report in ninety seconds instead of scrubbing through film for an hour, and when she did spot-check, it kept agreeing with her eye.

Hand sketched comparison diagram titled How many flagged swimmers got a film check. Left panel, a document icon labeled Week 2, caption 9 of 10 flagged swimmers checked. Right panel, a question mark box icon labeled Week 8, caption 1 of 10 flagged swimmers checked.
Nobody decided to stop checking. It just kept being right, week after week, until checking felt like a formality.

So she checked nine flagged swimmers on film one week, then six, then three, then one. Not out of carelessness. Out of a habit that kept getting confirmed as unnecessary.

Hand sketched metaphor scene titled A dial, or a switch. Left, a gauge icon labeled MANY SETTINGS, caption how sure, moment to moment. Right, a box icon labeled TWO SETTINGS, caption checks it, or doesn't.
Bekah never had a confidence percentage in her head. She had a feeling with exactly two settings: worth checking, or not.
Hand sketched flow diagram titled Bekah's season, in five beats. Five steps: trusted her own eye, StrokeLine kept agreeing, checks turned rare highlighted, audit found two misses, rule three flags means film.
The third beat is the quiet one. Nothing dramatic happened there, which is exactly why nobody noticed it happening.

Then, in month six, the program director ran the quarterly audit every coach knew about but nobody thought much of: ten swimmers pulled at random, film compared against a month of reports.

Hand sketched timeline titled One swim season, four milestones. Four milestones: launch month 1, good months trust builds, checks fade month 4, audit month 6 highlighted.
The audit found the problem. It did not cause it. The problem had been growing for two months already.

Two of the ten had been flagged for a "narrow kick," a flaw StrokeLine names often. On film, both kicks were unusual, wider stance, different tempo, but neither was actually a problem. One swimmer had already spent three weeks in practice narrowing a kick that had been working fine.

We did not lose accuracy. We lost the difference between a swimmer's stroke being wrong and a swimmer's stroke simply not looking like everyone else's.

The fix wasn't retraining the model on more strokes, though that helps eventually. The fix was the report itself, and a rule that didn't depend on Bekah remembering to double-check.

Hand sketched labeled parts diagram titled What a good stroke report shows. Center document icon labeled Stroke Report. Four callouts: matches common pattern, differs not wrong, a confidence note, a film-pull rule.
The second callout is the one the old report never had at all: a way to say "different" without saying "broken."

Now any swimmer flagged three times in a month gets an automatic film pull, no coach has to ask. The report itself separates "matches common patterns" from "differs from common patterns, not necessarily wrong." The same unusual kick this season got flagged, the rule triggered a film check within a week, and Bekah confirmed it was fine before a single practice got rebuilt around fixing it.

I wrote "validated against elite technique models" because it sounded confident and it was true, as far as it went. It took watching three weeks of a swimmer's practice get spent correcting a stroke that was never broken to understand that a true claim can still leave out the one thing a coach needed to hear.

The five steps, if you want to remember itNot a checklist. FLIPS is what forces you to name the actual switch, not just the number behind it.

F
Find the person.
Bekah Lindqvist, eleven years coaching, reading a StrokeLine report every morning.
Grounds the whole answer in one real reader of one real report.
L
Locate the habit.
She stopped pulling film to double-check flagged swimmers, from nine a week down to one.
Names the habit as rational, not careless: it kept being confirmed as unnecessary.
I
Identify the flip.
She went from checking flags sometimes to not checking at all. No middle setting, and the trigger was good news: the tool kept being right.
The over-trust flip: the one that fires when things are going well, not badly.
P
Pinpoint the old decision.
We claimed StrokeLine's flags were "validated," full stop, without naming how it treats a genuinely unusual stroke.
A specific, reasonable-at-the-time decision, not a vague oversight.
S
Show the replay.
Same flagged kick, new rule. A film pull triggers automatically at three flags, caught within a week instead of a quarter.
Ends in something countable: a week, not a season, and zero practice hours wasted.
Time to catch a wrong technique call
12 weeks 6 weeks 0 12 weeks 1 week Old: quarterly audit only New: three-flags rule
Same wrong call. The only thing that changed is how long it took anyone to notice.

The recap, one line per letter: find the person is Bekah reading a report each morning, locate the habit is her fading film checks, identify the flip is trusting the whole flag list instead of any of it, pinpoint the old decision is the unqualified "validated" claim, and show the replay is a wrong flag caught in a week instead of a quarter.

And if you want to be sure it really works, try it somewhere elseA different flip family this time, a solar-panel crew instead of a pool deck. The failure isn't over-trust, it's a workaround nobody asked them to build.

Andres Quilaqueo leads a solar-install crew that uses ThermaCheck, a handheld thermal-imaging tool, to spot faulty panel connections by their heat signature before a system goes live.

ThermaCheck's onboarding never set an expectation about lighting: heavy morning shadows can throw off a thermal reading enough to trigger false alarms. Nobody told the crew this directly. They just noticed, over a few installs, that early-morning scans kept flagging connections that turned out fine, and midday scans almost never did.

Hand sketched quadrant titled ThermaCheck, sorted by light and reliability. Axes how much shadow in the shot, and how reliable the call is. Midday scan sits top left, little shadow and strong reliability. Early morning sits bottom right, heavy shadow and poor reliability.
Nobody built this workaround on purpose. The crew found the safe corner of the chart by trial and error, and started living there.

So the crew quietly started waiting until midday to run every scan, even on jobs where an early start mattered. This is the pre-editing flip: instead of feeding the tool the real, messy conditions of the job, they learned its failure shape and started grooming the conditions to avoid it, which meant an early-morning fault, the exact case they'd been avoiding, would now never get scanned at all until midday, if the crew remembered to go back.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the specific failure, build a rule that catches it," and stop.
Cost: there's no budget for a rule engine this quarter. Say so honestly, and start with one line in the onboarding naming the lighting issue directly, since even a named risk beats a silent one.
The model gets better, for real: if StrokeLine's accuracy improves to 99 percent on typical strokes, the unusual-stroke gap still needs naming, a rarer miss on an atypical case is still a miss nobody is watching for.

Where people run it wrong.
They write a disclaimer once, at signup, and never repeat the specific failure shape anywhere the person will see it again.
They assume a quiet stretch with no complaints means expectations are working, when it often means the safety net already eroded.
They respond to a discovered miss by adding a general "please double-check the AI" reminder instead of a rule that triggers automatically.

How to use it live. When someone asks how you'd set error expectations, ask yourself one question first: what is the one specific, nameable way this tool gets it wrong. Say that, not "it's not always right."

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: checks sometimes, then stops checking at all. It fires when things are going well, not badly, which is what makes it easy to miss.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bekah Lindqvist, an eleven-year swim technique coach who reviewed film for her whole roster every week before StrokeLine.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Pulling film to double-check flagged swimmers, from nine a week down to one, because StrokeLine kept agreeing with her own eye.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Checking a flagged swimmer's film, or not checking at all. No middle setting, and the trigger was the tool being right for months straight.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Calling every StrokeLine flag "validated against elite technique models" without naming that a genuinely unusual stroke could still get flagged.
6 · THE NUMBER
Fill in the blank: Bekah's film-check rate fell from 90 percent in week one to ___ percent by week eight.
Tap to flip
ANSWER
10 percent. The audit that found the problem didn't run until week ten, two weeks after the number had nearly bottomed out.
7 · THE REPLAY
Same unusual kick flagged, new rule in place. What changes?
Tap to flip
ANSWER
A three-flags rule triggers an automatic film pull within a week, instead of a quarterly audit catching it twelve weeks later, and no practice time gets wasted fixing a stroke that was never broken.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product and a different flip family. Which product, and which family?
Tap to flip
ANSWER
ThermaCheck, a solar-install thermal-scan tool. The pre-editing flip: the crew learned to only scan at midday, avoiding the exact lighting condition where real early-morning faults would show up.

Check yourself Score: 0 / 0

Short answer, recall the flip
1. What was the flip in Bekah's story, and what were its two settings?
Show hint
Look at the line chart of her film-check rate by week.
Show answer
Model answer: Checking a flagged swimmer's film against not checking at all. She moved from the first setting to the second over eight weeks, with no stop in between.
Multiple choice
2. Why couldn't Bekah have just "checked a little more carefully" instead of stopping entirely?
  • A. She was told by management to stop checking.
  • B. The habit is a two-setting flip, not a dial. Once checking stopped being necessary week after week, there was no middle ground to settle into.
  • C. StrokeLine locked her out of the film archive.
  • D. She didn't have time to check any swimmers at all, ever.
Show hint
Look at "identify the flip" in the FLIPS recap.
Show answer
B. A flip has exactly two settings and no middle. "Checking a bit less" would be a dial, and this behavior wasn't one.
True or false
3. True or false: the fix in this answer was to retrain StrokeLine so it never flags an unusual stroke again.
  • True
  • False
Show hint
Look at "the choice I would take back" and the labeled parts diagram.
Show answer
False. The fix was naming the failure honestly in the report and adding an automatic film-pull rule, not retraining the model to stop flagging anything unusual.
Short answer, where it wouldn't matter
4. Name a place in StrokeLine where this exact problem would NOT show up.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Swimmers with textbook technique, the large majority of the roster, where StrokeLine's flags are reliably right and don't need this extra scaffolding.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think of something you stopped double-checking because an app kept being right.
Show answer
Model answer: Many people stop proofreading autocorrected text, or checking a GPS route against their own sense of direction, once the tool has been right often enough. That's the same over-trust flip.
Before you close the answer
Why this works
Tests whether you can set an expectation specific enough to actually change behavior later, rather than a disclaimer that reads well at signup and does nothing three months in.
Follow-up traps
"Wouldn't naming the failure shape just make coaches distrust the whole tool?" Response: no, the opposite. Coaches trusted StrokeLine more once they knew exactly where its blind spot was, because it meant the rest of the report had earned its confidence honestly.

"Isn't a three-flags rule just adding more review, which FLIPS says doesn't count as a real fix?" Response: it's not more review, it's automatic and it doesn't depend on anyone remembering to ask. The disqualified move is telling a person to "watch more closely"; this rule removes the need to remember at all.
If pressed
StrokeLine's real fix also logs whether a flagged swimmer's film pull confirmed or overturned the flag, so the training set for "differs, not wrong" grows every time a coach actually checks, instead of staying frozen at launch.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more