ConceptFoundationalDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #4

What is the risk of an onboarding that oversells capability?

GUARD the product is HullSense, an AI engine-diagnostic tool used at a boatyard

The report is one page. It says HullSense Verified Diagnosis, and nothing about what it didn't check. Imelda Reyes manages service at a boatyard that runs HullSense on every engine that comes in for a checkup.

The direct answer
The risk isn't that the tool is wrong sometimes. It's that a confident-sounding onboarding tells the person receiving the verdict there's no reason to check it, and takes away the one thing that would have caught the wrong calls: someone asking for a second look. Fix the onboarding, not just the model, by naming a confidence range and a way to contest it, out loud, before the first verdict ever gets delivered.
Do this, in order
  1. Never brand a verdict as fully verified when its real accuracy varies by case.Why: a flat "verified" label erases the difference between a 94 percent case and a 68 percent one.
  2. Give the person receiving the verdict a visible way to ask for a second look.Why: without one, they have no lever at all, only a bill.
  3. Name which cases the tool is weakest on, in the report itself.Why: the owner of an older, less common engine deserves to know that before agreeing to a repair.
  4. Track how often people actually ask for a second opinion.Why: a near-zero rate isn't proof nothing's wrong, it's often proof nobody was ever told they could ask.
  5. Don't fix this by adding a review board or a policy document.Why: a real design change, a visible range and a visible path, does the work a committee can't.

How to answer this, stage by stage

Six stages. Say each one plainly, without moralizing, and the answer holds up under a real follow-up.

Stage 1
Scope it to one real report
Say it like this
"I'll answer this for HullSense, an engine-diagnostic AI used at a boatyard, and for the one-page report it hands to a boat owner."
Why this works
Turns an abstract risk question into a specific document with a specific reader.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's involved. Unequal, where the harm lands hardest. Ability to contest, who can push back. Reduce, the design fix. Detect, how you'd know it's happening."
Why this works
Signals a structured risk answer instead of a general worry.
Stage 3
Name both people
Say it like this
"There's Imelda, who runs the boatyard and reads HullSense's output every day. And there's the boat owner, who reads one page, once, and has no engineering background to check it against."
Why this works
The strongest GUARD move: naming the operator and the person who can't push back.
Stage 4
Name the ability-to-contest gap
Say it like this
"The report says Verified Diagnosis. It doesn't say verified against what, or how sure. There's no line on the page that tells an owner they could ask for another look."
Why this works
This is the hard step in GUARD: who never gets a lever, and why.
Stage 5
Give the number
Say it like this
"On the four common outboard engines HullSense sees most, it's right about 94 percent of the time. On older diesel inboards, a smaller slice of the fleet, it's right about 68 percent of the time. Both get the same badge."
Why this works
Turns "it oversells capability" into a measured, specific gap.
Stage 6
Close on the design fix
Say it like this
"Name the confidence range on the report, name what wasn't checked, and put a second-look request one tap away. That's the whole fix, and none of it touches the model."
Why this works
Ends on a concrete design decision, not a call for more caution in general.

Let's learn

The report is one page. It says HullSense Verified Diagnosis, and nothing about what it didn't check.

HullSense is a diagnostic tool a boatyard runs on an engine during a checkup, reading sensor data and comparing it against known failure patterns to name what's wrong before a mechanic tears anything apart.

On the boatyard's most common engines, four outboard motor lines that make up most of the fleet, HullSense gets it right about 94 times in 100. On older diesel inboards, a smaller and less common slice of boats, it's right about 68 times in 100, because far fewer of those engines exist in the data it learned from.

HullSense diagnostic accuracy, by engine type
100% 50% 0 94% 68% Common outboards Older diesel inboards both get the same badge
A 26-point accuracy gap, and the report never lets an owner know which side of it they're standing on.
Knowledge spark: what's a confidence range? A model's own honest measure of how sure it is, stated as a range instead of one flat number. "Between 60 and 75 percent sure" is a confidence range. "Verified" is not, it hides whether the model is barely sure or nearly certain.

At its worst: a boat owner with an older diesel inboard nearly authorized a full engine teardown, a repair costing thousands, based on a confident HullSense call. A newer mechanic on shift that day happened to double-check the sensor readings by hand before the work order went out, and found a faulty temperature sensor, not the engine damage HullSense had named. Nothing on the report had told the owner, or the mechanic, that this engine type was one HullSense handled worse.

The risk was never that HullSense gets some diagnoses wrong. It's that the report gave nobody a reason to ask whether this one might be one of them.
The decision I would take back We badged every HullSense report "Verified Diagnosis," attributing the call to the tool with full, unqualified authority, no visible human sign-off, no scope note. That built fast trust when the boatyard was marketing HullSense to new customers early on. It stopped making sense the moment accuracy started varying this much by engine type, and the badge never changed to reflect it.

What I would leave alone: for the four common outboard lines, where HullSense really is right 94 percent of the time, a strong, confident badge is honest and it should stay. The fix is naming where that confidence doesn't hold, not softening it everywhere.

The lesson: "verified" is a claim about the whole system. If any part of that system is weaker for some customers than others, the badge is making a promise the model itself never made.

Now here is the same thing as a story

The short version above is the plain risk case. Read this one for how close the near miss actually came.

The clipboard behind Imelda's desk has a HullSense printout stapled to every third page, one for nearly every engine that's come through this year. She's run service at this boatyard for twelve years, long enough to know most of these engines by the sound they make idling.

Hand sketched metaphor scene titled Two people, one lever. Left, a gauge icon labeled THE MECHANIC, caption holds the gauge, and the lever. Right, a person icon labeled THE OWNER, caption holds an invoice, nothing else.
Imelda can read the sensor data behind a HullSense call. The owner across the counter only ever sees the one word: Verified.

HullSense launched with a clean, confident report design, on purpose. New customers were nervous about trusting an algorithm with their engine, and a flat "Verified Diagnosis" badge, without any hedging, was what got people comfortable fast. It worked. Business grew.

Hand sketched comparison diagram titled One verdict, or a verdict with its own edges. Left panel, a document icon labeled Verdict only, caption HullSense Verified Diagnosis. Right panel, a scale icon labeled Verdict with range, caption confident here, unsure there.
Same underlying model, same sensor data. The only difference is whether the report admits where its own confidence actually sits.

Nobody at the boatyard was hiding the accuracy gap between common outboards and older diesels. Nobody had measured it as a gap at all, until the near miss.

Hand sketched quadrant titled Sorting calls by stakes and real accuracy. Axes what's at stake for the owner, and how well HullSense actually performs. Common spark issue and common oil pressure sit top left, low stakes and well handled. Old diesel block crack and rare vintage inboard sit bottom right, major cost and poorly handled.
The bottom right corner is where the real danger lives: high stakes, paired with a real accuracy rate nobody had put a number on yet.

A boat owner with a 1994 diesel inboard brought his engine in running rough. HullSense's report named a cracked engine block, a repair running into the thousands, with the same confident badge every other report carried. The owner, no engineer, had no reason to question it. He was one signature away from authorizing the teardown.

A newer mechanic on shift, still in the habit of double-checking sensor readings by hand before big jobs, happened to catch it: a faulty temperature sensor was throwing a reading that looked exactly like block damage to the model. The real fix cost around 200 dollars, not several thousand.

Hand sketched flow diagram titled Where the appeal should sit, and doesn't. Four steps: engine scanned, verdict shown, part replaced, no second look offered highlighted in the fourth step.
The gap wasn't in step one, two, or three. It was the missing fourth step, the one nobody had ever built.

We did not almost lose an engine. We almost sent a customer thousands of dollars into a repair he had no way to question, because nothing on the page ever told him he was allowed to.

Hand sketched labeled parts diagram titled What a trustworthy diagnosis report needs. Center document icon labeled Diagnosis Report. Four callouts: a confidence range, the model version, what wasn't checked, a second-opinion path.
Four additions to one page. None of them touch the model itself, only what the page is honest about.
Hand sketched timeline titled The second-opinion path, added. Four milestones: launch, no appeal path. Near miss, teardown almost ordered highlighted. Link added, get a second look. Dashboard, requests now tracked.
The near miss is the milestone that mattered. Everything after it was a direct response to that one afternoon.

I built the confident badge because early customers needed reassurance, and it worked exactly as designed. It took one mechanic's habit of double-checking, a habit the product itself did nothing to encourage, to catch what the badge was quietly hiding from everyone else.

GUARD, applied plainlyNot a policy. GUARD is what forces the design to say who can push back, and who can't.

G
Groups. Who's involved.
Imelda's boatyard runs HullSense. The boat owner receives its verdict, once, with no engineering background of their own.
Names the operator and the subject before anything else.
U
Unequal. Where the harm lands hardest.
Owners of older, less common engines get the same confident badge as owners of the four common models, despite a 26-point accuracy gap.
Shows the risk isn't evenly spread, it concentrates on a specific group.
A
Ability to contest. Who never gets a lever.
The boat owner has no way, printed anywhere on the report, to ask for a second look before authorizing a repair.
The hardest step in GUARD, and the one this whole answer turns on.
R
Reduce. The design change.
Print a confidence range and a named second-opinion contact directly on the report, not in a separate policy document.
A product decision, not a training session or a review board.
D
Detect. How you'd know it's happening.
Track the rate of second-opinion requests. Near zero for months isn't proof of a healthy system, it's a sign nobody was ever told they could ask.
Catches the problem in production, before it takes an external near miss to surface it.
Second-opinion requests, by month
10% 5% 0 no visible path link added 7% by month 12 Month 1 Month 12
Eight months near zero was never a sign of trust well placed. It was a sign nobody had ever been shown the door.

The recap, one line per letter: groups is the mechanic and the owner, unequal is the 26-point gap between engine types, ability to contest is the missing second-look path, reduce is printing a range and a contact on the report itself, and detect is watching the request rate for a signal that's suspiciously calm.

And if you want to be sure it really works, try it somewhere elseSame five letters, a rental-property inspection tool instead of an engine. This time the subject is a tenant, not a boat owner.

Callum Osei manages a portfolio of rental units and uses an AI condition-assessment tool to inspect a unit at move-out and decide how much of a deposit gets withheld for damage.

Mapped onto GUARD: groups is Callum, who reviews the tool's report, and the tenant, who has already moved out and has no way to be present for a re-inspection. Unequal is that tenants in older units, where wear patterns look more like damage to a model trained mostly on newer units, lose more of their deposit on average. Ability to contest is the real gap: many tenants don't know a dispute process exists at all, since the deposit letter states a dollar figure with no mention of how it was reached. Reduce is printing the specific findings with photos and a plain dispute deadline on the letter itself, not buried in a lease clause. Detect is tracking dispute rates by unit age, watching for the same silence HullSense had.

Hand sketched icon list titled What a fair property-condition report needs. Four items: a gauge icon labeled a confidence range not one score, a document icon labeled what rooms weren't inspected, a person icon labeled a named contact to dispute it, a scale icon labeled which version made the call.
Different building, same missing lever: a tenant who never learns they could have asked for another look.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the confidence range, give a way to contest it," and stop.
Cost: there's no budget to redesign the report this quarter. Say so honestly, and start with one added sentence naming the engine types HullSense is weakest on, since even one honest line beats none.
The model gets better, for real: if HullSense's overall accuracy climbs to 97 percent, the ability-to-contest gap still matters, a rarer wrong call on a major repair is still a major repair nobody could question.

Where people run it wrong.
They treat this as a model-accuracy problem and try to fix it by retraining, when the actual gap is in what the report tells the reader.
They assume a low complaint rate means the system is trusted, when it often means nobody knew a complaint was possible.
They respond to a near miss with a review board instead of a specific, visible change to the document itself.

How to use it live. When someone asks about the risk of overselling capability, ask yourself one question first: who reads this output with no way to check it against anything else. Name that person before you name the model's accuracy number.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what's the risk of an onboarding that oversells capability"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Ability to contest is the hardest step, and the one this answer turns on.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Imelda Reyes, a twelve-year boatyard service manager, and the boat owner across the counter who reads HullSense's report once with no way to check it.
3 · THE UNEQUAL HARM
Where does the risk land unevenly?
Tap to flip
ANSWER
Owners of older diesel inboards, where HullSense is right only 68 percent of the time, versus 94 percent on the four common outboard lines, get the exact same confident badge.
4 · THE MISSING LEVER
What ability does the boat owner not have?
Tap to flip
ANSWER
A visible, printed way to request a second opinion before authorizing a repair based on HullSense's call.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Badging every report "Verified Diagnosis" with full, unqualified authority, which built fast trust early on but stopped being honest once accuracy started varying by engine type.
6 · THE NUMBER
Fill in the blank: HullSense's accuracy gap between common outboards and older diesels is ___ points.
Tap to flip
ANSWER
26 points: 94 percent versus 68 percent, both carrying the same badge.
7 · THE FIX IN ACTION
After the second-opinion link was added, what changed?
Tap to flip
ANSWER
Requests rose from near zero to about 7 percent of diagnoses by month twelve, catching real errors the flat badge had been hiding.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and who's the subject there?
Tap to flip
ANSWER
A rental-property condition-assessment tool. There, the subject is a tenant who has already moved out and often doesn't know a dispute process exists.

Check yourself Score: 0 / 0

Multiple choice
1. According to this answer, what is the actual risk of an onboarding that oversells capability?
  • A. The model will eventually be replaced by a better one.
  • B. The person receiving the output has no reason to question it, so nobody catches the cases where it's wrong.
  • C. Engineers will spend too much time improving accuracy.
  • D. Customers will stop using the product entirely.
Show hint
Look at the direct answer and the ability-to-contest step.
Show answer
B. The near miss happened because nothing on the report gave the owner, or anyone else, a reason to check the call.
True or false
2. True or false: HullSense's accuracy is the same across all engine types.
  • True
  • False
Show hint
Look at the bar chart of accuracy by engine type.
Show answer
False. 94 percent on common outboards versus 68 percent on older diesel inboards, a 26-point gap the report never mentions.
Fill in the blank
3. Fill in the blank: for eight months with no visible second-opinion path, requests stayed near ___ percent.
Show hint
Look at the line chart of second-opinion requests by month.
Show answer
Zero. A near-zero rate wasn't evidence of trust. It was evidence nobody had ever been shown a path to ask.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Badging every report "Verified Diagnosis" with no qualifier. It made sense while it built trust with nervous new customers, and stopped making sense once accuracy started varying widely by engine type.
Short answer, apply it yourself
5. Think of an AI-generated result you've received, a credit score explanation, a spam filter decision, a delivery estimate. Did you know you could question it, and how would you have found out?
Show hint
Think about whether the output came with any visible way to ask for a human review.
Show answer
Model answer: Most people can't recall being told a dispute path existed until they went looking for one themselves, which is exactly the ability-to-contest gap this answer is about.
Before you close the answer
Why this works
Tests whether you can separate a model-accuracy problem from a design-honesty problem, and whether you'll name the person who has no power in the interaction, not just the person operating the tool.
Follow-up traps
"Wouldn't printing a confidence range just confuse customers and reduce trust?" Response: only if it's poorly worded. A plain range, "confident here, less sure there," reads as honest, not alarming, and it's what actually earns durable trust instead of trust that collapses the first time a near miss goes public.

"Isn't a 94 percent accuracy rate already good enough to call it verified?" Response: for the common engines, maybe. But the badge doesn't distinguish those from the 68 percent case, and that's the actual problem, not the average number.
If pressed
The boatyard's real fix also logs which HullSense model version produced each report, so when accuracy on older diesels improves after a retrain, the badge language for that engine type can update without waiting for a full policy rewrite.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more