CaseIntermediateDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #2

How do you teach users what the AI can and cannot do without a manual?

SPARK the product is GlintCheck, a camera AI that checks stamped auto parts for defects at Ridgeline Stamping

Ridgeline Stamping runs a press line that turns sheet steel into car door panels. Priya Deshmukh has run that line for nine years. GlintCheck is the camera mounted over the last press, checking every panel as it comes off.

The direct answer
Teach the edge inside the first shift, not in a document. Show one real photo of a defect GlintCheck catches, and one real photo of a defect type it does not catch, on purpose, before the operator ever runs a panel alone. Pair every "checked" screen with one plain rule: if something looks off and GlintCheck didn't flag it, log it anyway. That's the whole manual, and it fits in the first four minutes.
Do this, in order
  1. Show one real catch and one real miss in the first shift, with actual photos, not words.Why: a picture of a miss teaches the edge faster than any paragraph.
  2. Give a standing rule: log anything odd GlintCheck didn't flag.Why: without it, silence reads as "all clear," not "unchecked."
  3. Split "here's the tool" from "here's your dashboard" into two separate moments.Why: merging them removes the one pause where a new operator could ask what's out of scope.
  4. Name the trained defect set out loud, not just its accuracy number.Why: a percentage hides which specific things it was never taught to see.
  5. Re-run the miss-example the day the line itself changes, a new coating, a new supplier.Why: the edge moves when the input does, and the lesson has to move with it.
  6. Leave the confidence-per-pixel overlay for later.Why: real depth, but day one only needs the edge, not the mechanism.

How to answer this, stage by stage

A student who reads only this section, and reads it out loud once, should be able to answer the question cold.

Stage 1
Scope it to one real product and one real person
Say it like this
"I'll answer this for GlintCheck, a camera AI that checks stamped door panels for defects, and for Priya, who runs one press line at Ridgeline Stamping."
Why this works
Grounds an abstract "how do you teach" question in one real screen a real person looks at.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, what she does today. Payoff, the habit I want. Anchor, the one design call. Risk, what breaks it. Keep out, what I won't build yet."
Why this works
Shows a plan before a story, so the interviewer knows you're building toward something.
Stage 3
Reframe the question
Say it like this
"Nobody reads a manual on day one. So the edge of what this thing can do has to live inside the product's first real moment, not in a document sitting next to it."
Why this works
Separates a strong answer from a list of training-materials features nobody would use.
Stage 4
Give the one decision
Say it like this
"In the first shift, before she runs a single panel alone, I'd show Priya one photo GlintCheck flags correctly, and one photo it doesn't flag at all, on purpose, so she learns the edge before she ever needs it."
Why this works
Concrete, checkable, and matches the one-sentence answer exactly.
Stage 5
Prove it with a failure
Say it like this
"When Ridgeline switched coating suppliers, GlintCheck kept passing every panel clean for three weeks, because streaks weren't a defect it had ever been shown. Nobody watched for them, because the first shift never said streaks were outside the tool's reach."
Why this works
Shows the real cost of a missing boundary lesson, not a hypothetical one.
Stage 6
Say what you'd measure
Say it like this
"I'd track how many odd-but-unflagged panels operators log each week. If that number sits near zero for weeks, the boundary lesson didn't stick, no matter how the accuracy number looks."
Why this works
Shows you think past launch day, and that you'd catch the failure before an audit does.
Stage 7
Say what you'd leave alone
Say it like this
"Crack and hole detection don't need this scaffolding. GlintCheck matches expert judgment on those almost every time, so I wouldn't slow the shift down teaching caution nobody needs there."
Why this works
Shows judgment, not blanket caution applied everywhere out of fear.
Stage 8
Close on the one line
Say it like this
"Show a catch, show a miss, give a log-it rule. That's the whole boundary lesson, and it fits in four minutes."
Why this works
Restates the decision in one breath, ready for whatever comes next.

Let's learn

For nine years, Priya Deshmukh walked the press line at Ridgeline Stamping at six in the morning, and could spot a hairline crack in a stamped door panel before the machine even finished its cycle.

GlintCheck is a camera mounted over the last press. It checks every panel the second it's stamped, for cracks, misaligned holes, and warping, the three defect types it was trained on.

Before GlintCheck, Priya and her crew spot-checked one panel in every ten, about 400 of the 4,000 stamped in a shift. They caught most of what they were looking for. Now GlintCheck checks all 4,000, and catches nearly every crack and misaligned hole, over 97 out of 100.

GlintCheck's catch rate, by defect type
100% 50% 0 98% 97% 81% 4% Crack Misaligned hole Pitting Coating streak
Coating streak was never in the trained set. GlintCheck didn't get worse at catching it, it was never built to.
Knowledge spark: what's a trained defect set? The exact list of defect types the model was shown thousands of examples of, so it learned to recognize them. Anything outside that list, it was never taught to see, no matter how good its overall accuracy number looks.

At its worst: three weeks after a coating supplier switch, a batch of 1,200 panels with coating streaks shipped past GlintCheck, which had never been shown that pattern, and past every operator, who had learned to trust a clean scan as the whole story. A routine skip-level quality audit caught it, not the line.

GlintCheck did not get worse. It was still doing exactly what it was built to do. The streaks slipped through the one gap nobody had ever named out loud.
The decision I would take back We built the first-shift walkthrough as one ten-minute screen: log in, watch a demo scan, see a green check, start the line. It was simple to build and easy to schedule everyone through in one sitting. That worked fine while the trained set covered nearly everything the line produced. It stopped working the day the coating changed and a whole new defect type showed up with nobody watching for it.

What I would leave alone: crack and hole detection do not need any of this scaffolding. GlintCheck matches an expert's judgment on those almost every time, tested across more than 50,000 panels. Adding caution there would only slow the shift down for no real gain.

The lesson: a tool that checks everything is not the same as a tool that checks everything you care about. If the onboarding never says which is which, the operator has no way to know, and neither did we, until an audit found out for us.

Hand sketched icon list titled What a boundary-teaching first shift needs. Four items: a gauge icon labeled show one real catch, a question mark box icon labeled show one real miss, a document icon labeled name the trained set, a person icon labeled give a log-it-anyway rule.
Four items, and the second one is the one the old ten-minute walkthrough never had.

Now here is the same thing as a story

The short version above is what you'd say defending the design in an interview. Read this one for how it actually happened on the floor.

The press line at Ridgeline Stamping starts loud at 5:45, when the first coil feeds through and the stamps begin their rhythm. Priya has run that rhythm for nine years. Ask her to find a hairline crack in a stack of a hundred panels and she'll have it before you've finished asking.

Hand sketched labeled parts diagram titled Priya's line, before GlintCheck. Center person icon labeled Priya, 6am. Four callouts: eyes every tenth panel, knows cracks by feel, logs bad ones on paper, trusts her own memory.
Before GlintCheck, one in ten panels got a real human look. The other nine rode on Priya's memory of what usually goes wrong.

GlintCheck arrived in February. For the first few months it was a genuine relief: every single panel got checked, not one in ten, and Priya's crew could spend their attention on the panels the camera actually flagged instead of guessing which ones to sample.

The first-shift walkthrough new operators got was short: log in, watch a demo scan run clean, see a green checkmark, start the line. Nobody asked what the checkmark didn't cover. Nobody had been given a reason to ask.

Hand sketched comparison diagram titled The day a new defect shows up. Left panel, a question mark box icon labeled Old onboarding, caption trusts every flag as the whole story. Right panel, a person icon labeled New onboarding, caption watches for what got skipped.
Same camera, same panel. The difference is only in what the operator was ever told to watch for beyond the screen.

Then Ridgeline switched coating suppliers, for a cheaper, faster-drying finish. Nobody flagged it as a training-data event, because nobody thought of GlintCheck as having training data at all. It was just "the checker."

Hand sketched quadrant titled Sorting defects by how well GlintCheck sees them. Axes how visual the defect is, and how well GlintCheck catches it. Hairline crack and misaligned hole sit top right, obvious and almost always caught. Surface pitting sits in the middle. Coating streak sits bottom left, subtle and rarely caught.
Coating streak sat in the one corner nobody had drawn on the map yet: subtle, and almost never caught.

For three weeks, the streaks kept coming, and GlintCheck kept passing every panel clean, because a streak was never in its trained set. Then a skip-level quality audit, the kind that happens twice a year, pulled a random sample and found 1,200 panels already shipped downstream with visible coating streaks.

We did not lose a camera's worth of accuracy. We lost three weeks nobody was watching the one thing the camera was never built to watch.

The fix wasn't a smarter camera. It was a four-minute redesign of that first shift.

Hand sketched flow diagram titled The new first-shift walkthrough. Four steps: see a real catch, see a real miss highlighted, hear the boundary, log anything odd.
Step two is the one that used to be missing entirely. A photo of a real miss teaches the edge faster than any sentence about it.
Hand sketched timeline titled The four-minute first shift. Four milestones: shown a catch, a real cracked panel. Shown a miss highlighted, a real coating streak. Told the edge, trained set named. Given the rule, log it if unsure.
Four minutes, four milestones. Nothing here needed a manual, only a screen and two real photos.

Now every new operator sees one real photo of a caught crack, and one real photo of a coating streak GlintCheck waved through, side by side, on their very first login. They're told, plainly: this camera has a list of things it looks for, and this isn't one of them yet. If something looks wrong and the screen says clean, log it anyway.

I built the old walkthrough because it was fast to run and easy to schedule everyone through on their first day. It took watching 1,200 panels ride past two separate sets of eyes, human and camera, before I understood that "checked" and "checked for what you care about" are not the same claim, and our onboarding had only ever taught the first one.

SPARK, the design behind the four minutesNot a checklist for operators. SPARK is what forces the design itself to survive the day it's wrong.

S
Situation. What she does today, without the redesign.
Priya's crew trusts a green checkmark on the tablet as "this panel is fine," full stop, with no sense of the camera's own edges.
Grounds the design in one real habit, not a hypothetical user.
P
Payoff. The habit worth building.
Operators stop treating a flag as the whole truth, and start noticing when something looks off that the screen never mentioned.
Names the habit as the real product, not the time saved.
A
Anchor. The one design call.
The first-shift screen shows one real catch and one real miss, using actual photos, plus a standing log-it-anyway rule.
Concrete enough to build, concrete enough to argue with.
R
Risk. What breaks the anchor.
The line changes a material or supplier, and a defect type nobody has ever photographed for training shows up.
Names the real failure mode, not a vague "accuracy drops."
K
Keep out. What we won't build yet.
A live per-pixel confidence overlay. Real depth, but day one only needs the edge of the tool named, not its inner mechanics.
Shows restraint instead of a wish list nobody asked for.
Operators logging an unflagged anomaly, by week
12 6 0 audit finds it new onboarding old onboarding Week 1 Week 8
Under the old onboarding, logging quietly died out by week five, three weeks before anyone else noticed anything was wrong.

The recap, one line per letter: situation is the trusted checkmark, payoff is noticing what got skipped, anchor is the catch-and-miss first screen, risk is a genuinely new defect type, and keep out is the confidence overlay we're saving for later.

And if you want to be sure it really works, try it somewhere elseSame five letters, a denim mill instead of a stamping plant. A completely different fabric, and this time the edge is a dye lot, not a coating.

Aldergate Mills runs LoomSight, a camera system over its looms that checks woven denim for flaws, snags, misthreads, tension lines, as it comes off each roll. Hollis Vance is the quality technician who signs off on every roll before it ships to a cutting house.

Mapped onto SPARK: situation is Hollis walking a physical roll by hand under a light table before LoomSight existed, catching snags by touch as much as sight. Payoff is the habit of noticing a weave pattern LoomSight has never priced in, instead of trusting a clean scan as final. Anchor is the same catch-and-miss first session, but built from fabric photos: one real snag caught, one real irregularity from a brand-new dye lot that LoomSight waves through because it has never seen that pattern before. Risk is a new dye supplier shipping a pattern the model has never priced in, exactly the coating-switch problem in a different material. Keep out is a full thread-count overlay, which is real information Hollis doesn't need on day one.

Hand sketched decision tree titled LoomSight's boundary, at Aldergate Mills. Root, weave flag raised, branching to four outcomes: matches a trained flaw leads to sort as defect, new dye lot pattern leads to show Hollis first, known false alarm leads to clear it, low confidence leads to hold for a check.
The branch that matters most is the second one: a pattern nobody trained the model on gets routed to a person, not waved through silently.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show a catch, show a miss, give a log-it rule," and stop there.
Cost: there's no budget for a proper first-shift redesign this quarter. Say so honestly, and start with a single laminated card taped by the tablet showing one miss photo, since even a low-effort version beats none.
The model gets better, for real: if GlintCheck's overall accuracy climbs to 99 percent, that's still not a reason to drop the miss-example, a rarer failure is a more dangerous one to have nobody watching for.

Where people run it wrong.
They write a manual nobody opens, instead of building the lesson into the product's own first moment.
They quote one overall accuracy number and assume it teaches the boundary, when it hides exactly which things were never in scope.
They treat "the model got better" as a reason to remove the caution the first design earned, instead of a fresh reason to re-check it.

How to use it live. When someone asks how you'd teach capability without a manual, ask yourself one question first: what's the one photo of a miss I'd want a brand-new operator to see before their first real panel. Say that photo out loud, not a policy about training materials.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you teach capability without a manual"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. The anchor step is the concrete design decision that answers the question.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Deshmukh, a nine-year floor supervisor at Ridgeline Stamping who could spot a hairline crack before a machine finished its cycle.
3 · THE HABIT
What did operators stop doing because GlintCheck worked?
Tap to flip
ANSWER
They stopped watching for anything the screen itself didn't flag, treating a clean scan as the whole story instead of one camera's opinion.
4 · THE ANCHOR
What's the one concrete design decision this answer commits to?
Tap to flip
ANSWER
Show a real catch and a real miss in the first shift, using actual photos, plus a standing rule to log anything odd the screen didn't flag.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging "here's the tool" and "here's your dashboard" into one ten-minute screen, which removed the pause where a new operator could ask what's out of scope.
6 · THE NUMBER
Fill in the blank: coating streaks went undetected for ___ before an audit caught them.
Tap to flip
ANSWER
Three weeks, during which 1,200 panels with visible streaks shipped downstream.
7 · THE REPLAY
Same coating switch, redesigned first shift. What changes?
Tap to flip
ANSWER
Operators keep logging odd panels steadily, six to nine a week, instead of the rate falling to zero, so a new defect type gets caught by a person within days, not weeks.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the new edge?
Tap to flip
ANSWER
LoomSight at Aldergate Mills, a denim-flaw scanner. There, the edge is a brand-new dye lot pattern the model has never priced in, not a coating change.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer recommends writing a longer, clearer manual for new GlintCheck operators.
  • True
  • False
Show hint
Look at the direct answer at the top of the page.
Show answer
False. The fix is inside the first shift itself, real photos and a standing rule, not a document nobody would read on day one.
Multiple choice
2. Why did coating streaks go undetected for three weeks?
  • A. GlintCheck's camera lens was dirty.
  • B. Coating streaks were never part of GlintCheck's trained defect set, and nobody had been told to watch for anything outside it.
  • C. Operators were told to ignore GlintCheck entirely during the supplier switch.
  • D. The new coating supplier shipped defective steel.
Show hint
Look at the catch-rate chart and the quadrant diagram.
Show answer
B. GlintCheck caught coating streaks only 4 percent of the time, because it was never shown that pattern, and the onboarding never named the gap.
Fill in the blank
3. Fill in the blank: under the old onboarding, operators logging an unflagged anomaly fell from twelve a week to ___ by week five.
Show hint
Look at the line chart of operators logging anomalies by week.
Show answer
Zero. The habit of logging odd panels quietly died out three weeks before the audit found the problem.
Short answer, where it wouldn't matter
4. Name a defect type at Ridgeline where this whole boundary-teaching problem does NOT apply, and why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Crack and misaligned-hole detection. GlintCheck matches expert judgment on those almost every time, so there's no real edge there worth teaching caution around.
Short answer, apply it yourself
5. Pick an AI product you use yourself. What's one thing it clearly cannot do that you only learned by it failing, rather than by being told upfront?
Show hint
Think of a time an app or assistant gave a wrong or strange answer, and how you found out it wasn't built for that case.
Show answer
Model answer: Most people can name a case, a chatbot inventing a fact, a photo app mistagging something unusual, that they only discovered by hitting it, exactly the gap this answer says onboarding should close early.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Merging the tool introduction and the dashboard tour into one ten-minute screen. It made sense while the trained set covered nearly everything the line produced, so there was no visible edge to teach yet.
Before you close the answer
Why this works
Tests whether you can design a capability lesson into a product's own first moment instead of defaulting to documentation nobody reads, and whether you can name a real AI-specific edge rather than a generic training problem.
Follow-up traps
"Couldn't you just retrain GlintCheck to catch coating streaks and skip the onboarding fix entirely?" Response: eventually, yes, but retraining takes time and new defect types will keep showing up faster than any model can be retrained for all of them, so the onboarding habit is the fallback that covers whatever training hasn't caught up to yet.

"Isn't naming the trained set just admitting the product is unreliable?" Response: no, it's the opposite. Naming the edge plainly is what lets people trust the parts that are reliable without over-trusting the parts that aren't.
If pressed
Ridgeline's real fix also timestamps every operator log against the exact model version running that shift, so when the trained set is updated to finally catch coating streaks, nobody has to guess which shift's data trained which fix.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more