CaseAdvancedDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #8

How do you onboard a user into an agent product where the AI acts on their behalf?

ORDER the product is ShopPilot, an agent that reschedules maintenance and reorders tooling at a CNC machine shop

What's the first thing you'd let an agent do on your behalf, before you've ever watched it work? Soledad Larrain manages Ferro Machine Works, a CNC shop where ShopPilot watches tooling wear and job schedules, and can act on both.

The direct answer
Sequence autonomy by how hard an action is to undo, never by how complete the feature is. Start every user in shadow mode, where the agent drafts every action but executes nothing. Unlock real autonomy action-type by action-type, the cheapest-to-reverse ones first, and require a track record before handing over anything expensive to undo.
Do this, in order
  1. Start every new user in shadow mode, where the agent drafts but never executes.Why: it's the only way to see what the agent would have done before it costs anything.
  2. Unlock autonomy for reversible actions first, expensive-to-undo ones last.Why: a mistake you can undo costs nothing; a mistake you can't is the one worth sequencing carefully around.
  3. Require a proven track record before any hard-to-undo action type earns autonomy.Why: a new action type has no history yet, and history is the only real evidence a graduated system has.
  4. Never grant blanket, one-switch full autonomy on day one, no matter what the vendor claims.Why: "autonomy" isn't one setting, it's a different bet for every different action type.
  5. Keep a full action log and an undo window even after autonomy is fully granted.Why: staged trust doesn't mean the record-keeping goes away once the training wheels come off.

How to answer this, stage by stage

Seven stages. Say each one out loud, and the answer holds together under a real follow-up.

Stage 1
Scope it to one real agent
Say it like this
"I'll answer this for ShopPilot, an agent that reschedules maintenance and reorders tooling at a CNC shop, and for Soledad, who manages that shop."
Why this works
Grounds an abstract "agent onboarding" question in one real set of actions with real consequences.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what we're actually protecting. Reversibility, which mistakes can't be undone. Dependency, what has to come first. Evidence, what we can learn cheaply. Rank, the actual sequence."
Why this works
Signals a structured sequencing answer instead of a single opinion about trust.
Stage 3
Reframe the question
Say it like this
"Onboarding into an agent isn't teaching someone to read a dashboard. It's deciding, action by action, how much of your own judgment you're willing to hand over before you've seen it prove itself."
Why this works
Separates a real design answer from a generic "build trust gradually" line.
Stage 4
Give the one decision
Say it like this
"Every new shop starts in shadow mode for two weeks, watching what ShopPilot would have done. Then reversible actions, rescheduling a slot, go live first. Auto-submitting a parts order, hard to undo, only earns autonomy after a proven track record on smaller, reversible calls."
Why this works
Concrete, sequenced, and matches the direct answer exactly.
Stage 5
Prove it with a failure
Say it like this
"A peer shop granted full autonomy on day one, one switch, every action type at once. A stale schedule sync caused ShopPilot to auto-submit a duplicate tooling order, and by the time anyone noticed, the supplier had already shipped it."
Why this works
Shows the real cost of treating autonomy as one setting instead of many.
Stage 6
Say what you'd measure
Say it like this
"I'd track how many actions in shadow mode the user would have overridden, if they could have. A high override rate on a reversible action type means it's not ready to graduate yet, no matter how long the shop's used it."
Why this works
Shows the graduation criteria is evidence-based, not just a calendar countdown.
Stage 7
Close on the one line
Say it like this
"Sequence by what's hard to undo, not by feature completeness. Shadow first, reversible next, expensive-to-undo last, and always logged."
Why this works
Restates the ranked decision plainly, ready for any follow-up.

Let's learn

What's the first thing you'd let an agent do on your behalf, before you've ever watched it work?

ShopPilot watches tooling wear sensors and the week's job schedule across a CNC shop, and it can act on both: reschedule a maintenance window, or reorder tooling and consumables before a machine runs dry.

Before ShopPilot, Soledad's team checked tooling wear by hand once a shift and reordered consumables from a running paper tally, about ninety minutes a week across a dozen machines. ShopPilot cuts that to a glance.

The glance was never the risk. The risk is what happens the one time ShopPilot acts on something that can't be quietly undone.

Cost to reverse a ShopPilot action, by type
$3,000 $1,500 0 ~$0 ~$0 $2,400 full wage Reschedule slot Draft PO Auto-submit order Approve overtime
Four action types, four completely different bets. Treating them as one autonomy switch ignores the whole right half of this chart.
Knowledge spark: what's shadow mode? A setting where an agent shows you exactly what it would have done, without actually doing it. It's the only way to see the agent's real judgment on real data before any of it can cost you anything.

At its worst: a shop in the same trade group as Soledad's granted ShopPilot full autonomy the day it launched, one switch, every action type at once. A stale schedule sync caused it to auto-submit a tooling order that duplicated one already placed by hand that morning. By the time anyone noticed, the supplier had shipped both, and the shop ate a 2,400-dollar restocking fee to send half of it back.

The shop didn't lose 2,400 dollars to a bad model. It lost it to a switch that treated "reschedule a slot" and "spend the shop's money" as the exact same kind of decision.
The decision I would take back We treated "grant autonomy" as one all-or-nothing switch: flip it on, and ShopPilot could act on anything from rescheduling a slot to submitting a five-figure tooling order, at the same permission level. That was harmless in the pilot, where it only ever touched scheduling. It stopped being harmless the week the same switch also covered auto-submitting real purchase orders.

What I would leave alone: once an action type has a real track record, rescheduling maintenance after months of accurate, reversible calls, there's no need to keep re-confirming it. The fix isn't permanent caution everywhere, it's sequencing the caution to where the real cost actually sits.

The lesson: "autonomy" sounds like one word, but it's never one decision. Every action type an agent can take is its own separate bet, and a single switch that covers all of them is a decision made once for a dozen different risks.

Now here is the same thing as a story

The short version above is the sequencing argument. This one is how Soledad actually built it, one tier at a time.

The shop floor at Ferro Machine Works goes quiet at 4:30, when the CNC lines finish their last cut of the day and ShopPilot starts drafting tomorrow's schedule and checking tooling levels against the week ahead.

Hand sketched flow diagram titled What unblocks what. Four steps: shadow mode drafts only, track record earned highlighted, reversible actions unlocked, hard to undo actions unlocked.
Nothing downstream of shadow mode can happen honestly without it. It's the one step every later tier depends on.

Soledad ran the pilot for two months with ShopPilot in shadow mode only, watching what it would have rescheduled and reordered without letting it touch anything real. It was right often enough that she started trusting the drafts without reading every line.

Hand sketched comparison diagram titled Reversible, or not. Left panel, a box icon labeled Swings both ways, caption reschedule a maintenance slot. Right panel, a scale icon labeled Bolted shut, caption submit a parts order.
Two doors, and the whole design question is which one a beginner gets handed first.

Then word came through the trade group's messaging channel: a shop two towns over, running the exact same ShopPilot rollout, had flipped full autonomy on for everything at once, on day one, because the vendor's default setup made that the easy path. A stale sync between ShopPilot and their supplier's order system caused it to place a tooling order that duplicated one their own buyer had already called in that morning.

Hand sketched quadrant titled Sorting ShopPilot's actions. Axes how often it happens, and how costly to undo. Reschedule slot and draft PO sit top left, frequent and cheap to undo. Auto-submit order and approve overtime sit bottom right, rare but expensive to undo.
The bottom right corner is small, only a few action types live there, but it's where nearly all the real risk sits.

Soledad hadn't granted ShopPilot autonomy over anything real yet. She heard about the near disaster before it had a chance to become her own.

We did not almost lose 2,400 dollars ourselves. We came close to making the exact same one-switch decision, just a few weeks later than the shop it actually happened to.

Hand sketched timeline titled The staged autonomy rollout. Four milestones: weeks 1-2 shadow mode only highlighted, weeks 3-5 reversible actions live, weeks 6-8 hard to undo actions live, ongoing every tier still logged.
Eight weeks, four tiers. Nothing here required waiting for a perfect model, only a sequence that matched the real cost of each mistake.

Soledad's shop instead unlocked autonomy one action type at a time. Rescheduling a maintenance slot went live first, reversible with one click if wrong. Drafting a purchase order for a human to approve came next. Auto-submitting an actual order, the expensive-to-undo one, only earned autonomy after six straight weeks of accurate, reversible calls on everything else.

Hand sketched labeled parts diagram titled What a good autonomy screen shows. Center person icon labeled Autonomy Settings. Four callouts: what it can do alone, what still needs a nod, an undo window, a full action log.
The second callout is the one a single autonomy switch has no room for at all: some actions still need a nod, even after months of trust.

I built the pilot's single switch because it was the fastest way to demo the product's full range at once. It took a peer shop's near-duplicate order, not a mistake of my own, to see that a switch covering every action type at once was really a dozen separate bets wearing one label.

ORDER, applied to autonomyNot a trust score. ORDER is what forces the sequence to follow what's actually hardest to undo, not what's easiest to demo.

O
Outcome. What every tier is protecting.
Keeping the cost of any single wrong autonomous action small, while still letting ShopPilot save real time.
Without a stated outcome, a sequencing answer is just a personal preference.
R
Reversibility. Which mistakes can't be undone.
Rescheduling costs one click to fix. An auto-submitted order costs a restocking fee and a phone call to a supplier.
The hardest step, and the one the whole sequence turns on.
D
Dependency. What has to come first.
Shadow mode has to exist before any real autonomy tier makes sense; an action log has to exist before any tier can be trusted.
Shows the sequence is forced by reality, not chosen by taste.
E
Evidence. What can be learned cheaply first.
Two weeks of shadow-mode drafts show what ShopPilot would have done, before any of it costs a dollar.
Cheap evidence beats an expensive mistake as a way to learn the same thing.
R
Rank. The actual sequence, defended.
Shadow mode, then reversible actions, then hard-to-undo actions only after a proven track record.
States the order plainly and defends the first tier: nothing skips shadow mode.
Fully-autonomous action volume, staged rollout
40/wk 20/wk 0 reversible live hard-to-undo live Week 1 Week 8
Volume never jumps to full autonomy in one step. Every rise on this line corresponds to a specific action type earning its own track record.

The recap, one line per letter: outcome is keeping any single wrong action cheap, reversibility is the real axis the sequence runs on, dependency is shadow mode and logging coming first, evidence is two weeks of free drafts before any real cost, and rank is the four-tier order Soledad actually shipped.

And if you want to be sure it really works, try it somewhere elseSame five letters, a home-services booking agent instead of a machine shop. This time the hardest-to-undo action is cancelling a job, not spending money.

Mira Castellano is a homeowner using HandyLoop, an agent that books and reschedules home-repair contractors on their behalf, based on quotes and availability it finds automatically.

Mapped onto ORDER: outcome is keeping any single wrong booking decision small and easy to walk back. Reversibility is rescheduling a visit, cheap and easy to undo, versus cancelling a booked job outright, which can cost a contractor's deposit and burn a relationship for future work. Dependency is a two-week shadow period where HandyLoop shows which contractors it would have booked, before booking anything for real. Evidence is watching how often Yuki would have overridden a suggested contractor during that shadow period. Rank is rescheduling first, booking a new, unfamiliar contractor second with a confirmation step, and cancelling an existing job last, always requiring a direct yes.

Hand sketched decision tree titled ShopPilot meets a new action type. Root, new action type seen, branching to four outcomes: cheap and reversible leads to auto-approve, hard to undo and new leads to ask first N times, hard to undo and proven leads to auto-approve logged, totally novel leads to always ask.
The same tree that sorts ShopPilot's tooling orders sorts HandyLoop's cancellations. The action changes. The logic doesn't.
Hand sketched icon list titled HandyLoop's staged autonomy, at home. Four items: a document icon labeled draft only first two weeks, a box icon labeled reschedule a visit alone, a scale icon labeled book a new contractor ask first, a person icon labeled cancel a job always confirm.
Cancelling a job sits at the bottom of this list on purpose. It's the one action a homeowner is least likely to forgive an agent for getting wrong alone.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "shadow mode first, reversible actions next, hard-to-undo actions last," and stop.
Cost: there's no budget to build a full shadow-mode simulation this quarter. Say so honestly, and start with a manual approval step on every action for the first month, since even a slower rollout beats a blanket switch.
The model gets better, for real: if ShopPilot's prediction accuracy improves to near-perfect, the sequencing still matters, a rarer mistake on an expensive-to-undo action is still expensive the one time it happens.

Where people run it wrong.
They treat "grant autonomy" as one on/off switch instead of a separate decision per action type.
They sequence autonomy by which features are built first, instead of by which mistakes are hardest to undo.
They stop logging and confirming once an action type earns full autonomy, instead of keeping the record-keeping in place indefinitely.

How to use it live. When someone asks how you'd onboard a user into an agent that acts on their behalf, ask yourself one question first: which of this agent's actions would be the hardest to undo if it got one wrong. Name that action, and put it last in the sequence, before you say anything else.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you onboard a user into an agent that acts on their behalf"?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Reversibility is the hardest step, and the axis the whole sequence is built on.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Soledad Larrain, who manages Ferro Machine Works, a CNC shop where ShopPilot reschedules maintenance and reorders tooling.
3 · THE OUTCOME
What is every autonomy tier actually protecting?
Tap to flip
ANSWER
Keeping the cost of any single wrong autonomous action small, while still letting the agent save real time for the shop.
4 · THE SEQUENCE
What's the four-tier order this answer commits to?
Tap to flip
ANSWER
Shadow mode first, drafts only. Then reversible actions like rescheduling. Then hard-to-undo actions like auto-submitting an order, only after a proven track record.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating "grant autonomy" as one all-or-nothing switch covering every action type at the same permission level, harmless in a scheduling-only pilot, costly once it also covered real purchase orders.
6 · THE NUMBER
Fill in the blank: the peer shop's duplicate tooling order cost about ___ dollars in restocking and cancellation fees.
Tap to flip
ANSWER
2,400 dollars, caused by a stale schedule sync under a one-switch, full-autonomy setup.
7 · THE REPLAY
Same stale-sync bug, staged rollout in place. What changes?
Tap to flip
ANSWER
Order-submission autonomy hasn't been earned yet in week two, so the duplicate order surfaces as a draft awaiting approval instead of shipping automatically.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the hardest-to-undo action there?
Tap to flip
ANSWER
HandyLoop, a home-services booking agent. There, the hardest-to-undo action is cancelling an already-booked job, not spending money directly.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the peer shop's duplicate tooling order cost about ___ dollars to reverse.
Show hint
Look at the grouped bar chart of cost to reverse each action type.
Show answer
2,400 dollars. Rescheduling a slot, by contrast, costs about zero to undo.
Multiple choice
2. According to this answer, what should decide the order in which an agent earns autonomy over different actions?
  • A. Whichever feature the engineering team finished building first.
  • B. How hard each action type is to undo if the agent gets it wrong.
  • C. Whichever action the user asks for most often.
  • D. A fixed thirty-day calendar countdown, the same for every action type.
Show hint
Look at the reversibility step and the quadrant diagram.
Show answer
B. Reversibility is the axis the entire sequence is built on, not feature completeness or a fixed timeline.
True or false
3. True or false: this answer recommends granting full autonomy to a new user on their first day, since ShopPilot's model is accurate.
  • True
  • False
Show hint
Look at the priority list and the staged-rollout timeline.
Show answer
False. Every user starts in shadow mode, with real autonomy earned tier by tier over roughly eight weeks, regardless of the model's accuracy.
Short answer, where it wouldn't matter
4. Name an action type at Ferro Machine Works where this careful sequencing matters less, and why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Rescheduling a maintenance slot, once it has a long track record. It's cheap to undo, so re-confirming it forever adds friction without reducing any real risk.
Short answer, apply it yourself
5. Think of an app or assistant that's ever taken an action on your behalf without asking first. Was that action easy to undo, or not, and would you have wanted a say first?
Show hint
Think of an auto-renewed subscription, an autocorrected message that sent, or an app that rebooked something automatically.
Show answer
Model answer: Many people can recall an auto-renewal or an autosent message that was hard to walk back, exactly the kind of action this answer says should earn autonomy last, not first.
Before you close the answer
Why this works
Tests whether you can treat "autonomy" as a set of separate, sequenced decisions instead of one trust dial, and whether you'll design the sequence around what's actually irreversible instead of what's easiest to demo.
Follow-up traps
"Doesn't shadow mode just delay the value the user is paying for?" Response: two weeks of delay on a multi-year relationship is a small cost next to a 2,400-dollar mistake in week one, and shadow mode still shows the user real value, just without real risk yet.

"What if a user wants to skip the staged rollout and turn everything on immediately?" Response: let them see the reversibility chart before they decide, the same one Soledad saw. Most people choose the staged path once the actual cost difference is in front of them, not behind a settings toggle.
If pressed
ShopPilot's real graduation criteria isn't just time elapsed, it's a minimum number of shadow-mode drafts the user would not have overridden, so a shop that uses the system rarely doesn't graduate to autonomy just because a calendar date passed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more