CaseAdvancedDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #17

How do you re-onboard existing users when the AI behaviour changes significantly?

FLIPS the product is Loomstaff, a freelance-hiring marketplace, and ScoutMatch is the AI that ranks candidate matches

Loomstaff connects small studios with freelance illustrators and motion designers. ScoutMatch ranks incoming applicants against a job post. Noor Delacroix runs a five-person design studio, and posts a new freelance job through Loomstaff most weeks.

The direct answer
When ScoutMatch's model changes enough that it wants different input than before, don't announce it. Show existing power users a personal before-and-after, built from their own last few posts, the moment the change ships. One worked example of what to change beats any changelog, because it corrects the actual habit instead of just describing that something is different.
Do this, in order
  1. Show existing users a personal before-and-after the moment a significant model change ships.Why: it fixes the actual gap, no way to compare old and new, that let the failure happen silently.
  2. Trigger this only for changes big enough to want different input, not every minor tune.Why: small updates don't break anyone's habits and don't deserve the ceremony.
  3. Aim it at the users whose habits are most shaped by the old behavior, not everyone equally.Why: power users like Noor built the deepest habits around the old model, so they have the most to lose.
  4. Use one real worked example from the user's own history, not a general description of the update.Why: an abstract changelog doesn't fix a habit; a concrete side-by-side score does.
  5. Watch for quiet declines in feature use, not just support tickets, after any model change.Why: abandonment never files a ticket. It just stops showing up in the numbers, late.

How to answer this, stage by stage

Nobody is grading whether you can define "re-onboarding." They're grading whether you can name the exact gap a silent model update leaves open.

Stage 1
Scope it to one product, one user
Say it like this
"I'll answer this for Loomstaff's ScoutMatch feature, and for Noor, a small studio owner who hires through it every week."
Why this works
Keeps the answer from turning into a generic "communicate changes to users" essay.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a method, not just a story you happen to remember.
Stage 3
Name the habit before the change
Say it like this
"Noor learned that short, keyword-heavy job posts scored just as well as long ones under the old model. So she stopped writing full creative briefs and started writing three-word posts."
Why this works
The habit is what the update actually breaks, not the model's raw accuracy number.
Stage 4
Identify the flip
Say it like this
"She goes from trusting the ranked match list every Monday to quietly ignoring it and scrolling every applicant by hand. No ticket, no complaint. She just stops using the ranking."
Why this works
Names the exact two-setting switch: trusts the list, or scrolls everyone herself. No middle ground.
Stage 5
Give the reversal
Say it like this
"We shipped the new model as a silent backend swap, with nothing telling existing users what changed. I'd replace that with a personal before-and-after, shown the day it ships."
Why this works
This is the concrete, defensible answer to the actual question being asked.
Stage 6
Replay it, and close
Say it like this
"With that comparison in front of her, Noor adjusts her posting habit in one morning instead of quietly giving up on ranking for three weeks. Re-onboarding isn't an email. It's a mirror held up to the user's own history."
Why this works
Ends on something countable, one morning versus three weeks, and restates the decision one last time.

Let's learn

Here is what happens when a tool works fine for months, changes for a good reason, and nobody who depends on it is told.

Loomstaff is a freelance-hiring marketplace, and ScoutMatch is the AI feature that ranks who applies to your job post.

Before ScoutMatch existed, Noor read every one of the 25 to 30 applicants on a typical post by hand, about 45 minutes of screening most Mondays. Once ScoutMatch's first model launched and proved itself, she trusted the top five ranked matches and picked from those in under six minutes.

Knowledge spark: what's a model migration? Swapping the AI model behind a feature for a newer one, usually trained differently or on new data. It can raise average quality a lot, and still behave very differently on inputs the old model liked, since it learned different patterns.

Loomstaff then migrated ScoutMatch to a model trained on real hire outcomes rather than keyword overlap, and platform-wide match quality genuinely improved. But Noor's terse, three-word posts, the exact habit the old model rewarded, are now the worst possible input for a model that wants richer context to reason about fit.

Noor's personal match quality, six weeks around the migration
100% 50% 0 model v2 ships Wk 1 Wk 3 Wk 6
Her own hire-satisfaction rate slid from 80% to 50% over three quiet weeks, with nothing on screen marking the moment it started.

At its worst: platform-wide match quality actually rose after the migration, while Noor's own results kept sliding, so the average dashboard leadership watched looked healthy the entire time she was losing trust in the tool.

The decision I would take back We shipped model updates as a silent backend swap, with nothing on screen telling existing users the model itself had changed. That made sense when updates were small tweaks nobody needed to notice. It stopped making sense once an update changed what kind of input the model actually wanted, which is exactly the case that punishes an established habit.

What I would leave alone: small, routine tuning updates, the kind that shift results by a point or two without changing what input the model rewards, don't need a comparison moment. Flagging every minor update would train users to ignore the flag entirely.

We did not just make the model better. We made three weeks of Noor's mornings worse, and told her nothing while it happened.

The lesson: a model update that raises the average can still break a specific person's habit completely. "Better on average" is not the same claim as "safe for everyone who already learned how the old version worked."

Now here is the same thing as a story

The short version above is what you'd say defending this decision to Loomstaff's product leadership. Read this one for how it actually happened to Noor.

Loomstaff's dashboard opens to a fresh stack of applicants most Monday mornings, right around eight, coffee still too hot to drink. That's when Noor does her hiring for the week.

Hand sketched flow diagram titled Noor's habit, before the flip. Four steps: write full brief, matches score well, trust grows, posts get terser.
Nobody told her to write terser posts. The old model just kept rewarding it, one Monday at a time, until it was simply how she worked.

For nearly a year, that habit served her well. She'd post three words, "motion designer, 2d, kids brand," and the ranked list handed her someone right for the job almost every time.

Then Loomstaff migrated ScoutMatch to a model trained on real hire outcomes instead of keyword overlap. It was, on the whole, a real improvement. Nobody at Loomstaff told existing users, because nothing about the interface had changed.

Hand sketched timeline titled The migration, marked. Three milestones: model v2 ships with no notice sent, matches slip silently over three weeks, she quits ranking and scrolls all applicants by hand.
The gap between the first milestone and the third one is the whole story. Nothing marked the middle of it as it happened.

Her ops manager mentioned it first, almost in passing: "Are you sure Loomstaff's still worth the subscription? These matches feel off lately." Noor had noticed too, but hadn't said anything, assuming it was just a run of unlucky posts.

Hand sketched comparison diagram titled Small move, big snap. Left panel, a gauge icon labeled Match score, caption drifts down gradual. Right panel, a circle icon labeled Her trust, caption fine then gone.
The score drifted. Her trust didn't drift. It held, and then it was simply gone.

That week, Noor stopped opening the ranked list at all. She went back to reading all 28 applicants by hand, the way she had before Loomstaff existed, except now she was also paying a marketplace fee for a ranking feature she no longer used.

Hand sketched metaphor scene titled What we assumed vs what she has. Left, a gauge icon labeled A DIAL, caption we assumed this. Right, a circle icon labeled A SWITCH, caption trust or none.
We designed the update assuming trust was a dial that would settle back into place. For Noor, it was a switch, and it only moved one way.

The old design assumed a model update was an internal detail, invisible by default. The new design we should have shipped treats a significant update as a moment that belongs to the user, not just the engineering team.

We built the silent-update habit because early on, updates really were small and invisible enough not to matter. It took watching Noor's ops manager notice before we did to see that "the model changed" and "nothing changed for the user" had quietly stopped being the same claim.

Five letters, and where the silence broke themNot a new model. A new kind of input it wanted, with nobody around to say so.

Hand sketched icon list titled FLIPS, the five letters. Five items: find the person, locate the habit, identify the flip, pinpoint the old call, show the replay.
The memory aid. The gap that broke Noor's Monday sits entirely inside letter three.
F
Find the person.
Noor Delacroix, hiring weekly through Loomstaff, trusting ScoutMatch's ranked list every Monday morning.
Grounds the whole answer in one real person's routine.
L
Locate the habit.
She stopped writing full creative briefs, since terse, keyword-heavy posts scored just as well under the old model.
Names the exact thing the update was about to punish.
I
Identify the flip.
Trusting the ranked list every Monday, to quietly scrolling every applicant by hand and never opening the ranking again.
The hardest step, and the one that names the actual two-setting switch.
P
Pinpoint the old decision.
Shipping the model migration as a silent backend swap, with no comparison point for existing users.
Names the specific, reversible choice that let the change land silently.
S
Show the replay.
With a personal before-and-after shown at launch, Noor adjusts her posting habit in one morning instead of three silent weeks.
Proves the reversal actually restores what the silent update took away.
Hire-satisfaction rate: platform average vs Noor, before and after v2
100% 50% 0 71% 79% Platform average 80% 50% Noor, personally
The average went the right direction. Noor's own number went the opposite way, and the average alone would never have shown anyone that.

The recap, one line per letter: find the person is Noor's Monday routine, locate the habit is her terse three-word posts, identify the flip is trusting the ranked list versus scrolling every applicant herself, pinpoint the old decision is the silent model swap, and show the replay is the personal before-and-after cutting three weeks down to one morning.

And if you want to be sure it really works, try it somewhere elseSame five letters, a B2B invoicing platform instead of a hiring marketplace. Different flip family this time, delegation instead of abandonment.

Fenwick Ledger is a B2B invoicing platform. Priya Sandoval is a senior collections specialist there, and used to write every overdue-payment reminder herself before an AI drafting tool let her hand first-drafts to two junior reps.

Mapped onto FLIPS: find the person is Priya, who delegated reminder drafts to junior reps once the AI's tone proved reliably firm and professional. Locate the habit is that the juniors stopped learning to write firm reminders themselves, since editing a good draft is faster than writing one from scratch. Identify the flip: Fenwick Ledger updates the drafting model to sound warmer, as part of a customer-empathy push, and the juniors, who never built the skill to firm up a soft draft, send the new gentler wording unedited. Overdue accounts go quiet. Priya notices two accounts stall and takes review back herself, checking every reminder before it goes out, undoing the delegation entirely. Pinpoint the old decision: the tone update shipped as a general style improvement, without flagging that it changed assertiveness specifically for overdue-collections use cases. Show the replay: with a comparison shown at launch, a real overdue-notice draft before and after, plus a toggle to keep the firmer tone for collections specifically, Priya spot-checks three drafts a week instead of reclaiming all of them.

Hand sketched quadrant titled Fenwick Ledger, sorting the re-onboard cases. Axes how senior and how much they delegated. Priya before sits senior and fully delegated. Priya after sits senior and little delegated. New hire sits junior and little delegated.
Priya didn't become more junior. She moved from fully delegating to barely delegating, and that's a different kind of loss than a skill gap.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show existing users a personal before-and-after the moment a significant model change ships," and stop.
Cost: there's no engineering time this quarter to build a live comparison screen. Say so, and start with a manually-curated email showing three real before-and-after examples pulled from the top twenty affected accounts.
The model gets better, for real: even when the average genuinely improves, as it did here, a specific user's habit can still be punished by exactly the change that helped everyone else.

Where people run it wrong.
They treat a model update as purely a backend concern, invisible by design, even when it changes what input the model wants.
They send a generic changelog email instead of a comparison built from the user's own history.
They watch the platform-wide average and miss the specific users whose numbers moved the opposite direction.

How to use it live. When someone asks how to re-onboard users after a big AI change, ask yourself one question: what would this person need to see about their own history to trust the new version as much as the old one? Build that, not an announcement.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Abandonment flip: uses it daily, then quietly stops opening it. No ticket, no complaint, just a slow decline that looks like seasonality on a dashboard.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Noor Delacroix, who runs a five-person design studio and hires freelancers through Loomstaff most weeks.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Writing full creative briefs in her job posts. Terse, keyword-heavy posts scored just as well under the old model.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Trusting the ranked match list every Monday, versus quietly scrolling every applicant by hand and never opening the ranking again.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping the model migration as a silent backend swap, with nothing telling existing users what kind of input the new model actually wanted.
6 · THE NUMBER
Fill in the blank: Noor's personal hire-satisfaction rate fell from 80% to ___ over three weeks.
Tap to flip
ANSWER
50%. Meanwhile the platform-wide average rose from 71% to 79% over the same period.
7 · THE REPLAY
Same migration, redesigned launch. What changes?
Tap to flip
ANSWER
Noor sees a personal before-and-after the day it ships, and adjusts her posting habit in one morning instead of quietly abandoning the ranked list over three weeks.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and which flip family?
Tap to flip
ANSWER
Fenwick Ledger, a B2B invoicing platform. There, the flip family is delegation, not abandonment: a senior specialist takes work back from junior reps once a tone update quietly breaks their unedited drafts.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer recommends flagging every single model update, no matter how small, with a full comparison screen.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Small, routine tuning updates don't need this. Flagging every minor update would train users to ignore the flag entirely.
Multiple choice
2. What was the flip in Noor's story, described correctly?
  • A. Her match scores dropped from 80% to 50%.
  • B. She went from trusting the ranked list every Monday to quietly scrolling every applicant by hand.
  • C. She started writing longer job posts again.
  • D. She switched to a competing marketplace.
Show hint
The flip is a behavior, not a metric.
Show answer
B. The score dropping is the cause; what she actually did about it is the flip. That's the two-setting switch: trust the list, or check everyone yourself.
Fill in the blank
3. Fill in the blank: it took about ___ weeks after the model migration before Noor stopped using the ranked list entirely.
Show hint
Look at the timeline diagram of the migration.
Show answer
Three. Three silent weeks of sliding results, with no signal anywhere that the model itself had changed.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Shipping model updates as a silent backend swap. It made sense while updates were small tweaks nobody needed to notice.
Short answer, why no middle setting
5. Why couldn't Noor have just "checked her matches a bit more carefully" instead of abandoning the ranked list entirely?
Show hint
Think about what "checking a bit more" would actually require her to do.
Show answer
Model answer: Checking a ranked list carefully still means trusting the ranking is roughly right and double-checking the top few. Once she stopped trusting it at all, the only honest option left was screening every applicant, which is a different, more expensive behavior entirely, not a dial turned up.
Short answer, apply it yourself
6. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if its underlying behavior changed a lot overnight?
Show hint
Think about a recommendation feed, a spam filter, or an autocomplete tool you've stopped double-checking.
Show answer
Model answer: Many people stop reading a spell-checker's suggestions closely once it's been reliable for months. If it suddenly started making different kinds of mistakes, most people would keep trusting it for a while before noticing the shift, exactly the gap this answer is about.
Before you close the answer
Why this works
Tests whether you understand that a model update is a UX event for existing users, not just a backend deployment, and whether you can design for a person whose habits are tied to the model's old behavior specifically.
Follow-up traps
"Isn't a personal before-and-after expensive to build for every user?" Response: target it at users whose recent behavior most resembles the old model's reward pattern, not everyone, which keeps the build small and the users who need it most get it.

"What if the model update is actually a strict improvement for this user too?" Response: still worth a lightweight version, since even a real improvement can look like a regression without a comparison point, and a quiet unexplained change costs trust either way.
If pressed
Loomstaff's real trigger for this flow compares each user's post style against a simple heuristic, average words per post, against the shift in what the new model weights most, so the comparison only fires for users actually at risk of the mismatch, not the whole user base.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more