ConceptIntermediateDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #7

How much should you tell users about which model powers a feature?

GUARD the product is WorkLatch, a gig-work marketplace, and the matching model behind its job feed

WorkLatch matches gig workers to available shifts and jobs. To control cost during busy hours, it sometimes quietly switches to a cheaper matching model. Denholm Ashworth, a WorkLatch product manager, approved that switch. Priska Boateng works full-time through WorkLatch and logs in every morning at 6am.

The direct answer
The real question isn't how much branding detail to share about a model's name. It's whether the person most affected by a quiet quality change has any way to notice it or push back. Tell users, plainly, when the system is running in a lower-quality mode, especially the users who depend on it for income, even if you never name which model is behind it.
Do this, in order
  1. Show a plain signal when match quality drops, before naming any model at all.Why: the harm isn't which model runs. It's a worker having no way to know their results just got worse.
  2. Protect full-time, high-dependency workers from the fallback first.Why: they're the ones who lose real income, not just a little convenience, when quality drops.
  3. Give every worker a way to check their own recent quality tier.Why: without it, a bad week reads as personal bad luck instead of a system decision.
  4. Track fallback exposure by worker-dependency segment in production.Why: waiting for complaints misses the workers least likely to complain and most affected.
  5. Skip naming the specific model by default; most users don't need that detail to be treated fairly.Why: the fix that actually matters is about impact, not branding.

How to answer this, stage by stage

Nobody's grading whether you'd disclose a model's name. They're grading whether you noticed that two very different people are asking two very different questions about "which model."

Stage 1
Scope it to one real system
Say it like this
"I'll answer this for WorkLatch's job-matching model, and the cost-driven fallback that sometimes swaps in a cheaper version during peak hours."
Why this works
Grounds an abstract disclosure question in one real backend decision with real consequences.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, unequal impact, ability to contest, reduce, and detect."
Why this works
Signals you're going to name the power imbalance directly, not just list disclosure best practices.
Stage 3
Name both groups
Say it like this
"Denholm decides when the fallback activates. Priska has no idea it exists, and no way to know when her matches came from it."
Why this works
GUARD's core move: naming the operator and the subject on the same page, not just "users" in the abstract.
Stage 4
Say where the harm lands unevenly
Say it like this
"Casual weekend users barely notice a quality dip. Full-time workers who log in at the exact peak hours the fallback runs lose real income, and they're the ones who need WorkLatch working well the most."
Why this works
Shows the unequal impact lands hardest on exactly the people least able to absorb it.
Stage 5
Ask who can't push back
Say it like this
"Priska can't inspect, appeal, or opt out of a fallback she's never been told exists. She just sees a bad week and assumes something's wrong with her, not the system."
Why this works
This is GUARD's strongest move: naming the exact person with no lever at all.
Stage 6
Give the reduce and detect steps
Say it like this
"Reduce: protect high-dependency workers from the fallback first, and show a plain quality-tier signal. Detect: track fallback exposure by worker-dependency segment, so we catch this in our own data before a worker has to complain."
Why this works
Turns the concern into a specific product change and a specific way to catch it happening again.
Stage 7
Close on the one line
Say it like this
"It's not about naming the model. It's about whether the person who depends on the system most has any way to know when it's quietly working worse for them, and right now, she doesn't."
Why this works
Restates the direct answer, ready for whatever gets pushed back on next.

Let's learn

WorkLatch's matching model looks at a gig worker's history, location, and availability, and suggests the shifts most likely to be a good fit. During peak system load, a cheaper fallback model quietly takes over to control inference cost, then hands back to the primary model once load drops.

The fallback runs briefly, usually under ten minutes, and was built to be invisible on purpose, so no worker would see a jarring change mid-session.

Knowledge spark: what's an inference fallback? A cheaper, faster backup model a system switches to automatically when the main model would cost too much or take too long under heavy demand. It's a real cost tool. It's also a real quality change, whether anyone announces it or not.
Weekly earnings during fallback-heavy weeks, by worker type
-10% -5% 0% -9% Full-time workers -1% Casual weekend workers
The same fallback, the same brief window, and a nine-times bigger hit for the workers who can least afford it.

The turn: the fallback itself isn't the real problem. Systems need a cheap backup for peak load, and that's a reasonable engineering call. The real problem is that the fallback was built to be invisible to everyone equally, when its actual cost lands on almost nobody equally.

A fallback built to be invisible on purpose is invisible to the casual user and the full-time worker in exactly the same way, even though it costs them completely different things.

At its worst: a full-time worker has three straight weeks of below-average earnings during a stretch when the fallback happened to run every morning during her usual login window. She assumes her account is being deprioritized, or her rating dropped, and starts working extra unpaid hours polishing her profile to compensate, chasing a problem that was never about her at all.

The decision I would take back We set the fallback to trigger uniformly based on system load, with no consideration of which specific workers were logged in during that load, because building it that way was operationally the simplest option. That made sense while the fallback ran rarely and briefly. It stopped making sense once peak hours expanded, and high-dependency workers, who log in during exactly those hours out of necessity, ended up disproportionately exposed.

What I would leave alone: for a casual user who logs in occasionally for weekend gig money, an undisclosed brief fallback genuinely doesn't matter much, and announcing every backend cost decision to them would just be noise nobody asked for.

The lesson: "how much to disclose" is the wrong first question. The right one is who bears the cost of not disclosing, and whether that person has any way to notice or object.

Now here is the same thing as a story

The short version above is what you'd say defending this fix to WorkLatch's leadership. Read this one for how the gap actually got found.

Denholm Ashworth manages WorkLatch's matching system. He approved the fallback policy eighteen months ago, confident it was a clean, sensible way to control cost without anyone noticing a thing.

Hand sketched timeline titled The fallback policy's quiet rollout. Three milestones: fallback built rare and brief, peak hours expand fallback more common highlighted, Priska's bad weeks no explanation given.
Nothing about the fallback itself changed. What changed was how often, and for whom, it was running.

The trigger was a near miss, not a disaster. An engineer proposed making the fallback trigger even more aggressively, to save further on cost during a projected demand spike. Before approving it, Denholm asked for the actual usage data behind the existing fallback, mostly out of caution, expecting nothing surprising.

What he found was Priska Boateng's login pattern, repeated across hundreds of other full-time workers: logged in every weekday at 6am, exactly the hour the fallback ran most often, because that's when the best shifts were posted and full-time workers needed to be first in line.

Hand sketched comparison diagram titled Two people, one lever. Left panel, a person icon labeled Denholm, the PM, caption controls the fallback switch. Right panel, a person icon labeled Priska, gig worker, caption no way to see or object.
Same feature, same switch. One of these two people knew it existed.

Priska herself had never filed a complaint. She'd had three unusually thin weeks a few months back, blamed her own rating, and quietly started spending an extra hour each morning polishing her profile before logging in, a habit she kept up long after those weeks ended.

Match quality score across a typical peak morning
100 50 0 4am: 92 6am: 71 7am: 92
The exact forty minutes Priska logs in every single day are the exact forty minutes the fallback runs most often.
Priska never got a worse account. She got the same forty minutes, every single morning, at exactly the moment the system quietly ran cheaper.

Here's the decision I'd take back. The fallback's trigger was built to respond to system load alone, with no view into who was logged in when it fired. That was a clean, simple rule to write. It just happened to hit the same specific workers every single day, because their schedule wasn't random. It was the schedule of someone depending on this for a living.

Hand sketched flow diagram titled Where the appeal should be, and isn't. Four boxes: load rises, fallback activates, match quality drops, worker never told highlighted.
Four steps, and the last one is the only one that was ever a choice rather than an engineering fact.

I'd add two things. First, a plain quality-tier signal, visible any time it changes, so a worker knows a bad stretch is the system, not them. Second, route the fallback away from workers flagged as high-dependency first, and only onto casual accounts, before it ever touches someone like Priska.

Hand sketched labeled parts diagram titled What a quality tier indicator needs. Center gauge icon labeled Quality Tier, with four callouts: current tier shown, why it changed, who is protected first, check your history.
None of this requires naming a single model. It just requires telling Priska what her own screen is actually showing her.

Replay the same peak morning under the new routing: the fallback activates for casual accounts first, and only reaches full-time workers like Priska if load stays high after that. Her quality score barely moves. She stops the extra hour of profile-polishing, because there was never anything wrong with her profile to begin with.

We built the fallback to be invisible because invisible felt like the safest, most respectful choice at the time. It took one worker's login pattern, pulled almost by accident during an unrelated review, to see that invisible to everyone equally is not the same thing as fair to everyone equally.

GUARD, naming who can't push backNot a policy memo. GUARD is what forces you to put both people, the one with the switch and the one without, on the same page.

G
Groups. Name the operator and the subject.
Denholm decides when the fallback runs. Priska has never been told it exists.
Puts a real name on both sides of the decision, not "the system" and "users."
U
Unequal. Where the harm actually lands.
Full-time workers lose 9% of weekly earnings during fallback-heavy weeks. Casual users lose about 1%.
Shows the harm concentrates on the people who can least absorb it.
A
Ability to contest. Who can't push back.
Priska can't inspect, appeal, or opt out of a fallback she's never been told exists. She just sees a bad week.
The hardest step, and the one that names the actual harm plainly.
R
Reduce. The specific design change.
A plain quality-tier signal, and routing the fallback onto casual accounts before high-dependency ones.
A real product decision, not a policy statement about transparency.
D
Detect. How you'd catch it in production.
Track fallback exposure by worker-dependency segment, so the pattern surfaces in data before a worker ever has to complain.
Catches the next version of this problem before it needs a near miss to find it.
Hand sketched quadrant titled Who actually gets exposed to the fallback. Axes depends on WorkLatch from casual to full-time, and exposure to the fallback from rarely active at peak to logs in at peak daily. Casual weekend only sits low on both. Full-time peak login sits high on both, the dangerous corner. Full-time off-peak login sits high dependency, low exposure. Casual peak hours sits low dependency, high exposure.
The dangerous corner isn't "uses the app a lot." It's depending on it and being exposed to the fallback at the same time.

The recap, one line per letter: groups is Denholm and Priska, unequal is the nine-times gap in real income impact, ability to contest is Priska having no way to even know the fallback exists, reduce is protecting high-dependency workers first, and detect is tracking exposure by segment before a complaint ever arrives.

And if you want to be sure it really works, try it somewhere elseSame five letters, an online tutoring marketplace instead of gig work. A different building, and the subject is a student, not a worker.

An online tutoring marketplace uses a matching model to pair students with tutors, and runs the same kind of cost-driven fallback during peak evening hours. Mapped onto GUARD: groups is the platform's engineering lead who approved the fallback, and a student studying for a make-or-break exam who books sessions every evening at 7pm, exactly peak time. Unequal is that a casual, occasional student barely notices a slightly less ideal tutor match, while the exam-prep student, matched repeatedly during the fallback window, gets a string of weaker tutor fits during the exact weeks that matter most. Ability to contest is the student having no way to know a session was matched by the fallback rather than the primary model, so a string of bad sessions reads as bad luck, or worse, as her own fault for not asking better questions. Reduce is routing the fallback toward casual, low-stakes bookings first. Detect is tracking match quality by whether a student has an upcoming exam flagged in their profile, not waiting for a support ticket.

Hand sketched icon list titled What should always be disclosed, reused here for the tutoring marketplace. Three items: when match quality is degraded, who gets protected first, how to check your own tier.
A gig-work app and a tutoring marketplace look nothing alike. The same three disclosures protect the same kind of person in both.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "protect who depends on it most, and let them see when quality drops," and stop.
Cost: there's no engineering time to build dependency-based routing this quarter. Start with the cheap fix: a plain quality-tier indicator, even before the routing logic changes.
The model gets better, for real: if the fallback model itself improves, that's still not a reason to skip disclosure. A better fallback can still be worse than the primary model, and the person depending on the difference still deserves to know it happened.

Where people run it wrong.
They treat "how much to disclose" as a branding question about naming the model.
They build invisibility as a single setting applied equally to every user, mistaking equal treatment for fair treatment.
They wait for complaints instead of checking whether the people least likely to complain are the ones being hit hardest.

How to use it live. When someone asks how much to tell users about which model powers a feature, don't reach for a branding rule. Ask who's affected when that model quietly gets worse, and whether that specific person has any way to notice or say something about it.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how much should you tell users about which model powers a feature"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real issue is a power imbalance between who controls a model swap and who bears its cost.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Denholm Ashworth, the WorkLatch PM who approved the fallback policy, and Priska Boateng, a full-time gig worker affected by it with no way to know.
3 · THE UNEQUAL IMPACT
How much more do full-time workers lose during fallback-heavy weeks, compared to casual workers?
Tap to flip
ANSWER
About nine times as much: a 9% weekly earnings dip for full-time workers, versus roughly 1% for casual weekend workers.
4 · WHO CAN'T PUSH BACK
Why can't Priska contest the fallback's effect on her income?
Tap to flip
ANSWER
She has never been told the fallback exists, so a bad week reads as personal bad luck or a rating problem, not a system decision she could appeal.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting the fallback to trigger uniformly by system load alone, with no view into which specific workers were exposed to it.
6 · THE NUMBER
Fill in the blank: match quality drops from 92 to ___ during the fallback window each morning.
Tap to flip
ANSWER
71. And that window lands at exactly 6am, Priska's daily login time.
7 · THE REPLAY
Same peak morning, redesigned routing. What changes?
Tap to flip
ANSWER
The fallback routes to casual accounts first. Priska's quality score barely moves, and she stops the extra hour of unnecessary profile-polishing.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and who's the subject there?
Tap to flip
ANSWER
An online tutoring marketplace. There, the subject is a student studying for a make-or-break exam, matched during the same peak-hour fallback window.

Check yourself Score: 0 / 0

Multiple choice
1. What does this answer say is the real question behind "how much to disclose about which model powers a feature"?
  • A. Whether the model's exact name should appear in the app's settings menu.
  • B. Whether the person most affected by a quality change has any way to notice or object to it.
  • C. Whether investors are told which vendor supplies the model.
  • D. Whether the model's training data is published.
Show hint
Look at the direct answer at the top of the page.
Show answer
B. The branding question is a distraction from the real one: does the affected person have any lever at all.
True or false
2. True or false: this answer recommends telling every WorkLatch user the specific name of the fallback model.
  • True
  • False
Show hint
Look at priority list item 5.
Show answer
False. It recommends a plain quality-tier signal instead, without needing to name any specific model.
Fill in the blank
3. Fill in the blank: full-time workers lose about ___% of weekly earnings during fallback-heavy weeks.
Show hint
Look at the first grouped bar chart.
Show answer
9%. Casual weekend workers, by contrast, lose only about 1% over the same weeks.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Triggering the fallback by system load alone, with no view into who was exposed. It made sense while the fallback ran rarely, before peak hours expanded.
Short answer, where it wouldn't matter
5. Name a kind of WorkLatch user for whom the undisclosed fallback genuinely doesn't matter much.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A casual user who logs in occasionally for weekend gig money, where a brief, rare quality dip barely registers.
Short answer, apply it yourself
6. Pick a platform you depend on for income or something important. Is there any backend decision it could make quietly that would affect you more than a casual user of the same platform?
Show hint
Think about ranking or recommendation systems where being seen or matched depends on a model you can't inspect.
Show answer
Model answer: Many people point to seller or creator platforms, where a quiet ranking-algorithm change can affect a full-time seller's income far more than an occasional hobbyist's.
Before you close the answer
Why this works
Tests whether you can look past the literal wording of "which model" and find the actual power imbalance underneath, and whether you'd design for the person least able to notice or object, not just the average user.
Follow-up traps
"Isn't protecting some workers from the fallback first just moving the cost onto casual users instead?" Response: yes, deliberately. Casual users lose almost nothing from an occasional dip, while full-time workers lose real income, so the same cost lands where it actually does the least damage.

"Won't a quality-tier signal just make workers anxious even when nothing's wrong?" Response: the signal only shows when quality has actually changed; the anxiety already exists today, aimed at the worker's own rating instead of the real cause.
If pressed
WorkLatch's real dependency flag isn't self-reported; it's inferred from login consistency and hours worked over the prior eight weeks, since asking workers to self-identify as "full-time" produced under-reporting from people worried it would affect their standing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more