Artifact critiqueFoundationalDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #4

Design the disclosure that tells a user they are talking to an AI.

AUDIT the product is Quillwave Bank's app, and Sable, its in-app chat assistant for account questions

Quillwave Bank is a digital-first bank. Sable is the chat assistant that answers account questions inside its app. Marguerite Salou runs quarterly trust audits for Quillwave's product team, reviewing screen recordings of real customer sessions.

The direct answer
Don't trust a disclosure until you've tested whether a real user, asked cold right after the conversation, can say they were talking to an AI. Sable's current disclosure, a small caption seen once at onboarding, reports 98% acknowledgment. A cold comprehension test found only 34% could actually say so afterward. The gap between those two numbers is the entire audit.
Do this, in order
  1. Test comprehension in the moment it matters, not at signup months earlier.Why: a tap-through at onboarding measures curiosity, not whether someone remembers it during a real conversation.
  2. Ask who wrote the current disclosure and whether it was ever user-tested at all.Why: a caption nobody tested is a guess wearing a compliance checkbox.
  3. Say "AI" in plain words in the first sentence of a chat, not a small badge nearby.Why: a badge can be present on screen and still never be read.
  4. Name what Sable can't do, next to the disclosure, not buried in a help article.Why: knowing it's a machine matters less than knowing where the machine's edges are.
  5. Re-run the comprehension test whenever the app or the caption changes.Why: a disclosure that worked when Sable was new can quietly stop working once it's routine.

How to answer this, stage by stage

Nobody's grading whether you can write a catchy disclosure line. They're grading whether you can tell a real disclosure from one that only looks like it works on a dashboard.

Stage 1
Scope it to one real screen
Say it like this
"I'll answer this for Sable, Quillwave Bank's in-app chat assistant, and the caption that's supposed to tell someone it's not a human."
Why this works
Turns a design prompt into a real artifact you can actually inspect and critique.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who made it, uncover what it's actually tested against, demand the version and date, isolate what's missing, then test it myself."
Why this works
Signals you're going to check the claim, not just admire or redesign the copy on faith.
Stage 3
Ask who wrote it, and whether it was tested
Say it like this
"The current caption was written by the launch team eighteen months ago. Nobody's asked since then whether a real user actually notices it."
Why this works
Most candidates skip straight to redesigning the words. This step asks whether anyone ever checked the old ones.
Stage 4
Uncover what "acknowledged" really measures
Say it like this
"The reported 98% comes from a one-time onboarding modal tapped through at first app open. It was never measured again, at the actual moment someone opens a chat."
Why this works
A metric is only trustworthy once you know exactly what moment it was measured at.
Stage 5
Isolate what's missing
Say it like this
"The word 'AI' never appears in the chat's first sentence. There's no line about what Sable can't do, and no visible path to reach a human agent from inside the chat."
Why this works
What a disclosure leaves out usually matters more than the words it does include.
Stage 6
Test it myself
Say it like this
"I'd pull five real users right after a live chat and ask them cold: who or what did you just talk to? That test found 34%, not 98%."
Why this works
AUDIT's strongest move: replicate the claim yourself instead of trusting the dashboard.
Stage 7
Close on the one line
Say it like this
"A number that says '98% acknowledged' isn't a fact about a conversation. It's a fact about a button someone tapped once, and those two things aren't the same claim."
Why this works
Restates the direct answer, ready for whatever gets pushed back on next.

Let's learn

Sable answers account questions inside Quillwave Bank's app, things like balance checks, transaction disputes, and card freezes.

At first app open, a one-time modal reads "This app uses AI features. Tap to continue." Ninety-one percent of users tap through in under a second. Quillwave's product dashboard reports that as "98% of users have acknowledged AI use."

Knowledge spark: what's a comprehension test? Asking someone, right after an experience, whether they understood a specific fact about it. It's different from asking whether they saw a screen; it asks whether the fact actually landed.
What Quillwave reported versus what a cold test found
100% 50% 0% 98% Onboarding modal, reported 34% Cold test, after a real chat
The 98% was never measured at the moment it claims to describe. Test it at the actual moment, and it drops by nearly two thirds.

The turn: the caption on screen isn't really the problem. It says the true word, "AI." The real problem is nobody had ever checked whether that word, seen once at onboarding, was still doing any work eighteen months later, at the one moment it actually mattered.

A tap on a modal, months before a chat even opens, was never a fact about that chat. It was a fact about a button, dressed up as a fact about a conversation.

At its worst: a customer messages Sable at 11pm trying to freeze a stolen card, assumes she's talking to a live agent, over-shares out of fear and stress, and gets a flatly worded response that reads as cold coming from a "person." She files a complaint about rude staff, and it lands on an actual employee's quarterly review.

The decision I would take back We set the AI badge's size and placement as a small caption when Sable first launched, back when every user opened the chat out of curiosity and read every new element on the screen closely. That default stayed exactly the same for eighteen months. It stopped making sense once opening a chat became routine, and nobody was reading a caption they'd already seen fifty times.

What I would leave alone: for a simple, low-stakes query, like checking today's balance, a small persistent badge is genuinely fine as-is. The emotional stakes of that interaction are low regardless of whether Sable or a person answers it.

The lesson: "acknowledged" is not a fact about understanding. It's a fact about an action someone took, at some point, in some context, and a real audit has to ask exactly which action, and exactly which context, before believing the number at all.

Now here is the same thing as a story

The short version above is what you'd say defending this audit to Quillwave's product leadership. Read this one for how the gap actually got found.

Marguerite Salou has run Quillwave's quarterly trust audits for three years. She reviews screen recordings the way some people watch weather patterns, looking for the thing that's changed too slowly for anyone else to notice.

For most of those three years, the 98% acknowledgment number sat quietly in every quarterly deck, unquestioned. It sounded solid. Nobody had a reason to look under it.

Hand sketched flow diagram titled Where the disclosure actually lives. Five boxes: app opens once, tap through modal, months pass, chat begins highlighted, assumes human.
Five steps, and the gap between the second and the fourth is where the whole disclosure quietly stopped doing its job.

The trigger was a new hire's plain question. A junior compliance analyst, sitting in on her first audit review, asked, "wait, has anyone actually tested whether people notice this while they're chatting, or just whether they tapped the modal once?"

Marguerite didn't have an answer. Nobody did.

Hand sketched decision tree titled Reading 98 percent acknowledged three ways. Root 98 percent acknowledged AI, branching to understood in the moment leads to real comprehension, tapped a modal once leads to stale consent, box needed ticking leads to compliance ritual, never opened the modal leads to silently excluded.
The same headline number supports at least three very different stories. Only a real test tells you which one is true.

She pulled twenty real chat sessions and cold-called each customer within an hour of their conversation ending, asking one plain question: who or what did you just talk to? Only seven said "an AI" or "a bot" without prompting. The other thirteen said "a person," "customer service," or simply weren't sure.

Comprehension test score, measured once a year since Sable launched
80% 40% 0% Year 1: 71% Year 2: 52% Year 3: 34%
Nothing about the caption changed across three years. What changed was how routine the chat had become, and the same words stopped landing.
The caption never lied. It just quietly stopped being read, and nobody had a number that would have told them that, until Marguerite went and asked.

Here's the decision I'd take back. We set the badge's size and placement once, when Sable was new and every user read the screen closely out of curiosity. Eighteen months later, the exact same caption sat in the exact same spot, on an app that had become routine.

Hand sketched labeled parts diagram titled What a real disclosure needs. Center document icon labeled Chat Disclosure, with four callouts: plain word AI, what it cannot do, reach a human, tested on real users.
Four parts, and Sable's live disclosure was missing three of them.

I'd put the word "AI" directly into Sable's first sentence, every single chat, not a badge nearby. I'd add one line about what Sable can't help with, and a visible tap to reach a human agent.

Hand sketched quadrant titled Prominent is not the same as timely. Axes prominence from easy to miss to hard to miss, and proximity to the moment from months earlier to right when it matters. Onboarding modal sits high prominence, low proximity. Small chat caption sits low prominence, high proximity. Verbal check-in line sits high on both. Terms of service sits high prominence, very low proximity.
The onboarding modal was loud and early. What Sable needed was quieter and exactly on time.

Replay the same 11pm stolen-card conversation under the new design: Sable's first line reads "I'm Sable, Quillwave's AI assistant. I can freeze your card right now; for anything else, tap here for a person." The customer knows exactly who she's talking to, gets her card frozen in under a minute, and never files the complaint that used to land on someone's review.

I built the caption once, when curiosity did the work for me. It took a new hire's honest question to see that curiosity doesn't last, and a disclosure that only works while a product is new isn't a disclosure. It's a grace period.

AUDIT, the actual readNot a gut check on the wording. AUDIT is what tells you whether a disclosure is evidence of understanding or a number wearing a checkmark.

A
Ask who made it.
The launch team wrote the caption eighteen months ago. Nobody outside that team has revisited it since.
The hardest step: most people never ask whether the original author ever checked their own work.
U
Uncover what it's tested against.
The 98% comes from a modal tap-through at first app open, never re-measured at the moment a chat actually starts.
A number means nothing until you know exactly what moment it was measured at.
D
Demand the version and the date.
The caption has been unchanged across three major app versions and three years, while the product itself became routine.
A disclosure with no revisit date is a decision nobody remembers making.
I
Isolate what's missing.
No plain "AI" in the first sentence, no stated limits, no visible path to a human.
What a disclosure leaves out usually matters more than the words it keeps.
T
Test it yourself.
A cold call to twenty real customers after a real chat found 34% comprehension, not 98%.
Replicate the claim before it reaches a leadership deck, not after.
Hand sketched timeline titled The caption that never changed. Four milestones: Sable launches caption written, app v3.0 caption unchanged, app v6.0 caption unchanged highlighted, today still unchanged.
Three major app versions, and the one thing that never got revisited was the exact line meant to keep users oriented.

The recap, one line per letter: ask is whether the launch team ever tested their own caption, uncover is that 98% measures a modal tap, not a chat moment, demand is three years with no version change, isolate is the missing "AI" word and missing limits, and test is the cold call that found 34%.

And if you want to be sure it really works, try it somewhere elseSame five letters, a public library's self-checkout kiosk instead of a bank chat. A different building, the same missing test.

A public library system added an AI-staffed help kiosk that answers questions about book holds and late fees, with a small printed sticker reading "Ask Libby, our virtual assistant" above the screen.

Mapped onto AUDIT: ask is whether the sticker was tested on actual patrons, especially older visitors less familiar with the phrase "virtual assistant," or whether it was approved by a marketing committee alone. Uncover is what "understood" was ever measured against, likely nothing, since no comprehension test exists at all. Demand is the sticker's print date, unchanged since installation two years ago, through one full kiosk software update. Isolate is the missing plain word "AI" (the sticker only says "virtual assistant," which many patrons read as a person on a screen elsewhere in the building) and the missing note that late-fee amounts can occasionally be wrong. Test is a simple hallway survey: stop ten patrons leaving the kiosk and ask who or what helped them, the same cold-test move Marguerite used at Quillwave.

Hand sketched icon list titled What Sable's live disclosure leaves out, reused here to show what the library kiosk sticker leaves out. Three items: no plain AI in the first sentence, no mention of what it cannot do, no easy path to a human.
Same three gaps, a very different building. The audit doesn't change; only the sticker does.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "test comprehension at the real moment, not at onboarding," and stop.
Cost: there's no budget to redesign the chat's first message this quarter. Start with a cheap fix: add one plain sentence, "You're chatting with Sable, our AI," to the very first message, no redesign required.
The model gets better, for real: if Sable's answers get more accurate, that's still not a reason to skip re-testing comprehension. A better model can still leave someone thinking they spoke to a person.

Where people run it wrong.
They treat a single onboarding tap-through as proof of ongoing understanding.
They measure whether a disclosure was shown, never whether it was understood.
They assume a badge that worked at launch still works once the product feels routine.

How to use it live. When someone hands you a disclosure design to judge, don't read the words first. Ask when it was last tested on a real person, in the real moment it's supposed to matter, and go from there.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "design the disclosure that tells a user they're talking to an AI"?
Tap to flip
ANSWER
AUDIT: ask who made it, uncover what it's tested against, demand the version, isolate what's missing, test it yourself. It fits because the real work is judging an existing disclosure, not inventing one from scratch.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Marguerite Salou, who has run Quillwave Bank's quarterly trust audits for three years.
3 · THE HEADLINE NUMBER
What did Quillwave report, and what did it actually measure?
Tap to flip
ANSWER
98% "acknowledged AI use," measured from a one-time onboarding modal tap, never re-measured at the moment a real chat begins.
4 · WHAT'S MISSING
Name two things Sable's live disclosure leaves out.
Tap to flip
ANSWER
The plain word "AI" in the chat's first sentence, and any mention of what Sable can't do or how to reach a human.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting the badge's size and placement once, when Sable was new and users read the screen closely, and never revisiting it once the app became routine.
6 · THE NUMBER
Fill in the blank: the cold comprehension test found only ___% could say they'd spoken with an AI.
Tap to flip
ANSWER
34%. Down from a claimed 98%, and down from 71% the very first year Sable launched.
7 · THE REPLAY
Same 11pm stolen-card chat, redesigned disclosure. What changes?
Tap to flip
ANSWER
Sable's first line plainly says "AI assistant" and offers a tap to reach a person; the customer knows who she's talking to and never files the complaint that used to land on a real employee's review.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the same missing piece?
Tap to flip
ANSWER
A public library's AI help kiosk. Same gap: a sticker that was never comprehension-tested on the actual visitors using it.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Quillwave's dashboard reported ___% acknowledgment, based on a one-time onboarding modal tap.
Show hint
Look at the grouped bar chart comparing reported versus tested numbers.
Show answer
98%. A cold test at the real moment found only 34% comprehension, less than half of what the reported number implied.
Multiple choice
2. Why does this answer say the 98% number is misleading, even though it's technically accurate?
  • A. Because 98% is mathematically impossible.
  • B. Because it measures a one-time modal tap, not whether users understood they were talking to an AI during a real chat.
  • C. Because users were forced to tap the modal against their will.
  • D. Because Sable's answers are frequently wrong.
Show hint
Look at the "uncover" step of AUDIT.
Show answer
B. A true fact measured at the wrong moment can still create a false impression about what's actually happening now.
True or false
3. True or false: this answer recommends removing the AI badge from the chat entirely, since it clearly isn't working.
  • True
  • False
Show hint
Look at the priority list and "what I would leave alone."
Show answer
False. It recommends moving the disclosure into the first sentence of the chat itself and testing it, not removing it.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Setting the badge's size and placement once at launch, when curiosity meant users read the screen closely. It stopped working once the app became routine.
Short answer, where it wouldn't matter
5. Name a kind of Sable interaction where the current small badge is genuinely fine as-is.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A simple, low-stakes query like checking today's balance, where it barely matters whether Sable or a person answered.
Short answer, apply it yourself
6. Pick an app you use that has some kind of AI disclosure. If someone asked you cold right now, could you say exactly where and how it told you?
Show hint
Think about a chat or feature labeled "AI" somewhere, and whether you actually remember reading that label recently.
Show answer
Model answer: Most people can name that a feature is "AI-powered" in general but struggle to say exactly where or when they were told, which is the same gap this answer found at Quillwave.
Before you close the answer
Why this works
Tests whether you can tell the difference between a disclosure that exists and a disclosure that actually works, and whether you'd trust a clean-looking metric without checking what moment it was actually measuring.
Follow-up traps
"Isn't a comprehension test just extra process for something that's obviously fine?" Response: it looked obviously fine for three years, right up until a cold test found 34%, which is exactly why the test has to happen before someone else finds the gap for you.

"Won't saying 'AI' in every first message get repetitive and annoying to regular users?" Response: fair, which is why the fix pairs it with something useful in the same sentence, like what Sable can help with right now, so the disclosure earns its place instead of just repeating itself.
If pressed
Quillwave's real fix doesn't run the disclosure line every single message; it shows the full line once per new chat session and a smaller persistent tag for the rest of that same conversation, so the plain word "AI" reappears exactly when a new conversation starts, not on every single reply.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more