Artifact critiqueIntermediateDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #19

Critique an interface that presents AI output with no indication of its origin.

AUDIT the artifact is QuickReply, Quenlow Telecom's support chat, where AI drafts and human agents both send under the same name

Marisol Roldan audits support quality at Quenlow Telecom. Every transcript she pulls up shows the same thing: a customer's question, and a reply signed "Agent Priya." Half the time, no human named Priya was anywhere near it.

The direct answer
Tag every message as human or AI, right in the transcript, permanently. Right now "Agent Priya" could be a person or QuickReply, and nobody, not the customer, not Marisol's own quality team, can look at a transcript later and tell which one answered, or what model version did it. That's not a missing disclosure. It's a missing audit trail, on the one document that will eventually get handed to someone asking hard questions.
Do this, in order
  1. Tag every message as human or AI, permanently, in the transcript itself.Why: this is the one change that makes every other audit question answerable.
  2. Record which model version answered, attached to the message, not buried in a separate backend log.Why: a version-pinned claim is checkable later. An unpinned one isn't.
  3. Mark any AI answer that was later corrected, right where the customer can still see it.Why: right now a wrong AI answer just sits there, uncorrected, forever, in the customer's own view.
  4. Stop reporting "Agent" resolution and satisfaction metrics as one blended number.Why: blending AI and human performance under one label hides exactly the gap someone will eventually go looking for.
  5. Check what "resolved" actually means before trusting the resolution rate.Why: if it just means the customer didn't reply again, silence is being counted as satisfaction.
  6. Leave the underlying QuickReply model alone; the interface is the actual problem.Why: this is a transparency failure, not an accuracy one, and fixing accuracy wouldn't touch it.

How to answer this, stage by stage

You're handed an interface, not a blank page. Spend your time finding what it hides, not redesigning it from scratch.

Stage 1
Scope the artifact
Say it like this
"I'll critique QuickReply, Quenlow Telecom's support chat, where an AI-drafted reply and a human agent's reply both get sent under the same 'Agent Priya' name."
Why this works
Names the exact artifact instead of critiquing "AI transparency" in the abstract.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Who benefits from this ambiguity, what's the resolution rate actually based on, is there a version pin, what's missing, and how would I test it myself."
Why this works
Shows a method for judging trust in an artifact, not just a list of complaints.
Stage 3
Reframe the question
Say it like this
"This isn't a missing disclaimer. It's that the one document everyone actually reads, the transcript, can't answer 'who or what said this' if it's ever handed to a regulator, a lawyer, or even Marisol's own team."
Why this works
Moves the critique from a UX nitpick to a real, specific failure.
Stage 4
Give the one fix
Say it like this
"Tag every message as human or AI, permanently, right in the transcript. Never share one name across both."
Why this works
Matches the direct answer exactly, which is what makes it hold up under follow-up.
Stage 5
Prove it with a failure
Say it like this
"Marisol pulled ten 'Agent Priya' transcripts flagged as wrong. She had to go into a separate backend log for each one just to find out four of the ten were QuickReply, not a person, and two of those four were never corrected in what the customer can still see."
Why this works
Shows the audit gap is real and costly, not a hypothetical worry.
Stage 6
Close on one line
Say it like this
"A transcript with no origin tag isn't neutral. It's a document that already failed the one question it will eventually be asked."
Why this works
Restates the direct answer in a form short enough to remember under pressure.

Let's learn

Quenlow Telecom runs QuickReply, an AI tool that drafts and sends replies to simple support questions, billing dates, outage status, plan details, under a shared agent name and avatar that a human agent also uses.

Before anyone looked closely, the interface seemed fine: one name, one avatar, one clean transcript per customer. Resolution rate for "Agent Priya" ran at 94 percent, a number the support team was proud of.

Knowledge spark: what does an audit trail actually mean here? A record detailed enough that someone who wasn't in the room can reconstruct exactly what happened later: who or what acted, when, and on what version. A transcript without an origin tag isn't a trail. It's a single frame with everything before and after it missing.

The turn. The shared name was never really about a clean-looking chat window. It meant nobody, not the customer, not Marisol's own quality team, could look at a transcript later and answer the one question an auditor always asks first: who, or what, actually said this.

Hand sketched metaphor scene titled What the transcript shows versus hides. Left panel labeled shown, a question mark icon, caption Agent Priya, one name. Right panel labeled hidden, a document icon, caption which model, what version.
Everything on the right exists somewhere in Quenlow's systems. None of it survives into the transcript anyone actually reads.

At its worst: a customer got a wrong answer about a fee waiver from QuickReply, acted on it, and the transcript they later screenshotted for a dispute still read "Agent Priya says you're covered," with no way for anyone, including Quenlow's own dispute team, to confirm whether a person or a model had said it.

The decision I would take back We gave QuickReply the same name and avatar as a human agent, since it made the chat window feel more consistent and less jarring for customers switching between AI and human replies. That made sense when QuickReply only handled a handful of simple, low-stakes questions. It stopped making sense once it started handling billing and eligibility questions that customers act on.

What I would leave alone: QuickReply's underlying answer quality for routine questions like outage status, which is genuinely accurate and doesn't need a redesign. The interface's silence about its own origin is the actual problem, not the model behind it.

Now here is the same thing as a story

The short version above is what you'd say to Quenlow's product leadership. Read this one for how the audit actually happened.

Marisol Roldan has audited Quenlow's support quality for five years, and can usually tell a scripted answer from a genuinely thoughtful one within the first two lines.

Her monthly audits used to be simple: pull twenty random "Agent Priya" transcripts, score them for accuracy and tone, done by lunch.

The five AUDIT questions started as a personal checklist Marisol used on vendor claims, not on her own company's product, until a routine pull turned up a transcript where "Agent Priya" told a customer their late fee was waived. It wasn't.

Hand sketched labeled parts diagram titled The five AUDIT checks. Center question mark icon labeled trust this transcript, with five callouts: who benefits, what's the claim based on, version pin, what's missing, test it yourself.
Marisol had used all five questions on other companies' claims for years. This was the first time she turned them on her own transcript.

She asked who benefited from the ambiguity first. The answer wasn't a person acting in bad faith, it was a metrics dashboard: blending AI and human replies under one "Agent" resolution rate made the team's overall numbers look stronger than either channel actually was on its own.

Hand sketched flow diagram titled Where the origin tag should sit, and doesn't. Four boxes: QuickReply drafts, sent as one name, no origin tag highlighted, customer reads it.
The gap sits in the third box. Everything before it and after it works exactly as designed.
A shared name didn't just blur who answered. It erased the one fact anyone would need later to figure out what actually went wrong.

She went looking for the version pin next, and found none in the transcript itself, only in a backend log disconnected from what the customer or a dispute reviewer would ever see. Testing it herself meant pulling ten flagged-wrong transcripts and manually cross-referencing backend logs for each one.

Support tickets citing confusion over who answered
350 175 0 340 Before the tag 60 After the tag
Tickets about "was I talking to a bot" fell 82 percent once the transcript itself answered the question, instead of leaving customers to guess.

The team considered simply retraining QuickReply to sound more clearly automated in tone, and rejected it. Tone is easy to fake convincingly either direction, and it still leaves nothing in the transcript itself for an auditor to check later.

Hand sketched icon list titled What a trustworthy version would show. Four items: a document icon labeled human or AI tagged on every message, a gauge icon labeled which model version answered, a question mark icon labeled a mark when AI was later corrected, a person icon labeled an easy way to ask for a human.
None of these four require touching QuickReply's actual answers. All four just require the transcript to stop hiding what already happened.
Customer trust score, before and during the rollout
5.0 2.5 0 wk2: 2.7 tag rolls out wide wk10: 3.8
Trust dipped slightly in week two, when customers first noticed the tag and had questions about it, then climbed steadily as they learned what it meant.

I let QuickReply share a name with human agents because it made the interface feel cleaner and more consistent at launch. It took Marisol tracing a wrong fee-waiver answer through a disconnected backend log, by hand, to see that "clean" had quietly become "unaccountable."

The five checks, run on our own productNot a story about a vendor's claim. AUDIT works just as well turned on your own interface.

A
Ask who benefits.
A blended "Agent" resolution rate makes the whole support team's numbers look better than either channel alone.
Finds the incentive behind the ambiguity, not just the ambiguity itself.
U
Uncover the basis for the claim.
"94% resolved" partly means the customer simply didn't reply again within ten minutes, not that the answer was confirmed correct.
A resolution number means nothing until you know what counts as resolved.
D
Demand the version pin. The hardest step.
No transcript names which model or version answered. Only a disconnected backend log does, if anyone thinks to check it.
Without this, nobody can reproduce or specifically fix what a customer actually saw.
I
Isolate what's missing.
No origin tag, no confidence signal, and no mark when a wrong AI answer was later corrected.
What the transcript leaves out is more informative than what it shows.
T
Test it yourself.
Marisol manually cross-referenced ten flagged-wrong transcripts against backend logs. Four of ten were QuickReply.
Verifying by hand is what actually confirmed the gap, not a claim on a dashboard.

The recap, one line per letter: ask is the blended metric that benefits from the ambiguity, uncover is what "resolved" quietly includes, demand is the missing version pin, isolate is the missing origin tag and correction mark, and test is Marisol's own manual cross-check that confirmed all of it.

And if you want to be sure it really works, try it somewhere elseSame five checks, a foodbank's eligibility chatbot instead of a telecom support chat. Higher stakes, same missing tag.

Grayhaven Foodbank runs AskHarbor, a chatbot that answers questions about pantry eligibility and pickup schedules, with no indication anywhere in the chat that answers come from a model rather than a staff member. Ollie Fenwick coordinates Grayhaven's programs and hears from clients directly when something goes wrong.

Mapped onto AUDIT: ask who benefits finds that an unlabeled chatbot lets the foodbank present "24/7 staff support" without disclosing that nights and weekends are AI-only. Uncover the basis finds that AskHarbor's own "helpful" rating comes from a thumbs-up button most clients never touch either way. Demand the version pin finds nothing recorded per answer at all, not even in a backend log, since the foodbank never built one. Isolate what's missing finds no origin tag and no easy path to a human when the answer affects whether someone eats that week. Test it yourself means Ollie sitting down and running real eligibility questions through AskHarbor himself, and finding it gave an outdated income threshold for two straight weeks.

Hand sketched quadrant titled Other AI touchpoints, sorted. Axes: how high the stakes, how clearly disclosed. Billing chat replies and contract cancellation sit high stakes and hidden. Outage status sits low stakes and hidden. Plan upsell reply sits moderate stakes and moderately disclosed.
The same audit applies to any AI touchpoint. The ones worth fixing first sit in the bottom right: high stakes, poorly disclosed.
Hand sketched timeline titled One client's session with AskHarbor. Four milestones: client asks about eligibility at 0:00, AskHarbor answers with no tag highlighted at 0:04, client acts on the wrong answer two days later, client turned away at pickup.
Four minutes to answer. Two days before the wrong answer actually cost someone a meal.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "tag every message as human or AI, permanently, in the transcript itself," and stop.
Cost: there's no budget for a full redesign this quarter. Say so honestly, and ship the tag alone first; it's the cheapest single fix and it closes most of the audit gap on its own.
The model gets better, for real: if QuickReply's accuracy improves next year, the missing origin tag doesn't improve with it. A more accurate model that still can't say who answered is still an unauditable transcript.

Where people run it wrong.
They treat a shared name as a small cosmetic choice instead of an audit trail decision.
They accept a high resolution or satisfaction number without checking what counts as a resolution.
They assume fixing the model's accuracy also fixes the transparency gap, when the two are unrelated.

How to use it live. When handed an interface to critique, don't start with what it does well. Ask first: if this exact transcript got pulled six months from now by someone hostile, could it answer who or what said this? If not, that's your critique, stated in one sentence.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "critique an interface that presents AI output with no indication of its origin"?
Tap to flip
ANSWER
AUDIT: ask who benefits, uncover the claim's basis, demand the version pin, isolate what's missing, test it yourself. Demand the version pin is the hardest step here.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Marisol Roldan, who has audited Quenlow Telecom's support quality for five years.
3 · WHO BENEFITS
Who benefits from QuickReply and human agents sharing one name?
Tap to flip
ANSWER
The support team's own metrics, since blending AI and human replies under one "Agent" resolution rate makes the overall number look stronger than either channel is alone.
4 · WHAT'S MISSING
Name the single biggest gap in the QuickReply transcript.
Tap to flip
ANSWER
No tag anywhere showing whether a human or QuickReply answered, and no version pin recorded alongside the message itself.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Giving QuickReply the same name and avatar as a human agent, since it made the chat window feel cleaner and more consistent when QuickReply only handled a few low-stakes questions.
6 · THE NUMBER
Fill in the blank: tickets about "was I talking to a bot" fell from 340 to ___ per month after the tag shipped.
Tap to flip
ANSWER
60 per month, an 82 percent drop, once the transcript itself answered the question instead of leaving customers to guess.
7 · THE REPLAY
Same flagged-wrong transcript, redesigned interface. What changes?
Tap to flip
ANSWER
Marisol can see instantly, from the transcript alone, that QuickReply answered, which version, and whether it was later corrected, with no backend log needed.
8 · CROSS PRODUCT TRANSFER
Section 4 runs the same five checks on a different product. Which one, and what raises the stakes?
Tap to flip
ANSWER
Grayhaven Foodbank's AskHarbor. There, a wrong unlabeled answer doesn't just confuse a customer, it can leave someone turned away from a meal.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision does this critique take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Giving QuickReply the same name as a human agent. It made sense while QuickReply only handled a few simple, low-stakes questions and consistency felt more important than disclosure.
Multiple choice
2. Why is the missing version pin specifically called out as the hardest AUDIT check here?
  • A. Because QuickReply has no version history at all.
  • B. Because without it, nobody can reproduce or specifically fix what a customer actually saw, even inside Quenlow's own team.
  • C. Because version pins are only required for regulated industries.
  • D. Because customers always ask for the version number directly.
Show hint
Look at the "demand the version pin" step in the recap.
Show answer
B. A version pin is what makes a claim about what happened checkable later. Without it, even the company's own audit team is stuck.
True or false
3. True or false: this critique recommends retraining QuickReply to sound more obviously robotic in tone.
  • True
  • False
Show hint
Look at the rejected alternative in the story.
Show answer
False. That option was considered and rejected. Tone is too easy to fake either direction and still leaves nothing checkable in the transcript itself.
Fill in the blank
4. Fill in the blank: of the ten flagged-wrong transcripts Marisol checked by hand, ___ turned out to be QuickReply, not a human agent.
Show hint
Look at "prove it with a failure" in the walkthrough.
Show answer
Four. And two of those four were never corrected in what the customer could still see.
Short answer, where it wouldn't matter
5. Name a part of QuickReply that this critique says doesn't need to change.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: QuickReply's underlying answer quality for routine, low-stakes questions like outage status, which is genuinely accurate. The transcript's silence about its own origin is the actual problem.
Short answer, apply it yourself
6. Think of a chat, email, or review you've read recently. Could you tell for certain whether it was AI-written? What would have told you?
Show hint
Think about customer service chats, product reviews, or auto-generated emails.
Show answer
Model answer: Most people can only guess from tone, which is unreliable. A small permanent tag on the message itself is the only thing that actually settles it.
Before you close the answer
Why this works
Tests whether you can look past a clean-looking interface and find the specific document, the transcript, that will eventually be handed to someone asking hard questions, and check whether it can actually answer them.
Follow-up traps
"Wouldn't tagging every message just make customers trust the chat less?" Response: the opposite happened here; trust dipped briefly in week two and then climbed past its starting point once customers understood what the tag meant.

"Isn't this just a labeling issue, not a real audit problem?" Response: no, because the missing tag is exactly what stopped Quenlow's own quality team from confirming what happened in a disputed transcript, which is a real operational cost, not a cosmetic one.
If pressed
Quenlow's actual fix stores the model version as a hidden field on the message record itself, not just the visible tag, so a dispute reviewer can pull the exact version even for messages sent before the visible tag existed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more