Artifact critiqueAdvancedDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #6

Critique a design that surfaces a chain of thought to end users.

AUDIT the product is Cordage Capital, a robo-advisor, and its See how I reasoned through this panel

Cordage Capital is a robo-advisor. When it recommends a portfolio rebalance, a panel labeled "See how I reasoned through this" shows the model's raw reasoning text. Halyna Marchetti was brought in as a trust researcher to review the panel before its rollout widened to every client.

The direct answer
Do not trust a raw reasoning panel just because it looks technical. Check whether the displayed text is the model's actual reasoning path or a separate summary written after the fact, whether it was ever tested against a plain-language version for real decision quality, and whether a mid-reasoning reversal, the model changing its mind partway through, gets flagged as a normal, healthy revision or just dumped on screen as unexplained noise. Raw and honest are two different axes, and this panel was built as if they were the same one.
Do this, in order
  1. Confirm the displayed reasoning is the real computation, not a bolted-on summary written afterward.Why: a reasoning panel that isn't faithful to the actual answer is decoration wearing a lab coat.
  2. Flag a mid-reasoning reversal as a labeled revision, not raw unexplained text.Why: a model changing its mind is normal. Shown without a label, it reads as instability instead.
  3. Version-pin every reasoning trace to the exact model that produced it.Why: without a pin, nobody can check an old, disputed trace against anything.
  4. Test the panel against real decisions, not against how technical it looks.Why: looking rigorous and actually improving a client's choice are two different outcomes.
  5. Keep the fully raw, unflagged trace available to internal advisors only.Why: trained advisors read a self-correction as diligence. Most clients, cold, read it as doubt.

How to answer this, stage by stage

Nobody's grading whether you know chain-of-thought panels exist. They're grading whether you can spot the exact assumption baked into this one that nobody actually checked.

Stage 1
Scope it to one real panel
Say it like this
"I'll answer this for Cordage Capital's 'See how I reasoned through this' panel, shown right under every portfolio rebalance recommendation."
Why this works
Anchors the critique in one real artifact instead of chain-of-thought UIs in general.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who built it, uncover what it's tested against, demand the version pin, isolate what's missing, then test it myself."
Why this works
Signals a method for checking a claim, not a list of opinions about the copy.
Stage 3
Ask who built it, and why
Say it like this
"An engineer shipped this fast after a new 'reasoning mode' model release, to differentiate from a competitor. No user researcher signed off on whether the panel was honest, just whether it was impressive."
Why this works
AUDIT's opening move, and the one most candidates skip when the artifact looks technically sophisticated.
Stage 4
Uncover what "reasoning" actually means here
Say it like this
"Nobody had confirmed whether the displayed text is the literal path the model used, or a separate, cleaned-up narration generated afterward. Those are not the same guarantee."
Why this works
A reasoning panel means nothing until you know whether it's faithful to the real computation.
Stage 5
Demand the version pin
Say it like this
"The panel doesn't record which exact model version produced a given trace. A client disputing a six-month-old recommendation can't be shown what the model actually saw."
Why this works
A reasoning trace nobody can trace back to a specific model isn't evidence, it's a snapshot with no timestamp.
Stage 6
Isolate what's missing, and test it
Say it like this
"There's no flag for a mid-reasoning reversal, so a normal self-correction reads as instability. I'd run half of clients against the raw panel and half against a plain summary, and measure which group makes better rebalancing decisions, not which group feels more impressed."
Why this works
AUDIT's strongest move: test the real effect yourself instead of trusting how technical the feature looks.
Stage 7
Close on the one line
Say it like this
"A raw reasoning panel isn't automatically more honest. It's just more raw. Whether it's honest depends on whether it's faithful, versioned, and flagged, and this one is none of the three."
Why this works
Restates the direct answer, ready for a follow-up push.

Let's learn

Cordage Capital's robo-advisor recommends portfolio rebalances. When it added a "See how I reasoned through this" panel, showing the model's raw reasoning text, it followed a new model release built to write out its own deliberation step by step.

Before the panel, clients saw one plain-language reason line under a recommendation. About 6% of them still called a human advisor asking "why," each call costing roughly twelve minutes of advisor time. After the panel shipped, those calls dropped to 2%.

Knowledge spark: what's faithfulness, in reasoning traces? Whether the reasoning a model displays is actually the path it used to reach its answer, rather than a separate, cleaned-up story generated afterward. A trace can read as detailed and still not be faithful.
Support contacts, before and after the reasoning panel shipped
8% 4% 0% 6% 2% "Why" calls, before / after 0.1% 3.4% "Contradicts itself," before / after
One complaint type fell. A different one, that barely existed before, rose thirty-four times over.

The turn: a model reversing itself mid-reasoning isn't actually the problem. Models revise their thinking constantly on the way to a good answer, and that revision is often exactly what makes the final recommendation solid. The real problem is showing that revision raw, with no label telling the reader they're looking at a normal correction rather than a system that can't make up its mind.

A raw reasoning trace isn't more honest just because it's less edited. It's just less edited. Whether that helps or scares someone depends entirely on whether they know what they're looking at.

At its worst: a client reads a trace where the model leans toward selling a stock, then writes "actually, let me reconsider given the client's stated risk tolerance" and reverses to holding it. She screenshots the contradiction and posts it publicly with the caption "this AI can't make up its mind about my retirement money." It spreads widely, and trust drops broadly even though the final recommendation was the better one.

The decision I would take back We shipped the raw reasoning trace with no persisted, versioned record of which model produced it, since the feature was rushed out fast to match a competitor's new "reasoning mode" release. That made sense for a quick, narrow beta. It stopped making sense the moment a client disputed a months-old recommendation and nobody could confirm what the model had actually seen that day.

What I would leave alone: for internal advisor use, the fully raw, unflagged trace is exactly right as-is. Advisors are trained to read a model second-guessing itself as diligence, not instability, so they need nothing softened.

The lesson: "we're showing our reasoning" is not one claim. It's at least three, faithful, versioned, and flagged, and a panel that satisfies none of them is still technically a reasoning panel. It's just not a trustworthy one.

Now here is the same thing as a story

The short version above is what you'd say defending this critique to Cordage's product leadership. Read this one for how the gap actually got found.

Halyna Marchetti reviews AI features before they roll out widely, brought in specifically because she's not on the team that built them. She reads an interface the way a proofreader reads a contract, looking for the sentence someone assumed nobody would test.

For the first two months after the reasoning panel launched to a small pilot group, it looked like a clean win. "Why" calls dropped. The product team called it a success in a leadership update, pointing at the falling call volume as proof.

Hand sketched flow diagram titled Where the reasoning trace comes from. Four boxes: client asks why, model writes reasoning highlighted, panel shows it raw, client reads it.
Four steps, and nobody had checked what actually happens between the second box and the third.

The trigger wasn't anything that happened at Cordage. A rival robo-advisor's chain-of-thought panel went viral for a different reason: a screenshot showing its model reasoning "the client is 61, so I should recommend more conservative bonds," then flatly contradicting that logic two lines later. Financial commentators mocked it for days.

Halyna watched the story spread and thought, quietly, "we have the exact same panel."

Hand sketched comparison diagram titled Faithful trace or theater. Left panel, a gauge icon labeled Real trace, caption the actual tokens used. Right panel, a question mark box icon labeled After the fact story, caption a separate summary call.
Cordage had never actually confirmed which of these two panels it was showing.

She pulled Cordage's own reasoning logs and found the same pattern quietly present, just not yet screenshotted. Traces containing a mid-reasoning reversal, "actually, let me reconsider," had been climbing for weeks, unnoticed because the complaint category "AI contradicts itself" hadn't existed as a support tag until someone finally added it.

"AI contradicts itself" complaints, by week since the panel launched
4% 2% 0% Week 1: 0.1% Week 5: 1.2% Week 9: 3.4%
The pattern was already climbing at Cordage before the rival's incident. It just hadn't been noticed yet.
Cordage's own version of the same failure was already climbing quietly. It took someone else's public embarrassment to make anyone go check.

Here's the decision I'd take back. We shipped the raw trace with no version pin and no revision flag, because the whole feature was rushed out to answer a competitor's release. That made sense for a fast pilot. It stopped making sense the moment real client trust was riding on it.

Hand sketched metaphor scene titled A diary versus a stage script. Left, a document icon labeled Diary, caption private, unfiltered. Right, a box icon labeled Script, caption written for an audience.
The panel was built like a diary and shown like a script, without deciding which one it actually was.

I'd add one plain label, "reconsidered," wherever the model's own trace shows a reversal, plus a version pin on every trace so a disputed recommendation from months ago can be checked against the model that actually made it.

Hand sketched labeled parts diagram titled What a trustworthy reasoning panel needs. Center document icon labeled Reasoning Panel, with four callouts: faithfulness check, version pin, revision flag, plain summary.
Four parts, and Cordage's live panel had none of them when it shipped.

Replay the same stock-hold recommendation under the new design: the client sees "reconsidered: kept this stock after weighing your stated risk tolerance more heavily," instead of a raw, unlabeled reversal. Nothing about the actual advice changes. What changes is whether it reads as a system thinking carefully or a system that can't decide.

We built the panel fast to answer a competitor's move, and treated "shows more of the model" as automatically the safer, more honest choice. It took watching a rival get mocked publicly, then finding the exact same fault line quietly growing in our own logs, to see that raw and honest were never the same promise.

AUDIT, reading a reasoning panel like evidenceNot a style critique. AUDIT is what separates a reasoning panel that's actually checkable from one that only looks like it is.

A
Ask who built it.
An engineer shipped it fast after a competitor's model release. No user researcher signed off before the pilot went live.
The hardest step: most people never ask who decided a technical-looking feature was ready.
U
Uncover what it's tested against.
Nobody confirmed whether the displayed text is the model's real reasoning path or a separate narration written afterward.
A reasoning panel means nothing until its faithfulness is established.
D
Demand the version pin.
No trace is tied to a specific model version, so a disputed recommendation from months ago can't be checked against anything.
A trace with no version pin can't be reproduced or defended later.
I
Isolate what's missing.
No flag for a mid-reasoning reversal, so a normal self-correction reads as instability to a cold reader.
What a reasoning panel doesn't label is usually the part that costs it trust.
T
Test it yourself.
Split clients between the raw panel and a plain summary, and measure decision quality, not how impressed they feel.
Replicate the real effect before someone else's viral screenshot does it for you.
Hand sketched quadrant titled Faithful is not the same as reassuring. Axes faithful to computation from low to high, and reassuring to read from unsettling to reassuring. Raw CoT unfiltered sits high faithful, low reassuring. Cleaned narration sits low faithful, high reassuring. Plain reason line sits low faithful, moderately high reassuring. Flagged marked CoT sits high on both.
The goal corner isn't raw, and isn't cleaned. It's faithful and flagged, at the same time.

The recap, one line per letter: ask is an engineer shipping fast with no research sign-off, uncover is the unconfirmed gap between real reasoning and after-the-fact narration, demand is the missing version pin, isolate is the missing revision flag, and test is splitting real clients between raw and plain versions to measure actual decisions.

And if you want to be sure it really works, try it somewhere elseSame five letters, a university admissions office instead of a robo-advisor. A different building, and the reversal that scares people is about a person, not a portfolio.

A university's admissions office uses an assistant that drafts a reasoning summary for why an applicant was waitlisted, shown to admissions staff before a final decision. Mapped onto AUDIT: ask is whether an admissions officer or only an engineering intern reviewed the panel before staff started relying on it daily. Uncover is whether the summary reflects the model's actual weighting of test scores, essays, and extracurriculars, or a separate, cleaned explanation generated after the ranking was already set. Demand is the version pin, since admissions cycles run for months and a challenged decision needs to be traceable to the exact model that produced it. Isolate is the missing flag for cases where the model's stated reasoning shifted between an early read of the file and its final rank. Test is pulling ten waitlisted files and checking, by hand, whether the summary text actually matches the features that moved the ranking, not just whether it sounds plausible.

Hand sketched icon list titled What Cordage's panel never checked, reused here for the admissions office's reasoning summary. Three items: is this the real reasoning path, which model version made it, does a revision get flagged.
A robo-advisor and an admissions office share almost nothing else. They share exactly these three gaps.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "raw isn't honest by default, check faithfulness, versioning, and revision flags," and stop.
Cost: there's no engineering time to rebuild the panel this quarter. Start with the cheapest fix: label mid-reasoning reversals with one word, "reconsidered," without touching anything else.
The model gets better, for real: if the new model version reasons more clearly overall, that's still not a reason to skip a faithfulness check. A clearer-sounding trace can still be an after-the-fact story instead of the real path.

Where people run it wrong.
They treat a technical-looking panel as automatically more trustworthy than a plain-language one.
They celebrate a drop in "why" calls without checking what new complaint category might be quietly rising instead.
They ship a reasoning feature to match a competitor's announcement before anyone tests what it actually does to real decisions.

How to use it live. When someone hands you a chain-of-thought panel to critique, don't start with the wording. Ask whether it's faithful to the real computation, whether it's pinned to a version, and whether a reversal gets labeled. If any of the three is missing, say so before anything else.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "critique a design that surfaces a chain of thought to end users"?
Tap to flip
ANSWER
AUDIT: ask who built it, uncover what it's tested against, demand the version pin, isolate what's missing, test it yourself. It fits because the task is judging an existing artifact, not designing a new one.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Halyna Marchetti, a trust researcher brought in to review Cordage Capital's reasoning panel before its full rollout.
3 · THE HEADLINE NUMBER
What number did the product team celebrate, and what did it hide?
Tap to flip
ANSWER
"Why" calls falling from 6% to 2%. It hid a new complaint type, "AI contradicts itself," climbing from 0.1% to 3.4% over the same weeks.
4 · WHAT'S MISSING
Name two things Cordage's live reasoning panel never checked.
Tap to flip
ANSWER
Whether the displayed reasoning is faithful to the real computation, and whether it's pinned to the exact model version that produced it.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping the raw trace with no version pin and no revision flag, since the feature was rushed to answer a competitor's model release.
6 · THE NUMBER
Fill in the blank: "contradicts itself" complaints reached ___% by week 9, up from 0.1% in week 1.
Tap to flip
ANSWER
3.4%. The pattern was already climbing at Cordage before a rival's public incident drew attention to it.
7 · THE REPLAY
Same stock-hold reversal, redesigned panel. What changes?
Tap to flip
ANSWER
The client sees a labeled "reconsidered" note instead of a raw, unexplained reversal. The advice is identical; only whether it reads as careful or unstable changes.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the same missing check?
Tap to flip
ANSWER
A university admissions office's waitlist reasoning summaries. Same gap: nobody confirmed the displayed reasoning is faithful to the real ranking, or flagged when it shifted mid-review.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer argues the raw reasoning panel should be removed entirely, since it caused new complaints.
  • True
  • False
Show hint
Look at the priority list and "what I would leave alone."
Show answer
False. It stays, with a label added for reversals and a version pin added; the raw form stays exactly as-is for internal advisor use.
Multiple choice
2. Why does this answer say a mid-reasoning reversal shouldn't automatically read as a problem?
  • A. Because clients never actually read the reasoning panel.
  • B. Because models revising their thinking is often exactly what produces a better final answer.
  • C. Because Cordage's model is never wrong.
  • D. Because reversals only happen once a year.
Show hint
Look at "the turn" paragraph in Let's learn.
Show answer
B. The problem isn't the revision itself, it's showing it with no label telling the reader it's a normal, healthy correction.
Fill in the blank
3. Fill in the blank: "why" calls to human advisors dropped from 6% to ___% after the reasoning panel launched.
Show hint
Look at the first grouped bar chart.
Show answer
2%. That drop is exactly what made the feature look like an uncomplicated success at first.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Shipping the raw trace with no version pin or revision flag. It made sense for a fast pilot answering a competitor's release, not for a feature real client trust would depend on.
Short answer, where it wouldn't matter
5. Name a group of users for whom the fully raw, unflagged reasoning trace is genuinely fine as-is.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Internal advisors, who are trained to read a self-correction as diligence rather than instability.
Short answer, apply it yourself
6. Pick an AI tool that shows its reasoning. Would you actually know if that reasoning were faithful to its real computation, or just a plausible-sounding story written afterward?
Show hint
Think about whether you've ever seen a tool explain a decision that didn't quite match the actual result you got.
Show answer
Model answer: Most people admit they have no way to check this at all, which is exactly the gap this answer's "uncover" step is built to catch.
Before you close the answer
Why this works
Tests whether you can tell the difference between a feature that looks rigorous and one that's actually checkable, and whether you'd accept a technical-looking panel on faith the way the product team initially did.
Follow-up traps
"Isn't labeling every reversal just adding more clutter to an already dense panel?" Response: one word, "reconsidered," is not clutter; it's the single cheapest fix that turns unexplained noise into a legible signal.

"Couldn't a faithfulness check itself be gamed, by training the model to look more faithful?" Response: possible, which is exactly why testing it yourself, by perturbing an input and confirming the displayed reasoning actually changes to match, matters more than trusting a vendor's claim about it.
If pressed
Cordage's real faithfulness test perturbs one input at a time, like a client's stated risk tolerance, and checks whether the displayed reasoning trace changes in a way that actually matches the resulting recommendation change, not just whether the trace reads differently.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more