CaseAdvancedDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #3

How do citations change user behaviour, and what happens when they are wrong?

FLIPS the product is Ferrytale Autos, a used-car marketplace, and its listing assistant that cites the vehicle history report

Ferrytale Autos is a used-car marketplace. Its listing assistant reads each car's official vehicle history report and writes a plain-language citation, like "No accidents reported," straight onto the listing. Innes Calloway has reviewed listing quality at Ferrytale for four years, working the back office mid-morning shift.

The direct answer
A citation doesn't make people check less carefully. It makes them stop checking at all, once it's been right enough times in a row. The fix isn't asking reviewers to try harder. It's flagging the specific citations the model itself is least sure about, so a human glance lands exactly where the risk actually is, instead of nowhere.
Do this, in order
  1. Route only the model's low-confidence citations to a human, not a random sample.Why: a flat spot-check rate misses exactly the cases most likely to be wrong.
  2. Never let a high accuracy rate justify removing the review step entirely.Why: 99.8% accurate still means someone gets the other 0.2%, with nobody watching for them.
  3. Track the spot-check rate itself, not just the error rate.Why: the spot-check rate can fall to zero for months before an error rate ever shows it.
  4. Leave simple, unambiguous fields, like an exact mileage number, uncontested.Why: not every citation carries the same risk, and treating them all the same wastes review time.
  5. Give the reviewer the exact source page next to the citation, not just the claim.Why: a citation you can't trace back to its source isn't checkable, it's just a second claim.

How to answer this, stage by stage

Nobody's grading whether you know citations can be wrong. Everyone knows that. They're grading whether you can name what actually happens to a person's behavior once citations are usually right.

Stage 1
Scope it to one real product
Say it like this
"I'll answer this for Ferrytale Autos, where a listing assistant cites the vehicle history report directly on every used-car listing."
Why this works
Turns "how do citations change behavior" from an abstract idea into one real screen.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a method for finding the behavior change, not just a list of pros and cons of citations.
Stage 3
Name the person and the habit
Say it like this
"Innes Calloway used to open the source report and check every AI-written citation against it. Over about four months, as the citations kept being right, she checked fewer and fewer, then stopped."
Why this works
Shows the habit forming gradually and reasonably, not as a lapse in judgment.
Stage 4
Identify the flip
Say it like this
"This is an over-trust flip. She didn't go from checking a lot to checking a little. She went from checking some to checking none. There's no middle setting once the habit's gone."
Why this works
Names the exact mechanism, and calls out that improvement itself can cause a flip.
Stage 5
Name the old decision
Say it like this
"We merged the human glance-check into the auto-publish pipeline once accuracy crossed 99%, because by then the pause looked like pure overhead on an already-accurate system."
Why this works
Names a real, reversible product decision, not a vague call to "add more review."
Stage 6
Show the replay
Say it like this
"With low-confidence citations flagged instead, Innes checks about fifteen a day out of five hundred. She catches the wrong 'no accidents' line before it publishes, and the buyer never drives three hours to a car with a hidden rebuilt title."
Why this works
Ends on a countable, specific outcome, not a vague "much better."
Stage 7
Say what you'd measure, and close
Say it like this
"I'd watch the spot-check rate itself, not just the error rate, since the spot-check rate can hit zero months before an error rate shows anything. A citation makes people trust more, not check less carefully, until one day they check not at all."
Why this works
Closes on the direct answer, with a leading indicator ready for a follow-up question.

Let's learn

Ferrytale Autos' listing assistant reads a used car's official vehicle history report and writes a plain-language citation onto the listing, like "No accidents reported."

Before the assistant, a copy editor read the full multi-page report by hand and wrote that line herself, about six minutes per listing. With the assistant, the same line appears instantly, and it's held at 99.8% accurate across its first six months, about one wrong citation in every 500.

Knowledge spark: what's a vehicle history report? A record pulled from title, registration, and repair databases showing a used car's past, accidents, rebuilt titles, prior owners. Buyers rely on it because they can't see that history just by looking at the car.
Citations sent for human spot-check, before and after the pipeline merge
100% 50% 0% 100% Before the merge 0% After the merge
The error rate barely moved. The number of citations anyone actually looked at went straight to zero.

The turn: the rare wrong citation isn't really the problem. Ferrytale's assistant is right on the overwhelming majority of listings. The real problem is that nobody was left checking specifically for the wrong ones, because spot-checks had stopped entirely once accuracy proved itself out.

Spot-check rate versus citation error rate, over 6 months
100% 50% 0% Month 1 Month 4 Spot-check rate Error rate: ~0.2%
One of these lines fell to zero in four months. The other one barely moved the whole time, which is exactly why nobody noticed the first one falling.
A 99.8% accuracy rate is not a plan for the 0.2%. It's just a very convincing reason nobody built one.

At its worst: one wrong "no accidents" citation on a car with a rebuilt salvage title sends a buyer driving three hours to see it, only to find the truth at a mechanic's lift. He posts a furious video that thousands of people see, and several unrelated sellers pull their own perfectly accurate listings out of sheer caution. A tool that was right 499 times out of 500 still manages to do real, visible harm.

The decision I would take back We merged the human spot-check into the auto-publish pipeline once citation accuracy crossed 99%, since the pause felt like pure overhead sitting on top of an already-accurate system. That made sense while accuracy was climbing and every check kept coming back clean. It stopped making sense the moment "clean every time" became the reason nobody was watching for the one time it wasn't.

What I would leave alone: a citation pulled from a single, unambiguous report field, like an exact mileage number, is fine with no human check at all. The real risk lives in claims like "no accidents," which require reading across several sections of the report and can be misjudged by the model on a messy or unusual one.

The lesson: an accuracy number tells you how often something is right. It never tells you who's standing there the day it isn't, unless you built that person a job to do.

Now here is the same thing as a story

The short version above is what you'd say defending the fix to Ferrytale's trust and safety lead. Read this one for how the gap actually got found.

Innes Calloway has reviewed listing quality at Ferrytale Autos for four years, working the back office mid-morning shift. She can spot a rewritten fact from a straight one just by the rhythm of the sentence.

When the listing assistant launched, she opened the source report and checked every single citation against it, all five hundred listings that crossed her desk each week. It was slow, but it was thorough, and she trusted her own eyes over the tool's.

Hand sketched icon list titled The five letters. Five items: F find the person, L locate the habit, I identify the flip, P pinpoint the decision, S show the replay.
Five steps, and the third one, the flip, is the one that actually explains what happened to Innes over four months.

The citations kept coming back clean. Week after week, she'd spot-check a handful and find nothing wrong. So she checked fewer. Then only the ones that looked unusual. Then, without ever deciding to, she stopped opening the source report at all.

Hand sketched timeline titled The four months before the wrong citation shipped. Four milestones: launch every citation checked, month 2 spot rate falling, month 4 checks merged away highlighted, month 6 one wrong ships.
Nothing dramatic happened at any single point on this line. That's exactly what made it easy to miss.

Around month four, the team formally merged the review step into the auto-publish pipeline. Accuracy had held above 99% for weeks, and the pause before publishing looked, on paper, like a bottleneck with no upside left.

Then, in month six, one catastrophe. A sedan with a rebuilt salvage title got citied as "no accidents reported," because the report's rebuild note sat in a section format the model hadn't seen much of before. A buyer named Desmond drove three hours to see it, and only found out the truth when his mechanic lifted it onto a rack.

Hand sketched comparison diagram titled Small move, big snap. Left panel, a gauge icon labeled Checks some, caption spot rate falling for months. Right panel, a question mark box icon labeled Checks none, caption one wrong one ships clean.
The gap between these two panels isn't a percentage. It's four quiet months with no name.

He posted a video about it that thousands of people watched. Three unrelated sellers pulled their own, entirely accurate listings off Ferrytale that same week, just to be safe.

Innes never lost trust in the tool. She simply ran out of reasons to keep checking something that was always right, until the one time it wasn't.

Here's the decision I'd take back. We removed the human glance-check because it looked like friction on a system that had already proven itself. Nobody asked what happens on the one day in five hundred it hasn't.

Hand sketched metaphor scene titled Switch not dial. Left, a gauge icon labeled Many Settings, caption what we assumed. Right, a box icon labeled Two Positions, caption verify all or none.
We designed for a dial, careful checking that eases off gradually. What Innes actually had was a switch.

I'd put a version of the check back, but a smarter one. Not every citation, just the ones the model itself flags as a thin match, maybe fifteen out of five hundred a day, the ones built from a messy or unusual report section.

Hand sketched labeled parts diagram titled What's inside a citation line. Center document icon labeled Citation Line, with four callouts: source page, confidence, plain claim, report date.
The confidence tag is the part that turns a flat spot-check into one aimed at the actual risk.

Replay the same sedan under the new design: its citation gets flagged low-confidence because the rebuild note sat in an unusual format. Innes opens the source page, catches the mismatch, and the listing corrects before it ever reaches a buyer. Desmond never makes the drive.

I merged the check away because it looked like it was slowing down a system that had already earned trust. It took one buyer's ruined afternoon to see that trust earned in bulk still needs somewhere to fail safely, one case at a time.

The five steps, if you want to remember itNot a checklist. FLIPS is what tells you the flip isn't in the model. It's in the person watching it.

F
Find the person.
Innes Calloway, four years reviewing listing quality, mid-morning shift, back office.
Grounds the answer in one specific person's hands, not "reviewers" in general.
L
Locate the habit.
She stopped opening the source report to check citations, over about four months, as they kept coming back clean.
The habit is the product working, not a lapse in her judgment.
I
Identify the flip.
Checks some, falling steadily, to checks none. An over-trust flip, fired by the citations getting good, not bad.
The hardest step, and the one that explains why improvement itself is a kind of risk.
P
Pinpoint the old decision.
Merging the human glance-check into the auto-publish pipeline once accuracy crossed 99%.
A specific, reversible product decision, not a vague call to "add more review."
S
Show the replay.
With low-confidence citations flagged, Innes catches the rebuild-title mismatch before it publishes. Desmond never drives three hours.
Ends on a countable outcome, not a vague improvement.
Hand sketched decision tree titled Reading a wrong citation three ways. Root Citation says no accidents, branching to report is clean leads to citation correct, rebuilt title missed leads to wrong ships anyway, field is ambiguous leads to needs a human read, model flags low match leads to held for review.
The same headline citation supports four very different stories. Only the confidence flag tells you which one you're in.

The recap, one line per letter: find the person is Innes on the mid-morning shift, locate the habit is checking fewer citations each month as they held clean, identify the flip is checks-some collapsing to checks-none, pinpoint the decision is merging the check into auto-publish, and show the replay is catching the mismatch before Desmond ever makes the drive.

And if you want to be sure it really works, try it somewhere elseA different flip family, a funeral home instead of a car lot. Different building, same lesson about what a family submits.

Halcyon Rest Funeral Home uses an assistant that drafts obituary text, citing facts from a family's intake form: dates, survivors, career details. Ottoline Baptiste, the home's director, used to read every family's raw handwritten notes herself before typing anything.

Mapped onto FLIPS, but a different family entirely: this is a pre-editing flip, not over-trust. Find the person is Ottoline. Locate the habit is families, once they noticed which kinds of notes the assistant handled cleanly, starting to write their own submissions shorter and plainer, stripping out the messy, specific memories that made a life distinct. Identify the flip is families feeding the assistant their raw, detailed memories versus a sanitized, easy version they've learned it likes. Pinpoint the decision is the assistant failing quietly on messy input instead of saying which part confused it, so families invented their own theory about what "worked." Show the replay is the assistant flagging the exact sentence it couldn't parse, so a family adds one clarifying line instead of cutting the memory entirely, and the obituary keeps the detail that mattered.

Hand sketched decision tree titled Reading a wrong citation three ways, reused here to show a family's note being read three different ways by the drafting assistant.
A different building, but the same shape of tree: one input, several very different readings, and only a flag that tells you which one you got.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "citations build trust until they don't, and the fix is flagging what the model itself is unsure about," and stop.
Cost: there's no budget to build a confidence-scoring system this quarter. Start with routing any citation pulled from more than one report section, since that's already a rough proxy for risk.
The model gets better, for real: if citation accuracy climbs to 99.95%, that's still not a reason to remove the flagged-review step. It's a reason to expect the flagged pool to shrink, not disappear.

Where people run it wrong.
They treat a rising accuracy number as permission to remove the human step entirely.
They spot-check a flat, random sample instead of routing by the model's own confidence.
They measure the error rate and never think to measure the spot-check rate, which is the number that actually moves first.

How to use it live. When someone asks how citations change behavior, don't reach for "people trust it more." Ask yourself what specific checking behavior a person used to do, and name the exact month it quietly stopped.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: checks sometimes, then stops checking at all. Fires when the thing being checked gets better, not worse.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Innes Calloway, four years reviewing listing quality at Ferrytale Autos, mid-morning back-office shift.
3 · THE HABIT
What did Innes stop doing because the citations worked?
Tap to flip
ANSWER
She stopped opening the source vehicle history report to check citations against it, over about four months.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Checking some citations (a falling spot-check rate) versus checking none at all. No middle setting once the habit's gone.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging the human spot-check into the auto-publish pipeline once accuracy crossed 99%, since the pause looked like pure overhead by then.
6 · THE NUMBER
Fill in the blank: citation accuracy held around ___% across the first six months.
Tap to flip
ANSWER
99.8%. About one wrong citation in every 500, the exact rate that made the review step look unnecessary.
7 · THE REPLAY
Same rebuilt-title sedan, redesigned pipeline. What changes?
Tap to flip
ANSWER
The citation gets flagged low-confidence, Innes catches the mismatch before publishing, and the buyer never drives three hours to a hidden problem.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and which flip family?
Tap to flip
ANSWER
Halcyon Rest Funeral Home's obituary drafting assistant. There, the flip family is pre-editing: families sanitize their own submissions once they learn what the assistant handles cleanly.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Innes stop checking citations, according to this answer?
  • A. She was told to stop by her manager.
  • B. The citations kept coming back clean, so checking felt like less and less of a useful step.
  • C. The tool stopped showing her the source report.
  • D. She didn't have time due to a staffing shortage.
Show hint
Look at "Locate the habit" in the walkthrough.
Show answer
B. The habit faded because it kept being unnecessary, which is exactly what an over-trust flip looks like from the inside.
True or false
2. True or false: Innes "checked a bit less carefully" as time went on, rather than stopping entirely.
  • True
  • False
Show hint
Look at the bar chart: spot-check rate goes from 100% to 0%.
Show answer
False. There's no middle setting. The checking fully stopped once the pipeline merge removed the step entirely.
Fill in the blank
3. Fill in the blank: under the redesign, roughly ___ citations out of 500 a day get flagged for a human check.
Show hint
Look at Stage 6 of the walkthrough.
Show answer
15. Only the model's own low-confidence citations get routed to Innes, not a flat random sample.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Merging the human glance-check into the auto-publish pipeline. It made sense while accuracy kept climbing and every check came back clean.
Short answer, where it wouldn't matter
5. Name a kind of citation where no human check is genuinely needed, even under the new design.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A citation from a single unambiguous field, like an exact mileage number, carries little risk of misreading.
Short answer, apply it yourself
6. Pick a product you use that shows citations or sources. What's a habit of double-checking you used to have that you've quietly stopped?
Show hint
Think about a search or research tool where you used to click through to the source, and don't anymore.
Show answer
Model answer: Many people say they used to click through to the linked source on an AI summary and stopped once the summaries kept matching what they found.
Before you close the answer
Why this works
Tests whether you understand that citations change trust, not accuracy, and whether you can name a specific behavior that fades to nothing instead of describing a vague feeling of increased confidence.
Follow-up traps
"Couldn't you just keep the spot-check rate flat, say 5%, instead of routing by confidence?" Response: a flat random sample mostly catches easy, already-clean citations, missing the specific messy cases where errors actually cluster.

"Isn't a flagged-review queue just the old review step with extra steps?" Response: no, the old step reviewed everything at a fixed rate regardless of risk; the new one scales down as the model improves, since only genuinely uncertain cases ever reach it.
If pressed
Ferrytale's real confidence flag isn't the model's own softmax score; it's a separate check for whether the cited fact came from a single report field or was stitched together across more than one section, since that stitching is where nearly all of the real errors clustered.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more