How do citations change user behaviour, and what happens when they are wrong?
Ferrytale Autos is a used-car marketplace. Its listing assistant reads each car's official vehicle history report and writes a plain-language citation, like "No accidents reported," straight onto the listing. Innes Calloway has reviewed listing quality at Ferrytale for four years, working the back office mid-morning shift.
- Route only the model's low-confidence citations to a human, not a random sample.Why: a flat spot-check rate misses exactly the cases most likely to be wrong.
- Never let a high accuracy rate justify removing the review step entirely.Why: 99.8% accurate still means someone gets the other 0.2%, with nobody watching for them.
- Track the spot-check rate itself, not just the error rate.Why: the spot-check rate can fall to zero for months before an error rate ever shows it.
- Leave simple, unambiguous fields, like an exact mileage number, uncontested.Why: not every citation carries the same risk, and treating them all the same wastes review time.
- Give the reviewer the exact source page next to the citation, not just the claim.Why: a citation you can't trace back to its source isn't checkable, it's just a second claim.
How to answer this, stage by stage
Nobody's grading whether you know citations can be wrong. Everyone knows that. They're grading whether you can name what actually happens to a person's behavior once citations are usually right.
Let's learn
Ferrytale Autos' listing assistant reads a used car's official vehicle history report and writes a plain-language citation onto the listing, like "No accidents reported."
Before the assistant, a copy editor read the full multi-page report by hand and wrote that line herself, about six minutes per listing. With the assistant, the same line appears instantly, and it's held at 99.8% accurate across its first six months, about one wrong citation in every 500.
The turn: the rare wrong citation isn't really the problem. Ferrytale's assistant is right on the overwhelming majority of listings. The real problem is that nobody was left checking specifically for the wrong ones, because spot-checks had stopped entirely once accuracy proved itself out.
At its worst: one wrong "no accidents" citation on a car with a rebuilt salvage title sends a buyer driving three hours to see it, only to find the truth at a mechanic's lift. He posts a furious video that thousands of people see, and several unrelated sellers pull their own perfectly accurate listings out of sheer caution. A tool that was right 499 times out of 500 still manages to do real, visible harm.
What I would leave alone: a citation pulled from a single, unambiguous report field, like an exact mileage number, is fine with no human check at all. The real risk lives in claims like "no accidents," which require reading across several sections of the report and can be misjudged by the model on a messy or unusual one.
The lesson: an accuracy number tells you how often something is right. It never tells you who's standing there the day it isn't, unless you built that person a job to do.
Now here is the same thing as a story
The short version above is what you'd say defending the fix to Ferrytale's trust and safety lead. Read this one for how the gap actually got found.
Innes Calloway has reviewed listing quality at Ferrytale Autos for four years, working the back office mid-morning shift. She can spot a rewritten fact from a straight one just by the rhythm of the sentence.
When the listing assistant launched, she opened the source report and checked every single citation against it, all five hundred listings that crossed her desk each week. It was slow, but it was thorough, and she trusted her own eyes over the tool's.
The citations kept coming back clean. Week after week, she'd spot-check a handful and find nothing wrong. So she checked fewer. Then only the ones that looked unusual. Then, without ever deciding to, she stopped opening the source report at all.
Around month four, the team formally merged the review step into the auto-publish pipeline. Accuracy had held above 99% for weeks, and the pause before publishing looked, on paper, like a bottleneck with no upside left.
Then, in month six, one catastrophe. A sedan with a rebuilt salvage title got citied as "no accidents reported," because the report's rebuild note sat in a section format the model hadn't seen much of before. A buyer named Desmond drove three hours to see it, and only found out the truth when his mechanic lifted it onto a rack.
He posted a video about it that thousands of people watched. Three unrelated sellers pulled their own, entirely accurate listings off Ferrytale that same week, just to be safe.
Here's the decision I'd take back. We removed the human glance-check because it looked like friction on a system that had already proven itself. Nobody asked what happens on the one day in five hundred it hasn't.
I'd put a version of the check back, but a smarter one. Not every citation, just the ones the model itself flags as a thin match, maybe fifteen out of five hundred a day, the ones built from a messy or unusual report section.
Replay the same sedan under the new design: its citation gets flagged low-confidence because the rebuild note sat in an unusual format. Innes opens the source page, catches the mismatch, and the listing corrects before it ever reaches a buyer. Desmond never makes the drive.
I merged the check away because it looked like it was slowing down a system that had already earned trust. It took one buyer's ruined afternoon to see that trust earned in bulk still needs somewhere to fail safely, one case at a time.
The five steps, if you want to remember itNot a checklist. FLIPS is what tells you the flip isn't in the model. It's in the person watching it.
The recap, one line per letter: find the person is Innes on the mid-morning shift, locate the habit is checking fewer citations each month as they held clean, identify the flip is checks-some collapsing to checks-none, pinpoint the decision is merging the check into auto-publish, and show the replay is catching the mismatch before Desmond ever makes the drive.
And if you want to be sure it really works, try it somewhere elseA different flip family, a funeral home instead of a car lot. Different building, same lesson about what a family submits.
Halcyon Rest Funeral Home uses an assistant that drafts obituary text, citing facts from a family's intake form: dates, survivors, career details. Ottoline Baptiste, the home's director, used to read every family's raw handwritten notes herself before typing anything.
Mapped onto FLIPS, but a different family entirely: this is a pre-editing flip, not over-trust. Find the person is Ottoline. Locate the habit is families, once they noticed which kinds of notes the assistant handled cleanly, starting to write their own submissions shorter and plainer, stripping out the messy, specific memories that made a life distinct. Identify the flip is families feeding the assistant their raw, detailed memories versus a sanitized, easy version they've learned it likes. Pinpoint the decision is the assistant failing quietly on messy input instead of saying which part confused it, so families invented their own theory about what "worked." Show the replay is the assistant flagging the exact sentence it couldn't parse, so a family adds one clarifying line instead of cutting the memory entirely, and the obituary keeps the detail that mattered.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "citations build trust until they don't, and the fix is flagging what the model itself is unsure about," and stop.
Cost: there's no budget to build a confidence-scoring system this quarter. Start with routing any citation pulled from more than one report section, since that's already a rough proxy for risk.
The model gets better, for real: if citation accuracy climbs to 99.95%, that's still not a reason to remove the flagged-review step. It's a reason to expect the flagged pool to shrink, not disappear.
Where people run it wrong.
They treat a rising accuracy number as permission to remove the human step entirely.
They spot-check a flat, random sample instead of routing by the model's own confidence.
They measure the error rate and never think to measure the spot-check rate, which is the number that actually moves first.
How to use it live. When someone asks how citations change behavior, don't reach for "people trust it more." Ask yourself what specific checking behavior a person used to do, and name the exact month it quietly stopped.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a flagged-review queue just the old review step with extra steps?" Response: no, the old step reviewed everything at a fixed rate regardless of risk; the new one scales down as the model improves, since only genuinely uncertain cases ever reach it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Trust, transparency and explainability in UX
- #1 What does a user need to see to trust an AI recommendation?
- #2 Explain the difference between explainability and transparency in a product context.
- #4 Design the disclosure that tells a user they are talking to an AI.
- #5 When does showing the model's reasoning help, and when does it reduce trust?
- #6 Critique a design that surfaces a chain of thought to end users.
- #7 How much should you tell users about which model powers a feature?