Artifact critiqueIntermediateDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #8

Critique a design that presents an AI answer with the same authority as a database lookup.

AUDIT a card that looks verified and a card that is verified should never wear the same suit

Brant & Vollmer is a mid-size litigation firm. Petra Falkenrath is a senior litigation associate there. Verdictly is an AI legal research tool the firm licenses alongside CourtWire, its paid, verified case-law database. Both tools show their answers inside the exact same card.

The direct answer
The design fails because it borrows a verified record's whole visual language, same card, same bold case name, same pin-cite format, for a paragraph nobody has checked yet. Fix it by stripping every earned signal until Verdictly names its own model version and run date on the card, shows the actual source passages behind the claim, and carries a plain, unmissable mark that no person has confirmed it.
Fix it in this order
  1. Strip the borrowed signals off the AI answer card: no bold case names, no pin-cite formatting, until they're earned.Why: those signals currently do the convincing, not the actual research behind them.
  2. Print the model version and the exact run date on the card itself.Why: a wrong answer with no version pin can't be traced back to a specific, fixable run.
  3. Show the real source passages the answer drew from, not just a citation number.Why: a citation number is a claim; a quoted passage is something an associate can actually check in ten seconds.
  4. Add a plain, visible "not yet confirmed" state until a person has checked it.Why: without this, the card's silence reads as confirmation by default.
  5. Test the redesigned card against real past research questions, not the vendor's demo set.Why: a demo is built to look good; a firm's own hard questions are what the card actually has to survive.

How to answer this, stage by stage

The interviewer isn't asking whether you can spot a UI that's a little busy. They're asking whether you can name the exact visual choice that's doing the lying.

Stage 1
Name the artifact you're actually critiquing
Say it like this
"I'll critique Verdictly, an AI legal research tool at Brant & Vollmer, specifically the screen where its AI-generated answers render in the same card as CourtWire's verified case-law pulls."
Why this works
Locks the critique to one concrete screen instead of a general complaint about AI in legal tools.
Stage 2
Say your method out loud
Say it like this
"I'll use AUDIT. Ask who paid for the claim, uncover the eval set behind it, demand the version pin, isolate what's missing from the screen, then test it myself against real cases."
Why this works
Shows a repeatable way to judge any artifact, not just a one-off gut reaction to this screen.
Stage 3
Name what's actually wrong, not just that something feels off
Say it like this
"The problem isn't that Verdictly can be wrong. Every model can be wrong. The problem is the screen gives a wrong answer the exact same authority as a fact nobody disputes."
Why this works
Separates the critique from a vague complaint about AI accuracy.
Stage 4
Deliver the verdict, before any evidence
Say it like this
"Strip the borrowed signals, the bold case name and pin-cite look, off the AI card until they're earned. Put the model version, the source passage, and an unconfirmed flag on it instead."
Why this works
This is the direct answer, and it names the fix, not just the flaw.
Stage 5
Back it with the audit itself
Say it like this
"When I ran forty of the firm's actual past research questions back through Verdictly, six answers cited cases that were either misquoted or didn't exist at all. Every one of those six looked identical to the true ones on screen."
Why this works
Turns "this design is risky" into a specific, checkable failure rate.
Stage 6
Close on the one line
Say it like this
"Never let a screen give an unverified paragraph the same authority as a verified record. Whatever separates a claim from a fact has to show up somewhere a reader can actually see it."
Why this works
Restates the verdict in one breath, ready for a live follow-up.

The card that borrowed someone else's authority

For six years, Petra verified every legal research answer by hand against CourtWire's records before it went anywhere near a client memo, about 25 minutes a question, across roughly 30 questions a week. That's close to 12 and a half hours a week just checking.

Hand sketched comparison titled Two boxes one look. Left, a green document icon labeled case citation, caption verified record. Right, a red question mark box icon labeled AI answer, caption generated unverified.
Two very different kinds of paragraph. One card style, worn by both.

Verdictly drafted a first-pass answer to the same kind of question in under 15 seconds. Because its answer card used CourtWire's exact layout, font, and citation style, associates began treating a Verdictly answer as pre-verified the same way a CourtWire pull already was. Verification time on a Verdictly answer fell to about 3 minutes.

Verdictly answers cited in a filed document without independent verification, tracked weekly
80% 40 0 Week 1 Week 10 8% 61%
Nobody was watching this number climb. By the week the fabricated case was caught, three out of five filed answers had never been independently checked.

Here's the turn: the twenty-two minutes Verdictly saved per question was never the real story. The real problem is that an unverified paragraph, once it wears a verified record's exact clothes, walks straight into a client memo without anyone stopping to ask whether it should.

Hand sketched labeled parts diagram titled What's missing from the answer screen. A red question mark box icon at the center labeled AI Answer Screen, with four callouts: no version pin, no eval set named, no confidence, no source excerpt.
Four missing facts. Any one of them would have been enough to slow Petra down.

At its worst, this design flaw doesn't cost a firm an afternoon of embarrassment. A fabricated case citation filed in a real brief can draw court sanctions, a malpractice complaint, and a client's whole file being reviewed from scratch for anything else the tool touched.

The decision I would take back Verdictly's design team reused the firm's existing "verified record" card component for AI answers too, so the whole research tool would feel like one consistent product. That made sense when Verdictly was new, rarely used, and every answer got checked anyway out of habit. It stopped making sense once associates trusted it enough to stop checking.

What I would leave alone: CourtWire's actual verified case-pull cards don't need any redesign. They're correctly earning the authority they display, because a person confirmed the record was real before it ever reached the screen.

The lesson: a screen can lie about how trustworthy it is without a single wrong word on it, just by choosing to look like something else that already earned its trust honestly.

Now here is the same thing as a story

The short version above is what you'd say defending this fix to the firm's managing partner. Read this one for how the near miss actually happened.

The Brant & Vollmer research floor goes quiet around 6 p.m., after the partners leave and the associates start second-drafting memos due the next morning. Petra had worked litigation research there for six years and could smell a weak citation before she finished reading the sentence around it.

Verdictly arrived in the spring, and the first months were good. Petra would pull a first-pass answer, then walk every citation back to CourtWire by hand before it went into anything real. The tool was fast. Her own habits didn't change at first.

Knowledge spark: what does "hallucinated" mean for a legal citation? It means the model generated something that reads exactly like a real case, a plausible name, a plausible court, a plausible year, that either doesn't exist at all or doesn't say what the model claims it says. It isn't a typo. It's a fully formed, confident-sounding fact with nothing real behind it.

By midsummer, she was spot-checking maybe half of Verdictly's answers instead of all of them, since the ones she'd checked kept coming back clean. By early fall, she was checking only the ones that felt unusual, because the card looked exactly like a CourtWire pull whether it was right or not, and nothing on the screen ever told her which was which.

Hand sketched quadrant titled Which claims deserve scrutiny. Axes stakes if wrong versus how it's displayed. Case citation sits low stakes, styled like fact. AI answer sits high stakes, styled like fact, the dangerous corner. Settlement estimate sits in the middle.
The AI answer sits in the one corner nothing should ever sit in: high stakes, dressed exactly like a fact.

The trigger was small. A junior associate asked Petra, in passing, whether a case Verdictly had cited that morning was "one of the real ones or one of the AI ones," and Petra realized she had no way to answer without opening a second window and searching for it herself.

The card had never once said which kind of paragraph it was. Petra had been supplying that judgment herself, out of habit, for months, and the habit had quietly stopped.

She pulled forty of the firm's real research questions from the past quarter and ran each one back through Verdictly, checking every citation by hand this time. Six answers, fifteen percent, cited a case that was either misquoted or did not exist. One of those six had already been forwarded into a client memo two weeks earlier, caught only because a partner happened to recognize the case name as unfamiliar.

Hand sketched flow diagram titled The AUDIT check step by step. Five boxes: ask who paid, find the eval set, pin the version, isolate gaps, test it yourself highlighted.
The fifth box is the one that actually caught the problem. The other four explain why nobody caught it sooner.

Brant & Vollmer didn't drop Verdictly. They rebuilt its answer card: a visible model version and run date, the actual source passage quoted beneath any citation, and a plain amber "not yet confirmed" strip that only clears once a person has checked the specific citation, not just opened the page.

AUDIT, five checks that separate a claim from a factNot a general AI-skepticism lecture. AUDIT is what tells you exactly which fact a good screen is hiding.

A
Ask who paid. Whose claim of reliability is this?
Verdictly's own marketing, not an independent source, claimed it was "as reliable as your case law database."
A vendor's own claim about its own accuracy is not evidence of that accuracy.
U
Uncover the eval set. What was this actually tested against?
No named eval set existed, only a handful of vendor demo queries chosen to look good.
A score with no named test set behind it is decoration, not evidence.
D
Demand the version pin. Which model, on which date?
The original card showed neither, so a bad answer couldn't be traced to a specific, fixable run of the model.
The hardest step, and the one this critique turns on: without a version pin, nothing about the failure is reproducible.
I
Isolate what's missing. What does the screen leave out?
No confidence signal, no source passage, no way to see what documents the answer actually drew from.
What a screen omits is usually more telling than what it shows.
Hand sketched icon list titled Signs an AI answer is dressed up as fact. Three items: a scale icon labeled same font weight as verified data, a gauge icon labeled no hedge language anywhere, a box icon labeled no export or version stamp.
Any one of these three, alone on a screen, is worth a second look before it ships.
T
Test it yourself. Replicate the claim on your own real cases.
Forty of the firm's own past research questions, checked by hand, turned up a fifteen percent fabrication rate the vendor's own materials never mentioned.
This is the move that actually caught the problem, not a review of the vendor's paperwork.

The recap, one line per letter: ask who paid means the vendor's own accuracy claim isn't evidence, uncover the eval set means no named test set makes a score meaningless, demand the version pin means an untraceable answer can't be fixed, isolate what's missing means the absence of confidence and sourcing is the real design flaw, and test it yourself means the firm's own forty questions are what finally surfaced the fifteen percent failure rate.

And if you want to be sure it really works, try it somewhere elseSame five checks, a hospital's medical coding AI instead of a law firm's research tool. A different missing signal breaks the second story.

Briarstone Health uses CodeAssure, an AI tool that suggests billing codes for patient encounters, displayed in the exact same grid and font as codes pulled straight from the verified CMS code database. Soren Halvorsen, a compliance auditor there, runs the same critique on CodeAssure's suggestion screen. Mapped onto AUDIT: ask who paid means CodeAssure's vendor cites its own internal accuracy number with no outside review; uncover the eval set means that number was never tied to a named, published set of real encounter notes; demand the version pin means the suggested-code grid shows no model version, so a wrong code can't be traced to a specific update.

The isolate step is where this story diverges from Petra's. Verdictly's screen was missing a confidence signal. CodeAssure's screen is missing something else: it never shows which part of the patient's encounter note the suggested code was actually drawn from, so a coder has no way to spot-check a code against the specific line that justified it. Soren's test-it-yourself step pulled fifty recent claims and hand-checked each suggested code against the full encounter note.

Hand sketched timeline titled How the redesign rolled out, reused here for CodeAssure. Four milestones: flaw flagged week 1, user test week 3, confidence added week 6 highlighted, trust calibrates week 10.
Same shape of fix, ten weeks later, in a hospital billing office instead of a law firm.
Miscoded claims caught before submission, shared-style screen versus flagged screen
100% 50 0 12% Shared-style screen 84% Flagged screen
Same coders, same fifty claims re-checked, seven times more miscodes caught once the screen stopped pretending every suggestion was equally certain.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "strip the borrowed authority, add a version pin, a source excerpt, and an unconfirmed flag," and stop.
Cost: there's no budget this quarter for a full redesign. Say so honestly, and ship the unconfirmed flag alone first, since a single visible flag costs almost nothing and closes the biggest gap.
The model gets better, for real: if Verdictly's accuracy genuinely improves to near-zero fabrication, that's still not a reason to remove the flag. It's a reason to watch the fabrication rate even more closely, since a better model makes the rare miss even easier to miss.

Where people run it wrong.
They critique the model's accuracy and never look at the screen it's rendered on.
They accept a vendor's own reliability claim without asking what it was tested against.
They add a confidence score but keep the same borrowed visual weight, so the number gets ignored anyway.

How to use it live. When someone hands you a screen to critique, ask one question before anything else: if I covered up every word and just looked at the shapes, fonts, and boxes, could I tell which claims have been checked? If not, that's the actual flaw, not whatever the words say.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that hands you a design and asks you to judge it?
Tap to flip
ANSWER
AUDIT: ask who paid, uncover the eval set, demand the version pin, isolate what's missing, test it yourself. Built for artifact critique, not for designing or measuring something yourself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Petra Falkenrath, a senior litigation associate at Brant & Vollmer, who verified research answers by hand for six years before Verdictly.
3 · THE FLAW
What's the actual design flaw this critique names?
Tap to flip
ANSWER
The AI answer card borrows CourtWire's exact verified-record styling, so an unchecked paragraph and a confirmed fact are visually indistinguishable.
4 · THE HARD STEP
Which AUDIT step is the hardest, and why?
Tap to flip
ANSWER
Demand the version pin. Without a specific model version and run date on the card, a bad answer can never be traced back to something fixable.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Reusing the firm's verified-record card component for Verdictly's AI answers too, to make the tool feel consistent, made when usage was low and every answer got checked anyway.
6 · THE NUMBER
Fill in the blank: when Petra tested 40 real research questions by hand, ___ answers cited a case that was misquoted or didn't exist.
Tap to flip
ANSWER
6 out of 40, a 15 percent fabrication rate, with every one of those six rendered identically to the true answers on screen.
7 · THE FIX, VERIFIED
Same forty questions, redesigned card. What changes?
Tap to flip
ANSWER
Every fabricated citation still shows up, but now it carries an amber, unconfirmed flag and a source passage a reader can check in seconds, instead of looking identical to a verified pull.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's missing there instead?
Tap to flip
ANSWER
Briarstone Health's CodeAssure medical coding tool. There, the missing piece isn't a confidence score, it's the source excerpt from the encounter note that would let a coder check a suggested code against the actual visit.

Check yourself Score: 0 / 0

True or false
1. True or false: the core flaw in this critique is that Verdictly's model is too inaccurate to use.
  • True
  • False
Show hint
Look at Stage 3's reframe.
Show answer
False. Every model can be wrong sometimes. The flaw is that the screen gives a wrong answer the same visual authority as a verified fact.
Multiple choice
2. Why does "demand the version pin" count as the hardest AUDIT step in this critique?
  • A. Version numbers are the hardest thing for a vendor to provide.
  • B. Without it, a bad answer can never be traced back to a specific, fixable run of the model.
  • C. Associates find version numbers confusing to read.
  • D. It's the only step that requires a lawyer, not a designer.
Show hint
Look at the D step in the AUDIT recap.
Show answer
B. An untraceable failure can't be diagnosed or fixed, which is what makes the missing version pin more than a cosmetic gap.
Fill in the blank
3. Fill in the blank: at Briarstone Health, coders using the redesigned, flagged screen caught ___ percent of miscoded claims, up from 12 percent on the shared-style screen.
Show hint
Look at the bar chart in Section 4.
Show answer
84 percent. Roughly seven times more miscodes caught, once the screen stopped treating every suggestion as equally certain.
Short answer, apply it yourself
4. Think of a screen you use that mixes a verified fact with an AI-generated guess in the same visual space. Could you tell them apart without clicking anything?
Show hint
Cover the words and look only at the shapes, fonts, and boxes.
Show answer
Model answer: Most screens fail this test, which is exactly the gap the AUDIT framework is built to surface.
Short answer, why no middle setting
5. Why wouldn't adding just a small confidence percentage next to the AI answer have been enough, on its own?
Show hint
Look at the isolate step and the follow-up trap about confidence scores.
Show answer
Model answer: A number rendered in the same borrowed visual weight gets ignored the same way the words were. The container has to change, not just add a number to it.
Short answer, where it wouldn't matter
6. Name a part of Verdictly's screen where this same scrutiny genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: CourtWire's own verified case-pull cards. A person already confirmed those records were real before they reached the screen.
Before you close the answer
Why this works
Tests whether you can name the specific visual mechanism that makes a design misleading, rather than offering a general worry about AI accuracy, and whether your proposed fix actually addresses that mechanism instead of bolting a number onto the existing problem.
Follow-up traps
"Wouldn't just adding a confidence score fix this?" Response: not if it's rendered in the same borrowed visual weight; the container has to visibly differ, not just carry an extra number nobody reads.

"Isn't stripping the styling going to make associates trust the tool less overall, even when it's right?" Response: that's the point, calibrated distrust on the fifteen percent that's wrong is worth more than blanket trust that occasionally files a fake case.
If pressed
The redesigned card also logs every time an associate marks an answer "confirmed," with their name and timestamp, so a later audit can see who actually checked a citation instead of just trusting that someone must have.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more