CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #9

How do you display uncertainty in a data extraction feature with 20 extracted fields?

ORDER rank the 20 fields by what a wrong one actually costs, not by where it sits on the form

Oakferry Mutual is a regional auto and home insurer. Marcus Villanueva is its Claims Operations Lead. ClaimScan is the AI tool that reads a photographed claim form and extracts 20 fields, from policy number to injury severity code, before an adjuster ever opens the file.

The direct answer
Don't give all 20 fields the same confidence treatment. Rank them by what a wrong value actually costs and how hard it is to undo, then only the top few, the ones that drive payout or eligibility, get a loud, unmissable low-confidence flag. The rest get a quiet inline dot, or nothing at all, because a wrong mailing address costs a phone call and a wrong injury code costs a reopened claim.
Design it in this order
  1. Rank all 20 fields by cost-if-wrong times how hard the error is to undo, not by their order on the form.Why: a form's layout was never designed around which mistakes actually hurt.
  2. Give the top tier, usually four or five fields, a loud, adjuster-facing flag whenever confidence dips.Why: these are the fields that change what the claimant actually gets paid.
  3. Give the middle tier a quiet, collapsed indicator, visible on hover, not by default.Why: worth checking eventually, not worth interrupting a review for.
  4. Let the bottom tier show no flag unless confidence is extremely low.Why: a wrong mailing address gets caught and fixed on the very next phone call anyway.
  5. Track confidence by field, monthly, after launch, not just once at rollout.Why: extraction quality can quietly drop on one field type long before anyone notices at the aggregate level.

How to answer this, stage by stage

Nobody's grading whether you can name all twenty fields off memory. They're grading whether you know which four actually deserve an adjuster's attention.

Stage 1
Scope it to one real form
Say it like this
"I'll answer this for ClaimScan at Oakferry Mutual, a tool that pulls 20 fields off a photographed auto and home claim form for Marcus Villanueva's adjusters to review."
Why this works
Stops "how do you show uncertainty on 20 fields" from turning into an abstract UI pattern with nothing behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what all 20 fields are actually competing to protect. Reversibility, which wrong field is hardest to undo. Dependency, what has to happen before what. Evidence, what I could learn cheaply first. Rank, the actual tiering."
Why this works
Shows a method for tiering the fields instead of a gut feeling about which ones look important.
Stage 3
Name what's really being protected
Say it like this
"This isn't about extraction accuracy in general. It's about which specific wrong field changes what a claimant actually gets paid, or whether the claim gets approved at all."
Why this works
Without naming this, "rank the fields" is just an opinion about which ones look scarier.
Stage 4
Give the tiering, before any reasoning
Say it like this
"Five fields that drive payout or eligibility get a loud flag. Eight fields that matter but are cheap to fix get a quiet, collapsed dot. Seven fields that are trivial to correct get nothing unless confidence is extremely low."
Why this works
This is the direct answer, said as an actual tiering instead of a vague promise to "show confidence."
Stage 5
Prove the tiering with the near miss
Say it like this
"ClaimScan's original screen showed one average confidence badge for the whole form. It read 94 percent healthy the same week the injury-code field specifically was running closer to 81 percent, and nothing told an adjuster to look there instead of everywhere."
Why this works
Turns "an average hides the field that matters" into a specific, checkable number.
Stage 6
Say what you'd learn cheaply before rolling it out fully
Say it like this
"Before flagging all 20 fields this way company-wide, I'd run the tiered design on one region's claims for a month and watch whether adjusters actually open the flagged fields more than the unflagged ones."
Why this works
Shows you're checking the design works before betting the whole rollout on it.
Stage 7
Say what you'd leave alone
Say it like this
"I wouldn't put a confidence flag on the claimant's mailing address or the date the form was submitted. Both get caught and corrected on the very next phone call regardless."
Why this works
Shows judgment about where uncertainty display matters instead of flagging every field out of caution.
Stage 8
Close on the one line
Say it like this
"Twenty fields don't deserve twenty equal flags. Rank them by what a wrong one costs, and spend the adjuster's attention only where a mistake actually changes the outcome."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

What happens when one badge has to speak for twenty different kinds of maybe

Before ClaimScan, an intake clerk typed all 20 fields off a photographed claim form by hand, about 14 minutes a claim, across roughly 45 claims a day per clerk. ClaimScan extracts the same 20 fields automatically in under 5 seconds, and an adjuster reviews the output before the claim moves forward.

Hand sketched icon list titled Not all 20 fields carry equal risk. Four items: a document icon labeled policy number easy to verify, a scale icon labeled injury code drives the payout, a gauge icon labeled date of loss timing sensitive, a question mark box icon labeled witness contact hard to confirm.
Twenty fields, four very different kinds of consequence if one comes out wrong.

Here's the turn: the extraction mistakes weren't the real problem. The real problem was one aggregate confidence badge sitting at the top of the form, telling adjusters the claim was "94 percent healthy" while hiding that one specific field, the one that actually decides the payout, was running far worse than the average.

Where adjuster review time actually goes, by field risk tier, before and after tiering
16 min 8 0 14 min, flat review Before tiering Tier 1, 6 min Tier 2, 3 min Tier 3, 0.5 min After tiering, 9.5 min
Tiering didn't just protect the fields that matter. It cut a third of the review time off fields that never needed the same scrutiny.
Hand sketched quadrant titled Which fields get a confidence flag first. Axes cost if wrong versus easy to verify. Mailing address sits cheap and easy to verify. Injury code sits expensive and hard to verify, the priority corner. Policy number sits cheap and easy. Date of loss sits in the middle.
The injury code sits alone in the corner that matters: expensive to get wrong, and hard to catch after the fact.

At its worst, an averaged confidence score doesn't just waste a reviewer's time. It hides the one field that decides how much a claimant actually gets paid, behind a badge that looks perfectly fine.

The decision I would take back ClaimScan's first version averaged all field confidences into one score shown once at the top of the form, to keep the review screen simple. That made sense when ClaimScan only extracted six straightforward fields at launch. It stopped making sense once the product grew to 20 fields of wildly different difficulty and consequence.

What I would leave alone: the claimant's mailing address and the date the form was submitted don't need this same tiered flagging. Both get caught and corrected on the very next phone call, regardless of what the screen shows.

The lesson: an average is a decision to treat twenty different risks as one, and nobody at Oakferry chose that on purpose. It just happened to be the simplest number to put on the screen.

Now here is the same thing as a story

The short version above is what you'd say defending this design to Oakferry's VP of Claims. Read this one for how close the actual near miss came.

Marcus Villanueva could eyeball a claim form and know within a minute which fields would give an adjuster trouble, nine years running Oakferry's claims floor had taught him that much. ClaimScan's rollout, at first, looked like it was making that instinct obsolete in a good way.

The first two months were smooth. Adjusters reviewed the extracted fields, corrected the occasional typo, and moved claims through faster than ever. The 94 percent confidence badge at the top of every form sat there, green, reassuring, unremarkable.

Knowledge spark: why would one field extract worse than the others? A handwritten injury code, squeezed into a small box and sometimes overwritten or circled, is a much harder read for a model than a printed policy number in a clean font. Averaging confidence across fields this different is like averaging a clear photo with a blurry one and calling the result "pretty sharp."

By month three, adjusters had started treating the green badge as a green light for the whole claim, not just an average. Nobody had decided that on purpose. It's just what a single number at the top of a form invites you to do.

Hand sketched comparison titled Reversible or not. Left, a green document icon labeled mailing typo, caption fixed in an email. Right, a red scale icon labeled wrong injury code, caption drives payout hard to claw back.
Both are extraction errors. Only one of them requires reopening a paid claim.

The near miss came on a claim where the handwritten injury code had been misread, a single character off, changing a moderate injury into a minor one. The claim's overall confidence still read 94 percent, since nineteen other fields had extracted cleanly. The claim was approved and paid at the lower amount.

Nobody skipped a step. The adjuster reviewed a form that said 94 percent healthy, on a claim where the one field that mattered most was wrong.

A follow-up call from the claimant's doctor's office, three weeks later, is what actually caught it, not the review screen. Marcus pulled ClaimScan's field-by-field logs and found the injury-code field specifically had been running at 81 percent accuracy for the whole quarter, buried inside an average that never dropped below 90.

Hand sketched flow diagram titled What unblocks what. Five boxes: scan arrives, OCR extracts, confidence scored highlighted, adjuster reviews, payout decision.
The third box is where the whole claim's fate actually gets decided. It used to be the least visible one on the screen.

Oakferry didn't roll ClaimScan back. They rebuilt the review screen around three tiers instead of one average: a loud flag on the five fields that drive payout whenever confidence dips below 92 percent, a quiet dot on eight middle fields, and nothing at all on the seven fields that barely matter if they're briefly wrong.

Hand sketched labeled parts diagram titled Inside a claim intake form. A document icon at the center labeled Claim Form, with four callouts: policy number, injury code, date of loss, witness contact.
Four of the twenty fields, each with a genuinely different reason to look, or not look, twice.

ORDER, ranked by what you can't undoNot a checklist of twenty fields. ORDER is what decides which four of them actually earn an adjuster's attention.

O
Outcome. What all 20 fields are competing to protect.
Whether the claimant gets paid the right amount, and whether the claim is even eligible, not a flat extraction-accuracy score.
Without this, tiering the fields is just a guess about which ones look important.
R
Reversibility. Which wrong field is hardest to undo.
A wrong mailing address gets fixed on a phone call. A wrong injury code, once a claim is paid, needs a reopened file and sometimes a regulatory notice.
The hardest step, and the one the whole tiering actually turns on.
D
Dependency. What has to happen before what.
Confidence scoring has to run before an adjuster can review anything, and the review has to clear before a payout decision gets made.
Some of the order is forced by the pipeline itself, not just by judgment.
E
Evidence. What to learn cheaply before a full rollout.
Running the tiered design on one region's claims for a month, watching whether adjusters actually open the flagged fields more often than the unflagged ones.
Shows the design gets tested against real behavior, not just shipped on a hunch.
R
Rank. State the tiers and defend the top one.
Five loud-flag fields, eight quiet-dot fields, seven unflagged fields, defended by which one changes the actual payout.
Turns "show uncertainty on 20 fields" into an arguable, specific design.
Hand sketched timeline titled The tiered rollout, timed. Four milestones: pilot tier ships month 1, adjusters trained month 2 highlighted, full coverage month 4, audit confirms it month 6.
Six months from one flat badge to a screen that spends attention where it's actually earned.

The recap, one line per letter: outcome is the actual payout and eligibility decision, not a generic accuracy number, reversibility is why an injury code outranks a mailing address, dependency is the pipeline order confidence scoring has to run in, evidence is a one-region pilot before full rollout, and rank is the three-tier design itself, loud where it counts and silent where it doesn't.

And if you want to be sure it really works, try it somewhere elseSame five letters, a shipping terminal's customs manifest instead of a claim form. A different dependency breaks the second story.

Brindlemark Terminal runs customs brokerage for container freight. Chinwe Obiora leads the brokerage desk there, using ManifestClear, an AI tool that extracts around 20 fields from a shipping manifest, container ID, weight, tariff code, hazardous-material flag, consignee, and more. Mapped onto ORDER: outcome is clearing containers without a customs hold or a compliance fine, not a generic field-accuracy score; reversibility puts the hazardous-material flag and the tariff code above everything else, since a wrong one there can trigger a seizure or an audit that a corrected contact name never could.

The dependency step works differently here than at Oakferry. A wrong hazardous flag has to be caught before the container is released, not after, so that field's review can't happen on the adjuster's own schedule the way an insurance claim's can. The evidence step reflects that: Chinwe ran ManifestClear silently alongside the existing manual customs check for three weeks, comparing its top-tier flags against what human officers actually caught, before trusting it to flag anything live.

Hand sketched labeled parts diagram titled Inside a claim intake form, reused here for a shipping manifest. Center document icon labeled Claim Form, with policy number, injury code, date of loss, and witness contact around it.
Same four-callout shape. Swap in a tariff code and a hazardous flag, and the ranking argument is identical.
Containers held for manual recheck per week, before and after tiered field flagging
50 25 0 Week 1 Week 10 46 12
Fewer containers held overall, because the officers' attention finally went to the tariff and hazmat fields first instead of being spread evenly across all twenty.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "rank the fields by cost-if-wrong and reversibility, loud flags on the few that matter, quiet or none on the rest," and stop.
Cost: there's no engineering time this quarter to build three separate flag styles. Say so honestly, and ship the top-tier loud flag alone first, since protecting the costliest fields matters more than a fully graded system on day one.
The model gets better, for real: if ClaimScan's injury-code extraction genuinely improves to 97 percent, that's still not a reason to drop the tier. It's a reason to watch that specific field even more closely, since a rare miss on an improved field is easier to overlook than ever.

Where people run it wrong.
They show one averaged confidence number and call the uncertainty problem solved.
They flag all 20 fields equally, which trains reviewers to stop looking at any of them.
They rank fields by how often they're wrong instead of by what a wrong one actually costs.

How to use it live. When someone asks how to show uncertainty across many fields, ask back: if only one of these fields could get a loud flag, which one would you protect? Build outward from that answer, not from the form's own layout.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "what would you build or show first" question across many fields like this one?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. It ranks fields by what's hardest to undo, not by how they're laid out on the form.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marcus Villanueva, Claims Operations Lead at Oakferry Mutual, nine years running the claims floor before ClaimScan.
3 · THE OUTCOME
What's the real outcome this tiering is protecting?
Tap to flip
ANSWER
Whether the claimant gets paid the right amount and whether the claim is eligible at all, not a generic field-accuracy percentage.
4 · THE RANK
State the three-tier design, top to bottom.
Tap to flip
ANSWER
Loud flag on 5 payout-driving fields, quiet collapsed dot on 8 middle fields, no flag on 7 low-risk fields unless confidence is extremely low.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Averaging all 20 fields into one confidence score at launch, to keep the screen simple, made sense back when ClaimScan only handled six easy fields.
6 · THE NUMBER
Fill in the blank: the claim's overall badge read 94 percent healthy the same quarter the injury-code field specifically was running at ___ percent.
Tap to flip
ANSWER
81 percent. Buried inside an average that never dropped below 90, so nothing on screen pointed an adjuster there.
7 · THE REPLAY
Same misread injury code, tiered screen instead of an average. What changes?
Tap to flip
ANSWER
The injury-code field carries its own loud flag whenever its confidence dips below 92 percent, regardless of how clean the other 19 fields extracted, so the adjuster is pointed straight at it instead of trusting one green average.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what dependency works differently there?
Tap to flip
ANSWER
Brindlemark Terminal's ManifestClear customs tool. There, the hazardous-flag and tariff-code fields must be checked before a container is released, not on the reviewer's own schedule, so the evidence step used a silent shadow run before trusting live flags.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: ClaimScan's original design showed one ___ confidence score across all 20 fields, instead of ranking them.
Show hint
Look at "the decision I would take back."
Show answer
Averaged. A single averaged number treated a hard-to-read injury code and an easy-to-read policy number as equally trustworthy.
Multiple choice
2. Why does the injury-code field outrank the mailing-address field in this answer's tiering?
  • A. The injury code appears first on the physical claim form.
  • B. The injury code is extracted less often overall.
  • C. A wrong injury code changes the payout and is hard to undo, while a wrong address gets fixed on the next phone call.
  • D. Adjusters find injury codes more interesting to review.
Show hint
Look at the Reversibility step.
Show answer
C. Rank by what's hardest to undo. A mailing address is trivially reversible; a paid claim on a wrong injury code is not.
True or false
3. True or false: the tiered redesign puts a loud confidence flag on all 20 fields, just with different colors per tier.
  • True
  • False
Show hint
Look at the priority list's fourth bullet.
Show answer
False. The bottom tier of seven fields gets no flag at all unless confidence is extremely low, on purpose.
Short answer, apply it yourself
4. Think of a form or dashboard you use with many fields or metrics at once. If you could only get a loud alert on one of them, which would you pick, and why?
Show hint
Ask which one, if wrong, would be hardest to undo later.
Show answer
Model answer: Usually the field tied most directly to money moving or a decision being finalized, the same logic that put the injury code first here.
Short answer, why no middle setting
5. Why wasn't "just lower the overall confidence threshold that triggers a flag" enough to fix ClaimScan's problem?
Show hint
Look at how the average hid the injury-code field's real number.
Show answer
Model answer: A lower threshold still averages 20 fields into one number. It would flag more claims overall without ever pointing an adjuster at the specific field that actually mattered.
Short answer, where it wouldn't matter
6. Name a field on the claim form where this tiered flagging genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The claimant's mailing address or the form's submission date. Both get corrected on the next phone call regardless of what the screen shows.
Before you close the answer
Why this works
Tests whether you'll default to a single confidence number because it's simple to build, or actually rank twenty different risks by what a wrong one costs and how hard it is to walk back.
Follow-up traps
"Why not just show all 20 confidence scores individually?" Response: twenty numbers on one screen is its own kind of noise; ranking into three tiers keeps the signal without asking an adjuster to read twenty percentages every claim.

"What if a low-risk field's error rate suddenly spikes?" Response: that's exactly why confidence gets tracked by field monthly, not just at launch, a spike would move that field up a tier, not stay invisible forever.
If pressed
The rebuilt screen also lets an adjuster manually promote any field to the loud-flag tier for a specific claim, since an unusual case, like a multi-vehicle accident, can make an otherwise low-risk field suddenly worth a second look.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more