ConceptAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #18

Describe how uncertainty display differs for expert versus novice users.

PICK the same score, shaped for who is actually looking at it

Alder Valley Department of Workforce Services runs the state's unemployment system. Marisol Quintero just lost her retail job and is filing a claim on her phone. Walt Huang has reviewed eligibility cases for the department for fourteen years. TrueLine is the eligibility score both of them now look at, from two different screens.

The direct answer
Show a novice a plain confidence band tied to one next step, never a bare number. Show an expert the full evidence and reasoning behind that same score, by default, not behind an extra click. Both screens show the same score in a different shape. Never a different score.
Do this, in order
  1. Show a novice a plain band with one next step, and show an expert the full evidence behind that same score, by default.Why: matching the detail to the reader is the one decision the rest of the design hangs on.
  2. Add one plain-language reason to a borderline band, not the raw number.Why: "your last quarter's hours are close to the cutoff" changes what a person does next; a number alone does not.
  3. Give caseworkers the evidence trail without an extra click.Why: making an expert re-derive the reasoning by hand wastes the exact time the tool was built to save.
  4. Keep both screens tied to one underlying score, never two.Why: the moment an applicant and a caseworker compare notes and the numbers do not match, trust in the whole tool breaks.
  5. Track the override rate on cases the model called confident.Why: a rising override rate is the real warning that experts need more shown by default, before it turns into burnout.
  6. Leave the scoring model itself untouched.Why: the actual gap was about who sees what shape of the number, not about how accurate that number is.

How to answer this, stage by stage

Nobody is grading whether you can name a color scheme. They are grading whether you can say, plainly, which reader gets which shape of the same number, and why.

Stage 1
Scope it to one product, one pair of readers
Say it like this
"I'll answer this for TrueLine, the eligibility score Alder Valley's Department of Workforce Services shows to unemployment applicants and to the caseworkers who review the flagged cases."
Why this works
Keeps "expert versus novice" from turning into a vague design lecture with no product in it.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my pick before any reasoning. Impact, who feels each kind of error. Cost asymmetry, which error is hidden and expensive. Kill criteria, what would change my mind."
Why this works
Shows a real method, not just a personal opinion about screens.
Stage 3
State the position before the reasoning
Say it like this
"Show the applicant a plain band and one next step. Show the caseworker the full evidence behind that same score, by default. That's the pick. Everything else is why."
Why this works
PICK's P step, said first, is what stops the answer from turning into "it depends."
Stage 4
Name who feels each kind of error
Say it like this
"A novice who gets a bare number reads it like a coin flip about her rent. An expert who gets no evidence has to redo the whole case by hand before he trusts an override. Both are real costs, just paid by different people."
Why this works
Shows you can name the actual person on each side, not just an abstract tradeoff.
Stage 5
Find the cost asymmetry, the hard part
Say it like this
"The expensive error isn't 'too little detail for the novice.' It's the applicant making a real decision, like quitting a shift, off a number she was never given the context to read. That's the one I'd design against first."
Why this works
This is the step most candidates skip, naming which side actually breaks something.
Stage 6
Give the kill criteria
Say it like this
"If the caseworker override rate on 'confident' cases stays flat for a few months, the split is working. If it climbs, that's the signal experts need more shown by default, not just available on request."
Why this works
Turns the design pick into something measurable instead of a permanent opinion.
Stage 7
Close on the one line
Say it like this
"Same score, two shapes. A plain band and one action for the person deciding her own life. Full evidence, no extra click, for the person deciding someone else's case."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

TrueLine is a score a state unemployment system runs on every claim, then shows to two different people: the applicant filing it, and the caseworker who reviews the ones the model flags.

Before TrueLine, an applicant called a hotline, waited on hold, and got a verbal guess from whoever picked up. A written decision arrived by mail two to three weeks later. Caseworkers built every eligibility case from scratch, digging through wage ledgers by hand, about forty minutes for a single flagged claim.

Hand sketched flow diagram titled One score, shaped two ways. Five boxes: Pull wage records, Run eligibility rules, Score confidence, Shape for the viewer highlighted, Show the screen.
The whole design question lives in that fourth box. Everything before it is identical for every reader.

Now an applicant gets an eligibility read in under two minutes on her phone. A flagged case lands in a caseworker's queue the same day, with a wage-ledger pull already attached.

Here's the turn: TrueLine's team, worried about overwhelming a first-time applicant, decided every reader would see the exact same screen. A plain three-band result, "Likely eligible," "Needs a closer look," or "Likely not eligible," with a "Learn more" link that only opened a paragraph, no real numbers behind it, for anyone. Applicant and caseworker got the identical view.

Average caseworker time on a flagged case, before and after separate screens
40 min 20 0 40 min One shared screen 6 min Two built for the reader
Same wage ledger, same rules, same score. Only the default screen changed. Six minutes instead of forty, once the evidence was already there.

At its worst, that single-screen choice cuts two ways at once. An applicant reads "Likely eligible" as settled fact and makes a real decision on it. A caseworker, staring at the same plain band on a flagged case, has no way to see why it flagged, and rebuilds the whole case by hand anyway, the exact work the tool was supposed to remove.

Applicants who quit qualifying part-time work within a week of seeing their score, by month
8% 4 0 3% concern line Month 1 Month 4 Month 5
The redesign shipped partway through month five. The rate crossed the concern line back in month two, nobody was watching a number that small yet.
Hand sketched comparison titled Two screens, one score. Left panel, a person icon labeled Applicant view, caption plain band, one next step. Right panel, a document icon labeled Caseworker view, caption full evidence, no extra click.
Same score feeding both panels. Only the shape changes, never the number underneath.
The decision I would take back TrueLine's team decided at launch to build one detail view behind the "Learn more" link, instead of two: a plain-language one for the applicant, and an evidence-and-reasoning one for the caseworker console. One template was cheaper to build and test. It stopped being cheap the day a caseworker's forty minutes of manual digging came right back, and an applicant's real decision hinged on a number she never had the context to read.

What I would leave alone: the scoring model underneath doesn't need a rebuild. Nothing here is about accuracy. It is about who sees what shape of an already-good number.

The lesson: detail is not a virtue by itself. Detail only earns its place if the person reading it can actually act on it.

Now here is the same thing as a story

The short version above is what you would say defending this split to Alder Valley's director. Read this one for how close the shift already came.

TrueLine lives on two screens: a phone Marisol Quintero holds in one hand at her kitchen table, and a wall-mounted terminal at the caseworker station where Walt Huang has worked for fourteen years.

Marisol lost her retail job on a Friday and filed her claim that Saturday morning. TrueLine returned "Likely eligible" in about ninety seconds. She had picked up a part-time gig folding shirts at a friend's shop to cover the gap between paychecks, and reading "Likely eligible" as a settled thing, she gave her friend two weeks' notice that same afternoon, worried that keeping both jobs might somehow mess up her claim.

Walt can tell a genuine wage-quarter problem from a plain paperwork mixup before he finishes reading the first line of a case. For months, the plain band worked fine on the clear cases, instant answers, no complaints. His flagged queue moved fast too, at first, on the obviously wrong claims.

Knowledge spark: why does a wage record even get flagged? An eligibility model checks whether someone's earnings in the right quarters clear a state cutoff. Payroll dates do not always land cleanly inside a calendar quarter, so a person who worked the exact same hours as always can still show up looking borderline, purely because of when a paycheck happened to post.

The trigger was small. A caseworker two desks over said, not unkindly, "You're clicking into every single 'closer look' case like you don't trust the tool at all." Walt started to say he didn't distrust it, then stopped. He didn't distrust the score. He had no way to see why it landed where it did, so every flagged case meant rebuilding it from the wage ledger up, the same forty minutes it always took.

Hand sketched metaphor scene titled Traffic light or cockpit panel. Left, a green circle icon labeled TRAFFIC LIGHT, caption one glance, one action. Right, an amber gauge icon labeled COCKPIT PANEL, caption every gauge, on purpose.
Alder Valley had built one screen and asked it to be both at once.

Marisol's case landed in Walt's queue nine days after she filed. Her last quarter's hours sat right on the eligibility cutoff, actually a payroll-date quirk, not a real problem. Walt caught it, but it took him the usual forty minutes to trace it back through the wage ledger, and by then Marisol had already worked her last shift at the shop.

We did not take a few minutes from Marisol. We took a job she never needed to give up.

The extra pay from that folding-shirts gig would have covered her through the two weeks it took her actual claim to clear. She learned the eligibility question, once corrected, had never really been a question at all. By then the shift was already gone.

Hand sketched labeled parts diagram titled What's in the caseworker evidence panel. A document icon at the center labeled Case File, with four callouts around it: wage ledger, rule cited, past rulings, override button.
Four things Walt actually needed to see. Before the redesign, none of them showed up without him rebuilding the case himself.

With the caseworker console now showing the wage ledger, the rule cited, and two similar past rulings by default, Walt's review of a borderline case like Marisol's drops from forty minutes to about six. Her applicant screen changed a little too, for cases like hers: instead of a bare "Needs a closer look," she now sees one plain line, "your last quarter's hours are close to the cutoff, we will check it and follow up within three days," instead of a silence that reads like a wall.

The old design gave both of them the exact same screen, and neither of them enough. The new one gives them different screens, built from the identical number.

I built one screen because it felt fair to give everyone the same experience. Fair is not the same as useful, and it took watching a real shift disappear over a payroll quirk to see the difference.

PICK, in one screenNot a lecture on being thorough with two audiences. PICK is what tells you which shape of the number goes where.

P
Position. The pick, before any reasoning.
Plain band and one next step for the applicant. Full evidence, by default, for the caseworker.
This is the hardest step and the answer to the question: commit to the split before defending it.
I
Impact. Who feels each kind of error.
Marisol feels a bare number like a fact she must act on. Walt feels a silent screen as forty minutes of rebuilding a case from scratch.
Names the real person on each side of the tradeoff, in units, not in the abstract.
C
Cost asymmetry. Which error is hidden and expensive.
An applicant's real-life decision made off a number with no context is hidden and expensive. A caseworker's wasted forty minutes is visible and, on its own, cheaper to absorb.
Explains why the applicant's simplicity and the caseworker's detail both matter, but for different reasons.
K
Kill criteria. What would change the pick.
If the override rate on "confident" cases climbs instead of holding flat, experts need more shown by default, not just available on request.
Turns the pick into something measurable instead of a permanent opinion.
Hand sketched quadrant titled Where full detail earns its place. Axes case ambiguity from clear cut to borderline, and detail shown by default from plain band to full evidence. Clean approve cases sit low on detail for applicants, high for caseworker audits. Borderline cases sit low for applicants and high for caseworkers.
The same borderline case earns a different default depending on who is reading it, not on the score itself.

The recap, one line per letter: position is the plain band for the applicant and the full evidence for the caseworker, impact is naming Marisol's real decision and Walt's real forty minutes, cost asymmetry is treating the applicant's hidden decision as the one to design against first, and kill criteria is watching the override rate on confident cases as the real warning sign.

And if you want to be sure it really works, try it somewhere elseSame four letters, a translation tool instead of an eligibility score. A different pair of readers, a different hidden cost.

Lingua Field is a translation tool used two ways: casual travelers checking a menu or a sign, and certified translators like Bram Voss doing paid legal-document work. Mapped onto PICK: position is a plain "should be fine" or "double check this" signal with one suggested rephrasing for the traveler, and the exact ambiguous source phrase plus its alignment score, by default, for the translator. Impact is a traveler who cannot read a raw 71 percent alignment score for a menu item, against a translator who bills by the hour and needs the ambiguous span named, not a vague flag he has to hunt for across a whole document.

The cost asymmetry runs differently here than at Alder Valley. For the traveler, the hidden and expensive error is trusting a wrong translation of something like an allergy warning, rare, real, and invisible to any dashboard since the traveler never reports what she never knew was wrong. For Bram, the hidden cost is minutes multiplied across every paid job: a vague "double check this" sends him re-scanning entire pages that were actually fine, at his hourly rate.

Hand sketched decision tree titled How much detail should this screen show. Root: who is looking, and how risky is the case. Three branches: applicant clear case leads to plain band only, applicant borderline case leads to plain band plus one reason, caseworker any flagged case leads to full evidence by default.
Reused here for Lingua Field, swap "caseworker" for "certified translator" and the same branches hold.

The old decision at Lingua Field was scoring every phrase the same way and showing one confidence label everywhere it appeared, traveler app or professional console, because one label was simpler to ship. That made sense before certified translators were even a paying customer segment. The kill criteria: if Bram's own correction rate on flagged spans stays under one in twenty, the split's calibration is right. If it climbs, the flag itself needs to carry more of the reason for the ambiguity, not just its presence.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "plain band and one action for the novice, full evidence by default for the expert, same score underneath," and stop.
Cost: there is no budget to build two full screens right now. Say so honestly, and start by adding just one plain-language reason to the borderline band, since that alone is cheap and stops the worst version of the novice's mistake.
The model gets better, for real: if the score's accuracy genuinely improves, that is still not a reason to merge the two screens back into one, a better score just means the caseworker's evidence trail gets shorter to read, not that it disappears.

Where people run it wrong.
They build one "honest" screen and call it fair, when fair for both readers actually means different, not identical.
They give the expert detail behind a click instead of by default, and then wonder why review time never drops.
They wait for a complaint to notice the novice-side cost, when the real damage never generates a complaint at all.

How to use it live. When someone asks how uncertainty should look for two kinds of users, ask yourself first: which one of them will make a real decision off this screen alone, with no one to check with? Design for that person's error first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question comparing uncertainty display for two different kinds of users?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Commit to one split, then show which error is hidden and expensive enough to design against.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Quintero, a first-time unemployment applicant, and Walt Huang, a caseworker who has reviewed eligibility cases at Alder Valley for fourteen years.
3 · THE IMPACT
Who feels each kind of error here, and how?
Tap to flip
ANSWER
Marisol reads a bare number and treats it like a settled fact. Walt gets no evidence and has to rebuild every flagged case by hand.
4 · THE COST ASYMMETRY
Which error is the hidden, expensive one to design against first?
Tap to flip
ANSWER
An applicant making a real decision, like quitting a shift, off a number she was never given the context to read. That cost never shows up as a support ticket.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building one detail view behind the "Learn more" link instead of two: one plain for applicants, one evidence-based for caseworkers.
6 · THE NUMBER
Fill in the blank: Walt's average review time on a flagged case dropped from 40 minutes to about ___ minutes after the caseworker view changed.
Tap to flip
ANSWER
About 6 minutes. Full evidence shown by default removed the need to rebuild the case by hand.
7 · THE REPLAY
Same borderline case, redesigned screens. What changes?
Tap to flip
ANSWER
Walt's review takes about 6 minutes instead of 40, and Marisol's screen adds one plain reason instead of silence, so she does not quit a shift on a guess.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the split there?
Tap to flip
ANSWER
Lingua Field's translation tool. Travelers get a plain "should be fine" or "double check this" signal; certified translators like Bram Voss get the exact ambiguous phrase and alignment score, by default.

Check yourself Score: 0 / 0

Multiple choice
1. Why does TrueLine show caseworkers the full evidence by default instead of behind a click?
  • A. State law requires it on every case.
  • B. So a caseworker does not have to rebuild a flagged case by hand every time.
  • C. It makes the applicant screen load faster.
  • D. The evidence view uses a different, more accurate model.
Show hint
Look at Walt's review time, before and after.
Show answer
B. Full evidence by default removes the forty minutes of rebuilding the case from scratch, the exact cost the tool was supposed to remove.
True or false
2. True or false: the applicant screen and the caseworker screen in this answer are built from two separately computed scores.
  • True
  • False
Show hint
Look at the direct answer's last two sentences.
Show answer
False. Both screens show the same underlying score in a different shape. A second, separately computed score would break trust the moment the two readers compared notes.
Fill in the blank
3. Fill in the blank: applicants quitting qualifying part-time work within a week of seeing their score rose to about ___ percent by month four, before the redesign shipped.
Show hint
Look at the line chart.
Show answer
6.2 percent. It had already crossed the 3 percent concern line back in month two, unnoticed.
Short answer, apply it yourself
4. Think of an app you use that shows the exact same screen to a total beginner and to someone who does this for a living. What's one thing an expert wishes it showed by default?
Show hint
Ask what the expert re-derives by hand every single time, that the tool already knows.
Show answer
Model answer: Most people can name a "why" they wish showed up automatically instead of a click away, exactly the gap Walt's console had.
Short answer, why no middle setting
5. Why not just add one medium-detail setting that works for everyone?
Show hint
Look at the Cost asymmetry step.
Show answer
Model answer: A medium setting still gives the applicant more than she needs to act safely and the caseworker less than he needs to trust an override. It moves the mismatch, it does not remove it.
Short answer, where it wouldn't matter
6. Name a case where this expert-versus-novice split genuinely doesn't matter.
Show hint
Look at the quadrant diagram's lower-left corner.
Show answer
Model answer: A completely clear-cut approval. Both the applicant and the caseworker just need "approved," and no amount of extra evidence changes what either of them does next.
Before you close the answer
Why this works
Tests whether you default to one honest-feeling screen for everyone, or actually ask who is looking and what they can do with more detail before deciding what to show.
Follow-up traps
"Isn't showing less detail to the applicant just hiding information from her?" Response: no, the plain band is the same true score in different clothes, not a different number. The caseworker's full evidence is available to any applicant who asks for a review, it is just not the default.

"What if the applicant wants the raw number anyway?" Response: let her ask for it behind a clearly labeled "show me the number" link, but never make a bare number with no reasoning the default, since that is worse than either the band or the full context on their own.
If pressed
Caseworker overrides get logged with which specific evidence field actually changed Walt's mind, not just that a human disagreed, so the model's retraining loop knows exactly what it was still missing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more