Describe how uncertainty display differs for expert versus novice users.
Alder Valley Department of Workforce Services runs the state's unemployment system. Marisol Quintero just lost her retail job and is filing a claim on her phone. Walt Huang has reviewed eligibility cases for the department for fourteen years. TrueLine is the eligibility score both of them now look at, from two different screens.
- Show a novice a plain band with one next step, and show an expert the full evidence behind that same score, by default.Why: matching the detail to the reader is the one decision the rest of the design hangs on.
- Add one plain-language reason to a borderline band, not the raw number.Why: "your last quarter's hours are close to the cutoff" changes what a person does next; a number alone does not.
- Give caseworkers the evidence trail without an extra click.Why: making an expert re-derive the reasoning by hand wastes the exact time the tool was built to save.
- Keep both screens tied to one underlying score, never two.Why: the moment an applicant and a caseworker compare notes and the numbers do not match, trust in the whole tool breaks.
- Track the override rate on cases the model called confident.Why: a rising override rate is the real warning that experts need more shown by default, before it turns into burnout.
- Leave the scoring model itself untouched.Why: the actual gap was about who sees what shape of the number, not about how accurate that number is.
How to answer this, stage by stage
Nobody is grading whether you can name a color scheme. They are grading whether you can say, plainly, which reader gets which shape of the same number, and why.
Let's learn
TrueLine is a score a state unemployment system runs on every claim, then shows to two different people: the applicant filing it, and the caseworker who reviews the ones the model flags.
Before TrueLine, an applicant called a hotline, waited on hold, and got a verbal guess from whoever picked up. A written decision arrived by mail two to three weeks later. Caseworkers built every eligibility case from scratch, digging through wage ledgers by hand, about forty minutes for a single flagged claim.
Now an applicant gets an eligibility read in under two minutes on her phone. A flagged case lands in a caseworker's queue the same day, with a wage-ledger pull already attached.
Here's the turn: TrueLine's team, worried about overwhelming a first-time applicant, decided every reader would see the exact same screen. A plain three-band result, "Likely eligible," "Needs a closer look," or "Likely not eligible," with a "Learn more" link that only opened a paragraph, no real numbers behind it, for anyone. Applicant and caseworker got the identical view.
At its worst, that single-screen choice cuts two ways at once. An applicant reads "Likely eligible" as settled fact and makes a real decision on it. A caseworker, staring at the same plain band on a flagged case, has no way to see why it flagged, and rebuilds the whole case by hand anyway, the exact work the tool was supposed to remove.
What I would leave alone: the scoring model underneath doesn't need a rebuild. Nothing here is about accuracy. It is about who sees what shape of an already-good number.
The lesson: detail is not a virtue by itself. Detail only earns its place if the person reading it can actually act on it.
Now here is the same thing as a story
The short version above is what you would say defending this split to Alder Valley's director. Read this one for how close the shift already came.
TrueLine lives on two screens: a phone Marisol Quintero holds in one hand at her kitchen table, and a wall-mounted terminal at the caseworker station where Walt Huang has worked for fourteen years.
Marisol lost her retail job on a Friday and filed her claim that Saturday morning. TrueLine returned "Likely eligible" in about ninety seconds. She had picked up a part-time gig folding shirts at a friend's shop to cover the gap between paychecks, and reading "Likely eligible" as a settled thing, she gave her friend two weeks' notice that same afternoon, worried that keeping both jobs might somehow mess up her claim.
Walt can tell a genuine wage-quarter problem from a plain paperwork mixup before he finishes reading the first line of a case. For months, the plain band worked fine on the clear cases, instant answers, no complaints. His flagged queue moved fast too, at first, on the obviously wrong claims.
The trigger was small. A caseworker two desks over said, not unkindly, "You're clicking into every single 'closer look' case like you don't trust the tool at all." Walt started to say he didn't distrust it, then stopped. He didn't distrust the score. He had no way to see why it landed where it did, so every flagged case meant rebuilding it from the wage ledger up, the same forty minutes it always took.
Marisol's case landed in Walt's queue nine days after she filed. Her last quarter's hours sat right on the eligibility cutoff, actually a payroll-date quirk, not a real problem. Walt caught it, but it took him the usual forty minutes to trace it back through the wage ledger, and by then Marisol had already worked her last shift at the shop.
The extra pay from that folding-shirts gig would have covered her through the two weeks it took her actual claim to clear. She learned the eligibility question, once corrected, had never really been a question at all. By then the shift was already gone.
With the caseworker console now showing the wage ledger, the rule cited, and two similar past rulings by default, Walt's review of a borderline case like Marisol's drops from forty minutes to about six. Her applicant screen changed a little too, for cases like hers: instead of a bare "Needs a closer look," she now sees one plain line, "your last quarter's hours are close to the cutoff, we will check it and follow up within three days," instead of a silence that reads like a wall.
The old design gave both of them the exact same screen, and neither of them enough. The new one gives them different screens, built from the identical number.
I built one screen because it felt fair to give everyone the same experience. Fair is not the same as useful, and it took watching a real shift disappear over a payroll quirk to see the difference.
PICK, in one screenNot a lecture on being thorough with two audiences. PICK is what tells you which shape of the number goes where.
The recap, one line per letter: position is the plain band for the applicant and the full evidence for the caseworker, impact is naming Marisol's real decision and Walt's real forty minutes, cost asymmetry is treating the applicant's hidden decision as the one to design against first, and kill criteria is watching the override rate on confident cases as the real warning sign.
And if you want to be sure it really works, try it somewhere elseSame four letters, a translation tool instead of an eligibility score. A different pair of readers, a different hidden cost.
Lingua Field is a translation tool used two ways: casual travelers checking a menu or a sign, and certified translators like Bram Voss doing paid legal-document work. Mapped onto PICK: position is a plain "should be fine" or "double check this" signal with one suggested rephrasing for the traveler, and the exact ambiguous source phrase plus its alignment score, by default, for the translator. Impact is a traveler who cannot read a raw 71 percent alignment score for a menu item, against a translator who bills by the hour and needs the ambiguous span named, not a vague flag he has to hunt for across a whole document.
The cost asymmetry runs differently here than at Alder Valley. For the traveler, the hidden and expensive error is trusting a wrong translation of something like an allergy warning, rare, real, and invisible to any dashboard since the traveler never reports what she never knew was wrong. For Bram, the hidden cost is minutes multiplied across every paid job: a vague "double check this" sends him re-scanning entire pages that were actually fine, at his hourly rate.
The old decision at Lingua Field was scoring every phrase the same way and showing one confidence label everywhere it appeared, traveler app or professional console, because one label was simpler to ship. That made sense before certified translators were even a paying customer segment. The kill criteria: if Bram's own correction rate on flagged spans stays under one in twenty, the split's calibration is right. If it climbs, the flag itself needs to carry more of the reason for the ambiguity, not just its presence.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "plain band and one action for the novice, full evidence by default for the expert, same score underneath," and stop.
Cost: there is no budget to build two full screens right now. Say so honestly, and start by adding just one plain-language reason to the borderline band, since that alone is cheap and stops the worst version of the novice's mistake.
The model gets better, for real: if the score's accuracy genuinely improves, that is still not a reason to merge the two screens back into one, a better score just means the caseworker's evidence trail gets shorter to read, not that it disappears.
Where people run it wrong.
They build one "honest" screen and call it fair, when fair for both readers actually means different, not identical.
They give the expert detail behind a click instead of by default, and then wonder why review time never drops.
They wait for a complaint to notice the novice-side cost, when the real damage never generates a complaint at all.
How to use it live. When someone asks how uncertainty should look for two kinds of users, ask yourself first: which one of them will make a real decision off this screen alone, with no one to check with? Design for that person's error first.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the applicant wants the raw number anyway?" Response: let her ask for it behind a clearly labeled "show me the number" link, but never make a bare number with no reasoning the default, since that is worse than either the band or the full context on their own.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on UX for uncertainty and confidence display
- #1 When should you show a confidence score to a user, and when should you hide it?
- #2 Describe three ways to communicate uncertainty without displaying a number.
- #3 What is the risk of showing a percentage confidence that users cannot interpret?
- #4 Design the UI for a feature that is 70 percent confident in its answer.
- #5 Explain how hedging language in generated text affects user trust.
- #6 How would you design an interface that encourages verification without being annoying?