CaseAdvancedDesigning for Uncertainty & Trust / Human-in-the-loop product design / #2

Design the review interface for a human checking 200 AI-generated outputs an hour.

SPARK the interface I'll design is TriageAI's claim review queue, an eighteen-second-a-claim screen

Harborcrest Mutual Insurance runs TriageAI, a tool that drafts a settlement recommendation for property damage claims after a photo, an estimate, and adjuster notes come in. Noor Rasheed is a claims adjuster on Harborcrest's catastrophe response team, and after a regional hailstorm, she has to clear about 200 flagged claims an hour to keep the queue from burying the team.

The direct answer
Sort the queue by risk, not by arrival time: a single claim card per screen, ordered so anything low-confidence or high-dollar always surfaces first, with the AI's reasoning shown in three lines, before/after photos side by side, and three keyboard-shortcut buttons: approve, adjust, escalate. A fixed dollar floor forces any large claim into view regardless of how confident the model sounds about it.
Do this, in order
  1. Sort the queue by risk, not by when the claim arrived.Why: at 200 an hour, the order she sees claims in decides which ones actually get her attention.
  2. Set a hard dollar floor that always surfaces a claim, no matter how confident the model sounds.Why: a confidence score can be wrong in exactly the cases where being wrong costs the most.
  3. Show the AI's reasoning in three lines, not a percentage.Why: a number tells her nothing to check. Three lines of reasoning gives her something to disagree with.
  4. Give her three keyboard-shortcut responses: approve, adjust, escalate.Why: at eighteen seconds a claim, a slow click is the same as no review at all.
  5. Hide a real audit claim inside the auto-clear tier about once every twenty claims.Why: this is the only way to know the tier she never sees is actually being handled correctly, not just quietly ignored.
  6. Leave the underlying settlement math untouched for small, routine claims well inside range.Why: this whole design is about the edges. Most storm-week claims never need any of this.

How to answer this, stage by stage

Nobody is grading whether you can sketch a pretty screen. They're grading whether your screen survives the one hour a week it actually gets tested.

Stage 1
Name the volume, out loud, first
Say it like this
"200 an hour is eighteen seconds a claim. Any design that assumes she can read carefully has already failed before I draw a single box."
Why this works
Anchors every later decision in the actual constraint, instead of a generic "clean dashboard" answer.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how the queue works today. Payoff, the habit I want her to build. Anchor, the actual screen. Risk, what breaks it. Keep out, what I won't build yet."
Why this works
Signals a method for a design question, instead of describing UI elements at random.
Stage 3
Say what's wrong with today's screen
Say it like this
"Right now every claim looks the same in the queue, whether it's a 400 dollar fence or a 40,000 dollar roof. She's spending the same six seconds on both."
Why this works
Shows you found the actual design flaw, not a hypothetical one.
Stage 4
Give the anchor: the one decision the screen hangs on
Say it like this
"Sort by risk, not arrival time. Show the reasoning in three lines, photos side by side, three shortcut buttons. And a dollar floor that always surfaces a large claim no matter what the confidence score says."
Why this works
This is the direct answer, and it's specific enough to argue with.
Stage 5
Prove the anchor survives its own risk
Say it like this
"Say the model's confidence is quietly wrong on water damage claims specifically. Sorting by confidence alone would bury exactly those. The dollar floor catches them anyway, because it doesn't trust the confidence score at all past a certain size."
Why this works
Shows the anchor was designed against its own failure, not just its best case.
Stage 6
Say what you're not building yet
Say it like this
"No auto-approval above the dollar floor, no AI-written denial letters, no leaderboard ranking adjusters by speed. That last one especially would just reward clicking fast."
Why this works
Shows scope judgment instead of a wish list of everything the tool could someday do.
Stage 7
Say how you'd know it's actually working
Say it like this
"I'd drop a real claim into her auto-clear tier disguised as normal, about once every twenty. If it stops getting caught, the tier isn't really being watched anymore."
Why this works
Shows the design accounts for its own decay under real fatigue, not just day one.
Stage 8
Close on the one line
Say it like this
"At this speed, the order she sees things in is the entire design. Everything else is decoration around that one decision."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Say we build a screen that has eighteen seconds to earn a person's real attention, two thousand times over.

Before TriageAI, a Harborcrest adjuster reviewed a property damage claim from scratch: measuring damage from photos, pricing repairs against a materials guide, writing up a settlement recommendation by hand. A single claim took about twenty-two minutes, and a normal week's volume of around 1,800 claims spread across forty adjusters, comfortably.

Hand sketched flow diagram titled A claim before TriageAI. Four boxes: Photo arrives, Adjuster measures damage, Estimates cost by hand, Mails settlement.
Slow, but nothing ever waited in a queue long enough to need sorting.

Now TriageAI drafts a full settlement recommendation, photos, estimate, and reasoning, in under two seconds, and most weeks an adjuster reviews and confirms a batch in a normal, comfortable rhythm.

Here's the turn: the problem isn't a bad week where the model gets more wrong. The problem is a good tool meeting a genuinely enormous queue, where the interface itself, the plain order things appear in, quietly decides which claims a person actually gets to think about.

Hand sketched timeline titled 48 hours after the hailstorm. Four milestones: Storm hits hour 0, Claims arrive hour 6, Queue hits 3,000 highlighted hour 20, CAT team activated hour 24.
The queue doesn't wait for the interface to be ready. It just arrives.
Claims later reopened for a missed detail, flat queue versus risk-sorted queue, per 1,000 storm claims
50 25 0 46 Flat queue 9 Risk-sorted queue
Same adjusters, same tool, same storm. The only thing that changed was the order claims appeared in.

At its worst, a genuinely large claim sits near the bottom of a two-thousand-item flat queue behind a wall of routine ones, and by the time anyone reaches it, the wrong number has already gone out in a settlement letter.

Hand sketched quadrant titled Where claims actually sit. Axes AI confidence from unsure to sure, and dollar amount from small to large. Roof collapse sits top left, large and unsure. Water intrusion sits middle left. Storm debris and fence damage sit bottom right, small and sure.
Most claims are small and easy. The ones that matter sit in the top-left corner, and a flat queue has no way to find them.
The decision I would take back We launched TriageAI's queue sorted by arrival time, the same as the paper-folder system it replaced. That made sense at a normal week's volume, where nothing sat long enough for order to matter. It stopped making sense the first storm week, when order became the entire game.

What I would leave alone: small, routine claims, a fence panel, a gutter, well inside TriageAI's tested range, don't need any of this. The whole design is aimed at the rare, large, uncertain claim sitting somewhere in a queue of two thousand.

The lesson: at real volume, the interface isn't a window onto the work. It is the work. The order is the decision.

Now here is the same thing as a story

The short version above is what you'd say defending this redesign to Harborcrest's claims operations lead. Read this one for how the old queue nearly let one slip through.

Noor Rasheed has adjusted property claims for six years, and she is the one people ask when a photo doesn't match the estimate, the kind of detail that saves the company real money on a claim nobody else double-checked.

Knowledge spark: why would sorting by confidence alone still fail? A confidence score reflects how the model rates itself, not how right it actually is. If the model is systematically overconfident on one type of claim, say water damage with ambiguous photos, sorting by that same confidence score will bury exactly the claims most worth a second look.

After the hailstorm hit, TriageAI's queue filled with claims in arrival order, same as every ordinary week, just at fifteen times the volume. Noor and her CAT team started clearing the flat list top to bottom, fast, the way the screen had always worked.

Hand sketched comparison titled The day confidence lies. Left, a red question mark box icon labeled Sorted by confidence alone, caption buried at the bottom. Right, a green scale icon labeled Dollar floor added, caption always surfaces.
The same claim, in two different queue designs, ends up in two very different places.

A colleague working the same queue leaned over around hour three and said, half-joking, "you've cleared ninety of these in the last twenty minutes, are you even reading the estimates?" Noor laughed, then actually checked her own pace. She had been.

Hand sketched labeled parts diagram titled What's on the claim card. Center document icon labeled Claim Card, with four callouts: risk tier, before slash after photos, reason three lines, approve adjust escalate.
Every claim showed the same four things. None of them told her which claim actually needed her attention first.

Deep in the flat queue, past hundreds of routine fence and siding claims, sat a roof-collapse claim TriageAI had flagged with unusually low confidence, forty thousand dollars, a genuinely uncertain read on structural damage from a partial photo set. Nothing about its position in the list made it look any different from the fence claim next to it.

The queue never lied about any single claim. It just never told her which one, out of two thousand, was actually worth slowing down for.

Noor reached it with about forty minutes left in her shift, tired, and nearly approved the draft settlement as written. A habit from years of careful work made her open the photo set full-screen one more time, and she caught a support beam crack the draft estimate hadn't priced in at all.

With the redesigned queue, that same claim, low confidence and high dollar amount, would have surfaced first, not four hundred items in. Run the storm week forward: Noor sees it inside her first hour, fresh, not on the tail end of a long shift, and the crack gets caught before any settlement number is ever written down.

The old queue asked her to be equally careful about everything. The new one asks the queue itself to know which claims need that from her.

I built the queue in arrival order because it matched the paper system it replaced, and paper had never needed sorting. It took a colleague's offhand comment, and a roof claim I nearly missed at hour six, to see that speed alone was never the actual danger. Order was.

SPARK, in one screenNot a UI style guide. SPARK is what tells you the queue's order is the whole design.

S
Situation. How the queue works today.
Claims appear in arrival order, same as the paper folders TriageAI replaced, with no signal for size or uncertainty.
Grounds the redesign in a real, currently-flawed screen, not a hypothetical one.
P
Payoff. The habit this should build.
Spend real attention on the rare, uncertain, expensive claim, and move fast, confidently, through the routine ones without either feeling rushed.
Names the actual behavior change, not just a faster click rate.
A
Anchor. The one decision everything hangs on.
A risk-sorted queue, one claim card at a time, reasoning in three lines, photos side by side, three keyboard-shortcut buttons, and a dollar floor that overrides confidence.
This is the hardest step and the direct answer: a concrete screen, not a design philosophy.
R
Risk. What breaks the first time it's wrong.
The model's confidence is quietly miscalibrated on one claim type. The dollar floor survives this by never trusting confidence past a certain size, regardless of what it says.
Proves the anchor was designed against its own failure mode, not just its best day.
K
Keep out. What we won't build, day one.
No auto-approval above the dollar floor, no AI-drafted denial letters, no adjuster speed leaderboard.
Shows judgment about scope, not a wish list dressed up as a roadmap.
Hand sketched icon list titled Not built day one. Three items: a box icon labeled No auto-approve above floor, a document icon labeled No AI-written denial letters, a gauge icon labeled No adjuster speed leaderboard.
Each of these would be a real feature someday. None of them fixes what today's flat queue actually gets wrong.

The recap, one line per letter: situation is a flat, arrival-order queue built for calm weeks, payoff is real attention on the rare hard claim instead of even attention on all of them, anchor is a risk-sorted single-claim card with a dollar floor, risk is a miscalibrated confidence score, and keep out is holding back on auto-approval, AI-written denials, and speed leaderboards.

And if you want to be sure it really works, try it somewhere elseSame five letters, an HVAC dispatch queue instead of an insurance claim. A different volume spike, the same order problem.

Cross-County Comfort runs a dispatch tool that reads a customer's service request and drafts a repair plan for a technician, ordinary filter swaps up through full system failures. Ezra Lindqvist dispatches technicians from a laptop with a cracked hinge in the back office, and during the first heat wave of the year, requests come in faster than any one dispatcher can read closely.

Mapped onto SPARK: situation is a flat dispatch queue in call order, payoff is real attention on the calls that are actually dangerous or unusual, anchor is a queue sorted so a reported gas smell or an active warranty claim always surfaces above routine filter swaps, regardless of how confident the draft plan sounds, risk is a technician getting sent alone to a job the draft plan underestimated, and keep out is holding off on fully automated scheduling for anything outside the routine tier.

Hand sketched decision tree titled Which dispatch call needs a person first. Root: AI drafts a repair plan. Four branches: routine filter swap leads to Auto-schedule, gas smell reported leads to Human calls first, warranty still active leads to Human calls first, repeat visit same unit leads to Human calls first.
Swap "claim" for "service call," and the same three warning conditions do the same job in a different queue.
AI confidence score versus reported job risk, heat-wave dispatch batch
AI confidence, low to high, left to right Risk if wrong Routine filter swaps Gas smell, warranty calls
High confidence didn't mean low risk here either. The rule that catches this has to check risk directly, not just trust the score.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "sort the queue by risk, not arrival time, and force a floor rule that ignores confidence past a certain size," and stop.
Cost: there's no engineering time to build risk scoring this quarter. Say so honestly, and start with the dollar floor alone, a single rule, before building the fuller sort.
The model gets better, for real: even if TriageAI's accuracy improves across the board, the dollar floor still earns its place, since the whole point was never trusting the score completely on the cases where being wrong costs the most.

Where people run it wrong.
They design the screen for a calm demo day, then ship it straight into a queue fifteen times that size with no changes.
They sort purely by the model's own confidence, which fails exactly when the model is confidently wrong about something specific.
They measure success by how fast the queue empties, which rewards exactly the rushed clicking this design is trying to prevent.

How to use it live. When someone hands you a review-interface question, ask the volume question before you draw anything: how many seconds does this person actually have per item? Design the sort order for that number, not for a demo.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "design this interface" question?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Design against the failure before you build.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Noor Rasheed, a claims adjuster on Harborcrest Mutual Insurance's catastrophe response team, six years in.
3 · SITUATION
What's wrong with the queue as it works today?
Tap to flip
ANSWER
Claims appear in arrival order, so a routine 400 dollar fence claim and a genuinely uncertain 40,000 dollar roof claim get the exact same visual treatment.
4 · THE ANCHOR
What's the one design decision everything else hangs on?
Tap to flip
ANSWER
A risk-sorted single-claim queue with a dollar floor that always surfaces a large claim, regardless of how confident the model sounds about it.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Launching the queue sorted by arrival time, copied straight from the paper system it replaced, which never needed sorting since nothing sat long enough to matter.
6 · THE NUMBER
Fill in the blank: at 200 claims an hour, Noor has about ___ seconds per claim.
Tap to flip
ANSWER
Eighteen seconds. The number that makes the whole design problem real instead of abstract.
7 · THE REPLAY
Same storm week, redesigned queue. What changes?
Tap to flip
ANSWER
The uncertain roof-collapse claim surfaces in Noor's first hour instead of four hundred items in, fresh instead of on the tail end of a long shift.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the anchor there?
Tap to flip
ANSWER
Cross-County Comfort's HVAC dispatch queue. The anchor is sorting so gas-smell and active-warranty calls always surface above routine filter swaps.

Check yourself Score: 0 / 0

Multiple choice
1. Why does a dollar floor need to override the AI's confidence score, rather than just trusting a low-confidence sort?
  • A. Confidence scores are always inaccurate by design.
  • B. The model can be confidently wrong on a specific claim type, which would bury exactly the claims worth a second look.
  • C. Dollar amounts are easier for the interface to display.
  • D. Regulators require a dollar-based sort.
Show hint
Look at the Risk step.
Show answer
B. A confidence score reflects the model's own self-rating, which can be systematically wrong on a whole category of claims without ever showing up as low confidence.
True or false
2. True or false: the fastest way to fix the flat queue is to just tell adjusters to slow down and read more carefully.
  • True
  • False
Show hint
Look at "the choice I would take back" and the eighteen-second constraint.
Show answer
False. At 200 claims an hour, telling someone to slow down isn't a real option. The fix has to be in the order the queue presents claims, not in asking for more human effort.
Fill in the blank
3. Fill in the blank: reopened claims dropped from 46 per 1,000 to ___ per 1,000 after the queue was risk-sorted.
Show hint
Look at the bar chart in Section 1.
Show answer
9. Roughly a fifth of the original rate, on the same tool and the same adjusters, from changing only the order claims appeared in.
Short answer, apply it yourself
4. Think of an app where you review a list of AI suggestions quickly, an inbox, a feed, a code review tool. What would a risk-sorted version of that list actually surface first?
Show hint
Think about what "high stakes and low confidence" would look like in that specific list.
Show answer
Model answer: Most review lists could be resorted around the item most likely to be both wrong and costly if missed, rather than the item that arrived most recently.
Short answer, where it wouldn't matter
5. Name a kind of claim in Harborcrest's system where this redesign changes nothing.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Small, routine claims well inside TriageAI's tested range, a fence panel or a gutter. They stay near the bottom of the sort and rarely need a close look either way.
Short answer, the number question
6. If the storm surge had brought only 400 claims instead of 3,000, would the risk-sorted queue still matter as much? Why or why not?
Show hint
Look at what made order matter in the first place.
Show answer
Model answer: Less so. At a smaller volume, an adjuster can still reach every claim inside a reasonable shift, so order matters less. The design earns its value specifically at volumes big enough that not everything gets equal attention.
Before you close the answer
Why this works
Tests whether you design for the actual constraint, seconds per item, instead of a generic clean-dashboard answer.
Follow-up traps
"Couldn't adjusters just game the dollar floor by underreporting claim size?" Response: the floor is set from the AI's own estimate, not the adjuster's input, so it isn't something an adjuster controls or could quietly lower.

"Isn't a hidden audit claim a little deceptive toward your own team?" Response: it's disguised only in placement, not in stakes, since it's a real claim needing a real answer either way, and it's the only honest way to know the auto-clear tier is actually still being checked.
If pressed
Harborcrest's actual build also logs how long a card stayed open before a decision, since a sub-two-second open-to-approve time on a flagged claim is a second, independent signal the sort order alone can't catch.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more