CaseAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #18

Design the internal triage process for incoming quality feedback.

SPARK the product is LineWatch, an AI copilot at Caldermill Fabrication that flags surface defects on a camera line before a part ships

Caldermill Fabrication machines precision metal parts for industrial equipment makers. LineWatch watches a camera feed on the line and flags anything that looks like a surface defect, a weld crack, a scuff, a dimension that drifted out of spec. Noor Al-Sayed is the plant's only quality engineer, on the floor every day shift.

The direct answer
Sort every incoming flag into one of three tiers the moment it lands, not into one shared inbox reviewed in arrival order. High confidence and low consequence goes to a same-day batch queue anyone on the QA team can clear. High confidence and high consequence pages the on-shift engineer immediately. Anything the model is unsure about, a pattern it hasn't seen before, gets held from shipping and routed straight to engineering, because uncertainty paired with an unfamiliar shape is exactly the case a busy inbox buries.
Do this, in order
  1. Sort every flag into a tier the instant it lands, instead of one shared inbox in arrival order.Why: arrival order rewards whichever flag showed up last, not the one that actually matters most.
  2. Give the low-confidence, unfamiliar-pattern tier its own queue, held from shipping by default.Why: this is exactly the case a busy person skims past, because it doesn't look urgent on the surface.
  3. Page a person immediately only for the high-confidence, high-consequence tier.Why: paging for everything trains people to ignore the page.
  4. Let anyone on the QA team clear the low-stakes tier, not only the one senior engineer.Why: one person triaging every flag alone is the bottleneck the whole redesign exists to remove.
  5. Leave auto root-cause tagging and cross-plant sharing for a later version.Why: day one only needs to sort by severity, not diagnose why each defect happened.

How to answer this, stage by stage

Nobody is grading how many tiers you can invent. They're grading whether your design actually survives the flag that doesn't fit any tier cleanly.

Stage 1
Ground it in one plant, one role
Say it like this
"I'll design this for Caldermill Fabrication, where LineWatch flags surface defects and Noor Al-Sayed is the only quality engineer reviewing what it catches."
Why this works
A triage design for "quality feedback" in general has nothing to push against. One real plant gives it a shape.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how it's done today. Payoff, the habit I want the design to build. Anchor, the one concrete decision. Risk, what breaks it. Keep out, what I won't build yet."
Why this works
Names the method before the design starts, so the interviewer isn't guessing at your structure as you go.
Stage 3
Show today, without the redesign
Say it like this
"Right now, about a hundred and fifty flags land overnight, all in one shared inbox. Noor reviews whichever one is on top. Older flags quietly stack up underneath."
Why this works
SPARK's situation step. You can't design a fix for a problem you haven't shown happening.
Stage 4
Name the habit you want to build
Say it like this
"I want the team to stop treating every flag as equally urgent, and start trusting that the queue itself already sorted the ones that actually need a person right now."
Why this works
This is the payoff step. The habit you're building, not the hours saved, is the actual product.
Stage 5
Give the one decision everything hangs on
Say it like this
"The anchor: every flag gets tiered on two axes the instant it's created, model confidence and consequence if missed, and routed automatically, before any human ever opens an inbox."
Why this works
This matches the direct answer, said as the concrete design decision an interviewer can push on.
Stage 6
Prove it survives being wrong
Say it like this
"Say LineWatch sees a defect shape it's never seen before and isn't confident about. That's not low priority just because confidence is low, it's its own tier, held from shipping automatically until engineering looks."
Why this works
The risk step. An anchor that only works when the model is confident isn't a design, it's a happy path.
Stage 7
Say what you're deliberately not building
Say it like this
"Day one doesn't try to guess the root cause of each defect, machine fault versus material fault. That's a real feature, just not this one, and building it now would slow down shipping the tiering itself."
Why this works
The keep-out step. Naming what you won't build shows judgment, not a shorter wish list.
Stage 8
Close on the one line
Say it like this
"So: tier every flag the instant it lands, by confidence and consequence, not by who happens to open the inbox first. The unfamiliar, uncertain ones get their own queue, not the bottom of everyone else's."
Why this works
Closes on the anchor again, in one breath, ready to survive a follow-up.

Let's learn

The clipboard hanging by Caldermill's inspection station has a crack down one corner, from years of being set down too hard between parts.

Before LineWatch, catching a subtle weld crack meant a physical inspection of every tenth part, about twenty minutes per hour of Noor's shift, and still missing the rare ones between samples. LineWatch checks every single part, in real time, and flags anything that looks off in under a second.

Hand sketched flow diagram titled Today, without the triage design. Four steps: 150 flags land overnight, Noor opens one shared inbox, reviews whichever's on top, older flags pile up unseen, with the last step highlighted.
Four steps, and the last one is the step nobody designed on purpose.

Now about a hundred and fifty flags land in one shared inbox every night, from four production lines running around the clock. Noor reviews them each morning in whatever order they arrived, which means whichever flag came in last night at 11pm gets seen before one that's been sitting since Tuesday.

Flags reviewed within their target time, before and after tiering
100% 50% 0% 44% Before tiering 95% After tiering
Arrival order was never a real triage rule. It just felt like one because someone was always reviewing something.

The turn: the flags Noor missed weren't the loud, obvious ones. They were the quiet, unfamiliar ones sitting under a pile of routine scuffs and scratches that all looked more urgent simply because they were newer.

The decision I would take back We merged flagging and triaging into one job, done by one person. That made sense when LineWatch watched a single line and produced a dozen flags a day, easily held in Noor's head. It stopped making sense the moment the plant added three more lines and the flag count grew past what any one morning review could actually cover.

At its worst: a genuinely new defect pattern, one LineWatch had never flagged before and wasn't confident about, sits unopened for two weeks under forty routine scuff flags, because nothing separated "unfamiliar and uncertain" from "ordinary."

What I would leave alone: the routine, high-confidence, low-consequence flags, a scuff on a non-critical surface, don't need a person at all most days. A quick same-day batch glance is enough, and building anything heavier for those wastes the redesign's whole point.

The lesson: a triage design isn't measured by how it handles the flags that already look obvious. It's measured by what happens to the one flag that doesn't look urgent yet but is.

Now here is the same thing as a story

The short version above is what you'd pitch to Caldermill's plant manager. Read this one for how the gap actually got found.

Noor Al-Sayed has been Caldermill's only quality engineer for four years, the kind of person who can spot a hairline weld crack from the far end of the floor.

For LineWatch's first year on one line, the mornings were easy. A dozen flags, all reviewed by nine, all closed by lunch.

Hand sketched quadrant titled Sorting flags into tiers. Axes, model confidence from unsure to sure, and consequence if missed from minor to severe. Surface scuff sits sure and minor. Weld crack sits sure and severe. New defect shape sits unsure and moderately severe. Dimension drift sits in the middle.
Four kinds of flag, and only one of them was ever getting lost.

Three more lines came online over eight months. The flag count crept from a dozen a day to over a hundred and fifty a night, and Noor's morning review quietly shrank from "everything" to "whatever's near the top when I open the screen."

Knowledge spark: why does an unfamiliar pattern deserve its own tier, not just a low-confidence label? A model that isn't confident is telling you something different depending on what it's unsure about. Low confidence on a defect it's seen a thousand times just means a blurry photo. Low confidence on a shape it's never seen before means the model is guessing, and that's the case worth a person's attention, not less of it.

The plant manager pulled ten open flags at random for a routine audit and found three sitting untouched for over two weeks, all three of them unfamiliar defect shapes LineWatch had flagged with low confidence, all three buried under newer, ordinary scuffs.

Nobody had ever decided that "low confidence" should mean "low priority." It just quietly became true, because nothing in the inbox told anyone otherwise.
Hand sketched comparison diagram titled The day it's wrong. Left panel, a gauge icon labeled Familiar defect, caption high confidence auto-tiered fast. Right panel, a question mark box icon labeled Never-seen pattern, caption held for engineering not shipped.
One of these needs speed. The other one needs to be stopped, not sped up.

Caldermill rebuilt the triage around two axes instead of one inbox: confidence, and consequence if the flag turns out to be right. A flag that's both unfamiliar and low-confidence now gets its own queue, held from shipping automatically, routed to engineering the same day it appears.

Hand sketched labeled parts diagram titled The anchor close up, a triage ticket. Center document icon labeled Triage ticket, with four callouts around it: confidence bucket, consequence tier, routed queue, SLA clock.
Four fields on every ticket, decided the instant it's created, not whenever someone gets around to opening it.
Flags open longer than two weeks, week by week
7 3 0 Wk1: 3 Wk3: 7 (audit) Wk6: 0
The audit didn't cause the backlog. It just found what four extra lines had already built, one quiet week at a time.

Run the same audit forward under the new design: a flag with an unfamiliar shape and low confidence lands in its own held queue the same morning, and engineering sees it before it's ever old enough to look routine.

I merged flagging and triaging into Noor's one job because for a year, on one line, it genuinely was small enough to hold in her head. It took an audit finding three buried flags to see that a design which depends on one person's morning glance isn't a triage process. It's a hope with a job title attached.

SPARK, in one screenNot a feature list. SPARK is what forces a design to survive the day the model is wrong, not just the day it's right.

S
Situation. How it's done today.
One hundred fifty flags a night, one shared inbox, reviewed in arrival order by one quality engineer.
Grounds the design in a real, current bottleneck instead of an imagined one.
P
Payoff. The habit the design should build.
Stop treating every flag as equally urgent. Start trusting the queue already did the sorting.
The habit, not the time saved, is what you actually shipped.
A
Anchor. The one design decision.
Tier every flag on confidence and consequence the instant it's created, routed automatically, before any human opens an inbox.
The hardest step, and the direct answer: this is the decision everything else hangs on.
R
Risk. What breaks it.
An unfamiliar defect shape with low model confidence isn't low priority, it's its own held tier, so it can't be mistaken for routine and buried.
Proves the anchor survives the exact case that broke the old design.
K
Keep out. What waits for later.
Auto root-cause tagging and cross-plant flag sharing. Real ideas, just not day one.
Shows judgment about scope instead of a longer wish list.
Hand sketched timeline titled Rolling out the tiered triage. Four milestones: design tiers week 1, pilot one line week 2 highlighted, all four lines week 4, drop the shared inbox week 6.
Six weeks, and the pilot week is where the design actually gets tested before it touches every line.

The recap, one line per letter: situation is one overloaded shared inbox, payoff is trusting a pre-sorted queue instead of checking everything, anchor is the two-axis tier applied the instant a flag is created, risk is the unfamiliar-and-uncertain case getting its own held queue, and keep out is root-cause tagging waiting for a later version.

And if you want to be sure it really works, try it somewhere elseSame five letters, a permits office instead of a factory floor. Nothing else about the two jobs is alike.

Fenwick Creek's municipal permits office uses an AI tool that flags incomplete or risky building-permit applications for a human reviewer before approval. Renata Vasquez has processed permits there for six years.

Mapped onto SPARK: situation is the same shape, every flagged application landing in one shared review queue in arrival order. Payoff is reviewers trusting that queue already separated the routine missing-signature cases from the ones that actually need a real look. Anchor: tier every flagged application by how confident the tool is and how serious the consequence of a wrong approval would be, same two axes, a different building entirely. Risk: an application involving a genuinely new zoning situation the tool has never seen, flagged with low confidence, gets its own held tier instead of vanishing under routine paperwork flags. Keep out: don't try to auto-approve the routine tier yet, day one only sorts, it doesn't decide.

Hand sketched icon list titled What we left for later. Three items: a funnel icon labeled auto root cause tagging, a box icon labeled cross plant flag sharing, a gauge icon labeled predicting tomorrow's flag volume.
Three real ideas, and none of them belong in the version that ships first.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "tier by confidence and consequence the instant a flag lands, and give the unfamiliar-uncertain case its own held queue," and stop.
Cost: building automatic routing takes an extra sprint the team doesn't have yet. Say so, and ship the tiers as a manual dropdown first, since a slower tiered process still beats a fast untiered one.
The model gets better, for real: if LineWatch's accuracy improves overall, the tiers still matter, because a rarer miss is an even easier one to mistake for routine and lose in the queue.

Where people run it wrong.
They build one shared queue and call reviewing it in order "triage."
They let low confidence quietly mean low priority, without ever deciding that on purpose.
They try to build root-cause diagnosis before the basic tiering even ships.

How to use it live. When someone asks you to design a triage process, ask yourself first: what happens to the flag the model is least sure about. If your design doesn't have a clear answer for that one, it isn't done yet.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Scope flip. Noor's review quietly shrank from covering every flag to only whatever was near the top of the queue as volume grew past what she could hold in her head.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Noor Al-Sayed, Caldermill Fabrication's only quality engineer, four years on the floor before LineWatch scaled to four lines.
3 · THE HABIT
What did Noor stop doing because LineWatch worked?
Tap to flip
ANSWER
She stopped physically inspecting every tenth part by hand, and started reviewing whatever the camera flagged instead.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Reviewing every single flag each morning, versus only reviewing whatever happened to be near the top of the inbox once volume outgrew her.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging flagging and triaging into one job, done by one person. Fine on one line, broken once three more lines came online.
6 · THE NUMBER
Fill in the blank: after tiering launched, ___% of flags were reviewed within their target time, up from 44%.
Tap to flip
ANSWER
95%. Arrival order was never a real rule, it just looked like one because someone was always reviewing something.
7 · THE REPLAY
Same audit, new tiered design. What changes?
Tap to flip
ANSWER
An unfamiliar, low-confidence flag lands in its own held queue the same morning it's created, instead of sitting buried for two weeks under routine scuffs.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what changed?
Tap to flip
ANSWER
Fenwick Creek's building-permit review queue. Same two-axis tiering, applied to zoning applications instead of surface defects.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the anchor decision tiers every flag on two axes, model confidence and ___, the instant the flag is created.
Show hint
Look at the A step.
Show answer
Consequence, or consequence if missed. Confidence alone can't tell you priority. A confident model on a routine scuff and an unconfident model on a serious crack need very different handling.
Multiple choice
2. Why does an unfamiliar, low-confidence defect shape get its own held queue instead of just a lower priority tag?
  • A. Because low-confidence flags are always wrong.
  • B. Because engineering prefers to review everything in one batch at the end of the week.
  • C. Because low confidence on an unfamiliar shape means the model is guessing, which is exactly the case a busy inbox tends to bury as routine.
  • D. Because it's required by a factory safety regulation.
Show hint
Look at the knowledge spark on confidence and unfamiliar shapes.
Show answer
C. Low confidence means something different depending on what's unfamiliar about the input. A guess on a genuinely new pattern deserves more attention, not less.
True or false
3. True or false: building automatic root-cause tagging for each defect belongs in the very first version of this triage design.
  • True
  • False
Show hint
Look at the K step, keep out.
Show answer
False. Day one only needs to sort flags by severity. Diagnosing why each defect happened is a real feature, just not this one, and building it first would slow down shipping the tiering itself.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Merging flagging and triaging into one person's job. It made sense on a single line producing a dozen flags a day, small enough to hold in one person's head, and broke once three more lines multiplied the volume.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think about a notification or inbox you now check less carefully than you used to.
Show answer
Model answer: Many people stop reading every notification in full once a tool proves reliable enough, which is exactly the habit that breaks quietly when volume or accuracy shifts.
Short answer, where it wouldn't matter
6. Name a kind of flag in this story that genuinely doesn't need a person reviewing it most days.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine, high-confidence, low-consequence flag, like a scuff on a non-critical surface. A same-day batch glance is enough; building anything heavier for those defeats the point of tiering.
Before you close the answer
Why this works
Tests whether your triage design has a real answer for the flag that doesn't look urgent yet, not just a plan for the obvious ones.
Follow-up traps
"Isn't holding low-confidence flags from shipping going to slow the line down?" Response: only for the rare unfamiliar case, which is a small fraction of flags; the routine high-confidence tier still clears same day, faster than the old shared inbox did.

"What if the tiers themselves are wrong for some new kind of defect?" Response: that's exactly why "unfamiliar and uncertain" is its own tier instead of a rule buried in the confidence score alone, it catches the case the tiering logic itself hasn't seen yet.
If pressed
Caldermill's actual rollout kept a fourth, informal channel for a full week, floor supervisors could still flag anything by radio regardless of tier, specifically to catch cases the new design hadn't anticipated yet.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more