ConceptAdvancedDesigning for Uncertainty & Trust / Human-in-the-loop product design / #13

Describe the labour implications of designing a review queue.

GUARD a recycling sort line, and whoever gets stuck with the hard batch

Baycrest Regional MRF runs an AI vision system over its recycling sort line, flagging photos of likely contamination for a person to confirm before a bale gets rejected. Teodoro Reyes has reviewed those flagged photos for three years, working through a queue built on a number nobody ever checked again after launch day.

The direct answer
Build the review quota around each batch's real difficulty, not a flat items-per-hour target. Give reviewers a logged button to flag a genuinely hard batch and pause the clock, and base any warning on that adjusted rate, never the raw count. A queue that can't tell an easy batch from a hard one will always end up punishing whoever got stuck with the hard one.
Do this, in order
  1. Weight the review quota by predicted batch difficulty, not a flat average.Why: a flat quota is only fair while the mix behind it stays even, and the mix never stays even for long.
  2. Give reviewers a logged pause button for a genuinely hard batch.Why: without a record of difficulty, a careful slow hour looks identical to a lazy one.
  3. Route any written warning through the difficulty-adjusted rate, never the raw count.Why: a warning built on the wrong number turns diligence into a disciplinary file.
  4. Track the gap between assigned difficulty and shift performance, by shift and by person.Why: this is the signal that catches unfair drift months before six warnings pile up.
  5. Publish the quota math to the reviewers themselves.Why: a rule nobody can see is a rule nobody can push back on.

How to answer this, stage by stage

Nobody is grading whether you can list labour concerns. They're grading whether you can name who actually gets hurt, and the one design change that protects them.

Stage 1
Scope it to one queue, not an industry essay
Say it like this
"I'll ground this in one place, a recycling facility where an AI flags photos of contamination on the sort line, and a person confirms each flag before a bale gets rejected."
Why this works
Keeps the answer from turning into a labour policy essay with nobody in it.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's affected. Unequal, where the harm lands hardest. Ability to contest, who can push back. Reduce, the actual design change. Detect, how you'd catch it in production."
Why this works
Signals a method for naming who bears the cost, not just a list of worries.
Stage 3
Name the two groups
Say it like this
"There are two groups. The reviewer, whose day is measured by a quota. And whoever set that quota, using the AI's average time per item, without ever seeing what's actually hiding inside that average."
Why this works
A labour question with no named person in it is a policy memo, not an answer.
Stage 4
Find who can't push back
Say it like this
"The reviewer can't contest a warning, because the system only logs how many items got done, never how hard those items actually were. There's nothing on record to point to."
Why this works
This is GUARD's hardest step, and the one most answers skip.
Stage 5
Give the one decision
Say it like this
"I'd weight the quota by each batch's predicted difficulty, and give the reviewer a logged pause button for a genuinely hard batch, so a slow hour on hard material looks different from a slow hour of just going slow."
Why this works
This is the direct answer, stated as a real interface change, not a policy statement.
Stage 6
Prove it with the six-month drift
Say it like this
"As the model got better at clearing easy contamination on its own, the human queue quietly filled with only hard cases. The quota never moved. The afternoon shift, who kept getting the harder batches, picked up eleven written warnings in six months. The morning shift got two."
Why this works
Makes the harm checkable with a real number, not just asserted.
Stage 7
Close on the one line
Say it like this
"If you design a queue and can't tell me whether a slow hour was hard work or slack work, you didn't design a review process. You designed a way to blame the wrong person."
Why this works
Leaves the interviewer with the one sentence the whole answer stands on.

Let's learn

What happens when a review queue quietly turns into someone's whole day, and nobody designed it to stay fair?

Baycrest Regional MRF runs an AI vision system over its recycling sort line. It watches photos of material moving past mid-sort and flags anything that looks like contamination, a garden hose, a car battery, a bag of food scraps, so a whole bale doesn't get rejected downstream. A person confirms or clears each flag before it becomes a real contamination notice.

Before the AI, a crew walked the line pulling contamination by eye, on their feet, all shift. It was slow and it wore people out fast.

Hand sketched comparison titled Two people, one lever. Left panel, a person icon labeled Reviewer, caption measured by the count. Right panel, a gauge icon labeled Quota Owner, caption sets the number.
Two people over one decision. Only one of them can change the number.

Now a reviewer sits at a screen and works through the AI's flagged photos instead, at a rate the AI itself predicted at launch, about 220 items an hour, on average.

Here's the turn: that average was never the real workload. Some flags are an easy call, a car battery sitting in plain sight. Others are genuinely hard, a torn bag with three material types tangled inside it, and take three or four times as long to judge right. A flat quota built on one average only stays fair while the queue keeps being an even mix of both.

Average difficulty score of items reaching the human queue, month over month
100 50 0 50, quota built for this level Month 1, 32 Month 4, crosses 50 Month 6, 67
Nobody had a reason to watch this line. The quota was set once, at launch, and never looked at again.

At its worst, the model gets better at clearing easy contamination on its own and stops sending it to a person at all. What's left in the human queue is only the hard cases. The quota never moves to match, and the reviewer who happens to get the harder queue that week looks like the slow one, and gets written up for it.

Written warnings for missed quota, by shift, before and after a difficulty-weighted fix
12 6 0 2 11 Before fix 1 2 After fix green = morning amber = afternoon
Same reviewers, same skill. The gap was the quota's fault, not theirs.
We didn't design a review queue. We designed a way to blame whoever got the hard batch that week.
The decision I would take back Baycrest set the review quota using the AI's average predicted time per item, one flat number for the whole queue. That made sense at launch, when the queue was a genuine mix of easy and hard cases. It stopped making sense once the model got better and began clearing easy contamination on its own, leaving only the hard cases for a person to see.

What I would leave alone: the raw throughput number is still useful for one thing, planning how many people to staff on a shift. It should just never be the number a warning gets built on.

The lesson: whoever decides what a review queue measures is also deciding what kind of day a person has. A number that looks fair on day one can turn deeply unfair once the mix behind it shifts, and nothing in a flat quota is built to notice that shift happening.

Now here is the same thing as a story

The short version above is what you'd say to a plant manager. Read this one for how the quota actually turned against Teodoro.

Teodoro Reyes has worked the contamination review screen at Baycrest for three years. He can tell a genuine hazard from a false flag in about two seconds, just from the shape of the photo. When the AI system launched, he liked it. It caught things a tired eye at the end of a shift might miss.

Hand sketched labeled parts diagram titled What is inside a quota. Center gauge icon labeled Quota, with four callouts: raw item count, average time assumption, batch difficulty mix, appeal path missing.
Three of these four parts existed on day one. The fourth was never built.
Knowledge spark: why would the queue's difficulty rise on its own? As the model kept training on confirmed contamination photos, it got better at the common, obvious cases, a car battery, a hose, and started clearing those on its own without sending them to a person. What stayed in the human queue slowly became the leftover hard cases the model still wasn't sure about.

For the first year, the mix stayed roughly even. Teodoro's afternoon shift and the morning shift both saw a similar blend of easy and hard flags. Then, quietly, over about six months, the model's own confidence on common contamination climbed high enough that it started auto-clearing those cases before they ever reached a person.

Hand sketched timeline titled Six months to the audit. Five milestones: Model auto clears easy cases month one, Queue difficulty rising month two, Crosses fifty point line month four highlighted, Warnings pile up month five, Corporate audit month six.
The line crossed its own warning point two months before anyone outside the queue found out why.

Teodoro didn't get slower. His afternoon shift just kept drawing the harder leftover batch, the same three or four material types tangled in one torn bag, over and over. He started missing the 220-item mark most days. His supervisor noticed the number first, not the reason behind it.

A corporate audit in month six pulled six months of shift records looking for a pattern in performance write-ups. It found one, just not the one anyone expected: eleven of the site's thirteen warnings for missed quota belonged to the afternoon shift.

Teodoro wasn't slacking. He was reading harder photos, correctly, at the same careful pace he'd always worked at. The quota simply never knew the difference between a hard hour and a lazy one, because nothing in the system had ever been built to tell them apart.

With the redesigned quota, each batch carries the AI's own confidence-derived difficulty score, and the target adjusts to match it. A hard afternoon costs Teodoro fewer items, on record, not a warning. Run the same six months forward with that fix in place, and the warning count on the afternoon shift drops from eleven to two, the same two you'd expect from any shift on any given week.

I built the quota off one number because it was the easiest thing to point to on launch day. It took a corporate audit to see that the number never asked whether the work behind it had changed.

GUARD, in one screenNot a lecture on workplace fairness. GUARD is what tells you who can't push back, and why.

G
Groups. Who is affected.
The reviewer, measured by the quota. The person who set that quota, using the AI's own average, without seeing inside it.
Puts a name on both sides of the decision, not just a policy.
U
Unequal. Where the harm lands hardest.
Whichever shift keeps drawing the harder leftover batch, once the model starts auto-clearing the easy cases on its own.
Names the specific group the design change protects.
A
Ability to contest. Who can push back.
The reviewer can't. The system logs raw item counts only, never batch difficulty, so a written warning has nothing to argue against.
The hardest step, and the direct answer to the question.
R
Reduce. The actual design change.
Weight the quota by predicted batch difficulty, and give the reviewer a logged pause button for a genuinely hard batch.
A product decision, not a training session or a policy memo.
Hand sketched icon list titled What a fair queue needs. Four items: a gauge icon labeled Difficulty score visible, a document icon labeled Pause logged and auditable, a scale icon labeled Quota adjusts to mix, a question mark box icon labeled Appeal path exists.
Four small parts. None of them require a new model, only a new way to read the one already running.
D
Detect. How you'd catch it in production.
Track the gap between a shift's assigned difficulty and its recognized rate, and watch for warnings clustering by shift.
Catches the drift months before it becomes a stack of disciplinary files.

The recap, one line per letter: groups is the reviewer and whoever set the quota, unequal is the shift that keeps drawing the harder leftover batch, ability to contest is a warning with no record to argue against, reduce is a difficulty-weighted quota with a logged pause button, and detect is watching for warnings clustering by shift before they pile up.

And if you want to be sure it really works, try it somewhere elseSame five letters, a textile mill instead of a recycling line. A different fabric, the same quiet unfairness.

Corrymoor Textile Mill runs an AI fabric scanner down its weaving line, flagging likely defects, slubs, holes, a broken weave pattern, for a quality inspector to confirm before a bolt gets pulled. Priyantha Fernando has inspected fabric there for five years, and used to see a fairly even mix of easy, obvious flags and harder, ambiguous ones.

Hand sketched decision tree titled Where the appeal path branches. Root labeled Inspector flagged for missed quota. Two branches: Difficulty weighted, leads to Fair record. Raw count only, leads to No way to contest.
Same fork in the road as Baycrest's. One branch has a place to stand. The other doesn't.

Mapped onto GUARD: groups is Priyantha and whoever set the mill's inspection quota off the scanner's average confidence; unequal is that as the scanner got sharper at catching common slubs on its own, her queue slowly filled with only the rare, hard-to-classify weave faults, and her recognized-defect rate on record fell from a healthy number near her old average down toward a level that read as underperforming, even though she was catching just as much real trouble, only harder trouble; ability to contest is that the mill's system, like Baycrest's, only ever logged bolts inspected per hour, with nothing to show which bolts were the hard ones; reduce is the same fix, a difficulty-weighted target and a logged pause for a genuinely hard bolt; detect is watching the gap between assigned defect complexity and recognized rate, by inspector, every month, not waiting for a coaching conversation to surface it first.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "weight the quota by real difficulty, not a flat average, and give reviewers a logged way to flag a hard batch," and stop.
Cost: there's no budget to build a full difficulty model this quarter. Say so honestly, and start by simply logging the AI's own existing confidence score as a stand-in for difficulty, since that number already exists and costs nothing new to read.
The model gets better, for real: if the AI keeps improving and clears even more of the easy cases on its own, the human queue's average difficulty will keep climbing, which makes a difficulty-weighted quota more necessary over time, not less.

Where people run it wrong.
They set a review quota once, at launch, and never revisit it as the underlying model changes what's actually left in the queue.
They treat a missed quota as a performance problem before ever checking whether the batch itself got harder.
They build the appeal process as a conversation with a manager instead of a number the reviewer can point to.

How to use it live. When someone asks about the labour side of a review queue, ask one question back to yourself first: can this system tell a hard hour from a lazy one? If the honest answer is no, that's the whole design problem, before you get to anything else.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question about the labour side of a review queue's design?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Name who bears the cost of a design decision, and whether they can push back on it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Teodoro Reyes, a contamination review clerk at Baycrest Regional MRF, three years into reading the AI's flagged photos.
3 · THE UNEQUAL HARM
Where does the harm from a flat quota land hardest?
Tap to flip
ANSWER
On whichever shift keeps drawing the harder leftover batch, once the model starts clearing easy contamination on its own.
4 · THE HARD STEP
What can't Teodoro do when he gets a written warning?
Tap to flip
ANSWER
Show that his batch was harder than average. The system only logs raw item counts, never batch difficulty.
5 · THE DECISION
What's the actual design fix?
Tap to flip
ANSWER
Weight the quota by predicted batch difficulty, and give reviewers a logged pause button for a genuinely hard batch.
6 · THE NUMBER
Fill in the blank: the afternoon shift picked up ___ written warnings in six months, versus two for the morning shift.
Tap to flip
ANSWER
Eleven. The same flat quota, applied to two very different real workloads.
7 · THE REPLAY
Same six months, redesigned quota. What changes?
Tap to flip
ANSWER
The quota adjusts down automatically on hard batches, and the warning gap between shifts nearly disappears, down to two on each side.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the same failure?
Tap to flip
ANSWER
Corrymoor Textile Mill's fabric-defect inspector. Same pattern: the scanner auto-clears easy defects, leftover flagged cases get harder, and the quota never moves to match.

Check yourself Score: 0 / 0

Multiple choice
1. Why does a flat items-per-hour quota become unfair over time in Teodoro's queue?
  • A. The AI's confidence scores are wrong from the start.
  • B. As the AI gets better at clearing easy contamination itself, the human queue is left with a harder mix, but the quota never adjusts.
  • C. Reviewers naturally get slower as they gain experience.
  • D. Baycrest lowered the quota on purpose to save money.
Show hint
Look at the Unequal step.
Show answer
B. The quota was set for a mix that no longer exists. The model quietly changed what's left for a person to see.
True or false
2. True or false: the raw throughput number is still useful for something, just not for judging one reviewer's performance.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
True. It is fine for planning how many people to staff on a shift. It should never be the number a warning is built on.
Fill in the blank
3. Fill in the blank: the average difficulty score of items reaching the human queue rose from 32 to ___ over six months.
Show hint
Look at the line chart in "Let's learn."
Show answer
67. More than double, and it crossed the level the quota was built for two months before the audit found the warning pattern.
Short answer, apply it yourself
4. Think of a queue you've worked in, support tickets, grading, an inbox. If the easy items started getting handled some other way over time, what would a fair quota need to track instead of raw count?
Show hint
Think about what actually predicts how long an item will take, not just that it exists.
Show answer
Model answer: Some measure of expected difficulty or time per item, so a quota tracks real workload rather than just a count of items that no longer represent an even mix.
Short answer, where it wouldn't matter
5. Name a place in Baycrest's system where this exact fix would not matter.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The raw throughput number used purely to plan shift staffing. It is a fine number for capacity planning, just the wrong one for judging a single person.
Short answer, the number question
6. If the model had gotten worse instead of better, would the same difficulty-weighted quota fix still make sense? Why or why not?
Show hint
Think about what a worse model would send to the human queue.
Show answer
Model answer: Yes, arguably more so. A worse model would flood the queue with harder, more ambiguous cases directly, so tracking real difficulty rather than raw count would matter even sooner.
Before you close the answer
Why this works
Tests whether you see that a metric design decision is also a job design decision, not just a UX detail bolted onto the AI feature.
Follow-up traps
"Won't a pause button just get abused, to duck the quota?" Response: it's logged and auditable, so a spike in pause use is itself a visible signal, and the rare abuse case is far cheaper than punishing genuinely hard, careful work every week.

"Why not just lower the quota for everyone instead of weighting it?" Response: a flat lower quota wastes capacity on the genuinely easy stretches. Weighting keeps output up when the batch is easy and fair when it isn't.
If pressed
Baycrest's real fix logged the AI's own confidence-derived difficulty score at the moment each item was scored, so the weighting used a number the system already produced, rather than requiring a new model to be trained.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more