Describe the labour implications of designing a review queue.
Baycrest Regional MRF runs an AI vision system over its recycling sort line, flagging photos of likely contamination for a person to confirm before a bale gets rejected. Teodoro Reyes has reviewed those flagged photos for three years, working through a queue built on a number nobody ever checked again after launch day.
- Weight the review quota by predicted batch difficulty, not a flat average.Why: a flat quota is only fair while the mix behind it stays even, and the mix never stays even for long.
- Give reviewers a logged pause button for a genuinely hard batch.Why: without a record of difficulty, a careful slow hour looks identical to a lazy one.
- Route any written warning through the difficulty-adjusted rate, never the raw count.Why: a warning built on the wrong number turns diligence into a disciplinary file.
- Track the gap between assigned difficulty and shift performance, by shift and by person.Why: this is the signal that catches unfair drift months before six warnings pile up.
- Publish the quota math to the reviewers themselves.Why: a rule nobody can see is a rule nobody can push back on.
How to answer this, stage by stage
Nobody is grading whether you can list labour concerns. They're grading whether you can name who actually gets hurt, and the one design change that protects them.
Let's learn
What happens when a review queue quietly turns into someone's whole day, and nobody designed it to stay fair?
Baycrest Regional MRF runs an AI vision system over its recycling sort line. It watches photos of material moving past mid-sort and flags anything that looks like contamination, a garden hose, a car battery, a bag of food scraps, so a whole bale doesn't get rejected downstream. A person confirms or clears each flag before it becomes a real contamination notice.
Before the AI, a crew walked the line pulling contamination by eye, on their feet, all shift. It was slow and it wore people out fast.
Now a reviewer sits at a screen and works through the AI's flagged photos instead, at a rate the AI itself predicted at launch, about 220 items an hour, on average.
Here's the turn: that average was never the real workload. Some flags are an easy call, a car battery sitting in plain sight. Others are genuinely hard, a torn bag with three material types tangled inside it, and take three or four times as long to judge right. A flat quota built on one average only stays fair while the queue keeps being an even mix of both.
At its worst, the model gets better at clearing easy contamination on its own and stops sending it to a person at all. What's left in the human queue is only the hard cases. The quota never moves to match, and the reviewer who happens to get the harder queue that week looks like the slow one, and gets written up for it.
What I would leave alone: the raw throughput number is still useful for one thing, planning how many people to staff on a shift. It should just never be the number a warning gets built on.
The lesson: whoever decides what a review queue measures is also deciding what kind of day a person has. A number that looks fair on day one can turn deeply unfair once the mix behind it shifts, and nothing in a flat quota is built to notice that shift happening.
Now here is the same thing as a story
The short version above is what you'd say to a plant manager. Read this one for how the quota actually turned against Teodoro.
Teodoro Reyes has worked the contamination review screen at Baycrest for three years. He can tell a genuine hazard from a false flag in about two seconds, just from the shape of the photo. When the AI system launched, he liked it. It caught things a tired eye at the end of a shift might miss.
For the first year, the mix stayed roughly even. Teodoro's afternoon shift and the morning shift both saw a similar blend of easy and hard flags. Then, quietly, over about six months, the model's own confidence on common contamination climbed high enough that it started auto-clearing those cases before they ever reached a person.
Teodoro didn't get slower. His afternoon shift just kept drawing the harder leftover batch, the same three or four material types tangled in one torn bag, over and over. He started missing the 220-item mark most days. His supervisor noticed the number first, not the reason behind it.
A corporate audit in month six pulled six months of shift records looking for a pattern in performance write-ups. It found one, just not the one anyone expected: eleven of the site's thirteen warnings for missed quota belonged to the afternoon shift.
Teodoro wasn't slacking. He was reading harder photos, correctly, at the same careful pace he'd always worked at. The quota simply never knew the difference between a hard hour and a lazy one, because nothing in the system had ever been built to tell them apart.
With the redesigned quota, each batch carries the AI's own confidence-derived difficulty score, and the target adjusts to match it. A hard afternoon costs Teodoro fewer items, on record, not a warning. Run the same six months forward with that fix in place, and the warning count on the afternoon shift drops from eleven to two, the same two you'd expect from any shift on any given week.
I built the quota off one number because it was the easiest thing to point to on launch day. It took a corporate audit to see that the number never asked whether the work behind it had changed.
GUARD, in one screenNot a lecture on workplace fairness. GUARD is what tells you who can't push back, and why.
The recap, one line per letter: groups is the reviewer and whoever set the quota, unequal is the shift that keeps drawing the harder leftover batch, ability to contest is a warning with no record to argue against, reduce is a difficulty-weighted quota with a logged pause button, and detect is watching for warnings clustering by shift before they pile up.
And if you want to be sure it really works, try it somewhere elseSame five letters, a textile mill instead of a recycling line. A different fabric, the same quiet unfairness.
Corrymoor Textile Mill runs an AI fabric scanner down its weaving line, flagging likely defects, slubs, holes, a broken weave pattern, for a quality inspector to confirm before a bolt gets pulled. Priyantha Fernando has inspected fabric there for five years, and used to see a fairly even mix of easy, obvious flags and harder, ambiguous ones.
Mapped onto GUARD: groups is Priyantha and whoever set the mill's inspection quota off the scanner's average confidence; unequal is that as the scanner got sharper at catching common slubs on its own, her queue slowly filled with only the rare, hard-to-classify weave faults, and her recognized-defect rate on record fell from a healthy number near her old average down toward a level that read as underperforming, even though she was catching just as much real trouble, only harder trouble; ability to contest is that the mill's system, like Baycrest's, only ever logged bolts inspected per hour, with nothing to show which bolts were the hard ones; reduce is the same fix, a difficulty-weighted target and a logged pause for a genuinely hard bolt; detect is watching the gap between assigned defect complexity and recognized rate, by inspector, every month, not waiting for a coaching conversation to surface it first.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "weight the quota by real difficulty, not a flat average, and give reviewers a logged way to flag a hard batch," and stop.
Cost: there's no budget to build a full difficulty model this quarter. Say so honestly, and start by simply logging the AI's own existing confidence score as a stand-in for difficulty, since that number already exists and costs nothing new to read.
The model gets better, for real: if the AI keeps improving and clears even more of the easy cases on its own, the human queue's average difficulty will keep climbing, which makes a difficulty-weighted quota more necessary over time, not less.
Where people run it wrong.
They set a review quota once, at launch, and never revisit it as the underlying model changes what's actually left in the queue.
They treat a missed quota as a performance problem before ever checking whether the batch itself got harder.
They build the appeal process as a conversation with a manager instead of a number the reviewer can point to.
How to use it live. When someone asks about the labour side of a review queue, ask one question back to yourself first: can this system tell a hard hour from a lazy one? If the honest answer is no, that's the whole design problem, before you get to anything else.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just lower the quota for everyone instead of weighting it?" Response: a flat lower quota wastes capacity on the genuinely easy stretches. Weighting keeps output up when the batch is easy and fair when it isn't.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Human-in-the-loop product design
- #1 When should a human be required to approve an AI action rather than merely able to?
- #2 Design the review interface for a human checking 200 AI-generated outputs an hour.
- #3 Explain how review fatigue undermines a human-in-the-loop design.
- #4 What is the difference between human-in-the-loop and human-on-the-loop?
- #5 How do you decide which cases get routed to a human?
- #6 Describe a confidence-based routing policy and its failure mode.