ConceptIntermediateDesigning for Uncertainty & Trust / Human-in-the-loop product design / #15

What is the right batch size for human review, and why does it matter?

PICK a service dispatch queue, and the ticket buried at position eighty-four

Northfield Comfort Systems runs an AI model that reads incoming HVAC service tickets and drafts a diagnosis and urgency level for each one. Bettina Kirsch is the dispatcher who reviews the AI's calls before technicians get sent out, working through however many tickets happen to be sitting in her queue.

The direct answer
Review in small, frequent batches, about eight to twelve tickets at a time, every thirty to forty five minutes, not one giant pile at the end of the day and not one ticket at a time either. Batch size isn't a throughput setting. It's how long a real emergency gets to hide before anyone looks at it.
Do this, in order
  1. Review in small batches, every thirty to forty five minutes, not once at the end of the day.Why: a single end-of-day batch means an emergency that arrived at 9am can sit unseen for hours by design, before anyone even gets tired.
  2. Cap batch size below the point where review quality starts to drop.Why: attention fades inside a long batch, and the tickets reviewed last get the least careful look.
  3. Give true emergencies a separate, immediate-alert lane outside the normal batch entirely.Why: the batch size argument is about everything else. A real emergency should never have to wait for its batch's turn.
  4. Don't shrink batches to one ticket at a time either.Why: constant interruption costs Bettina re-orientation time on every single ticket, and that adds up faster than people expect.
  5. Let batch size grow only as the AI's emergency-detection accuracy genuinely improves.Why: the risk that sets the ceiling on batch size is a missed emergency, not review speed, so that's the number that should move it.

How to answer this, stage by stage

Nobody is grading whether you can name a number. They're grading whether you know which kind of mistake that number is protecting against.

Stage 1
Scope it to one queue
Say it like this
"I'll ground this in an HVAC dispatch queue, where an AI drafts a diagnosis and urgency for each incoming ticket, and a dispatcher reviews the batch before sending a technician."
Why this works
Keeps "batch size" from becoming an abstract operations question.
Stage 2
Take a position before reasoning
Say it like this
"My pick is small, frequent batches, eight to twelve tickets, reviewed every thirty to forty five minutes. Not one ticket at a time, and not one big pile at close of day."
Why this works
This is PICK's P step, and it's the direct answer, stated before any justification.
Stage 3
Name who feels each kind of error
Say it like this
"Too small, and Bettina loses time re-orienting after every interruption. Too large, and whatever's sitting near the bottom of the pile gets the least careful look, right when it might be the one that matters most."
Why this works
Puts a real cost, in real units, on both sides of the tradeoff.
Stage 4
Find the cost asymmetry
Say it like this
"A slightly delayed easy ticket is cheap and visible, someone calls to ask where their technician is. A buried emergency is hidden and expensive, nobody complains about a ticket they don't know is late, until it's a no-heat call in freezing weather."
Why this works
This is PICK's hardest step, and the actual reason batch size matters at all.
Stage 5
Say what would change your pick
Say it like this
"If the AI's emergency-detection accuracy got close to perfect, I'd loosen the batch size, since the risk that sets the ceiling would have genuinely dropped. Until then, I'm optimizing against the buried emergency, not review speed."
Why this works
Shows a confident pick that isn't stubborn, PICK's kill criteria.
Stage 6
Close on the one line
Say it like this
"Batch size isn't a throughput setting. It's how long a real emergency gets to hide before a person looks at it."
Why this works
Restates the direct answer, framed as the actual stakes, ready for a follow-up.

Let's learn

Picture ninety service tickets sitting in one pile, all waiting for a single look at five o'clock.

Northfield Comfort Systems' AI reads each incoming HVAC ticket, drafts a likely diagnosis, no heat, a failing compressor, a refrigerant leak, and flags an urgency level. Bettina Kirsch reviews the AI's calls before a technician gets dispatched, and can bump, confirm, or reclassify any ticket in seconds.

Before the AI, every ticket got a phone call and a few minutes of Bettina's own judgment, one at a time, all day. Slower, but nothing sat unseen.

Hand sketched metaphor scene titled One big batch versus small batches. Left, a box icon labeled ONE BIG BATCH, caption buried at the bottom. Right, a gauge icon labeled SMALL BATCHES, caption seen while it is fresh.
Same ninety tickets. Only the shape of the pile changes.

Now Northfield reviews tickets in one batch, at the end of the day. On a slow day, twenty tickets, that's fine. On a busy day, ninety or more, that same habit means whatever came in at 9am waits for a 5pm look, whether or not anyone's tired by then.

Average wait before an emergency ticket gets escalated, by review pattern
300 min 150 0 22 min Small batches, every 45 min 296 min One end-of-day batch
That gap exists before anyone's even tired. It's just how long the pile sits.

Here's the turn: the real cost of a bad batch size was never the extra minutes at the end of a long batch. It's that everything near the bottom of a big pile gets the least careful look, right when it might be the one call that can't wait.

Miss rate on urgent tickets, by position in the review batch
15% 7 0 Tickets 1-10, 2% Tickets 40-50, 6% Tickets 80-90, 11%
Same dispatcher, same skill, the whole shift. Position in the pile did the damage.
The decision I would take back Northfield set up end-of-day batch review when ticket volume ran about fifteen a day, and a single pass at five o'clock was genuinely fine at that size. It stopped being fine once volume climbed past ninety a day in peak season, and nobody revisited the batch size when the number that justified it had already changed.

What I would leave alone: routine maintenance tickets, filter checks, seasonal tune-ups, don't need this at all. A slightly stale review on one of those costs a scheduling inconvenience, nothing more.

The lesson: a batch size decided once, at one volume, doesn't stay right as volume grows. Nobody thinks of "how many tickets pile up before I look" as a safety setting, until it is one.

Now here is the same thing as a story

The short version above is what you'd say defending this to Northfield's operations lead. Read this one for the Friday it nearly went wrong.

Bettina Kirsch has dispatched HVAC technicians for Northfield for five years. She can tell a nuisance call from a real emergency before the ticket's even fully loaded, from the words a customer used and the time of year.

Hand sketched comparison titled The asymmetry, drawn. Left panel, a small document icon labeled Easy ticket, caption delayed a bit, low cost. Right panel, a larger gauge icon labeled Buried emergency, caption hours late, high cost.
Both boxes are the same size on a screen. They are not the same size in real life.

For most of the year, the end-of-day batch worked fine. Twenty, thirty tickets, reviewed once at close. She'd never once missed something that mattered.

Hand sketched timeline titled A Friday afternoon in July. Four milestones: Walk in freezer ticket arrives 9:40am, Batch keeps growing all day, Ninety four tickets by 5pm review highlighted, Owner calls back angry next morning.
The ticket didn't get lost. It got buried under everything that arrived after it.

Then came a Friday in July. A commercial customer's walk-in freezer ticket landed at 9:40 in the morning, flagged urgent by the AI. Nothing looked wrong yet. Bettina still reviewed once a day, at five, the way she always had.

By five o'clock, ninety four tickets sat in that day's batch. She worked through them in order, the way she always did, and somewhere past ticket eighty, tired and moving fast to clear the queue before heading home, she confirmed the freezer ticket's diagnosis but didn't register how long it had already been sitting. The technician wasn't dispatched until the next morning.

The owner called back angry before any inventory actually spoiled. Nothing broke. It was close enough that it stuck with her.

The freezer ticket was never lost. It was just the eighty fourth thing anyone looked at that day.

With batches capped at ten and reviewed every forty five minutes, that same Friday plays out differently. The freezer ticket sits in a batch of eight others at 10am, gets Bettina's full attention, the same attention every ticket in a small batch gets, and a technician is on site by early afternoon.

I built the review process around one look a day because that was the simplest thing to ship, and it was true at fifteen tickets a day. I never asked what would happen to the same design once volume tripled.

PICK, in one screenNot a lecture on batch processing. PICK is what tells you which error you're actually protecting against.

P
Position. The pick, before reasoning.
Small, frequent batches, eight to twelve tickets, every thirty to forty five minutes.
The direct answer, stated first, not argued into.
I
Impact. Who feels each error.
Too small costs Bettina re-orientation time on every interruption. Too large costs whoever's ticket lands near the bottom of the pile the least careful look.
Names both sides in real, comparable units.
C
Cost asymmetry. The one that matters more.
A delayed easy ticket is cheap and visible. A buried emergency is hidden and expensive, and it's the one to optimize against.
This is PICK's hardest step, and the reason batch size is a real decision, not a preference.
K
Kill criteria. What would flip the pick.
If the AI's emergency-detection accuracy got close to perfect, batch size could safely grow, since the risk it's protecting against would have genuinely dropped.
Separates a confident pick from a stubborn one.
Hand sketched labeled parts diagram titled What is inside a batch review. Center gauge icon labeled Batch Review, with four callouts: ticket count, review cadence, fatigue point, escalation lane.
Batch size looks like one number. It is really four smaller decisions wearing one number.

The recap, one line per letter: position is small, frequent batches over one big daily pile, impact is re-orientation cost against buried-ticket cost, cost asymmetry is optimizing against the hidden emergency rather than the visible delay, and kill criteria is loosening batch size only as real emergency-detection accuracy actually improves.

And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary telehealth line instead of an HVAC dispatch desk.

Cragmoor Veterinary Telehealth runs an AI model that triages pet owner symptom submissions, flagging likely urgency, from a mild skin rash to suspected toxin ingestion. A vet triage nurse reviews the AI's calls in batches before deciding which cases get a callback within the hour versus a routine morning slot.

Hand sketched icon list titled What makes a batch size right. Four items: a gauge icon labeled Small enough to stay fresh, a document icon labeled Frequent enough to catch outliers, a scale icon labeled Big enough to stay efficient, a question mark box icon labeled Separate lane for true emergencies.
The same four checks apply whether the queue is service tickets or vet symptom reports.

Mapped onto PICK: position is the same, small frequent batches over one large end-of-shift pile; impact is a routine rash case waiting a bit longer versus a suspected toxin case sitting unreviewed for hours; cost asymmetry is that a delayed rash costs an annoyed pet owner, while a delayed toxin case costs an animal's life, so the batch design has to optimize against the second one even though it's rarer; kill criteria is the same, loosen batch size only once the AI's own toxin-flagging accuracy is trusted enough that a missed case becomes genuinely unlikely rather than merely rare.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "small, frequent batches, because batch size decides how long an emergency gets to hide, not how fast the queue clears," and stop.
Cost: there's no staffing budget to review every forty five minutes around the clock. Say so honestly, and shrink batch frequency overnight only, when true emergencies are rarer, while keeping it tight during peak daytime hours.
The model gets better, for real: if the AI's urgency flagging becomes highly reliable, the safe batch size can grow, since the buried-emergency risk that set the original ceiling has genuinely gone down.

Where people run it wrong.
They set batch size once, at a lower volume, and never revisit it as volume grows past what made that size safe.
They treat batch size purely as a throughput or staffing question, with no thought to what's hiding inside a large pile.
They shrink to reviewing one ticket at a time to "be safe," and lose more time to constant interruption than the risk was ever worth.

How to use it live. When someone asks about batch size, ask one question back to yourself first: what's the worst thing that could be sitting at the bottom of the biggest batch I'd allow? Size the batch to that answer, not to whatever feels efficient.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an "A or B" tradeoff question like batch size?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Commit to a pick, then show the asymmetry that justifies it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bettina Kirsch, a dispatcher at Northfield Comfort Systems who reviews the AI's HVAC ticket diagnoses before technicians get sent out.
3 · THE IMPACT
Who feels each kind of batch-size error, and how?
Tap to flip
ANSWER
Too small costs Bettina re-orientation time on every interruption. Too large costs whoever's ticket lands near the bottom of the pile the least careful look.
4 · THE COST ASYMMETRY
Which kind of error should batch size actually protect against?
Tap to flip
ANSWER
The buried emergency, hidden and expensive, not the delayed easy ticket, which is cheap and visible and gets noticed on its own.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting one end-of-day batch review when volume was fifteen tickets a day, and never revisiting it once volume tripled in peak season.
6 · THE NUMBER
Fill in the blank: the miss rate on urgent tickets rose from 2 percent early in a batch to ___ percent for tickets 80 through 90.
Tap to flip
ANSWER
11 percent. More than five times higher, from fatigue alone, on the same dispatcher's same shift.
7 · THE REPLAY
Same Friday, capped batches of ten, reviewed every 45 minutes. What changes?
Tap to flip
ANSWER
The freezer ticket sits in a batch of nine others at 10am, gets full attention, and a technician is on site by early afternoon instead of the next morning.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the same cost asymmetry?
Tap to flip
ANSWER
Cragmoor Veterinary Telehealth's symptom triage line. Same asymmetry: a delayed routine case is cheap, a buried toxin-ingestion case can cost an animal's life.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: a single end-of-day batch review averages a ___ minute wait before an emergency ticket gets escalated, versus 22 minutes for small frequent batches.
Show hint
Look at the grouped bar chart in "Let's learn."
Show answer
296 minutes. Almost five hours, and that gap exists purely from the review schedule, before fatigue even enters the picture.
True or false
2. True or false: shrinking review to one ticket at a time, reviewed the instant it arrives, is the safest possible batch size.
  • True
  • False
Show hint
Look at priority list item 4.
Show answer
False. Constant interruption costs real re-orientation time on every single ticket, which adds up and can cost more than the risk it's meant to prevent.
Multiple choice
3. Why does batch size matter beyond how long tickets sit in the queue?
  • A. Larger batches always take longer for the AI to generate.
  • B. Attention fades across a long batch, so tickets reviewed later get a less careful look, right when one of them might be the one that matters.
  • C. Technicians refuse to take tickets from large batches.
  • D. Larger batches cost more to store in the database.
Show hint
Look at the line chart on miss rate by position.
Show answer
B. The miss rate on urgent tickets climbs across a batch, purely from fatigue, independent of the underlying volume problem.
Short answer, apply it yourself
4. Think of a queue you clear in one big batch, an inbox, a stack of grading, a moderation queue. What's likely hiding near the bottom that a smaller, more frequent batch would have caught sooner?
Show hint
Think about what you review last when you're tired or rushing to finish.
Show answer
Model answer: Most people can name something they've skimmed too fast near the end of a long batch, a flagged comment, a tail email, a report detail, precisely because attention fades the same way it does in any long review batch.
Short answer, where it wouldn't matter
5. Name a kind of ticket at Northfield where batch size genuinely doesn't matter.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Routine maintenance tickets, filter checks, seasonal tune-ups. A stale review there costs a scheduling inconvenience, not a real emergency sitting unseen.
Short answer, the number question
6. If Northfield's ticket volume tripled again next year, would the same eight-to-twelve batch size still be right? Why or why not?
Show hint
Think about what actually sets the batch size ceiling.
Show answer
Model answer: Not necessarily the same number, but the same principle, review often enough that nothing sits unseen for long and batches stay small enough that attention doesn't fade before the last ticket.
Before you close the answer
Why this works
Tests whether you'll commit to a real number and defend it with an asymmetry, instead of hedging with "it depends on the team."
Follow-up traps
"Why not just review every ticket the instant it arrives?" Response: constant interruption costs real re-orientation time on every single ticket, and at high volume that adds up to more lost time than the delay it prevents.

"Couldn't the AI just sort urgent tickets to the top of every batch?" Response: it should, and does, but that only helps if the AI's own urgency flag is right, which is exactly why a separate immediate-alert lane exists for the cases the model itself is unsure about.
If pressed
Northfield's real fix also logs the AI's own confidence on the urgency flag, so a low-confidence urgent call skips the batch entirely and pages Bettina directly, rather than waiting even forty five minutes.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more