CalculationIntermediateDesigning for Uncertainty & Trust / Human-in-the-loop product design / #19

What SLA does a human review step impose on the overall product?

BOUND the product is Codewatch, the AI compliance assistant Brindlewood County's permits office runs ahead of a human sign-off

Brindlewood County's permits office uses Codewatch, an assistant that reads a building permit application and drafts a recommend, flag, or deny before a person opens the file. Ottilie Marsh runs the permits program and has to answer for how long a decision actually takes.

The direct answer
A human review step doesn't add a flat few minutes to the SLA. It adds a queueing delay that depends on reviewer headcount against application volume, and that queueing term dominates the total far more than the review itself ever does. For Codewatch, the honest SLA is under an hour on a normal day and up to four hours at the office's known monthly peak, and the number that actually protects that promise is headcount, not review speed.
Do this, in order
  1. State the SLA as a range tied to a known peak, never as one single number.Why: a single number hides the one week a month that actually breaks it.
  2. Break the total time into its real parts before quoting any figure at all.Why: intake, AI draft, queue wait, human review, and publish are five different things, not one blob called "review."
  3. Model queue wait as a function of headcount against volume, not a fixed constant.Why: that's the one part of the total that actually swings under load.
  4. Sanity-check the range against both the office's stated target and the old manual baseline.Why: a number with nothing to compare against is just a guess with confidence attached.
  5. Watch reviewer headcount against volume as the assumption that would break this first.Why: it's the single biggest lever on the total, by a wide margin.
  6. Don't spend effort speeding up Codewatch's own draft time.Why: it's under a minute already, a rounding error next to the queue.

How to answer this, stage by stage

Nobody is grading whether you can multiply. They're grading whether you find the term that actually moves, instead of dressing up a guess in decimal points.

Stage 1
Scope it to one real pipeline
Say it like this
"I'll estimate this for Codewatch, the AI assistant Brindlewood County's permits office runs ahead of a human sign-off on every building permit application."
Why this works
Grounds an abstract question in one real pipeline instead of "a product" in general.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break the equation down, own each number, give a range instead of one figure, sanity-check it, and say which assumption would move it most."
Why this works
Tells the interviewer you have a method before a single number gets said out loud.
Stage 3
Break down the equation
Say it like this
"Total time to a decision equals intake time, plus Codewatch's own draft time, plus how long the application waits for a free reviewer, plus the reviewer's own time on it, plus the time to sign off and publish."
Why this works
States the arithmetic before touching a number, so nothing later feels invented on the spot.
Stage 4
Own the numbers, and give the range
Say it like this
"I'll assume eighty five applications a day on average, three reviewers at six productive hours each, forty seconds for Codewatch to draft a recommendation, and about six minutes of human review. On a normal day that's under an hour, start to finish. In the last week of the month, when volume nearly doubles, the queue alone can push that past four hours."
Why this works
A single number here would have hidden the one week that actually matters, which is the whole point of the question.
Stage 5
Sanity-check it
Say it like this
"The county's own target is same-day for standard permits, and the old manual process took three to five business days. Even the slow end of my range beats both, so this number isn't a red flag, it's a fact worth planning around."
Why this works
Compares the estimate to something real, instead of leaving it floating with no context.
Stage 6
Say what moves it, and close
Say it like this
"The one assumption that changes this the most is reviewer headcount. Double the reviewers on peak days and the queue term nearly disappears. Double Codewatch's own draft speed and the total barely moves at all, because the draft was never where the time was hiding."
Why this works
Ends by naming where a real hiring or engineering dollar should actually go, not just where the story is dramatic.

Let's learn

Every evening, Brindlewood County's three permit reviewers used to read every application cover to cover by hand, checking it against zoning code themselves. Three to five business days per decision, no matter how straightforward the application was.

Hand sketched comparison diagram titled Manual review versus AI-assisted, same decision. Left panel, a document icon labeled Manual before, caption 3 to 5 business days. Right panel, a gauge icon labeled AI-assisted now, caption under an hour most days.
Same decision, same reviewers. What changed is how long it waits before a person ever sees it.

Codewatch drafts a recommendation in under a minute now, and on a normal day, the whole thing, application in, decision out, takes under an hour.

The extra speed itself was never the risk. The risk was treating "under an hour" as the SLA, when it was really "under an hour, when reviewers have room," and reviewers stopped having room the one week a month that mattered most.
Where the total time goes, normal day versus peak week (minutes)
250 min 125 0 57 min Normal day 250 min Peak week
The clay-colored band is queue wait. On a normal day it's a sliver. At peak, it is almost the entire bar, while intake, AI draft, review, and publish barely change at all.

At its worst, during the last week of the month, when applications spike ahead of a fee-schedule change, the queue alone stretches the total past four hours, and nobody had written that number down anywhere as the actual SLA.

The decision I would take back Brindlewood County sized reviewer headcount off the daily average application volume from the original planning document, not the known peak week every fee change causes. That made sense when average was the only number anyone had measured. It stopped making sense the first fee change after Codewatch shipped, since it was the first time the queue actually became visible to an applicant.

What I would leave alone: Codewatch's own draft time, under a minute, isn't worth optimizing further. It was never where the delay was hiding, and chasing it would be effort spent on the wrong term of the equation.

The lesson: a human review step's real cost to an SLA isn't how long the review itself takes. It's how long the line in front of the reviewer gets once volume outpaces headcount, and that number moves a lot more than the review ever does.

Now here is the same thing as a story

The short version above is what you'd say defending this number to the county council. Read this one for how the gap actually got found.

Ottilie Marsh has run Brindlewood County's permits program for four years. She's the one who has to explain, in front of the council, why some applicants wait an hour and others wait most of an afternoon for what should be the same kind of decision.

Hand sketched labeled parts diagram titled What builds the total SLA. Center gauge icon labeled Total SLA, with five callouts around it: intake, AI draft, queue wait, human review, publish.
Five parts. Four of them barely move. One of them is the whole story.

For most of the year, Codewatch had made her job easier. Applications that used to take days now cleared same-day, and the council had started pointing to the permits office as a model for the rest of the county's services.

Hand sketched flow diagram titled Codewatchs permit pipeline. Five boxes: application filed, Codewatch drafts, queue for reviewer highlighted, human reviews, decision published.
Four boxes barely changed when Codewatch arrived. The middle one is where the whole SLA actually lives.

Then came the last week of a fee-schedule change, when applications nearly doubled as people rushed to file before new rates took effect. Codewatch kept drafting recommendations in under a minute, same as always. But the queue in front of Brindlewood County's three reviewers, whose headcount hadn't grown since the original volume plan was written two years earlier, stretched past four hours by Thursday afternoon.

Hand sketched timeline titled A normal-day application, start to finish. Four milestones: filed at 0 minutes, Codewatch drafts at plus 1 minute, queue clears at plus 46 minutes highlighted, decision published at plus 57 minutes.
This is what the same walk looked like on a normal day. Peak week, only one of these four numbers changed.

A local contractor called the office asking why an application that "used to take an hour" was still sitting after lunch. Petra pulled the logs and found the honest answer: nothing about Codewatch or the reviewers had gotten slower. The line of applications waiting for a person had simply gotten longer than anyone had sized for.

Knowledge spark: why does queue wait grow faster than volume does? When arrivals outpace how fast reviewers can clear them, even for a short stretch, the backlog doesn't grow steadily, it compounds, because every new application waits behind every one still stuck in the line ahead of it.

Nothing about the AI's own part of the job had changed by a single second. It was the line of applications waiting for a person that grew, and nobody had written that line into the SLA at all.

Hand sketched decision tree titled What Codewatch routes, and to where. Root Codewatch drafts a recommendation, branching to three leaves: clean low complexity leads to fast-track review, flagged zoning conflict leads to full reviewer sign-off, missing documents leads to returned to applicant.
Not every application needs the same amount of a reviewer's time. The queue math changes depending on which branch it takes.

With headcount sized to the known peak week instead of the average, and the SLA quoted as a range instead of one figure, the same Thursday call from the contractor gets a different answer: still filed after lunch, but now inside a number the office actually promised, instead of one that quietly slipped past a promise nobody had written down.

I built the staffing plan off the number that looked calm on a spreadsheet. It took one fee-change week and one contractor's phone call to see that the number that actually mattered was the one that only shows up the week volume spikes.

BOUND, the arithmetic behind the promiseNot a guess dressed up with decimals. BOUND is what tells you which term of the equation is actually worth watching.

B
Break it down. State the equation first.
Total time to decision equals intake time, plus Codewatch's draft time, plus queue wait for a free reviewer, plus the reviewer's own time, plus the time to sign off and publish.
Nothing that follows is invented on the spot, since the shape of the answer was said out loud first.
O
Own numbers. Say where each one came from.
Eighty five applications a day, from the office's own filing logs. Three reviewers at six productive hours each, from the staffing roster. Forty seconds to draft, from Codewatch's own timing logs. About six minutes of human review, from a sample of recent cases.
Every figure is checkable, not a number that sounds plausible and isn't backed by anything.
U
Use a range, not a single point.
Under an hour on a normal day. Up to about four hours at the office's known monthly peak, almost all of it the queue, not the review itself.
The hardest step and the direct answer: a single number would have hidden the exact week that broke the promise.
N
Nail the sanity check.
The county's own target is same-day for standard permits, and the old manual process took three to five business days. Even the slow end of this range clears both comfortably.
If the peak number had come out worse than the old manual process, that would be a sign the model was wrong, not just a sign of a busy week.
D
Direction. What moves the estimate most.
Reviewer headcount against volume, by a wide margin. Codewatch's own draft speed barely moves the total at all, no matter how much faster it gets.
Points a real hiring or engineering decision at the term that actually matters.
Hand sketched icon list titled What actually determines the SLA. Three items: a person icon labeled Reviewer headcount, a funnel icon labeled Application volume, a gauge icon labeled AI draft speed.
Only the first two of these three actually move the number that matters.
How much the total SLA swings if each assumption changes
Reviewer headcount halved 180 min Volume up 50% 140 min Review time up 50% 9 min Codewatch draft doubled <1 min
The two bars about reviewers and volume dwarf the other two. That's the whole argument for where the SLA risk actually lives.

The recap, one line per letter: break it down is the five-part equation, own numbers is each figure traced to a real log, use a range is under an hour normally against four hours at peak, nail the sanity check is beating both the county's target and the old manual baseline, and direction is reviewer headcount, not draft speed, as the lever that actually matters.

And if you want to be sure it really works, try it somewhere elseSame five letters, a prescription prior authorization instead of a building permit. A different queue, the same hidden term.

Bramblecote Pharmacy Group uses ScriptCheck, an assistant that reviews a prior-authorization request against a patient's plan rules and drafts an approve, deny, or escalate before a pharmacist signs off. Folasade Nwachukwu is a pharmacist who reviews ScriptCheck's flagged cases each shift.

Mapped onto BOUND: break it down is the same shape, request intake plus ScriptCheck's draft plus queue wait for a free pharmacist plus the pharmacist's own time plus the pharmacy's own release step. Own numbers looks different here, about 40 requests an hour at a busy pharmacy, two pharmacists on shift, ScriptCheck drafting in about 20 seconds, and roughly 3 minutes of pharmacist review per escalated case. Use a range separates a quiet Tuesday morning, where a request clears in under ten minutes, from a Monday after a holiday, where refill requests pile up and the queue alone can add forty minutes. Nail the sanity check compares that against the pharmacy chain's own promise of same-visit fills for routine prescriptions, which the slow end of the range still meets. Direction is, once again, staffing against volume, not how fast ScriptCheck drafts its recommendation.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the SLA is a range driven by queue wait, and queue wait is driven by headcount against volume," and stop.
Cost: there's no budget to hire more reviewers before the next peak. Say so honestly, and shift a spare reviewer onto the known peak days instead of spreading headcount evenly across the month.
The model gets better, for real: if Codewatch's draft accuracy improves further, that still doesn't touch the queue term at all, since the draft was never the part of the total that was slow.

Where people run it wrong.
They quote a single average number as the SLA, and the one week it breaks becomes a crisis nobody saw coming.
They spend engineering time speeding up the AI step, because it's the part that's visible and interesting, while the queue quietly does all the damage.
They size headcount off an average from a planning document instead of the known peak the business itself already causes, like a fee change or a holiday rush.

How to use it live. When someone asks what SLA a review step imposes, ask yourself one question first: is the number they want a flat cost, or a queue. If it's a queue, headcount against volume is almost always the number that actually decides it, not how fast any one step runs.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "what SLA does this impose" calculation question?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. The range step is the answer, since a single number hides the peak.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ottilie Marsh, who runs Brindlewood County's permits program and has to explain the office's turnaround times to the county council.
3 · THE EQUATION
What are the five parts of the total time to a decision?
Tap to flip
ANSWER
Intake time, Codewatch's draft time, queue wait for a free reviewer, the reviewer's own time, and the time to sign off and publish.
4 · THE HIDDEN TERM
Which part of the equation actually swings between a normal day and peak week?
Tap to flip
ANSWER
Queue wait. It goes from about 46 minutes to about 235 minutes, while intake, draft time, review time, and publish barely change.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Sizing reviewer headcount off the average daily volume in the original planning document, instead of the known peak week every fee change causes.
6 · THE NUMBER
Fill in the blank: on a normal day, the total time to a decision is under ___ minutes.
Tap to flip
ANSWER
About 57 minutes. At peak week, that same total stretches to about 250 minutes, almost all of it queue wait.
7 · THE REPLAY
Same fee-change week, headcount sized to the known peak. What changes?
Tap to flip
ANSWER
The contractor's application still files after lunch, but now inside a promised range instead of quietly slipping past a number nobody had written down.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
ScriptCheck, a pharmacy prior-authorization assistant. Same shape: pharmacist headcount against request volume drives the SLA, not how fast ScriptCheck drafts its recommendation.

Check yourself Score: 0 / 0

True or false
1. True or false: Codewatch's own draft time is the main reason the total SLA gets longer during peak week.
  • True
  • False
Show hint
Look at the stacked bar chart.
Show answer
False. Codewatch's draft time is under a minute in both cases. The queue wait for a free reviewer is what balloons at peak, from about 46 minutes to about 235.
Multiple choice
2. Why does the answer give the SLA as a range instead of a single number?
  • A. Because ranges sound more impressive in an interview.
  • B. Because Codewatch's draft time itself varies a lot day to day.
  • C. Because a single average number hides the one peak week where the queue actually breaks the promise.
  • D. Because BOUND requires exactly two numbers in every answer.
Show hint
Look at the "use a range" step.
Show answer
C. The peak week is where the real risk to the SLA lives, and a single average number would erase it entirely.
Fill in the blank
3. Fill in the blank: halving reviewer headcount swings the total SLA by about ___ minutes, the largest swing of any assumption tested.
Show hint
Look at the sensitivity chart.
Show answer
About 180 minutes. Far larger than the swing from review time or Codewatch's own draft speed, which is why headcount is the lever worth watching.
Short answer, apply it yourself
4. Think of a product you use that has a support or review queue. What's the one assumption that would blow up its wait time first, if it changed?
Show hint
Think about what happens to a help desk or a delivery service during a holiday rush.
Show answer
Model answer: Most queue-based services name a similar answer, staff or capacity against a demand spike, since demand rarely rises as a smooth average, it rises as a sudden peak.
Short answer, where it wouldn't matter
5. Name a part of this same pipeline where speeding things up would not meaningfully change the total SLA.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Codewatch's own draft time. It's already under a minute, so even doubling its speed barely moves the total, since the queue is where almost all the time actually sits.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Sizing reviewer headcount off average daily volume rather than the known peak week. It made sense while average was the only number anyone had measured, and stopped making sense the first real peak after Codewatch shipped.
Before you close the answer
Why this works
Tests whether you'll model a review step as a queue, not a flat cost, and whether you can point to the one assumption that actually drives it. Most candidates just estimate the review time itself.
Follow-up traps
"Couldn't you just make Codewatch faster to fix the peak-week problem?" Response: no, the draft step is already under a minute; doubling its speed swings the total by less than a minute, since the queue, not the draft, is where the time actually sits.

"Isn't hiring more reviewers expensive just for one week a month?" Response: that's a real tradeoff, and it's why the answer names headcount against volume as the lever, not a blanket hire, a temporary or cross-trained reviewer for the known peak week is cheaper than promising an SLA the office can't keep.
If pressed
The queue math here follows a simple congestion pattern: once arrivals outpace how fast reviewers clear cases, even for a few days, the wait doesn't grow in a straight line with the extra volume, it compounds, because each new application waits behind everyone still stuck ahead of it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more