How would you phase out the human as model quality improves?
BOUNDthe arithmetic behind letting go of a review, one tier at a time
Wrenfield Virtual Care runs an AI symptom checker that recommends how urgently a patient should be seen: emergency care now, a same-day video visit, or self-care at home. Nnamdi Okafor supervises the triage nursing team and, under today's policy, reviews every single AI recommendation over a headset before it reaches a patient.
The direct answer
Phase out review tier by tier, not all at once, and let each tier's own measured error rate set the pace, not a launch date. Self-care recommendations can drop to a small sampled review once accuracy holds steady for two months. Moderate cases follow later, on a slower proof. Anything involving a severe symptom like chest pain never drops below full human review, no matter how good the model's accuracy number gets, because the cost of one missed case there is too high to trade for review-labor savings.
Do this, in order
Split review policy by severity tier, not by one blended accuracy number.Why: a 98 percent overall accuracy can hide a dangerous error rate on the one tier where a miss actually kills someone.
Never let the severe tier drop below full human review, regardless of measured accuracy.Why: the cost of a single missed stroke or chest-pain case is high enough that no realistic labor savings justifies the trade.
Require a sustained accuracy threshold over months, not a single good week, before reducing review on any tier.Why: a short hot streak can look like proof and isn't. A model can get lucky over two weeks and unlucky over two months.
Build an automatic rollback: any tier that crosses its error floor again goes straight back to full review.Why: phasing out review is not a one-way door. The moment the evidence changes, the policy has to change back automatically, not after a committee meeting.
Size the reduced review sample so a real drift is still statistically detectable, not just a token spot-check.Why: a 2 percent sample that's too small to notice a real problem is decoration, not oversight.
Leave self-care recommendations for the most routine, lowest-severity symptoms at the loosest review level first.Why: this is genuinely the tier where being wrong costs the least, so it's the correct place to spend the first round of trust.
How to answer this, stage by stage
Nobody is grading whether you can say "phase it out gradually." They're grading whether you can show the actual numbers behind where the line gets drawn, and why one tier never moves.
Stage 1
Scope it to one concrete queue
Say it like this
"I'll ground this in a real queue: a telehealth symptom checker recommending emergency, same-day, or self-care, with a nurse supervisor reviewing every call today."
Why this works
Keeps the answer from becoming a vague policy statement with no real numbers behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down into an equation. Own each number. Use a range instead of one guess. Nail a sanity check. Say which assumption swings it most."
Why this works
Signals a real estimation method, not a story pretending to be math.
Stage 3
State the equation out loud
Say it like this
"Review cost equals review percentage times daily volume times minutes per review times a nurse's loaded cost per minute. Risk equals the miss rate on unreviewed cases times the cost of a missed case. I'm trading one against the other, tier by tier."
Why this works
Shows the arithmetic before touching a single number, so the estimate is checkable, not asserted.
Stage 4
Own the numbers behind each tier
Say it like this
"Say self-care cases are 55 percent of volume with a measured 0.4 percent error rate, moderate cases are 37 percent at 3 percent error, and severe cases are 8 percent at 1.2 percent error. Those percentages, not intuition, decide the order."
Why this works
Anchors the plan in actual measured numbers instead of a gut feeling about which tier "seems fine."
Stage 5
Give the phased answer, with a range
Say it like this
"Self-care drops to a 10 percent sample once its error rate holds under half a percent for two to four months. Moderate cases follow at 50 percent review, eight to fourteen months out. Severe cases stay at full review, indefinitely, full stop."
Why this works
This is the direct answer, stated as a range rather than a single false-precise date.
Stage 6
Run the sanity check
Say it like this
"The blended plan saves around sixty percent of review labor cost within a year. A single missed severe case, in cost and harm, would wipe out months of those savings on its own. That's why severe never moves."
Why this works
Tests the plan against a known real cost, instead of trusting the arithmetic blind.
Stage 7
Name what would change the answer most
Say it like this
"The severe tier's error rate is the number I'd watch hardest. Even a small rise there swings the whole plan, because that's the one tier where the harm side of the equation is enormous."
Why this works
Shows you know which assumption is load-bearing, not treating every number as equally important.
Stage 8
Close on the one line
Say it like this
"You don't phase out the human by trusting a model more overall. You phase out review tier by tier, at the pace the evidence earns, and you leave the highest-harm tier alone no matter how good the number looks."
Why this works
Restates the direct answer in one breath, ready for a follow-up.
Let's learn
Picture four thousand symptom checks a day, and one nurse team reviewing every single one before a patient hears back.
Wrenfield's model reads a patient's described symptoms and recommends emergency care, a same-day visit, or self-care, in seconds. Before the model, a nurse triaged every call by phone, slower, but nobody's recommendation reached a patient unreviewed.
Confidence alone doesn't decide the order. The severe tier sits high on harm no matter how confident the model gets.
Now Nnamdi's team still reviews every recommendation, at every severity level, because the policy never distinguished between a wrong self-care suggestion and a wrong emergency call. Both cost the same nurse-minute today, even though they're nowhere near the same risk.
Here's the turn: the real question was never "is the model good enough overall." It's "which specific tier is good enough, proven over time, to earn less review, and which tier's harm ceiling makes that trade a bad one no matter what the number says."
Daily review-labor cost build-up, by severity tier, under the phased plan
Red's share of the cost barely shrinks. That's the point, not a bug: it's the tier that stays fully reviewed.
At its worst, review effort keeps getting spent evenly across all three tiers forever, wasting a team's attention on thousands of low-risk self-care calls while giving no more scrutiny to the handful of severe ones that actually deserve it.
Three different clocks for three different tiers. Only one of them never finishes.
The decision that mattered
Set review policy by measured error rate per severity tier, with the severe tier held at full review permanently regardless of accuracy. This isn't caution for its own sake, it's the sanity check: no plausible labor savings on that tier outweighs the cost of even one missed severe case.
What I would leave alone: nothing about the severe tier's policy changes, ever, on this plan. That's the one place where "leave it alone" means never touching the review requirement at all, not a temporary caution.
The lesson: a single blended accuracy number hides exactly the tier where a miss matters most. Split the number before you trust it.
Now here is the same thing as a story
The short version above is what you'd say pitching this plan to Wrenfield's chief medical officer. Read this one for how the tiered thinking actually got built.
Nnamdi Okafor has supervised triage nurses for seven years, and before the symptom checker, his team fielded every call by phone, sorting genuine emergencies from routine worry with nothing but a script and a lot of practice.
Knowledge spark: why not just pick one accuracy number for the whole model?
A single blended accuracy can sit at 98 percent while hiding a 1.2 percent error rate on the rare severe cases that actually cause harm, because those cases are such a small share of total volume that they barely move the average at all.
When the phase-out conversation first came up, the instinct in the room was to pick one accuracy threshold, say 97 percent, and reduce review across the board once the model crossed it. Nnamdi pushed back: "a 97 percent average tells me almost nothing about the one call where someone's having a heart attack."
Three branches, three very different endings. The third one never actually ends.
The team split the accuracy number by tier instead, and the picture changed completely. Self-care recommendations were already extremely reliable. Moderate cases needed more evidence. Severe cases, even at a respectable 1.2 percent error rate, meant roughly one missed serious case for every eighty-three reviewed, out of a small but real daily volume.
What moves the phase-out timeline most, if the assumption is wrong
The red tier's own error rate dwarfs every other assumption. That's the one number worth checking twice.
Nobody was arguing to trust the model less. They were arguing that trust has to be earned separately for the tier where being wrong actually costs a life, not averaged in with the tiers where it costs an extra phone call.
With the tiered plan running, self-care review drops to a small sample after two straight months under the error floor, moderate cases follow later on a longer proof, and severe cases stay fully reviewed, forever, on principle backed by the sanity check: no labor savings there is worth trading against even one preventable miss.
The old framing asked "is the model good enough yet." The new one asks "good enough for which tier, and never for which one."
I proposed the single blended threshold first because it was the simplest policy to write down. It took Nnamdi's one sentence about a heart attack to see that simple and safe were never the same policy.
BOUND, in one screenNot a lecture on rolling out automation. BOUND is what makes the phase-out plan a real estimate instead of a guess with a confident tone.
B
Break it down. State the equation.
Review cost equals review percentage times volume times minutes times cost per minute. Risk equals miss rate times harm cost, computed separately per tier.
Makes the tradeoff checkable before a single number gets picked.
O
Own numbers. State each assumption.
Self-care, 55 percent of volume, 0.4 percent error. Moderate, 37 percent, 3 percent error. Severe, 8 percent, 1.2 percent error, all measured against real follow-up outcomes.
The hardest step: real, checkable numbers instead of assumed ones.
U
Use a range, not a single date.
Self-care in two to four months, moderate in eight to fourteen, severe never, given as ranges tied to a sustained accuracy window, not a guessed calendar date.
Avoids the false precision of a single confident-sounding timeline.
N
Nail the sanity check.
Blended savings of about sixty percent of review labor within a year, compared against the cost of a single missed severe case, which alone would erase months of that saving.
Tests the whole plan against a known real cost instead of trusting the arithmetic blind.
D
Direction. What swings it most.
The severe tier's own error rate. A small rise there swings the entire plan, because the harm side of that equation is so large.
Names the one assumption worth checking twice before trusting the rest.
Phasing out review is not a one-way door. Any of these three flips a tier straight back to full review.
The recap, one line per letter: break it down is the review-cost and harm-cost equation, own numbers is the measured error rate per tier, use a range is a two-to-four and eight-to-fourteen month window rather than one date, nail the sanity check is comparing savings against the cost of one missed severe case, and direction is the severe tier's own error rate as the number that swings everything.
None of these four assume the model stays good forever. They assume it might slip, and build the catch in from day one.
And if you want to be sure it really works, try it somewhere elseSame five letters, a content moderation queue instead of a telehealth line. A completely different kind of harm, the same tiered logic.
Bracklowe Media runs an AI model that scores posts for policy violations before a human moderator acts. Solenne Kadiri leads the trust and safety team and, under the current policy, has every flagged post reviewed by a person before any action, removal or otherwise, is taken.
Mapped onto BOUND: break it down is moderator-minutes times volume times review percentage, weighed against the harm cost of a wrongly removed post or a missed genuine violation; own numbers means splitting posts into low-harm categories like spam, at 62 percent of volume with a 0.3 percent error rate, and high-harm categories like credible threats of violence, at 2 percent of volume with a 4 percent error rate; use a range means spam review can drop to a 15 percent sample within two months of sustained accuracy, while anything flagged as a credible threat stays at full human review indefinitely, the same logic as the severe telehealth tier, just wearing a different industry's numbers.
Swap "symptom" for "post," and the same tiered arithmetic decides which category earns less review, and which never does.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "split the phase-out by harm tier, and never fully automate the tier where a miss is catastrophic, no matter the accuracy number," and stop.
Cost: there's no budget to build a full tiered rollback system this quarter. Say so honestly, and start with the single highest-leverage move, holding the highest-harm tier at full review while the rest of the system is still being built.
The model gets better, for real: even a genuinely excellent model doesn't change the sanity check on the highest-harm tier, because the argument was never about today's accuracy, it was about what a single miss costs no matter how rare it becomes.
Where people run it wrong.
They set one blended accuracy threshold for the whole system, letting a good overall number hide a dangerous error rate on the one tier that matters most.
They treat a phase-out as a one-way door, with no rollback trigger if a tier's error rate creeps back up after review was already reduced.
They pick a phase-out date off a launch deadline instead of a sustained accuracy window measured over real months.
How to use it live. When someone asks how you'd phase out a human reviewer, ask one question first: what does a single missed case in the worst tier actually cost, and does any plausible labor saving really outweigh it? Build the plan around that answer, not around a single confidence number.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "phase out the human" question?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. Show the arithmetic behind the phase-out, don't just assert it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Nnamdi Okafor, triage nurse supervisor at Wrenfield Virtual Care, seven years supervising the team.
3 · THE EQUATION
What's the core tradeoff this plan is built on?
Tap to flip
ANSWER
Review-labor cost per tier, against the cost of a missed case on that same tier, computed separately instead of blended into one number.
4 · OWN NUMBERS
What are the three tiers and their measured error rates?
Tap to flip
ANSWER
Self-care, 55% of volume at 0.4% error. Moderate, 37% at 3% error. Severe, 8% at 1.2% error.
5 · THE RANGE
What's the phase-out timeline, per tier?
Tap to flip
ANSWER
Self-care in 2 to 4 months, moderate in 8 to 14 months, severe never, on a permanent full-review policy.
6 · THE NUMBER
Fill in the blank: the phased plan is estimated to cut daily review-labor cost from 900 dollars to about ___ dollars.
Tap to flip
ANSWER
396 dollars, roughly a 56 percent cut, while the severe tier's own cost barely moves.
7 · DIRECTION
Which single assumption would change this plan the most if it were wrong?
Tap to flip
ANSWER
The severe tier's own error rate. A small rise there swings the entire plan, because the harm side of the equation is so large.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the never-automate tier there?
Tap to flip
ANSWER
Bracklowe Media's content moderation queue. Posts flagged as credible threats of violence stay at full human review indefinitely.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: the severe tier makes up 8 percent of daily volume with a measured error rate of ___ percent.
Show hint
Look at Stage 4's "own the numbers" example.
Show answer
1.2 percent. Small in share of volume, but the tier where the harm cost of a miss is highest.
Multiple choice
2. Why does this plan phase out review differently for each severity tier instead of using one accuracy threshold for the whole system?
A. Different tiers require different software to run.
B. A single blended accuracy number can hide a dangerous error rate on the small tier where a miss actually causes serious harm.
C. Regulators require different thresholds by law in every case.
D. It's cheaper to build three separate dashboards than one.
Show hint
Look at the knowledge spark.
Show answer
B. A 98 percent blended accuracy can sit comfortably above the severe tier's real, more dangerous error rate, since severe cases are such a small share of total volume.
True or false
3. True or false: once a tier's review is reduced, the policy should stay fixed even if that tier's error rate later rises.
True
False
Show hint
Look at priority list item 4.
Show answer
False. An automatic rollback returns any tier to full review the moment its error rate crosses the floor again. Phase-out is not a one-way door.
Short answer, apply it yourself
4. Think of a product you use where some actions matter far more than others, a bank app, a home security system, a work tool. What would the "never automate this tier" category be, and why?
Show hint
Think about which action, if wrong, would be the hardest or costliest to undo.
Show answer
Model answer: A common answer: a banking app's wire transfers above a large threshold, since a wrongly approved one is very hard to reverse and can be very costly.
Short answer, where it wouldn't matter
5. Name a category in this plan where phasing out review is genuinely the right call, not just a cost-cutting shortcut.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Self-care recommendations, where the measured error rate is already very low and the cost of being wrong is genuinely small compared to the severe tier.
Short answer, the number question
6. If the severe tier's error rate rose from 1.2 percent to 4 percent, would the "review is capped at 8 percent of volume, so it's cheap to keep" argument still justify full review there? Why or why not?
Show hint
Think about whether the argument for full review depends on volume share at all.
Show answer
Model answer: Yes, and even more strongly. The argument for full review on this tier was never about its share of volume, it was about the cost of a single miss, which only gets worse as the error rate rises.
Before you close the answer
Why this works
Tests whether you can turn "phase out the human" into real, checkable arithmetic instead of a confident-sounding rollout story.
Follow-up traps
"Isn't keeping the severe tier at 100 percent review forever just avoiding the hard problem?" Response: it's the sanity check speaking, not avoidance. No plausible labor saving on 8 percent of volume outweighs the cost of one missed severe case, so the math itself says to leave it alone.
"What if the model becomes so accurate that even the severe tier's error rate drops to near zero?" Response: the policy would still hold, because the argument is about the cost of a single miss, not the frequency, and a rare miss on the highest-harm tier is still the most expensive kind there is.
If pressed
Wrenfield's actual plan also requires the reduced-review sample size on each tier to stay large enough to detect a real doubling of that tier's error rate within one month, not just a token handful of spot checks, since a sample too small to notice drift is not really oversight at all.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.