ConceptAdvancedAI Opportunity & Model Strategy / When NOT to use AI / #17
Describe a workflow where partial automation is worse than none.
TRACEa green checkmark that only ever meant half of what Tavish thought it meant
Procurely builds OrderGate, a purchase order approval tool. It added a "Pre-Checked" badge that confirms a PO's format and vendor ID are filled in correctly before it reaches a human approver. Tavish Bellhurst is the procurement manager who signs off on POs. Ondine Marchetti is the PM who has to work out why a fraudulent-looking $42,000 PO made it all the way through.
The direct answer
Partial automation is worse than none when the automated slice creates a false sense that the whole job got done. Here, a "Pre-Checked" badge only confirms format and vendor ID, never budget limits or vendor legitimacy, but it reads to a busy human as full sign-off. The fix isn't removing the badge, it's naming exactly what it checked and what it didn't, every single time it shows up.
Do this, in order
Measure the real end-to-end catch rate, not the automated step's own accuracy.Why: the pre-check being 99 percent accurate at its own narrow job says nothing about whether the whole workflow still catches what matters.
Name exactly what a partial automation badge covers, every time it appears.Why: "Pre-Checked" without a scope reads as full approval to a busy human, which is exactly the gap that let a real PO through.
Watch for the metric that actually matters drifting weeks after a "faster" launch.Why: review time improving immediately can hide a real catch-rate decline that takes weeks to show up in the numbers.
Check whether the automated step adds a mode-switch cost, not just time saved.Why: trusting a badge sometimes and verifying from scratch other times is its own new source of error a fully manual process never had.
Train explicitly on what the automation does NOT cover, not just how to use it.Why: nobody had ever told Tavish, in writing, that budget and vendor legitimacy were still entirely his job.
If the scope can't be named clearly, keep the step fully manual.Why: a badge that can't honestly describe its own limits is more dangerous than no badge at all.
How to answer this, stage by stage
Nobody is scoring whether you can define partial automation. They're scoring whether you'd have caught a green checkmark quietly meaning less than everyone assumed.
Stage 1
Scope it to one real workflow, not automation in the abstract
Say it like this
"Let me give you a real case. A purchase order tool adds a 'Pre-Checked' badge that only confirms format and vendor ID. A human approver starts treating that badge as full sign-off, and a fraudulent-looking PO gets through because nobody was checking budget or vendor legitimacy anymore."
Why this works
Grounds the answer in something checkable instead of a general warning about half-measures.
Stage 2
Say your structure out loud before any content
Say it like this
"I'll run this as TRACE. Timeline, what shipped and when the real number actually dropped. Recut, why partial automation specifically causes this. Assume nothing about what one improving number means. Cause candidates, the real suspects. Evidence test, the one check that confirms it."
Why this works
Signals a repeatable diagnostic method, not a single incident recalled from memory.
Stage 3
Reframe the question: this isn't about the badge being wrong
Say it like this
"The pre-check itself works fine, it correctly flags format and vendor ID nearly every time. The real problem is that a human reading 'Pre-Checked' has no way to know it means only that, not 'this whole PO is fine.' That gap is where the risk actually lives."
Why this works
This is where a strong answer separates itself from just blaming the automation for being inaccurate.
Stage 4
Give the one decision: name the scope, every time the badge shows up
Say it like this
"Here's what I'd actually do. Change the badge from 'Pre-Checked' to something like 'Format verified, budget and vendor still need your review.' Same automation, same speed, but now it can't be mistaken for something it never was."
Why this works
This is the direct answer, stated as something you'd actually build, not just a warning to be careful.
Stage 5
Prove it with the compressed failure
Say it like this
"This is exactly what happened. Review time dropped from 14 minutes to 4 the week the badge launched, looked like a clean win. Eight weeks later, a $42,000 PO with a fake-looking vendor sailed through, because it was correctly formatted, so it got the badge, and Tavish had quietly stopped checking anything the badge didn't cover."
Why this works
This is where the story lives, compressed to the one number and the one moment that actually proves the decision.
Stage 6
Say what you'd measure, and name the AI-specific risk
Say it like this
"I'd track real catch rate for budget and vendor issues specifically, not the pre-check's own format accuracy, because those are two completely different numbers and only the second one tells you if the workflow is actually safe. A model that's accurate at a narrow task can still make the whole system less safe if a human silently over-trusts what it covers."
Why this works
Names the real failure mode, silent over-trust of a partial signal, and pairs it with the one measurement that would catch it.
Stage 7
Say what you'd leave alone, then close on one line
Say it like this
"This isn't an argument against automating any part of PO review. The format and vendor ID check genuinely saves real time and it's reliable at that specific job. The fix is naming its edges honestly, not tearing it out. Say what a partial check covers, every time, or don't ship the badge at all."
Why this works
Closes with real judgment instead of blanket suspicion of automation, and restates the direct answer in one breath.
Let's learn
OrderGate is a tool that reviews purchase orders before a human approver sees them, catching obvious formatting problems so approvers spend their time on judgment calls instead of typos.
Nothing about the launch week predicted week eight. That gap is exactly what makes this failure mode dangerous.
Before the badge, Tavish reviewed every PO the same way, checking format, budget limits, and vendor legitimacy together, every time. It took about 14 minutes per PO on average, and manual review caught roughly 92 percent of problematic POs before approval.
Catch rate for budget and vendor issues specifically, before and after the badge shipped
Manual review onlyPre-check badge in place
The pre-check itself never touched budget or vendor legitimacy. This drop happened entirely because a human quietly stopped checking what the badge was never built to cover.
After the badge shipped, the average review time fell to 4 minutes, a real, immediate, celebrated improvement. Nobody was watching the one number that actually mattered: how often a genuinely problematic PO still got caught.
We didn't make Tavish's review faster. We made it look complete when it had quietly stopped covering most of what it used to.
Here's the turn: the pre-check's own accuracy was never the problem, it correctly validated format and vendor ID almost every time. The turn is that "Pre-Checked" read as a broader claim than it was ever built to make, and Tavish, busier that quarter than usual, took the shortcut the badge seemed to offer.
Knowledge spark: what makes partial automation different from full automation here?
Full automation removes the human step entirely, so nobody's relying on a mental model of what got checked. Full manual review keeps one consistent mode. Partial automation asks a human to remember, every single time, exactly which slice a badge covers and which slice is still theirs, and that memory quietly erodes under real workload.
At its worst, this cost a real, near-miss loss: a $42,000 PO to a vendor with a barely-plausible name and a properly formatted request got the "Pre-Checked" badge, sailed through Tavish's queue in under a minute, and was only caught in a later audit, not by the workflow that was supposed to catch it.
Catch rate for budget and vendor issues, week by week after the badge shipped
Still close to manual rateSliding, unnoticed
No single bad day drove this. It slid a little each week as Tavish's trust in the badge grew a little at a time, until a $42,000 PO finally tested it directly.
The choice I would take back
Launching the "Pre-Checked" badge with no explicit statement of its scope, on the assumption that approvers would naturally understand it only covered format. That made sense when the team was focused on shipping something useful fast. It stopped making sense the moment real workload gave Tavish a reason to lean on any signal that saved him time.
What I would leave alone: the underlying format and vendor ID check stays exactly as it is, it's fast, reliable, and genuinely saves real minutes. The fix was never removing the automation, only naming what it does and doesn't cover, every time it appears.
The lesson: a badge that automates part of a job and stays silent about the rest doesn't split the work cleanly, it invites a human to quietly assume the rest got handled too. Name the edges of what got automated, out loud, every time.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like to watch a fast queue turn into a quiet blind spot.
Tavish Bellhurst had approved purchase orders at his company for five years. He was good at it in the unglamorous way: methodical, a little slow, and almost never wrong. Every PO got the same three checks, format, budget, vendor, whether it was PO number 4 or PO number 400 that week.
OrderGate's pre-check badge launched on a Tuesday. The first time Tavish saw a green "Pre-Checked ✓" sitting at the top of a PO, he still ran his usual three checks anyway, out of habit. It took him 13 minutes, barely faster than before.
Two of four things Tavish used to check, quietly narrowed to two, without anyone ever saying so out loud.
By week three, his queue had nearly tripled, a headcount gap nobody had backfilled yet. The badge started doing something nobody designed it to do: it became permission. A "Pre-Checked" PO felt safe enough to move through faster, and faster started meaning skipping the budget and vendor steps entirely, just this once, just for the ones with the green check.
There was no single bad Tuesday where Tavish decided to stop checking. It thinned out in small moves: skimming past the badge instead of reading it, trusting it a little more each week his queue grew a little longer.
Three real candidates for what caused the drift. Ondine's evidence test pointed hardest at the first one.
The PO that surfaced it was for $42,000, properly formatted, correct cost center code, vendor name close enough to a real supplier's that nobody glanced twice. It carried the green badge. It took Tavish under a minute to approve.
We didn't lose 56 percent of our catch rate. We lost the one habit that used to catch exactly this kind of PO, one badge at a time, over eight quiet weeks.
Ondine Marchetti, brought in to work out what happened, didn't start by blaming Tavish. She started by asking three suspects: was it the badge's wording, his climbing workload, or a training gap. She ran a quick test, seeding ten known-bad POs, budget overruns and shaky vendors, all correctly formatted, into a review batch, and watched what got caught.
Nine of the ten sailed through with the green badge attached. When she asked Tavish afterward what "Pre-Checked" meant to him, he said, without hesitating, "That it's good to go." Nobody had ever told him otherwise, in writing, once.
Full manual review never asked Tavish to remember which mode he was in. Partial automation asked him to remember it every single time, and that memory wore down under real workload.
Back when the badge shipped, skipping an explicit scope statement wasn't an unreasonable call. The team was focused on shipping something useful fast, and "obviously it just means format" felt self-evident in the room. It stopped being self-evident the moment a real, busy human needed a shortcut and the badge quietly offered him one.
Here's the replay: same badge, same launch week, but with its scope named plainly from day one.
The fork that should have run before the badge ever shipped, drawn out plainly.
With the badge relabeled "Format verified, budget and vendor still yours," Tavish's review time settles around 9 minutes, slower than the celebrated 4, but the seeded-error test now catches 9 of 10 known-bad POs instead of 1. The workflow gets a little slower and a lot safer, on purpose.
One version of this story spends eight quiet weeks looking like a win before a $42,000 near-miss proves it wasn't. The other spends those same eight weeks a little less impressive on a dashboard, and a lot more honest about what actually got checked.
What I'd tell myself, watching Ondine's seeded test come back nine wrong: a badge that doesn't say what it means will always get read as meaning more than it does. Say the edges out loud, or don't ship the badge.
TRACE, run on a checkmark that meant less than everyone assumedNot a script for blaming the human who trusted the badge. TRACE is what finds the real gap between what shipped and what actually broke, weeks apart.
T
Timeline. What shipped, and when did the real metric actually move?
The pre-check badge shipped and review time dropped immediately, a visible win. The real catch rate for budget and vendor issues didn't visibly crater until an audit eight weeks later.
The gap between the launch and the real drop is the whole point, not a footnote.
R
Recut. Why does partial automation specifically cause this, not automation in general?
Full automation removes the human's mental model entirely. Full manual review keeps one mode. Partial automation asks a human to remember which slice got automated, every time, and that memory erodes under real workload.
This is the mechanism, not just the symptom, of why partial specifically is the dangerous middle ground.
A
Assume nothing. What did the team wrongly assume the improving number meant?
Review time dropping from 14 to 4 minutes looked like a clean efficiency win. Nobody checked whether the whole workflow still caught what it used to.
A number getting better is not proof the thing it stands in for got better too.
C
Cause candidates. What are the real suspects, named honestly?
Three candidates: the badge's wording reading as broader than its scope, Tavish's own climbing workload, and a total absence of training on what the badge excludes.
This is the direct answer's evidence, laid out as real, checkable candidates instead of a single assumed cause.
E
Evidence test. What one check actually confirms the cause?
Seeding ten known-bad POs into a real review batch and watching how many got caught. Nine of ten sailed through with the green badge, and Tavish confirmed he read "Pre-Checked" as full sign-off.
A real test against known outcomes, not a guess, is what turns a suspicion into a confirmed diagnosis.
The recap, one line per letter: timeline is the real gap between the launch and the eight-week-later drop, recut is why partial specifically, not automation broadly, causes this, assume nothing is the trap of reading a faster review time as a safer one, cause candidates is the honest shortlist of what might explain the drift, and evidence test is the seeded-PO check that confirmed which one actually did.
And if you want to be sure it really works, try it somewhere elseSame five letters, an insurance claims queue instead of purchase orders. The judgment call this time is whether a fraud score without a label reads as a full decision.
Delano Whitfield adjusts claims at Stonecrest Mutual. A new AI tool flags each claim with a fraud-risk score, "Low Risk," shown at the top of every file. It only ever scores documentation consistency, whether dates and numbers line up, never whether the claimed damage itself is medically or physically plausible. Mapped onto TRACE: timeline is that claim processing time dropped the week the score launched, and the real over-payment rate on implausible claims didn't visibly rise until a quarterly audit two months later. Recut is the same mechanism: adjusters started reading "Low Risk" as a broader signal than documentation consistency, the one thing it actually measured. Assume nothing means not trusting the faster processing time as proof nothing was slipping through. Cause candidates are the label's own vague wording, a real backlog that quarter, and no training on the score's narrow scope. Evidence test is seeding ten claims with medically implausible damage but perfectly consistent paperwork into the queue, seven got a "Low Risk" score and sailed through Delano's review in under two minutes each.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "name the scope of a partial automation step, every time it appears," and stop.
Cost: no time to run a seeded-error test before the meeting. Say so honestly, and propose it as the next concrete step instead of guessing at the cause.
The model got better, for real: say the pre-check's format accuracy climbs to 99.9 percent. The problem doesn't shrink, because the gap was never about the model's accuracy at its own job, it was about what a human silently assumed that job covered.
Where people run it wrong.
They measure the automated step's own accuracy and call the whole workflow safe, without checking the end-to-end catch rate.
They ship a badge or a score with no stated scope, assuming users will naturally infer the narrow thing it actually means.
They treat rising throughput as proof of a healthy workflow, when it can just as easily be proof that real checks quietly stopped happening.
How to use it live. The moment an interviewer describes a workflow where AI handles "part" of a decision, ask yourself first: does the human downstream know exactly which part, every time, or are they filling that gap with an assumption. That question buys real thinking time, and it's usually exactly where the risk is hiding.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits diagnosing why partial automation made a workflow worse?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. It separates when something shipped from when the real metric actually broke, then tests the real cause instead of guessing.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Tavish Bellhurst, the procurement manager who approves POs using Procurely's OrderGate tool. Ondine Marchetti is the PM who diagnoses the near-miss afterward.
3 · THE MECHANISM
Why is partial automation specifically riskier than either full automation or no automation?
Tap to flip
ANSWER
It asks a human to remember, every single time, exactly which slice got automated and which slice is still theirs. That memory erodes under real workload in a way a single consistent mode never has to.
4 · THE FIX
What's the one concrete thing this answer says to actually do?
Tap to flip
ANSWER
Name exactly what a partial automation badge covers, every single time it appears, e.g. "Format verified, budget and vendor still yours," instead of a scope-free "Pre-Checked."
5 · THE OLD DECISION
What old decision would this answer take back?
Tap to flip
ANSWER
Launching the "Pre-Checked" badge with no stated scope, assuming approvers would naturally infer it only meant format. Reasonable while shipping fast felt urgent. Wrong the moment real workload gave someone a reason to lean on the shortcut it seemed to offer.
6 · THE NUMBER
Fill in the blank: catch rate for budget and vendor issues dropped from ___ percent before the badge to ___ percent after.
Tap to flip
ANSWER
90 percent, down to 34 percent. The pre-check itself never touched either of those issue types, the drop was entirely from a human quietly over-trusting what it did cover.
7 · THE EVIDENCE TEST
What one check confirmed the real cause, instead of just guessing?
Tap to flip
ANSWER
Seeding ten known-bad, correctly formatted POs into a real review batch. Nine of ten sailed through with the green badge, confirming Tavish was reading the badge as full sign-off, not just a format check.
8 · CROSS PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what's the equivalent evidence test?
Tap to flip
ANSWER
Stonecrest Mutual's claims queue, where a "Low Risk" fraud score only measures documentation consistency. The equivalent test seeded ten medically implausible but consistently documented claims, seven got a "Low Risk" score and sailed through.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: average PO review time dropped from 14 minutes to ___ minutes the week the pre-check badge launched.
Show hint
Look at "Let's learn."
Show answer
4 minutes. A real, immediate, celebrated improvement that hid a much slower-moving decline in what actually got caught.
Multiple choice
2. Why did partial automation make this workflow worse than either full automation or no automation at all?
A. The AI model's format check was frequently wrong.
B. It asked Tavish to remember which slice of the review was automated, and that memory eroded under real workload.
C. Procurely charged more for the partial automation feature.
D. It made review take longer than before.
Show hint
Look at the R step, "Recut," in the framework recap.
Show answer
B. Full manual review never required remembering a mode. Full automation removes the human's mental model entirely. Only the partial version needs a human to track a scope, and that tracking wore down.
True or false
3. True or false: the fix Ondine's team chose was to remove the pre-check badge entirely.
True
False
Show hint
Look at "what I would leave alone" and the replay in the story.
Show answer
False. The format check stayed. The fix was relabeling the badge to state its scope plainly, "Format verified, budget and vendor still yours," not removing the automation.
Short answer, where it wouldn't matter
4. Name a part of the PO review process where this scope-labeling problem does NOT apply, and say why not.
Show hint
Think about a fully automated step versus a fully manual one, not a partial one.
Show answer
Model answer: A step that's fully automated with no human review at all, like auto-rejecting a PO with a malformed vendor ID field, doesn't create this risk, because there's no human forming an assumption about what got checked.
Short answer, apply it yourself
5. Think of a tool you use that shows a badge, score, or checkmark covering only part of a decision. What do you actually assume it means, and is that assumption written down anywhere?
Show hint
Think of a spam filter, a spell checker, or a credit-card fraud alert.
Show answer
Model answer: An email spam filter's "clean" inbox reads as "no phishing risk," when it usually only checks known spam patterns, not whether a legitimate-looking sender address is actually spoofed, a gap most users have never had explained to them.
Short answer, work the number
6. If the seeded test had caught 8 of 10 known-bad POs instead of 1, would this still count as a case of partial automation being worse than none?
Show hint
Compare 8 of 10 against the 90 percent manual-only catch rate from before the badge shipped.
Show answer
Model answer: Likely not as clearly. 8 of 10 is closer to the original 90 percent manual rate, so the badge wouldn't have caused a meaningful decline. The direct answer's concern is specifically about a real drop below what manual review alone achieved, not partial automation in general.
Before you close the answer
Why this works
Tests whether you can trace a slow, quiet failure back to a specific design choice, a badge with no stated scope, rather than blaming either the model's accuracy or the human's diligence, and whether you'd measure the real end-to-end outcome instead of trusting a faster surface metric.
Follow-up traps
"Isn't this really Tavish's fault for not reading carefully?" Response: a design that relies on sustained careful reading of an unlabeled badge, under real and growing workload, will fail this way eventually for almost anyone, that's a design gap, not a character flaw.
"Couldn't more training have fixed this without changing the badge?" Response: training helps at the moment it happens, but a scope-free badge keeps re-teaching the wrong lesson every day after, the label itself needs to carry the scope, not rely on memory holding indefinitely.
If pressed
The seeded-error test used ten POs deliberately chosen to be format-perfect but substantively wrong, isolating exactly the gap between what the pre-check covers and what a human assumes it covers, rather than testing the pre-check's own accuracy, which was never in question.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.