ConceptFoundationalAI Opportunity & Model Strategy / Opportunity identification for AI / #9
Explain the difference between automating a task and augmenting a person doing it.
PICK · Auristem's ad model went from flagging its unsure calls for a person to shipping every call alone, and the one bonus episode it was never sure about went out anyway
Auristem places the ads you hear inside podcast episodes: it finds the natural break, picks an ad that fits the topic, and checks whether the segment around it is safe for that advertiser's name. Melaine Osgerby owns the brand-safety and placement model. Hendrina Rutherfoord runs partner growth, and needed the long tail of smaller shows publishing ads as fast as a rival network that had just started shipping with no review at all. Melaine had to decide which parts of that job a model could own outright, and which parts still needed a person standing between the model's answer and an advertiser's name.
The direct answer
Automating means the model does the whole job and the person is gone. Augmenting means a person still makes the one call that needs real judgment, and the model hands them the fast, mechanical part already done. Keep a person in the loop wherever a wrong call could do real harm before anyone notices, like a toy ad landing next to explicit content. Let the model run alone wherever a wrong call is cheap and fast to catch, like picking the slightly less relevant of two already safe ads. At Auristem, brand safety on an unusual episode needed a person. Choosing between two safe ads never did.
Do this, in order
Keep a person in the loop wherever a wrong call is hard to undo.Why: this is the whole decision, and it's the exact gap that let the bonus episode ship blind.
Let the model run alone wherever a wrong call is cheap and fast to catch.Why: forcing a person to check something that was never going to need catching wastes the reason you built the model.
Set the line at how hard a mistake is to undo, not at how good the model scored this week.Why: a good model can still be unsure about a case it has barely seen, and the harm doesn't care what the average score says.
Keep the confidence threshold as a real gate, even inside a fast, automated flow.Why: a number under the line should send something to a person, whatever you call the flow around it.
Treat the mistake that slips through unseen as the expensive one, and pay to stop it, not to clean it up.Why: that's the real cost asymmetry, and it only points one direction.
Never let a speed target quietly delete a safety step instead of shrinking it.Why: nobody at Auristem meant to remove the check completely. That's just what "match the rival's speed" turned into once nobody wrote the line down.
How to answer this, stage by stage
Nobody in the room is grading whether you can define "automate" and "augment" like a dictionary. They're grading whether you can point at the one call in a real job that still needs a person, and defend it with a number.
1
Put a real product under the question
Say it like this
"Let's ground this. Auristem places ads inside podcast episodes: it finds the ad break, matches an ad to the topic, and checks the segment for brand safety. Melaine owns that last check. Her partner-growth lead, Hendrina, wants the whole long tail of shows publishing ads in under two minutes, to match a rival that's already doing it with zero review."
Why this works
Keeps the interviewer grading a real decision, not a dictionary definition.
2
Name the method before you use it
Say it like this
"I'll use PICK. Position, what actually separates automating from augmenting. Impact, what's lost on each side if you get it wrong. Cost asymmetry, which mistake is cheap and which one is expensive. Kill criteria, the one test for which fits here."
Why this works
Two seconds of structure tells the room a method is coming, not a mood.
3
Draw the actual line between the two words
Say it like this
"My position: automating means the model does the whole job and the person is out of it, completely. Augmenting means a person is still there making the one call that actually needs judgment, and the model just hands them the fast part, already done. Those are two different products wearing the same label, 'an AI feature.'"
Why this works
This is the direct answer, said before the story can blur it into "be careful with AI."
4
Bring the number that makes it real
Say it like this
"Auristem's old flow sent anything under ninety five percent brand-safety confidence to a person, about six percent of the long tail's fourteen thousand weekly episodes. The new flow sent nothing. It just shipped. Ten days after the switch, a Night Ledger bonus episode scored eighty nine percent confident 'clean,' which isn't very sure, and it went out anyway. It ran three days and about forty one thousand downloads before anyone caught it."
Why this works
A real number turns "this feels risky" into something the room can actually check.
5
Walk both directions before picking one
Say it like this
"Keep a person on everything, and you're paying someone to check ad picks that were never going to go wrong in a way that mattered, which wastes a perfectly good model. Remove the person from everything, and the one time the model's unsure, on a case it's barely seen, nothing stops it, and forty one thousand people hear the mistake before you do."
Why this works
Naming both losses stops the answer from collapsing into "always keep a human," which isn't actually a decision.
6
Weigh what each mistake actually costs
Say it like this
"Keeping the ninety five percent threshold gate on the long tail costs about a hundred and forty reviewer hours over six weeks. Cheap, and forgotten by the next sprint. Skipping it cost Auristem an advertiser's whole quarter, about a hundred and eighty thousand dollars pulled, plus four hundred and twenty hours reclassifying the entire backlog by hand to prove nothing else had slipped through. That's the asymmetry. It only points one way."
Why this works
This is the center of PICK: the two mistakes aren't the same size, and saying which one is bigger is what makes the call defensible.
7
Hand over the one test, not a feeling
Say it like this
"Here's my test: if a wrong output slipped through with nobody watching, would it do real harm that's hard to undo? For brand safety on an unusual episode, yes, so a person stays in the loop. For picking between two already safe ads, no, a wrong pick just underperforms a little and gets caught in next week's numbers, so the model runs alone. Same company, same model, two different answers, and the test is what tells you which."
Why this works
A kill test with no real check behind it is just an opinion wearing a framework's clothes.
Let's learn
Ten weeks ago, nothing published on Auristem's long tail without a number attached to it. A brand-safety score, and if that score dropped under ninety five percent, a person looked before the ad went live.
Four steps for every episode, and only the last one used to have anywhere to send a number nobody trusted yet.
Auristem listens to a podcast episode, picks an ad slot, matches an ad to the topic, and scores whether the segment around it is safe for that advertiser's name. Across the whole network, that's about forty thousand episodes a week.
The long tail, about nineteen hundred smaller shows and fourteen thousand episodes a week, used to work like this: anything the model scored ninety five percent confident or higher published straight away, in under two minutes. Anything under that line went into a queue, and a person looked before it aired, sometimes the same hour, sometimes not until the next morning if the queue backed up.
Then a rival ad network started shipping with no review anywhere, live in under two minutes on every episode, no exceptions. Three of Auristem's long-tail partner shows switched over inside a month, chasing the speed. Hendrina set a goal: match that on the whole long tail, this quarter. The fix that shipped didn't just speed up the queue. It removed the queue. Every long-tail episode now publishes in about ninety seconds, whatever the score says.
Same model either way. The only thing that changed is whether its "not sure" has anywhere left to go.
Here's the turn. The extra mistakes were never really the problem. Most episodes score well above ninety five percent, and the model is right about almost all of them. The real problem is what happens on the ones it isn't sure about, now that "not sure" doesn't send anything anywhere.
We didn't make the queue faster. We took away the one moment somebody got to say, wait, let me look at that first.
Knowledge spark: what does "eighty nine percent confident" actually mean here?
It's the model's own guess about its own guess. If it says ninety five percent a hundred times, it should be right about ninety five of those times. Eighty nine percent isn't a bad number. It's just not a sure one, and a sure number and an unsure number used to be treated very differently before the queue disappeared.
Night Ledger runs mostly as a straight true-crime interview show, tagged family safe for its regular episodes. A few times a year it puts out a bonus episode: darker, looser, with explicit language and blunt descriptions the regular show never uses. One of those bonus episodes uploaded on a Tuesday morning. The model scored it eighty nine percent confident clean, which isn't very sure. Under the old rule, that would have landed in someone's queue that morning. Under the new one, it just went out, with a Pinwheel Kids toy ad sitting at the top of it.
The choice I would take back
Auristem's rollout removed the human review step for the whole long tail, instead of keeping the ninety five percent threshold gate and only removing review for episodes that cleared it. That made sense in the planning meeting: building a fast path and a slow path into the same low-latency publishing pipeline looked like nearly double the engineering work, for a check that had barely caught anything on the shows everyone was watching that week. I would take that back and keep the gate, even inside the fast flow.
Two of these were always safe to automate. Two of them were exactly the ones that needed a person left in the room.
What I would leave alone: Auristem's ad relevance model, which of two already brand-safe ads best fits a listener, was fully automated on the long tail long before any of this, and it should stay that way. A wrong pick there just means a slightly less relevant ad played. It shows up in next week's numbers and gets swapped. Nobody needs to watch it happen.
The lesson: a model's confidence number can be honest on the cases it's seen a thousand times, and quietly miscalibrated on the rare ones it's barely seen at all. That's exactly why a threshold that sends "not sure" answers to a person exists in the first place. Auristem kept computing the number. It just stopped answering it.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for what it cost Auristem to learn it the slow way.
Every Monday morning, Melaine Osgerby pulled the same report: how many long-tail episodes had landed under the ninety five percent line the week before, and how many of those a reviewer had actually caught something on. Most weeks it was small, a handful out of thousands. She still read every row, because the handful was the whole point of the number.
She'd owned the brand-safety model for a year and a half, long enough to know which advertisers cared how much, and which shows ran hot on language without ever meaning any harm by it.
The rival showed up in a trade newsletter first, a small item about an ad network shipping podcast spots "instantly, no manual review, ever." Within a month, three of Auristem's smaller partner shows had quietly moved their inventory over, chasing the speed. Hendrina Rutherfoord, who ran partner growth, brought it to Melaine's team with a number of her own: match under two minutes, network wide, by the end of the quarter.
Two weeks between the rival's first mention and the day the queue disappeared for good. What actually broke arrived ten days after that.
The team built it fast, and mostly it built the right thing. Ad slotting got quicker. Matching got quicker. The one place where the plan and the calendar disagreed was the review queue, since a separate fast path and slow path inside the same new low-latency pipeline looked, on a whiteboard, like nearly double the work, for a check that had flagged a real problem on maybe one show in fifty over the past year. So the queue came out. Not on purpose, exactly. It just stopped being anyone's job to rebuild it.
For the first nine days after the switch, nothing happened. Every long-tail episode went live in under two minutes. Ad revenue on the long tail ticked up. Nobody complained.
Then Night Ledger uploaded a bonus episode at six in the morning, the kind it ran three or four times a year: looser, funnier in a way that could turn dark fast, with language and content warnings its regular episodes never carried. The model scored the whole episode eighty nine percent confident clean. Under the old rule, eighty nine would have landed in a queue by breakfast. Under the new one, there was no queue to land in. It published at six oh four, with a Pinwheel Kids ad, bright and cheerful, sitting right at the top.
Neither of them was wrong about what they were holding. The target was real. So was the number. The question was which one still had somewhere to go.
The model wasn't lying about being unsure. Eighty nine isn't a sure number. There just wasn't anyone left for it to tell.
By the time a parent posted a screenshot, three days had passed and about forty one thousand people had downloaded that episode. The screenshot showed the ad, mid roll, right before a blunt, ten-second content warning the show itself had recorded. It found its way into a parenting forum, then a couple of larger accounts, tagging Pinwheel Kids directly. Pinwheel Kids pulled its entire quarterly campaign that same week, not just from Night Ledger, from the whole network, about a hundred and eighty thousand dollars of spend. Auristem spent the next several weeks hand-reclassifying every episode in the long-tail backlog, just to be able to tell its remaining advertisers nothing else had slipped through.
I want to say the problem was that the model got something wrong. It did. But that's not really the story. Nobody at Auristem decided, on purpose, that unsure outputs should ship anyway. They decided that building two paths through one pipeline wasn't worth the sprint. The number kept getting computed. It just stopped landing anywhere.
Months earlier, in the actual planning meeting, an engineer raised the two-path idea and someone above him said, reasonably, that six percent didn't feel worth slowing the whole quarter's roadmap for. Nobody in that room was being careless. They were weighing a real cost against a risk that hadn't happened yet, which is exactly the kind of bet that looks fine right up until it doesn't.
Here's the replay. Same six weeks, same rival, same deadline. This time Melaine keeps the ninety five percent gate inside the fast pipeline, and pairs it with a small change: flagged episodes go to a reviewer's phone as a push alert instead of an email digest, so the wait drops from up to a day to under ten minutes on a normal morning. Ninety four percent of episodes still publish in ninety seconds, automatically. The Night Ledger bonus episode scores eighty nine percent again, gets flagged again, and this time a reviewer sees it in six minutes, catches the mismatch between a kids' toy ad and an explicit warning, and reroutes the ad before a single download happens.
One version of this story trades a queue for a network-wide incident. The other trades it for six minutes on one Tuesday morning nobody outside the review team ever hears about. Same model, same deadline pressure, same rival. The only thing that changed was whether "not sure" still had somewhere to go.
What I'd tell myself, watching that queue quietly disappear in a planning meeting nobody thought was a big decision: the sprint estimate for keeping a safety gate is not the same question as whether you can afford to lose it. Nobody ever asked the second one out loud.
PICK: who's still watching when the model is wrong
Not a rule about trusting AI less. PICK only earns its keep when it names the one decision that actually needs a person, and lets the model run free on everything else.
One path costs a scheduled hundred and forty hours. The other already sent Auristem its real bill, in a single week, from a single advertiser.
PPosition. The real distinction.
Automating removes the person completely. The model's answer is the final answer, and it ships on its own. Augmenting keeps a person as the one making the actual call, with the model doing the fast, mechanical setup underneath them.
This isn't a ladder where automating is always the more grown-up version of augmenting. They're two different designs, built for two different kinds of mistake.
State the position before any story, so it doesn't look reverse-engineered from what already went wrong.
IImpact. What's lost each way.
Automate something that actually needed judgment, and you remove the one check that would have caught it, and the mistake scales invisibly. At Auristem that meant forty one thousand downloads before anyone looked.
Augment something that was already safe to run alone, and you waste the model's whole point, paying a person to check work that was never going to need catching.
Naming both losses stops the answer from collapsing into "always keep a human," which isn't a real decision either.
CCost asymmetry. The heart of it.
Keeping the ninety five percent threshold gate on the long tail costs about a hundred and forty reviewer hours over six weeks. Cheap, and it's forgotten by the next sprint review. Skipping it cost Auristem a hundred and eighty thousand dollars of a single advertiser's quarter, plus four hundred and twenty hours reclassifying the whole backlog by hand afterward. Start from the cheap mistake. Only accept the expensive one once real evidence says it's actually safe to.
KKill criteria. The one test.
Does a wrong output, unreviewed, do real harm that's hard to undo. Melaine's team considered one shortcut instead of rebuilding the queue: score an episode's "unusualness" against the show's own history, and only route the unusual ones to a person, skipping the threshold check for everything that looked typical. It got rejected, because the unusualness score was just as unproven on rare episode types as the brand-safety score itself. Trusting one unverified number to decide whether to trust another one doesn't remove the risk. It just hides it a layer deeper.
Four branches, one question: if nobody's watching and it's wrong, how hard is that to undo.
Cost, by the numbers: keeping the review gate versus what skipping it actually cost
Cheap, paid on a scheduleExpensive, paid after the fact
Keeping a person on the six percent of long-tail episodes that scored under ninety five percent cost about a hundred and forty hours over six weeks, work Auristem could plan around. Skipping it cost about four hundred and twenty hours of reclassifying and account repair, on top of the hundred and eighty thousand dollars Pinwheel Kids pulled.
The kill line, charted: confidence-versus-accuracy gap on bonus episodes, during recalibration
Above the kill lineCleared the kill line
The team tracked bonus and one-off episode types specifically, not the network average, because the average was already sitting near ninety six percent confident and would have looked fine on a dashboard the whole time.
The trade worth saying out loud: keeping a person in that loop costs real minutes on six percent of the long tail, every single week, forever, not just during a rollout. That's slower, and it's worth paying, because the alternative, a confident-sounding model shipping to nobody's supervision on a case it's barely seen, only shows its real cost after thousands of people have already heard it.
And if you want to be sure it really works, try it somewhere else
Same four letters, a mail-order pharmacy instead of a podcast network, and the missing check is a self-reported pill, not a bonus episode.
Cobbleworth Pharmacy Group runs a refill-approval model across its mail-order arm: check the new refill against the patient's history, flag anything with a possible drug interaction, and clear the rest. After a staffing shortage left refills backing up for days, someone proposed auto-approving any refill the model scored as a "routine, no new interaction" case, no pharmacist look, to clear the backlog fast.
Position: automating removes the pharmacist from approving a refill entirely, once it clears the "routine" bar. Augmenting keeps the pharmacist approving every refill, with the interaction check already run and attached, so they're deciding, not searching. Impact: automate a refill that actually needed a look, say a patient who self-reported a new over-the-counter drug that weakly interacts with what they're already on, a case the model had barely seen labeled either way, and a real interaction ships with nobody watching. Augment a refill that's identical to last month's, same drug, same dose, nothing new reported, and you keep a pharmacist doing work a model could safely own, which is exactly the backlog they were trying to clear. Cost asymmetry: keeping a pharmacist's look on anything with a new self-reported drug costs a few extra minutes per case, cheap, the kind of minutes a pharmacy budgets for anyway. An interaction that ships unreviewed costs a real adverse event, and there's no version of that bill that's small. Kill criteria: does a wrong approval, unreviewed, risk real harm that's hard to undo. If the refill has anything new attached to it, a pharmacist stays in the loop. If it's identical to last month's, the model clears it alone.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the one test, unreviewed and hard to undo, before anything else.
Cost: no time to build a full fast path before a real deadline. Fine, but keep the threshold gate on the riskiest slice first, brand safety or new drug interactions, not a faster path for the easy cases that were never the risk.
The model got better, for real: say a new version scores near perfect on average. Still don't drop the gate. A good average can hide one rare case nobody separated out.
Where people run it wrong.
They automate a whole category because most of it is safe, instead of automating the safe part and keeping the risky sliver augmented.
They keep a person on everything out of habit, long after the model has earned the easy calls.
They cut the safety gate to hit a speed target without ever deciding, out loud, that they were cutting it.
How to use it live. If you're ever asked whether something should run on its own, buy yourself a second by asking one plain question out loud: "if this is wrong and nobody looks, how hard is that to undo." That question is the whole method, asked instead of stated.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question about automating a task versus augmenting a person?
Tap to flip
ANSWER
PICK: state the real line between the two, name what's lost on each side, find which mistake is cheap versus expensive, then give the one test for which one fits.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Melaine Osgerby, who owns the brand-safety and ad-placement model at Auristem, an AI tool that places ads inside podcast episodes.
3 · THE POSITION
What actually separates automating from augmenting?
Tap to flip
ANSWER
Automating removes the person completely, the model's answer ships on its own. Augmenting keeps a person making the real call, with the model handing them the fast, mechanical part already done.
4 · THE GAP
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
The old flow sent anything under ninety five percent confidence to a person. The new flow sent nothing. A bonus episode scored eighty nine percent, not very sure, and shipped anyway.
5 · THE REVERSAL
What old decision would Melaine take back?
Tap to flip
ANSWER
Auristem removed the whole human-review step for the long tail instead of keeping the ninety five percent threshold gate inside the new fast pipeline. It made sense when a separate fast and slow path looked like double the engineering work for a check that rarely caught anything.
6 · THE NUMBER
Fill in the blank: the bonus episode scored ___ percent confident clean, ran ___ days across about ___ downloads, and the advertiser pulled about $___ in spend.
Tap to flip
ANSWER
89 percent. 3 days. About 41,000 downloads. About $180,000 in Pinwheel Kids' pulled quarterly campaign.
7 · THE KILL TEST
What's the one test that tells you whether to automate or augment?
Tap to flip
ANSWER
If a wrong output slipped through with nobody watching, would it do real harm that's hard to undo. Yes means keep a person in the loop. No, if it's cheap and fast to catch, means let the model run alone.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of Auristem's bonus episode there?
Tap to flip
ANSWER
Cobbleworth Pharmacy Group's refill-approval model. The role goes to a refill with a newly self-reported over-the-counter drug, a case the model has barely seen labeled either way, standing in for Night Ledger's rare bonus episode.
Check yourself Score: 0 / 0
True or false
1. True or false: since Auristem's ad-relevance matching had already run fully automated for months with no complaints, brand-safety scoring was just as safe to automate the same way.
True
False
Show hint
Think about what makes a wrong relevance pick different from a wrong brand-safety call.
Show answer
False. Relevance matching and brand safety are different kinds of task. A wrong relevance pick just underperforms a little and gets caught in next week's numbers. A wrong brand-safety call can put a family brand's name next to explicit content before anyone looks. Being safe to automate in one place says nothing about the other.
Multiple choice
2. What made the Night Ledger incident possible?
A. The brand-safety model had never been trained on any true-crime content.
B. The threshold that used to send unsure scores to a person had been removed for the whole long tail, so an 89 percent score published without anyone seeing it.
C. Pinwheel Kids asked to run its ad next to mature content on purpose.
D. Auristem's ad-relevance matching picked the wrong ad category.
Show hint
Check what changed about the review queue before the bonus episode ever uploaded.
Show answer
B. Nothing was wrong with the ad pick or the training data on its own. The gap that mattered was that a low-confidence score had nowhere left to go once the human check was removed.
Fill in the blank
3. The bonus episode scored ___ percent confident "clean." It ran for ___ days and about ___ downloads before a parent's screenshot caught it.
Show hint
Look at the numbers in the direct answer's stage-by-stage walkthrough, stage 4.
Show answer
89 percent. 3 days. About 41,000 downloads. The gap between those numbers is exactly why an unreviewed judgment task scales a mistake invisibly, instead of stopping it early and small.
Short answer, name the rejected alternative
4. What old decision would Melaine take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back" in the Let's learn section.
Show answer
Model answer: Auristem removed the entire human-review step for the long tail instead of keeping the ninety five percent confidence gate inside the new fast pipeline. It made sense in the planning meeting because building a separate fast and slow path looked like nearly double the engineering work, for a check that had flagged a real problem on maybe one show in fifty over the past year.
Short answer, where it wouldn't matter
5. Name a place in Auristem's own system where this exact scrutiny would NOT be needed, and say why.
Show hint
Look at "what I would leave alone" in the Let's learn section.
Show answer
Model answer: Ad-relevance matching, picking which of two already brand-safe ads best fits a listener, was already fully automated and should stay that way. A wrong pick there just means a slightly less relevant ad played, caught in next week's numbers and swapped. Nothing about it can put an advertiser's name somewhere it shouldn't be.
Short answer, apply it yourself
6. Think of an AI feature you've used yourself that makes a decision without asking you first. What would need to be true about a wrong decision there for you to want a person back in the loop?
Show hint
Think about whether a wrong call there is cheap and fast to catch, or hard to undo.
Show answer
Model answer: A photo app auto-deletes what it thinks are duplicate or blurry photos from a phone's camera roll. That's fine, a wrongly deleted photo is annoying but usually recoverable. If the same app auto-deleted photos it thought were sensitive instead, a wrong call there could permanently lose something that mattered, and I'd want to approve that one myself first.
Before you close the answer
Why this works
Tests whether you understand that "automate or augment" is really a question about which mistakes a person can absorb and which ones scale invisibly. Most candidates answer with generic advice about trusting AI less. The real judgment is knowing the same model can be perfectly safe to run alone on one part of a job and dangerous to run alone on another part of the exact same job.
Follow-up traps
"Isn't keeping a person in the loop just slower and more expensive, always?" Response: only on the part of the job where a mistake is cheap to catch. On the part where a mistake is hard to undo, the real expensive path is the one that skips review, about a hundred and eighty thousand dollars and four hundred and twenty hours worth, in this case.
"What if the classifier gets good enough that 89 percent is basically never wrong?" Response: then the fix is proving that on bonus and edge-case episodes specifically, not on the average. An average sitting near ninety six percent hid a gap that only showed up on the rare episode type nobody had separately checked.
If pressed
The brand-safety score was computed from only the first ninety seconds or so of an episode's audio, a shortcut that kept scoring fast across forty thousand weekly episodes. Night Ledger's bonus episodes open with the same chatty, ordinary tone as its regular show and only turn explicit a few minutes in, so the model was reading exactly the part of the audio that looked most like every other episode it had ever scored well.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.