How does the human-in-the-loop design change for irreversible actions?
GUARD the product is Duotrack, Skerrymist Broadcasting's live AI captioning system
Skerrymist Broadcasting runs Duotrack, a system that turns a live satellite news feed into captions and dubbed audio, in forty seven languages at once, seconds after someone speaks. Ines Calloway reviews Duotrack's output from the control room, six years into the job.
The direct answer
Move the human check from after the action to before it, but only for the slice of actions nobody can undo once they leave the building. A caption that has already reached air cannot be recalled, so that one needs a real, active human confirm before it airs, not a sample check after the fact. Everything reversible, a routine caption a viewer can see corrected seconds later, keeps the lighter review it already had.
Do this, in order
Put the human check before the irreversible step, not after.Why: once a caption airs, it cannot be pulled back, so the check has to happen on the way out the door, not once it's already gone.
Draw the actual line between reversible and irreversible for this product, in writing.Why: without a named line, the team finds it by accident, usually during an incident review.
Give the reviewer a real pause, not a countdown they can beat by clicking fast.Why: a timer that runs whether or not anyone looked isn't a check, it's a delay wearing a check's clothes.
Track how often that pause gets skipped or rushed through.Why: a confirm step that's always tapped through on autopilot has quietly stopped catching anything.
Log what a person confirmed and when, for every irreversible action.Why: a mistake needs to trace back to a decision someone made, not a default nobody remembers choosing.
Leave the lighter, after-the-fact sample review in place for reversible, low-stakes captions.Why: forcing routine weather and sports captions through the same hard gate slows down the vast majority of the feed that carries no real risk.
How to answer this, stage by stage
Nobody is grading whether you know the word "irreversible." They're grading whether your review design actually changes shape once it applies.
Stage 1
Scope it to one real feature
Say it like this
"I'll answer this for Duotrack, Skerrymist's live AI captioning system, which pushes captions and dubbed audio to air in forty seven languages during breaking news."
Why this works
Gives the interviewer something concrete to push on, instead of "AI systems" in the abstract.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, who's affected. Unequal, where the harm lands unevenly. Ability to contest, who never gets a lever. Reduce, the actual design change. Detect, how I'd know."
Why this works
Signals a method, so the answer doesn't turn into a loose list of worries.
Stage 3
Reframe the question
Say it like this
"This isn't really about how much review to add. It's about whether review happens before the action or after it, and that only starts to matter once the action can't be undone."
Why this works
Separates a real answer from the generic "add more review" answer this question is built to catch.
Stage 4
Name who can't push back
Say it like this
"The person named wrongly by a caption has no way to correct the record before millions of people already saw it. Ines can catch it before it airs. The person named in it can't do anything once it has."
Why this works
This is GUARD's hardest step, and the one most answers skip entirely.
Stage 5
Give the one decision
Say it like this
"For anything irreversible, I'd require an explicit human tap before it airs, not a countdown. For anything reversible, I'd keep the lighter, after-the-fact sample review it already has."
Why this works
This is the direct answer, stated as one concrete design rule, not a philosophy.
Stage 6
Prove it with the near miss
Say it like this
"A caption once misattributed a quote to a named politician, and the catch came with about a tenth of a second left on a four-second clock. That's not a design working. That's luck, dressed up as a design."
Why this works
Shows the current design's real edge, compressed to a few sentences instead of a whole case study.
Stage 7
Close on the one line
Say it like this
"How much review a feature needs isn't the real design question. Whether it happens before the thing becomes permanent, or after, is the question, and the honest answer flips completely once an action can't be taken back."
Why this works
Restates the direct answer in one breath, ready for whatever the interviewer pushes on next.
Let's learn
Here is what happens when a feature is fast and right nearly every time, and the one time it isn't, nobody can take it back.
Duotrack is Skerrymist Broadcasting's live captioning system. It listens to a raw satellite feed and pushes captions or dubbed audio to air, in forty seven languages at once, seconds after someone speaks.
Ines holds a lever. The person a caption names, wrongly or not, holds nothing at all.
Before Duotrack, a human interpreter sat on every language desk, translating live and reading captions onto air through a fixed eight-second broadcast delay. That covered six languages a night, reliably, but slowly.
Now Duotrack covers forty seven languages, and pushes captions to air in under two seconds on routine segments, getting the words right close to every time.
The occasional wrong word was never the real danger. The real danger is a caption that names a real person and gets it wrong, once, live, in front of millions, with no way to call it back.
Named-person captions with an explicit human pause, before and after the redesign
Before the fix, the pause was a formality eight times out of nine. After, it's an actual decision every time.
At its worst: a caption falsely quotes a named public figure during a breaking story, it reaches air before anyone reacts, and there is no broadcast version of deleting a bad post.
The decision I would take back
We merged caption generation and the broadcast push into a single four-second countdown timer, so review became "catch it before the clock runs out" instead of an actual human decision point. That made sense when the model rarely touched a named individual at all. It stopped making sense the day it did.
What I would leave alone: routine, non-attributed captions, weather, sports scores, general commentary, can keep shipping on the lighter, after-the-fact sample review. Forcing those through a hard stop would slow down the vast majority of the feed that carries no real risk at all.
The lesson: how much review a feature needs isn't really the design question. Whether that review happens before the action becomes permanent, or after, is the actual question, and the honest answer changes completely the moment the action can't be undone.
Now here is the same thing as a story
The short version above is what you'd say defending this design to Skerrymist's standards desk. Read this one for how close the actual near miss came.
Ines Calloway has reviewed Duotrack's live output from Skerrymist's control room for six years. She can catch a mistranslated idiom before it finishes appearing on screen, an instinct she built watching two earlier systems fail before this one.
Five parts. Only one of them used to require a person to actually decide anything.
For the better part of a year, Duotrack had been steady. Fast, mostly right, and on the rare miss, small: a wrong verb tense, a place name slightly off. Nothing that reached air ever needed an on-screen correction.
Then, on a Tuesday breaking-news segment out of the capital, Duotrack's caption attributed an inflammatory line to a sitting member of parliament who had, in the raw feed, said close to the opposite thing two sentences earlier.
Knowledge spark: what makes a caption like this more dangerous than a wrong verb tense?
A wrong verb tense is obviously small. A caption that puts different words in a named person's mouth looks exactly as confident as a correct one, and once it airs, there's no way to un-say it to the people who already heard it.
The four-second buffer had already started counting down when Ines caught the mismatch, mid-glance at a confidence trend she'd started watching out of habit rather than instinct.
The buffer was four seconds long. She pulled it at 3.9.
She pulled it from the queue with about a tenth of a second left on the clock. Nobody outside the control room ever knew how close it came.
Nothing about that caption ever looked uncertain on Ines's screen. The clock was the only thing that was actually running out.
An internal review after the near miss found the same auto-push, no real human confirm required, had let nine other named-person captions through unreviewed that same month. None as damaging. All of them the kind of thing that could have been.
The missing box was never drawn on any diagram. That's usually how a gap like this survives a launch review.
With the redesigned gate, any caption naming a real person requires Ines to actively tap a confirm, not just let a clock run out. Run the same Tuesday forward: Duotrack still drafts the same wrong attribution, but it never reaches air until Ines actively signs off on it, and this time she doesn't, because the pause gives her the seconds she actually needs to notice the mismatch instead of racing a clock.
The old design asked Duotrack to always make its own deadline. The new one asks a person to actually decide, every single time, for the captions that can't be taken back.
I built the countdown because a hard stop on every caption felt like it would slow the whole broadcast down. It took a member of parliament almost getting misquoted live to see that the four-second clock was never actually a check. It was a countdown to whether Ines happened to be looking at the right moment.
GUARD, in one screenNot a policy document. GUARD is what tells you why the review design itself has to change once an action can't be undone.
G
Groups. Who is affected.
The operator, Ines, reviewing captions in the control room. The subject, whoever a caption names, on or off screen, who has no seat in that room at all.
Names both people the design has to answer for, not just the one Skerrymist employs.
U
Unequal. Where the harm lands unevenly.
A routine caption error costs a viewer a moment of confusion, corrected on screen seconds later. A named-person error costs that person their reputation, live, with no on-screen correction that reaches everyone who saw the original.
Shows the harm isn't spread evenly across every kind of caption, which is exactly why one review rule can't cover all of them.
A
Ability to contest. Who never gets a lever.
The person named wrongly in a caption gets no warning before it airs and no way to stop it once it does. The review, as built, protected Skerrymist from an error. It never protected the person the error was actually about.
The hardest step, and the direct answer: irreversibility is exactly the condition where this question stops being theoretical.
R
Reduce. The specific design change.
Require an explicit, active confirm tap for any caption naming a real individual, before it reaches air, replacing a countdown that ran whether or not anyone was looking.
A concrete product change, not a review board or a policy memo.
D
Detect. How you'd know in production.
Track how often the confirm step is skipped, rushed, or waved through in under a second, since a confirm that's always instant has quietly stopped being a check.
Catches the moment the fix itself becomes a formality, before an incident does.
The most common caption problem is also the least dangerous one. The rarest is the one that needed the gate.
Three things Duotrack is deliberately not trusted to decide by itself, even now.
The recap, one line per letter: groups is the reviewer and the person a caption names, unequal is a routine error against a reputation with no way back, ability to contest is the subject who never got a lever, reduce is an explicit confirm replacing a blind countdown, and detect is watching for the day that confirm itself goes quiet.
And if you want to be sure it really works, try it somewhere elseSame five letters, a parole risk score instead of a live caption. A completely different feature, the same irreversible edge.
Quarryfield Parole Board uses an AI risk-assessment score to help decide who gets released early. Talia Munro is a caseworker who reviews the score before it reaches the board's decision meeting each week.
Mapped onto GUARD: groups is Talia, who can flag a score she doesn't trust, and the person up for parole, who used to never see the score or the factors behind it at all. Unequal is that a wrong score for a low-risk case costs someone months of freedom, while a wrong score in the other direction costs the board a rare, visible mistake, two harms that don't land on the same person or at the same weight. Ability to contest is the actual design gap: for years, nobody up for parole could see which factors drove their own score, let alone dispute one that was wrong. Reduce is showing the named factors behind every score to the person it's about, before the hearing, with a real channel to flag one as inaccurate. Detect is tracking how often a flagged factor gets corrected on review, which tells the board whether the score is actually earning trust or just being rubber-stamped.
Caseworker override rate of the parole risk score, after factors became visible and contestable
The override rate didn't fall once contesting a score became possible, it rose, because the flaws were always there, just invisible to the one person they mattered most to.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "for actions nobody can undo, put the human check before it happens, not after," and stop.
Cost: there's no budget to build a full confirm gate this quarter. Say so honestly, and ship the visible-factors change first, since it costs almost nothing and it's the part that gives the subject a lever at all.
The model gets better, for real: if the parole risk score becomes noticeably more accurate, that's still not a reason to remove the contest channel, an accurate score that nobody can question is still a score nobody can catch when it's wrong about the one case that matters.
Where people run it wrong.
They design the review around protecting the company from a mistake, and never ask who the mistake actually happens to.
They treat "irreversible" as a single bucket, when a low-stakes irreversible action can often still ride the lighter path.
They wait for a lawsuit or a headline to find the gap, instead of asking upfront which action, if wrong, has genuinely no way back.
How to use it live. When someone hands you a feature and asks how the review should change for the riskiest actions, ask one question first: once this happens, can the person it happened to ever actually undo it. If the answer is no, that action gets its own gate, not a bigger version of the one everything else already has.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "how does human review change for X" design question?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. The ability-to-contest step is what makes irreversible actions different from reversible ones.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ines Calloway, a control-room reviewer at Skerrymist Broadcasting who has watched Duotrack's live captions for six years.
3 · THE OLD DESIGN'S GAP
What did the countdown timer actually let happen?
Tap to flip
ANSWER
A caption could reach millions of viewers without any human decision actually being required, since the four-second buffer counted down whether or not a person was looking.
4 · THE UNEQUAL SPLIT
Who can push back when a caption is wrong, and who can't?
Tap to flip
ANSWER
Ines can catch it before it airs. The person named wrongly in the caption has no lever at all once it's live.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging caption generation and the broadcast push into one countdown timer, so review became racing a clock instead of an actual human decision.
6 · THE NUMBER
Fill in the blank: before the redesign, only ___ of every 9 weekly named-person captions got an explicit human pause.
Tap to flip
ANSWER
1 of 9. After the redesign, all 9 require an explicit tap, every time, regardless of how routine the shift has been.
7 · THE REPLAY
Same Tuesday, redesigned gate. What changes?
Tap to flip
ANSWER
Duotrack still drafts the same wrong attribution, but it can't reach air until Ines actively confirms it, and the pause gives her the seconds she needs to catch the mismatch instead of racing the clock.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
Quarryfield Parole Board's AI risk score. Same GUARD shape: the person up for parole couldn't see or contest the score's factors until that ability was built in.
Check yourself Score: 0 / 0
Multiple choice
1. Why does an irreversible action need a human check before it happens, rather than a sample check after?
A. Because AI models are always less accurate on irreversible actions.
B. Because once the action is out, there's no way to undo it, so the check has to happen while it can still be stopped.
C. Because irreversible actions always happen less often, so review is cheaper to add upfront.
D. Because interviewers expect every answer to mention a review gate.
Show hint
Look at the direct answer.
Show answer
B. A check after an irreversible action is a postmortem, not a review. It has to sit before the thing becomes permanent.
True or false
2. True or false: the most common kind of caption error Duotrack made was also the most dangerous one.
True
False
Show hint
Look at the quadrant diagram.
Show answer
False. Routine wording issues were the most common and the least dangerous. Named-person misattribution was the rarest and the most dangerous, which is exactly why it needed its own gate.
Fill in the blank
3. Fill in the blank: Ines pulled the near-miss caption from the queue with about ___ of a second left on the four-second buffer clock.
Show hint
Look at the timeline diagram.
Show answer
A tenth. The buffer ran to 4.0 seconds; she pulled it at 3.9. That margin is luck, not a working design.
Short answer, where it wouldn't matter
4. Name a case in Duotrack where this exact redesign would not need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine weather or sports caption with no named individual in it. It's low-stakes and correctable on screen seconds later, so it can stay on the lighter, sample-based review.
Short answer, apply it yourself
5. Pick a product you use yourself. Name one action it lets you take that can't be undone, and say whether it makes you confirm before or just tells you after.
Show hint
Think of sending an email, posting publicly, or an automatic payment.
Show answer
Model answer: Many messaging apps let you unsend within a short window, which is itself a quiet admission that the send action needed a second chance. A public social post usually gets no such window at all.
Short answer, name the reversal
6. What old design decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Merging caption generation and the broadcast push into one countdown timer. It made sense while the model rarely touched a named individual at all, and stopped making sense the day it did.
Before you close the answer
Why this works
Tests whether you'll actually change the review's timing, not just its intensity, once an action can't be undone. Most candidates just say "add more review."
Follow-up traps
"Isn't a four-second buffer already a form of human-in-the-loop review?" Response: only if a person is required to act inside it. A countdown that fires by default whether or not anyone looked is a delay, not a review.
"What about actions that are irreversible but genuinely low stakes?" Response: the real line isn't reversible-versus-not alone, it's reversible-or-low-harm versus irreversible-and-high-harm. A low-stakes irreversible action, like a routine on-screen correction, can often still ride the lighter path.
If pressed
Duotrack's actual named-individual detector runs as a separate, lightweight classifier ahead of the main translation model, purely to decide which captions need the hard gate, since running a full check on every caption would blow the two-second latency budget the routine feed depends on.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.