The direct answer
Before launch, write the rollback trigger as a number, not a mood. For every metric that matters, name the exact value that ends the rollout, something like "if the policy-exceeding reply rate goes above 3 per 1,000 accepted drafts in any rolling 24 hours, the rollout ends," and put one person's name on who can act on that number the second it's crossed, with no approval needed. A criterion with no number and no owner isn't a rollback plan. It's a hope that someone notices in time.
Do this, in order
Write a named numeric threshold for every metric that matters, before launch, not "we'll keep an eye on it."Why: this is the one practice the whole rollback plan turns on.
Name, in writing, exactly who can act on that number, and take away the need to ask anyone first.Why: a threshold nobody's allowed to pull alone is a suggestion, not a trigger.
Write a clean path for both groups a rollback touches: the agents already mid-conversation with it on, and the agents still waiting for their rollout wave.Why: rollback isn't one switch, both groups need to know what happens to them the moment it fires.
Set the threshold on the metric that drifts quietly, not just the one that would spike loudly.Why: a rate creeping from 1.6 to 7.2 over three weeks never once looks like an emergency on any single day.
After launch, check whether the threshold itself ever fired on noise, or missed a real incident that stayed under it.Why: the number you wrote before launch is a guess. You only find out if it was the right guess afterward.
Leave low-stakes surfaces, like a sign-off line suggestion, without their own threshold.Why: writing a rollback number for something that can't cause harm just slows down the launches that actually need one.
How to answer this, stage by stage
Eight moves. Naming who can't act without a written number, and then actually writing the number, are where the real answer lives.
1
Ground it in one real product before naming a framework
Say it like this
"Say Verlane, a customer-messaging platform, builds a tool that drafts a reply for a support agent the moment a customer message comes in. The agent can accept it in one click, edit it, or type their own. It's about to go from a 5 percent internal alpha to a 25 percent rollout across real client brands."
Why this works
Grounds a broad question in one specific product before naming a method, so the answer can't drift into a generic checklist.
2
State your structure in one line
Say it like this
"I'd use GUARD here, because 'what rollback criteria would you set' is really asking who a launch decision protects if things go wrong, and who's stuck waiting on someone else's judgment call. Who's affected, where a missing number hurts worst, who can't act without one, the actual number and owner I'd write down, and how I'd know later if that number was even the right one."
Why this works
Two seconds that show you have a plan before you say a single specific thing.
3
Name both groups, not just "the users"
Say it like this
"There are two groups a rollback actually touches. The agents and customers already on the 25 percent rollout, mid-conversation, using AI-drafted replies right now. And the other 75 percent, still waiting for their rollout wave, who inherit whatever gets decided about whether that wave still happens."
Why this works
This is GUARD's G step. The same rollback decision gives the two groups two completely different questions to answer.
4
Show where the missing number lands hardest
Say it like this
"The metric that matters here is the rate of AI-drafted replies that promise something outside policy, a refund, an extension, that an agent accepted without catching. Week one, that's 1.6 per 1,000 accepted replies. Week two, 2.4. Week three, 4.8. Nobody had written down what number should end the rollout, so nothing about that climb looked like an emergency on any single day."
Why this works
This is U. It names the actual rate that drifted, not "the model might make risky promises" in the abstract.
5
Name who can't act without a pre-agreed trigger
Say it like this
"Rasmus Kjeld is the on-call engineer who actually sees the week-three number at two in the morning. He thinks it's bad. But 'criteria' was never written down as a real number with his name on it, so the most he can do is post it in a Slack channel and wait for someone with the authority to agree. That takes until Thursday morning."
Why this works
This is A, GUARD's hardest step, and the one a rollback-criteria question usually skips.
6
Give the number, not a general "watch it closely"
Say it like this
"Here's what I'd write before launch, not after: for every metric that matters, a real number. Something like, 'if the policy-exceeding reply rate goes above 3 per 1,000 accepted drafts in any rolling 24 hours, the rollout ends.' Not 'we'll keep an eye on quality.' A number."
Why this works
This is the first half of R. A threshold you could point to on launch day, not a mood everyone agreed to feel.
7
Give the number an owner who doesn't have to ask
Say it like this
"And the number needs an owner who doesn't have to ask. Whoever's on call, Rasmus included, gets pre-agreed authority to flip the feature flag back to manual replies the second that number's crossed. No meeting, no waiting for Thursday."
Why this works
This is the second half of R, and it's the part most rollback plans skip. A number nobody's allowed to act on alone isn't a trigger.
8
Say how you'd know the criteria itself was wrong, then close
Say it like this
"After it's run for a while, I'd check the threshold two ways. How many times did it trigger a rollback that turned out, on review, to be noise, that's too strict. And how many real incidents happened while the number stayed under it, that's too loose. So: write the number, name who can act on it without asking, and check afterward whether that number was actually the right one to write."
Why this works
Closes on the direct answer in one breath, and shows the criteria itself is a guess you go back and grade, not a document you file and forget.
Let's learn
The tool is a draft reply that appears the moment a support agent opens a customer message, before the agent has typed a word.
This is Verlane's auto-reply suggestion tool, built for the support teams of the e-commerce brands that run their customer messaging through Verlane. The agent can accept the draft in one click, edit it, or ignore it and write their own.
Before this tool, an agent typed every reply from scratch. That took about 6 minutes a ticket, and a full shift handled roughly 70 tickets.
Ticket handle time, before and after the suggestion tool
Same agents, same client brands. The only thing that changed was who typed the first draft.
Before: agent types every reply
6.0 min
After: agent accepts, edits, or rewrites a draft
3.5 min
Faster handle time is the whole business case. That part of the tool works exactly as sold.
Knowledge spark: what's a phased rollout?
Turning a feature on for a small slice of users first, watching it, then turning it on for more. Instead of everyone getting it on day one, it goes to 5 percent, then 25 percent, then everyone, so a problem shows up small before it shows up everywhere.
By week three of the 25 percent rollout, the drafts were mostly fine. But some drafts promised things outside Verlane's client policy: an instant refund above the $25 limit, or a two-week subscription extension when policy caps a goodwill extension at 3 days. Under ticket pressure, agents sometimes accepted a confident-sounding draft without checking it against policy, and Verlane's client brands were then stuck honoring the promise, or awkwardly walking it back with a real customer who'd already been told yes.
Policy-exceeding reply rate, by rollout week
Per 1,000 accepted drafts. The dashed line is the number nobody had actually written down.
Week 3 (Rasmus flags it, 2am)
4.8
Week 4 (rollout finally paused, Thursday)
7.2
The rate crossed 3 per 1,000 partway through week two. Nothing happened, because no one had written "3" down as the number that ends the rollout, or said whose job it was to act on it.
We didn't lose control of the rollout because the model got worse. We lost it because nobody had written down the number that was supposed to stop it.
At its worst, this costs a few hundred customers across Verlane's client brands a promise the agent never should have made, some in writing, that a client's support team then has to honor at a loss or retract from a customer who already believed it. One client brand nearly pulled its contract over it.
The choice I would take back
The rollout plan said "we'll keep an eye on quality metrics and pull back if something looks off," instead of a written number with a name attached to who could act on it. That was fine at the 5 percent internal alpha, where the whole team read every draft by hand. It stopped being fine the day real client brands and real customers were in it, and "keep an eye on it" had no owner and no number to actually trigger on.
What I would leave alone. The tool's suggestions for a closing line, "Thanks for reaching out, have a great day!", genuinely can't create a policy-exceeding promise no matter how much traffic they see. That surface doesn't need its own threshold. Save the sharp numbers for the drafts that can actually cost someone money.
The lesson. "Watch it closely" isn't a plan. It's a feeling with no owner and no number, and by the time a feeling turns into an actual decision, the number it should have caught is already three times worse.
Now here is the same thing as a story
Read the short version above if you're pressed for time. Read this one when you want to feel why a written number matters, not just know that it does.
Btissam Amrani has run product for Verlane's support-agent tools team for three years. She's the one people ask to turn a vague "make agents faster" request into a feature real client brands will actually trust with their customers.
When the auto-reply suggestion tool was ready to leave its internal alpha, Btissam wrote a launch plan that read, in the risk section: "we'll keep an eye on quality metrics and pull back if something looks off." At 5 percent, tested only on Verlane's own support inbox, that line was almost redundant. Every draft got read by someone on her team before it shipped.
The 25 percent rollout went out on a Monday. For the first two weeks, the numbers looked fine in the weekly review. A few odd drafts here and there, nothing that looked like a pattern. Her analytics lead called the launch "clean" in a standup.
Then, on a Wednesday night in week three, Rasmus Kjeld, the on-call engineer, was going through the backend policy-checker that flags accepted drafts against Verlane's written refund and extension rules, a tool that already existed for a different reason and had never been wired to anything urgent. He wasn't looking for a problem. He was closing out a routine dashboard pass before handing off at shift change.
The number for that week was 4.8 policy-exceeding drafts per 1,000, more than triple where it had started.
Same rising number. Only one of them was ever allowed to act on it.
Rasmus posted it in the on-call channel at 2:14am. He thought it was bad. He didn't know if it was "rollback bad," because nobody had ever told him what number counted as rollback bad, or told him he was allowed to be the one who decided that. So he wrote a message, tagged Btissam, and went home.
Btissam saw it at 8am. She agreed it looked wrong, but wanted her analytics lead to confirm the trend before recommending anything to the VP who'd signed off on the rollout. The VP wanted fifteen minutes on his calendar, which didn't open until Thursday morning. By the time the rollout actually paused, the rate had reached 7.2 per 1,000, and two more days of client complaints had come in.
We didn't quietly lose a little accuracy. We quietly built a rollback plan with no number in it, and then acted surprised when nobody could tell us it was time to use it.
I want to say the problem is that Btissam wrote a bad launch plan. She didn't, not really. She wrote the plan everyone writes: watch it, and act if it looks off. The trouble is that "looks off" isn't a number, and it doesn't come with anyone's name on it. Rasmus had a feeling on Wednesday night. A feeling isn't authority.
Back when the plan got signed off, in a meeting with her analytics lead, the question on the table was simple: is the alpha ready for a bigger slice of users. The answer, sensibly, was yes, the alpha numbers were clean. Nobody in that room asked a second question: if this goes wrong at 25 percent, what number ends it, and who's allowed to pull it.
I would go back and put that second question in the room. Btissam writes the threshold and the owner into the launch plan itself, the same day the rollout percentage gets decided. Rasmus sees the number cross 3 per 1,000 sometime Tuesday of week two, and he doesn't post a message and go home. He flips the flag himself, because the plan already said he could.
One version of the plan tells you to watch the metric. The other tells you exactly what to do the moment it moves.
What I'd tell the Btissam who wrote that first launch plan: "we'll keep an eye on it" and "we've decided what ends it" read like the same sentence in a review meeting, and I only ever wrote the first one.
GUARD, run against a rollout with no number written down
This is a risk question, so the framework is GUARD. "What rollback criteria would you set" sounds like it's asking for a checklist item, which is exactly why it's easy to answer with "we'll monitor closely" instead of an actual number.
G, groups. The agents and customers already on Verlane's 25 percent rollout, using AI-drafted replies mid-conversation right now. And the other 75 percent, still waiting for their rollout wave, who inherit whatever gets decided about whether that wave still happens.
U, unequal. The harm doesn't land as one bad reply. It lands as a rate that quietly climbs, 1.6, then 2.4, then 4.8 per 1,000, past a number nobody had named, so nothing about any single day looks like an emergency until it's well past the point a real threshold would have caught.
The step that should sit here, and doesn't.
A, ability to contest. Rasmus Kjeld, the on-call engineer, sees the bad number first, at 2am, three weeks into the rollout. He has no authority to act on it, because "criteria" was never written down as a real number with his name attached, so the fastest anyone can act is Thursday, once a VP's calendar opens.
R, reduce. Write a named numeric threshold for every metric that matters before launch, for example, "policy-exceeding reply rate above 3 per 1,000 accepted drafts in any rolling 24 hours." Pre-agree, in writing, who can act on it the second it's crossed, no approval chain required.
D, detect. After the criteria has run for a while, check it two ways: how many times it fired on something that turned out, on review, to be noise, which means it was set too strict, and how many real incidents happened while the number stayed under it, which means it was set too loose.
Where this answer would fail
If the fix is "monitor it closely" or "add a review meeting," none of it counts. A written number and a named owner are things you can point to on launch day and check whether they exist. A feeling that something looked off is not.
And if you want to be sure it really works, try it somewhere else
Cartway, a grocery delivery app, runs into the same gap rolling out an AI substitution tool, in a warehouse with nothing to do with customer messaging.
G, groups. The warehouses already running the substitution tool on 20 percent of orders, where the model auto-picks a replacement for an out-of-stock item and texts the customer a short window to reject it. And the warehouses still waiting their turn, who inherit whatever the rollout committee decides once something goes wrong.
U, unequal. The harm concentrates in orders where a customer's profile flags an allergy, and the model substitutes something that crosses it, peanut butter for almond butter, dairy cream for oat cream. That error rate, per 10,000 orders, climbs from 0.4 to 1.1 to 2.6 over three weekends, with no named number saying when it should stop.
A, ability to contest. The regional ops lead on weekend duty spots the spike Saturday night. She has no authority to disable substitutions without the VP of operations, who's unreachable until Monday.
R, reduce. The same practice: a named threshold, "allergen-flagged substitution errors above 1 per 10,000 orders in a rolling 48 hours," with the weekend on-duty lead pre-authorized to disable substitutions immediately, no VP required.
D, detect. A quarterly review of every rollback the threshold triggered, checking which ones were real versus noise, and a separate check for any allergen complaint that came in while the number stayed under the line.
Swap the trigger and it still runs
- Speed: a model that updates weekly instead of daily just means the missing number bites less often, but when it does, the same gap, no threshold, no owner, still decides how long the damage runs.
- Cost: a cheaper model tempts more teams to launch solo on a "we'll watch it" plan, which means more surfaces going live with no written number behind them, not fewer.
- The model gets better: a sharper model makes the weekly numbers look even cleaner, which makes it easier to skip writing a real threshold, not harder, because a good-looking dashboard is exactly what talks a launch plan out of naming one.
Where people run it wrong
- Treating "we'll monitor it closely" as a rollback plan, when it's actually a promise with no number and no owner behind it.
- Writing a threshold, but never naming who's allowed to act on it alone, so it still needs a meeting to fire.
- Setting a threshold only for the loud, catastrophic failure, and missing the quiet one that drifts for weeks before anyone notices.
How to use it live
Ask one question before you answer: "who's actually allowed to pull this back at 2am, and what number tells them to?" That buys you a second to think, and it's usually the exact question the interviewer wanted asked.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits a question about the rollback criteria you'd set before launch, and why?
Tap to flip
ANSWER
GUARD, for risk. The real question isn't whether the model is safe enough, it's who a rollback decision protects, where a missing number hurts worst, who can't act without one, the actual number and owner you'd write down, and how you'd know later if that number was right.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Btissam Amrani, a product manager at Verlane. She's run the support-agent tools team for three years, and wrote the launch plan for the auto-reply suggestion tool's 25 percent rollout.
3 · THE HABIT
What did Btissam never write down, even as the rollout's numbers kept looking clean?
Tap to flip
ANSWER
A real number for when the rollout should end, and a named owner who could act on it. Her launch plan said "we'll keep an eye on quality metrics and pull back if something looks off," which was fine at a 5 percent hand-reviewed alpha and stopped being fine at 25 percent.
4 · THE SWITCH
What's the two-setting switch this answer turns on?
Tap to flip
ANSWER
Either a rollback criterion is a written number with a named owner who can act alone, or it isn't really a criterion, it's a mood everyone agreed to feel. There's no version in between where "looks off" quietly does the same job as a number.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
The launch plan's risk section said "we'll keep an eye on quality metrics and pull back if something looks off," instead of a written number with an owner attached. That made sense at the 5 percent alpha, where every draft got hand-reviewed. It stopped making sense once real client brands and customers were in the 25 percent rollout.
6 · THE NUMBER
Fill in the blank: the policy-exceeding reply rate went from ______ per 1,000 in week one to ______ per 1,000 by week four, when the rollout was finally paused.
Tap to flip
ANSWER
1.6 and 7.2. It crossed 3 per 1,000, the number this answer proposes as the actual threshold, partway through week two, more than a week before anyone with the authority to act had even seen it.
7 · THE REPLAY
Same rollout, the threshold and the owner written in from day one. What changes?
Tap to flip
ANSWER
Rasmus sees the rate cross 3 per 1,000 sometime Tuesday of week two, and flips the feature flag back himself that night, because the plan already said he could. The rollout pauses at roughly 3.2 per 1,000 instead of climbing to 7.2 over two more days of waiting for a calendar to open.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the reduce step become?
Tap to flip
ANSWER
Cartway's grocery-delivery substitution tool. Reduce: a named threshold, allergen-flagged substitution errors above 1 per 10,000 orders in a rolling 48 hours, with the weekend on-duty ops lead pre-authorized to disable substitutions immediately, no VP approval required.
Check yourself Score: 0 / 0
Short answer
1. What specific rollback criteria does this answer propose, and why didn't the original launch plan have anything like it?
Show hint
Look for the exact sentence the original risk section used, then compare it to what a real threshold would say instead.
Show answer
Model answer: "The launch plan said 'we'll keep an eye on quality metrics and pull back if something looks off,' which had no number and no owner. I'd replace it with: 'if the policy-exceeding reply rate goes above 3 per 1,000 accepted drafts in any rolling 24 hours, the rollout ends,' with the on-call engineer pre-authorized to flip the feature flag back the moment it's crossed, no approval needed."
Multiple choice
2. Which two groups does the G step name in Btissam's story, and what gives them different stakes in a rollback?
- A. Agents and customers already on the 25 percent rollout, mid-conversation, and the other 75 percent still waiting for their rollout wave.
- B. Btissam and Rasmus, who disagree about the launch date.
- C. The analytics lead and the VP, who disagree about the Thursday meeting.
- D. Verlane and the client brands, who disagree about the refund policy.
Show hint
Look for who's already using the feature versus who's still waiting on it.
Show answer
A. B, C, and D name real people or organizations in the story, but not the two groups the G step separates: the one already living with a rollback if it happens, and the one whose wave just gets decided for them.
Fill in the blank
3. The policy-exceeding reply rate was ______ per 1,000 in week one, and had climbed to ______ per 1,000 by the time the rollout paused in week four.
Show hint
Both numbers are in the chart, "Policy-exceeding reply rate, by rollout week."
Show answer
1.6 and 7.2. Same tool, same rollout. What changed was only how long it took for a number nobody had written down to finally get acted on.
True or false
4. True or false: the risk in this story was that Rasmus, the on-call engineer, failed to notice the problem in time.
Show hint
Ask what Rasmus actually did once he saw the number, and what stopped him from doing more.
Show answer
False. Rasmus caught it early, at 2am in week three, well before the worst of it. The risk was that he had no written number to check it against and no authority to act on his own, so noticing wasn't enough.
Short answer, apply it yourself
5. Pick an AI feature you've seen roll out gradually. What would the rollback criteria for it need to name, a metric, a number, and an owner, to actually work?
Show hint
Think about who would be the first person to see it go wrong, and whether they'd actually be allowed to do anything about it.
Show answer
Model answer: "A ride-share app's AI fare-estimate feature, rolled out to a slice of cities. The metric would be the rate of trips where the final fare beat the estimate by more than 20 percent, with a threshold like 'above 4 percent of trips in any city, in any rolling day.' The owner would need to be the on-call pricing engineer, given authority to pause fare estimates in that city alone, without waiting on a regional manager."
Multiple choice
6. In this story, where would a strict rollback threshold matter least?
- A. The tool's suggested closing lines, like "Thanks for reaching out, have a great day!"
- B. Drafts that suggest a refund above the client's stated policy limit.
- C. Drafts that suggest a subscription extension longer than policy allows.
- D. Any draft the backend policy-checker would flag as exceeding written policy.
Show hint
Ask which kind of suggestion genuinely can't create a policy-exceeding promise, no matter how often it's used.
Show answer
A. B, C, and D are exactly the drafts the threshold exists to catch. A is what this answer would leave alone: wording with no way to promise anything a client would need to honor or retract.