Present the rollout plan for a launch I describe.
- Size each phase by confirmed real catches of its newest edge-case category, not by what percent of the company that phase covers.Why: a phase length with no arithmetic behind it is a guess wearing a schedule.
- Hold every phase to at least a one-week floor, even when the math says less would do.Why: a floor shorter than a week can promote a phase before a weekday and a weekend both show up in the reports.
- Check the flag volume each phase would produce against how many reports the finance team can actually hand-review in a week.Why: a plan the reviewers can't keep up with turns into a growing queue, not a launch.
- Give the plan a range, low if the rarest new category is as common as assumed, high if it isn't.Why: one number hides a fifteen-week swing between the two ends.
- Watch the phase introducing the rarest new edge case hardest, not the biggest phase.Why: population size doesn't decide how long a phase needs, the rarest thing inside it does.
- When the honest number and the review capacity disagree with the calendar, stretch the phase, never quietly skip the floor to hit a date.Why: skipping the floor is how a category gets trusted before the finance team has actually seen it.
How to answer this, stage by stage
Nobody is grading whether you land on exactly ten weeks. They're grading whether you can defend the arithmetic behind "three weeks for phase one," whether the range is honest, and whether you close on something the room can check. Seven moves get you there.
Let's learn
Before this tool existed, six people on Corriedale's finance team read every one of the roughly eight hundred expense reports that come in each week, by hand, checking each one against a policy binder nobody had fully memorized. That took most of two working days, every week, and it still missed things, because nobody can hold four different kinds of policy exception in their head report after report.
Now the tool reads all eight hundred in about a minute and flags the roughly forty a week that look wrong. The team reviews forty instead of eight hundred.
The real risk isn't the extra speed. It's what happens the first time a new department introduces a violation pattern the model has only seen a handful of real examples of, and the finance team either waves through something real or drowns in something the model is still guessing at.
At its worst, a category nobody has actually earned trust on rolls out to all eleven hundred employees at once. A rare pattern that's shown up six times a week in Sales shows up sixty times a week company-wide, and the finance team, buried in flags they can't yet tell apart from noise, starts waving them through unread. The tool built to catch missed violations ends up hiding them behind a stamp of approval nobody has time to actually check.
The choice I would take back. Corriedale's rollout playbook, built for a dozen past internal tools, a new expense-report screen, a new time-tracking app, gives every phase two weeks, then doubles the population, then opens it to everyone. That felt careful. It also assumed every new tool changes a workflow, not a judgment. None of those past tools ever had to earn trust on a specific kind of mistake before more people saw it.
What I would leave alone. The simple flags, a missing receipt, a charge with no attached approval, don't need this staged confidence-building. The model has been reliably right about those since week one, because every department produces that kind of report the same way. Turn those on company-wide, day one. Only the categories the model hasn't seen enough of yet need the staged rollout.
The lesson. How long a phase should run was never really a calendar question. It's a question about whether the team has actually seen the thing they're being asked to trust, and Corriedale's two-week habit had been quietly answering a question nobody had actually asked.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why phase one, the one everyone assumed was already done, almost wasn't.
Every Friday afternoon, Renato Aguirre pulls the week's flag count before anyone else sees it. He's run rollouts of internal tools at Corriedale Textiles for four years, and he can usually tell within an hour of a week's numbers coming in whether something is actually wrong or just noisy.
The expense flagging tool was the biggest thing on his roadmap that year. It started with Sales and Client Relations, two hundred people, the department with the most travel and the most client dinners, and so the most ways to bend a receipt.
For the first ten days it looked clean. A steady trickle of flags, a few dozen a week, nothing that lined up into a pattern. Corriedale's rollout playbook, the same one used for a screen redesign two years earlier, called for two weeks per phase, then double the department, then open it to everyone. Leadership wanted Marketing and Regional Ops added by day fourteen, and by day ten the numbers already looked quiet enough to promise it.
Renato started phase one checking every flag by hand with the team. By day six he was skimming the weekly summary instead. By day ten, with the promotion date circled and the numbers calm, nobody on the finance team was cross-checking the split-transaction flags against each other anymore, just clearing them one at a time.
On day thirteen, one of the reviewers noticed something while closing out an unrelated report: a client dinner from week two, fourteen hundred dollars, had gone through as three separate charges of under five hundred each, filed two days apart by the same sales rep. None of the three had been flagged. Not because the model missed the pattern outright. It hadn't yet seen enough real splits to be confident calling one, so it had stayed quiet on the borderline cases rather than flood the team with guesses.
She wasn't alarmed. She mentioned it in passing, the way you'd mention a typo.
Only eleven confirmed splits by day thirteen, not the eighteen the arithmetic actually needed before the model could be trusted on that category. Phase one had never been ready. The calendar said ready. The numbers didn't.
Renato pulled the promotion. Phase one ran eight more days, to day twenty-one, and picked up the seven more confirmed splits it needed along the way, eighteen total, right on the arithmetic. Only then did Marketing and Regional Ops go live.
Three weeks in, not two, isn't a disaster on its own. What stayed with Renato was what would have happened on the version where he hadn't pulled the promotion. The gift-mixing pattern Marketing was about to introduce ran even rarer than split transactions. If phase one had opened on schedule at day fourteen, phase two would have layered an even less-tested category on top of a barely-tested one, in front of four hundred fifty people instead of two hundred, the same week leadership's board update was due.
The team never had a real number for how long phase one needed. They had a habit, two weeks, that had worked for tools that only ever changed a screen. Eleven days in, with the flags quiet, the habit said done. The arithmetic, run properly, said phase one actually needed three weeks to see the category the required eighteen times, and that gap, ten days, was the whole margin the near miss lived inside.
Weeks earlier, when the rollout plan first got signed off, someone on the call had asked whether two weeks was really enough for something that had to learn a pattern, not just a new screen. Renato said it had always been enough before. Nobody pushed back. It was the only number in the room.
The second version of the plan didn't touch the floor under phase three, that was already governed by review capacity either way. What changed was phase one: three weeks, not two, sized to the confirmed count instead of the calendar habit. Phase two, watching its own rarer category from the start this time, ran its full five weeks before Renato let anyone add the rest of the company. No second near miss.
One design counted days. The other counted catches. The first can look calm right up until four hundred fifty people inherit a category nobody actually finished testing.
The thing I'd tell myself, back on that day-ten call: a phase that looks quiet isn't the same thing as a phase that's actually been proven. I'd been counting the calendar. I should have been counting the catches.
BOUND: what actually decided Corriedale's ten weeks
This is a sizing question about how many confirmed catches it actually takes to trust a rare category, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. How long a phase should run isn't one number. It's an equation: expected confirmed violations in a week equal the reports in that phase, times the share that fall into the risky spend category, times the share of those that are a genuine violation. Weeks needed equal the confirmed catches we want to see, divided by that weekly number.
O, own the numbers. Phase one, Sales, two hundred people, four hundred reports a week. Thirty percent are travel and client-entertainment claims. Five percent of those are real split transactions, six a week. We want eighteen confirmed before trusting the category. That's three weeks. Phase two adds two hundred fifty people and a rarer pattern, gift-mixing, five weeks. Phase three, the rest of the company, is governed by the one-week floor, call it two weeks for safety.
U, use a range. If eight percent of phase two's reports are gift-adjacent, five weeks holds. If the real rate is closer to two percent, four times rarer, phase two alone stretches to twenty weeks, because two hundred fifty people can't hand you eight confirmed gift-mixing cases any faster than that. The whole plan runs ten weeks on the low end, twenty-five on the high end.
N, nail the sanity check. Six reviewers, each able to spare about thirty minutes a day for this, is nine hundred minutes a week team-wide. The phased plan peaks at four hundred eighty minutes a week, well inside that. A flat jump to twenty-five percent of the company, mixed across every department before any category is calibrated, runs about eleven hundred minutes a week, past what the team can sustain.
D, direction. Two things could move this number, and they don't move it the same amount. Whether gift-mixing is really as rare as assumed swings phase two from five weeks to twenty, a fifteen-week swing on the whole plan. Losing one reviewer barely touches it, the phased plan's peak load sits well under capacity either way. Rarity is the bigger lever. Headcount is the smaller one, and it's the one people reach for first because it feels like doing something.
And if you want to be sure it really works, try it somewhere else
Tallowmere Mutual, a regional home and auto insurer, runs an AI tool that reads incoming claims and flags ones that look like a policy violation, before a payout goes out. The pilot team is eight adjusters on the auto desk, out of forty across the claims department.
B, break it down. Same shape, different work. Weeks needed equal the confirmed catches we want to see for a category, divided by claims in the phase times the share touching more than one policy rider times how often that turns out to be a genuine double-filing.
O, own the numbers. The pilot team handles three hundred claims a week. Five percent touch more than one rider within ten days, fifteen a week. Eight percent of those are genuine double-filings, about one a week. The team wants six confirmed catches before trusting the rate. That's about five weeks.
U, use a range. If the real double-filing rate is closer to two percent, four times rarer, the pilot phase alone stretches past nineteen weeks.
N, nail the sanity check. Eight adjusters, each able to spare twenty minutes a day for this, is about eight hundred minutes a week team-wide. The pilot's flag volume runs well under that, even the rare-case scenario doesn't breach it, the wait is in weeks, not review minutes.
D, direction. Same tension as Corriedale's expense tool. How common the double-filing pattern really is swings the plan by fifteen weeks. Adding a ninth adjuster adds real capacity, but it doesn't change how many double-filings are sitting in those three hundred weekly claims.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: size each phase by confirmed real catches of its newest category, check the flag volume against real review capacity, don't run every phase on a company percentage.
Cost: the team can't add another reviewer this year. Don't cut the floor, cut the pace: hold the rarer-category phase longer instead of promoting it on schedule.
The model got better: a newer version rarely misses the rare category anymore. That doesn't remove the need to check rarity, it just moves where the real risk sits, from today's rare category to whatever pattern the new model hasn't seen yet.
Where people run it wrong.
They give every phase the same length because it's the company default, without checking whether that phase's population can actually produce the confirmed-catch count that fast.
They treat the one-week floor as the real statistical answer, when the floor only matters once a phase is already big enough for the floor to be the binding constraint.
They reach for more headcount as the first fix, when the number actually breaking the plan is a rarity rate nobody has confirmed yet.
How to use it live. Say the equation before naming a single number: "the phase length isn't the real question, how many confirmed catches of its newest category it takes to trust the model is, so before I set a calendar, I'd want to know how rare that pattern actually is." That buys the room to ask a real question instead of repeating a two-week habit that sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Rollout strategy and phased launches
- #1 Design the rollout plan for an AI feature going to two million users.
- #2 What percentage would you start a canary at, and how do you decide?
- #3 Explain the difference between a feature flag rollout and a model rollout.
- #4 What metrics gate each stage of a phased rollout?
- #5 How do you choose which users go first?
- #6 Describe the rollback criteria you would set before launch.