How long should each rollout phase last, and what determines it?
- Size each phase by how many home-days it takes to see the rare failure pattern happen enough times to trust it, not by a flat calendar rule.Why: a phase length with no arithmetic behind it is a guess wearing a schedule.
- Hold every phase to at least a one-week floor, even when the math says less would do.Why: a floor shorter than a week can promote a phase before a weekday and a weekend both show up in it.
- Check the whole phase plan against the real date the business needs the full launch live by, before promising it.Why: that date is a hard sanity check, not a formality to mention at the end.
- Give the plan a range, low if the rare pattern behaves the way you assumed, high if it doesn't.Why: one number hides a forty-six day swing between the two ends.
- Watch the smallest phase hardest, not the biggest one.Why: the internal phase has the fewest homes to draw a rare failure from, so it's the one that actually breaks under a wrong guess about rarity.
- When the numbers don't fit the date, shrink the confidence bar or cut a phase, never quietly skip the floor.Why: skipping the floor is how a phase gets promoted before it's actually been tested.
How to answer this, stage by stage
Nobody is grading whether you land on exactly forty-five days. They're grading whether you can defend the arithmetic behind "two weeks a phase," whether the range is honest, and whether you close on something the room can check. Eight moves get you there.
Let's learn
Northloft's smart thermostat has a feature that watches when people are actually home, and moves the heating and cooling schedule to match, instead of running on a fixed clock someone set once and never touched again.
Northloft's rollout playbook, built over a dozen past feature launches, gives every phase a flat two weeks. It worked for a new color theme, a notification tweak, a redesigned settings screen. None of those had a failure mode that only showed up in one narrow slice of homes.
The scheduling optimizer's real failure only shows up in those split-schedule homes, and only while the model is still learning that particular home, about one bad night in twenty during its first two weeks. A flat two-week default doesn't know the difference between a phase with four hundred homes and a phase with two hundred thousand.
At its worst, that gap shows up as a phase promoted before the real rate is known. Thousands of new homes inherit the same unnoticed problem all at once, in the exact week the team has the least practice spotting it, right as the marketing push for the season turns the volume up.
The choice I would take back. Northloft's rollout playbook set every phase at a flat two weeks, because two weeks had always been enough for a dozen past launches. That felt responsible, a known number instead of a guess. It also assumed every future feature would fail the same way those did, evenly, not tucked inside one narrow slice of homes.
What I would leave alone. The one-week floor under early access and broad beta doesn't need to change even if split-schedule homes turn out to be more common than we think. That floor exists so a weekday and a weekend both show up in the data, and that reason doesn't get stronger or weaker with rarity. Only the internal phase, the one small enough that the math actually binds, needs to move.
The lesson. How long each rollout phase should last was never really a calendar question. It's a question about how fast a phase's own size lets a real problem show itself, and the company's two-week habit had been quietly answering a question nobody had actually asked.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why the internal phase, the smallest one, is the one that almost got missed.
Tevita Faleolo has run feature rollouts at Northloft for three years. He's the one who decides when a phase has run long enough to promote to the next, and he can usually tell within a day of a phase's numbers coming in whether something is actually wrong or just noisy.
The scheduling optimizer was the biggest thing on his roadmap that year. Internal testing started with four hundred of Northloft's own employees running it on their own homes. For the first ten days it looked clean. A handful of flagged nights, nothing that lined up with anything, the kind of noise every rollout has.
Marketing had already booked the launch push for the first hard freeze of the season, about nine weeks out. Everyone wanted the internal phase done in the two weeks the playbook always gave it, and by day ten it already looked fine.
Tevita started the internal phase checking every flagged night by hand. By day six he was skimming the daily summary instead. By day ten, with the freeze push looming and the numbers quiet, he'd stopped opening the flag list some evenings altogether.
On day eleven, one of Northloft's own engineers mentioned in the team chat that her thermostat had run the heat for six hours overnight while she was on a night shift and the house sat empty. She wasn't upset. She thought it was funny. She dropped a laughing emoji and moved on.
He pulled the full flag list back up. Two more split-schedule homes had the same kind of night in the same week, quiet, unremarked, sitting in a queue nobody had reread since day six. Four hundred employee homes had produced exactly the twelve bad nights the team needed to trust the rate. Nobody had counted them. They'd counted days instead.
Three bad nights, out of four hundred homes, isn't a disaster on its own. What scared Tevita was the arithmetic sitting right behind it. If the internal phase had been promoted on schedule, at day ten as planned, the same pattern would have shown up next in broad beta, forty thousand homes, during the same week Northloft's marketing push went live. Not three homes noticing. Hundreds.
The team never had a real number for how long the internal phase needed. They had a habit, two weeks, that had worked for a dozen launches that didn't have this kind of failure. Ten days in, with the numbers quiet, the habit said done. The arithmetic, run properly, said the internal phase actually needed fifteen days to see this problem the required twelve times, and that gap, five days, was the whole margin the near miss lived inside.
Weeks earlier, when the rollout plan first got signed off, someone on the call had asked whether two weeks was really enough for something this new. Tevita said it had always been enough before. Nobody pushed back. It was the only real number in the room.
The second version of the plan didn't change the two-week default for early access or broad beta, both were already governed by the one-week floor either way. What changed was the internal phase: fifteen days, not ten, sized to the actual arithmetic instead of the calendar habit. The freeze deadline still had eighteen days of room under the honest plan. Broad beta went live on schedule. No second near miss.
One design asked how many days had passed. The other asked how many bad nights the phase had actually produced. The first can look calm right up until forty thousand homes inherit the same quiet problem at once.
The thing I'd tell myself, back on that day-ten call: a phase that looks quiet isn't the same thing as a phase that's actually been tested. I'd counted the calendar. I should have been counting the nights.
BOUND, sized in home-days instead of a flat calendar
This is a sizing question about how many home-days it actually takes to trust a rare failure rate, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. How long a phase should run isn't one number. It's an equation: bad nights expected in a day equals the homes in that phase, times the share running an odd schedule, times how often the model gets one of those nights wrong while it's still learning that home. Days needed equals the bad nights we want to see, divided by that daily number.
O, own the numbers. For Northloft, about four percent of homes are split-schedule: a night-shift worker, someone home all day while the house empties, two people on different clocks. The model gets about one night in twenty wrong on those homes during its first two weeks. We want to see that twelve times before trusting the rate. At four hundred employee homes, that's about fifteen days.
U, use a range. If four percent holds, fifteen days for the internal phase is right, and the one-week floor governs every phase after it. If the real rate is closer to one percent, four times rarer, the internal phase alone needs about sixty days, because four hundred homes can't hand you twelve bad nights any faster than that. Every later phase has enough homes that the swing barely reaches it.
N, nail the sanity check. Marketing needs the full launch live before the first freeze, about sixty-three days out. The honest plan, fifteen plus seven plus ten plus fourteen, adds up to forty-five days, eighteen days of room. The rarer scenario runs to ninety-one days, twenty-eight days past the freeze.
D, direction. Two things could move this number, and they don't move it the same amount. Whether split-schedule homes are really as common as assumed swings the internal phase from fifteen days to sixty, a forty-six day swing on the whole plan. Recruiting more employee homes into that phase helps, doubling it from four hundred to eight hundred cuts the swing roughly in half, but it doesn't change whether the four percent guess was right in the first place. Rarity is the bigger lever. Recruiting is the smaller one, and it's the one people reach for first because it feels like doing something.
And if you want to be sure it really works, try it somewhere else
Verdant Row Co-op runs an AI system that decides when to open and close irrigation valves per field zone, reading soil moisture and the forecast, instead of a fixed timer a farmer sets once each season. The co-op is rolling it out from its own twenty research plots to all three thousand member farms.
B, break it down. Same shape, different work. Weeks needed equals the bad irrigation cycles we want to see, divided by farms in the phase times the share running mixed crops on adjoining zones times how often the system waters the wrong zone while it's still learning that farm's map.
O, own the numbers. About eight percent of member farms mix crops with very different water needs on neighboring zones, an orchard block next to a row-crop block. On those farms, the system gets about one cycle in ten wrong during its first two weeks, and cycles run twice a week. The co-op wants to see that eight times before trusting the rate. On the twenty research plots, that's about five weeks.
U, use a range. If eight percent holds, the honest plan reaches every farm in about ten weeks. If mixed-crop farms are really closer to two percent, four times rarer, the research-plot phase alone stretches past twelve weeks, because twenty plots don't hand you eight bad cycles any faster than that.
N, nail the sanity check. The co-op needs full coverage before spring planting, about sixteen weeks out. The common scenario uses ten weeks, six weeks of room. The rare scenario runs past nineteen weeks, three weeks past planting.
D, direction. Same tension as Northloft's thermostat. How common mixed-crop farms really are swings the plan by nine weeks. Adding a third agronomist to the rollout team adds real capacity, but it doesn't change how many mixed-crop farms are sitting in those twenty research plots.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: size each phase from real home-days, check it against the actual deadline, don't run every phase on a company default.
Cost: the team can't add another reviewer this year. Don't cut the floor, cut the pace: space out the phases most likely to hit the rare pattern, rather than promoting on schedule.
The model got better: a newer version of the optimizer rarely mispredicts split-schedule homes anymore. That doesn't remove the need to check rarity, it just moves where the real risk sits, from split-schedule homes to whatever pattern the new model hasn't seen yet.
Where people run it wrong.
They give every phase the same length because it's the company default, without checking whether that phase's population can even produce the signal that fast.
They treat the calendar floor as the actual statistical answer, when the floor only matters once a phase is already big enough for the floor to be the binding constraint.
They reach for a new hire as the first fix, when the number actually breaking the plan is a rate nobody has measured yet.
How to use it live. Say the equation before naming a single number: "the phase length isn't the real question, the home-days behind it are, so before I set a calendar, I'd want to know how rare the failure pattern this is protecting against actually is." That buys the room to ask a real question instead of repeating a two-week habit that sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Rollout strategy and phased launches
- #1 Design the rollout plan for an AI feature going to two million users.
- #2 What percentage would you start a canary at, and how do you decide?
- #3 Explain the difference between a feature flag rollout and a model rollout.
- #4 What metrics gate each stage of a phased rollout?
- #5 How do you choose which users go first?
- #6 Describe the rollback criteria you would set before launch.