ConceptAdvancedShipping & Model Lifecycle / Rollout strategy and phased launches / #7

How long should each rollout phase last, and what determines it?

The direct answer
Size each phase by how many home-days it takes to see the rare failure pattern often enough to trust it, not by a flat number of weeks. Hold every phase to a one-week floor so a weekday and a weekend both show up, then check the whole plan against the real date the business needs the launch live by. If those two numbers clash, shrink the confidence bar or cut a phase, never quietly skip the floor to hit the date.
Do this, in order
  1. Size each phase by how many home-days it takes to see the rare failure pattern happen enough times to trust it, not by a flat calendar rule.Why: a phase length with no arithmetic behind it is a guess wearing a schedule.
  2. Hold every phase to at least a one-week floor, even when the math says less would do.Why: a floor shorter than a week can promote a phase before a weekday and a weekend both show up in it.
  3. Check the whole phase plan against the real date the business needs the full launch live by, before promising it.Why: that date is a hard sanity check, not a formality to mention at the end.
  4. Give the plan a range, low if the rare pattern behaves the way you assumed, high if it doesn't.Why: one number hides a forty-six day swing between the two ends.
  5. Watch the smallest phase hardest, not the biggest one.Why: the internal phase has the fewest homes to draw a rare failure from, so it's the one that actually breaks under a wrong guess about rarity.
  6. When the numbers don't fit the date, shrink the confidence bar or cut a phase, never quietly skip the floor.Why: skipping the floor is how a phase gets promoted before it's actually been tested.

How to answer this, stage by stage

Nobody is grading whether you land on exactly forty-five days. They're grading whether you can defend the arithmetic behind "two weeks a phase," whether the range is honest, and whether you close on something the room can check. Eight moves get you there.

1
Scope it to one real product and one real starting population
Say it like this
"Let's ground this. Say Northloft, a home automation company, sells a smart thermostat with an AI scheduling optimizer built in. It watches when people are actually home and moves the heating and cooling schedule to match, instead of running on a fixed clock someone set once and forgot about. Northloft has four hundred thousand homes running that app today. A phased rollout means moving from Northloft's own four hundred employees, up through all four hundred thousand, in stages."
Why this works
Stops the answer from staying abstract before a single number gets attached to it.
2
Say the structure out loud before naming any numbers
Say it like this
"The question isn't how many weeks a phase should get. It's how many home-days it takes to see the real failure often enough to trust it, at whatever size that phase actually is. A bigger phase hits that count faster. A smaller phase needs more calendar time to get there. Phase length is an output of the phase's own size, not a company default."
Why this works
Shows the equation before the arithmetic, so the numbers that follow read as a plan, not a habit.
3
Break down the equation
Say it like this
"Here's the shape. Expected bad nights in a day equals the homes in that phase, times the share of homes with an odd schedule, times how often the model gets one of those nights wrong. Days needed equals however many bad nights we want to see before we trust the rate, divided by that daily number."
Why this works
This is the B step of BOUND, the equation stated before a single number gets attached.
4
Own the real numbers behind each term
Say it like this
"For Northloft, about four percent of homes run what we call a split-schedule pattern, a night-shift worker, someone home all day while the rest of the house is out, two people on different wake times under one roof. While the model is still learning one of those homes, it gets a night wrong about one time in twenty. We want to see that happen at least twelve times before we trust the rate. At four hundred employee homes, that's about fifteen days. Call it two weeks."
Why this works
This is the O step, a real proposed number, not "we'll watch it closely."
5
Give the range, not one number
Say it like this
"If four percent is right, that math holds for every phase. If the real number is closer to one percent, four times rarer, the internal phase alone needs about sixty days, not fifteen, because four hundred homes just don't hand you twelve split-schedule bad nights any faster than that. Every phase after it has enough homes that this swing barely shows. The internal phase is the one place a wrong guess about rarity actually breaks the plan."
Why this works
A single number here would claim a confidence the plan doesn't have yet.
6
Sanity check the plan against the real deadline
Say it like this
"Marketing wants this live for everyone before the first hard freeze of the season, about sixty-three days out. The plan above adds up to forty-five days to reach everyone, eighteen days of room. If the rarer number is the true one, the same plan runs to ninety-one days, twenty-eight days past the freeze. That's the real test, whether the honest number fits inside the date, not the hopeful one."
Why this works
This is the N step, and it's the step most rollout plans skip.
7
Name what moves the number most
Say it like this
"If I had to bet on one thing to watch, it's not the deadline. It's how common split-schedule homes actually turn out to be. That single assumption can add forty-six days to the plan, almost all of it sitting in the smallest phase. Recruiting more employee homes into that first phase helps, but it only closes about half the gap. Adding a reviewer barely moves it at all."
Why this works
This is the D step, what a good estimator says out loud and a bad one skips.
8
Close on the one line
Say it like this
"So here's what I'd actually say: I'd size each phase by how many home-days it takes to see the real failure enough times to trust it, hold a one-week floor under every phase regardless, and check the whole plan against the real freeze date before I promise it to anyone. If the honest number doesn't fit, I'd rather shrink the confidence bar or cut a phase than quietly skip the floor."
Why this works
Closes on the literal ask, a claim someone could check, not a vibe about moving carefully.
If you remember one thing The internal phase is the one that breaks first if the rare pattern turns out rarer than assumed, not the big ones. Check that assumption before you trust any phase's length, because a one-week floor and a big population protect every other phase from it.

Let's learn

Northloft's smart thermostat has a feature that watches when people are actually home, and moves the heating and cooling schedule to match, instead of running on a fixed clock someone set once and never touched again.

Northloft's rollout playbook, built over a dozen past feature launches, gives every phase a flat two weeks. It worked for a new color theme, a notification tweak, a redesigned settings screen. None of those had a failure mode that only showed up in one narrow slice of homes.

Knowledge spark: what's a split-schedule home? A home where people are awake and around on a pattern the model doesn't expect. A night-shift worker. Someone home all day while the rest of the house is out. Two people on different wake times under one roof. About four in every hundred Northloft homes.

The scheduling optimizer's real failure only shows up in those split-schedule homes, and only while the model is still learning that particular home, about one bad night in twenty during its first two weeks. A flat two-week default doesn't know the difference between a phase with four hundred homes and a phase with two hundred thousand.

We didn't lose a few bad nights. We lost the chance to notice them while only four hundred homes were watching, instead of forty thousand.

At its worst, that gap shows up as a phase promoted before the real rate is known. Thousands of new homes inherit the same unnoticed problem all at once, in the exact week the team has the least practice spotting it, right as the marketing push for the season turns the volume up.

The decision that mattered Size every phase by how many home-days it actually takes to trust the failure rate, not by whatever the company's default happens to be. That's the one check that decides whether a bad pattern gets caught at four hundred homes or forty thousand.

The choice I would take back. Northloft's rollout playbook set every phase at a flat two weeks, because two weeks had always been enough for a dozen past launches. That felt responsible, a known number instead of a guess. It also assumed every future feature would fail the same way those did, evenly, not tucked inside one narrow slice of homes.

What I would leave alone. The one-week floor under early access and broad beta doesn't need to change even if split-schedule homes turn out to be more common than we think. That floor exists so a weekday and a weekend both show up in the data, and that reason doesn't get stronger or weaker with rarity. Only the internal phase, the one small enough that the math actually binds, needs to move.

The lesson. How long each rollout phase should last was never really a calendar question. It's a question about how fast a phase's own size lets a real problem show itself, and the company's two-week habit had been quietly answering a question nobody had actually asked.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why the internal phase, the smallest one, is the one that almost got missed.

Tevita Faleolo has run feature rollouts at Northloft for three years. He's the one who decides when a phase has run long enough to promote to the next, and he can usually tell within a day of a phase's numbers coming in whether something is actually wrong or just noisy.

The scheduling optimizer was the biggest thing on his roadmap that year. Internal testing started with four hundred of Northloft's own employees running it on their own homes. For the first ten days it looked clean. A handful of flagged nights, nothing that lined up with anything, the kind of noise every rollout has.

Marketing had already booked the launch push for the first hard freeze of the season, about nine weeks out. Everyone wanted the internal phase done in the two weeks the playbook always gave it, and by day ten it already looked fine.

Tevita started the internal phase checking every flagged night by hand. By day six he was skimming the daily summary instead. By day ten, with the freeze push looming and the numbers quiet, he'd stopped opening the flag list some evenings altogether.

On day eleven, one of Northloft's own engineers mentioned in the team chat that her thermostat had run the heat for six hours overnight while she was on a night shift and the house sat empty. She wasn't upset. She thought it was funny. She dropped a laughing emoji and moved on.

It wasn't the bad night that stopped Tevita. It was realizing nobody would have noticed, if she hadn't happened to mention it out loud.

He pulled the full flag list back up. Two more split-schedule homes had the same kind of night in the same week, quiet, unremarked, sitting in a queue nobody had reread since day six. Four hundred employee homes had produced exactly the twelve bad nights the team needed to trust the rate. Nobody had counted them. They'd counted days instead.

Three bad nights, out of four hundred homes, isn't a disaster on its own. What scared Tevita was the arithmetic sitting right behind it. If the internal phase had been promoted on schedule, at day ten as planned, the same pattern would have shown up next in broad beta, forty thousand homes, during the same week Northloft's marketing push went live. Not three homes noticing. Hundreds.

The team never had a real number for how long the internal phase needed. They had a habit, two weeks, that had worked for a dozen launches that didn't have this kind of failure. Ten days in, with the numbers quiet, the habit said done. The arithmetic, run properly, said the internal phase actually needed fifteen days to see this problem the required twelve times, and that gap, five days, was the whole margin the near miss lived inside.

Weeks earlier, when the rollout plan first got signed off, someone on the call had asked whether two weeks was really enough for something this new. Tevita said it had always been enough before. Nobody pushed back. It was the only real number in the room.

Hand-sketched number line running from 0 to 100 days. A green marked point at 45 days labeled if odd-schedule homes are as common as we think. A red-orange marked point at 91 days labeled if they're four times rarer than we think. An amber flag pinned at 63 days labeled 63-day freeze deadline, sitting closer to the 45-day point than the 91-day point.
The 63-day freeze deadline sits comfortably past the honest plan's 45 days. It sits nowhere near the 91-day plan if split-schedule homes turn out rarer than assumed.

The second version of the plan didn't change the two-week default for early access or broad beta, both were already governed by the one-week floor either way. What changed was the internal phase: fifteen days, not ten, sized to the actual arithmetic instead of the calendar habit. The freeze deadline still had eighteen days of room under the honest plan. Broad beta went live on schedule. No second near miss.

One design asked how many days had passed. The other asked how many bad nights the phase had actually produced. The first can look calm right up until forty thousand homes inherit the same quiet problem at once.

The thing I'd tell myself, back on that day-ten call: a phase that looks quiet isn't the same thing as a phase that's actually been tested. I'd counted the calendar. I should have been counting the nights.

BOUND, sized in home-days instead of a flat calendar

This is a sizing question about how many home-days it actually takes to trust a rare failure rate, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. How long a phase should run isn't one number. It's an equation: bad nights expected in a day equals the homes in that phase, times the share running an odd schedule, times how often the model gets one of those nights wrong while it's still learning that home. Days needed equals the bad nights we want to see, divided by that daily number.
O, own the numbers. For Northloft, about four percent of homes are split-schedule: a night-shift worker, someone home all day while the house empties, two people on different clocks. The model gets about one night in twenty wrong on those homes during its first two weeks. We want to see that twelve times before trusting the rate. At four hundred employee homes, that's about fifteen days.
U, use a range. If four percent holds, fifteen days for the internal phase is right, and the one-week floor governs every phase after it. If the real rate is closer to one percent, four times rarer, the internal phase alone needs about sixty days, because four hundred homes can't hand you twelve bad nights any faster than that. Every later phase has enough homes that the swing barely reaches it.
N, nail the sanity check. Marketing needs the full launch live before the first freeze, about sixty-three days out. The honest plan, fifteen plus seven plus ten plus fourteen, adds up to forty-five days, eighteen days of room. The rarer scenario runs to ninety-one days, twenty-eight days past the freeze.
D, direction. Two things could move this number, and they don't move it the same amount. Whether split-schedule homes are really as common as assumed swings the internal phase from fifteen days to sixty, a forty-six day swing on the whole plan. Recruiting more employee homes into that phase helps, doubling it from four hundred to eight hundred cuts the swing roughly in half, but it doesn't change whether the four percent guess was right in the first place. Rarity is the bigger lever. Recruiting is the smaller one, and it's the one people reach for first because it feels like doing something.

The build-up: days elapsed as the rollout climbs toward everyone
Internal complete14 days
Early access complete21 days
Broad beta complete31 days
Half rollout complete45 days
Northloft's freeze deadline sits at day sixty-three. The honest plan reaches everyone with eighteen days to spare. The rarer scenario, day ninety-one, would miss the freeze by twenty-eight days.
What moves the plan's length most (swing in total days from the 45-day baseline)
Split-schedule homes rarer than assumed (1% instead of 4%)+46
Confidence bar raised from 12 bad nights to 20+11
Recruit more homes into the internal phase (400 to 800)−6
Add a second reviewer, cut broad beta's buffer to the floor−3
Rarity swings the plan the most, and nearly all of it sits in the smallest phase. Recruiting and hiring both help, but neither one touches whether the four percent guess was right.

And if you want to be sure it really works, try it somewhere else

Verdant Row Co-op runs an AI system that decides when to open and close irrigation valves per field zone, reading soil moisture and the forecast, instead of a fixed timer a farmer sets once each season. The co-op is rolling it out from its own twenty research plots to all three thousand member farms.

B, break it down. Same shape, different work. Weeks needed equals the bad irrigation cycles we want to see, divided by farms in the phase times the share running mixed crops on adjoining zones times how often the system waters the wrong zone while it's still learning that farm's map.
O, own the numbers. About eight percent of member farms mix crops with very different water needs on neighboring zones, an orchard block next to a row-crop block. On those farms, the system gets about one cycle in ten wrong during its first two weeks, and cycles run twice a week. The co-op wants to see that eight times before trusting the rate. On the twenty research plots, that's about five weeks.
U, use a range. If eight percent holds, the honest plan reaches every farm in about ten weeks. If mixed-crop farms are really closer to two percent, four times rarer, the research-plot phase alone stretches past twelve weeks, because twenty plots don't hand you eight bad cycles any faster than that.
N, nail the sanity check. The co-op needs full coverage before spring planting, about sixteen weeks out. The common scenario uses ten weeks, six weeks of room. The rare scenario runs past nineteen weeks, three weeks past planting.
D, direction. Same tension as Northloft's thermostat. How common mixed-crop farms really are swings the plan by nine weeks. Adding a third agronomist to the rollout team adds real capacity, but it doesn't change how many mixed-crop farms are sitting in those twenty research plots.

Weeks to full coverage: common scenario vs rare scenario
Mixed-crop farms as common as assumed10 weeks
Mixed-crop farms four times rarer19 weeks
The planting deadline sits at sixteen weeks. The common scenario clears it with room. The rare scenario misses it by three weeks, the same shape as the thermostat's freeze deadline, a different crop, the same arithmetic.
Same shape, different lever At Northloft, split-schedule homes were the swing lever, and recruiting more employees only closed half the gap. At Verdant Row, mixed-crop farms play the same role, and a third agronomist only closes part of that gap too. The equation is identical. Which assumption to check first isn't obvious until you've done the arithmetic.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: size each phase from real home-days, check it against the actual deadline, don't run every phase on a company default.
Cost: the team can't add another reviewer this year. Don't cut the floor, cut the pace: space out the phases most likely to hit the rare pattern, rather than promoting on schedule.
The model got better: a newer version of the optimizer rarely mispredicts split-schedule homes anymore. That doesn't remove the need to check rarity, it just moves where the real risk sits, from split-schedule homes to whatever pattern the new model hasn't seen yet.

Where people run it wrong.
They give every phase the same length because it's the company default, without checking whether that phase's population can even produce the signal that fast.
They treat the calendar floor as the actual statistical answer, when the floor only matters once a phase is already big enough for the floor to be the binding constraint.
They reach for a new hire as the first fix, when the number actually breaking the plan is a rate nobody has measured yet.

How to use it live. Say the equation before naming a single number: "the phase length isn't the real question, the home-days behind it are, so before I set a calendar, I'd want to know how rare the failure pattern this is protecting against actually is." That buys the room to ask a real question instead of repeating a two-week habit that sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a rollout phase-length sizing question, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question, how many home-days it takes to trust a rare failure rate, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tevita Faleolo, who has run feature rollouts at Northloft, a home automation company, for three years. He decides when a phase has run long enough to promote to the next one.
3 · WHAT THE FIRST PLAN GOT WRONG
What did Northloft's original rollout plan assume that turned out to be the wrong basis for sizing a phase?
Tap to flip
ANSWER
It assumed a flat two weeks, the company's rollout default, would work for every phase, when only the small internal phase's home count actually needed that long to see the rare failure enough times.
4 · THE STRUCTURE IN THIS STORY
What's the difference between sizing a phase by the calendar and sizing it by home-days?
Tap to flip
ANSWER
The calendar habit asks how many days have passed. Home-days sizing asks how many times the rare failure has actually shown up. A phase can look calm on the calendar while it hasn't produced enough real evidence to trust yet.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Setting every phase to a flat two weeks, because two weeks had worked for a dozen past launches. It made sense when past features didn't have a failure mode tucked inside one narrow slice of homes.
6 · THE NUMBER
Fill in the blank: the team wanted to see the rare failure at least ___ times before trusting it. That took about ___ days in the ___-home internal phase, and the plan's real freeze deadline sat ___ days out.
Tap to flip
ANSWER
12 times. 15 days. 400-home. 63 days out.
7 · THE REPLAY
Same rollout, same team, second design. What changes?
Tap to flip
ANSWER
The internal phase runs fifteen days instead of ten, sized to the actual arithmetic instead of the calendar habit. Broad beta still goes live on schedule, with eighteen days of room before the freeze deadline, and there's no second near miss.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same sizing question for a different product. Which product, and which lever swings it most?
Tap to flip
ANSWER
Verdant Row Co-op's irrigation-scheduling AI, rolling out to farms instead of homes. How common mixed-crop farms actually are swings the plan more than adding another agronomist does.

Check yourself Score: 0 / 0

Fill in the blank
1. To trust the split-schedule failure rate was real, the team wanted to see it happen at least ___ times. At the ___-home internal phase, that took about ___ days, but the company's default plan only gave it ___ days.
Show hint
Check the O and D steps, and the story's day-ten call.
Show answer
12 times; 400-home; 15 days; 10 days. That five-day gap, between what the arithmetic needed and what the habit gave it, is the whole margin the near miss lived inside.
Multiple choice
2. Why does the internal phase swing so much harder under the rarer scenario than early access or broad beta do, even though all three are checking for the same failure?
  • A. The internal phase has the fewest homes, so it takes the rare pattern the longest to show up in, while the bigger phases already have enough homes that the one-week floor decides their length instead.
  • B. The model runs a different version of the optimizer during the internal phase.
  • C. Employees get a special calibration setting the rest of the customer base doesn't.
  • D. The internal phase only checks for the pattern on weekends.
Show hint
Compare which number, home count or the floor, actually decides each phase's length.
Show answer
A. Once a phase has enough homes, the one-week floor becomes the real constraint, not the arithmetic. The internal phase is the only one small enough that the arithmetic still wins.
True or false
3. True or false: because the internal phase looked quiet for ten days, it had already produced enough evidence to trust the failure rate and promote to early access.
  • True
  • False
Show hint
Compare "looked quiet" against the actual count of bad nights, not the number of days that passed.
Show answer
False. Ten days of a quiet flag list isn't the same as twelve real bad nights. The phase had actually produced exactly the number needed by day eleven, nobody had counted them, they'd counted days instead.
Short answer
4. Tevita's manager says, "Just add a second reviewer and we can promote every phase faster." Why doesn't adding a reviewer shorten how long the internal phase actually needs to run?
Show hint
Compare what a reviewer changes against what actually produces a trustworthy rate.
Show answer
Model answer: A reviewer changes how fast flagged nights get looked at, not how many home-days it takes for twelve real bad nights to happen in a four-hundred-home population. Reviewing faster doesn't make split-schedule homes appear faster. Only more homes, or a truly higher failure rate, does that.
Short answer, apply it yourself
5. Think of a rollout or pilot you could run at your own job, scaling from a small test group to everyone. What real, rare failure would you want to see happen a set number of times before trusting each phase's numbers, and what would decide how long that takes?
Show hint
Name something that only shows up in a narrow slice of users or cases, not something that happens evenly across everyone.
Show answer
Model answer: A customer support team piloting an AI reply-drafting tool for ticket responses, scaling from five agents to sixty. The rare failure would be the tool drafting a confidently wrong answer on a ticket type it rarely sees, maybe a refund tangled with a shipping delay. I'd want to see that happen at least ten times before trusting the rate, and how long that takes depends on how many of that rare ticket type each phase's agents actually handle in a week.
Short answer, the number question
6. If the confidence bar were raised from twelve bad nights to twenty for every phase, not just the internal one, would broad beta's ten-day plan still hold? Show the math.
Show hint
Recompute broad beta's forty-thousand-home phase at the higher threshold and compare it to the one-week floor plus the three-day review buffer.
Show answer
Yes, it would still hold. Broad beta's forty thousand homes produce twelve bad nights in well under a day at the original rate, so even twenty bad nights would still arrive in under a day. The ten-day plan was never set by the arithmetic there, it was set by the one-week floor plus a three-day review buffer, and raising the confidence bar doesn't touch either of those.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more