InterviewAdvancedShipping & Model Lifecycle / Rollout strategy and phased launches / #22

Present the rollout plan for a launch I describe.

The direct answer
Size each phase by how many confirmed real violations of that phase's newest edge-case category you need to see before trusting the model on it, not by a flat percentage of the company. Check the resulting flag volume against how many reports the finance team can actually hand-review each week. If a category turns out rarer than assumed, stretch that phase instead of promoting a category nobody has actually seen enough real examples of yet.
Do this, in order
  1. Size each phase by confirmed real catches of its newest edge-case category, not by what percent of the company that phase covers.Why: a phase length with no arithmetic behind it is a guess wearing a schedule.
  2. Hold every phase to at least a one-week floor, even when the math says less would do.Why: a floor shorter than a week can promote a phase before a weekday and a weekend both show up in the reports.
  3. Check the flag volume each phase would produce against how many reports the finance team can actually hand-review in a week.Why: a plan the reviewers can't keep up with turns into a growing queue, not a launch.
  4. Give the plan a range, low if the rarest new category is as common as assumed, high if it isn't.Why: one number hides a fifteen-week swing between the two ends.
  5. Watch the phase introducing the rarest new edge case hardest, not the biggest phase.Why: population size doesn't decide how long a phase needs, the rarest thing inside it does.
  6. When the honest number and the review capacity disagree with the calendar, stretch the phase, never quietly skip the floor to hit a date.Why: skipping the floor is how a category gets trusted before the finance team has actually seen it.

How to answer this, stage by stage

Nobody is grading whether you land on exactly ten weeks. They're grading whether you can defend the arithmetic behind "three weeks for phase one," whether the range is honest, and whether you close on something the room can check. Seven moves get you there.

1
Scope it to one real company and one starting department
Say it like this
"Let's ground this. Say Corriedale Textiles, a wool and textile manufacturer with about eleven hundred employees, just built an AI tool that reads every expense report the second it's filed and flags the ones that look like they break policy, before anyone gets reimbursed. Corriedale's finance team is six people. A phased rollout means starting the flagger in one department and expanding it, not turning it on for all eleven hundred people at once."
Why this works
Stops the answer from staying abstract before a single number gets attached to it.
2
Say the structure out loud before naming any numbers
Say it like this
"Here's how I'd frame it. The question isn't what percent of the company gets the tool each week. It's whether the finance team has actually seen enough real examples of that department's kind of policy break to trust what the model is telling them about it. Different departments break policy in different ways, so each phase is really asking the model to earn trust on a new category, not just asking it to reach more people."
Why this works
Reframes the question before touching a number, so the interviewer knows a real method is coming, not a percentage habit.
3
Break down the equation
Say it like this
"Here's the shape. Weeks a phase needs equals the confirmed real violations we want to see in its newest category, divided by the expected confirmed violations per week. And expected confirmed violations per week equals the reports coming from that phase, times the share that fall into the risky spend type, times the share of those that turn out to be an actual violation and not an explainable one."
Why this works
This is the B step of BOUND, the equation stated before a single number touches it.
4
Own the real numbers, phase by phase
Say it like this
"Phase one is Sales and Client Relations, two hundred people, about four hundred reports a week. Thirty percent of those are travel and client-entertainment claims, the ones prone to getting split into two smaller charges to duck the five-hundred-dollar pre-approval line. About five percent of those are real splits, so six a week. We want to see eighteen confirmed ones before trusting the model on that category. That's three weeks. Phase two adds Regional Ops and Marketing, two hundred fifty more people, and a new pattern, a client gift folded into a meal claim. That one's rarer, so it needs five weeks. Phase three is everyone else, and by then the one-week floor, not the arithmetic, decides it. Call it two weeks for safety."
Why this works
This is the O step, real proposed numbers, not "we'll monitor it closely."
5
Give the range, not one number
Say it like this
"If eight percent of Marketing and Ops's reports are really gift-adjacent, phase two takes five weeks. If the real number is closer to two percent, four times rarer, phase two alone stretches to twenty weeks, and the whole plan runs ten weeks on the low end, twenty-five on the high end. Phase one and phase three barely move, their populations are big enough that the swing doesn't reach them the same way."
Why this works
A single number here would claim a confidence the plan doesn't actually have yet.
6
Sanity check the plan against what the finance team can actually review
Say it like this
"Six reviewers, each able to spare about half an hour a day for this on top of their regular accounts-payable work, that's nine hundred minutes a week, team-wide. My phased plan peaks around four hundred eighty minutes a week, comfortably inside that. Compare that to what the standard playbook would have done: jump straight to twenty-five percent of the company, spread evenly across every department at once, before any single category is calibrated. That version runs about eleven hundred minutes a week. The team can't sustain that, and the queue just grows."
Why this works
This is the N step, and it's the step a flat percentage rollout never runs.
7
Name the biggest lever, then close on the one line
Say it like this
"If I had to bet on one thing, it's whether the gift-mixing pattern is really as common as I assumed, not the review capacity. Losing a reviewer barely moves this plan, because the phased design never gets close to the ceiling. Getting the gift-pattern rate wrong adds fifteen weeks by itself. So here's what I'd actually say: size each phase by confirmed violations of its newest category, hold the one-week floor no matter what, and check the flag volume against real review capacity before promising a date. If the honest number and the capacity check disagree with the calendar, I'd rather stretch a phase than promote a category the finance team hasn't actually seen enough of yet."
Why this works
Closes on the literal ask, a claim someone could check, not a vibe about moving carefully.
If you remember one thing A flat "some percent, then more, then everyone" schedule sizes a phase by how many people it reaches. The real question is how many confirmed catches of that phase's newest violation type it takes before the model earns trust on it, and that number depends on the category, not the calendar.

Let's learn

Before this tool existed, six people on Corriedale's finance team read every one of the roughly eight hundred expense reports that come in each week, by hand, checking each one against a policy binder nobody had fully memorized. That took most of two working days, every week, and it still missed things, because nobody can hold four different kinds of policy exception in their head report after report.

Now the tool reads all eight hundred in about a minute and flags the roughly forty a week that look wrong. The team reviews forty instead of eight hundred.

Knowledge spark: what's a split-transaction violation? Breaking one expense that should need pre-approval into two or more smaller charges, so each one slides in under the five-hundred-dollar line on its own. A fourteen-hundred-dollar client dinner filed as three separate charges of under five hundred, two days apart, is a split transaction.

The real risk isn't the extra speed. It's what happens the first time a new department introduces a violation pattern the model has only seen a handful of real examples of, and the finance team either waves through something real or drowns in something the model is still guessing at.

We didn't lose two weeks. We lost the eighteen catches those two weeks were supposed to produce.
The decision that mattered Size every phase by how many confirmed real violations of its newest category the team has actually seen, not by whatever percent of the company the company's playbook happens to schedule next. That's the one check that decides whether a category gets trusted at two hundred people or at eleven hundred.

At its worst, a category nobody has actually earned trust on rolls out to all eleven hundred employees at once. A rare pattern that's shown up six times a week in Sales shows up sixty times a week company-wide, and the finance team, buried in flags they can't yet tell apart from noise, starts waving them through unread. The tool built to catch missed violations ends up hiding them behind a stamp of approval nobody has time to actually check.

The choice I would take back. Corriedale's rollout playbook, built for a dozen past internal tools, a new expense-report screen, a new time-tracking app, gives every phase two weeks, then doubles the population, then opens it to everyone. That felt careful. It also assumed every new tool changes a workflow, not a judgment. None of those past tools ever had to earn trust on a specific kind of mistake before more people saw it.

What I would leave alone. The simple flags, a missing receipt, a charge with no attached approval, don't need this staged confidence-building. The model has been reliably right about those since week one, because every department produces that kind of report the same way. Turn those on company-wide, day one. Only the categories the model hasn't seen enough of yet need the staged rollout.

The lesson. How long a phase should run was never really a calendar question. It's a question about whether the team has actually seen the thing they're being asked to trust, and Corriedale's two-week habit had been quietly answering a question nobody had actually asked.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why phase one, the one everyone assumed was already done, almost wasn't.

Every Friday afternoon, Renato Aguirre pulls the week's flag count before anyone else sees it. He's run rollouts of internal tools at Corriedale Textiles for four years, and he can usually tell within an hour of a week's numbers coming in whether something is actually wrong or just noisy.

The expense flagging tool was the biggest thing on his roadmap that year. It started with Sales and Client Relations, two hundred people, the department with the most travel and the most client dinners, and so the most ways to bend a receipt.

For the first ten days it looked clean. A steady trickle of flags, a few dozen a week, nothing that lined up into a pattern. Corriedale's rollout playbook, the same one used for a screen redesign two years earlier, called for two weeks per phase, then double the department, then open it to everyone. Leadership wanted Marketing and Regional Ops added by day fourteen, and by day ten the numbers already looked quiet enough to promise it.

Renato started phase one checking every flag by hand with the team. By day six he was skimming the weekly summary instead. By day ten, with the promotion date circled and the numbers calm, nobody on the finance team was cross-checking the split-transaction flags against each other anymore, just clearing them one at a time.

On day thirteen, one of the reviewers noticed something while closing out an unrelated report: a client dinner from week two, fourteen hundred dollars, had gone through as three separate charges of under five hundred each, filed two days apart by the same sales rep. None of the three had been flagged. Not because the model missed the pattern outright. It hadn't yet seen enough real splits to be confident calling one, so it had stayed quiet on the borderline cases rather than flood the team with guesses.

She wasn't alarmed. She mentioned it in passing, the way you'd mention a typo.

It wasn't the missed charge that stopped Renato. It was realizing the count.

Only eleven confirmed splits by day thirteen, not the eighteen the arithmetic actually needed before the model could be trusted on that category. Phase one had never been ready. The calendar said ready. The numbers didn't.

Renato pulled the promotion. Phase one ran eight more days, to day twenty-one, and picked up the seven more confirmed splits it needed along the way, eighteen total, right on the arithmetic. Only then did Marketing and Regional Ops go live.

Three weeks in, not two, isn't a disaster on its own. What stayed with Renato was what would have happened on the version where he hadn't pulled the promotion. The gift-mixing pattern Marketing was about to introduce ran even rarer than split transactions. If phase one had opened on schedule at day fourteen, phase two would have layered an even less-tested category on top of a barely-tested one, in front of four hundred fifty people instead of two hundred, the same week leadership's board update was due.

The team never had a real number for how long phase one needed. They had a habit, two weeks, that had worked for tools that only ever changed a screen. Eleven days in, with the flags quiet, the habit said done. The arithmetic, run properly, said phase one actually needed three weeks to see the category the required eighteen times, and that gap, ten days, was the whole margin the near miss lived inside.

Weeks earlier, when the rollout plan first got signed off, someone on the call had asked whether two weeks was really enough for something that had to learn a pattern, not just a new screen. Renato said it had always been enough before. Nobody pushed back. It was the only number in the room.

Hand-sketched number line running from 0 to 25 weeks. An amber flag marked at 6 weeks labeled the old two-week-per-phase habit. A green dot marked at 10 weeks labeled the honest plan. A red-orange dot marked at 25 weeks labeled if vendor-gift cases are 4x rarer than assumed.
The old habit would have promised six weeks. The honest arithmetic runs ten. If the gift-mixing pattern turns out four times rarer than assumed, it runs twenty-five.

The second version of the plan didn't touch the floor under phase three, that was already governed by review capacity either way. What changed was phase one: three weeks, not two, sized to the confirmed count instead of the calendar habit. Phase two, watching its own rarer category from the start this time, ran its full five weeks before Renato let anyone add the rest of the company. No second near miss.

One design counted days. The other counted catches. The first can look calm right up until four hundred fifty people inherit a category nobody actually finished testing.

The thing I'd tell myself, back on that day-ten call: a phase that looks quiet isn't the same thing as a phase that's actually been proven. I'd been counting the calendar. I should have been counting the catches.

BOUND: what actually decided Corriedale's ten weeks

This is a sizing question about how many confirmed catches it actually takes to trust a rare category, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. How long a phase should run isn't one number. It's an equation: expected confirmed violations in a week equal the reports in that phase, times the share that fall into the risky spend category, times the share of those that are a genuine violation. Weeks needed equal the confirmed catches we want to see, divided by that weekly number.
O, own the numbers. Phase one, Sales, two hundred people, four hundred reports a week. Thirty percent are travel and client-entertainment claims. Five percent of those are real split transactions, six a week. We want eighteen confirmed before trusting the category. That's three weeks. Phase two adds two hundred fifty people and a rarer pattern, gift-mixing, five weeks. Phase three, the rest of the company, is governed by the one-week floor, call it two weeks for safety.
U, use a range. If eight percent of phase two's reports are gift-adjacent, five weeks holds. If the real rate is closer to two percent, four times rarer, phase two alone stretches to twenty weeks, because two hundred fifty people can't hand you eight confirmed gift-mixing cases any faster than that. The whole plan runs ten weeks on the low end, twenty-five on the high end.
N, nail the sanity check. Six reviewers, each able to spare about thirty minutes a day for this, is nine hundred minutes a week team-wide. The phased plan peaks at four hundred eighty minutes a week, well inside that. A flat jump to twenty-five percent of the company, mixed across every department before any category is calibrated, runs about eleven hundred minutes a week, past what the team can sustain.
D, direction. Two things could move this number, and they don't move it the same amount. Whether gift-mixing is really as rare as assumed swings phase two from five weeks to twenty, a fifteen-week swing on the whole plan. Losing one reviewer barely touches it, the phased plan's peak load sits well under capacity either way. Rarity is the bigger lever. Headcount is the smaller one, and it's the one people reach for first because it feels like doing something.

The build-up: weeks elapsed as the rollout reaches everyone
Phase one complete (Sales)3 weeks
Phase two complete (+ Ops, Marketing)8 weeks
Phase three complete (everyone)10 weeks
Phase one holds to three weeks because that's what eighteen confirmed splits actually take. Phase two is the plan's real bottleneck, five weeks, the smallest true-case rate of any phase.
What moves the plan's length most (swing in weeks from the 10-week baseline)
Gift-mixing cases rarer than assumed (2% instead of 8%)+15
Phase one's confidence bar raised from 18 catches to 30+2
One fewer reviewer during phase two's peak, confirmations lag+1
Recruit more Sales travelers into phase one early (200 to 400)−1.5
Rarity swings the plan by fifteen weeks on its own, far more than losing a reviewer or raising the confidence bar. Recruiting more travelers into phase one helps a little, but it doesn't touch whether the eight percent guess about gift-mixing was right in the first place.

And if you want to be sure it really works, try it somewhere else

Tallowmere Mutual, a regional home and auto insurer, runs an AI tool that reads incoming claims and flags ones that look like a policy violation, before a payout goes out. The pilot team is eight adjusters on the auto desk, out of forty across the claims department.

B, break it down. Same shape, different work. Weeks needed equal the confirmed catches we want to see for a category, divided by claims in the phase times the share touching more than one policy rider times how often that turns out to be a genuine double-filing.
O, own the numbers. The pilot team handles three hundred claims a week. Five percent touch more than one rider within ten days, fifteen a week. Eight percent of those are genuine double-filings, about one a week. The team wants six confirmed catches before trusting the rate. That's about five weeks.
U, use a range. If the real double-filing rate is closer to two percent, four times rarer, the pilot phase alone stretches past nineteen weeks.
N, nail the sanity check. Eight adjusters, each able to spare twenty minutes a day for this, is about eight hundred minutes a week team-wide. The pilot's flag volume runs well under that, even the rare-case scenario doesn't breach it, the wait is in weeks, not review minutes.
D, direction. Same tension as Corriedale's expense tool. How common the double-filing pattern really is swings the plan by fifteen weeks. Adding a ninth adjuster adds real capacity, but it doesn't change how many double-filings are sitting in those three hundred weekly claims.

Weeks to trust the double-filing pattern: common scenario vs rare scenario
Double-filings as common as assumed5 weeks
Double-filings four times rarer19 weeks
Same shape as the expense tool's phase two, a different report, the same arithmetic, and the same fifteen-week swing sitting on the same kind of guess.
Same shape, different lever At Corriedale, gift-mixing rarity was the swing lever, worth fifteen weeks on its own. At Tallowmere, double-filing rarity plays the same role, worth about the same swing, on a completely different kind of report. The equation doesn't change. Which assumption to check first isn't obvious until you've actually run the arithmetic.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: size each phase by confirmed real catches of its newest category, check the flag volume against real review capacity, don't run every phase on a company percentage.
Cost: the team can't add another reviewer this year. Don't cut the floor, cut the pace: hold the rarer-category phase longer instead of promoting it on schedule.
The model got better: a newer version rarely misses the rare category anymore. That doesn't remove the need to check rarity, it just moves where the real risk sits, from today's rare category to whatever pattern the new model hasn't seen yet.

Where people run it wrong.
They give every phase the same length because it's the company default, without checking whether that phase's population can actually produce the confirmed-catch count that fast.
They treat the one-week floor as the real statistical answer, when the floor only matters once a phase is already big enough for the floor to be the binding constraint.
They reach for more headcount as the first fix, when the number actually breaking the plan is a rarity rate nobody has confirmed yet.

How to use it live. Say the equation before naming a single number: "the phase length isn't the real question, how many confirmed catches of its newest category it takes to trust the model is, so before I set a calendar, I'd want to know how rare that pattern actually is." That buys the room to ask a real question instead of repeating a two-week habit that sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a "present the rollout plan" question, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question, how many confirmed real catches of each department's newest violation type it takes to trust the model, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Renato Aguirre, who has run internal tool rollouts at Corriedale Textiles, a wool and textile manufacturer, for four years. He decides how fast each phase expands to the next department.
3 · WHAT THE FIRST PLAN GOT WRONG
What did Corriedale's original rollout plan assume that turned out to be the wrong basis for sizing a phase?
Tap to flip
ANSWER
It assumed a flat two weeks per phase, the company's default for past internal tools, would work here too, when only the confirmed-catch count for that phase's newest violation type actually decided whether it was ready.
4 · THE STRUCTURE IN THIS STORY
What's the difference between sizing a phase by the calendar and sizing it by confirmed catches?
Tap to flip
ANSWER
The calendar habit asks how many days have passed. Confirmed-catch sizing asks how many real violations of that phase's newest category have actually been seen and checked. A phase can look quiet on the calendar while it hasn't produced enough real evidence to trust yet.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Setting every phase to a flat two weeks, because two weeks had always worked for past internal tools. It made sense when those tools only changed a screen. It stopped making sense once a tool had to earn trust on a specific kind of mistake before more people saw it.
6 · THE NUMBER
Fill in the blank: phase one wanted to see ___ confirmed split-transaction catches before trusting the model on that category. That took about ___ weeks in the ___-person Sales department, and only ___ had actually been confirmed by day thirteen.
Tap to flip
ANSWER
18 catches. 3 weeks. 200-person. 11 confirmed.
7 · THE REPLAY
Same rollout, same team, second design. What changes?
Tap to flip
ANSWER
Phase one runs three weeks instead of two, sized to the confirmed-catch count instead of the calendar habit. Phase two runs its full five weeks before Marketing and Regional Ops get added, and there's no second near miss at four hundred fifty people.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same rollout-sizing question for a different product. Which product, and which lever swings it most?
Tap to flip
ANSWER
Tallowmere Mutual's claims-fraud flagging tool, rolling out to insurance adjusters instead of a finance team. How rare the double-filing pattern actually is swings the plan more than adjuster headcount does.

Check yourself Score: 0 / 0

Fill in the blank
1. Phase one wanted to see ___ confirmed split-transaction catches before trusting the model on that category. Only ___ had been confirmed by day thirteen, against a company default of ___ weeks per phase.
Show hint
Check the O step and the story's day-thirteen moment.
Show answer
18 catches; 11 confirmed; 2 weeks. That seven-catch gap is exactly why Renato pulled the promotion instead of letting it go through on schedule.
Multiple choice
2. Why does phase two swing so much harder under the rarer scenario than phase one or phase three do, even though all three are checking for a real policy violation?
  • A. Phase two's gift-mixing pattern starts out rarer than phase one's split-transaction pattern, so it takes longer to build confirmed catches, while phase three's bigger population and the one-week floor already govern its length.
  • B. Phase two runs an older version of the flagging model.
  • C. Marketing employees get a separate review process the rest of the company doesn't.
  • D. Phase two only checks flags on weekdays.
Show hint
Compare which number, catch rate or population size, actually decides each phase's length.
Show answer
A. Once a phase's population is big enough, the floor or its own arithmetic settles quickly. Phase two is the one place a rarer-than-assumed pattern still has room to stretch the plan.
True or false
3. True or false: because phase one's flag count looked quiet for ten days, it had already produced enough confirmed splits to trust the category and promote to phase two.
  • True
  • False
Show hint
Compare "looked quiet" against the actual count of confirmed splits, not the number of days that passed.
Show answer
False. Ten quiet days isn't the same as eighteen confirmed real splits. Phase one had only eleven by day thirteen, nobody had counted them, they'd counted days instead.
Short answer
4. Renato's leadership says, "Just add a second reviewer and we can promote every phase faster." Why doesn't adding a reviewer shorten how long phase one actually needs to run?
Show hint
Compare what a reviewer changes against what actually produces a confirmed catch.
Show answer
Model answer: A reviewer changes how fast flagged reports get checked, not how many split-transaction cases actually happen in a two-hundred-person department in a given week. Reviewing faster doesn't make the rare pattern show up faster. Only more travelers, or a genuinely higher violation rate, does that.
Short answer, apply it yourself
5. Think of a rollout or pilot you could run at your own job, scaling from a small test group to everyone. What real, rare mistake would you want to see happen a set number of times before trusting each phase's numbers, and what would decide how long that takes?
Show hint
Name something that only shows up in a narrow slice of cases, not something that happens evenly across everyone.
Show answer
Model answer: A customer support team piloting an AI tool that drafts refund decisions, scaling from five agents to fifty. The rare mistake would be the tool confidently approving a refund on an order type it rarely sees, say a bundled subscription plus a one-time add-on. I'd want to see that happen at least ten times before trusting the rate, and how long that takes depends on how many of that order type each phase's agents actually handle in a week.
Short answer, the number question
6. If the confidence bar for phase one were raised from eighteen confirmed splits to thirty, would the honest ten-week plan still hold? Show the math.
Show hint
Recompute phase one's true weekly rate against the new target, then add the difference back into the total.
Show answer
No, not quite. At six confirmed splits a week, thirty catches take five weeks instead of three, two weeks longer. The honest plan would run twelve weeks instead of ten. Phase two and phase three don't change, the higher bar was only set for phase one's category.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more