CaseAdvancedShipping & Model Lifecycle / Rollout strategy and phased launches / #1
Design the rollout plan for an AI feature going to two million users.
The direct answer
Phase the rollout by how well the assistant's assumptions match the people arriving in each phase, not by a fixed rollout percentage. And put a required look-before-you-accept step back in front of any booking the pilot group never tested: more than one traveler on it, or a connection under the minimum time. The assistant does not fail at two million users because it gets less accurate. It fails because two million ordinary travelers do not look like the forty thousand who tested it, and the one-tap habit the pilot built in people follows them straight into the day it stops being safe.
Do this, in order
Put a required comparison step back for any booking with more than one traveler, or a connection under the minimum time.Why: this is the fix. It restores the pause the pilot's own small size used to provide for free, without slowing down the millions of cases the tool actually gets right.
Phase the rollout by population match, not by headcount.Why: two hundred fifty thousand new users who fly like the pilot group teach you nothing the pilot didn't already know. The real test is the people who don't.
Watch the share of proposals going to more than one traveler, phase by phase.Why: that number climbs for weeks before a single complaint does. It tells you the moment your test group stops looking like your real one.
Leave the one-tap Accept alone for solo travelers with a normal connection.Why: that is exactly the case the model is good at. Slowing it down there costs millions of people time for nothing.
Do not fix this with a human reviewer or a stricter confidence score.Why: both are new dials sitting on top of the old decision. Neither one is the decision that actually broke.
How to answer this, stage by stage
Seven moves. This is a case question at a size most interview stories never reach, so the hard part is saying what changes with size, not what stays the same. Each stage has the words you'd actually say.
1
Scope the huge word down to one screen
Say it like this
"'Rollout' is too big to answer directly, so let me pick one piece of it. I'm going to design the moment a flight gets cancelled and the assistant proposes a new one. That's the screen where this whole thing either works or doesn't."
Why this works
A huge word like "rollout" has no shape until you pick one inspectable moment inside it. This is that moment.
2
Say your five-part structure out loud
Say it like this
"I'll cover five things. Who's actually receiving this. What habit the tool builds in them. Where that habit snaps once we're at two million people instead of forty thousand. Which old decision I'd take back. And what the same bad day looks like once it's fixed."
Why this works
Two seconds of structure stop you rambling through a huge question and tell the interviewer you have a plan for it.
3
Reframe what "rollout" is actually testing
Say it like this
"The real question here isn't whether the model is accurate enough to go wide. It's what quietly worked because the pilot group was small and self-selected, and stops working the day it's ordinary people at ordinary scale."
Why this works
Separates you from a candidate who lists rollout percentages and rollback triggers without ever asking what the pilot group was hiding.
4
Give the one decision, and only one
Say it like this
"Here's what I'd actually do. Ship the fast, one-tap accept to the population it's tuned for. But bring back a required look-before-you-accept step for anything the pilot never tested: more than one traveler on the same booking, or a connection under the minimum time."
Why this works
This is the direct answer, said out loud. Small and specific beats "add more monitoring," which nobody can picture.
5
Prove it with the failure
Say it like this
"Say Selin was in the pilot for seven months, always flying alone for work, and the assistant never once got it wrong. She stops reading the comparison card. Then she flies with her mother for a family wedding, their connection gets cancelled, and the Accept button she doesn't even look at anymore books them onto two different flights, three hours apart, into a city neither of them knows."
Why this works
Four sentences, and it's the exact spot where a habit built by a small, forgiving group met a case it was never built for.
6
Name the number you'd watch before the next phase
Say it like this
"Before I widen the rollout further, I'd watch the share of proposals going to more than one traveler, phase by phase, not just the overall accept rate. That number tells you when your test population stops looking like your real one."
Why this works
Shows you think past launch day, and it names the number that moves early instead of the one everyone checks after the fact.
7
Land the whole answer in one breath
Say it like this
"So: phase the rollout by who the product has actually been tested on, not by headcount, and put the comparison step back for the one case the pilot group never had. That's how it survives meeting two million people who don't look like the forty thousand who tried it first."
Why this works
Restates the decision and why, in one breath. That's the line an interviewer remembers.
If you remember one thing
Stages 3 and 5 are what the interviewer is really grading. Reframe "rollout" as "who did we actually test this on," then prove it with one specific person meeting one specific case the pilot group never had. Everything else in the answer is proof.
Let's learn
Here is what happens when a system that always worked meets exactly the group it was never tested on.
Say an airline builds a helper that lives in its app. When your flight gets cancelled or badly delayed, it looks at what's left and proposes a new flight for you, right there on your phone.
The number that barely moved
Before the assistant, a cancelled flight meant calling the airline or finding a gate agent. Selin, who flies for work about twice a month, once waited forty minutes on hold for a rebooking, twice in the same year. With the assistant working the way it was built to, that same cancelled flight turns into a new boarding pass in about ninety seconds. Tap the card, done.
The airline ran this with a pilot group first: forty thousand of its most frequent flyers, people who opted in, for seven months. Then it went to general availability, all two million app users, in three widening phases.
Rollout, as a share of the two million
From forty thousand people who chose to be tested on, to two million who never signed up for anything. Once every phase is live, that's about 120,000 rebooking-eligible itineraries a month across the whole app.
Knowledge spark: what's a minimum connection time?
The shortest time the airport says you need to get from one gate to another to make a connecting flight. Sometimes it's twenty minutes. In a big, spread-out airport, it can be over an hour.
Here's what the pilot group never showed anyone. Frequent flyers who fly for work mostly fly alone. The forty thousand people in the pilot were the airline's most reliable, most solo travelers. General availability is a different crowd entirely: families, couples, people flying with their kids or their parents.
Solo, or traveling with someone: pilot cohort vs. GA population
Pilot cohort
Flying solo
91%
With a companion
9%
GA population
Flying solo
54%
With a companion
46%
Pilot cohortGA population
Nine percent of the pilot ever needed the assistant to keep two travelers together. Forty six percent of the real population does.
The assistant did not get slower or less accurate at two million users. It got faster at giving the wrong answer to exactly the people who needed it to slow down and ask.
At its worst, that's worse than never building the tool. A tired gate agent, calling out two names at the counter, pauses and asks, "are you two traveling together?" The assistant never asks. It answers in ninety seconds, and the answer can be wrong in a way a person would have caught on instinct, without even trying.
The decision that mattered
Bring back the required comparison card for any itinerary with more than one traveler, or a connection under the minimum time. Not a human reviewer. Not a stricter confidence score. The pause the pilot's own small size gave us for free, built back in on purpose.
What I would leave alone. Solo travelers with a normal connection are exactly the case the pilot tested, ninety one times out of a hundred. The one-tap Accept stays for them. Slowing that down to protect against a case that almost never touches them would cost millions of people time for nothing.
The lesson. We tested the assistant on the group most likely to forgive it, and least likely to ever need it to think about anyone but themselves. That group proved the model works. It never had the chance to show us who the model was missing.
Now here is the same thing as a story
Use this version when you have room to let it land, not just list it.
The rebooking card lives in an app Selin already had open, one tap behind the boarding pass she pulls up forty times a year. She's a marketing director for a regional distributor, and she flies for client visits about twice a month, the kind of traveler an airline builds a whole loyalty tier around.
She joined the pilot in March. The first few times a flight got cancelled, she read the whole card: her old flight on the left, the proposed new one on the right, gate and time and seat all lined up so she could check it before tapping "Use this." It was always right. By June she was skimming the card, glancing at the new departure time and tapping through. By August she wasn't opening it at all. The notification would land, and her thumb would tap Accept before she'd even read the airline's name on it.
Competent alone, the good months, the habit thinning, the trigger
It had never once been wrong. Not in five months, not in nine cancelled flights. So why would this one be different.
In October, general availability finished rolling out, and Selin's mother flew with her for the first time, for her niece's wedding in Portland. Chicago to Denver to Portland, connecting through a hub that gets thunderstorms most October afternoons.
Thursday, 4:50pm. A ground stop at Denver. Their connecting flight, and about six thousand other people's, cancelled at once.
The notification landed the way it always did. Selin's thumb was already moving.
We did not get her flight wrong. We put her mother down alone, in a city neither of them had ever seen.
The assistant had found the fastest way to get one traveler named Selin Aydin to Portland that evening. It was a good answer, for one traveler. It put her on a 6:40 connection through Salt Lake City. It put her mother, booked separately in the same reservation but not read as "this person is with her," on the next open seat anywhere: a 9:15 through Phoenix, landing three hours later, alone, at a gate her mother had never stood at before.
People are switches, not dials
I want to say the problem was one wrong booking. It wasn't, not really. Selin never had a percentage in her head. She had a feeling, and the feeling only had two settings: read the card, or trust the button. There was no setting in between, no "skim it a little more carefully this time." Once the tool had been right for seven months straight, the only lever left was whether she opened the card at all.
So here's the decision I'd take back. In July, once pilot satisfaction passed ninety eight percent with almost no complaints, someone in a planning meeting pointed out that the comparison card felt like a leftover from when the model needed watching. Nobody had asked for a mandatory extra tap on something that was reliably right. It looked like friction with no upside. So the default changed: one big Accept button, the comparison tucked under a small "view details" link almost nobody tapped.
That was a sensible call in July. A comparison card built for a model that still needed checking is, in fact, clutter once the model stops needing checking. It stopped being sensible the day two million people, most of them not traveling alone, started tapping that same button.
I'd put the comparison back, not for everyone, just for the itineraries the pilot group never had: more than one traveler on the reservation, or a connection under the minimum time. If Selin's booking gets flagged the moment it includes her mother, she sees the two options side by side, taps the one that keeps them on the same connection, and they land in Portland together by 6:40, not three hours apart in two different airports.
That's the whole difference. One design hands two million people a single button and asks them to trust it the way forty thousand hand-picked people learned to. The other hands the ninety one percent who fly alone the same fast button, and gives the other nine percent one honest look before it moves them.
And the part I'd tell myself, if I could go back: we measured whether the flight the assistant picked was a good flight. We never asked whether it was picking for one traveler when two were standing at the same gate.
The five steps, sized for two million people
This is a Perturbation question wearing a design question's clothes: something changed size, not shape, and FLIPS runs straight down the line, F to S.
FLIPS, in five rows
FFind the person
Whose morning is this?
Not "GA users." One pilot member, one booking, one Thursday afternoon.
In this answer: Selin Aydin, in Windrose Airlines' rebooking pilot for seven months, flying with her mother for the first time.
LLocate the habit
What did they stop doing because it worked, and what did the pilot's own size quietly cover for?
The habit is the product working. What the pilot group's small, solo-heavy shape kept off the assistant's desk.
In this answer: She stopped reading the comparison card. The pilot never showed the tool a party of more than one traveler often enough to matter.
IIdentify the flip
What verb snaps, with no middle setting?
Not "more errors." A specific behavior with exactly two settings, and no drift back once trust is gone.
In this answer: Reads every proposed flight before accepting it, or taps Accept without reading it at all.
PPinpoint the old decision
Which choice only made sense before?
Small, specific, reasonable at the time. Never "add more review."
In this answer: Replacing the required comparison card with a single one-tap Accept button, once pilot accuracy passed ninety eight percent.
SShow the replay
Same bad day, new design. Better ending?
Run the same trigger through the fixed product. End on something you can count.
In this answer: The flag catches the two-traveler booking. Selin and her mother land in Portland together, by 6:40, instead of three hours apart.
A small move in the model. A big snap in what she does with her thumb.
Why I is the hard step
Anyone can say "she got careless." The hard part is naming what her thumb did differently at 4:50pm than it did in March, and proving there's no setting in between. "Skimmed the card" is a dial. "Tapped Accept without opening it" is a switch. If your flip has a middle, keep looking.
And if you want to be sure it really works, try it somewhere else
A hospital's chest X-ray tool flags scans by severity for a radiologist's second read. Same question shape, a different product, a different flip entirely.
Only one letter changes
F. Kenji Sato, staff radiologist at a mid-size hospital, eleven years reading chest films, about ninety a shift. L. He stopped writing down his own first impression before checking the tool's severity score. The two agreed often enough that his own note started to feel like paperwork. I. A different flip. He doesn't check less or delegate the read. He starts keeping a private spreadsheet of patients he's personally worried about, cross-checking each new flagged film against his own notes by hand, off the record. No middle setting: either the hospital's system tracks a patient over time, or Kenji does it himself on paper. P. We built the tool to score each film on its own, with no link back to the same patient's earlier scans, because at launch it only had to handle a single read with no history to draw on. S. Give the tool a patient's last three scans, and let it flag a small nodule that's individually unremarkable each time but trending up across visits. Kenji doesn't need his own spreadsheet. A hospital-wide chart audit finds the same slow-growing cases even for doctors who never built a private system at all.
A second decision worth taking back
Scoring every scan as if the patient has no past is a decision, not a limitation. Adding a memory of the same patient's last visit is a small, specific fix, and it replaces a shadow spreadsheet nobody signed off on.
Swap the trigger and it still runs
Speed: under two million users' worth of load, the proposal screen takes twenty seconds to load instead of two. People start hanging up and calling an agent before they ever see a comparison card. Same missing pause, a different mechanism.
Cost: to control costs at scale, the airline quietly shifts rebooking proposals onto a cheaper, less curated seat pool. Nobody widens the flagging rules to match. Worse options start moving through the same trusted one-tap path.
The model got better: the case on this page. Accuracy climbing through the pilot is exactly what got the comparison card removed in the first place.
Where people run it wrong
Blaming the model's accuracy instead of the screen that stopped asking anyone to look.
Fixing it with a mandatory second approval on every booking, which slows down the ninety one percent of cases that were never the problem.
Waiting for a viral complaint instead of watching the multi-traveler share of proposals, which was already climbing every week of the rollout.
How to use it live
Say the scale gap out loud before anything else. "So we went from forty thousand people who chose to be tested on, to two million who didn't." Naming that costs five seconds, and it isn't stalling, it's where the real answer starts, because "design the rollout" only means something once you've said who the pilot group actually was.
Flashcards (click a card to flip it)
Eight fixed slots, pulled straight from the answer above.
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Checks sometimes, then stops checking at all. It usually fires when the product is getting better, not worse, so nobody expects the snap.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Selin Aydin, a marketing director who flies for work about twice a month. Seven months in Windrose Airlines' rebooking pilot, always flying alone, never once had a bad proposal.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She stopped opening the two-flight comparison card before accepting a proposed rebooking. She started tapping Accept without reading it.
4 · THE FLIP, HERE
What's the two-setting switch in this story?
Tap to flip
ANSWER
Reads every proposed flight before accepting it, or taps Accept without reading it at all. No in-between setting once the mandatory comparison step was removed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Replacing the required comparison card with a single one-tap Accept button, once pilot satisfaction passed ninety eight percent and the card started to feel like leftover clutter.
6 · THE NUMBER
The pilot cohort was ___% solo travelers. The general population at rollout is only ___% solo.
Tap to flip
ANSWER
91% pilot, 54% general availability. That gap is why a booking with two travelers on it never showed up as a real problem until scale brought ordinary families in.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The app flags the two-traveler, tight-connection booking before it books anything. Selin sees both options, picks the one that keeps her and her mother together, and they land in Portland by 6:40pm, not three hours apart.
8 · CROSS-PRODUCT
Section 4 answers this same question for a different product, with a different flip family. Which product, which family?
Tap to flip
ANSWER
A hospital chest X-ray triage tool, using the workaround flip: a radiologist starts keeping a private spreadsheet because the tool has no memory of a patient's earlier scans.
Check yourself Score: 0 / 0
Multiple choice
1. What was the flip in Selin's story, and what were its two settings?
A. She reads the comparison card a little faster than she used to.
B. She reads every proposed flight before accepting it, or she taps Accept without reading it at all.
C. The assistant's accuracy dropped once general availability hit.
D. She asks her mother to double-check the booking before she accepts it.
Show hint
A flip is a verb the person does, not a change in the model, and it has exactly two settings.
Show answer
B. C describes the model, not a person's behavior. A is a dial, not a flip, there was no "reads it faster" setting she actually landed on. D is a fix, not what happened in the story.
True or false
2. True or false: the fast one-tap Accept button should be removed for every traveler, including solo fliers with plenty of time between flights.
True
False
Show hint
Look at the "what I would leave alone" paragraph. Who was the pilot actually testing?
Show answer
False. Solo travelers with a normal connection are the exact case the pilot proved out, ninety one percent of the time. Slowing that path down to guard against a case that almost never touches them would cost millions of people time for no reason.
Fill in the blank
3. The decision this answer takes back is replacing the required ______ card with a single one-tap ______ button.
Show hint
One of these showed both flight options side by side. The other just had one big word on it.
Show answer
Comparison, Accept. The comparison card felt like clutter once the model was reliably right. It stopped being clutter the moment two million ordinary travelers, not forty thousand hand-picked ones, started tapping past it.
Multiple choice
4. Why couldn't Selin have just "read the card a bit faster" instead of not reading it at all?
A. She did read it faster, that's exactly what happened.
B. Once the app removed the required comparison step, there was no partial version of checking left. Only tap-and-trust, or open-and-compare.
C. The assistant told her not to read it.
D. Reading the card was against airline policy.
Show hint
Look at what the redesigned screen actually offered her: one big button, and a small link almost nobody tapped.
Show answer
B. Once the interface itself only had two paths, an instant tap or a deliberate tap on "view details," there was no calibrated middle setting for her thumb to land on. That's what makes it a flip and not a dial.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think of a spell checker, a map app, a spam filter. Something whose whole job is letting you stop doing something by hand.
Show answer
Model answer: "Text message autocorrect. I stopped rereading messages before I send them, because it's almost always right. If it started guessing wrong more often, I wouldn't 'proofread a bit more carefully.' I'd go back to typing every word out and checking it myself, because there's no calibrated middle between trusting autocorrect completely and rereading every message." Any honest answer works if it names a real two-setting switch, not just "I'd be more careful."
Fill in the blank, do the math
6. Windrose expects about 120,000 rebooking-eligible itineraries a month once every phase is live. At the pilot's 9 percent companion rate, about how many would include more than one traveler? At the real general-availability rate of 46 percent, how many actually do? About how many more is that, a month?
Show hint
120,000 times 0.09, and 120,000 times 0.46.
Show answer
About 10,800 at the pilot's rate, versus about 55,200 at the real rate. That's roughly 44,400 more multi-traveler itineraries a month than the pilot ever had to handle. That gap is the exact case the comparison card got removed for.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.