CaseAdvancedShipping & Model Lifecycle / Rollout strategy and phased launches / #19

How would you handle a rollout where cost scales faster than expected?

The direct answer
Pause Skyfill's rollout at 25 percent and fix the confidence threshold driving the cost overrun before expanding further, instead of rolling forward while "optimizing it in parallel." The gap here is 3.5 times the forecast, caused by the model escalating to a far pricier path on messy real canvases the launch test set never included, and that problem gets worse, not better, as more real users arrive. Resume expansion only once a threshold recalibrated on real files holds the escalation rate under 15 percent for three straight days.
How to handle it, in order
  1. Pause the rollout at its current stage instead of rolling forward while "optimizing in parallel."Why: the overrun is 3.5 times forecast and it's driven by something that gets worse as more real, messy input arrives, not something usage will smooth out on its own.
  2. Recalibrate the confidence threshold against real user canvases, not the clean internal test set.Why: the clean 400-file set never contained the mid-project, mixed-style canvases that actually trip the expensive escalation, so it was measuring the wrong thing from day one.
  3. Set the resume line as a real number: escalation rate under 15 percent for three straight days.Why: without a number, "the fix feels done" quietly becomes the whole justification for expanding again.
  4. Test the recalibrated threshold against output quality before shipping it, not just against cost.Why: a cheaper threshold that also makes fills worse just trades one hidden problem for another.
  5. Leave plain, already-tidy canvas fills alone.Why: the cheap model reads those confidently almost every time, so there's no cost problem to chase there, at any scale.
  6. Tell finance the real number before the board deck, not after an audit finds it.Why: a monitoring email that trusts the launch-day assumption instead of measuring live data is part of what let this run for ten weeks unnoticed.

How to answer this, stage by stage

This is a yes or no about whether to pause or keep shipping while you fix it, not a rule for every cost line an AI feature ever runs over, so PICK carries the answer.

1
Scope it to one product and one specific overrun
Say it like this
"Let's make this real. Say Palette Foundry is a design tool, and inside it there's a feature called Skyfill: pick an empty patch of canvas, type a prompt, and it paints something that matches your design. Twenty five percent of users have it. The team just found out it costs three and a half times what the launch plan said it would. I need to decide what happens next."
Why this works
Stops the answer floating at "it depends on the cost" and gives the interviewer one real number to push on.
2
Say your structure out loud
Say it like this
"I'll pick a position first, then say who feels it on each side if I'm wrong, then name which kind of miss actually costs more, then say what evidence would change my mind. That's PICK, and I'll go in that order."
Why this works
Shows a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Reframe what the question is really testing
Say it like this
"This isn't really about whether three and a half times sounds bad. It's one question: is this a cost problem more users will make worse, or one that settles down on its own. That's the whole split."
Why this works
Shows the interviewer you see past the surface number to the judgment being tested.
4
State the position, with the number in it
Say it like this
"My pick: pause Skyfill's rollout at twenty five percent right now, and fix the actual cost driver before expanding to the next stage. Not roll it out to fifty percent while we 'optimize in the background.' The gap is three and a half times forecast, and it's caused by something that gets worse, not better, as more real users show up."
Why this works
PICK rewards a real committed line drawn at a specific number, not a hedge dressed up as caution.
5
Name who feels each side
Say it like this
"Here's the split. Pause it, and the people who feel it are the users still waiting for a feature that's already working for a quarter of the app, plus a roadmap that slips about a week. Keep expanding it while we fix it in parallel, and the people who feel it are finance, explaining a bill nobody budgeted for, three months from now, once far more users are running through the expensive path."
Why this works
Turns "the stakes are different" from a claim into something the interviewer can picture happening.
6
Name the cost asymmetry, plainly
Say it like this
"The pause is cheap and it's loud: a delayed feature, an annoyed product team, gone the moment the fix ships. The keep-going mistake is quiet and it's expensive: nothing crashes, nothing errors, the feature keeps looking like a success story, and the bill keeps climbing underneath that success story until someone in finance finally reads it. I'm optimizing against the one nobody would catch by watching the usual dashboard."
Why this works
Names which miss is which instead of leaving "asymmetry" as an unexplained word.
7
Name the kill criteria
Say it like this
"I'd flip this back to expanding once a threshold recalibrated on real, messy user files holds the actual escalation rate under fifteen percent for three days straight. Right now it's running at fifty percent, so it stays paused, on real evidence, not on how confident engineering sounds about the fix."
Why this works
Shows the pick can move, and states exactly what would move it.
8
Close on the countable difference
Say it like this
"Here's the number that makes the case. Pausing and fixing it properly costs about twelve thousand dollars over the plan and about a week. Rolling forward while fixing it in the background costs about forty seven thousand, and it still takes three weeks, because now the fix has to be tested against a population that keeps growing underneath it. Pausing isn't the cautious option here. It's the cheap one."
Why this works
Ends on the line the interviewer remembers, not a feeling about risk.

One more thing before the walkthrough ends: this isn't a blanket rule that every cost overrun needs a pause. Most candidates hear "costs more than planned" and reach for the same caution every time. Name whether the overrun gets better or worse as more people use the feature, and you've shown judgment instead of reciting a rule of thumb.

Let's learn

Here's what happens when a feature works so well that success itself becomes the expensive part.

Palette Foundry is a design tool. Inside it sits a feature called Skyfill: pick an empty patch of your canvas, type what you want there, and it paints something that matches the rest of your design.

The team built Skyfill with two models working together. A fast, cheap model paints the first draft. If that model isn't sure the fill will match the rest of the canvas, it quietly hands the job to a slower, much pricier model for a second pass, no button, no ask, it just happens. Before launch, the team tested this hand-off on 400 of their own design files, clean, finished, one style each. On those files, the hand-off fired 8 times in 100. At that rate, the cost model said Skyfill would run about $18,000 a month once a quarter of all users had it.

Real users are not working from finished files. A canvas mid-project is messy: three unfinished layers, two clashing styles, a placeholder nobody has replaced yet. The cheap model reads all of that as low confidence far more often, so the hand-off isn't firing 8 times in 100. It's firing on about half of every real fill. At the 25 percent stage, that puts the actual bill at about $63,000 a month. Three and a half times the plan.

Here is the turn. Three and a half times isn't the real problem. The real problem is that the more Skyfill gets used exactly the way it was meant to be used, on real, messy, in-progress work, the worse this number gets, because that's exactly the kind of canvas that trips the hand-off. Usage climbing was supposed to be the win. Instead, usage climbing is what makes it more expensive, at every stage still ahead.

We didn't build a feature that costs more as it grows. We built one that costs more as it succeeds.

At its worst: if nothing changes and Skyfill reaches everyone, the plan expected about $72,000 a month. The real pace works out to about $252,000 a month, for a feature that was supposed to be a normal line item, not the biggest number on the infrastructure bill.

The choice I would take back We calibrated the hand-off using our own tidy design files, because that's what we had before real users showed up. That was a fair call in week one, there wasn't anything else to test against. I'd take back leaving that calibration in place once real fills started coming in. The moment beta users had made a few thousand messy canvases, that data should have replaced the clean set. It never did, because the numbers still looked fine for the first few weeks.

What I would leave alone. The part of Skyfill that never needs any of this: filling a plain, mostly-empty background on a simple, already-tidy canvas. The cheap model reads those confidently almost every time, the hand-off barely ever fires there, and it never will, no matter how many more people use it.

The lesson. A cost forecast built from a clean test set only tells you the cost of clean work. The whole reason to ship an AI feature to real users is that their work isn't clean. Any cost model that skips that step is measuring the wrong thing from day one.

A hand-drawn comparison of two boxes with deliberately unequal weight. Left, a small plain grey box labeled pause, eight days, one slide moved, with a note underneath reading who feels it, the launch calendar. Right, a much larger jagged red-orange box labeled keep going, burn doubles for three weeks straight, with a note underneath reading who feels it, the company's margin.
Same rollout, two very different sizes of what it costs to get this wrong

Two numbers that were never supposed to disagree

You don't need this to answer the question. Read it if you want to feel why holding the line at 25 percent beats fixing it while you keep climbing.

Julen Kastrup can read a cost model the way some people read a recipe. He knows exactly which line will run over before anyone's bought a single ingredient.

He's run the Skyfill pod at Palette Foundry for two years and priced the team's last AI feature himself, spreadsheet and all. Skyfill launched at 5 percent. Every Monday for the first six weeks, Julen opened the full weekly infra report line by line: fills served, hand-off rate, blended cost per fill, next to the number the launch model had promised. Off by a rounding error, every single week.

By week seven, at the 25 percent stage, the report started arriving with one line at the top: "Skyfill spend: within range." Julen started reading only that line. Six good weeks does that to a person, and there was a launch deck due.

Then came the quarterly audit. Oskar Weng, from finance, runs the infra reconciliation ahead of every board deck. He pulled Skyfill's actual GPU billing, compared it to what the automated dashboard had been reporting all along, and messaged Julen: "Your Skyfill number's off by a lot. The email's been marking it 'within range' using the 8 percent hand-off rate from launch. Real rate's running at half of all fills. Did you know that?"

The alarm never rang, because it was still set to a number we made up on launch day.

Julen didn't stop the rollout. Launch was one week from the board deck, adoption numbers were the best story in the room, and pulling the ramp back felt like turning a win into a confession. So he told engineering to recalibrate the threshold without blocking the ramp for it, and Skyfill moved from 25 percent to 50 percent on schedule the next Monday, the fix running underneath it in parallel.

Recalibrating against real, messy user canvases turned out to take about three weeks, not because the fix itself was hard, but because it now had to be tested carefully against a population that kept growing underneath it. Every extra day of testing meant one more day of a bigger population paying the old, broken rate. By the time the new threshold shipped, those three weeks had cost Palette Foundry about $47,000 more than the plan ever budgeted for that stretch, on top of everyone already exposed to it.

Back at the original launch readiness review, an engineer had actually asked whether they should hold off calibrating until real user files existed. The answer, reasonable at the time, was that there weren't any real files yet, the clean set was all there was, and they'd revisit it once the beta numbers came in. Nobody put a date on revisiting it. The first few weeks looked fine, so nobody did.

Here is the replay. Same message from Oskar, new design: the rollout holds flat at 25 percent the moment Julen reads it. Recalibrating the threshold against real canvases, without a growing population underneath it, takes the team about a week. Total extra spend over the plan during that week: about $12,000, against the $47,000 the way it actually happened.

One design pays $12,000 to learn the truth. The other pays four times that to keep pretending it already knew it. The board deck moves eight days. Nobody outside the room ever needs to know the number was wrong at all.

And the thing I'd go back and tell myself, in that first launch review: the system we built to watch for trouble was only ever as honest as the number we typed into it on day one.

PICK, with a real dollar figure on both sides

This is a yes or no about whether to pause or keep shipping while you fix it, not a rule for every rollout Palette Foundry runs, so PICK carries the weight here.

P, position. Pause Skyfill's rollout at 25 percent the moment the real hand-off rate is known to be running at 3.5 times the modeled cost, and fix the confidence threshold before expanding further. Don't keep ramping while "optimizing in parallel" once the gap is this size.
I, impact. A paused rollout is felt by the users still waiting for Skyfill and by the roadmap: a real but bounded cost, a slipped date, a launch deck that says "in one more sprint" instead of "live today." A rollout that keeps expanding with the cost problem unresolved is felt by the whole company's unit economics, and it worsens with every stage, because more of the audience means more of the exact messy files that trip the expensive path.
C, cost asymmetry. Pausing is cheap and it's loud: a delayed feature everyone can see and complain about, gone the moment the fix ships. Continuing is hidden and expensive: nothing about it shows up as an error, a crash, or an unhappy user, it shows up three months later as a line item finance has to explain to a board that assumed an AI feature would cost what the roadmap said it would.
K, kill criteria. Resume expanding once a recalibrated threshold holds the real hand-off rate under 15 percent for three straight days, measured against real user files, not the old test set. Above that, it stays paused, because right now the only reason anyone caught this was a quarterly audit, not the rollout's own numbers.
Knowledge spark: what is the model's confidence score doing here? Before Skyfill paints anything, the cheap model checks its own guess: how sure am I this will match the rest of the canvas? A high number, it keeps its own result. A low number, it quietly hands the job to the pricier model instead. Nobody sees this happen. It just changes which model gets paid for.
Knowledge spark: why did a 400-file test set miss this? The team tested the hand-off on their own finished design files, because that's all they had before launch. Finished files are tidy, one style, no clashing layers. Real users' canvases mid-project are not. The model was calibrated on the wrong kind of file, so it never saw the pattern that would actually decide the real cost.
Cost, by the numbers: pausing versus rolling forward while you fix it
$12,000 $47,000 Pause at 25%, fix first (about 1 week) Keep rolling to 50%, fix in parallel (about 3 weeks)
Holding the rollout at 25 percent while the threshold gets recalibrated costs about $12,000 over the plan, and takes about a week. Letting the rollout keep expanding while the same fix runs in the background costs about $47,000 over the plan, nearly four times as much, because the fix now has to be tested carefully against a population that's still growing underneath it.
The kill line, charted: real hand-off rate against the 15% resume line
At or below the kill line, safe to expand
Above the kill line, hold the rollout
Kill line: 15% hand-off rate
0% 50% 100% 15% kill line 12% Week 1, 5% rollout 22% Week 6, ramping to 25% 50% Week 10, audit day 11% Week 11, after the fix
The real rate crossed the 15 percent kill line as early as week six, four weeks before the audit found it, because nothing was actually measuring the live rate against a real line. The dashboard was measuring it against the 8 percent number from launch day instead.

Try the same four letters on a shoebox of old photographs

Iwan Dace runs product for Kinfolk Archive, an app three hundred thousand people use to restore and colorize old family photographs. Upload a scan, and a model repairs the damage and brings back the color.

Two changes are queued for the same release. A new album layout that groups restored photos by decade, pure interface, no model involved. And an expansion of the restoration model's escalation logic: a fast single-pass model handles typical scans, but a torn, badly faded, or multi-face photo gets handed to a slow, expensive multi-pass model instead, the one built to actually repair real damage.

Iwan split the release the way Julen eventually wished he had.

P. Fast, light rollout for the album layout, no model touched, ship it in a week. Full pause-and-fix treatment for the escalation logic: the threshold was calibrated on employees' own well-kept family photos, and real users are uploading exactly the water-damaged, decades-old, attic-stored photos the app was built for, at a rate the test set never captured.
I. A wrong album layout is felt by one user who regroups a folder by hand, a minute of annoyance. An unresolved escalation cost is felt by the whole company's per-photo economics, and it worsens specifically as more people upload the damaged photos Kinfolk Archive exists to fix.
C. The layout miss is cheap and loud, someone flags it the same day. The cost miss is quiet and expensive, restorations keep succeeding, users stay happy, and the wrong number sits inside a monthly infrastructure bill nobody reads line by line until finance closes the quarter.
K. Resume expanding restoration access once a threshold recalibrated against real degraded photos holds the escalation rate under a set line for three straight days. Below that line, the risk is close enough to normal photo processing. Above it, it stays paused.

What I would leave alone, at Kinfolk Archive The album layout doesn't need any of this. It never touches the model, and getting it wrong costs one person one minute with a mouse.

Swap the trigger and it still runs

  • Speed: if Kinfolk Archive needed the fix live in three days instead of a full quarter's notice, the pick doesn't move, the pause still holds, the test just runs on a smaller sample of real damaged photos, not skipped.
  • Cost: if recalibrating the threshold took more engineering time than expected, the pick still doesn't move, that cost was never the question, whether the company could keep affording its own most-used feature was.
  • The model got better: if the new threshold already held the escalation rate under the line on 500 real damaged photos, that's exactly the evidence that flips it back to expanding.

Where people run it wrong

  • Treating "nothing crashed" as proof a rollout is healthy, when a cost problem never shows up on a crash dashboard.
  • Blending a harmless UI change and a model-cost change into one release plan, so the safe half makes the risky half look safer than it is.
  • Waiting for the automated report to say something's wrong, when the report was built using the same assumption that turned out to be false.

Buy yourself two seconds, out loud

Say the reframe before reaching for a rule of thumb. "Give me a second, I want to check whether this cost problem is something usage will smooth out, or something usage makes worse." That's true, it's already the reframe from stage three, and it buys you time to find the real asymmetry instead of reciting "AI costs are just unpredictable."

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's the hardest step to nail?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a quiet, compounding cost overrun costs more than a loud, cheap, visible pause.
2 · THE PERSON
Who is this answer about, and what does he already do well?
Tap to flip
ANSWER
Julen Kastrup, who has run the Skyfill pod at Palette Foundry for two years and priced the team's last AI feature himself.
3 · THE HABIT
What did Julen stop doing once the weekly cost report never let him down?
Tap to flip
ANSWER
Reading the full weekly infra report line by line. By week seven he was reading only the automated summary line, "Skyfill spend: within range."
4 · THE ASYMMETRY
What are the two ways to get this pick backwards, and who gets hurt by each?
Tap to flip
ANSWER
Rolling forward with an unresolved cost driver hurts finance and the company's margin, quietly, for months. Pausing every minor cost blip hurts the roadmap and the team's credibility, for no real reason, when the overrun isn't the kind that compounds.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Pause Skyfill's rollout at 25 percent and fix the confidence threshold before expanding, because the overrun is 3.5 times forecast and gets worse as real usage grows, not better.
6 · THE NUMBER
At the 25 percent stage, Skyfill's real monthly cost was running at about $______, against a forecast of $18,000.
Tap to flip
ANSWER
$63,000. About 3.5 times the plan, driven by a hand-off rate of roughly 50 percent instead of the modeled 8 percent.
7 · THE KILL CRITERIA
What evidence would flip the pick back to expanding the rollout?
Tap to flip
ANSWER
A threshold recalibrated on real user files holding the hand-off rate under 15 percent for three straight days. Right now it's running at about 50 percent, so it stays paused.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the position land there?
Tap to flip
ANSWER
Kinfolk Archive, a photo restoration app. The new album layout ships fast, no model touched. The restoration model's escalation logic gets paused and recalibrated first, because the real damaged photos it needs to handle are exactly what trips the expensive path.

Check yourself Score: 0 / 0

True or false
1. True or false, with why: since Skyfill's fills at 25 percent rollout still looked great with no errors, the cost overrun wasn't really a problem worth pausing for.
  • True
  • False
Show hint
Think about what a crash dashboard is built to catch, and what it isn't.
Show answer
False. Good-looking output with no errors is exactly what let this run for ten weeks unnoticed. The cost problem was real and it was getting worse with scale, it just never touched any of the signals the team was watching.
Multiple choice
2. Which of these is the actual mechanism behind this answer's pick?
  • A. Ship the rollout forward on schedule, since nothing has crashed or errored.
  • B. Cancel Skyfill entirely until the model stops needing the expensive path at all.
  • C. Hold the rollout at its current stage, recalibrate the threshold against real user files, and resume once a real evidence line is crossed.
  • D. Keep rolling forward while engineering "looks into" the cost in the background.
Show hint
Three of these either ignore the growing exposure or give up on the feature completely.
Show answer
C. A and D both let more real users hit the broken path while a fix is pending. B throws away a feature that works fine for most canvases. Only C matches the process to what's actually driving the overrun, and sets a real bar for resuming.
Fill in the blank
3. Fill in the blank: the rollout only resumes expanding once the real hand-off rate holds under ______ percent for three straight days.
Show hint
It's the K step from the PICK recap, the number that turns "the fix feels done" into a real bar.
Show answer
15. Well under the 50 percent rate found at the audit. Below 15 percent, the risk is close enough to what the launch plan already assumed. Above it, the rollout stays paused.
Short answer
4. If the real hand-off rate had only doubled, from 8 percent to 16 percent, instead of climbing to 50 percent, would the same pause-first pick still hold? Walk through it.
Show hint
Think about what actually drives the position: the size of the gap, or the direction it's heading as usage grows.
Show answer
Probably not, and that's the point. At 16 percent, the cost overrun is small and just barely above the 15 percent kill line already set for resuming. That's close enough to fix inside a normal sprint without freezing the ramp, especially if the trend is flat rather than climbing. The pause was earned by the size of the gap and the fact that it gets worse with scale, not by the mere existence of an overrun.
Multiple choice
5. Why did the automated weekly cost email keep saying "within range" even as the real hand-off rate climbed to 50 percent?
  • A. Because Oskar had turned off the alert thresholds for that quarter by mistake.
  • B. Because the email compared spend to the 8 percent hand-off rate assumed at launch, instead of measuring the real, live rate.
  • C. Because the premium model had gotten cheaper to run, which offset the higher hand-off rate.
  • D. Because Skyfill's user count had actually gone down that quarter.
Show hint
Ask what number the dashboard was actually checking spend against.
Show answer
B. The dashboard was built using the launch-day assumption, not a live measurement of the real hand-off rate. It could only ever confirm the number it had already been told was true, which is exactly why a quarterly audit caught this before the dashboard ever did.
Short answer, apply it yourself
6. Pick a product you use yourself with an AI feature in it. Name one part of it where inference cost could scale faster than usage does, and why.
Show hint
Look for a place where the model sometimes takes a cheap path and sometimes a pricier one, and ask what decides which path a real, messy request takes.
Show answer
Model answer: "A writing app I use has a grammar-fix mode that sometimes just fixes a sentence and sometimes rewrites the whole paragraph if it thinks the fix alone won't read well. If more of what people actually paste in is messy, half-finished drafts instead of clean text, that 'rewrite the whole thing' path probably fires far more than whatever clean writing sample they tested it on, and it would get more expensive exactly as more real, messy writing shows up." Any answer works if it names a cheap-versus-expensive path inside the model and a reason real input trips the expensive one more than a test set did.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more