What is the cost profile of running two models in parallel during a migration?
- Budget both models' full per-image cost stacked together for the whole window, not a doubled guess.Why: the parallel run adds the new model's own bill on top of what you already pay, it does not just double one number.
- Size the window by how many dual-tagged uploads the rarest search category needs, not by a fixed number of weeks.Why: a category that only shows up a dozen times a day decides how long you actually need to trust the diff, not a habit borrowed from the last migration.
- Give the cost a range, low if the new model is only a little pricier per image, high if it is a lot pricier.Why: one number hides a swing of several thousand dollars.
- Check the total against the team's normal monthly inference budget before promising a figure.Why: a spike that looks big on its own might be a small, tolerable bump against the real budget, or it might not be, and you will not know until you check.
- Watch the rare category's upload rate hardest, not the new model's price tag.Why: if that rate is slower than assumed, it swings the total further than the new model turning out pricier ever could.
- Track tag quality by category during the run, not just an overall average.Why: a rare category can quietly get worse while the aggregate score still looks fine, and that is exactly what the parallel run exists to catch.
How to answer this, stage by stage
Nobody is grading whether you land on exactly $5,400. They are grading whether you can defend the arithmetic behind it, whether the range is honest, and whether you close on something the room could go check. Seven moves get you there.
Let's learn
Say we build a model that reads every new photo a contributor uploads and writes the keywords buyers search by. Before any tagging model existed, contributors typed their own keywords by hand, and most photos got eight or nine of them, missing half the terms a buyer might actually type into search.
Then Cormorant went live. It reads a photo and writes about forty keywords in under a second, for ninety cents per thousand images. A landscape shot now turns up under "hillside," "golden hour," and "hiking trail," not just whatever the contributor happened to type. Search got wider, overnight, for almost nothing.
Now Shutterbrook is replacing Cormorant with Petrel, a bigger model that catches keywords Cormorant misses, especially in small, specific categories: aerial drone work, macro insect shots, a handful of niche clusters that barely show up in daily volume but matter a lot to the buyers searching for exactly that. Petrel is slower and pricier per image than Cormorant. That trade is on purpose, more cost and more turnaround time for three weeks, in exchange for tag quality the parallel run has to actually prove before the switch, not a free upgrade nobody checked.
Say plainly: the two hundred four extra dollars a day Petrel costs is not the real problem. A few thousand dollars of inference is nothing next to what a stock marketplace makes in a month. The real problem is what happens the day the parallel run ends. Cut over on a calendar guess, and a category the two models have barely tagged together gets promoted to production with nobody having actually checked whether Petrel is any good at it.
At its worst, Petrel goes live company-wide having tagged industrial drone photos only a handful of times, ever, next to Cormorant. If it turns out quietly worse at that one narrow category, those photos stop surfacing in search the week the model flips, and nobody notices for a month, because drone photography is a sliver of daily volume, easy to miss inside an overall accuracy number that still reads fine.
The choice I would take back. When this migration was first sized, the spend alert was set as one flat number, two hundred sixty dollars a day, worked out from the average sixty thousand uploads. That felt careful. It also assumed daily volume would stay close to that average for three straight weeks, and nothing in the alert would notice if it didn't.
What I would leave alone. Shutterbrook also migrates a much cheaper spam-detection classifier the same quarter, both models together cost about six dollars a day. A volume spike there might double the bill for an afternoon and nobody would need to care. Only a migration where the two models' combined cost is big enough to actually move a budget line needs this kind of watching.
The lesson. A migration budget built off an average day doesn't know what to do on the one day that isn't average. If the number can only survive a quiet three weeks, it was never really a number. It was a hope with a dollar sign on it.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why the fifteen hundred dollar afternoon almost went unnoticed.
Kunal Bhandari has run three model migrations at Shutterbrook in as many years, and each time he built the cost model himself, in a spreadsheet, before anyone signed off on a start date. He's the one people ask when a rollout feels rushed. He usually says whether it is, within a day.
The Cormorant-to-Petrel migration started clean. Day one, the billing dashboard read almost exactly what his spreadsheet predicted, two hundred fifty-something dollars. He checked it every morning that first week, cross-referencing actual spend against the plan line by line.
By day eight, it had tracked to plan for a week straight. He started checking every few days instead. By day eleven, with the numbers still calm, he'd stopped opening the dashboard at all unless someone asked. The two earlier migrations had gone exactly this way, quiet after the first week, and this one looked no different.
Marketing had scheduled a "Photographer of the Month" contest for that same quarter, entirely unrelated to the tagging migration, nobody had connected the two calendars. The contest's submission deadline landed on day twelve. For four days, daily uploads ran two and a half times higher than normal, contributors racing to get entries in before the cutoff.
Both models tag every upload during a parallel run. A volume spike doesn't just add to Cormorant's small bill, it multiplies Petrel's much bigger one right alongside it. For those four days, the combined daily cost ran close to six hundred forty-five dollars instead of two hundred fifty-eight, an overage of about fifteen hundred dollars that the flat two-hundred-sixty-dollar alert never crossed, because two hundred sixty was never linked to what a day's uploads actually were, only to what they'd averaged in the plan.
Kunal caught it by accident, on day thirteen, opening the dashboard for an unrelated reason, a question from someone in finance about a different line item entirely. The spike jumped out at him immediately, four days already past.
Thirteen hundred dollars over four days isn't a disaster. What stayed with Kunal was the shape of the near miss. The alert had no idea what a normal day actually looked like in real time. It only knew what the plan had assumed a normal day would look like, three weeks earlier, on a spreadsheet. If the contest had run for two weeks instead of four days, or if a second, bigger spike had landed on top of it later in the window, nothing built into the migration would have said a word until someone happened to open the dashboard for some other reason.
Kunal never really had a number for what a spike would do to the bill. He had a habit, a flat two-hundred-sixty-dollar alert, that had worked fine on two calmer migrations before this one. Weeks earlier, at the kickoff meeting, someone had actually asked whether the alert should scale with live volume instead of sitting at one number. The answer at the time was that it would be over-engineering a three-week migration. Nobody pushed back further. It was the only design in the room.
The second version of the alert isn't a bigger flat number. It's pegged to volume itself, roughly four point three cents combined per hundred images, times whatever that day's real upload count actually is, with a flag if the day's total crosses one point four times the rolling baseline. Run the same contest week through that design and the flag fires the first afternoon, not two days after the fact, and the overage stops near four hundred dollars instead of running the full fifteen hundred before anyone looks.
The thing I'd tell myself, back at that kickoff meeting: a budget built off the average day is a budget that has never met a real one.
BOUND: the arithmetic behind the parallel-run invoice
This is a cost build-up and a sizing question, how much running two models on every upload actually costs and for how long, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Daily parallel-run cost is not one number, it's two added together. Uploads per day times Cormorant's cost per image, plus uploads per day times Petrel's cost per image. Total cost is that daily figure times how many days the window actually needs, and the window itself is sized by how many dual-tagged examples the rarest search category needs before its tags are trustworthy, not by a number of weeks picked off a calendar.
O, own the numbers. Sixty thousand uploads a day. Cormorant costs ninety cents per thousand images, fifty-four dollars a day. Petrel, at our planning assumption of about four times Cormorant's price, costs three dollars forty per thousand, two hundred four dollars a day. Combined, two hundred fifty-eight dollars a day. We looked at shadow-sampling only five percent of uploads through Petrel instead, to cut that cost by most of it, and rejected it: at five percent, the rarest category, industrial drone shots, would produce about half a tagged example a day, and it would take most of a year to gather the two hundred fifty we need. Full parallel tagging on every upload is the only way to gather that category fast enough. The planned window is twenty-one days, sized to that same category at its real rate of about twelve uploads a day. Total: about five thousand four hundred dollars.
U, use a range. If Petrel turns out only twice as pricey as Cormorant instead of four times, the twenty-one-day run costs about thirty-four hundred dollars. If it's six times as pricey, closer to seventy-nine hundred. The window length swings the number harder still: if the drone-shot category shows up at half the assumed rate, six a day instead of twelve, reaching two hundred fifty confirmed examples takes about forty-two days, not twenty-one, and the total roughly doubles, to about ten thousand eight hundred dollars.
N, nail the sanity check. Shutterbrook's whole ML team spends about forty-two thousand dollars a month on inference, across every model it runs. Five thousand four hundred dollars for this one migration is about thirteen percent of a month's total budget. Even the worst case, near eight thousand dollars, stays under a fifth of it. Tolerable, and worth telling finance about in advance rather than letting it show up as a surprise line on next month's invoice.
D, direction. Two things could move this number, and they don't move it equally. The drone-shot category showing up at half its assumed daily rate stretches the window from twenty-one days to about forty-two and adds roughly fifty-four hundred dollars. Petrel turning out pricier than assumed, six times Cormorant instead of four, adds about twenty-five hundred dollars. The rare category's real rate is the bigger lever, and it's the number nobody at Shutterbrook had actually confirmed before the migration started. The four-times assumption for Petrel's price came straight off a vendor pricing page.
One more thing the arithmetic alone doesn't show: cutover isn't a single clean pass either. Search only switches to Petrel once it matches or beats Cormorant's precision on each category's own eval set on at least nine of the last ten weekly checks, not the first week it happens to look fine. An aggregate accuracy number can sit near ninety-eight percent for the whole window while quietly missing keywords on a category too small to move that average, so precision gets tracked by category, not just as one overall number, for as long as the two models run side by side.
And if you want to be sure it really works, try it somewhere else
Rentford County's permit office runs an AI tool that reads incoming building-permit applications and routes each one to the right review queue: residential, commercial, historic district, and a handful of narrower categories. The office is migrating that router from an older classifier to a newer one, and both models classify every application for a stretch before the county cuts over.
B, break it down. Same shape, smaller numbers. Weekly parallel-run cost equals applications per week times the old model's cost per document, plus applications per week times the new model's cost per document. Weeks needed equals the confirmed catches wanted for the rarest permit type, divided by how often that type actually arrives.
O, own the numbers. Eight hundred applications a week. The old classifier costs four tenths of a cent per document, about three dollars twenty a week. The new one, assumed three and a half times pricier, costs about eleven dollars twenty a week. Combined, fourteen dollars forty a week. Accessory-dwelling-unit permits, the rarest category, arrive about five times a week. Rafiq Sandoval, the PM running this migration, wants forty confirmed dual-classified examples of that category before trusting the new model on it. At five a week, that's eight weeks. Total: about one hundred fifteen dollars.
U, use a range. If the new model turns out six times pricier instead of three and a half, the eight-week run costs about one hundred seventy-nine dollars. If the accessory-dwelling-unit rate is half what's assumed, two and a half a week instead of five, the window stretches to sixteen weeks and the total roughly doubles, to about two hundred thirty dollars.
N, nail the sanity check. Rentford's whole permitting-software budget for the year runs about six thousand dollars. Even the worst case here, two hundred thirty dollars, is under four percent of that. Nowhere near a number that needs a finance conversation, the office could run this migration twice over without anyone noticing the spend.
D, direction. Same tension as Shutterbrook's migration, at a different scale. How often the rare permit type really arrives swings the total more than the new model's price does, because it's the one number nobody had actually measured before the migration started, only assumed from a rough sense of last year's mix.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: budget both models stacked, not doubled, size the window by how many examples the rarest category needs, check it against the real monthly budget.
Cost: finance says the migration budget is frozen this quarter. Don't shrink the rare-category threshold to fit, shrink the daily upload scope instead, run the parallel tag on a fixed, deliberately over-sampled slice of the rare category's own traffic rather than every upload, and say plainly that the tradeoff is a slower window, not a shakier bar for trusting the new model.
The model got better: Petrel turns out to need barely any correction versus Cormorant. That doesn't remove the need to check the rare category's real rate, it just means the eventual bill leans toward the low end of the range, not that the range stops mattering.
Where people run it wrong.
They price the parallel run by doubling one number, when it's actually two separate bills, of very different sizes, stacked on top of each other.
They pick the window length off the last migration's calendar, without checking whether this migration's rarest category can even produce enough examples that fast.
They set a spend alert off the planned average day and never revisit it, so a real event, a spike, a promo, a seasonal surge, can blow past the plan for days before anyone notices.
How to use it live. Say the equation before naming a single number: "the parallel-run cost isn't the new model's price times two, it's the old model's bill plus the new model's bill, stacked for as long as the rarest category needs to earn enough evidence to trust." That buys the room to ask a real question instead of guessing a lump sum that sounds cautious.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't tracking accuracy by category overkill, why not just watch the overall number?" Response: an aggregate score near 98 percent can hide a category too small to move it, which is exactly the silent, per-category regression the tracking exists to catch before cutover.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Model migration and version changes for users
- #1 Your provider deprecates the model behind your main feature in 60 days. Write the plan.
- #2 How do you test a replacement model against the behaviour users have come to expect?
- #3 Explain why a strictly better model can still be a bad migration.
- #4 What should you tell users when model behaviour changes underneath them?
- #5 Describe a dual-running strategy for a model migration.
- #6 How do you handle customers who tuned their prompts to the old model?