CalculationAdvancedShipping & Model Lifecycle / Model migration and version changes for users / #11

What is the cost profile of running two models in parallel during a migration?

The direct answer
Budget for both models' full per-image cost stacked on top of each other for the whole window, not a padded guess and not a simple doubling of one bill. At the marketplace's upload volume that is about $5,400 for a planned three-week run, roughly 13 percent of a month's total inference budget. Size the window itself by how many dual-tagged uploads the rarest search category needs before its tags are trustworthy, not by a calendar guess, because that rate is what actually blows the number, more than the new model turning out pricier per image.
Do this, in order
  1. Budget both models' full per-image cost stacked together for the whole window, not a doubled guess.Why: the parallel run adds the new model's own bill on top of what you already pay, it does not just double one number.
  2. Size the window by how many dual-tagged uploads the rarest search category needs, not by a fixed number of weeks.Why: a category that only shows up a dozen times a day decides how long you actually need to trust the diff, not a habit borrowed from the last migration.
  3. Give the cost a range, low if the new model is only a little pricier per image, high if it is a lot pricier.Why: one number hides a swing of several thousand dollars.
  4. Check the total against the team's normal monthly inference budget before promising a figure.Why: a spike that looks big on its own might be a small, tolerable bump against the real budget, or it might not be, and you will not know until you check.
  5. Watch the rare category's upload rate hardest, not the new model's price tag.Why: if that rate is slower than assumed, it swings the total further than the new model turning out pricier ever could.
  6. Track tag quality by category during the run, not just an overall average.Why: a rare category can quietly get worse while the aggregate score still looks fine, and that is exactly what the parallel run exists to catch.

How to answer this, stage by stage

Nobody is grading whether you land on exactly $5,400. They are grading whether you can defend the arithmetic behind it, whether the range is honest, and whether you close on something the room could go check. Seven moves get you there.

1
Scope it to one real product and one real migration
Say it like this
"Let's ground this. Say Shutterbrook, a stock-photo marketplace, gets sixty thousand new photo uploads a day from contributors. Every upload needs auto-tags, the keywords buyers search by, before it is even visible in results. Shutterbrook's PM, Kunal Bhandari, is migrating the tagging model from an older one, Cormorant, to a newer one, Petrel. During the migration, both models tag every single upload, side by side, for a few weeks before search actually flips over to the new one."
Why this works
Stops the answer floating in the abstract before a single dollar gets attached to it.
2
Say the structure out loud before naming any numbers
Say it like this
"Here's how I'd frame it. The question isn't 'how much does the new model cost.' It's what we spend running both models on every single upload, for as long as the dual run actually lasts. That's two costs stacked on top of each other, not one cost doubled, and the length of that window is its own number to size, not a guess borrowed from the last migration."
Why this works
Reframes the question before touching a figure, so the interviewer hears a method coming, not a vibe.
3
Break down the equation
Say it like this
"Here's the shape. Daily parallel-run cost equals uploads per day times Cormorant's cost per image, plus uploads per day times Petrel's cost per image. Total cost equals that daily figure times how many days the dual run needs to last, and the day count itself comes from how many dual-tagged examples the rarest search category needs before its tags are trustworthy."
Why this works
This is the B step of BOUND, the equation stated before a single number touches it.
4
Own the real numbers
Say it like this
"Sixty thousand uploads a day. Cormorant costs about ninety cents to tag a thousand images, so fifty-four dollars a day. Petrel is a bigger model, we're assuming close to four times the price, about three dollars forty per thousand, so two hundred four dollars a day. Run them together and that's two hundred fifty-eight dollars a day. Our planning window is twenty-one days, so about five thousand four hundred dollars total."
Why this works
This is the O step, real proposed numbers, not "we'll keep an eye on the spend."
5
Give the range, sized by evidence, not the calendar
Say it like this
"If Petrel turns out only twice as pricey as Cormorant instead of four times, the whole run costs about thirty-four hundred dollars. If it's six times as pricey, closer to seventy-nine hundred. And the twenty-one days isn't a guess either. Our rarest search category, industrial drone shots, only gets about a dozen uploads a day, and I want at least two hundred fifty of those tagged by both models before I trust Petrel's tags on that category specifically. At twelve a day, that's about three weeks on its own."
Why this works
A single number here claims a confidence the plan doesn't have, and ties the window to real evidence instead of a habit.
6
Sanity check the total against the normal monthly budget
Say it like this
"The whole ML team's normal inference spend, every model, every system, runs about forty-two thousand dollars a month. Five thousand four hundred dollars over three weeks is about thirteen percent of that, for one migration. Even the worse case, near eight thousand dollars, stays under a fifth of a month's budget. That's a real number to walk into finance with, not a surprise on next month's invoice."
Why this works
This is the N step, and it's the step a flat "just run it and see" plan always skips.
7
Name the biggest lever, then close on the one line
Say it like this
"If I had to bet on one thing moving this number, it's not Petrel's price tag, it's how often that rare drone-shot category actually shows up. If it shows up half as often as assumed, the window stretches from three weeks to six and the total roughly doubles, further than Petrel coming in pricier ever would. So here's what I'd actually say: budget the two models stacked, not doubled, size the window by how many dual-tagged examples the rarest category needs, check the total against the real monthly budget, and watch that category's upload rate harder than the model's price tag."
Why this works
Closes on a number someone in the room could actually go check, not a promise to be careful.
If you remember one thing Running two models in parallel does not double one bill. It adds the new model's own bill on top of what you already pay, for as long as the window lasts, and the window's real length comes from how fast the rarest category earns enough tagged examples to trust, not from how many weeks feel safe.

Let's learn

Say we build a model that reads every new photo a contributor uploads and writes the keywords buyers search by. Before any tagging model existed, contributors typed their own keywords by hand, and most photos got eight or nine of them, missing half the terms a buyer might actually type into search.

Then Cormorant went live. It reads a photo and writes about forty keywords in under a second, for ninety cents per thousand images. A landscape shot now turns up under "hillside," "golden hour," and "hiking trail," not just whatever the contributor happened to type. Search got wider, overnight, for almost nothing.

Knowledge spark: what's a parallel run? Running the old model and the new model on the same real uploads at the same time, so their tags can be compared before anyone trusts the new one alone. Nothing goes live on Petrel yet. Search still reads Cormorant's tags. Petrel just tags along, quietly, so its answers can be checked.

Now Shutterbrook is replacing Cormorant with Petrel, a bigger model that catches keywords Cormorant misses, especially in small, specific categories: aerial drone work, macro insect shots, a handful of niche clusters that barely show up in daily volume but matter a lot to the buyers searching for exactly that. Petrel is slower and pricier per image than Cormorant. That trade is on purpose, more cost and more turnaround time for three weeks, in exchange for tag quality the parallel run has to actually prove before the switch, not a free upgrade nobody checked.

Say plainly: the two hundred four extra dollars a day Petrel costs is not the real problem. A few thousand dollars of inference is nothing next to what a stock marketplace makes in a month. The real problem is what happens the day the parallel run ends. Cut over on a calendar guess, and a category the two models have barely tagged together gets promoted to production with nobody having actually checked whether Petrel is any good at it.

We weren't risking a bigger bill. We were risking a category nobody had actually watched.
The decision that mattered Size the parallel-run window by how many dual-tagged examples the rarest category needs, not by a flat number of weeks or a flat dollar alert built off the average day. That's the one check that decides whether Petrel gets trusted on drone shots, or just assumed to be fine because the aggregate number looked calm.

At its worst, Petrel goes live company-wide having tagged industrial drone photos only a handful of times, ever, next to Cormorant. If it turns out quietly worse at that one narrow category, those photos stop surfacing in search the week the model flips, and nobody notices for a month, because drone photography is a sliver of daily volume, easy to miss inside an overall accuracy number that still reads fine.

The choice I would take back. When this migration was first sized, the spend alert was set as one flat number, two hundred sixty dollars a day, worked out from the average sixty thousand uploads. That felt careful. It also assumed daily volume would stay close to that average for three straight weeks, and nothing in the alert would notice if it didn't.

What I would leave alone. Shutterbrook also migrates a much cheaper spam-detection classifier the same quarter, both models together cost about six dollars a day. A volume spike there might double the bill for an afternoon and nobody would need to care. Only a migration where the two models' combined cost is big enough to actually move a budget line needs this kind of watching.

The lesson. A migration budget built off an average day doesn't know what to do on the one day that isn't average. If the number can only survive a quiet three weeks, it was never really a number. It was a hope with a dollar sign on it.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why the fifteen hundred dollar afternoon almost went unnoticed.

Kunal Bhandari has run three model migrations at Shutterbrook in as many years, and each time he built the cost model himself, in a spreadsheet, before anyone signed off on a start date. He's the one people ask when a rollout feels rushed. He usually says whether it is, within a day.

The Cormorant-to-Petrel migration started clean. Day one, the billing dashboard read almost exactly what his spreadsheet predicted, two hundred fifty-something dollars. He checked it every morning that first week, cross-referencing actual spend against the plan line by line.

By day eight, it had tracked to plan for a week straight. He started checking every few days instead. By day eleven, with the numbers still calm, he'd stopped opening the dashboard at all unless someone asked. The two earlier migrations had gone exactly this way, quiet after the first week, and this one looked no different.

Marketing had scheduled a "Photographer of the Month" contest for that same quarter, entirely unrelated to the tagging migration, nobody had connected the two calendars. The contest's submission deadline landed on day twelve. For four days, daily uploads ran two and a half times higher than normal, contributors racing to get entries in before the cutoff.

Both models tag every upload during a parallel run. A volume spike doesn't just add to Cormorant's small bill, it multiplies Petrel's much bigger one right alongside it. For those four days, the combined daily cost ran close to six hundred forty-five dollars instead of two hundred fifty-eight, an overage of about fifteen hundred dollars that the flat two-hundred-sixty-dollar alert never crossed, because two hundred sixty was never linked to what a day's uploads actually were, only to what they'd averaged in the plan.

Hand-sketched number line running from zero to eight thousand dollars. A green dot at three thousand four hundred dollars labeled new model only two times pricier. An amber dot at five thousand four hundred dollars labeled the planned three week run. A red orange dot at seven thousand nine hundred dollars labeled new model six times pricier. A separate bracket off to the side marks forty two thousand dollars, labeled one month of the team's whole inference budget, for scale.
The plan called for about $5,400. Even the worst case sits well inside a single month's normal spend, which is exactly why the four-day spike was a scare, not a crisis.

Kunal caught it by accident, on day thirteen, opening the dashboard for an unrelated reason, a question from someone in finance about a different line item entirely. The spike jumped out at him immediately, four days already past.

It wasn't the fifteen hundred dollars that stopped him. It was realizing the alert would have let it run to fifteen thousand and never said a word.

Thirteen hundred dollars over four days isn't a disaster. What stayed with Kunal was the shape of the near miss. The alert had no idea what a normal day actually looked like in real time. It only knew what the plan had assumed a normal day would look like, three weeks earlier, on a spreadsheet. If the contest had run for two weeks instead of four days, or if a second, bigger spike had landed on top of it later in the window, nothing built into the migration would have said a word until someone happened to open the dashboard for some other reason.

Kunal never really had a number for what a spike would do to the bill. He had a habit, a flat two-hundred-sixty-dollar alert, that had worked fine on two calmer migrations before this one. Weeks earlier, at the kickoff meeting, someone had actually asked whether the alert should scale with live volume instead of sitting at one number. The answer at the time was that it would be over-engineering a three-week migration. Nobody pushed back further. It was the only design in the room.

The second version of the alert isn't a bigger flat number. It's pegged to volume itself, roughly four point three cents combined per hundred images, times whatever that day's real upload count actually is, with a flag if the day's total crosses one point four times the rolling baseline. Run the same contest week through that design and the flag fires the first afternoon, not two days after the fact, and the overage stops near four hundred dollars instead of running the full fifteen hundred before anyone looks.

The thing I'd tell myself, back at that kickoff meeting: a budget built off the average day is a budget that has never met a real one.

BOUND: the arithmetic behind the parallel-run invoice

This is a cost build-up and a sizing question, how much running two models on every upload actually costs and for how long, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. Daily parallel-run cost is not one number, it's two added together. Uploads per day times Cormorant's cost per image, plus uploads per day times Petrel's cost per image. Total cost is that daily figure times how many days the window actually needs, and the window itself is sized by how many dual-tagged examples the rarest search category needs before its tags are trustworthy, not by a number of weeks picked off a calendar.
O, own the numbers. Sixty thousand uploads a day. Cormorant costs ninety cents per thousand images, fifty-four dollars a day. Petrel, at our planning assumption of about four times Cormorant's price, costs three dollars forty per thousand, two hundred four dollars a day. Combined, two hundred fifty-eight dollars a day. We looked at shadow-sampling only five percent of uploads through Petrel instead, to cut that cost by most of it, and rejected it: at five percent, the rarest category, industrial drone shots, would produce about half a tagged example a day, and it would take most of a year to gather the two hundred fifty we need. Full parallel tagging on every upload is the only way to gather that category fast enough. The planned window is twenty-one days, sized to that same category at its real rate of about twelve uploads a day. Total: about five thousand four hundred dollars.
U, use a range. If Petrel turns out only twice as pricey as Cormorant instead of four times, the twenty-one-day run costs about thirty-four hundred dollars. If it's six times as pricey, closer to seventy-nine hundred. The window length swings the number harder still: if the drone-shot category shows up at half the assumed rate, six a day instead of twelve, reaching two hundred fifty confirmed examples takes about forty-two days, not twenty-one, and the total roughly doubles, to about ten thousand eight hundred dollars.
N, nail the sanity check. Shutterbrook's whole ML team spends about forty-two thousand dollars a month on inference, across every model it runs. Five thousand four hundred dollars for this one migration is about thirteen percent of a month's total budget. Even the worst case, near eight thousand dollars, stays under a fifth of it. Tolerable, and worth telling finance about in advance rather than letting it show up as a surprise line on next month's invoice.
D, direction. Two things could move this number, and they don't move it equally. The drone-shot category showing up at half its assumed daily rate stretches the window from twenty-one days to about forty-two and adds roughly fifty-four hundred dollars. Petrel turning out pricier than assumed, six times Cormorant instead of four, adds about twenty-five hundred dollars. The rare category's real rate is the bigger lever, and it's the number nobody at Shutterbrook had actually confirmed before the migration started. The four-times assumption for Petrel's price came straight off a vendor pricing page.

One more thing the arithmetic alone doesn't show: cutover isn't a single clean pass either. Search only switches to Petrel once it matches or beats Cormorant's precision on each category's own eval set on at least nine of the last ten weekly checks, not the first week it happens to look fine. An aggregate accuracy number can sit near ninety-eight percent for the whole window while quietly missing keywords on a category too small to move that average, so precision gets tracked by category, not just as one overall number, for as long as the two models run side by side.

The build-up: what $5,400 is actually made of
Cormorant's own cost, 21 days at business as usual$1,134
Petrel's added cost, running alongside for 21 days$4,284
Total parallel-run cost$5,418
The parallel run doesn't double Cormorant's own bill. It adds Petrel's whole bill, close to four times bigger on its own, on top of a cost that was already there.
What moves the total most (swing in dollars from the $5,418 baseline)
Drone-shot category shows up at half the assumed rate (6/day instead of 12/day)+$5,418
Petrel turns out 6x Cormorant's price instead of the assumed 4x+$2,520
Photo-contest volume spike, 2.5x uploads for 4 days inside the window+$1,548
Two extra buffer days added "just in case"+$516
The rare category's real rate swings the total more than Petrel's price does, more than a real volume spike did, and far more than padding the schedule for comfort.

And if you want to be sure it really works, try it somewhere else

Rentford County's permit office runs an AI tool that reads incoming building-permit applications and routes each one to the right review queue: residential, commercial, historic district, and a handful of narrower categories. The office is migrating that router from an older classifier to a newer one, and both models classify every application for a stretch before the county cuts over.

B, break it down. Same shape, smaller numbers. Weekly parallel-run cost equals applications per week times the old model's cost per document, plus applications per week times the new model's cost per document. Weeks needed equals the confirmed catches wanted for the rarest permit type, divided by how often that type actually arrives.
O, own the numbers. Eight hundred applications a week. The old classifier costs four tenths of a cent per document, about three dollars twenty a week. The new one, assumed three and a half times pricier, costs about eleven dollars twenty a week. Combined, fourteen dollars forty a week. Accessory-dwelling-unit permits, the rarest category, arrive about five times a week. Rafiq Sandoval, the PM running this migration, wants forty confirmed dual-classified examples of that category before trusting the new model on it. At five a week, that's eight weeks. Total: about one hundred fifteen dollars.
U, use a range. If the new model turns out six times pricier instead of three and a half, the eight-week run costs about one hundred seventy-nine dollars. If the accessory-dwelling-unit rate is half what's assumed, two and a half a week instead of five, the window stretches to sixteen weeks and the total roughly doubles, to about two hundred thirty dollars.
N, nail the sanity check. Rentford's whole permitting-software budget for the year runs about six thousand dollars. Even the worst case here, two hundred thirty dollars, is under four percent of that. Nowhere near a number that needs a finance conversation, the office could run this migration twice over without anyone noticing the spend.
D, direction. Same tension as Shutterbrook's migration, at a different scale. How often the rare permit type really arrives swings the total more than the new model's price does, because it's the one number nobody had actually measured before the migration started, only assumed from a rough sense of last year's mix.

Total cost: assumed accessory-dwelling-unit rate vs. half the assumed rate
ADU permits arrive as often as assumed (5/week, 8-week window)$115
ADU permits arrive half as often (2.5/week, 16-week window)$230
A much smaller office, a much smaller bill, and the exact same lever doing the swinging: how often the rare category actually shows up, not what the new model costs per document.
Same shape, different lever size At Shutterbrook, the drone-shot category's real rate was worth $5,418 on its own, bigger than the entire baseline plan. At Rentford, the accessory-dwelling-unit rate plays the same role at a scale that barely registers against a six-thousand-dollar annual budget. The equation doesn't change. Whether the swing actually matters depends entirely on how big the underlying bill already is.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: budget both models stacked, not doubled, size the window by how many examples the rarest category needs, check it against the real monthly budget.
Cost: finance says the migration budget is frozen this quarter. Don't shrink the rare-category threshold to fit, shrink the daily upload scope instead, run the parallel tag on a fixed, deliberately over-sampled slice of the rare category's own traffic rather than every upload, and say plainly that the tradeoff is a slower window, not a shakier bar for trusting the new model.
The model got better: Petrel turns out to need barely any correction versus Cormorant. That doesn't remove the need to check the rare category's real rate, it just means the eventual bill leans toward the low end of the range, not that the range stops mattering.

Where people run it wrong.
They price the parallel run by doubling one number, when it's actually two separate bills, of very different sizes, stacked on top of each other.
They pick the window length off the last migration's calendar, without checking whether this migration's rarest category can even produce enough examples that fast.
They set a spend alert off the planned average day and never revisit it, so a real event, a spike, a promo, a seasonal surge, can blow past the plan for days before anyone notices.

How to use it live. Say the equation before naming a single number: "the parallel-run cost isn't the new model's price times two, it's the old model's bill plus the new model's bill, stacked for as long as the rarest category needs to earn enough evidence to trust." That buys the room to ask a real question instead of guessing a lump sum that sounds cautious.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a "cost profile of running two models in parallel" question, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a cost build-up and a sizing question, how much running two models on every upload actually costs and for how long, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Kunal Bhandari, who has run three model migrations at Shutterbrook, a stock-photo marketplace, in as many years. He owns the Cormorant-to-Petrel tagging migration and built the cost model himself.
3 · WHAT THE FIRST PLAN GOT WRONG
What did Kunal's original spend alert assume that turned out to be the wrong basis?
Tap to flip
ANSWER
It set one flat daily-dollar alert, two hundred sixty dollars, worked out from the average sixty thousand uploads, when a real event, a photo contest, could push a day's uploads two and a half times higher and the alert had no way to notice.
4 · THE STRUCTURE IN THIS STORY
What's the difference between sizing the parallel-run window by the calendar and sizing it by evidence?
Tap to flip
ANSWER
The calendar habit picks a number of weeks that felt safe on the last migration. Evidence sizing asks how many dual-tagged examples of the rarest category are needed before trusting the new model on it, which came out to twenty-one days here because that category shows up about twelve times a day.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Setting the spend alert as a flat two hundred sixty dollars a day, built off the average volume. It made sense because two earlier, calmer migrations never saw a volume swing big enough to matter. It stopped making sense once a marketing contest could push a day's uploads well past that average.
6 · THE NUMBER
Fill in the blank: the planned parallel run costs about $___ over ___ days, at ___ uploads a day, and the contest-week spike cost about $___ extra before anyone caught it.
Tap to flip
ANSWER
$5,400. 21 days. 60,000 uploads. About $1,548.
7 · THE REPLAY
Same migration, same team, second design. What changes?
Tap to flip
ANSWER
The spend alert is pegged to live volume instead of a flat number, roughly 4.3 cents combined per hundred images times the day's real upload count, firing if a day crosses 1.4x the rolling baseline. The same contest-week spike gets caught the first afternoon, costing under $400 instead of the $1,548 it actually cost.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same parallel-run cost question for a different product. Which product, and which lever swings it most?
Tap to flip
ANSWER
Rentford County's permit-routing migration, classifying building-permit applications instead of tagging photos. How often the rarest permit type, accessory-dwelling-unit permits, actually arrives swings the total more than the new model's assumed price.

Check yourself Score: 0 / 0

Fill in the blank
1. The planned parallel run costs about $___ over ___ days. The contest-week volume spike added about $___ before Kunal caught it, against a flat daily alert of $___.
Show hint
Check the O step and the story's day-thirteen moment.
Show answer
$5,400; 21 days; $1,548; $260. That gap is exactly why a flat, volume-blind alert never fired while the spike was happening.
Multiple choice
2. Why does the window length swing the total cost harder than Petrel turning out pricier per image?
  • A. The drone-shot category's real upload rate directly decides how many days it takes to gather enough confirmed examples, and that day count multiplies the entire daily cost, while the price ratio only scales one term inside a single day.
  • B. Petrel's price is fixed by contract and can never actually change.
  • C. Cormorant's cost is always the larger of the two costs.
  • D. The sanity check only looks at window length, never at price.
Show hint
Compare what each assumption actually multiplies: a single day's cost, or the number of days themselves.
Show answer
A. A pricier model raises one day's cost. A slower-arriving rare category raises the number of days, and that number multiplies every day's cost at once.
True or false
3. True or false: because the daily cost tracked close to plan for the first week, the parallel run was safe from a volume-driven overrun for the rest of the window.
  • True
  • False
Show hint
Compare "quiet for a week" against what the alert was actually built to notice.
Show answer
False. A quiet first week said nothing about a later, unrelated event. The alert was never linked to live volume, only to the plan's average, so a real spike on day twelve was invisible to it no matter how calm day one through eight had looked.
Short answer
4. Someone on the team says, "Just cap Petrel to a random 10 percent sample of uploads, that cuts the cost almost to nothing." Why does that undercut the whole point of the parallel run?
Show hint
Think about how often the rarest category shows up, not the overall upload volume.
Show answer
Model answer: At ten percent of about twelve drone-shot uploads a day, that's roughly one a day. Reaching the two hundred fifty confirmed examples needed would take close to nine months, far past any reasonable migration window, leaving the exact category most likely to hide a regression essentially untested.
Short answer, apply it yourself
5. Think of a migration or upgrade you could run at your own job, an old and a new version side by side before switching over. What's the rarest real case you'd want enough examples of before trusting the new version, and what would decide how long that takes?
Show hint
Name something that only shows up in a narrow slice of your own work, not something that happens every day.
Show answer
Model answer: A support team piloting a new auto-reply drafter alongside the old one, scaling from one queue to all of them. The rare case would be a bundled-subscription refund dispute, a ticket type that only arrives a few times a week. Trusting the new drafter there would take as long as that ticket type takes to show up enough times to check, not however long the rollout calendar happens to allot.
Short answer, the number question
6. If Shutterbrook's daily upload volume were actually 90,000 instead of 60,000, would the honest 21-day, $5,400 plan still hold? Show the math.
Show hint
Recompute the daily combined cost at the new volume, then check what happens to the rare category's rate too.
Show answer
Close to yes, for a surprising reason. At 90,000 uploads a day, combined daily cost rises to about $387 (Cormorant $81 plus Petrel $306). But the drone-shot category would also show up more often, about 18 a day instead of 12, so the window shrinks to about 14 days instead of 21. Total cost: roughly $387 times 14, about $5,400, almost exactly the original number. The higher daily rate and the shorter window cancel each other out.
Before you close the answer
Why this works
Tests whether you can price an AI migration by its real mechanics, cost stacked not doubled, a window sized by evidence, rather than reciting a lump sum that merely sounds cautious. Most candidates guess a number and stop there.
Follow-up traps
"Why not just throttle Petrel to a random sample of traffic instead of running it on every upload?" Response: we sized a 5 percent sample and rejected it, the rarest category would take most of a year to gather enough examples at that rate, so full parallel tagging is what actually buys the confidence.

"Isn't tracking accuracy by category overkill, why not just watch the overall number?" Response: an aggregate score near 98 percent can hide a category too small to move it, which is exactly the silent, per-category regression the tracking exists to catch before cutover.
If pressed
The "nine of the last ten weekly checks" cutover bar only means something if each category's own eval set is big enough to trust. We required at least 30 examples per category in that eval set before a weekly check counted at all, otherwise a single lucky week on five examples could pass by chance alone.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more