ConceptAdvancedShipping & Model Lifecycle / Model migration and version changes for users / #5
Describe a dual-running strategy for a model migration.
The direct answer
Run the new model next to the old one on the same live traffic, for a set window: the old model's call still reaches the customer, the new model's score only gets logged and checked. Write the pass bar before that window opens, checked by traffic segment and not one overall number, and cut over only once it clears that bar. A win on a test set is a prediction. Real customers, scored side by side, are what prove it holds.
Do this, in order
Run the new model next to the old one on the same live traffic, for a set window, before switching anything.Why: a win on an offline test is a prediction about live traffic, not proof of it, and only scoring both on the same real requests settles which one is true.
Keep the old model's output as the one that actually reaches the customer during that window.Why: the new model gets to prove itself on every real interaction without a single customer ever seeing its call if it's wrong.
Set the pass bar, checked by segment, before the window opens, not after.Why: an overall agreement number can look healthy while hiding a bad miss on one smaller, newer slice of traffic, which is exactly what a single aggregate score can't show.
Fix the window at a length long enough for a rare segment to show up, and not a day longer than that.Why: too short and a rare miss never gathers enough samples to prove itself; too long and the business pays for two models while the cheaper one's savings sit on the table.
Never let a subjective read, "it just sounds better," override the pass bar.Why: a nicer-sounding answer and a correct one are not the same claim, and the whole point of the window is checking that difference with real evidence, not a feeling.
Don't build a full new comparison for traffic where the two models already agree at a stable, high rate.Why: that's caution with no target, and it slows down segments that were never actually at risk.
How to answer this, stage by stage
Seven moves. Anchor it to the loyalty complaint a fast launch almost let through, not a general pitch for "test it carefully."
1
Scope it and say your structure in one breath
Say it like this
"So I'm Teodora, I own model quality for Tenor, that's the sentiment tool inside Halstead's contact center, a home-goods chain with three hundred stores. I'm not going to give you a general 'test it carefully' answer. I'll walk through one real design: run the new model next to the old one on the same live traffic for three weeks, old model still makes the call, new model just gets scored. Then I'll show you the week that almost proved why."
Why this works
A scoped example gives the interviewer something to picture, and the structure tells them you have a plan before you've said a single detail.
2
Reframe what an offline win is actually worth
Say it like this
"Everyone wants to hear 'the new model scored ninety against the old one's eighty-two on our test set, ship it.' That's real, but it's not proof. That test was built sixteen months ago. A model can win an old exam and still be wrong about something the exam never asked."
Why this works
This is the actual insight being tested. Skip it and you're describing a QA checklist, which nobody would disagree with and nobody learns from.
3
Give the anchor
Say it like this
"Here's the design. Both models score every one of our twelve thousand weekly calls and chats. The old model's flag is still what pings a store manager, so nothing changes for a customer during the window. The new model's flag just gets logged. Before we start, we write down the pass bar: it has to match a hand-graded review, checked by traffic segment, not just overall, or it doesn't graduate."
Why this works
Naming the specific mechanism, out loud, is what separates a real rollout design from a vague "we'll compare them" gesture.
4
Walk through the near miss
Say it like this
"Week one, the two models agreed on ninety-four percent of calls, which looked healthy. But Nnamdi's team hand-reviews a sample of the disagreements every week, and fourteen of the sixty they checked that week mentioned our new loyalty program. Eleven of those fourteen, the new model called it neutral. A person reading the same transcript called it a clear escalation, someone furious about points that vanished off her account."
Why this works
A real, specific example makes the risk concrete instead of theoretical, and it shows exactly what an aggregate agreement number is blind to.
5
Say what changed because of it
Say it like this
"So we didn't ship on the ninety-four percent. We retrained the new model on fresh loyalty-complaint transcripts and reran the check the next week on a wider sample. Forty-five loyalty-related disagreements, only two were still wrong. That's the number I wanted on the wall before the escalation flag changed at all three hundred stores."
Why this works
It shows the rollout actually responding to what it found, with a number that moved because of a real fix, not a report nobody acted on.
6
Say the trade-off you're accepting, out loud
Say it like this
"I'll name the trade-off, because it's real. The new model is forty percent cheaper per call and about twice as fast, that's the whole reason we're moving. Running both models for three weeks means paying for both at once. I'm accepting that cost on purpose, because it's a lot cheaper than finding out about the loyalty gap after all three hundred stores are already on the new model."
Why this works
Naming the cost you're accepting, not pretending the decision is free, is what separates a real trade-off from a wish that speed, cost, and quality could all win at once.
7
Close on the number and what's kept off the table
Say it like this
"So: three weeks, same live traffic, old model still makes the call, new model gets checked by segment before it earns one. Eleven of fourteen wrong before the fix, two of forty-five after. And the thing I never do is sign off because the new model's answers just read better. It graduates on the bar, not on the vibe."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.
Let's learn
What happens the first time a model that already won on your test set gets a real complaint wrong, and nobody's checking, because the test already said it was the better one?
Halstead runs three hundred home-goods stores. Tenor is the tool built into its contact center: it reads every customer call and chat and flags the ones where the customer sounds ready to leave a bad review, or just leave, so a store manager can call back before either happens.
Today, without dual-running
Before Tenor, a manager found out about an unhappy customer the slow way: a bad review posted days later, or a canceled account showing up in a monthly report. Tenor has run for sixteen months on its current model, and on its own five-hundred-ticket hand-graded test set, it catches a real escalation correctly eighty-two times out of a hundred.
Knowledge spark: what's dual-running?
Running two models on the same real traffic at the same time, but only letting one of them actually decide anything. The other one just gets watched and graded, quietly, until it's earned the right to take over.
A newer model scored ninety out of a hundred on that same test set. It also costs forty percent less to run and answers in about half the time. On paper, that's an easy switch.
What the test set said, and what dual-running found instead
the number that looked like a winthe number that carried the real riskthe same segment, after the fix
The old test set never once suggested a problem. But eleven of the first fourteen loyalty-related disagreements in week one were the new model missing a real escalation, and a segment-level check, not luck, is what took that down to two of the next forty-five.
We didn't test the new model on our customers. We tested it on last year's complaints.
At its worst, this doesn't show up as an outage. It shows up as thousands of angry customers a month getting a calm, confident, wrong answer about their points, and the tool built to catch that mood swing sailing right past it, because a passing grade on an old test looked like proof.
The decision that mattered
Run the new model next to the old one on live traffic for a set window, with the old model still making the real call, and require a pass bar checked by segment before cutover. A test-set win says nothing about a kind of complaint the test never saw.
The choice I would take back. The plan graduated the new model off one number: does it beat the old score on the standing test set. That made sense for every earlier update, because nothing about Tenor's traffic had really changed between tests. It stopped making sense the moment a whole new kind of complaint existed that the old test had barely seen, six loyalty-related tickets out of five hundred, next to about one in eleven live calls today.
What I would leave alone. Return and delivery complaints, the bulk of Tenor's traffic, already sat close to the old model's numbers going into the window, and a light spot check was enough there. Building a full new segment-level check for traffic that was never actually at risk would have slowed the whole migration down for no real reason.
The lesson. A test set is only as current as the day it was built. If the product it grades keeps meeting new kinds of complaints, the test has to keep up too, or a model can pass a real exam and still be wrong about the thing customers are calling about this week.
Now here is the same thing as a story
The short version is above. Read this one when you've got a few minutes, for why it mattered.
Teodora Vasileva had owned Tenor's model quality for a year and a half by the time this decision came up, long enough to know the standing number cold: eighty-two out of a hundred real escalations caught, checked every quarter against the same five-hundred-ticket test set. It was a real number, earned slowly, and it was the reason nobody at Halstead questioned Tenor anymore.
Finance wanted the switch fast. The new model was forty percent cheaper to run and answered in half the time, and it had just scored ninety on the same test set Tenor had always been graded against. The plan on the table did what had worked with every earlier model update: run it, check the score, ship it.
The anchor: two models, one live traffic, only one of them deciding anything
Teodora didn't ship on the score alone. She set a three-week window. Both models would score every one of Tenor's twelve thousand weekly calls and chats. The old model's flag would keep going to store managers, exactly as before. The new model's flag would just get logged, quietly, next to it.
Week one, the overall agreement between the two models read ninety-four percent. Close, healthy-looking, the kind of number that makes a launch date start to feel like a formality.
Nnamdi Okoye runs Tenor's quality review team, and he doesn't trust an aggregate number he hasn't opened himself. Every week, his team pulls a sample of the calls where the two models disagreed and reads the transcript the way a manager would, not the way a percentage counts it.
Fourteen of that week's sixty sampled disagreements mentioned Halstead Plus, the loyalty points program Halstead had launched two months earlier. One was a customer who'd had two hundred points vanish off her account and called in furious about it. Tenor's new model scored the call neutral. The old model, still running underneath, had flagged it correctly, so a manager called her back that same afternoon.
The day it's wrong, and what catches it before a customer ever sees it
We didn't test the new model on our customers. We tested it on last year's complaints.
Nnamdi didn't stop at one. His team pulled every disagreement that week that touched loyalty points, fourteen in all. Eleven of them, the new model had missed. Every one of those eleven would have gone straight past a manager if the new model's flag had been the one that counted that week.
That's the part the overall agreement number could never have shown. Ninety-four percent agreement sounds like two models mostly seeing the same thing. It says nothing about which six percent of the disagreements were the ones that actually mattered.
So here is the decision Teodora took back.
The plan graduated a model update off one number: does it beat the old score on the standing test set. That made sense for every update before this one, because nothing about Tenor's traffic had really changed between tests. It stopped making sense the moment a new kind of complaint existed that the standing test had barely seen.
Teodora didn't pull the new model. She had it retrained on a fresh batch of loyalty complaints and reran the check the next week on a wider sample. Forty-five loyalty-related disagreements, checked by hand. Two were wrong, not eleven. That's the number that let Tenor's new model earn its cutover, segment by segment, not on the strength of a single quarter-old test.
And the thing I'd want to tell myself, back when ninety on the old test looked like enough to greenlight a switch for three hundred stores: a model beating an old exam only tells you it learned that exam. It never tells you whether the world sitting behind it moved.
SPARK, in one screen
This question sounds like it wants a QA checklist, "test the new model carefully before switching." It's really asking for one rollout design that survives a model winning on paper and still being wrong about something that paper never saw, so SPARK fits. A question asking how Teodora would keep watching Tenor a year after the cutover would reach for LEAD instead.
S, situation. Today, before this rollout design, a model update graduates off one number: does the new model beat the old one on Tenor's standing five-hundred-ticket test set. That test was built sixteen months ago and reflects the traffic from back then. It has no real way to tell you how the new model handles the complaints Halstead's customers are calling about this month.
P, payoff. Not "ship a cheaper, faster model." The habit worth building: catching a live-traffic-only regression, a wrong call on a kind of complaint the old test never saw, while it's still sitting quietly in a log next to the old model's correct answer, not after it's the only answer three hundred stores are getting.
A, anchor. Both models score every live call and chat for three weeks. The old model's flag still reaches the store manager; the new model's flag is only logged and scored. A pass bar is set before the window opens: the new model has to match hand-graded review at or above a set rate, checked by traffic segment, not by one overall number. Cutover happens once every segment worth watching clears its own bar.
R, risk. Run the window too long, and Halstead pays for both models at once while the cheaper one's savings sit on the table for no added safety. Run it too short, and a rare pattern, a complaint type that only shows up a handful of times a week, never gathers enough samples to prove itself either way, and it ships broken to every store before anyone finds out.
K, keep out. Don't let dual-running decide the cutover on a subjective read, "the new model's replies sound sharper," or a launch date already on the calendar. The window exists to produce a number checked against real evidence. If that number hasn't cleared its bar, a nicer-sounding answer isn't a reason to switch early.
What we left for later, kept visibly separate from day one
Why the anchor survives the risk
Check it against the near miss. Does a segment-level pass bar still catch the danger even when the overall agreement number looks healthy? Yes, that's the whole reason it's checked by segment and not just once, in total. Does it avoid slowing down traffic that was never actually at risk? Yes, because the bar only demands a fresh, hand-checked look where a segment is genuinely new or thin in the old test, not everywhere equally.
And if you want to be sure it really works, try it somewhere else
A fraud-detection model at a regional credit union runs on a completely different clock, but the same gap between a good test score and a good live call shows up the moment a new pattern arrives that the test never saw.
S. Ana Petrova leads fraud model quality at Cordell, a regional credit union with about sixty thousand members. Today, without this rollout design, an analyst reviews roughly four hundred transactions a week that the current model flags, out of sixty thousand total, a model that's run for three years with no live comparison ever built for a model update. P. The habit worth building: catch a live-traffic fraud pattern the offline test never saw, a new scam type moving real money, while it's still only costing a handful of members, not after a month of losses shows up in a report nobody can undo. A. Same shape, higher stakes. The new fraud model scores every live transaction for four weeks, a longer window than Tenor's because real money is moving. The old model's flag still triggers the analyst's review. The new model's flag is only logged, and it has to clear its own hand-checked bar on every fraud type before it earns the call. R. Too cautious, and members wait an extra month for a model that already proved itself by week three, while Cordell pays for two models the whole time. Too loose, and a slow-building scam pattern, gift-card and romance-scam transfers, keeps slipping through, because the new model's real performance on that pattern was never actually checked, only assumed from an old test. K. Compliance doesn't get to sign off because the new model's fraud explanations read more convincingly in a demo. Sign-off happens when the fraud-type bar clears, not when the write-up sounds good.
Catch rate on gift-card and romance-scam transfers, week by week
before the fraud-specific check existedthe check rolling out, still being tunedafter it's tuned to this scam pattern
A test score built on last year's scams would have called week one "ready." Catching this specific pattern only climbed once real transactions, not the old test, started grading the new model.
Swap the trigger and it still runs
Speed: even if the new model answered in a tenth of the time, that wouldn't tell you whether it catches a gift-card scam the old test never saw. Speed and live accuracy are different questions.
Cost: if the new model cost nothing at all to run, that still wouldn't prove it handles Halstead Plus complaints correctly. Free doesn't mean checked.
The model gets better: if the new model's overall test score climbed to ninety-nine, that still wouldn't guarantee it caught this month's version of a pattern it has never actually been graded against.
Where people run it wrong
Treating a win on the standing test set as proof the new model is ready for everything the old one handled.
Watching only the overall agreement number during the window and never opening a sample of the actual disagreements.
Running the same long comparison on every segment equally, including ones that were never actually at risk, which just slows the whole migration down.
How to use it live
If you're asked this cold, pick one real thing your product started handling only recently, a new program, a new complaint type, a new transaction pattern, and ask whether your last eval was built before or after it existed. That question, asked of yourself, finds the real gap faster than describing a general testing plan.
Flashcards (click a card to flip it)
1 · THE SITUATION
What's the situation, before this rollout design?
Tap to flip
ANSWER
Tenor already catches 82 of 100 real escalations on a 16-month-old test set, but that set has only 6 loyalty-related tickets out of 500, while live loyalty complaints now make up about 1 in 11 calls.
2 · THE PAYOFF
What's the real habit this rollout is trying to build?
Tap to flip
ANSWER
Catching a live-traffic-only regression while it's still sitting quietly in a log next to the old model's correct answer, not after it's the only answer three hundred stores are getting.
3 · THE ANCHOR
What's the one design decision everything else hangs on?
Tap to flip
ANSWER
Run both models on every live call for three weeks. The old model's flag still reaches the manager; the new model's flag only gets logged and must clear a segment-level, hand-checked pass bar before it earns the cutover.
4 · THE RISK
What breaks if the dual-run window is too long, or too short?
Tap to flip
ANSWER
Too long, Halstead pays for two models at once and delays the cheaper model's savings with no added safety. Too short, a rare pattern like loyalty complaints never gathers enough samples to prove itself, and ships broken to every store.
5 · THE PROOF
What did the hand review find that the overall agreement number never would have?
Tap to flip
ANSWER
Overall agreement read 94 percent, healthy-looking. But 11 of 14 loyalty-related disagreements that week were the new model missing a real escalation, hidden completely inside a fine-looking aggregate.
6 · THE NUMBER
___ of ___ loyalty-related disagreements were wrong before the fix. ___ of ___ were wrong after.
Tap to flip
ANSWER
11 of 14 before. 2 of 45 after.
7 · THE REPLAY
Same loyalty complaint, the segment-level pass bar already in place. What changes?
Tap to flip
ANSWER
The new model's miss gets caught in the log and fixed before cutover, checked on purpose against a bar built for that segment, instead of found by luck after all three hundred stores are already relying on it.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor test?
Tap to flip
ANSWER
Cordell Credit Union's fraud-detection model migration. Its anchor tests whether the new model actually catches a live scam pattern, not just whether it scored well on an old fraud test.
Check yourself Score: 0 / 0
True or false
1. True or false: because the two models agreed on 94 percent of calls in week one of the dual run, that was good evidence Tenor's new model was ready to take over the escalation flag at all 300 stores.
True
False
Show hint
Think about what an overall agreement rate actually measures, and what Nnamdi's hand review found hiding inside it.
Show answer
False. A 94 percent agreement rate says nothing about which disagreements matter. 11 of 14 loyalty-related disagreements that same week were the new model missing a real escalation.
Multiple choice
2. Which rollout design matches the anchor this answer argues for?
A. Compare the new model's score to the old model's score on the standing test set, and switch once it's higher.
B. Run both models on the same live traffic for a set window, keep the old model's flag live, and require the new model to clear a segment-level, hand-checked bar before cutover.
C. Wait until the new model matches the old model's score on every single segment before turning it on anywhere.
D. Skip the comparison window and switch straight to the new model everywhere, since it already won on the standing test.
Show hint
The anchor needs a check built for live traffic segments, not just a check for whether the new model beat an old score.
Show answer
B. A checks only the old test's signal, which is exactly what almost hid the loyalty gap. C is the over-applied risk, needlessly slow-walking every segment to one shared bar. D removes any check before the harm reaches customers.
Fill in the blank
3. ___ of ___ loyalty-related disagreements were wrong before the Halstead team retrained the new model, which fell to ___ of ___ after.
Show hint
This number shows up twice, once in the story, once in the chart.
Show answer
11 of 14; 2 of 45. A segment-level check built for loyalty complaints, not the standing test score, is what took the team from catching problems by luck to catching them on purpose.
Short answer
4. What old decision does this rollout design take back, and why did it make sense when it was first made?
Show hint
Think about which signal every earlier model update had already trusted, before anyone had reason to check whether it still meant something on a new kind of traffic.
Show answer
Model answer: The original plan graduated any model update off one number, whether it beat the old model on the standing test set. That made sense for every earlier update, because Tenor's traffic hadn't really changed between tests. It stopped making sense once a new kind of complaint, loyalty points, existed that the standing test had barely seen.
Short answer, apply it yourself
5. Pick a product you use yourself, or one your team is building. What's one place a good score on an old test could be hiding a real gap in how it handles something newer?
Show hint
Look for something the product started handling only recently, a new feature, a new policy, a new kind of request, that the last real test might predate.
Show answer
Model answer: "Our internal support bot scores well on last year's ticket set, but we launched a new subscription tier two months ago, and the bot's never really been checked against questions about it, so a good overall score could just be an old product wearing a new one's name."
Multiple choice
6. Based on this answer's own numbers, if week one's loyalty disagreement count had been 2 of 14 wrong instead of 11 of 14, would running the three-week dual run still have been the right call?
A. Yes, because the point of the window is knowing the real rate on a segment the old test never covered, whether it turns out high or low.
B. No, at 2 of 14 the team should have just trusted the standing test score and skipped the window.
C. No, a lower miss rate means the loyalty segment was never really a risk in the first place.
D. Yes, but only because a higher number would have looked worse to Finance.
Show hint
Compare what a lower miss rate changes about the need to know it, against what it changes about whether a real, live check was still worth running.
Show answer
A. A lower miss rate doesn't remove the need to know it, or the need for a real segment-level check. It would still have been the wrong rollout to find that out by watching the standing test score alone.
If the interviewer pushes back
Why this works
Tests whether you'll trust a model that already won on paper, or insist on watching it work on today's traffic before it touches a real customer. Most candidates stop at the eval score because it's the number that's easiest to defend in a meeting.
Follow-up traps
"Isn't a three-week window just slowing down a model you already know is better?"
Response: The eval score proves it's better on last year's traffic. The window is what proves it's better on this month's, and that's not the same claim.
"What if the two models never fully agree, even once the new one is genuinely fine? Doesn't that mean the bar is unreachable?"
Response: The pass bar isn't zero disagreement, it's a hand-checked rate per segment. Two good models will still disagree sometimes, which is exactly why the bar is a rate and not a perfect match.
If pressed
The segment-level bar itself was set at needing the new model's hand-checked accuracy within about three points of the old model's, on any segment worth at least two percent of weekly volume, not matched exactly, because two independently trained models will never agree completely, even when both are working fine.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.