InterviewAdvancedShipping & Model Lifecycle / Model migration and version changes for users / #22

Walk me through a model migration you would run for a product I name.

The direct answer
Run the migration in parallel with a mandatory, non-negotiable spot-check quota on the new model's passes, not just its fails, sized around the categories where a false pass is dangerous, and don't let a benchmark win talk anyone out of it. The habit that breaks a migration isn't supervisors checking less, it's supervisors checking a different, smaller set than they think they are, because "better on the vendor's benchmark" quietly gets heard as "safe to stop watching."
Do this, in order
  1. Set a mandatory spot-check quota on the new model's passes, not just its fails, before migration starts.Why: nothing in a pass/fail dashboard tells anyone a pass was wrong, so a habit of only reviewing fails will always miss a model that's quietly under-flagging.
  2. Weight that quota toward the categories where a false pass is dangerous, not spread it evenly.Why: a missed disclosure to a vulnerable customer costs far more than a missed tone note, and an even spot-check spread would barely touch the category that matters most.
  3. Never let a vendor benchmark number replace a live, category-specific eval on this call center's own calls.Why: the benchmark that sold this migration was built on a different center's call mix, and it said nothing about this center's most common vulnerable-customer scenario.
  4. Make the spot-check quota structural, part of the tool's workflow, not a habit supervisors maintain on their own.Why: an informal habit is the first thing dropped under time pressure, which is exactly what happened here.
  5. Track pass-review findings by category, not as one blended "spot-check accuracy" number.Why: a rare category's regression can hide completely inside a number built mostly from common, low-stakes calls.
  6. Treat a benchmark win as a reason to look harder at what the benchmark didn't test, not a reason to look less.Why: "the new model is better" is exactly the kind of good news that talks people out of the checking habit that would have caught the one place it isn't.

How to answer this, stage by stage

Nobody is grading whether you can list generic migration steps. They are grading whether you can find the specific human habit that breaks when a model gets swapped out from under it. Seven moves get you there.

1
Take the named product and ground it in one real workflow
Say it like this
"Let's say you've named Verdict, Corley Mutual's call-quality scoring tool. It listens to every agent call, about 3,400 a day, and scores it against fourteen compliance checks plus a tone read, passing or flagging each call for a supervisor. Corley's PM, Peregrine Brandão, is migrating Verdict from its first scoring model to a newer one the vendor says catches twelve more compliance issues per hundred calls than the old one."
Why this works
Grounds the named product in a real workflow with real numbers before the migration plan starts.
2
Name the person and their habit before the model, not after
Say it like this
"Before I talk about the migration steps, I want to name who this actually affects day to day. Corley's QA supervisors don't just review the calls Verdict flags as failing, they spot-check a random slice of the calls it passes too, about five percent a week, specifically because a pass is a claim nobody's double-checking by default."
Why this works
This is the F step. Naming the habit before the migration plan is what stops the answer from being a generic checklist.
3
Say what breaks, and why it's a flip and not a dial
Say it like this
"Here's the flip I'd watch for. The new model isn't just an upgrade, it's marketed as better, tested and proven on a vendor benchmark. Good news like that is exactly what makes supervisors quietly drop the pass spot-check. Not slow it down, drop it entirely, because 'the new model catches more' gets heard as 'the passes don't need checking anymore.' That's two settings, checking some passes or checking none, with no middle ground once the habit lapses."
Why this works
This is the I step. Improvement is a perturbation too, and it's the one people forget to guard against.
4
Name the old decision that made the habit fragile
Say it like this
"The decision I'd point to: the five-percent pass spot-check was never built into Verdict's own workflow, it was a habit supervisors kept on their own initiative, with no quota enforced by the tool. That was fine when nothing pressured it. The moment a migration gave everyone a good reason to trust the tool more, the informal habit was the first thing to go, because nothing in the system would have noticed if it quietly stopped."
Why this works
This is the P step, a decision taken back, not a dial turned up. The fix isn't "remind supervisors to keep checking," it's making the check structural.
5
Say what it costs, concretely
Say it like this
"If the pass spot-check quietly stops, the cost doesn't show up as a dip in Verdict's pass rate, it shows up as calls that should have been flagged sitting marked 'pass' with nobody looking. At Corley the sharpest edge is the mandatory cooling-off disclosure to vulnerable customers, a category that's rare, maybe fifteen calls a week, but a miss there is a regulatory problem, not just a coaching note."
Why this works
Names where the cost actually lands, not a vague "quality could suffer."
6
Design the migration to protect the habit, not just the model
Say it like this
"So the migration plan isn't just 'run both models in parallel and compare accuracy.' It's that, plus a mandatory pass spot-check quota built into Verdict's own workflow, weighted so the vulnerable-customer disclosure category gets checked at a much higher rate than five percent, because that's the category where a false pass costs the most and shows up the least."
Why this works
Turns the flip finding into an actual design change, not just an observation.
7
Show the replay, then close
Say it like this
"Run the same migration with that structural quota in place, and a regression in the disclosure category gets caught inside the first two weeks of the parallel run, on the strength of the mandatory check, not three weeks later after a regulator complaint. So here's what I'd actually say: build the spot-check quota into the tool itself, weight it toward the categories a false pass hurts most, and treat a benchmark win as a reason to look harder at what the benchmark didn't cover, not a reason to look less."
Why this works
Closes on a countable result, not a promise to "stay vigilant."
If you remember one thing A model migration doesn't just risk the model being worse. It risks the people around it trusting it more than they should, right when a benchmark win gives them the best reason yet to stop checking.

Let's learn

Say we build a model that listens to a customer-service call and scores it against a compliance checklist: did the agent verify identity correctly, disclose required terms, avoid prohibited language, stay within a required call-closing script. Before a tool like this existed, compliance review meant a supervisor randomly sampling maybe two percent of calls a week by hand, missing almost everything.

Then Verdict went live at Corley Mutual, an insurance call center. Every one of the roughly 3,400 calls a day gets scored automatically against fourteen checks. Supervisors review every flagged fail, and, on their own initiative, spot-check five percent of the passes too, because a tool that only gets reviewed when it says something's wrong has no way to catch a pass it got wrong.

Knowledge spark: what's a false pass? A call the model scores as compliant when it actually wasn't. Unlike a false fail, which gets a human's eyes immediately because it's flagged, a false pass is invisible by default. Nobody looks at a call the system says is fine.

Now Corley is migrating Verdict to a newer scoring model. The vendor's own benchmark shows it catching twelve more compliance issues per hundred calls than the old one, a real, meaningful improvement on paper. But a benchmark built from someone else's call mix says nothing specific about Corley's own rarest, highest-stakes category: the mandatory cooling-off disclosure owed to customers flagged as vulnerable, only about fifteen calls a week region-wide.

Say plainly: the twelve-point benchmark improvement is not the problem. The problem is what a benchmark win does to a habit that was never actually required. The moment supervisors hear "the new model is proven better," the informal five-percent pass spot-check, the only thing standing between a false pass and nobody noticing, is exactly the kind of extra effort that quietly stops.

Good news is a perturbation too. It just doesn't feel like one while it's happening.
The decision that mattered Build the pass spot-check into Verdict's own workflow as a mandatory quota, weighted toward the categories where a false pass is dangerous, instead of leaving it as an informal habit supervisors maintain on their own. A habit with no structure behind it is the first thing dropped under the pressure of a migration everyone's excited about.

At its worst, the new model quietly under-catches the vulnerable-customer disclosure category for weeks, invisible inside a benchmark-driven aggregate improvement, until a missed disclosure surfaces through a regulator complaint instead of an internal check, at which point it's no longer a coaching conversation.

The choice I would take back. When Verdict first launched, the five-percent pass spot-check was set up as a supervisor best practice, documented in the onboarding guide but never built into the tool's own workflow or tracked as a required metric. That felt like enough structure at the time, supervisors were conscientious and the practice held for over a year.

What I would leave alone. Corley also runs a much lower-stakes model that scores call hold-music selection against a brand-consistency checklist. If supervisors stopped spot-checking that one entirely during a migration, nothing regulatory is at stake, and it genuinely wouldn't matter.

The lesson. An informal habit survives right up until something gives everyone a good reason to relax it. A migration billed as an improvement is exactly that kind of reason, which is precisely why the checking habit needs to be built into the system, not left to survive on goodwill.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how close the missed disclosures came to going unnoticed entirely.

Peregrine Brandão has owned Verdict's roadmap at Corley Mutual for two years. Before that, they spent six years as a QA supervisor themselves, which is why the pass spot-check habit existed at all, they built it into the team's culture from the start, even though nothing in Verdict required it.

The migration to the new scoring model was approved on the strength of the vendor's benchmark, a clean twelve-point improvement in overall compliance-issue detection. The rollout plan had both models scoring every call in parallel for three weeks before cutover, which felt thorough. Supervisors kept doing their usual work throughout: review every fail, spot-check some passes.

By the second week, though, something had quietly changed. With the new model flagging fewer calls as failing overall, supervisor review queues were shorter, and several supervisors, without anyone deciding this out loud, stopped pulling their weekly pass sample. It felt reasonable. The new model was proven better. Reviewing its passes started to feel like double-checking a colleague everyone already trusted.

Hand-sketched comparison diagram, two panels. Left panel labeled SPOT-CHECKED with a document icon, captioned one call in twenty every week before the upgrade. Right panel labeled TRUSTED, UNCHECKED with a question-mark box icon, captioned the newer model was assumed better so the checking stopped.
The switch wasn't a decision anyone made in a meeting. It was five separate supervisors, independently, quietly deciding the same thing was safe to stop doing.

Three weeks in, a regulator's office contacted Corley's compliance team about a customer complaint: a vulnerable customer had not received the required cooling-off disclosure on a policy cancellation call. Verdict's new model had scored that call as fully compliant. It was not an isolated case. A retroactive audit of the parallel-run window found six missed disclosures across fifteen vulnerable-customer calls in the new model's flagged-pass bucket, a forty-percent miss rate on the exact category that mattered most, sitting invisible inside an overall compliance-detection number that looked twelve points better than before.

The new model wasn't dishonest about being better. It was honest about being better at everything the benchmark happened to measure.

What struck Peregrine hardest wasn't the six missed calls, it was how ordinary the reason was. Nobody had decided to stop checking. Five supervisors had each, separately, quietly let a habit lapse because the tool gave them a good reason to. If the parallel-run window had been one week shorter, or if the regulator complaint had landed one week later, the migration would have completed on schedule with nobody the wiser.

At the original launch of Verdict, someone on the QA team had actually proposed making the pass spot-check a required, tracked metric inside the tool itself, not just a documented best practice. It was set aside at the time as unnecessary process for a team that was already conscientious. That judgment held for a year, right up until a migration gave everyone a shared, reasonable-sounding excuse to relax at the same time.

The redesigned rollout now builds a mandatory spot-check quota directly into Verdict's supervisor workflow: fifteen percent of passes in the vulnerable-customer disclosure category specifically, every week, non-optional, tracked as its own metric on the team dashboard, separate from the general five-percent quota on everything else. Run the same migration through that design, and the missed-disclosure pattern surfaces inside the first two weeks of the parallel run, through the mandatory check, not through a regulator's phone call in week three.

The thing Peregrine would tell their past self, back when that first proposal was set aside: a habit that depends on everyone staying equally careful, forever, isn't a safeguard. It's a countdown to the first good reason not to be.

The five steps, if you want to remember it

This is a person's checking habit flipping between two settings once a migration gives them a reason to stop, not a sizing or design question, so FLIPS fits and BOUND or SPARK don't.

F, find the person. Corley Mutual's QA supervisors, the ones who review Verdict's flagged fails and, informally, spot-check a slice of its passes.
L, locate the habit. A roughly five-percent weekly spot-check of calls Verdict scored as passing, maintained as an unwritten best practice, not a required, tracked metric.
I, identify the flip. Checking some passes, versus checking none. The trigger wasn't bad news, it was good news: a vendor benchmark showing the new model twelve points better, which made the informal check feel unnecessary and let it lapse without anyone deciding to drop it.
P, pinpoint the old decision. Leaving the pass spot-check as a documented best practice instead of a structural, tracked requirement inside Verdict's own workflow. Reasonable when the team was small and conscientious, fragile the moment a shared, plausible excuse to relax showed up for everyone at once.
S, show the replay. With a mandatory, category-weighted spot-check quota built into the tool, the same missed-disclosure regression surfaces inside the first two weeks of the parallel run instead of via a regulator complaint in week three, six real misses caught by structure instead of caught by luck.

One more thing worth naming: the rejected alternative here was leaning on the vendor's benchmark alone as the migration's quality gate, no live parallel-run spot-checking at all. It was considered, briefly, as a way to speed up the rollout timeline, and rejected because a benchmark built from another center's call mix can't speak to Corley's own rarest, highest-stakes category, exactly the gap that turned out to matter.

Miss rate on the vulnerable-customer disclosure category, old model vs. new
Old scoring model, same category, same 3-week window a year earlier7%
New scoring model, vulnerable-customer disclosure calls, parallel-run window40%
A category the vendor's overall twelve-point benchmark improvement said nothing about, because it wasn't the category the benchmark was built to test.
Pass spot-check rate across the 3-week parallel run
5% 2.5% 0% Week 1, day 1 Week 3, day 5
Nobody decided to stop. The rate just drifted from 5% toward zero as the migration went on and the new model's benchmark reputation did the rest.

And if you want to be sure it really works, try it somewhere else

Thermex Field Services runs an AI tool that recommends a likely fault and repair steps to HVAC technicians before they open up a unit, based on the customer's description and the system's service history. Thermex just migrated that model to a newer version.

F, find the person. Thermex's field technicians, who run the tool's recommendation before every job and, until now, always ran it, regardless of how complex the job looked.
L, locate the habit. Using the tool on every single dispatched job, simple or complex, trusting it as a starting point even for a multi-system fault with several plausible causes.
I, identify the flip. Runs it on every job, versus runs it only on the simple ones. This is a scope flip, not the over-trust flip from Corley's story: the trigger was the new model taking noticeably longer to return a recommendation on multi-system jobs, a real latency cost technicians felt on the clock.
P, pinpoint the old decision. The tool was built to always return one ranked recommendation, simple or complex, with no way to signal "this one's going to take longer to reason through." Technicians had no way to anticipate the delay, so a slow multi-system job felt like the tool being unreliable, and they started skipping it for exactly those jobs, the ones where a second opinion would have helped most.
S, show the replay. A version that flags upfront, before running, whether a job looks complex enough to take longer, and gives technicians a "quick read" option that trades some accuracy for speed on those cases, keeps technicians using the tool on complex jobs instead of abandoning it there entirely. Alaric Marchbanks, the PM who owns this migration, tracked usage on multi-system jobs at 31 percent before the fix and 74 percent after.

Same shape, opposite direction At Corley, good news made a checking habit shrink to nothing. At Thermex, a slowdown made a usage habit shrink to nothing, on exactly the hardest jobs where the tool mattered most. Different flip families, same underlying shape: the case that most needs the tool's help is often the first one people stop trusting it on.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the flip: a benchmark win is exactly the kind of good news that talks people out of the checking habit that would catch the model's real gap, so migrate with a mandatory, category-weighted pass spot-check built into the workflow, not left as a habit.
Cost: the QA team is short-staffed this quarter and can't add review hours. Don't drop the quota to fit the headcount, narrow its scope instead, keep the mandatory check on the highest-stakes category only and drop the general five-percent check temporarily, and say plainly that's a real tradeoff, not a free efficiency gain.
The model got better: the new scoring model turns out to beat the old one on every category, including the vulnerable-customer disclosure. That doesn't make the structural spot-check unnecessary, it's still the thing that would have caught the regression if it had happened, and this time it just confirms there wasn't one.

Where people run it wrong.
They treat a vendor benchmark as proof the migration is safe, when a benchmark built from someone else's data can't speak to this product's own rarest, highest-stakes category.
They plan the migration around the model's accuracy and never ask what happens to the human habits sitting around it.
They leave a safeguard as an informal best practice instead of building it into the tool's own workflow, so it survives right up until something gives everyone a shared reason to let it lapse.

How to use it live. Say the flip before naming a single migration step: "the real risk in this migration isn't the model, it's what a benchmark win does to the humans checking it. Good news is a perturbation too." That buys the room to ask what informal habit around this specific tool might quietly stop, instead of reciting a generic parallel-run checklist.

Flashcards (click a card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Checks sometimes, then stops checking at all, triggered by good news (a benchmark win) rather than bad news, which is what makes it easy to miss.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Peregrine Brandão, who owns Verdict's roadmap at Corley Mutual and, as a former QA supervisor themselves, built the pass spot-check habit into the team's culture in the first place.
3 · THE HABIT
What did supervisors stop doing because it worked?
Tap to flip
ANSWER
Spot-checking a random five percent of Verdict's passing calls every week, an informal habit that had worked fine for over a year before the migration.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Spot-checking some passes every week, versus spot-checking none. Five supervisors independently let the habit lapse once the new model's benchmark reputation made it feel unnecessary.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Leaving the pass spot-check as a documented best practice instead of a mandatory, tracked quota built into Verdict's own workflow. It made sense when the team was small and conscientious, and stopped holding once a migration gave everyone a shared, reasonable-sounding excuse to relax at the same time.
6 · THE NUMBER
Fill in the blank: the vendor benchmark showed a ___-point improvement, but the retroactive audit found a ___% miss rate on the vulnerable-customer disclosure category, versus ___% on the old model.
Tap to flip
ANSWER
12-point improvement. 40% miss rate on the new model. 7% on the old model.
7 · THE REPLAY
Same migration, same team, second design. What changes?
Tap to flip
ANSWER
A mandatory 15% spot-check quota on the vulnerable-customer disclosure category specifically, built into Verdict's workflow and tracked as its own metric. The regression surfaces inside the first two weeks of the parallel run, through the structural check, instead of through a regulator complaint in week three.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same migration question for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Thermex Field Services' HVAC fault-diagnosis tool. Scope flip: technicians stopped using the tool on complex, multi-system jobs once the new model got noticeably slower on exactly those cases.

Check yourself Score: 0 / 0

Multiple choice
1. What was the actual flip in this story?
  • A. Verdict's scoring accuracy dropped after migration.
  • B. Supervisors went from spot-checking about 5% of passing calls every week to checking effectively none, once the new model's benchmark reputation made the habit feel unnecessary.
  • C. Corley Mutual stopped using Verdict entirely during the migration.
  • D. The vendor's benchmark numbers turned out to be fabricated.
Show hint
The flip is a human behavior change, not a claim about the model's own accuracy.
Show answer
B. The model's benchmark numbers were real. What flipped was a human habit, in response to good news about the model, not bad news.
Fill in the blank
2. Verdict scores roughly ___ calls a day against ___ compliance checks. The rare, high-stakes category at the center of this story is the ___ disclosure, which showed up in about ___ calls a week.
Show hint
Check the walkthrough's first stage and the "Let's learn" section.
Show answer
3,400 calls; 14 checks; mandatory cooling-off disclosure; about 15 calls a week. A category small enough to hide almost entirely inside an aggregate improvement number.
True or false
3. True or false: because the new model's overall compliance-detection rate was 12 points better than the old model, it was reasonable to trust it more on every category, including the vulnerable-customer disclosure.
  • True
  • False
Show hint
Think about what the benchmark that produced the 12-point number actually measured.
Show answer
False. The benchmark was built on a different call mix and said nothing specific about Corley's own rarest category. An aggregate improvement is compatible with a real regression on a small slice, exactly what happened here.
Short answer
4. Someone proposes fixing this by just telling supervisors "please keep spot-checking passes even after the migration." Why does that not actually fix the problem?
Show hint
Think about why the habit lapsed in the first place, and whether a reminder changes that.
Show answer
Model answer: The habit was already informal and undocumented as a requirement before the migration, and it lapsed anyway under a plausible, shared excuse. A verbal reminder doesn't change the fact that nothing in the system would notice or enforce it if the habit lapsed again under the next plausible excuse. The fix has to be structural: a tracked, mandatory quota inside the tool itself.
Short answer, apply it yourself
5. Think of a tool or process at your own job that got a real, well-earned upgrade recently. What habit around it might have quietly gotten looser simply because the upgrade felt trustworthy?
Show hint
Look for a check or review step that existed mainly out of caution, not because anyone was told to do it.
Show answer
Model answer: A finance team that upgraded its expense-report anomaly detector to a version with a much lower false-positive rate. The habit that likely loosened: a manager spot-checking a sample of reports the old tool didn't flag, since the old tool's high false-positive rate made "check some unflagged ones too" feel necessary, and the new tool's calmer flag rate makes that same habit feel like redundant effort.
Short answer, the number question
6. If the vulnerable-customer disclosure category had shown up in 150 calls a week instead of 15, would the same regression have been as likely to hide inside the aggregate 12-point improvement number? Why or why not?
Show hint
Think about what share of total call volume the category would represent at each size.
Show answer
Less likely. At 15 calls a week out of 3,400 daily calls, the category is under half a percent of total volume, easily swamped by an aggregate number. At 150 calls a week it would be closer to five percent of volume, large enough that a 33-point swing in its own miss rate (from 7% to 40%) would very likely have visibly dented the overall compliance-detection number, making the regression harder to miss even without the structural spot-check.
Before you close the answer
Why this works
Tests whether you can name the human habit a migration puts at risk, not just the model's own accuracy. Most candidates plan a migration around the model; the strong answer plans it around what the people relying on the model are likely to quietly stop doing once it looks trustworthy.
Follow-up traps
"Isn't a mandatory spot-check quota just extra process that slows the team down?" Response: the quota replaces an informal habit that already existed and already cost the same review time, it just makes that time non-optional on the category where skipping it is most dangerous.

"Couldn't you catch this with a bigger, better benchmark instead of relying on live spot-checks?" Response: a benchmark, however large, is still frozen at the moment it was built; a live spot-check catches whatever the model does on this center's actual calls today, including drift a static benchmark can't see.
If pressed
The redesigned quota isn't flat across every category either, it scales with how expensive a false pass is estimated to be: fifteen percent on the vulnerable-customer disclosure category, five percent on general compliance checks, and one percent on low-stakes checks like call-closing script adherence, where a false pass costs almost nothing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more