Walk me through a model migration you would run for a product I name.
- Set a mandatory spot-check quota on the new model's passes, not just its fails, before migration starts.Why: nothing in a pass/fail dashboard tells anyone a pass was wrong, so a habit of only reviewing fails will always miss a model that's quietly under-flagging.
- Weight that quota toward the categories where a false pass is dangerous, not spread it evenly.Why: a missed disclosure to a vulnerable customer costs far more than a missed tone note, and an even spot-check spread would barely touch the category that matters most.
- Never let a vendor benchmark number replace a live, category-specific eval on this call center's own calls.Why: the benchmark that sold this migration was built on a different center's call mix, and it said nothing about this center's most common vulnerable-customer scenario.
- Make the spot-check quota structural, part of the tool's workflow, not a habit supervisors maintain on their own.Why: an informal habit is the first thing dropped under time pressure, which is exactly what happened here.
- Track pass-review findings by category, not as one blended "spot-check accuracy" number.Why: a rare category's regression can hide completely inside a number built mostly from common, low-stakes calls.
- Treat a benchmark win as a reason to look harder at what the benchmark didn't test, not a reason to look less.Why: "the new model is better" is exactly the kind of good news that talks people out of the checking habit that would have caught the one place it isn't.
How to answer this, stage by stage
Nobody is grading whether you can list generic migration steps. They are grading whether you can find the specific human habit that breaks when a model gets swapped out from under it. Seven moves get you there.
Let's learn
Say we build a model that listens to a customer-service call and scores it against a compliance checklist: did the agent verify identity correctly, disclose required terms, avoid prohibited language, stay within a required call-closing script. Before a tool like this existed, compliance review meant a supervisor randomly sampling maybe two percent of calls a week by hand, missing almost everything.
Then Verdict went live at Corley Mutual, an insurance call center. Every one of the roughly 3,400 calls a day gets scored automatically against fourteen checks. Supervisors review every flagged fail, and, on their own initiative, spot-check five percent of the passes too, because a tool that only gets reviewed when it says something's wrong has no way to catch a pass it got wrong.
Now Corley is migrating Verdict to a newer scoring model. The vendor's own benchmark shows it catching twelve more compliance issues per hundred calls than the old one, a real, meaningful improvement on paper. But a benchmark built from someone else's call mix says nothing specific about Corley's own rarest, highest-stakes category: the mandatory cooling-off disclosure owed to customers flagged as vulnerable, only about fifteen calls a week region-wide.
Say plainly: the twelve-point benchmark improvement is not the problem. The problem is what a benchmark win does to a habit that was never actually required. The moment supervisors hear "the new model is proven better," the informal five-percent pass spot-check, the only thing standing between a false pass and nobody noticing, is exactly the kind of extra effort that quietly stops.
At its worst, the new model quietly under-catches the vulnerable-customer disclosure category for weeks, invisible inside a benchmark-driven aggregate improvement, until a missed disclosure surfaces through a regulator complaint instead of an internal check, at which point it's no longer a coaching conversation.
The choice I would take back. When Verdict first launched, the five-percent pass spot-check was set up as a supervisor best practice, documented in the onboarding guide but never built into the tool's own workflow or tracked as a required metric. That felt like enough structure at the time, supervisors were conscientious and the practice held for over a year.
What I would leave alone. Corley also runs a much lower-stakes model that scores call hold-music selection against a brand-consistency checklist. If supervisors stopped spot-checking that one entirely during a migration, nothing regulatory is at stake, and it genuinely wouldn't matter.
The lesson. An informal habit survives right up until something gives everyone a good reason to relax it. A migration billed as an improvement is exactly that kind of reason, which is precisely why the checking habit needs to be built into the system, not left to survive on goodwill.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how close the missed disclosures came to going unnoticed entirely.
Peregrine Brandão has owned Verdict's roadmap at Corley Mutual for two years. Before that, they spent six years as a QA supervisor themselves, which is why the pass spot-check habit existed at all, they built it into the team's culture from the start, even though nothing in Verdict required it.
The migration to the new scoring model was approved on the strength of the vendor's benchmark, a clean twelve-point improvement in overall compliance-issue detection. The rollout plan had both models scoring every call in parallel for three weeks before cutover, which felt thorough. Supervisors kept doing their usual work throughout: review every fail, spot-check some passes.
By the second week, though, something had quietly changed. With the new model flagging fewer calls as failing overall, supervisor review queues were shorter, and several supervisors, without anyone deciding this out loud, stopped pulling their weekly pass sample. It felt reasonable. The new model was proven better. Reviewing its passes started to feel like double-checking a colleague everyone already trusted.
Three weeks in, a regulator's office contacted Corley's compliance team about a customer complaint: a vulnerable customer had not received the required cooling-off disclosure on a policy cancellation call. Verdict's new model had scored that call as fully compliant. It was not an isolated case. A retroactive audit of the parallel-run window found six missed disclosures across fifteen vulnerable-customer calls in the new model's flagged-pass bucket, a forty-percent miss rate on the exact category that mattered most, sitting invisible inside an overall compliance-detection number that looked twelve points better than before.
What struck Peregrine hardest wasn't the six missed calls, it was how ordinary the reason was. Nobody had decided to stop checking. Five supervisors had each, separately, quietly let a habit lapse because the tool gave them a good reason to. If the parallel-run window had been one week shorter, or if the regulator complaint had landed one week later, the migration would have completed on schedule with nobody the wiser.
At the original launch of Verdict, someone on the QA team had actually proposed making the pass spot-check a required, tracked metric inside the tool itself, not just a documented best practice. It was set aside at the time as unnecessary process for a team that was already conscientious. That judgment held for a year, right up until a migration gave everyone a shared, reasonable-sounding excuse to relax at the same time.
The redesigned rollout now builds a mandatory spot-check quota directly into Verdict's supervisor workflow: fifteen percent of passes in the vulnerable-customer disclosure category specifically, every week, non-optional, tracked as its own metric on the team dashboard, separate from the general five-percent quota on everything else. Run the same migration through that design, and the missed-disclosure pattern surfaces inside the first two weeks of the parallel run, through the mandatory check, not through a regulator's phone call in week three.
The thing Peregrine would tell their past self, back when that first proposal was set aside: a habit that depends on everyone staying equally careful, forever, isn't a safeguard. It's a countdown to the first good reason not to be.
The five steps, if you want to remember it
This is a person's checking habit flipping between two settings once a migration gives them a reason to stop, not a sizing or design question, so FLIPS fits and BOUND or SPARK don't.
F, find the person. Corley Mutual's QA supervisors, the ones who review Verdict's flagged fails and, informally, spot-check a slice of its passes.
L, locate the habit. A roughly five-percent weekly spot-check of calls Verdict scored as passing, maintained as an unwritten best practice, not a required, tracked metric.
I, identify the flip. Checking some passes, versus checking none. The trigger wasn't bad news, it was good news: a vendor benchmark showing the new model twelve points better, which made the informal check feel unnecessary and let it lapse without anyone deciding to drop it.
P, pinpoint the old decision. Leaving the pass spot-check as a documented best practice instead of a structural, tracked requirement inside Verdict's own workflow. Reasonable when the team was small and conscientious, fragile the moment a shared, plausible excuse to relax showed up for everyone at once.
S, show the replay. With a mandatory, category-weighted spot-check quota built into the tool, the same missed-disclosure regression surfaces inside the first two weeks of the parallel run instead of via a regulator complaint in week three, six real misses caught by structure instead of caught by luck.
One more thing worth naming: the rejected alternative here was leaning on the vendor's benchmark alone as the migration's quality gate, no live parallel-run spot-checking at all. It was considered, briefly, as a way to speed up the rollout timeline, and rejected because a benchmark built from another center's call mix can't speak to Corley's own rarest, highest-stakes category, exactly the gap that turned out to matter.
And if you want to be sure it really works, try it somewhere else
Thermex Field Services runs an AI tool that recommends a likely fault and repair steps to HVAC technicians before they open up a unit, based on the customer's description and the system's service history. Thermex just migrated that model to a newer version.
F, find the person. Thermex's field technicians, who run the tool's recommendation before every job and, until now, always ran it, regardless of how complex the job looked.
L, locate the habit. Using the tool on every single dispatched job, simple or complex, trusting it as a starting point even for a multi-system fault with several plausible causes.
I, identify the flip. Runs it on every job, versus runs it only on the simple ones. This is a scope flip, not the over-trust flip from Corley's story: the trigger was the new model taking noticeably longer to return a recommendation on multi-system jobs, a real latency cost technicians felt on the clock.
P, pinpoint the old decision. The tool was built to always return one ranked recommendation, simple or complex, with no way to signal "this one's going to take longer to reason through." Technicians had no way to anticipate the delay, so a slow multi-system job felt like the tool being unreliable, and they started skipping it for exactly those jobs, the ones where a second opinion would have helped most.
S, show the replay. A version that flags upfront, before running, whether a job looks complex enough to take longer, and gives technicians a "quick read" option that trades some accuracy for speed on those cases, keeps technicians using the tool on complex jobs instead of abandoning it there entirely. Alaric Marchbanks, the PM who owns this migration, tracked usage on multi-system jobs at 31 percent before the fix and 74 percent after.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the flip: a benchmark win is exactly the kind of good news that talks people out of the checking habit that would catch the model's real gap, so migrate with a mandatory, category-weighted pass spot-check built into the workflow, not left as a habit.
Cost: the QA team is short-staffed this quarter and can't add review hours. Don't drop the quota to fit the headcount, narrow its scope instead, keep the mandatory check on the highest-stakes category only and drop the general five-percent check temporarily, and say plainly that's a real tradeoff, not a free efficiency gain.
The model got better: the new scoring model turns out to beat the old one on every category, including the vulnerable-customer disclosure. That doesn't make the structural spot-check unnecessary, it's still the thing that would have caught the regression if it had happened, and this time it just confirms there wasn't one.
Where people run it wrong.
They treat a vendor benchmark as proof the migration is safe, when a benchmark built from someone else's data can't speak to this product's own rarest, highest-stakes category.
They plan the migration around the model's accuracy and never ask what happens to the human habits sitting around it.
They leave a safeguard as an informal best practice instead of building it into the tool's own workflow, so it survives right up until something gives everyone a shared reason to let it lapse.
How to use it live. Say the flip before naming a single migration step: "the real risk in this migration isn't the model, it's what a benchmark win does to the humans checking it. Good news is a perturbation too." That buys the room to ask what informal habit around this specific tool might quietly stop, instead of reciting a generic parallel-run checklist.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't you catch this with a bigger, better benchmark instead of relying on live spot-checks?" Response: a benchmark, however large, is still frozen at the moment it was built; a live spot-check catches whatever the model does on this center's actual calls today, including drift a static benchmark can't see.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Model migration and version changes for users
- #1 Your provider deprecates the model behind your main feature in 60 days. Write the plan.
- #2 How do you test a replacement model against the behaviour users have come to expect?
- #3 Explain why a strictly better model can still be a bad migration.
- #4 What should you tell users when model behaviour changes underneath them?
- #5 Describe a dual-running strategy for a model migration.
- #6 How do you handle customers who tuned their prompts to the old model?