Design the regression suite you would run before any model swap.
- Size the suite as categories times cases per category, never a flat sample count pulled from habit.Why: a sample size with no arithmetic behind it is a habit wearing a checklist.
- Count every distinct copy format as its own category, and never merge two formats just because they used to get graded together.Why: a merged category hides the one format where the actual regression lives.
- Run enough cases per category to hit a real confidence bar, about 25 to catch a regression hitting one case in ten with 90 percent confidence, not a round number that feels thorough.Why: too few cases per category means the suite can miss a real regression and still call itself clean.
- Score every case against the actual client's brand-voice guide and hard requirements, not a generic quality skim.Why: a sentence that reads fine in general can still be off-voice, or missing a required line, and a generic check waves it through.
- Check the total against how many days your reviewers really have before the old model gets switched off.Why: a suite too big to finish before the cutover doesn't protect anyone, it just delays when the same regression gets found.
- When it doesn't fit the runway, tighten the bar only on the highest-stakes categories and loosen it elsewhere, never cut every category by the same amount.Why: treating every category as equally risky is how a suite grows too big to finish before the deadline forces your hand anyway.
How to answer this, stage by stage
Nobody is grading whether you land on exactly 210 cases. They're grading whether the number came from real arithmetic, whether the range is honest, and whether you close on something the room can check. Eight moves get you there.
Let's learn
Thornwood Creative is a mid-size marketing agency. Its in-house AI copy tool, Wordloom, drafts first-pass copy for forty-six client accounts, close to nine hundred pieces a week: headlines, captions, product descriptions, email subject lines, and more.
Wordloom launched a year and a half ago doing exactly one thing: ad headlines. Before every model change back then, someone pulled fifty recent headlines and one reviewer skimmed them for anything odd. It worked, because there was only one thing to check.
Wordloom now writes in twelve different client-facing shapes. The regression check never grew with it. Fifty outputs, pulled at random, one reviewer, same as year one.
Split fifty outputs evenly across twelve categories and each one gets about four cases. Four cases cannot tell you much of anything about whether one specific format, say, LinkedIn captions for a client with strict disclosure rules, still works the way it used to.
At its worst, that gap ships quietly. A model swap goes through, the flat fifty-sample skim looks fine because it never really tested the category where the actual regression lives, and three weeks later a client's account manager gets a call asking why their supplement brand's captions suddenly read like a startup's landing page, casual where it used to be careful, or why a required disclosure line stopped showing up. Nobody flagged it, because nobody's check was built to.
The choice I would take back. Thornwood kept the flat fifty-sample check because it had always worked, back when Wordloom only wrote one thing. Nobody re-sized it as the tool grew, because nothing about it ever visibly broke. That felt like evidence it was still enough. It wasn't evidence of anything, since a check spread too thin across too many categories can pass clean for years and never actually be testing most of them.
What I would leave alone. Wordloom also spits out internal content-idea brainstorm lists for copywriters' own use, never sent to a client. Nobody's brand voice is riding on those. A quick five-case skim is plenty; the full 25-case, scored-against-a-guide treatment would be effort spent protecting something that costs a few wasted minutes to fix, not a client relationship.
The lesson. A regression suite sized for a one-format tool doesn't resize itself when the tool grows to twelve. The category count was never really about how many outputs to pull. It was about how many separate promises Thornwood had made to its clients, and only the arithmetic showed how many of them a habit-sized check was quietly skipping.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why a fifty-sample habit almost went unquestioned one more time.
Soojin Baek has managed Wordloom at Thornwood Creative for two years. She's the one who signs off on any change to the model underneath it, and she's good at her job in a specific way: she never lets a migration go out on vibes. Every swap gets a written check, filed, dated.
For most of those two years, the written check was the same fifty-sample skim the tool had used since launch. It always came back clean. Thornwood swapped model versions three times in Soojin's tenure, and each time, fifty outputs looked fine, and the migration shipped on schedule.
Then, on a Tuesday in September, the vendor's deprecation notice landed. Nothing dramatic. One email: the model version Wordloom had been running on gets switched off in fifteen working days, on a fixed date, no extensions.
Soojin opened the usual template. Pull fifty recent outputs. Book time with a reviewer. Same as always.
She stopped halfway through filling it in. Wordloom didn't write one thing anymore. It wrote twelve. Fifty outputs, spread evenly, was around four cases per format. Four cases for search headlines. Four for product descriptions. Four for the LinkedIn captions that Thornwood's supplement and wellness clients relied on to carry a legally required disclosure line.
She worked out the real numbers instead. Twelve categories, about 25 cases each to catch a regression hitting one case in ten with 90 percent confidence: 300 cases. Fifteen working days sounded like room, until she checked with engineering, who needed nine of them just to point Wordloom's pipeline at the new model safely. Six real days left. Two reviewers, forty cases a day between them. Six days bought 240 case-reviews. Three hundred needed seven and a half.
It didn't fit. She could have shrunk every category evenly to make it fit, twenty cases instead of twenty-five, everywhere. She rejected that. A flat cut protects the categories that were never at risk exactly as much as it starves the ones that are.
Instead she tiered it. The six categories that touch the most client accounts and the most revenue, full 25-case treatment, 90 percent confidence. The other six, a lighter 10-case check, enough to catch a regression only if it's hitting one case in five, not one in ten. Two hundred and ten cases total. Five and a quarter days. It fit, with room to spare.
The suite ran. In the LinkedIn captions category, one of the six full-bar categories, for a wellness client called Fennroot, four of the twenty-five test cases came back missing the required paid-partnership disclosure tag. Sixteen percent, well past the one-in-ten bar the suite was built to catch.
Soojin flagged it to the vendor's support team before cutover. It turned out to be a real regression in how the new model handled a specific instruction buried in the system prompt, one that hadn't shown up in the vendor's own benchmark because their benchmark never tested a client-specific compliance requirement in the first place. It was fixed and re-tested before a single account switched over. No client ever saw it.
The thing I'd tell myself, back on that Tuesday with the old template half-filled-in: a check that's always come back clean isn't proof it's still testing the right thing. Sometimes it's proof it stopped testing most of it years ago, and nobody noticed because nothing broke loudly enough to ask.
BOUND, run on the categories a copy tool actually writes
This is a sizing question: how many test cases catch a real regression at a confidence level worth trusting. Not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Total cases equal the number of distinct copy categories the tool actually writes, times how many real cases per category it takes to catch a chosen defect rate at a chosen confidence level. Not a flat sample size borrowed from when the tool did one thing.
O, own the numbers. Wordloom writes in twelve client-facing categories today. To catch a regression hitting one case in ten with 90 percent confidence, the math needs about 22 cases a category, minimum; round up to 25 for real-world noise. Twelve times 25 is 300. I considered leaning on the vendor's own published benchmark instead, the one where they claim the new model wins most head-to-head comparisons. I rejected it: their benchmark scores generic writing quality, not whether a specific client's disclosure line survives, so a model can win their leaderboard and still fail Thornwood's actual categories.
U, use a range. Merge related formats back into six broad buckets and it's 150 cases. Count every format on its own, the honest read, and it's 300. The range depends entirely on how many genuinely distinct things the tool is asked to do.
N, nail the sanity check. Fifteen working days until cutover, minus nine for engineering to wire up the new model safely, leaves six real days. Two reviewers clear about 240 case-reviews in that window. The 150-case low estimate fits easily. The 300-case honest count needs seven and a half days, past the runway. The fix: tier the bar, full 25-case rigor on the six highest-stakes categories, a lighter 10-case check elsewhere, for 210 cases, inside the six-day window. Silent brand-voice drift, a technically fine sentence that quietly stops sounding like the client, is the failure mode this whole suite exists to catch, and scoring against each client's actual brand-voice guide, not a generic quality skim, is the guardrail.
D, direction. Two assumptions could move this number, and they don't move it equally. Redefining categories, six broad versus twelve real ones, swings the total by 150 cases. Raising the confidence bar from 90 to 95 percent only adds 60. Category count is the bigger lever, and the one Soojin actually controls, since it's a product call about what counts as a distinct promise, not a statistics argument.
And if you want to be sure it really works, try it somewhere else
Caldbrook County's planning department uses an AI tool called Draftwell to draft the letters that go out on every permit application: approvals, requests for more information, rejections, appeal responses, inspection notices. The state is switching the shared AI platform every county runs on to a new vendor.
B, break it down. Same shape, a different set of promises. Total cases equal how many distinct letter categories Draftwell writes, times how many real cases per category catch a chosen defect rate at a chosen confidence.
O, own the numbers. Broadly, Draftwell writes five kinds of letters. Split rejections and appeal responses by residential versus commercial, since a wrong citation in a commercial appeal carries real legal exposure, and it's nine. Catching a regression hitting one case in about seven, with 85 percent confidence, needs about 12 cases a category, minimum; round to 15.
U, use a range. Five broad categories times 15 is 75 cases. Nine full categories times 15 is 135.
N, nail the sanity check. The state mandates cutover in six working days. County IT needs two of them to point Draftwell at the new environment. Four real days left. One permit supervisor reviews about 25 cases a day, so four days is 100 case-reviews. The 75-case low estimate fits with room to spare. The 135-case honest count needs five and a half days, past the four-day runway.
D, direction. Here the lever flips. Merging back to five categories only saves 60 cases. But raising the confidence bar to 99 percent across all nine categories, because a wrong statement in a legal appeal letter carries real exposure, adds 135, doubling the suite. For Caldbrook, the confidence bar moves the total more than the category count does, because the risk that matters here isn't breadth, it's a single wrong sentence holding up in an appeal.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: count every distinct copy format as its own category, run about 25 cases a category to hit a real confidence bar, and check the total against how many days you actually have before the old model gets switched off.
Cost: the team can't add a second reviewer. Don't shrink the whole suite evenly, tier it: full bar on the highest-stakes categories, a lighter bar on the rest, so the suite still finishes on time.
The model got better: the new model wins the vendor's own benchmark comfortably. Still run the full suite. A model that's better on average can still regress one narrow, uncelebrated instruction, like a disclosure tag, that never showed up on anyone's leaderboard.
Where people run it wrong.
They keep the old flat sample size because it's "always worked," without asking whether the number of things being tested has grown since it was set.
They treat the vendor's own before/after benchmark as sufficient proof, when it never tested the categories or the client requirements that actually matter to them.
Under time pressure, they cut every category by the same amount, instead of tiering the bar by which categories are actually high-stakes.
How to use it live. Say the equation before naming a single number: "the sample size isn't one flat count, it's how many separate things the model has to keep doing right, times how many real cases it takes to trust each one, so before I give you a number I'd want to know how many of those there really are." That buys the room to ask a real question instead of reciting a number that only sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Model migration and version changes for users
- #1 Your provider deprecates the model behind your main feature in 60 days. Write the plan.
- #2 How do you test a replacement model against the behaviour users have come to expect?
- #3 Explain why a strictly better model can still be a bad migration.
- #4 What should you tell users when model behaviour changes underneath them?
- #5 Describe a dual-running strategy for a model migration.
- #6 How do you handle customers who tuned their prompts to the old model?