Artifact critiqueAdvancedShipping & Model Lifecycle / Model migration and version changes for users / #8

Design the regression suite you would run before any model swap.

The direct answer
Size the suite as categories times cases per category, not a flat sample count pulled from habit. Count every distinct copy format the tool writes as its own category, and run about 25 real cases in each one, enough to catch a regression that hits one case in ten with 90 percent confidence, not a promise every case comes out perfect. Check that total against how many days your reviewers actually have before the old model gets switched off, and if it does not fit, tier which categories get the full bar instead of shrinking the whole suite evenly.
Do this, in order
  1. Size the suite as categories times cases per category, never a flat sample count pulled from habit.Why: a sample size with no arithmetic behind it is a habit wearing a checklist.
  2. Count every distinct copy format as its own category, and never merge two formats just because they used to get graded together.Why: a merged category hides the one format where the actual regression lives.
  3. Run enough cases per category to hit a real confidence bar, about 25 to catch a regression hitting one case in ten with 90 percent confidence, not a round number that feels thorough.Why: too few cases per category means the suite can miss a real regression and still call itself clean.
  4. Score every case against the actual client's brand-voice guide and hard requirements, not a generic quality skim.Why: a sentence that reads fine in general can still be off-voice, or missing a required line, and a generic check waves it through.
  5. Check the total against how many days your reviewers really have before the old model gets switched off.Why: a suite too big to finish before the cutover doesn't protect anyone, it just delays when the same regression gets found.
  6. When it doesn't fit the runway, tighten the bar only on the highest-stakes categories and loosen it elsewhere, never cut every category by the same amount.Why: treating every category as equally risky is how a suite grows too big to finish before the deadline forces your hand anyway.

How to answer this, stage by stage

Nobody is grading whether you land on exactly 210 cases. They're grading whether the number came from real arithmetic, whether the range is honest, and whether you close on something the room can check. Eight moves get you there.

1
Scope it to one real product and one real deadline
Say it like this
"Let's ground this. Say Thornwood Creative, a marketing agency, built an in-house AI copy tool called Wordloom. It drafts everything from ad headlines to product descriptions for forty-six client accounts, close to nine hundred pieces of copy a week. The vendor behind Wordloom's model just sent the standard notice: the version they've been running for a year and a half stops being served in fifteen working days. Every account switches over automatically on that date."
Why this works
Stops the answer from staying abstract before a single number gets attached to it.
2
Say the real question out loud before naming any numbers
Say it like this
"The question isn't 'how many outputs should we check.' It's 'how many different things does this tool have to keep doing right, and how many real cases of each one does it take before I'd trust that it still does them.' A tool that writes one thing can get away with a small flat sample. A tool that writes twelve different things can't, no matter how thorough fifty outputs sounds."
Why this works
Shows the equation before the arithmetic, so the numbers that follow read as a plan, not a habit.
3
Break down the equation
Say it like this
"Here's the shape. Total cases equal the number of copy categories the tool actually handles, times how many real cases per category it takes to catch a regression at a confidence level I'm willing to stand behind. I'm not picking a flat number. I'm picking a defect rate I want to catch and a confidence I want in catching it, and letting those two numbers set the sample size."
Why this works
This is the B step of BOUND, the equation stated before a single number gets attached.
4
Own the real numbers behind each term
Say it like this
"Wordloom writes in twelve client-facing shapes now: search and social ad headlines, captions for three platforms, product descriptions, email subject lines, landing-page hero copy, blog intros, button copy, video scripts, and a brand-voice spot-check. To catch a regression hitting one case in ten with 90 percent confidence, the math says about 22 cases a category, minimum. I round that up to 25, because real client copy is messier than the clean math assumes. Twelve categories times 25 is 300 cases."
Why this works
This is the O step, a real proposed number with a stated reason, not "we'll check a bunch and see."
5
Give the range, not one number
Say it like this
"If I merge related formats back into six broad buckets, the way we used to before Wordloom grew, that's 150 cases. If I count every format on its own, the honest count, that's 300. I'd argue for the honest count, because merging is exactly the habit that let a regression hide in the first place."
Why this works
A single number here would claim a precision about "how many things this tool does" that a shortcut doesn't earn.
6
Sanity check against the real runway, and name the trade-off you're accepting
Say it like this
"Fifteen working days sounds like room to breathe, but engineering needs nine of them just to wire the new model into a safe test environment. That leaves six real days. Two reviewers can get through about 40 cases a day between them, so six days is 240 case-reviews. Three hundred cases needs seven and a half days. It doesn't fit. So I'd tier it: the full 25-case, 90-percent bar on the six categories that touch the most client accounts and the most revenue, a lighter 10-case check on the other six. That's 210 cases, five and a quarter days, inside the window. The trade-off is real: on the lower-stakes categories we're only guaranteed to catch a regression if it's hitting one case in five, not one in ten. I'd rather know that going in than pretend every category got the same certainty."
Why this works
This is the N step, and it names the quality-for-time trade-off out loud instead of hoping nobody asks.
7
Say which assumption moves the total most
Say it like this
"Two things could change this number, and they don't move it the same amount. How I define a category, six broad buckets versus twelve real ones, swings the total by 150 cases. Raising the confidence bar from 90 to 95 percent only adds 60. So the category question is the one worth getting right first, and it's the one I actually control, it's a product decision about what counts as a distinct promise, not a statistics debate."
Why this works
This is the D step, the direction a good estimator names and a bad one skips.
8
Close on the one line
Say it like this
"So here's what I'd actually do: size the suite as categories times cases per category, run about 25 cases a category to hit a real confidence bar, and check that total against how many days my reviewers have before the old model gets switched off. Where it doesn't fit, I'd tier the bar by stakes, not cut the whole suite evenly and hope."
Why this works
Closes on the literal ask, a claim someone could check, not a vibe about being careful.
If you remember one thing A regression suite's real size is categories times cases per category, and the category count almost always moves the total more than the confidence bar does. Counting twelve things as six doesn't make the tool simpler. It just makes six of its promises untested.

Let's learn

Thornwood Creative is a mid-size marketing agency. Its in-house AI copy tool, Wordloom, drafts first-pass copy for forty-six client accounts, close to nine hundred pieces a week: headlines, captions, product descriptions, email subject lines, and more.

Wordloom launched a year and a half ago doing exactly one thing: ad headlines. Before every model change back then, someone pulled fifty recent headlines and one reviewer skimmed them for anything odd. It worked, because there was only one thing to check.

Wordloom now writes in twelve different client-facing shapes. The regression check never grew with it. Fifty outputs, pulled at random, one reviewer, same as year one.

Knowledge spark: what does "90 percent confidence" mean here? It means: if a regression is really hitting one case in ten, running enough test cases gives you a 90 percent chance of seeing at least one bad one. It is not a promise every case will be perfect. It's a promise about how likely you are to catch a real problem if it's actually there.
Hand-sketched comparison diagram. Left panel, a single sheet of paper labeled 50 outputs, one skim, captioned pulled at random, no categories, one reviewer eyeballs them. Right panel, a scale icon labeled 12 categories, 25 each, captioned every copy format counted on its own, scored against the client's voice guide.
Fifty outputs spread across twelve real categories is about four cases each. That is not enough to trust any single one of them.

Split fifty outputs evenly across twelve categories and each one gets about four cases. Four cases cannot tell you much of anything about whether one specific format, say, LinkedIn captions for a client with strict disclosure rules, still works the way it used to.

We were not checking whether the new model could write good copy. We were checking whether it could still sound like forty-six different clients.

At its worst, that gap ships quietly. A model swap goes through, the flat fifty-sample skim looks fine because it never really tested the category where the actual regression lives, and three weeks later a client's account manager gets a call asking why their supplement brand's captions suddenly read like a startup's landing page, casual where it used to be careful, or why a required disclosure line stopped showing up. Nobody flagged it, because nobody's check was built to.

The decision that mattered Size the suite by how many distinct copy formats Wordloom actually writes, not by how many outputs feel like enough to check in an afternoon. That's the one call that decides whether a regression gets caught in testing, or in a client's inbox.

The choice I would take back. Thornwood kept the flat fifty-sample check because it had always worked, back when Wordloom only wrote one thing. Nobody re-sized it as the tool grew, because nothing about it ever visibly broke. That felt like evidence it was still enough. It wasn't evidence of anything, since a check spread too thin across too many categories can pass clean for years and never actually be testing most of them.

What I would leave alone. Wordloom also spits out internal content-idea brainstorm lists for copywriters' own use, never sent to a client. Nobody's brand voice is riding on those. A quick five-case skim is plenty; the full 25-case, scored-against-a-guide treatment would be effort spent protecting something that costs a few wasted minutes to fix, not a client relationship.

The lesson. A regression suite sized for a one-format tool doesn't resize itself when the tool grows to twelve. The category count was never really about how many outputs to pull. It was about how many separate promises Thornwood had made to its clients, and only the arithmetic showed how many of them a habit-sized check was quietly skipping.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why a fifty-sample habit almost went unquestioned one more time.

Soojin Baek has managed Wordloom at Thornwood Creative for two years. She's the one who signs off on any change to the model underneath it, and she's good at her job in a specific way: she never lets a migration go out on vibes. Every swap gets a written check, filed, dated.

For most of those two years, the written check was the same fifty-sample skim the tool had used since launch. It always came back clean. Thornwood swapped model versions three times in Soojin's tenure, and each time, fifty outputs looked fine, and the migration shipped on schedule.

Then, on a Tuesday in September, the vendor's deprecation notice landed. Nothing dramatic. One email: the model version Wordloom had been running on gets switched off in fifteen working days, on a fixed date, no extensions.

Soojin opened the usual template. Pull fifty recent outputs. Book time with a reviewer. Same as always.

She stopped halfway through filling it in. Wordloom didn't write one thing anymore. It wrote twelve. Fifty outputs, spread evenly, was around four cases per format. Four cases for search headlines. Four for product descriptions. Four for the LinkedIn captions that Thornwood's supplement and wellness clients relied on to carry a legally required disclosure line.

Hand-sketched number line diagram. Left point labeled Low estimate, 150 cases, 6 core categories. Middle point, highlighted in blue, labeled Runway, roughly 240 reviews possible in 6 working days. Right point labeled High estimate, 300 cases, 12 full categories.
The honest count, 300 cases, sits past what the real six-day runway can absorb. The runway sits comfortably past the merged, six-category count.

She worked out the real numbers instead. Twelve categories, about 25 cases each to catch a regression hitting one case in ten with 90 percent confidence: 300 cases. Fifteen working days sounded like room, until she checked with engineering, who needed nine of them just to point Wordloom's pipeline at the new model safely. Six real days left. Two reviewers, forty cases a day between them. Six days bought 240 case-reviews. Three hundred needed seven and a half.

It didn't fit. She could have shrunk every category evenly to make it fit, twenty cases instead of twenty-five, everywhere. She rejected that. A flat cut protects the categories that were never at risk exactly as much as it starves the ones that are.

Instead she tiered it. The six categories that touch the most client accounts and the most revenue, full 25-case treatment, 90 percent confidence. The other six, a lighter 10-case check, enough to catch a regression only if it's hitting one case in five, not one in ten. Two hundred and ten cases total. Five and a quarter days. It fit, with room to spare.

The suite ran. In the LinkedIn captions category, one of the six full-bar categories, for a wellness client called Fennroot, four of the twenty-five test cases came back missing the required paid-partnership disclosure tag. Sixteen percent, well past the one-in-ten bar the suite was built to catch.

Four captions out of twenty-five were missing Fennroot's disclosure tag. A fifty-sample check, spread across twelve categories, would have caught less than one, on average.

Soojin flagged it to the vendor's support team before cutover. It turned out to be a real regression in how the new model handled a specific instruction buried in the system prompt, one that hadn't shown up in the vendor's own benchmark because their benchmark never tested a client-specific compliance requirement in the first place. It was fixed and re-tested before a single account switched over. No client ever saw it.

The thing I'd tell myself, back on that Tuesday with the old template half-filled-in: a check that's always come back clean isn't proof it's still testing the right thing. Sometimes it's proof it stopped testing most of it years ago, and nobody noticed because nothing broke loudly enough to ask.

BOUND, run on the categories a copy tool actually writes

This is a sizing question: how many test cases catch a real regression at a confidence level worth trusting. Not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. Total cases equal the number of distinct copy categories the tool actually writes, times how many real cases per category it takes to catch a chosen defect rate at a chosen confidence level. Not a flat sample size borrowed from when the tool did one thing.
O, own the numbers. Wordloom writes in twelve client-facing categories today. To catch a regression hitting one case in ten with 90 percent confidence, the math needs about 22 cases a category, minimum; round up to 25 for real-world noise. Twelve times 25 is 300. I considered leaning on the vendor's own published benchmark instead, the one where they claim the new model wins most head-to-head comparisons. I rejected it: their benchmark scores generic writing quality, not whether a specific client's disclosure line survives, so a model can win their leaderboard and still fail Thornwood's actual categories.
U, use a range. Merge related formats back into six broad buckets and it's 150 cases. Count every format on its own, the honest read, and it's 300. The range depends entirely on how many genuinely distinct things the tool is asked to do.
N, nail the sanity check. Fifteen working days until cutover, minus nine for engineering to wire up the new model safely, leaves six real days. Two reviewers clear about 240 case-reviews in that window. The 150-case low estimate fits easily. The 300-case honest count needs seven and a half days, past the runway. The fix: tier the bar, full 25-case rigor on the six highest-stakes categories, a lighter 10-case check elsewhere, for 210 cases, inside the six-day window. Silent brand-voice drift, a technically fine sentence that quietly stops sounding like the client, is the failure mode this whole suite exists to catch, and scoring against each client's actual brand-voice guide, not a generic quality skim, is the guardrail.
D, direction. Two assumptions could move this number, and they don't move it equally. Redefining categories, six broad versus twelve real ones, swings the total by 150 cases. Raising the confidence bar from 90 to 95 percent only adds 60. Category count is the bigger lever, and the one Soojin actually controls, since it's a product call about what counts as a distinct promise, not a statistics argument.

The build-up: cases stacking toward the tiered total, against the real runway
Tier A, 6 categories at full bar (25 each)150 cases
+ Tier B, 6 categories at light bar (10 each)210 cases
Runway ceiling, 6 days × 2 reviewers240 case-reviews
The tiered suite lands at 210 cases, thirty case-reviews under the ceiling. The full, untiered 300-case count would have needed 60 more than the runway allows.
What moves the suite size most (swing in total cases from the 300-case baseline)
Merge 12 categories back to 6 broad buckets−150
Raise confidence bar from 90% to 95%+60
Loosen the defect rate to catch (1-in-10 to 1-in-6)−120
Add a third reviewer to the regression team0
Redefining categories swings the total more than any statistics knob does. A third reviewer helps the review finish faster; it does nothing to how many cases actually need running.

And if you want to be sure it really works, try it somewhere else

Caldbrook County's planning department uses an AI tool called Draftwell to draft the letters that go out on every permit application: approvals, requests for more information, rejections, appeal responses, inspection notices. The state is switching the shared AI platform every county runs on to a new vendor.

B, break it down. Same shape, a different set of promises. Total cases equal how many distinct letter categories Draftwell writes, times how many real cases per category catch a chosen defect rate at a chosen confidence.
O, own the numbers. Broadly, Draftwell writes five kinds of letters. Split rejections and appeal responses by residential versus commercial, since a wrong citation in a commercial appeal carries real legal exposure, and it's nine. Catching a regression hitting one case in about seven, with 85 percent confidence, needs about 12 cases a category, minimum; round to 15.
U, use a range. Five broad categories times 15 is 75 cases. Nine full categories times 15 is 135.
N, nail the sanity check. The state mandates cutover in six working days. County IT needs two of them to point Draftwell at the new environment. Four real days left. One permit supervisor reviews about 25 cases a day, so four days is 100 case-reviews. The 75-case low estimate fits with room to spare. The 135-case honest count needs five and a half days, past the four-day runway.
D, direction. Here the lever flips. Merging back to five categories only saves 60 cases. But raising the confidence bar to 99 percent across all nine categories, because a wrong statement in a legal appeal letter carries real exposure, adds 135, doubling the suite. For Caldbrook, the confidence bar moves the total more than the category count does, because the risk that matters here isn't breadth, it's a single wrong sentence holding up in an appeal.

Days to a trustworthy suite: standard confidence vs legal-exposure confidence
9 categories, standard 85% confidence bar5.4 days, 135 cases
9 categories, 99% confidence for legal exposure10.8 days, 270 cases
Caldbrook's four-day runway can't absorb either scenario in full. Unlike Thornwood, where category count was the dominant lever, here the confidence bar drives the swing, because the thing at risk is one wrong sentence surviving a legal appeal, not a range of client voices.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: count every distinct copy format as its own category, run about 25 cases a category to hit a real confidence bar, and check the total against how many days you actually have before the old model gets switched off.
Cost: the team can't add a second reviewer. Don't shrink the whole suite evenly, tier it: full bar on the highest-stakes categories, a lighter bar on the rest, so the suite still finishes on time.
The model got better: the new model wins the vendor's own benchmark comfortably. Still run the full suite. A model that's better on average can still regress one narrow, uncelebrated instruction, like a disclosure tag, that never showed up on anyone's leaderboard.

Where people run it wrong.
They keep the old flat sample size because it's "always worked," without asking whether the number of things being tested has grown since it was set.
They treat the vendor's own before/after benchmark as sufficient proof, when it never tested the categories or the client requirements that actually matter to them.
Under time pressure, they cut every category by the same amount, instead of tiering the bar by which categories are actually high-stakes.

How to use it live. Say the equation before naming a single number: "the sample size isn't one flat count, it's how many separate things the model has to keep doing right, times how many real cases it takes to trust each one, so before I give you a number I'd want to know how many of those there really are." That buys the room to ask a real question instead of reciting a number that only sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about sizing a regression suite before a model swap, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question, how many real cases per copy category catch a regression at a stated confidence level, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Soojin Baek, product manager for Wordloom at Thornwood Creative, a marketing agency. She's owned every model migration for the tool for two years and files a written check for each one.
3 · WHAT THE OLD CHECK ACTUALLY TESTED
What did Thornwood's fifty-sample check actually cover, and why did it stop being enough?
Tap to flip
ANSWER
It was sized for Wordloom's launch version, which only wrote ad headlines. Spread across twelve categories, fifty outputs average about four cases each, nowhere near enough to trust any one format.
4 · THE EQUATION
What's the equation behind the regression suite's size?
Tap to flip
ANSWER
Total cases equal the number of distinct copy categories the tool writes, times how many real cases per category it takes to catch a chosen defect rate at a chosen confidence level.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Keeping the flat fifty-sample check after Wordloom grew from one copy format to twelve. It made sense at launch, when one format was the whole product, and nobody re-sized it as the tool grew because it never visibly broke.
6 · THE NUMBER
Fill in the blank: the honest, full-granularity count was ___ cases across ___ categories, but the real runway only allowed ___ case-reviews, so the tiered suite that actually shipped came to ___ cases.
Tap to flip
ANSWER
300 cases. 12 categories. 240 case-reviews. 210 cases.
7 · THE REPLAY
Same migration, new suite, what changed?
Tap to flip
ANSWER
The tiered 210-case suite ran before cutover. In the LinkedIn captions category, 4 of 25 cases for client Fennroot came back missing a required disclosure tag, well past the 1-in-10 bar. It got fixed and re-tested before any client account switched over.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same sizing question for a different product. Which product, and which assumption swings its total most?
Tap to flip
ANSWER
Caldbrook County's permit-letter tool, Draftwell. There, the confidence bar (raised for legal-exposure categories) swings the total more than the category count does, the opposite of Wordloom's case.

Check yourself Score: 0 / 0

Short answer
1. Why couldn't Soojin just keep testing Wordloom with the same fifty random outputs Thornwood had always used?
Show hint
Compare how many copy formats Wordloom wrote at launch against how many it writes now.
Show answer
Model answer: The fifty-sample check was sized for Wordloom's launch version, which only wrote ad headlines. Now spread across twelve categories, fifty outputs give each one only about four cases, not enough to trust that any single format still works the way it used to.
Multiple choice
2. Why did the suite land on about 25 cases per category instead of, say, 10 or 50?
  • A. 25 is a round number that felt thorough without taking too long to review.
  • B. It's the number of cases needed to catch a regression hitting one case in ten with 90 percent confidence, rounded up for real-world noise.
  • C. It matches the number of client accounts Thornwood serves, divided by two.
  • D. It's the maximum number of cases one reviewer can check in a single day.
Show hint
Check the O step: it names a defect rate and a confidence level, then derives the sample size from those.
Show answer
B. The number comes from a stated defect rate (one case in ten) and a stated confidence level (90 percent), not from a feeling of thoroughness or a scheduling constraint.
True or false
3. True or false: once Soojin worked out that the honest suite needed 300 cases, she ran all 300 before the cutover.
  • True
  • False
Show hint
Check the N step: what did the six-day runway actually allow?
Show answer
False. The 300-case suite needed seven and a half days, past the real six-day runway. She tiered it instead: full 25-case rigor on the six highest-stakes categories, a lighter 10-case check on the rest, for 210 cases total.
Fill in the blank
4. The vendor gave Thornwood ___ working days until cutover. Engineering needed ___ of those just to wire up the new model, leaving a real runway of ___ working days, or ___ case-reviews at two reviewers' pace.
Show hint
Check the N step and the build-up chart's final row.
Show answer
15 working days; 9 working days; 6 working days; 240 case-reviews. That 240-case ceiling is exactly why the 300-case honest count had to be tiered down to 210.
Short answer, apply it yourself
5. Think of an AI tool you use that does more than one kind of task. If you had to test it before a model swap, what would count as one "category" for it, and would a single flat sample size actually cover all of them?
Show hint
Pick a tool where the tasks feel genuinely different from each other, not just cosmetic variations on one task.
Show answer
Model answer: A coding assistant that both writes new functions and explains existing code. Those are different categories: one is generation, one is comprehension, and a regression could hit one and not the other. A single flat sample split across both would under-test whichever one has more edge cases, the same problem Wordloom had with client-specific formats.
Short answer, the number question
6. If Thornwood's real runway had been four working days instead of six, at the same 40 cases a day, would the 210-case tiered suite still fit? Show the math.
Show hint
Work out how many case-reviews four days actually buys, and compare it to 210.
Show answer
No, it would not fit. Four days at 40 cases a day is 160 case-reviews, short of the 210-case tiered suite by 50. Soojin would need to tier further, likely trimming Tier B's lighter categories down from 10 cases each to 5 or 6, accepting an even looser bar on the lowest-stakes formats to fit the shorter window.
Follow-up footer
Why this works
This question tests whether a candidate can turn "run a regression suite" into real arithmetic, categories times cases per category, rather than a vague promise to "test thoroughly." Most candidates name a number. Few can defend where it came from, or say what they'd cut first when it doesn't fit the calendar.
Follow-up traps
"Why not just run the full 300-case suite and push the deadline?"Response: the deadline isn't ours to move, the vendor sets the cutover date and the old model stops being served on it regardless of whether testing is done.
"Isn't tiering the bar just a way of accepting known risk?"Response: yes, openly, on the categories that touch the least revenue and the fewest accounts, which is a better place to accept it than spreading the same risk evenly across categories that matter far more.
If pressed
The sample-size math itself: n equals ln(1 minus confidence) divided by ln(1 minus the defect rate you're trying to catch). For a 10 percent defect rate at 90 percent confidence, that's ln(0.1) over ln(0.9), about 21.9, which is why 25 cases a category, not 20 or 30, was the real number behind the suite.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more