Artifact critiqueAdvancedShipping & Model Lifecycle / Rollout strategy and phased launches / #13
Critique a plan that rolls out to 100 percent on day one because the evals passed.
The direct answer
Don't treat a passed eval as proof the plan is safe for every document type. Launch the part the eval set actually covers right away, and hold back the part it barely covers, routing every automatic freeze on that part to a person to check, for a fixed window, before it reaches a member. Only turn on full automatic decisions for that part once it proves itself against real, unlabeled production documents, not the eval sample.
Do this, in order
Don't flip DocSentry to full automatic decisions on every document type because the eval passed.Why: a 97 percent score is a fact about the documents in the eval set, not about a document type it has never been shown.
Turn DocSentry on for scanned documents right away, and hold phone-photographed documents back.Why: the eval set is almost entirely scanned documents, so that channel has earned its launch and the photo channel hasn't.
Route every auto-freeze on a phone-photographed document to a fraud analyst first, for eight weeks or five hundred decisions, whichever comes first.Why: a wrongful freeze can lock a member out of money they need that week, so that's the direction that actually needs a person watching.
Before day one, run DocSentry silently against a fresh batch of real, unlabeled documents already in the queue, and compare its flag rate to the eval set's.Why: this is the only way to see the eval-to-production gap before a single member is touched by it.
Track the phone-photo channel's override rate on its own dashboard, watched weekly.Why: a spike in one channel barely moves the company-wide number and stays invisible if it's measured blended in.
Set a real number that ends the human check, not "once it feels safe."Why: without one, the check either never comes off or comes off before it's earned.
How to answer this, stage by stage
Seven moves, from naming one real plan to the line you'd close on.
1
Ground it in one real plan and one real number before naming the framework
Say it like this
"Say a credit union built a tool called DocSentry. It reads a loan document, a pay stub, an ID scan, a bank statement, and flags the ones that look altered before a human opens the file. The team tested it against thirty two hundred real documents from the last two years, and it got ninety seven out of a hundred right. The plan on the table is: since the eval passed, turn it on for every single loan document starting Monday. That's the plan I want to pull apart."
Why this works
Grounds "the evals passed" in a real score and a real plan before any framework talk starts.
2
State the plan in one breath
Say it like this
"I'd run this through GUARD, because the real problem with this plan isn't whether the score is real, it is. It's what a passed eval can't tell you: who meets the gap between the eval and the real documents first, whether they'd ever know it happened, the actual rollout I'd write instead, and how I'd catch the gap before day one instead of after."
Why this works
Two seconds naming the plan before diving in, instead of reciting an acronym.
3
Reframe it as a coverage question, not an accuracy question
Say it like this
"This isn't really a question about whether DocSentry is accurate. Ninety seven percent is a real number. The question is what that ninety seven percent was measured against, and whether that's the same shape of document showing up Monday morning."
Why this works
Separates a checklist answer from one that understands what an eval actually proves.
4
Give the one decision
Say it like this
"I'd split the launch by document type instead of by percentage of members. Scanned documents go live Monday, that's what the eval set actually is. Phone-photographed documents wait. Every auto-freeze DocSentry puts on a phone-photographed document goes to a person first, for eight weeks or five hundred decisions, whichever comes first. Full automatic decisions on that channel only turn on once it's earned it."
Why this works
Names an actual mechanism, tied to what's actually untested, not a value everyone already agrees with.
5
Prove it with the compressed failure
Say it like this
"Here's what the eval didn't catch. Of the thirty two hundred documents in the eval set, about one percent were phone photos. Right now, a third of Cedar Fork's real loan documents come in as phone photos, mostly gig and freelance income that doesn't look like a normal pay stub. We ran DocSentry quietly against five hundred of those real photos already sitting in the queue, without acting on it. It flagged twenty nine percent of them as altered, against a three percent flag rate on documents shaped like the eval set. Almost all of it was glare and cropped corners, not fraud. Blended across everyone applying that week, that pushes the flag rate from three percent to almost twelve percent, company-wide, and it hits every phone-photo applicant the same week DocSentry goes live for everyone, not one person catching it first in a small watched batch."
Why this works
Real numbers, a concrete cost, proves the eval-to-production gap instead of asserting it.
6
Say what you'd detect, and when
Say it like this
"Before day one, before any document type goes fully live, I'd run DocSentry silently against last month's real incoming documents, the ones nobody's labeled yet, and check whether its flag rate on that live batch matches its flag rate on the eval set. If it doesn't, that's the eval-to-production gap showing up before a single member has been touched. Once the phone-photo channel is live in its held-back mode, I'd track its override rate on its own dashboard, not folded into one number, every week."
Why this works
Turns detection into something that happens before launch, not a promise to watch closely afterward.
7
Land the answer in one breath
Say it like this
"So: a passed eval tells you the model is good at the documents you tested it on. It doesn't tell you anything about the documents you didn't. You don't earn the right to skip the small check by scoring well on the big test. You earn it by proving the model works on what you haven't checked yet."
Why this works
Restates the critique in one breath, the line an interviewer remembers on the way out.
Let's learn
What does a test passing actually promise you?
Before any of this, two people on Cedar Fork Credit Union's fraud team checked every loan document by eye. About twenty minutes each, for around eight hundred loan applications a month. That's most of a working month, every month, just holding pay stubs and bank statements up to the light.
Knowledge spark: what's an eval set
A batch of documents someone has already labeled, altered or clean, used to check a model before it goes live. It can only include what someone thought to put in it. A pattern nobody added never gets tested at all.
Then Cedar Fork built DocSentry, a tool that reads a loan document and flags the ones that look altered before a human ever opens the file. The team checked it against thirty two hundred real documents pulled from the last two years, and it got ninety seven out of a hundred right. Fast, cheap, a lot better than the light box.
So the plan going into launch day was simple. Since the eval passed, turn DocSentry on for every single loan application, every document type, all at once, starting Monday.
DocSentry's flag rate, scanned documents vs. phone-photographed documents
From a shadow test against 500 real phone-photographed documents already in the queue. Same model, same rules, never counted in the eval.
Scanned documents (the eval set's shape)
3%
Phone-photographed documents (shadow test)
29%
Phone photos are already a third of Cedar Fork's real loan documents. Blended across everyone applying, that pushes the company-wide flag rate from three percent to almost twelve percent, and it lands the same week DocSentry goes live for everyone, not on one watched batch first.
Here's the turn. The extra flags on scanned documents were never the real risk, those look almost exactly like the eval set. The real risk is a document shape the eval never saw, hitting the model at full volume on day one, before a single person has watched even one of them go wrong.
Ninety seven percent on the test and ninety seven percent on Monday are two different numbers.
At its worst, that costs a member their loan the exact week they need it, a car repair, a rent deposit, and there's no flag anywhere saying this document type was never in the test, just a generic dispute line and a five to seven day wait.
Knowledge spark: what's a shadow test
Running the model against real, live documents without acting on what it says. You watch how it would have called them, so a gap between the eval and reality shows up on a screen before it ever touches a member.
The decision I would take back
Reading "97 percent on our eval" as the same thing as "97 percent in production," and using that one number to sign off on launching every document type to everyone at once. It's a real number. It's just not a number about the documents the eval never saw.
What I would leave alone. DocSentry's bank statement check, the part that adds up the listed deposits and flags the ones that don't total right, doesn't care whether the document came in as a phone photo or a scan. It's arithmetic, not a learned pattern in an image. That part can launch to everyone on day one with no canary, because the input shift that breaks the fraud-image model doesn't touch it at all.
The lesson. An eval score is a report on the documents you gave it, not a promise about the documents you didn't. If the eval passing is the whole justification for skipping a phased rollout, the plan has skipped the actual test, which is what real, unlabeled documents look like once they start arriving for real.
Now here is the same thing as a story
The short version sits above. Read this one for the Monday Grady's account got frozen for driving normally.
Grady Munsell spent nine years on a warehouse floor before he started driving nights for two rideshare apps, stacking shifts around a schedule that suits him better now. He's good at it, he knows which streets pay and which ones waste gas.
He took out a small car loan through Cedar Fork two years ago, back when he still had the warehouse job. He scanned his pay stub at a branch, the loan cleared in three days, and he never thought about the process again. It just worked.
One side can flip the switch. The other only finds out what it did.
This spring he applied again, a home repair loan this time, and Cedar Fork's app let him snap a photo of his rideshare earnings screen right on his phone instead of driving to a branch. Easier for him. He didn't think twice about it.
DocSentry had never really met a document like that. Among the scanned pay stubs it learned from, glare and cropped corners and a screen's reflection almost always meant one thing: someone had tampered with the page. So on the Monday it went live for every document type at once, it did what it was built to do when it sees that shape. It flagged Grady's screenshot as altered and froze the loan, the same week his contractor needed a deposit.
We didn't catch one bad document. We proved, to everyone applying that week, that DocSentry had never met a paycheck like theirs.
He called support. Nobody who picked up could tell him why. There was no flag on his file saying DocSentry had never priced a document shaped like this before, just the general dispute script, the same one for a stolen card, and a promise someone would look in five to seven days. He didn't have five to seven days. He put the deposit on a credit card instead.
The step that should have caught a new document type, and never got built
None of this was carelessness. Isabela Petrosyan, the product manager who owned the launch, had a draft of the channel split sitting in her notes weeks before Monday, the scanned documents going live first, the photo channel held for a human check. She just hadn't gotten anyone to sign off on it, and the roadmap had a date on it that nobody wanted to move.
Two days before launch, she'd raised it with Denys Sotelo, the data science lead who'd built DocSentry. He wasn't wrong about the model. He was right that it was the same weights, the same weekly scoring job, running against a document that had simply arrived through a different button on the app. What he hadn't done, what nobody had done, was point DocSentry at even one real photo from that button before Monday. The eval said 97 percent, and 97 percent was the number the room trusted.
Play the same Monday forward with the channel split already built. Scanned documents go live on schedule, exactly as before. Grady's photo still routes to a fraud analyst instead of straight to a decision, because the photo channel hasn't earned full automation yet. Same screenshot, same glare, same cropped corner, DocSentry still tags it as altered. But this time that tag lands in a queue instead of on his account. An analyst reads the earnings screen, sees it's real, and clears it in under two minutes. Grady's loan moves on schedule. He never finds out any of this happened.
The model is identical on both Mondays. What changes is who stands between DocSentry and the first phone-photographed paycheck it ever meets: a fraud analyst with a queue, or Grady, alone, with a frozen loan and a five to seven day wait.
Here's what I'd want to go back and say, in that two-days-out conversation: a passed eval is a fact about the documents we already showed the model. It is not a fact about the documents we hadn't. We treated them as the same fact, and bet every document type on that mix-up before Monday even started.
GUARD, for a plan that only tested yesterday's documents
This reads like a launch-timing question. The real test is who meets the gap between the eval and the real documents first.
G, groups. The eval set's thirty two hundred covered documents, almost all of them scanned, mostly traditional pay stubs, built over the last two years, versus the roughly eight hundred real loan applications Cedar Fork processes every month now, a third of them phone-photographed gig or freelance income the eval sample barely has an example of.
U, unequal. The harm lands on the document shape the eval never saw. On day one, every member who submits a phone-photographed, gig-style income document that week gets flagged the same Monday, all at once, not caught first in a small watched batch.
A, ability to contest. A member like Grady has no way to know DocSentry has never really been tested on the kind of document he's submitting. He just gets a freeze and a generic support queue. Nobody upstream even knows there's a gap until the complaint volume makes it impossible to miss.
R, reduce. An eval score proves the model is good at last year's documents, not this year's. Launch scanned documents on day one, they match the eval set. Hold phone-photographed documents back, route every auto-freeze on that channel to a person for eight weeks or five hundred decisions, and widen only once the override rate proves the gap has closed.
D, detect. Before day one, run DocSentry silently against a fresh, unlabeled sample of last month's real incoming documents, and compare its flag rate on that live sample to its flag rate on the eval set. Once live, track the held-back channel's override rate by document type, on its own dashboard, watched weekly, not folded into one company-wide number.
How often a person overturned DocSentry's freeze on a phone-photographed document, week 1 to week 8
The dashed line is the 15% bar that lets the human check come off. It falls under it in week eight.
Week one, an analyst overturned seven in ten of DocSentry's phone-photo freezes, almost all glare and cropped corners. Engineering shipped an image cleanup step that runs before scoring, and the override rate fell each week as it rolled out. By week eight it's under the bar, and the channel is cleared to go fully automatic.
What doesn't count as fixing this
If the fix here is "launch it and watch the dashboard closely" or "add a second review pass after rollout," it doesn't count. That's a bigger dial on the same all-at-once bet, not a smaller one. The only version that closes the gap is testing the model against the documents it's never seen, at a size small enough that being wrong shows up in a queue, not a headline.
And if you want to be sure it really works, try it somewhere else
Windrow Cooperative built CropGuard, a tool that scores a photo of a crop leaf and flags early blight before it spreads. It was built and checked against leaf photos from the co-op's home valley, pale clay-free soil, even overhead light. The eval passed at 94 percent, and the plan is to turn it on for every grower the co-op serves, including a foothill region that joined this season, on the same day.
G, groups. Windrow's agronomy team, led by Tobiah Quintrell, who built and checked CropGuard against leaf photos from the home valley's soil and light, versus every grower in the newly joined foothill region, about to get the same automatic spray alerts with none of that checking behind them. U, unequal. Barely matters for growers still in the home valley. It lands hardest on the foothill growers, where reddish soil dust on a leaf's underside reads to CropGuard as the same rust-colored blotching that signals blight in the home valley's data. A, ability to contest. A foothill grower has no way to know CropGuard has never scored a leaf photographed against their soil. They just get a spray alert and a bill for a treatment they didn't need, and only learn why after a season of arguing with the app. R, reduce. Launch CropGuard for the home valley on day one, that's what the eval set is. Roll it out to one foothill cluster first, about forty growers, and hold every automatic spray alert there for an agronomist to check against a real photo, for one full growing season. D, detect. Before day one, run CropGuard silently against last season's real, unlabeled leaf photos from the foothill region, and compare its alert rate there to its alert rate in the home valley. Once live, track the foothill cluster's override rate on its own, not folded into the co-op-wide number.
Swap the trigger and it still runs
Speed: leadership wants DocSentry live before the next board update, so the eight-week hold gets squeezed to two days instead.
Cost: the held-back channel means paying fraud analysts to review flagged photos by hand, so it keeps getting proposed as "just watch the dashboard closely" instead, because that's free.
The model gets better: the override rate keeps dropping every week, which makes it tempting to go fully automatic on the photo channel at week two instead of finishing the eight, because the numbers already look nice.
Where people run it wrong
Treating "97 percent on our eval" as proof it'll work on every document type. The real test is whether the model has ever seen this shape of document before, and average accuracy doesn't answer that.
Writing the channel split and the window after watching how the first numbers came in. A safeguard set after you already know the answer is just a label on what already happened.
Letting the team that owns the launch date also decide when the human check comes off.
How to use it live
Ask what the eval set actually contains, and what's showing up in production that isn't in it, before asking whether the eval score itself is real. That's usually where the gap is hiding, in about five seconds.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
What framework fits critiquing a plan that skips a phased rollout because the eval passed, and why?
Tap to flip
ANSWER
GUARD, for risk and fairness. The eval score is real. The real question is who meets the gap between the eval and the real documents first, and whether they'd ever know.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Grady Munsell, a Cedar Fork Credit Union member who switched from warehouse work to rideshare driving and applied for a loan using a phone-photographed earnings screen.
3 · THE HABIT
What did Grady trust about Cedar Fork's loan process, and why?
Tap to flip
ANSWER
His last loan, a scanned pay stub from his old job, cleared in three days with no trouble. So he expected the same the second time, with no reason to think a photographed screenshot would be treated any differently.
4 · THE GAP
What's the document shape DocSentry had never actually been tested against?
Tap to flip
ANSWER
A phone-photographed gig-income screenshot, glare and cropped corners and all. About one percent of the eval set was phone photos, against a third of Cedar Fork's real submissions.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Reading "97 percent on the eval" as proof the plan was safe to launch every document type to everyone at once. It made sense because the score was real. It just wasn't a score about the documents nobody had shown the model yet.
6 · THE NUMBER
Fill in: the shadow test against 500 real phone-photographed documents flagged ______ percent of them as altered.
Tap to flip
ANSWER
29 percent, against a 3 percent flag rate on documents shaped like the eval set. Almost all of it was glare and cropped corners, not real fraud.
7 · THE REPLAY
Same Monday, with the channel split and a human check this time. What changes for Grady?
Tap to flip
ANSWER
DocSentry still flags his screenshot. This time it sits in a queue instead of freezing his loan. An analyst clears it in under two minutes, and Grady's loan moves on schedule. He never finds out it happened.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the eval-to-production gap become?
Tap to flip
ANSWER
CropGuard, a crop-disease detection tool at a farming cooperative, rolled out to a newly joined foothill region on day one. The gap: reddish soil dust on a leaf reads as the same rust-colored blight the model learned to catch in a different valley's soil.
Check yourself Score: 0 / 0
Short answer
1. What's the specific gap between DocSentry's eval set and the documents it met on launch day?
Show hint
Think about what fraction of the eval set was a phone photo, versus what fraction of real submissions are.
Show answer
Model answer: "The eval set was thirty two hundred documents, almost all scanned, built over the last two years. Only about one percent were phone photos. By launch day, a third of Cedar Fork's real loan documents were coming in as phone-photographed gig or freelance income. The eval never really tested the document type that was about to be a third of production."
True or false
2. True or false: the phased hold should also apply to DocSentry's bank statement math check, which adds up listed deposits, the same as it applies to the fraud-image check.
True
False
Show hint
Ask whether adding up numbers cares what kind of photo the document arrived as.
Show answer
False. The math check doesn't rely on a learned visual pattern, it just adds numbers. It doesn't care if the document is a phone photo or a scan, so the same distribution shift that breaks the image-based fraud check genuinely doesn't threaten it. That part can launch to everyone on day one.
Multiple choice
3. Why isn't "launch it to everyone and watch the dashboard closely" enough on its own?
A. Watching more often catches the problem just as well as holding freezes for a human check.
B. A false-flag spike inside one document channel barely moves a single company-wide number, so watching the blended average wouldn't show it in time.
C. DocSentry can't be monitored once it's running in production.
D. Phone-photographed documents already get reviewed by a separate team.
Show hint
Do the blended math: two thirds of documents at 3 percent, a third at 29 percent. What does the total look like against a company-wide alarm bar?
Show answer
B. A, C, and D all treat "watch it closely" as a real fix. The real problem is that the blended number they'd be watching doesn't isolate the one channel that's actually broken.
Fill in the blank
4. The shadow test against 500 real phone-photographed documents flagged ______ percent of them as altered, against a 3 percent flag rate on scanned, eval-shaped documents.
Show hint
It's the number that proves DocSentry wasn't broken, it just had never met a document shaped this way.
Show answer
29 percent. Almost all of it from phone glare and cropped corners. That gap is the whole argument: the model wasn't wrong on purpose, it just had no real example of this document shape in its eval.
Short answer, apply it yourself
5. Think of a product you use that was clearly built and tested on one shape of input. What's a different shape it might quietly be getting wrong, and how would you find out?
Show hint
Look for a product whose defaults assume a device, a document format, a language, or a lighting condition you don't always have.
Show answer
Model answer: "A receipt-scanning app I use flags handwritten receipts as 'unreadable' way more than printed ones, because it was almost certainly trained mostly on printed receipts. I'd know it was quietly failing on handwritten ones if that failure rate showed up as its own number, instead of getting buried in the app's overall 'success rate,' which stays high because most receipts are still printed."
Short answer
6. What old decision would you take back here, and why did it make sense when it was made?
Show hint
Think about what the 97 percent score was assumed to prove, and by whom, in the room where the launch got signed off.
Show answer
Model answer: "Reading '97 percent passed the eval' as proof the plan was safe to ship every document type to every member at once. It made sense in that Monday meeting, the score was real, the deadline was real, and nobody had separated 'the model is good' from 'the model has been shown this.' Those are two different claims, and it's easy to launch on the first one while believing you've earned the second."
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.