The direct answer
Run the readiness review segment by segment, not as one blended score. Build the test sample to match production's real mix of document types, especially the kind that got skipped in the POC, and require every segment to clear the accuracy bar on its own before anyone signs off. A tool that's right 98 times out of 100 can still be wrong for one in three of the people it was never actually tested on, and one clean average will hide that every time.
Do this, in order
Test the model segment by segment before you trust the blended number.Why: this is the one practice the whole readiness review turns on.
Build the test sample to match production's real mix of document types, not whatever files were fastest to pull.Why: the POC's own sample skewed toward simple, fast-closing loans, and quietly left the harder ones out.
Require every segment, not just the average, to clear the accuracy bar before sign-off.Why: a segment that's failing only shows up if someone actually checks it on its own.
After launch, watch the underwriter override rate broken out by document type, including types nobody thought to test at launch.Why: this is how you catch the blind spot the readiness review itself didn't know to name.
Leave the standardized, salaried-file segment alone.Why: rebuilding a segment that was never wrong just slows the launch down for nothing.
If a launch date is already set on a blended number, say plainly that it hasn't been checked segment by segment yet.Why: the honest delay costs a week now; the alternative costs a wrongly declined mortgage later.
How to answer this, stage by stage
Seven moves. Naming who can't tell the call was made on POC-size evidence, and writing the actual checklist item, are where the real answer lives.
1
Ground it in one real product before naming a framework
Say it like this
"Say a bank builds a tool that reads a mortgage file's income documents, pay stubs, W-2s, 1099s, bank statements, and tells the underwriter whether the numbers actually line up. Right now it's a POC. It's been tested in-house on 150 loan files, and it agreed with what the underwriter actually decided 98 percent of the time."
Why this works
Grounds a broad question in one specific product before naming a method, so the answer can't drift into generic checklist language.
2
State your structure in one line
Say it like this
"I'd use GUARD here, because 'what readiness review would you run' is really asking who a promotion decision protects and who it exposes. Who's affected, where does a skipped check land hardest, who can't tell the decision was made on POC-size evidence, the actual checklist item you'd add, and how you'd catch what that checklist still missed."
Why this works
Two seconds that show you have a plan before you say a single specific thing.
3
Name both groups, not just "the bank"
Say it like this
"There's Northbridge Bank, who gets a faster loan decision the day this tool goes live. And there's every borrower whose file the tool checks after that, most of whom don't look anything like the 150 files it was actually tested on."
Why this works
This is GUARD's G step. The same promotion decision gives the two groups two completely different starting points.
4
Show where the skipped check lands hardest
Say it like this
"The POC's 150 files were last quarter's closed loans, and self-employed and gig-income borrowers only made up 8 of them, about 5 percent, because those loans usually take longer to close and most hadn't finished yet. But they're 27 percent of Northbridge's real loan book. Their income proof is a 1099 and a bank statement instead of a clean W-2, and that's exactly the file type the tool never really saw."
Why this works
This is U. It names the real document pattern the sample skipped, not "the model might be biased" in the abstract.
5
Name who can't tell the call was made on POC-size evidence
Say it like this
"Emil Solberg, the underwriting director, signs off on the rollout because the tool scored 98 percent. Nobody tells him that number came from a sample that was 95 percent W-2 salaried files. He's approving a number he has no way to check against the loan book it's about to run on."
Why this works
This is A, GUARD's hardest step, and the one a readiness-review question usually skips.
6
Give the one checklist item, not a general sign-off
Say it like this
"Here's what I'd add to the checklist: a segmented accuracy sign-off. Before promotion, test the model against a sample sized to match each document type's real share of the loan book, W-2, self-employed, everything else, and require every one of those segments to clear the bar on its own. Not the blended number. Every segment."
Why this works
This is R, and it's the actual answer. A named checklist item you could point to, not "test it more."
7
Say how you'd catch what even that checklist missed, then close
Say it like this
"Once it's live, I'd track the underwriter override rate by document type, every quarter, including types the launch checklist never singled out, foreign income, trust income, whatever's next. If one type's override rate climbs while the blended rate looks fine, that's the flag. So: segment the readiness review, require every segment to pass on its own, and keep watching by segment after launch, because the review will always miss a category nobody thought to name yet."
Why this works
Closes on the direct answer in one breath, and shows the fix isn't a one-time checklist, it's an ongoing watch.
Let's learn
The tool is a check that runs on a mortgage file the moment an underwriter opens it, before they've read a single document themselves.
This is Northbridge Bank's newest underwriting tool. It reads the income documents in a loan file, pay stubs, W-2s, 1099s, bank statements, tax returns, and tells the underwriter whether the numbers on them actually agree with each other.
Before this tool, an underwriter checked those documents by hand. That took about 50 minutes a file, on top of everything else in a loan review, and a loan usually took 9 business days to close.
Knowledge spark: what's a POC?
A proof of concept. A small, early version of a tool, tested on a small pile of files to see whether the idea even works, before anyone builds it for every file the bank actually handles.
In its first round of testing, the new tool checked a file in under 6 minutes, and it agreed with what the underwriter actually decided on 98 percent of the 150 files in the test.
Here's the catch. Those 150 files weren't a random slice of Northbridge's loan book. They were last quarter's closed loans, and self-employed and gig-income borrowers barely showed up in them, because those loans usually take longer to close and most hadn't finished yet. When the team later ran the tool against a wider set that actually matched the loan book, on ordinary salaried files it still held at 98 percent. On self-employed and gig-income files, the ones with 1099s and bank statements instead of a clean W-2, its accuracy dropped to 69 percent.
Agreement with the underwriter's decision, by document type
Same tool, same launch date on the table. The only thing that changed was which document type the file actually was.
Ordinary salaried files, W-2 and pay stubs
98%
Self-employed and gig-income files
69%
The tool never got worse. The test just never asked it about the borrower the sample forgot to include.
The real gap wasn't 29 points of accuracy. It was a borrower nobody had built a test file for.
At its worst, this costs Northbridge a self-employed borrower who has a perfectly good loan, and gets wrongly told their documents don't add up, from a tool that looked, on paper, ready to launch.
The choice I would take back
The eval plan tested the tool against whatever 150 files were sitting closed and ready, because that data was fastest to pull and the launch date was already set. That made sense when nobody had stopped to ask whether "closed last quarter" looked anything like the loan book the tool would actually run on. It stopped making sense the day that 98 percent became the number the committee used to approve a full rollout.
What I would leave alone. For an ordinary, salaried borrower, a POC file and a production file read almost the same. Rebuilding that whole segment's test set wouldn't move its score at all. Save the segmented effort for the document types that actually vary.
The lesson. One blended number can be exactly right and still hide a group it was never tested on. A readiness review that only checks the average is checking the easiest part of the question.
Now here is the same thing as a story
Read the short version above if you're pressed for time. Read this one when you want to feel why segmenting a number matters, not just know that it does.
Ronke Halton has run product for Northbridge Bank's underwriting team for four years. She's the one people ask to turn a vague compliance requirement into something an underwriter can actually click through in a morning.
When the mortgage-document tool was ready for its first real test, Ronke pulled the 150 most recently closed loan files. It was the fastest data to get to, the pilot review was three weeks out, and every file in that batch had already gone through a full underwriting cycle, so there was something real to grade the tool against.
For three weeks, this went fine. Every batch she reran, the tool agreed with the underwriter's actual decision 97, 98 percent of the time. Her analytics lead called the result strong in a review meeting. The underwriting committee started asking when it could touch a live file.
At first, Ronke checked which document types made up each new batch before she reran the test, a W-2 here, a 1099 there, roughly matching what she remembered from the loan book. Then she stopped checking the mix and just reran the same 150 files, since the number kept holding. Then, in her final readiness memo to the committee, she reported one line: 98 percent agreement with underwriter decisions. No mention of what kind of file that 98 percent came from.
Then came a Tuesday, three months after launch.
The bank's quality team ran its regular quarterly audit, pulling 40 declined loans at random to check the underwriters' work, a routine the bank had done for years, long before the tool existed. Fourteen of the forty were self-employed or gig-income borrowers. Most of them had been flagged "inconsistent" over things a person would have waved through, a 1099 total that didn't quite match a bank deposit because of a processing fee, a bank statement covering a different month than the pay period.
Same loan decision. Two very different starting points.
Ronke didn't wave it off. She pulled three months of override logs and broke them out by document type herself, something the launch checklist had never asked anyone to do. Ordinary salaried files: the tool's calls matched what an underwriter would have done, 98 percent of the time. Self-employed and gig-income files: 69 percent.
We didn't quietly lose 29 points of accuracy. We quietly wrote the hardest borrowers out of the review that was supposed to protect them.
I want to say the problem is that Ronke cut a corner. She didn't, not really. She never had a number in her head either, just a feeling, the pilot is working, or it isn't. Three straight weeks of 97, 98 percent had switched her feeling to "it's working," and there was no setting in between where she kept asking who that number actually covered.
Back when the eval plan got signed off, in a meeting with her analytics lead and the launch date already on the calendar, the question on the table was simple: is the tool accurate enough to trust with real files. The answer everyone reached, sensibly, was to test it against whatever closed loans were sitting ready, since that was real underwriting data and nothing needed to be invented. Nobody in that room asked whether "closed last quarter" matched the loan book the tool would run on once it was live.
I would go back and put that question in the room. Ronke tests against a sample that matches the real mix from the start, and the 69 percent shows up in week one of the pilot instead of month three of production, before a single self-employed borrower gets a declined loan they never really understand.
One version of the checklist tells you the number is good. The other tells you what the number is actually good at.
What I'd tell the Ronke who wrote that first eval plan: "fast to pull" and "looks like our loan book" were never the same test, and I only ever built the first one.
GUARD, run against a launch decision built on one blended number
This is a risk question, so the framework is GUARD. "What readiness review would you run" sounds like a project-management checklist question, which is exactly why it's easy to answer with process instead of a real check.
G, groups. Northbridge Bank's underwriting team, who get a faster loan decision the day this tool launches. And every borrower whose file the tool checks after that, most of whom look nothing like the 150 files it was tested on.
U, unequal. A salaried borrower with a clean W-2 gets checked by a tool tested almost entirely on files like theirs, and it reads them fine. A self-employed or gig-income borrower, with a 1099 and a bank statement instead, gets checked by the same tool, tested on almost none of them.
The check that should sit here, and doesn't.
A, ability to contest. Emil Solberg, the underwriting director, signs the rollout off because the tool scored 98 percent. Nobody shows him that the sample behind that number was 95 percent salaried files, so he has no way to check whether it says anything about the borrowers it's about to run on.
R, reduce. Add a segmented accuracy sign-off to the launch checklist. Test the model against a sample sized to match each document type's real share of the loan book, and require every segment, not the blended average, to clear the accuracy bar before promotion.
D, detect. After launch, track the underwriter override rate by document type every quarter, including types the launch checklist never singled out. A type whose override rate climbs while the blended rate looks fine is the flag that the review missed something.
Where this answer would fail
If the fix is "run the pilot a bit longer" or "add a review step," none of it counts. A segmented sample sized to the real loan book, and a checklist that requires every segment to pass on its own, are things you can point to on launch day and check whether they happened.
And if you want to be sure it really works, try it somewhere else
A county benefits office's eligibility-screening tool runs into the same gap, in a room with nothing to do with mortgages.
G, groups. Vesna Castro, a program analyst who builds a solo tool to screen benefit applications for Northfield County, using the 200 most recently closed cases because they were the fastest data to pull before her pilot review. And the county's applicants, who inherit whatever the tool decides once the board sets a countywide rollout date.
U, unequal. An applicant who filed everything in English, in person, at the counter, gets screened by a tool tested almost entirely on cases like theirs. An applicant who needed an interpreter or submitted documents in another language, a group that's 22 percent of the county's real caseload but only 4 percent of Vesna's 200 test cases, gets screened by a tool that barely saw a case like theirs.
A, ability to contest. The county board approves the rollout because the pilot scored 95 percent agreement with the caseworkers' own decisions. Nobody on the board asks what share of that pilot came from interpreter-needed cases, so nobody there can tell the number was never really tested on them.
R, reduce. The same practice: build a segmented sample sized to the county's real caseload mix, English-only and interpreter-needed alike, and require both segments to clear the bar on their own before the tool goes live.
D, detect. After launch, track the appeal rate by whether the case needed an interpreter, every month, not just overall. A rising appeal rate in one group while the countywide rate looks steady is the flag.
Swap the trigger and it still runs
- Speed: a faster model just means the untested segment's failure shows up two weeks sooner instead of three months, but the fix stays the same, a segmented sample and a segmented sign-off before the blended number gets trusted.
- Cost: a cheaper model means more teams run a POC solo on whatever data is fastest to grab, which means more segments quietly left out of testing before anyone asks about the mix.
- The model gets better: a sharper model makes the blended pilot score look even better, which makes it easier to wave through a full rollout on one number, not harder, because a better-looking average is exactly what gets a launch date approved.
Where people run it wrong
- Treating a high blended score as proof the tool is ready, when it's actually proof the test matched the easy part of the population.
- Waiting until production to find the segment that fails, instead of stress-testing it against the real mix before launch.
- Turning the fix into "test on more files," which just makes the same skewed sample bigger, not more representative.
How to use it live
Ask one question before you answer: "what share of the test set actually matches what production will see?" That buys you a second to think, and it's usually the exact question the interviewer wanted asked.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits a question about the readiness review before promoting a POC to production, and why?
Tap to flip
ANSWER
GUARD, for risk. The real question isn't whether the POC's number was good, it's who the promotion decision protects, where a skipped check lands hardest, who can't tell the number came from POC-size evidence, the actual checklist item you'd add, and how you'd catch what even that missed.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ronke Halton, a product manager at Northbridge Bank. She can turn a vague compliance requirement into something an underwriter can use in a morning, and built the mortgage-document tool's 150-file test set from the loans that happened to be fastest to pull.
3 · THE HABIT
What did Ronke stop checking as the pilot's number kept looking good?
Tap to flip
ANSWER
Which document types made up each test batch. She went from checking the mix before every rerun, to reusing the same 150 files, to reporting one blended number in her final readiness memo with no mention of what it was built from.
4 · THE SWITCH
What's the two-setting switch in this story?
Tap to flip
ANSWER
Either the readiness review checks every document type on its own, or a launch decision quietly stands on a number that only really describes the easy files. There's no partial version of "we checked."
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
The eval plan tested the tool against whatever 150 loans were already closed and ready, because that data was fastest to pull and the launch date was already set. That made sense before anyone asked whether "closed last quarter" matched the loan book the tool would run on. It stopped making sense the day 98 percent became the number the committee approved a rollout with.
6 · THE NUMBER
Fill in the blank: agreement on ordinary salaried files was ______ percent. Agreement on self-employed and gig-income files was ______ percent.
Tap to flip
ANSWER
98 and 69. Same tool, same launch date on the table. What changed was only which document type the file actually was.
7 · THE REPLAY
Same eval plan, a segmented sample built in from the start. What changes?
Tap to flip
ANSWER
The 69 percent number shows up in week one of the pilot instead of month three of production, before three months of self-employed borrowers get quietly declined, and before a routine audit has to stumble onto it by chance.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the reduce step become?
Tap to flip
ANSWER
Northfield County's benefits-eligibility tool, built solo by analyst Vesna Castro on the 200 closed cases fastest to pull. Reduce: the same segmented practice, a sample sized to the county's real caseload mix, English-only and interpreter-needed alike, with both required to clear the bar before rollout.
Check yourself Score: 0 / 0
Short answer
1. What specific readiness-review checklist item does this answer add, and why did the original eval plan skip it?
Show hint
Look for the decision made when the eval plan was written, not the audit that caught the problem later.
Show answer
Model answer: "The eval plan tested the tool against whatever 150 loans were already closed and ready, because that data was fastest to pull and the launch date was already set. It stopped making sense once that number got used to approve a full rollout. I'd add a segmented accuracy sign-off: test the model against a sample sized to match each document type's real share of the loan book, and require every segment to clear the bar on its own before promotion."
Multiple choice
2. Which two groups does the G step name in Ronke's story, and what gives them different starting points?
- A. Northbridge Bank's underwriting team, who get a faster decision the day the tool launches, and every borrower whose file it checks after that, most of whom don't look like the 150 test files.
- B. Ronke and Emil Solberg, who disagree about the launch date.
- C. The underwriting committee and the quality team, who disagree about the audit result.
- D. The tool and the loan files it was tested on.
Show hint
Look for who controls the promotion decision and who's on the receiving end of it.
Show answer
A. B, C, and D name real people or things in the story, but not the two groups the G step separates: the one the launch speeds up, and the one who inherits whatever the test sample left out.
Fill in the blank
3. The tool agreed with underwriters ______ percent of the time on ordinary salaried files, and ______ percent of the time on self-employed and gig-income files.
Show hint
Both numbers are in the chart, "Agreement with the underwriter's decision, by document type."
Show answer
98 and 69. Same tool, same launch date. What changed was only which document type the file actually was.
True or false
4. True or false: the risk in this story was that the tool's 98 percent accuracy number was calculated wrong.
Show hint
Ask what the 98 percent was actually measured against.
Show answer
False. The number was correct for the sample it was measured against. The risk was that the sample was 95 percent salaried files, so 98 percent measured how well the tool matched the easy part of the loan book, not how it would handle every borrower running through it.
Short answer, apply it yourself
5. Pick an AI tool you've seen tested in a small pilot before a full rollout. What group might that pilot's test data have skipped?
Show hint
Think about which cases were fastest or easiest to grab for the pilot, and who doesn't look like those cases.
Show answer
Model answer: "A resume-screening tool piloted on a hiring team's last 100 hires. Those were the applicants who made it all the way through, mostly people with a standard resume format and a steady job history. It probably never saw a resume with an employment gap, a career change, or a format that doesn't fit a template, exactly the applicants a pilot like that would quietly screen out of its own test."
Multiple choice
6. In this story, where would the segmented readiness review matter least?
- A. Ordinary salaried files, where the tool's accuracy already held at 98 percent in the wider test.
- B. Self-employed and gig-income files, where accuracy dropped to 69 percent.
- C. Any file Emil Solberg would use to decide whether to expand the rollout.
- D. Any file the quarterly override audit would flag.
Show hint
Ask which document type the POC and production would score almost the same on anyway.
Show answer
A. B, C, and D are exactly where the segmented review earns its place. A is what this answer would leave alone: a document type where the wider test already showed the sample was representative enough.