ConceptAdvancedAI Opportunity & Model Strategy / When NOT to use AI / #13
Explain why regulated decision-making often demands explainability that current models cannot provide.
GUARD · the reason nobody could actually give
Whitmarsh Mortgage is an online home lender. Trueline is the model that scores every mortgage application and decides, on its own, which files never reach a person. Fritjof Meskell is the product manager who owns Trueline's roadmap. Halfdan Solbakken is the senior underwriter who took a phone call he could not finish. Gwynneth Loures runs a landscape gardening business, six years in, and was buying her first home when Trueline turned her down.
The direct answer
Regulated decisions like a mortgage denial legally need a specific, checkable reason, not a rough sense of what the model thinks. A complex model's real driver is often a tangled mix of features nobody can cleanly translate into that reason, so the notice it generates is a plausible guess dressed up as an answer. The fix is not a better-written notice. For the band of decisions close enough to the line that the reason actually matters, either use a model simple enough to be its own explanation, or hand the decision to a person who can read the file and say, specifically, why. Some model capability gets traded away on purpose, because a guess that sounds confident is worse than no guess at all.
Do this, in order
Never let a model generate a regulated adverse-action reason it cannot actually verify against its own decision.Why: this is the direct answer. Everything below just protects it.
For the decision band close to the cutoff, use an interpretable model or a human underwriter, not a fallback guess mapped from a complex score.Why: GUARD's Reduce step, the actual design fix, not a better-written form letter.
Name both people before reacting to a denial: the team that decided which files get a human look, and the applicant who gets whatever that decision produced.Why: the Groups step. Skip it and you fix a form letter instead of the real gap.
Watch for where the harm concentrates. It lands hardest on whoever already got the "no" and has the least power to check it.Why: the Unequal step. An underwriter reading the same file would catch the real driver in a minute.
Ask, before launch, whether your own team can reconstruct the true reason for a denial from the file alone, without the generated notice. A "no" is the risk signal.Why: the Detect step. This is what catches it before a regulator or a denied applicant does.
Leave the fast, automated path alone anywhere a decision is clearly good news or clearly, cleanly bad, like an approval that needs no reason at all.Why: gating every decision the same way just teaches the team the whole rule is decoration.
How to answer this, stage by stage
Nobody is grading whether you feel bad for Gwynneth. They're grading whether "explainability current models can't provide" turns into an actual mechanism, with a fix specific enough to defend.
1
Nail the sentence, not the category
Say it like this
"Let me make this one file. A mortgage model scores an application and turns down anyone below a line. The applicant this time runs a small landscaping business, good credit, income that easily covers the loan averaged across the year, but her deposits swing hard between the summer season and the winter. The model scored her just under the cutoff. The denial letter said 'insufficient income.' That's not exactly false, but it's not really the reason either, and I'll answer around that gap, because 'explainability regulation needs' means nothing until there's a real letter a real person is holding."
Why this works
Keeps the answer from turning into a lecture about AI ethics in general.
2
Name your method, fast
Say it like this
"I'll run this as GUARD. Groups, who's actually affected here. Unequal, where the harm concentrates. Ability to contest, can she actually get a real answer if she asks. Reduce, the specific design fix. Detect, how the team would know this gap exists before a regulator finds it for them."
Why this works
Two seconds of structure tells the interviewer you have a method, not a hunch.
3
Name the two people before anything else
Say it like this
"There are two people in this. Whitmarsh's product and underwriting team, who decided which score band gets a human look and which one doesn't. And Gwynneth, who gets whatever that decision produced, with no way to see inside it. One of them is holding the switch. The other one has empty hands."
Why this works
This is the sharpest move GUARD has, and it stops the answer from drifting into "the model was wrong" instead of "someone decided she wouldn't get a real answer."
4
Give the one decision
Say it like this
"Here's what I'd actually do. Below the auto-decline cutoff, Trueline doesn't get to decide alone and doesn't get to generate the reason alone either. Those files go to an underwriter, same as the middle band already does, so a person reads the file and writes the specific reason. Above the clear auto-approve line, nothing changes, because an approval doesn't legally need a reason at all."
Why this works
This matches the direct answer word for word. If it doesn't, the interviewer notices before you do.
5
Prove it with the four sentences that break
Say it like this
"Here's what happens without the fix. About 470 applications a month get auto-declined with nobody ever opening the file. A sample check found that for roughly three in ten of the ones near the cutoff, the reason on the letter wasn't actually the model's real top driver, it was the nearest match from an old list that never got rebuilt. Gwynneth called to ask why. The underwriter who took the call couldn't give her one honest, specific number, because nobody, including him, had ever looked at her actual file."
Why this works
The story below is the long version. This is the same shape, compressed so the interviewer hears the whole thing first.
6
Say the cost out loud
Say it like this
"This isn't free. Sending about 470 files a month back to a person instead of an instant decision adds a few business days and spends real underwriter hours, hours we'd freed up specifically by letting the model decide that band alone. I'd take that trade for the slice where a wrong or unexplainable 'no' is a real, uncontestable harm, and keep the fast path everywhere a decision is clean or good news."
Why this works
Naming the cost is what makes this a real decision, not a wish that speed and compliance were both free.
7
Close on a number they could go check
Say it like this
"You'll know it's working when compliance can pull any denial near the cutoff and get a reason that matches the applicant's actual file, every time, not most of the time. You'll know it's still broken when the only honest answer to 'why, specifically' is a shrug about what the model was probably weighing."
Why this works
Ends on something checkable, not a promise that it's handled.
Let's learn
Trueline is the model behind Whitmarsh Mortgage's home loans. Every application gets a score, and the score decides three things: an instant approval, an instant decline, or a hand-off to a person.
Four of Trueline's inputs. Two of them are ordinary. Two are the kind of interaction term that doesn't belong to any single, nameable reason.
Whitmarsh built Trueline fourteen months ago to replace an older scorecard, because the old one kept turning down applicants like Gwynneth for looking, on paper, less steady than they actually were. Trueline reads bank cash flow, gig and seasonal income patterns, and the usual credit file, and it approves real applicants the old model missed. It scores about 3,400 applications a month, on a 300 to 850 scale. Below 615, Trueline declines the file on its own, no underwriter ever opens it. From 615 to 664, an underwriter decides either way. Above 664, it approves on its own.
Knowledge spark: what does the law actually ask for?
A federal rule called Regulation B says a lender who turns someone down has to give the real, specific reason, not a vague one. "Insufficient income" only counts if that's genuinely what drove the decision. A close-enough guess isn't the same thing as the true reason, even if it reads like one.
Here's the turn. Whitmarsh never rebuilt the tool that writes that reason. It still runs off a list of fifteen canned codes built five years ago, for the old scorecard, back when every input mapped cleanly to one code each. Trueline doesn't work that way. It weighs over 300 engineered features, many of them interaction terms, one factor multiplied against another, that were never single, nameable things to begin with. When Trueline's real top driver is one of those, the notice generator can't use it. It falls back to whichever of the fifteen old codes sits numerically closest, and prints that instead.
Denial reason matches the model's real top driver, by review path
Nobody checked the fileA person read the file
Both numbers come from Constanza Prosek's sample of 200 files in each band, checked against what a person could independently confirm from the applicant's actual file.
We didn't send Gwynneth a slightly-off letter. We sent her a reason nobody at Whitmarsh could actually stand behind.
At its worst, this is a real applicant reading a confident, official-looking sentence, believing it's the true reason, and having no way to check whether it is.
The choice I would take back
When Trueline replaced the old scorecard, Whitmarsh kept the same fifteen-code reason library unchanged, and pulled underwriter review off the near-cutoff band to free up hours for the growing volume. Both calls were fine for a simple, cleanly-mapped model. Neither one was ever revisited once the model underneath them got a great deal more complicated.
What I would leave alone
Trueline also ranks and pre-screens files above the auto-approve line, and nothing about that needs this fix. An approval doesn't legally require a specific reason, so speeding up the good news carries none of the same risk. Gating that the same way as a denial would just teach the team the whole rule is decoration.
The lesson: a model getting better at predicting risk and a model getting better at explaining itself are two separate improvements, and building one doesn't hand you the other for free. Regulated decisions need the second one specifically, and a bigger, more accurate model can quietly move further away from it while everyone celebrates the accuracy number going up.
Now here is the same thing as a story
Stage five above compresses this into four sentences. Here's the six weeks underneath, the part a stand-up answer skips.
Gwynneth Loures has run her own landscaping business for six years. Spring through early fall, she's booked solid, planting beds, laying stone paths, keeping a dozen properties looking cared for. Winter is quiet, some months almost nothing comes in. She'd built her whole financial life around that rhythm: a cushion set aside every August for the slow months, bills paid on time every single year, credit that any lender would call solid.
The old scorecard and its reason list were built for each other. Trueline inherited the list without anyone asking whether it still fit.
Under the old scorecard, an application like hers was actually a coin flip, the model saw the winter dip and read it as risk, full stop. Trueline was supposed to fix exactly that. It looks at cash flow across the whole year, not just any single month, and it approved plenty of seasonal businesses the old model would have turned away. For over a year, that was the whole story: a better model, more real applicants getting a fair look.
Gwynneth's file scored 604. The cutoff was 615.
Whitmarsh's team decided, months earlier, which score band would ever get a person's eyes on it. Gwynneth found out what that decision meant on the day it applied to her.
The letter that reached her inbox that evening read: "Your application was declined. Reason: insufficient stated income relative to obligations." Eleven words, official, final. Nothing about it looked like a guess.
It was one anyway. What actually drove Trueline's score wasn't her income level at all, income covered her mortgage payment comfortably averaged across the year, the same way it had for four years running. The real top factor was a single interaction term: the swing between her summer deposits and her winter ones, multiplied against a seasonal volatility figure tied to her declared industry code. That term had no matching code in Whitmarsh's list of fifteen. The system reached for the nearest one it had, "insufficient income," and printed it as if it were the answer.
Three days later, Gwynneth called Whitmarsh's appeal line. Halfdan Solbakken, a senior underwriter, took the call. He read her the letter's wording back to her, slower, kinder, but it was still the same eleven words.
Halfdan wasn't hiding anything from her. He genuinely didn't have it. Her file had never crossed a person's desk before that call.
"Which number," she asked him. "If my income is the problem, tell me the number I'd have needed, and I'll show you I clear it." Halfdan pulled up her file while she waited on the line. There was a score. There was a printed reason. There was nothing underneath either of those he could turn into an answer to her actual question. He told her, honestly, that he'd have to look into it and call her back. He hung up knowing he had nothing to call her back with.
We didn't give Gwynneth the wrong reason. We gave her a reason nobody at Whitmarsh had ever actually checked.
Halfdan couldn't let it go. He'd underwritten mortgages for eleven years, and he'd never once told an applicant "I don't know why" and meant it literally. He flagged it to Fritjof Meskell, Trueline's product owner, the next morning: "I don't think we could answer a regulator on this one any better than I answered her."
Fritjof brought it to Constanza Prosek, Whitmarsh's fair-lending analyst. She pulled 200 auto-declined files from the two quarters before Gwynneth's, all scored between 590 and 614, and did by hand what the system was supposed to do automatically: read each file and work out the real top driver herself. She compared her answer against the letter each applicant had actually received.
138 of 200 matched. In the other 62, close to one in three, the generated reason cited something that provably wasn't the model's real top factor, and in most of those, the true driver was an interaction term with no code to its name.
Two separate questions got treated as one. A file can be close to the cutoff and still have a clean reason. Gwynneth's had neither going for it.
I keep coming back to the meeting fourteen months earlier, when Whitmarsh decided how Trueline would launch. Rebuilding the reason-code list to match Trueline's real complexity was on the table. It got shelved, not out of carelessness, the fifteen codes had covered every case the old scorecard ever produced, and nobody expected the new model's inputs to be different in kind, only in accuracy. That was a reasonable read of a model that hadn't shipped yet. It stopped being reasonable the day Trueline's real driver for a real file was something the list had no way to say.
Here's the redo, run properly. Gwynneth's file, same score, same 604. But now anything below 615 goes to an underwriter, the same as the 615 to 664 band already did. Halfdan reads her actual numbers: strong personal credit, steady annual income that covers the payment, and a debt calculation that penalizes her seasonal swing more than her real risk warrants. He writes the true reason himself, ties it to an adjustable line item, and tells her exactly what would move her score over. She gets an honest answer in four business days instead of an instant, hollow one.
One design let a fifteen-item list stand in for the truth, because it had always been close enough before. The other one checked what actually drove the decision, every time it mattered.
What I'd tell myself, in that first meeting: we asked whether Trueline would be more accurate than the old model. We never asked whether its reasons would still be real ones.
GUARD, run against a reason nobody could stand behind
This was never really about whether Trueline's score was wrong. It probably wasn't, by its own numbers. GUARD is for naming who pays when a model's real reason and its printed reason stop being the same thing.
GGroups. Who actually decides, and who's stuck with it.
Gwynneth, and every Whitmarsh applicant whose file lands below the auto-decline cutoff with no underwriter ever seeing it. And on the other side, Whitmarsh's product and underwriting leadership, who set where that cutoff sits and which files skip a human look entirely. Halfdan sits in between: he inherited a system he had no part in designing, and he's the one who has to sound confident on the phone about a number he can't actually stand behind.
Nobody in this chain acted carelessly. Gwynneth applied in good faith. Halfdan told her the truth as far as he had it. The gap sits in what the system handed him to work with.
UUnequal. Where the harm actually lands.
The 31 percent mismatch rate is the same for every file that lands in the near-cutoff, auto-declined band. The harm from it isn't. An applicant with a simple salaried income, one clean factor driving their score, gets a reason that's almost always genuinely theirs. An applicant like Gwynneth, whose real financial picture is a tangled interaction term Trueline was built to weigh but never built to explain, gets a plausible-sounding reason that may not be the true one at all, and has no way to tell the difference from where she's standing.
The exact people Trueline was built to serve better, seasonal and self-employed applicants with real but uneven income, are the same people most likely to trip the reason-code gap. The model's win and the explainability failure share a cause.
AAbility to contest. Could Gwynneth actually get a checkable answer.
No. She asked the single most reasonable follow-up question there is, "which number," and the honest answer required opening a file nobody had opened before her call. The printed reason carried no hedge, no citation back to an actual figure in her application, nothing distinguishing a person-verified reason from a fallback guess. A confident, official sentence reads as authoritative whether or not it's checkable, and Whitmarsh gave her no way to find out, on her own, that hers might not be.
GUARD's sharpest question isn't whether the reason was wrong. It's whether the person reading it had any real path to find out. For Gwynneth, until Halfdan escalated it, the honest answer was no.
Four of these five steps happened exactly as designed. The step that mattered, a person reading the actual file, was never in the path at all.
RReduce. The actual fix, not a policy memo.
Below the 615 cutoff, Trueline no longer decides alone and no longer generates the notice alone. Those files route to an underwriter, exactly like the 615-to-664 band already did, so a person reads the real file and writes a specific, checkable reason. Above 664, nothing changes, an approval carries no legal reason requirement, so there's no gap to close there. Trueline's score still feeds the underwriter as one signal among several; it just stops being allowed to write the final sentence on its own for the band where that sentence is legally load-bearing.
The alternative on the table was cheaper and got rejected: have a second model turn Trueline's feature-weight output into a fluent, natural-language explanation instead of routing to a person. Rejected because a fluent explanation from a second model doesn't verify that the explanation is true, it just makes an unverified guess sound more convincing, and harder for anyone, including Whitmarsh, to catch when it's wrong.
DDetect. How you'd know this gap exists before it hurts someone.
The honest pre-launch question was never asked: could Whitmarsh reconstruct, from the file alone, the true reason behind a Trueline denial without looking at the generated notice. That question was the risk signal, sitting there unasked for fourteen months. Now Constanza's team runs that exact check quarterly on a random sample of near-cutoff declines, comparing what a person independently finds against what the notice actually said, and tracks the match rate as its own number, separate from Trueline's accuracy score.
The failure worth saying plainly: if the only thing standing between an applicant and an unverifiable "no" is whether an underwriter happens to escalate a phone call, that isn't detection. That's a compliance gap waiting for someone outside the company to find it first.
Share of near-cutoff denials with a reason that matches the real driver, month by month
Before the fixAfter the fix
Seven months flat around 69 percent, unwatched because nobody was measuring reason accuracy on its own. Gwynneth's call in month 7 is what finally made someone look.
The trade-off, said out loud: routing the roughly 470 auto-declined files a month back to an underwriter instead of an instant decision adds a few business days of turnaround and costs real underwriter hours, hours that had been freed up specifically to handle Whitmarsh's growing volume. Whitmarsh took that trade on purpose, for the slice of decisions where an unverifiable reason is a genuine, uncontestable harm, and left the fast, fully automated path in place for the clear majority of files where the reason was never in doubt.
And if you want to be sure it really works, try it somewhere else
Same five letters, a trucking company's hiring model instead of a mortgage, and the same explainability gap shows up under a different regulation.
Same fork Whitmarsh faced, a completely different industry. The line an applicant can't check is still the one that decides.
Vantage Freight, a trucking carrier, uses Roadscore to auto-screen driver applicants against safety and reliability signals before a human recruiter ever sees a file. A federal law called the Fair Credit Reporting Act works a lot like Regulation B here: reject someone based on a background-style score, and they're owed a specific reason, not a vague one.
Same rank, mapped onto Roadscore: size it by what the applicant can never check for themselves. Roadscore auto-rejected candidates whose score sat just under its hiring line, and the rejection notice cited "insufficient recent driving history," a code built for Roadscore's earlier, simpler version. After Vantage retrained the model on richer telematics data, the real top factor for a growing share of near-cutoff applicants became an interaction term, hard-braking frequency weighted against the specific mix of routes they'd driven, something the old code never covered. The fix runs the same shape as Whitmarsh's: near-cutoff rejections, and anything scored soon after a model retrain, get held for a human reviewer instead of an automatic notice, so the reason on the letter is the real one.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: never let a model generate a regulated reason it can't verify against its own decision, full stop.
Cost: no budget this quarter to route every near-cutoff file to a person. Do the cheaper version first: show the applicant the actual computed values behind the top two or three factors, not just a category name, so at least what's printed is checkable even without a human rewrite.
The model got better, for real: Trueline's accuracy on seasonal-income applicants improves next year, catching more real approvals than before. That's a reason to re-test whether the reason-mapping gap has narrowed, with fresh audit data, not a reason to assume a better model quietly fixed its own explainability on the way up.
Where people run it wrong.
They treat a confident, official-sounding sentence as proof it's also the true reason, when sounding certain and being verified are two different things a model can produce separately.
They let overall model accuracy stand in for "explainable," when nobody ever checked whether the stated reason for a specific decision actually matches what drove it.
They wait for a denied applicant's phone call, or a regulator's exam, to reveal the gap, instead of asking before launch whether the team can reconstruct a real reason from the file alone.
How to use it live. Before answering, ask out loud: "if this exact applicant asked for the specific number behind this decision, could anyone here actually give it to them?" If the honest answer is no, that question alone is most of the real answer.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
Which framework fits "why does regulated decision-making need explainability current models can't provide"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real test isn't whether the model's score was right, it's whether anyone can turn that score into the specific, checkable reason the law actually requires.
2 · THE PEOPLE
Who are the people this answer names?
Tap to flip
ANSWER
Fritjof Meskell, the product manager who owns Trueline. Halfdan Solbakken, the underwriter who took Gwynneth's call. Gwynneth Loures, the applicant who got an unverifiable reason. Constanza Prosek, the analyst whose sample found the pattern.
3 · THE OLD DEFAULT
What did Whitmarsh keep using as its reason-generator after Trueline replaced the old scorecard?
Tap to flip
ANSWER
The same fifteen canned reason codes built five years earlier for the old, simple scorecard, where every input mapped cleanly to one code. Trueline's complex, interaction-heavy score never got its own mapping built.
4 · THE MECHANISM, IN ONE LINE
What's the actual mechanism this question is testing?
Tap to flip
ANSWER
A regulated decision needs a specific, verifiable reason. A complex model's real driver is often a tangled interaction term that doesn't reduce to one, so the generated reason becomes a plausible-sounding guess instead of a genuine, checkable answer.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pulling underwriter review off the near-cutoff band when Trueline launched, and keeping the old scorecard's reason-code list unchanged, on the assumption a more accurate model would only need more accurate reasons, not a different way of generating them.
6 · THE NUMBER
Fill in the blank: Constanza's sample found the printed reason matched the model's real top driver in ___ of 200 near-cutoff auto-declines.
Tap to flip
ANSWER
138 of 200, about 69 percent. The other 62, close to one in three, cited a reason that provably wasn't the model's actual top factor.
7 · THE REPLAY
Same file, new design. What changes?
Tap to flip
ANSWER
Anything below the 615 cutoff routes to an underwriter, same as the referred band already did. Halfdan reads Gwynneth's real numbers, writes the true reason himself, and tells her exactly what would move her score. Four business days, and it's honest.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and what's the equivalent explainability gap?
Tap to flip
ANSWER
Vantage Freight's Roadscore, a driver-hiring safety model bound by the Fair Credit Reporting Act instead of Regulation B. Its equivalent gap: a rejection reason built for an earlier model version, no longer matching the retrained model's real top driver after a retrain added a new interaction term.
Check yourself Score: 0 / 0
Multiple choice
1. Why couldn't Halfdan give Gwynneth a real number when she asked which one she'd needed?
A. He wasn't allowed to discuss denied applications with applicants.
B. Her file had never been opened by a person, and Trueline's real driver wasn't something the reason-code system could turn into a specific, checkable answer.
C. Trueline's servers were down that week.
D. Gwynneth's credit score was too low to qualify under any circumstance.
Show hint
Check the Ability to contest step in the GUARD recap.
Show answer
B. Her credit and income were solid. The system had never generated a genuinely verified reason for her file in the first place, so there was nothing honest for Halfdan to hand her.
True or false, with why
2. True or false: the core problem here was that Trueline's risk score was inaccurate.
True
False
Show hint
Check the lede of the GUARD recap section.
Show answer
False. The score itself was likely fine by its own numbers. The failure was that the printed reason for the score wasn't verifiably the score's real driver, which is a different problem from accuracy.
Fill in the blank
3. Constanza's sample found that about ___ percent of near-cutoff auto-declines had a printed reason that didn't match the model's real top driver.
Show hint
Check flashcard 6, and the bar chart in Let's learn.
Show answer
31 percent. 62 of the 200 sampled files, roughly one in three, mostly because the real top driver was an interaction term with no matching code.
Short answer, name the rejected alternative
4. Whitmarsh considered one other fix besides routing near-cutoff files to a person. What was it, and why was it rejected?
Show hint
Check the Reduce step in the GUARD recap.
Show answer
Model answer: Using a second model to turn Trueline's feature-weight output into a fluent, natural-language explanation. Rejected because a fluent explanation from another model doesn't verify that the explanation is actually true, it just makes an unverified guess sound more convincing and harder to catch.
Short answer, apply it yourself
5. Think of a decision an AI tool has made about you, or someone you know, where you were given a reason but had no way to check whether it was the real one. What was it?
Show hint
Look for a moment where a system gave you a confident-sounding reason with no number or detail you could actually verify against.
Show answer
Model answer: A job application rejected by an automated screener with the message "your experience doesn't match our current needs." No specific skill, requirement, or gap was named, so there was no way to know if the stated reason was the real one or a generic default.
Fill in the blank, work the number
6. Whitmarsh processes about 3,400 applications a month. If 14 percent fall below the auto-decline cutoff, roughly how many files a month move to underwriter review under the fix, and is that more or less than the 62 that had a genuinely wrong reason each month before it?
Show hint
3,400 times 0.14, then compare to 62.
Show answer
About 476 a month, far more than the 62 that were actually wrong. A real fix always reviews more files than the number that turn out to have a genuine problem. That's the cost of catching the ones that do.
Before you close the answer
Why this works
Tests whether you can name the actual mechanism, a regulated decision needs a specific, verifiable reason, and a complex model's real driver often doesn't reduce to one, instead of just saying "AI should be explainable" and stopping there. Most candidates stop at "add an explanation."
Follow-up traps
"Couldn't you just use SHAP or feature importance to generate the reason automatically?" Response: that's close to what Whitmarsh's old system already did, mapping the top computed factor to a code, and it's exactly what produced the 31 percent mismatch. Feature importance describes what correlated with the score, not a stable, causal reason a person can act on, and Constanza's team also found the top-ranked factor could flip between two features on nearly identical files.
"Isn't sending files to a person just slower and more expensive for no real gain?" Response: it's genuinely slower and it genuinely costs underwriter hours, about 470 files a month. Whitmarsh accepted that specifically for the slice where a wrong or unverifiable "no" is a real compliance exposure and a real harm to a real applicant, not across the whole portfolio.
If pressed
The instability Constanza found is a known weak spot in post-hoc, feature-attribution methods on ensemble models: two applications with materially identical inputs can produce different top-ranked "reasons" from tiny numeric differences, because the attribution reflects the model's local sensitivity, not a fixed, singular cause. Whitmarsh's fix sidesteps the whole question by not asking a post-hoc method to answer it: for the regulated band, a person reads the file directly instead of trusting an attribution method to have found the one true reason.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.