ConceptIntermediateQuality, Cost & Token Economics / Measuring ROI and business impact / #9

Explain the difference between cost avoidance and cost reduction to a finance partner.

PICK · risk scoring for workers' comp underwriting

Underlay is a risk-scoring engine that Northloom built for workers' comp underwriting. Prairiestone Mutual, a mid-market insurer covering manufacturing and logistics employers, runs it across every new application. Wojciech Szymanski owns Underlay's product story at Northloom. Bogdana Onyekwere runs Finance at Prairiestone. One quarter, she booked a number Underlay never actually promised her, and had to explain the gap to her own board.

The direct answer
Split the savings by where the dollar actually comes from. The drop in underwriter overtime and temp staffing is cost reduction: it is already missing from a real invoice, and finance can check it against last year's bill. The claims Underlay's flagged tier is expected to prevent are cost avoidance: a modeled range built from a claims model, not a subtraction from this year's paid-loss ledger. Never let the second number get booked as if it were the first.
Do this, in order
  1. Split every savings claim by where the dollar comes from, before finance sees one blended figure.Why: a single "AI savings" line hides which part is already in the bank and which part is a bet on a bad year that didn't happen.
  2. Book the realized reduction with its own receipt attached, and nothing else.Why: the temp-staffing cut is checkable against last year's invoice. Tie it to that invoice and it survives any audit.
  3. Carry the avoided-claims number as a range, with the word "modeled" printed next to it, not a point estimate.Why: a single dollar figure invites finance to treat a probability as a fact. A range is honest about what the model actually knows.
  4. Never let a probabilistic avoidance claim get booked against this year's loss ratio target.Why: workers' comp claims take years to develop. This year's paid losses can't confirm or deny a claim about the future, so booking it there only sets up a gap someone has to explain later.
  5. Set a real bar for when the avoided-claims range gets to graduate into a reduction.Why: without a kill line, "modeled" becomes a permanent hedge instead of a real position that can update.
  6. Don't undersell the temp-staffing cut just to be cautious about the bigger number.Why: it's real money, already gone from a real bill. Downplaying it starves the product of credit it actually earned.

How to answer this, stage by stage

Nobody is grading whether you know the dictionary definitions of avoidance and reduction. They're grading whether you can point at one specific AI feature's benefit and say, honestly, which kind of dollar it is, and defend that split when someone tries to blend it back together.

1
Anchor it in one product before defining any terms
Say it like this
"Let me ground this in one case. Underlay is Northloom's risk-scoring tool for workers' comp underwriting. Prairiestone Mutual runs it, and Bogdana Onyekwere is the finance partner who has to sign off on what it's actually worth."
Why this works
A dictionary answer to "what's the difference" stays abstract. One real account keeps every claim checkable against a real number.
2
Name the method before naming a dollar figure
Say it like this
"I'll run this as PICK. Take a position on which framing this feature earns, name who feels the confusion between the two, say which mistake is the expensive one, then say what would change my mind."
Why this works
Tells the interviewer a structure is already running, so the next four minutes read as a plan, not a ramble.
3
Give the position, committed, before any reasoning
Say it like this
"Here's my position. For Underlay, I split the savings claim in two on purpose. The drop in temp-staffing spend is cost reduction, because it's already missing from an invoice. The claims the flagged tier is expected to prevent are cost avoidance, because that's a modeled range, not a subtraction from anything that's already been paid."
Why this works
This is the direct answer, said in one breath, before the interviewer has to dig for it.
4
Put a real person and a real number on each side
Say it like this
"If Bogdana books the avoided-claims midpoint, $2.75 million, as a hard cut to this year's loss ratio, she's the one standing at the Q3 reserve review explaining why paid losses barely moved. If she discounts the temp-staffing number too, out of caution, Prairiestone stops crediting Underlay for the $12,800 a month it's already saving, checkable against last year's bill."
Why this works
Turns "there's a difference" into two people who each pay a real price for getting it wrong.
5
Name which mistake is the expensive one, and say why
Say it like this
"Underselling a real reduction is cheap and loud. Someone points at the invoice, corrects you on the spot, and it's fixed in the room. Booking a modeled avoidance as a hard reduction is quiet and expensive. It sits in a budget line for months, looking like a fact, until a reserve review asks where the money went and nobody has a receipt."
Why this works
This is the hardest part of PICK. Naming which error actually costs more is what makes it a real stance instead of a shrug.
6
Say what evidence would move the number from one column to the other
Say it like this
"This isn't a permanent hedge. Once we have two full accident years of matched claims data on the flagged cohort, fully developed and checked against what the model predicted, the avoided-claims number can move from modeled avoidance to a defensible reduction. Until then, it stays a range with 'modeled' printed on it."
Why this works
Shows the split is a live position tied to evidence, not a rule you'd defend forever regardless of what the data does.
7
Prove it with the near miss, compressed to four sentences
Say it like this
"Here's what happens without the split. Prairiestone's dashboard blended both numbers into one line, $3.19 million, no tag saying which part was real. Bogdana built this year's budget around it. At the Q3 review the loss ratio hadn't moved, because workers' comp claims take years to pay out, and the CFO asked her to explain a gap that was never actually a gap, just a range that got mistaken for a fact."
Why this works
Shows the real, countable cost of skipping the split, not just "it could get confusing."
8
Close on the decision, in one breath
Say it like this
"So: split it by where the dollar comes from. A reduction has a receipt. An avoidance has a range. Underselling the receipt is a mistake you fix on the spot. Booking the range as a receipt is the one that costs you a board meeting."
Why this works
Restates the direct answer plainly, so the interviewer leaves with the decision, not just the story behind it.

Let's learn

Underlay is a score. Feed in a business's paperwork, its payroll class codes, its injury history, its OSHA records, and in about twenty seconds it hands back a number from 0 to 100, saying how much workers' comp risk that business carries.

Hand sketched numbered icon list titled Before Underlay, every application got the same 45 minutes. Three rows: a document icon, 2,400 applications a month, each read by hand. A gauge icon, 45 minutes a review, routine or not. A person icon, temp underwriters hired every renewal season to keep up.
Before Underlay, an underwriter at Prairiestone spent 45 minutes on every application, whether it needed that much or not.

Prairiestone Mutual processes 2,400 new applications a month. Before Underlay, each one took an underwriter about 45 minutes: pulling injury history, checking OSHA violations, reading old safety-inspection notes, verifying payroll class codes. Every renewal season, the team hired temp underwriters just to keep the pile from growing.

Hand sketched decision tree titled How Underlay splits an application. Root box, application scored 0 to 100, branching into three leaves. Under 30 leads to routine, 6 minute glance. 30 to 69 leads to standard, 25 minute review. 70 plus leads to flagged, deep review plus inspection, this leaf outlined in red-orange.
One score, three very different next steps. The score doesn't just speed things up, it decides which applications get a mandatory safety inspection before Prairiestone will bind them.

With Underlay, about 60 percent of applications score under 30 and get a 6-minute glance instead of 45 minutes. Another 25 percent land in the middle and still get the normal 25-minute review. The last 15 percent score 70 or higher, and now get a deep manual review plus a required on-site safety inspection before Prairiestone will bind the policy at all, a step that never existed before.

Here's the turn. The routine-tier time savings adds up to about 936 underwriter-hours a month. Some of that shows up as a real, checkable number: Prairiestone used to spend $19,200 a month on temp underwriters every renewal season, and now spends $6,400, a $12,800-a-month cut that matches last year's actual invoice line for line. The rest of those hours meant Prairiestone never had to hire the three underwriters it had budgeted for a Gulf Coast expansion, about $288,000 a year that simply never left the ground. That second number is real too. It's just not the same kind of real.

Knowledge spark: what's a long-tail claim? A workers' comp claim isn't paid all at once. An injury reported this year can take three to five years to fully settle, as medical treatment, disability payments, and legal costs keep coming in. So this year's total paid claims are mostly old policies, not the ones written this year.

The flagged tier is where the bigger number lives. Northloom's claims model compares the flagged, now-inspected applications against similar risk profiles Prairiestone bound without inspection in past years. It estimates the inspection step, and the pricing or decline decisions that follow it, prevents somewhere between $2.1 million and $3.4 million a year in future claims payouts.

Nobody can point to a bill that shrank. This is a claim about a bad year that, statistically, mostly won't happen now.
Hand sketched comparison diagram titled Two kinds of dollars, drawn to scale. Left panel, a plain document icon labeled reduction, caption temp staffing cut, $12,800 a month, checked against last year's bill. Right panel, a filled box with a large question mark labeled avoidance, caption avoided claims, $2.1M to $3.4M a year, modeled, no bill to check it against.
One of these dollars has a receipt behind it. The other has a confidence interval. They don't belong on the same line.
Where Underlay's savings actually come from, by kind
$3.0M $1.5M 0 $0.15M Temp staffing cut (realized) $0.29M Hires never made (avoidance) $2.75M Avoided claims (avoidance)
Realized, checkable against an invoiceAvoidance, a budget line that never openedAvoidance, a modeled claims range
The avoided-claims midpoint is about eighteen times the size of the one number Bogdana can actually check against a bill. That's exactly why it can't share a line with it.
The choice that mattered When Northloom and Prairiestone first built the finance dashboard, they gave it one line: "Underlay Impact." No tag telling anyone which part was cash that already didn't get spent and which part was a claims model's own guess about a year that didn't happen. That was fine when the whole number was small, mostly the temp-staffing piece. It stopped being fine once the avoided-claims estimate grew to ten times the size of everything else on the line.
Hand sketched metaphor scene titled The one idea to remember. Left panel, a plain document icon labeled reduction, caption a bill that stopped arriving. Right panel, a box with a large question mark labeled avoidance, caption a bill that might never have come.
If you remember one line from this whole answer, remember this one.

What I would leave alone: the standard tier, the 25 percent of applications in the middle. Underlay assists there, but nothing about pricing or timing changed, so there's no savings claim to split in the first place. Not every tier needs this fight.

The lesson: a savings number is really a claim about where the dollar came from. If it came from a bill that already didn't get sent, it's a reduction. If it came from a bad year that didn't happen, and you can't yet prove it wouldn't have happened anyway, it's an avoidance, and it belongs on a different line, defended with different evidence.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a single dashboard cell, not a bad model, is what actually cost Bogdana a board meeting.

Ask anyone at Prairiestone Mutual who catches a bad number first, and they'll point at Bogdana Onyekwere. Six years running Finance there, and she's the one who spots a line about to become a problem before the meeting has even settled into its chairs.

Underlay rolled out eighteen months ago. The first two quarters were good ones. The temp-staffing line dropped, clean and visible, right where she expected it to. She could point at last year's invoice, this year's invoice, and the gap between them, and nobody in a budget meeting ever asked her to defend that number twice.

Then Northloom's claims team finished the model behind the flagged tier, and the "Underlay Impact" dashboard line started climbing past anything the temp-staffing cut could explain on its own. Bogdana didn't ask where the extra size was coming from. The number had been trustworthy for two quarters. She let it keep being trustworthy without checking why.

It thinned in three beats, and none of them looked careless. Beat one: a strong quarter for the whole line of business put "AI ROI" on the agenda for the annual planning offsite, and someone asked Bogdana to put a number on Underlay for next year's budget. Beat two: she pulled the one number the dashboard gave her, $3.19 million, and dropped it straight into the loss ratio target for next year, because it was the number the tool had been showing her all along. Beat three: she presented it to the board as a committed reduction, not a modeled estimate, because the dashboard had never once suggested it was anything else.

Hand sketched timeline titled The board deck that treated a range like a receipt. Five milestones: Underlay launches, one blended line, Underlay Impact. Good quarters, the number looks great, nobody splits it. Board deck, this milestone emphasized in red, Bogdana books $2.75M as this year's cut. Q3 review, loss ratio barely moves. The split, one realized line, one modeled line.
Nothing looked wrong at the board deck. It took a full reserve cycle for the gap to show up.

The trigger wasn't a bad quarter. It was a routine one. At the Q3 reserve review, six months later, the CFO ran the usual comparison: budgeted loss ratio against actual. The gap was almost exactly the size of the avoided-claims piece of "Underlay Impact." Nothing had gone wrong with the model. The flagged applications really were being priced up, declined, or inspected the way they were supposed to be. But workers' comp claims take three to five years to fully pay out, so this year's actual paid losses were still almost entirely policies written before Underlay ever scored a single application.

We did not lose Bogdana's trust to a bad model. We lost it to a spreadsheet cell with no unit attached.

The CFO wasn't accusing anyone of cooking a number. He just wanted to know why a committed $2.75 million reduction hadn't shown up anywhere in the actual paid-claims data. Bogdana didn't have an answer in the room, because she'd never had to build one. Worse than the gap itself: for the next two budget cycles, the board discounted every number Underlay produced by half, on principle, including the temp-staffing cut that had never once been wrong.

The decision that opened the door traced back to a fifteen-minute stretch of the original dashboard-build meeting, back when Underlay first launched. Someone from Northloom's data team had asked whether the finance dashboard needed two lines instead of one, a realized line and a modeled line. The room's answer was no. At the time, the modeled piece didn't exist yet. The dashboard only had the temp-staffing number to show, and one line was the whole truth.

Run the same year again, with the split in place instead of one blended cell. Bogdana books $153,600 as a confirmed reduction against this year's temp-staffing budget, backed by two real invoices. She carries $2.1 million to $3.4 million as a separate, clearly labeled avoidance range, footnoted "modeled, confirmable once the flagged cohort's claims fully develop." At the Q3 review, the realized line matches plan exactly. The avoidance range isn't expected to show up in paid losses yet, and nobody in the room is surprised that it hasn't.

One dashboard asked a spreadsheet cell to make a promise it had no way of keeping. The other asked it to just say, honestly, what kind of promise each number actually was.

What Wojciech would tell himself, back in that fifteen-minute meeting: the blended line wasn't a simplification. It was a decision to let finance treat a probability like a fact, and nobody in the room had meant to make that decision on purpose.

PICK: telling a real dollar from a modeled one

Not a way to dress up "it's complicated" in four letters. PICK forces a real commitment about which framing this feature earns, then makes you say, out loud, which kind of mistake actually costs someone money.

PPosition. Your pick, in one sentence, before any reasoning.
For Underlay, the routine-tier time savings that already left a real invoice is cost reduction. The flagged-tier claims Underlay is expected to prevent are cost avoidance, a modeled range, never booked as a hard line-item cut.
Say the split before the reasoning, or the interviewer spends the next two minutes waiting to find out what you'd actually tell finance.
IImpact. Who feels each kind of error, in what units.
Book the avoidance as a reduction, and Bogdana stands at a reserve review explaining a $2.75 million gap that was never actually a gap. Undersell the reduction as avoidance, and Prairiestone stops crediting Underlay for $12,800 a month it's already keeping off the books.
Naming both people, the one who overclaims and the one who underclaims, keeps this from turning into a one-sided caution story.
CCost asymmetry. The heart of it.
Underselling a real reduction is cheap and loud. Someone points at the invoice and corrects you on the spot, in the room. Booking a modeled avoidance as a hard reduction is quiet and expensive. It sits inside a budget line for two quarters, looking like a fact, until a reserve review has to explain why the paid-claims number never moved.
This is the step that earns the pick. Anyone can say "there's a difference." Naming which mistake actually costs more is what survives a follow-up question.
KKill criteria. What evidence would flip the framing.
Once the flagged cohort has two full accident years of matched claims data, fully developed and checked against the model's own prediction, the avoided-claims range can move from "avoidance, modeled" to a defensible reduction. Until then, it stays a range, not a point number, no matter how good the story sounds.
A pick with no kill criteria is just a caution you're defending forever. This makes it a position you'd actually update, on purpose, when the evidence earns it.

And if you want to be sure it really works, try it somewhere else

Same four letters, a warehouse network instead of an insurer, and this time the modeled dollar isn't a claim, it's a breakdown that didn't happen.

Trestlewatch is a predictive-maintenance model. It watches vibration and temperature sensors on conveyor and sortation equipment across Portrune Logistics' 40 warehouses, and flags any machine likely to fail within the next 14 days. Before Trestlewatch, every machine got checked on a fixed weekly rota, whether it needed it or not, costing about 1,800 technician-hours a month across the network.

Hand sketched quadrant diagram titled Portrune's two dollars, sorted the same way. X axis how visible the mistake is, from hidden to obvious. Y axis how expensive it gets, from cheap to costly. Fewer rota inspections plotted obvious and cheap. Avoided downtime booked as fact plotted hidden and costly. A conveyor failure caught in time plotted in the middle.
Same PICK, a different domain, a different modeled dollar. Both times, a flat dashboard line was the thing that let it hide.
The decision Portrune would take back Trestlewatch's rollout deck had one number: "$550,000 saved a year." Nkosana Odogwu, who runs maintenance finance at Portrune, booked it as a straight cut to this year's maintenance budget. The real, checkable piece, about $25,000 a month in inspection overtime that stopped, was true from month one. The modeled piece, a range of $410,000 to $690,000 in avoided unplanned downtime, was a guess about breakdowns that didn't happen, and a quiet quarter, with abnormally low downtime everywhere, made it impossible to reconcile against anything on an actual ledger.

Same rank, different lever: Lenka Grishin, who leads maintenance at Portrune, didn't design the rollout deck. The fix isn't a smarter sensor. It's splitting the claim: the overtime drop is a reduction, checkable on a timesheet every month. The avoided downtime is an avoidance, a range, only provable after enough real months pass without the failures the model predicted.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the split: any hour that's already off a timesheet is reduction. Any breakdown that didn't happen is avoidance, always a range, never a line-item cut.
Cost: no budget this quarter for a full matched-cohort study. Ship the cheap version first: two separate dashboard numbers with two different colors, not two decimal places crammed onto one number.
The model got better, for real: say Trestlewatch's failure-prediction accuracy doubles overnight. The split barely changes. The avoided-downtime range gets narrower and more trustworthy, but it still isn't a reduction until real downtime data over a matched period confirms it.

Where people run it wrong.
They let one dashboard cell hold both kinds of dollars, so nobody downstream can tell which one they're looking at.
They wait for a perfect matched-cohort study before booking anything, and the product never gets credit for the real, checkable reduction that was sitting there the whole time.
They treat one quiet quarter, or one bad one, as proof the modeled range was wrong, when a single quarter is never enough data to confirm or kill a multi-year estimate.

How to use it live. Ask the split question before naming a number: "Is this dollar already missing from an invoice, or is it a claim about a year that didn't happen, because those two get defended with completely different evidence." That buys real thinking time, and it reframes the whole question before you have to guess at a figure.

The flagged cohort's claims development, tracked against the kill line
100% 50% 0% kill line: 80% developed 10% 35% 65% 88% Year 1 Year 2 Year 3 Year 4
Share of flagged cohort's claims fully developedKill line, 80% developed and confirmed
Prairiestone is at Year 2, about 35 percent developed. The avoided-claims range doesn't graduate to a reduction until it crosses the 80 percent line, somewhere between Year 3 and Year 4. Until then, it stays a range.

Three things worth stating directly, since this is where the real judgment sits. The alternative Northloom considered, and rejected, was building one "confidence-weighted savings" number, letting the model's own certainty automatically discount the avoided-claims figure into a single blended dollar. It lost, because collapsing it back into one number, even a smarter, discounted one, still hands finance a single point estimate they can book, the exact habit that caused the Q3 gap in the first place. The AI-specific failure worth naming is treating a calibrated probability as a verdict: Underlay's score says a business resembles others that filed claims at a known higher rate, it does not say this specific business would have filed a claim. The guardrail is refusing to let the avoided-claims figure move off "range, modeled" until the K step's matched-cohort evidence actually clears the bar. And the trade-off is real: the mandatory safety inspection on the flagged tier adds an average of 9 business days before Prairiestone can bind a policy, and Northloom accepted that delay on purpose, only for the 15 percent of applications where a wrong score is the expensive kind of wrong.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
PICK: commit to a position on a tradeoff, then show which of the two mistakes actually costs more, and to whom.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bogdana Onyekwere, who has run Finance at Prairiestone Mutual for six years, and booked a modeled number as a hard reduction without ever meaning to.
3 · THE POSITION
What's the P step here, in one line?
Tap to flip
ANSWER
The temp-staffing time savings is cost reduction, already missing from a real invoice. The flagged tier's prevented claims are cost avoidance, a modeled range, never booked as a hard cut.
4 · THE COST ASYMMETRY
Which mistake is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
Underselling a real reduction is cheap and visible: someone corrects it on the spot against an invoice. Booking a modeled avoidance as a reduction is hidden and expensive: it sits in a budget line for months until a reserve review finds the gap.
5 · THE OLD DECISION
What decision would Wojciech take back?
Tap to flip
ANSWER
Building the finance dashboard with one blended line, "Underlay Impact," instead of two tagged lines. It made sense at launch, when the whole number was small and mostly the realized temp-staffing piece.
6 · THE NUMBER
Fill in the blank: the realized reduction is $___ a month. The modeled avoidance range is $___ to $___ a year.
Tap to flip
ANSWER
$12,800 a month in realized temp-staffing cuts. $2.1 million to $3.4 million a year in modeled avoided claims, midpoint $2.75 million.
7 · THE REPLAY
Same Q3 review, new dashboard, what changes?
Tap to flip
ANSWER
The realized line, $153,600 a year, matches plan exactly. The avoidance range stays separate and unconfirmed, and nobody in the room is surprised it hasn't shown up in paid losses yet.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the modeled dollar there?
Tap to flip
ANSWER
Trestlewatch, Portrune Logistics' predictive-maintenance model. The modeled dollar is avoided unplanned downtime, a range that only a quiet quarter made impossible to reconcile.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Bogdana have to explain a gap at the Q3 reserve review?
  • A. Underlay's risk score got less accurate that quarter.
  • B. She had booked the modeled avoided-claims midpoint as a hard reduction, and workers' comp claims take years to develop, so this year's paid losses couldn't confirm it either way.
  • C. Prairiestone's temp-staffing invoice went up unexpectedly.
  • D. Northloom raised Underlay's contract price mid-year.
Show hint
Check what kind of evidence the avoided-claims number was actually built from, and how long a workers' comp claim takes to pay out.
Show answer
B. The model didn't get worse. Bogdana had treated a modeled range as a booked fact, and the long claims-development timeline meant this year's numbers were never going to confirm it either way.
Fill in the blank
2. Underlay scores an application in about ___ seconds. Manual review used to take about ___ minutes.
Show hint
It's stated where Underlay is first described, in Let's learn.
Show answer
20 seconds and 45 minutes. The scoring speed enabled the tiering. The manual-review time was the number the routine tier's savings got measured against.
True or false
3. True or false: the $12,800-a-month temp-staffing cut and the $2.1 million to $3.4 million avoided-claims range belong on the same dashboard line, since Underlay caused both.
  • True
  • False
Show hint
Ask what kind of evidence each number is backed by, and whether that evidence could ever contradict a single blended figure.
Show answer
False. One is realized cash, checkable against an invoice. The other is a modeled counterfactual with a confidence range. Blending them let Prairiestone book a range as if it were a fact.
Short answer, name the old decision
4. What old decision would Wojciech take back, and why did it make sense when Northloom first built the dashboard?
Show hint
Look at the key point box titled "The choice that mattered," right after the first chart.
Show answer
Model answer: Building the finance dashboard with one blended "Underlay Impact" line instead of two tagged lines. It made sense at launch, because the modeled avoided-claims piece didn't exist yet, so one line was the whole truth at the time.
Short answer, apply it yourself
5. Think of a tool at your own job, or a service you use, that claims to "save" you something. Name one part of that claim you could check against a real bill or timesheet, and one part that's really a guess about a bad thing that didn't happen.
Show hint
Ask whether the savings would still show up if you pulled last month's actual invoice.
Show answer
Model answer: A spam filter's "saves you an hour a week" claim is a guess, avoidance, about junk mail you never had to read. A cheaper mailbox storage plan you were able to downgrade to because of it is checkable against the actual bill, a reduction.
Short answer, work the number
6. If the flagged cohort's claims development reaches 60 percent by Year 3, per the kill line in this answer, should Prairiestone start calling the avoided-claims number a reduction?
Show hint
Check the K step in the framework recap, and the line chart's kill line.
Show answer
No, not yet. The kill line is 80 percent developed and confirmed against the model's prediction. 60 percent at Year 3 is progress, but the range stays a range until it actually crosses 80 percent, which this answer places somewhere between Year 3 and Year 4.
Before you close the answer
Why this works
Tests whether you treat a probabilistic model output as a probability when it talks to money, not just when it talks to engineering, and whether you'll split a savings claim by its evidence instead of handing finance one convenient number.
Follow-up traps
"Isn't splitting it just a way to dodge accountability for the bigger number?" Response: no, it's the opposite. A range with a kill line is a claim you can be held to; a blended point estimate is a claim nobody can actually check until it fails.

"What if finance just wants one number for the board deck?" Response: give them one headline number if they insist, but require the footnote splitting realized from modeled, so when someone asks the follow-up question in the room, there's an honest answer sitting right there instead of a scramble.
If pressed
The claims model behind the avoided-claims range isn't comparing this year's flagged applications to a guess. It's comparing them to a matched set of similarly-scored applications Prairiestone actually bound without inspection in prior years, so the range has a real control group behind it, not just a formula.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more