CaseIntermediateQuality, Cost & Token Economics / Measuring ROI and business impact / #1

How do you build the ROI case for an AI feature before it ships?

BOUND · sizing a no-show feature before a quarter of engineering gets spent on it

Every month, about 133 of Brackenfell Dental Group's booked appointments turned into empty chairs. Recallwise, Stonewick Health's scheduling and reminder tool, was already running there, sending the same reminder to every patient, 48 hours out, no matter how that patient usually behaved. Osareme Duskmere owns Recallwise's roadmap at Stonewick, and before asking for a full quarter of two engineers' time to build something smarter, she had to answer one question with real numbers: would it actually pay for itself.

The direct answer
Write the money equation out loud before naming a single number: what comes in, times how many practices pay, minus what it costs to build once and run forever. Price a low case and a high case with your own company's numbers, not just the pilot's. Then find the one assumption that swings the answer hardest. Here it is the voice-escalation rate buried inside the running cost, not the adoption guess everyone argues about in the room, and you go get the real version of that number before committing the whole quarter, not after.
Do this, in order
  1. Write the equation out loud before a single number gets attached to it.Why: an estimate nobody can see the shape of is a guess wearing a spreadsheet's clothes.
  2. Price a low case and a high case with the company's own numbers, not just the pilot's.Why: one confident total claims a certainty nobody in the room actually has.
  3. Find which single assumption swings the total hardest, and check that one first.Why: here it's the escalation rate, not the adoption guess, and it moves the answer almost twice as much.
  4. Never trust a small pilot's number over years of production data already sitting in a dashboard.Why: the real number was one query away, and nobody ran it before the case went to review.
  5. Set a real payback bar before you build, and hold the honest range up against it.Why: a bar you check after the quarter is spent isn't a bar, it's a postmortem.
  6. Don't ship it free to the whole install base just to skip a pricing conversation.Why: the running cost scales with every appointment on the platform, whether or not that practice ever needed the feature.

How to answer this, stage by stage

Nobody is grading whether Osareme can multiply. They're grading whether she'll hand over one confident total, or show which number inside it she trusts least, and go fix that one before spending real engineering time.

1
Scope it to one real feature, one real pilot, before the quarter gets spent
Say it like this
"Let's ground this in one real case. Recallwise is Stonewick Health's scheduling and reminder tool for dental practices. Osareme Duskmere owns its roadmap. Brackenfell Dental Group ran a six-week pilot of a new feature, Predictive Reminder Timing, at its flagship location, 1,400 appointments a month."
Why this works
An abstract "how do you build an ROI case" answer stays a slogan. One real feature with one real pilot behind it is something you can actually show the arithmetic for.
2
Say your structure out loud before touching a single number
Say it like this
"I'm going to run BOUND. Break the equation into its real terms, own where every number came from, give a low and a high range instead of one guess, run it against something I already know, then say which input would move the answer most."
Why this works
Signals a method already in motion, not five numbers arriving in whatever order they occur to you.
3
Break the equation down, out loud, before touching a single number
Say it like this
"Money in equals how many practices turn the feature on, times what they pay a month, times twelve. Money out equals what it costs to build once, plus what it costs to run forever after: every appointment, times what one reminder cycle costs, times twelve, times however many practices are using it."
Why this works
An estimate with no equation stated out loud is a guess wearing a spreadsheet's clothes. The interviewer needs to see the shape before they can trust a single figure inside it.
4
Own every number, and say plainly where each one came from
Say it like this
"340 practices run Recallwise today, that's our own account list. About 1,100 appointments a month each, from our own usage data. The premium price, sixty dollars a month, I picked because Brackenfell's flagship location recovered over five thousand dollars a month in the pilot, so sixty is an easy yes for them. The one number I'm least sure of is the escalation rate, how often a reschedule needs a live call instead of a text, and I only have a six-week pilot at one location behind that guess."
Why this works
Naming the weakest source out loud, before anyone else finds it, is what makes the rest of the numbers trustworthy.
5
Give a range, not one confident number
Say it like this
"At thirty-five percent adoption, this pays back in about thirty months. At fifty-five percent, about nineteen. I'm not going to stand here and tell you it's exactly one number, because it isn't, not yet."
Why this works
A single number this early implies a confidence nobody in the room, including the person saying it, actually has.
6
Run the sanity check against something the room already trusts
Say it like this
"The last feature I built like this, the waitlist auto-backfill, cost about a hundred eighty thousand dollars and paid back in fourteen months. This one costs less to build, ninety-five thousand, but even in the best case here, nineteen months, it's slower than the last one. That's the number that made me stop and recheck my inputs before I brought this to you."
Why this works
Comparing to a known result is what catches an estimate that looks fine on its own but is actually worse than it sounds.
7
Name what would move the answer most, and go get that number first
Say it like this
"It's not the adoption rate. Swinging that from thirty-five to fifty-five percent barely moves the payback. It's the escalation rate. I used the pilot's guess, ten percent, because that's all I had. Our own production dashboard, on the reminder feature we already ship, shows sixteen percent. At sixteen, this slides past our own two-year bar in every adoption case. That's the number I need before I ask you for the quarter."
Why this works
This is the hardest and most useful step in BOUND. Naming which single input would sink the case, and getting that number real before committing real engineering time, is what separates a serious estimate from a hopeful one.
8
Close on the decision, not the arithmetic
Say it like this
"So: I'm not asking for the quarter yet. I'm asking for two more weeks, at three more locations, about four thousand dollars, to measure this feature's own escalation rate directly instead of borrowing a number from a different one. If it lands close to ten percent, this clears our bar again, same as before. If it's still near sixteen, we either raise the price or start with the practices where it pays off fastest."
Why this works
Restates the actual decision in one breath, the thing an interviewer is grading, not a recap of the math that got you there.

Let's learn

What does it actually take to know, before you build it, whether a feature will pay for itself?

Recallwise is Stonewick Health's tool for dental practices. It sends appointment reminders, answers routine scheduling texts, and offers a new time when someone says they can't make it, without a front-desk person picking up the phone first.

Hand sketched numbered icon list titled Before Predictive Reminder Timing, one flagship month. Three rows: a document icon, every patient got the same reminder, 48 hours out. A gauge icon, 133 no-shows a month, out of 1,400 booked. A scale icon, the waitlist backfilled some slots, not most.
Before the new feature, Brackenfell's flagship location sent one reminder, at one time, to every patient, no matter how that patient usually behaved.

Before Predictive Reminder Timing, every one of Brackenfell's 1,400 monthly appointments got the same text, 48 hours out. No-show rate: 9.5 percent, about 133 empty chairs a month. An existing feature, a same-day waitlist, already backfilled a quarter of those slots. The rest just sat empty.

Predictive Reminder Timing learns, per patient, when they actually tend to read a message and reply, and offers a new slot automatically if they say they can't make it. Stonewick ran it as a pilot at Brackenfell's flagship location for six weeks. No-show rate fell to 6.6 percent, about 92 a month. Counting the same waitlist backfill, that's 31 real slots recovered a month, at roughly $172 a slot, Brackenfell's own blended rate across hygiene and restorative visits.

Brackenfell's own math was never in question. Over five thousand dollars a month, recovered, at one location. What was in question was whether Stonewick could make money building it for the other 339.

Here's the turn. A feature that recovers real money for a dental practice isn't automatically a good bet for the company that builds it. Osareme doesn't get to keep Brackenfell's five thousand dollars. She gets whatever price practices will pay for the feature, minus whatever it costs Stonewick to run it, forever, across every practice that turns it on. That's a completely different sum, built from completely different numbers, and most of them didn't exist yet.

Knowledge spark: what's a confidence threshold? Recallwise doesn't just decide "rebook" or "don't." It scores how sure it is that a new slot is a real match: right provider, right insurance, right visit length. Above the cut-off, it books the new time on its own. Below it, a person at the front desk looks first.
Hand sketched left to right flow diagram titled What happens after a patient says can't make it. Five boxes connected by arrows: Reminder sent, Patient replies, Confidence checked, this box emphasized to mark the routing decision, Auto-rebook, Or flag desk.
One step in the middle, the confidence check, decides whether a reschedule costs a text message or a live phone call.

So Osareme wrote the equation out loud, the way BOUND asks for, before trusting a single figure inside it.

Hand sketched labeled parts diagram titled What is inside the money equation. A central gauge icon labeled Annual value, with six labeled callouts radiating around it: Practices that pay, Monthly fee, Appointments run, Cost per cycle, Cost to build once, Escalation rate.
Six numbers feed one total. Only one of them turned out to be the risky one.

Money in: how many of Stonewick's 340 practices turn the feature on, times a $60 monthly add-on fee, times twelve. The price wasn't a guess. Brackenfell's own recovered revenue was over eighty times that fee, so $60 was always going to be an easy yes for a practice that had seen the pilot's numbers.

Money out: how many appointments run through the model across the whole install base, average 1,100 a month per practice, times what one reminder cycle actually costs to run, times twelve. Built up from the ground: a base text message, $0.008, already a known cost. A small model call to decide the best time and channel, $0.004. And an escalation cost, a live AI voice call, for the reschedules a text alone can't close or the model isn't confident enough to book on its own, at $0.18 a call.

That last piece is where the case actually lived or died, and Osareme's first pass used the only number she had: the pilot's own escalation rate, 10 percent of appointments. That gave a running cost of about three cents an appointment, and a build cost of $95,000, two engineers for one quarter.

Naive case: annual net value by adoption rate (escalation rate assumed at 10%, from the pilot)
$70,000 $35,000 0 $38,556 Low, 35% adopt $49,572 Mid, 45% adopt $60,588 High, 55% adopt
Low adoptionMid adoptionHigh adoption
Across the whole adoption range, the naive case clears its $95,000 build cost in 19 to 30 months. It looked like a case worth bringing to roadmap review.

What that costs at its worst: if Stonewick greenlights the full quarter and ships this to the whole install base on a running-cost number that's actually wrong, the mistake doesn't show up on day one. It shows up a year later, buried inside a blended margin line, the same way a wrong pricing assumption hid inside Nettlebright's account for months in a different case. By then the quarter is already spent and the feature is already live everywhere.

The choice Osareme would take back Trusting the six-week pilot's escalation-rate guess, 10 percent, as the running-cost input for the whole install base, instead of first pulling the escalation rate Stonewick already had, years of it, sitting in the production dashboard for the base reminder feature Recallwise has shipped since day one.

What I would leave alone: the no-show reduction itself, 9.5 percent down to 6.6, and the base text-message cost, $0.008. Both held steady through everything that came next. The risk was never in the value side of the equation. It was hiding in one line of the cost side.

The lesson: an estimate is only as trustworthy as its shakiest number. Find the shakiest one before you trust the total, not after the quarter is already spent.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a case that sailed through review almost got built on a number nobody had actually checked.

Osareme Duskmere has never once brought a build request to roadmap review without a spreadsheet behind it. Three years into owning Recallwise's roadmap, she still builds the numbers by hand before trusting anyone else's summary of them, her own included.

The pilot at Brackenfell's flagship location ran clean. Six weeks, one location, no-show rate down from 9.5 to 6.6 percent, real dollars recovered, Bertram Aldercress, Brackenfell's COO, telling her outright he'd pay double the fee she eventually asked for. Osareme built her first pass of the case the same week the pilot closed. Adoption between 35 and 55 percent, a fee that undercut the value ten times over, a build cost of $95,000. The payback came out between nineteen and thirty months. She sent it to the roadmap deck.

It looked airtight in three quiet beats. Beat one: the first pass used the pilot's own escalation rate, 10 percent, because it was the only number that existed for a feature that hadn't shipped yet, and nobody in the review room asked where it came from. Beat two: a colleague flagged, almost as an aside, that six weeks at one location was a small sample to hang a company-wide number on. Osareme noted it, agreed in principle, and moved on, because the direction of the whole case was so clearly positive that the caveat felt like a formality. Beat three: the case sailed through the roadmap review with barely a follow-up question. Everyone liked the story. Nobody asked to see the escalation number's own history.

Hand sketched horizontal timeline titled The quarter the case almost shipped on the wrong number. Five milestones: Pilot ends, caption 6 weeks, 1 location. First case built, caption looks like 23 months. Roadmap review, caption case sails through. Real number found, this milestone emphasized in red, caption sitting in a dashboard. Case rebuilt, caption 42 months, not 23.
No single bad decision. A number nobody checked, sitting quietly inside a case everyone liked.

The trigger wasn't a dashboard alarm or an angry email. It was a data engineer, two days before the build was meant to start, mentioning almost in passing that the base reminder feature's own escalation rate, the one already running in production across all 340 practices, had been sitting at 16 percent for six months. Not 10. Osareme had never pulled it, because the case wasn't about the base feature. It just hadn't occurred to her that a close cousin's real number was one query away the whole time.

Hand sketched two panel comparison titled Two sources for one number. Left panel, a person icon labeled Six-week pilot, captioned one location, a guess dressed as a rate. Right panel, a gauge icon labeled Production dashboard, captioned years of data, already sitting there.
One of these numbers had six weeks behind it. The other had six months, and Osareme had never asked it a question.

She rebuilt the case that afternoon with the real number. Running cost per appointment moved from three cents to about four. Not a dramatic jump on paper. But at every adoption level, the payback slid past two years, and at the low end, it moved out past four.

Cumulative net value vs. the $95,000 build cost, at mid adoption, naive guess vs. real production rate
$150k $75k 0 build cost, $95,000 24-month bar naive: crosses ~23mo corrected: crosses ~42mo 0 12mo 24mo 36mo
Naive: 10% escalation guessCorrected: 16% real rateBuild cost line
Same feature, same adoption, one number changed. The naive line clears the build cost well inside the two-year bar. The corrected line doesn't clear it until well past three and a half years.
We almost spent a quarter of engineering time on a number that was never actually tested. We tested the direction. We never tested the part that could sink it.

Run it again, with the real number in hand from the start. Osareme doesn't take the case back to review as a "build it" or "don't." She takes it back with a third option: extend the pilot two more weeks, three more locations, about $4,000, and measure Predictive Reminder Timing's own escalation rate directly, instead of borrowing the base feature's number one more time. If the real rate lands close to 10 percent, the original case holds and the quarter gets approved next sprint. If it lands near 16, the fee moves to somewhere around $75 to $80, or the first rollout goes only to the practices with the worst no-show rates today, where the value per practice is highest and the model has the least room to guess wrong.

One version of this case would have committed a full quarter of engineering time off a number that had six weeks of evidence behind it, sitting three rows down in a spreadsheet from a number with six months. The other spent $4,000 and two weeks finding out which one was real before the bigger bet got made.

What Osareme would tell herself, back in that first roadmap review: the case didn't fail because the estimate was wrong. It almost failed because nobody, herself included, had asked out loud which number in it was the one doing all the work.

BOUND, or how not to greenlight a quarter off a borrowed number

Not a way to dress up a spreadsheet in five letters. BOUND is what stops an estimate from being a confident total with no visible seams, by forcing every number in it to say where it came from and how much it's actually trusted.

BBreak it down. State the equation out loud before touching a number.
Money in: practices that turn it on, times $60 a month, times twelve. Money out: the one-time build, $95,000, plus appointments times the cost of one reminder cycle, times twelve, times however many practices are running it, forever after.
An answer with no visible equation is a guess with a confident tone. State the shape first, so every number after it has somewhere to sit.
OOwn numbers. Say where each one actually came from.
340 practices and 1,100 appointments a month, both from Stonewick's own usage data. The $60 fee, anchored against Brackenfell's own recovered revenue. Adoption, 35 to 55 percent, anchored on the waitlist auto-backfill feature's real 41 percent year-one attach rate. The escalation rate, 10 percent, from a six-week pilot at one location, the weakest source of the five.
Naming the weakest source out loud, before anyone else finds it, is what makes the rest of the case trustworthy.
UUse a range. A low case and a high case, not one number.
Naive case: 19 to 30 months payback across the adoption range. Corrected case, real escalation rate: 34 to 54 months. The range itself is the honest answer, not a rounding of it to one figure.
A single number this early claims a confidence nobody in the room actually has.
NNail the sanity check. Compare the answer to something already known.
The last comparable feature, the waitlist auto-backfill, cost $180,000 and paid back in 14 months. This feature costs less to build, $95,000, but even its best naive case, 19 months, is slower than that. Against the fleet's own known history, the case looked weaker than its own headline number suggested.
This is the step that catches an estimate that reads fine in isolation but is actually worse than the last real result it should be compared to.
DDirection. Which single assumption moves the answer most.
Not adoption. Swinging adoption across its whole realistic range moves net value by about $12,000 a year. Swinging the escalation rate from the pilot's 10 percent guess to the real 16 percent moves it by about $22,000, nearly twice as much, and flips a case that clears the bar into one that doesn't.
Naming which single cell would sink the whole case, and getting that number real before committing the build, is what a good estimator does that a confident guess never bothers to.
Hand sketched decision tree titled When Recallwise books it, and when it asks first. Root box: Patient says can't make it. Three branches: confidence 90% or higher leads to Auto-rebook the slot. Confidence under 90% leads to Flag the front desk. No reply after two tries leads to Flag the front desk.
The same threshold that protects a patient from a wrong booking is also the lever that decides how expensive the feature is to run.

Three things worth stating directly, since this is where the real judgment sits. Stonewick considered, and rejected, bundling Predictive Reminder Timing into the base subscription for free, shipping it to all 340 practices at once with no opt-in and no fee. It lost, because the running cost scales with every appointment on the platform whether or not that practice had a no-show problem worth solving, guaranteeing the case could never recoup its build cost inside any reasonable window. The AI-specific failure worth naming is a confidently wrong auto-reschedule: the model books a new slot with the wrong provider, the wrong insurance coverage, or the wrong visit length, and the patient only finds out at the counter. The guardrail is the 90 percent confidence threshold itself, the bar the model has to clear before it acts alone, most of the time, by design, not a promise that it's always right. And the trade-off is real and accepted on purpose: keeping that threshold at 90 percent, instead of lower, means more reschedules escalate to a live call, which is exactly what makes the running cost higher than the naive case assumed. Stonewick chose that cost on purpose, because a wrongly auto-booked appointment costs more in a frustrated patient and a support ticket than the extra escalation calls ever save.

And if you want to be sure it really works, try it somewhere else

Same five letters, an industrial laundry instead of a dental office, and this time the swing factor isn't a dollar figure at all. It's whether the model's own catch rate can be trusted off the number of real events it's actually seen.

Driftguard is Grishwold Systems' tool for commercial laundries: it reads vibration data off a washer-extractor's bearings and flags a likely failure five to ten days before it happens, instead of the machine going down mid-shift with no warning. Copperfen Linen Services runs 24 of these machines across three plants, and had been losing about $84,000 a year to unplanned bearing failures, six a year, fleet-wide, at roughly $14,000 each in lost throughput, expedited parts, and overtime to catch up.

The decision Grishwold's team would take back Grishwold's first estimate treated the pilot's two-for-two catch rate as a real number to build a business case around, when two events is barely enough to know the model works at all, let alone how often.

Driftguard ran on six of Copperfen's machines for four months. It flagged three real events; two were confirmed bearing wear caught in time, one was a false alarm. Grishwold's engineer, sizing the build case, used an industry-benchmark catch rate, 70 to 85 percent, rather than the pilot's own small-sample 100 percent, because two real events aren't enough to trust a rate at all.

Hand sketched quadrant diagram titled Which machines actually needed watching. X axis how often it runs, from light duty to heavy duty. Y axis how costly one failure is, from cheap fix to a lost shift. Flatwork ironer plotted low on both axes. Folding line plotted low on cost, mid on use. Washer-extractor plotted high on both axes, top right.
Not every machine in the fleet carries the same risk. The washer-extractors, run hardest and costliest to lose, are where the model's judgment actually pays for itself.

At a 70 percent catch rate: 4.2 failures caught a year, worth $58,800, against roughly $756 in false-alarm inspections and $3,600 a year to run the sensors and the model. Net, about $54,444 a year, against a one-time build and sensor cost of $20,400. Payback: about four and a half months. At 85 percent: payback closer to three and a half months. Both fast, both a clear yes.

Same method, different lever: at Brackenfell, the number that swung the whole case was a dollar figure, the escalation rate. At Copperfen, the dollar swing barely matters, four and a half months versus three and a half either way clears any reasonable bar. What actually matters here is whether the catch rate can be trusted at all off two real failures. A model's own precision and recall aren't safe to trust from a handful of events, no matter how good they look, and that's the number Grishwold would go get more of before scaling to the rest of the fleet, not the dollar total.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the equation: failures caught, times cost avoided, minus false alarms and running cost, minus the one-time build.
Cost: no budget this quarter to extend either pilot. Pull whatever adjacent production number already exists, the way the base reminder feature's real escalation rate was already sitting in a dashboard, before assuming a small pilot's guess is the only number available.
The model got better, for real: say Driftguard's catch rate turns out to be 95 percent instead of 70 to 85. The direction step barely moves, because the swing was never really about the catch rate's size, it was about whether two events were enough to trust any number at all.

Where people run it wrong.
They let a small pilot's number stand in for a rate that needs a much bigger sample before anyone should trust it.
They lead with one confident total instead of a range, and let the room assume more precision than the estimate actually earned.
They skip the sanity check against the last comparable feature, and miss that a case is quietly worse than the company's own track record.

How to use it live. Ask yourself, out loud if you have to: "which number in this equation have I actually tested, and which one am I just hoping holds?" That question alone is usually the difference between an estimate and a guess with a spreadsheet attached.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "build the ROI case before it ships" question, and why not FLIPS?
Tap to flip
ANSWER
BOUND: show the arithmetic, own the assumptions, use a range, sanity-check it, name what swings it most. FLIPS is built for a behavior that snaps between two settings, and sizing a feature's payback isn't that shape.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Osareme Duskmere, Senior Product Manager at Stonewick Health, who owns Recallwise's roadmap and has to defend a full quarter of engineering time before it gets spent.
3 · THE EQUATION
What's the B step's equation here, in plain words?
Tap to flip
ANSWER
Money in equals how many practices pay, times the monthly fee, times twelve. Money out equals the one-time build cost, plus appointments times the cost of one reminder cycle, times twelve, forever after.
4 · THE SWING
Which single assumption swings this estimate hardest, and why?
Tap to flip
ANSWER
The escalation rate, how often a reschedule needs a live call. Moving it from the pilot's 10 percent guess to the real 16 percent production number nearly halves the annual net value, a bigger swing than the whole adoption-rate range.
5 · THE OLD DECISION
What decision would Osareme take back?
Tap to flip
ANSWER
Trusting the six-week pilot's escalation-rate guess over the years of production data Stonewick already had on the base reminder feature's own escalation rate, sitting in a dashboard nobody had queried for this case.
6 · THE NUMBER
Fill in the blank: the naive case pays back in about ___ months at high adoption. The corrected case pays back in about ___ months at the same adoption.
Tap to flip
ANSWER
About 19 months naive. About 34 months corrected, using the real 16 percent escalation rate instead of the pilot's 10 percent guess.
7 · THE REPLAY
Same case, real escalation number, what actually changes?
Tap to flip
ANSWER
Osareme doesn't ask for the quarter yet. She asks for two more weeks and about $4,000 to measure the new feature's own escalation rate directly, with a clear rule for what happens at each possible result.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what swings that estimate most?
Tap to flip
ANSWER
Driftguard, Grishwold Systems' predictive-maintenance tool for Copperfen Linen Services. What swings it there isn't a dollar figure, it's whether a catch rate built on just two real failure events can be trusted at all.

Check yourself Score: 0 / 0

Multiple choice
1. Why did the payback estimate shift from about 19 to 30 months (naive) to about 34 to 54 months (corrected)?
  • A. The no-show reduction from the pilot turned out to be smaller than reported.
  • B. The escalation rate used in the running cost was a six-week guess, not the real production number.
  • C. Stonewick raised the build cost estimate from $95,000 to a higher figure.
  • D. Brackenfell Dental Group asked for a lower monthly fee.
Show hint
Check the D step in the framework recap, and the line chart in the story section.
Show answer
B. The pilot's 10 percent escalation guess was replaced with the real production rate, 16 percent, which raised the running cost and pushed payback well past the two-year bar.
Fill in the blank
2. The pilot showed Brackenfell's no-show rate fall from 9.5% to ___%, over six weeks.
Show hint
It's stated early in Let's learn, right after the before-state diagram.
Show answer
6.6%. A drop of 2.9 points, worth 31 recovered slots a month at Brackenfell's flagship location alone.
True or false
3. True or false: the no-show reduction number is the one Osareme would recheck before trusting the case again.
  • True
  • False
Show hint
Check "what I would leave alone" in Let's learn.
Show answer
False. The no-show number held steady through everything that followed. It was the escalation-rate assumption inside the running cost that was wrong, not the value side of the equation.
Short answer, name the old decision
4. What old decision would Osareme take back, and why did it make sense at first?
Show hint
Look at the key point box titled "The choice Osareme would take back," right after the first chart.
Show answer
Model answer: Trusting the six-week pilot's escalation-rate guess instead of pulling the years of production data Stonewick already had on the base reminder feature's own escalation rate. It made sense at first because the pilot was the only number that existed for this specific new feature, and the production number belonged to a different, if related, feature.
Short answer, apply it yourself
5. Think of a feature or tool you use that was probably justified with an ROI estimate before it shipped. What's one number in that estimate you'd want measured directly instead of borrowed from something similar?
Show hint
Think about which number in the pitch sounds confident but was probably copied from a similar, not identical, feature.
Show answer
Model answer: A grocery delivery app's "estimated arrival time" feature was probably first sized using how often drivers run late on a restaurant-delivery estimate, borrowed rather than measured for grocery orders specifically, which take longer to pack and load. That's the number worth measuring directly before trusting the estimate.
Short answer, work the number
6. If the real escalation rate had matched the pilot's original 10% guess instead of the corrected 16%, would the case clear a 24-month payback bar at 45% adoption? Show the check.
Show hint
Compare the naive case's mid-adoption bar in the first chart against the 24-month line in the second one.
Show answer
Yes, barely. At 45% adoption, the naive case's net value is $49,572 a year against a $95,000 build cost, a payback of about 23 months, just inside the 24-month bar. That's exactly why the corrected 16% rate, which pushes it to about 42 months, was the number worth catching before the build started.
Before you close the answer
Why this works
Tests whether you'll write the actual equation and go looking for the weakest number inside it, or lead with one confident total that quietly assumes every input is equally solid.
Follow-up traps
"Isn't a three to four year payback still fine for a platform feature, why not just build it?" Response: a slower payback than the last comparable feature, which cleared in 14 months, is exactly the signal that the estimate, not necessarily the feature, needs more checking before the quarter gets spent.

"Why not just skip the estimate and build a cheap version to find out?" Response: a two-week, four-thousand-dollar extension of the existing pilot answers the one open question directly, for a fraction of what a full quarter costs to find out the hard way.
If pressed
The escalation rate isn't one number system-wide. It splits sharply by whether a patient has ever replied to a text before: patients with zero reply history escalate to a voice call at closer to 40 percent, patients with a reply history at closer to 9. The real fix might be routing by that split, not tuning one blended threshold for everyone.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more