CalculationAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #8

How do you account for the cost of maintaining an AI feature in its ROI?

BOUND · recipe and nutrition planning for a meal-kit subscription

Provender reads a Bramblewick Kitchen subscriber's allergies, macros, and goals, then writes three personalized recipe-and-nutrition-plan options into that week's box. Idrissa Frobisher owns what Provender actually costs to run, a number that had said the same thing, 452% ROI, on every renewal deck for three straight years while the subscriber base grew forty times over. Then Bramblewick's CFO, Stavros Delacorte, asked a question the deck had never been built to answer: whose number is this, this year's, or launch day's?

The direct answer
Count five ongoing costs into Provender's ROI, not just the one-time build: inference at real volume, eval and monitoring labor, periodic retraining against new dietary rules, incident response for model regressions, and data-pipeline upkeep. Then re-run the number every year, against that year's real usage, not the figure computed at launch. Do that here and the frozen 452% becomes a real, still-healthy 136%, and that's the number that belongs in front of a CFO.
Do this, in order
  1. Add the five ongoing cost lines to the ROI, every year, not once at launch.Why: a build cost divided into value forever is what made 452% look real when this year's honest number is 136%.
  2. Re-run the number on this year's real usage, not launch day's.Why: usage-driven costs grow with the exact thing that makes leadership love the feature, so a frozen formula gets more wrong every year nobody touches it.
  3. Give a range on the lines that are genuinely uncertain, not one clean figure.Why: retraining and incident response can swing two to three times over based on one real event, so a single number claims confidence nobody has.
  4. Watch retraining cadence as the line most likely to actually double, not just the biggest one.Why: inference moves the ROI the most in raw dollars, but it's the most forecastable; one regulation change can double retraining overnight.
  5. Sample eval coverage instead of reviewing every plan Provender writes.Why: full review would cost more than the feature is worth; a calibrated sample against a golden set still catches drift.
  6. Don't freeze the model to dodge the retraining line.Why: dietary guidance and allergen law keep moving even when the code doesn't, so a frozen model drifts out of safety quietly instead of costing a line item.

How to answer this, stage by stage

Nobody is grading whether you can name a cost per API call. They're grading whether you know a number that never changes is a warning sign, not proof nothing's wrong, and whether you can rebuild it honestly with real arithmetic.

1
Anchor it to one real feature and one real number owner
Say it like this
"Let's ground this in one feature. Provender is Bramblewick Kitchen's recipe and nutrition planner. It reads a subscriber's allergies, macros, and goals, and writes three personalized options into that week's box. Idrissa Frobisher owns what it actually costs to run, and that's the number I want to build."
Why this works
An abstract "how do you account for maintenance cost" answer stays a slogan. One real feature keeps every number checkable.
2
Name the two sheets before doing any math
Say it like this
"There are two ROI numbers sitting in this company right now. The one on the renewal deck, 452%, computed once at launch. And the one nobody's built yet, this year's value against this year's real cost. I'm going to build the second one."
Why this works
Says what's actually wrong, a stale number, not "AI is expensive," before a single figure appears.
3
Lay the equation out loud, term by term
Say it like this
"Provender's real annual cost is five things added together: inference for every plan it writes, the labor to review and monitor its output, periodic retraining against new dietary and allergen rules, on-call time when a bad version ships, and keeping the nutrition and ingredient data current. The build cost doesn't belong in this sum, that's paid once. This is what gets paid every single year after."
Why this works
Naming the terms before any figure stops "the model's basically free" from quietly standing in for real arithmetic.
4
Give the decision, in one breath
Say it like this
"Here's what I'd do. Count all five lines into the ROI, every year, against that year's real usage. For Bramblewick, that turns the frozen 452% into a real 136%. Still a good return, just an honest one."
Why this works
This is the direct answer, said before any story about how the number went stale.
5
Own each line, with where the number came from
Say it like this
"Inference: 18.72 million plans a year at just over a cent each, call it $262,000, that one's solid, it's a rate card times a known volume. Eval and monitoring: two reviewers spending part of their week checking output against a golden set, about $57,000. Retraining: a quarterly refresh cycle, about $75,000 a year. On-call: a handful of incidents a year, about $18,000. Data pipeline: a quarter of a data engineer, about $32,500."
Why this works
Shows the estimate is built from parts a listener could check themselves, not one vague "maintenance is expensive" line.
6
Range the lines that actually swing
Say it like this
"Inference I'll treat as a point number, it's driven by subscriber count, which the business forecasts anyway. Retraining, eval labor, and incident response are the ones I'd range: $45,000 to $110,000 for retraining, depending on whether a guideline change forces an off-cycle rebuild, $45,000 to $70,000 for eval, $15,000 to $60,000 for incidents. All in, the honest range on the whole maintenance line is about $392,000 to $542,000 a year, not one clean figure."
Why this works
A single number here would claim more certainty than three of the five lines actually have.
7
Run it against something real, and say what got turned down
Say it like this
"Even the fully-loaded number, about two and a half cents a plan, is roughly eight hundred times cheaper than paying a dietitian to build that plan by hand, about nineteen dollars each. So this was never a reason to stop running Provender. We did talk about freezing the model to dodge the retraining line entirely. We turned it down, because dietary guidance and allergen labeling keep changing whether or not the code does, and a frozen model just drifts out of compliance quietly instead of costing us a line item."
Why this works
Naming the sanity check and the rejected alternative together is what makes this a judgment call, not a spreadsheet exercise.
8
Name the line that actually bites, then close on one line
Say it like this
"If I had to name the single line that changes this picture most if it doubled, it's inference, it's the biggest dollar amount by far. But it's also the one I trust most to stay forecastable. The one I'd actually lose sleep over is retraining, because it's cheap today and one regulatory change away from doubling overnight. So: count all five lines, re-run the number yearly, range the ones that swing, and watch retraining, not the biggest line, for the surprise."
Why this works
Closing on the decision and the thing to actually watch is what makes this sound rehearsed, not like a story that trailed off.

Let's learn

For three years, Idrissa Frobisher's renewal slide said the same thing: 452% ROI. Nobody had thought to ask if that number still meant anything.

Provender is Bramblewick Kitchen's recipe and nutrition planner. A subscriber tells it what they're allergic to and what they're eating toward, keto, diabetic-friendly, high-protein, and each week it writes three personalized recipe-and-nutrition-plan options into their box.

Before Provender, six recipe developers and dietitians built eight fixed weekly menus by hand. A subscriber picked one of the eight, filtered only by a coarse tag, vegetarian or gluten-free, nothing more specific than that. A brand-new menu variant took about three weeks of a dietitian's time to build and test.

Hand sketched icon list titled Bramblewick's kitchen desk before Provender. Three rows: a person icon, six developers hand-built eight fixed weekly menus. A document icon, one coarse filter tag, vegetarian or gluten-free only. A box icon, a new menu variant took about three weeks.
Before Provender, personalizing a plan to one subscriber's real allergies and goals wasn't something the kitchen desk could do at all, let alone every week.

With Provender, a subscriber gets three options built just for them, in about four seconds, at roughly a cent and a half each to run. Bramblewick has 120,000 active subscribers now, each getting three plans a week. That's 18.72 million plans a year. At $0.014 a plan, inference alone runs about $262,000 a year.

Knowledge spark: what's actually in one Provender plan? Provender reads a subscriber's allergy list, macro targets, and that week's available ingredients, then writes a recipe, a nutrition breakdown, and plain-language cooking steps. Every part of that costs money to run, every single time, whether the subscriber ever opens the box or not.
Hand sketched labeled parts diagram titled What Provender really costs a year. A central gauge icon labeled Provender fully loaded, with five labeled callouts around it: inference every plan generated, eval and monitoring labor, retraining and prompt-tuning, on-call for model regressions, and data pipeline upkeep.
Five real lines feed the true cost. The launch-day sheet only ever knew about the first one, and only barely, at pilot scale.

Here's the turn. Provender's original pitch didn't get the inference number wrong. It just never added anything else. The ROI sheet built at launch divided a year of value, $1,050,000, made up of about $640,000 in subscriptions that would have churned without the personalization and $410,000 in recipe-development headcount Bramblewick didn't have to add, by one number, the $190,000 it cost to build Provender in the first place, back when it served one diet program and 3,000 pilot subscribers. That was a fair sheet the week it was written. There was almost nothing else to count yet: no retraining had happened, no incident had happened, monitoring was one person glancing at a dozen plans most Fridays.

The sheet wasn't wrong. It just never grew up.
Provender's real annual cost, five lines stacked, against what the naive sheet ever counted
$500k $250k 0 $0 counted Naive sheet (build cost only) $444,580 total Fully-loaded sheet (five lines)
InferenceEval & monitoringRetrainingOn-callData pipeline
The naive sheet's cost line never moved after launch. The real one is $444,580 a year, made of five parts, only one of which even existed yet at pilot scale.

What it costs at its worst: three years later, Bramblewick runs four diet programs and 120,000 subscribers, and Provender genuinely does need retraining, monitoring, on-call, and a maintained data pipeline, real costs that never made it onto that sheet. One ordinary Tuesday, a routine model update quietly started favoring a flavor-boost instruction over the hard allergen-exclusion rule, and a handful of nut-free-tagged recipes picked up a walnut-oil finishing drizzle, an ingredient the rule-based filter never caught, because walnut oil isn't tagged as a whole nut in the ingredient database, and the drizzle only ever showed up in the free-text finishing touches the filter never scanned. Branwell Gower, on the eval team, happened to have that exact recipe type in her weekly six-percent sample. She caught it Thursday morning. Nothing shipped. Nobody outside the eval team ever knew how close it came.

Hand sketched comparison diagram titled Two cost lines, same feature. Left panel, a document icon labeled the pitch-deck sheet, caption value minus the one build cost, about 452 percent, frozen at launch. Right panel, a scale icon labeled the fully-loaded sheet, caption value minus everything it takes to keep Provender safe and current, about 136 percent.
Two sheets, same feature, same year's value. Only one of them ever asked what it costs to keep Provender honest.
ROI, naive vs. fully-loaded, as Provender scaled
500% 250% 0% 452%, frozen 228% 165% 145% 136%, real Pilot 3,000 subs Year 1 40,000 subs Year 2 80,000 subs Now 120,000 subs
Naive ROI, never recomputedFully-loaded ROI, real cost each year
The naive line never moved because nobody recomputed it. The real return fell as scale arrived, then settled around 136%, still a strong number, just not the one on the deck.
The choice that mattered The ROI sheet's cost cell was a single hard-coded number, $190,000, set the week Provender launched. That was the sensible call for a one-diet-program pilot with 3,000 subscribers, there was barely anything ongoing yet worth adding. Nobody ever came back and gave that cell a reason to grow as the product did.

What I would leave alone: the internal test-kitchen tool, a small assistant the two remaining recipe developers use to search for ingredient ideas, about 50 uses a week, no subscriber ever sees its output directly. Its naive ROI, just inference cost, a few dollars a month, is basically the honest number, because it never scaled past a handful of people and it never touches a customer-facing safety claim.

The lesson: a yearly ROI number isn't a fact you compute once and frame. It's a habit you keep, because the exact growth that makes a feature look good on a renewal deck is the growth that makes the ongoing costs real.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a number that never moves is itself the warning sign.

Idrissa Frobisher has owned Provender's numbers since the pilot, back when it served one diet program and about 3,000 subscribers who'd signed up specifically to test it. Idrissa kept the launch pitch deck bookmarked, mostly out of habit: value minus one build cost, $1,050,000 minus $190,000, divided by $190,000. 452%. It was a genuinely honest number, the week it was worked out.

For a year, nothing about that math needed to change. Provender served one diet program. Retraining had never happened, there was nothing yet worth retraining against. Monitoring was Idrissa, personally, glancing at a dozen plans most Fridays. Bramblewick's board loved the slide. So did every renewal deck after it.

It thinned in three quiet beats, and none of them looked like a mistake. Bramblewick launched a keto program, then diabetic-friendly, then a fourth built around GLP-1 users' specific nutrition needs, each one adding real, ongoing work: new dietary rules to encode, new golden-set cases to build, more plans written every single week as the subscriber base climbed past 40,000, then 80,000, then 120,000. Somewhere in there, a real reviewer role got created, a real quarterly retraining cycle got built, a real on-call rotation started catching bad versions before they shipped. Every one of those was the right call. None of them ever made it back into that one cost cell on the renewal slide, because nobody had ever been asked to open the slide and check.

Then came the Tuesday that almost changed that on its own.

A routine model update shipped, meant to make Provender's recipe write-ups read a little more appetizing, a small nudge toward warmer, more flavorful language. What nobody caught in testing: that nudge occasionally out-competed the hard allergen-exclusion rule sitting underneath it, in the free-text finishing touches Provender wrote for a recipe, not the structured ingredient list the safety filter actually scanned. For a handful of nut-free-tagged boxes, that meant a walnut-oil drizzle, suggested as a garnish, in a recipe meant to contain no tree nuts at all.

Hand sketched timeline titled The three years nobody re-opened the cost sheet. Four milestones: pilot launch, 1 diet program 3,000 subscribers. Scaled up, 4 programs 120,000 subscribers. The near miss, walnut oil in a nut-free box, this milestone emphasized. Budget review, Stavros asks for the real number.
No single bad year. A number that held for three of them, until an ordinary Tuesday and an ordinary budget review both landed close together.

Branwell Gower reviews six percent of every week's plans against a golden set, mostly a spot check, mostly uneventful. That week, by ordinary luck, three of the flagged recipes happened to be nut-free boxes carrying the new garnish line. She caught it Thursday morning. Nothing shipped. Nobody outside the eval team ever knew how close it came.

Idrissa sat with that for a few days. Not because Provender had failed, it hadn't, the review caught exactly what it was built to catch. What sat wrong was smaller: if that six-percent sample had landed on a different set of recipes that week, the walnut oil would have shipped, and there would have been no line item anywhere that explained why the company had only a six-percent chance of catching it, because the ROI sheet had never once said reviewing Provender's output cost anything at all.

Three weeks later, in the ordinary course of an annual budget review, Stavros Delacorte, Bramblewick's CFO, pulled the last three years of renewal decks side by side. He wasn't looking for a reason to cut Provender. He was doing what a CFO does: noticing that one number, 452%, appeared identically on all three, while the subscriber count behind it had grown forty times over.

"Whose number is this," he asked Idrissa, in the flat tone that isn't an accusation yet. "This year's, or the year we launched?"

We didn't lose the number. We lost the habit of asking whether it was still true.

Idrissa didn't have an honest answer. That was the real problem, not the number itself.

Hand sketched full page metaphor scene titled Free on the sticker was never free on the ledger. Left panel, a document icon labeled the sticker, caption 452 percent, computed once at launch, never revisited. Right panel, a scale icon labeled the ledger, caption the bill that keeps arriving, every single year since.
The whole answer to this question in one picture. A number printed once was never the same thing as a bill that keeps arriving.

What Idrissa built over the following week was the sheet that should have existed the whole time: five real ongoing lines, inference at 120,000 subscribers' actual usage, eval and monitoring labor, quarterly retraining, on-call, data-pipeline upkeep, adding up to about $444,580 a year, with honest ranges on the three lines that genuinely swing. Run against this year's real value, $1,050,000, that's a 136% return. Not 452%. A real 136%, worked out, defensible, and still comfortably worth running, about eight hundred times cheaper per plan than paying a dietitian to write it by hand.

The decision that opened the door traced back to the very first pricing meeting, the week Provender launched. Someone asked whether the cost cell should be a formula that recalculated each year or a fixed figure. A fixed figure was faster to ship, and at the time, genuinely correct, there wasn't yet anything ongoing worth modeling. Nobody wrote down a date to come back and check.

Run that meeting again, with one line added: review the cost cell every year, against that year's real usage, not just the arithmetic it started with. Same forty times growth, same near miss, same CFO's question. This time Idrissa's answer is immediate: "This year's. Here it is." The number is smaller than the one on the wall. It's also true.

One design let "the ROI" mean whatever number happened to be sitting on the slide. The other lets it mean whatever the year actually cost.

What Idrissa would tell their past self, back in that first pricing meeting: a number that never changes isn't a sign nothing changed. It's a sign nobody was asked to look.

BOUND, or the five lines a launch-day sheet never grew

Not a story wearing a framework's clothes. This is an estimation problem with a stale number sitting on top of a real, moving one, and BOUND is what turns "AI upkeep costs something" into a figure a CFO can actually check.

BBreak it down. What's the actual equation?
Provender's real annual cost is five lines added together: inference for every plan written, eval and monitoring labor, periodic retraining against new dietary and allergen rules, on-call time for model regressions, and data-pipeline upkeep. None of that is the one-time build cost, that's paid once. This is what gets paid every year after, whether anyone's watching or not.
Say the five lines before naming a figure, or "AI maintenance" quietly stays a category nobody can total.
OOwn the numbers. Where did each one come from?
Inference: 18.72 million plans a year at $0.014 each, from Bramblewick's own model rate card, about $262,000. Eval and monitoring: two reviewers at 30 percent of their time each, loaded at $95,000, about $57,000. Retraining: a quarterly refresh cycle, about $75,000. On-call: roughly three incidents a year at $6,000 each, about $18,000. Data pipeline: a quarter of a data engineer at $130,000 loaded, about $32,500. This is also where the rejected alternative sits: freezing Provender's model to dodge the retraining line, turned down because dietary guidance and allergen labeling law keep moving whether or not the code does, so a frozen model just drifts out of compliance quietly instead of costing a line item.
Owning the number means saying where it came from and what got turned down instead, not just stating a figure.
UUse a range, not one number.
Inference is solid enough to treat as a point figure, it's a rate card times a forecastable subscriber count. Retraining, eval labor, and incident response are genuinely uncertain: $45,000 to $110,000, $45,000 to $70,000, and $15,000 to $60,000. All in, the honest range on the whole maintenance line is $392,000 to $542,000 a year, not the single $444,580 a slide would round to.
The whole case for recomputing yearly instead of trusting one figure lives inside that range.
Hand sketched quadrant diagram titled Which cost line actually deserves the worry. X axis how likely it doubles on its own, from steady and forecastable to can spike overnight. Y axis how much the ROI moves if it does, from small move to big move. Inference plotted top left, big move but steady. Retraining plotted middle right, moderate move but can spike overnight. Eval labor in the middle. Data pipeline and on-call plotted lower.
Inference swings the number hardest. Retraining is the one actually likely to move without warning.
NNail the sanity check. Does the number survive being compared to something real?
Even the fully-loaded cost, about 2.4 cents a plan, is roughly 800 times cheaper than paying a dietitian to build the same personalized plan by hand, about $19 each. That's the check that matters: this was never a reason to stop running Provender. The number that should have worried Idrissa three years earlier is a different one: 452%, comfortably impressive by any standard Bramblewick would set, sat unchanged on three straight renewal decks while the subscriber base it was supposed to describe grew forty times over.
The hardest step, and the one most answers skip. A number that looks calm can still be sitting on top of a real, moving cost underneath it.
DDirection. Which assumption would move the answer most?
By raw dollars, inference swings it hardest: double it and the 136% return falls to about 49%. But inference is also the line Bramblewick can already forecast, it tracks subscriber count, which the business already models. The line actually worth watching is retraining. It's a fraction of inference's size, but it's one dietary-guideline change or one new allergen-labeling rule away from doubling overnight, the way it very nearly did the week Branwell caught the walnut oil.
Naming the line that's both uncertain and consequential, not just the biggest one, is what a good estimator does that a bad one skips.

Three things worth stating directly, since this is where the real judgment sits. The alternative Bramblewick seriously considered, freezing Provender's model version to avoid the retraining line entirely, lost because the world it reasons about, dietary science and allergen law, keeps moving even when the code doesn't, so a frozen model doesn't stay safe, it just stops telling anyone when it stops being safe. The AI-specific failure worth naming is silent constraint drift: a routine update nudged Provender's language toward more appetizing write-ups, and that nudge occasionally beat the hard allergen-exclusion rule in the free-text part of a recipe the structured filter never scans, so a nut-free box nearly got a walnut-oil garnish from a model that had no idea it had done anything wrong. The guardrail is the eval sample, six percent of every week's plans checked against a golden set, gated on a calibrated bar, allergen-exclusion accuracy has to clear 99.5 percent on that set, not 100, because a handful of genuinely ambiguous ingredient calls will never resolve to a clean yes or no. And the trade-off is real: reviewing every single plan instead of sampling six percent would catch a systematic error faster, but it would cost more than Provender's entire eval and monitoring line already does, several times over, so Bramblewick is trading a small, known chance of a slower catch for keeping that line anywhere near $57,000 instead of ten times that.

And if you want to be sure it really works, try it somewhere else

Same five letters, a physical-therapy app instead of a recipe box, and this time the thing that almost drifted wasn't an ingredient. It was a patient's own healing timeline.

Gaitline is Northrop Health's AI rehab planner. A patient recovering from knee or shoulder surgery checks in weekly with their pain and mobility numbers, and Gaitline writes that week's exercise plan, built against the specific limits their surgeon set for that stage of recovery. Northrop runs it across 8,500 active patients, each getting one regenerated plan a week, about 442,000 plans a year.

Hand sketched comparison diagram titled Same BOUND, a rehab app instead of a recipe box. Left panel, a box icon labeled Idrissa, Bramblewick Kitchen, caption a nut-free recipe drifted toward walnut oil. Right panel, a person icon labeled Ozgur, Northrop Health, caption a rehab plan drifted past a knee patient's real limit.
Same BOUND, a different domain, a different thing that almost drifted past a hard limit nobody had priced the checking of.

Ozgur Vasquez-Thorn owns Gaitline's numbers, and the ROI sheet Northrop pitched investors with had the same shape as Bramblewick's: value against one build cost, computed at launch, never revisited. What almost drifted here wasn't an allergen. A version update loosened how strictly Gaitline weighted a surgeon's specific "no full weight-bearing before week six" note buried in a patient's intake file, and for one post-op knee patient, a plan quietly progressed toward exercises the surgeon hadn't cleared yet. A clinical reviewer caught it in a routine weekly audit, days before the patient would have seen it.

The decision Ozgur would take back Gaitline's ROI sheet, like Provender's, priced in the one-time build and never added the ongoing clinical-review labor that keeps a written plan inside a specific surgeon's real limits, because the pilot cohort was 40 patients under one surgeon's protocol, and there was nothing yet worth reviewing at scale.

Same rank, different lever: Northrop's fix isn't a bigger clinical team or a smarter model. It's the same five-line habit, run against Gaitline's own numbers, retraining tied to how often post-op protocols actually change across specialties, not a calendar guess, and eval sampling weighted toward the patients furthest from a routine recovery, not an even random slice.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: count all five lines, re-run the ROI yearly against real usage, watch the line most likely to double, not the biggest one.
Cost: there's no engineering time this quarter to build a real range. Ship the cheap version first, a single memo estimating a low and high bound from last year's actuals, revisited next quarter.
The model got better, for real: say inference cost drops by half overnight. The fully-loaded ROI improves, but the habit doesn't change. A frozen sheet was already wrong at the old inference price; a cheaper model just moves one term, it doesn't excuse never re-running the sum.

Where people run it wrong.
They build the ROI sheet once, at launch, and never give it a reason to be revisited.
They range every line evenly instead of separating the forecastable ones from the genuinely uncertain ones.
They watch the line with the biggest dollar figure and miss the smaller one that's actually about to move.

How to use it live. Ask the age question before naming a number: "Is this ROI figure computed against this year's real cost, or is it the number from the year this feature launched?" That question alone usually tells you whether you're about to defend a real number or a stale one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: show the arithmetic, own the assumptions. Built for estimation and cost or ROI questions like this one, not a habit-flip story.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Idrissa Frobisher, who owns Provender's ROI at Bramblewick Kitchen, and inherited a cost sheet built the week Provender launched, back when it served one diet program and 3,000 pilot subscribers.
3 · THE BLIND SPOT
What did the launch-day ROI sheet get built for, that three years of real growth didn't match?
Tap to flip
ANSWER
A single-diet, 3,000-subscriber pilot with almost no ongoing cost yet. The $190,000 build-cost cell never grew as Bramblewick added three more diet programs and forty times the subscribers, so 452% kept appearing on decks that no longer described the real product.
4 · THE EQUATION
What five lines make up Provender's real annual cost?
Tap to flip
ANSWER
Inference at real volume, eval and monitoring labor, retraining and prompt-tuning, on-call for model regressions, and data-pipeline upkeep.
5 · THE OLD DECISION
What decision would Idrissa take back?
Tap to flip
ANSWER
Hard-coding the ROI sheet's cost cell to the one-time $190,000 build cost, with no plan to revisit it, decided the week Provender launched, before there was anything ongoing worth adding.
6 · THE NUMBER
Fill in the blank: the naive ROI on Provender's renewal deck said ___%. The fully-loaded, honest number is about ___%.
Tap to flip
ANSWER
452% naive. About 136% fully-loaded, still a strong return, about 800 times cheaper per plan than a dietitian building it by hand.
7 · THE REPLAY
Same near miss, new habit, what changes?
Tap to flip
ANSWER
The cost cell gets worked out again every year against real usage. When Stavros asks whose number it is, Idrissa already has this year's: 136%, not three-year-old arithmetic.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the different near miss?
Tap to flip
ANSWER
Gaitline, Northrop Health's AI rehab planner. The near miss there is a post-op knee patient's plan drifting past the surgeon's cleared limit, not an allergen.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Provender's ROI sheet keep saying 452% for three straight years, even as Bramblewick scaled forty times over?
  • A. Because inference cost never changed the whole time.
  • B. Because the sheet's cost line was a fixed, one-time build figure that nobody had a reason to revisit.
  • C. Because Bramblewick stopped tracking subscriber growth.
  • D. Because Stavros personally approved the same number every year.
Show hint
Check "the choice that mattered" key point box in Let's learn.
Show answer
B. The sheet divided value by a fixed $190,000 build cost. Nothing in its design ever added the real ongoing costs that showed up as the product scaled.
Fill in the blank
2. Provender writes about ___ million personalized plans a year, at roughly $___ each in inference cost alone.
Show hint
Look at the subscriber count and plans-per-week figure in Let's learn, right after the "before" diagram.
Show answer
18.72 million plans a year, at $0.014 each. About $262,000 in total inference cost, the one line the naive sheet's own math would have gotten right if it had ever included it.
True or false
3. True or false: the walnut-oil near miss happened because Provender's rule-based allergen filter had a bug in how it scanned the ingredient list.
  • True
  • False
Show hint
Check where in the recipe the walnut oil actually showed up, in the "what it costs at its worst" paragraph.
Show answer
False. The structured filter worked exactly as built. The problem sat in the free-text finishing-touch language a model update started favoring, a part of the output the filter never scanned at all.
Short answer, name the rejected alternative
4. What alternative did Bramblewick consider instead of building the five-line fully-loaded model, and why was it turned down?
Show hint
Look at the O step in the framework recap.
Show answer
Model answer: Freezing Provender's model version to avoid the retraining cost line. Turned down because dietary guidance and allergen labeling law keep changing whether or not the code does, so a frozen model quietly drifts out of safety instead of costing a real line item.
Short answer, apply it yourself
5. Think of a subscription product you use, or would pitch, that got sold against one upfront cost. What's one ongoing cost a naive version of that pitch would leave out?
Show hint
Think about what has to keep happening after launch for the product to stay safe or accurate, not just what it cost to build.
Show answer
Model answer: A home security app's AI camera alerts get pitched against the price of the camera hardware alone, but reviewing false alarms, retraining the person-detection model for a new season's lighting, and the cloud storage for weeks of clips are all real costs a one-time hardware price never counts.
Short answer, work the number
6. If retraining cost doubled from $75,000 to $150,000 a year, with every other line unchanged, would the fully-loaded ROI fall below 100 percent?
Show hint
Check the D step's sensitivity figures, and compare retraining's swing to inference's.
Show answer
No. Total maintenance would rise to about $519,580, giving an ROI of about 102%, still just above 100%. That's a real drop from 136%, but retraining isn't the line with the biggest raw swing, inference is; doubling inference instead drops the ROI to about 49%.
Before you close the answer
Why this works
Tests whether you'll treat the cost of maintaining an AI feature as real, ongoing arithmetic, not a line you gesture at, and whether you know a number that never changes is itself a warning sign, not proof everything's fine.
Follow-up traps
"Isn't $444,580 a year kind of a rounding error next to $1,050,000 in value?" Response: in dollars, yes, that's exactly why it's still worth running. The point was never that Provender is unprofitable. It's that 452% was never a real number, and a CFO deserves the real one, 136%, not the one that happens to look better.

"Why not just review every single plan Provender writes instead of sampling six percent?" Response: full review would cost several times more than the whole eval line already does, for a return that mostly buys speed on catching a systematic error, not certainty; the six-percent sample against a calibrated golden set was the trade Bramblewick chose to make instead.
If pressed
The 99.5 percent allergen-exclusion bar on the golden set isn't arbitrary, it's set just above Provender's measured natural variance on genuinely ambiguous ingredient calls, an oil pressed from a legume that shares a processing line with tree nuts, for instance, so the bar can catch a real regression without also failing on cases no reasonable system could resolve to a clean yes or no.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more