How do you account for the cost of maintaining an AI feature in its ROI?
Provender reads a Bramblewick Kitchen subscriber's allergies, macros, and goals, then writes three personalized recipe-and-nutrition-plan options into that week's box. Idrissa Frobisher owns what Provender actually costs to run, a number that had said the same thing, 452% ROI, on every renewal deck for three straight years while the subscriber base grew forty times over. Then Bramblewick's CFO, Stavros Delacorte, asked a question the deck had never been built to answer: whose number is this, this year's, or launch day's?
- Add the five ongoing cost lines to the ROI, every year, not once at launch.Why: a build cost divided into value forever is what made 452% look real when this year's honest number is 136%.
- Re-run the number on this year's real usage, not launch day's.Why: usage-driven costs grow with the exact thing that makes leadership love the feature, so a frozen formula gets more wrong every year nobody touches it.
- Give a range on the lines that are genuinely uncertain, not one clean figure.Why: retraining and incident response can swing two to three times over based on one real event, so a single number claims confidence nobody has.
- Watch retraining cadence as the line most likely to actually double, not just the biggest one.Why: inference moves the ROI the most in raw dollars, but it's the most forecastable; one regulation change can double retraining overnight.
- Sample eval coverage instead of reviewing every plan Provender writes.Why: full review would cost more than the feature is worth; a calibrated sample against a golden set still catches drift.
- Don't freeze the model to dodge the retraining line.Why: dietary guidance and allergen law keep moving even when the code doesn't, so a frozen model drifts out of safety quietly instead of costing a line item.
How to answer this, stage by stage
Nobody is grading whether you can name a cost per API call. They're grading whether you know a number that never changes is a warning sign, not proof nothing's wrong, and whether you can rebuild it honestly with real arithmetic.
Let's learn
For three years, Idrissa Frobisher's renewal slide said the same thing: 452% ROI. Nobody had thought to ask if that number still meant anything.
Provender is Bramblewick Kitchen's recipe and nutrition planner. A subscriber tells it what they're allergic to and what they're eating toward, keto, diabetic-friendly, high-protein, and each week it writes three personalized recipe-and-nutrition-plan options into their box.
Before Provender, six recipe developers and dietitians built eight fixed weekly menus by hand. A subscriber picked one of the eight, filtered only by a coarse tag, vegetarian or gluten-free, nothing more specific than that. A brand-new menu variant took about three weeks of a dietitian's time to build and test.
With Provender, a subscriber gets three options built just for them, in about four seconds, at roughly a cent and a half each to run. Bramblewick has 120,000 active subscribers now, each getting three plans a week. That's 18.72 million plans a year. At $0.014 a plan, inference alone runs about $262,000 a year.
Here's the turn. Provender's original pitch didn't get the inference number wrong. It just never added anything else. The ROI sheet built at launch divided a year of value, $1,050,000, made up of about $640,000 in subscriptions that would have churned without the personalization and $410,000 in recipe-development headcount Bramblewick didn't have to add, by one number, the $190,000 it cost to build Provender in the first place, back when it served one diet program and 3,000 pilot subscribers. That was a fair sheet the week it was written. There was almost nothing else to count yet: no retraining had happened, no incident had happened, monitoring was one person glancing at a dozen plans most Fridays.
What it costs at its worst: three years later, Bramblewick runs four diet programs and 120,000 subscribers, and Provender genuinely does need retraining, monitoring, on-call, and a maintained data pipeline, real costs that never made it onto that sheet. One ordinary Tuesday, a routine model update quietly started favoring a flavor-boost instruction over the hard allergen-exclusion rule, and a handful of nut-free-tagged recipes picked up a walnut-oil finishing drizzle, an ingredient the rule-based filter never caught, because walnut oil isn't tagged as a whole nut in the ingredient database, and the drizzle only ever showed up in the free-text finishing touches the filter never scanned. Branwell Gower, on the eval team, happened to have that exact recipe type in her weekly six-percent sample. She caught it Thursday morning. Nothing shipped. Nobody outside the eval team ever knew how close it came.
What I would leave alone: the internal test-kitchen tool, a small assistant the two remaining recipe developers use to search for ingredient ideas, about 50 uses a week, no subscriber ever sees its output directly. Its naive ROI, just inference cost, a few dollars a month, is basically the honest number, because it never scaled past a handful of people and it never touches a customer-facing safety claim.
The lesson: a yearly ROI number isn't a fact you compute once and frame. It's a habit you keep, because the exact growth that makes a feature look good on a renewal deck is the growth that makes the ongoing costs real.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a number that never moves is itself the warning sign.
Idrissa Frobisher has owned Provender's numbers since the pilot, back when it served one diet program and about 3,000 subscribers who'd signed up specifically to test it. Idrissa kept the launch pitch deck bookmarked, mostly out of habit: value minus one build cost, $1,050,000 minus $190,000, divided by $190,000. 452%. It was a genuinely honest number, the week it was worked out.
For a year, nothing about that math needed to change. Provender served one diet program. Retraining had never happened, there was nothing yet worth retraining against. Monitoring was Idrissa, personally, glancing at a dozen plans most Fridays. Bramblewick's board loved the slide. So did every renewal deck after it.
It thinned in three quiet beats, and none of them looked like a mistake. Bramblewick launched a keto program, then diabetic-friendly, then a fourth built around GLP-1 users' specific nutrition needs, each one adding real, ongoing work: new dietary rules to encode, new golden-set cases to build, more plans written every single week as the subscriber base climbed past 40,000, then 80,000, then 120,000. Somewhere in there, a real reviewer role got created, a real quarterly retraining cycle got built, a real on-call rotation started catching bad versions before they shipped. Every one of those was the right call. None of them ever made it back into that one cost cell on the renewal slide, because nobody had ever been asked to open the slide and check.
Then came the Tuesday that almost changed that on its own.
A routine model update shipped, meant to make Provender's recipe write-ups read a little more appetizing, a small nudge toward warmer, more flavorful language. What nobody caught in testing: that nudge occasionally out-competed the hard allergen-exclusion rule sitting underneath it, in the free-text finishing touches Provender wrote for a recipe, not the structured ingredient list the safety filter actually scanned. For a handful of nut-free-tagged boxes, that meant a walnut-oil drizzle, suggested as a garnish, in a recipe meant to contain no tree nuts at all.
Branwell Gower reviews six percent of every week's plans against a golden set, mostly a spot check, mostly uneventful. That week, by ordinary luck, three of the flagged recipes happened to be nut-free boxes carrying the new garnish line. She caught it Thursday morning. Nothing shipped. Nobody outside the eval team ever knew how close it came.
Idrissa sat with that for a few days. Not because Provender had failed, it hadn't, the review caught exactly what it was built to catch. What sat wrong was smaller: if that six-percent sample had landed on a different set of recipes that week, the walnut oil would have shipped, and there would have been no line item anywhere that explained why the company had only a six-percent chance of catching it, because the ROI sheet had never once said reviewing Provender's output cost anything at all.
Three weeks later, in the ordinary course of an annual budget review, Stavros Delacorte, Bramblewick's CFO, pulled the last three years of renewal decks side by side. He wasn't looking for a reason to cut Provender. He was doing what a CFO does: noticing that one number, 452%, appeared identically on all three, while the subscriber count behind it had grown forty times over.
"Whose number is this," he asked Idrissa, in the flat tone that isn't an accusation yet. "This year's, or the year we launched?"
Idrissa didn't have an honest answer. That was the real problem, not the number itself.
What Idrissa built over the following week was the sheet that should have existed the whole time: five real ongoing lines, inference at 120,000 subscribers' actual usage, eval and monitoring labor, quarterly retraining, on-call, data-pipeline upkeep, adding up to about $444,580 a year, with honest ranges on the three lines that genuinely swing. Run against this year's real value, $1,050,000, that's a 136% return. Not 452%. A real 136%, worked out, defensible, and still comfortably worth running, about eight hundred times cheaper per plan than paying a dietitian to write it by hand.
The decision that opened the door traced back to the very first pricing meeting, the week Provender launched. Someone asked whether the cost cell should be a formula that recalculated each year or a fixed figure. A fixed figure was faster to ship, and at the time, genuinely correct, there wasn't yet anything ongoing worth modeling. Nobody wrote down a date to come back and check.
Run that meeting again, with one line added: review the cost cell every year, against that year's real usage, not just the arithmetic it started with. Same forty times growth, same near miss, same CFO's question. This time Idrissa's answer is immediate: "This year's. Here it is." The number is smaller than the one on the wall. It's also true.
One design let "the ROI" mean whatever number happened to be sitting on the slide. The other lets it mean whatever the year actually cost.
What Idrissa would tell their past self, back in that first pricing meeting: a number that never changes isn't a sign nothing changed. It's a sign nobody was asked to look.
BOUND, or the five lines a launch-day sheet never grew
Not a story wearing a framework's clothes. This is an estimation problem with a stale number sitting on top of a real, moving one, and BOUND is what turns "AI upkeep costs something" into a figure a CFO can actually check.
Three things worth stating directly, since this is where the real judgment sits. The alternative Bramblewick seriously considered, freezing Provender's model version to avoid the retraining line entirely, lost because the world it reasons about, dietary science and allergen law, keeps moving even when the code doesn't, so a frozen model doesn't stay safe, it just stops telling anyone when it stops being safe. The AI-specific failure worth naming is silent constraint drift: a routine update nudged Provender's language toward more appetizing write-ups, and that nudge occasionally beat the hard allergen-exclusion rule in the free-text part of a recipe the structured filter never scans, so a nut-free box nearly got a walnut-oil garnish from a model that had no idea it had done anything wrong. The guardrail is the eval sample, six percent of every week's plans checked against a golden set, gated on a calibrated bar, allergen-exclusion accuracy has to clear 99.5 percent on that set, not 100, because a handful of genuinely ambiguous ingredient calls will never resolve to a clean yes or no. And the trade-off is real: reviewing every single plan instead of sampling six percent would catch a systematic error faster, but it would cost more than Provender's entire eval and monitoring line already does, several times over, so Bramblewick is trading a small, known chance of a slower catch for keeping that line anywhere near $57,000 instead of ten times that.
And if you want to be sure it really works, try it somewhere else
Same five letters, a physical-therapy app instead of a recipe box, and this time the thing that almost drifted wasn't an ingredient. It was a patient's own healing timeline.
Gaitline is Northrop Health's AI rehab planner. A patient recovering from knee or shoulder surgery checks in weekly with their pain and mobility numbers, and Gaitline writes that week's exercise plan, built against the specific limits their surgeon set for that stage of recovery. Northrop runs it across 8,500 active patients, each getting one regenerated plan a week, about 442,000 plans a year.
Ozgur Vasquez-Thorn owns Gaitline's numbers, and the ROI sheet Northrop pitched investors with had the same shape as Bramblewick's: value against one build cost, computed at launch, never revisited. What almost drifted here wasn't an allergen. A version update loosened how strictly Gaitline weighted a surgeon's specific "no full weight-bearing before week six" note buried in a patient's intake file, and for one post-op knee patient, a plan quietly progressed toward exercises the surgeon hadn't cleared yet. A clinical reviewer caught it in a routine weekly audit, days before the patient would have seen it.
Same rank, different lever: Northrop's fix isn't a bigger clinical team or a smarter model. It's the same five-line habit, run against Gaitline's own numbers, retraining tied to how often post-op protocols actually change across specialties, not a calendar guess, and eval sampling weighted toward the patients furthest from a routine recovery, not an even random slice.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: count all five lines, re-run the ROI yearly against real usage, watch the line most likely to double, not the biggest one.
Cost: there's no engineering time this quarter to build a real range. Ship the cheap version first, a single memo estimating a low and high bound from last year's actuals, revisited next quarter.
The model got better, for real: say inference cost drops by half overnight. The fully-loaded ROI improves, but the habit doesn't change. A frozen sheet was already wrong at the old inference price; a cheaper model just moves one term, it doesn't excuse never re-running the sum.
Where people run it wrong.
They build the ROI sheet once, at launch, and never give it a reason to be revisited.
They range every line evenly instead of separating the forecastable ones from the genuinely uncertain ones.
They watch the line with the biggest dollar figure and miss the smaller one that's actually about to move.
How to use it live. Ask the age question before naming a number: "Is this ROI figure computed against this year's real cost, or is it the number from the year this feature launched?" That question alone usually tells you whether you're about to defend a real number or a stale one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just review every single plan Provender writes instead of sampling six percent?" Response: full review would cost several times more than the whole eval line already does, for a return that mostly buys speed on catching a systematic error, not certainty; the six-percent sample against a calibrated golden set was the trade Bramblewick chose to make instead.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.