How does ROI differ for a feature that unlocks a new customer segment?
Loomtalk is Cinderwood's real-time AI voice tutor. A learner talks, the model answers out loud, the conversation keeps moving. Calla Onwordi has owned Cinderwood's shared ROI calculator for two years, ever since it scored one segment: working professionals rehearsing calls in a language they half know. Then Cinderwood shipped Loomtalk Lite, a short-burst mode meant to catch five minutes on a commute, and a very different kind of learner showed up behind it. Soline, three weeks into the growth team, asked why Lite used the same per-minute rate as everything else. Calla did not have an answer. That was the whole problem.
- Build the new segment's unit economics from its own usage logs, not the shared template.Why: the shared model assumed cost scales only with minutes talked; Loomtalk Lite's real cost came mostly from reconnects and retries, which do not.
- Re-measure predicted cost against real cost before scaling a new segment's marketing spend.Why: the model on screen said $3.92 profit while the real number was a $3.74 loss, and nobody sees the second number until it shows up in the whole company's ledger.
- Fix the AI-specific driver first, not just the price.Why: an accent adapter and a longer silence timeout cut real cost from $11.74 to $7.94 a user before a single price change, which proves the problem was the model, not the market.
- Reprice only after the cost side is fixed.Why: raising the price first just makes a broken assumption more profitable to be wrong about; fixing cost first, then repricing, is what actually gets to a real $3.06 margin.
- Leave the shared calculator alone for features that share the old segment's shape.Why: grammar drills and pronunciation coaching have no live voice pipeline, so the old formula is still true for them; not every new feature needs its own model.
- Watch the gap between predicted and real cost, not the total spend line.Why: total spend only showed the problem in month six; predicted-versus-real cost per session would have shown it in week two.
How to answer this, stage by stage
Nobody is grading whether you know what a connection floor is. They are grading whether you know a shared cost formula is a claim about the past, not a law about the future, and whether you can rebuild it honestly when a new kind of user shows up.
Let's learn
Loomtalk is Cinderwood's real-time AI voice tutor. A learner opens the app, talks, and the model answers back out loud, correcting a mistake or nudging the conversation forward, the way a patient human tutor would.
For two years, Loomtalk served one kind of learner: working professionals on the Fluency Track plan, $54 a month, rehearsing calls in a language they mostly know but don't quite trust yet. They talked in long stretches, three or four sessions a week, about twenty-four minutes each. Cinderwood's shared Feature ROI Calculator was built off exactly that usage, and it worked. Every new feature Calla scored against it, grammar drills, a pronunciation coach, a pricing test, came back accurate, because every one of those features served the same kind of session.
Then Cinderwood shipped Loomtalk Lite: a short-burst mode, $8 a month, built for five minutes on a bus or a lunch break, and tuned for accents the original Fluency Track voice model barely touched. A very different learner showed up behind it. Not professionals polishing an existing skill. Beginners, texting distance from the language, practicing four times a day for five and a half minutes at a stretch, on phones with patchy signal.
Calla ran Loomtalk Lite through the same calculator that had been right for two years. It took the one rate that formula has always used, just over half a cent a minute, multiplied it by Lite's shorter sessions, and came back with a number that looked, if anything, better than the original segment's: a predicted profit of $3.92 a user, every month.
Here's the turn. The extra beginners weren't the problem. What Loomtalk Lite actually did was talk to Calla's spreadsheet in a shape it had never been tested against. A twenty-four minute professional session barely notices a reconnect or two. A five-and-a-half minute beginner session, full of pauses while someone hunts for a word, reconnects two or three times, and each reconnect bills that flat minimum again. On top of that, the accent Lite's users mostly speak wasn't one the model handled confidently yet, so it asked people to repeat themselves on more than a third of their turns, and every re-ask is another partial call to the model. None of that scales with minutes talked. All of it scaled with how new, and how nervous, a learner was.
What it cost at its worst: nobody was watching for a gap that small-looking a formula could hide. Loomtalk Lite grew fast, exactly the way a cheap, easy-to-try feature should. Every new user made the real loss a little bigger, and the shared calculator kept saying, cheerfully, that the segment was working.
What I would leave alone: Loomtalk's grammar drills and pronunciation coach, both text-based or short pre-recorded clips with no live voice connection at all. They share the Fluency Track's shape closely enough that the shared calculator has never once been wrong about them, whichever segment ends up using them.
The lesson: a shared cost formula isn't a fact, it's a measurement that happened to be true for the segment it was built from. The day a genuinely new kind of session shows up, the formula's shape needs checking before its output gets trusted, not after.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a boring spreadsheet assumption cost real money for six months before anyone noticed.
Calla Onwordi built Cinderwood's Feature ROI Calculator the year Loomtalk launched, back when there was exactly one segment to build it for. She pulled the rate straight from real sessions: professionals on the Fluency Track, twenty-four minutes at a stretch, half a cent a minute, give or take. For two years, that number was simply true. She could plug any new feature into it and trust what came out.
For two good years, that trust never once cost her anything. Grammar drills, a text feature, went into the calculator and came back cheap and right. Pronunciation coaching, built on short pre-recorded clips, went in and came back cheap and right. A pricing test on the Fluency Track itself went in and came back right down to the dollar. Every time, reusing the shared rate took ten minutes. Building a new model from scratch would have taken a week. She never had a reason to spend the week.
The habit thinned in three small beats, and none of them looked like carelessness. Grammar drills: she typed the new feature's expected minutes into the calculator without asking whether its usage pattern matched the rate's own origin. Pronunciation coaching: same thing, faster this time, because the last two had been fine. Then Loomtalk Lite's pricing test: she ran it through the same calculator, saw a predicted $3.92 profit a user, told the launch review it looked healthy, and moved on to the next slide.
Then came a Tuesday, three weeks into a new hire's first month.
Soline was building a simple chart, predicted cost against actual sessions per feature, mostly to learn the codebase behind the numbers. She noticed the same rate, half a cent a minute, sitting in Loomtalk Lite's row as in the Fluency Track's row above it. "Wait," she asked, not accusing anyone of anything, "why does Lite use the same per-minute rate? Its sessions are like a fifth of the length." Calla started to give the obvious answer, that cost scales with minutes, and stopped halfway through the sentence. She had never actually checked whether that was true for a session shape this different.
What Calla did next was not tweak the rate. She pulled Loomtalk Lite's real usage logs, all of them, for the past six months, and built its cost from the ground up: base talk time at the same half-cent rate, plus every reconnect at its own flat charge, plus every accent-driven retry at its own small cost. The real number came out to $11.74 a user a month. Against an $8 price, that wasn't a thin margin. It was a loss of $3.74 a user, every user, every month, since the month Loomtalk Lite launched.
The decision she'd take back traced to the week the calculator was built. Someone, and it might as well have been her, decided the ROI tool's cost input would be one hard number, typed in once, rather than a live pull from each feature's own usage logs. At the time, with one segment and one session shape, that was obviously the right call. Building a flexible per-segment cost engine for a company with a single product line would have been pure overhead nobody needed yet.
Run that decision again, with one thing added: a note, attached to the rate itself, saying re-measure this the day a feature's sessions stop looking like the ones it came from. Same Tuesday. Same question from Soline. This time Calla doesn't need six months of real usage to answer her, because the calculator already flagged Loomtalk Lite's predicted-versus-real gap the week it crossed twenty-five percent, back in month one, at a $2,992 loss instead of a $127,160 one.
The replay didn't stop at finding the number. Cinderwood shipped a lightweight accent adapter tuned to the three accents driving the worst retry rates, cutting re-asks from 38 percent of turns to 16 percent. They loosened the silence timeout from twelve seconds to twenty-five, giving beginners room to think without forcing a reconnect, and cutting reconnects from 2.6 a session to 1.4. That alone brought real cost down to $7.94 a user. Only then did Cinderwood raise Loomtalk Lite's price, from $8 to $11, still a fraction of the Fluency Track's $54. The real margin flipped from a $3.74 loss to a $3.06 profit, a swing worth about $231,200 a month across the segment's 34,000 users.
One design let "the ROI" mean whatever number a two-year-old rate happened to produce. The other lets it mean whatever the segment actually costs, checked against real use, before the marketing budget finds out the hard way.
What Calla would tell her past self, back in the week that calculator was built: a number that's right for two straight years isn't proof it will still be right on the third. It's proof nobody has yet asked it to describe something different from what it was built to describe.
FLIPS, or the five steps if you want to remember them
Not a story wearing a framework's clothes. This is what to actually run, in order, any time a question asks how ROI changes when a genuinely new segment shows up.
Two things worth naming directly. The alternative Cinderwood actually considered instead of rebuilding Loomtalk Lite's cost model from scratch was simpler and cheaper: apply a flat thirty percent discount to the shared per-minute rate for short sessions and call it done. That got rejected, because the real cost driver, reconnect floors and accent retries, isn't proportional to session length at all, so a flat discount keeps the same wrong shape of error; at short enough sessions it can make the estimate even more wrong, since the floor charge is a bigger share of a short session's true cost, not a smaller one. The AI-specific failure worth naming is silent unit-economics drift: a cost formula that assumes inference scales linearly with minutes talked will quietly mislead you the moment a new segment's accent or connection pattern breaks that assumption, and the guardrail is a per-segment dashboard that pulls real cost per session off usage logs, not the modeled formula, refreshed weekly, flagging anything more than twenty-five percent off. There's a real trade-off buried in the fix, too: the accent adapter Cinderwood shipped is fast and cheap but leaves a 16 percent residual retry rate; a larger, more accurate ASR model could have cut that further, at roughly three times the latency and cost per call, which would have timed out constantly on the patchy mobile networks this segment actually uses. Cinderwood chose the smaller, faster, imperfect fix on purpose.
And if you want to be sure it really works, try it somewhere else
Same five letters, a crop-disease app instead of a language tutor, and this time the thing that broke wasn't a cost formula's shape. It was what the users fed it, once feeding it stopped being free.
Blightline is Loambright's AI crop-disease tool: a farmer photographs a leaf, and the model says what's wrong with it, if anything. Loambright built it for commercial cooperatives on a flat monthly seat fee, unlimited scans, mostly routine checks on healthy-looking plants. A new offline capture mode unlocked a completely different segment: smallholder farmers, paying per scan because they can't afford a flat fee, on a plan Loambright priced at a few cents a photo.
F Zorah, agronomy PM at Loambright, owns Blightline's eval set and its accuracy sheet. L She stopped rechecking whether the golden eval set's mix of easy and hard cases still matched what real users were sending in, since it had held steady for the cooperative segment for a year. I A different flip entirely: once a scan costs money, smallholders stop scanning routinely and start scanning only the worst-looking, most alarming plants, exactly the substitution flip, using something for the easy cases and saving it for the hard ones, in reverse. P The old decision: the golden eval set was drawn once from the cooperative segment's routine, mostly-healthy scan mix, and never rebuilt for a segment that scans selectively. S The replay: rebuild the eval set from the smallholder segment's actual, harder scan mix, and the accuracy bar recalibrates from an apparent 71 percent, measured against the wrong mix, to a true and still solid 89 percent against the mix it's really serving. The model never got worse. The yardstick was measuring the wrong thing.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: new segment, new cost model, built off real usage, checked against the price before you scale spend behind it.
Cost: there's no time this quarter to build a full usage-log pipeline. Ship the cheap version first, a manual spot-check of fifty real sessions against the shared formula's prediction, revisited next quarter.
The model got better, for real: say the per-minute rate drops by half overnight. The predicted number improves, but the habit doesn't change. A formula built for one segment's shape was already the wrong thing to trust for another's; a cheaper model just moves one input, it doesn't excuse skipping the recheck.
Where people run it wrong.
They edit one number in the old formula instead of asking whether its shape still applies.
They watch total spend instead of the gap between predicted and real cost per unit, so the problem only shows up once it's expensive.
They fix the price before fixing the mechanism, which just makes a wrong assumption more profitable to keep being wrong about.
How to use it live. Ask one question before trusting any ROI number for a new segment: "Was this number measured from this segment's own usage, or inherited from a different one?" That question alone usually tells you whether you're looking at a real number or a borrowed one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't the real fix just 'don't build for beginners, they're too expensive'?" Response: no. After the accent adapter and the longer silence timeout, real cost dropped to $7.94 a user, close enough that a modest, honest price move made the segment profitable. The expensive part was never the segment. It was measuring it wrong.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.