CaseAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #10

What would make you conclude an AI feature has negative ROI despite good usage?

FLIPS · churn-prediction retention offers for a subscription SaaS business

Holdline scores every one of Cascadeworks' 11,400 subscription accounts for churn risk, once a day. Score high enough, and the account drops into a Save Queue: a Customer Success rep gets nudged to send a discount or a free upgrade before the account can cancel. Shirin Kestleman runs retention analytics at Cascadeworks. Nanami Aoyagi joined her team seven months into the program and asked one plain question nobody had an answer for.

The direct answer
Run a real holdout: hold back a random slice of flagged accounts from any retention offer, then compare their churn rate to the accounts that got one. If the two groups leave at close to the same rate, the offer isn't preventing churn, it's just money spent on people who were staying anyway. That gap, not the number of offers sent, is what tells you whether the feature has negative ROI.
Do this, in order
  1. Build a real holdout before trusting any "saved" number.Why: without a group that got no offer, "we saved 90 accounts" is a guess wearing a fact's clothes.
  2. Measure spend against the incremental lift only, not the whole flagged group.Why: paying for a discount on an account that was staying anyway is pure cost with nothing coming back.
  3. Check who the model never flagged, not just who it flagged.Why: the churn hiding in the false negatives can cost more than the false positives ever save.
  4. Put the audit back on the same dashboard as usage, as a running number, not a one-time test.Why: the coverage number climbed for six straight months with nobody re-checking what it actually meant.
  5. Route the expensive offer to accounts the model is actually confident about, not the whole queue.Why: most flagged accounts never needed the full discount to stay; a cheaper play protects the same revenue for a fraction of the cost.
  6. Leave the bottom of the queue alone.Why: accounts that never get flagged and never get an offer aren't costing the program anything, so there's nothing there to audit.

How to answer this, stage by stage

Nobody is grading whether you can say "usage isn't the same as value." They're grading whether you'd actually build the missing control group, or just gesture at the idea and move on.

1
Scope it to one real product before answering in the abstract
Say it like this
"Let's make this concrete. Say we're Cascadeworks, a project-management SaaS with about 11,400 accounts. We license Holdline, it scores every account for churn risk daily, and anything over 72 lands in a Save Queue that Customer Success works."
Why this works
An abstract "usage isn't value" answer stays a slogan. One real account lets you actually walk the interviewer through it.
2
Say the plan out loud before naming a number
Say it like this
"I'm going to run this as FLIPS. Find the person whose habit changes, find what they stopped doing, name the flip, name the old decision behind it, then replay the same story with that decision reversed."
Why this works
Signals a method already in motion, not five thoughts arriving in whatever order they occurred to you.
3
Reframe what the question is actually testing
Say it like this
"Here's the real question underneath this one: can you tell the difference between people using the flag, and the flag actually changing what happens to a customer? Those are two different numbers, and only one of them shows up on most dashboards."
Why this works
Separates a real answer from a list of "signs to watch for" that never commits to anything.
4
Give the one decision, plainly
Say it like this
"Concretely: I'd hold back a random slice of flagged accounts from any offer, on purpose, and compare their churn rate to the ones that got treated. If the two numbers land close together, the offer isn't saving anyone, it's a cost sitting on top of people who were staying anyway."
Why this works
This is the direct answer, said in one breath, before the interviewer has to go looking for it.
5
Prove it with the compressed story
Say it like this
"There was a retention lead, Shirin, who used to audit twenty-five flagged accounts by hand every Friday. It kept checking out, so the sample shrank to fifteen, then eight, then nobody checked at all, because the coverage number on the dashboard kept climbing anyway. A new hire asked what the holdout group was. There wasn't one. Once they built it, flagged accounts left at almost the same rate whether they got an offer or not."
Why this works
Shows the real, countable cost of trusting usage, not just "it could go wrong somewhere."
6
Say what you'd measure going forward
Say it like this
"I'd keep a small holdout running permanently, not just for one test, and put the true lift number on the same dashboard as coverage, right next to it, so nobody can read one without seeing the other."
Why this works
Shows you're designing a permanent check, not running a one-off fire drill after the fact.
7
Say what you'd leave alone
Say it like this
"The bottom of the queue, accounts scoring under 40, never gets an offer and never needs an audit. There's no spend there to check, so there's nothing to protect."
Why this works
Shows judgment instead of blanket suspicion of the whole feature.
8
Close on the decision, not the story
Say it like this
"So: good usage tells you people are picking up the flag. Only a holdout tells you the flag is worth what it costs. Until I've run that comparison, I don't know if this feature has positive or negative ROI, I just know it's popular."
Why this works
Leaves the interviewer with the decision, not just the story behind it.

Let's learn

Holdline is an AI tool that scores every subscription account for how likely it is to cancel, and pushes the riskiest ones into a queue a support team can act on before it happens.

Before Holdline, Cascadeworks' 11,400 accounts got flagged for churn risk by gut feel: a rep noticed a slow login streak, or a support ticket that sounded final. Maybe a third of the accounts that actually left had ever been flagged by anyone at all.

Now Holdline scores every account daily, and about 640 accounts a month cross the line into the Save Queue, where a rep sends a retention offer, a 25 percent discount for two billing cycles or a free upgrade, averaging $270 an account. Cascadeworks' average account is worth about $3,900 a year.

Hand sketched comparison diagram titled Small move, big snap. Left panel, a funnel icon labeled The small move, caption sample audit: 25 accounts a week, then 15, then 8. Right panel, a balance-scale icon labeled The big snap, caption then 0, she trusts the coverage number whole, no middle setting.
The audit didn't fade evenly. It held for months, then dropped to nothing in one step.
Knowledge spark: what's a holdout group? A slice of the exact same population you'd normally treat, picked at random, that gets nothing on purpose. It's the only honest way to know what would have happened without you.

Here's the turn. More offers going out, and more of the queue getting touched inside 48 hours, isn't proof the program is working. It's proof people are using it. Those are two different questions, and Cascadeworks was only ever answering the first one.

Good usage was never proof. It was the reason nobody looked again.
Save Queue coverage, nine months
100% 50% 0% audit line cut, month 6 68% m6 93% m7: Nanami asks 98% m1 m2 m3 m4 m5 m8
Save Queue coverageAudit line item cut, month 6Nanami's question, month 7
The number leadership actually watched climbed every single month, straight through the point where the only thing checking it stopped existing.

Over the six months after the sample audit got cut, Cascadeworks sent retention offers to about 3,840 flagged accounts, at $270 each: $1,036,800 total. A holdout test run later found the offer's true, causal effect: about 58 of those accounts were saved because of it. At $3,900 a year each, that's roughly $226,000 of real value, against $1,036,800 spent, a net loss of $810,800, while the dashboard's headline number sat at 98 percent, the best it had ever looked.

The choice I would take back: cutting the weekly sample audit from the monthly retention review to make the meeting shorter. The audit had come back clean for two straight quarters, so it looked like the safe thing to cut.

What I would leave alone The bottom of the queue, accounts scoring under 40, never gets an offer and never needs a check. There's no spend riding on those accounts, so there's nothing there an audit could catch.

The lesson: a feature that's popular is not the same as a feature that's working. Usage tells you people picked the tool up. Only a group you deliberately held back tells you whether picking it up changed anything.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a coverage number climbing for six months in a row was the actual problem, not proof against one.

Every Friday afternoon, Shirin Kestleman pulled the same twenty-five rows. Twenty-five accounts Holdline had flagged that week, twenty-five retention offers Customer Success had already sent, and one afternoon to find out whether any of it was real.

She'd built Cascadeworks' whole retention report from a single spreadsheet, five years back, before there was a Holdline to plug in. She knew which numbers lied by omission and which ones told the truth on their own, and she'd never once let a number onto a leadership slide without tracing it back to an actual account first.

Hand sketched horizontal timeline titled Nine months on the Save Queue. Four milestones left to right: Before Holdline, caption her own spreadsheet. The good months, caption 25 checked a week, clean. The habit thins, caption 25, then 15, then 8. The question, this milestone marked in red-orange, caption what's our holdout.
Nine months, four real moments. None of them looked like a crisis while they were happening.

Holdline arrived in the spring. It scored all 11,400 of Cascadeworks' accounts for churn risk, once a day, and anything over 72 dropped into the queue Shirin called the Save List. For the first two quarters, her Friday habit didn't change: twenty-five accounts, pulled at random from that week's flagged list, checked line by line against their usage history and support tickets, to see whether the offer had actually landed on someone who was really about to leave. Almost every week, it checked out. Real risk, real offer, real save.

So the sample shrank. Twenty-five became fifteen. Fifteen became eight. Nobody decided this on purpose. It just kept checking out clean, and a Friday afternoon is worth more spent on something that might actually be broken.

By month six, she'd stopped pulling the sample at all. The Save List's own weekly numbers had taken over the job the audit used to do: offers sent, up. Coverage, the share of flagged accounts a rep actually reached within 48 hours, up to 91 percent. It read like proof. It wasn't a number anyone was worried about.

We didn't stop the audit because the model got worse. We stopped because the usage number kept climbing, and climbing looked exactly like winning.

The trigger, when it came, wasn't a bad quarter. It was a question. Nanami Aoyagi joined the retention team in month seven, straight out of a program built around causal inference, and in her second sprint planning she asked the kind of question a newcomer asks because nobody's told her not to: "What's our holdout? The accounts we deliberately don't send an offer to, so we know how many would've stayed on their own?"

Nobody had one. Nobody had ever built one.

Hand sketched comparison diagram titled Trust in a save number was never a dial. Left panel, a dial gauge icon labeled What we assumed, caption trust fades a little as the sample shrinks. Right panel, a two-position switch icon labeled What actually happens, caption she audits, or she doesn't, it holds then it flips.
This is the whole answer, in one picture. Trust in a save number doesn't fade with the sample. It holds, then it flips, on the day someone asks a plain question.

Shirin spent the next two months building it properly instead of feeling bad about it. For eight weeks, one in ten newly flagged accounts, 128 out of about 1,280, got randomly withheld from any offer at all, tracked exactly the same as the other 1,152. At the 90-day mark, the treated accounts, the ones who got a discount or a free upgrade, had churned at 8.6 percent. The untouched accounts, the ones who got nothing, had churned at 10.1 percent.

A real gap. Just a small one. A true lift of a point and a half, on the population the whole program existed to save.

And underneath that number sat a worse one. Of every ten accounts that actually canceled that quarter, company-wide, four had never shown up on the Save List at all. Their churn scores had sat quietly under 72 right up until the day they left, because whatever pushed them out, a champion who quit, a competitor chosen at renewal, a budget line cut two departments up, wasn't the kind of thing Holdline had ever been trained to see. It watched product usage. Nobody had ever told it to watch anything else.

Hand sketched quadrant diagram titled Two ways a churn flag goes wrong. X axis how visible the mistake is, from hidden to obvious. Y axis how much it costs, from cheap to costly. Real churner never flagged plotted hidden and costly. Discount sent to a stayer plotted obvious and moderate cost. Holdout confirmed true save plotted obvious and worthwhile.
A wasted discount shows up on next month's invoice. A real churner nobody flagged shows up nowhere, until the renewal date that never comes.

The decision that opened the door was small, and it happened in a meeting nobody would call a mistake, even now. Six months in, the monthly retention review kept running past an hour, and the VP running it asked for it back down to thirty minutes. The line item that went first was "sample audit: real save or false flag." It got replaced with a single usage slide. That was a fair trade in the room that day. The audit had come back clean for two straight quarters. Nobody was worried about a number that already looked good.

Run the same nine months again, with the holdout live from month one instead of never. The gap that took a new hire's question to find shows up on its own, in month seven, sitting right next to the coverage number on the same dashboard. Shirin redesigns the offer before the spend gets any worse: the full discount goes only to accounts scoring above 85 with a usage-decline pattern, the exact signal Holdline is actually good at reading. Everyone else in the queue gets a cheap automatic nudge instead, a check-in email and a feature tip, twelve dollars instead of two hundred and seventy.

Program spend over the same six months falls from $1,036,800 to about $195,000. True value protected barely moves, holding near $210,000. For the first time since Holdline launched, the math tips the other way: about $15,000 ahead instead of $810,800 behind.

What I'd tell myself, back in that thirty-minute meeting: cutting the audit line item didn't feel like a decision about the model at all. It felt like a decision about a meeting running long. It just happened to also be the decision that let a $270 offer get mailed to ninety-eight percent of a queue for six months before anyone checked whether it was doing anything.

FLIPS, or the five questions a holdout group answers for you

Not a way to dress up "seems like it's working" in five letters. FLIPS is what forces you to name what a person stopped checking, and why that felt completely fine at the time.

FFind the person. Whose morning is this?
Shirin Kestleman, Head of Retention Analytics at Cascadeworks, five years in, the one who wouldn't let a number onto a slide until she'd traced it back to something real.
Name a real person before naming a real number, or the answer stays a slogan about usage versus value.
LLocate the habit. What did they stop doing because it worked?
The Friday sample audit. Twenty-five flagged accounts checked by hand every week, for two straight quarters, because it kept coming back clean.
The habit she lost is the actual thing that used to catch this. Name it precisely, or the story has no engine.
IIdentify the flip. What verb snaps?
Audits the sample, or trusts the coverage number whole. No setting in between; by month six, nobody was checking at all.
This is the hardest step. Everything the rest of the answer proves runs back through this one line.
90-day churn rate, treated vs. held back
12% 6% 0% 8.6% Treated (got an offer) 10.1% Held back (no offer)
Treated: 8.6% left anywayHeld back: 10.1% left with nothing
The true, causal lift is only 1.5 points. Every account above that line was leaving, or staying, on its own.
PPinpoint the old decision. Which choice only made sense before?
The monthly retention review ran long, so the "real save or false flag" line item got cut, replaced with a single usage slide: offers sent, coverage, saves logged.
Small, reasonable, in a real meeting. That's what makes it worth taking back instead of blaming anyone in the room.
SShow the replay. Same story, new design, better ending?
With the holdout live from month one, the gap shows up in month seven instead of never. Spend on offers falls from $1,036,800 to about $195,000 over the same six months, and the program's own math finally works in its favor.
A replay that only saves money is thin. This one also gives Shirin something to say when someone asks "saved compared to what."
Hand sketched numbered icon list titled FLIPS, the five letters. Five rows: F, find the person, Shirin, Head of Retention Analytics. L, locate the habit, the Friday sample audit, 25 accounts checked by hand. I, this row in red-orange, identify the flip, audits the sample or trusts the number whole. P, pinpoint the old decision, the audit line got cut from the monthly review. S, show the replay, the holdout goes back on the same dashboard.
The memory aid. Five rows, one color break, at the step that actually decides the whole answer.
Where six months of Save Queue spend actually went
$1.04M $0.5M 0 $1,036,800 Spent on offers $226,000 True value protected -$810,800 Net program cost
What went outWhat the offer actually causedThe real result
Coverage read 98 percent the whole time this bar was getting worse. Usage and ROI were never the same chart.

Three things worth stating directly, since this is where the real judgment sits. The alternative Shirin considered and rejected was simpler: just raise the flag threshold from 72 to something stricter, hoping fewer, more confident flags would fix the spend problem on their own. It lost, because a flag can be confident and still not need a $270 offer to keep the account; only a holdout tells you that, and a higher threshold alone never answers it. The AI-specific failure worth naming is the blind spot in what the model was trained to see: Holdline reads in-product usage decline well and reads almost nothing else, so it stays confidently silent about the four in ten churners a support ticket or a contract renewal date would have flagged instead. The guardrail is the running holdout itself, plus a standing review of which reason codes the false negatives share, feeding back into the next retrain. And the trade-off is real: narrowing the full discount to only the highest-confidence segment protects nearly the same revenue for a fraction of the cost, but it also means a handful of real, lower-confidence risks get the cheap nudge instead of the strong offer, and some of them will still leave. That's accepted on purpose, because the alternative, the expensive offer for the whole queue, is what cost $810,800 in the first place.

And if you want to be sure it really works, try it somewhere else

Same five letters, a walk-in freezer instead of a subscription database, and a different flip family entirely: this time the person doesn't stop checking, he stops being asked to.

ChillGuard is a predictive-maintenance tool built for commercial refrigeration. Frostgate Provisions, a regional grocery chain, runs it across the compressor units in about 60 stores. Every unit gets scored daily for failure risk, and anything flagged pushes an automatic work order onto a technician's route.

Hand sketched comparison diagram titled Same five letters, only the I changes. Left panel, a person icon labeled Shirin, Cascadeworks, caption I: audits a sample, then trusts the number whole, over-trust. Right panel, a person icon labeled Piotrek, Frostgate Provisions, caption I: signs off every dispatch, then lets the model dispatch alone, delegation.
Same method, a different flip family entirely. Only the I column actually changes.

Piotrek Iwuchukwu runs field service for Frostgate's northern district, 40 technicians across 60 stores. Before ChillGuard, every emergency work order over a set cost needed his personal sign-off, checked against the unit's service history, because the district's emergency-dispatch budget was thin enough to matter.

The decision Piotrek would take back ChillGuard shipped one flat output, dispatch: yes, with no reason code visible to anyone before the truck rolled. Nobody downstream could sanity-check a flagged order, so nobody did.

When ChillGuard's flagged volume climbed, so did a genuinely good-looking usage number: jobs closed per week, up nearly every month. Leadership pushed Piotrek to stop signing off personally and let the model dispatch on its own, the same way Cascadeworks let a coverage number replace an audit. He handed the decision down. That's the flip here, not over-trust: he goes from approving every emergency order himself to letting the model dispatch with nobody checking at all, and there's no setting in between.

Six months later, work orders closed were still climbing. The count of real, confirmed mid-shift compressor failures avoided wasn't. Piotrek pulled the reason codes behind every flagged order and found one, "temperature drift," driving most of the volume, and, it turned out, most of the false alarms.

Same rank, different lever: he takes the decision back, not the whole system. Any order flagged on temperature drift alone now routes through a two-minute phone check with the store before a truck gets sent. Every other reason code still auto-dispatches, same as before. Work orders fall by about 22 percent. Confirmed failures avoided hold flat. The 22 percent that disappeared were never going to prevent anything.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: route only the model's shakiest reason code through a human check, everything else keeps auto-dispatching.
Cost: no budget this quarter to build the reason-code breakdown. Ship the cheap version first, track "orders closed" and "confirmed failures avoided" as two separate numbers on the same weekly report, even before you can explain why they've split.
The model got better, for real: say ChillGuard's temperature-drift accuracy doubles overnight. The fix barely moves. You don't know it's doubled until the false-alarm rate on that code actually drops in your own numbers. Until then, the phone check stays.

Where people run it wrong.
They watch jobs closed and call it proof, without ever counting confirmed failures avoided as its own separate number.
They treat every reason code a model outputs as equally trustworthy, when usually one or two are doing most of the damage.
They read one good week as proof the process works, instead of proof it got lucky once.

How to use it live. Ask the split question before naming a fix: "Is usage counting people touching the tool, or counting the outcome the tool's supposed to cause, because those are two different numbers, and only one of them tells you if it's working." That buys you room to actually answer, instead of guessing at a fix on the spot.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust: checks sometimes, then stops checking at all, usually because the change looks like an improvement. Here, a climbing usage number is the "good news" that triggers it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Shirin Kestleman, Head of Retention Analytics at Cascadeworks, who built the company's whole retention-reporting stack from a spreadsheet before Holdline ever existed.
3 · THE HABIT
What did they stop doing because it worked?
Tap to flip
ANSWER
Auditing a random sample of flagged, "saved" accounts every Friday to confirm they were real churn risk and not false flags.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Audits the sample every week, or trusts the coverage number completely and never checks again. No middle setting; by month six it was fully off.
5 · THE OLD DECISION
What decision would Shirin take back?
Tap to flip
ANSWER
Cutting the "real save or false flag" audit line item from the monthly retention review to shorten a meeting that kept running long.
6 · THE NUMBER
Fill in the blank: treated accounts churned at ___%. Held-back accounts with no offer churned at ___%.
Tap to flip
ANSWER
8.6% and 10.1%. A true lift of only 1.5 points on the exact population the offer was supposed to be saving.
7 · THE REPLAY
Same nine months, new design, what changes?
Tap to flip
ANSWER
The holdout runs from month one, the gap surfaces in month seven, and spend falls from $1,036,800 to about $195,000 while true value protected barely moves, about $15,000 ahead instead of $810,800 behind.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and which flip family?
Tap to flip
ANSWER
ChillGuard, Frostgate Provisions' refrigeration predictive-maintenance tool. The flip family is delegation: a manager hands a decision to the model, then has to take it back.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Shirin stop running the Friday audit by month six?
  • A. Holdline's accuracy dropped and the audits stopped being useful.
  • B. The audit line item got cut from the monthly review after the sample kept coming back clean.
  • C. She was reassigned off the retention team.
  • D. A new privacy policy required the audits to stop.
Show hint
Look at what happened in the monthly retention review meeting, in the story's "the decision that opened the door" paragraph.
Show answer
B. The meeting kept running long, so the audit line item got cut to shorten it, and replaced with a single usage slide instead.
True or false
2. True or false: the Save Queue's 98 percent coverage number was evidence that Holdline was working.
  • True
  • False
Show hint
Check what coverage actually measures, versus what the holdout test measured.
Show answer
False. Coverage only measures how many flagged accounts got touched within 48 hours. It says nothing about whether the touch changed anything; only the holdout comparison measured that.
Fill in the blank
3. The true incremental lift from the retention offer, measured by the holdout test, was only ___ percentage points.
Show hint
Subtract the treated churn rate from the held-back churn rate, both stated in the story.
Show answer
1.5 percentage points. 10.1 percent (held back) minus 8.6 percent (treated). Small, real, and far short of what the spend assumed.
Short answer, name the reversal
4. What old decision would Shirin take back, and why did it make sense at the time?
Show hint
Look at "the choice I would take back" in Let's learn, and "the decision that opened the door" in the story.
Show answer
Model answer: Cutting the weekly sample audit from the monthly retention review, to bring an over-long meeting back down to thirty minutes. It made sense because the audit had come back clean for two straight quarters, so nobody in the room thought they were cutting anything risky.
Short answer, apply it yourself
5. Think of a tool you use yourself that reports "usage" as its main success number. Name one way that number could climb while the tool actually does nothing for you.
Show hint
Look for a number that counts an action being taken, not the outcome the action was supposed to cause.
Show answer
Model answer: A budgeting app might report "transactions categorized" climbing every month, while actual overspending never changes, because the app is good at sorting money, not at stopping you from spending it.
Short answer, work the number
6. If the holdout test had found treated and held-back accounts churning at the exact same rate, what would that tell you, and what should happen to the program?
Show hint
Think about what a zero-point gap means for the offer's causal effect, not just its usage numbers.
Show answer
Model answer: A zero-point gap would mean the offer has no true causal effect at all. Every dollar spent on it would be pure cost with nothing coming back, worse than what actually happened here, and the offer would need a full redesign or should stop for that segment entirely.
Before you close the answer
Why this works
Tests whether you can tell a usage number apart from a causal one, and whether you'd actually build the missing control group instead of just noting the gap. Most candidates stop at noticing.
Follow-up traps
"Isn't running a holdout unfair to the customers you deliberately don't help?" Response: it's temporary and small, 10 percent of one segment for eight weeks, and it's the only honest way to find out if the other 90 percent is actually being helped. Refusing to measure doesn't protect anyone, it just hides the cost.

"What if the holdout group is just unlucky that quarter?" Response: run it on a rolling basis instead of once, and size it against the base churn rate so an eight-week window holds enough accounts that one unlucky month can't swing the number alone.
If pressed
The four-in-ten churners Holdline never flagged at all is the deeper problem: the model was trained mostly on in-product usage decline, so it's blind to a champion leaving or a budget cut at the buyer's company, signals that show up in support tickets and contract metadata, not login counts.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more