What would make you conclude an AI feature has negative ROI despite good usage?
Holdline scores every one of Cascadeworks' 11,400 subscription accounts for churn risk, once a day. Score high enough, and the account drops into a Save Queue: a Customer Success rep gets nudged to send a discount or a free upgrade before the account can cancel. Shirin Kestleman runs retention analytics at Cascadeworks. Nanami Aoyagi joined her team seven months into the program and asked one plain question nobody had an answer for.
- Build a real holdout before trusting any "saved" number.Why: without a group that got no offer, "we saved 90 accounts" is a guess wearing a fact's clothes.
- Measure spend against the incremental lift only, not the whole flagged group.Why: paying for a discount on an account that was staying anyway is pure cost with nothing coming back.
- Check who the model never flagged, not just who it flagged.Why: the churn hiding in the false negatives can cost more than the false positives ever save.
- Put the audit back on the same dashboard as usage, as a running number, not a one-time test.Why: the coverage number climbed for six straight months with nobody re-checking what it actually meant.
- Route the expensive offer to accounts the model is actually confident about, not the whole queue.Why: most flagged accounts never needed the full discount to stay; a cheaper play protects the same revenue for a fraction of the cost.
- Leave the bottom of the queue alone.Why: accounts that never get flagged and never get an offer aren't costing the program anything, so there's nothing there to audit.
How to answer this, stage by stage
Nobody is grading whether you can say "usage isn't the same as value." They're grading whether you'd actually build the missing control group, or just gesture at the idea and move on.
Let's learn
Holdline is an AI tool that scores every subscription account for how likely it is to cancel, and pushes the riskiest ones into a queue a support team can act on before it happens.
Before Holdline, Cascadeworks' 11,400 accounts got flagged for churn risk by gut feel: a rep noticed a slow login streak, or a support ticket that sounded final. Maybe a third of the accounts that actually left had ever been flagged by anyone at all.
Now Holdline scores every account daily, and about 640 accounts a month cross the line into the Save Queue, where a rep sends a retention offer, a 25 percent discount for two billing cycles or a free upgrade, averaging $270 an account. Cascadeworks' average account is worth about $3,900 a year.
Here's the turn. More offers going out, and more of the queue getting touched inside 48 hours, isn't proof the program is working. It's proof people are using it. Those are two different questions, and Cascadeworks was only ever answering the first one.
Over the six months after the sample audit got cut, Cascadeworks sent retention offers to about 3,840 flagged accounts, at $270 each: $1,036,800 total. A holdout test run later found the offer's true, causal effect: about 58 of those accounts were saved because of it. At $3,900 a year each, that's roughly $226,000 of real value, against $1,036,800 spent, a net loss of $810,800, while the dashboard's headline number sat at 98 percent, the best it had ever looked.
The choice I would take back: cutting the weekly sample audit from the monthly retention review to make the meeting shorter. The audit had come back clean for two straight quarters, so it looked like the safe thing to cut.
The lesson: a feature that's popular is not the same as a feature that's working. Usage tells you people picked the tool up. Only a group you deliberately held back tells you whether picking it up changed anything.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a coverage number climbing for six months in a row was the actual problem, not proof against one.
Every Friday afternoon, Shirin Kestleman pulled the same twenty-five rows. Twenty-five accounts Holdline had flagged that week, twenty-five retention offers Customer Success had already sent, and one afternoon to find out whether any of it was real.
She'd built Cascadeworks' whole retention report from a single spreadsheet, five years back, before there was a Holdline to plug in. She knew which numbers lied by omission and which ones told the truth on their own, and she'd never once let a number onto a leadership slide without tracing it back to an actual account first.
Holdline arrived in the spring. It scored all 11,400 of Cascadeworks' accounts for churn risk, once a day, and anything over 72 dropped into the queue Shirin called the Save List. For the first two quarters, her Friday habit didn't change: twenty-five accounts, pulled at random from that week's flagged list, checked line by line against their usage history and support tickets, to see whether the offer had actually landed on someone who was really about to leave. Almost every week, it checked out. Real risk, real offer, real save.
So the sample shrank. Twenty-five became fifteen. Fifteen became eight. Nobody decided this on purpose. It just kept checking out clean, and a Friday afternoon is worth more spent on something that might actually be broken.
By month six, she'd stopped pulling the sample at all. The Save List's own weekly numbers had taken over the job the audit used to do: offers sent, up. Coverage, the share of flagged accounts a rep actually reached within 48 hours, up to 91 percent. It read like proof. It wasn't a number anyone was worried about.
The trigger, when it came, wasn't a bad quarter. It was a question. Nanami Aoyagi joined the retention team in month seven, straight out of a program built around causal inference, and in her second sprint planning she asked the kind of question a newcomer asks because nobody's told her not to: "What's our holdout? The accounts we deliberately don't send an offer to, so we know how many would've stayed on their own?"
Nobody had one. Nobody had ever built one.
Shirin spent the next two months building it properly instead of feeling bad about it. For eight weeks, one in ten newly flagged accounts, 128 out of about 1,280, got randomly withheld from any offer at all, tracked exactly the same as the other 1,152. At the 90-day mark, the treated accounts, the ones who got a discount or a free upgrade, had churned at 8.6 percent. The untouched accounts, the ones who got nothing, had churned at 10.1 percent.
A real gap. Just a small one. A true lift of a point and a half, on the population the whole program existed to save.
And underneath that number sat a worse one. Of every ten accounts that actually canceled that quarter, company-wide, four had never shown up on the Save List at all. Their churn scores had sat quietly under 72 right up until the day they left, because whatever pushed them out, a champion who quit, a competitor chosen at renewal, a budget line cut two departments up, wasn't the kind of thing Holdline had ever been trained to see. It watched product usage. Nobody had ever told it to watch anything else.
The decision that opened the door was small, and it happened in a meeting nobody would call a mistake, even now. Six months in, the monthly retention review kept running past an hour, and the VP running it asked for it back down to thirty minutes. The line item that went first was "sample audit: real save or false flag." It got replaced with a single usage slide. That was a fair trade in the room that day. The audit had come back clean for two straight quarters. Nobody was worried about a number that already looked good.
Run the same nine months again, with the holdout live from month one instead of never. The gap that took a new hire's question to find shows up on its own, in month seven, sitting right next to the coverage number on the same dashboard. Shirin redesigns the offer before the spend gets any worse: the full discount goes only to accounts scoring above 85 with a usage-decline pattern, the exact signal Holdline is actually good at reading. Everyone else in the queue gets a cheap automatic nudge instead, a check-in email and a feature tip, twelve dollars instead of two hundred and seventy.
Program spend over the same six months falls from $1,036,800 to about $195,000. True value protected barely moves, holding near $210,000. For the first time since Holdline launched, the math tips the other way: about $15,000 ahead instead of $810,800 behind.
What I'd tell myself, back in that thirty-minute meeting: cutting the audit line item didn't feel like a decision about the model at all. It felt like a decision about a meeting running long. It just happened to also be the decision that let a $270 offer get mailed to ninety-eight percent of a queue for six months before anyone checked whether it was doing anything.
FLIPS, or the five questions a holdout group answers for you
Not a way to dress up "seems like it's working" in five letters. FLIPS is what forces you to name what a person stopped checking, and why that felt completely fine at the time.
Three things worth stating directly, since this is where the real judgment sits. The alternative Shirin considered and rejected was simpler: just raise the flag threshold from 72 to something stricter, hoping fewer, more confident flags would fix the spend problem on their own. It lost, because a flag can be confident and still not need a $270 offer to keep the account; only a holdout tells you that, and a higher threshold alone never answers it. The AI-specific failure worth naming is the blind spot in what the model was trained to see: Holdline reads in-product usage decline well and reads almost nothing else, so it stays confidently silent about the four in ten churners a support ticket or a contract renewal date would have flagged instead. The guardrail is the running holdout itself, plus a standing review of which reason codes the false negatives share, feeding back into the next retrain. And the trade-off is real: narrowing the full discount to only the highest-confidence segment protects nearly the same revenue for a fraction of the cost, but it also means a handful of real, lower-confidence risks get the cheap nudge instead of the strong offer, and some of them will still leave. That's accepted on purpose, because the alternative, the expensive offer for the whole queue, is what cost $810,800 in the first place.
And if you want to be sure it really works, try it somewhere else
Same five letters, a walk-in freezer instead of a subscription database, and a different flip family entirely: this time the person doesn't stop checking, he stops being asked to.
ChillGuard is a predictive-maintenance tool built for commercial refrigeration. Frostgate Provisions, a regional grocery chain, runs it across the compressor units in about 60 stores. Every unit gets scored daily for failure risk, and anything flagged pushes an automatic work order onto a technician's route.
Piotrek Iwuchukwu runs field service for Frostgate's northern district, 40 technicians across 60 stores. Before ChillGuard, every emergency work order over a set cost needed his personal sign-off, checked against the unit's service history, because the district's emergency-dispatch budget was thin enough to matter.
When ChillGuard's flagged volume climbed, so did a genuinely good-looking usage number: jobs closed per week, up nearly every month. Leadership pushed Piotrek to stop signing off personally and let the model dispatch on its own, the same way Cascadeworks let a coverage number replace an audit. He handed the decision down. That's the flip here, not over-trust: he goes from approving every emergency order himself to letting the model dispatch with nobody checking at all, and there's no setting in between.
Six months later, work orders closed were still climbing. The count of real, confirmed mid-shift compressor failures avoided wasn't. Piotrek pulled the reason codes behind every flagged order and found one, "temperature drift," driving most of the volume, and, it turned out, most of the false alarms.
Same rank, different lever: he takes the decision back, not the whole system. Any order flagged on temperature drift alone now routes through a two-minute phone check with the store before a truck gets sent. Every other reason code still auto-dispatches, same as before. Work orders fall by about 22 percent. Confirmed failures avoided hold flat. The 22 percent that disappeared were never going to prevent anything.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: route only the model's shakiest reason code through a human check, everything else keeps auto-dispatching.
Cost: no budget this quarter to build the reason-code breakdown. Ship the cheap version first, track "orders closed" and "confirmed failures avoided" as two separate numbers on the same weekly report, even before you can explain why they've split.
The model got better, for real: say ChillGuard's temperature-drift accuracy doubles overnight. The fix barely moves. You don't know it's doubled until the false-alarm rate on that code actually drops in your own numbers. Until then, the phone check stays.
Where people run it wrong.
They watch jobs closed and call it proof, without ever counting confirmed failures avoided as its own separate number.
They treat every reason code a model outputs as equally trustworthy, when usually one or two are doing most of the damage.
They read one good week as proof the process works, instead of proof it got lucky once.
How to use it live. Ask the split question before naming a fix: "Is usage counting people touching the tool, or counting the outcome the tool's supposed to cause, because those are two different numbers, and only one of them tells you if it's working." That buys you room to actually answer, instead of guessing at a fix on the spot.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the holdout group is just unlucky that quarter?" Response: run it on a rolling basis instead of once, and size it against the base churn rate so an eight-week window holds enough accounts that one unlucky month can't swing the number alone.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.