What business impact does an AI feature have on churn, and how would you measure it?
Candor drafts and summarizes performance reviews inside Fenhollow's people-management platform. Hartsbridge Retail Group, a 140-store home-goods chain, rolled it out to every store manager because Candor was the reason Hartsbridge picked Fenhollow's higher tier over a cheaper competitor two years ago. Chiamaka Adebayo owns Candor's product story at Fenhollow. Franklin Vandenburg runs People at Hartsbridge, and he's the one who has to decide, at renewal, whether Candor is still worth the eighteen percent it adds to the bill.
- Track the weekly rewrite rate, not adoption or completion.Why: completion is mandatory paperwork that stays near 100 percent no matter what. Only the share of drafts a manager throws out tells you whether Candor still earns the price Hartsbridge pays for it.
- Tie the rewrite rate to a real dollar, the add-on revenue at risk, before finance asks.Why: a leading indicator nobody has priced is a chart nobody acts on. Pricing it is what turns a metric into a decision.
- Set real thresholds and act on them, don't wait for the number to become a renewal conversation.Why: under 15 percent is healthy, 15 to 30 earns a sample audit, over 30 sustained earns a Customer Success call. Each threshold buys weeks the annual renewal date never gives you.
- Never let "the feature is switched on" count as "the feature is being used."Why: that's the easiest way this exact metric gets gamed, and it hides a worsening account behind a number that can only tick up.
- Only draft when there's enough source material to draft from.Why: a draft built on three thin notes is a guess wearing a paragraph's clothing. The guardrail against a made-up detail is refusing to draft past that point, not catching it after the fact.
- Keep the draft-to-final history, for every review, on purpose.Why: without it nobody can compute the rewrite rate at all, which is exactly the gap that let Hartsbridge's number climb for eight months unnoticed.
How to answer this, stage by stage
Nobody is grading whether you can define churn. They're grading whether you can name a number that would have looked fine right up until the morning the renewal call went badly, and say what you'd have watched instead.
Let's learn
Candor reads a manager's notes on an employee, logged over a review period, and writes a first draft of that employee's performance review, plus a short summary of themes across the manager's whole team. The manager can keep it, edit it, or ignore it and start over.
Before Candor, a Hartsbridge store manager wrote every review of every employee from a blank page, 45 to 60 minutes each. With around nine employees per store and 140 stores, reviews land on a rolling schedule tied to each employee's hire date rather than one big cycle, so roughly 25 to 30 reviews get written across the whole account every week, all year.
With Candor turned on, a manager opens a review and finds a draft already sitting there, built from the notes they logged that period. In the first four weeks after rollout, managers kept the bulk of that draft 92 percent of the time, only 8 percent of reviews got written from scratch instead. Reviews that used to take an hour were taking twelve minutes.
Here's the turn. Ten months later, that 8 percent had become 41 percent. Almost half of every week's reviews were being written as if Candor didn't exist. And the number everyone at Fenhollow was actually watching, the share of reviews submitted on time, never moved. It sat at 97 percent in week one and 97 percent in week forty-four. Completion is mandatory. A manager writes the review whether they trust the draft or not. It just never told anyone which one was happening.
What that costs at its worst is not just a worse review. It's a false sense that everything is fine, dressed up as data. When Candor was still in private beta, Fenhollow's first internal report on it looked at the twelve customers who signed up earliest, all of them already Fenhollow's biggest champions. Every one of them renewed. The report said "100 percent retention on Candor accounts, versus 91 percent baseline," which was true and told nobody anything, because those twelve were never going to leave regardless of what Candor's drafts were like.
That's the second, quieter way the number gets gamed. "Candor-active" meant an admin had switched the feature on for the account, nothing about whether the store managers were actually keeping the drafts. An account can only turn a feature on, rarely off, so a metric built that way can only ever look healthy or flat, never honestly bad, even while the real thing it's supposed to stand for is falling apart underneath it.
Candor's add-on costs Hartsbridge $27,000 a year, on top of a $150,000 base platform contract. At renewal, once Franklin actually looked past the completion number, he asked to drop the add-on and keep the base platform. That's $27,000 gone on one account, and a warning sign on the base contract too, since Candor was the reason Hartsbridge chose the pricier tier over a competitor two years earlier.
What I would leave alone: Candor also lets a manager paste in a messy pile of one-on-one notes and get back a tidy paragraph, purely for their own reference, never submitted as an official review. That path doesn't need this instrumentation. Nobody's trust in it decides whether Hartsbridge renews. It's a convenience, not the thing the contract is actually paying for.
The lesson: a metric that can only improve or hold flat is not measuring health, it's measuring participation. If the number you're watching would look the same whether people trust the tool or are quietly working around it, you haven't found the metric yet, you've found the paperwork.
Now here is the same thing as a story
The short version is above. Read this one when you want to feel why eleven hours of Franklin's own time, not a bad model, is what almost cost Fenhollow the account.
Franklin Vandenburg built Hartsbridge's whole review process himself, six years ago, out of a shared spreadsheet nobody trusted and a filing cabinet nobody opened. By the time Candor arrived, 140 store managers ran reviews the way he'd taught them to, on time, on a schedule, without him having to chase anyone.
The first months with Candor were good ones. Managers who used to grumble about review season started mentioning, almost in passing, that it wasn't so bad anymore. Franklin watched the completion dashboard every Monday morning, the way he always had. It sat at 97, 98 percent, week after week, exactly what he expected from a process he'd built to be reliable.
He stopped opening a sample of the actual reviews to read them. He'd done that for years out of habit, a dozen or so a month, just to get a feel for quality. Once Candor was writing most of the first draft, reading them felt like reading his own instructions back to himself. So he checked the one number instead, and the one number kept saying everything was fine.
The thing that cracked it wasn't a complaint. On a regional call, a store manager mentioned, almost as an aside, that she'd stopped using "the AI thing" a while back, easier to just write it herself. Nobody else on the call reacted. Franklin barely reacted either, at first.
But it stuck with him. A few days later he pulled up her store's last three reviews and read them properly for the first time in months. Then he pulled a few more stores. Then, over the next three weeks, with the renewal call approaching, he read ninety-one reviews by hand, about eleven hours he didn't actually have, because he had no number that could tell him which stores to worry about and which ones were fine. He was checking everything because he could no longer trust that checking nothing was safe.
What he found in those ninety-one reviews wasn't a disaster. It was a pattern. Roughly one in six of the rewritten ones had a specific reason behind them, a claim in Candor's draft that wasn't anywhere in the manager's own notes, an achievement invented out of a gap the model filled the way a confident guesser fills a gap. The rest were softer: a tone that read like a corporate handbook, not like a store manager talking about someone she actually supervised. Once a manager caught the tool inventing one detail, or just sounding wrong once, they didn't go back and check draft by draft after that. They just stopped starting from it.
The decision that let this hide traced back to a short conversation at Fenhollow, more than a year earlier, when the review editor was first being built. An engineer asked whether the system should keep both versions of a review, Candor's original draft and whatever the manager finally saved, or just keep the final one and let the draft disappear. Keeping both meant a bigger, messier schema for a feature that, at the time, had a few hundred users total. The room chose simple. Nobody in it was deciding, on purpose, that Fenhollow would spend a year unable to answer whether Candor was still worth what it charged for it. They were deciding to ship a smaller feature faster, which was the right call, for exactly as long as it stayed true.
Run the same forty-four weeks again, with the history kept and the thresholds in place instead. The reversion rate crosses 15 percent in week fourteen, same as before. This time it's caught the same week, not the same year. Chiamaka's team pulls a sample, finds the hallucinated specifics and the tone mismatch, and ships a grounding fix by week eighteen: Candor now only drafts against material it can point back to, and it says less when a manager has logged less. By week twenty-four the rate is back under 15 percent. Franklin never opens a spreadsheet of ninety-one reviews. He spends zero hours on it. The renewal call runs eleven minutes, and it's about next year's store count, not about whether Candor still works.
One design waits for a person to notice something feels off, on a regional call, months after it started. The other design notices for you, and hands Chiamaka's team a reason to act while there's still a renewal left to save.
What Chiamaka would tell herself, back in that short meeting about the schema: it wasn't really a storage decision. It was a decision that, for a year, nobody would be able to prove whether Candor was still earning what Hartsbridge paid for it, and nobody in the room meant to make that call on purpose.
LEAD: telling a dashboard from a warning
Not a way to dress up "watch engagement" in four letters. LEAD forces you to name the real dollar, find the number that moves before it does, and say out loud how that number gets faked.
And if you want to be sure it really works, try it somewhere else
Same four letters, a regional internet provider instead of a retail chain, and this time the thing that breaks trust isn't an invented achievement, it's an outdated promise.
Firstline is Corvallo's AI feature for helpdesk software. It drafts an agent's reply to an incoming support ticket, pulling from the customer's account history and the company's own policy pages. Baymarsh Broadband, a regional internet provider, runs it across its whole support team. Meskerem Haile runs Support there, and Firstline was part of why Baymarsh signed Corvallo's higher tier in the first place.
L, link: Firstline's own add-on revenue on the Baymarsh account, and the base helpdesk contract it was sold to justify. E, early signal: the weekly full-rewrite rate, tickets where an agent deletes Firstline's draft entirely and writes a new reply from nothing. A, abuse: "sent using Firstline" got counted even when an agent kept only the opening line and rewrote everything after it, a looser definition of adoption that let real reversion hide inside a metric that looked like usage. D, decision: under 10 percent, leave it alone; 10 to 25, sample the rewrites for policy claims that don't match the current policy page; past 25, sustained, flag the account and check whether Firstline's grounding source has gone stale, since that's a different fix than retraining the model.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to it: "Don't watch whether a ticket got answered, watch whether the agent kept Firstline's own words in the reply. That number moves before a renewal does."
Cost: no budget this quarter for a policy-audit pipeline. Ship the cheap version: a support lead spot-checks ten rewritten replies a week by hand, not a full tool.
The model got better, for real: say Firstline's grounding gets fixed and the stale-policy problem goes away entirely. The full-rewrite rate should fall. If it doesn't within a few weeks, trust was never only about accuracy, and it's worth asking what else agents stopped believing along the way.
Where people run it wrong.
They watch average reply time instead of rewrite rate, and reply time can get faster even as trust craters, because a fast rewrite from scratch still beats a slow one.
They let "sent using Firstline" cover a one-line edit, so real reversion hides inside a number that looks like adoption.
They wait for a CSAT survey to say something's wrong, and CSAT is slow and sparse, and by the time it moves the account's decision is basically already made.
How to use it live. Ask which number would have looked perfectly healthy right up until the morning it broke, before you name a metric. That buys real thinking time, and it reframes the whole question before you have to guess at a number.
Two things worth stating directly, since this is where the real judgment sits. The alternative Fenhollow considered and rejected for the E step was the simplest possible one, the percent of managers who ever opened Candor's draft at all. It lost, because open rate sat near 95 percent from week one regardless of what happened after, everyone was curious once, so it carried no predictive power. The AI-specific failure worth naming twice, once per product, is a model stating something false with full confidence, an invented achievement at Hartsbridge, a stale policy promise at Baymarsh, and the guardrail in both cases is the same shape: sample real output against real source material on a schedule, don't wait for a customer to catch it first. The trade-off Fenhollow accepted on purpose: Candor now drafts less often than it used to, refusing to write anything when a manager's notes are too thin to draft from safely, which means fewer employees get the time-saving draft in exchange for every draft that does get generated being one a manager can actually trust.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a manager rewrites the draft just because they're a strong writer who prefers their own voice, not because they distrust it?" Response: that's exactly why the 15-to-30 tier samples drafts before anyone acts on them. A rising rate gets checked against real notes before it's treated as a trust problem, never assumed to be one on sight.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.