What early signal predicts churn from an AI feature?
The number that predicts a lost subscriber isn't how often someone opens a money app. It's whether they still bother telling it when it's wrong.
- Track the abandoned-fix rate, not the open rate.Why: it's the earliest tell. Someone who gives up mid-correction has already checked out weeks before their app opens admit it.
- Make one correction stick for the merchant, not just the one transaction.Why: this is the fix under the fix. A correction that resets itself is a reason to stop correcting at all.
- Put one line of "why" on every nudge that flags real money.Why: a nudge with no explanation gives a person nothing to do except close the app.
- Gate any change to the category model behind a golden set of irregular-income transactions.Why: a model trained mostly on steady paychecks breaks quietly on gig and freelance spending, and that gap stays invisible until it costs a real subscriber.
- Split the signal by income pattern, steady against irregular.Why: the same wrong nudge is a shrug for one group and a reason to quit for the other.
Six moves, if you only get one shot
Nobody is grading whether you can name a metric. They're grading whether you can name the one that would have caught this before it was too late to catch.
Let's learn
Nudgewell is an app that reads your bank feed, sorts your spending into categories, and nudges you when one of them runs hot.
Before it, Talveen Bhoja tracked her own business spending in a spreadsheet, about a hundred and thirty transactions a quarter, sorted by hand for her freelance tax filing. That took her close to five hours each quarter, done at her kitchen table after client work was done.
With Nudgewell, that same tax prep dropped to under an hour a quarter, and her daily check became a thirty second glance before bed. In her first weeks, Nudgewell got about fifteen transactions wrong a week, tagging them oddly because her spending doesn't look like a monthly paycheck. She fixed nearly all of them by hand, about twelve or fourteen a week, checking the receipt before trusting the tag.
Here is the turn. The rising abandon rate was not really about the tags getting worse. Say it plainly: the app's accuracy barely moved. What changed is what Talveen did the second time it got something wrong the same way twice.
At its worst, this cost more than one lost user. Talveen still opened the app most nights, just to glance at the balance, so Ambervale's own churn model, which watches app opens, never flagged her account as at risk. It read her as engaged right up until her subscription lapsed at renewal, four weeks after she'd stopped fixing anything at all.
The choice I would take back is not the nudge itself. It's that a fix she made once never held for the next bill from the same place. That was a fine choice back when Nudgewell mostly served people with one paycheck a month and few repeat oddball expenses. Nobody building that screen was picturing a freelancer's client roster.
What I would leave alone: the nudge for someone on a steady salary who overspends on takeout twice in one week. There, a falling correction rate really does mean the model is doing fine and the person trusts it. That part of the design never broke.
The lesson: a fix that doesn't hold teaches a person not to bother fixing anything. And a person who stops fixing your mistakes is already most of the way to a person who stops opening your app.
Now here's the same thing, slower
The short version is above. Read on for the Thursday a six hundred dollar repair bill quietly cost Ambervale a subscriber.
Ten at night is when Talveen finally sits with her phone, after the shoot's wrapped and the invoice is sent.
She's run her own textile design studio for six years. Client work pays in bursts, one big order some months, almost nothing the next, so every quarter she used to sit at that same table and sort five hours of receipts by hand into a spreadsheet, one column for materials, one for software, one for travel. She never once missed a filing deadline.
Nudgewell arrived in the spring, and for two months it was the best five minutes of her evening. She'd check the day's spending on the way to bed, tap through anything that looked off, and be done before the kettle boiled. In week one it mistagged a nudge or two most days, and she fixed nearly all of them, twelve or so a week, checking each receipt against the tag before she trusted it.
The checking thinned out in three beats. By week three, about eight fixes, mostly quick ones. By week six, four, only the ones that looked plainly wrong. By week nine, one, if that. Nothing had gone wrong. The app was simply right often enough that fixing felt like a chore she no longer needed.
Then, on a Thursday in week five, her laptop died mid project and the repair bill came to six hundred and forty dollars. Nudgewell tagged it "Shopping" and sent a nudge: forty five percent over her weekly budget. She opened the transaction, changed the tag to "Business, equipment," and moved on. Small thing.
Five weeks later, a software license renewed for five hundred and eighty dollars. Same wrong tag. Same nudge. Same forty five percent line, word for word.
This time Talveen didn't open the transaction to fix it. She read the nudge, closed the app, and mostly didn't open it again.
The real cost wasn't a second wrong tag. It was that she stopped believing a fix would hold. She never had a number in her head for how much she trusted Nudgewell. She had a habit, and the habit only had two settings, worth opening or not, and the second wrong nudge flipped it for good.
A year before any of this, when Ambervale built that correction screen, the team made it fix one transaction only, never the vendor behind it. In the room where that got decided, it looked like the safe, simple call. Most of Nudgewell's users were on one monthly paycheck, and repeat miscategorizations were rare enough that fixing them one at a time was plenty. Nobody in that meeting had a freelance client roster in mind.
Run that same Thursday through the fixed design. Talveen taps the wrong tag once, and Nudgewell remembers it for that vendor going forward. Five weeks later, the software renewal lands tagged correctly. No nudge. No forty five percent line. She's still checking the app every night at ten, and by the time her subscription comes up for renewal, Ambervale's own numbers read her account as growing, not at risk.
One design remembered a mistake for one receipt. The other remembered it for her.
What I'd tell myself, before any of this: the day a fix stops holding is the day to ask what it actually taught the app, not just whether the screen looked fixed.
FLIPS, letter by letter
This is a metric question wearing a perturbation question's clothes, but the honest answer only shows up once you find the moment Talveen's behavior actually snapped, so FLIPS carries the weight here.
Two things worth naming straight out, since this is where the real judgment sits. The easy fix on offer was a visible confidence number next to every tag, the kind of fix that looks like it does something. That got turned down on purpose. A number doesn't stop the same wrong tag from coming back, it just gives Talveen one more thing to glance at and still lose faith in. The real risk worth naming by name is silent categorization drift for spending that doesn't look like a steady paycheck, because Nudgewell's model learned its categories mostly from users paid the same amount on the same day each month. The guardrail is a golden set built from freelance and gig income transaction histories, checked by hand, and any change to the category model has to correctly tag at least ninety two of every hundred transactions in that set before it ships, not just the overall average across every user. That check costs something too. Applying Talveen's own merchant rule the moment a transaction lands, instead of overnight in a batch job, adds a small delay and a little more compute to every categorize call. Worth it. It's cheaper than one lost year of her subscription, and far cheaper than the freelancers she talks budgeting apps with at her co-working space.
Run it on a completely different problem
Same five letters, a farm advisory tool instead of a money app, and a flip that runs on what a person feeds the model rather than what they open.
Cultiva is a growers' co-op that runs FieldSense, a tool that answers crop and pest questions from a farmer's written field notes. Coralee Penhale is the agronomist it pays to advise about forty smallholder growers.
F, find the person. Coralee can spot a nitrogen deficiency by leaf color from ten feet away, before FieldSense ever loads.
L, locate the habit. She used to write full field notes into the tool, weeds spotted, pest sightings, odd soil readings, everything, before asking for a recommendation.
I, identify the flip. Feeds the tool her full, messy note, or trims it down to just the clean numbers first. Two settings, no middle, once the messy version starts getting a shrug instead of an answer.
P, pinpoint the old decision. When confidence was low, FieldSense showed a flat "consult a local expert" message instead of asking a follow up question, to keep the first release simple.
S, show the replay. The messy notes kept tripping that low confidence bail out, so Coralee learned to leave the odd details out, and FieldSense started answering more, on cleaner input it never should have trusted as complete.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to the fix: whenever a model's fallback is a flat bail-out message, swap it for one targeted question, for any tool whose whole value is in messy input.
Cost: engineering says the clarifying question can't ship for two months. Don't read a falling fallback rate as improvement in the meantime. It's more likely input getting sanitized than a model getting smarter.
The model got better, for real: say FieldSense's accuracy genuinely rises next season. That still isn't proof the fallback rate should be trusted on its own, a better model just makes it harder to notice growers still trimming their notes out of habit.
Where people run it wrong.
They read a falling "not sure" rate as the model improving, without checking whether the input changed instead.
They fix it with a training session telling people to "give the model more detail," instead of fixing why detail made the tool bail out in the first place.
They wait for a season's yield numbers to notice, instead of comparing note length against outcome every month.
How to use it live. Say the reframe before the fix: "A fallback rate that's falling can mean the model got better, or it can mean people quietly stopped giving it the hard parts." That buys room to give the real answer instead of taking a falling number at face value.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if she abandons a fix because the tag was actually already right, and she just didn't need to change anything?" Response: a tap that confirms a tag as fine and closes counts as completed, not abandoned. Only an edit that gets opened and then discarded without saving counts, which keeps the signal from being triggered by people who are simply satisfied.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #6 How do you build an early warning system for silent quality degradation?
- #7 Describe the relationship between refusal rate and downstream satisfaction.