Critique this metric: number of AI queries per user per week.
A number that used to tell the truth about a customer's account can start lying the day the product finally gets good.
- Stop reading a falling query count as decline. Track the share of answers used with no follow up question instead.Why: this is the real fix, not a number to glance at once a quarter.
- Split that share by how new the account is, thirty days and under versus everyone else.Why: for a brand new account a falling query count still means someone gave up. Ten weeks in, it can mean the opposite.
- Turn off the automatic nudge that fires the moment weekly queries drop by a third.Why: right now it punishes the exact behavior the product exists to build.
- Gate every new model or prompt version behind a golden set of hand checked surveys before it ships.Why: once a person stops asking the same question four ways, that habit is no longer there to catch a quiet drop in quality.
- Chart the one shot answer share next to the raw query count every week, never one without the other.Why: either number alone tells a story that is not true. Together they show whether trust is real.
- Keep the thirty day nudge alive for brand new accounts.Why: that is the one place the old warning sign is still honest.
How to answer this, stage by stage
Nobody is grading whether you can name what is wrong with a metric. They are grading whether you can say what a falling number actually means before you react to it. Seven moves get you there.
Let's learn
What happens when the number a team built to protect a customer's account quietly stops meaning what it used to mean?
Distillery is a tool inside Cobble Insights. It reads a pile of open ended survey answers and hands a product team back the themes hiding inside them, real quotes attached, instead of a spreadsheet nobody has time to read.
Before Bryce Sabatini used it, he ran product research at Palmerhouse, a mattress and bedroom furniture company, and every quarter he read a random sample of about two hundred and twenty survey comments by hand. That took him close to fourteen hours across a week, on top of his real job.
In his first weeks with Distillery, the same report took him under ninety minutes. But he still ran the tool hard, an average of thirty four queries a week while a survey wave was live. He would ask it to show a theme, then ask again with the quotes pulled out, then ask it to compare against last quarter, then rephrase the same question to see if the answer held.
Here is the turn. Ten weeks in, Bryce was down to nine queries a week. Say this plainly: that drop was not the problem. Distillery's first answer had started matching what he used to find by checking it four ways, so he stopped checking it four ways. He read one summary and wrote it straight into the deck.
At its worst, this cost more than an awkward email. Cobble's customer health dashboard flags any account whose weekly queries drop by a third or more, two weeks running, as churn risk. It fired for Bryce the same week Palmerhouse's hundred and sixty four thousand dollar contract came up for its renewal review, and the flag landed in front of his own VP before anyone on Cobble's side thought to ask why.
The choice I would take back is not the alert itself. It is that nobody put a date on it. Cobble built that rule in Distillery's first year, when a falling query count really did mean someone had given up on the tool. Nobody ever came back to ask whether that was still true once the product got better.
What I would leave alone: the same alert, for any account in its first thirty days. There, a falling query count still usually means the old thing, someone tried it twice, did not trust what came back, and quietly stopped. That part of the metric never broke.
The lesson: a number that used to tell you the truth does not stay true forever. The day people finally trust a tool is exactly the day the number you built to protect them can start lying, and nobody tells you it happened.
Now here is the same thing as a story
The short version is above. Read on for the Thursday morning this nearly cost Palmerhouse's renewal.
Bryce has run product research at Palmerhouse for four years. Hand him an angry review and he can tell you inside ten seconds whether it is really about the mattress or about the delivery driver who showed up late. That instinct is the whole job.
When Distillery arrived, the good months looked like this: Monday mornings, coffee still hot, he would open a fresh batch of survey answers and start pulling themes by nine. The tool was fast, and for the first few weeks he still checked it hard out of habit, the same instinct that made him good at the job in the first place.
The habit thinned out in three beats. By week two, he was down to two or three follow up queries per theme instead of four, mostly just checking the quotes. By week five, one, a single rephrase, just to hear the same answer twice. By week eight, none. He would read the first summary Distillery gave him and start typing it into the slide.
There was not one morning where it changed. It just thinned out, the way a habit does when nothing bad ever happens to bring it back. The only morning anyone noticed was a Thursday at 9:40, when an email landed in Bryce's inbox while he was two slides into a board deck due at eleven: "Your Distillery usage is down. Here are five prompts to try."
He did not need the prompts. He knew that. But now he had a choice, ignore an automated flag his own VP could also see, or spend twenty minutes running queries he did not need, just to make the number look healthy again. He ran them. The deck slipped.
The real cost was not twenty minutes. Cobble's customer success team saw the same flag and booked a check in call with Palmerhouse for the following week, the same week the renewal conversation was supposed to start. Bryce's VP asked him, in front of the deck he had just finished late, whether Palmerhouse was "still getting value" out of a tool that had just saved him twelve hours a week.
A year earlier, when Cobble's growth team built that alert, the logic was sound. In Distillery's first months, a query count that dropped by a third really did predict a churned account, because back then a drop meant someone had tried the tool twice, gotten a bad answer, and quietly gone back to reading comments by hand. Nobody in that meeting pictured a power user tripping the same wire from the opposite direction, because at the time, nobody was one yet.
Run the same Thursday through the fixed design. The alert only fires now if a mature account's resolution share drops, not its raw query count, and Bryce's share that week was eighty one percent, well above the healthy line. No email. No twenty lost minutes. His deck is done at 10:35, same as always, twenty five minutes to spare. Two weeks later, the renewal review cites his rising resolution share as an expansion signal, and Cobble's account team pitches a second research seat instead of booking a save call.
One design watched a number that used to matter. The other watched what the number had actually started to mean.
What I would tell myself, before any of this: the day a number stops changing for the reason you built it to catch is the day to go find out what it is actually measuring now.
The five FLIPS moves, and where each one shows up here
This is a critique question, but the honest answer to "what does this metric miss" turns out to be a real flip, so FLIPS is doing the work here, not a formula bolted on afterward.
Two things worth naming directly, since this is where the real judgment lives. First, the easy fix on offer was to just raise the alert's threshold so it fires less often. That got ruled out on purpose: it is a dial, not a decision, it just delays the same wrong alarm instead of fixing what it actually measures. Second, the risk worth naming by name is silent quality drift: the moment people stop triangulating an answer with follow up queries, that habit stops catching a model or prompt update that quietly got worse. The guardrail is a golden set, a batch of past surveys with themes a person already checked by hand, and a rule that a new model version only ships once it matches those hand checked themes at least nine times in ten, checked before rollout, not after a customer notices. That check costs something too. Fewer queries per user already means lower cost to run the model every week, a real saving, but the golden set gate adds a few days before a cheaper or faster version reaches anyone, and that delay is the price of not quietly trading Bryce's trust for a lower bill.
And if you want to be sure it really works, try it somewhere else
Same five letters, a fleet maintenance tool instead of a survey tool, and a flip running the opposite direction, so a rising query count is the one that fools you this time.
Ferrow Fleet runs TorqueSight, a tool that answers diagnostic questions about engine fault codes for repair technicians. Deslin Vashti runs product for it.
F, find the person. A senior mechanic at Ferrow Fleet, fifteen years in, who can diagnose a fault code by the sound of the engine before TorqueSight even loads.
L, locate the habit. He used to run every diagnostic query himself. Once TorqueSight proved reliable, he handed the tool to his apprentice and stopped opening it at all.
I, identify the flip. Apprentice runs the diagnostic queries alone, or the senior mechanic takes them back and runs every one again himself. Two settings: delegated, or reclaimed.
P, pinpoint the old decision. Ferrow's product team built its engagement dashboard to treat a rising query count as adoption working as intended, because early on, more queries always meant more trucks fixed faster.
S, show the replay. The apprentice's answers started drifting on a batch of newer engine models TorqueSight had not seen much of. The senior mechanic noticed and quietly started running every one of the apprentice's queries again himself. Query count per active account doubled. Deslin's dashboard called it the best adoption month all year.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the fix: whatever raw usage number you are watching, pair it with a share that shows whether the work is being trusted, not just repeated.
Cost: engineering says a real resolution share metric cannot ship for two months. Do not read the raw query count alone in the meantime and call it a stopgap. Hold off on any alert built from that number until the real one exists.
The model got better, for real: say TorqueSight's diagnostic accuracy genuinely improves next quarter. That still is not the same claim as "rising queries mean people trust it more." A better model just makes a delegation flip like Ferrow's harder to spot, because the apprentice's raw numbers stay steady too.
Where people run it wrong.
They treat a falling number as decline and a rising number as growth, on reflex, without asking who is actually behind either one.
They fix the metric's mistake with a bigger review board, instead of a second number that would have caught it automatically.
They wait for a renewal call to ask why usage moved, instead of charting the real signal every week whether anyone remembers to look or not.
How to use it live. Say the reframe before naming a single fix: "A usage number that only counts how often someone asks can't tell a person who trusts the tool apart from one who gave up on it, or in this case, apart from a second person quietly doing the same work twice." That buys you room to give the real answer, instead of reciting "vanity metric" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't a team just game single shot resolution rate by making the model refuse follow up questions?" Response: that is what the golden set gate is for. A model change that raises resolution share by getting worse at handling follow ups would fail the golden set's accuracy check before it ever reached a user.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?