ConceptIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #3

Why is usage a weak success metric for an AI feature?

The direct answer
Usage counts how many times a person poked the model, not whether what came back actually held up. On an AI feature those two numbers can move in opposite directions at once: a model that starts breaking a rule it used to keep gets used more, not less, because every failed try drives another one. Watch the plan a person keeps without a fight, split by whether the app already knew their preferences, not the raw count of times someone hit generate.
Do this, in order
  1. Stop reporting a raw trigger count, on its own, as the success number.Why: it climbs the same way whether people love the product or are stuck fighting it.
  2. Recut the number by cohort before trusting the average.Why: an average can blend a flat, healthy group with a group that is quietly failing, and hide the second one completely.
  3. Rule out an instrumentation bug and rule out honest exploration before naming a real cause.Why: skip this step and you either call a tracking glitch a product win, or call a real product win a crisis.
  4. Name the specific AI failure behind the extra usage: a retry, a forced step, or a workaround habit.Why: each one needs a different fix, and guessing wrong spends a quarter fixing the wrong thing.
  5. Run the one evidence check that separates real engagement from inflated usage before changing anything.Why: it is the cheapest way to be sure, and it is the line that actually wins the argument in the room.
  6. Only once the real cause is confirmed, decide whether to fix the model or fix the workflow around it.Why: a permanent fix built on an unconfirmed cause usually gets rebuilt once the real one turns up.

How to answer this, stage by stage

Nobody is grading whether you can say "usage is a bad metric." They are grading whether you can catch a number lying to you before you act on it. Eight moves get you there.

1
Scope it to one real product and one real number everyone was already proud of
Say it like this
"Let's ground this. Larderly is an app that builds a week of dinners and a matching grocery list from what a person likes, what they can spend, and what is already in their kitchen. Iwo Callow runs product for it, and the one number the whole team was proud of was plans generated per active user per week."
Why this works
Grounds the diagnosis in a real product and a real dashboard number before naming a single suspect.
2
Say the diagnostic rule out loud, before naming a cause
Say it like this
"Here's how I'd frame it. A number going up is not proof of anything by itself. Before I trust it, I rule out the boring explanations first, a tracking bug, then honest growth, before I go looking for a real cause underneath it."
Why this works
States the method, rule out then narrow, before touching this specific app, so it reads as a rule you could reuse anywhere.
3
Reframe the question before answering it
Say it like this
"The real question isn't 'why did usage go up.' It's 'does this number count the thing that helps someone, or does it count anyone who touched the model, for any reason at all, including a reason that means the model failed them.'"
Why this works
This is the reframe. It separates a trigger count from an outcome count before a single figure gets debated.
4
Walk the timeline, mark what shipped, including the thing that looked like a win
Say it like this
"Ten weeks before this number climbed, Larderly shipped a bigger recipe pool to stop plans from repeating so much. Everyone read the launch notes as good news. The usage climb started the very same week."
Why this works
This is the T step. The real cause almost always ships weeks before the metric moves, and it often wears the clothes of an improvement.
5
Recut the number by cohort before believing the average
Say it like this
"Split it by new users and returning users. New users kept about two out of every three plans without touching regenerate, same as always. Returning users, the ones who had already told Larderly about an allergy, dropped from about seven in ten kept to under four in ten."
Why this works
This is the R step. A twenty seven point average drop turned out to be a flat line for one group and a cliff for another.
6
Rule out the boring explanations before trusting the real one
Say it like this
"First I'd check the event log for double counting a click. It wasn't that. Then I'd check whether the extra generations landed on new meal slots, which would be a genuinely good sign, or the exact same slot again, which would not be. It was the same slot, almost every time."
Why this works
This is the A step, the one most candidates skip. Ruling out the innocent answer is what makes the real one believable.
7
Name the confirmed cause, and the ones you checked and cleared
Say it like this
"The confirmed cause is retries. The bigger recipe pool got worse at holding onto a stored 'never include' flag, so returning users kept hitting a plan that broke their own rule and had to ask again, two or three times, for one they could actually cook. I also checked whether the app forced a fresh plan before unlocking the list, it doesn't, and whether people had quietly given up and started writing their own list by hand, which happens sometimes, but nowhere near enough to explain a jump this size."
Why this works
This is the C step. Naming the causes you ruled out, not just the one you confirmed, is what makes this a real diagnosis instead of a lucky guess.
8
Close on the one check that ends the argument
Say it like this
"The test I'd run: kept-without-a-retry rate, returning users only, week over week since the swap. If that number holds flat while total generations climb, the app is genuinely more loved. If it drops while generations climb, the extra usage is the model failing quietly, and the raw count is lying to the dashboard."
Why this works
This is the E step, the single query that settles the debate in the room instead of two people trading opinions about what usage "really means."
If you remember one thing A trigger count and an outcome count look identical on a dashboard. Only a recut by cohort, and one honest check for retries, tells you which one you are actually looking at.

Let's learn

Every week, before any app got involved, someone sat down with about sixty five minutes and a notepad to plan five dinners and turn them into a grocery list, crossing something out, adding it back, checking what was already in the fridge.

Larderly does that job now. Tell it what you like, what you can spend, and what is already sitting in your kitchen, and it builds a week of dinners with a matching grocery list in under a minute.

Knowledge spark: what does "kept without a retry" mean? A plan a person sends straight to their grocery list, in the same sitting it was made, with no regenerate and no manual edit. It is the plain stand in for "this one actually worked."

In its first months, the average active user asked Larderly for about 1.3 plans a week, mostly one, sometimes a quick swap of a single dinner. Seven plans out of ten went straight to the grocery list with no changes at all.

Ten weeks ago, Larderly shipped a bigger recipe pool, meant to stop the same twelve dinners from showing up every single week. Since then, the average active user asks for 2.1 plans a week, a jump of sixty one percent. Leadership put that number straight on the dashboard as proof the change had worked.

Plans generated per active user, and the rate kept without a retry, week by week
Bigger recipe pool ships 80% 60% 40% 73% 71% 66% 58% 49% kept 44% kept 1.2 plans/wk 1.3 plans/wk 1.5 plans/wk 1.7 plans/wk 1.9 plans/wk 2.1 plans/wk wk -2 wk 0 wk 2 wk 4 wk 7 wk 10
Two lines of the same story. The italic numbers above each point are how many plans people generated, climbing the whole time. The bold percent below each point is how many of those plans got kept without a retry, falling the whole time. The dashboard only ever showed the top line.
The extra plans were not extra love. Split those same users by whether they had already told Larderly about an allergy or a dislike, and the real drop landed almost entirely on the people the app was supposed to already know.
Kept-without-a-retry rate, new users versus returning users
Before the swap After the swap 67% 65% New users 73% 38% Returning users
New users barely moved, sixty seven percent down to sixty five, well inside normal week-one exploring. Returning users, the ones with a stored allergy or dislike on file, fell from seventy three percent to thirty eight. The average of seventy one down to forty four hid this completely.

At its worst, this almost cost Larderly the thing that was actually working. Leadership was ready to double the budget for the bigger recipe pool, reading the climbing count as proof people loved it more. Meanwhile, word of mouth signups were quietly going soft, because the people telling their friends about Larderly were also the ones typing "have to regenerate three times before it remembers I don't eat shellfish" into a one star review.

The decision that mattered The team's whole dashboard carried one number, plans generated per active user per week, and nothing next to it asking whether the plan survived being looked at by an actual fridge and an actual allergy.

The choice I would take back is not shipping the bigger recipe pool. It's putting one raw trigger count alone on the team's main dashboard, with nothing beside it checking whether the thing it triggered actually held up. That was a fine choice back when the app was small enough that a person could read every support message by hand. It stopped being fine the day ten thousand people a week started quietly retrying without ever filing one.

What I would leave alone: a new user generating two or three plans in their first week, while Larderly is still learning what they like, is not the same problem. That is normal, healthy poking around, and it should never get folded into the same alarm as a returning user regenerating for the fourth time on a rule the app should already know.

The lesson: a number that only counts how many times someone poked the model can climb for the best reason there is, or the worst one, and it looks exactly the same on a dashboard either way.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how ordinary the meeting was where nobody had an answer.

Iwo Callow has a habit that makes them slightly annoying in meetings: whenever a chart is climbing and everyone else is happy, they ask what exactly is being counted, before they clap.

For its first year, Larderly earned that habit a rest. The app grew almost entirely on word of mouth, one grocery list photographed and shared at a time. Support tickets stayed low. Sunday nights became the app's best hour, people opening Larderly after dinner, planning the coming week in about ninety seconds, and watching the grocery list sync to their phone before they even stood up from the couch.

The habit thinned in three small beats, none of them loud enough to notice on their own. First, a support message came in from someone who had marked shellfish as a hard no, asking, politely, whether the app "ever actually remembers that." It got tagged a one off and closed. The following week, almost the same question arrived from someone else, phrased almost the same way. Also closed, also filed as a one off. And through both of those tickets, plans generated per active user kept climbing, so nobody connected either one to the chart everyone was already proud of.

The trigger came from Costas Sonnen, two weeks into the job as Larderly's first dedicated data analyst, sitting through his first cohort review. He looked at the sixty one percent jump on the slide and asked one plain question: are we sure people love hitting regenerate, or are we just sure they're not getting what they want the first time? Nobody in the room had an answer ready.

We had not found people falling in love with the new recipe pool. We had found people arguing with a memory the app was supposed to keep for them.

Iwo spent the rest of that afternoon and most of the next morning pulling the raw event log apart by hand, about a day and a half in total, splitting every extra generation by whether it landed on a brand new meal slot or the exact same slot a person had already tried. It landed on the same slot, over and over.

Hand sketched comparison diagram titled Three reasons a trigger count can climb. Three panels. Left, a red and grey gauge icon labelled Retry-driven, caption Plan breaks a rule, user hits regenerate two to three times. Middle, a plain grey box labelled Forced step, caption App requires a new plan before the list unlocks. Right, a grey box with a white question mark labelled Workaround habit, caption Tapped from habit, then ignored, list written by hand.
Only the first panel is colored in the real chart. Larderly's own numbers point at retries, not a mandatory step and not a habit done out of ritual and then ignored.

The original call happened eighteen months earlier, in a room with Iwo and two engineers and no data team at all. Plans generated per active user was the one number anyone could pull from the database in an afternoon, so it went on the dashboard, and it stayed there because it kept climbing and nobody had a reason to argue with a chart that only ever went up.

Three weeks after Costas's question, Larderly shipped a held out check that any future model swap has to clear before it ever reaches live traffic, a set of two hundred flagged allergy and dislike cases the plan is not allowed to break, tested before shipping, not after. Returning user kept-without-a-retry rate climbed back to sixty eight percent inside that same three weeks. Total generations settled to about 1.6 a week, still higher than the original 1.3, this time for an honest reason, people actually swapping in a dinner they wanted, not fighting to get one they could eat.

What I would tell myself, back in that first dashboard meeting eighteen months earlier: the easiest number to pull is never automatically the right one to trust, and a chart that only ever climbs is exactly the one worth a second look, not the one you stop checking.

Reading the climb correctly: TRACE against Larderly's own chart

This is a question about whether a healthy-looking number is actually healthy, so TRACE fits, a rule-out-then-narrow method, not a design or a tradeoff framework.

T
Timeline. When did it start, and what shipped right before.
The generation count started climbing the same week the bigger recipe pool shipped, ten full weeks before anyone thought to ask why.
R
Recut. Slice it, do not just average it.
New users held flat at roughly two in three plans kept. Returning users fell from about seven in ten kept to under four in ten.
A
Assume nothing. Rule out the boring explanations first.
Checked the event log for double counting, checked whether extra generations were new meal slots being explored or the same slot retried. It was the same slot.
C
Cause candidates. Name three, not one.
Retry-driven, confirmed. A forced regeneration step, ruled out, Larderly never required one. A workaround habit of writing the list by hand instead, present, but far too small to explain a jump this size.
E
Evidence test. The one check that ends the argument.
Kept-without-a-retry rate, returning users only, tracked week over week since the swap. One query, and the raw count stops being able to hide anything.

Two things worth saying here, since this is exactly where an AI PM question earns its name. First, the alternative most people reach for is quietly rolling the recipe pool back to the old, narrower one, so the kept rate recovers and the whole mess disappears from the dashboard. That got rejected. Rolling back does not fix anything, it just swaps which failure nobody can see, the same way muting a smoke alarm is not the same as putting out a fire. Second, the real fix could never be "the model must never break a stored rule again." Nobody can promise that about a model. The bar has to be a threshold: the new recipe pool has to clear ninety eight percent adherence on a held out set of two hundred flagged allergy and dislike cases before it ships, checked again after every future swap, not a promise of zero mistakes forever.

Knowledge spark: what is a held out set? A pile of tricky, already checked cases the model never trained on, kept aside just to test it before anything ships. If a plan breaks even one flagged allergy case in that pile, it does not go live.

That trade is real too. The bigger recipe pool was cheaper to run per call and it genuinely fixed the repetition complaints, but it was worse at holding onto a stored rule. Larderly ended up paying for more generations to land one a person could actually use, so the true cost per kept plan went up, even though the cost per single call went down.

And if you want to be sure it really works, try it somewhere else

Same five checks, a completely different kind of business, so the method proves itself instead of repeating a story you happened to prepare.

Kellerman Field Services runs RouteCast, an AI feature that builds each technician's daily stop order and drive route every morning before their first job.

T, timeline. RouteCast switched to a faster route-planning model to cut morning planning time from four minutes to under thirty seconds. Routes generated held flat at one per technician per morning, but "regenerate route" taps per technician per week climbed soon after the swap.
R, recut. Yannick Deighton, who runs product for RouteCast, split the tap count by tenure. Technicians under six months on the job kept trusting the generated route, taps stayed flat. Technicians with five or more years climbed sharply.
A, assume nothing. Ruled out that veterans simply had more stops that day, stop counts were flat across both groups. Ruled out a logging bug double counting a tap.
C, cause candidates. A retry, the route proposed back to back stops on opposite ends of the county inside one hour. A forced step, real here, RouteCast requires a generated route before a technician can mark the first stop started. A workaround habit, confirmed, veterans generated the route, glanced at it, then called dispatch on the radio and built their own from memory anyway.
E, evidence test. Percent of generated routes actually driven start to finish with no radio call to dispatch, split by tenure.

Same shape, different stakes At Larderly, the hidden failure was a plan breaking someone's own allergy. At RouteCast, it is a technician quietly trusting their own memory over a route with a bad drive time estimate baked in. The rank does not change: a trigger count never proves the trigger did its job.
Hand sketched metaphor scene titled What the route counter sees, and what the technician actually does. Left, a green document icon labelled Route made, caption Counted as usage, looks healthy. A red VS in the middle. Right, a person icon in red labelled Drives own way, caption Calls dispatch on the radio instead.
The counter on the left only ever sees a route get made. It never sees the technician on the right pick up the radio five minutes later.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule, a trigger count and a kept count are not the same thing, then name what you would recut by.
Cost: engineering says a held out set of real routes is too expensive to build this quarter. Don't skip the check, shrink it. Hand pull twenty flagged bad routes and check adherence by eye, slower, but still real evidence.
The model got better, for real: say RouteCast's overall drive time accuracy jumps twenty percent. That still doesn't prove the workaround habit went away. A genuinely better model can still be ignored by someone who already learned, the hard way, not to trust it.

Where people run it wrong.
They treat any rising trigger count as proof the feature is loved, without ever asking what a person did right after triggering it.
They fix the dashboard instead of the model, reporting a friendlier number rather than chasing why the honest one moved.
They wait for a complaint before checking, but a retry or a workaround habit files no ticket, a person just quietly starts doing more work around the tool.

How to use it live. Say the rule before naming a single cause: "A trigger count and an outcome count are not the same thing on an AI feature, and before I trust either one going up, I want to know which one actually moved." That buys the room to give a real diagnosis instead of reciting "more usage is good" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about whether a climbing usage number is telling the truth, and why?
Tap to flip
ANSWER
TRACE. It is a diagnosis framework built to rule out the boring explanations first, then narrow to the real cause, not a story about a habit with two settings.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Iwo Callow, product lead for Larderly, an AI meal planning app, with data analyst Costas Sonnen asking the question that started the recut.
3 · THE FALSE SIGNAL
What number did Larderly's dashboard show that looked healthy?
Tap to flip
ANSWER
Plans generated per active user per week, up from 1.3 to 2.1, a sixty one percent jump, read by leadership as proof a recent change had worked.
4 · THE RECUT
What did splitting that number by cohort actually show?
Tap to flip
ANSWER
New users stayed roughly flat, about two in three plans kept without a retry. Returning users, who had already set an allergy or dislike, fell from about seven in ten kept to under four in ten.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Putting one raw trigger count alone on the team's dashboard, with nothing checking whether the plan actually held up. It made sense eighteen months earlier, when it was the only number three people could pull from the database in an afternoon.
6 · THE NUMBER
Fill in the blank: plans generated per active user rose from ___ to ___ a week, while returning users' kept-without-a-retry rate fell from about ___ to ___.
Tap to flip
ANSWER
1.3 to 2.1 plans a week; about 73 percent down to 38 percent for returning users.
7 · THE REPLAY
Same climbing chart, new design, what changes?
Tap to flip
ANSWER
A held out check of two hundred flagged allergy and dislike cases has to clear before any future model swap ships. Three weeks after it shipped, returning user kept rate climbed back to 68 percent, and generations settled near 1.6 a week for an honest reason.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the confirmed cause there?
Tap to flip
ANSWER
RouteCast, the daily technician routing tool from Kellerman Field Services. The confirmed cause is a workaround habit, veteran technicians generate the route, then call dispatch and build their own anyway.

Check yourself Score: 0 / 0

True or false
1. True or false: Larderly's climbing "plans generated per active user" number, on its own, was enough evidence that people liked the app more after the recipe pool change.
  • True
  • False
Show hint
Check what the cohort recut showed once the average got split apart.
Show answer
False. A raw trigger count does not say whether a person kept what got made. Recutting by cohort showed most of the climb was returning users retrying, not new love for the product.
Fill in the blank
2. Returning users' kept-without-a-retry rate fell from about ___ percent to about ___ percent in the ten weeks after the recipe pool shipped.
Show hint
Check the grouped bar chart in "Let's learn."
Show answer
73 percent; 38 percent. New users barely moved in the same window, which is what made this a cohort problem, not a whole-product problem.
Multiple choice
3. Why did ruling out an instrumentation bug and honest exploration matter before naming retries as the cause?
  • A. Because a data team always has to check code before touching a product number.
  • B. Because without ruling those out first, a tracking glitch or genuine delight could get misdiagnosed as a serious model failure, or a real failure could get waved away as growth.
  • C. Because Larderly's engineers refused to believe the number without a formal bug report.
  • D. Because the model's own confidence score already flagged every retry automatically.
Show hint
This is the A step in TRACE, assume nothing.
Show answer
B. Skipping the boring explanations means a tracking bug can get read as a crisis, or a real crisis can get waved off as growth. Neither mistake is cheap.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory from eighteen months earlier, not a setting anyone could just adjust.
Show answer
Model answer: Putting one raw trigger count, plans generated per active user, alone on the team's dashboard, with nothing next to it checking whether the plan actually got kept. It made sense when it was the only number three people could pull from the database in an afternoon, long before there was a dedicated data team to build anything better.
Short answer, apply it yourself
5. Think of an AI tool you use yourself, something that answers, drafts, or generates for you. Name one way its own "usage" count could climb for a bad reason instead of a good one.
Show hint
Look for a retry, a required step, or a habit that keeps firing after someone has stopped trusting the output.
Show answer
Model answer: An AI email drafting tool where "drafts generated" climbs because the model keeps writing in a tone the user has to regenerate away from, not because people are actually writing more emails. The trigger count goes up while the number of drafts sent unedited quietly goes down.
Short answer, the number question
6. If Larderly's overall kept-without-a-retry rate had only dropped from 71 percent to 65 percent, instead of 44 percent, would recutting by cohort still matter? Show the reasoning.
Show hint
Think about what an average can still be hiding, even when it barely moves.
Show answer
Yes, it would still matter. A small overall drop can still be hiding a much bigger one inside a single cohort. The average never tells you which group is actually in trouble, only recutting does, and it costs almost nothing to check.
Before you close the answer
Why this works
Tests whether you'll trust a trigger count that's climbing, or go looking for what happens right after the trigger fires. Most candidates stop at the dashboard.
Follow-up traps
"Isn't more usage still generally a good sign, even if some of it is retries?" Response: not on its own. A retry and a delighted repeat use produce the exact same tick on a trigger counter, you cannot tell them apart without recutting by whether the output got kept.

"Couldn't the returning-user drop just be noise, not a real pattern?" Response: check the volume before dismissing it. A twenty seven point swing across most of the weekly actives, held for ten straight weeks, with the extra taps landing on the same slot instead of a new one, is not noise.
If pressed
The held out set that shipped after this used two hundred real cases pulled from actual flagged allergy and dislike tags, not synthetic ones, and it reruns on every future model swap, not just this one, with the release blocked if adherence drops below ninety eight percent on that set.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more