Artifact critiqueIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #20

Critique this metric: number of AI queries per user per week.

A number that used to tell the truth about a customer's account can start lying the day the product finally gets good.

The direct answer
Queries per user per week is not a health number, it is a mood ring: it climbs when people do not trust a tool and are checking their own work, and it falls for the same reason, because they finally do. Stop reading a drop as decline. Track the share of answers used with no follow up question instead, split by how new the account is, and only worry when that share falls.
Do this, in order
  1. Stop reading a falling query count as decline. Track the share of answers used with no follow up question instead.Why: this is the real fix, not a number to glance at once a quarter.
  2. Split that share by how new the account is, thirty days and under versus everyone else.Why: for a brand new account a falling query count still means someone gave up. Ten weeks in, it can mean the opposite.
  3. Turn off the automatic nudge that fires the moment weekly queries drop by a third.Why: right now it punishes the exact behavior the product exists to build.
  4. Gate every new model or prompt version behind a golden set of hand checked surveys before it ships.Why: once a person stops asking the same question four ways, that habit is no longer there to catch a quiet drop in quality.
  5. Chart the one shot answer share next to the raw query count every week, never one without the other.Why: either number alone tells a story that is not true. Together they show whether trust is real.
  6. Keep the thirty day nudge alive for brand new accounts.Why: that is the one place the old warning sign is still honest.

How to answer this, stage by stage

Nobody is grading whether you can name what is wrong with a metric. They are grading whether you can say what a falling number actually means before you react to it. Seven moves get you there.

1
Scope it to one real product, one real account
Say it like this
"Let's ground this. Cobble Insights makes Distillery, a tool that reads open ended survey answers and hands a product team back the themes and quotes instead of a spreadsheet. Bryce Sabatini runs product research at Palmerhouse, a mattress and bedroom furniture company, and he has been using it for about ten weeks."
Why this works
Grounds the critique in a real product and a real person before naming a single number.
2
Say your structure out loud
Say it like this
"Here's how I'd take this apart. I want to find what a person actually does with their hands when they start trusting an answer, then ask what a raw count of questions can and can't tell you about that."
Why this works
Two seconds of structure tells the interviewer you have a plan, before you say a single word about the metric itself.
3
Reframe what a falling number can mean
Say it like this
"A falling number here is not obviously bad. It is what happens the moment someone stops needing to check an answer four different ways before they will use it. Raw query count cannot tell the difference between a person who gave up and a person who finally trusts the tool."
Why this works
This is the reframe the whole critique turns on. A candidate who skips this just says "vanity metric" and moves on.
4
Give the one fix, not a category of fix
Say it like this
"So here's what I'd track instead. The share of answers a person uses with no follow up question, split by how new their account is. Watch that number, not the raw count, and only react when it falls."
Why this works
This matches the direct answer. A named, buildable number beats "we'd look at engagement more holistically."
5
Prove it with the failure that actually happened
Say it like this
"Here's what happens without that fix. Bryce's queries dropped from thirty four a week to nine over ten weeks, because Distillery had earned his trust. Cobble's dashboard read that as churn risk and fired an automatic email asking if he needed help, the same week his account came up for renewal."
Why this works
A four sentence failure story does more work here than a paragraph of theory.
6
Say what stays, and what you'd measure going forward
Say it like this
"I'd keep the old alert alive for accounts under thirty days old, where a query drop still usually means someone is stuck. Everywhere else I'd swap it for the one shot answer share, and chart both numbers weekly by account age, so I can tell the two apart before a renewal call, not during one."
Why this works
Shows judgment instead of blanket caution. Not every account gets the same read.
7
Close on the option you ruled out and what it costs
Say it like this
"We looked at just raising the alert's threshold so it fires less often, and ruled that out. It's a dial, not a fix, it just moves the same wrong alarm further down the road. The real cost here is a short delay before a cheaper model version can ship, because it has to clear a quality check first. That's the trade I'd take over losing the one thing that used to catch a bad answer by accident."
Why this works
Naming a rejected option and a real cost turns "track something else" into a defensible decision.
If you remember one thing A number that falls when trust rises and falls when trust breaks cannot tell you which one is happening. Read something else before you react to it.

Let's learn

What happens when the number a team built to protect a customer's account quietly stops meaning what it used to mean?

Distillery is a tool inside Cobble Insights. It reads a pile of open ended survey answers and hands a product team back the themes hiding inside them, real quotes attached, instead of a spreadsheet nobody has time to read.

Hand sketched comparison titled A dial or a switch. Left a dial with a needle labeled DIAL, many settings, more is always better. Right a two position switch labeled SWITCH, trust it or check it, no middle.
A metric built like a dial assumes more is always healthier. Trust does not work like a dial. It works like a switch.

Before Bryce Sabatini used it, he ran product research at Palmerhouse, a mattress and bedroom furniture company, and every quarter he read a random sample of about two hundred and twenty survey comments by hand. That took him close to fourteen hours across a week, on top of his real job.

In his first weeks with Distillery, the same report took him under ninety minutes. But he still ran the tool hard, an average of thirty four queries a week while a survey wave was live. He would ask it to show a theme, then ask again with the quotes pulled out, then ask it to compare against last quarter, then rephrase the same question to see if the answer held.

Ten weeks, and the number Cobble was watching kept dropping
34 21 9 week 0 week 10
Queries per week, Bryce's account. This is the number Cobble's health dashboard actually watches.

Here is the turn. Ten weeks in, Bryce was down to nine queries a week. Say this plainly: that drop was not the problem. Distillery's first answer had started matching what he used to find by checking it four ways, so he stopped checking it four ways. He read one summary and wrote it straight into the deck.

We did not lose an engaged user. We finally built a tool good enough that he stopped needing to check it.
The number nobody's dashboard was built to show
80% 35% Week 0 Week 10 35% 81%
Single shot resolution share: how often Bryce used Distillery's first answer with no follow up query. This climbed while the raw query count fell.
Knowledge spark: what is a single shot resolution share? The share of times a person uses the model's first answer with no follow up question. It is a way to tell "trusts the answer" apart from "stopped asking." A rising share is a good sign. A metric that cannot make that split is just counting how busy someone's week was.

At its worst, this cost more than an awkward email. Cobble's customer health dashboard flags any account whose weekly queries drop by a third or more, two weeks running, as churn risk. It fired for Bryce the same week Palmerhouse's hundred and sixty four thousand dollar contract came up for its renewal review, and the flag landed in front of his own VP before anyone on Cobble's side thought to ask why.

The decision that mattered An alert wired to fire automatically whenever a weekly query count drops by a third or more, two weeks running, with no person reviewing it before it goes out.

The choice I would take back is not the alert itself. It is that nobody put a date on it. Cobble built that rule in Distillery's first year, when a falling query count really did mean someone had given up on the tool. Nobody ever came back to ask whether that was still true once the product got better.

What I would leave alone: the same alert, for any account in its first thirty days. There, a falling query count still usually means the old thing, someone tried it twice, did not trust what came back, and quietly stopped. That part of the metric never broke.

The lesson: a number that used to tell you the truth does not stay true forever. The day people finally trust a tool is exactly the day the number you built to protect them can start lying, and nobody tells you it happened.

Now here is the same thing as a story

The short version is above. Read on for the Thursday morning this nearly cost Palmerhouse's renewal.

Bryce has run product research at Palmerhouse for four years. Hand him an angry review and he can tell you inside ten seconds whether it is really about the mattress or about the delivery driver who showed up late. That instinct is the whole job.

When Distillery arrived, the good months looked like this: Monday mornings, coffee still hot, he would open a fresh batch of survey answers and start pulling themes by nine. The tool was fast, and for the first few weeks he still checked it hard out of habit, the same instinct that made him good at the job in the first place.

The habit thinned out in three beats. By week two, he was down to two or three follow up queries per theme instead of four, mostly just checking the quotes. By week five, one, a single rephrase, just to hear the same answer twice. By week eight, none. He would read the first summary Distillery gave him and start typing it into the slide.

Hand sketched comparison titled The habit that faded. Left an amber funnel labeled Triangulate, ask it four ways before trusting one answer. Right a teal document labeled Trust it, ask once, write it into the report.
Nine weeks apart, same task, same person. One side is not lazier than the other. It is just further along.

There was not one morning where it changed. It just thinned out, the way a habit does when nothing bad ever happens to bring it back. The only morning anyone noticed was a Thursday at 9:40, when an email landed in Bryce's inbox while he was two slides into a board deck due at eleven: "Your Distillery usage is down. Here are five prompts to try."

He did not need the prompts. He knew that. But now he had a choice, ignore an automated flag his own VP could also see, or spend twenty minutes running queries he did not need, just to make the number look healthy again. He ran them. The deck slipped.

The real cost was not twenty minutes. Cobble's customer success team saw the same flag and booked a check in call with Palmerhouse for the following week, the same week the renewal conversation was supposed to start. Bryce's VP asked him, in front of the deck he had just finished late, whether Palmerhouse was "still getting value" out of a tool that had just saved him twelve hours a week.

It was never about how many questions he asked. He never had a number for how much he trusted Distillery. He had a habit, and it had already changed weeks before any dashboard noticed.

A year earlier, when Cobble's growth team built that alert, the logic was sound. In Distillery's first months, a query count that dropped by a third really did predict a churned account, because back then a drop meant someone had tried the tool twice, gotten a bad answer, and quietly gone back to reading comments by hand. Nobody in that meeting pictured a power user tripping the same wire from the opposite direction, because at the time, nobody was one yet.

Run the same Thursday through the fixed design. The alert only fires now if a mature account's resolution share drops, not its raw query count, and Bryce's share that week was eighty one percent, well above the healthy line. No email. No twenty lost minutes. His deck is done at 10:35, same as always, twenty five minutes to spare. Two weeks later, the renewal review cites his rising resolution share as an expansion signal, and Cobble's account team pitches a second research seat instead of booking a save call.

One design watched a number that used to matter. The other watched what the number had actually started to mean.

What I would tell myself, before any of this: the day a number stops changing for the reason you built it to catch is the day to go find out what it is actually measuring now.

The five FLIPS moves, and where each one shows up here

This is a critique question, but the honest answer to "what does this metric miss" turns out to be a real flip, so FLIPS is doing the work here, not a formula bolted on afterward.

F
Find the person. Whose morning is this.
Bryce Sabatini, senior product researcher at Palmerhouse, four years in, could place a complaint by ear before Distillery ever loaded.
In this story: Bryce, not "product teams" in general.
L
Locate the habit. What did they stop doing because it worked.
He stopped asking Distillery to prove the same theme four different ways, quotes, a rephrase, a quarter over quarter check, before he would trust it enough to write it down.
Five weeks to thin from four checks to one, three more to thin to none.
I
Identify the flip. The verb that snaps, two settings, no middle.
Triangulate every answer with follow up queries, or accept one answer and write it down. Once trust builds, he does not drift back to asking four times.
This is the whole critique. A raw count cannot tell you which setting someone is in.
P
Pinpoint the old decision. Which choice only made sense before.
The churn risk alert wired to any big drop in weekly queries, built in Distillery's first year, when that drop reliably meant someone had given up.
Nobody put an expiry date on a rule that stopped being true.
S
Show the replay. Same bad day, new design.
Pair the alert with single shot resolution share, split by account age. The Thursday nudge never fires, the deck ships on time, the renewal call turns into an upsell pitch.
Deck done at 10:35 instead of derailed. A contract flagged "at risk" instead reads "expand."
Hand sketched diagram titled FLIPS, the whole method. A person labeled Bryce in the center with five labeled call outs around him: F, find the person. L, locate the habit. I, identify the flip. P, pinpoint the old decision. S, show the replay.
Five moves, one person at the center. Skip the person and the method turns back into a formula.

Two things worth naming directly, since this is where the real judgment lives. First, the easy fix on offer was to just raise the alert's threshold so it fires less often. That got ruled out on purpose: it is a dial, not a decision, it just delays the same wrong alarm instead of fixing what it actually measures. Second, the risk worth naming by name is silent quality drift: the moment people stop triangulating an answer with follow up queries, that habit stops catching a model or prompt update that quietly got worse. The guardrail is a golden set, a batch of past surveys with themes a person already checked by hand, and a rule that a new model version only ships once it matches those hand checked themes at least nine times in ten, checked before rollout, not after a customer notices. That check costs something too. Fewer queries per user already means lower cost to run the model every week, a real saving, but the golden set gate adds a few days before a cheaper or faster version reaches anyone, and that delay is the price of not quietly trading Bryce's trust for a lower bill.

And if you want to be sure it really works, try it somewhere else

Same five letters, a fleet maintenance tool instead of a survey tool, and a flip running the opposite direction, so a rising query count is the one that fools you this time.

Ferrow Fleet runs TorqueSight, a tool that answers diagnostic questions about engine fault codes for repair technicians. Deslin Vashti runs product for it.

F, find the person. A senior mechanic at Ferrow Fleet, fifteen years in, who can diagnose a fault code by the sound of the engine before TorqueSight even loads.
L, locate the habit. He used to run every diagnostic query himself. Once TorqueSight proved reliable, he handed the tool to his apprentice and stopped opening it at all.
I, identify the flip. Apprentice runs the diagnostic queries alone, or the senior mechanic takes them back and runs every one again himself. Two settings: delegated, or reclaimed.
P, pinpoint the old decision. Ferrow's product team built its engagement dashboard to treat a rising query count as adoption working as intended, because early on, more queries always meant more trucks fixed faster.
S, show the replay. The apprentice's answers started drifting on a batch of newer engine models TorqueSight had not seen much of. The senior mechanic noticed and quietly started running every one of the apprentice's queries again himself. Query count per active account doubled. Deslin's dashboard called it the best adoption month all year.

Same shape, opposite direction A query count that rises can be two people doing one job just as easily as it can be one person doing it twice as well. Volume alone never says which.
Queries per week, by who is actually asking
19 1 20 22 Before drift After drift
ApprenticeSenior mechanic
Total queries per active account nearly doubled, from twenty to forty two a week. Ferrow's dashboard read that as growth. It was really one person quietly redoing another person's work.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the fix: whatever raw usage number you are watching, pair it with a share that shows whether the work is being trusted, not just repeated.
Cost: engineering says a real resolution share metric cannot ship for two months. Do not read the raw query count alone in the meantime and call it a stopgap. Hold off on any alert built from that number until the real one exists.
The model got better, for real: say TorqueSight's diagnostic accuracy genuinely improves next quarter. That still is not the same claim as "rising queries mean people trust it more." A better model just makes a delegation flip like Ferrow's harder to spot, because the apprentice's raw numbers stay steady too.

Where people run it wrong.
They treat a falling number as decline and a rising number as growth, on reflex, without asking who is actually behind either one.
They fix the metric's mistake with a bigger review board, instead of a second number that would have caught it automatically.
They wait for a renewal call to ask why usage moved, instead of charting the real signal every week whether anyone remembers to look or not.

How to use it live. Say the reframe before naming a single fix: "A usage number that only counts how often someone asks can't tell a person who trusts the tool apart from one who gave up on it, or in this case, apart from a second person quietly doing the same work twice." That buys you room to give the real answer, instead of reciting "vanity metric" on reflex.

Flashcards (click a card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over trust. The change here is an improvement: Distillery got good enough that Bryce needed fewer follow up queries, not more review.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bryce Sabatini, senior product researcher at Palmerhouse, using Cobble Insights' Distillery for about ten weeks.
3 · THE HABIT
What did they stop doing because it worked?
Tap to flip
ANSWER
Asking Distillery to prove the same theme four different ways, quotes, a rephrase, a quarter over quarter check, before trusting it enough to write it down.
4 · THE FLIP
What's the two setting switch here?
Tap to flip
ANSWER
Triangulate every answer with follow up queries, or accept one answer and write it into the report. No middle setting once trust builds.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
The alert that flags any account whose weekly queries drop by a third or more, two weeks running, as churn risk. It made sense while the tool was new, when a drop really did mean someone had given up.
6 · THE NUMBER
Fill in the blank: over ten weeks, Bryce's weekly queries fell from thirty four to ___, while his single shot resolution share rose from thirty five percent to ___.
Tap to flip
ANSWER
Nine; eighty one percent. The second number is the one that shows the drop was a good sign.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The alert watches resolution share instead of raw query count. No nudge fires. Bryce's deck ships on time, and the renewal review reads the numbers as an upsell case instead of a churn risk.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
TorqueSight, Ferrow Fleet's diagnostic tool. Delegation flip: a senior mechanic reclaiming diagnostic queries from an apprentice once the model's answers started drifting.

Check yourself Score: 0 / 0

Multiple choice
1. What is the real problem with "queries per user per week" as a health metric for a tool like Distillery?
  • A. It costs too much to compute for large accounts.
  • B. It cannot tell someone who gave up apart from someone who now trusts the tool.
  • C. It only makes sense for tools with a free tier.
  • D. It ignores how many themes a survey wave has.
Show hint
Think about what a falling number and a rising number can each mean.
Show answer
B. A falling raw count can mean disengagement or mastery, and the metric alone cannot say which one happened.
True or false
2. True or false: Bryce's falling query count was itself something Cobble's product team should have fixed.
  • True
  • False
Show hint
Check what the falling number actually meant for Bryce, not for the dashboard.
Show answer
False. The falling count was Distillery working as intended. The thing that needed fixing was the alert reading the drop as a problem, not the drop itself.
Fill in the blank
3. Over ten weeks, Bryce's single shot resolution share rose from thirty five percent to ___ percent.
Show hint
Check the second chart in "Let's learn."
Show answer
81 percent. It climbed the whole time his raw query count was falling, which is exactly why the two numbers need to be read together.
Multiple choice
4. What does this answer say to track instead of raw query count?
  • A. Total time spent inside the app each week.
  • B. The share of answers used with no follow up question, split by account age.
  • C. The product team's own satisfaction score for the tool.
  • D. The number of themes Distillery flags per survey wave.
Show hint
It has to separate a person who trusts the answer from a person who is still checking it.
Show answer
B. Splitting by account age is what stops a brand new, genuinely struggling account from being read the same way as a ten week veteran.
Short answer, apply it yourself
5. Think of an app you use a lot less than you used to, without ever actually losing trust in it. What's one number that app's own dashboard would misread as you losing interest?
Show hint
Look for a habit that got shorter, not a habit that stopped.
Show answer
Model answer: A grocery delivery app. Once someone learns the one order that always works, they stop browsing and just reorder the same list. App opens per week would fall, and a team watching only that number would read it as churn risk, when it is actually the most loyal kind of user, the automatic repeat kind.
True or false
6. True or false: if Bryce's query count had fallen from thirty four to nine in two weeks instead of ten, this answer would still call it pure good news with nothing further to check.
  • True
  • False
Show hint
A drop that fast looks different from a drop that builds slowly over ten weeks.
Show answer
False. A drop that fast looks more like someone hit a wall and gave up, not someone building trust gradually. The resolution share still settles it, but a fast drop deserves a closer look before you call it good news.
Before you close the answer
Why this works
Tests whether you will read a number's direction on reflex or ask what changed underneath it. Most candidates treat a falling number as bad by habit, the same way Cobble's own dashboard did.
Follow-up traps
"Doesn't a rising resolution share just mean people are settling for a worse answer instead of pushing back?" Response: that is why account age and a floor both matter. If resolution share rises alongside a drop in edits and complaints, it is a good sign. If it rises while support tickets rise too, that is exactly what this metric would still catch.

"Couldn't a team just game single shot resolution rate by making the model refuse follow up questions?" Response: that is what the golden set gate is for. A model change that raises resolution share by getting worse at handling follow ups would fail the golden set's accuracy check before it ever reached a user.
If pressed
The golden set is not static either. Cobble rotates ten percent of each quarter's hand checked surveys back into it, because a set trained against last year's product language starts missing new complaint themes the moment the product line changes, and a stale golden set would stop catching that.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more