CaseIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #11

Describe metrics for an internal AI tool where there is no revenue signal.

The direct answer
Pair average handle time with a reopen number: the share of tickets that come back within five business days because the first fix didn't actually hold, split by ticket category, checked every week. The reopen number is the signal that moves first. The moment a category's reopen rate crosses twice its normal level for two weeks running, freeze that category's suggestions and send those tickets to a person, no matter how good the average looks that week.
Do this, in order
  1. Pair average handle time with the reopen-within-5-days rate, split by category, and gate rollout on the reopen number.Why: this is the leading signal, not a nice thing to glance at once a quarter.
  2. Set a real freeze rule: twice a category's normal reopen rate for two straight weeks pulls that category's suggestions.Why: a climbing number nobody acts on is a chart, not a metric.
  3. Treat suggestion-acceptance rate as a compliance number whenever usage is mandated, not a trust number.Why: a mandate inflates usage without proving the tool actually helped anyone.
  4. Only gate the categories where a wrong suggestion is hidden and expensive, like VPN or system access, not every category.Why: a wrong password-reset suggestion is loud and cheap; gating it just slows the tool down for nothing.
  5. Check how fresh the knowledge base article behind a flagged category is before blaming the model.Why: the usual cause is a confident answer that used to be true and never got re-checked.
  6. Report the reopen number to the same budget conversation that decides headcount, not a separate engineering dashboard.Why: a number only changes a decision if it reaches the person who makes it.

How to answer this, stage by stage

Nobody is grading whether you can name a metric off the top of your head. They are grading whether you can find the number that would have caught this six weeks before the budget review did. Seven moves get you there.

1
Scope it to one real tool, one real team
Say it like this
"Let's ground this. Corvan Systems runs an internal helpdesk, fourteen agents, about thirty-one hundred tickets a month. They built Ridgepath into the ticketing system: it reads an incoming ticket and hands the agent a suggested fix pulled from the internal knowledge base and past resolved tickets."
Why this works
Grounds the answer in a real internal tool before naming a single number.
2
Name the outcome that actually matters, not the model's score
Say it like this
"There's no revenue line here, so I'm not starting with accuracy or a thumbs-up score. I'm starting with what Corvan already tracks: how many more support agents they'd need to hire as ticket volume grows. That's the L in LEAD, the link to something real the business already watches."
Why this works
Stops the answer from measuring the model instead of the business it sits inside.
3
Say why that outcome is too slow to act on
Say it like this
"Corvan reviews headcount once a quarter. If I wait for that review to tell me Ridgepath isn't working, I'm three months behind whatever already went wrong. I need a number that moves weeks earlier than that."
Why this works
Sets up the actual answer instead of just defining what a leading indicator is.
4
Give the leading signal, the real answer
Say it like this
"Here's the signal. Average handle time drops the moment an agent clicks a suggestion instead of typing their own answer, whether or not the suggestion actually fixed anything. What I'd watch instead is the reopen rate: how many tickets come back within five business days because the first fix didn't hold, split by category, checked every week."
Why this works
Matches the direct answer. This is the E step, and it's the part most candidates skip past on their way to a vanity number.
5
Name how the metric gets gamed
Say it like this
"If leadership sets an adoption target, say eighty percent of tickets have to touch a Ridgepath suggestion, agents will hit that number by clicking accept and quietly fixing the ticket themselves anyway. A high acceptance rate sitting next to a high reopen rate isn't proof the tool works. It's proof of a mandate."
Why this works
This is the A step. Naming it before an interviewer asks is what makes the answer feel earned instead of lucky.
6
Say what changes at each threshold
Say it like this
"Under one and a half times a category's normal reopen rate, keep rolling out. Cross two times for two weeks straight, freeze that category, route it to a person, and don't turn it back on until it clears a test set of the tickets that broke it. Adoption rate alone never makes that call. Only the reopen number does."
Why this works
This is the D step. A metric with no threshold attached to it is a chart nobody actually acts on.
7
Close on the trade you accepted and the option you turned down
Say it like this
"We looked at a thumbs-up survey after every ticket instead, and turned it down. Response rates run under ten percent, and it only tells you how the agent felt that minute, not whether the fix held five days later. The reopen gate costs something too. Pulling in more ticket history to catch a stale knowledge base reference before it ships adds real delay and roughly doubles what each suggestion costs to generate. I'd only pay that for the categories where a confident wrong answer is expensive."
Why this works
Naming a rejected alternative and a real cost is what turns this into a decision instead of a wish list.
If you remember one thing A number that keeps dropping is not proof a tool is working. It's proof you haven't built the number that would tell you when the drop stops meaning what you think it means.

Let's learn

Before any AI touched a single ticket, an internal helpdesk ran the old way: fourteen agents closing about thirty-one hundred tickets a month, each one taking about eighteen minutes from open to close, everything from a forgotten password to a broken VPN connection.

Ridgepath is the tool that got built into that same ticketing system. When a ticket lands, it reads it, checks the internal knowledge base and the history of resolved tickets, and hands the agent a suggested fix. The agent can use it as written, edit it, or ignore it and type their own.

Six weeks after Ridgepath went live, average handle time had dropped from eighteen minutes to eleven. Leadership loved the number. It went straight into the slide that said the tool was working.

Six weeks, two lines, only one of them had a chart
Average handle time (what the slide showed)
18m 14m 11m week 0 week 6
VPN category reopen rate (nobody's slide)
19% 12% 6% week 0 week 6 week 4, already loud
Both lines moved for the same six weeks. The left one reached a slide deck. The right one didn't exist as a chart until someone went looking for it, four weeks after it had already crossed twice its normal level.

Here is the turn. The six minutes that came off the average were not actually gone. They had just stopped showing up in the number everyone was watching.

We didn't shrink the work by six minutes. We moved it somewhere the dashboard never checked.

In the VPN and network-access category specifically, the share of tickets reopened within five business days climbed from six percent to nineteen percent over that same six weeks. A ticket would close, marked fixed, and the same employee would be back within a day or two, still locked out.

Knowledge spark: why would an AI suggest a fix that's already broken? Ridgepath wasn't guessing at random. It was confidently repeating a step that used to work: a script pointing at a VPN client version Corvan had retired two months earlier. The knowledge base article behind that step never got re-checked after the client was pulled, so the model kept handing out an answer that was true right up until it wasn't.

At its worst, this cost more than wasted agent minutes. During a new facility's opening week, about forty field technicians spent the better part of two days locked out of the inventory system, chasing the same VPN fix through two more tickets before anyone caught the pattern. The facility's go-live slipped by two days.

The choice I would take back Picking one number, average handle time, as the whole story, with nothing standing next to it. That was fine while Ridgepath only handled simple, cheap-to-fix categories. It stopped being fine the day it started suggesting fixes for problems complicated enough that a wrong answer doesn't announce itself.

What I would leave alone: password resets and software-install requests never needed a reopen gate. If Ridgepath gets one of those wrong, the employee just asks again in the same conversation, and it costs almost nothing. Gating every category the same way would only have slowed down the parts of the tool that were never broken.

The lesson: a number that keeps dropping is not proof a tool is working. It's proof you haven't built the number that would tell you when the drop stops meaning what you think it means.

Now here is the same thing as a story

The short version is above. Read on for how ordinary the week this nearly went unnoticed looked from inside the support team.

The helpdesk floor at Corvan's headquarters gets loud right around nine, when the overnight tickets land on top of the queue all at once. Danika Braddock runs support operations there, nine years in, most of them spent turning a stack of angry tickets into a calm one before lunch. She's the person who can look at a week of ticket data and tell you which category is about to become a problem before anyone else has noticed.

Ridgepath went live in March. For the first month it looked like exactly what Danika had been asking two budget cycles running for. Average handle time fell from eighteen minutes to fourteen within two weeks, then to eleven by week six. She put that number in front of the CFO's office herself, proof the team could absorb two new facilities' worth of tickets without adding headcount.

For those first two months, Danika still pulled ten closed tickets a week by hand and read through them, the way she always had before any tool touched a ticket. Every one she checked looked fine. Suggested fix, ticket closed, no complaints. By week five she was down to three checks a week. By week seven, none. The dashboard was doing the checking for her now, or so it seemed.

Then a message showed up in the team's chat from a facilities manager at the new site, not a complaint exactly, more a question. "Anyone else still fighting with VPN? This is the third ticket I've opened this month for the same thing."

Danika pulled the ticket. Then the two before it. Same employee, same suggested fix each time, a script pointing at a VPN client version Corvan had retired two months earlier. Ridgepath had been handing out that same wrong, confident step for six weeks, and every one of those tickets had closed on time, inside the handle-time number everyone was proud of.

We didn't lose six minutes off the average. We spent it hunting the same VPN ticket three times over, on a man who couldn't do his job in between.

It was never really about the average dropping. Danika didn't have a number in her head for "safe." She had a habit, checking ten tickets a week, and once the dashboard looked good enough for long enough, the habit quietly stopped. Nothing dramatic flipped it off. It just wasn't there by week seven.

Two weeks after the call about the VPN tickets, leadership had already set a company goal: eighty percent of tickets should touch a Ridgepath suggestion. Agents didn't need telling twice. The acceptance rate climbed to ninety-one percent within a month, and for a while it sat on the same slide as the falling handle time, both looking like proof of the same thing. What the acceptance number never showed was how many agents clicked accept and then quietly fixed the ticket the old way anyway, because the suggestion was wrong and nobody had time to argue with a form field.

Hand sketched scene titled What adoption rate actually caught. Left, a clipboard with a large checkmark labelled Suggestion used, yes. Right, a small stick figure walking away from a screen labelled Fixed it herself anyway.
The acceptance number said yes on every one of these. It had no way to show the agent walking back to fix it by hand a minute later.

The decision I'd take back happened back in February, before launch, in a half-hour meeting about what to put on the rollout dashboard. Average handle time was the obvious choice, it was already the number everyone reported, and adding a second one felt like extra work for a launch that was already running late. Nobody in that room was picturing a stale VPN script running quietly for six weeks under a number that kept looking fine.

What Danika actually built, once she went looking: a weekly reopen-rate chart, split by category, with a real rule attached. VPN and network access got frozen the same afternoon she found the second reopened ticket, routed straight to a person, and flagged for a knowledge base fix. Three weeks later, with the retired client reference pulled from Ridgepath's grounding documents and the category retested against the tickets that had broken it, VPN suggestions went back out to ten percent of that category's tickets first. By the following Monday, the reopen rate on that slice was back under seven percent, and it held there through the next facility's opening, the one that didn't slip.

What I'd tell myself, back in that February meeting: a number that only measures how fast a ticket closes was never going to catch a fix that doesn't actually work. You have to go build the number that watches what happens after.

LEAD, before the quarter closes

This is a metric question with no revenue line to point to, so LEAD fits: find the signal that moves first. Not a story about a flip, and not a fairness question about who can't push back.

L
Link. The business outcome that actually matters, not the model's own score.
Headcount avoided: the number of extra support agents Corvan would need to hire to keep pace as ticket volume grows, the figure that shows up in the quarterly budget review.
In this story: the three hires the team would have needed for two new facilities.
E
Early signal. What moves weeks before the outcome does.
The reopen-within-5-business-days rate, split by ticket category, checked weekly. It doesn't wait for a quarterly review to tell the truth.
It climbed from 6 percent to 19 percent in the VPN category, six weeks before any headcount review would have shown a thing.
A
Abuse. How this metric gets gamed, by the team or the tool.
A company-wide adoption target pushed acceptance to 91 percent. That number measured compliance with a mandate, not whether the fix actually held.
A high acceptance rate sitting next to a high reopen rate is a warning, not a win.
D
Decision. What actually changes at each threshold.
Under 1.5 times a category's baseline, keep shipping. Over 2 times for two weeks straight, freeze the category, route to a person, retest before it comes back.
VPN froze the same afternoon the second reopened ticket was found, came back at a 10 percent canary three weeks later.

Two things worth naming directly, since this is where the AI-specific judgment actually lives. First, the alternative most people reach for is a quick thumbs-up survey after each ticket. That got turned down on purpose: response rates on these usually sit under ten percent, and even the ones that come in only say how the agent felt in the moment, not whether the fix held five days later, which is exactly as slow and exactly as thin as the quarterly review it was supposed to replace. Second, the real bar for a category is not "must always give a correct answer." It's calibrated: a category clears the gate when its reopen rate holds under one and a half times its own baseline across a rolling two-week window, checked against real ticket outcomes, not one good day. The failure mode worth naming by name is stale grounding, a model confidently repeating a fix that used to be true because the document behind it never got re-checked after the thing it described changed. The guardrail is the reopen-rate gate plus a freshness check on any grounding document past a set age, not a person spot-checking suggestions by hand, since Danika's own three-a-week habit had already proven too thin to catch this at real volume. The trade is real too: pulling in more ticket history to catch a stale reference before it ships adds real delay, something like a second and a half more per suggestion, and roughly doubles what each suggestion costs to generate. Worth paying for VPN and system access, where a wrong answer is expensive and hidden. Not worth paying for a password reset, where a wrong answer is loud and cheap.

And if you want to be sure it really works, try it somewhere else

Same four letters, an HR policy tool instead of a helpdesk, so the method proves itself instead of repeating a story I happened to prepare.

Halewood Manufacturing runs Beaconline, an assistant that answers employee questions about benefits, leave, and internal policy, pulled from the handbook and HR's own guidance. Farid Delfina runs HR operations across Halewood's four plants.

L, link. The number Halewood already tracks: headcount for HR generalists. Two acquisitions in eighteen months doubled the number of employees each generalist supports, and Beaconline's whole job is keeping that ratio from breaking without another hire.
E, early signal. The escalation-within-24-hours rate: how often an employee books time with a live generalist to ask again, on a question Beaconline had already answered, split by policy topic, checked weekly.
A, abuse. Leadership tells employees to "ask Beaconline first" before a generalist will take the meeting. Self-serve resolution looks like 88 percent, but that number is measuring the mandate, not whether the answer was right.
D, decision. During open enrollment, benefits-topic escalation crossing twice its normal rate for even one week freezes Beaconline's benefits answers and routes straight to a person, because a wrong answer there can cost someone a missed deadline, not just a redo.

Same shape, different stakes At Corvan, the unwatched cost was a technician locked out for two days. At Halewood, it's an employee who missed an enrollment window because a confident answer was wrong and nobody escalated it in time. The early signal doesn't change: find the rate that shows the fix didn't hold, before the slow review would have caught it.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the pairing: whatever outcome the business already tracks, pair it with a reopen or escalation rate, and gate rollout on that number, not the average alone.
Cost: engineering says the category-level reopen dashboard can't ship for two months, not two weeks. Don't roll every category out in the meantime and call the delay a formality. Hold the riskiest categories, VPN, system access, benefits, out of full rollout until the dashboard exists.
The model got better, for real: say Ridgepath's suggestions get meaningfully more accurate on average. That's still not the same claim as "the reopen gate is now unnecessary." A better model just makes a stale-grounding failure rarer, not impossible, and rare-and-hidden is exactly the kind of failure a reopen gate exists to catch.

Where people run it wrong.
They treat a falling average as proof nothing is wrong, instead of asking what a fast, wrong fix would even look like on that same dashboard.
They celebrate an adoption or acceptance number without ever checking whether it was hit by choice or by mandate.
They put the whole guardrail in one person's habit, a manager spot-checking tickets by hand, instead of a number that gets checked every week whether anyone remembers to look or not.

How to use it live. Say the split before naming a single metric: "For an internal tool with no revenue line, I look for the number the business already tracks, headcount, cost, hours, and then I go find the thing that would move before that number does." That buys you room to give the real answer instead of reaching for "user satisfaction" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a metric question with no revenue signal, and why?
Tap to flip
ANSWER
LEAD. It's a metric question, find the signal that moves first, not a question about a flip or about who can't push back.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Danika Braddock, nine years running support operations at Corvan Systems, owner of the internal helpdesk's metrics.
3 · THE OUTCOME
What's the real outcome behind a number with no revenue line?
Tap to flip
ANSWER
Headcount avoided: the extra support agents Corvan would need to hire as ticket volume grows, the figure the quarterly budget review actually watches.
4 · THE EARLY SIGNAL
What's the leading signal that moved before the outcome did?
Tap to flip
ANSWER
The reopen-within-5-business-days rate, split by ticket category, checked weekly. It climbed six weeks before any quarterly review would have shown a problem.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Picking one number, average handle time, with nothing standing next to it. It made sense while Ridgepath only handled simple, cheap-to-fix categories.
6 · THE NUMBER
Fill in the blank: over six weeks, handle time dropped from 18 minutes to ___, while the VPN category's reopen rate climbed from 6 percent to ___.
Tap to flip
ANSWER
11 minutes; 19 percent. The second number is the one nobody had built a chart for.
7 · THE THRESHOLD
What's the actual rule for freezing a category, and what happens next?
Tap to flip
ANSWER
Reopen rate crossing twice a category's normal rate for two weeks straight freezes suggestions there, routes tickets to a person, and only reopens at a 10 percent canary once the cause is fixed and retested.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the early signal?
Tap to flip
ANSWER
Beaconline, Halewood Manufacturing's HR policy assistant. Its early signal is the escalation-within-24-hours rate on questions it already answered, split by policy topic.

Check yourself Score: 0 / 0

Multiple choice
1. Which number does this answer say is the real leading indicator for Ridgepath, the internal helpdesk assistant?
  • A. The percentage of tickets Ridgepath resolves with no human agent involved at all.
  • B. The reopen-within-5-business-days rate, split by ticket category.
  • C. The average star rating agents give Ridgepath's suggestions.
  • D. The total number of tickets Ridgepath answers per day.
Show hint
It has to catch a fix that looked done but didn't actually hold.
Show answer
B. The reopen rate catches hidden rework that average handle time, or a star rating given in the moment, cannot see.
Fill in the blank
2. Over six weeks, the VPN category's reopen rate climbed from 6 percent to ___ percent, even while average handle time kept dropping.
Show hint
Check the two-line chart in "Let's learn."
Show answer
19 percent. It moved for the whole six weeks that handle time was falling, and nobody had built the chart that would have shown it.
True or false
3. True or false: average handle time dropping from eighteen minutes to eleven minutes is proof Ridgepath was working.
  • True
  • False
Show hint
Check what the average number actually blends together.
Show answer
False. The average blended a real time saving on simple tickets with a hidden cost, tickets closing marked fixed when they weren't, that never showed up until someone built a separate number for it.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory, not a setting anyone could just turn up.
Show answer
Model answer: Picking average handle time as the only number on the launch dashboard, with nothing standing next to it. It made sense while Ridgepath only handled simple categories where a wrong suggestion was cheap and obvious; it stopped making sense once Ridgepath started suggesting fixes for problems like VPN access, where a wrong answer doesn't announce itself.
True or false
5. True or false: the password-reset ticket category needs the same reopen-rate gate as VPN and network access before its suggestions can be trusted.
  • True
  • False
Show hint
Check "what I would leave alone" in "Let's learn."
Show answer
False. A wrong password-reset suggestion is obvious right away, the employee just asks again in the same conversation, so gating that category would only slow down a part of the tool that was never actually broken.
Short answer, apply it yourself
6. Think of a tool at your own job with no revenue number attached to it, an internal wiki, a shared scheduling tool, an IT self-service portal. What would a good early signal look like for it, one that would move weeks before anyone official noticed a problem?
Show hint
Look for a sign the first fix didn't really hold, not just that people used the tool.
Show answer
Model answer: An internal wiki search tool. A good early signal would be how often someone searches the same term twice within an hour, since that usually means the first result didn't actually answer their question, well before any formal usage review would ever catch it.
Before you close the answer
Why this works
Tests whether you'll pair a fast-looking average with a signal that catches the rework hiding underneath it, instead of trusting a number that only measures how quickly a ticket got closed. Most candidates stop at "we'd track user satisfaction."
Follow-up traps
"Isn't checking reopens by category just going to slow down every rollout?" Response: no, only categories that actually cross the threshold get frozen; a category with a normal reopen rate keeps shipping at full speed.

"Couldn't an agent dodge the reopen number by just opening a brand new ticket instead of reopening the old one?" Response: that's why the reopen check tracks the same employee and the same underlying issue, not just the ticket ID. A new ticket from the same person about the same VPN drop inside the window still counts.
If pressed
The canary re-launch isn't ten percent of everything Ridgepath handles. It's ten percent of that specific category's tickets, chosen at random each day, so the retest has real coverage of the exact thing that broke instead of diluting into categories that were never at risk.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more