Describe metrics for an internal AI tool where there is no revenue signal.
- Pair average handle time with the reopen-within-5-days rate, split by category, and gate rollout on the reopen number.Why: this is the leading signal, not a nice thing to glance at once a quarter.
- Set a real freeze rule: twice a category's normal reopen rate for two straight weeks pulls that category's suggestions.Why: a climbing number nobody acts on is a chart, not a metric.
- Treat suggestion-acceptance rate as a compliance number whenever usage is mandated, not a trust number.Why: a mandate inflates usage without proving the tool actually helped anyone.
- Only gate the categories where a wrong suggestion is hidden and expensive, like VPN or system access, not every category.Why: a wrong password-reset suggestion is loud and cheap; gating it just slows the tool down for nothing.
- Check how fresh the knowledge base article behind a flagged category is before blaming the model.Why: the usual cause is a confident answer that used to be true and never got re-checked.
- Report the reopen number to the same budget conversation that decides headcount, not a separate engineering dashboard.Why: a number only changes a decision if it reaches the person who makes it.
How to answer this, stage by stage
Nobody is grading whether you can name a metric off the top of your head. They are grading whether you can find the number that would have caught this six weeks before the budget review did. Seven moves get you there.
Let's learn
Before any AI touched a single ticket, an internal helpdesk ran the old way: fourteen agents closing about thirty-one hundred tickets a month, each one taking about eighteen minutes from open to close, everything from a forgotten password to a broken VPN connection.
Ridgepath is the tool that got built into that same ticketing system. When a ticket lands, it reads it, checks the internal knowledge base and the history of resolved tickets, and hands the agent a suggested fix. The agent can use it as written, edit it, or ignore it and type their own.
Six weeks after Ridgepath went live, average handle time had dropped from eighteen minutes to eleven. Leadership loved the number. It went straight into the slide that said the tool was working.
Here is the turn. The six minutes that came off the average were not actually gone. They had just stopped showing up in the number everyone was watching.
In the VPN and network-access category specifically, the share of tickets reopened within five business days climbed from six percent to nineteen percent over that same six weeks. A ticket would close, marked fixed, and the same employee would be back within a day or two, still locked out.
At its worst, this cost more than wasted agent minutes. During a new facility's opening week, about forty field technicians spent the better part of two days locked out of the inventory system, chasing the same VPN fix through two more tickets before anyone caught the pattern. The facility's go-live slipped by two days.
What I would leave alone: password resets and software-install requests never needed a reopen gate. If Ridgepath gets one of those wrong, the employee just asks again in the same conversation, and it costs almost nothing. Gating every category the same way would only have slowed down the parts of the tool that were never broken.
The lesson: a number that keeps dropping is not proof a tool is working. It's proof you haven't built the number that would tell you when the drop stops meaning what you think it means.
Now here is the same thing as a story
The short version is above. Read on for how ordinary the week this nearly went unnoticed looked from inside the support team.
The helpdesk floor at Corvan's headquarters gets loud right around nine, when the overnight tickets land on top of the queue all at once. Danika Braddock runs support operations there, nine years in, most of them spent turning a stack of angry tickets into a calm one before lunch. She's the person who can look at a week of ticket data and tell you which category is about to become a problem before anyone else has noticed.
Ridgepath went live in March. For the first month it looked like exactly what Danika had been asking two budget cycles running for. Average handle time fell from eighteen minutes to fourteen within two weeks, then to eleven by week six. She put that number in front of the CFO's office herself, proof the team could absorb two new facilities' worth of tickets without adding headcount.
For those first two months, Danika still pulled ten closed tickets a week by hand and read through them, the way she always had before any tool touched a ticket. Every one she checked looked fine. Suggested fix, ticket closed, no complaints. By week five she was down to three checks a week. By week seven, none. The dashboard was doing the checking for her now, or so it seemed.
Then a message showed up in the team's chat from a facilities manager at the new site, not a complaint exactly, more a question. "Anyone else still fighting with VPN? This is the third ticket I've opened this month for the same thing."
Danika pulled the ticket. Then the two before it. Same employee, same suggested fix each time, a script pointing at a VPN client version Corvan had retired two months earlier. Ridgepath had been handing out that same wrong, confident step for six weeks, and every one of those tickets had closed on time, inside the handle-time number everyone was proud of.
It was never really about the average dropping. Danika didn't have a number in her head for "safe." She had a habit, checking ten tickets a week, and once the dashboard looked good enough for long enough, the habit quietly stopped. Nothing dramatic flipped it off. It just wasn't there by week seven.
Two weeks after the call about the VPN tickets, leadership had already set a company goal: eighty percent of tickets should touch a Ridgepath suggestion. Agents didn't need telling twice. The acceptance rate climbed to ninety-one percent within a month, and for a while it sat on the same slide as the falling handle time, both looking like proof of the same thing. What the acceptance number never showed was how many agents clicked accept and then quietly fixed the ticket the old way anyway, because the suggestion was wrong and nobody had time to argue with a form field.
The decision I'd take back happened back in February, before launch, in a half-hour meeting about what to put on the rollout dashboard. Average handle time was the obvious choice, it was already the number everyone reported, and adding a second one felt like extra work for a launch that was already running late. Nobody in that room was picturing a stale VPN script running quietly for six weeks under a number that kept looking fine.
What Danika actually built, once she went looking: a weekly reopen-rate chart, split by category, with a real rule attached. VPN and network access got frozen the same afternoon she found the second reopened ticket, routed straight to a person, and flagged for a knowledge base fix. Three weeks later, with the retired client reference pulled from Ridgepath's grounding documents and the category retested against the tickets that had broken it, VPN suggestions went back out to ten percent of that category's tickets first. By the following Monday, the reopen rate on that slice was back under seven percent, and it held there through the next facility's opening, the one that didn't slip.
What I'd tell myself, back in that February meeting: a number that only measures how fast a ticket closes was never going to catch a fix that doesn't actually work. You have to go build the number that watches what happens after.
LEAD, before the quarter closes
This is a metric question with no revenue line to point to, so LEAD fits: find the signal that moves first. Not a story about a flip, and not a fairness question about who can't push back.
Two things worth naming directly, since this is where the AI-specific judgment actually lives. First, the alternative most people reach for is a quick thumbs-up survey after each ticket. That got turned down on purpose: response rates on these usually sit under ten percent, and even the ones that come in only say how the agent felt in the moment, not whether the fix held five days later, which is exactly as slow and exactly as thin as the quarterly review it was supposed to replace. Second, the real bar for a category is not "must always give a correct answer." It's calibrated: a category clears the gate when its reopen rate holds under one and a half times its own baseline across a rolling two-week window, checked against real ticket outcomes, not one good day. The failure mode worth naming by name is stale grounding, a model confidently repeating a fix that used to be true because the document behind it never got re-checked after the thing it described changed. The guardrail is the reopen-rate gate plus a freshness check on any grounding document past a set age, not a person spot-checking suggestions by hand, since Danika's own three-a-week habit had already proven too thin to catch this at real volume. The trade is real too: pulling in more ticket history to catch a stale reference before it ships adds real delay, something like a second and a half more per suggestion, and roughly doubles what each suggestion costs to generate. Worth paying for VPN and system access, where a wrong answer is expensive and hidden. Not worth paying for a password reset, where a wrong answer is loud and cheap.
And if you want to be sure it really works, try it somewhere else
Same four letters, an HR policy tool instead of a helpdesk, so the method proves itself instead of repeating a story I happened to prepare.
Halewood Manufacturing runs Beaconline, an assistant that answers employee questions about benefits, leave, and internal policy, pulled from the handbook and HR's own guidance. Farid Delfina runs HR operations across Halewood's four plants.
L, link. The number Halewood already tracks: headcount for HR generalists. Two acquisitions in eighteen months doubled the number of employees each generalist supports, and Beaconline's whole job is keeping that ratio from breaking without another hire.
E, early signal. The escalation-within-24-hours rate: how often an employee books time with a live generalist to ask again, on a question Beaconline had already answered, split by policy topic, checked weekly.
A, abuse. Leadership tells employees to "ask Beaconline first" before a generalist will take the meeting. Self-serve resolution looks like 88 percent, but that number is measuring the mandate, not whether the answer was right.
D, decision. During open enrollment, benefits-topic escalation crossing twice its normal rate for even one week freezes Beaconline's benefits answers and routes straight to a person, because a wrong answer there can cost someone a missed deadline, not just a redo.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the pairing: whatever outcome the business already tracks, pair it with a reopen or escalation rate, and gate rollout on that number, not the average alone.
Cost: engineering says the category-level reopen dashboard can't ship for two months, not two weeks. Don't roll every category out in the meantime and call the delay a formality. Hold the riskiest categories, VPN, system access, benefits, out of full rollout until the dashboard exists.
The model got better, for real: say Ridgepath's suggestions get meaningfully more accurate on average. That's still not the same claim as "the reopen gate is now unnecessary." A better model just makes a stale-grounding failure rarer, not impossible, and rare-and-hidden is exactly the kind of failure a reopen gate exists to catch.
Where people run it wrong.
They treat a falling average as proof nothing is wrong, instead of asking what a fast, wrong fix would even look like on that same dashboard.
They celebrate an adoption or acceptance number without ever checking whether it was hit by choice or by mandate.
They put the whole guardrail in one person's habit, a manager spot-checking tickets by hand, instead of a number that gets checked every week whether anyone remembers to look or not.
How to use it live. Say the split before naming a single metric: "For an internal tool with no revenue line, I look for the number the business already tracks, headcount, cost, hours, and then I go find the thing that would move before that number does." That buys you room to give the real answer instead of reaching for "user satisfaction" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't an agent dodge the reopen number by just opening a brand new ticket instead of reopening the old one?" Response: that's why the reopen check tracks the same employee and the same underlying issue, not just the ticket ID. A new ticket from the same person about the same VPN drop inside the window still counts.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?