Design the metric tree for an AI-powered support deflection feature.
- Split "did not reach a human" into resolved and abandoned at the very first leaf, before any other branch.Why: a raw deflection number blends a customer who got real help with a customer who just quit. Those are opposite outcomes wearing the same number.
- Define "resolved" with a real proxy signal, no repeat contact on the same issue within 48 hours and no negative signal, then check that proxy against a hand-labeled set of real chats.You cannot ask every customer if the bot actually helped them, so the proxy has to earn trust before anyone puts it in a slide.
- Report the resolved leaf as the number leadership acts on, not the raw deflection total.The raw total is what gets read out loud in a review. If it is not the honest number, the honest number never gets heard.
- Add one layer under resolved for task type, billing, outage, plan change, since that is the ISP's own real work split.A tree with no second layer under its safest leaf still cannot tell anyone which part of the product is actually earning the number.
- Sample a slice of "resolved" chats by hand every week and check for confident, wrong answers.A bot that sounds sure and is wrong looks exactly like a resolved chat in the data. Only a human reading the transcript catches it.
- Leave out root-cause tagging, channel splits, and a dollar-per-resolution branch on day one.Stacking more branches onto an unproven split just multiplies whatever is wrong with the split underneath them.
How to answer this, stage by stage
Nobody is grading whether you can draw boxes and arrows. They are grading whether the top number in your tree survives the first time someone in a leadership review asks "why did it move." Eight moves get you there.
Let's learn
Every evening, Rian Pellham, VP of Support at Cordwell Fiber, an internet provider that serves about ninety thousand homes, opens a slide with one number on it. Tether, Cordwell's chat assistant, answers billing questions, logs outage reports, and processes plan changes before a customer ever reaches a person. Before Tether existed, every one of those requests went straight to a phone queue, an average hold of eleven minutes at peak hours.
With Tether running, about forty thousand customers a month start a chat instead of calling. Sixty one percent of those chats never touch a human agent at all. That number, deflection rate, is the one Rian has reported for a year. It has always gone up. It has always looked good.
Here is the turn. A chat that never reaches a human is not automatically a good outcome. Some of those sixty one points are customers whose billing question got answered correctly. Some of them are customers who typed "representative" three times, got looped back to the same menu, and closed the tab. Both count as deflected. Only one of them is actually a win.
At its worst, this costs more than a wrong slide. Say the raw number goes into the review unsplit, and Rian recommends letting a staffing contract for twelve agents lapse, since deflection alone seems to justify a smaller team. The fifteen points of abandoned customers do not disappear. Most of them call back, now angrier, now with a longer story to tell, right as the team meant to handle them is gone.
The choice I would take back is not the sixty one percent itself. It's that Tether's dashboard was built, back when Tether first launched, to report one ratio at the top and nothing underneath it. That made sense in month one, when the whole team was just trying to prove the chatbot could take load off the phone queue at all. It stopped making sense the day that one ratio became the number a headcount decision leaned on.
What I would leave alone: the escalated branch, chats that do reach a human, does not need this same scrutiny. A chat that reaches a person either gets handled by that person or it doesn't, and that outcome is already tracked the same way it was before Tether existed. The ambiguity only lives inside the branch where nobody human ever looked.
The lesson: a single number at the top of a dashboard is not a metric, it's a summary. A metric tree is what you get when you refuse to let that summary travel further than it can be trusted.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how close this one came to shipping wrong on an ordinary Tuesday.
Tomi Idowu has run analytics for Cordwell Fiber's support org for three years. She is the person Rian calls before any number goes on a slide, because she is the one who actually knows what a metric is built from, not just what it says.
Tether launched eighteen months ago and it worked, plainly and quickly. Hold times on the phone queue dropped from eleven minutes to six within two months. Tomi built one dashboard tile for it: deflection rate, and it climbed steadily, forty two percent in month one, past fifty by month six, sixty one by the eighteenth month. Every one of Rian's reviews opened with that tile, and every one of those reviews went well.
The trigger was almost nothing. On a Wednesday afternoon, a support supervisor named Aveline stopped by Tomi's desk to complain about something unrelated and mentioned, almost in passing, "I keep getting calls from people who say they already tried the chatbot for this exact thing." Tomi wrote it down and moved on. It was one sentence, said while someone was looking for a stapler.
Two days later, prepping the slide for that month's review, Tomi pulled a sample of two hundred and fifty "deflected" chats and actually read them. Not the summary, the transcripts. Forty six of the sixty one points held up, real questions, real answers, no comeback. But close to fifteen points were customers who typed "agent," got looped through the same three menu options, and closed the tab. Deflected, technically. Solved, not even close.
Rian's slide, drafted the night before, said this: "Deflection at sixty one percent supports letting our twelve-agent staffing contract lapse at renewal." He was two days from presenting it.
Tomi went back to Rian with the split, not a hunch, the actual read-through. Forty six percent really resolved. Fifteen points was people giving up. Cutting twelve agents on the strength of the raw number would have meant those fifteen points of frustrated customers landing back on a smaller team weeks later, right when nobody was staffed to catch them.
Rian pulled the slide the morning of the review. What went up instead was smaller and less flattering: forty six percent, resolved and checked, next to a note that the team was building a permanent split so this never had to be caught by hand again. The staffing contract got renewed for one more quarter while the real number proved itself.
The replay, six weeks later: the resolved leaf held steady at forty five to forty seven percent, checked weekly against fresh transcripts. The abandoned leaf, once visible, started shrinking on its own, because a chatbot's builders started fixing what abandoned chats had in common, mostly customers stuck in a loop asking for a person and never getting routed to one. By the next review, raw deflection had actually dropped to fifty eight percent, but resolved had climbed to fifty one. A worse-looking top number, a better real one.
What I would tell my past self, the one who built a dashboard with one tile on it: the number that goes in the deck is never really the number itself. It's whatever the number was built to hide.
SPARK, in one screen
This is a design question, not a story about a habit with two settings, so SPARK fits, built forward instead of run backward like a recovery framework.
Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. First, the alternative Tomi actually considered and rejected was using the post-chat thumbs-up rating as the resolved-versus-abandoned split, since it was already collected and looked like a clean signal. She ruled it out. Response rate on that rating sat around eight percent, and the people who bother to answer skew toward the furious and the delighted, not the average case, so building the tree's anchor split on it would have made the whole tree noisy and easy to nudge by design. Second, the real failure mode worth naming plainly: a bot can sound completely sure and still be wrong, a confident answer about an outage ETA or a billing charge that the customer doesn't push back on, not because it was right, but because they didn't trust themselves enough to argue with it. That chat reads as "resolved, no repeat contact" in the data and is actually a hallucination sitting quietly inside your best-looking number. The guardrail is the same weekly transcript sample that built the ninety percent trust bar in the first place, specifically flagging any resolved chat where the bot stated a number, a date, or a charge, and checking it against the real record. There's a real cost to this trade too: the 48 hour wait window means the resolved leaf can't be finalized same day, so the tree reports a number tagged "provisional" until the window closes, trading same-day freshness for a number worth defending in a room.
And if you want to be sure it really works, try it somewhere else
Same five letters, a different vertical, so the method proves itself instead of repeating a story that happened to work once for an ISP.
Sable Health runs a patient portal assistant that answers billing questions, reschedules appointments, and checks insurance coverage before routing anything to a human scheduler.
S, situation. Before this design, Sable's ops team reported one number to the hospital's finance committee: "self-service completion rate," forty eight percent of portal chats that never reached a scheduler.
P, payoff. The habit worth building: the finance committee stops treating that one rate as proof of savings and starts asking which branch of it is real.
A, anchor. Split "self-service completion" into confirmed (the appointment or billing question actually resolved, checked against whether the patient shows up or calls back) versus silent drop-off (the patient closed the app mid-flow, no confirmation, no record of what they needed next).
R, risk. A budget review where the forty eight percent gets used to justify not filling two open scheduler positions, while a chunk of that number is actually patients who gave up trying to book and simply didn't show for care they needed.
K, keep out. Day one skips splitting by insurance type, skips a no-show cost estimate, and skips flagging which specialty gets the most drop-off. All real, none trustworthy yet.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, resolved split from abandoned at the first leaf, and say in one line why the raw number can't be trusted without it.
Cost: engineering says the transcript-checked proxy can't ship for six weeks, only a simpler one can ship this week. Don't report the unchecked version as final. Mark it provisional and say plainly what confidence it hasn't earned yet.
The model got better, for real: say Tether's answer accuracy jumps ten points overnight. That still doesn't mean the abandoned branch shrinks to zero on its own, a better model can still get looped by a routing bug or a confusing menu. Keep checking the split; don't retire it because the model improved.
Where people run it wrong.
They report the raw top-line number because it is the one number everyone already understands, and skip the split that would make it honest.
They build every branch on day one, root cause, channel, cost, all at once, and end up trusting none of them because none were checked properly.
They treat the bot's own "chat closed" flag as ground truth instead of building a proxy that gets checked against real transcripts.
How to use it live. Say the anchor before naming anything else: "The tree only earns its name if the top number splits into resolved and abandoned before anything else. That's the one branch a leadership deck can't survive without." That buys you room to walk through the rest calmly instead of listing metrics off the top of your head.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the resolved proxy is wrong sometimes too?" Response: it will be, which is why it's checked weekly against a hand-read sample and reported with its own precision number, about 89 percent, instead of being presented as certain.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?