Artifact critiqueAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #8

Design the metric tree for an AI-powered support deflection feature.

The direct answer
Build the metric tree so the top number, deflection rate, splits into two leaves the moment it leaves the model: resolved (the customer's issue actually got handled, checked against a proxy signal, not just handled by the bot) and abandoned (the customer gave up, closed the tab, or came back within 48 hours). Report the resolved leaf as the real number, not the raw total. Everything else on the tree, channel, cost, root cause, can wait. That one split cannot, because a raw deflection rate can look great while abandonment is doing most of the work underneath it.
Do this, in order
  1. Split "did not reach a human" into resolved and abandoned at the very first leaf, before any other branch.Why: a raw deflection number blends a customer who got real help with a customer who just quit. Those are opposite outcomes wearing the same number.
  2. Define "resolved" with a real proxy signal, no repeat contact on the same issue within 48 hours and no negative signal, then check that proxy against a hand-labeled set of real chats.You cannot ask every customer if the bot actually helped them, so the proxy has to earn trust before anyone puts it in a slide.
  3. Report the resolved leaf as the number leadership acts on, not the raw deflection total.The raw total is what gets read out loud in a review. If it is not the honest number, the honest number never gets heard.
  4. Add one layer under resolved for task type, billing, outage, plan change, since that is the ISP's own real work split.A tree with no second layer under its safest leaf still cannot tell anyone which part of the product is actually earning the number.
  5. Sample a slice of "resolved" chats by hand every week and check for confident, wrong answers.A bot that sounds sure and is wrong looks exactly like a resolved chat in the data. Only a human reading the transcript catches it.
  6. Leave out root-cause tagging, channel splits, and a dollar-per-resolution branch on day one.Stacking more branches onto an unproven split just multiplies whatever is wrong with the split underneath them.

How to answer this, stage by stage

Nobody is grading whether you can draw boxes and arrows. They are grading whether the top number in your tree survives the first time someone in a leadership review asks "why did it move." Eight moves get you there.

1
Scope it to one real product, one real report
Say it like this
"Let's ground this. Tether is Cordwell Fiber's chat assistant. It handles billing questions, outage reports, and plan changes before a human agent ever picks up. Tomi Idowu owns the metric tree that feeds Rian Pellham's monthly leadership review, and that review is where headcount decisions for the contracted support team actually get made."
Why this works
Grounds the design in a real product, a real report, and a real decision, not a hypothetical dashboard nobody acts on.
2
Say the real question you're designing for
Say it like this
"A metric tree isn't really a chart. It's the thing that lets you answer 'why did this number move' without a follow-up meeting. If the tree can't answer that on its own, it's not a tree, it's one number wearing a nicer font."
Why this works
States the design goal before naming a single branch, so the anchor decision reads as earned, not arbitrary.
3
Name the habit today, without the tree
Say it like this
"Right now Tomi pulls one number, deflection rate, chats that never reached a human divided by total chats, and sends it up. Sixty one percent this month. Nobody below that number can say what it's actually made of."
Why this works
This is the S step. Shows the gap the design is filling, a single unexamined ratio standing in for a real answer.
4
Give the one decision the whole tree hangs on
Say it like this
"The anchor is the split right under that top number. Of the chats that never reached a human, how many actually got resolved, no repeat contact within 48 hours, no thumbs-down, versus how many were really abandoned, the customer closed the tab or came back on the same issue within two days. Everything else on the tree can wait. That split can't."
Why this works
This is the A step, the one concrete design decision a reader can point at and argue with.
5
Show what breaks the first time the split is missing
Say it like this
"Two days before a review, Rian was ready to walk in with sixty one percent deflection as the reason not to renew a contract for twelve support agents. When Tomi actually ran the split, only forty six of those sixty one points were real resolutions. Fifteen points was people giving up. Cut the twelve agents on the raw number, and those fifteen points of abandoned customers show up as hold-time spikes the very next month."
Why this works
This is the R step, made concrete with an actual headcount decision instead of an abstract warning about bad data.
6
Say how you'd earn trust in the split, not just declare it
Say it like this
"We don't get to ask every customer if the bot actually helped them. So we built a proxy, no repeat contact in 48 hours, and checked it against two hundred and fifty hand-read transcripts. It called a chat resolved correctly about eighty nine times out of a hundred. That's the number we're willing to report. A proxy nobody's checked is just a guess with a clean UI."
Why this works
Shows the eval-set thinking that separates a real metric design from a metric that sounds precise but was never checked against anything.
7
Say what you're deliberately leaving off, and why that's safe
Say it like this
"Day one, the tree stops at task type under resolved. It does not yet tag why each chat got abandoned, it does not split by app versus web, and it does not carry a dollar figure per resolution. Those are all real questions. None of them are trustworthy questions until the resolved and abandoned split underneath them has been live long enough to believe."
Why this works
This is the K step. It shows restraint is a design choice, not a gap nobody noticed.
8
Close on the anchor, defended in one line
Say it like this
"So the tree is: total chats, then reached-a-human versus didn't, then inside didn't, resolved versus abandoned, checked against real transcripts, then task type under resolved. The split under the top number outranks every other branch, because it's the one place a good-looking number can hide a bad outcome."
Why this works
Closes on the rule itself, something a reader can apply to a metric tree they've never seen, not just this one.
If you remember one thing A deflection number that doesn't separate a customer who got helped from a customer who gave up isn't a metric. It's a guess wearing a metric's clothes.

Let's learn

Every evening, Rian Pellham, VP of Support at Cordwell Fiber, an internet provider that serves about ninety thousand homes, opens a slide with one number on it. Tether, Cordwell's chat assistant, answers billing questions, logs outage reports, and processes plan changes before a customer ever reaches a person. Before Tether existed, every one of those requests went straight to a phone queue, an average hold of eleven minutes at peak hours.

With Tether running, about forty thousand customers a month start a chat instead of calling. Sixty one percent of those chats never touch a human agent at all. That number, deflection rate, is the one Rian has reported for a year. It has always gone up. It has always looked good.

Knowledge spark: what does "deflection" actually count? A deflected chat is any conversation that ends without a customer reaching a human agent. That's it. It says nothing about whether the customer's problem got solved, or whether they just stopped trying.

Here is the turn. A chat that never reaches a human is not automatically a good outcome. Some of those sixty one points are customers whose billing question got answered correctly. Some of them are customers who typed "representative" three times, got looped back to the same menu, and closed the tab. Both count as deflected. Only one of them is actually a win.

Hand sketched decision tree titled The anchor close up, the split under the top number. Root box says forty thousand chats this month. Three branches lead down: reached a human leads to escalated thirty nine percent, no repeat contact no bad signal leads to resolved forty six percent, closed tab looped or came back in forty eight hours leads to abandoned fifteen percent.
The raw sixty one percent is escalated subtracted from the whole. The real question sits one layer lower, inside that number, resolved against abandoned.
Deflection rate, raw versus adjusted for real resolution
61% 46% 0 Raw deflection Resolved only 61% 46%
Fifteen of those sixty one points are chats that never reached a human but were never really solved either. They sit inside the raw number with nowhere to be seen until the tree splits them out.
We did not just measure how many chats skipped a human. We measured how many customers quietly gave up, and called it a win.

At its worst, this costs more than a wrong slide. Say the raw number goes into the review unsplit, and Rian recommends letting a staffing contract for twelve agents lapse, since deflection alone seems to justify a smaller team. The fifteen points of abandoned customers do not disappear. Most of them call back, now angrier, now with a longer story to tell, right as the team meant to handle them is gone.

Hand sketched comparison diagram titled The day it's wrong. Left panel, a gauge icon labelled What the deck showed, caption Deflection sixty one percent, looks like a clean win. Right panel, a red question mark box labelled What was actually true, caption Only forty six percent resolved, fifteen points was abandonment hiding inside the count.
The deck and the truth used the same number to say two different things. The tree exists to stop that from happening twice.

The choice I would take back is not the sixty one percent itself. It's that Tether's dashboard was built, back when Tether first launched, to report one ratio at the top and nothing underneath it. That made sense in month one, when the whole team was just trying to prove the chatbot could take load off the phone queue at all. It stopped making sense the day that one ratio became the number a headcount decision leaned on.

What I would leave alone: the escalated branch, chats that do reach a human, does not need this same scrutiny. A chat that reaches a person either gets handled by that person or it doesn't, and that outcome is already tracked the same way it was before Tether existed. The ambiguity only lives inside the branch where nobody human ever looked.

The lesson: a single number at the top of a dashboard is not a metric, it's a summary. A metric tree is what you get when you refuse to let that summary travel further than it can be trusted.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how close this one came to shipping wrong on an ordinary Tuesday.

Tomi Idowu has run analytics for Cordwell Fiber's support org for three years. She is the person Rian calls before any number goes on a slide, because she is the one who actually knows what a metric is built from, not just what it says.

Tether launched eighteen months ago and it worked, plainly and quickly. Hold times on the phone queue dropped from eleven minutes to six within two months. Tomi built one dashboard tile for it: deflection rate, and it climbed steadily, forty two percent in month one, past fifty by month six, sixty one by the eighteenth month. Every one of Rian's reviews opened with that tile, and every one of those reviews went well.

The trigger was almost nothing. On a Wednesday afternoon, a support supervisor named Aveline stopped by Tomi's desk to complain about something unrelated and mentioned, almost in passing, "I keep getting calls from people who say they already tried the chatbot for this exact thing." Tomi wrote it down and moved on. It was one sentence, said while someone was looking for a stapler.

It was not a complaint. It was a crack, and Tomi almost let it close over.

Two days later, prepping the slide for that month's review, Tomi pulled a sample of two hundred and fifty "deflected" chats and actually read them. Not the summary, the transcripts. Forty six of the sixty one points held up, real questions, real answers, no comeback. But close to fifteen points were customers who typed "agent," got looped through the same three menu options, and closed the tab. Deflected, technically. Solved, not even close.

Rian's slide, drafted the night before, said this: "Deflection at sixty one percent supports letting our twelve-agent staffing contract lapse at renewal." He was two days from presenting it.

Hand sketched labeled diagram titled Today, without you, one number, no leaf. A gauge in the center reads deflection sixty one percent. Four labels point outward: one number a month, no split beneath it, straight into the deck, nobody asks why it moved.
This was the dashboard Tomi had built and trusted for a year, one honest-looking number with nothing holding it up from underneath.

Tomi went back to Rian with the split, not a hunch, the actual read-through. Forty six percent really resolved. Fifteen points was people giving up. Cutting twelve agents on the strength of the raw number would have meant those fifteen points of frustrated customers landing back on a smaller team weeks later, right when nobody was staffed to catch them.

Rian pulled the slide the morning of the review. What went up instead was smaller and less flattering: forty six percent, resolved and checked, next to a note that the team was building a permanent split so this never had to be caught by hand again. The staffing contract got renewed for one more quarter while the real number proved itself.

The replay, six weeks later: the resolved leaf held steady at forty five to forty seven percent, checked weekly against fresh transcripts. The abandoned leaf, once visible, started shrinking on its own, because a chatbot's builders started fixing what abandoned chats had in common, mostly customers stuck in a loop asking for a person and never getting routed to one. By the next review, raw deflection had actually dropped to fifty eight percent, but resolved had climbed to fifty one. A worse-looking top number, a better real one.

What I would tell my past self, the one who built a dashboard with one tile on it: the number that goes in the deck is never really the number itself. It's whatever the number was built to hide.

SPARK, in one screen

This is a design question, not a story about a habit with two settings, so SPARK fits, built forward instead of run backward like a recovery framework.

S
Situation. Who does the job today, and how, without this design.
Tomi reports one ratio, chats that skipped a human over total chats, with nothing underneath it a reader can check.
P
Payoff. The habit this should build.
Whoever reads the number stops quoting the top line alone and starts asking which leaf is actually driving it, every time, before it goes further up the chain.
A
Anchor. The one decision everything else hangs on.
Split "resolved" from "abandoned" at the very first leaf under the top number, defined by a checked proxy, not a guess.
R
Risk. What breaks the first time you're wrong.
A leadership review where the raw number justifies cutting real staff, while a hidden abandonment branch is actually doing the work behind that number.
K
Keep out. What day one deliberately skips.
Root-cause tagging, channel splits, and a dollar figure per resolution, all real, none trustworthy until the resolved and abandoned split has proven itself.

Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. First, the alternative Tomi actually considered and rejected was using the post-chat thumbs-up rating as the resolved-versus-abandoned split, since it was already collected and looked like a clean signal. She ruled it out. Response rate on that rating sat around eight percent, and the people who bother to answer skew toward the furious and the delighted, not the average case, so building the tree's anchor split on it would have made the whole tree noisy and easy to nudge by design. Second, the real failure mode worth naming plainly: a bot can sound completely sure and still be wrong, a confident answer about an outage ETA or a billing charge that the customer doesn't push back on, not because it was right, but because they didn't trust themselves enough to argue with it. That chat reads as "resolved, no repeat contact" in the data and is actually a hallucination sitting quietly inside your best-looking number. The guardrail is the same weekly transcript sample that built the ninety percent trust bar in the first place, specifically flagging any resolved chat where the bot stated a number, a date, or a charge, and checking it against the real record. There's a real cost to this trade too: the 48 hour wait window means the resolved leaf can't be finalized same day, so the tree reports a number tagged "provisional" until the window closes, trading same-day freshness for a number worth defending in a room.

Knowledge spark: why not just trust the bot's own "chat closed successfully" flag? Because that flag only reports what the bot thinks happened, not what actually happened to the customer. A model that hands out a wrong outage ETA with total confidence will still log the chat as closed successfully. The flag measures the bot's certainty, not the customer's outcome, and those are not the same thing.
Hand sketched numbered list titled Kept off the tree, day one. Four items with icons: one, why each chat got abandoned. Two, channel, app web or text. Three, dollar cost per resolution. Four, the model's own confidence score.
Every one of these is a real future branch. None of them earns a place on the tree until the resolved and abandoned split under them has run long enough to trust.

And if you want to be sure it really works, try it somewhere else

Same five letters, a different vertical, so the method proves itself instead of repeating a story that happened to work once for an ISP.

Sable Health runs a patient portal assistant that answers billing questions, reschedules appointments, and checks insurance coverage before routing anything to a human scheduler.

S, situation. Before this design, Sable's ops team reported one number to the hospital's finance committee: "self-service completion rate," forty eight percent of portal chats that never reached a scheduler.
P, payoff. The habit worth building: the finance committee stops treating that one rate as proof of savings and starts asking which branch of it is real.
A, anchor. Split "self-service completion" into confirmed (the appointment or billing question actually resolved, checked against whether the patient shows up or calls back) versus silent drop-off (the patient closed the app mid-flow, no confirmation, no record of what they needed next).
R, risk. A budget review where the forty eight percent gets used to justify not filling two open scheduler positions, while a chunk of that number is actually patients who gave up trying to book and simply didn't show for care they needed.
K, keep out. Day one skips splitting by insurance type, skips a no-show cost estimate, and skips flagging which specialty gets the most drop-off. All real, none trustworthy yet.

Same shape, different stakes At Cordwell, the hidden branch was a customer quietly furious about a bill. At Sable, it's a patient who quietly never got rebooked. The anchor doesn't change: whatever the raw number is hiding outranks the number itself, every time.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, resolved split from abandoned at the first leaf, and say in one line why the raw number can't be trusted without it.
Cost: engineering says the transcript-checked proxy can't ship for six weeks, only a simpler one can ship this week. Don't report the unchecked version as final. Mark it provisional and say plainly what confidence it hasn't earned yet.
The model got better, for real: say Tether's answer accuracy jumps ten points overnight. That still doesn't mean the abandoned branch shrinks to zero on its own, a better model can still get looped by a routing bug or a confusing menu. Keep checking the split; don't retire it because the model improved.

Where people run it wrong.
They report the raw top-line number because it is the one number everyone already understands, and skip the split that would make it honest.
They build every branch on day one, root cause, channel, cost, all at once, and end up trusting none of them because none were checked properly.
They treat the bot's own "chat closed" flag as ground truth instead of building a proxy that gets checked against real transcripts.

How to use it live. Say the anchor before naming anything else: "The tree only earns its name if the top number splits into resolved and abandoned before anything else. That's the one branch a leadership deck can't survive without." That buys you room to walk through the rest calmly instead of listing metrics off the top of your head.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits designing a metric tree for a new AI feature, and why?
Tap to flip
ANSWER
SPARK. It's a design question, building the tree forward before anything ships, not a reversal of a habit that already broke.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tomi Idowu, analytics lead for Cordwell Fiber's support org, who owns the metric tree behind VP Rian Pellham's leadership review.
3 · THE HABIT
What habit should a good metric tree build in the person reading it?
Tap to flip
ANSWER
Tracing the top-line number down to its real driver instead of quoting one deflection number in isolation. The P step.
4 · THE ANCHOR
What's the one design decision the whole tree hangs on?
Tap to flip
ANSWER
Splitting "resolved" from "abandoned" at the very first leaf under the top number, defined by a checked proxy signal, not a guess.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building the dashboard with one raw ratio and nothing underneath it. It made sense at launch, when the only goal was proving the chatbot took load off the phone queue at all.
6 · THE NUMBER
Fill in the blank: raw deflection ran ___ percent, but only ___ percent of chats actually held up as resolved once checked.
Tap to flip
ANSWER
61 percent raw; 46 percent resolved. The other 15 points was abandonment hiding inside a good-looking number.
7 · THE REPLAY
Same review, new tree, what actually changed by the next one?
Tap to flip
ANSWER
Raw deflection dropped to 58 percent but resolved climbed to 51 percent, a worse top-line number sitting on top of a genuinely better real one.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the anchor?
Tap to flip
ANSWER
Sable Health's patient portal assistant. Its anchor is splitting "confirmed" from "silent drop-off" inside its self-service completion rate.

Check yourself Score: 0 / 0

True or false
1. True or false: a chat that "deflects," meaning it never reaches a human agent, always means the customer's issue got solved.
  • True
  • False
Show hint
Check what "deflected" actually counts, in the knowledge spark in "Let's learn."
Show answer
False. Deflected only means no human was reached. It says nothing about whether the customer's problem got solved or they simply gave up.
Multiple choice
2. Why does the resolved-versus-abandoned split outrank every other branch on the tree?
  • A. Because it is the easiest branch to build with the data already collected.
  • B. Because a raw deflection number blends a customer who got real help with a customer who quit, and only this split separates them.
  • C. Because leadership specifically requested it by name.
  • D. Because it removes the need for a human review team entirely.
Show hint
Think about what the raw sixty one percent was actually made of.
Show answer
B. The raw number treats a solved chat and a chat someone gave up on as the same outcome. This split is the only branch that tells them apart.
Fill in the blank
3. Tether's raw deflection ran about ___ percent, but once checked against real transcripts, only about ___ percent held up as truly resolved.
Show hint
Check the chart in "Let's learn."
Show answer
61 percent; 46 percent. The gap, 15 points, was abandonment sitting quietly inside a number that looked like a clean win.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for a design choice made at launch, not a setting anyone could just turn back up.
Show answer
Model answer: Building Tether's dashboard as one raw ratio with nothing underneath it. It made sense in month one, when the whole point was proving the chatbot could take any load off the phone queue at all, before that number became the basis for a staffing decision.
Short answer, apply it yourself
5. Think of a tool you use that reports one summary number, an app's "battery health," a bank's "spending score," anything like it. What two real outcomes might that single number be blending together that you'd want split apart?
Show hint
Look for a number that could be produced by two very different underlying situations.
Show answer
Model answer: A fitness app's "recovery score" blends genuinely rested sleep with sleep that was just long but low quality, like restless, interrupted hours. Both could produce the same score, but only one of them actually means the person is ready to train hard.
Short answer, the number question
6. If the abandoned leaf had been 2 percent instead of 15 percent that month, would the resolved-versus-abandoned split still be worth building on day one? Show the reasoning.
Show hint
Think about what the size of the gap changes and what it doesn't.
Show answer
Yes, the split is still worth building. A small gap today doesn't mean it stays small. Without the split, nobody would even know it was 2 percent instead of 15, and the whole value of the anchor is catching the month it quietly grows, not just the month it's already large.
Before you close the answer
Why this works
Tests whether you design a metric that survives a follow-up question in a real leadership review, or hand over a single number that looks clean until someone asks what it's made of.
Follow-up traps
"Isn't building the whole split just slower than shipping the raw number now?" Response: the raw number ships immediately either way. The split just adds a "provisional" tag and a 48 hour window before a number gets called final, a small delay against a decision that could cost twelve real jobs if it's wrong.

"What if the resolved proxy is wrong sometimes too?" Response: it will be, which is why it's checked weekly against a hand-read sample and reported with its own precision number, about 89 percent, instead of being presented as certain.
If pressed
The resolved proxy actually runs two checks, not one: no repeat contact on the same issue within 48 hours, and no negative signal from a thumbs-down or a reopened billing ticket. A chat only counts as resolved if it clears both, which is part of why the 48 hour window at 89 percent precision beat a same-day 24 hour window, which only cleared about 74 percent against the same hand-read sample.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more