CalculationAdvancedResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #14

What internal metrics prove an AI tool saved real time rather than shifted it?

LEAD the scenario: Briarcombe Dental Plan, a dental insurance claims processor whose reviewers use an AI tool to draft adjudication notes

Interviewer's question: "What internal metrics prove an AI tool saved real time rather than shifted it?" Dalton Osei supervises a claims-review team at Briarcombe Dental Plan. His reviewers use an AI tool that drafts the adjudication note for each flagged claim, and the team's official handle-time number has looked great for months.

The direct answer
Watch the clarification queue, not the handle-time average. Handle time only measures the step reviewers are willing to show you. If ambiguous claims are quietly being kicked to a supervisor for manual re-review, that queue's length is the real number, and it moves weeks before the official metric ever admits there's a problem. A tool that saved time shrinks that queue. A tool that just shifted the work grows it, while handle time keeps looking fine.
Do this, in order
  1. Track clarification-queue length by claim type, not just average handle time.Why: this is the leading signal, it drifts weeks before the official metric shows any sign of trouble.
  2. Never trust self-reported time-savings surveys as the primary evidence.Why: reviewers report what looks good, especially if they worry the tool might get pulled without a good showing.
  3. Set a real threshold on the clarification queue that triggers a pause, not just a conversation.Why: a metric with no action attached to it is a wall decoration, not a control.
  4. Break every metric out by claim type, not as one blended number.Why: an average hides the one claim type where work is quietly piling up somewhere else.
  5. Sample real claims periodically, the way an audit would, instead of trusting the dashboard alone.Why: the dashboard only shows what someone chose to log there, and the audit is what actually finds the gap.
  6. Expand the tool's scope only once both handle time and the clarification queue improve together.Why: one metric moving while the other worsens is exactly the pattern that means time got shifted, not saved.

How to answer this, stage by stage

Nobody is grading whether you can name "average handle time." They're grading whether you can name the metric that hides the problem right up until it's a crisis.

Stage 1
Scope it to one team and one real tool
Say it like this
"I'll answer this for Dalton Osei's claims-review team at Briarcombe Dental Plan, where reviewers use an AI tool to draft adjudication notes."
Why this works
Anchors an abstract metrics question in one real team's actual dashboard.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link to the real outcome, find the early signal, name how it gets gamed, then say what decision each reading should trigger."
Why this works
Signals a method for a question that otherwise invites naming a metric with no reasoning behind it.
Stage 3
Name the real outcome, not the model's score
Say it like this
"The real outcome is claims genuinely processed faster, without the extra work just landing somewhere nobody's counting."
Why this works
Grounds the answer in the business result, not the AI tool's own accuracy number.
Stage 4
Name the early signal, and why it's the hard part
Say it like this
"The clarification queue, the ambiguous claims kicked to a supervisor for manual re-review. It grows quietly, weeks before handle time ever looks bad."
Why this works
This is the direct answer, and the hardest step in LEAD, since it's tempting to just say "watch handle time" and stop.
Stage 5
Name how the metric gets gamed
Say it like this
"A survey asking reviewers how much time they saved gets gamed easily, some worry the tool gets pulled if the number looks bad, so they overstate it, even while their real timestamps show the work just moved."
Why this works
Shows you know every metric has a way to be hit without doing the real work.
Stage 6
Say the decision each threshold triggers
Say it like this
"If the clarification queue for a claim type crosses a set point, we pause auto-drafting there and investigate. If it stays flat or falls while handle time also falls, we expand."
Why this works
A metric nobody acts on is a dashboard decoration. This makes it a real control.
Stage 7
Close on the one line
Say it like this
"Watch the clarification queue, not the handle-time average, because that's the number that would look perfectly healthy right up until an audit found the work that never disappeared, it just moved."
Why this works
Restates the direct answer, ready for whatever gets pushed on next.

Let's learn

What happens the first time a number looks perfectly healthy right up until the morning someone actually checks it?

Briarcombe Dental Plan's AI tool reads a flagged claim and drafts the adjudication note, findings, coverage decision, next steps, for a human reviewer to check and sign off on. Before the tool, a reviewer wrote every note from scratch, about eighteen minutes a claim.

Knowledge spark: what's a clarification queue? A holding area for claims a reviewer isn't confident enough to sign off on alone. Instead of guessing, they route it to a supervisor for a second, manual look. It's a completely normal safety valve, right up until it starts absorbing work that should have been handled the first time.

With the AI tool drafting first, average handle time dropped to about seven minutes a claim. On the dashboard, that number looked like a clean win, month after month.

Hand sketched flow diagram titled Where the hidden work goes. Four steps: claim drafted, reviewer unsure, clarify queue highlighted, supervisor redoes it.
Four steps, and the third one never shows up in the number everyone was watching.

The turn: the extra speed was never really extra speed. A share of claims a reviewer wasn't confident enough to sign off on alone were quietly getting routed to a supervisor's clarification queue, work that still had to happen, just somewhere the handle-time dashboard never looked.

The decision I would take back Briarcombe tracked average handle time as the tool's one success metric from day one. That felt like the obvious, clean number to watch at launch. It quietly missed the one place the real cost was hiding, a queue that isn't anyone's official job to measure.

At its worst: leadership expands the AI tool to every claim type based on a handle-time chart that looks great, the clarification queue keeps growing underneath it, and a state regulator's routine audit eventually finds a real backlog of unresolved claims nobody upstream ever saw coming.

What I would leave alone: the AI tool's actual drafting speed and wording quality don't need to change at all. This isn't a model problem, it's a measurement problem, and fixing the metric fixes the visibility without touching the tool itself.

The lesson: a metric that only measures the part everyone's willing to show you will always look healthy, right up until someone finally checks the part nobody was measuring.

Now here is the same thing as a story

The short version above is what you'd say to Briarcombe's leadership team. Read this one for how the actual audit found it.

Every Monday, Dalton Osei pulls the week's handle-time report before anyone else is at their desk, a habit from years of running the floor before the AI tool ever existed.

Hand sketched timeline titled The audit that found it. Four milestones: handle time drops week one, queue quietly grows week six highlighted, audit samples twenty claims week ten, six of twenty were hidden found.
Ten weeks of a great-looking number, and it took an outside sample to see what it was hiding.

For months, the Monday number kept looking better. Dalton didn't think much of the clarification queue ticking up slightly, week over week, since it had always existed as a normal safety valve for genuinely hard claims.

Then Briarcombe's internal audit team pulled a random sample of twenty claims marked "fast" in the system, the kind of claim the handle-time report would call a clean success. Six of the twenty had actually spent hours sitting in the clarification queue before a supervisor redid the work by hand, time the official handle-time number never counted at all.

The claims weren't fast. They were fast on a clock that stopped counting the moment the real work moved somewhere else.

Dalton pulled the reviewers' own self-reported survey answers next, the ones leadership had been citing as proof the tool was a hit. Reviewers had reported saving, on average, over eleven minutes a claim, several minutes more than the timestamps could actually account for.

Hand sketched comparison diagram titled How it gets gamed. Left panel, a document icon labeled Survey answer, caption checks the box, says huge time saved. Right panel, a question mark icon labeled Real timestamps, caption the same work, just moved to another queue.
Only one of these two numbers can actually be checked against a clock.
Clarification-queue share of claims, week by week
30% 15% 0% Audit: 26% Week 1 Week 10
The queue climbed for ten straight weeks while the handle-time chart on the same dashboard kept looking fine.
Real claim backlog age, before and after the audit's fix
10d 5d 0 9 days Before the fix 2 days After the threshold
This was the number the handle-time report was never built to catch at all.

Dalton put the clarification-queue share on the same dashboard as handle time, with a hard rule attached: if any claim type's queue share crossed 20 percent, auto-drafting paused there until a supervisor reviewed why. The real backlog, the one that mattered to an actual patient waiting on a claim, dropped from nine days to two within a month.

LEAD, in one screenNot a dashboard redesign. LEAD is what tells you which number to trust before an audit has to find the real one.

L
Link. The real outcome, not the model's score.
Claims genuinely processed faster, without the extra work landing somewhere nobody's counting.
Anchors everything that follows in the business result, not a vanity number.
E
Early signal. The hardest step, and the answer.
Clarification-queue share by claim type, rising weeks before handle time shows any sign of trouble.
Accuracy at 96 percent tells you nothing. A queue share climbing from 4 to 26 percent tells you everything.
A
Abuse. How the metric gets gamed.
Self-reported time-savings surveys overstate real gains, since reviewers report what looks good, especially under any pressure to justify the tool.
Every metric has a cheap way to be hit without doing the real work, name it before someone finds it for you.
D
Decision. What each reading actually triggers.
A queue share crossing 20 percent for a claim type pauses auto-drafting there. A queue share flat or falling alongside falling handle time triggers expansion.
A metric nobody acts on is a decoration. This makes it a real lever.
Hand sketched icon list titled Three signals leadership actually watches. Three items: average handle time the visible one, clarification queue length the hidden one, self reported time saved the gameable one.
Only one of the three actually tells you the truth before it's too late.

The recap, one line per letter: link is real claims processed faster without hidden extra work, early signal is the clarification-queue share climbing weeks ahead of trouble, abuse is a self-reported survey overstating real savings, and decision is a named threshold that pauses or expands the rollout.

And if you want to be sure it really works, try it somewhere elseSame four letters, a school district's IT helpdesk instead of a dental insurer. Nothing else about the two jobs is alike.

Marrow Bend School District's IT helpdesk uses an AI tool to draft responses to tech-support tickets. Tickets get marked "resolved" quickly, while a hidden "escalated tickets" queue, the ones a technician wasn't confident enough to close alone, quietly grows underneath.

Hand sketched quadrant titled Visible against hidden, cheap against expensive. Axes cost to notice and cost if missed. Handle time sits low cost to notice, low cost if missed. Clarification queue sits high cost to notice, high cost if missed. Survey response sits moderate on both.
The number worth watching sits alone in the corner nobody checks by default.

Mapped onto LEAD: link is tickets genuinely resolved faster, not just marked closed. Early signal is the escalated-ticket share by issue type, climbing quietly while "time to close" looks great. Abuse is technicians marking tickets resolved to hit a closure-rate target, even when the real issue reopens days later. Decision is pausing AI-drafted responses for any issue type whose escalation share crosses a set point, and expanding only where both numbers improve together.

Hand sketched labeled parts diagram titled What handle time does not count. Center gauge icon labeled Handle time, with four callouts: clarification queue hours, supervisor redo time, appeal follow up calls, off system notes.
Four kinds of real work, and none of them show up in the one number everyone was watching.

Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "watch the clarification queue, not the handle-time average, because that's the number that hides the real problem," and stop.
Cost: if there's no budget to build a proper queue-tracking dashboard, start by hand-sampling twenty claims a week like the audit did, cheap and honest until the real system exists.
The model gets better, for real: if the AI tool's draft quality genuinely improves, the queue share should fall on its own, and that's exactly the confirming signal that time really is being saved, not just shifted.

Where people run it wrong.
They treat handle time as the whole story, since it's the number that's easiest to put on a dashboard.
They trust a self-reported survey as real evidence instead of a marketing number.
They never break a metric out by claim type, so one bad category hides inside a healthy-looking average.

How to use it live. When someone asks you how to know if an AI tool saved real time, ask yourself first: where would the extra work go if it didn't actually disappear. Go measure that place, not the place that's easy to measure.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what metric proves time was saved rather than shifted"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Built for metric questions, not a design or estimation method.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dalton Osei, who supervises a claims-review team at Briarcombe Dental Plan, and pulls the week's handle-time report every Monday morning.
3 · THE LINK
What's the real business outcome here, not the model's own score?
Tap to flip
ANSWER
Claims genuinely processed faster, without the extra work quietly landing in a queue nobody was measuring.
4 · THE EARLY SIGNAL
What's the leading indicator here, and what does it catch before the lagging metric does?
Tap to flip
ANSWER
Clarification-queue share by claim type. It climbed from 4 percent to 26 percent over ten weeks while handle time kept looking perfectly healthy.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Tracking average handle time as the tool's one success metric from launch, which quietly missed the clarification queue where the real cost was hiding.
6 · THE NUMBER
Fill in the blank: the audit sampled 20 claims marked "fast," and ___ of them had actually spent hours in the clarification queue.
Tap to flip
ANSWER
6 of 20. Nearly a third of the "fast" claims had hidden work the handle-time number never counted.
7 · THE REPLAY
Same audit finding, new dashboard. What changes?
Tap to flip
ANSWER
Auto-drafting pauses automatically once a claim type's queue share crosses 20 percent, and the real backlog age drops from 9 days to 2 within a month.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different org. Which one, and what's the hidden queue there?
Tap to flip
ANSWER
Marrow Bend School District's IT helpdesk. The escalated-ticket queue is the hidden signal, growing while "time to close" looks great.

Check yourself Score: 0 / 0

Short answer, apply it yourself
1. Think of a tool at your own job with a metric everyone watches on a dashboard. What hidden queue or step might be absorbing work that metric never counts?
Show hint
Think about the step someone takes when they aren't confident enough to finish the task alone.
Show answer
Model answer: Something like a "needs manager approval" queue on an approvals dashboard, which can quietly absorb the exact cases the automated step couldn't actually handle.
Multiple choice
2. Why is the clarification-queue share a stronger metric here than average handle time?
  • A. It's easier to calculate than handle time.
  • B. It moves weeks before handle time shows any problem, catching hidden work handle time never counts.
  • C. Reviewers prefer being measured on it.
  • D. It's required by dental insurance regulations.
Show hint
Look at the line chart and its caption.
Show answer
B. That's the whole point of a leading indicator: it tells you the truth before the lagging metric does.
True or false
3. True or false: reviewers' self-reported time-savings survey was more reliable evidence than the real system timestamps.
  • True
  • False
Show hint
Look at the comparison diagram in the story.
Show answer
False. The survey overstated real savings by several minutes a claim, exactly the abuse pattern LEAD's A step is built to catch.
Fill in the blank
4. Fill in the blank: the real claim backlog age dropped from 9 days to ___ days within a month of enforcing the queue-share threshold.
Show hint
Look at the second bar chart.
Show answer
2 days. This was the number the original handle-time dashboard was never built to catch.
Short answer, where it wouldn't matter
5. Name a part of the AI tool this fix genuinely didn't need to touch.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The tool's actual drafting speed and wording quality. This was a measurement problem, not a model problem, so the fix lived entirely in the dashboard.
Short answer, the number question
6. If the clarification-queue share had stayed flat at 4 percent the whole ten weeks instead of climbing to 26, would the audit likely have found the same problem? Why or why not?
Show hint
Think about what the audit was actually checking for.
Show answer
Model answer: No, a flat 4 percent queue share alongside falling handle time would be the real signal that time was genuinely saved, not shifted, and the audit would likely have found a healthy system instead.
Before you close the answer
Why this works
Tests whether you'll trust a dashboard number that's already easy to see, or go looking for the hidden step where real work quietly piles up instead.
Follow-up traps
"Isn't a growing clarification queue just a sign the tool is being appropriately cautious?" Response: a small, stable queue is healthy caution. A queue that keeps climbing month over month while handle time stays flat is a sign work is accumulating somewhere nobody's acting on it.

"What if reviewers just stop routing to the queue to make the number look good?" Response: that's exactly why the queue share gets paired with a periodic real-claim audit, the same kind that found the original gap, so gaming one number alone doesn't go unnoticed.
If pressed
Briarcombe's actual threshold isn't a flat 20 percent for every claim type. Complex claim categories get a higher allowed queue share from the start, since some ambiguity there is expected and healthy, while routine claim types get a much tighter bar.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more