CaseIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #5

How do you measure whether an AI feature saved users time?

The direct answer
Measure it as an equation, not a feeling: the old write-up time minus the new review-and-fix time, per meeting, times only the meetings that actually finish through the AI draft instead of getting abandoned and rewritten by hand. Count the minutes a person spends fixing the draft as part of the "after" number, never just how fast the model finishes. Give the answer as a range, and check it against the team's real headcount before it goes in a deck.
Do this, in order
  1. State the equation before touching a single number: old time minus new time, per meeting, times meetings that actually finish through the draft.Why: without the equation stated first, every number that follows is a guess wearing data's clothes.
  2. Count review-and-fix time as part of the "after" number, never how fast the model finishes.Why: leaving that out is the single easiest way to overstate what an AI feature saved.
  3. Only multiply the per-meeting saving by meetings that finish through the draft, and count the abandoned ones at their old cost, not zero.Why: a meeting somebody rewrote from scratch saved nothing, and pretending otherwise inflates the total.
  4. Give a range, not one confident number.Why: a single number pretends the review time and the completion rate never move.
  5. Sanity-check the range against the team's real headcount and against what they used to spend in total.Why: catches a number that would mean more freed-up time than the team could possibly have had.
  6. Watch the review-time assumption hardest, since it swings the total more than any other single number.Why: it is also the one most likely to quietly creep up without anyone noticing.

How to answer this, stage by stage

Nobody is grading whether you can say a feature "saves time." They are grading whether you can build that claim out of real numbers, in front of them, and still trust it after you multiply it out. Seven moves get you there.

1
Scope it to one real product, one real team
Say it like this
"Let's ground this. Gistly is Tidewell's tool that listens to a video call and writes up the recap and the action items right after it ends. Xiomara Trevane runs product for it, and the team we're actually measuring is Kastellan Advisory, forty consultants, with Ynez Blackwood running client services on their side."
Why this works
Grounds the whole equation in one real team before a single number gets used, instead of a guess about AI features in general.
2
Say the equation out loud, before touching a number
Say it like this
"Here's how I'd frame it: time saved equals the old write-up time minus the new review-and-fix time, per meeting, times how many meetings actually finish through the AI draft instead of getting abandoned and rewritten by hand."
Why this works
States the method before a single number, so the interviewer hears a formula, not a feeling.
3
Own the "before" number, and say exactly where it came from
Say it like this
"Before Gistly, a Kastellan consultant spent about twelve minutes writing the recap and assigning owners by hand. That's not a guess. Tidewell's customer success team shadowed thirty consultants for two weeks before the rollout and timed it."
Why this works
A number with no source is a guess in a confident tone. This one has a source anyone could go check.
4
Own the "after" number, and don't let review time hide
Say it like this
"The draft itself shows up about ninety seconds after the call ends. But the real 'after' number is four minutes, that's how long a consultant spends checking it, fixing a wrong name or a task the model invented, and sending it. If I measured against ninety seconds instead of four minutes, I'd be claiming we erased the whole job, and nobody's calendar backs that up."
Why this works
This is the trap the whole question is testing. Leaving out review time is the easiest way to lie to yourself with real data.
5
Multiply by meetings that actually finish this way, not by every meeting
Say it like this
"About one in five of Kastellan's calls come out too rough to use, usually a noisy call with people talking over each other, and the consultant rewrites it from scratch instead. Those meetings save nothing, so I only multiply the eight minutes by the six hundred twenty four meetings a month that actually finish through the draft, not all eight hundred."
Why this works
Counting a meeting as a win just because the tool touched it is how you get a number that can't survive an audit.
6
Give a range, then check it against real headcount
Say it like this
"Depending on how messy the calls run that month, I'd put it somewhere between fifty six and a hundred and two hours a month across the team, call it eighty, give or take twenty. Split over forty consultants, that's about two hours a person a month, under two percent of anyone's month, and about half of what the team used to spend on write-ups entirely. That's a number I can defend."
Why this works
A single confident number is a guess in a suit. A range that survives being divided by headcount is a real estimate.
7
Name the assumption that would swing it most, and close
Say it like this
"If I had to watch one number closely, it's the four minutes of review time, not the completion rate. Let that creep up to eight minutes and the whole eighty-three-hour estimate falls to about forty two, more damage than even a bad drop in completion rate could do. So that's what I'd check every month, not just whether people are still opening the tool."
Why this works
Closes on the one lever that actually moves the answer, which is what a real estimator tracks after the meeting ends.
If you remember one thing "Did it save time" is never one number. It's an equation with at least four moving parts, and the part people forget to put in is the minutes a person spends fixing what the model got wrong.

Let's learn

Every week, a Kastellan Advisory consultant used to spend about an hour just writing up the notes from five client calls, by hand, before anyone had heard of Gistly.

Gistly is Tidewell's tool that listens in on a video call and writes the recap and the action items itself, right after the call ends. Before it existed, a consultant at Kastellan spent about twelve minutes per meeting typing up who said what and who owns what next. Tidewell's customer success team timed this properly, shadowing thirty consultants for two weeks before Kastellan's pilot ever started, so the twelve minutes is a real number, not a guess.

With Gistly running, the draft appears about ninety seconds after the call ends. But a consultant still spends about four minutes reading it, fixing a name it got wrong or an action item it invented out of two people talking at once, before sending it on.

Kastellan's team, eight hundred meetings a month, old way versus new way
160h 80h 0 Old way New way Given back 160 hrs 77 hrs 83 hrs
All eight hundred meetings, old process, cost the team about a hundred sixty hours a month. The same eight hundred meetings under the new process, four minutes review on the ones that finish through the draft plus the old twelve minutes on the ones that get rewritten by hand, cost about seventy seven hours. The gap is eighty three hours a month given back.
Knowledge spark: why does one in five meetings get abandoned? When three or more people talk over each other on a noisy call, Gistly sometimes invents a task nobody actually said, or hands it to the wrong person. Every action item links back to the exact clip it came from, so a consultant can check it in one click. Edit more than about fifteen percent of a meeting's items and it gets pulled out of the "completed" count and routed back for review, instead of quietly counted as a win.

Here is the turn. The four minutes a consultant spends reviewing each draft is not the problem. That review time is supposed to be there, it's how a wrong name or an invented task gets caught before it reaches a client. The real trap sits somewhere else: it is very easy to build the "after" number without those four minutes in it at all, because the model's own draft appears in about ninety seconds, and ninety seconds looks like the obvious number to put on a slide.

The four minutes was never the problem. Forgetting to count it was.

At its worst, this costs more than an awkward renewal call. If a team starts trusting the inflated number, someone might tell consultants to skip the review step to "capture the full time savings," and then a wrong action item, assigned to the wrong person, actually reaches a client with nobody catching it first.

The choice I would take back is how Tidewell first measured the "after" number, against how fast the model finishes a draft, not against how long a person actually spends with it. That was a fine shortcut in the earliest pilot, with one friendly customer and a demo to show a board. It stopped being fine once eight hundred real meetings a month were running through it and a renewal depended on the number holding up.

What I would leave alone: quick daily standups that nobody forwards to a client. Nobody was ever going to hold Gistly to a strict, defendable time-saved number there, so it doesn't need a shadow study or a sanity check. Just watch whether people keep opening it.

The lesson: an AI feature's time-saved number is not one number. It's an equation with at least four moving parts, old time, new time, review time, and how many actually finish this way, and the part most people forget to put in is the minutes a person spends fixing what the model got wrong.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how small the question was that Xiomara couldn't answer.

Xiomara Trevane can tell inside about ten seconds whether a number on a slide is going to survive a hard question. Two years running product for Gistly, Tidewell's meeting summarizer, taught her that the hard way.

For most of her first year on the job, the number on every slide was the same one: Gistly turns twelve minutes of write-up into about ninety seconds. It came from an early demo, one clean recording of one calm meeting, and it was true, as far as it went. Sales loved it. Customers loved it. Kastellan Advisory, a forty-person consulting firm, signed on as a pilot that spring, and by summer Ynez Blackwood, who ran client services there, was telling anyone who asked that Gistly had given her team their afternoons back.

Nobody checked the number again. Why would they. It kept working.

Then, on an ordinary Wednesday in October, Josiah Merriweather, three weeks into the job as a sales engineer, sat down to build the slide for Kastellan's renewal. He read the line about ninety seconds, then typed one question into the team channel before sending the deck out: does that number include the time people spend fixing the draft, or not?

Nobody answered right away. Xiomara looked at the question for a long moment before she realized she genuinely did not know.

It was not that the ninety-second number was a lie. It was that nobody had ever built the other half of it, the part where a person reads the draft and fixes what's wrong.

The real cost was never going to be an angry customer catching the gap on a call. It was going to be someone at Kastellan quietly running their own numbers before Tidewell did, and finding the truth first.

So Xiomara pulled the real data instead of guessing. Tidewell's usage logs already had a timestamp for when a draft appeared and a timestamp for when a consultant marked it sent. She lined those up against the twelve-minute shadow study from before the rollout, the one nobody had opened since launch. She also pulled the meetings that never got sent as a draft at all, the ones a consultant clearly gave up on and rewrote from a blank page instead.

The old decision she wanted to take back was small, and it made sense at the time. In that first demo, with one customer and a board meeting two days away, measuring against the draft's own speed was the fastest way to show the model worked at all. Review time wasn't even being logged yet. Nobody was lying. They just hadn't built the other half of the number.

The replay: by the following Monday, Xiomara had a real range, fifty six to a hundred and two hours a month for Kastellan's whole team, call it eighty, and a slide that said so, with the twelve minutes, the four minutes, and the six hundred twenty four completed meetings sitting right on it. Ynez Blackwood read it on the renewal call and said it matched what her team actually felt, closer to an afternoon a month than an afternoon a week. The deal renewed anyway. It renewed on a number that could survive somebody asking about it twice.

The old slide and the new slide were never really about ninety seconds versus eighty hours. They were about whether Tidewell could say a true thing under a follow-up question, which is the only kind of number worth putting in front of a customer.

What I would tell myself, sitting where Xiomara sat that Wednesday: the demo number was never dishonest on purpose. It just never had anyone ask it a second question. Build the second question in before somebody new has to ask it for you.

BOUND, walked through Kastellan's real numbers

This is a question about sizing a real claim with real assumptions, not a story about a habit switching, so BOUND fits, not a diagnosis or a design framework.

B
Break it down. State the equation before a single number.
Time saved per meeting equals the old write-up time minus the new review-and-fix time, times how many meetings finish through the draft instead of getting rewritten by hand.
O
Own numbers. Say where each one came from.
Twelve minutes, from a two-week shadow study before rollout. Four minutes and the six hundred twenty four completed meetings, both straight from Tidewell's usage logs, not a guess.
U
Use a range, not one confident figure.
Fifty six to a hundred and two hours a month for Kastellan's team, depending on how messy that month's calls run.
N
Nail the sanity check. Does the number survive being multiplied out?
Eighty three hours over forty consultants is about two hours a person a month, and about half of the hundred sixty hours the team used to spend on write-ups entirely. Both numbers are believable. A claim near a hundred forty hours, built by skipping review time, would mean the whole job nearly vanished. It didn't.
D
Direction. Which single assumption would move the answer most.
The four minutes of review time. Let it creep to eight and the estimate falls from eighty three hours to about forty two, more damage than even a bad drop in completion rate could do.

Two things worth saying here, since this is where an AI PM question earns its name. First, the alternative most teams reach for is just asking people: how many hours did Gistly save you this month? Xiomara's team tried that early on, and Kastellan's consultants self-reported about twenty five minutes saved a meeting, roughly three times the eight minutes the logs actually showed. People remember the one meeting where the draft was perfect and forget the one they quietly rewrote. That's why the real number came from timestamps, not a survey. Second, "completed" is not a rule that the draft must be perfect. A meeting counts once a consultant edits fewer than about fifteen percent of the drafted action items, checked against a held-out set of real recorded meetings, refreshed every quarter. Edit more than that and the meeting gets flagged and pulled out of the count instead of quietly logged as a win. There's a real trade-off sitting underneath all of it too: a cheaper, faster model was on the table, and it tested fine on clean recordings, but on Kastellan's actual noisy, multi-speaker calls it pushed the abandon rate from twenty two percent to about forty. Xiomara's team kept the slower, pricier model on purpose, because a model that costs less to run but saves less real time is a worse trade for a product that gets sold on hours saved.

What would change the eighty three hour claim the most, if it's wrong
what we'd claim: 83 hrs Review time creeps to 8 min 42 hrs Old baseline was really 9 min 52 hrs Completion rate drops to 60% 64 hrs Bars show hours saved if that one assumption alone turns out wrong. Shorter bar, bigger miss.
A slip in review time does more damage than a slip in the completion rate. That's why it's the one number worth checking every month, not just whether people keep opening the tool.

And if you want to be sure it really works, try it somewhere else

Same five letters, a completely different product and industry, so the method proves itself instead of repeating a story you happened to prepare.

Ossington Mutual sells ClaimBrief, a tool that reads an adjuster's site-visit notes, photos, and phone summary of a claim, and drafts the claim-file write-up and the next steps.

B, break it down. Time saved per file equals the old write-up time minus the new review time, times how many files actually close through the draft instead of getting redone by hand.
O, own numbers. Eighteen minutes a file by hand, from Ossington's own adjuster time logs. Seven minutes to review and fix the AI draft, mostly catching a misread damage estimate or the wrong policy clause.
U, use a range. About three hundred eighty to five hundred twenty hours a month across sixty adjusters, depending on how many files touch a rare policy exclusion the model handles badly.
N, nail the sanity check. Four hundred fifty one hours over sixty adjusters is about seven and a half hours a person a month, under five percent of anyone's working month. Believable.
D, direction. Not review time this time. The reopened-claim rate. A file paid out wrong because the model missed a coverage exclusion costs far more, weeks later in rework and in a customer's trust, than any amount of review time saved up front. Quentin Nazarenko, who runs product for ClaimBrief, watches that number harder than the raw minutes.

Same shape, different stakes At Kastellan, the hidden trap was leaving review time out of the "after" number. At Ossington, it's leaving the cost of a reopened claim out entirely, since that cost shows up weeks later, not the same day the file closes.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the equation, old time minus new time, times meetings that finish through the draft, not every meeting the tool touched.
Cost: the estimate has to ship today with no time for a real shadow study. Use the vendor's own logged timestamps as a floor, and say plainly it's a floor, not a finished number, until a real time study backs it up.
The model got better, for real: Gistly's accuracy jumps ten points overnight. That still doesn't excuse skipping review time in the equation. A better model just moves the completion rate up. It doesn't make the four minutes of checking disappear.

Where people run it wrong.
They measure "after" against how fast the model answers, not how long a person spends with what it gave them.
They count every meeting the tool touched as a win, even the ones somebody quietly rewrote by hand.
They report one confident number instead of a range, so the first hard question breaks it.

How to use it live. Say the equation out loud before a single number: old time minus new time, times how many actually finish this way. That buys you the room to ask what "finish" even means, instead of guessing at a headline number on the spot.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits measuring whether an AI feature actually saved users real time, and why?
Tap to flip
ANSWER
BOUND. It's an estimation question, you build an equation and own the assumptions out loud, not a story about a habit that flips.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Xiomara Trevane, senior product manager for Gistly at Tidewell, checking the real time-saved number for the Kastellan Advisory account, led on the client side by Ynez Blackwood.
3 · THE EQUATION
What's the equation for time saved here?
Tap to flip
ANSWER
Time saved per meeting equals the old write-up time minus the new review-and-fix time, times how many meetings actually finish through the AI draft instead of getting rewritten by hand.
4 · THE NUMBERS OWNED
Where did the twelve-minute "before" number come from?
Tap to flip
ANSWER
A two-week shadow study Tidewell's customer success team ran on thirty Kastellan consultants before Gistly ever rolled out, not a guess pulled from a demo.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Measuring the "after" number against how fast Gistly's draft appears, about ninety seconds, not against the four minutes a consultant spends fixing it. That was the fastest way to show the model worked in an early demo with one customer and a board meeting two days out. It stopped making sense once eight hundred real meetings a month depended on the number being true.
6 · THE NUMBER
Fill in the blank: Kastellan's consultants save about ___ minutes on each of the roughly ___ meetings a month that actually finish through the AI draft.
Tap to flip
ANSWER
8 minutes; 624 meetings. That works out to about 83 hours a month across the whole team.
7 · THE SANITY CHECK
How do you know the eighty three hour a month number isn't inflated?
Tap to flip
ANSWER
Check it two ways: it works out to about two hours a consultant a month, under two percent of anyone's month, and it's a little over half of the hundred sixty hours the team used to spend on write-ups before, not the whole thing.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product, in a different industry. Which one, and what's the equation's version there?
Tap to flip
ANSWER
ClaimBrief, the AI claim-file drafting tool at Ossington Mutual. Its equation is the old file write-up time minus the new review time, per claim, times how many files actually close through the draft instead of getting redone by hand.

Check yourself Score: 0 / 0

True or false
1. True or false: the fastest way to measure whether Gistly saved time is to compare the old write-up time against how quickly the model generates a draft.
  • True
  • False
Show hint
Think about what the model actually hands the consultant, and what the consultant still has to do with it.
Show answer
False. The draft appears in about ninety seconds, but the real "after" number is the four minutes a consultant spends checking and fixing it. Leaving that out inflates the time saved.
Multiple choice
2. Which single assumption would swing Kastellan's monthly hours-saved number the most if it turned out wrong?
  • A. The number of consultants on the team
  • B. The four minutes of review-and-fix time
  • C. How long each video call runs
  • D. Whether the draft appears in ninety seconds or two minutes
Show hint
Check the sensitivity chart in Section 3.
Show answer
B. Doubling the review time from four to eight minutes cuts the estimate from about 83 hours to about 42, more damage than a drop in completion rate would do.
Fill in the blank
3. Kastellan's consultants used to spend about ___ minutes writing a recap by hand. With Gistly, they spend about ___ minutes reviewing and fixing the draft.
Show hint
Check the chart in "Let's learn."
Show answer
12; 4. The eight-minute gap between those two numbers is what the whole estimate is built on.
Short answer, apply it yourself
4. Think of an AI tool you use yourself, something that drafts, sorts, or summarizes for you. What review or fix-up step do you do afterward that a simple "time saved" claim about that tool would probably leave out?
Show hint
Look for the step between the tool finishing and you actually trusting the result.
Show answer
Model answer: An AI email-drafting tool. The claim usually counts only how fast it writes a draft, not the minutes spent rereading it for tone and checking the facts before it actually gets sent. That checking time is real, and it's exactly the part a headline number tends to skip.
Short answer, the number question
5. If Kastellan's abandon rate doubled from 22 percent to 44 percent next month, but review time stayed at four minutes, would counting only completed meetings still be the right way to measure it? Show the reasoning.
Show hint
Separate the method from the outcome it produces.
Show answer
Yes, the method still holds. Completed meetings would drop from about 624 to about 448, and the monthly total would fall from around 83 hours to around 60. The number gets smaller, which is correct, since more meetings are getting rewritten by hand. The equation isn't broken by that, it's doing exactly its job: showing a true number is smaller than a hopeful one.
Multiple choice
6. Why does Xiomara reject asking Kastellan's consultants to self-report how many minutes Gistly saved them?
  • A. Because consultants aren't allowed to give feedback on internal tools.
  • B. Because self-reported estimates came in at about three times the number the actual usage logs showed.
  • C. Because Kastellan's contract forbids collecting any usage data.
  • D. Because Gistly's model cannot generate a report without a person's input.
Show hint
Check the rejected alternative in the BOUND recap.
Show answer
B. People remember the meeting where the draft was perfect and forget the one they quietly rewrote, so self-reports ran about three times higher than the logged, real number.
Before you close the answer
Why this works
Tests whether you can turn "did this save time" into real arithmetic instead of a vibe, and whether you know to count the minutes a person spends fixing the model's output, not just how fast the model answers.
Follow-up traps
"Isn't four minutes of review basically as fast as writing it yourself?" Response: no, writing twelve minutes of notes from a blank page and spending four minutes fixing a mostly right draft are not the same task. The eight-minute gap is real, and it's what the shadow study actually measured, not a guess.

"Why not count all eight hundred meetings, since Gistly touches every one?" Response: touching a meeting isn't the same as finishing it. The hundred seventy six meetings a month that get rewritten by hand saved nobody anything, and counting them as wins is exactly the kind of number that falls apart under a real audit.
If pressed
The fifteen percent edited-action-items threshold that decides whether a meeting counts as "completed" is checked against a held-out set of about 200 recorded Kastellan meetings, refreshed every quarter, so a model swap or a new call format, like a five-person video call versus a two-person phone call, gets tested before it can quietly change which meetings count.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more