How do you measure whether an AI feature saved users time?
- State the equation before touching a single number: old time minus new time, per meeting, times meetings that actually finish through the draft.Why: without the equation stated first, every number that follows is a guess wearing data's clothes.
- Count review-and-fix time as part of the "after" number, never how fast the model finishes.Why: leaving that out is the single easiest way to overstate what an AI feature saved.
- Only multiply the per-meeting saving by meetings that finish through the draft, and count the abandoned ones at their old cost, not zero.Why: a meeting somebody rewrote from scratch saved nothing, and pretending otherwise inflates the total.
- Give a range, not one confident number.Why: a single number pretends the review time and the completion rate never move.
- Sanity-check the range against the team's real headcount and against what they used to spend in total.Why: catches a number that would mean more freed-up time than the team could possibly have had.
- Watch the review-time assumption hardest, since it swings the total more than any other single number.Why: it is also the one most likely to quietly creep up without anyone noticing.
How to answer this, stage by stage
Nobody is grading whether you can say a feature "saves time." They are grading whether you can build that claim out of real numbers, in front of them, and still trust it after you multiply it out. Seven moves get you there.
Let's learn
Every week, a Kastellan Advisory consultant used to spend about an hour just writing up the notes from five client calls, by hand, before anyone had heard of Gistly.
Gistly is Tidewell's tool that listens in on a video call and writes the recap and the action items itself, right after the call ends. Before it existed, a consultant at Kastellan spent about twelve minutes per meeting typing up who said what and who owns what next. Tidewell's customer success team timed this properly, shadowing thirty consultants for two weeks before Kastellan's pilot ever started, so the twelve minutes is a real number, not a guess.
With Gistly running, the draft appears about ninety seconds after the call ends. But a consultant still spends about four minutes reading it, fixing a name it got wrong or an action item it invented out of two people talking at once, before sending it on.
Here is the turn. The four minutes a consultant spends reviewing each draft is not the problem. That review time is supposed to be there, it's how a wrong name or an invented task gets caught before it reaches a client. The real trap sits somewhere else: it is very easy to build the "after" number without those four minutes in it at all, because the model's own draft appears in about ninety seconds, and ninety seconds looks like the obvious number to put on a slide.
At its worst, this costs more than an awkward renewal call. If a team starts trusting the inflated number, someone might tell consultants to skip the review step to "capture the full time savings," and then a wrong action item, assigned to the wrong person, actually reaches a client with nobody catching it first.
The choice I would take back is how Tidewell first measured the "after" number, against how fast the model finishes a draft, not against how long a person actually spends with it. That was a fine shortcut in the earliest pilot, with one friendly customer and a demo to show a board. It stopped being fine once eight hundred real meetings a month were running through it and a renewal depended on the number holding up.
What I would leave alone: quick daily standups that nobody forwards to a client. Nobody was ever going to hold Gistly to a strict, defendable time-saved number there, so it doesn't need a shadow study or a sanity check. Just watch whether people keep opening it.
The lesson: an AI feature's time-saved number is not one number. It's an equation with at least four moving parts, old time, new time, review time, and how many actually finish this way, and the part most people forget to put in is the minutes a person spends fixing what the model got wrong.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how small the question was that Xiomara couldn't answer.
Xiomara Trevane can tell inside about ten seconds whether a number on a slide is going to survive a hard question. Two years running product for Gistly, Tidewell's meeting summarizer, taught her that the hard way.
For most of her first year on the job, the number on every slide was the same one: Gistly turns twelve minutes of write-up into about ninety seconds. It came from an early demo, one clean recording of one calm meeting, and it was true, as far as it went. Sales loved it. Customers loved it. Kastellan Advisory, a forty-person consulting firm, signed on as a pilot that spring, and by summer Ynez Blackwood, who ran client services there, was telling anyone who asked that Gistly had given her team their afternoons back.
Nobody checked the number again. Why would they. It kept working.
Then, on an ordinary Wednesday in October, Josiah Merriweather, three weeks into the job as a sales engineer, sat down to build the slide for Kastellan's renewal. He read the line about ninety seconds, then typed one question into the team channel before sending the deck out: does that number include the time people spend fixing the draft, or not?
Nobody answered right away. Xiomara looked at the question for a long moment before she realized she genuinely did not know.
The real cost was never going to be an angry customer catching the gap on a call. It was going to be someone at Kastellan quietly running their own numbers before Tidewell did, and finding the truth first.
So Xiomara pulled the real data instead of guessing. Tidewell's usage logs already had a timestamp for when a draft appeared and a timestamp for when a consultant marked it sent. She lined those up against the twelve-minute shadow study from before the rollout, the one nobody had opened since launch. She also pulled the meetings that never got sent as a draft at all, the ones a consultant clearly gave up on and rewrote from a blank page instead.
The old decision she wanted to take back was small, and it made sense at the time. In that first demo, with one customer and a board meeting two days away, measuring against the draft's own speed was the fastest way to show the model worked at all. Review time wasn't even being logged yet. Nobody was lying. They just hadn't built the other half of the number.
The replay: by the following Monday, Xiomara had a real range, fifty six to a hundred and two hours a month for Kastellan's whole team, call it eighty, and a slide that said so, with the twelve minutes, the four minutes, and the six hundred twenty four completed meetings sitting right on it. Ynez Blackwood read it on the renewal call and said it matched what her team actually felt, closer to an afternoon a month than an afternoon a week. The deal renewed anyway. It renewed on a number that could survive somebody asking about it twice.
The old slide and the new slide were never really about ninety seconds versus eighty hours. They were about whether Tidewell could say a true thing under a follow-up question, which is the only kind of number worth putting in front of a customer.
What I would tell myself, sitting where Xiomara sat that Wednesday: the demo number was never dishonest on purpose. It just never had anyone ask it a second question. Build the second question in before somebody new has to ask it for you.
BOUND, walked through Kastellan's real numbers
This is a question about sizing a real claim with real assumptions, not a story about a habit switching, so BOUND fits, not a diagnosis or a design framework.
Two things worth saying here, since this is where an AI PM question earns its name. First, the alternative most teams reach for is just asking people: how many hours did Gistly save you this month? Xiomara's team tried that early on, and Kastellan's consultants self-reported about twenty five minutes saved a meeting, roughly three times the eight minutes the logs actually showed. People remember the one meeting where the draft was perfect and forget the one they quietly rewrote. That's why the real number came from timestamps, not a survey. Second, "completed" is not a rule that the draft must be perfect. A meeting counts once a consultant edits fewer than about fifteen percent of the drafted action items, checked against a held-out set of real recorded meetings, refreshed every quarter. Edit more than that and the meeting gets flagged and pulled out of the count instead of quietly logged as a win. There's a real trade-off sitting underneath all of it too: a cheaper, faster model was on the table, and it tested fine on clean recordings, but on Kastellan's actual noisy, multi-speaker calls it pushed the abandon rate from twenty two percent to about forty. Xiomara's team kept the slower, pricier model on purpose, because a model that costs less to run but saves less real time is a worse trade for a product that gets sold on hours saved.
And if you want to be sure it really works, try it somewhere else
Same five letters, a completely different product and industry, so the method proves itself instead of repeating a story you happened to prepare.
Ossington Mutual sells ClaimBrief, a tool that reads an adjuster's site-visit notes, photos, and phone summary of a claim, and drafts the claim-file write-up and the next steps.
B, break it down. Time saved per file equals the old write-up time minus the new review time, times how many files actually close through the draft instead of getting redone by hand.
O, own numbers. Eighteen minutes a file by hand, from Ossington's own adjuster time logs. Seven minutes to review and fix the AI draft, mostly catching a misread damage estimate or the wrong policy clause.
U, use a range. About three hundred eighty to five hundred twenty hours a month across sixty adjusters, depending on how many files touch a rare policy exclusion the model handles badly.
N, nail the sanity check. Four hundred fifty one hours over sixty adjusters is about seven and a half hours a person a month, under five percent of anyone's working month. Believable.
D, direction. Not review time this time. The reopened-claim rate. A file paid out wrong because the model missed a coverage exclusion costs far more, weeks later in rework and in a customer's trust, than any amount of review time saved up front. Quentin Nazarenko, who runs product for ClaimBrief, watches that number harder than the raw minutes.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the equation, old time minus new time, times meetings that finish through the draft, not every meeting the tool touched.
Cost: the estimate has to ship today with no time for a real shadow study. Use the vendor's own logged timestamps as a floor, and say plainly it's a floor, not a finished number, until a real time study backs it up.
The model got better, for real: Gistly's accuracy jumps ten points overnight. That still doesn't excuse skipping review time in the equation. A better model just moves the completion rate up. It doesn't make the four minutes of checking disappear.
Where people run it wrong.
They measure "after" against how fast the model answers, not how long a person spends with what it gave them.
They count every meeting the tool touched as a win, even the ones somebody quietly rewrote by hand.
They report one confident number instead of a range, so the first hard question breaks it.
How to use it live. Say the equation out loud before a single number: old time minus new time, times how many actually finish this way. That buys you the room to ask what "finish" even means, instead of guessing at a headline number on the spot.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not count all eight hundred meetings, since Gistly touches every one?" Response: touching a meeting isn't the same as finishing it. The hundred seventy six meetings a month that get rewritten by hand saved nobody anything, and counting them as wins is exactly the kind of number that falls apart under a real audit.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?
- #7 Explain the problem with measuring acceptance rate of AI suggestions.