Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- Track whether attorneys stop rereading the flagged clauses, not whether they keep opening the tool.Why: raw usage can be mandated or habitual. A falling recheck rate is the first place real trust would show up, weeks before anyone says so out loud.
- Track the unedited-redline share on high-stakes matters specifically, not the firm-wide average.Why: a rising unedited rate on routine NDAs proves nothing. The same rate on a nine-figure acquisition is either real trust or a quiet, expensive miss waiting to happen.
- Track voluntary reuse on jobs a person could do faster by hand.Why: someone who runs the tool when nothing makes them is telling you something a mandated click never can.
- Before crediting any of the three, check it against an outside judgment: a senior partner's sample review or a golden set.Why: every one of these three signals has an honest, non-trust explanation. Fatigue, boilerplate, and cover-your-self reuse all produce the same graph.
- Set the actual threshold and the actual reaction at each one, before the numbers move.Why: a metric nobody has agreed to act on is a chart on a wall, not a way to run the product.
- Never let the lagging number, hours billed per matter, be the first thing you check.Why: by the time hours per matter confirms trust, the leading signals have already been telling you for a month.
How to answer this, stage by stage
Nobody is grading whether you can name three metrics off the top of your head. They are grading whether you can catch yourself before you hand them a metric that just moves with a mandate. Seven moves get you there.
Let's learn
Say we build a tool that reads a company's contracts and flags the clauses worth worrying about: an indemnity that runs too wide, a termination right buried in a side letter, a liability cap that quietly favors the other side. Coverstone does this for Ondine Analytics' law firm customers, and drafts a first-pass redline for each flagged clause too. Before a tool like this, a mid-level associate at a firm like Havenwright read every page of a contract by hand, maybe eighty pages a day across a busy week, flagging risk from memory and training.
With Coverstone running, that same associate can move through three or four times as many pages, because the tool clears the routine language on its own and only slows her down on the clauses that actually carry risk.
Four months into the pilot, Purnima pulled the numbers she actually cared about. The share of flagged clauses that a second attorney reread from scratch, once close to thirty in every hundred, had fallen under six. On the firm's largest matters, deals above fifty million dollars, the share of Coverstone's redlines that shipped with no edits at all had climbed past a third, up from almost nothing in month one.
Here is the turn. Neither of those two falling and rising lines was, by itself, proof of trust. The falling recheck rate could mean the associates believed Coverstone. It could just as easily mean they were exhausted three months into a live deal and stopped double-checking out of fatigue, not confidence. The rising unedited-redline share on the largest matters could mean the tool's drafts were genuinely good. It could also mean Coverstone had learned to write safe, lawyerly-sounding boilerplate that nobody caught, because it read like something a partner would write, even where it missed the deal-specific risk entirely.
At its worst, this costs far more than fifteen extra hours of partner review. Say Purnima had simply celebrated the falling recheck rate and the rising unedited share as proof the product worked, and Ondine's sales team started quoting both numbers to prospective firms. If even one of those unedited, high-stakes redlines had missed a real liability trap on a fifty-million-dollar deal, the cost lands on Havenwright's client, not on a dashboard, and it lands long after anyone thought to check.
The choice I would take back is not which three metrics to watch. It is that Coverstone's dashboard, built in the first sprint, only ever logged how many contracts ran through the tool. Nobody built a way to log whether a flagged clause got a second, independent read, so for the pilot's first two months, the one number that would have shown trust building was simply not being kept anywhere.
What I would leave alone: the raw usage count still has a real job, just not this one. It tells Ondine's support team when a firm has gone quiet and might be about to churn. That's a fine use for it. It was only ever the wrong number for the question "do they trust it."
The lesson: usage is the easiest number to move and the least honest one to trust. The real signal is always upstream of it, in a habit nobody thought to log until someone went looking for it on purpose.
Now here is the same thing as a story
The short version is above. Read on for how ordinary the week looked from Purnima's side when the mandate almost fooled her.
Purnima Rathgar has spent three years building products for law firms who do not want to hear the word "AI" anywhere near a client-facing document. She is good at the part of the job most product managers avoid: sitting with a number that looks great and asking what else, besides the thing she wants it to mean, could have produced it.
For the first two months of the Havenwright pilot, the numbers Purnima watched were the easy ones. Sessions per week, contracts run, minutes saved per matter. All three climbed steadily, and every one of them made a nice slide.
Then, in the ninth week, Havenwright's managing partner sent a firm-wide memo: every new matter, every contract, gets run through Coverstone, no exceptions, for the audit trail. Usage jumped twenty percent that same week. Purnima's first instinct was to be pleased.
What stopped her was a smaller number she had started tracking almost as an afterthought, buried in a spreadsheet nobody else at Havenwright had asked to see: how often a second attorney reread a flagged clause from scratch before signing off on it. That number had already been sliding, slowly, for weeks before the mandate ever landed.
Because the recheck rate did not fall the week usage jumped. It had been falling since week one, quietly, while the usage graph sat flat. By the time the managing partner's memo went out, real associates had already, on their own, mostly stopped rereading what Coverstone flagged. The mandate did not create the trust. It just made the usage graph finally catch up to something that had been true for two months.
Aldric Thane, the senior associate running most of Havenwright's M&A pilot work, put it to her plainly on a call that same week: "I stopped double-checking the small stuff around week three. I just never told anyone, because nobody asked." That was the near miss. If Purnima had waited for a dashboard number to move before believing anything, she would have learned what Aldric already knew a full month and a half after the fact.
What Purnima did that week: she did not touch the mandate, and she did not touch Coverstone's model. She pulled forty of the unedited, high-stakes redlines and asked two senior partners, who had never seen the tool's output before, to score each one blind against Ondine's own golden set of real risky clauses. Nineteen in every hundred cleared as genuinely sound. The rest were safe-sounding language that a tired associate, three months into a live deal, had rubber-stamped without really reading.
Two years earlier, when Coverstone's dashboard first shipped, the engineering team built exactly one number into it: contracts processed per week. It was the fastest thing to log and the easiest thing to put on a slide for Ondine's own investors. Nobody in that early sprint asked what a second, independent read would look like, because in the earliest pilots, firms were small enough that Purnima could just ask people directly.
What she built after that Havenwright week: a real recheck-rate metric, logged automatically every time a second attorney opened a flagged clause Coverstone had already scored, and a standing quarterly sample where senior partners blind-score a batch of unedited high-stakes redlines against the golden set. The replay, one quarter later: recheck rate held near six percent, and the golden-set agreement on unedited high-stakes redlines climbed from nineteen to eighty-eight, once Ondine tightened the model's prompt to flag its own uncertainty on non-boilerplate language instead of writing confidently over it.
What I would tell myself, back on that call with Aldric: the usage graph was never lying. It just was never going to be early enough to matter.
LEAD, the four questions that separate trust from a mandate
This is a metric question, what would show trust building before the lagging outcome confirms it, so LEAD fits, not a story about a single habit with two settings.
Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. First, the alternative most candidates reach for is a satisfaction survey, asking associates to rate how much they trust Coverstone out of ten. That got ruled out on purpose. A survey score moves with how convenient a tool feels, and Coverstone can feel wonderfully convenient, fast, well formatted, easy to read, while an associate still quietly reruns everything it says by hand, or worse, has simply stopped noticing when it writes something plausible and wrong. Second, the real bar for believing any of the three metrics is not "the number moved in the right direction." It is calibrated: the unedited-redline claim only counts once a blind sample clears roughly ninety percent agreement with senior partner judgment against the golden set, not one hundred, because Ondine is not claiming the model never misses, only that it misses rarely enough, on the matters that matter, to be worth the trust it's earning.
The quality and cost trade-off sits right underneath the D step too. Running a second, independent pass, asking the model to explain its own reasoning and flag where it is least sure before a redline ships unedited, costs real tokens and adds a few seconds to every high-stakes matter. That is worth paying on a fifty-million-dollar acquisition. It is not worth paying on a routine two-page NDA, where a missed nuance costs an afternoon, not a client.
And if you want to be sure it really works, try it somewhere else
Same four letters, a different product, a different industry, so the method proves itself instead of repeating a story you happened to prepare.
Talvera Mutual runs Castellan, an AI tool that reads incoming crop insurance claims and flags which ones carry a real chance of fraud or misvaluation, for a team of field adjusters who used to inspect every claim by hand.
L, link. Not claims processed per week. Payout accuracy held steady while adjuster time per claim keeps falling, across at least two full growing seasons, with no fraud case surfacing later that Castellan should have caught.
E, early signal. The share of Castellan-cleared claims an adjuster still drives out to inspect in person, even though nothing requires it. That fell from near forty percent in month one to under ten by month five at Talvera's largest regional office.
A, abuse. A falling in-person inspection rate can mean real trust in the model's read. It can also mean an adjuster with a backlog of two hundred claims simply stopped having time to drive anywhere, which looks identical on a chart.
D, decision. Moreau, who leads claims operations at Talvera, cross-checks the falling inspection rate against each adjuster's open caseload. If the caseload is flat and inspections still fall, that's trust. If the caseload is climbing at the same time, that's burnout wearing a trust number's clothes, and the fix is more staffing, not a victory lap.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the three signals and name the one thing usage can never tell you, that a mandate moves it in a week and trust does not.
Cost: Ondine's engineering team says the quarterly blind-scoring sample is too expensive to run every quarter. Don't quietly drop it to save money. Shrink the sample size before you drop the check itself, because an unchecked trust metric is worse than no metric at all.
The model got better, for real: say Coverstone's redline accuracy jumps fifteen points overnight. That still does not make a rising unedited-redline share proof of trust on its own. A better model can still be handed to an exhausted associate who was never really checking.
Where people run it wrong.
They report usage as the headline trust metric because it is the easiest number already sitting in the dashboard.
They see a leading signal move the right direction and call it proven, without ever checking it against an outside judgment.
They wait for the lagging outcome, hours billed per matter or claims processed per adjuster, to confirm trust, and by then the real story is a month or two stale.
How to use it live. Open by naming the wrong answer before the right one: "The obvious metric here is usage, and it's the wrong one, because a mandate can move it in a week." That single sentence buys you the room to give the real three-metric answer next, instead of reciting a features list on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the recheck rate falls because the model actually got worse and people stopped noticing?" Response: that's exactly why the recheck rate alone never gets credited as trust. It only counts once it's checked against an outside judgment, like the golden-set blind score, that isn't the tool grading itself.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?
- #7 Explain the problem with measuring acceptance rate of AI suggestions.