ConceptAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #4

Describe three metrics that would tell you an AI feature is trusted rather than merely used.

The direct answer
Watch three things, none of which is how often the tool gets opened: whether attorneys quietly stop rereading the clauses it flags, whether its unedited redlines start shipping untouched on the matters with real money on the line, and whether people run it on jobs they could easily do by hand. All three can be faked. A falling recheck rate can just mean people are tired, not convinced. A rising unedited-redline share can mean the tool writes safe-sounding boilerplate nobody double-checks, not good advice. Reuse on easy tasks can be a paper trail people want for themselves, not confidence in the tool. So every one of the three only counts as trust once you test it against an outside judgment, not just against itself.
Do this, in order
  1. Track whether attorneys stop rereading the flagged clauses, not whether they keep opening the tool.Why: raw usage can be mandated or habitual. A falling recheck rate is the first place real trust would show up, weeks before anyone says so out loud.
  2. Track the unedited-redline share on high-stakes matters specifically, not the firm-wide average.Why: a rising unedited rate on routine NDAs proves nothing. The same rate on a nine-figure acquisition is either real trust or a quiet, expensive miss waiting to happen.
  3. Track voluntary reuse on jobs a person could do faster by hand.Why: someone who runs the tool when nothing makes them is telling you something a mandated click never can.
  4. Before crediting any of the three, check it against an outside judgment: a senior partner's sample review or a golden set.Why: every one of these three signals has an honest, non-trust explanation. Fatigue, boilerplate, and cover-your-self reuse all produce the same graph.
  5. Set the actual threshold and the actual reaction at each one, before the numbers move.Why: a metric nobody has agreed to act on is a chart on a wall, not a way to run the product.
  6. Never let the lagging number, hours billed per matter, be the first thing you check.Why: by the time hours per matter confirms trust, the leading signals have already been telling you for a month.

How to answer this, stage by stage

Nobody is grading whether you can name three metrics off the top of your head. They are grading whether you can catch yourself before you hand them a metric that just moves with a mandate. Seven moves get you there.

1
Scope it to one real product, one real reader
Say it like this
"Let's ground this in something real. Coverstone is a tool that reads a company's contracts, flags the clauses that carry real risk, and drafts a first-pass redline. Purnima Rathgar runs product for it at Ondine Analytics, and her main pilot account is Havenwright, a mid-size firm whose M&A group leans on it every week."
Why this works
Grounds the answer in a real product and a real reader before naming a single metric, so nothing that follows floats free of a person.
2
Say what "trusted" is not, before saying what it is
Say it like this
"First I'd throw out the obvious one. Query volume, sessions per week, that whole family. Havenwright's managing partner told every associate to run Coverstone on every new contract, for audit reasons. Volume went up the same week and it means nothing about whether anyone believes what it says."
Why this works
Naming and rejecting the obvious wrong answer first is what tells the interviewer you've actually thought about the failure mode, not just recited a metric.
3
Name the outcome all three metrics are chasing
Say it like this
"The thing that actually matters is durable reliance without a miss reaching a client. Fewer partner-review hours per matter, held over two or three renewal cycles, with nothing dangerous slipping through. That's the L, the link. Everything else is either a leading signal for it or a way it gets faked."
Why this works
This is the L step. Without a named outcome, the three metrics are just three numbers, not evidence of anything.
4
Give the three early signals, with the real numbers behind each
Say it like this
"One: the share of flagged clauses a second attorney rereads from scratch. It ran near thirty percent for Havenwright's first two months and fell under six by month four. Two: on matters above fifty million dollars, the share of Coverstone's redlines that ship with zero edits. That was almost zero in month one and crossed a third by month five. Three: how often an associate runs it on a contract short enough to just read by hand, no partner told them to."
Why this works
This is the E step, the actual answer to the question. Real numbers with a timeframe make it checkable instead of asserted.
5
Say how each one gets gamed or misread
Say it like this
"Here's the part most people skip. The falling recheck rate could just mean the associates are burned out at month four of a deal, not convinced. The rising unedited-redline share could mean Coverstone writes safe, generic language that looks fine and nobody catches what's missing. And the voluntary reuse could be an associate covering herself, wanting a timestamped run on file, not trusting a word of what it says."
Why this works
This is the A step, the hard one in LEAD. A metric interview answer that skips how the metric gets gamed hasn't really understood the metric.
6
Say the actual decision at each threshold
Say it like this
"If the recheck rate falls without billable hours also easing off, I treat that as fatigue, not trust, and I put a mandatory spot-check back in for a sample. If the unedited-redline share on the fifty-million-dollar matters crosses a third, I pull those redlines and score them against senior partner judgment before I believe it. If it clears ninety percent agreement, that's real. If it doesn't, I've caught silent drift before a client ever sees it."
Why this works
This is the D step. A metric with no stated action at a threshold is decoration, not a way to run the product.
7
Close on the three, defended in one line
Say it like this
"So: falling recheck rate, rising unedited high-stakes redlines, and voluntary reuse on easy jobs, every one of them checked against an outside read before I believe it. None of the three is usage. Usage is the thing a mandate can fake in a week."
Why this works
Closes on the actual three-metric answer the question asked for, in a shape a reader can repeat cold.
If you remember one thing A number that climbs because someone was told to click it is not trust. Trust is the number that climbs on its own, and even then, check it against a judgment that isn't the tool's own.

Let's learn

Say we build a tool that reads a company's contracts and flags the clauses worth worrying about: an indemnity that runs too wide, a termination right buried in a side letter, a liability cap that quietly favors the other side. Coverstone does this for Ondine Analytics' law firm customers, and drafts a first-pass redline for each flagged clause too. Before a tool like this, a mid-level associate at a firm like Havenwright read every page of a contract by hand, maybe eighty pages a day across a busy week, flagging risk from memory and training.

With Coverstone running, that same associate can move through three or four times as many pages, because the tool clears the routine language on its own and only slows her down on the clauses that actually carry risk.

Knowledge spark: why is usage such a bad stand-in for trust here? A law firm can mandate usage in an afternoon. Havenwright's managing partner can tell every associate to run every contract through Coverstone, for the audit trail alone, and the query count climbs the same week whether or not a single person believes what the tool says.

Four months into the pilot, Purnima pulled the numbers she actually cared about. The share of flagged clauses that a second attorney reread from scratch, once close to thirty in every hundred, had fallen under six. On the firm's largest matters, deals above fifty million dollars, the share of Coverstone's redlines that shipped with no edits at all had climbed past a third, up from almost nothing in month one.

Recheck rate on flagged clauses, week 1 to week 18
30% 15% 0% wk 1 wk 9 wk 18 managing partner mandates full usage 28% 6%
The recheck rate was already sliding for nine weeks before the mandate ever landed. If Purnima had waited for the usage graph to tell her something, she would have missed the real signal entirely.

Here is the turn. Neither of those two falling and rising lines was, by itself, proof of trust. The falling recheck rate could mean the associates believed Coverstone. It could just as easily mean they were exhausted three months into a live deal and stopped double-checking out of fatigue, not confidence. The rising unedited-redline share on the largest matters could mean the tool's drafts were genuinely good. It could also mean Coverstone had learned to write safe, lawyerly-sounding boilerplate that nobody caught, because it read like something a partner would write, even where it missed the deal-specific risk entirely.

A number that looks exactly like trust and a number that looks exactly like exhaustion can be the same line on the same chart.

At its worst, this costs far more than fifteen extra hours of partner review. Say Purnima had simply celebrated the falling recheck rate and the rising unedited share as proof the product worked, and Ondine's sales team started quoting both numbers to prospective firms. If even one of those unedited, high-stakes redlines had missed a real liability trap on a fifty-million-dollar deal, the cost lands on Havenwright's client, not on a dashboard, and it lands long after anyone thought to check.

Hand sketched comparison diagram titled The reuse number can lie. Left panel, a document icon labelled Looks like trust, caption she reruns Coverstone on a simple NDA nobody made her check. Right panel, a person icon labelled Might just be cover, caption the timestamped run protects her if a clause turns out wrong later.
The same voluntary rerun on an easy contract can mean two opposite things. One is real confidence. The other is an associate building herself a paper trail in case a partner asks later why she missed something.

The choice I would take back is not which three metrics to watch. It is that Coverstone's dashboard, built in the first sprint, only ever logged how many contracts ran through the tool. Nobody built a way to log whether a flagged clause got a second, independent read, so for the pilot's first two months, the one number that would have shown trust building was simply not being kept anywhere.

What I would leave alone: the raw usage count still has a real job, just not this one. It tells Ondine's support team when a firm has gone quiet and might be about to churn. That's a fine use for it. It was only ever the wrong number for the question "do they trust it."

The lesson: usage is the easiest number to move and the least honest one to trust. The real signal is always upstream of it, in a habit nobody thought to log until someone went looking for it on purpose.

Now here is the same thing as a story

The short version is above. Read on for how ordinary the week looked from Purnima's side when the mandate almost fooled her.

Purnima Rathgar has spent three years building products for law firms who do not want to hear the word "AI" anywhere near a client-facing document. She is good at the part of the job most product managers avoid: sitting with a number that looks great and asking what else, besides the thing she wants it to mean, could have produced it.

For the first two months of the Havenwright pilot, the numbers Purnima watched were the easy ones. Sessions per week, contracts run, minutes saved per matter. All three climbed steadily, and every one of them made a nice slide.

Then, in the ninth week, Havenwright's managing partner sent a firm-wide memo: every new matter, every contract, gets run through Coverstone, no exceptions, for the audit trail. Usage jumped twenty percent that same week. Purnima's first instinct was to be pleased.

What stopped her was a smaller number she had started tracking almost as an afterthought, buried in a spreadsheet nobody else at Havenwright had asked to see: how often a second attorney reread a flagged clause from scratch before signing off on it. That number had already been sliding, slowly, for weeks before the mandate ever landed.

We did not almost mistake a mandate for trust. We almost mistook it for the wrong nine weeks.

Because the recheck rate did not fall the week usage jumped. It had been falling since week one, quietly, while the usage graph sat flat. By the time the managing partner's memo went out, real associates had already, on their own, mostly stopped rereading what Coverstone flagged. The mandate did not create the trust. It just made the usage graph finally catch up to something that had been true for two months.

Aldric Thane, the senior associate running most of Havenwright's M&A pilot work, put it to her plainly on a call that same week: "I stopped double-checking the small stuff around week three. I just never told anyone, because nobody asked." That was the near miss. If Purnima had waited for a dashboard number to move before believing anything, she would have learned what Aldric already knew a full month and a half after the fact.

Unedited redlines on matters above $50M, before and after month four's audit
35% 18% 0% Claimed, unaudited After senior review 34% 19%
Purnima had two senior partners independently reread a random sample of the "unedited" redlines against the golden set. Fifteen points of that thirty-four percent turned out to be safe boilerplate a rushed associate had waved through, not language a partner would actually have signed off on.

What Purnima did that week: she did not touch the mandate, and she did not touch Coverstone's model. She pulled forty of the unedited, high-stakes redlines and asked two senior partners, who had never seen the tool's output before, to score each one blind against Ondine's own golden set of real risky clauses. Nineteen in every hundred cleared as genuinely sound. The rest were safe-sounding language that a tired associate, three months into a live deal, had rubber-stamped without really reading.

Two years earlier, when Coverstone's dashboard first shipped, the engineering team built exactly one number into it: contracts processed per week. It was the fastest thing to log and the easiest thing to put on a slide for Ondine's own investors. Nobody in that early sprint asked what a second, independent read would look like, because in the earliest pilots, firms were small enough that Purnima could just ask people directly.

What she built after that Havenwright week: a real recheck-rate metric, logged automatically every time a second attorney opened a flagged clause Coverstone had already scored, and a standing quarterly sample where senior partners blind-score a batch of unedited high-stakes redlines against the golden set. The replay, one quarter later: recheck rate held near six percent, and the golden-set agreement on unedited high-stakes redlines climbed from nineteen to eighty-eight, once Ondine tightened the model's prompt to flag its own uncertainty on non-boilerplate language instead of writing confidently over it.

What I would tell myself, back on that call with Aldric: the usage graph was never lying. It just was never going to be early enough to matter.

LEAD, the four questions that separate trust from a mandate

This is a metric question, what would show trust building before the lagging outcome confirms it, so LEAD fits, not a story about a single habit with two settings.

L
Link. The business outcome all three metrics exist to predict.
Not sessions per week. Durable reliance without a miss reaching a client, held across two or three renewal cycles, seen in fewer partner-review hours per matter over time.
E
Early signal. The three things that move before the outcome does.
Falling recheck rate on flagged clauses. Rising unedited-redline share on matters above fifty million dollars. Voluntary reuse on contracts short enough to read by hand.
A
Abuse. How each one gets gamed or misread.
A falling recheck rate can be fatigue, not trust. A rising unedited-redline share can be safe boilerplate nobody catches. Voluntary reuse can be an associate building herself a record, not confidence in the tool.
D
Decision. What actually changes at each threshold.
Recheck rate falling without billable hours easing: reinstate a mandatory spot-check sample. Unedited share crossing a third on high-stakes matters: blind-score a sample against the golden set before believing it. Below ninety percent agreement, treat the rise as drift, not a win.

Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. First, the alternative most candidates reach for is a satisfaction survey, asking associates to rate how much they trust Coverstone out of ten. That got ruled out on purpose. A survey score moves with how convenient a tool feels, and Coverstone can feel wonderfully convenient, fast, well formatted, easy to read, while an associate still quietly reruns everything it says by hand, or worse, has simply stopped noticing when it writes something plausible and wrong. Second, the real bar for believing any of the three metrics is not "the number moved in the right direction." It is calibrated: the unedited-redline claim only counts once a blind sample clears roughly ninety percent agreement with senior partner judgment against the golden set, not one hundred, because Ondine is not claiming the model never misses, only that it misses rarely enough, on the matters that matter, to be worth the trust it's earning.

Knowledge spark: what is a golden set, in this context? A stack of real, past contract clauses that Ondine's senior partners have already scored by hand for risk. Every new version of Coverstone's model or prompt gets run against that same stack before it ships, so a change that quietly gets worse at catching a real risk gets caught before a live client contract ever sees it.

The quality and cost trade-off sits right underneath the D step too. Running a second, independent pass, asking the model to explain its own reasoning and flag where it is least sure before a redline ships unedited, costs real tokens and adds a few seconds to every high-stakes matter. That is worth paying on a fifty-million-dollar acquisition. It is not worth paying on a routine two-page NDA, where a missed nuance costs an afternoon, not a client.

And if you want to be sure it really works, try it somewhere else

Same four letters, a different product, a different industry, so the method proves itself instead of repeating a story you happened to prepare.

Talvera Mutual runs Castellan, an AI tool that reads incoming crop insurance claims and flags which ones carry a real chance of fraud or misvaluation, for a team of field adjusters who used to inspect every claim by hand.

L, link. Not claims processed per week. Payout accuracy held steady while adjuster time per claim keeps falling, across at least two full growing seasons, with no fraud case surfacing later that Castellan should have caught.
E, early signal. The share of Castellan-cleared claims an adjuster still drives out to inspect in person, even though nothing requires it. That fell from near forty percent in month one to under ten by month five at Talvera's largest regional office.
A, abuse. A falling in-person inspection rate can mean real trust in the model's read. It can also mean an adjuster with a backlog of two hundred claims simply stopped having time to drive anywhere, which looks identical on a chart.
D, decision. Moreau, who leads claims operations at Talvera, cross-checks the falling inspection rate against each adjuster's open caseload. If the caseload is flat and inspections still fall, that's trust. If the caseload is climbing at the same time, that's burnout wearing a trust number's clothes, and the fix is more staffing, not a victory lap.

Same shape, different stakes At Havenwright, the unlogged risk was a liability clause slipping past a rushed sign-off. At Talvera, it is a fraudulent claim clearing because an exhausted adjuster stopped driving out to look. The check does not change: never credit a leading signal as trust until you have ruled out exhaustion as the reason it moved.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the three signals and name the one thing usage can never tell you, that a mandate moves it in a week and trust does not.
Cost: Ondine's engineering team says the quarterly blind-scoring sample is too expensive to run every quarter. Don't quietly drop it to save money. Shrink the sample size before you drop the check itself, because an unchecked trust metric is worse than no metric at all.
The model got better, for real: say Coverstone's redline accuracy jumps fifteen points overnight. That still does not make a rising unedited-redline share proof of trust on its own. A better model can still be handed to an exhausted associate who was never really checking.

Where people run it wrong.
They report usage as the headline trust metric because it is the easiest number already sitting in the dashboard.
They see a leading signal move the right direction and call it proven, without ever checking it against an outside judgment.
They wait for the lagging outcome, hours billed per matter or claims processed per adjuster, to confirm trust, and by then the real story is a month or two stale.

How to use it live. Open by naming the wrong answer before the right one: "The obvious metric here is usage, and it's the wrong one, because a mandate can move it in a week." That single sentence buys you the room to give the real three-metric answer next, instead of reciting a features list on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits naming the metrics that show trust building before a lagging outcome confirms it, and why?
Tap to flip
ANSWER
LEAD. It is a metric question, find the signal that moves first, not a story about a single habit with two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Purnima Rathgar, who runs product for Coverstone at Ondine Analytics, working with Havenwright LLP's M&A group and senior associate Aldric Thane.
3 · THE WRONG METRIC
What metric does this answer reject as the north star, and why?
Tap to flip
ANSWER
Raw usage, sessions or contracts processed per week. A managing partner's firm-wide mandate moved it twenty percent in a single week, with no change in whether anyone actually believed the output.
4 · THE THREE SIGNALS
Name the three leading signals this answer actually gives.
Tap to flip
ANSWER
Falling recheck rate on flagged clauses, rising unedited-redline share on matters above fifty million dollars, and voluntary reuse on contracts short enough to read by hand.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building Coverstone's dashboard to log only contracts processed per week. It made sense in the earliest pilots, when firms were small enough that Purnima could just ask people directly whether they trusted it.
6 · THE NUMBER
Fill in the blank: the recheck rate on flagged clauses fell from about ___ percent to about ___ percent by week eighteen.
Tap to flip
ANSWER
28 percent down to 6 percent, and the fall started nine weeks before the managing partner's usage mandate ever landed.
7 · THE REPLAY
Same bad week, new dashboard. What changes?
Tap to flip
ANSWER
A real recheck-rate metric and a standing quarterly blind-score sample against the golden set. One quarter later, golden-set agreement on unedited high-stakes redlines climbed from 19 to 88 percent.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the early signal?
Tap to flip
ANSWER
Castellan, Talvera Mutual's crop-claim fraud tool. Its early signal is the falling share of cleared claims an adjuster still drives out to inspect in person, unprompted.

Check yourself Score: 0 / 0

True or false
1. True or false: a rising query volume, once a managing partner mandates that every associate use the tool, is strong evidence that trust is building.
  • True
  • False
Show hint
Check stage 2 of the walkthrough.
Show answer
False. A mandate can move usage in a week regardless of trust. It went up the same week the recheck rate had already been sliding, quietly, for nine weeks.
Multiple choice
2. Why doesn't a rising unedited-redline share on its own prove trust?
  • A. Because attorneys are legally required to edit every AI-drafted redline.
  • B. Because unedited redlines are always a sign of a bug in the model.
  • C. Because it can also mean the model writes safe, generic language that looks fine but nobody actually checked.
  • D. Because redlines are never shown to senior partners at all.
Show hint
Look at what the blind-scoring sample against the golden set actually found.
Show answer
C. Nineteen in every hundred "unedited" redlines cleared as genuinely sound against the golden set. The rest were boilerplate a rushed associate waved through without really reading it.
Fill in the blank
3. When Purnima had two senior partners blind-score a sample of unedited redlines against the golden set, only ___ in every hundred cleared as genuinely sound.
Show hint
Check the chart in "Now here is the same thing as a story."
Show answer
19. That gap, between a 34 percent unedited-redline claim and 19 percent real agreement, is exactly why the raw metric needed an outside check before anyone believed it.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for what Coverstone's dashboard logged in its very first sprint.
Show answer
Model answer: Building the dashboard to log only contracts processed per week, with no way to see whether a flagged clause got a real second read. It made sense in the earliest pilots, when firms were small enough that Purnima could just ask people directly whether they trusted the tool.
Short answer, apply it yourself
5. Think of an AI tool you use yourself, a coding assistant, a writing tool, anything. What would a rising "I didn't check that" behavior look like for that tool, and how would you tell if it meant real trust instead of fatigue?
Show hint
Look for a habit you could log automatically, and a reason that habit could change without trust changing.
Show answer
Model answer: For a coding assistant, it might be how often someone runs the suggested code without reading the diff first. That could mean real trust in its output. It could also mean the person is behind on a deadline and skipping review out of time pressure, which you would only tell apart by checking whether their overall workload had also spiked.
Short answer, the number question
6. If the recheck rate had fallen from 28 percent to 6 percent in the same three weeks the mandate landed, instead of nine weeks earlier, would that change which explanation you'd reach for first?
Show hint
Think about what the timing, not just the size, of the drop actually tells you.
Show answer
Yes, it would. A recheck rate that falls in lockstep with a usage mandate is much more likely to be compliance, associates doing the minimum required, than a rate that was already falling on its own for weeks beforehand. Timing relative to the mandate is itself part of the evidence, not a detail to skip past.
Before you close the answer
Why this works
Tests whether you will hand an interviewer the first metric that comes to mind, usage, or actually interrogate whether a number can be moved by something other than trust before you believe it.
Follow-up traps
"Couldn't you just ask associates directly if they trust it?" Response: you can, and it's worth doing, but a self-report moves with how convenient a tool feels, not with whether someone is quietly rerunning everything it says by hand. Behavior is harder to fake than a survey answer.

"What if the recheck rate falls because the model actually got worse and people stopped noticing?" Response: that's exactly why the recheck rate alone never gets credited as trust. It only counts once it's checked against an outside judgment, like the golden-set blind score, that isn't the tool grading itself.
If pressed
The quarterly blind-scoring sample at Ondine isn't random. It's stratified by matter size and clause type, so a rare, high-severity clause type isn't drowned out by a flood of routine boilerplate the way a simple random sample would understate it. That's what let Purnima catch the boilerplate problem in month four instead of a full year later.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more