What metric captures the value of an AI feature that prevents work rather than performs it?
- Report violations avoided, not flags issued.Why: a flag count climbs the moment the flagging threshold drops, whether or not one extra real violation exists.
- Run a real holdback before reporting any prevention number to anyone outside the team.Why: no prevention number means anything until you know what ships when nothing stops it.
- Read the holdback slice in full, not by sample.Why: a small sample of a small slice is too noisy to defend under audit; a full read of a small slice is not.
- Check flag precision against a fixed eval set on a schedule.Why: an over-flagging model looks more valuable on a rising chart while actually costing agents more time on flags that turn out to be nothing.
- Convert the gap into a dollar figure using a real cost per shipped violation.Why: "violations avoided" only means something to an audit committee once it is in the same unit as everything else on their report.
- Keep flags issued as an internal engineering number only.Why: it still helps tune the model day to day, it should just never leave the building as proof of anything.
How to answer this, stage by stage
Nobody is grading whether you can name a metric. They are grading whether you can rank the candidate numbers by which one survives being checked later, not by which one looks best on a slide this quarter. Seven moves get you there.
Let's learn
Every week for three years, Northgate Assurance's compliance team pulled a random sample of a hundred and fifty agent emails, out of about six thousand sent to policyholders, and read them after they had already reached an inbox.
Calder is the tool Northgate built to change that. It reads all six thousand emails before an agent hits send, and it holds back the ones that break a rule: a quote with no "subject to underwriting" line, a cancellation notice missing the state's required wording, coverage described as approved when it is still pending.
In the old sample, compliance found on average three real violations a week inside those hundred and fifty emails, a rate too small a sample to trust as the true number across all six thousand. Once Calder went live, it started flagging about three hundred and forty emails a week, roughly one in eighteen, and holding each one for a person to fix or clear before it sent.
Here is the turn. Three hundred and forty flags a week sounds like a strong number, and for the first six months Northgate's leadership reported exactly that figure to the audit committee as proof Calder was working. But a flag count does not say what it replaced. Turn Calder's threshold down one notch and the count climbs, whether or not a single extra real violation existed. Nobody outside the team that built Calder could tell the difference from that number alone.
At its worst, this costs more than an awkward meeting. If leadership kept reporting flags issued as the value metric, two bad things could happen at once. The model team, chasing a rising number, quietly tunes for more flags, more false alarms, more agents overriding the tool out of habit. Meanwhile the real question, whether Calder actually keeps violations off a customer's screen, goes unanswered. Worst case, a state examiner asks Northgate's compliance officer to show that Calder reduced real exposure, and there is no defensible number in the building, only a flag count everyone privately admits could be moved by one setting.
The choice I would take back: when Calder first shipped, eighteen months earlier, the team picked "flags raised per week" as the launch dashboard's headline number, because it was the only figure the model produced without anyone doing extra work. No baseline needed, ships day one, easy to read on a slide. Nobody in that room was picturing a state examiner asking for the math behind it. Why would they. Calder was still a prototype then.
Once Soledad ran the real comparison, the gap between the two numbers was not small.
What I would leave alone: Calder's own before-send delay, two to three seconds while it scores an email, does not need this same rigor. Nobody reports that number outside the engineering team, it is just a normal speed target, not a claim made to a regulator.
The lesson: a metric for something that prevents work has to answer "compared to what." A count the model can move by itself can never answer that question honestly, no matter how good it looks climbing on a slide.
Now here is the same thing as a story
The short version is above. Read on for the Tuesday a rising number almost became the only thing anyone checked.
Soledad Delacourt has run product for Calder for a year and a half. She is good at the part of the job most product managers dodge: sitting with a compliance report and reading it line by line, the way an underwriter reads a claim, instead of skimming for the headline number.
When Calder launched, the number on the Monday product review slide only ever went one direction. Two hundred and ten flags the first month. Two hundred and eighty by month three. Three hundred and forty by month six. Everyone in that room liked watching it climb, Soledad included.
For the first two months she opened a random dozen flagged emails herself every Friday afternoon, just to feel out whether they were real violations or noise. By month four she was down to two or three. By month six she had stopped opening any of them. She watched the number on the slide instead, the way everyone else did.
The trigger was small. Cassia Nettleton, three weeks into the compliance team, sat in on a Monday review and asked one question nobody in the room had an answer for. "Three hundred forty flags. Compared to what?"
Soledad started to answer with the obvious thing, that flags had climbed from two hundred and ten to three hundred and forty, proof Calder caught more each month. Then she stopped mid sentence. A climbing flag count could mean Calder was finding more real violations. It could just as easily mean the model team had nudged the flagging threshold down after a complaint about missed cases the quarter before. She did not actually know which one it was. Neither did anyone else in the room.
Eighteen months earlier, in a half hour meeting when Calder first shipped, the team picked "flags raised per week" for the launch dashboard because it was the only number the model produced without anyone doing extra work. No baseline needed. Ships day one. Easy to read from the back of the room. Nobody there was picturing a state examiner asking to see the math behind it. Calder was still a prototype then, and this was just the fastest number to put on a chart.
What Soledad actually did was not rebuild the model. She ran a holdback. For three weeks, Calder scored a random eight percent slice of Northgate's email, about four hundred and eighty a week, but never showed the score to the agent and never held anything back. She asked Selwyn Sarto's compliance team to read every one of those emails by hand for the full three weeks, not a sample, a full read, the only way to get a real number for what ships when nothing stops it.
Three weeks later, the real numbers came in. Across the holdback slice, without Calder, 2.1 percent of email carried a real violation through to a policyholder's inbox, close to what the old sample had always implied but never confirmed at real volume. A follow-up audit of a thousand Calder-cleared emails from the protected 92 percent of traffic found a real violation rate of 0.3 percent slipping past review. The gap, 1.8 points, applied to Northgate's full six thousand emails a week, comes out to roughly a hundred and eight violations avoided every week, worth about forty three thousand dollars a week at Northgate's average cost per shipped violation, just over two million dollars a year.
The flag count was a number the model could move by itself. The violations-avoided number needed a person to sit with an unprotected slice of real customer email for three weeks and count. That is the whole difference between a metric and a guess with a chart under it.
What I would tell myself, watching that slide climb for six months straight: a number going up is not the same thing as a number meaning something, and I let the first one stand in for the second for a lot longer than I should have.
ORDER, run against a number instead of a decision
This is a question about which candidate metric to trust, not a story about a habit with two settings, so ORDER fits: rank the numbers by which one survives being checked later.
Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. The alternative most candidates reach for by reflex is trusting Calder's own confidence score as the value proxy, since the model already produces it and it needs no extra work to report. That got ruled out on purpose: a confidence score is the model grading itself, and the whole reason a measured number like this matters is that it comes from a check outside the model, not another number the model can quietly move. Second, the flag threshold itself is not a fixed rule that fires or does not. It is a cut-off on the model's own risk score, calibrated against a fixed golden set of six hundred labeled emails so it catches real violations most of the time, by design, not every time and not on a hand-written keyword list. That calibration is also where the real trade-off sits. Turn the threshold down and Calder catches more real violations, but it also holds up more clean emails for a person to clear, real minutes an agent spends waiting on a flag that turns out to be nothing. Northgate accepted a slower send on the highest-risk email types, quotes and cancellations, and left the threshold looser on routine, low-risk mail.
And if you want to be sure it really works, try it somewhere else
Same five checks, a different industry, so the method proves itself instead of repeating a story you happened to prepare.
Hollowell Apothecary runs Ossory, an AI tool that reads outgoing prescription emails, refill denials, formulary substitutions, prior-authorization notices, before they reach a patient, for about nine hundred thousand mail-order patients.
O, outcome. What every candidate metric competes to capture: real patient harm and appeal complaints kept off a pharmacist's desk, not how many emails Ossory flagged this week.
R, reversibility. "Denials flagged" can be inflated any time by loosening what counts as missing language, and nobody can tell after the fact whether that means more real gaps or a jumpier model. A harm rate measured on a logged holdback cannot be rebuilt once the window closes.
D, dependency. Florian Havens, who runs product for Ossory, first needed a real baseline: the rate of denial emails that go out missing a required appeal notice or plain-language summary when nothing stops them.
E, evidence. A two-week, five percent holdback, read in full by Hollowell's own compliance pharmacist instead of the usual spot check.
R, rank. Report harm avoided, converted into appeal complaints avoided, as the number that goes to Hollowell's board. Flags raised stays internal, a dial the model team watches, never the proof anyone shows outside the building.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule: reject flags issued first, then say the real number needs a measured baseline, not the other way round.
Cost: the compliance team says a three-week holdback is too expensive to run every quarter. Do not cave and report the flag count in the meantime. Run a shorter, smaller holdback instead, and say plainly that a rougher real number still beats a precise fake one.
The model got better, for real: say Calder's precision jumps ten points overnight. That still is not the same claim as "the flag count now means something." A better model can still be handed a threshold nobody remembers tuning.
Where people run it wrong.
They report the number the model produces for free, because it needs no extra work, instead of asking what that number is actually measuring.
They build a permanent measurement program off one holdback, then discover the real rate moved once a new state rule shipped and their baseline is a year stale.
They let the model's own confidence score stand in for an outside check, because it is the fastest number to put on a slide, not because it answers the question.
How to use it live. Say the rule before naming a single number: "I rank candidate metrics by which one survives being checked later, not by which one needs the least work to produce." That buys you the room to reject the easy answer out loud, instead of reciting "track the flag count" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just trust the model's own confidence score instead of running a live holdback?" Response: a confidence score is the model grading itself. Only a check that lives outside the model, a holdback read by a person, can catch the model quietly drifting.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #7 Explain the problem with measuring acceptance rate of AI suggestions.