ConceptAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #6

What metric captures the value of an AI feature that prevents work rather than performs it?

The direct answer
Report violations avoided, not flags issued. Get that number by holding the tool back from a small, logged slice of traffic long enough to measure the real rate of violations that ship without it, then compare that rate against the traffic the tool protects. A flag count cannot do this. It moves the moment someone nudges a threshold, with no way for anyone outside the model team to tell why.
Do this, in order
  1. Report violations avoided, not flags issued.Why: a flag count climbs the moment the flagging threshold drops, whether or not one extra real violation exists.
  2. Run a real holdback before reporting any prevention number to anyone outside the team.Why: no prevention number means anything until you know what ships when nothing stops it.
  3. Read the holdback slice in full, not by sample.Why: a small sample of a small slice is too noisy to defend under audit; a full read of a small slice is not.
  4. Check flag precision against a fixed eval set on a schedule.Why: an over-flagging model looks more valuable on a rising chart while actually costing agents more time on flags that turn out to be nothing.
  5. Convert the gap into a dollar figure using a real cost per shipped violation.Why: "violations avoided" only means something to an audit committee once it is in the same unit as everything else on their report.
  6. Keep flags issued as an internal engineering number only.Why: it still helps tune the model day to day, it should just never leave the building as proof of anything.

How to answer this, stage by stage

Nobody is grading whether you can name a metric. They are grading whether you can rank the candidate numbers by which one survives being checked later, not by which one looks best on a slide this quarter. Seven moves get you there.

1
Ground it in one product, one number
Say it like this
"Let's make this real. Calder is Northgate Assurance's tool. Every week their four hundred agents send about six thousand emails to policyholders, and Calder reads each one before it goes out, holding back the ones that break a disclosure rule. Soledad Delacourt runs product for it."
Why this works
Grounds a metric question in a real product and a real volume before touching a single candidate number.
2
Reject the easy number, out loud, before defending anything
Say it like this
"The obvious answer is 'how many emails did Calder flag this week.' I would throw that number out before I ever reported it to anyone outside the team."
Why this works
Naming and rejecting the obvious metric first shows judgment, not just recall of a definition.
3
Name what a flag count actually rewards
Say it like this
"Turn Calder's flagging threshold down one notch and the flag count climbs. Nobody outside the model team can tell if that means more real violations or just a jumpier model. A count you can move by turning one dial isn't a metric, it's a vanity number."
Why this works
Names the AI-specific failure mode, a model whose own threshold can quietly inflate the very number meant to prove it works.
4
Say what every candidate number is actually trying to capture
Say it like this
"Every candidate metric on the table is trying to answer one question: how much real regulatory exposure did we keep off a customer's screen. Not how many times the model raised its hand."
Why this works
This is the O step. Without a named outcome, ranking candidate metrics is just picking whichever one sounds biggest.
5
Say what a real prevention number depends on
Say it like this
"You can't say 'we avoided a violation' unless you know what would have shipped without Calder. That number doesn't exist yet. It has to be measured, not assumed."
Why this works
This is the D step, the dependency every other candidate metric quietly needs and usually skips.
6
Spend the cheap evidence before the expensive metric
Say it like this
"So for three weeks we held Calder back from eight percent of traffic, chosen at random, still scored but never shown to the agent. Compliance read every single one of those by hand, not a spot check. That gave us a real number: how many violations ship when nothing stops them."
Why this works
This is the E step, a real baseline paid for on purpose, in days, not a guess dressed up as data.
7
Land the ranked call, in a number a regulator would accept
Say it like this
"So the number I would actually report is violations avoided, built from that baseline gap, applied to full volume, turned into dollars using our average cost per shipped violation. Flags issued stays on an internal engineering dashboard. It never goes near the audit committee again."
Why this works
Closes on the ranked, defended call, the exact thing a reader can repeat under follow-up.
If you remember one thing A metric for something that prevents work has to answer "compared to what." A number the model can move on its own, by turning one dial, can never answer that question honestly.

Let's learn

Every week for three years, Northgate Assurance's compliance team pulled a random sample of a hundred and fifty agent emails, out of about six thousand sent to policyholders, and read them after they had already reached an inbox.

Calder is the tool Northgate built to change that. It reads all six thousand emails before an agent hits send, and it holds back the ones that break a rule: a quote with no "subject to underwriting" line, a cancellation notice missing the state's required wording, coverage described as approved when it is still pending.

Knowledge spark: what is shadow mode? A setting where Calder still reads and scores every email, but the score is only logged, never shown to the agent and never used to hold anything back. It lets you measure what happens without the tool, on real traffic, without actually turning the tool off.

In the old sample, compliance found on average three real violations a week inside those hundred and fifty emails, a rate too small a sample to trust as the true number across all six thousand. Once Calder went live, it started flagging about three hundred and forty emails a week, roughly one in eighteen, and holding each one for a person to fix or clear before it sent.

Here is the turn. Three hundred and forty flags a week sounds like a strong number, and for the first six months Northgate's leadership reported exactly that figure to the audit committee as proof Calder was working. But a flag count does not say what it replaced. Turn Calder's threshold down one notch and the count climbs, whether or not a single extra real violation existed. Nobody outside the team that built Calder could tell the difference from that number alone.

At its worst, this costs more than an awkward meeting. If leadership kept reporting flags issued as the value metric, two bad things could happen at once. The model team, chasing a rising number, quietly tunes for more flags, more false alarms, more agents overriding the tool out of habit. Meanwhile the real question, whether Calder actually keeps violations off a customer's screen, goes unanswered. Worst case, a state examiner asks Northgate's compliance officer to show that Calder reduced real exposure, and there is no defensible number in the building, only a flag count everyone privately admits could be moved by one setting.

The choice I would take back: when Calder first shipped, eighteen months earlier, the team picked "flags raised per week" as the launch dashboard's headline number, because it was the only figure the model produced without anyone doing extra work. No baseline needed, ships day one, easy to read on a slide. Nobody in that room was picturing a state examiner asking for the math behind it. Why would they. Calder was still a prototype then.

Once Soledad ran the real comparison, the gap between the two numbers was not small.

Real violation rate: held back versus protected by Calder
2.5% 1.25% 0% Held back, no Calder Calder protected 2.1% 0.3%
Three weeks, full manual read of both slices: 2.1 percent of email carried a real violation through when Calder was held back, against 0.3 percent on the traffic Calder protects. Applied to all six thousand emails a week, that 1.8-point gap is about a hundred and eight violations avoided a week, close to a third of the three hundred and forty flags Calder raises, and the number Northgate now reports.
Hand sketched flow diagram titled What has to happen before the number means anything. Four boxes connected by arrows in order. Hold back a slice. Log the real rate, shown highlighted in blue. Compare the groups. Convert to dollars.
Nothing after the second box means anything until the real rate gets logged on a slice nobody protected. Comparing the groups and turning the gap into dollars both sit downstream of that one measurement.

What I would leave alone: Calder's own before-send delay, two to three seconds while it scores an email, does not need this same rigor. Nobody reports that number outside the engineering team, it is just a normal speed target, not a claim made to a regulator.

The lesson: a metric for something that prevents work has to answer "compared to what." A count the model can move by itself can never answer that question honestly, no matter how good it looks climbing on a slide.

Now here is the same thing as a story

The short version is above. Read on for the Tuesday a rising number almost became the only thing anyone checked.

Soledad Delacourt has run product for Calder for a year and a half. She is good at the part of the job most product managers dodge: sitting with a compliance report and reading it line by line, the way an underwriter reads a claim, instead of skimming for the headline number.

When Calder launched, the number on the Monday product review slide only ever went one direction. Two hundred and ten flags the first month. Two hundred and eighty by month three. Three hundred and forty by month six. Everyone in that room liked watching it climb, Soledad included.

For the first two months she opened a random dozen flagged emails herself every Friday afternoon, just to feel out whether they were real violations or noise. By month four she was down to two or three. By month six she had stopped opening any of them. She watched the number on the slide instead, the way everyone else did.

The trigger was small. Cassia Nettleton, three weeks into the compliance team, sat in on a Monday review and asked one question nobody in the room had an answer for. "Three hundred forty flags. Compared to what?"

Soledad started to answer with the obvious thing, that flags had climbed from two hundred and ten to three hundred and forty, proof Calder caught more each month. Then she stopped mid sentence. A climbing flag count could mean Calder was finding more real violations. It could just as easily mean the model team had nudged the flagging threshold down after a complaint about missed cases the quarter before. She did not actually know which one it was. Neither did anyone else in the room.

We had been showing the audit committee a number that could rise for the best possible reason or the worst possible one, and it would look exactly the same on the slide.

Eighteen months earlier, in a half hour meeting when Calder first shipped, the team picked "flags raised per week" for the launch dashboard because it was the only number the model produced without anyone doing extra work. No baseline needed. Ships day one. Easy to read from the back of the room. Nobody there was picturing a state examiner asking to see the math behind it. Calder was still a prototype then, and this was just the fastest number to put on a chart.

What Soledad actually did was not rebuild the model. She ran a holdback. For three weeks, Calder scored a random eight percent slice of Northgate's email, about four hundred and eighty a week, but never showed the score to the agent and never held anything back. She asked Selwyn Sarto's compliance team to read every one of those emails by hand for the full three weeks, not a sample, a full read, the only way to get a real number for what ships when nothing stops it.

Hand sketched comparison diagram titled Which number survives being checked later. Left panel, a gauge icon labelled Flags issued, caption Turn the threshold down, this climbs, nobody outside the model team can tell why. Right panel, a scale icon labelled Violations avoided, caption Needs a holdback rate logged at the time, cannot be rebuilt after the fact.
A flag count can be moved after the fact by anyone with access to one setting. A rate measured on a logged holdback cannot be reconstructed once the three weeks are over, which is exactly what makes it the number to trust.

Three weeks later, the real numbers came in. Across the holdback slice, without Calder, 2.1 percent of email carried a real violation through to a policyholder's inbox, close to what the old sample had always implied but never confirmed at real volume. A follow-up audit of a thousand Calder-cleared emails from the protected 92 percent of traffic found a real violation rate of 0.3 percent slipping past review. The gap, 1.8 points, applied to Northgate's full six thousand emails a week, comes out to roughly a hundred and eight violations avoided every week, worth about forty three thousand dollars a week at Northgate's average cost per shipped violation, just over two million dollars a year.

The flag count was a number the model could move by itself. The violations-avoided number needed a person to sit with an unprotected slice of real customer email for three weeks and count. That is the whole difference between a metric and a guess with a chart under it.

What I would tell myself, watching that slide climb for six months straight: a number going up is not the same thing as a number meaning something, and I let the first one stand in for the second for a lot longer than I should have.

ORDER, run against a number instead of a decision

This is a question about which candidate metric to trust, not a story about a habit with two settings, so ORDER fits: rank the numbers by which one survives being checked later.

O
Outcome. What every candidate metric is competing to capture.
Not "how many times the model raised its hand." Real regulatory exposure kept off a policyholder's screen, stated as a number nobody can move by adjusting one setting.
R
Reversibility. Rank by which measurement is hardest to fake or rebuild after the fact.
Flags issued can be inflated any time by turning the threshold down, and nobody can tell after the fact why the count moved. A rate measured on a holdback, logged at the time it happened, cannot be rebuilt once the window has closed.
D
Dependency. What actually has to exist before what.
A prevention number means nothing without a real baseline rate of what ships when nothing stops it. That baseline depends on a holdback, logged and read in full, not assumed from the old spot check.
E
Evidence. What you could learn cheaply before committing to a permanent measurement program.
A time-boxed holdback, eight percent of traffic for three weeks, read in full by compliance instead of sampled, priced in real hours but small next to a permanent audit function.
R
Rank. State the order, defend the top pick in one line.
Run the holdback first, compute the baseline gap, convert it to violations avoided in dollars, report that externally. Flags issued stays an internal engineering number, never the proof of value.

Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. The alternative most candidates reach for by reflex is trusting Calder's own confidence score as the value proxy, since the model already produces it and it needs no extra work to report. That got ruled out on purpose: a confidence score is the model grading itself, and the whole reason a measured number like this matters is that it comes from a check outside the model, not another number the model can quietly move. Second, the flag threshold itself is not a fixed rule that fires or does not. It is a cut-off on the model's own risk score, calibrated against a fixed golden set of six hundred labeled emails so it catches real violations most of the time, by design, not every time and not on a hand-written keyword list. That calibration is also where the real trade-off sits. Turn the threshold down and Calder catches more real violations, but it also holds up more clean emails for a person to clear, real minutes an agent spends waiting on a flag that turns out to be nothing. Northgate accepted a slower send on the highest-risk email types, quotes and cancellations, and left the threshold looser on routine, low-risk mail.

Knowledge spark: why not just trust the model's own confidence score instead of running a live holdback? Because a confidence score is the model's opinion of itself. If the model has quietly drifted, its confidence drifts with it, and nothing outside the model would notice. A holdback checked by a person, on real traffic, catches that drift because it does not depend on the model being right about being right.

And if you want to be sure it really works, try it somewhere else

Same five checks, a different industry, so the method proves itself instead of repeating a story you happened to prepare.

Hollowell Apothecary runs Ossory, an AI tool that reads outgoing prescription emails, refill denials, formulary substitutions, prior-authorization notices, before they reach a patient, for about nine hundred thousand mail-order patients.

O, outcome. What every candidate metric competes to capture: real patient harm and appeal complaints kept off a pharmacist's desk, not how many emails Ossory flagged this week.
R, reversibility. "Denials flagged" can be inflated any time by loosening what counts as missing language, and nobody can tell after the fact whether that means more real gaps or a jumpier model. A harm rate measured on a logged holdback cannot be rebuilt once the window closes.
D, dependency. Florian Havens, who runs product for Ossory, first needed a real baseline: the rate of denial emails that go out missing a required appeal notice or plain-language summary when nothing stops them.
E, evidence. A two-week, five percent holdback, read in full by Hollowell's own compliance pharmacist instead of the usual spot check.
R, rank. Report harm avoided, converted into appeal complaints avoided, as the number that goes to Hollowell's board. Flags raised stays internal, a dial the model team watches, never the proof anyone shows outside the building.

Same shape, different stakes At Northgate, the ungrounded number was a flag count that could climb for the wrong reason. At Hollowell, it is the same shape wearing a different name. The rank does not change: whatever number the model can move on its own outranks nothing, and the number a person had to measure by hand always outranks it.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule: reject flags issued first, then say the real number needs a measured baseline, not the other way round.
Cost: the compliance team says a three-week holdback is too expensive to run every quarter. Do not cave and report the flag count in the meantime. Run a shorter, smaller holdback instead, and say plainly that a rougher real number still beats a precise fake one.
The model got better, for real: say Calder's precision jumps ten points overnight. That still is not the same claim as "the flag count now means something." A better model can still be handed a threshold nobody remembers tuning.

Where people run it wrong.
They report the number the model produces for free, because it needs no extra work, instead of asking what that number is actually measuring.
They build a permanent measurement program off one holdback, then discover the real rate moved once a new state rule shipped and their baseline is a year stale.
They let the model's own confidence score stand in for an outside check, because it is the fastest number to put on a slide, not because it answers the question.

How to use it live. Say the rule before naming a single number: "I rank candidate metrics by which one survives being checked later, not by which one needs the least work to produce." That buys you the room to reject the easy answer out loud, instead of reciting "track the flag count" on reflex.

Flashcards (click a card to flip it)

1 · FRAMEWORK
What framework fits ranking candidate metrics for a feature that prevents work, and why?
Tap to flip
ANSWER
ORDER. It ranks candidate numbers by which one survives being checked later, not a story about a habit with two settings.
2 · PERSON
Who is this answer about?
Tap to flip
ANSWER
Soledad Delacourt, product manager for Calder at Northgate Assurance, working with compliance lead Selwyn Sarto.
3 · REJECTED METRIC
What's the obvious metric here, and why does it fail?
Tap to flip
ANSWER
Flags issued per week. It climbs when the flagging threshold drops, so nobody outside the model team can tell if it means more real violations or just a jumpier model.
4 · OUTCOME
What is every candidate metric actually trying to capture?
Tap to flip
ANSWER
Real regulatory exposure kept off a customer's screen, not how many times the model raised its hand.
5 · DEPENDENCY
What has to exist before a prevention number means anything?
Tap to flip
ANSWER
A real baseline: the rate of violations that ship when nothing stops them, measured on a logged, fully-read holdback slice.
6 · THE NUMBER
Fill in the blank: the holdback found a real violation rate of ___ percent; the Calder-protected traffic came in at ___ percent.
Tap to flip
ANSWER
2.1 percent held back; 0.3 percent protected. The 1.8-point gap, applied to full volume, is the violations-avoided number.
7 · RANKED CALL
What number actually gets reported, and how is it built?
Tap to flip
ANSWER
Violations avoided: the 1.8-point baseline gap, applied to full weekly volume, converted to dollars using the average cost per shipped violation.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its ranked call?
Tap to flip
ANSWER
Ossory, the outgoing prescription-email compliance tool at Hollowell Apothecary. Its ranked call is harm avoided, converted into appeal complaints avoided, not flags raised.

Check yourself Score: 0 / 0

True or false
1. True or false: reporting "three hundred forty flags a week" is a fair way to show Calder's value, since flags and real violations are basically the same thing.
  • True
  • False
Show hint
Check what happens to the flag count when the threshold moves.
Show answer
False. Flags include false positives and can be inflated by turning the threshold down. The real number, violations avoided, needs a measured baseline, and it comes out to about a hundred eight a week, close to a third of the raw flag count.
Multiple choice
2. Why does a flag-issued count fail as a value metric for a feature that prevents work rather than performs it?
  • A. The model cannot be trusted to run in production at all.
  • B. It can be inflated by turning the flagging threshold down, and nobody outside the model team can tell why it moved.
  • C. Flags are private company data and cannot be reported to anyone.
  • D. Compliance teams do not trust automated tools of any kind.
Show hint
Think about what one dial can do to the count with no change in real violations.
Show answer
B. A count that moves whenever the model team changes one setting cannot answer "compared to what," which is the whole question a prevention metric exists to answer.
Fill in the blank
3. The three-week holdback found a real violation rate of ___ percent when Calder was held back, against ___ percent on the traffic Calder protected.
Show hint
Check the bar chart in "Let's learn."
Show answer
2.1 percent; 0.3 percent. The 1.8-point gap between them, applied to full volume, is where the violations-avoided number comes from.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the memory of the half-hour meeting eighteen months earlier, not a dial anyone could just turn back up.
Show answer
Model answer: Choosing "flags raised per week" as Calder's launch dashboard metric, because it was the only number the model produced without anyone doing extra work, no baseline needed, easy to ship day one. Nobody in that meeting expected a state examiner to ever ask for the math behind it.
Short answer, apply it yourself
5. Think of a tool you use that stops something rather than doing something for you, a spam filter, a spell-check underline, a car's lane-departure beep. What's the honest "value" metric for it, and what's the easy-to-fake number people probably report instead?
Show hint
Look for a count the tool itself can move versus a rate someone would have to measure by hand.
Show answer
Model answer: A spam filter's easy fake number is "emails blocked per day," which climbs the moment the filter gets stricter. The honest metric is the real spam rate reaching the inbox with the filter off, on a small logged slice, compared against the rate with it on.
Short answer, the number question
6. If Northgate's average cost per shipped violation were 100 dollars instead of 400, would the ranked call, report violations avoided in dollars rather than flags issued, still be the right one? Show the reasoning.
Show hint
Think about what the cost figure changes and what it does not.
Show answer
Yes, the call would not change. The cost per violation changes the size of the dollar figure reported, not which measurement approach survives being checked later. Flags issued is still the number one dial can move; violations avoided still needs a holdback either way.
Before you close the answer
Why this works
Tests whether you'll settle for a metric the model can move on its own, or insist on a real, measured baseline before reporting a number that people outside the team will act on.
Follow-up traps
"Isn't a shadow-mode holdback letting real violations reach real customers on purpose?" Response: yes, a small, bounded, logged slice, eight percent for three weeks, monitored the whole time, which is close to the same exposure the old spot-check method already accepted forever, just finally measured honestly.

"Why not just trust the model's own confidence score instead of running a live holdback?" Response: a confidence score is the model grading itself. Only a check that lives outside the model, a holdback read by a person, can catch the model quietly drifting.
If pressed
The golden eval set used to calibrate Calder's threshold holds six hundred labeled emails, split across eleven known violation types. It gets re-scored every time Calder's model changes or a new state disclosure rule ships, and the flagging threshold is only allowed to move if precision on that set stays inside a two-point band, not whenever the flag count looks low that week.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more