CaseIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #23

Describe the metrics you would show a board versus the ones you would show your team.

A single trust number can be true, climbing, and hiding the one category about to break, all at once. Rank the metrics by what a wrong number would cost, not by who is in the room when you show it.

The direct answer
Give the board the blended trust number and the worst clause category sitting right beside it, never the blend by itself, gated behind a weekly eval score before any new contract type reaches a sales deck. Give the team the same numbers broken all the way down to clause type and model version, because that is where they can actually fix something this week. Build both from one pipeline, so the board's number can never say more than the team's own data can prove.
Do this, in order
  1. Put the worst clause-category floor next to the blended average on every board slide, never the blend alone.Why: this is the number a board would approve an unrecoverable expansion on top of, so it can't hide.
  2. Give the team the full clause-by-clause acceptance and edit numbers, tagged by model and prompt version.Why: this is where they find the exact place to fix this week, not somewhere inside a blend.
  3. Build the board's number as a straight roll-up of the team's own tagged data, one pipeline, never two.Why: two separately built numbers drift apart quietly, and nobody notices until an incident forces the comparison.
  4. Run a weekly golden-set eval, hand-checked by outside counsel, sampled by clause type.Why: cheap to run, and it would have caught the weak category before a customer did.
  5. Gate any new contract category behind that eval score, not behind the blended acceptance rate.Why: this is the actual board decision the whole ranking exists to protect.
  6. Leave the payment-terms and confidentiality categories alone.Why: real, stable trust already sits there. Adding a review gate would only slow the team down for nothing.

How to answer this, stage by stage

Eight moves. The trap in this question is answering with two separate lists, a fancy one and a detailed one, when the real test is whether you know which number would let someone approve something they can't take back.

1
Anchor it in one product, one board, one team
Say it like this
"Let me make this concrete. Anchorpoint sells ClauseGuard, a tool that reads a vendor contract and drafts the redline a procurement person would have written by hand. Dana Tolan runs product for it, and reports to a board every quarter and to an engineering and legal team every week."
Why this works
Grounds an abstract split, board versus team, in a real product before naming a single metric.
2
Name the one outcome every metric has to serve
Say it like this
"I'd use ORDER here. Before ranking any metric, name the outcome they're all competing for: only let ClauseGuard touch a new kind of contract when it's actually good enough there, not when the average number looks good enough."
Why this works
Without a stated outcome, ranking metrics is just opinion. This is the line the rest of the answer has to keep proving.
3
Rank by what's hardest to undo, not by audience comfort
Say it like this
"The split isn't 'simple for the board, detailed for the team.' It's: which mistake can't be taken back. A board that never sees the weak category can approve selling something the model isn't ready for. A team missing a segment number loses a week, not a signed contract."
Why this works
Reframes the question away from politeness and toward the actual cost of getting either deck wrong.
4
Say the two numbers side by side, out loud
Say it like this
"So here's what I'd actually put in front of each room. The board gets the blended acceptance rate and the worst clause floor, next to each other, always. The team gets that same floor broken open by clause type, model version, and prompt version, because that's the level where a fix actually happens."
Why this works
Matches the direct answer. Names two concrete, buildable numbers instead of a philosophy about transparency.
5
Prove it with the near miss that almost signed
Say it like this
"Here's what happens without that split. Anchorpoint's blended acceptance rate hit 94 percent. Nobody on the board saw that ClauseGuard's liability-clause floor, on real master service agreements, had already fallen to 58 percent. A logistics customer, Presswick Logistics, nearly signed an MSA with an uncapped liability clause ClauseGuard had suggested striking, and their own contracts lead caught it the night before signature."
Why this works
A short, specific failure story does more work here than a paragraph about why averages can mislead.
6
Name the cheap check that would have caught it first
Say it like this
"A weekly golden-set eval would have caught this before Presswick did. Take a rotating sample of real liability clauses, have outside counsel hand-check what ClauseGuard drafted, and require a 90 percent pass rate before that contract type gets sold as no-review. That check costs a few hours of legal time a week."
Why this works
This is the evidence step. It shows the fix isn't a bigger dashboard, it's a small, cheap, recurring test.
7
Say which decision each deck actually protects
Say it like this
"The board's deck protects one decision: whether to expand into a new contract category or keep selling the current one harder. The team's dashboard protects a different one: which clause type to retrain or re-prompt this week. Same underlying data, two different decisions on top of it."
Why this works
Shows the split isn't about detail level, it's about which choice each audience is actually about to make.
8
Close on the option you rejected, and why
Say it like this
"We considered just giving the board the same full dashboard the team uses. Ruled that out. A board looking at forty clause-type rows can't find the one that matters either, they'll either ignore all of it or fixate on the wrong row. The fix isn't full transparency, it's an honest roll-up, built from the same pipeline the team already trusts."
Why this works
Naming a rejected option turns "board gets less detail" into a defended design choice instead of a shortcut.

If you remember one thing: a number that's climbing on average and falling in the one place that matters is not two separate stories. It's one story the board never got to hear.

Let's learn

ClauseGuard is a tool inside Anchorpoint's procurement software. It reads an incoming vendor contract, flags the clauses worth arguing about, and drafts the redline a person would have written by hand.

Before ClauseGuard, Anchorpoint's own procurement team redlined every contract themselves, about three hours of work on each one, on top of the actual negotiating. Dana Tolan's team built ClauseGuard to cut that down. It worked. Average time to a first redline draft dropped to eleven minutes, and the team's blended first-pass acceptance rate, the share of redlines a procurement person used with no material edit, climbed steadily: 71 percent in the first quarter, 94 percent by the third.

Hand sketched flow diagram titled What has to be true before the board sees a number. Four boxes connected left to right: Tag each clause, Roll into a floor, Clear the eval bar highlighted, Board sees the gap.
The board's number only means something once it's built on top of the team's own clause-level data and a real eval check, in that order.

Ninety-four percent is a real number. It is also not the whole story, and here is why. Most of ClauseGuard's volume, about eight in ten contracts, are simple order forms and standard NDAs, and on those, ClauseGuard is genuinely excellent. But sales had started pushing it into master service agreements, the longer contracts with indemnification and liability clauses, months before the model had been tuned or checked against enough of them.

First-pass acceptance rate: the blend the board saw, versus the floor nobody showed them
100% 0% 94% 58% Blended, all contracts MSA liability clauses
The blend was climbing because low-stakes categories dominate the count. The floor, the category a board decision would actually rest on, sat forty points lower.
Knowledge spark: what is a golden set? A batch of past clauses a human expert has already checked by hand. A new model or prompt version has to match those hand-checked answers before it ships to anyone. It is a cut-off point, not a promise the model will always get it right.

What that costs, at its worst, is not a bad redline. Dana's team pitched the board on selling ClauseGuard's MSA redlining as a paid, no-review tier, off the strength of that 94 percent. Two weeks into a pilot with Presswick Logistics, ClauseGuard suggested striking a liability cap in their MSA and replacing it with language pulled from a mismatched template, effectively uncapped liability. Callix Doran, Presswick's contracts lead, caught it the night before signature, going through the file one more time out of habit.

Nobody lied with that 94 percent. It just never got asked the one question that mattered: 94 percent of what.
The decision I would take back Computing one blended acceptance number for the whole product, with no category floor attached. It made sense when nearly all volume was low-stakes order forms and NDAs. It stopped being safe the moment MSAs entered the mix and a slide got built off that same blended number.

What I would leave alone: the payment-terms and confidentiality-boilerplate categories, which have sat at 96 percent or higher for nine straight months, stable, no drift. Adding a review gate there would slow the team down for a problem that doesn't exist in that category.

The lesson: a metric that's true on average and false at the edge is the most dangerous kind, because the edge is exactly where a board decision gets made.

Now here is the same thing as a story

The short version is above. Keep reading for the Tuesday night Presswick's own contracts lead caught what a board slide never showed.

Dana Tolan had spent five years in procurement software before ClauseGuard, most of it watching product metrics get built the ordinary way, one number, a green arrow, a slide. Dana was good at the part everyone undervalues: making a number defensible enough that nobody in the room would ask a second question about it. For the first two quarters, that instinct served the product well.

The good months looked like this. Every Monday, Dana pulled the acceptance rate, watched it climb, and dropped one line into the weekly update. Every quarter, the same number went into the board deck, a little higher each time. Procurement customers loved ClauseGuard. Nobody complained. The number kept saying so.

The blind spot grew in three beats, and nobody felt any of them happen. First, sales started quoting ClauseGuard for master service agreements, not just NDAs and order forms, because customers kept asking for it and the blended number gave no reason to say no. Second, the team's own weekly clause-level review, the one place a category floor would have shown up, quietly stopped happening every week and started happening "when there's time," because the blended number always looked fine. Third, nobody updated the board deck's single acceptance line to say which categories it was actually made of.

Then came a Tuesday night that wasn't even inside Anchorpoint's own walls. Callix Doran, Presswick Logistics' contracts lead, was doing a last pass on an MSA before signature the next morning, the kind of pass most people skip when a deal has already been reviewed twice. Callix noticed the liability clause read wrong, pulled the original template, and found ClauseGuard's redline had swapped in language that removed the cap entirely.

Hand sketched comparison titled Reversible or not. Left, a green box labeled Adjust this week's queue, caption swings back, undo it Friday. Right, a red-orange box labeled Sell MSAs with no review, caption bolted, once it is signed.
A team can re-prioritize a review queue by Friday. A board can't unsell a no-review tier a customer already signed against.

Callix didn't call Anchorpoint's support line. Callix called Presswick's own procurement lead, who called Dana directly, at 9:40 the next morning, an hour before signature. The question wasn't angry. It was worse than angry. It was simply: "Does anyone at Anchorpoint actually know how often this happens?"

Dana didn't know. Nobody did. The 94 percent everyone had been reporting had never been asked to answer that question, because nobody had ever built it to.

The real cost wasn't the one MSA, which got fixed by hand that morning in twenty minutes.

We didn't almost lose one contract clause. We almost let a customer's own contracts lead find the gap in our number before we did.

A year earlier, in the meeting where the team decided how to report ClauseGuard's health to the board, someone had asked whether one number was enough or whether the board needed a breakdown by contract type. The answer, at the time, was reasonable: almost all volume was NDAs and order forms, a single number was accurate enough, and a category breakdown would just be noise nobody had time to read. Nobody put a date on that decision or a trigger for revisiting it once the mix of contracts changed.

Run the same Tuesday through the fixed design. The board's deck already carries the MSA liability floor beside the blend, and MSA auto-redline was never sold as no-review in the first place, because the golden-set eval never cleared 90 percent on that category. Callix's late pass still finds nothing wrong, because there was never an uncapped clause to find. Dana's 9:40 call that morning is a routine renewal check-in instead of an emergency.

One design let a real number keep quietly meaning less than everyone thought it meant. The other one made sure the board's number could never promise more than the team's own data had actually earned.

What I'd tell myself, back in that first meeting about reporting: the moment "one number is enough" stops being tested against what's actually flowing through it, is the moment to write down when to check again, not to assume the answer stays yes forever.

ORDER, ranked for two audiences instead of one

This is a metric-split question, but the real judgment is a ranking call, which audience needs which signal by what's hardest to walk back, so ORDER does the work here, not LEAD.

O, outcome. Every metric on either deck competes to serve one thing: only let ClauseGuard act on its own in a contract category once it's actually earned that, not once an average number looks earned. The board and the team serve that outcome at different altitude.
R, reversibility. A board that never sees the category floor can approve selling a capability that isn't ready, and once a customer signs against it, that's not a bug fix, it's a legal exposure someone has to disclose. A team missing a segment number for a week is a slower fix, not a signed mistake.
D, dependency. The board's number only means anything if it's a genuine roll-up of the team's own clause-tagged, model-version-tagged data. Build them from two separate pipelines and they drift apart silently, the exact failure that let 94 and 58 sit unconnected on two different desks.
E, evidence. Cheap to learn before either deck ships: run the golden-set eval weekly, by clause category, hand-checked by outside counsel. It costs a few hours of legal time and would have flagged the liability-clause gap before Presswick's contracts lead had to.
R, rank. Board gets: contract value and margin under ClauseGuard-assisted negotiation, the acceptance trend with the worst category floor beside it always, the golden-set pass rate as the gate on any new category, and the override rate as a leading quality signal. Team gets: clause-by-clause acceptance and edit distance by model and prompt version, a near-miss log reviewed weekly, golden-set failure detail by clause type, and latency. The board's deck is four numbers, none of them alone. The team's dashboard is the forty rows those four numbers are built from.
The check that keeps this ranking honest Swap the outcome and the order should move. If a missed clause in an MSA only ever cost a person a second read before mailing, the floor number could sit further down the board deck. It ranks first here because Anchorpoint would be carrying a real liability exposure it might have to disclose to a customer, not a typo.

ORDER again, at a clinic where the near miss has a heartbeat

Fauna Diagnostics builds VetSight, a tool that reads an X-ray or scan an on-call vet takes at a rural clinic and flags what needs a closer look before anyone treats the animal.

O. Every version of VetSight's board-versus-team split protects one thing: a vet trusts a flagged case is genuinely urgent, and an unflagged one genuinely isn't.
R. A missed emergency case, a bloat or a torsion that VetSight reads as routine, can cost an animal's life within hours. There is no undoing that once a clinic has acted on the wrong read. A team missing a segment number for a week is a slower fix, nothing more.
D. The board's headline override rate has to be built from the same per-case category tags the team already keeps, or it will quietly stop matching what's actually happening on emergency cases.
E. Cheap to check weekly: a supervising vet reviews a rotating sample of confirmed emergency-triage cases against what VetSight flagged, and that pass rate gates whether VetSight can be sold as solo overnight coverage for a small clinic.
R. Board gets the blended override rate and the emergency-triage floor beside it, plus the eval pass rate gating any "replaces a second vet" sales claim. Team gets per-case-type override logs, near misses, and model version by release.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the ranked call: never let a blended trust number reach a board deck without the worst category's floor sitting beside it.
Cost: engineering says a true per-category eval pipeline can't ship for six weeks. Don't sell the blended number as proof of readiness in the meantime, hold the expansion claim until the floor exists.
The model got better, for real: say ClauseGuard's liability-clause handling genuinely improves next quarter. That still doesn't make a blended number the right thing to show a board, it just means the floor number, shown honestly, finally earns the expansion instead of assuming it.

Where people run it wrong.
They treat "board sees less" as meaning less accurate, instead of a different roll-up of the same real data.
They build the board's number by hand each quarter instead of pulling it straight from the team's own pipeline, which is exactly how the two numbers drift apart.
They wait for an incident to ask which category is actually weak, instead of running a cheap eval every week whether anyone remembers to ask or not.

How to use it live. Say the outcome before naming a single metric: "Every number I put in either deck protects one decision, whether this thing is ready to act on its own in a given category." That buys room to give the real ranking, instead of reciting "board gets summary, team gets detail" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a board-versus-team metric split, and why not LEAD?
Tap to flip
ANSWER
ORDER. The real job is ranking which audience needs which signal by what's hardest to undo. LEAD picks one leading metric, it doesn't split audiences by decision reversibility.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dana Tolan, who runs product for ClauseGuard at Anchorpoint, reporting to a board quarterly and to an engineering and legal team weekly.
3 · WHAT THE BLEND HID
What did the blended acceptance number stop being able to show, once volume shifted?
Tap to flip
ANSWER
Which single clause category was quietly getting worse. It kept climbing because low-stakes categories dominated the count, while the MSA liability floor fell to 58 percent.
4 · THE DEPENDENCY
What has to be true before the board's number can be trusted?
Tap to flip
ANSWER
It has to be a straight roll-up of the same clause-tagged, model-version-tagged data the team already tracks, never computed on a separate path.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Computing one blended acceptance number for the whole product with no category floor. It made sense when nearly all volume was NDAs and order forms, and stopped being safe once MSAs entered the mix.
6 · THE NUMBER
Fill in the blank: ClauseGuard's blended acceptance rate reached ___ percent, while its MSA liability-clause floor had fallen to ___ percent.
Tap to flip
ANSWER
94 percent; 58 percent. That gap is what almost let an uncapped-liability MSA reach signature.
7 · THE REPLAY
Same near miss, new design, what changes?
Tap to flip
ANSWER
With the floor number on the board deck and a golden-set gate at 90 percent, MSA auto-redline never gets sold as no-review. Callix Doran's catch becomes a normal queue flag instead of a last-night scramble.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what plays the role of the MSA liability floor there?
Tap to flip
ANSWER
VetSight at Fauna Diagnostics. The equivalent floor is the override rate on emergency-triage cases, hidden inside a blended override rate that routine wellness scans dominate.

Check yourself Score: 0 / 0

True or false
1. True or false: the fact that ClauseGuard's blended acceptance rate kept climbing proves the product was getting safer across every contract type.
  • True
  • False
Show hint
Check what dominates the count behind that blend.
Show answer
False. The blend climbed because low-stakes categories dominated volume, while the MSA liability-clause floor fell from a healthier number to 58 percent underneath it.
Multiple choice
2. What does this answer say the board should see next to the blended acceptance rate?
  • A. Nothing, the blended number is sufficient on its own.
  • B. The worst clause category's floor, always shown beside the blend.
  • C. The full clause-by-clause dashboard the team uses.
  • D. A satisfaction score from the procurement team.
Show hint
Think about what a board would approve an unrecoverable decision on top of.
Show answer
B. The floor is the number a board would otherwise never see, and the one an expansion decision actually rests on.
Fill in the blank
3. ClauseGuard's blended acceptance rate reached ______ percent, while its MSA liability-clause floor had fallen to ______ percent.
Show hint
Check the bar chart in "Let's learn."
Show answer
94 percent, and 58 percent. The forty-point gap between them is the whole reason a single blended number was the wrong thing to build a board decision on.
Multiple choice
4. What does the dependency step in ORDER argue for here?
  • A. The team's dashboard should be shown to the board unchanged.
  • B. The board's number should be simplified as much as possible.
  • C. The board's number has to be built as a genuine roll-up of the team's own tagged data, not a separate pipeline.
  • D. Only engineering should ever see the clause-level numbers.
Show hint
Ask what happens when two numbers about the same thing get computed two different ways.
Show answer
C. A separately built board number can drift from what the team actually sees, and nobody notices until an incident forces the comparison.
Short answer, apply it yourself
5. Pick a product you use that reports one big score, a rating, a health score, a percentage. What's one segment that score could be quietly hiding?
Show hint
Look for whatever makes up most of the count behind that score.
Show answer
Model answer: A ride-share app's overall on-time rate. It could be 96 percent citywide while airport pickups, a small slice of trips but the ones people care most about, run late half the time, buried inside the average by sheer volume of easy short rides.
Short answer, the number question
6. If the MSA floor had fallen to 40 percent instead of 58 percent, would gating expansion behind a 90 percent eval bar still be the right call, or would you also pull the category out of the sales deck entirely while the fix happens? Explain.
Show hint
Ask whether the gate alone stops the wrong claim from reaching a customer today, not just tomorrow's release.
Show answer
Model answer: At 40 percent, the gate alone isn't enough. A category that far below the bar shouldn't be sitting in an active sales deck at all while it's fixed, since a rep could still pitch it before the next eval run catches the drift. Pull it from the deck, keep the eval and the floor number, and only put it back once it clears the bar.
Before you close the answer
Why this works
Tests whether you understand that an aggregate metric can be true and dangerous at the same time, and whether you'd build the board's number as a real roll-up instead of a friendlier, separate story.
Follow-up traps
"Isn't showing the board a lower floor number just going to make them nervous and slow everything down?" Response: that's exactly why the eval gate sits next to it. A board that sees "58 percent today, needs 90 before we sell it" can keep investing with eyes open. A board that only sees 94 percent can't tell "handled" apart from "not tested yet."

"Why not just make the team's full dashboard the single source of truth for both audiences?" Response: a board looking at forty clause-type rows can't find the one that matters either. It gets ignored or fixated on the wrong row. The fix is an honest roll-up, not transparency dumped on the wrong audience.
If pressed
The golden set itself isn't static. Anchorpoint rotates ten percent of it each quarter with newly hand-reviewed MSA clauses, because a set built against last year's contract language starts missing the newer liability-cap phrasing outside counsel has started using, and a stale golden set would stop catching that drift too.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more