CaseAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #19

How would you measure whether an AI feature improved decision quality, not just speed?

The direct answer
Measure decision quality by what happened to the actual decision, not by how fast it got made. Tag every position call where the AI brief was used, then score the reasoning behind it against what really occurred, months later, against a small group who also read the primary source. Time to decide only gets to stand alone as a metric once that quality gap holds under a set line for two quarters running, and never for a call that moves real money.
Do this, in order
  1. Score decision quality against what actually happened to the decision, not against how fast it got made.Why: this is the actual position, and every bullet below only exists to make it real.
  2. Build the outcome check itself: tag every AI-brief decision, then score its reasoning against reality three to six months later, using a reviewer who was not in the room.Why: without this pipeline, "measure the outcome" is a sentence, not a system.
  3. Run a compression audit on a random sample of briefs every quarter, checked against the source transcript by a person, not the model grading itself.Why: this catches the specific way the model fails, dropping a conditional clause, before it shows up in a client's numbers.
  4. Split the quality number by how much the decision actually risks, never one blended average.Why: a screening pass and a position that moves size are not the same decision, and averaging them hides which one is actually bleeding.
  5. Set the line that lets speed alone stand in: a low-stakes, reversible call, or a quality gap under a fixed cut-off held for two straight quarters.Why: without a real stop line, the audit runs forever and nobody trusts the tool or ever stops checking it.
  6. Skip the confidence survey. Do not ask analysts how sure they feel about a brief.Why: feeling sure and being right are not the same thing, and a fast, clean brief is built to make people feel sure.

How to answer this, stage by stage

Nobody is grading whether you can say "measure the outcome." They are grading whether you can build the actual pipeline that measures it, and say what would ever let you stop. Seven moves, in order.

1
Scope it to one real brief, one real call
Say it like this
"Let's ground this in something specific. Sorel Capital is a mid-cap fund. Their analysts used to read the full earnings call, the filing, and three or four sell-side notes before deciding what to do with a position. They built Strand, a tool that turns all of that into a one-page brief. Dariusz Kavanagh covers industrials, and he reads that brief instead of the source material now."
Why this works
Grounds the answer in a real tool and a real job before a single number gets named.
2
State the position before any of the reasoning
Say it like this
"Here's my pick. Measure decision quality by what happens to the decision, not by how fast it got made. I would rather have a slower, noisier answer that is actually true than a fast, clean number that only proves the brief is quick to read."
Why this works
This is the P step. Interviewers are testing whether you can commit before you've earned the right to hedge.
3
Say why time-to-decide is the wrong proxy, in plain terms
Say it like this
"Time to decide only tells you the brief is short. It cannot tell you the brief is right. Sorel's median time from earnings call to position decision dropped from about four and a half days to under a day. That number climbed every single quarter, right through the quarter it cost them money."
Why this works
Naming a metric that looked perfect right up until it wasn't is the strongest move available in a measurement question.
4
Name who feels each kind of error
Say it like this
"If you only track speed, you reward a brief that helps someone decide fast and wrong. Nobody sees that cost that day. It shows up in the portfolio a quarter later, and by then nobody connects it back to a twenty-two minute read. If you only track the outcome, you're stuck waiting months for a confirmation, and the market moves for a hundred other reasons in that time, so you can't cleanly blame or credit the brief either."
Why this works
This is the I step. Both sides of the tradeoff get a real cost and a real person feeling it, not just the one the question is obviously fishing for.
5
Find the actual cost asymmetry, and say which one you're optimizing against
Say it like this
"A fast wrong call compounds. It sits in the book, unflagged, while the next ten decisions get made the exact same way, off the exact same kind of brief. An outcome measurement lag just delays how sure you feel. It doesn't cost you anything by itself, it just makes you wait. Those are not the same size of problem, so I optimize against the one that compounds, even though it's slower to catch."
Why this works
This is the C step, the actual heart of PICK. Naming which error is cheap and which is expensive is what makes the position defensible instead of just stated.
6
Give the real mechanism, not a promise to "monitor outcomes"
Say it like this
"Every brief-assisted decision gets tagged. Three to six months out, a reviewer who wasn't in the room scores the thesis against what actually happened, using a fixed rubric, checked against the transcript. On top of that, we sample twenty-five briefs a quarter and diff them against the source call by hand, specifically checking whether a guidance number kept its conditions attached. That's how we catch the model quietly dropping a hedge, which is the actual failure mode here, not a generic reading-comprehension slip."
Why this works
A named, buildable pipeline beats a vague commitment to "track it." This is also where the AI-specific failure mode lives: a summarization model compressing away a low-salience conditional clause attached to a high-salience number.
7
Give the kill criteria, and close on what you rejected
Say it like this
"Speed alone becomes fine to trust in exactly one place: a low-stakes, reversible call, like which names get a deeper read this week. No money moves, and a miss gets caught next cycle. For anything that sizes a position, the outcome check never fully goes away, even once the quality gap clears a set line for two quarters running. We also looked at just asking analysts how confident they felt in a brief, and dropped it. Feeling sure isn't the same as being right, and a fast, clean brief is exactly the kind of thing that makes people feel more sure than they should."
Why this works
This is the K step plus the rejected alternative in one breath. A pick with no kill criteria sounds stubborn. A pick with no rejected option sounds lucky.
If you remember one thing A dropping time-to-decide number is not proof the decision got better. It is only proof the decision got faster, and those are only the same thing once you've gone and checked.

Let's learn

What happens when a team gets faster at deciding and never actually checks whether the deciding got better?

Strand is a feature inside Sorel Capital's research desk. Before a position review, it reads the earnings call transcript, the latest filing, and a stack of sell-side notes, and hands back a one-page brief: the numbers that moved, the guidance, the risks flagged on the call, in order of what matters most.

Before Strand, an analyst spent about five and a half hours reading the source material ahead of each name's review. Sorel's team covered maybe three names a week at that pace. After Strand, the brief takes about twenty-two minutes to read. The team started covering names the same day they reported, sometimes the same hour.

The number everyone watched, and the number nobody built
Prep time per name (minutes)
330 22 Before After
Decision-quality score, out of 10 (later, by review)
8.3 6.1 Full read Brief only
Prep time fell 93 percent and kept falling, quarter after quarter. Nobody had a chart for the right-hand number until a reviewer went back and scored 240 decisions by hand.

Here is the turn. The 93 percent drop in prep time was real, and it was good. But nobody at Sorel had ever built a number for whether the decisions coming out of that fast brief were actually right. When a reviewer finally went back, months later, and scored 240 of them against what really happened, brief-only decisions averaged 6.1 out of 10 on thesis accuracy. Decisions where the analyst also pulled the primary transcript averaged 8.3.

We did not make Sorel's analysts faster at deciding. We made them faster at deciding wrong, and nobody had a dashboard for that.

Strand's brief was not sloppy. It read every transcript, every filing, and it almost always got the headline number right. What it did, quietly and often, was drop the condition sitting next to a guidance number. A CFO says margins will hold flat, assuming a cost pressure from last quarter doesn't come back. Strand's brief keeps "margins expected flat" and loses the "assuming." That clause sits far from the number it qualifies, and a model built to compress a call into six bullets treats it as the part that's safe to cut.

Two ways to be wrong about a brief. Left, a gauge labelled measurement lag, you wait longer to be sure, nothing lost while you wait. Right, a scale tipping labelled a fast wrong call, ships today, compounds in the book before anyone checks it.
Waiting to be sure costs you time. A fast wrong call costs you money you don't notice yet.
Knowledge spark: what is compression loss? When a model turns a long passage into a short one, it has to decide what to keep. A number usually survives. A condition attached to that number, sitting in its own clause a sentence later, often does not, because on its own it looks like a minor detail rather than the thing that changes what the number means.

At its worst, this cost more than a bad grade on a rubric. Sorel's position in Ferrand Machine Works, a mid-cap industrial supplier, got sized up the same day Dariusz read a brief that said margins would hold flat. Six weeks later the cost pressure the CFO had actually flagged, a steel surcharge, came back. Ferrand's next quarter showed margin compression of about 180 basis points, exactly the range the CFO had warned about on the call Strand had summarized. The position cost the fund roughly 60 basis points at the portfolio level before it got trimmed.

The choice I would take back is not building Strand. It's that Strand shipped able to compress a guidance number away from the condition attached to it, with nothing checking whether that condition survived the trip. That was invisible while the brief mostly handled routine calls with nothing conditional in them. It stopped being invisible the first time a real hedge sat inside a real number.

What I would leave alone: Strand's summary of a company's cost breakdown or segment mix. Those are just numbers restated, nothing conditional to lose, and a reviewer checking them against the filing takes thirty seconds. Auditing something that never drops a clause is thirty seconds spent on nothing.

The lesson: a metric that only measures how fast someone can act was never going to tell you whether they acted on the truth. You have to go build the second number on purpose, and be honest that it arrives late.

Now here is the same thing as a story

The short version is above. Read on for how ordinary the week looked from inside Sorel, right up until it wasn't.

Dariusz Kavanagh has covered industrials at Sorel Capital for six years. He built a reputation the slow way, by being the analyst who actually finished the 10-K instead of skimming the earnings-call transcript and calling it done. Before Strand, his prep for a name reporting that week looked like a Tuesday spent entirely inside a PDF.

Strand arrived in the winter, and for months it was the best change to his job in years. He'd open the brief at eight, done reading by eight-twenty-two, and be in a position review by nine with time to actually think about sizing instead of just catching up on facts. He still pulled the full transcript on names he was worried about. He just stopped needing to on the routine ones.

Then routine started meaning something different, without anyone deciding it should.

He never noticed the brief getting worse. It never got worse. It just kept being confidently, cleanly wrong about the one sentence that mattered.

Ferrand Machine Works reported on a Thursday. The CFO spent about forty seconds on margins, said they'd hold roughly flat in the back half, then added, almost as an aside, that this assumed a steel input surcharge from the spring wouldn't recur, and that if it did, they'd see another 150 to 200 basis points of compression. Strand's brief carried one bullet: "Guidance: margins expected flat in H2." Clean. Confident. Missing the second half of the sentence.

Flow diagram, four boxes. The CFO's catch, brief drops it, bullet says flat highlighted, fast call wrong.
Three of these four steps happened inside twenty-two minutes. Nobody was in the room to catch the second one.

Dariusz read it, agreed with the desk that flat margins meant the stock was cheap at its current multiple, and added to the position before lunch. Under an hour, start to finish. By the old process that decision would have taken most of a day, with the transcript open the whole time and that "assuming" landing right where the CFO put it.

Six weeks later the surcharge came back. Ferrand's next print showed the 180 basis points of compression the CFO had actually described on the call, word for word, in the part of the sentence the brief never carried. The position had already cost the fund about 60 basis points at the portfolio level by the time it got trimmed.

Nobody blamed Dariusz. He'd read the brief the way the brief was built to be read, confidently, once, and moved on. The old decision that mattered was made two years earlier, when Strand shipped with no rule requiring a conditional clause to travel with the number it qualified. That was fine right up until the first time a hedge sat inside a headline number on a name someone actually sized.

What Sorel did next: built the outcome pipeline, the compression audit, and a rule that any bullet built off a guidance number now carries the sentence it came from, in full, one click away. The next time a CFO added a condition to a number, the brief kept it. Dariusz's own gap between brief-assisted and full-read decisions, tracked from that point forward, closed from 2.2 points to 0.4 within three quarters.

What I'd tell myself, before any of this: a number that only measures the part of the trade you already like was never going to warn you. You have to go build the other one, and admit up front that it's going to be slow and a little noisy, because that's the price of it actually being true.

PICK, for tradeoff questions

This is a tradeoff question dressed as a metric question, speed versus quality, so PICK fits better than LEAD. LEAD is for finding the earliest signal of a known outcome. Here the real work is committing to which number to trust when they disagree.

PICK, for tradeoff questions. Position, pick first. Impact, who feels it. Cost asymmetry, which error hides. Kill criteria, what flips you.
Four letters, one call to defend
P
Position. State the pick before the reasoning.
Measure decision quality by the outcome tied to the actual decision, not by time to decide.
Said first, no hedging, before a single number gets named.
I
Impact. Who feels each kind of error.
Speed-only tracking rewards a fast wrong call, felt by the whole portfolio, months later, never traced back to the brief. Outcome-only tracking makes the desk wait months for confirmation the tool works, confounded by everything else moving the market.
Neither group is imaginary. Both get named.
C
Cost asymmetry. Which error is cheap, which is hidden.
A fast wrong call compounds silently across every decision made the same way before anyone checks. A measurement lag just delays your confidence, it doesn't cost anything on its own.
Optimize against the one that compounds, even though it's slower to catch.
K
Kill criteria. What would flip the pick.
A low-stakes, reversible call, like which names get a deeper read this week, where speed alone is fine. Or a quality gap that holds under 0.5 points for two quarters running, for the screening tier only, never for a call that sizes a position.
This is the line, charted, below.
The quality gap, by quarter, against the line that would flip the pick
2.2 pts 0.5 pts 0.0 kill line Q1, gap found Q2, audit live Q3 Q4
The gap fell from 2.2 points to 0.4 in three quarters once the compression audit was in place. It crossed the kill line in Q4, for the low-stakes screening tier only, position-sizing calls still get the outcome check regardless.

One thing worth being plain about, since this is where the AI-specific judgment actually lives. The alternative most teams reach for first is asking analysts how confident they feel in a brief, a quick pulse survey after each read. That got rejected on purpose: confidence and correctness are not the same measurement, and a clean, well-formatted brief is exactly the kind of output that makes a reader feel more certain than the underlying reasoning earns, which is the specific trap here, not a generic trust problem. The bar for the compression audit is calibrated, not a hard rule: a brief passes when it preserves the conditional attached to a quantified guidance number on a sampled eval set of past calls, checked most of the time, not promised to be perfect every time. The failure mode worth naming directly is compression loss, a summarization model dropping a low-salience clause next to a high-salience number, and the guardrail is the quarterly audit plus the citation-back requirement, not a person eyeballing the brief once and moving on. The trade is real: building the outcome pipeline costs analyst and data time, and it takes a full quarter or two before you can honestly say the tool's decisions hold up, against a same-day dashboard of time saved that could ship on day one.

And if you want to be sure it really works, try it somewhere else

Same four letters, a hospital radiology desk instead of a research desk, so the method proves itself instead of repeating a story I happened to prepare.

Cedarbrook Radiology uses Talbryn, an AI tool that pre-reads chest CT scans and flags the ones most likely to need a second look, with a short written summary of what it found. Toma Iskander is one of six radiologists who sign off on flagged cases each week.

P, position. Measure decision quality by what the follow-up imaging or biopsy actually showed, weeks later, never by how fast a radiologist clears the queue.
I, impact. Speed-only tracking rewards clearing the queue fast on a wrong call, felt by a patient who gets told it's nothing. Outcome-only tracking means waiting weeks for a biopsy result, confounded by everything else about that patient's case.
C, cost asymmetry. A missed finding compounds, the patient doesn't come back until it's worse. A slow outcome check just delays confidence in the tool, it costs nothing by itself.
K, kill criteria. Speed alone is fine for the first-pass triage sort, deciding which scans get read today versus tomorrow, fully reversible, nobody's diagnosis rests on it. It's never fine for the sign-off itself.

Same flagged case, two different Tuesdays. Left, a document, pulls the scan, checks the original image series before signing off. Right, a gauge, trusts the flag, signs off on the AI summary alone, scan unopened.
Same tool, same flag. One habit checks the picture. The other checks the caption.
Same shape, different stakes At Sorel, the unmeasured cost was a mispriced position. At Cedarbrook, it's a finding nobody caught in time. The C step doesn't change: name which error compounds silently, and optimize against that one, even when it's the slower number to collect.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pick: tie the metric to the outcome of the actual decision, not to how fast the decision got made, then name the kill criteria in one line.
Cost: the outcome pipeline needs a data engineer for two quarters, not two weeks. Don't ship the speed-only dashboard as a stand-in and call it "close enough." Hold position-sizing decisions to the old, slower process until the pipeline exists.
The model got better, for real: say Strand's summarization genuinely improves. That still isn't proof it stopped dropping conditionals specifically. A better model finds new ways to sound confident about the wrong thing, it doesn't retire the need to check.

Where people run it wrong.
They treat a falling time-to-decision number as proof the tool is working, instead of asking what it would look like if the tool were making people fast and wrong.
They build the outcome check once, as a one-time study, instead of a running pipeline that keeps sampling every quarter.
They set no kill criteria at all, so the audit either runs forever or gets quietly dropped the first time someone complains it's slowing things down.

How to use it live. Open with the split before naming a fix: "speed tells you the brief is short, it can't tell you the brief is right, so I'd measure the second thing on purpose." That buys you the room to give the real answer instead of reaching for "we'd track engagement with the feature," which is speed wearing a different name.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits deciding whether an AI feature's speed win is also a quality win, and why?
Tap to flip
ANSWER
PICK. It's a tradeoff question, speed versus quality, so it needs a committed position and a named cost asymmetry, not a leading-indicator search like LEAD.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dariusz Kavanagh, a senior research analyst covering industrials at Sorel Capital, using Strand, the fund's in-house AI research brief.
3 · THE POSITION
What's the P step here, in one line?
Tap to flip
ANSWER
Measure decision quality by the outcome tied to the actual decision, not by time to decide, even though the outcome measure is slower and noisier to collect.
4 · THE ASYMMETRY
What's the C step, the actual cost asymmetry?
Tap to flip
ANSWER
A fast wrong call compounds silently across every decision made the same way before anyone checks. A measurement lag just delays your confidence, it costs nothing by itself.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Shipping Strand with no rule requiring a conditional clause to stay attached to the guidance number it qualifies. It was invisible while most calls had nothing conditional in them.
6 · THE NUMBER
Fill in the blank: brief-only decisions scored ___ out of 10 on thesis accuracy. Full-read decisions scored ___.
Tap to flip
ANSWER
6.1 and 8.3. A 2.2 point gap that nobody had measured until a reviewer scored 240 decisions by hand.
7 · THE KILL CRITERIA
What would make time-to-decide an acceptable metric on its own?
Tap to flip
ANSWER
A low-stakes, reversible call, like a screening pass with no money committed. Or a quality gap held under 0.5 points for two quarters running, and even then, never for a decision that sizes a position.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the outcome check?
Tap to flip
ANSWER
Talbryn, Cedarbrook Radiology's flagging tool. Its outcome check is comparing flagged reads against what the follow-up imaging or biopsy actually showed, weeks later.

Check yourself Score: 0 / 0

Multiple choice
1. What is the direct answer to "how would you measure whether an AI feature improved decision quality, not just speed?"
  • A. Track how many analysts open the AI brief each week.
  • B. Ask analysts to rate how confident the brief made them feel.
  • C. Tag every AI-assisted decision and score it against what actually happened, months later, against a comparison group who read the primary source.
  • D. Time how long each brief takes to read and set a target under thirty minutes.
Show hint
The answer has to check the decision, not the reading experience.
Show answer
C. Only an outcome tied to the real decision, checked against a comparison group, actually tests quality instead of speed or self-reported trust.
True or false
2. True or false: once Sorel's median time-to-decision started dropping every quarter, that alone was good evidence the AI brief was working.
  • True
  • False
Show hint
Check what that number kept doing right through the quarter Ferrand cost the fund money.
Show answer
False. The time-to-decision number climbed steadily the whole way through, including the quarter a dropped conditional clause cost the fund about 60 basis points on one position. Speed was never proof of quality.
Fill in the blank
3. Over three quarters after the compression audit went live, the decision-quality gap between brief-only and full-read decisions fell from 2.2 points to ___ points.
Show hint
Check the kill-line chart in the PICK recap.
Show answer
0.4 points. Close to, but still short of, the point where the fund would let time-to-decision stand alone even for the low-stakes tier.
Short answer
4. Name one place in Strand's own workflow where this exact quality check would be a waste of time, and say why.
Show hint
Look for "what I would leave alone" in "Let's learn."
Show answer
Model answer: A brief's restatement of a cost breakdown or segment mix. There's no conditional clause to lose, it's just numbers pulled straight from the filing, so checking it against the source takes thirty seconds and rarely finds anything.
Short answer, apply it yourself
5. Think of an AI tool you've used that made some task faster. What would a real decision-quality check look like for it, one that couldn't be satisfied just by the task getting done quicker?
Show hint
Look for something you'd only find out was wrong after some time had passed.
Show answer
Model answer: An AI code-review assistant that flags issues fast. A real check would not be how many pull requests get approved per day, but how many of the bugs it missed showed up as production incidents in the following month, compared against a sample of pull requests reviewed the old way.
Multiple choice
6. Why did this answer reject asking analysts how confident a brief made them feel, as the quality signal?
  • A. Analysts refused to fill out surveys.
  • B. Feeling confident and being correct are not the same thing, and a clean brief is built to make people feel more certain than the reasoning earns.
  • C. Surveys are too expensive to run every quarter.
  • D. Confidence data can't be stored alongside outcome data.
Show hint
Check the paragraph right after the PICK recap's four rows.
Show answer
B. This is the actual AI-specific risk: a fluent, confident-sounding output trains the reader toward overconfidence, which is the opposite of a useful quality signal.
Before you close the answer
Why this works
Tests whether you will commit to measuring something slower and harder to collect because it's actually true, instead of reporting the number that's easy to graph. Most candidates stop at "we'd track time saved and user satisfaction."
Follow-up traps
"Isn't waiting months for outcome data just too slow to be useful in practice?" Response: that's exactly why it's paired with the quarterly compression audit, a much faster proxy that catches the specific failure mode directly, without waiting for a position to mature.

"Couldn't the reviewer scoring the rubric just be biased toward finding problems?" Response: that's why the reviewer is never the analyst who made the original call, and the rubric is checked against the transcript, not against their own judgment of what a good decision looks like.
If pressed
The compression audit's twenty-five-brief sample isn't picked at random from all briefs. It's weighted toward calls where a company's guidance included a stated condition, because that's where the failure mode actually lives, and a flat random sample would mostly draw briefs with nothing to drop in the first place.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more