How would you measure whether an AI feature improved decision quality, not just speed?
- Score decision quality against what actually happened to the decision, not against how fast it got made.Why: this is the actual position, and every bullet below only exists to make it real.
- Build the outcome check itself: tag every AI-brief decision, then score its reasoning against reality three to six months later, using a reviewer who was not in the room.Why: without this pipeline, "measure the outcome" is a sentence, not a system.
- Run a compression audit on a random sample of briefs every quarter, checked against the source transcript by a person, not the model grading itself.Why: this catches the specific way the model fails, dropping a conditional clause, before it shows up in a client's numbers.
- Split the quality number by how much the decision actually risks, never one blended average.Why: a screening pass and a position that moves size are not the same decision, and averaging them hides which one is actually bleeding.
- Set the line that lets speed alone stand in: a low-stakes, reversible call, or a quality gap under a fixed cut-off held for two straight quarters.Why: without a real stop line, the audit runs forever and nobody trusts the tool or ever stops checking it.
- Skip the confidence survey. Do not ask analysts how sure they feel about a brief.Why: feeling sure and being right are not the same thing, and a fast, clean brief is built to make people feel sure.
How to answer this, stage by stage
Nobody is grading whether you can say "measure the outcome." They are grading whether you can build the actual pipeline that measures it, and say what would ever let you stop. Seven moves, in order.
Let's learn
What happens when a team gets faster at deciding and never actually checks whether the deciding got better?
Strand is a feature inside Sorel Capital's research desk. Before a position review, it reads the earnings call transcript, the latest filing, and a stack of sell-side notes, and hands back a one-page brief: the numbers that moved, the guidance, the risks flagged on the call, in order of what matters most.
Before Strand, an analyst spent about five and a half hours reading the source material ahead of each name's review. Sorel's team covered maybe three names a week at that pace. After Strand, the brief takes about twenty-two minutes to read. The team started covering names the same day they reported, sometimes the same hour.
Here is the turn. The 93 percent drop in prep time was real, and it was good. But nobody at Sorel had ever built a number for whether the decisions coming out of that fast brief were actually right. When a reviewer finally went back, months later, and scored 240 of them against what really happened, brief-only decisions averaged 6.1 out of 10 on thesis accuracy. Decisions where the analyst also pulled the primary transcript averaged 8.3.
Strand's brief was not sloppy. It read every transcript, every filing, and it almost always got the headline number right. What it did, quietly and often, was drop the condition sitting next to a guidance number. A CFO says margins will hold flat, assuming a cost pressure from last quarter doesn't come back. Strand's brief keeps "margins expected flat" and loses the "assuming." That clause sits far from the number it qualifies, and a model built to compress a call into six bullets treats it as the part that's safe to cut.
At its worst, this cost more than a bad grade on a rubric. Sorel's position in Ferrand Machine Works, a mid-cap industrial supplier, got sized up the same day Dariusz read a brief that said margins would hold flat. Six weeks later the cost pressure the CFO had actually flagged, a steel surcharge, came back. Ferrand's next quarter showed margin compression of about 180 basis points, exactly the range the CFO had warned about on the call Strand had summarized. The position cost the fund roughly 60 basis points at the portfolio level before it got trimmed.
The choice I would take back is not building Strand. It's that Strand shipped able to compress a guidance number away from the condition attached to it, with nothing checking whether that condition survived the trip. That was invisible while the brief mostly handled routine calls with nothing conditional in them. It stopped being invisible the first time a real hedge sat inside a real number.
What I would leave alone: Strand's summary of a company's cost breakdown or segment mix. Those are just numbers restated, nothing conditional to lose, and a reviewer checking them against the filing takes thirty seconds. Auditing something that never drops a clause is thirty seconds spent on nothing.
The lesson: a metric that only measures how fast someone can act was never going to tell you whether they acted on the truth. You have to go build the second number on purpose, and be honest that it arrives late.
Now here is the same thing as a story
The short version is above. Read on for how ordinary the week looked from inside Sorel, right up until it wasn't.
Dariusz Kavanagh has covered industrials at Sorel Capital for six years. He built a reputation the slow way, by being the analyst who actually finished the 10-K instead of skimming the earnings-call transcript and calling it done. Before Strand, his prep for a name reporting that week looked like a Tuesday spent entirely inside a PDF.
Strand arrived in the winter, and for months it was the best change to his job in years. He'd open the brief at eight, done reading by eight-twenty-two, and be in a position review by nine with time to actually think about sizing instead of just catching up on facts. He still pulled the full transcript on names he was worried about. He just stopped needing to on the routine ones.
Then routine started meaning something different, without anyone deciding it should.
Ferrand Machine Works reported on a Thursday. The CFO spent about forty seconds on margins, said they'd hold roughly flat in the back half, then added, almost as an aside, that this assumed a steel input surcharge from the spring wouldn't recur, and that if it did, they'd see another 150 to 200 basis points of compression. Strand's brief carried one bullet: "Guidance: margins expected flat in H2." Clean. Confident. Missing the second half of the sentence.
Dariusz read it, agreed with the desk that flat margins meant the stock was cheap at its current multiple, and added to the position before lunch. Under an hour, start to finish. By the old process that decision would have taken most of a day, with the transcript open the whole time and that "assuming" landing right where the CFO put it.
Six weeks later the surcharge came back. Ferrand's next print showed the 180 basis points of compression the CFO had actually described on the call, word for word, in the part of the sentence the brief never carried. The position had already cost the fund about 60 basis points at the portfolio level by the time it got trimmed.
Nobody blamed Dariusz. He'd read the brief the way the brief was built to be read, confidently, once, and moved on. The old decision that mattered was made two years earlier, when Strand shipped with no rule requiring a conditional clause to travel with the number it qualified. That was fine right up until the first time a hedge sat inside a headline number on a name someone actually sized.
What Sorel did next: built the outcome pipeline, the compression audit, and a rule that any bullet built off a guidance number now carries the sentence it came from, in full, one click away. The next time a CFO added a condition to a number, the brief kept it. Dariusz's own gap between brief-assisted and full-read decisions, tracked from that point forward, closed from 2.2 points to 0.4 within three quarters.
What I'd tell myself, before any of this: a number that only measures the part of the trade you already like was never going to warn you. You have to go build the other one, and admit up front that it's going to be slow and a little noisy, because that's the price of it actually being true.
PICK, for tradeoff questions
This is a tradeoff question dressed as a metric question, speed versus quality, so PICK fits better than LEAD. LEAD is for finding the earliest signal of a known outcome. Here the real work is committing to which number to trust when they disagree.
One thing worth being plain about, since this is where the AI-specific judgment actually lives. The alternative most teams reach for first is asking analysts how confident they feel in a brief, a quick pulse survey after each read. That got rejected on purpose: confidence and correctness are not the same measurement, and a clean, well-formatted brief is exactly the kind of output that makes a reader feel more certain than the underlying reasoning earns, which is the specific trap here, not a generic trust problem. The bar for the compression audit is calibrated, not a hard rule: a brief passes when it preserves the conditional attached to a quantified guidance number on a sampled eval set of past calls, checked most of the time, not promised to be perfect every time. The failure mode worth naming directly is compression loss, a summarization model dropping a low-salience clause next to a high-salience number, and the guardrail is the quarterly audit plus the citation-back requirement, not a person eyeballing the brief once and moving on. The trade is real: building the outcome pipeline costs analyst and data time, and it takes a full quarter or two before you can honestly say the tool's decisions hold up, against a same-day dashboard of time saved that could ship on day one.
And if you want to be sure it really works, try it somewhere else
Same four letters, a hospital radiology desk instead of a research desk, so the method proves itself instead of repeating a story I happened to prepare.
Cedarbrook Radiology uses Talbryn, an AI tool that pre-reads chest CT scans and flags the ones most likely to need a second look, with a short written summary of what it found. Toma Iskander is one of six radiologists who sign off on flagged cases each week.
P, position. Measure decision quality by what the follow-up imaging or biopsy actually showed, weeks later, never by how fast a radiologist clears the queue.
I, impact. Speed-only tracking rewards clearing the queue fast on a wrong call, felt by a patient who gets told it's nothing. Outcome-only tracking means waiting weeks for a biopsy result, confounded by everything else about that patient's case.
C, cost asymmetry. A missed finding compounds, the patient doesn't come back until it's worse. A slow outcome check just delays confidence in the tool, it costs nothing by itself.
K, kill criteria. Speed alone is fine for the first-pass triage sort, deciding which scans get read today versus tomorrow, fully reversible, nobody's diagnosis rests on it. It's never fine for the sign-off itself.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pick: tie the metric to the outcome of the actual decision, not to how fast the decision got made, then name the kill criteria in one line.
Cost: the outcome pipeline needs a data engineer for two quarters, not two weeks. Don't ship the speed-only dashboard as a stand-in and call it "close enough." Hold position-sizing decisions to the old, slower process until the pipeline exists.
The model got better, for real: say Strand's summarization genuinely improves. That still isn't proof it stopped dropping conditionals specifically. A better model finds new ways to sound confident about the wrong thing, it doesn't retire the need to check.
Where people run it wrong.
They treat a falling time-to-decision number as proof the tool is working, instead of asking what it would look like if the tool were making people fast and wrong.
They build the outcome check once, as a one-time study, instead of a running pipeline that keeps sampling every quarter.
They set no kill criteria at all, so the audit either runs forever or gets quietly dropped the first time someone complains it's slowing things down.
How to use it live. Open with the split before naming a fix: "speed tells you the brief is short, it can't tell you the brief is right, so I'd measure the second thing on purpose." That buys you the room to give the real answer instead of reaching for "we'd track engagement with the feature," which is speed wearing a different name.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't the reviewer scoring the rubric just be biased toward finding problems?" Response: that's why the reviewer is never the analyst who made the original call, and the rubric is checked against the transcript, not against their own judgment of what a good decision looks like.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?