How does success get measured differently for a research-adjacent PM versus an applied PM?
Thrumcast builds SpliceCast, the tool that swaps in a different ad for each listener inside the same podcast episode, matched while the file streams. Cedric Loomis owns SpliceCast's live targeting. Zoraida Krauss owns Undertone, a research bet on a model that reads an episode's raw audio, not just its transcript, to catch when an ad's tone clashes with the moment it lands in. Four quarters in, Octavian Duskwater, who runs product for both, put their two OKRs on the same slide for the first time.
- Grade the research-adjacent PM against a stated, reproducible threshold, not a shipped number that does not exist yet.Why: get this backwards and the whole review room ends up comparing two different kinds of number as if they were one.
- Require the threshold to hold on every real case the bet needs to work on, not one blended average.Why: an average can clear the bar while two out of three real cases sit far below it.
- Once the threshold clears and holds, force the bet through a live test inside the real product, not another lab report.Why: a number that agrees with human raters still might not predict what a real listener actually does.
- Kill or shrink the bet if the live test breaks the lab number, or the threshold sits flat for two quarters straight.Why: "still researching" is not a status. Left unchecked, it is a decision nobody ever makes, on purpose.
- Keep the two PMs on separate clocks in the same review, and say so out loud.Why: one number moves every week, the other moves every quarter. Putting them on one slide with no label is what starts the confusion.
- Leave small, cheap exploratory work out of this whole system.Why: a two-week spike does not need a threshold or a live test. That machine is for a bet already asking for a team and real budget.
How to answer this, stage by stage
Nobody is grading whether you can name two job titles. They are grading whether you can say, in numbers, how far each one sits from something a listener actually feels.
Let's learn
SpliceCast is Thrumcast's tool for podcast ads. It swaps in a different ad for each listener inside the same episode, matched while the file streams, so two people hearing the same show can hear two different ads in the same slot.
Before dynamic insertion, one ad got baked into every copy of an episode. Everybody heard it, whether it fit them or not, and about 54 out of 100 listeners skipped past it inside the first few seconds. With SpliceCast matching ads by the words in the transcript and a listener's own profile, completion climbed fast: 74, then 78, then 82, then 85 out of 100, across four straight releases.
At the same time, on a completely different clock, Zoraida's team was chasing its own climbing number. Undertone is a research bet: a model that listens to an episode's real audio, not just its transcript, to catch when an ad's tone clashes with the moment it drops into, an upbeat coffee ad landing right after a segment about a missing child, for example. Its benchmark, how well its tonal-mismatch score agreed with a panel of human listeners, went 0.31, 0.41, 0.51, 0.58 across those same four quarters. A clean climb. By every measure the research team owned, a real and honest win.
Here's the turn. A climbing research number was never the problem. The problem was that everyone in the room started reading it the exact same way they read Cedric's completion rate: as proof something had already gotten better for a real listener. It hadn't. Undertone had never picked a single real ad. Not once.
Octavian asked the obvious next question anyway: what did 0.58 change this quarter, for one listener? Nobody had an answer, so Zoraida split the benchmark by genre. True crime scored 0.70, comfortably over the 0.60 bar. Comedy scored 0.38. Business news scored 0.42. Neither one was close. The eval set that produced the blended 0.58 was 60 percent true-crime examples, because that's where labeled audio already existed from an older transcription project. The whole climb had been carried by the one genre the set was already stacked toward.
What it costs at its worst: comedy and business-news shows make up about 60 percent of SpliceCast's episode volume. Shipping Undertone-driven targeting broadly on the strength of the blended 0.58 would have pushed a barely-tested signal onto most of the catalog, at real risk to completion rate on accounts like Pinerow Coffee's, whose quarterly renewal already runs on completion rate and cost per completed listen, not a research score nobody outside the lab had ever heard of.
What I would leave alone: Cedric's own review stays exactly as it is. A completion rate that's already shipping to real listeners every week doesn't need a threshold or a shadow test bolted on, it's already being tested, every single day, by people deciding whether to skip.
The lesson: a research number that keeps climbing is not proof of anything except that the research number keeps climbing. It only means something once it's been checked against the real thing it claims to predict, on every case it needs to work for, not the one case it happens to already be good at.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one for the four quarters that actually got Octavian to pull both slides side by side.
Zoraida Krauss spent two years building acoustic QA models for a call-center vendor before Thrumcast hired her onto its small applied-research team. She can hear a tonal mismatch before a waveform even finishes rendering on her screen, the half-second where a chipper read lands wrong against whatever just happened in the room. That ear is the whole reason Undertone exists.
Undertone started as a two-person side project fourteen months ago: could a model listen to an episode's actual audio, not just its transcript, and catch the exact thing Zoraida's ear already caught, an ad's tone clashing with the moment it drops into. Nobody at Thrumcast expected much from it. It got a slice of compute, a small labeled set borrowed from an older transcription project, and one line on the company's quarterly OKR slide, right under Cedric Loomis's completion rate for SpliceCast.
For the first few quarters, that line was easy to like. The correlation between Undertone's tonal-mismatch score and a panel of human listeners went 0.31, then 0.41, then 0.51. Every quarter, Zoraida reported it the same way Cedric reported his: current value, target, a green arrow pointing up. Nobody asked to see it any other way, because a green arrow was the entire language that slide spoke.
So the checking thinned. First Zoraida stopped re-verifying the labeled set's genre mix every quarter, since the old transcription project had already been through one review, two years back. Then she stopped asking whether a rising correlation actually meant anything for a real ad decision, because the number kept doing the one thing a research number is supposed to do: go up. By the fourth quarter, when it hit 0.58, the team had started saying "almost there" out loud in meetings, meaning almost ready to matter, without anyone quite deciding what "there" was supposed to be.
Then came the QBR. Nothing dramatic. Octavian Duskwater, who runs product for both lines, was building one slide with both OKRs on it for the first time, mostly because a board member had asked for "the AI roadmap in one page." He put Cedric's completion rate, 85 percent, next to Zoraida's correlation, 0.58, and asked the question anyone building that slide eventually asks: "What did this one actually change, this quarter, for a listener?"
Cedric had an answer in one sentence. Zoraida did not.
She went back to her desk and did the thing she'd stopped doing: pulled the labeled set apart by genre. Sixty percent of it, twelve hundred of two thousand examples, was true crime, because that was the genre with existing labeled audio sitting around from the old project. Split by genre, true crime scored 0.70, clean, comfortably past the 0.60 bar the team had quietly been aiming for. Comedy scored 0.38. Business news scored 0.42. The number that had been climbing for four straight quarters had never really been climbing everywhere. It had been climbing in the one place it already worked, and that one place was carrying the whole average.
Zoraida rebalanced the set to an even four hundred examples a genre and reran the same model overnight. The fair blended number came back at 0.50, not 0.58. Eight points of the quarter's progress had never been about Undertone getting better. It had been about which cases got counted.
The decision Octavian would take back sits in a five-minute conversation fourteen months earlier, when Undertone first got a line on the OKR slide. Someone asked whether a two-person research bet needed its own reporting format, separate from shipped product lines. The answer, reasonable at the time: no, keep one template, don't make leadership learn two formats for a project that small. It made sense when Undertone was small enough that nobody was really watching that line anyway. It stopped making sense the day a board member asked for the AI roadmap on one page, and a lab number and a real one sat there looking identical.
Run the same fourteen months again, with the rule this answer argues for already in place. Undertone's line on the OKR slide carries a small label, "research, threshold-gated," next to Cedric's "shipped." The correlation still climbs the same way, 0.31 to 0.58, because that part of the story doesn't change. But nobody calls it "almost there" without a stated bar, and the genre split happens every quarter, not just the one where someone finally asked. True crime clears 0.60 by quarter three. Zoraida converts it into a real shadow test inside SpliceCast two weeks into the next quarter, scored against real skip data on real episodes, not just against human raters. Comedy and business news stay in research, smaller team, no promise of a ship date, until they clear the same bar on their own.
What I'd tell myself, back in that five-minute conversation about the reporting template: a number that climbs is not the same claim as a number that shipped, and the two will always look identical on a slide that doesn't say which one it is. Somebody has to say so, on purpose, every single quarter, or the slide will say it for you, wrong.
LEAD: the four checks a climbing number has to survive
Not a way to prove research is slower than product. LEAD is what forces you to say which number would have told the truth in month three, and what you would have actually done about it.
Three things worth stating directly, since the real judgment sits here. The alternative Zoraida's team considered, and rejected, was simply raising the threshold from 0.60 to 0.75, to be extra safe on a blended number. It lost, because a higher bar on the same skewed set still gets carried by the genre with the most labeled examples, it just delays the same false confidence by another quarter instead of fixing it. The AI-specific failure worth naming by name is eval-set skew posing as model progress: a benchmark can climb honestly, with no one cheating, while the set behind it quietly overrepresents the one case the model already handles well. The guardrail is a rebalanced, even-weighted eval set checked every quarter, plus a live shadow test before any threshold crossing counts as proof. And the trade-off is real, and accepted on purpose: converting only true crime into a real roadmap item, instead of shipping Undertone everywhere the blended number looked ready, means comedy and business-news advertisers keep running on keyword-only targeting for longer, a real cost, paid on purpose, rather than risk completion rate on the two out of three genres the bet had not actually earned yet.
And if you want to be sure it really works, try it somewhere else
Same four letters, a collection truck instead of a recording booth, and this time the thing nobody separates is a route that shipped from a route that's still just a promising number.
Grovehaul runs two products. RouteKeel plans a truck's stops for the day and ships a route every morning, no research needed anymore, just real trucks and real miles. Spillwatch is the research bet: a model reading truck-mounted camera footage, aiming to flag a bin that will overflow before its scheduled pickup, so the truck swings back before it does instead of after a resident calls to complain. Virgil Bramscott owns Spillwatch, and ran into the exact same slide problem Thrumcast did.
RouteKeel's link is short: fewer missed pickups, fewer complaint calls, a contract renewal with the city measured every quarter. Spillwatch's link ran through steps that did not exist yet: an overflow-prediction score had to agree with what a human reviewing the footage would call, had to run fast enough to change a route mid-shift, had to actually cut missed pickups once it touched a real truck. Its benchmark, prediction accuracy against a labeled set of past overflow photos, climbed the same clean way Undertone's did, quarter over quarter, and the labeled set behind it was built almost entirely from residential routes, because that's where Grovehaul had the most photos on file. Split by route type, residential scored well past the bar. Commercial routes, dumpsters behind restaurants and small offices, sat far under it, because a commercial bin fills on a completely different rhythm and the model had barely seen one.
Mapped straight onto LEAD: the link is a truck actually turning around before an overflow, not a prediction score agreeing with old photos. The early signal is accuracy holding on an evenly split set, by route type, not the blended number the team had been reporting. The abuse is the same shape exactly, a labeled set stacked toward the case that was already easy to collect, dressed up as steady progress. The decision: Virgil converted only the residential slice into a live pilot on real trucks, kept commercial routes in research with a smaller team, and stopped reporting Spillwatch's number on the same slide as RouteKeel's until the label said which kind of number it was.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: grade a research bet against a threshold that holds on every real case, not a shipped number, and never call it proven without a live test.
Cost: no budget for a full shadow test yet. Log the research score inside the real product's decision path without acting on it for a few weeks, free, before spending anything on a full pilot.
The model got better, for real: say Undertone's blended correlation hits 0.80 next quarter. Still check it by genre first, because a number that's honestly better on average can still be badly wrong on the one slice you're about to ship it to.
Where people run it wrong.
They read a climbing research number the same way they read a shipped one, and never ask what it's actually changed for a real case yet.
They fix the whole benchmark by raising the bar, instead of fixing the set that's quietly stacked toward the case that was already easy.
They keep a research bet on the same reporting clock as a shipped product, so a lab number and a real one sit on one slide looking interchangeable.
How to use it live. Before answering, ask one question out loud: does this number's climb come from getting better everywhere, or from the set behind it happening to already lean toward where it works? That question alone is usually the whole diagnosis a LEAD question about research is testing for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"If Undertone eventually clears 0.60 on every genre, doesn't that prove research metrics work fine as-is?" Response: it proves the threshold approach works. It doesn't prove the old blended-average approach was ever safe, since it would have let a broad ship decision happen a full quarter earlier, on a number that was still failing two out of three real cases.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on AI PM role variants: platform, applied, infra, research
- #1 Describe the difference between an applied AI PM and a platform AI PM in terms of who their customer is.
- #2 What does an AI infrastructure PM own that an applied AI PM does not?
- #4 Give an example roadmap item for a model platform PM and explain why it would never appear on an applied roadmap.
- #5 Which role variant would you assign to owning the internal prompt library, and why?
- #6 An AI platform PM's users are internal engineers. How does that change discovery?
- #7 Describe the tension between a platform PM's abstraction goals and an applied PM's shipping deadline.