ConceptIntermediateModel Fluency & the AI PM Role / AI PM role variants: platform, applied, infra, research / #3

How does success get measured differently for a research-adjacent PM versus an applied PM?

LEAD · a quarterly review at Thrumcast, where a research bet's climbing number finally met an applied product's real one

Thrumcast builds SpliceCast, the tool that swaps in a different ad for each listener inside the same podcast episode, matched while the file streams. Cedric Loomis owns SpliceCast's live targeting. Zoraida Krauss owns Undertone, a research bet on a model that reads an episode's raw audio, not just its transcript, to catch when an ad's tone clashes with the moment it lands in. Four quarters in, Octavian Duskwater, who runs product for both, put their two OKRs on the same slide for the first time.

The direct answer
Grade an applied PM against a number the product already moves this quarter: completion rate, renewal, revenue. Grade a research-adjacent PM against a stated, reproducible threshold that has to hold on every real case the bet needs to work on, not one blended average, and never call the bet proven until it survives a live test inside the real product. A research number that keeps climbing but never has to clear that bar, or never gets checked against real behavior, is not being measured. It is being protected.
Do this, in order
  1. Grade the research-adjacent PM against a stated, reproducible threshold, not a shipped number that does not exist yet.Why: get this backwards and the whole review room ends up comparing two different kinds of number as if they were one.
  2. Require the threshold to hold on every real case the bet needs to work on, not one blended average.Why: an average can clear the bar while two out of three real cases sit far below it.
  3. Once the threshold clears and holds, force the bet through a live test inside the real product, not another lab report.Why: a number that agrees with human raters still might not predict what a real listener actually does.
  4. Kill or shrink the bet if the live test breaks the lab number, or the threshold sits flat for two quarters straight.Why: "still researching" is not a status. Left unchecked, it is a decision nobody ever makes, on purpose.
  5. Keep the two PMs on separate clocks in the same review, and say so out loud.Why: one number moves every week, the other moves every quarter. Putting them on one slide with no label is what starts the confusion.
  6. Leave small, cheap exploratory work out of this whole system.Why: a two-week spike does not need a threshold or a live test. That machine is for a bet already asking for a team and real budget.

How to answer this, stage by stage

Nobody is grading whether you can name two job titles. They are grading whether you can say, in numbers, how far each one sits from something a listener actually feels.

1
Anchor it in one real review, not two job titles
Say it like this
"Let me make this real. Thrumcast builds SpliceCast, which swaps in a different ad for each listener inside the same episode. Cedric owns SpliceCast's live targeting. Zoraida owns Undertone, a research bet on reading an episode's actual audio instead of just its transcript. Four quarters in, our head of product put both OKRs on the same slide for the first time, and that's the day this question got real."
Why this works
A named product and a real slide keep this out of a debate about two abstract job titles.
2
Say the structure out loud
Say it like this
"I'll run this as LEAD. Link, what real outcome each number is supposed to protect. Early signal, what moves before that outcome does. Abuse, how a number gets used to claim a win it never earned. Decision, what actually changes at each stage."
Why this works
Two seconds of structure tells the room you have a method, not just an opinion about two teams.
3
Reframe the real question
Say it like this
"This isn't really asking who has the harder job. It's asking how far each PM's number sits from something a listener actually feels. Cedric's number is three steps from a listener's ear. Zoraida's is five steps from a listener's ear, and two of those steps don't exist as real product yet."
Why this works
This is the whole answer in miniature. Skip it and the rest sounds like a turf fight over titles.
4
Give the decision, committed
Say it like this
"So here's what I'd actually do. I would not let Zoraida's climbing correlation sit on the same slide as Cedric's completion rate with no label. I'd grade her against a threshold that has to hold on every genre the bet needs to work on, and I would not call it proven until it survives a real shadow test inside SpliceCast."
Why this works
This is the direct answer to the question, said out loud, before a single chart distracts from it.
5
Prove it with the real numbers
Say it like this
"Here's what actually happened. Undertone's correlation with human tonal-fit ratings went 0.31, 0.41, 0.51, 0.58 across four quarters, closing in on our 0.60 bar. Split by genre, true crime scored 0.70. Comedy scored 0.38. Business news scored 0.42. The blended number was carried almost entirely by one genre."
Why this works
A real number, broken open, beats any amount of talk about research being different in the abstract.
6
Name the abuse before the interviewer does
Say it like this
"Two ways that number was lying to us. Our labeled set was 60 percent true crime, because that's where the labeled data already existed from an older project, so the blend leaned toward the genre where Undertone already worked. And every quarter the OKR slide just said 'benchmark climbing, on track,' with nobody asking whether it had ever touched a real ad decision."
Why this works
Naming the failure yourself beats waiting for a follow-up question to expose it.
7
Say what you would leave alone
Say it like this
"I wouldn't touch a small spike. If two researchers want to spend a week checking whether raw audio even carries a usable signal, that doesn't need a threshold or a shadow test. That whole system is for a bet already big enough to be sitting on the same OKR slide as a shipped product."
Why this works
Naming a place you would not add process shows judgment instead of blanket caution.
8
Close on the one line
Say it like this
"So: an applied PM's number gets graded against what shipped this quarter. A research-adjacent PM's number gets graded against a threshold that holds on every real case, proven with a live test, not a lab report. Until Undertone clears that bar for real, on a fair set, it isn't measuring success. It's measuring how convinced we've gotten."
Why this works
Leaves the room with an actual rule, not just a well-told story about one quarter.

Let's learn

SpliceCast is Thrumcast's tool for podcast ads. It swaps in a different ad for each listener inside the same episode, matched while the file streams, so two people hearing the same show can hear two different ads in the same slot.

Hand sketched left to right flow diagram titled How one ad decision moves through SpliceCast. Five connected boxes reading: Episode audio, Keyword scan, Listener match, this box emphasized in violet, Ad slot filled, Ad plays.
Five steps. The third one, matching the ad to the actual listener, is the step SpliceCast is graded on every week.

Before dynamic insertion, one ad got baked into every copy of an episode. Everybody heard it, whether it fit them or not, and about 54 out of 100 listeners skipped past it inside the first few seconds. With SpliceCast matching ads by the words in the transcript and a listener's own profile, completion climbed fast: 74, then 78, then 82, then 85 out of 100, across four straight releases.

Knowledge spark: what's a benchmark threshold? A number a research team decides has to be true before a bet counts as real. Not a guess made after the fact. A line drawn in advance, so nobody gets to move it once they see how close the result landed.

At the same time, on a completely different clock, Zoraida's team was chasing its own climbing number. Undertone is a research bet: a model that listens to an episode's real audio, not just its transcript, to catch when an ad's tone clashes with the moment it drops into, an upbeat coffee ad landing right after a segment about a missing child, for example. Its benchmark, how well its tonal-mismatch score agreed with a panel of human listeners, went 0.31, 0.41, 0.51, 0.58 across those same four quarters. A clean climb. By every measure the research team owned, a real and honest win.

Undertone's benchmark correlation, by quarter
0.8 0.4 0.0 threshold: 0.60 0.31 0.41 0.51 0.58, Q4 Q1 Q2 Q3 Q4
Blended correlation, by quarterThreshold, 0.60
Four straight quarters of real, honest progress. Still two points under the bar the team had quietly been aiming for.

Here's the turn. A climbing research number was never the problem. The problem was that everyone in the room started reading it the exact same way they read Cedric's completion rate: as proof something had already gotten better for a real listener. It hadn't. Undertone had never picked a single real ad. Not once.

We didn't build a model that lied. We built a number that looked exactly like progress, right up until someone asked what it had actually changed.

Octavian asked the obvious next question anyway: what did 0.58 change this quarter, for one listener? Nobody had an answer, so Zoraida split the benchmark by genre. True crime scored 0.70, comfortably over the 0.60 bar. Comedy scored 0.38. Business news scored 0.42. Neither one was close. The eval set that produced the blended 0.58 was 60 percent true-crime examples, because that's where labeled audio already existed from an older transcription project. The whole climb had been carried by the one genre the set was already stacked toward.

Correlation by genre, quarter four, against the same 0.60 threshold
0.8 0.4 0.0 0.60 0.70 True crime 0.38 Comedy 0.42 Business news
Clears the thresholdWell under the threshold
Rebalanced to an even 400 examples a genre, the fair blended number came back at 0.50, not 0.58. Eight points of that climb had never been about the model getting better. They were about which cases got counted.

What it costs at its worst: comedy and business-news shows make up about 60 percent of SpliceCast's episode volume. Shipping Undertone-driven targeting broadly on the strength of the blended 0.58 would have pushed a barely-tested signal onto most of the catalog, at real risk to completion rate on accounts like Pinerow Coffee's, whose quarterly renewal already runs on completion rate and cost per completed listen, not a research score nobody outside the lab had ever heard of.

The choice I would take back Octavian folded Undertone's benchmark into the same quarterly OKR template as SpliceCast's shipped numbers, one line, current value, target, trend arrow, with nothing marking which kind of number it was. That made sense when Undertone was a two-person side bet nobody scrutinized closely. It stopped making sense once budget and headcount started depending on that same line looking good next to Cedric's.

What I would leave alone: Cedric's own review stays exactly as it is. A completion rate that's already shipping to real listeners every week doesn't need a threshold or a shadow test bolted on, it's already being tested, every single day, by people deciding whether to skip.

The lesson: a research number that keeps climbing is not proof of anything except that the research number keeps climbing. It only means something once it's been checked against the real thing it claims to predict, on every case it needs to work for, not the one case it happens to already be good at.

Hand sketched comparison diagram titled Two desks, two clocks. Left panel, a gauge icon labeled Cedric, applied PM, caption checks completion rate every Monday. Right panel, a document icon labeled Zoraida, research-adjacent PM, caption checks Undertone's benchmark once a quarter.
Same company, two different clocks. Putting both numbers on one slide with no label is what made them look like the same kind of thing.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one for the four quarters that actually got Octavian to pull both slides side by side.

Zoraida Krauss spent two years building acoustic QA models for a call-center vendor before Thrumcast hired her onto its small applied-research team. She can hear a tonal mismatch before a waveform even finishes rendering on her screen, the half-second where a chipper read lands wrong against whatever just happened in the room. That ear is the whole reason Undertone exists.

Undertone started as a two-person side project fourteen months ago: could a model listen to an episode's actual audio, not just its transcript, and catch the exact thing Zoraida's ear already caught, an ad's tone clashing with the moment it drops into. Nobody at Thrumcast expected much from it. It got a slice of compute, a small labeled set borrowed from an older transcription project, and one line on the company's quarterly OKR slide, right under Cedric Loomis's completion rate for SpliceCast.

For the first few quarters, that line was easy to like. The correlation between Undertone's tonal-mismatch score and a panel of human listeners went 0.31, then 0.41, then 0.51. Every quarter, Zoraida reported it the same way Cedric reported his: current value, target, a green arrow pointing up. Nobody asked to see it any other way, because a green arrow was the entire language that slide spoke.

So the checking thinned. First Zoraida stopped re-verifying the labeled set's genre mix every quarter, since the old transcription project had already been through one review, two years back. Then she stopped asking whether a rising correlation actually meant anything for a real ad decision, because the number kept doing the one thing a research number is supposed to do: go up. By the fourth quarter, when it hit 0.58, the team had started saying "almost there" out loud in meetings, meaning almost ready to matter, without anyone quite deciding what "there" was supposed to be.

Then came the QBR. Nothing dramatic. Octavian Duskwater, who runs product for both lines, was building one slide with both OKRs on it for the first time, mostly because a board member had asked for "the AI roadmap in one page." He put Cedric's completion rate, 85 percent, next to Zoraida's correlation, 0.58, and asked the question anyone building that slide eventually asks: "What did this one actually change, this quarter, for a listener?"

Cedric had an answer in one sentence. Zoraida did not.

She went back to her desk and did the thing she'd stopped doing: pulled the labeled set apart by genre. Sixty percent of it, twelve hundred of two thousand examples, was true crime, because that was the genre with existing labeled audio sitting around from the old project. Split by genre, true crime scored 0.70, clean, comfortably past the 0.60 bar the team had quietly been aiming for. Comedy scored 0.38. Business news scored 0.42. The number that had been climbing for four straight quarters had never really been climbing everywhere. It had been climbing in the one place it already worked, and that one place was carrying the whole average.

Zoraida rebalanced the set to an even four hundred examples a genre and reran the same model overnight. The fair blended number came back at 0.50, not 0.58. Eight points of the quarter's progress had never been about Undertone getting better. It had been about which cases got counted.

The correlation itself was never fake. It just was never the same kind of number as a completion rate, and the slide that put them side by side never said so.

The decision Octavian would take back sits in a five-minute conversation fourteen months earlier, when Undertone first got a line on the OKR slide. Someone asked whether a two-person research bet needed its own reporting format, separate from shipped product lines. The answer, reasonable at the time: no, keep one template, don't make leadership learn two formats for a project that small. It made sense when Undertone was small enough that nobody was really watching that line anyway. It stopped making sense the day a board member asked for the AI roadmap on one page, and a lab number and a real one sat there looking identical.

Run the same fourteen months again, with the rule this answer argues for already in place. Undertone's line on the OKR slide carries a small label, "research, threshold-gated," next to Cedric's "shipped." The correlation still climbs the same way, 0.31 to 0.58, because that part of the story doesn't change. But nobody calls it "almost there" without a stated bar, and the genre split happens every quarter, not just the one where someone finally asked. True crime clears 0.60 by quarter three. Zoraida converts it into a real shadow test inside SpliceCast two weeks into the next quarter, scored against real skip data on real episodes, not just against human raters. Comedy and business news stay in research, smaller team, no promise of a ship date, until they clear the same bar on their own.

What I'd tell myself, back in that five-minute conversation about the reporting template: a number that climbs is not the same claim as a number that shipped, and the two will always look identical on a slide that doesn't say which one it is. Somebody has to say so, on purpose, every single quarter, or the slide will say it for you, wrong.

LEAD: the four checks a climbing number has to survive

Not a way to prove research is slower than product. LEAD is what forces you to say which number would have told the truth in month three, and what you would have actually done about it.

LLink. What is the real outcome?
Not whether a benchmark number is climbing. The real outcome, for both PMs, is whether a real listener does something different because of the work, this quarter for Cedric, eventually for Zoraida. Cedric's link is short: better targeting, higher completion, an advertiser like Pinerow Coffee renews. Zoraida's link is long and mostly imaginary right now: a tonal-mismatch score has to predict a real skip, has to run inside SpliceCast's actual decision path in time, has to lift completion over keyword-only targeting, has to be worth what an advertiser pays for it. None of those four steps existed yet at 0.58.
Grading a lab benchmark and a shipped number as if they sat the same distance from the outcome is the mistake this whole answer exists to stop.
Hand sketched left to right flow diagram titled The chain from Undertone to a real renewal. Five connected boxes reading: Raw audio, Tonal score, Ad decision (not yet), this box emphasized in red-orange, Skip or stay, Renewal.
The middle box is the whole gap. Undertone's chain has a step that simply does not exist as real product yet.
EEarly signal. What moves first?
For Cedric, the leading signal is close and fast: a dip in completion rate on one ad category shows up inside a week, well before a renewal decision ever gets made. For Zoraida, the honest early signal isn't the blended correlation at all. It's the per-genre correlation holding above 0.60 on a fairly weighted eval set, reproduced quarter over quarter, not once. That's the number that would have told the truth in month three: not 0.58 climbing, but 0.38 sitting flat on comedy the whole time.
This is the answer to the question, in one line. A blended research average does not warn you when it's quietly failing on two out of three real cases. A number checked per case does.
AAbuse. How does it get gamed?
Two ways this number lied without anyone deciding to lie. The labeled set was 60 percent true crime, because that's where old labeled audio already sat around, so the blend leaned toward the genre where Undertone already worked, without a single person choosing that on purpose. And every quarter the OKR slide reported one word, "climbing," standing in for a claim nobody had actually tested: that the model was getting closer to something real.
Every metric can be hit without doing the real work. An unevenly stacked eval set is the quietest way to hit this one, because nobody has to cheat. The deck is already stacked before anyone runs the eval.
Hand sketched comparison scene titled What the number doesn't say. Left panel, a document icon labeled On paper, caption 0.58 correlation, threshold nearly cleared. Right panel, a question mark icon labeled In reality, caption Undertone has never picked a real ad yet.
A number can be honestly measured and still say almost nothing about the thing it's supposed to predict.
DDecision. What actually changes?
Three real changes, not one dashboard tweak. Rebalance the eval set to an even weight across every genre the bet needs to work on, before trusting any blended number again. Gate any "proven" claim on a live shadow test inside SpliceCast's real decision path, scored against real skip data, not just human raters. And convert only the slice that actually clears the bar, true crime, into a real roadmap item, while comedy and business news stay in research at reduced scope until they earn the same conversion.
A metric nobody acts on is decoration. This is what actually ships, what actually gets held back, and who gets to say which is which.
Hand sketched decision tree titled What Zoraida does at each threshold. Root box reads Undertone's quarterly benchmark, branching into three outcomes. Below 0.60, one genre only leads to Keep funding, small team. Clears 0.60, holds across genres leads to Convert, run a live shadow test. Shadow test breaks, or flat two quarters leads to Kill or narrow the bet.
Three doors, one condition each. The blended 0.58 alone never should have opened any of them.

Three things worth stating directly, since the real judgment sits here. The alternative Zoraida's team considered, and rejected, was simply raising the threshold from 0.60 to 0.75, to be extra safe on a blended number. It lost, because a higher bar on the same skewed set still gets carried by the genre with the most labeled examples, it just delays the same false confidence by another quarter instead of fixing it. The AI-specific failure worth naming by name is eval-set skew posing as model progress: a benchmark can climb honestly, with no one cheating, while the set behind it quietly overrepresents the one case the model already handles well. The guardrail is a rebalanced, even-weighted eval set checked every quarter, plus a live shadow test before any threshold crossing counts as proof. And the trade-off is real, and accepted on purpose: converting only true crime into a real roadmap item, instead of shipping Undertone everywhere the blended number looked ready, means comedy and business-news advertisers keep running on keyword-only targeting for longer, a real cost, paid on purpose, rather than risk completion rate on the two out of three genres the bet had not actually earned yet.

And if you want to be sure it really works, try it somewhere else

Same four letters, a collection truck instead of a recording booth, and this time the thing nobody separates is a route that shipped from a route that's still just a promising number.

Grovehaul runs two products. RouteKeel plans a truck's stops for the day and ships a route every morning, no research needed anymore, just real trucks and real miles. Spillwatch is the research bet: a model reading truck-mounted camera footage, aiming to flag a bin that will overflow before its scheduled pickup, so the truck swings back before it does instead of after a resident calls to complain. Virgil Bramscott owns Spillwatch, and ran into the exact same slide problem Thrumcast did.

Hand sketched quadrant diagram titled Same shape, a waste route instead of a podcast ad. X axis, how much real route data exists, from sparse to plenty. Y axis, cost if the call is wrong, from small to a missed pickup. Spillwatch bet sits high and to the left, sparse data and high cost. New collection zone sits mid left. Routine RouteKeel run sits low and to the right, plenty of data and low cost.
Different truck, same shape of danger. The routes that matter most are exactly the ones the model has seen the least.

RouteKeel's link is short: fewer missed pickups, fewer complaint calls, a contract renewal with the city measured every quarter. Spillwatch's link ran through steps that did not exist yet: an overflow-prediction score had to agree with what a human reviewing the footage would call, had to run fast enough to change a route mid-shift, had to actually cut missed pickups once it touched a real truck. Its benchmark, prediction accuracy against a labeled set of past overflow photos, climbed the same clean way Undertone's did, quarter over quarter, and the labeled set behind it was built almost entirely from residential routes, because that's where Grovehaul had the most photos on file. Split by route type, residential scored well past the bar. Commercial routes, dumpsters behind restaurants and small offices, sat far under it, because a commercial bin fills on a completely different rhythm and the model had barely seen one.

The decision Grovehaul would take back Spillwatch's benchmark went on the same weekly ops dashboard as RouteKeel's missed-pickup count, no flag marking which number was shipped and which was still a lab result. It made sense for a small pilot with a handful of cameras. It stopped making sense once the benchmark started deciding how much budget Spillwatch got next quarter.

Mapped straight onto LEAD: the link is a truck actually turning around before an overflow, not a prediction score agreeing with old photos. The early signal is accuracy holding on an evenly split set, by route type, not the blended number the team had been reporting. The abuse is the same shape exactly, a labeled set stacked toward the case that was already easy to collect, dressed up as steady progress. The decision: Virgil converted only the residential slice into a live pilot on real trucks, kept commercial routes in research with a smaller team, and stopped reporting Spillwatch's number on the same slide as RouteKeel's until the label said which kind of number it was.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: grade a research bet against a threshold that holds on every real case, not a shipped number, and never call it proven without a live test.
Cost: no budget for a full shadow test yet. Log the research score inside the real product's decision path without acting on it for a few weeks, free, before spending anything on a full pilot.
The model got better, for real: say Undertone's blended correlation hits 0.80 next quarter. Still check it by genre first, because a number that's honestly better on average can still be badly wrong on the one slice you're about to ship it to.

Where people run it wrong.
They read a climbing research number the same way they read a shipped one, and never ask what it's actually changed for a real case yet.
They fix the whole benchmark by raising the bar, instead of fixing the set that's quietly stacked toward the case that was already easy.
They keep a research bet on the same reporting clock as a shipped product, so a lab number and a real one sit on one slide looking interchangeable.

How to use it live. Before answering, ask one question out loud: does this number's climb come from getting better everywhere, or from the set behind it happening to already lean toward where it works? That question alone is usually the whole diagnosis a LEAD question about research is testing for.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question about how success gets measured differently for a research-adjacent PM versus an applied PM?
Tap to flip
ANSWER
LEAD: link to the real outcome, find the early signal that moves before it does, name how the number gets abused, decide what actually changes at each stage. Built for a bet that hasn't shipped yet, not a fixed spec.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Zoraida Krauss, who owns Undertone, Thrumcast's research bet on reading an episode's raw audio. Cedric Loomis, who owns SpliceCast's shipped live targeting. Octavian Duskwater, who runs product for both.
3 · THE HABIT
What did the team stop doing once Undertone's correlation kept climbing quarter over quarter?
Tap to flip
ANSWER
They stopped re-checking the labeled set's genre mix, and stopped asking whether a rising correlation had ever touched a real ad decision. A climbing number became reassurance enough on its own.
4 · THE GAP
What was hiding inside the blended 0.58, once Zoraida split it by genre?
Tap to flip
ANSWER
True crime scored 0.70, past the 0.60 bar. Comedy scored 0.38. Business news scored 0.42. Two of three real genres sat well under the line the blended number made look almost cleared.
5 · THE OLD DECISION
What decision would Octavian take back?
Tap to flip
ANSWER
Folding Undertone's benchmark into the same quarterly OKR template as SpliceCast's shipped numbers, with no label marking which kind of number it was. It made sense when Undertone was a tiny side bet nobody scrutinized closely.
6 · THE NUMBER
Fill in the blank: Undertone's blended correlation climbed from ___ to ___ across four quarters. Rebalanced evenly by genre, the fair number was actually ___.
Tap to flip
ANSWER
0.31 to 0.58. Rebalanced, the fair blended number was 0.50, eight points lower than what the OKR slide reported.
7 · THE REPLAY
Same four quarters, gated on LEAD's own decision rule this time, what changes?
Tap to flip
ANSWER
The genre split happens every quarter, not just the one where someone asked. True crime clears 0.60 by quarter three and converts into a real shadow test inside SpliceCast. Comedy and business news stay in research, smaller team, no promised ship date.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the equivalent gap?
Tap to flip
ANSWER
Spillwatch, Grovehaul's research bet on predicting bin overflow from truck camera footage, owned by Virgil Bramscott. Its benchmark climbed too, but its labeled set skewed almost entirely toward residential routes, hiding a real gap on commercial ones.

Check yourself Score: 0 / 0

Multiple choice
1. Why can't Zoraida's climbing Undertone correlation be graded the same way as Cedric's climbing completion rate?
  • A. Because research work is inherently harder to measure than product work.
  • B. Because Cedric's number reflects real listeners this quarter, while Zoraida's number is a lab benchmark that has never touched a real ad decision, and the two sit different distances from the actual outcome.
  • C. Because Zoraida's team doesn't have access to real listener data at all.
  • D. Because correlation is a less trustworthy statistic than a percentage.
Show hint
Check stage 3 of the walkthrough, the reframe.
Show answer
B. The difference isn't difficulty, it's distance from the real outcome. Cedric's number already reflects real listener behavior. Zoraida's is still a prediction about what real behavior would be, unproven.
True or false
2. True or false: Zoraida's team deliberately cherry-picked easy examples to make Undertone's correlation look better than it was.
  • True
  • False
Show hint
Look at the Abuse step in the LEAD recap, and where the 60 percent true-crime skew actually came from.
Show answer
False. Nobody chose the skew on purpose. The labeled set was already 60 percent true crime because that's where old labeled audio happened to exist. The number lied without anyone deciding to lie.
Fill in the blank
3. Split by genre at quarter four, true crime scored ___, comedy scored ___, and business news scored ___, against a threshold of ___.
Show hint
Check the bar chart titled "Correlation by genre, quarter four."
Show answer
0.70, 0.38, 0.42, against a threshold of 0.60. One genre cleared the bar easily. Two genres, making up most of the catalog, sat far under it.
Short answer, where it wouldn't matter
4. Name a place in this same product where blending numbers across categories would NOT be hiding a real problem.
Show hint
Think about a slice of the business where every category is already roughly the same, not wildly different like the three podcast genres.
Show answer
Model answer: Ad file delivery success, whether the ad audio itself loaded and played correctly, barely varies by genre at all, it's a technical step, not a judgment call. Blending that number across genres hides nothing, because there's no real difference to hide.
Short answer, apply it yourself
5. Think of an AI product you use yourself. Name one place its team probably reports a research or model number as if it were already a real, shipped outcome.
Show hint
Look for a number that sounds like a win, "accuracy improved," "the model got smarter," with no mention of whether it was ever tested against real behavior.
Show answer
Model answer: A photo-editing app announcing its object-removal model's benchmark score improved, without saying whether real users' edits actually needed fewer manual touch-ups afterward. The benchmark and the thing a user actually feels can move in different directions.
Short answer, work the number
6. If Zoraida had rebalanced the eval set to an even split from quarter one instead of waiting until quarter four, would the blended number still have looked close to ready to ship by quarter four?
Show hint
Think about what an even blend of 0.70, 0.38, and 0.42 actually comes out to, versus the reported 0.58.
Show answer
No, it would have read as about 0.50, not 0.58. A fair, even-weighted blend never gets close to the 0.60 bar on these numbers. The skewed set wasn't just optimistic by a little, it was the entire reason the number looked ready.
Before you close the answer
Why this works
Tests whether you'll grade a research bet by a real, stated bar, or let a climbing number stand in for proof it hasn't earned. Most candidates either flatten research and product into the same kind of metric, or wave the whole question off as "research just takes longer" without saying what would actually prove it's working.
Follow-up traps
"Isn't rebalancing the eval set just moving the goalposts after the fact?" Response: no, the goal never moved, the set behind it was simply uneven before anyone ran the eval. Rebalancing doesn't change the bar, it makes the number that's being checked against the bar honest.

"If Undertone eventually clears 0.60 on every genre, doesn't that prove research metrics work fine as-is?" Response: it proves the threshold approach works. It doesn't prove the old blended-average approach was ever safe, since it would have let a broad ship decision happen a full quarter earlier, on a number that was still failing two out of three real cases.
If pressed
The correlation itself is measured as a Spearman rank correlation between Undertone's predicted tonal-mismatch score and a panel of three human raters per clip, scoring how well an ad's tone fit the moment it landed in on a five-point scale. The 0.60 threshold wasn't picked arbitrarily, it's the same correlation level Thrumcast's applied-science team had separately shown, on an unrelated project, to reliably predict a measurable shift in real skip behavior once a signal actually reached production.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more