Why does the AI PM role pull the PM further into the technical stack than most PM roles?
Deadair is Hushframe's auto-cut editor. Feed it a raw podcast or video recording and it cuts the dead air, the filler words, and the bad takes on its own. Yejide Wrentham owns Deadair's product calls. Bogumil Cavanagh, three weeks into the job, just asked her a question about one clip that she cannot answer. This is what changes about the PM job once the "how" stops being engineering's alone to decide.
- Own the threshold yourself, don't just sign off on it.Why: "cut filler words with high confidence" isn't a real decision until someone can say what number that confidence is, and defend it.
- Sit in on the model's eval review as a standing habit, not a launch-week visit.Why: the threshold that shipped clean can quietly stop fitting the product without a single line of code changing.
- Split the number by the real shape of the input instead of forcing one setting to cover every case.Why: the same 0.71 threshold was right for one calm voice and wrong for three people talking over each other.
- Track the error rate by segment every week, never as one blended average.Why: interview shows climbed from 1.8 to 6.4 percent over eight weeks while the blended number stayed boring the whole time.
- Know where to leave engineering alone.Why: chasing technical depth into every decision is its own failure; a choice that never changes what the product does isn't a product decision.
How to answer this, stage by stage
Nobody's grading whether an AI PM can write code. They're grading whether you'll notice the moment a number quietly became the whole product spec, and go learn it before someone asks.
Let's learn
Here is what happens when a single number, chosen once and never looked at again, quietly becomes the entire editorial judgment of a product.
Deadair is Hushframe's auto-cut editor. You give it a raw recording, a podcast or a video, and it cuts the silences, the ums and likes, and the bad retakes on its own, so a creator doesn't have to scrub through two hours of footage by hand.
When the filler-word feature shipped, engineering picked a confidence threshold of 0.71 after their own internal testing. On solo, scripted shows, one calm voice reading from notes, that number worked well. The false-cut rate, real words wrongly cut as if they were filler, sat at about 1.2 percent. Creators stopped checking Deadair's cuts before publishing. Yejide wrote one line into the PRD: "remove filler words the model is highly confident about," and left the actual number to engineering. That was normal. She'd done it on every PM job before this one.
Then Hushframe pushed Deadair into interview shows: two or three people talking over each other, real accents, real crosstalk. Same threshold, 0.71. On those shows, the false-cut rate crept from 1.8 percent to 6.4 percent over eight weeks. A guest would say "actually, that's a great point," and Deadair would cut "actually" clean out of the sentence, while leaving three genuine ums sitting right next to it.
The extra mistakes were never the real problem. A 6.4 percent false-cut rate is annoying but fixable. The real problem showed up when Bogumil Cavanagh, three weeks into the job, watched that exact clip and asked Yejide why the tool kept the ums and cut the real word. She didn't know. She had never opened the eval dashboard, never seen the threshold, never asked what "highly confident" actually meant as a number. She had delegated the entire question the day she wrote that PRD line, and never noticed she'd done it.
At its worst, this costs more than one awkward silence in a meeting. It means the product's actual editorial voice, what Deadair is willing to cut and what it protects, was never decided by anyone who owns the product. It was decided by whatever number engineering's early tests happened to land on, for a kind of show that made up half the catalog a year later. When a customer asks why the tool mangled their guest's sentence, there's no PM answer waiting. There's only "let me go check with engineering," which is the exact sentence a PM's job exists to make unnecessary.
What I would leave alone: Deadair's audio codec, the format it compresses a file into for cloud storage, is genuinely just an engineering call. It has zero effect on what gets kept or cut. I wouldn't sit in on that meeting, and I wouldn't want to. Getting closer to the stack means picking the parts that are actually product decisions in disguise, not treating every technical choice as one.
The lesson: a decision that looks like an implementation detail is only safe to hand off completely if it can never change what the product actually does. A confidence threshold decides what gets cut. That was never engineering's call alone to make. It just took a new hire's honest question to notice nobody outside engineering had ever looked at the number.
Now here is the same thing as a story
The short version is above, for when you're in the room. This one is for feeling why a spec line with no number in it is really just a promise nobody wrote down.
The eval dashboard has a slider nobody outside engineering had ever dragged, not once, in the two years since Deadair shipped its filler-word feature.
Yejide Wrentham has run product at Hushframe for four years. She came from a project-management tool before this, and she was good at the part of the job most PMs are good at: clean PRDs, tight acceptance criteria, engineers who trusted her because she never second-guessed their "how." Tell Yejide the outcome you want, and by Friday there's a spec on it, numbered and dated.
Deadair's filler-word feature shipped in her second year. Ruxandra Sterrenberg's team built it: a small model that listens for ums, likes, and false starts, scores each one, and cuts anything above a confidence line they set at 0.71. Yejide's whole PRD for it was one sentence: remove filler words the model is highly confident about. She never asked what 0.71 meant, and nobody expected her to. It shipped clean. Solo, scripted shows sounded better than any human editor could manage in the time. For months, the false-cut rate sat under one and a half percent, and Yejide stopped thinking about the number at all, because there was nothing pulling her back to it.
Then Hushframe went after interview shows. Two hosts, a guest on a bad mic, real crosstalk. Nobody touched the threshold, because nobody had a reason to think they should. The false-cut rate on those shows started climbing, quietly, the kind of climb you only see if you're looking at the right slice of the data instead of the average. Nobody was.
The trigger wasn't a catastrophe. It was Bogumil Cavanagh, three weeks into the job, sitting next to Yejide reviewing a flagged clip. A guest says, "actually, that's a great point," and Deadair cuts "actually" clean out, mid-sentence, while three real ums from the same guest sit untouched two lines later. Bogumil asks the obvious question: why did it keep the ums and cut the real word? They don't sound that different to me.
Yejide opened her mouth to answer and had nothing. Not a vague answer. Nothing. She didn't know the threshold. She didn't know why 0.71 instead of 0.6 or 0.8. She didn't know the model scored "actually" at 0.74 and those particular ums at 0.68, quiet and mumbled in that guest's voice, just under the line. She had been treating a live product decision as a settled implementation detail for two years.
I want to say the problem was the threshold. The threshold was fine, once, for the show it was tuned on. The real story is that Yejide had a switch, not a dial.
There was no setting where she checked in on it occasionally. Once she wrote that PRD line and it worked, the switch flipped to handed off, for good, until a new hire's honest question flipped it back.
Here's the decision I'd take back. Months before, in the meeting where the filler-word spec got signed off, Yejide wrote "remove filler words the model is highly confident about" and moved on to the next line item. Nobody argued. It read like a clean, decisive spec. Ruxandra's team picked 0.71 after their own tests, on their own show, and shipped it. That was reasonable, then, when Deadair only did one kind of show and the number quietly matched the audio it was built for.
I'd put the number in the spec. Not as decoration, as the actual decision: 0.71, reviewed against a real eval set, with Yejide's name on why. And I'd write in a standing rule: any time the show mix changes meaningfully, she sits down with Ruxandra and walks fifty disputed clips before the threshold ships unchanged.
Run the same Tuesday again, with that rule already in place. Bogumil asks his question. Yejide answers it in one breath: the interview threshold is 0.84, split six weeks ago after walking fifty clips together, one afternoon, and the false-cut rate came down from 6.4 to 1.9 percent within two release cycles. No trip to engineering. No awkward pause. Just an answer, because she'd already done the work of having one.
One version of that Tuesday ends in a shrug and a promise to follow up. The other ends in three sentences and moves on to the next clip.
What I'd tell my past self, the one who wrote a spec with no number in it and called it done: if a line in your PRD would need a different sentence for every possible number someone could put there, you haven't written a decision. You've written a question and hoped somebody else would answer it before it mattered.
FLIPS, or the line item that was never actually a decision
Not a way to make "getting technical" sound like a virtue on its own. FLIPS is what forces you to notice which technical number is secretly the product's whole editorial judgment, and put your name on it.
Three things worth being direct about, since this is where the real judgment sits. The AI-specific failure here is a quiet distribution shift: the confidence threshold was tuned once, on one kind of audio, and nothing about the number itself changes when the input distribution does, so a setting that was safe stays technically unchanged while it gets steadily wrong for a growing share of the catalog. The guardrail is the per-segment eval walk, not a smarter model. We also considered the lazier fix: raise the global threshold everywhere until interview shows behaved. Rejected, because that would have made Deadair miss real filler on every solo show to fix a problem that only lived in one segment, trading a small, contained cost for a bigger, invisible one. And there's a real trade-off in the fix we did ship: a higher threshold on interview shows means Deadair also lets more real filler words through uncut there, so creators on those shows get slightly less automatic cleanup in exchange for far fewer wrongly cut sentences. That's a real cost, not a free upgrade.
And if you want to be sure it really works, try it somewhere else
Same five letters, a completely different kind of document. This time the model isn't cutting a sentence. It's reading a building-permit application and putting its name on the law.
Parapet builds Flagstone, a tool that reads building-permit applications and flags anything that might violate the local code, citing the exact section, before a human reviewer signs off. Ionut Weatherstone runs product for it. For the first year, he sat in on every weekly accuracy review, checking Flagstone's citations against the actual code book himself, because early on the model got sections wrong often enough to matter. Citation accuracy climbed steadily, month over month, until it crossed 97 percent. Ionut stopped attending. Everyone did. The number said it had earned the trust.
F · Ionut Weatherstone, product lead at Parapet, who ran Flagstone's launch and used to personally check every flagged citation against the code book.
L · He stopped cross-checking citations himself once monthly accuracy climbed past 95 percent. He had four other launches competing for his attention, and the number kept saying it was fine.
I · A different family entirely: the over-trust flip. Old setting: spot-checks citations sometimes, stays skeptical. New setting: stops checking anything at all, because the model got better, not worse. This one fires on good news, and it still has no middle.
P · The team let the manual quarterly re-certification review lapse once accuracy crossed 95 percent, and never built any check tied to whether the underlying code book itself had changed since the model last saw it. A high confidence score was treated as permanent, when the thing it was confident about could be amended out from under it at any time.
S · With a "possibly stale" flag tied to the code's last-amended date, the same near miss gets caught automatically: any citation older than the jurisdiction's most recent amendment gets a visible flag and a required human glance, ten seconds or less. In the quarter before the near miss, three citations would have tripped that flag. All three were stale. All three would have been caught before a human ever saw them.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: with a model in the loop, "how" is the spec, and a rising confidence number is not the same claim as "the thing it's confident about is still true."
Cost: no budget this quarter to build the staleness flag. Put the quarterly re-certification review back on the calendar as a person's real job, not a nice-to-have, until the automated version exists.
The model got better, for real: say Flagstone's citation accuracy climbs to 99 percent next year. Still not a reason to stop checking whether the underlying code moved. A better model answers "is this usually right" better. It says nothing new about "is this specific citation still true today."
Where people run it wrong.
They treat a rising accuracy number as permission to stop watching entirely, instead of watching a narrower, different thing.
They build the confidence score once and never ask what it's actually confident about: the citation's wording, or its continued existence in the current code.
They let the review cadence quietly become "whenever someone remembers," instead of a standing habit with a name attached to it.
How to use it live. Ask one question before trusting any "the model's gotten really good" claim: good at what, exactly, and does that thing change underneath it without the model knowing? A model can be excellent at reading a document and still know nothing about whether the document it read is still current.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't walking fifty clips herself just micromanaging engineering?" Response: no, because she's not reviewing their code or their model's architecture. She's reviewing the one number that decides what the product is allowed to cut, which is a product decision wearing an engineering costume.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on AI PM vs traditional PM vs technical PM
- #1 List four responsibilities an AI PM holds that a traditional PM does not.
- #2 Which parts of the classic PM toolkit transfer unchanged to AI products, and which do not?
- #3 Explain why an AI PM often owns the evaluation set while a traditional PM would not own a test plan.
- #4 How does the discovery phase differ when feasibility is genuinely unknown until you build?
- #5 Describe the difference between an AI PM and an ML PM at a company that has both.
- #7 A traditional PM writes user stories. What is the AI equivalent artifact and why?