What are the failure modes of a research PM who has no product surface to ship into?
Foldout is Errata Sciences' tool for reading science: give it a stack of new papers, get back the two or three sentences a researcher actually needs. Amabel Abernathy owns Foldout's oncology roadmap. Thibault Cyprian led a research pod that was funded, fourteen months earlier, to build the one thing her team was missing. By the time anyone checked, the pod's own numbers had never looked better, and Amabel's team still had nothing.
- Pin every open research track to one named applied team and a real evaluation set tied to their actual need.Why: without a real destination, "progress" is just a benchmark number nobody outside the pod can check.
- Put a quarterly checkpoint on the calendar, not on whenever a budget review happens to ask.Why: nobody outside Thibault's own reporting line had standing to ask until the QBR forced it, fourteen months in.
- Ask the applied team directly whether they can name one thing that would change on their roadmap if the research succeeded.Why: if no team can answer that, the drift has already happened, whether anyone's noticed yet or not.
- Require every reported research number to state which eval set produced it.Why: a score climbing on paper hid the fact that oncology's real 36 cases hadn't been checked in eight months.
- Leave the first six to eight weeks of a genuinely new idea unpinned.Why: real early exploration needs room to not know yet what it's for; checkpointing it that early kills good bets the same way it kills real drift.
How to answer this, stage by stage
Nobody is grading whether you can define GUARD. They're grading whether you can name, plainly, what actually goes wrong when a smart person's work has nowhere real to land, and what you'd do before it costs a year.
Let's learn
Foldout reads a new science paper and hands a researcher the two or three sentences that actually matter, so they don't have to read the whole thing to know if it's worth their time.
Before Foldout, an oncology researcher spent about nine hours a week just reading trial papers to stay current. Foldout cut that to about two hours a week. That part of the product works, and it has worked for two years straight.
One task never got any easier. Catching two trial papers that quietly disagree, one says a drug combination is safe up to a certain dose, another says the same combination causes real harm below that dose, still takes an oncology researcher about five hours a week, done entirely by hand. The same five hours it took before Foldout ever launched.
Fourteen months ago, Errata Sciences funded a research pod to fix the five-hour problem directly. Thibault Cyprian led it. His original mandate was narrow and real: build a model that flags contradicting dosage thresholds across oncology trial papers, checked against 36 real cases Amabel Abernathy's team hand-labeled with him in the first two months.
Here's the turn. What Thibault's pod actually did for the next twelve months was not lazy, and it was not bad science. It climbed steadily, by every measure the pod itself was tracking: a broader "claim graph" benchmark spanning nine scientific fields, accuracy up from 61 to 89 percent. The problem was never that the research stalled. The problem is that nobody outside the pod was checking whether it was still walking toward oncology's 36 real cases, and it quietly wasn't.
What it costs at its worst: fourteen months, 2.3 million dollars in loaded research cost, and when Errata's quarterly business review finally asked what the pod's work would change for oncology, Amabel had no answer, because nobody had shown her a result against her team's real cases in eight months. Run cold against those same 36 cases, Thibault's current model caught nine. A simple rule-based check the pod had scrapped in month two, for being "not interesting enough," would have caught more.
What I would leave alone: the first six to eight weeks of a genuinely new research idea. Real exploration needs room to not know yet what it's for, and checkpointing it that early kills good bets at the same rate as real drift.
The lesson: a research pod doesn't usually fail because the science was too hard. It fails because nobody wrote down who was supposed to ask, on a fixed date, whether it was still walking toward something real, so nobody did, for over a year.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one slower, for the version that explains why nobody caught it for over a year.
Thibault Cyprian can tell you, without checking a slide, which trial papers in oncology's queue are worth a second read. He built that instinct over six years as a computational biologist before Errata Sciences hired him to lead research, and it was real from the start: in his first month on the pod, he sat with Amabel Abernathy's team for two full days, reading real trial papers with them, until he understood exactly what a dosage-conflict contradiction actually looks like on the page.
Amabel's team built the eval set with him. Thirty-six real pairs of oncology trial papers, hand-labeled, each one a genuine case where two papers disagreed about how much of a drug combination a patient could safely receive. It was slow, careful work, and by month two, Thibault's first model caught eleven of the thirty-six. Good enough to get excited about. Not good enough to ship.
For a while, that was the whole story, and it was a good one. Thibault ran his weekly check against the 36 cases every Friday, same time, same set, and posted the number in Amabel's channel without being asked. Eleven became fourteen became eighteen. Amabel started sketching, on her own time, what a "contradiction flag" feature in Foldout might actually look like on a researcher's screen.
Then a paper Thibault's team was reading for unrelated reasons pointed at something bigger. Not a fix for the 36 cases specifically, but a general architecture, a graph over claims pulled from any paper in any field, that could in theory catch this kind of contradiction everywhere, not just in oncology. It was a genuinely interesting idea. Errata's VP of Research thought so too.
The habit thinned in three moves, none of them dramatic. First, the Friday post against the 36 cases slipped to every other week, because the team was heads-down building the graph architecture and "the number hadn't really changed anyway." Then it slipped to monthly. Then, sometime around month seven, it just stopped, without anyone deciding to stop it. What replaced it was a different number, posted just as regularly: accuracy on a new benchmark the pod had built itself, 2,400 claim pairs pulled from nine scientific fields, climbing every month. Sixty-one percent. Seventy-four. Eighty-two. Eighty-nine.
Then came the quarterly business review, month fourteen, a Tuesday like any other. Someone from finance asked the question that gets asked at every QBR eventually: what has the research org's spend actually produced. Amabel was in the room. Someone turned to her and asked, plainly, what Thibault's pod's work would change for Foldout's oncology customers if it shipped tomorrow.
She said she didn't know. Nobody had shown her a result against her team's real cases in eight months.
That was the whole trigger. One direct question, in a room that wasn't trying to catch anyone out.
Thibault went back that afternoon and ran his current model, the one scoring 89 percent on the broad benchmark, cold, against the original 36 oncology cases nobody had checked in eight months. It caught nine. A simple rule-based heuristic the team had scrapped in month two, for being "not interesting enough" to keep building on, would have caught eleven, on its own, with no graph architecture at all.
The real cost was never the twelve months of research. It was the eight months where Amabel's team kept checking dosage conflicts by hand, five hours a week, because nobody told them the thing that was supposed to replace that work had quietly stopped aiming at it.
The decision I would take back sits in a meeting fourteen months earlier, the one where the pod got its charter. Errata's leadership offered Thibault two versions of the same job. Embed with Amabel's team, chartered around the 36 real cases, slower, narrower, harder to put in a board deck. Or run the pod as a "frontier" bet, chartered to advance claim synthesis broadly across the sciences, the bigger, more fundable idea, with a paper-worthy architecture attached. Thibault picked the second one, and honestly, so would most people in that room. A pod built around 36 hand-labeled examples sounds small next to "advance the science." Nobody in that meeting was wrong to think so. Nobody ever came back to check whether that was still true once the need got more pressing, once a new class of combination therapies started entering oncology trials with even more room for exactly this kind of contradiction.
Run the same fourteen months again, with one change: a checkpoint on the calendar every quarter, not tied to anyone remembering, where the pod has to show Amabel's team a result against the 36 real cases or explain why it can't. At the month-four checkpoint, the graph architecture's overlap with the oncology set has already fallen from 100 percent to 70 percent, enough to notice, not enough to panic. Leadership doesn't kill the broader idea. It just pins half the pod's remaining cycles back to the 36 cases from that point on. By month fourteen, instead of a QBR nobody can answer, Foldout ships a contradiction-flag feature that catches thirty-one of the thirty-six known cases, and Amabel's team's manual check drops from five hours a week to about forty minutes reviewing what got flagged.
One design let "interesting" stand in for "useful" for fourteen months. The other asked the useful question every three months, and only let interesting keep the runway it could still justify.
What I'd tell myself, back in that first meeting: a research pod doesn't fail because nobody can build the science. It fails because we never wrote down who was supposed to ask if it worked, so for over a year, nobody did.
GUARD, on a research pod with nowhere to ship
Not a way to prove Thibault did something wrong. GUARD is what forces you to say who actually pays, who could have raised a flag, and what you'd build so the next pod doesn't take fourteen months to find out.
The recap, one line per letter: three groups carry this cost, not one. The harm lands hardest on the specific team the research was originally promised to. Nobody had standing to ask until a budget review forced it. The fix is a named team and a real eval set, checked on a calendar. And the signal that would have caught it, weekly overlap with a real eval set, existed the whole time, for free.
Three things worth saying plainly, since this is where the real judgment sits. Errata's leadership considered capping every research track's runway at a fixed six months instead of tying it to a named applied team, and rejected it: a hard time cap can't tell genuine early-stage exploration from real drift, since both look identical in month two, while an applied-team checkpoint asks a question about direction, not duration. The AI-specific failure worth naming by name is silent evaluation drift: the pod's own reported accuracy kept climbing because it was being measured against a benchmark that had quietly stopped being the applied team's real eval set, so the number looked healthiest exactly when it was telling the applied team the least. The guardrail is requiring every research update to state which eval set produced its numbers, checked by the named applied team at each checkpoint. And the trade-off is real, accepted on purpose: pinning the pod back to oncology's 36 cases narrows the model, it may not transfer cleanly to cardiology's version of the same contradiction problem, and that's a smaller, real win worth taking over a broader, unproven one.
And if you want to be sure it really works, try it somewhere else
Same five letters, a different applied team, and this time the thing nobody separates is a model getting genuinely more impressive from a model getting any closer to something a lawyer could actually use.
Calque, a legal-document translation company, is used by immigration attorneys to translate visa-sponsorship agreements. Farrokh Doryan leads a research pod there, chartered eleven months ago around a real, narrow need: catching the exact clauses where two language versions of the same sponsorship contract quietly say different things, checked against 40 real cases the immigration team had hand-labeled with him.
Farrokh's pod drifted the same way Thibault's did, toward a broader, more publishable idea: a general "idiom and register preservation" model, meant to keep tone and meaning intact across twelve languages, not just legal clauses in two. It's a genuinely harder, more interesting problem, and the pod's own benchmark score climbed from 58 to 84 percent over eleven months. Checked against the immigration team's original 40 contract-clause pairs, the pod's contradiction-catch rate never moved. It sat at three of forty the whole time, exactly where it started, still fully manual for every attorney using Calque.
Mapped onto GUARD: the groups are Farrokh, the immigration applied team, and Calque itself, footing an eleven-month research bill. The harm lands hardest on the immigration team specifically, the one whose real, named need the pod was chartered around before drifting. Nobody outside Farrokh's reporting line had standing to ask, since the immigration team doesn't sit in research's chain of command. The fix: pin the pod back to the 40 real clause pairs, with a quarterly checkpoint owned jointly by research and the immigration team. The detect signal: ask the immigration team whether they can name one clause type they'd trust Calque to flag if the research shipped tomorrow, before that answer was no.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: pin every open research track to a named applied team and a real eval set, checked on a calendar, not a budget review.
Cost: no budget for a formal checkpoint process this quarter. Borrow the applied team's own weekly roadmap sync, fifteen minutes, ask one question: "what would change here if the research succeeded?"
The model got better, for real: say Farrokh's idiom-preservation model hits 95 percent next release. Attorneys still need the exact clause flagged, because sounding more natural across twelve languages and catching a contradictory sponsorship clause are two different problems, and only one of them protects a client.
Where people run it wrong.
They let "the research is going well," measured by the pod's own chosen benchmark, stand in for "the research is going somewhere," without ever asking an applied team to confirm it.
They wait for the budget review to be the first real check, instead of putting one on the calendar from day one.
They kill all open-ended research the moment it doesn't fit a roadmap yet, punishing genuine early exploration the same as genuine drift, when only a named-team-and-checkpoint tells the two apart.
How to use it live. When an interviewer asks about research-PM risk, ask one thing back before answering: "which applied team is this research supposed to eventually reach, and what would they do differently if it succeeded?" That question alone is usually the exact judgment a GUARD question about research is listening for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the broad benchmark really was the better long-term bet?" Response: then the quarterly checkpoint is exactly where that case gets made, to the applied team, with real numbers, instead of assumed for fourteen months with no one checking either way.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on AI PM role variants: platform, applied, infra, research
- #1 Describe the difference between an applied AI PM and a platform AI PM in terms of who their customer is.
- #2 What does an AI infrastructure PM own that an applied AI PM does not?
- #3 How does success get measured differently for a research-adjacent PM versus an applied PM?
- #4 Give an example roadmap item for a model platform PM and explain why it would never appear on an applied roadmap.
- #5 Which role variant would you assign to owning the internal prompt library, and why?
- #6 An AI platform PM's users are internal engineers. How does that change discovery?