ConceptAdvancedModel Fluency & the AI PM Role / AI PM role variants: platform, applied, infra, research / #8

What are the failure modes of a research PM who has no product surface to ship into?

GUARD · a research pod with no applied home at Errata Sciences' paper reader, Foldout

Foldout is Errata Sciences' tool for reading science: give it a stack of new papers, get back the two or three sentences a researcher actually needs. Amabel Abernathy owns Foldout's oncology roadmap. Thibault Cyprian led a research pod that was funded, fourteen months earlier, to build the one thing her team was missing. By the time anyone checked, the pod's own numbers had never looked better, and Amabel's team still had nothing.

The direct answer
Pin every open research track to one named applied team and a real evaluation set tied to their actual need, checked on a fixed calendar. If a track can't point to one applied roadmap it would change, redirect it or fold the person back onto a team that ships, before the drift costs a year and a budget review has to be the thing that finally asks.
Do this, in order
  1. Pin every open research track to one named applied team and a real evaluation set tied to their actual need.Why: without a real destination, "progress" is just a benchmark number nobody outside the pod can check.
  2. Put a quarterly checkpoint on the calendar, not on whenever a budget review happens to ask.Why: nobody outside Thibault's own reporting line had standing to ask until the QBR forced it, fourteen months in.
  3. Ask the applied team directly whether they can name one thing that would change on their roadmap if the research succeeded.Why: if no team can answer that, the drift has already happened, whether anyone's noticed yet or not.
  4. Require every reported research number to state which eval set produced it.Why: a score climbing on paper hid the fact that oncology's real 36 cases hadn't been checked in eight months.
  5. Leave the first six to eight weeks of a genuinely new idea unpinned.Why: real early exploration needs room to not know yet what it's for; checkpointing it that early kills good bets the same way it kills real drift.

How to answer this, stage by stage

Nobody is grading whether you can define GUARD. They're grading whether you can name, plainly, what actually goes wrong when a smart person's work has nowhere real to land, and what you'd do before it costs a year.

1
Scope it to one research PM and one applied team
Say it like this
"Let's ground this. Thibault runs a research pod at Errata Sciences, the company behind Foldout, a tool that reads science papers for researchers. Amabel owns Foldout's oncology roadmap. Fourteen months ago, his pod was supposed to be building toward the one thing her team actually needed."
Why this works
Naming a real research PM and a real applied PM stops the question from staying an abstract debate about funding research.
2
Say your structure out loud
Say it like this
"I'll run this as GUARD. Groups, who actually gets hurt. Unequal, where it lands hardest. Ability to contest, who has no way to raise a flag. Reduce, the actual fix. Detect, how you'd know it's happening before a budget review forces the question."
Why this works
Two seconds of structure signals a method, not a hot take about research culture.
3
Reframe the question in one breath
Say it like this
"This isn't really a question about whether Thibault is smart, or working hard. He's both. It's a question about whether his work has anywhere real to land, and who pays when it doesn't."
Why this works
This is the whole answer in miniature. Skip it and the rest sounds like a complaint about lazy researchers, which is the opposite of what happened.
4
Give the decision, committed
Say it like this
"So here's what I'd actually do. I'd pin every open research track to one named applied team and a real evaluation set tied to their actual need, and put a quarterly checkpoint on the calendar where the pod has to show a credible path back to that eval set, or the track gets redirected."
Why this works
This is the direct answer to the question, said out loud before a single number distracts from it.
5
Name the failure modes plainly, since that's the actual question
Say it like this
"Three failure modes, in order. Thibault loses over a year with nothing shipped, and that shows up on his own record. Amabel's oncology team gets exactly nothing for the need they were promised something for. And Errata funds fourteen months and 2.3 million dollars that nobody can point to a product from."
Why this works
The interviewer asked "what are the failure modes." Naming them by name, out loud, beats any amount of framework talk here.
6
Prove it with the real case, numbers first
Say it like this
"Here's what actually happened. Thibault's pod reported accuracy climbing from 61 to 89 percent, on a broad benchmark they'd built themselves. Run cold against the 36 real oncology cases the pod was originally chartered around, that same model caught nine. A simple rule-based check the team had scrapped in month two, for being 'not interesting enough,' would have caught more."
Why this works
A real number going the wrong direction beats any amount of talk about research culture in the abstract.
7
Say what you'd leave alone, and the trade-off
Say it like this
"I wouldn't put a checkpoint on the first six or eight weeks of a brand-new idea. Real exploration needs room to not know yet what it's for. And the trade-off is real: pinning the pod to oncology's 36 cases makes the model narrower, it may not transfer cleanly to cardiology's version of the same problem, and that's a trade worth taking on purpose."
Why this works
Naming a real trade-off and a place you'd leave alone shows judgment instead of blanket caution.
8
Close on the one line
Say it like this
"So: a research PM doesn't fail because the science was too hard. They fail when nobody wrote down who was supposed to ask if it was working, so nobody did, for fourteen months."
Why this works
Leaves the room with the actual decision, not just a well-told story about a benchmark.

Let's learn

Foldout reads a new science paper and hands a researcher the two or three sentences that actually matter, so they don't have to read the whole thing to know if it's worth their time.

Before Foldout, an oncology researcher spent about nine hours a week just reading trial papers to stay current. Foldout cut that to about two hours a week. That part of the product works, and it has worked for two years straight.

One task never got any easier. Catching two trial papers that quietly disagree, one says a drug combination is safe up to a certain dose, another says the same combination causes real harm below that dose, still takes an oncology researcher about five hours a week, done entirely by hand. The same five hours it took before Foldout ever launched.

Hand sketched numbered icon list titled Three groups holding the bag. Row one, a person icon, Thibault, two years in, no shipped work to show for it. Row two, a document icon, Oncology's team, a real named need, still unserved. Row three, a scale icon, Errata Sciences, 2.3 million spent, nothing to sell.
Nobody set out to hurt any of these three. That's exactly what makes the failure mode worth naming ahead of time.

Fourteen months ago, Errata Sciences funded a research pod to fix the five-hour problem directly. Thibault Cyprian led it. His original mandate was narrow and real: build a model that flags contradicting dosage thresholds across oncology trial papers, checked against 36 real cases Amabel Abernathy's team hand-labeled with him in the first two months.

Hours an oncology researcher spends each week, before Foldout and 14 months in
10h 5h 0h 9h 2h General paper reading 5h 5h Dosage-conflict cross-check
Before14 months later
General reading genuinely got faster. The dosage-conflict check, the exact task the research pod was funded to fix, has not moved at all.

Here's the turn. What Thibault's pod actually did for the next twelve months was not lazy, and it was not bad science. It climbed steadily, by every measure the pod itself was tracking: a broader "claim graph" benchmark spanning nine scientific fields, accuracy up from 61 to 89 percent. The problem was never that the research stalled. The problem is that nobody outside the pod was checking whether it was still walking toward oncology's 36 real cases, and it quietly wasn't.

We didn't lose the research. We lost track of what it was for.
Hand sketched left to right flow diagram titled Where the pod's direction went over 14 months. Five connected boxes reading Month 1, aligned. Month 4, split. Month 7, broad wins. Month 10, untouched. Month 14, QBR, this last box emphasized in rust.
No single month looks like a decision. Read across all five and the drift is the whole story.

What it costs at its worst: fourteen months, 2.3 million dollars in loaded research cost, and when Errata's quarterly business review finally asked what the pod's work would change for oncology, Amabel had no answer, because nobody had shown her a result against her team's real cases in eight months. Run cold against those same 36 cases, Thibault's current model caught nine. A simple rule-based check the pod had scrapped in month two, for being "not interesting enough," would have caught more.

The choice I would take back In the meeting that spun the pod up, Errata's leadership offered Thibault two mandates. Embed with Amabel's team, chartered around the 36 real cases, concrete but modest. Or run as a "frontier" pod chartered to advance claim synthesis broadly, bigger, more publishable, easier to sell to the board as a research bet. They picked the second one. It was a defensible call in the room: a pod built around 36 examples sounds small next to "advance the science." Nobody ever came back to check whether narrow-but-real had quietly become the better bet as the need got more pressing.

What I would leave alone: the first six to eight weeks of a genuinely new research idea. Real exploration needs room to not know yet what it's for, and checkpointing it that early kills good bets at the same rate as real drift.

The lesson: a research pod doesn't usually fail because the science was too hard. It fails because nobody wrote down who was supposed to ask, on a fixed date, whether it was still walking toward something real, so nobody did, for over a year.

Now here is the same thing as a story

The short version above is what you'd actually say out loud. Read this one slower, for the version that explains why nobody caught it for over a year.

Thibault Cyprian can tell you, without checking a slide, which trial papers in oncology's queue are worth a second read. He built that instinct over six years as a computational biologist before Errata Sciences hired him to lead research, and it was real from the start: in his first month on the pod, he sat with Amabel Abernathy's team for two full days, reading real trial papers with them, until he understood exactly what a dosage-conflict contradiction actually looks like on the page.

Amabel's team built the eval set with him. Thirty-six real pairs of oncology trial papers, hand-labeled, each one a genuine case where two papers disagreed about how much of a drug combination a patient could safely receive. It was slow, careful work, and by month two, Thibault's first model caught eleven of the thirty-six. Good enough to get excited about. Not good enough to ship.

For a while, that was the whole story, and it was a good one. Thibault ran his weekly check against the 36 cases every Friday, same time, same set, and posted the number in Amabel's channel without being asked. Eleven became fourteen became eighteen. Amabel started sketching, on her own time, what a "contradiction flag" feature in Foldout might actually look like on a researcher's screen.

Then a paper Thibault's team was reading for unrelated reasons pointed at something bigger. Not a fix for the 36 cases specifically, but a general architecture, a graph over claims pulled from any paper in any field, that could in theory catch this kind of contradiction everywhere, not just in oncology. It was a genuinely interesting idea. Errata's VP of Research thought so too.

The habit thinned in three moves, none of them dramatic. First, the Friday post against the 36 cases slipped to every other week, because the team was heads-down building the graph architecture and "the number hadn't really changed anyway." Then it slipped to monthly. Then, sometime around month seven, it just stopped, without anyone deciding to stop it. What replaced it was a different number, posted just as regularly: accuracy on a new benchmark the pod had built itself, 2,400 claim pairs pulled from nine scientific fields, climbing every month. Sixty-one percent. Seventy-four. Eighty-two. Eighty-nine.

Hand sketched horizontal timeline titled How the weekly oncology check quietly stopped. Four milestones. Both run, caption 36 cases checked weekly. Nod not check, caption benchmark wins, weeks skipped. No one recalls, caption last real oncology run. QBR asks, this milestone emphasized in rust, caption what would ship, no answer.
Nobody lied about any of these numbers. Every one Thibault posted was real. The number that stopped got quietly replaced by a different one.

Then came the quarterly business review, month fourteen, a Tuesday like any other. Someone from finance asked the question that gets asked at every QBR eventually: what has the research org's spend actually produced. Amabel was in the room. Someone turned to her and asked, plainly, what Thibault's pod's work would change for Foldout's oncology customers if it shipped tomorrow.

She said she didn't know. Nobody had shown her a result against her team's real cases in eight months.

Hand sketched comparison diagram titled Same pod, two moments. Left panel, a gauge icon, Month 2, caption Early demo catches 3 real oncology cases, the room is thrilled. Right panel, a question mark icon, Month 14, caption QBR asks what would ship, nobody can answer.
Same pod, same people, twelve months apart. Nothing about Thibault changed. What changed was who was still checking.

That was the whole trigger. One direct question, in a room that wasn't trying to catch anyone out.

Thibault went back that afternoon and ran his current model, the one scoring 89 percent on the broad benchmark, cold, against the original 36 oncology cases nobody had checked in eight months. It caught nine. A simple rule-based heuristic the team had scrapped in month two, for being "not interesting enough" to keep building on, would have caught eleven, on its own, with no graph architecture at all.

The real cost was never the twelve months of research. It was the eight months where Amabel's team kept checking dosage conflicts by hand, five hours a week, because nobody told them the thing that was supposed to replace that work had quietly stopped aiming at it.

The decision I would take back sits in a meeting fourteen months earlier, the one where the pod got its charter. Errata's leadership offered Thibault two versions of the same job. Embed with Amabel's team, chartered around the 36 real cases, slower, narrower, harder to put in a board deck. Or run the pod as a "frontier" bet, chartered to advance claim synthesis broadly across the sciences, the bigger, more fundable idea, with a paper-worthy architecture attached. Thibault picked the second one, and honestly, so would most people in that room. A pod built around 36 hand-labeled examples sounds small next to "advance the science." Nobody in that meeting was wrong to think so. Nobody ever came back to check whether that was still true once the need got more pressing, once a new class of combination therapies started entering oncology trials with even more room for exactly this kind of contradiction.

Run the same fourteen months again, with one change: a checkpoint on the calendar every quarter, not tied to anyone remembering, where the pod has to show Amabel's team a result against the 36 real cases or explain why it can't. At the month-four checkpoint, the graph architecture's overlap with the oncology set has already fallen from 100 percent to 70 percent, enough to notice, not enough to panic. Leadership doesn't kill the broader idea. It just pins half the pod's remaining cycles back to the 36 cases from that point on. By month fourteen, instead of a QBR nobody can answer, Foldout ships a contradiction-flag feature that catches thirty-one of the thirty-six known cases, and Amabel's team's manual check drops from five hours a week to about forty minutes reviewing what got flagged.

One design let "interesting" stand in for "useful" for fourteen months. The other asked the useful question every three months, and only let interesting keep the runway it could still justify.

What I'd tell myself, back in that first meeting: a research pod doesn't fail because nobody can build the science. It fails because we never wrote down who was supposed to ask if it worked, so for over a year, nobody did.

GUARD, on a research pod with nowhere to ship

Not a way to prove Thibault did something wrong. GUARD is what forces you to say who actually pays, who could have raised a flag, and what you'd build so the next pod doesn't take fourteen months to find out.

Hand sketched labeled parts diagram titled GUARD, on a research pod with nowhere to ship. A question mark icon at the center labeled GUARD, with five callouts arranged around it: G, groups affected. U, unequal landing. A, ability to contest. R, reduce, the fix. D, detect, the signal.
Five checks, run on one research pod. Skip any one of them and a real drift can run fourteen months before anyone names it.
GGroups. Who is actually affected?
Three groups, not one. Thibault himself, whose two years at Errata now show no shipped work on his record. Amabel's oncology team, promised a specific fix for a specific, real problem. And Errata Sciences, which funded fourteen months and 2.3 million dollars of work with no product at the end of it.
A risk question about research usually stops at "the company wastes money." That's real, but it's the smallest of the three harms here.
UUnequal. Where does the harm land hardest?
On Amabel's oncology team, specifically, not evenly across every applied team Foldout serves. The pod's original charter was aimed at their dosage-conflict need. When it drifted, it drifted away from the one team that had been told, explicitly, that this was being built for them.
This is the answer to the question, in one line. The failure mode isn't generic waste. It's a promise made to one named team that quietly stopped being kept.
Hand sketched decision tree titled The one decision that decides which failure mode you get. Root box, New research track spins up. Left branch, pinned to one applied team plus a real eval set, leading to Drift caught at the first quarterly check. Right branch, given a broad general mandate instead, leading to Drift goes uncaught until a budget review forces it.
One meeting, one choice of charter. Everything else in this answer follows from which branch gets picked.
AAbility to contest. Who can't raise a flag?
Nobody outside Thibault's own reporting line, his manager, the VP of Research, had visibility into whether the pod's direction still tracked toward anything oncology could use. Amabel had no standing to ask; research didn't report to her. Thibault had no forum to be asked, until the QBR forced the question fourteen months and 2.3 million dollars in.
Not because anyone hid anything. Every number Thibault posted was true. Nobody had a reason, or a scheduled moment, to ask the one question that mattered.
Share of the pod's weekly checks run against oncology's real 36 cases, by month
100% 50% 0% 100% 35% 0% · QBR Mo 1 Mo 2 Mo 4 Mo 6 Mo 8 Mo 10 Mo 12 Mo 14
Weekly checks run against oncology's 36 real cases
This number was available every single week, for free, the whole time. It hit zero around month ten, four months before anyone in a QBR asked the question it would have already answered.
RReduce. What's the actual fix?
Pin every open research track to one named applied team and a real evaluation set tied to their actual need, from the day it's chartered. Put a checkpoint on the calendar every quarter where the pod has to show that team a credible path back to that eval set, or the track gets redirected, not left to run until a budget review happens to ask.
A product decision, not a policy memo. It changes who has a standing meeting on their calendar, not who signs an acknowledgment.
DDetect. How would you know, before someone else asks?
Ask any applied team the research is meant to eventually reach whether they can name one concrete way it would change their roadmap if it succeeded. If no team can answer that, the drift has already happened, whether anyone downstream has noticed yet or not. The leading number here was the pod's own weekly-check overlap with oncology's 36 cases, and it was heading to zero months before the QBR forced the conversation.
A metric nobody's watching is decoration. This one existed for free and nobody had to build anything new to see it fall.

The recap, one line per letter: three groups carry this cost, not one. The harm lands hardest on the specific team the research was originally promised to. Nobody had standing to ask until a budget review forced it. The fix is a named team and a real eval set, checked on a calendar. And the signal that would have caught it, weekly overlap with a real eval set, existed the whole time, for free.

Three things worth saying plainly, since this is where the real judgment sits. Errata's leadership considered capping every research track's runway at a fixed six months instead of tying it to a named applied team, and rejected it: a hard time cap can't tell genuine early-stage exploration from real drift, since both look identical in month two, while an applied-team checkpoint asks a question about direction, not duration. The AI-specific failure worth naming by name is silent evaluation drift: the pod's own reported accuracy kept climbing because it was being measured against a benchmark that had quietly stopped being the applied team's real eval set, so the number looked healthiest exactly when it was telling the applied team the least. The guardrail is requiring every research update to state which eval set produced its numbers, checked by the named applied team at each checkpoint. And the trade-off is real, accepted on purpose: pinning the pod back to oncology's 36 cases narrows the model, it may not transfer cleanly to cardiology's version of the same contradiction problem, and that's a smaller, real win worth taking over a broader, unproven one.

And if you want to be sure it really works, try it somewhere else

Same five letters, a different applied team, and this time the thing nobody separates is a model getting genuinely more impressive from a model getting any closer to something a lawyer could actually use.

Calque, a legal-document translation company, is used by immigration attorneys to translate visa-sponsorship agreements. Farrokh Doryan leads a research pod there, chartered eleven months ago around a real, narrow need: catching the exact clauses where two language versions of the same sponsorship contract quietly say different things, checked against 40 real cases the immigration team had hand-labeled with him.

Hand sketched quadrant chart titled Calque's research bets, by publishability and applied tie. X axis, tie to an applied roadmap, from tied to no one to pinned to a named team. Y axis, how novel or publishable, from incremental to novel and publishable. Three plotted items: Farrokh's idiom-preservation model, high novelty, low applied tie. General translation quality scorer, in the middle. Contract-clause contradiction flagger, high applied tie, lower novelty.
Different service, same shape of trap. The most publishable bet and the most useful bet sit in opposite corners of the chart.

Farrokh's pod drifted the same way Thibault's did, toward a broader, more publishable idea: a general "idiom and register preservation" model, meant to keep tone and meaning intact across twelve languages, not just legal clauses in two. It's a genuinely harder, more interesting problem, and the pod's own benchmark score climbed from 58 to 84 percent over eleven months. Checked against the immigration team's original 40 contract-clause pairs, the pod's contradiction-catch rate never moved. It sat at three of forty the whole time, exactly where it started, still fully manual for every attorney using Calque.

The decision Farrokh would take back Calque's leadership let the research pod choose its own eval set once the initial 40 cases had been "used up" for the first model version, reasoning that a broader benchmark would prove more about the pod's general capability. It made sense as a way to avoid over-fitting to forty examples. It stopped making sense once nobody was checking whether the broader benchmark still said anything about the forty real cases attorneys were actually waiting on.

Mapped onto GUARD: the groups are Farrokh, the immigration applied team, and Calque itself, footing an eleven-month research bill. The harm lands hardest on the immigration team specifically, the one whose real, named need the pod was chartered around before drifting. Nobody outside Farrokh's reporting line had standing to ask, since the immigration team doesn't sit in research's chain of command. The fix: pin the pod back to the 40 real clause pairs, with a quarterly checkpoint owned jointly by research and the immigration team. The detect signal: ask the immigration team whether they can name one clause type they'd trust Calque to flag if the research shipped tomorrow, before that answer was no.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: pin every open research track to a named applied team and a real eval set, checked on a calendar, not a budget review.
Cost: no budget for a formal checkpoint process this quarter. Borrow the applied team's own weekly roadmap sync, fifteen minutes, ask one question: "what would change here if the research succeeded?"
The model got better, for real: say Farrokh's idiom-preservation model hits 95 percent next release. Attorneys still need the exact clause flagged, because sounding more natural across twelve languages and catching a contradictory sponsorship clause are two different problems, and only one of them protects a client.

Where people run it wrong.
They let "the research is going well," measured by the pod's own chosen benchmark, stand in for "the research is going somewhere," without ever asking an applied team to confirm it.
They wait for the budget review to be the first real check, instead of putting one on the calendar from day one.
They kill all open-ended research the moment it doesn't fit a roadmap yet, punishing genuine early exploration the same as genuine drift, when only a named-team-and-checkpoint tells the two apart.

How to use it live. When an interviewer asks about research-PM risk, ask one thing back before answering: "which applied team is this research supposed to eventually reach, and what would they do differently if it succeeded?" That question alone is usually the exact judgment a GUARD question about research is listening for.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question about who gets hurt when a research PM has no product surface to ship into?
Tap to flip
ANSWER
GUARD: name the groups affected, find where the harm lands unevenly, ask who can't push back, name the concrete fix, name how you'd detect it before a budget review forces the question.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Thibault Cyprian, a research PM at Errata Sciences leading an unaimed research pod, and Amabel Abernathy, who owns Foldout's oncology roadmap.
3 · THE THREE GROUPS
Name the three groups GUARD says are harmed here.
Tap to flip
ANSWER
Thibault himself (two years, no shipped track record), Amabel's oncology team (their real need goes unserved), and Errata Sciences (funds 14 months and 2.3 million with no product to show for it).
4 · WHERE IT LANDS HARDEST
Which group's harm is worst here, and why that one specifically?
Tap to flip
ANSWER
Amabel's oncology team. The pod's original mandate was explicitly aimed at their dosage-conflict need before it quietly drifted toward a general benchmark, so the team promised something specific got nothing.
5 · THE FIX
What's the concrete reduce-step fix?
Tap to flip
ANSWER
Pin every open research track to one named applied team and a real evaluation set tied to their actual need, with a quarterly checkpoint that shows a credible path to a product surface or ends the track.
6 · THE NUMBER
Fill in the blank: the research pod ran ___ months and cost about $___ million before a budget review finally asked what it would ship.
Tap to flip
ANSWER
14 months, 2.3 million dollars.
7 · THE REPLAY
Same 14 months, quarterly checkpoint from day one, what changes?
Tap to flip
ANSWER
The drift gets caught at month 4, when eval overlap with oncology's real cases starts falling. Thibault re-anchors, and by month 14 the model flags 31 of the 36 known cases instead of shipping nothing.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which company, and what's the equivalent gap?
Tap to flip
ANSWER
Calque, a legal-document translation company. Farrokh Doryan's research pod chases a general idiom-preservation benchmark while the immigration-law team's real need, catching contradictory sponsorship-clause translations, goes unaddressed.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Thibault's climbing benchmark score, 61 to 89 percent, not mean the research pod was on track?
  • A. Because the benchmark was graded by a different team than Thibault's own.
  • B. Because the benchmark had quietly stopped being oncology's real eval set, and never proved the model still worked on their 36 cases.
  • C. Because 89 percent isn't a high enough score to matter to anyone.
  • D. Because Errata Sciences doesn't allow research pods to publish benchmark results.
Show hint
Check the D, detect, step in the GUARD recap.
Show answer
B. A benchmark score can climb honestly while measuring something further and further from the applied team's real need. That gap is invisible unless someone checks the original eval set directly.
True or false
2. True or false: Amabel could have stopped the drift earlier if she'd simply asked Thibault more often how the research was going.
  • True
  • False
Show hint
Look at the A, ability to contest, step.
Show answer
False. Research didn't report to her, and nobody had put a standing checkpoint on the calendar. The fix isn't more asking, it's a structural checkpoint that doesn't depend on Amabel remembering to ask.
Fill in the blank
3. The pod's weekly checks against oncology's 36 real cases fell from ___ percent in month 1 to ___ percent by around month ___.
Show hint
Check the line chart in the GUARD recap section.
Show answer
100 percent, 0 percent, month 10. That number was available every week for free, and it hit zero four months before the QBR forced the question it would have already answered.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Giving the pod a broad "advance claim synthesis" mandate instead of pinning it to Amabel's oncology team from day one. It made sense in the room: a narrow pod built around 36 examples sounded small next to "advance the science," and a broad mandate was easier to sell to the board.
Short answer, apply it yourself
5. Think of an exploratory or R&D project you've seen, AI or not. Name one applied team that should have had a standing checkpoint on it, and didn't.
Show hint
Look for a project where "progress" was measured on the project's own terms, never checked against a specific team's real need.
Show answer
Model answer: A data-science team spending a quarter building a new recommendation algorithm, with no named merchandising team reviewing it against real conversion data. Nobody could say what would ship for merchandising if the algorithm worked.
Fill in the blank, work the number
6. If Errata had put the quarterly checkpoint in place from month 1, roughly how many months of drift would have gone uncaught at most, and why does that number matter?
Show hint
The fix runs on a fixed quarterly cadence. Work out when the first checkpoint would land.
Show answer
About 3 to 4 months at most. The first checkpoint would land around month 4, catching a drift that instead ran uncaught for 14. The fix doesn't need perfect foresight, only a check nobody skips.
Before you close the answer
Why this works
Tests whether you can separate "the research is technically impressive" from "the research is going somewhere real," and whether you'd design a structural check instead of trusting people to remember to ask. Most candidates blame the researcher; the stronger answer blames the missing checkpoint.
Follow-up traps
"Isn't this just micromanaging research, forcing everything to justify itself every quarter?" Response: no, the first six to eight weeks of a genuinely new idea stay unpinned on purpose. The checkpoint applies to a track that's already been running long enough to have something to show.

"What if the broad benchmark really was the better long-term bet?" Response: then the quarterly checkpoint is exactly where that case gets made, to the applied team, with real numbers, instead of assumed for fourteen months with no one checking either way.
If pressed
The rule-based heuristic Thibault's team scrapped in month two wasn't thrown away for being wrong. It was a simple keyword-and-threshold matcher that caught eleven of the 36 cases on its own, cheap to run, easy to explain to an oncologist. The graph architecture that replaced it was more sophisticated and, checked against the same cases fourteen months later, worse.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more