InterviewFoundationalModel Fluency & the AI PM Role / AI PM role variants: platform, applied, infra, research / #21
Tell me which AI PM variant you want and defend the choice against the other three.
PICK · a candidate commits to one AI PM seat and defends it against three panelists pulling for the other three, on Quillhaven Systems' enterprise search tool Undercroft
Undercroft is Quillhaven Systems' enterprise search tool: type a plain question, get an answer pulled from every wiki, contract, ticket, and old Slack thread the company owns. Four real jobs sit around it: applied, platform, infra, and research. Folasade Osunkoya is in the final round for the infra seat. Faramond Winterhalter, running the panel, has one line on Folasade's resume circled in pen. She's about to ask the question that line is there to answer.
The direct answer
I want the infra seat: owning Undercroft's embedding index, its versioning, and the reindex pipeline that decides whether a query and a document are even speaking the same language. I want it because I already broke exactly that once, badly, and I know precisely what I never want to happen silently again. Applied gets me closer to a searcher's daily pain, platform gets me more leverage, research gets me the next architecture, but none of those seats can protect what an unowned index nearly cost my last company: eleven weeks where the search looked fine and wasn't.
Do this, in order
Say the infra seat, first, before any reasoning.Why: a PICK answer that reasons its way to a pick sounds discovered, not decided, and the question is testing whether you can commit.
Name the real reason: you already broke this once.Why: a preference with no scar behind it sounds like a guess dressed up as a decision.
Answer each of the other three seats by name, not as one group.Why: waving off "the other three" together sounds like you never seriously considered any of them.
State what you're not optimizing for, honestly.Why: a pick with no acknowledged cost isn't a decision, it's marketing.
Give kill criteria for all three rejected seats, not just one.Why: three specific conditions read as judgment; one vague "if things change" reads as an excuse.
Back the whole thing with one real number, not a vibe.Why: "54 percent recall on a third of the corpus, for eleven weeks" survives follow-up questions that "I care about reliability" does not.
How to answer this, stage by stage
Nobody is grading whether you can name four job titles. They're grading whether you can hold one of them under three people, one at a time, trying to talk you out of it.
1
Scope it to one seat, one panel, one real scar
Say it like this
"Let's ground this. I'm in the final round for AI PM, Search Infrastructure, at Quillhaven Systems, on Undercroft, their enterprise search tool. Three people across the table each think I should want their seat instead. Here's the one I'm picking, and it's not a guess."
Why this works
A named company, product, and panel turn "which variant do you want" from a personality quiz into a real, defensible call.
2
Say the structure out loud
Say it like this
"I'll run this as PICK. Position, my actual pick, before any reasons. Impact, what I give up by not taking each of the other three. Cost asymmetry, what I'm openly not optimizing for. Kill criteria, what would actually change my mind."
Why this works
Two seconds of structure tells the panel you've thought this through, not that you're picking your favorite on the spot.
3
Give the position, committed
Say it like this
"My answer is infra. I want to own Undercroft's embedding index: the versioning, the reindex pipeline, the thing that decides whether a query and a document are even speaking the same language. I want it because I already broke exactly that, once, badly, at my last job. I know precisely what I never want to happen silently again."
Why this works
This is the direct answer, said out loud, tied to something she's actually good at, not a vibe about caring about quality.
4
Defend it against Applied, the pull toward the user
Say it like this
"Osazee went first. 'Applied gets you a real searcher's failed query today, and a fix by Friday. Infra, you're three steps removed from anyone who'd ever thank you for it.' I told him he's right about that loop, and I'm giving it up on purpose. Applied owns the relevance eval set and the confidence ranking a searcher actually sees. I'd be the one making sure the ground under that screen hasn't quietly moved. Somebody has to own that even when nobody claps."
Why this works
Naming what's genuinely lost, instead of pretending applied has no case, makes the pick sound considered instead of defensive.
5
Defend it against Platform, the pull toward leverage
Say it like this
"Kentigern's case was leverage. One decision on the retrieval API's contract reaches three products this year: a ticketing tool, a Slack bot, a CRM plugin. He's right that's a bigger lever than one index. I said: a shared API sitting on top of an index that's quietly wrong ships the same wrong answer to all three products at once, just faster. Platform's leverage only means something if what's underneath it can be trusted. I want to be the reason it can."
Why this works
Turning platform's own strongest argument, leverage, into the reason infra has to exist under it beats trying to out argue the leverage claim directly.
6
Defend it against Research, the pull toward what's next
Say it like this
"Uzoamaka's case: 'in two years, index versioning might just be a solved, boring problem. Don't you want to help build what replaces it?' Maybe, someday. But nobody at my last job got to explore a better retrieval architecture that spring, because we spent eleven weeks finding out a third of our corpus was speaking a different language than the queries hitting it. Research needs a stable floor to even know if a new architecture is actually better. I'd rather build that floor first."
Why this works
This doesn't dismiss research. It states plainly why research's own claims depend on infra being solid, which lands harder than "research sounds less practical to me."
7
Name the kill criteria, then close on one line
Say it like this
"Three things would move me. If the reindex pipeline gets genuinely boring, and the real lever for growth becomes people not trusting the answer they already get, I'd take applied. If two more teams ship against the retrieval API this year and each one hits the same contract gap platform hasn't fixed, I'd take platform. If recall plateaus for reasons that have nothing to do with reliability, because the architecture itself is the ceiling, I'd take research. Until one of those is true, I want infra, because I'm the person who already found out what happens when nobody owns it."
Why this works
Closing on three named, checkable conditions is what makes a strong opinion sound like judgment instead of stubbornness, and it's the line the whole room remembers.
Let's learn
Undercroft is Quillhaven Systems' enterprise search tool. An employee types a plain question. Undercroft pulls the answer from every wiki, contract, support ticket, and old Slack thread the company owns. It answers about 900,000 employee questions a week, across every customer that runs it.
Four real jobs sit around one product. Each one is a genuine seat, not a made up title to fill out an org chart.
An applied PM owns the answer box: the words a searcher reads, the confidence ranking, the eval set that scores whether an answer is actually right. A platform PM owns the retrieval API three other Quillhaven products already call: a ticketing tool, a Slack bot named Ask Undercroft, and a CRM plugin. An infra PM owns the embedding index underneath all of it, the thing that turns a document into a list of numbers a computer can compare. A research PM chases a better way to find the right document at all, instead of the way Undercroft uses today.
Knowledge spark: what's an embedding?
A model reads a document and turns it into a long list of numbers, a kind of fingerprint for what the document means. A search question gets turned into the same kind of list, then the system finds the documents whose numbers sit closest to it. Two documents that mean similar things end up with similar numbers, even if they don't share a single word.
Folasade Osunkoya is in the final round for the infra seat. Before this, she spent two years at Vaultkeep Discovery, a legal document search company, running exactly this kind of index for 1.2 million case filings. The panel is asking her this question because of one line on her resume: eleven weeks, once, the index looked completely fine and wasn't.
The third box is the one nobody was watching. That's the whole story in one picture.
Vaultkeep Discovery swapped its embedding model to cut inference cost, from about $38,000 a month to $19,000. Every document that arrived through the live intake path, new filings, about 1,900 a day, got re-embedded under the new model automatically. But 410,000 older filings, about a third of the whole corpus, had come in years earlier through a separate bulk import pipeline. Nothing in that pipeline knew to check which model had embedded a document. It just assumed "already embedded" meant "done."
Recall at 10, live path versus stale path, during the eleven-week gap
Recall@10, freshly embedded documentsRecall@10, documents on the old embedding
The eval set Vaultkeep trusted only ever sampled the fresh path. It stayed green through all eleven weeks, because the only part of the corpus it ever checked was the part that was fine.
The gap didn't announce itself. Comparing a new model's numbers to an old model's numbers is like comparing two different rulers and hoping they happen to agree. Nothing in the output said so. The search just kept returning a confident top result every time, and for a third of the corpus, it was quietly the wrong one.
The model never got worse at reading a document. It just started comparing two different rulers and calling them the same one.
What it costs at its worst: across those eleven weeks, queries a week quietly landing in the stale third of the corpus climbed from 340 to 825, and none of it showed up on the dashboard the team actually watched. Vaultkeep issued about $210,000 in service credits to twelve affected customers, and lost one account worth $86,000 a year. Roughly $296,000, and nobody saw a single number move on the screen they trusted, because that screen was built to sample the part of the corpus that was never broken.
The choice I would take back
When the bulk import pipeline was first built, nobody tagged a document with which embedding model produced its vectors. That was fine when only one model existed. It stopped being fine the day a second model shipped, and the version tag that would have caught the mismatch instantly simply didn't exist to check.
What I would leave alone: a one-off sandbox instance a team spins up for a hackathon, throwaway data nobody depends on for real work. It doesn't need this kind of versioning rigor. Let that one stay simple; nothing real is riding on it.
The lesson: a gap in coverage isn't a missing feature. It's a debt that sits quietly until the one week a real question happens to land exactly where nobody was looking.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one for the two years, then the eleven weeks, that actually built the answer.
The resume sitting in front of three panelists at Quillhaven Systems had one line circled in pen: "led embedding index recovery, Vaultkeep Discovery, eleven weeks." Faramond Winterhalter had circled it herself before Folasade Osunkoya ever sat down.
Folasade had spent two years at Vaultkeep Discovery keeping the machinery under its legal search product running: the pipeline that turned scanned court filings into an index a lawyer could actually search. She was good at it in the boring way that matters, the kind of good where nothing breaks long enough that nobody thinks about the job at all. New filings landed, got embedded, got indexed, and showed up in search within the hour. By her second year, she mostly watched dashboards that stayed green.
Then Vaultkeep swapped its embedding model to cut inference cost. Folasade's team retested the live intake path, the 1,900 new filings a day, and it looked perfect. Nobody retested the bulk import pipeline, the one that had carried 410,000 older filings in from an old acquisition, because nobody had ever needed to touch it since it was built. It had no owner, no test, and no reason, on paper, to be any different from the rest of the index.
It was different. The old filings kept their old model's embeddings. New queries got the new model's embeddings. Nothing about the mismatch looked broken from the outside. The search kept returning a confident top result every single time. It was just, for a third of the corpus, quietly the wrong one.
Eleven weeks between the migration and the first real question that happened to land in the wrong third of the corpus.
It surfaced on a Thursday, in a business review nobody expected to be memorable. A compliance associate had spent forty minutes searching for a specific precedent ahead of a regulatory audit and gotten nothing back. She found the case herself, through a competitor's tool, and asked, in front of her whole team, why Vaultkeep's own search had missed something that had been sitting in Vaultkeep's own corpus the entire time.
Folasade got the ticket that afternoon. It took two days to find the actual cause, because the first instinct was to blame the new model's quality, not its coverage. The break wasn't in what the model knew. It was in what had never been re-embedded at all.
The decision Folasade would take back sits in a fifteen minute standup, eighteen months before any of this. An engineer building the bulk import pipeline mentioned, almost as an aside, that it didn't tag which model had embedded each vector. Someone suggested adding one. It got shelved, because at the time there was exactly one embedding model in production, and tagging a single, unchanging value felt like busywork nobody had a sprint for.
Run the same eleven weeks again, with that tag in place and a shadow eval sampled across every ingestion path, not just the live one. The new model ships on the same Tuesday. The shadow eval flags the bulk import path within a day: recall on that segment craters to 54 percent before a single real customer query ever sees it. The fix ships before the compliance associate's search ever comes up empty. Two days, not seventy seven.
The missing tag never felt like a decision. It felt like nothing, because at the time it changed nothing. It only became a decision the day a second model made it matter, and by then nobody remembered it had ever been made.
That's the whole story behind one circled line on a resume. In the room, across the table from Osazee, Kentigern, and Uzoamaka, it took her about four minutes to say. It took Vaultkeep Discovery eleven weeks and roughly $296,000 to teach her.
PICK, said about your own seat instead of somebody else's feature
Not a way to rank four job titles by prestige. PICK is what forces you to say which seat you want, then defend it against the best case for the other three, one at a time, without hedging and without flinching.
PPosition. The pick, in one sentence, before any reasons.
I want the infra seat: owning Undercroft's embedding index, its versioning, and the reindex pipeline. Not applied, not platform, not research. I want it because I already broke exactly this once, badly, and I know precisely what I never want to happen silently again.
Say the pick before the story. A position that only shows up after the evidence sounds reverse engineered from it.
Four real reasons, four real seats. Only one of them is backed by a scar she can point to.
IImpact. What each of the other three seats actually costs her, if she doesn't take it.
Not choosing applied gives up the closest, fastest feedback loop there is: watching a real searcher's query fail today and shipping a fix by Friday. Not choosing platform gives up real leverage: one decision on the retrieval API's contract reaches three products, and three teams, at once, this year. Not choosing research gives up the chance to work on the next retrieval architecture while it's still new, instead of joining once it's already the standard. All three losses are real. None of them protect what an unowned embedding index nearly cost Vaultkeep Discovery.
Naming all three losses by name, not as one wave off, tells the panel she considered each seat seriously before rejecting it.
CCost asymmetry. What she's openly not optimizing for.
Shipping a model upgrade carefully, with a shadow eval across every ingestion path, adds about a week to every migration, and someone will notice the roadmap slip by Monday. That's the cheap, visible mistake, if you can even call it one. Skipping that same shadow eval to move faster is the hidden, expensive one: a third of a corpus can drift silently for eleven weeks, and the bill lands as $210,000 in credits and one lost account, not as a missed sprint goal.
This is the hardest move in PICK. Naming which mistake is loud and cheap, and which one is quiet and expensive, is what turns a preference into an actual decision.
One mistake shows up on a calendar the same week. The other one hides inside a search that looked, on the surface, like it was working.
Queries a week quietly landing in the stale segment, Vaultkeep Discovery
Queries a week landing in the stale segment
Nobody was watching this number, because it lived outside the eval set entirely. It more than doubled across eleven weeks and nobody knew to look.
KKill criteria. What would actually change her mind.
Three things would move her. The reindex pipeline gets genuinely boring, and the real growth lever becomes people not trusting an answer they already have; that's applied's moment. Two more teams ship against the retrieval API this year, on top of the three already calling it, and each hits the same contract gap platform hasn't closed; that's platform's moment. Recall plateaus for reasons that have nothing to do with reliability, because the architecture itself is the ceiling; that's research's moment. Until one of those is true, infra is the seat that needs someone who already knows what happens when nobody's watching it.
A position with no way to be proven wrong is just an opinion held tightly. Naming three separate, checkable conditions, one per rejected seat, before anyone asks, is what makes this a real case instead of a favorite.
Infra sits in the one corner nobody applauds and everybody pays for, if it's wrong. That's exactly why she wants it.
What would flip this pick
A boring, solved reindex pipeline plus a trust problem, applied. Two more teams hitting the same platform contract gap, platform. Recall stuck for architecture reasons, not reliability ones, research. Any one of those, and she'd take a different seat, out loud, without pretending she never would.
Three things worth stating directly, since the real judgment sits here. The alternative worth naming and rejecting isn't only the other three seats, it's also "let each product team own its own slice of embedding versioning independently." That option loses because embedding spaces have to stay globally consistent: you can't have applied's eval set say the index is fine while platform's three downstream consumers are quietly pulling mismatched results from a corpus slice that eval set never touched. The AI-specific failure worth naming by name is an embedding-space mismatch hiding behind confidence: a model migration leaves part of a corpus embedded in an old vector space while new queries arrive in a new one, and nearest-neighbor search returns a result anyway, with nothing in the output flagging that the comparison never meant anything. The guardrail is a version tag on every vector, model id plus a checksum of the embedding config, and a shadow eval sampled across every ingestion path, not just the live one, run before any embedding model ships. And the trade-off is real and accepted on purpose: that shadow eval adds roughly a week to every model migration, on purpose, because a fast migration that silently breaks a third of the corpus costs far more than the week anyone saved.
And if you want to be sure it really works, try it somewhere else
Same four letters, a vet clinic instead of a law firm, and this time the honest answer isn't infra at all.
Clarivue, built by Vetraine Diagnostics, reads X-rays and ultrasounds for 340 vet clinics and flags likely fractures, tumors, and foreign objects for a vet to confirm. Ozioma Amberidge is interviewing for AI PM, Platform, there. Corentine Wrackham, Vetraine's Head of Applied Diagnostics, pushes the same question from the other side of the table: why not applied, where you'd catch a missed fracture the same day a vet does?
Different office, different scan, a genuinely different honest answer this time. That's the point of the method, not a flaw in it.
Ozioma's answer runs the same four letters and lands somewhere different, on purpose. Position: the platform seat, owning the shared scoring contract that Vetraine's imaging app, its referral portal, and its insurance-claims tool all call. Impact: not choosing applied gives up seeing a missed fracture the same day it happens; not choosing infra gives up ownership of the raw scan-to-vector pipeline; not choosing research gives up the newer, multi-view architecture Vetraine's research team is testing. Cost asymmetry: Ozioma is openly not optimizing for the fastest individual bug fix. A contract change ships slower, on purpose, because it has to be checked against three consumers instead of one. Kill criteria: if any single consumer's needs diverge so far from the shared contract that platform decisions stop being able to serve all three at once, that's the sign to split the API, not to keep forcing one contract to fit three jobs.
Same method, a different, equally defensible answer. That's the whole point: PICK doesn't hand out one universal right seat. It makes you defend the one you actually want, on the evidence you actually have.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to it: name the seat, name the one scar or reason that earns it, name one thing you're giving up, done.
Cost: no time to build the full case live. Lead with the number: "I broke this once, it cost $296,000 and eleven weeks, and I don't want to do that again" earns you the follow-up question that lets you say the rest.
The model got better, for real: say the new embedding model genuinely outperforms the old one on every benchmark. Keep the shadow eval anyway. A better model can still land its new documents in a different corner of vector space than the old one, and "better" was never the same claim as "compatible."
Where people run it wrong.
They pick the seat that sounds most impressive instead of the one they can actually defend with evidence.
They wave off the other three seats as a group instead of naming what's genuinely lost by not choosing each one.
They give a position with no kill criteria, which reads as rigid instead of confident the moment a good interviewer pushes back twice.
How to use it live. Before answering, ask yourself one question: what's the one scar or fact that makes this pick actually mine, not just a job description I liked the sound of? If you can't name it, you haven't picked yet, you've just recited an org chart.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "tell me which AI PM variant you want and defend the choice against the other three"?
Tap to flip
ANSWER
PICK: commit to a position, name the impact of not choosing each alternative, find the cost asymmetry, then say what evidence would flip you. Used here in its most personal form: a candidate's own career pick, defended live.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Folasade Osunkoya, interviewing for AI PM, Search Infrastructure, at Quillhaven Systems. Faramond Winterhalter runs the panel. Osazee Groenwald (applied), Kentigern Sithole (platform), and Uzoamaka Falchester (research) each push back for their own seat.
3 · THE POSITION
What's the actual pick, in one sentence?
Tap to flip
ANSWER
The infra seat: owning Undercroft's embedding index, its versioning, and the reindex pipeline, because she already broke exactly that once and knows what silence costs.
4 · THE IMPACT
Name one thing lost by not choosing platform.
Tap to flip
ANSWER
Real leverage: one decision on the shared retrieval API's contract would reach three products, and three teams, at once. The infra seat doesn't get that reach.
5 · THE COST ASYMMETRY
What is Folasade openly not optimizing for?
Tap to flip
ANSWER
Visible, immediate impact. A careful model migration with a shadow eval adds about a week, and someone notices by Monday. Skipping that eval is the hidden, expensive mistake: eleven weeks of silent drift and about $296,000.
6 · THE NUMBER
Fill in the blank: recall on the fresh, live-path documents held at ___ percent. On the stale third of the corpus it was really ___ percent, for ___ weeks, undetected.
Tap to flip
ANSWER
91 percent fresh, 54 percent stale, 11 weeks. The eval set that stayed green only ever sampled the fresh path.
7 · THE KILL CRITERIA
Name one of the three conditions that would flip Folasade's pick.
Tap to flip
ANSWER
If two more teams ship against the retrieval API this year and each hits the same contract gap platform hasn't closed, that's platform's moment, and she'd take it. (Reindexing turning boring plus a trust problem moves her to applied; recall plateauing for architecture reasons moves her to research.)
8 · CROSS-PRODUCT TRANSFER
Section 4 runs PICK again for a different product and a different candidate. Which one, and which seat do they pick instead?
Tap to flip
ANSWER
Clarivue, Vetraine Diagnostics' vet imaging tool. Ozioma Amberidge picks platform, not infra, defending a shared scoring contract that three consumer apps call.
Check yourself Score: 0 / 0
Multiple choice
1. Why doesn't platform's bigger leverage argument win, even though Kentigern is right that one decision there reaches three products at once?
A. Because platform PMs aren't allowed to make contract decisions without infra's sign off.
B. Because a shared API built on top of a silently wrong index ships the same wrong answer to all three downstream products at once, just faster.
C. Because Quillhaven's platform team is short staffed this quarter.
D. Because Folasade doesn't think leverage counts as a real form of impact.
Show hint
Check stage 5 of the walkthrough, and the I step in the PICK recap.
Show answer
B. Leverage is real, and named as a real loss. But leverage built on an untrustworthy index just spreads the same failure faster and wider.
Fill in the blank
2. Recall on the freshly embedded live-path documents held at ___ percent. On the stale third of the corpus it was actually ___ percent, for ___ weeks, while the dashboard stayed green.
Show hint
Check the bar chart in Let's learn, and the eleven-week timeline.
Show answer
91 percent, 54 percent, 11 weeks. The eval set only ever sampled the fresh path, so it had no way to see the gap.
True or false
3. True or false: picking the infra seat means Folasade thinks applied, platform, and research don't really matter.
True
False
Show hint
Check the I step, and stages 4 through 6 of the walkthrough.
Show answer
False. She names a real, specific loss for not choosing each of the other three. PICK requires naming genuine impact for the paths not taken, not dismissing them.
Short answer, name the rejected alternative
4. Besides the other three seats, what alternative does this answer name and reject, and why does it lose?
Show hint
Look at the "three things worth stating directly" paragraph after the K step.
Show answer
Model answer: Letting each product team own its own slice of embedding versioning independently. It loses because embedding spaces have to stay globally consistent across every consumer, not just the one whose eval set happens to say the index is fine.
Short answer, apply it yourself
5. Think of a role or specialization you'd pick within a field you know. What's the one thing you're honestly not optimizing for by picking it, and what would have to be true for you to switch?
Show hint
Name the real cost of your pick, not just its upside, and a specific, checkable condition that would move you.
Show answer
Model answer: An engineer who picks database reliability work over customer-facing features gives up quick, visible wins and the credit that comes with shipping something users see. They'd switch if the database work became genuinely stable and boring, and the real lever for the product became something only a features engineer could move.
Short answer, work the number
6. If the stale segment had been 10 percent of the corpus instead of 34 percent, with the same 91 versus 54 percent recall gap, would Folasade's pick still hold up as strongly? Why or why not?
Show hint
Scale 410,000 documents down by the same ratio, and think about whether the underlying mechanism changes.
Show answer
Model answer: probably still worth fixing, but the case is less dramatic. At 10 percent of 1.2 million documents, that's about 120,000 affected instead of 410,000, a smaller but still real blast radius. The mechanism, an unversioned path invisible to the eval set, is exactly as dangerous regardless of what share of the corpus it happens to cover.
Before you close the answer
Why this works
Tests whether a candidate can commit to a real preference and defend it with evidence instead of a diplomatic non-answer, and whether they can name real losses for the paths not taken instead of pretending the other three have no case. Most candidates either dodge the question or pick a seat with nothing behind it.
Follow-up traps
"Isn't naming three kill criteria just hedging your bet so you can claim any outcome later?" Response: no, hedging is refusing to name a condition. Naming three specific, checkable ones before anyone asks is what makes the position falsifiable, which is the opposite of hedging.
"Couldn't Vaultkeep's problem have been caught with better monitoring instead of a whole seat dedicated to it?" Response: a shadow eval sampled across every ingestion path is exactly that monitoring. Someone still has to own building it, running it before every migration, and being the person the company can point to when it's the reason nothing broke.
If pressed
The version tag isn't just a model name. It's the model id plus a checksum of the embedding config itself, chunk size, normalization, pooling method, because two runs of the "same" model with a different chunking setting produce vectors just as incompatible with each other as two completely different models.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.