Advance Case: A Written Memo or PRD Delivered as Round One
Transcript
Read the full transcript (2,049 words)
[INTERVIEWER] Advance case: a written memo or PRD delivered as round one. Here's a hard truth about the written round. This memo is your entire first impression. There's no small talk, no chance to read the room, no way to talk your way out of a weak page. Someone busy opens the doc, gives it thirty seconds, and decides whether you're worth a real conversation.
So the one thing that gets you to round two is simple to say and hard to do: a reader who is thirty seconds in already knows what you're recommending, who it's for, and how you'd know it worked. If they have to hunt for your point, you've already lost, no matter how good the thinking underneath is. A round one memo is testing whether you can hold a real position in writing, for a reader who has never met you and has ten other things to do.
It's the closest thing an interview has to your actual job, because most PM decisions get made off a doc, not a whiteboard. They want to see that you can size a problem with a number, commit to one recommendation, cut scope on purpose, and prove you'd measure the model and not just the funnel. In the next few minutes I'll give you the memo structure that survives a skim, the six things that go in it in order, a fully worked example you can copy the shape of, and the specific ways strong candidates quietly signal they've shipped an AI feature before.
The mental model that changes everything: write for a skim first, a deep read second. Your reader skims to decide if it's worth a deep read. So the structure has to reward the skim. Open with the recommendation, in bold, one sentence. Not background. Not "in this memo I will explore the opportunity space around". The decision, stated flat, and then the whole memo argues for it.
This is the single biggest lever on a written round, and it's the one most people get wrong because it feels aggressive to lead with the answer. It isn't aggressive. It's a courtesy to a busy reader, and it signals you actually have a position. Next, state the problem in two sentences with a number in it. Something like "support handles 4,000 tickets a week, and 60% of them are password resets and order status lookups an assistant could resolve." A number does two jobs at once.
It makes the problem feel real, and it sizes the prize, so the reader already sees why this is worth doing before you've pitched a thing. Then name the user and the goal metric together, and keep it to one primary metric. For most AI features that metric is a task success or resolution rate, not a vanity engagement number, because an AI feature can drive engagement while actively doing harm.
Add one guardrail so you show you know the feature can backfire. If your assistant "deflects" a ticket but the customer just comes back angrier the next day, that's a loss dressed up as a win, and the guardrail is where you prove you see that. Solution block: state v1 scope, and state what you cut. A memo that ships everything ships nothing, because it's a wish list, not a plan.
Say what's in v1, and then explicitly, what's deferred and why. Cutting scope on paper, with a reason, is the clearest single signal of PM judgement there is. Anyone can add features. Deciding what not to build is the job. Give the eval plan its own section, about half a page. The tasks the feature has to handle, a golden set, a metric per task, and the human review loop.
This is the part that separates an AI PM memo from a generic product memo. If a hiring manager reads your eval section and thinks "this person has watched a model fail and built the evaluation loop to catch it," you're through. Close with risks and open questions, and be honest about them. Three real risks, each with a mitigation.
One genuine open question you'd resolve with data rather than pretend you've already solved. Total certainty on paper reads as naivety. A calibrated "here's what I don't yet know, and here's how I'd find out" reads as someone who's shipped before. Let me build one so you can see it land. The memo prompt: "Should we add an AI resolver to the support inbox?" Line one, in bold: recommend we ship an AI resolver scoped to password resets and order status lookups, targeting 35% automated resolution within one quarter.
That's the whole answer, in one sentence, before anything else. **On-screen reference block (the memo):** - **Problem:** 4,000 tickets a week, 60% are two repetitive intents, median human response time is 6 hours. - **User + metric:** customers waiting on simple lookups. Primary metric is automated resolution rate, meaning a ticket closed with no human touch AND no re-open within 72 hours.
Target 35%. Guardrail: CSAT on AI-handled tickets stays within 3 points of human-handled. - **Solution v1:** retrieval over the help centre plus the order API, two intents only, hard handoff to a human on low confidence or any refund or billing intent. Deferred: refunds, multi-turn troubleshooting, other languages. - **Eval plan:** 500-ticket golden set from real history, labelled with the correct resolution.
Metrics: resolution accuracy, false-resolution rate, handoff precision. LLM judge on transcripts, human review on every false resolution plus a weekly 5% audit. - **Risks:** false resolutions erode trust, help centre drift breaks retrieval, measured deflection hides repeat contacts. Walk it through. The problem line does the sizing for me: 4,000 tickets, 60% in two intents, six-hour waits. The reader now knows the prize is big and the target is narrow, which is exactly the combination that makes a scoped v1 look smart instead of timid.
The metric is where I'd spend my defence, because it's built to resist gaming. Automated resolution rate isn't just "ticket closed by the bot." It's closed with no human touch AND no re-open in 72 hours. Why bolt the 72-hour window on? Because a naive deflection metric goes up every time the bot fobs someone off, and then that customer opens a fresh ticket the next morning, angrier, and your dashboard still says you won.
The re-open window closes that loophole. When you show a metric you've deliberately hardened against gaming, you're telling the reader you've watched a real metric get gamed and you're not going to let it happen again. The scope cut is doing real work too. I'm shipping two intents and hard handing off on anything involving money or a refund. That deferral isn't laziness, it's risk management: refunds carry financial and trust risk, and they need a bigger, more careful eval set before I let a model near them.
Saying "I cut this because of risk, not because I ran out of time" is the sentence that reads as judgement. Now the eval section, because on an AI memo this is the whole game. Five hundred real tickets, labelled with the right resolution, frozen so I can compare versions. I track three numbers. Resolution accuracy, the obvious one. False-resolution rate, which is the dangerous one, because a ticket closed wrongly does more damage than one left open.
And handoff precision, whether the low-confidence cases actually get to a human. The LLM judge reads transcripts at scale, and a human reviews every false resolution, because that's the failure that hurts, plus a weekly 5% random audit to keep the judge calibrated. If someone challenges the whole approach, the false-resolution metric is my anchor: I'm not optimising for "closed the most tickets," I'm optimising for "closed them right, and caught myself when I didn't." And the risks, stated straight.
False resolutions erode trust, so I bias low-confidence cases toward a human handoff. Help centre content drifts and quietly breaks retrieval, so I re-index monthly and add a staleness check. And measured deflection can hide repeat contacts, which is exactly why the 72-hour re-open window is baked into the primary metric rather than bolted on as an afterthought. Notice the risks and the metric design are talking to each other.
That coherence is what makes a memo feel like one mind wrote it. Let me do a second one, faster, so you see the skeleton isn't tied to support tickets. The prompt asks if we should add AI-generated first drafts to our marketing content tool. Line one, in bold: recommend we ship AI first drafts scoped to short-form social posts only, targeting 40% of drafts published with light edits within a quarter.
Then the problem, with a number. Our users create roughly 12,000 social posts a month, and time to first draft is the step they abandon most. Next up, the user and metric. We're targeting solo marketers and small teams. The primary metric is the light-edit publish rate, meaning a draft that goes live with under 30% of its text changed.
That measures useful, not just used. The guardrail is that the brand-safety flag rate stays flat, so we're not shipping speed by loosening what the tool will say. For the v1 solution, it's short-form posts only, grounded in the user's own past posts for voice, with a hard block on anything the safety classifier flags. We explicitly defer long-form blogs and ads, because those carry legal and brand risk and need a proper eval set.
Finally, the eval plan and the risks. We use 400 real briefs, human-rated for on-brand voice and factual grounding, with an LLM judge doing a first pass and a human reviewing every safety flag. The main risk is that a confident off-brand draft trains users to distrust the whole tool, so the voice-grounding metric gates the release. Same six blocks, different product, and I got there in under a minute because the structure does the thinking for me.
That's the whole point of having a fixed skeleton. Under time pressure you're filling a shape, not inventing one. Here's what makes them lean in. The recommendation is the first thing on the page, in one sentence they could quote back to a colleague without re-reading. The primary metric has an anti-gaming definition built in, that no reopen within 72 hours clause, which shows you've watched a metric get gamed in real life.
And the scope is cut on purpose, with the cut justified by risk rather than by "I ran out of time." Those three signals together say: this person has run a real feature, not just written about one. There's a quieter signal too, which is that the memo is easy to argue with. A good reader should be able to disagree with your recommendation precisely, because you stated it precisely.
That's a strength, not a weakness. A vague memo can't be argued with, but it also can't be believed. A sharp one invites the exact conversation you want to be having in round two. Now the ways people tank it. The first is burying the recommendation under two paragraphs of context, so the skim reader gets to the bottom of the page and still doesn't know what you're actually proposing.
The second is a metric that goes up when the feature does harm: raw deflection, sessions, messages sent. Pick one of those and you've told the reader you don't understand how AI features fail. And the third, the fatal one for this round: no eval section at all. On an AI memo, a missing eval plan reads as "has never shipped a model in production," and that's usually the end of it.
So, the whole shape. Decision in bold on line one. Problem in two sentences with a number that sizes the prize. One user, one metric, one guardrail, and make the metric hard to game. A v1 with explicit, risk-justified cuts. An eval section that proves you've measured a model and not just a funnel. And honest risks that tie back to the choices you made.
Written for a skim first, a deep read second. If you carry one line into the round, make it this: decision in bold on line one, one gaming-resistant metric, and an eval section that proves you've measured a model, not just a funnel.