…
AI Product Case Questions

Two Written Prompts, a 2-Pager Each, Plus a Panel Presentation

How to handle an AI PM interview format with two written prompts, a two-page answer for each, and a panel presentation.

Transcript

Read the full transcript (2,139 words)

[INTERVIEWER] Two written prompts, a 2-pager each, plus a panel presentation. So this is the take home that scares people. You get two prompts, you write two pages on each, and then you sit in front of a panel and defend them. It feels like three jobs stacked on top of each other, and most people treat it that way, which is exactly why they run out of time.

Here is the thing that actually separates a strong submission from an average one. It is not the writing. It is that each pager has a decision at the very top, one metric that would prove you right, and an eval plan that shows you already know AI features fail quietly. Get those three things on the page and you are most of the way there.

Miss them, and it does not matter how polished the prose is. The panel is really testing whether you can think like an owner under a deadline, and then hold your position when three smart people poke at it. They are not marking your grammar. They want to see judgement: can you take an open prompt, decide what to build, cut the rest, and say out loud how you would know it worked.

In the next few minutes I will give you the exact structure for both pagers, the six blocks that go in each one, how to budget a two hour take home so you do not blow it on page one, and how to walk into that panel and lead so their interruptions cannot knock you off course. Your first move is to stop thinking of this as one big blur.

Two pagers means two identical passes. Same structure, same clock, twice. That single reframe is what keeps you from pouring ninety minutes into the first prompt and stapling a rushed second one on at the end. Start each pass by reading the prompt for the real ask. Underline the noun and the verb, because they tell you what kind of problem this is.

"Improve trust in AI summaries" is a metrics problem at heart. "Ship a coding assistant for junior devs" is a scoping problem. Different problem types want a different emphasis in your pager, so name the type before you write a word. Then write one sentence: the decision you are recommending. Not the topic, the decision. That sentence goes at the top of page one, in bold, and everything below it argues for it.

Now the core. Use the same six block PRD for both pagers, every time. Block one, the problem, stated with a number so it feels real. Block two, the users, and I mean a named segment, not "everyone". Block three, the goal plus one primary metric and its guardrail. One metric. Block four, the solution: what actually ships in version one, and just as important, what you are explicitly cutting.

Block five, the eval plan, which is how you measure the model itself, not just the funnel. Block six, risks, with what you would do about each. Roughly half a page of prose, half a page for the eval table. Here is where you win or lose. Make the eval plan the differentiator. Most candidates write a tidy product spec and stop, and every one of those specs looks the same.

For an AI feature, the eval plan is the spec. Name the task set, the golden set size, the metric per task, and the point where a human reviews. That block is the thing that quietly shows you have actually shipped an LLM feature before, and you know it can hallucinate on you. A hiring manager can smell the difference between someone who has read about evals and someone who has had one fail in production.

Write the panel narrative last, and keep it to three slides in your head. The decision. The one number that would prove it. The biggest risk and your mitigation. Why only three? Because panels interrupt. If you lead with the decision, an interruption at minute two still leaves your main point already delivered. If you save the decision for a big reveal at the end, and they cut you off, you have said nothing.

And when they do interrupt, treat the question as a gift, not an ambush. If someone asks "what if the citations are wrong?", you do not get defensive, you say "good, that is exactly what the citation precision metric catches, and here is the threshold I would hold it to." A panel is scoring how you handle pressure as much as what you built, so the calm answer that ties back to your own metric is worth more than a perfect page.

Last piece, and it is the one people ignore: time box hard. Say it is a two hour take home. That is forty minutes per pager, twenty minutes to build the panel story, twenty minutes of buffer for the thing that always goes wrong. A beautiful first pager next to a thin second one scores worse than two solid, slightly rough ones.

The panel reads consistency as judgement. Let me walk one all the way through so you can see it fill in. Prompt A: "Design an AI feature that increases trust in Claude's document summaries." Decision on line one, in bold: ship inline citations that link each summary claim back to the source span, gated behind a confidence threshold. That is the sentence.

Now the six blocks fill in underneath it. **On-screen reference block (Prompt A pager):** * **Problem:** users do not trust summaries because they cannot verify a claim without re-reading the whole document. * **Users:** knowledge workers summarising twenty page contracts and reports, not casual chat users. * **Goal + metric:** raise summary trust. Primary metric is verified claim rate, the share of summary sentences a user clicks to check that resolve to a correct source span.

Target 85%. Guardrail: summary latency stays under 4 seconds at p95. * **Solution v1:** every summary sentence carries a citation chip to its source paragraph. Sentences the model cannot ground above 0.7 confidence get dropped, never shown unsourced. Cut from v1: multi doc summaries and editing. * **Eval plan:** 300 example golden set of document and summary pairs, human labelled for claim support.

Metrics: citation precision, hallucination rate, coverage. LLM as judge for a first pass, human review on every disagreement plus a 10% random audit. * **Risks:** over dropping unsourced claims makes summaries feel thin. Mitigation: track claims shown versus claims generated, and A/B the threshold. Walk the panel through that top to bottom. The problem is verification cost, and I put a number on the document length so it is concrete.

A twenty page contract is not something you re-read to check one line, so an unverifiable summary is worse than no summary, because it looks authoritative and might be wrong. The metric is the clever bit: I am not measuring "trust" as a survey, I am measuring whether the claims a user bothers to check actually hold up, which is trust you can count.

The guardrail stops me from shipping a slow, citation heavy summary that technically scores well and nobody waits for. And notice the tradeoff I am putting on the table on purpose: dropping claims I cannot ground costs me coverage. I would rather say that out loud than pretend my design is free. Now spend real time on the eval block, because that is the one the panel will push on.

Three hundred document and summary pairs, human labelled, is enough to get a stable read without being so big you cannot refresh it. I am measuring three things, and they pull against each other on purpose. Citation precision asks: when the model cites a span, does that span actually support the claim? Hallucination rate asks: how many claims have no support at all?

Coverage asks: of everything worth saying, how much could I ground and keep? A model can score beautifully on precision by only ever citing the one sentence it is sure about, which is why I watch coverage next to it. For the grading itself, I run an LLM as judge first because it is cheap and covers all 300, then a human reviews every case where the judge and the label disagree, plus a flat 10% random audit to keep the judge honest even when it is confident.

If a panellist asks "why trust the judge?", that random audit is your answer: I do not trust it blindly, I check it against humans on a fixed slice every cycle. Now the second prompt, quickly, because you get two and I want you to see the same skeleton hold. Prompt B asks you to ship an AI feature that helps junior developers learn from a senior codebase.

The decision on line one is to ship an inline explain this change feature. When a junior accepts a suggestion, it generates a short grounded rationale tied to the specific files it touched. Then we run the same six blocks. The problem is that juniors accept suggestions they do not understand, so they fail to learn and cannot defend the change in review.

The users are developers with under two years of experience on an unfamiliar codebase. The metric is the rationale helpful rate, measured by a thumbs signal plus whether the developer edits the explanation before it lands in the pull request description. The guardrail is no measurable increase in time to merge. The eval uses 200 real change sets, human rated for whether the rationale is correct and actually grounded in those files.

The risk is that a confident wrong explanation teaches the wrong thing, so I gate it on the same groundedness check as Prompt A. Same structure, different domain, and that is the point. You do not reinvent the pager for each prompt, you just run the skeleton twice. When you present both to the panel, you connect them. You tell them both features live or die on the same groundedness check, so if you had one week of engineering time you would build that check once and reuse it.

That kind of line tells them you think in systems rather than one off features. It shows the difference between a product manager they would hire and a candidate who just wrote two nice documents. Here is what makes a panel lean in. The decision is on line one of each pager, not buried in a conclusion they have to dig for.

The eval plan names a real golden set size and a per task metric, so it reads as something operational you would actually run on Monday, not a paragraph of good intentions. There is one primary metric with a guardrail, not a dashboard of ten metrics that lets you avoid committing. And in the panel itself, you defend a tradeoff out loud.

When you say "dropping unsourced claims costs me coverage, and here is why I will take that trade," that is the moment they mark you as someone who has made this call before, not someone auditioning for it. Now for the ways people sink this one. The first and most common: writing two product specs with no eval plan at all, so the AI specific judgement never shows up, and your pager reads exactly like every other candidate's.

The second: spending ninety minutes making page one gorgeous and then bolting a thin, apologetic page two on at the very end. The panel notices the drop off instantly. And the third, which happens live in the room: reciting your pager at the panel line by line instead of leading with the decision and letting them drive. If you read it to them, you have turned a conversation into a monologue, and you have given up the one advantage of being in the room, which is that you can watch what they care about and go deeper there.

So put the whole thing together. Two prompts means two identical passes on the clock, not one big effort and one afterthought. You use the same six blocks each time: a problem with a number, a named user, one metric with a guardrail, a version one with explicit cuts, an eval plan, and honest risks. The panel narrative is just three things: the decision, the number, and the risk.

You put the decision first so an interruption cannot rob you of your point. Run the same clock on both so neither gets shortchanged, and treat the panel as a conversation you steer rather than a test you survive. If you remember one line walking in, make it this: two pagers, same six blocks each, decision on line one, and the eval table is the part that proves you've done this for real.

Keep learning