…
AI Product Case Questions

Prototype a Prompt Chain and Discuss Measuring Its Reliability

A worked answer to a real AI PM interview task: prototype a prompt chain and explain how you would measure its reliability.

Transcript

Read the full transcript (1,923 words)

[INTERVIEWER] Prototype a prompt chain and discuss measuring its reliability. Anyone can draw a chain of prompts on a whiteboard. Box, arrow, box, arrow, done. That is not the signal. The signal is whether you can say how you would know it is reliable, and exactly where in the chain it breaks. Here is why that matters so much. A chain's failure rate compounds.

If each of four stages is 90% reliable, the whole thing isn't 90%, it's more like 66%, because the reliability at each stage multiplies through. So the strong answer doesn't just draw the chain. It measures each stage and the end-to-end performance, and it can point at the one box that is dragging the whole system down. This is testing whether you think about LLM systems the way an engineer thinks about a pipeline, or the way a hopeful thinks about a demo.

A demo works once and you cheer. A pipeline you measure, because you know it is non-deterministic and it will fail differently tomorrow. The interviewer wants to see you decompose the chain, name what each stage can get wrong, measure them separately, and then report reliability honestly with variance, not as a single lucky number. Over the next few minutes I will build a real chain, show you how to measure every stage of it, walk through a worked example with actual numbers, and give you the two or three phrases that tell a hiring manager you have debugged one of these for real.

Start by designing the chain for a real task, not in the abstract. Take answering a question from your internal docs. A sensible chain has four stages. Retrieve, where you pull the candidate passages. Draft, where you write an answer grounded in those passages. Critique, where a second model call checks the draft against the passages for any claim the text doesn't support.

Finalize, where you revise the flagged bits, or flag low confidence and abstain. You also need to say why each stage exists, because that is half the answer. The critique stage exists because a single draft call will hallucinate, and a self-check catches a real chunk of those hallucinations before they reach the user. If you cannot justify a box, delete it.

Next, name what each stage can get wrong. You cannot measure a chain you haven't decomposed. Retrieve can miss the relevant passage entirely, which is low recall, or it can drag in noise that distracts the draft. Draft can add claims that are not in the passages, which is the classic grounding failure. Critique has two failure modes. It can miss a real hallucination, which is a false negative, or it can flag a perfectly good answer, which is a false positive.

Finalize can over-hedge, turning a correct answer into a mealy "I'm not sure." Four stages, and at least six distinct ways to fail. This is exactly why one end-to-end number tells you nothing about what to fix. This is where the compounding really bites, so let me put numbers on it. Suppose retrieval finds the right passage 92% of the time, the draft grounds correctly 90% of the time, and critique catches its errors 85% of the time.

Multiply those through and your end-to-end reliability is nowhere near 92%. It is dragged down every time you pass through a leaky stage, meaning the final compound score is always lower than your worst individual stage. So if a stakeholder says to just make the chain more reliable, the useless move is to tweak all four boxes at once. The useful move is to find the lowest number and fix that one, because that is the box actually costing you.

You can only do that if you have measured the stages separately, which is the entire reason this decomposition exists. Now measure each stage separately, because that is how you find the weakest stage instead of guessing. Retrieval gets recall-at-k and precision-at-k, scored against a labelled set of query-to-passage pairs. The draft gets groundedness, meaning what share of its claims are actually supported by the retrieved passages.

The critique gets an agreement rate, tracking how often it agrees with a human's judgement of the same draft. When you have these three numbers, you stop arguing about the chain in the abstract and you point at the number that is low. Then measure end-to-end task success, which is the number the product actually cares about. Does the final answer correctly and groundedly answer the question on a golden set?

Track it as a rate. Here is the part that makes it useful. Track the failure modes as categories underneath it. Wrong retrieval, hallucination that critique missed, correct answer wrongly hedged. So when the end-to-end number is 86% and not 95%, you can immediately say which bucket the missing 9 points are in, and that tells you what to build next.

Last, and this is the one people skip, is reliability over time, not just one run. LLM calls are non-deterministic, so a single pass of your golden set is an anecdote, not a measurement. Run the set n times, say five, and report the variance, not just the mean. A chain that averages 90% but swings from 80 to 98 between runs is not a reliable chain.

It just got a good roll the day you demoed it. While you are at it, set the critique stage to a low temperature so your safety check isn't the flakiest part. Log every intermediate output so a failure is actually debuggable, and add a confidence gate at finalize that abstains rather than guesses. Let me make it concrete. The task is an internal HR policy assistant, the kind of thing where a wrong answer about parental leave is genuinely bad.

The chain goes like this. Retrieve the top five policy chunks with embedding search. Draft an answer that cites chunk IDs. Run a critique call that asks if every sentence traces to a cited chunk and to list any that do not. Then finalize by revising the unsupported sentences or returning a message saying you could not find this in policy.

Here's the measurement plan on a 200-question golden set. **On-screen reference block (measurement, 200-question golden set):** - **Retrieval:** recall@5 = 0.92. The right chunk is in the top 5 for 92% of questions. This sets the ceiling for the whole chain, so it is measured first. - **Draft groundedness:** 0.81 of claims supported before critique. - **Critique catch rate:** of drafts with an unsupported claim, critique flags 74%.

- **End-to-end:** correct-and-grounded rate = 0.86, with a safe abstention on 8% and a wrong-but-confident rate of 6%. - **Stability:** across 5 runs, end-to-end ranges 0.83 to 0.88, reported as 0.86 plus or minus 0.02. Walk it through, because the story these numbers tell is the whole point. Start with retrieval at 0.92 recall-at-5. I measure it first on purpose, because it sets the hard ceiling for everything downstream.

If the right chunk isn't in the top 5, no amount of clever drafting or critiquing can save that question. The answer simply isn't in front of the model. 0.92 is decent, so retrieval isn't my problem. Draft groundedness is 0.81 before critique, meaning about one claim in five isn't properly supported by the passages. That is exactly why the critique stage exists.

The critique catch rate tells us that of the drafts that had an unsupported claim, critique caught 74% of them. Not all. A quarter slip through, and that quarter is where my residual risk lives. Now for the end-to-end numbers. 86% correct and grounded. 8% safe abstentions, where the assistant said the information was not in policy rather than guess.

I count that as a good outcome, not a failure. Then there is 6% wrong-but-confident, which is the number I lose sleep over, because that is a confident wrong answer about company policy going to an employee. I track that 6% completely separately from the 8% abstention, because they are opposite behaviours and blending them would hide the dangerous one.

For stability, across five runs it moves between 0.83 and 0.88, so I report 0.86 plus or minus 0.02. I don't report a clean-looking 0.86 that pretends the chain is deterministic. The read from all of this is clear. Retrieval is fine, the draft leaks some claims, and critique isn't catching enough of them. So my next iteration isn't a vague push to improve the chain.

It is specifically to strengthen the critique rubric and lower its temperature, because the numbers point right at that box. That is the practical difference decomposition buys you. I would also be concrete about what strengthening the critique actually means, because that is the follow-up an interviewer asks. Right now critique catches 74% of unsupported claims. To lift that, I would rewrite its prompt to check sentence by sentence rather than judging the answer as a whole.

A whole-answer check lets a bad sentence hide next to three good ones. I would give it a couple of worked examples of the exact hallucination pattern I am seeing. I would also drop its temperature so the same draft gets judged the same way twice. Then I re-run the same 200-question golden set and I am looking for two things.

I want the catch rate going up, and the wrong-but-confident rate coming down, without the abstention rate ballooning. An over-eager critique that flags everything just turns into hedging. That last check matters, because it is easy to fix one number by quietly breaking another, and the golden set is what stops me fooling myself. Here is what makes them lean in.

You measured each stage, so you can name the weakest stage out loud. In this case it is counterintuitive. Retrieval is fine, critique is the problem, and you would never know that from an end-to-end number. You reported variance across runs, which tells them you actually understand these chains are non-deterministic and don't trust a single pass. You tracked wrong-but-confident as its own number, separate from safe abstentions, which shows you know which failure actually hurts.

You also built in a confidence gate that abstains, which says you would rather the system say it doesn't know than hallucinate a policy. Those are the reflexes of someone who has shipped one of these. Now for the ways people fall down. The first is reporting one end-to-end number and no per-stage metrics. When it is 86%, you have absolutely no idea whether to fix retrieval, drafting, or critique, and you are reduced to guessing.

The second is ignoring that retrieval recall sets the ceiling for everything downstream. People pour effort into prompt-tuning the draft while the real answer isn't even being retrieved. The third is measuring one run and calling it reliability, which is just a demo with a fancier word attached. All three feel like progress and none of them are. So pull it together.

Design the chain and justify every box. Name what each stage can get wrong. Measure every stage separately plus the end-to-end, so a failure is attributable to a specific box, not a vague shrug. Run the golden set n times and report variance, because one run is an anecdote. Always track wrong-but-confident apart from safe abstentions, because one of those is a disaster and the other is the system behaving well.

The core concept to carry in is to decompose the chain, measure every stage plus end-to-end, run it n times for variance, and track wrong-but-confident separately from safe abstentions.

Keep learning