…
AI Product Case Questions

How do you evaluate a prompt? (live practical probe)

A worked answer to a live practical AI PM interview question: how do you evaluate a prompt?

Transcript

Read the full transcript (1,518 words)

[INTERVIEWER] How do you evaluate a prompt? (live practical probe) This one's a live probe, and it's checking a single thing. Do you evaluate a prompt with a repeatable method, or do you eyeball two outputs and say "yeah, looks good"? That's the whole test. Because "looks good" is how amateurs work, and it's exactly the reflex the interviewer is trying to catch.

The strong answer has a shape they can hear coming: define success criteria, test against a fixed set, compare to a baseline number, and change one variable at a time. If you can describe that loop and then run one example through it, you've passed the probe. Prompts are the surface a PM touches most on an AI product, and they're deceptively easy to fiddle with badly.

You tweak a word, the output looks nicer on the one example you tried, you ship it, and you've quietly broken three edge cases you never looked at. This question is testing whether you treat a prompt like code you'd test, or like a lucky charm you keep rubbing. Over the next few minutes I'll give you the five-step method, walk a real extraction example through it with actual scores, and flag the moment where a smart candidate says "this failure isn't even a prompt problem," which is the line that lands.

The method is simple enough to say in one breath, and that's deliberate: an interviewer wants to hear that you carry it around as a habit, not that you'd invent it fresh each time. Step one, define success criteria first, before you touch the prompt. What does a good output actually look like, concretely? For an extraction prompt that might be: correct fields, valid format, and no invented values.

Write it down. You cannot evaluate against "good vibes," and the act of writing the criteria often shows you what you didn't know you cared about. Step two, build a small fixed test set. Twenty to fifty representative inputs, and here's the important part, include the boring middle and the awkward edges: empty input, ambiguous input, adversarial input. Fixed, so every version of the prompt runs the exact same inputs.

This single move is what turns vibes into a method, because now "better" means the same set scored higher, not "the one example I retried feels nicer." Step three, establish a baseline. Run the current prompt, or a naive first draft, against the set and score it. Now every change you make gets measured as better or worse than a real number, instead of against a fuzzy memory of what the outputs used to look like.

Without a baseline you're just telling yourself a story. Step four, score repeatably. Deterministic where you can: format valid, exact field match, those are free and objective. For subjective quality, use a written rubric applied consistently, or an LLM-judge with that rubric plus a human spot-check. And run each input a few times, because the model is stochastic, and you want to notice the failures that only show up sometimes.

A prompt that's right four times out of five isn't "right," and one run would hide that. And be ready for the scoring follow-up, because on a subjective task the interviewer will ask "how do you even score that?" The answer is you turn the fuzzy quality into a written rubric with levels. For a summary prompt that might be a 1-to-5 scale where 5 is accurate and complete, 3 has a minor omission, and 1 invents something.

Once the rubric exists, two people, or an LLM-judge, grade the same output the same way, and the score means something. Without the rubric, "quality 4 out of 5" is just a feeling with a number stapled to it. The rubric is what makes a subjective score repeatable, which is the entire game here. Step five, change one thing, re-run, compare.

Edit the prompt, run the same set, compare to baseline. One variable at a time, so when the number moves you actually know what moved it. And keep the edge cases in view the whole time, because the classic mistake is a change that lifts the average while quietly breaking a corner case you stopped watching. Let me run one all the way through.

The prompt: "Extract the invoice total, currency, and due date from this email." **On-screen reference block (prompt eval run):** - **Criteria:** total is the correct number, currency is a valid ISO code, due date is ISO-8601, and the model returns null rather than guessing when a field is absent. - **Test set:** 30 real emails, including 5 with no due date, 3 in non-USD currencies, 2 with two candidate totals (subtotal and total), 1 with a scanned-image body and no parsable text.

- **Baseline:** current prompt scores 22/30 fully correct. Misses cluster on the two-totals case (picks subtotal) and the missing-date case (invents a date). - **Iterate:** add "if two amounts appear, take the one labelled total or the larger final amount" and "if a field is absent, return null, never guess." Re-run the same 30. Score 27/30, invented-date failures go to zero.

Walk it through. The criteria are specific and testable, and notice the sneaky one: the model must return null when a field is missing, not guess. That's the criterion that catches the most dangerous failure, a confident invented due date, and I wrote it down before I ever looked at an output. The test set is 30 real emails, but it's built, not scraped at random.

Five have no due date, three are in other currencies, two have both a subtotal and a total to trip the model up, and one is a scanned image with no readable text at all. Those aren't the happy path, they're the exact places extraction breaks. The baseline is 22 out of 30 fully correct, and because I have a real test set, I can see where the 8 misses cluster: it picks the subtotal on the two-total emails, and it invents a date when there isn't one.

That clustering is the whole value of the fixed set. It doesn't just tell me the score, it tells me what to fix. So I iterate, and I change exactly two instructions, one for each failure cluster: take the amount labelled total or the larger final amount, and return null for absent fields instead of guessing. I re-run the same 30, and the score goes to 27, with the invented dates dropping to zero.

Now here's the line that impresses. The remaining 3 misses are all the scanned-image email. And I say out loud: that's not a prompt problem, that's a retrieval problem, there's no text for the model to read, so I'd route those to OCR upstream rather than keep torturing the prompt. Knowing when to stop tuning the prompt is as much a signal as the tuning itself.

Here's what makes them lean in. You brought a fixed test set with deliberate edge cases, not two hand-picked examples that happen to work. You established a baseline number, so your improvement is measured, not asserted. You changed one variable and re-ran the same set, so you can actually attribute the gain. And you knew when the residual failure wasn't a prompt problem at all and routed it elsewhere.

That last one is the tell that you've done this for real, because rookies keep polishing a prompt against a problem no prompt can solve. There's a bonus signal too, if it comes up naturally: mentioning that you'd version the prompt and its test set together. When you change the prompt, the version bumps, and you keep the score attached to that version.

So six months later, when someone asks "why is this prompt written this weird way," the test set answers it: because version three fixed the two-totals bug and here's the run that proves it. Treating a prompt like a tracked artifact rather than a sticky note is a quiet mark of someone who's maintained one in production. Now the ways people fail the probe.

The first is literally the thing it's testing for: "I try a few inputs and read the outputs." Say that and you've confirmed you don't have a method. The second is no baseline, so every change is judged against a vague memory of before, which means you can't actually tell if you improved anything. And the third is optimising the average while never checking whether an edge case regressed, so you ship a prompt that's better on paper and worse on the cases that were already hard.

So the loop, clean. Criteria first, written down. Then a fixed test set that includes the awkward edges, not just the middle. Then a baseline number, so improvement is measured. Then change one variable, re-run the same set, and compare, keeping the edge cases in view the whole way. And know when a failure has left prompt territory and belongs somewhere else.

Carry this line in: criteria, then a fixed test set with edge cases, then a baseline number, then change one variable and re-run the same set.

Keep learning