…
AI Product Case Questions

What metrics did you use to evaluate your LLM's performance?

A worked answer to a real AI PM interview question: which metrics did you use to evaluate your LLM's performance?

Transcript

Read the full transcript (1,298 words)

[INTERVIEWER] What metrics did you use to evaluate your LLM's performance? "What metrics did you use to evaluate your model?" And the answer that quietly fails is "accuracy." One number, no task, done. The strong answer splits the world in two: offline evals on a fixed golden set before you ship, and online metrics on real users after. You keep those two separate, you tie every metric to the actual task, and you know, out loud, that the two can disagree with each other.

That disagreement is the whole reason you run both. Let me walk it. This question is checking whether you evaluate like an engineer or wave your hands. So the first move, before you name a single metric, is to clarify the task. "Performance" is meaningless on its own. A classifier, a summariser, a RAG question-answering system, and an autonomous agent are all scored completely differently.

So you pin the task first, then you pick metrics that fit it. Naming metrics before the task is the tell of someone who's memorised a list. So start with offline. This runs on a versioned golden set, say five hundred to a thousand labelled examples, and it should cover both the common head cases and the nasty edge cases, the ones that actually break.

Now match the metric to the task. For deterministic tasks, where there's a right answer, you use accuracy, precision, recall, F1, exact match, the classic set. For free-form generation, the old reference metrics like ROUGE and BLEU are weak proxies for real quality, so you lean on LLM-as-judge scored against a written rubric instead. For RAG specifically, you want faithfulness, is the answer actually supported by the retrieved context, plus answer relevance, and context precision and recall.

That's the RAGAS set, and naming it signals you've done this. And for safety, you track policy-violation rate, toxicity rate, and whether refusals are correct, because a model that refuses everything is safe and useless. Now flag a caveat, because using a model to grade a model is a real thing you have to handle carefully. LLM-as-judge is cheap and it scales, which is why everyone uses it.

But it has known biases. Position bias, where it favours whichever answer came first. Verbosity bias, where it rewards longer answers regardless of quality. And self-preference, where a judge model rates outputs from its own family higher. So you don't trust it blind. You calibrate it against human labels on a sample first, confirm it agrees with people, and only then let it grade at scale.

Mentioning this is a strong signal, because it shows you know the tool's failure mode. Then online, once real traffic is flowing. These come in tiers. First, product outcomes, the ones the business actually cares about: task success or completion rate, deflection rate, time to resolution. Second, explicit feedback the user hands you: thumbs-up rate, regenerate rate, edit or correction rate, how often they fix your answer.

Third, implicit signals: acceptance rate for code or suggestions, copy rate, retention. And fourth, guardrail metrics that stop a quality win from hiding a cost disaster: p95 latency, cost per query, hallucination-report rate, escalation rate. Without those guardrails a change that doubled your latency looks like a clean win. Now connect them, because this is the judgement they're really after.

Offline is your fast gate. It runs on every prompt change and every model swap, it's cheap, it's repeatable, and it catches regressions before they ever reach a user. Online is the truth, it's real value on real people, but it's slow, it's noisy, and it needs traffic to read anything. So the workflow is: regression-check offline first, then ship the change behind an A/B test, then read the online numbers.

Offline gates, online decides. And say the tradeoff plainly. Offline is fast, cheap, and repeatable, but it's a proxy, and proxies drift from real value. That's Goodhart's law, the moment you optimise the metric it stops measuring the thing. Plus benchmark contamination, where the test data has leaked into training and the score is fake. Online is real value, but it's slow and noisy and you can't iterate on it hour by hour.

Neither one alone is enough, which is exactly why you run both. Let me make it concrete with a Snap-style assistant, think My AI. Offline, you build an eight-hundred-example golden set, you use an LLM-judge rubric for helpfulness, and you add a safety classifier for policy violations with a target of ninety-nine percent policy-safe, because on a consumer product a safety miss is the thing that ends up in the press.

Online, you run an A/B test on conversation-continuation rate, thumbs-down rate, and messages per conversation, with guardrails of p95 latency under two seconds and a hard ceiling on cost per message. And here's the reason you insist on keeping both. On one real change, offline faithfulness jumped from ninety-two percent to ninety-six percent, looked like a clear win. But online, thumbs-up only crept from seventy-one to seventy-three percent.

That gap, big offline move, tiny online move, is exactly the divergence that stops you ever trusting the offline number on its own. If you'd shipped on offline alone, you'd have celebrated a win the users barely felt. A sharp interviewer will push on the golden set itself, because it's the foundation everything else sits on. The follow-up is usually "how do you build it, and how do you keep it honest?" So have an answer.

You build it from real production traffic, not made-up examples, because synthetic cases don't hit the edges that actually break. You deliberately over-sample the hard and rare cases, the ones a random sample would barely include, since those are where regressions hide. You version it, so a score from March is comparable to a score from June. And you refresh it, because every production failure you catch, every user complaint, gets added back in, which is how the set stops a shipped bug from ever shipping twice.

The second thing they may probe is inter-rater agreement: if humans labelling your golden set only agree with each other seventy percent of the time, then a model scoring seventy-five percent is already at the ceiling of what the labels can even measure, and chasing higher is chasing noise. Knowing that ceiling exists, and measuring it, is a strong signal that you've run a real eval and not just quoted one.

So what makes them lean in? First, you separated offline evals from online metrics and explained why you genuinely need both. Second, you knew the RAG-specific metrics, faithfulness and context recall, not just accuracy, so it's clear you've evaluated a real generative system. And third, you raised the LLM-judge bias problem and said you'd calibrate against humans. Those three together read as someone who's actually run an eval loop, not read about one.

The traps are the mirror image. Trap one, one number for everything, usually "accuracy," with no task attached. Trap two, trusting offline benchmarks as if they equal user value, when Goodhart and contamination both say they don't. And trap three, no guardrail metrics, so a change that doubled latency or cost sails through looking like a pure quality win. If your evaluation can't catch a regression it created, it isn't really evaluation.

So pull it together. Clarify the task first. Offline, run a golden set with task-matched metrics, use LLM-as-judge for free-form but calibrate it, and add the RAG and safety metrics where they apply. Online, watch product outcomes, explicit and implicit feedback, and guardrails. Gate on offline, decide on online, and expect them to disagree. The one line to carry in: offline evals on a golden set gate every change, online metrics on real users are the truth, and you keep both because they disagree.

Keep learning