…
Interview questions by skill

LLM Evaluation Interview Questions

How to tell whether an LLM feature is good enough to ship: eval specs, eval design, golden datasets and acceptance criteria for output that varies. These questions come up for AI engineers, data scientists and AI product managers alike. The answers focus on the decisions behind an eval, not on code. 104 questions across 4 subtopics, each with a model answer.

Asked in interviews for:AI EngineerData ScientistData AnalystAI Product Manager

Writing an LLM Eval Spec

How to define what good looks like for an AI feature: test cases, human baselines, how often evals run and how to stop them being gamed by prompt tuning. Asked of AI PMs, AI engineers and data scientists.

  1. What is an eval spec and who is its audience?
  2. List the components of a complete eval spec.
  3. How do you define a task-level success criterion for a summarization feature?
  4. Write a scoring rubric for the quality of a generated customer support reply.
  5. Describe the difference between an eval spec and a test plan.
  6. How many examples belong in a first eval set and how do you choose them?
  7. Explain how you would build an eval set that includes adversarial cases.
  8. What is the role of a human baseline in an eval spec?

All 27 Writing an LLM Eval Spec questions

LLM Eval Design

How to design evaluations a product team can run: what makes an eval relevant to users, how to know a suite is good enough, and how to evaluate safety on a small labelling budget. Asked of AI PMs, AI engineers and data scientists.

  1. What makes an eval product-relevant rather than research-relevant?
  2. Design an eval for a feature that drafts email replies.
  3. How do you decide between automated evals and human review?
  4. Explain the tradeoffs of LLM-as-judge for a product team.
  5. How do you validate that your judge model agrees with human raters?
  6. Describe a rubric that a non-technical reviewer could apply consistently.
  7. What is inter-rater reliability and why should a PM care?
  8. How do you design evals that catch regressions rather than just measuring level?

All 25 LLM Eval Design questions

Golden Datasets and Test Sets

What a golden dataset is, how big it should be, who owns it and how it leaks into development. A core topic for AI PMs, data scientists and ML engineers.

  1. What is a golden dataset and why does the PM usually own it?
  2. How do you construct a first golden set with no production traffic?
  3. Describe the composition of a golden set: what proportion should be edge cases?
  4. How do you keep a golden set representative as your user base changes?
  5. Explain the risk of a golden set that engineering can see during development.
  6. What is a holdout set and when would you use one for an AI product?
  7. How do you handle labelling disagreement inside a golden set?
  8. Describe a process for adding newly discovered failure cases to the golden set.

All 26 Golden Datasets and Test Sets questions

Acceptance Criteria for LLM Output

How to decide when an AI feature is ready to ship when its output varies: pass bars, comparison with human performance, and criteria that can be met on paper while the product is still bad.

  1. Rewrite this criterion to be testable: the model should not hallucinate.
  2. How do you express an acceptance criterion as a rate rather than an absolute?
  3. What is the difference between a threshold criterion and a distributional criterion?
  4. Write acceptance criteria for an AI feature that extracts fields from an invoice.
  5. How do you set a pass bar when human performance on the same task is 92 percent?
  6. Describe acceptance criteria that account for the severity of different error types.
  7. How would you write criteria for a feature where the worst case matters more than the average?
  8. What acceptance criteria would you set for latency at the 95th percentile and why not the mean?

All 26 Acceptance Criteria for LLM Output questions