…
AI Product Case Questions

Walk through eval pipeline design and human-in-the-loop setup

A worked answer to a real AI PM interview task: walk through eval pipeline design and a human-in-the-loop setup.

Transcript

Read the full transcript (1,954 words)

[INTERVIEWER] Walk through eval pipeline design and human-in-the-loop setup. This is the full loop from raw data all the way to a decision, and there is one part the interviewer is really watching for, which is the human-in-the-loop. It is not about whether you mention humans, but where they sit, on which cases, and how their judgement flows back into the model or the prompt.

Here is the trap on both sides. An answer with no humans fails because you are trusting an automated grader nobody ever checked. An answer with humans on everything also fails because it does not scale past a demo. The strong answer puts humans in exactly the right place, focusing on the cases automation is worst at and nowhere else.

This question tests whether you can design a standing process rather than a one-off measurement. Anyone can score a model once. Running an evaluation pipeline that samples the right data, grades it cheaply at scale, spends scarce human attention wisely, and feeds what it learns back in is the actual job of owning an AI feature. The interviewer wants to see it drawn as a cycle with the feedback arrow closing and a real rule for who reviews what.

Over the next few minutes I will walk through the five stages, show you a worked example that routes a thousand cases a week down to about 120 for humans, and point out the moves that prove you have run one of these instead of just sketching it. Draw it as a cycle rather than a line because the output feeds the next round.

That framing alone puts you ahead of most answers that just draw a straight arrow stopping at a score. Stage one is sampling the right data. Do not evaluate on a random dump because a random dump is ninety percent easy head traffic and your numbers drown out the cases that matter. Stratify it by taking the head intents by volume plus a deliberate over-sample of the rare and high-risk cases so they are not buried.

Pull from real production logs, keep a frozen held-out set so you can compare versions, and add a small adversarial set you wrote on purpose. Say the nuance out loud. The sample has to match production distribution for your online metric, but it should over-weight the tail for your safety metric. Those are two different sampling goals and a good answer keeps both slices.

Stage two is running auto-eval first at scale. Run every sampled case through automated checks. Use deterministic ones wherever you can, like schema validation, tool-call correctness, and exact-match, because those are free and exact. Use an LLM-as-judge with a written rubric for the semantic cases where there is no single right string. The point of auto-eval is coverage. It is cheap so it covers the entire sample, and crucially it produces not just a score but a confidence or an agreement signal.

That signal is what you will use to decide who needs a human. Stage three is routing to humans by disagreement and risk. Humans are the scarce resource so you spend them where automation is weakest and only there. There are three routing rules. First, cases where the LLM-judge is low-confidence or disagrees with a reference. Second, every safety-flagged case with no exceptions.

Third, a fixed random audit of five to ten percent to keep the judge honest even where it is confident. A judge that is confidently wrong is the one that hurts you and confidence alone will not surface it. That is the human-in-the-loop and the word that matters is targeted rather than blanket. Stage four is feeding the human labels back, which is the arrow most people forget to draw.

Human judgements do three jobs. They become new golden-set labels to grow your test set over time. They surface failure clusters of similar mistakes that drive the next actual change. And they recalibrate the automated grader. Draw the arrow from human labels back to both the golden set and the judge. Say both out loud because feeding the set is obvious while recalibrating the judge shows depth.

Let me push on that recalibration point because it is where an interviewer will test you. You have an LLM grading thousands of cases, so how do you know it is any good? You measure judge-versus-human agreement on the slice humans reviewed and treat it as a real metric with a threshold. Say you hold it at 0.85. When it drops below that, you do not panic and rip the judge out.

You read the disagreements to find the pattern because there is almost always a pattern. Maybe the judge is too harsh on short answers or it is rewarding a confident tone over correctness. You fix the rubric with a couple of concrete examples of the disagreement, re-run, and confirm agreement climbs back. That loop keeps the grader trustworthy. Being able to describe it is the difference between simply using an LLM-judge and running one you actually validate.

Stage five is adjusting and deciding. The output of the pipeline is not a number but a decision to change the prompt, fine-tune, add a retrieval source, or tighten a guardrail. Then re-run the same frozen set to confirm the change helped and did not quietly regress another task. Gate the release on the frozen set and then A/B test it online.

Set a cadence of weekly offline and continuous online sampling so this is a standing process the team runs instead of a heroic one-off before a launch. Let me ground it. The feature is an AI that categorises and routes incoming sales leads, so a wrong route means a hot enterprise lead lands in the wrong queue and goes cold.

Here is the pipeline, stage by stage. **On-screen reference block (lead-routing eval pipeline):** - **Sample:** 1,000 leads a week, stratified so rare "enterprise" and "partner" intents are 20% of the sample despite being 3% of volume, plus 100 adversarial edge cases (ambiguous, multi-intent). - **Auto-eval:** exact-match on the routing label against the 600-lead labelled golden set, LLM-judge with a rubric for the unlabelled remainder, schema check on the structured output.

- **Human loop:** two reviewers handle all judge-vs-reference disagreements (about 8% of volume), every "enterprise" route (money at stake), and a 5% random audit. Roughly 120 leads a week to humans, not 1,000. - **Feedback:** reviewer labels grow the golden set, judge-vs-human agreement tracked weekly (alert if it drops below 0.85), mis-routes clustered by cause. - **Adjust:** last cycle found "partner" leads mis-routed to "SMB", fixed with two rubric examples plus a retrieval field for partner domains.

Re-ran the frozen 600, routing accuracy went 0.88 to 0.93, no regression on other intents. Shipped, then A/B on live routing. Walk it through. The sampling is the first tell. I am pulling a thousand leads a week, but I am deliberately making enterprise and partner leads twenty percent of the eval even though they are three percent of real traffic.

Those are the leads worth real money and I need a stable read on them rather than four noisy examples. I also keep the hundred adversarial cases separate. Those are the genuinely ambiguous ones where routing quietly falls apart. Auto-eval covers all of it cheaply. I use exact-match where I have a labelled reference, an LLM-judge with a rubric for the rest, and a schema check so the output is even parseable.

Now for the human loop, and watch the numbers because this is the core of the answer. Two reviewers do not see a thousand leads. They see the eight percent where the judge and the reference disagree, plus every enterprise route because money is on the line, plus a five percent random audit to keep the judge honest. That is roughly 120 leads a week instead of a thousand.

These specific cases give us the most critical signal because they target the exact blind spots of the automated grader. Then the loop closes. Those reviewer labels grow the golden set so the test next month is richer. I track judge-versus-human agreement every week and alert if it falls below 0.85 because a drifting judge is a silent failure. I also cluster the mis-routes by cause.

Last cycle that clustering found something specific where partner leads were getting mis-routed to SMB. The fix was not a retrain. It was two example rows in the rubric plus a retrieval field for partner domains. I re-ran the frozen 600 and routing accuracy went from 0.88 to 0.93. I confirmed no other intent regressed, then shipped and A/B tested it live.

That is the full cycle. Notice the change was cheap and precise because the pipeline told me exactly what was broken. One more thing to have ready because interviewers love this follow-up is how you staff it. The honest answer is you do not need an army. Two part-time reviewers can handle 120 cases a week comfortably and their time is worth it because those are the labels that keep the whole system calibrated.

If human volume ever creeps up, that is a signal in itself. It usually means the judge got worse or the model started failing in a new way, so rising human load becomes an alarm rather than just a cost. I would also give the reviewers a tight rubric of their own so two humans grade the same borderline case the same way.

Human graders drift too and an inconsistent human label poisons the golden set just as badly as a bad model output. Here is what makes them lean in. Your humans are targeted at disagreement, risk, and a small random audit instead of being thrown at every case, so the design actually scales. The loop closes properly. You use human feedback to correct grader drift and expand your test set, and you monitor judge-versus-human agreement instead of assuming the grader stays right forever.

Your sampling over-weights the risky tail while keeping a production-matched slice for the online metrics, which shows you understand those are two different jobs. You also keep a frozen set for release gating separate from the growing labelled set so your version comparisons stay honest. Each one is a small signal that you have operated a pipeline instead of just diagramming one.

Now for the ways it falls apart. The first trap is saying a human reviews the outputs with no rule for which outputs. That sounds responsible but dies the moment volume goes past a demo because you cannot review everything. The second trap is never checking the LLM-judge against humans. The entire pipeline ends up trusting an unvalidated grader, and if that judge is quietly wrong, every number downstream is wrong too.

The third trap is random-only sampling. It feels fair but never surfaces the rare high-risk cases where the model actually fails because they are too rare to show up in a random draw. All three are comfortable and all three leave you blind exactly where it counts. So here is the whole cycle. Sample stratified, over-weighting the risky tail and keeping a frozen held-out set.

Auto-eval everything cheaply with a confidence signal. Send humans only the disagreements plus the risk cases plus a small random audit. Feed their labels back to expand the test set and correct grader drift, and watch judge-versus-human agreement like a metric in its own right. Then adjust, re-run the frozen set, gate the release, and A/B test online on a standing cadence.

The line to carry in is to sample stratified, auto-eval everything, send humans only the disagreements plus risk plus a small audit, and then feed their labels back to expand the set and correct the grader.

Keep learning