Whiteboard an Eval Framework for a Generative Feature
Transcript
Read the full transcript (1,760 words)
[INTERVIEWER] Whiteboard an eval framework for a generative feature. Whiteboarding an eval framework is really one test in disguise. Can you turn the word good into something you can actually measure? That's the whole thing. A generative feature does something fuzzy, someone says make sure it's good, and your job is to build the machinery that decides whether it is.
The strong answer builds a pipeline on the board from left to right. A task taxonomy, a golden set, a metric per task type, and a human loop for the cases automation can't judge. The weak answer waves at it and says we'll check quality and get some user feedback. One of those gets hired. The other doesn't. This question is testing whether you can operationalise quality for something that has no single right answer.
Anyone can eval a classifier because there's a label and you check against it. A generative feature is harder because good splits into correctness, tone, safety, and format. Each one needs a different yardstick. The interviewer wants to see you decompose the problem, commit to specific metrics, and build a loop that gets better over time instead of a single spot check.
Over the next few minutes I'll give you the four stage skeleton you draw on the board, a worked example with a real metric table, and the moves that tell a hiring manager you've actually run an eval rather than just read about one. First thing before you draw anything. Pick a concrete feature or ask for one. Don't eval in the abstract.
Say it's an AI that drafts customer service replies. Now the board has real tasks on it and everything you say gets specific. Then build left to right. Stage one. Define the tasks the feature must do. Give this a couple of minutes because it's the foundation. Break the feature into task types because different tasks need different metrics. This is the insight the weak answer misses.
For a reply generator, the task types are things like factual answers, tone matching, refusing unsafe requests, and correct formatting. Draw those as a column down the left of the board. Here's the point you say out loud. Good means something different for each of these. A good factual answer is correct. A good refusal is that it actually refused.
You can't score them with one number and pretending you can is the first mistake. Stage two. Build the golden set. This is your fixed and versioned set of inputs with reference answers or labels. Where does it come from? Two places. Mine it from real production logs, stratified so the rare intents are represented and not just the popular head.
Then add adversarial cases you write on purpose to probe the edges. Size it to hundreds per task type to start, not tens, or your numbers are noise. And version it because this is the part people forget. If your eval set moves every week you can never tell whether the model got better or the test just got easier.
A frozen set is what makes two model versions comparable. Let me be concrete about the stratification because it's the part people hand wave. If eighty percent of your real traffic is simple factual questions and three percent is safety critical, and you sample the golden set at production frequency, you'll get four safety examples and your safety number will be pure noise.
So you deliberately over sample the tail. You might make safety cases twenty percent of the golden set even though they're three percent of traffic. That's where the model actually hurts you and you need a stable read on it. You keep a separate production matched slice for the online metrics, but the offline safety set is weighted toward the danger on purpose.
Stage three. This is the core of the board. Pick a metric per task type. Spend real time here. Deterministic tasks get exact checks. Does it output valid JSON, does it call the right tool, exact match or a schema check. Semantic tasks where there's no single right string get either reference based scoring or an LLM as judge with a written rubric.
Safety tasks get two numbers, not one. A refusal rate and a false refusal rate, because a feature that refuses everything is safe and useless. The rule to say out loud is never one blended number for everything. A single quality score hides exactly the failure you need to see. Stage four. Add the human in the loop and be specific about where humans sit.
Automation can't be trusted on the cases that matter most, so route to humans in three situations. Every disagreement between the LLM judge and the reference. A random audit sample, say five to ten percent, to keep the judge honest even where it's confident. And every safety critical case. Then crucially draw the arrow back. Human labels recalibrate the judge and feed the next version of the golden set.
The humans aren't just graders. They're how the whole framework learns. And then close the loop. Offline eval on the frozen set gates a release. No task type ships below its threshold. Then you go online, shadow or A B test, with an online proxy metric. Something like a thumbs signal or the edit distance between what the model drafted and what the human actually sent.
Online disagreements get sampled back into the golden set. Say the honest bit out loud. Offline and online can disagree and when they do the online result wins because that's real users. Let me put it on the board properly. The feature is an AI that drafts customer service replies. Here's the eval framework, one row per task type. **On screen reference block:** Walk the panel across it row by row.
Top row, factual answers. Four hundred real tickets with reference answers, scored for correctness by an LLM judge against the reference. A human reviews every case where the judge and the reference disagree because those disagreements are where the judge is either wrong or the reference is stale. Both are worth knowing. Second row, tone and empathy. There's no right string for empathetic so I use a one to five rubric and an LLM judge applying it.
I gate on a mean of at least 4.0. A human audits ten percent at random every week because a tone rubric drifts and you want to catch the drift before it ships. Third row, the safety row. This is the one that separates people. One hundred and fifty adversarial cases I wrote on purpose. Two metrics, not one. Refusal rate targeting above ninety nine percent and false refusal rate targeting under two percent.
I say out loud that the false refusal number is a real cost. A model that refuses a legitimate angry customer because their message sounded harsh is a bad product. I refuse to hide that behind a single safety score. Bottom row, format and handoff. This one's deterministic so it's easy. Schema exact match on the structured output and a correct handoff rate for whether low confidence cases actually reach a human.
Spot check the failures. Then the release gate, said plainly. No task type below its threshold on the frozen golden set or it doesn't ship. Once it ships I A B test online. My online proxy for quality is agent edit rate, meaning how much a human agent changes the draft before sending it. If they're barely touching it the drafts are good.
If they're rewriting every one my offline scores were lying to me and those cases go straight back into the golden set. One more thing worth saying at the board because a panellist will ask it. Does this scale? The trick is that the human loop is targeted, not blanket. Say I run two thousand evals a week across the four task types.
The LLM judge covers all two thousand for pennies. Humans only touch the judge versus reference disagreements, which might be eight percent of volume, plus a ten percent random audit, plus every safety miss. That's three to four hundred cases a week going to humans, not two thousand. And those few hundred are the highest value labels in the whole system because they're exactly where automation is weakest.
When you can put a rough number on the human load like that you've shown the framework is something a real team could actually staff. It's not just a lovely diagram that dies on contact with a budget. Here's what makes them lean in. You used different metrics for different task types instead of one blended score. That's the whole insight of the question.
Your golden set is versioned and includes adversarial cases, not just happy path samples pulled at random. The human loop feeds labels back into the eval set so the framework improves itself instead of staying frozen at day one. And you called out that a false refusal is a genuine cost, not just a line item under safety. Each of those is a small tell that you've run an eval that mattered, where a wrong number cost something real.
Now the ways this goes wrong. The first is saying we'll use accuracy for a generative feature where there's no single right answer. That tells the interviewer you're thinking about classifiers rather than generation. The second is an LLM as judge with no human anchor. Nobody ever checks whether the judge itself is right and you've built a tower on an unverified grader.
The third is an eval set that changes every week. This quietly makes every model comparison meaningless because you can never tell if the model improved or the test got easier. All three feel fine in the moment and fall apart the second someone pushes. So the whole board one more time. Split the feature into task types because good means something different for each.
Freeze a golden set that's stratified and salted with adversarial cases. Pick a fit for purpose metric per type. Exact checks for deterministic tasks, a rubric or LLM judge for semantic ones, refusal and false refusal for safety. Build a human loop that both audits the judge and refeeds the set. Then gate the release offline and confirm it online where real users win any tie.
Carry this one line into the room. Split into task types, freeze a golden set with adversarial cases, use a fit for purpose metric per type, and build a human loop that both audits and refeeds the set.