How do you handle hallucinations in a GenAI model in production?
Transcript
Read the full transcript (1,281 words)
[INTERVIEWER] How do you handle hallucinations in a GenAI model in production? "How do you handle hallucinations in production?" And the trap is to reach for the one clever trick, the perfect prompt, the magic setting. There isn't one. A real production answer is a layered defence. You ground the model, you constrain it, you verify it, you let it abstain, and you catch whatever's left with humans and evals.
The strong signal here is that you talk in layers, with a measurable target attached, not one silver bullet. Let me walk the stack. This question is really asking whether you think like an operator or a hobbyist. So set the goal before you touch a single technique. You are not zeroing hallucinations out. You're reducing them and containing them.
Pick a target, say faithfulness of at least ninety-eight percent on a golden set, and pick a policy for the residual, the ones that slip through. Everything after this is in service of that number. So layer one, and this is the biggest lever by a distance: grounding. Attach retrieval. RAG. You fetch authoritative passages from your own trusted source and you instruct the model to answer only from those passages, and to cite them.
This is the move that kills most extrinsic hallucination, the invented stuff, because now the answer has an actual source behind it instead of coming from the model's memory. If you only remember one layer, remember this one. Layer two is constraints. Your system instruction does real work here. Tell it plainly: if the answer isn't in the provided sources, say you don't know.
Then drop the temperature for factual tasks, somewhere from zero to about zero point three, so the model stops getting creative where you don't want creativity. And ask for structured output where you can, because a structured answer is far easier to check automatically than a free paragraph. None of these are glamorous, but they're cheap and they stack. Layer three is verification.
This is a second pass that checks the claim against the retrieved context before the user ever sees it. A few ways to do it. You can run an NLI or faithfulness classifier that asks "does the context actually support this sentence." You can do a citation check, literally confirming the cited passage supports the claim it's attached to. Or you can sample the model a few times and check for self-consistency, because a fact it's sure of tends to stay stable across samples, and a hallucination wobbles.
Layer four, confidence and abstention. Take the retrieval score, or the verifier's score, and use it as a gate. Below your threshold, the model does not answer confidently. It abstains, or it hands off. This is the layer people forget, and it's the one that turns "the model guessed" into "the model knew when to stop." A system that knows what it doesn't know is worth far more than one that's occasionally brilliant and occasionally makes things up with the same tone.
Layer five is the human fallback. Low-confidence or high-stakes cases route to a person, rather than shipping a guess to the user. And layer six is evals and monitoring, which is what makes the whole thing a system instead of a hope. You keep a versioned golden set and you measure it offline, faithfulness and groundedness, before every single change.
Then online you watch the live signals: thumbs-down rate, correction rate, user-reported errors. And crucially, every failure you catch in production goes back into the eval set, so the same mistake can't ship twice. There's a seventh piece too, UX honesty: show the sources, surface uncertainty, and make correcting the answer a single click. Now say the tradeoff out loud, because it's real.
Every one of these layers adds latency and cost, and abstention lowers your coverage, the share of questions you'll actually answer. A verifier pass can roughly double your per-query latency and cost, because you're running the model's work twice. So you don't apply the heavy layers everywhere. You apply them where the stakes justify them. A medical answer gets the full stack.
A tone suggestion in a writing tool does not. Let me make it concrete. Picture a customer-support assistant grounded on the help centre. You run RAG over four thousand articles. Temperature at zero point two. The instruction is "answer only from these articles and cite the article id." Every draft goes through an NLI faithfulness check. And anything under zero point seven retrieval confidence gets routed to a human agent instead of answered.
Offline, you score a five-hundred-pair golden set every week against that ninety-eight percent faithfulness target, and you watch thumbs-down live. On a setup like this, hallucinated answers typically fall from somewhere around twelve percent down to about two percent, while the bot still deflects roughly forty percent of incoming tickets on its own. That's the shape of a real win: you didn't kill hallucination, you contained it to a level the business can live with, and you kept enough coverage to be useful.
Now a good interviewer pushes here, so be ready for it. The follow-up is usually: "your RAG grounding is on, and it still hallucinated. Why?" And the answer is that grounding stops extrinsic hallucination, the invented stuff, but it doesn't fully stop intrinsic hallucination, where the model contradicts the passage you actually gave it. It can misread the source, blend two passages, or over-generalise from them.
That's precisely why the verification layer exists, and why it checks faithfulness against the retrieved context specifically, not against the world. The second follow-up is "how do you know the retrieval was even good?" And there your answer is context recall on the eval set: if the right passage wasn't in what you retrieved, no amount of grounding instruction saves you, the model was working from the wrong source.
So grounding and verification handle different failures, and you need the evals to tell you which one broke. Naming that split, extrinsic versus intrinsic, retrieval failure versus generation failure, is what shows you've actually debugged one of these in the wild. So what makes them lean in? First, you led with grounding and evals, the two layers that do the most work, instead of opening with prompt tweaks.
Second, you named a measurable target and a monitoring loop, so this reads as a system with a number, not a vibe. And third, you acknowledged the latency and cost the defence buys, and you applied it by stakes rather than slapping the full stack on everything. That's an operator's answer. The traps here are common. The first is "we'll add a better prompt" as the entire answer.
One layer is not a defence, and it tells them you've never actually shipped this. The second is having no eval set, which means you have no way to know whether any of it is working, you're just hoping. And the third is forgetting the fallback: what does the product actually do at the moment the model shouldn't answer? If you can't answer that, you don't have a production system, you have a demo.
So the whole picture. Ground it with retrieval. Constrain it with instructions and temperature. Verify each claim against the context. Let it abstain below a confidence threshold. Fall back to a human on the hard cases. And measure all of it on a golden set with a live monitoring loop. Apply the expensive layers by stakes, not everywhere. Here's the line to carry in.
Ground it, constrain it, verify it, let it abstain, and measure it on a golden set. Hallucination handling is a layered system with a number attached, not a trick.