Define Hallucinations in LLMs
Transcript
Read the full transcript (1,015 words)
[INTERVIEWER] Define hallucinations in LLMs. Someone asks you to define hallucinations in a large language model, and the weak answer is that it's when the model gets something wrong. That's not incorrect, but it won't move a hiring manager either. What separates a strong answer is explaining why the model does it, and then showing that the risk isn't one flat number.
It scales with the stakes of the feature you've put it in. Let me give you the version that lands. Look, this is a knowledge check, but it's really testing whether you understand the machine well enough to make product calls on top of it. If you think a hallucination is a bug you can patch, you'll design the wrong product.
If you understand it's a property of how the model works, you design guardrails instead. That's the judgement they're probing. So start with a clean definition. A hallucination is model output that reads as confident and coherent, but is either unsupported by the source you gave it, or just false about the world. Fluent and wrong at the same time.
That's the whole idea. It also helps to name two flavours. Intrinsic hallucination is when the answer contradicts the source text you handed the model. Extrinsic is when it invents something you can't check against any source at all. Naming those two shows you've thought about it as more than a single blob of mistakes. Now for the mechanism, and this is the part most candidates skip.
A language model is a next token predictor. It's trained to produce the most plausible continuation of the text, not the most true one. By default it has no lookup step, and no internal sense of not actually knowing this. So when it hits a gap, it doesn't stop. It fills the gap with whatever sounds most likely. The confidence you see isn't confidence in the facts.
It's just fluent text. Once you say that out loud, the interviewer knows you understand the model rather than just its symptoms. Here's the part that separates you. The exact property that makes the model useful, its willingness to generalise and fill in gaps smoothly, is the same property that makes it hallucinate. You can't have maximum fluency and full coverage and zero hallucination all at once.
If you push the model to abstain more, to say it doesn't know whenever it's unsure, you do cut hallucinations. But you also raise refusals, and the thing gets less helpful. It behaves like a precision versus recall dial. Turn one up, the other slides down. So hallucination isn't a bug waiting for a fix. It's a setting on a dial you're choosing where to put.
Let me make that concrete, because the stakes point is the whole game. Legal research is the clean case. The Stanford RegLab study by Magesh and colleagues in 2024 tested purpose built legal AI tools like Lexis and Westlaw. These have retrieval bolted on specifically to stop this, yet they still hallucinated on roughly 17 to 33 percent of queries.
You've probably heard of Mata versus Avianca in 2023, where a lawyer filed a brief citing six cases ChatGPT invented wholesale. He got sanctioned. Hold that next to a brainstorming feature. Same model behaviour, but a made up idea in an ideation tool costs nothing. You shrug and move on. A confident wrong citation in a legal filing ends a career.
Same failure, wildly different cost. That gap is the product manager's point. If the interviewer wants to go deeper, the follow up is usually about whether you can detect it. The honest answer is partly. You can catch a lot with grounding checks to see if the claim matches the source, and with self consistency by sampling the model to see if the fact stays stable.
But you can't fully detect it, because the model's own confidence is not a reliable signal. It sounds exactly as sure when it's right as when it's making something up. That flat confidence makes hallucination dangerous. A model that hedged whenever unsure would be easy to handle. This one doesn't, so the burden falls on the product to add the checks the model can't add for itself.
There are three specific things that make a hiring manager lean in. First, you explained the mechanism, next token prediction with no grounding, not just the symptom. Second, you framed it as a tradeoff against helpfulness, not a defect to be fully removed. Third, you tied the risk to the stakes of the specific use case, and then named the mitigation those stakes justify.
In a high stakes feature you gate the output with citations, a confidence threshold, or human review. In a low stakes one you don't bother. That connection, from mechanism to product decision, is the signal they're listening for. Let's look at the traps, because they're easy to fall into. The first is defining it as the model making mistakes and stopping there, with no mechanism and no tradeoff.
Flat and forgettable. The second is claiming it can be fully eliminated with a better prompt or a bigger model. It can't, and saying so tells them you don't get the mechanism. Bigger models hallucinate less, but they still hallucinate. The third trap is treating every feature as equally risky, instead of tying the risk to the use case. If you'd defend a medical answer the same way you'd defend a brainstorm, you've missed the point of being the PM in the room.
So pull it all together. A hallucination is confident, fluent output that isn't grounded in your source or in fact. It happens because the model predicts plausible text, not true text. It trades off against helpfulness on a precision versus recall dial. The risk scales with the stakes of the feature, which is what decides your guardrails. The core takeaway to carry into the room is that a hallucination is confident text lacking grounding, and the product question is never about removing it entirely, but rather how much a wrong answer costs in that specific context.