…
Skip to content
Topics
On this page

LLM Hallucination: Causes and Fixes

LLM hallucination is output from a large language model that is fluent and confident but factually wrong, unsupported or invented. It occurs because the model generates statistically likely text rather than retrieving verified facts, so a plausible answer and a correct answer can look identical.

  • Fabricated facts: The model states names, statistics, citations or events that do not exist in any reliable source.
  • Invented APIs: In generated code, it calls functions, parameters or packages that no library actually provides.
  • Unfaithful summaries: It introduces details into a summary that the original document never contained.
  • Confident tone: The wording rarely signals uncertainty, so inaccurate statements are difficult to identify by reading alone.
  • Probabilistic origin: It is a side effect of next-token prediction, described in how LLMs work, not a software defect in one component.
Catching hallucinated function names in generated codeAn LLM suggests Python code that calls four functions. A validator checks each name against the real standard library. json.load and os.listdir exist and pass. json.safe_load and os.path.filesize do not exist, so they are flagged as hallucinations before the code is accepted. The example is illustrative.LLMSuggested callsjson.loadjson.safe_loados.listdiros.path.filesizeCheck namesResultexistsflaggedexistsflagged
Catching hallucinated function names in generated code

For example, a code assistant asked to parse a JSON string may suggest json.safe_load(), a function that does not exist in the Python standard library.

Key Characteristics of LLM Hallucination

  • Intrinsic hallucination: The output contradicts the prompt or a source document supplied in the context window.
  • Extrinsic hallucination: The output adds claims that cannot be verified from the supplied sources at all.
  • Plausible structure: Invented function names usually follow real naming conventions, such as combining two genuine library terms.
  • Topic dependence: Hallucination is more frequent for uncommon libraries, recent releases and specialised domains with limited training data.
  • Inconsistency: Repeating the same question several times can produce different fabricated details, which is a useful detection signal.
  • Variable detectability: Invented code usually fails during execution, while invented explanations can remain unnoticed for a long time.

How LLM Hallucination Happens

  1. Training gaps: The training data contains limited or outdated information about a particular library, version or technical fact.
  2. Knowledge cut-off: The model has no information about APIs released or modified after its training data was collected.
  3. Pattern completion: The model combines familiar fragments, such as safe_ from YAML and load from JSON, into a plausible new name.
  4. Sampling randomness: Higher temperature and top-p values make less likely, and often less accurate, tokens more common.
  5. Missing grounding: Without retrieved documentation in the prompt, the model answers entirely from patterns stored in its parameters.
  6. Pressure to answer: Models are optimised to produce helpful responses, so they generally attempt an answer instead of stating that information is unavailable.

Example: Checking Generated Function Names in Python

The program below parses code suggested by an assistant and checks every module.function call against the real installed modules.

Python
import ast
import json
import os

# Code an assistant suggested for "read a JSON file and list a folder"
generated = """
import json, os
data = json.load(open("config.json"))
text = json.dumps(data, indent=2)
files = os.listdir("logs")
safe = json.safe_load(text)
size = os.path.filesize("config.json")
"""

# Walk the syntax tree and collect every module.function call
calls = []
for node in ast.walk(ast.parse(generated)):
    if isinstance(node, ast.Call) and isinstance(node.func, ast.Attribute):
        target = node.func
        if isinstance(target.value, ast.Attribute):
            name = f"{target.value.value.id}.{target.value.attr}.{target.attr}"
        else:
            name = f"{target.value.id}.{target.attr}"
        calls.append(name)

# Check each name against the real installed modules
modules = {"json": json, "os": os}
for name in sorted(set(calls)):
    obj = modules[name.split(".")[0]]
    for part in name.split(".")[1:]:
        obj = getattr(obj, part, None)
    print(f"{name:22} {'ok' if obj is not None else 'NOT FOUND'}")
Output
json.dumps             ok
json.load              ok
json.safe_load         NOT FOUND
os.listdir             ok
os.path.filesize       NOT FOUND
  • Static validation: The ast module reads the code without executing it, so the check is safe to run on untrusted suggestions.
  • Two hallucinations found: json.safe_load borrows a name from YAML libraries, and os.path.filesize resembles the real os.path.getsize.
  • Real tools go further: Linters, type checkers and unit tests catch the same class of error in production developer tools.

LLM Hallucination Examples in Developer Tools

  • Code completion: Nonexistent functions, incorrect parameter names and deprecated APIs presented as current.
  • Stack trace explanations: Plausible but incorrect root causes for an exception, a use case of AI agents that read logs.
  • Pull request summaries: Descriptions of modifications that the actual diff does not contain.
  • Dependency suggestions: Package names that are not published, which attackers can register as malicious packages.
  • Documentation answers: Invented configuration options, environment variables or version numbers.
  • Citations: References to research papers, issue numbers or URLs that do not exist.

How to Reduce Hallucination in LLMs

  • Grounding with RAG: Retrieve the relevant API documentation and include it in the prompt, as described in what is RAG.
  • Output validation: Verify generated identifiers, schemas and imports against the real API, as demonstrated in the example above.
  • Tests and execution: Execute generated code in an isolated sandbox with unit tests before accepting it.
  • Explicit instructions: Instruct the model to say when information is unavailable, a technique covered in prompt engineering.
  • Lower temperature: Reduce sampling randomness for factual questions and code generation tasks.

Limitations of Current Fixes

  • No complete solution: Every mitigation reduces the frequency of hallucination, but none eliminates it entirely.
  • Retrieval errors: RAG can retrieve the wrong document, and the model can still misread the correct one.
  • Validation coverage: Name checks detect invented functions but not incorrect logic that uses genuine functions.
  • Review cost: Human review of important outputs remains necessary, which adds time and expense to every workflow.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What is the main reason an LLM hallucinates?

Frequently Asked Questions

Why do LLMs hallucinate?

LLMs generate the most likely continuation of the text, not a verified fact. When training data is missing or outdated, the model still produces a plausible answer, which can be wrong.

Can AI hallucination be completely eliminated?

No current method eliminates hallucination completely. Grounding, validation, testing and human review reduce it to a level that is acceptable for many applications.

Does RAG stop hallucination?

RAG reduces hallucination by placing relevant documents in the prompt, so the model can answer from real sources. The model can still misread those sources, and retrieval can return the wrong document.

How can developers detect hallucinated code?

Hallucination in code generation is usually caught by static checks, type checkers, linters and unit tests. Validating imported packages and called functions against the real library is a simple first step.

Does a lower temperature prevent hallucination?

A lower temperature makes output more consistent, but it does not add missing knowledge. A model can repeat the same wrong answer confidently at temperature 0.