Skip to content

Hallucination: Causes and Mitigation

At a glance Hallucination is the phenomenon where LLMs "confidently talk nonsense" — the core trust issue for generative AI. This article systematically covers hallucination taxonomy (factuality/faithfulness, knowledge/reasoning), five categories of causes, a layered mitigation system from data to evaluation, RAG's boundary, and the "AI lying vs hallucination" intent distinction.

Hallucination: Causes and Mitigation ​

Hallucination refers to a language model generating fluent, confident content that contradicts factual reality or conflicts with source material — the model "confidently talks nonsense." It is the most common trust barrier when LLMs enter real business: wrong policy answers from customer service, fabricated case law in legal summaries, code comments referencing non-existent APIs — all are hallucination.

This article first defines and categorizes hallucination, then breaks down causes, then presents a five-layer mitigation system from data/training/inference/retrieval/evaluation, and finally discusses RAG's relationship and the "lying vs hallucination" boundary. Honesty evaluation and Evaluation & Benchmarks; safety responsibility is in Safety & Risks.

1. What Hallucination Is: Definition and Classification ​

Academically (Ji et al., 2023 survey), hallucination is usually split into two major categories by "inconsistent with what":

Major CategoryInconsistent withTypical manifestation
Factuality hallucinationObjective world factsFabricated biographies, incorrect statistics, fictional references
Faithfulness hallucinationInput source material (source documents/context)Conclusions in summaries that don't appear in the original text, rewriting deviating from original meaning

A commonly used split is by cause:

TypeMeaningExample
Knowledge hallucinationThe model's "knowledge" is itself wrong/missingFabricating details about long-tail knowledge or obscure figures
Reasoning hallucinationThe steps or logic chain are wrongCalculating math wrong while pretending to know how

Why distinguish?

Factuality hallucination is fixed by "giving the right knowledge" (retrieval, knowledge bases); faithfulness hallucination is fixed by "constraining generation" (making it follow only given context); reasoning hallucination is fixed by "giving thinking time" (CoT, step-by-step verification). Different causes require different cures — which is why this article unfolds mitigation plans by classification.

Hallucination's real-business impact is far more serious than "getting the answer wrong":

Business ScenarioHallucination consequenceSeverity
Customer service QAWrong policy/pricing, user complaintsMedium
Medical/legal/financial adviceMisleading decisions, health and property damageHigh
Summaries and compliance docsFabricated clauses, compliance riskHigh
Code generationNon-existent APIs, fake interfaces, compilation failureMedium
Education and knowledge productsWrong knowledge "confidently" taughtHigh
News/content generationFalse details, misleading the publicHigh

This is why "hallucination misleading" is simultaneously listed by Safety & Risks as a risk category — it's not just a quality defect; it's a liability issue.

2. Causes: Why Models "Confidently Get It Wrong" ​

Hallucination isn't a "bug" but a structural byproduct of how language models work. Five core causes:

1. The Training Objective Doesn't Distinguish Fact from Fiction ​

The pretraining objective is "predict the probability of the next word"; it learns statistical patterns of text distribution, not "which sentences correspond to the real world." Pretraining corpus mixes fiction, rumors, misinformation, and real reports — the model has no signal to distinguish "fact" from "fiction." It only learns "this reads like what a human would write."

2. Knowledge Cutoff and Information Obsolescence ​

A model's learned knowledge is fixed at the training corpus collection timepoint (knowledge cutoff). New events, policies, and version APIs after the cutoff are unknown to the model, yet it continues answering out of a "linguistic fluency" instinct — naturally fabricating. This is one of the most typical sources of hallucination.

3. Weak Long-Tail Knowledge ​

The frequency of knowledge in corpus varies enormously. High-frequency knowledge ("the capital of China is Beijing") the model learns firmly; long-tail knowledge (some detail of a small city's regulation) appears only a few times in corpus — the model can only "guess the most plausible." The lower the frequency, the higher the hallucination rate — this is scaling laws' direct reflection on knowledge coverage (see Scaling Laws).

4. Decoding Randomness ​

Sampling strategies during generation introduce randomness (see Inference Fundamentals: Autoregression and Sampling): the higher the temperature, the larger the top-p, the more diverse the output — and the more likely it "drifts" to wrong answers. The same prompt generated multiple times may yield one correct and one wrong.

5. Alignment Side Effect: Sycophancy ​

Alignment training teaches models to "please the user" — when a user says "my plan is definitely right," the model may agree. This "following user expectations" tendency amplifies fabrication, especially when users present false premises.

Hallucination can't be eliminated 100%

The root of hallucination lies in the "predict the next word" objective itself: the model has no innate ability to distinguish "has been said" from "is true." Any scheme claiming "zero hallucination" deserves skepticism. The mitigation goal is to push the hallucination rate to a business-acceptable level and make the model know when to say "I don't know."

3. Mitigating Hallucination: A Five-Layer Governance System ​

Hallucination mitigation must be layered; single-point approaches (like adding only RAG) are far from enough:

LayerMethodSolves whatCost/trade-off
Data layerPretraining corpus cleaning, deduplication, factuality filteringReduce error information density at the sourceHigh offline cost
Training layerSFT teaches "say I don't know"; RLHF builds honesty into reward; increase long-tail knowledge coverageReduce fabrication at the behavior levelHigh training cost, has alignment tax
Inference layerLower temperature, constrained decoding, self-check (model self-review), prompting "indicate if uncertain"Controllability of individual generationReduces diversity, increases latency
Retrieval layerRAG: retrieve evidence from external knowledge sources then generateInject reliable facts, traceableDepends on retrieval quality and coverage
Evaluation layerTruthfulQA, faithfulness metrics, manual spot-check, hallucination rate monitoringQuantify hallucination, drive iterationRequires continuous labeling investment

Inference-Layer Specific Methods ​

text
Low-risk prompting example (constraining the model to admit uncertainty):
"If you're uncertain about the answer, explicitly state 'I'm not sure,' and provide what you can confirm.
  Do not fabricate data, cite non-existent literature, or invent people."

Self-check mode:
  Round 1: let the model generate response A
  Round 2: let the model independently review A — "check the following response item by item,
           flag any statement you may be uncertain about or that may be inaccurate," getting revised A'

Training-Layer Honesty Reward ​

Encode "honesty" into preference during RLHF/DPO: give positive preference to "admitting not knowing" responses, negative preference to "fabricating confidently" ones (see Alignment: RLHF and DPO). This is the most fundamental approach at the behavior layer, but balance with "helpfulness" is crucial — overly conservative models degenerate into "I don't know" repeaters.

The implementation order for mitigation methods

For most teams, the cost-effective implementation order is: first inference layer (prompt constraints + "allowed to say I don't know," zero cost) → then retrieval layer (RAG injects reliable facts) → then evaluation layer (quantify hallucination rate, drive iteration) → finally training layer (encode honesty into preference, highest cost). Don't start with large-scale training changes — use cheap methods first to suppress most hallucination, then use eval data to decide whether training investment is worthwhile.

Product-Layer Hallucination Defense Strategies ​

Beyond technical layers, product design can significantly reduce the actual damage of hallucination:

StrategyApproach
Scenario tieringDegraded scenarios to "retrieval only/template only" output, giving the model no free rein
Evidence chainKey responses must include citation sources; if no source, mark as "model-generated, unverified"
Confidence expressionGuide the model to use qualifiers like "based on available information," "I'm not sure"
Disclaimers and fallbackMandatory human review for medical/legal/financial scenarios
User educationExplicitly state "AI may make mistakes," cultivate verification habits

These strategies don't change the model, but they contain "hallucination damage" within product boundaries — combined with technical-layer methods, they form the complete hallucination governance puzzle.

4. RAG and Hallucination: Symptom relief or root cure? ​

Retrieval-Augmented Generation (RAG) is currently the most mainstream hallucination mitigation approach in industry: first retrieve relevant knowledge, then let the model generate based on retrieved results, making output "grounded in evidence."

DimensionWhat RAG helps with hallucinationRAG's boundary
Fact sourceProvides trustworthy facts, replacing model guessingThe retrieval library itself may be incomplete, wrong, or outdated
Knowledge cutoffSolves "post-cutoff" new informationRequires continuous knowledge base maintenance
FaithfulnessConstrains the model to generate based on contextRetrieving irrelevant content can "mislead" the model instead
TraceabilityCan point to "where the evidence is"The model may cite details not actually in the retrieved content

RAG is not a free pass

RAG only shifts hallucination from "model fabrication" to "retrieval and context" — source quality determines generation quality: if retrieved results are incomplete, truncated, or irrelevant to the question, the model will still fabricate based on incomplete context. The correct use of RAG pairs with "faithfulness verification" (checking whether every output claim can be found in the retrieved docs). System architecture is in RAG: Retrieval-Augmented Generation, hands-on in RAG in Practice.

RAG and Context & Long Context are two complementary paths: feeding more context can cover more knowledge, but attention dilution and positional decay (visible in needle-in-a-haystack tests) let the model "miss" critical facts — so long-context products also need faithfulness evaluation.

RAG also isn't "set and forget." Common failure modes:

Failure ModeManifestationCountermeasure
Irrelevant retrievalRetrieved results unrelated to question, model forced to fabricateImprove retrieval quality, query rewriting
Context truncationKey evidence chopped off by chunkingOptimize chunk size and overlap
Multi-hop questionsAnswers scattered across multiple documents, single retrieval misses themMulti-round retrieval, Agent-style RAG
Outdated knowledge baseInfo in the DB is staleUpdate cycle management

Detailed troubleshooting is in rag-in-practice.

5. "AI Lying" vs "Hallucination": Intent and Liability ​

This is a pair of concepts often confused in public discussion:

  • Hallucination is an "unintentional error": the model has no "knowing it's lying" intent; it's simply maximizing the probability of the next word. It's mechanistic, probabilistic.
  • Lying/deception implies "knowingly outputting falsehood" — current LLMs lack stable "knowing" ability, so strictly speaking we can't say models are lying; but models can "appear to lie": they can be induced (when users assert false premises), and can be shaped by alignment training to "accommodate," outputting content conflicting with what they've learned.

Who bears the responsibility?

Hallucination responsibility lies with the human, not the model: it's the system designer who chooses whether to let the model "freestyle" or "stay grounded in evidence." Exposing an LLM directly to high-risk scenarios (medical advice, legal opinions, financial decisions) without evidence layers and human fallback is a design failure. This is also why "hallucination misleading" is listed as a risk category in Safety & Risks.

From a trust engineering perspective, four more realistic things than "eliminating hallucination" are: let the model say "I don't know" (calibration), make outputs traceable (cite evidence), make processes verifiable (human review/evidence check), make risks fallback-able (scenario tiering).

6. Honesty Evaluation: How to Quantify Hallucination ​

Hallucination needs governance, but first it must be quantifiable. Main methods:

Eval MethodWhat it testsNote
TruthfulQA (2021)Proportion of common misconceptions/misbeliefs the model spotsClassic honesty benchmark
Fact checkingConsistency between generated content and authoritative sourcesVerify item by item against retrieval results
Faithfulness metricsConsistency between summary/rewriting and source documentse.g., RAGAS's faithfulness
Hallucination rate monitoringProportion of manual spot-checks in productionRequires manual annotation pipeline
Model self-assessmentHave a strong model check the credibility of responsesLow cost, can serve as initial filter

Measurement criteria: hallucination rate = proportion of sampled responses containing "at least one unverifiable claim." Sampling should be random, sample size sufficient (at least 50–100 for statistical significance), and distinguish between "minor imprecision" and "severe fabrication" — the governance cost for the two levels is entirely different.

The relationship between honesty evaluation and the overall eval system is in Evaluation & Benchmarks; engineering implementation is in Evaluation in Practice.

Minimum config for hallucination governance

The minimum requirement for production systems: (1) connect RAG or knowledge base for sensitive scenarios; (2) allow the model to say "I don't know" in prompts; (3) spot-check hallucination rate and set thresholds; (4) human fallback for high-risk scenarios. All four are essential — "just switching to a smarter model" won't solve hallucination.

7. Boundaries and Open Questions in Hallucination Research ​

A few open questions in hallucination research to watch:

  • Reliable hallucination detection: automatic detection (NLI models, model self-assessment) still lags behind manual detection; at insufficient automatic accuracy, it can't replace manual spot-checks;
  • Hallucination in eval sets themselves: incorrect annotations in eval sets also distort scores — hallucination eval must itself be validated;
  • The boundary between calibration and honesty: models that "admit not knowing" too much hurt usability; calibrating "knowing how much you know" is still unsettled;
  • Entanglement with memory: whether models can establish a reliable boundary between "remembering facts" and "fabricating facts" is a core question for the future.

The eval and governance sides of these problems link respectively to Evaluation & Benchmarks and Safety & Risks.

Further Reading ​

References ​