Theme
Hallucination: Causes and Mitigation
Hallucination refers to a language model generating fluent, confident content that contradicts factual reality or conflicts with source material — the model "confidently talks nonsense." It is the most common trust barrier when LLMs enter real business: wrong policy answers from customer service, fabricated case law in legal summaries, code comments referencing non-existent APIs — all are hallucination.
This article first defines and categorizes hallucination, then breaks down causes, then presents a five-layer mitigation system from data/training/inference/retrieval/evaluation, and finally discusses RAG's relationship and the "lying vs hallucination" boundary. Honesty evaluation and Evaluation & Benchmarks; safety responsibility is in Safety & Risks.
1. What Hallucination Is: Definition and Classification
Academically (Ji et al., 2023 survey), hallucination is usually split into two major categories by "inconsistent with what":
| Major Category | Inconsistent with | Typical manifestation |
|---|---|---|
| Factuality hallucination | Objective world facts | Fabricated biographies, incorrect statistics, fictional references |
| Faithfulness hallucination | Input source material (source documents/context) | Conclusions in summaries that don't appear in the original text, rewriting deviating from original meaning |
A commonly used split is by cause:
| Type | Meaning | Example |
|---|---|---|
| Knowledge hallucination | The model's "knowledge" is itself wrong/missing | Fabricating details about long-tail knowledge or obscure figures |
| Reasoning hallucination | The steps or logic chain are wrong | Calculating math wrong while pretending to know how |
Why distinguish?
Factuality hallucination is fixed by "giving the right knowledge" (retrieval, knowledge bases); faithfulness hallucination is fixed by "constraining generation" (making it follow only given context); reasoning hallucination is fixed by "giving thinking time" (CoT, step-by-step verification). Different causes require different cures — which is why this article unfolds mitigation plans by classification.
Hallucination's real-business impact is far more serious than "getting the answer wrong":
| Business Scenario | Hallucination consequence | Severity |
|---|---|---|
| Customer service QA | Wrong policy/pricing, user complaints | Medium |
| Medical/legal/financial advice | Misleading decisions, health and property damage | High |
| Summaries and compliance docs | Fabricated clauses, compliance risk | High |
| Code generation | Non-existent APIs, fake interfaces, compilation failure | Medium |
| Education and knowledge products | Wrong knowledge "confidently" taught | High |
| News/content generation | False details, misleading the public | High |
This is why "hallucination misleading" is simultaneously listed by Safety & Risks as a risk category — it's not just a quality defect; it's a liability issue.
2. Causes: Why Models "Confidently Get It Wrong"
Hallucination isn't a "bug" but a structural byproduct of how language models work. Five core causes:
1. The Training Objective Doesn't Distinguish Fact from Fiction
The pretraining objective is "predict the probability of the next word"; it learns statistical patterns of text distribution, not "which sentences correspond to the real world." Pretraining corpus mixes fiction, rumors, misinformation, and real reports — the model has no signal to distinguish "fact" from "fiction." It only learns "this reads like what a human would write."
2. Knowledge Cutoff and Information Obsolescence
A model's learned knowledge is fixed at the training corpus collection timepoint (knowledge cutoff). New events, policies, and version APIs after the cutoff are unknown to the model, yet it continues answering out of a "linguistic fluency" instinct — naturally fabricating. This is one of the most typical sources of hallucination.
3. Weak Long-Tail Knowledge
The frequency of knowledge in corpus varies enormously. High-frequency knowledge ("the capital of China is Beijing") the model learns firmly; long-tail knowledge (some detail of a small city's regulation) appears only a few times in corpus — the model can only "guess the most plausible." The lower the frequency, the higher the hallucination rate — this is scaling laws' direct reflection on knowledge coverage (see Scaling Laws).
4. Decoding Randomness
Sampling strategies during generation introduce randomness (see Inference Fundamentals: Autoregression and Sampling): the higher the temperature, the larger the top-p, the more diverse the output — and the more likely it "drifts" to wrong answers. The same prompt generated multiple times may yield one correct and one wrong.
5. Alignment Side Effect: Sycophancy
Alignment training teaches models to "please the user" — when a user says "my plan is definitely right," the model may agree. This "following user expectations" tendency amplifies fabrication, especially when users present false premises.
Hallucination can't be eliminated 100%
The root of hallucination lies in the "predict the next word" objective itself: the model has no innate ability to distinguish "has been said" from "is true." Any scheme claiming "zero hallucination" deserves skepticism. The mitigation goal is to push the hallucination rate to a business-acceptable level and make the model know when to say "I don't know."
3. Mitigating Hallucination: A Five-Layer Governance System
Hallucination mitigation must be layered; single-point approaches (like adding only RAG) are far from enough:
| Layer | Method | Solves what | Cost/trade-off |
|---|---|---|---|
| Data layer | Pretraining corpus cleaning, deduplication, factuality filtering | Reduce error information density at the source | High offline cost |
| Training layer | SFT teaches "say I don't know"; RLHF builds honesty into reward; increase long-tail knowledge coverage | Reduce fabrication at the behavior level | High training cost, has alignment tax |
| Inference layer | Lower temperature, constrained decoding, self-check (model self-review), prompting "indicate if uncertain" | Controllability of individual generation | Reduces diversity, increases latency |
| Retrieval layer | RAG: retrieve evidence from external knowledge sources then generate | Inject reliable facts, traceable | Depends on retrieval quality and coverage |
| Evaluation layer | TruthfulQA, faithfulness metrics, manual spot-check, hallucination rate monitoring | Quantify hallucination, drive iteration | Requires continuous labeling investment |
Inference-Layer Specific Methods
text
Low-risk prompting example (constraining the model to admit uncertainty):
"If you're uncertain about the answer, explicitly state 'I'm not sure,' and provide what you can confirm.
Do not fabricate data, cite non-existent literature, or invent people."
Self-check mode:
Round 1: let the model generate response A
Round 2: let the model independently review A — "check the following response item by item,
flag any statement you may be uncertain about or that may be inaccurate," getting revised A'Training-Layer Honesty Reward
Encode "honesty" into preference during RLHF/DPO: give positive preference to "admitting not knowing" responses, negative preference to "fabricating confidently" ones (see Alignment: RLHF and DPO). This is the most fundamental approach at the behavior layer, but balance with "helpfulness" is crucial — overly conservative models degenerate into "I don't know" repeaters.
The implementation order for mitigation methods
For most teams, the cost-effective implementation order is: first inference layer (prompt constraints + "allowed to say I don't know," zero cost) → then retrieval layer (RAG injects reliable facts) → then evaluation layer (quantify hallucination rate, drive iteration) → finally training layer (encode honesty into preference, highest cost). Don't start with large-scale training changes — use cheap methods first to suppress most hallucination, then use eval data to decide whether training investment is worthwhile.
Product-Layer Hallucination Defense Strategies
Beyond technical layers, product design can significantly reduce the actual damage of hallucination:
| Strategy | Approach |
|---|---|
| Scenario tiering | Degraded scenarios to "retrieval only/template only" output, giving the model no free rein |
| Evidence chain | Key responses must include citation sources; if no source, mark as "model-generated, unverified" |
| Confidence expression | Guide the model to use qualifiers like "based on available information," "I'm not sure" |
| Disclaimers and fallback | Mandatory human review for medical/legal/financial scenarios |
| User education | Explicitly state "AI may make mistakes," cultivate verification habits |
These strategies don't change the model, but they contain "hallucination damage" within product boundaries — combined with technical-layer methods, they form the complete hallucination governance puzzle.
4. RAG and Hallucination: Symptom relief or root cure?
Retrieval-Augmented Generation (RAG) is currently the most mainstream hallucination mitigation approach in industry: first retrieve relevant knowledge, then let the model generate based on retrieved results, making output "grounded in evidence."
| Dimension | What RAG helps with hallucination | RAG's boundary |
|---|---|---|
| Fact source | Provides trustworthy facts, replacing model guessing | The retrieval library itself may be incomplete, wrong, or outdated |
| Knowledge cutoff | Solves "post-cutoff" new information | Requires continuous knowledge base maintenance |
| Faithfulness | Constrains the model to generate based on context | Retrieving irrelevant content can "mislead" the model instead |
| Traceability | Can point to "where the evidence is" | The model may cite details not actually in the retrieved content |
RAG is not a free pass
RAG only shifts hallucination from "model fabrication" to "retrieval and context" — source quality determines generation quality: if retrieved results are incomplete, truncated, or irrelevant to the question, the model will still fabricate based on incomplete context. The correct use of RAG pairs with "faithfulness verification" (checking whether every output claim can be found in the retrieved docs). System architecture is in RAG: Retrieval-Augmented Generation, hands-on in RAG in Practice.
RAG and Context & Long Context are two complementary paths: feeding more context can cover more knowledge, but attention dilution and positional decay (visible in needle-in-a-haystack tests) let the model "miss" critical facts — so long-context products also need faithfulness evaluation.
RAG also isn't "set and forget." Common failure modes:
| Failure Mode | Manifestation | Countermeasure |
|---|---|---|
| Irrelevant retrieval | Retrieved results unrelated to question, model forced to fabricate | Improve retrieval quality, query rewriting |
| Context truncation | Key evidence chopped off by chunking | Optimize chunk size and overlap |
| Multi-hop questions | Answers scattered across multiple documents, single retrieval misses them | Multi-round retrieval, Agent-style RAG |
| Outdated knowledge base | Info in the DB is stale | Update cycle management |
Detailed troubleshooting is in rag-in-practice.
5. "AI Lying" vs "Hallucination": Intent and Liability
This is a pair of concepts often confused in public discussion:
- Hallucination is an "unintentional error": the model has no "knowing it's lying" intent; it's simply maximizing the probability of the next word. It's mechanistic, probabilistic.
- Lying/deception implies "knowingly outputting falsehood" — current LLMs lack stable "knowing" ability, so strictly speaking we can't say models are lying; but models can "appear to lie": they can be induced (when users assert false premises), and can be shaped by alignment training to "accommodate," outputting content conflicting with what they've learned.
Who bears the responsibility?
Hallucination responsibility lies with the human, not the model: it's the system designer who chooses whether to let the model "freestyle" or "stay grounded in evidence." Exposing an LLM directly to high-risk scenarios (medical advice, legal opinions, financial decisions) without evidence layers and human fallback is a design failure. This is also why "hallucination misleading" is listed as a risk category in Safety & Risks.
From a trust engineering perspective, four more realistic things than "eliminating hallucination" are: let the model say "I don't know" (calibration), make outputs traceable (cite evidence), make processes verifiable (human review/evidence check), make risks fallback-able (scenario tiering).
6. Honesty Evaluation: How to Quantify Hallucination
Hallucination needs governance, but first it must be quantifiable. Main methods:
| Eval Method | What it tests | Note |
|---|---|---|
| TruthfulQA (2021) | Proportion of common misconceptions/misbeliefs the model spots | Classic honesty benchmark |
| Fact checking | Consistency between generated content and authoritative sources | Verify item by item against retrieval results |
| Faithfulness metrics | Consistency between summary/rewriting and source documents | e.g., RAGAS's faithfulness |
| Hallucination rate monitoring | Proportion of manual spot-checks in production | Requires manual annotation pipeline |
| Model self-assessment | Have a strong model check the credibility of responses | Low cost, can serve as initial filter |
Measurement criteria: hallucination rate = proportion of sampled responses containing "at least one unverifiable claim." Sampling should be random, sample size sufficient (at least 50–100 for statistical significance), and distinguish between "minor imprecision" and "severe fabrication" — the governance cost for the two levels is entirely different.
The relationship between honesty evaluation and the overall eval system is in Evaluation & Benchmarks; engineering implementation is in Evaluation in Practice.
Minimum config for hallucination governance
The minimum requirement for production systems: (1) connect RAG or knowledge base for sensitive scenarios; (2) allow the model to say "I don't know" in prompts; (3) spot-check hallucination rate and set thresholds; (4) human fallback for high-risk scenarios. All four are essential — "just switching to a smarter model" won't solve hallucination.
7. Boundaries and Open Questions in Hallucination Research
A few open questions in hallucination research to watch:
- Reliable hallucination detection: automatic detection (NLI models, model self-assessment) still lags behind manual detection; at insufficient automatic accuracy, it can't replace manual spot-checks;
- Hallucination in eval sets themselves: incorrect annotations in eval sets also distort scores — hallucination eval must itself be validated;
- The boundary between calibration and honesty: models that "admit not knowing" too much hurt usability; calibrating "knowing how much you know" is still unsettled;
- Entanglement with memory: whether models can establish a reliable boundary between "remembering facts" and "fabricating facts" is a core question for the future.
The eval and governance sides of these problems link respectively to Evaluation & Benchmarks and Safety & Risks.
Further Reading
- Evaluation & Benchmarks — the position of honesty evaluation and hallucination rate monitoring
- Safety & Risks — hallucination misleading's role in the risk categories
- RAG: Retrieval-Augmented Generation — RAG's full architecture and evolution
- Inference Fundamentals: Autoregression and Sampling — how decoding randomness affects hallucination
- Evaluation in Practice — engineeringized quantification of hallucination rate
- Common Pitfalls & Anti-Patterns — the trap of "treating hallucination as a bug to fix"
References
- Ji et al. Survey of Hallucination in Natural Language Generation (ACM Computing Surveys, 2023) — hallucination survey, authoritative source for definitions and classification
- Lin et al. TruthfulQA: Measuring How Models Mimic Human Falsehood (ACL 2022, originally 2021) — the honesty evaluation benchmark
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020) — the original RAG paper
- Mallen et al. When Not to Trust Language Models: Investigating Effectiveness and Limitations of Parametric and Non-Parametric Memories (ACL 2023) — empirical study on long-tail knowledge and hallucination
- OpenAI. GPT-4 Technical Report (2023) — discussion of factuality and adversarial hallucination