Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
Garima AgrawalTharindu KumarageZeyad AlghamdiHuan Liu
Presents a systematic taxonomy and evaluation of knowledge graph augmentation techniques across inference, training, and validation stages to effectively curb hallucinations and improve factual reasoning in large language models.
Large language models often generate plausible yet incorrect or fabricated information, commonly referred to as hallucinations. These errors stem from training data gaps, ambiguous context, and the models' statistical text-generation mechanisms. As organizations increasingly deploy these models for decision-making and operational tasks, hallucinations introduce significant reliability, safety, and compliance risks. The article comprehensively evaluates how integrating structured external databases, known as knowledge graphs, can mitigate hallucinations and enhance the reasoning capabilities of large language models.
The article conducts a systematic literature review of studies published between 2019 and 2023 across leading natural language processing venues. It categorizes knowledge graph augmentation methodologies into three distinct lifecycle stages: inference, training, and validation. In knowledge-aware inference, structured facts are retrieved at run time, integrated into step-by-step reasoning prompts, or used to enforce operational constraints without retraining model parameters. Knowledge-aware training incorporates structured relational facts directly into models via pre-training or fine-tuning. Knowledge-aware validation leverages graphs as post-generation reference sources to verify claims and guide fact-checking.
The analysis reveals several key findings regarding performance and operational trade-offs. First, for question-answering tasks, augmenting smaller models with retrieved knowledge graph facts improves answer correctness by more than 80%, outperforming the baseline performance gains achieved merely by increasing model size. Second, integrating structured graph pathways into multi-step reasoning significantly enhances larger models; for example, knowledge graph-augmented reasoning boosted baseline accuracy from 66.8% to 85.7% on complex reasoning tasks and reached 88.2% accuracy in clinical diagnosis benchmarks. Third, while training-stage integrations effectively specialize models, pre-training and fine-tuning demand substantial computational resources and produce rigid, task-specific systems with limited transferability. Consequently, the research landscape has shifted away from resource-intensive pre-training toward inference-time retrieval, reasoning, and validation frameworks that avoid additional training costs.
These findings suggest that organizations can achieve higher accuracy and reduce hallucination risks without undertaking costly model training from scratch. Implementing inference-time retrieval and reasoning guardrails allows enterprises to maintain smaller, cost-effective models while enforcing factual consistency through structured business data. However, post-generation validation frameworks introduce additional computational overhead and may still fail to catch subtle errors if the underlying knowledge base is incomplete or biased.
Decision-makers should prioritize modular, inference-time knowledge graph integrations for general and knowledge-intensive workflows, reserving parameter fine-tuning only for narrow domains with stable data. Future development should focus on building dynamic, multi-modal, and bias-resistant knowledge graphs, exploring causality-aware representations, and establishing bidirectional synergies where models and graphs refine each other. Readers should maintain cautious confidence in these evaluations, as baseline benchmarks and model capabilities evolve rapidly, and retrieval-based performance remains strictly bounded by the scope and quality of the underlying graph data.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This paper establishes the foundational taxonomy and conceptual roadmap for integrating structured knowledge graphs with large language models to overcome factual hallucinations.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). This survey provides a comprehensive taxonomy of hallucination types and general mitigation pipelines in LLMs, establishing essential conceptual background for targeted knowledge graph interventions.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). This survey outlines the core paradigms and mechanics of retrieval-augmented generation (RAG), which forms the operational baseline for knowledge-graph-augmented LLM architectures.
- Paper: Survey of Hallucination in Natural Language Generation, Ziwei Ji et al. (2022). This foundational review defines the standard definitions, causes, and evaluation metrics for hallucinations in natural language generation systems.
- Paper: A Survey on Knowledge Graphs: Representation, Acquisition, and Applications, Shaoxiong Ji et al. (2020). This comprehensive survey details knowledge graph representations, completion techniques, and relational reasoning structures that serve as the external factual source for LLMs.
- Paper: JointLK: Joint Reasoning with Language Models and Knowledge Graphs for Commonsense Question Answering, Yueqing Sun et al. (2022). This paper presents pioneering methods for joint bidirectional reasoning and dynamic graph filtering between language models and structured knowledge graphs.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This study demonstrates why LLMs fail on long-tail factual knowledge, quantifying the theoretical necessity of external non-parametric retrieval to prevent factual errors.
- Paper: HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models, Junyi Li et al. (2023). This work introduces HaluEval, providing the standardized evaluation benchmarks and metrics widely referenced to measure LLM hallucination frequency and detection capabilities.
- Paper: Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting, Xinyan Guan et al. (2024). This paper advances graph-based hallucination mitigation by introducing an autonomous retrofitting framework that verifies and rewrites intermediate multi-step reasoning using structured subgraphs.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). This work extends static graph retrieval paradigms by developing an adaptive, self-correcting planning framework that navigates knowledge graphs dynamically during LLM reasoning.
- Paper: Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction, Bowen Zhang et al. (2024). This paper applies LLM reasoning in the reciprocal direction, utilizing language models to construct, define, and canonicalize knowledge graphs directly from unstructured text.
- Paper: RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, Cheng Niu et al. (2024). This research provides a fine-grained, word-level hallucination corpus specifically evaluating the failure modes and ungrounded generation that persist in retrieval-augmented workflows.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). This study investigates the mechanistic resolution of conflicts between external retrieved context and internal parametric memory using sparse autoencoder representation engineering.
- Paper: Multimodal Reasoning with Multimodal Knowledge Graph, Junlin Lee et al. (2024). This work generalizes knowledge graph augmentation into multimodal domains by integrating multimodal knowledge graphs with language models to mitigate cross-modal hallucinations.
