Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting
Xinyan GuanYanjiang LiuHongyu LinYaojie LuBen HeXianpei HanLe Sun
Proposes an autonomous framework that iteratively extracts, verifies, and retrofits intermediate factual claims in draft responses using knowledge graphs, effectively eliminating multi-step reasoning hallucinations in large language models.
Large language models frequently generate unsupported or factually incorrect statements, a challenge known as factual hallucination. While incorporating structured data from knowledge graphs offers a promising remedy, traditional approaches only query external sources using the user's initial prompt. This standard strategy fails during multi-step reasoning, where errors often emerge in intermediate steps involving entities that were not mentioned in the original question.
To address this limitation, the article presents and evaluates Knowledge Graph-based Retrofitting (KGR), an autonomous framework designed to correct factual errors within an artificial intelligence model's generated reasoning process. The primary objective is to demonstrate that refining initial draft responses using structured knowledge graphs significantly improves factual accuracy and reliability.
Researchers evaluated the approach using three prominent language models—ChatGPT, text-davinci-003, and the open-source Vicuna 13B—across three benchmark datasets representing varying levels of reasoning difficulty: Simple Question, Mintaka (complex multi-step questions), and HotpotQA (open-domain, multi-hop reasoning). The framework operates in a continuous, five-stage cycle: extracting individual factual claims from a model's draft answer, identifying key entities, retrieving relevant local subgraphs from Wikidata, selecting critical facts while filtering out noise, and validating and rewriting the draft response. This entire workflow executes autonomously using the language model itself, without requiring manual intervention or separate architectural components.
The findings confirm substantial performance gains across multiple evaluation metrics. First, retrofitting draft responses systematically outperformed traditional baselines, including standard prompting, detailed reasoning prompts (Chain of Thought), web retrieval correction (CRITIC), and query-only knowledge retrieval methods. Second, the framework achieved its most notable advantages in complex multi-step reasoning tasks, delivering an F1 score improvement of at least 6.2 points on Mintaka and 1.1 points on HotpotQA compared to query-focused retrieval methods. Third, the system demonstrated strong versatility across diverse model types, significantly boosting performance on aligned commercial models as well as compact open-source models like Vicuna 13B, where exact match accuracy rose from 14.0% to 46.0% on Simple Question.
These results show that fact-checking the model's intermediate reasoning, rather than merely retrieving context before generation, is critical for real-world reliability. Grounding outputs in structured graphs minimizes the risk of propagating false intermediate logic, reducing operational risk in sensitive question-answering applications. Furthermore, because the framework relies entirely on prompt-driven execution, it avoids expensive model retraining.
Organizations developing or deploying language models for complex knowledge retrieval should consider adopting response-retrofitting architectures. Technical teams should focus on optimizing the trade-offs between precision and recall when selecting retrieved knowledge chunks. Further research and development should specifically target improvements in entity detection and fact filtering to prevent irrelevant data from cluttering verification steps.
While the results demonstrate clear efficacy, the evaluation was conducted on sample subsets of 50 instances per dataset validation set. Additionally, the framework's overall accuracy remains constrained by the language model's ability to extract precise entities and filter out noisy triples during information retrieval. Stakeholders can be confident in the structural advantages of iterative post-generation retrofitting, while noting that production deployments will require ongoing refinement of the fact-selection pipeline.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). Provides a comprehensive architectural foundation and taxonomy for unifying large language models with knowledge graphs to mitigate factual hallucinations.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). Introduces the iterative reading-then-reasoning framework for querying structured knowledge sources with language models that informs multi-step retrofitting pipelines.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Establishes foundational retrieval-augmented generation paradigms that standard query-only approaches utilize and that intermediate response retrofitting directly builds upon.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). Offers a systematic categorization of hallucination types and post-generation factuality verification strategies that contextualize the need for autonomous retrofitting.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Demonstrates the failure of parametric memory on long-tail entity facts, motivating the extraction and retrieval of structured graph subgraphs.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). Introduces interleaved reasoning and tool-based action steps that serve as the conceptual precursor to executing multi-stage autonomous verification cycles.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). Establishes the prompt-driven iterative feedback and refinement formulation that KGR adapts for structured knowledge base retrofitting.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). Pioneers the decomposition of generated text into discrete atomic claims for granular fact verification, a core mechanism of the KGR cycle.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Extends knowledge graph verification into active, dynamic path planning and adaptive backtracking during multi-hop reasoning.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). Addresses the noise-filtering bottlenecks identified in retrieval retrofitting by introducing adaptive adversarial training against imperfect retrieved context.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). Unifies context filtering and generation into a single instruction-tuned model, providing a potential solution to KGR's reliance on separate prompt-driven fact selection.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Applies reinforcement learning to train models to autonomously interleave reasoning and external knowledge retrieval rather than relying solely on post-hoc prompt retrofitting.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). Generalizes graph-based retrieval techniques from local entity-triple verification to hierarchical community extraction for high-level summarization.
