Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs
Liyi ChenPanrong TongZhongming JinYing SunJieping YeHui Xiong
Presents a self-correcting adaptive planning framework for knowledge-graph-augmented large language models that dynamically adjusts path exploration breadth and backtracks from errors through task decomposition, memory updating, and reflection.
Large language models often struggle with outdated knowledge, factual hallucinations, and opaque decision-making processes. While pairing language models with structured knowledge graphs offers a promising remedy by providing explicit, editable facts, current graph-augmented methods suffer from critical rigidities. Existing approaches rely on fixed exploration breadths and follow strictly unidirectional paths, leaving models unable to adaptively navigate knowledge networks, retain complex multi-part query constraints, or backtrack when an exploration path fails.
The article demonstrates a novel framework called Plan-on-Graph (PoG), which enables large language models to adaptively explore knowledge graphs and autonomously self-correct erroneous reasoning paths. To achieve this, the article evaluates a four-stage process combining task decomposition, adaptive path exploration, dynamic memory tracking, and reflective evaluation.
To evaluate the framework, the authors conducted experiments across three standard multi-hop question answering benchmarks: ComplexWebQuestions, WebQSP, and GrailQA, using the Freebase knowledge graph. The system was tested using both GPT-3.5 and GPT-4 engines and compared against standard language model prompting methods, fine-tuned specialized models, and existing prompting-based graph reasoning architectures.
The findings show that Plan-on-Graph consistently outperforms existing methods in both accuracy and operational efficiency. When powered by GPT-4, the framework achieved top accuracy scores across all benchmarks (75.0% on ComplexWebQuestions, 87.3% on WebQSP, and 84.7% on GrailQA), surpassing all competing fine-tuned and prompting-based baselines. Furthermore, the framework reduced model interaction calls by at least 40.8%, cut output token consumption by approximately 76.2%, and delivered more than a fourfold execution speedup compared to leading graph prompting baselines. Analysis also revealed that 24% of test queries required path reversals, with self-correction successfully resolving up to 64% of those challenged queries.
These results demonstrate that integrating structured guidance, memory retention, and backtracking reflection allows language models to handle complex, multi-constraint reasoning tasks with significantly reduced computational cost and latency. By eliminating the need for expensive task-specific model fine-tuning while outperforming fine-tuned alternatives, this approach provides a cost-effective, transparent, and accurate architecture for enterprise-grade knowledge retrieval and automated reasoning systems.
Organizations deploying graph-augmented language models should adopt dynamic exploration breadths and reflection-based backtracking to prevent compounding reasoning failures. Before large-scale deployment, further development should focus on calibrating model self-confidence thresholds to optimize early stopping and exploring query rewriting techniques to handle non-standardized user inputs.
The primary limitations of this study include the model's occasional uncertainty regarding when retrieved information is fully sufficient, a fixed maximum exploration depth cap of four steps to prevent infinite looping, and potential performance drops on ambiguous or non-standard queries. Confidence in the reported performance and efficiency gains remains high across standard multi-hop benchmarks.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This comprehensive roadmap outlines the integration paradigms between large language models and knowledge graphs that Plan-on-Graph builds upon to address hallucinations and factual navigation.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). It introduces verbal self-reflection and dynamic memory updates for language model agents, which provide the direct conceptual basis for Plan-on-Graph's self-correcting reflection mechanism.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). It establishes deliberate multi-path exploration, evaluation, and backtracking in language models, directly preceding Plan-on-Graph's adaptive search over graph structures.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). It formalizes language model reasoning as an arbitrary graph with dynamic transformations and feedback loops, laying the foundational framework for planning over graph representations.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational work establishes retrieval-augmented generation to mitigate hallucinations via external factual sources, motivating the need for structured graph-based planning.
- Paper: A Survey on Knowledge Graphs: Representation, Acquisition, and Applications, Shaoxiong Ji et al. (2020). It provides essential background on knowledge graph representations and multi-hop relational path reasoning leveraged during graph-based sub-goal exploration.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). It highlights the compositionality gap and sub-question decomposition techniques in language models that motivate Plan-on-Graph's sub-objective decomposition strategy.
- Paper: Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, Shuo Ji et al. (2026). It extends adaptive graph-guided reasoning into agent memory by replacing static retrieval with iterative, step-by-step reconstruction over a graph structure.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This survey provides a broader architectural synthesis of agentic planning, memory, and self-evolution, situating graph-planning paradigms like Plan-on-Graph within autonomous systems.
- Paper: Efficient Rectification of Neuro-Symbolic Reasoning Inconsistencies by Abductive Reflection, Wen-Chao Hu et al. (2025). It advances the self-correction principle of Plan-on-Graph by employing targeted abductive reflection to rectify neuro-symbolic reasoning inconsistencies over graph problems.
- Paper: Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents, Theo Rusu et al. (2026). It investigates how graph-based memory and dynamic pruning scale to long-term agent interactions, evaluating multi-hop relational reasoning over extended contexts.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). It shifts from heuristic prompting to end-to-end reinforcement learning for autonomous search, evidence extraction, and self-verification in multi-step reasoning.
