Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs

Liyi ChenPanrong TongZhongming JinYing SunJieping YeHui Xiong

article2024NeurIPS102 citations

Presents a self-correcting adaptive planning framework for knowledge-graph-augmented large language models that dynamically adjusts path exploration breadth and backtracks from errors through task decomposition, memory updating, and reflection.

Listen

Large language models often struggle with outdated knowledge, factual hallucinations, and opaque decision-making processes. While pairing language models with structured knowledge graphs offers a promising remedy by providing explicit, editable facts, current graph-augmented methods suffer from critical rigidities. Existing approaches rely on fixed exploration breadths and follow strictly unidirectional paths, leaving models unable to adaptively navigate knowledge networks, retain complex multi-part query constraints, or backtrack when an exploration path fails.

The article demonstrates a novel framework called Plan-on-Graph (PoG), which enables large language models to adaptively explore knowledge graphs and autonomously self-correct erroneous reasoning paths. To achieve this, the article evaluates a four-stage process combining task decomposition, adaptive path exploration, dynamic memory tracking, and reflective evaluation.

To evaluate the framework, the authors conducted experiments across three standard multi-hop question answering benchmarks: ComplexWebQuestions, WebQSP, and GrailQA, using the Freebase knowledge graph. The system was tested using both GPT-3.5 and GPT-4 engines and compared against standard language model prompting methods, fine-tuned specialized models, and existing prompting-based graph reasoning architectures.

The findings show that Plan-on-Graph consistently outperforms existing methods in both accuracy and operational efficiency. When powered by GPT-4, the framework achieved top accuracy scores across all benchmarks (75.0% on ComplexWebQuestions, 87.3% on WebQSP, and 84.7% on GrailQA), surpassing all competing fine-tuned and prompting-based baselines. Furthermore, the framework reduced model interaction calls by at least 40.8%, cut output token consumption by approximately 76.2%, and delivered more than a fourfold execution speedup compared to leading graph prompting baselines. Analysis also revealed that 24% of test queries required path reversals, with self-correction successfully resolving up to 64% of those challenged queries.

These results demonstrate that integrating structured guidance, memory retention, and backtracking reflection allows language models to handle complex, multi-constraint reasoning tasks with significantly reduced computational cost and latency. By eliminating the need for expensive task-specific model fine-tuning while outperforming fine-tuned alternatives, this approach provides a cost-effective, transparent, and accurate architecture for enterprise-grade knowledge retrieval and automated reasoning systems.

Organizations deploying graph-augmented language models should adopt dynamic exploration breadths and reflection-based backtracking to prevent compounding reasoning failures. Before large-scale deployment, further development should focus on calibrating model self-confidence thresholds to optimize early stopping and exploring query rewriting techniques to handle non-standardized user inputs.

The primary limitations of this study include the model's occasional uncertainty regarding when retrieved information is fully sufficient, a fixed maximum exploration depth cap of four steps to prevent infinite looping, and potential performance drops on ambiguous or non-standard queries. Confidence in the reported performance and efficiency gains remains high across standard multi-hop benchmarks.

Cover for Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs

Abstract

Large Language Models (LLMs) have shown remarkable reasoning capabilities on complex tasks, but they still suffer from out-of-date knowledge, hallucinations, and opaque decision-making. In contrast, Knowledge Graphs (KGs) can provide explicit and editable knowledge for LLMs to alleviate these issues. Existing paradigm of KG-augmented LLM manually predefines the breadth of exploration space and requires flawless navigation in KGs. However, this paradigm cannot adaptively explore reasoning paths in KGs based on the question semantics and self-correct erroneous reasoning paths, resulting in a bottleneck in efficiency and effect. To address these limitations, we propose a novel self-correcting adaptive planning paradigm for KG-augmented LLM named Plan-on-Graph (PoG), which first decomposes the question into several sub-objectives and then repeats the process of adaptively exploring reasoning paths, updating memory, and reflecting on the need to self-correct erroneous reasoning paths until arriving at the answer. Specifically, three important mechanisms of Guidance, Memory, and Reflection are designed to work together, to guarantee the adaptive breadth of self-correcting planning for graph reasoning. Finally, extensive experiments on three real-world datasets demonstrate the effectiveness and efficiency of PoG.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 Methodology
  • 3.1 Task Decomposition
  • 3.2 Path Exploration
  • 3.3 Memory Updating
  • 3.4 Evaluation
  • 4 Experiments
  • 4.1 Experimental Setups
  • 4.1.1 Datasets & Evaluation Metrics
  • 4.1.2 Comparison Methods
  • 4.2 Performance Comparison
  • 4.3 Ablation Study
  • 4.4 Efficiency Study
  • 4.5 Case Study
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References
  • Appendix
  • A Prompts
  • A.1 Task Decomposition
  • A.2 Path Exploration
  • A.2.1 Relation Exploration
  • A.2.2 Entity Exploration
  • A.3 Memory Updating
  • A.4 Evaluation
  • A.4.1 Answer Question
  • A.4.2 Reflection
  • B Search SPARQL
  • B.1 Relation Search
  • B.2 Entity Search
  • B.3 Entity Name Search
  • C Datasets
  • D Baseline Descriptions
  • LLM-Only Methods
  • Finetuned KG-Augmented LLM Methods
  • Prompting KG-Augmented LLM Methods
  • E Implementation Details
  • H Broader Impact & Limitation
  • NeurIPS Paper Checklist
  • 1. Claims
  • 2. Limitations
  • 3. Theory Assumptions and Proofs
  • 4. Experimental Result Reproducibility
  • 5. Open access to data and code
  • 6. Experimental Setting/Details
  • 7. Experiment Statistical Significance
  • 8. Experiments Compute Resources
  • 9. Code Of Ethics
  • 10. Broader Impacts
  • 11. Safeguards
  • 12. Licenses for existing assets
  • 13. New Assets
  • 14. Crowdsourcing and Research with Human Subjects
  • 15. Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects

Knowls

  1. Knowl 1 — Plan-on-Graph Reasoning Algorithm

    algorithm

    Plan-on-Graph (PoG) is a knowledge graph question answering (KGQA) framework that couples large language models (LLMs) with knowledge graphs (G={(e,r,e′)∣e,e′∈E,r∈R}G = \{(e, r, e') \mid e, e' \in E, r \in R\}) through task decomposition, adaptive graph exploration, memory maintenance, and self-correcting reflection.

    Input: Question qq, Knowledge Graph G=(E,R)G = (E, R), Topic Entities Tq⊆ET_q \subseteq E, Maximum Search Depth DmaxD_{max}
    Output: Answer AqA_q
    Decompose qq via LLM into sub-objectives O={o1,o2,…,om}O = \{o_1, o_2, \dots, o_m\}
    Initialize candidate tail entities E0←TqE^0 \leftarrow T_q
    Initialize reasoning paths P←∅P \leftarrow \emptyset
    Initialize historical subgraph GSub←∅G_{Sub} \leftarrow \emptyset
    Initialize sub-objective status S←∅S \leftarrow \emptyset
    for d=1d = 1 to DmaxD_{max} do
        Retrieve connected candidate relations RcanddR^d_{cand} for entities in Ed−1E^{d-1}
        Rd←R^d \leftarrow LLM selects relevant relations from RcanddR^d_{cand} based on q,O,Ed−1q, O, E^{d-1}
        Retrieve connected candidate entities EcanddE^d_{cand} via queries (e,r,?)(e, r, ?) or (?,r,e)(?, r, e) for e∈Ed−1,r∈Rde \in E^{d-1}, r \in R^d
        Recall top entities from EcanddE^d_{cand} using dense similarity with qq
        Ed←E^d \leftarrow LLM selects relevant entities and extends reasoning paths PP
        Update GSub←GSub∪Rcandd∪EcanddG_{Sub} \leftarrow G_{Sub} \cup R^d_{cand} \cup E^d_{cand}
        S←S \leftarrow LLM updates status of each sub-objective oi∈Oo_i \in O using q,O,P,Sq, O, P, S
        IsSufficient,Answer←IsSufficient, Answer \leftarrow LLM evaluates if SS and PP suffice to answer qq
        if IsSufficientIsSufficient is True then
            return Answer
        end if
        NeedCorrection,Reason←NeedCorrection, Reason \leftarrow LLM reflects on whether EdE^d requires correction based on q,S,P,Edq, S, P, E^d
        if NeedCorrectionNeedCorrection is True then
            Eaddd←E^d_{add} \leftarrow LLM selects backtrack entities from ⋃i=1dEcandi\bigcup_{i=1}^d E^i_{cand} based on SS and ReasonReason
            Ed←Ed∪EadddE^d \leftarrow E^d \cup E^d_{add}
        end if
    end for
    return LLM generates fallback answer from P,SP, S
  2. Knowl 2 — Adaptive Path Exploration on Knowledge Graphs

    model/method

    Adaptive Path Exploration is a two-step iterative procedure that dynamically expands reasoning paths in a knowledge graph without imposing a rigid beam width or predefined path breadth constraint.

    Let ED−1={e1D−1,e2D−1,…,eND−1D−1}E^{D-1} = \{e^{D-1}_1, e^{D-1}_2, \dots, e^{D-1}_{N_{D-1}}\} denote the set of active tail entities at depth D−1D-1, and PP denote the set of current reasoning paths. The exploration proceeds in two stages:

    1. Relation Exploration: All relations adjacent to tail entities in ED−1E^{D-1} are retrieved as candidate relations RcandD={rcand,1D,…,rcand,ND−1D}R^D_{cand} = \{r^D_{cand,1}, \dots, r^D_{cand,N_{D-1}}\}. The LLM filters RcandDR^D_{cand} to select an unconstrained number of relevant relations RDR^D conditioned on the question qq, the active entities ED−1E^{D-1}, and the task sub-objectives OO.

    2. Entity Exploration: For each active path ending in entity enD−1e^{D-1}_n and relation rnD∈RDr^D_n \in R^D, candidate entities Ecand,nDE^D_{cand,n} are fetched via knowledge graph queries (enD−1,rnD,?)(e^{D-1}_n, r^D_n, ?) or (?,rnD,enD−1)(?, r^D_n, e^{D-1}_n). When candidate entity sets are large, a pre-trained DistilBERT dual-encoder (msmarco-distilbert-base-tas-b) computes semantic similarity between candidate entity names and qq to recall a compact pool EcandDE^D_{cand}. The LLM then selects an adaptive subset of entities ED⊆EcandDE^D \subseteq E^D_{cand} based on the question semantics and retrieved knowledge triplets, extending paths PP to depth DD.

  3. Knowl 3 — Guidance via Sub-Objective Task Decomposition

    model/method

    In the Plan-on-Graph paradigm, Guidance is established prior to graph navigation by prompting the LLM to perform semantic decomposition of the input question qq into an ordered list of atomic sub-objectives:

    O={o1,o2,…,om}O = \{o_1, o_2, \dots, o_m\}

    Each sub-objective oio_i encapsulates an explicit condition or constraint present in qq (e.g., locating an entity, identifying a relation, or filtering candidates by attributes). Sub-objectives can refer to the expected intermediate outputs of previous sub-objectives, defining relational dependencies. The set OO guides path exploration by directing relation selection at each hop toward satisfying individual conditions rather than processing the composite query in an unstructured manner.

  4. Knowl 4 — Dynamic Memory Architecture for Graph Navigation

    model/method

    The memory component in Plan-on-Graph maintains structured historical retrieval and reasoning state across exploration hops to enable reflection and backtracking. It continuously updates three data representations:

    1. Retrieved Subgraph (GSubG_{Sub}): Stores the union of all candidate relations and entities explored across all iterations, GSub=⋃d=1D(Rcandd∪Ecandd)G_{Sub} = \bigcup_{d=1}^D (R^d_{cand} \cup E^d_{cand}). This full historical graph serves as the candidate space from which the model can select backtracked entities during self-correction.

    2. Reasoning Paths (PP): Retains the active relational chains pn={(es,nd,rj,nd,eo,nd)}d=1Dpnp_n = \{(e^d_{s,n}, r^d_{j,n}, e^d_{o,n})\}_{d=1}^{D_{p_n}}, where es,nde^d_{s,n} and eo,nde^d_{o,n} are subject and object entities and rj,ndr^d_{j,n} is the linking relation. These preserve the factual subgraph context required for LLM deduction.

    3. Sub-Objective Status (SS): Maintains a state vector S={s1,s2,…,s∣O∣}S = \{s_1, s_2, \dots, s_{|O|}\} matching the sub-objective list OO, where each sis_i summarizes the currently validated factual findings for sub-objective oio_i. The LLM updates SS dynamically after each entity exploration step using qq, OO, historical SS, and updated paths PP.

  5. Knowl 5 — Reflection-Driven Self-Correction and Backtracking

    model/method

    The Reflection mechanism evaluates whether the current graph exploration trajectory is sufficient, dead-ended, or erroneous, and dynamically executes backtracking.

    At each iteration, after memory updating, the LLM determines whether the current evidence in reasoning paths PP and sub-objective status SS is sufficient to answer the question qq. If the evidence is insufficient, PoG enters the reflection stage:

    1. Trajectory Assessment: Conditioned on question qq, sub-objective status SS, reasoning paths PP, and planned next-hop entities EDE^D, the LLM predicts whether additional entities outside EDE^D are required, outputting a boolean decision and a textual rationale ReasonReason.

    2. Backtrack Selection: If self-correction is triggered, the LLM inspects the cumulative pool of all candidate entities previously stored in memory, Ecand=⋃i=1DEcandiE_{cand} = \bigcup_{i=1}^D E^i_{cand}. Guided by SS and ReasonReason, the LLM identifies a focused set of alternative entities EaddD⊆EcandE^D_{add} \subseteq E_{cand} to restart exploration.

    3. Path Augmentation: The exploration frontier is updated via ED←ED∪EaddDE^D \leftarrow E^D \cup E^D_{add}, allowing the framework to branch out from earlier nodes without restarting from scratch.

  6. Knowl 6 — Hits@1 Performance Comparison on CWQ, WebQSP, and GrailQA

    data/table

    The table below compares the exact match accuracy (Hits@1, in %) of Plan-on-Graph (PoG) against LLM-only baselines, fine-tuned KG-augmented LLM baselines, and prompting-based KG-augmented LLM baselines across ComplexWebQuestions (CWQ), WebQuestionsSP (WebQSP), and GrailQA (including Overall, I.I.D., Compositional, and Zero-shot subsets).

    Method CWQ WebQSP GrailQA
    Overall I.I.D. Comp. Zero-shot
    LLM-Only
    IO Prompt 37.6 63.3 29.4 - - -
    CoT 38.8 62.2 28.1 - - -
    Self-Consistency 45.4 61.1 29.6 - - -
    Fine-Tuned KG-Augmented LLM
    UniKGQA 51.2 79.1 - - - -
    TIARA - 75.2 73.0 87.8 69.2 68.0
    RE-KBQA 50.3 74.6 - - - -
    DeCAF 70.4 82.1 - - - -
    RoG 62.6 85.7 - - - -
    RnG-KBQA - - 68.8 86.2 63.8 63.0
    FC-KBQA - - 73.2 88.5 70.0 67.6
    Pangu - - 75.4 84.4 74.6 71.6
    FlexKBQA - - 62.8 71.3 59.1 60.6
    GAIN - - 76.3 88.5 73.7 71.8
    Prompting KG-Augmented LLM (GPT-3.5/Other)
    KD-CoT 50.5 73.7 - - - -
    KB-BINDER - 74.4 50.6 - - -
    StructGPT 54.3 72.6 - - - -
    ToG 57.1 76.2 68.7 70.1 56.1 72.7
    PoG (Ours) 63.2 82.0 76.5 76.3 62.1 81.7
    Prompting KG-Augmented LLM (GPT-4)
    InteractiveKBQA 59.2 72.5 - - - -
    ToG 67.6 82.6 81.4 79.4 67.3 86.5
    PoG (Ours) 75.0 87.3 84.7 87.9 69.7 88.6

    PoG consistently outperforms previous prompting-based methods (such as ToG and StructGPT) across all benchmarks. When powered by GPT-3.5, PoG achieves 81.7% accuracy on the zero-shot split of GrailQA, surpassing all fine-tuned baselines on that subset.

  7. Knowl 7 — Ablation Study of Plan-on-Graph Components

    data/table

    To analyze the contribution of each component in Plan-on-Graph, ablation experiments were conducted using GPT-3.5 on CWQ, WebQSP, and GrailQA. Hits@1 accuracy (in %) is reported when removing specific mechanisms:

    Method Variant CWQ WebQSP GrailQA
    PoG (Full Model) 63.2 82.0 76.5
    w/o Guidance 60.1 80.3 72.4
    w/o Memory 58.9 77.5 69.3
    w/o Reflection 59.4 78.1 70.5
    w/o Adaptive Breadth 61.3 80.2 73.8
    • w/o Guidance: Removes task decomposition into sub-objectives.
    • w/o Memory: Removes structured tracking of the subgraph, reasoning paths, and sub-objective statuses.
    • w/o Reflection: Continues unidirectional path extension without backtracking when information is insufficient.
    • w/o Adaptive Breadth: Fixes the maximum width of explored relations and entities to a static threshold instead of letting the LLM decide.

    Removing the Memory module causes the largest performance drop (4.3% on CWQ, 4.5% on WebQSP, and 7.2% on GrailQA), followed by the Reflection module.

  8. Knowl 8 — Computational and Efficiency Comparison Between PoG and ToG

    data/table

    Efficiency metrics measured with GPT-3.5 on CWQ, WebQSP, and GrailQA demonstrate the computational reduction achieved by adaptive breadth and reflection over the beam-search baseline Think-on-Graph (ToG):

    Dataset Method LLM Calls Input Tokens Output Tokens Total Tokens Time (s)
    CWQ ToG 22.6 8,182.9 1,486.4 9,669.4 96.5
    PoG 13.3 7,803.0 353.2 8,156.2 23.3
    WebQSP ToG 15.9 6,031.2 987.7 7,018.9 63.1
    PoG 9.0 5,234.8 282.9 5,517.7 16.8
    GrailQA ToG 11.1 4,066.0 774.6 4,840.6 50.2
    PoG 6.5 3,372.8 202.8 3,575.6 11.5

    PoG reduces the average number of LLM API calls by at least 40.8% across all datasets, decreases generated output tokens by up to 76.2% on CWQ, and achieves a wall-clock execution speedup of over 4×4\times on CWQ and GrailQA.

  9. Knowl 9 — Quantitative Impact of Backtracking and Search Depth

    empirical result

    Empirical analysis reveals the operational characteristics of reflection-driven backtracking and depth scaling in Plan-on-Graph:

    1. Backtracking Frequency: In the CWQ benchmark, 24% of all test questions triggered reverse occurrences (backtracking to earlier candidate entities via reflection).

    2. Post-Correction Success Rate: Following reflection and self-correction, PoG successfully recovers the correct answer in 48% of corrected cases on CWQ, 64% on WebQSP, and 36% on GrailQA.

    3. Search Depth Sensitivity: On the CWQ dataset, increasing the maximum exploration depth from 1 to 5 yields accuracy improvements from 53.8% at depth 1 up to ~64.5% at depth 5, with performance gains plateauing beyond depth 4 (reaching 63.2% at depth 4). Due to exponential compute scaling at higher depths, depth Dmax=4D_{max} = 4 offers the optimal balance between accuracy and efficiency.

  10. Knowl 10 — Limitations of the Plan-on-Graph Framework

    limitation

    Plan-on-Graph exhibits three primary limitations:

    1. Uncertainty in LLM Self-Confidence: LLMs struggle with accurately assessing their own certainty regarding whether retrieved information is sufficient or when to terminate exploration, necessitating heuristic exploration depth limits (e.g., setting maximum depth to 4) to prevent infinite loops.

    2. Multi-Step Overhead for Simple Queries: Executing multi-stage decomposition, path exploration, memory tracking, and evaluation loops incurs latency and token overhead on simpler questions where direct single-step answering would suffice.

    3. Vulnerability to Non-Standardized Natural Language Queries: When questions feature irregular syntax, ambiguous phrasing, or non-standard entity references, LLM zero-shot semantic parsing during task decomposition can degrade, cascading errors to subsequent relation and entity exploration stages.

Coverage note — All core contributions of the paper (task decomposition guidance, adaptive two-stage path exploration, dynamic 3-part memory updating, reflection and backtracking, experimental accuracy tables, ablation studies, efficiency benchmarks, self-correction frequency analysis, and stated limitations) are fully covered.

References

  1. 1.Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. Dbpedia: A nucleus for a web of open data. In international semantic web conference, pages 722–735. Springer, 2007.
  2. 2.Agnes Axelsson and Gabriel Skantze. Using large language models for zero-shot natural language generation from knowledge graphs. In Proceedings of the Workshop on Multimodal, Multilingual Natural Language Generation and Multilingual WebNLG Challenge (MM-NLG 2023), pages 39–54, 2023.
  3. 3.Jinheon Baek, Alham Aji, and Amir Saffari. Knowledge-augmented language model prompting for zero-shot knowledge graph question answering. In The 61st Annual Meeting Of The Association For Computational Linguistics, 2023.
  4. 4.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17682–17690, 2024.
  5. 5.Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data, pages 1247–1250, 2008.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  7. 7.Yong Cao, Xianzhi Li, Huiwen Liu, Wen Dai, Shuai Chen, Bin Wang, Min Chen, and Daniel Hershcovich. Pay more attention to relation exploration for knowledge base question answering. In Findings of the Association for Computational Linguistics: ACL 2023, pages 2119–2136, 2023.
  8. 8.Liyi Chen, Zhi Li, Weidong He, Gong Cheng, Tong Xu, Nicholas Jing Yuan, and Enhong Chen. Entity summarization via exploiting description complementarity and salience. IEEE Transactions on Neural Networks and Learning Systems, 34(11):8297–8309, 2023.
  9. 9.Liyi Chen, Zhi Li, Yijun Wang, Tong Xu, Zhefeng Wang, and Enhong Chen. Mmea: Entity alignment for multi-modal knowledge graph. In International Conference on Knowledge Science, Engineering and Management, pages 134–147. Springer, 2020.
  10. 10.Liyi Chen, Zhi Li, Tong Xu, Han Wu, Zhefeng Wang, Nicholas Jing Yuan, and Enhong Chen. Multi-modal siamese network for entity alignment. In Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining, pages 118–126, 2022.
  11. 11.Liyi Chen, Chuan Qin, Ying Sun, Xin Song, Tong Xu, Hengshu Zhu, and Hui Xiong. Collaboration-aware hybrid learning for knowledge development prediction. In Proceedings of the ACM on Web Conference 2024, pages 3976–3985, 2024.
  12. 12.Liyi Chen, Ying Sun, Shengzhe Zhang, Yuyang Ye, Wei Wu, and Hui Xiong. Tackling uncertain correspondences for multi-modal entity alignment. In Proceedings of the 38th Conference on Neural Information Processing Systems, 2024.
  13. 13.Xi Chen, Xinjiang Lu, Haoran Xin, Wenjun Peng, Haoyang Duan, Feihu Jiang, Jingbo Zhou, and Hui Xiong. A table-to-text framework with heterogeneous multidominance attention and self-evaluated multi-pass deliberation. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 607–620, 2023.
  14. 14.Xi Chen, Chuan Qin, Zhigaoyuan Wang, Yihang Cheng, Chao Wang, Hengshu Zhu, and Hui Xiong. Pre-dygae: Pre-training enhanced dynamic graph autoencoder for occupational skill demand forecasting. In Proceedings of the 33th International Joint Conference on Artificial Intelligence, 2024.
  15. 15.Zheng Gong and Ying Sun. Graph reasoning enhanced language models for text-to-sql. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2447–2451, 2024.
  16. 16.Yu Gu, Xiang Deng, and Yu Su. Don’t generate, discriminate: A proposal for grounding language models to real-world environments. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4928–4949, 2023.
  17. 17.Yu Gu, Sue Kase, Michelle Vanni, Brian Sadler, Percy Liang, Xifeng Yan, and Yu Su. Beyond iid: three levels of generalization for question answering on knowledge bases. In Proceedings of the Web Conference 2021, pages 3477–3488, 2021.
  18. 18.Bin Ji, Huijun Liu, Mingzhe Du, and See-Kiong Ng. Chain-of-thought improves text generation with citations in large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18345–18353, 2024.
  19. 19.Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Wayne Xin Zhao, and Ji-Rong Wen. Structgpt: A general framework for large language model to reason over structured data. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9237–9251, 2023.
  20. 20.Jinhao Jiang, Kun Zhou, Xin Zhao, and Ji-Rong Wen. Unikgqa: Unified retrieval and reasoning for solving multi-hop question answering over knowledge graph. In The International Conference on Learning Representations, 2023.
  21. 21.Byoungjip Kim, Youngsoo Jang, Lajanugen Logeswaran, Geon-Hyeong Kim, Yu Jin Kim, Honglak Lee, and Moontae Lee. Prospector: Improving llm agents with self-asking and trajectory ranking. 2023.
  22. 22.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  23. 23.Tianle Li, Xueguang Ma, Alex Zhuang, Yu Gu, Yu Su, and Wenhu Chen. Few-shot in-context learning on knowledge base question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6966–6980, 2023.
  24. 24.Xiaonan Li and Xipeng Qiu. Mot: Memory-of-thought enables chatgpt to self-improve. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6354–6374, 2023.
  25. 25.Xiaoxi Li, Yujia Zhou, and Zhicheng Dou. Unigen: A unified generative framework for retrieval and question answering with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 8688–8696, 2024.
  26. 26.Zhenyu Li, Sunqi Fan, Yu Gu, Xiuxing Li, Zhichao Duan, Bowen Dong, Ning Liu, and Jianyong Wang. Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18608–18616, 2024.
  27. 27.Linhao Luo, Yuan-Fang Li, Reza Haf, and Shirui Pan. Reasoning on graphs: Faithful and interpretable large language model reasoning. In The Twelfth International Conference on Learning Representations, 2024.
  28. 28.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36, 2024.
  29. 29.Ning Miao, Yee Whye Teh, and Tom Rainforth. Selfcheck: Using llms to zero-shot check their own step-by-step reasoning. In The Twelfth International Conference on Learning Representations, 2024.
  30. 30.Xuefei Ning, Zinan Lin, Zixuan Zhou, Zifu Wang, Huazhong Yang, and Yu Wang. Skeleton-of-thought: Large language models can do parallel decoding. In The Twelfth International Conference on Learning Representations, 2024.
  31. 31.V Sanh. Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  32. 32.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36, 2024.
  33. 33.Yiheng Shu and Zhiwei Yu. Distribution shifts are bottlenecks: Extensive evaluation for grounding language models to knowledge bases. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 71–88, 2024.
  34. 34.Yiheng Shu, Zhiwei Yu, Yuhan Li, Börje Karlsson, Tingting Ma, Yuzhong Qu, and Chin-Yew Lin. Tiara: Multi-grained retrieval for robust question answering over large knowledge base. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8108–8121, 2022.
  35. 35.Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Lionel Ni, Heung-Yeung Shum, and Jian Guo. Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph. In The Twelfth International Conference on Learning Representations, 2024.
  36. 36.Ying Sun, Hengshu Zhu, Lu Wang, Le Zhang, and Hui Xiong. Large-scale online job search behaviors reveal labor market shifts amid covid-19. Nature Cities, 1(2):150–163, 2024.
  37. 37.Alon Talmor and Jonathan Berant. The web as a knowledge-base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 641–651, 2018.
  38. 38.Yanchao Tan, Hang Lv, Xinyi Huang, Jiawei Zhang, Shiping Wang, and Carl Yang. Musegraph: Graph-oriented instruction tuning of large language models for generic graph mining. arXiv preprint arXiv:2403.04780, 2024.
  39. 39.Jiabin Tang, Yuhao Yang, Wei Wei, Lei Shi, Lixin Su, Suqi Cheng, Dawei Yin, and Chao Huang. Graphgpt: Graph instruction tuning for large language models. arXiv preprint arXiv:2310.13023, 2023.
  40. 40.Jianing Wang, Qiushi Sun, Nuo Chen, Xiang Li, and Ming Gao. Boosting language models reasoning with chain-of-knowledge prompting. arXiv preprint arXiv:2306.06427, 2023.
  41. 41.Keheng Wang, Feiyu Duan, Sirui Wang, Peiguang Li, Yunsen Xian, Chuantao Yin, Wenge Rong, and Zhang Xiong. Knowledge-driven cot: Exploring faithful reasoning in llms for knowledge-intensive question answering. arXiv preprint arXiv:2308.13259, 2023.
  42. 42.Lei Wang, Yi Hu, Jiabang He, Xing Xu, Ning Liu, Hui Liu, and Heng Tao Shen. T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19162–19170, 2024.
  43. 43.Shuyao Wang, Yongduo Sui, Chao Wang, and Hui Xiong. Unleashing the power of knowledge graph for recommendation via invariant learning. In Proceedings of the ACM on Web Conference 2024, pages 3745–3755, 2024.
  44. 44.Shuyao Wang, Yongduo Sui, Jiancan Wu, Zhi Zheng, and Hui Xiong. Dynamic sparse learning: A novel paradigm for efficient recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, pages 740–749, 2024.
  45. 45.Tianfu Wang, Qilin Fan, Chao Wang, Leilei Ding, Nicholas Jing Yuan, and Hui Xiong. Flagvne: A flexible and generalizable rl framework for network resource allocation. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence, 2024.
  46. 46.Tianfu Wang, Li Shen, Qilin Fan, Tong Xu, Tongliang Liu, and Hui Xiong. Joint admission control and resource allocation of virtual network embedding via hierarchical deep reinforcement learning. IEEE Transactions on Services Computing, 17(03):1001–1015, 2024.
  47. 47.Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. Kepler: A unified model for knowledge embedding and pre-trained language representation. Transactions of the Association for Computational Linguistics, 9:176–194, 2021.
  48. 48.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations, 2023.
  49. 49.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  50. 50.Shiwei Wu, Joya Chen, Tong Xu, Liyi Chen, Lingfei Wu, Yao Hu, and Enhong Chen. Linking the characters: Video-oriented social graph generation via hierarchical-cumulative gcn. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4716–4724, 2021.
  51. 51.Wei Wu, Chao Wang, Dazhong Shen, Chuan Qin, Liyi Chen, and Hui Xiong. Afdgcf: Adaptive feature de-correlation graph collaborative filtering for recommendations. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1242–1252, 2024.
  52. 52.Guanming Xiong, Junwei Bao, and Wen Zhao. Interactive-kbqa: Multi-turn interactions for knowledge base question answering with large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers, pages 10561–10582, 2024.
  53. 53.Liang Yao, Chengsheng Mao, and Yuan Luo. Kg-bert: Bert for knowledge graph completion. arXiv preprint arXiv:1909.03193, 2019.
  54. 54.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 2024.
  55. 55.Xi Ye, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou, and Caiming Xiong. Rng-kbqa: Generation augmented iterative ranking for knowledge base question answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6032–6043, 2022.
  56. 56.Wen-tau Yih, Matthew Richardson, Christopher Meek, Ming-Wei Chang, and Jina Suh. The value of semantic parse labeling for knowledge base question answering. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 201–206, 2016.
  57. 57.Donghan Yu, Sheng Zhang, Patrick Ng, Henghui Zhu, Alexander Hanbo Li, Jun Wang, Yiqun Hu, William Yang Wang, Zhiguo Wang, and Bing Xiang. Decaf: Joint decoding of answers and logical forms for question answering over knowledge bases. In The International Conference on Learning Representations, 2023.
  58. 58.Lingxi Zhang, Jing Zhang, Yanling Wang, Shulin Cao, Xinmei Huang, Cuiping Li, Hong Chen, and Juanzi Li. Fc-kbqa: A fine-to-coarse composition framework for knowledge base question answering. In The 61st Annual Meeting Of The Association For Computational Linguistics, 2023.
  59. 59.Shengzhe Zhang, Liyi Chen, Chao Wang, Shuangli Li, and Hui Xiong. Temporal graph contrastive learning for sequential recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 9359–9367, 2024.
  60. 60.Yuting Zhang, Ying Sun, Fuzhen Zhuang, Yongchun Zhu, Zhulin An, and Yongjun Xu. Triple dual learning for opinion-based explainable recommendation. ACM Transactions on Information Systems, 42(3):1–27, 2023.
  61. 61.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. Ernie: Enhanced language representation with informative entities. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1441–1451, 2019.
  62. 62.Lili Zhao, Qi Liu, Linan Yue, Wei Chen, Liyi Chen, Ruijun Sun, and Chao Song. Comi: Correct and mitigate shortcut learning behavior in deep neural networks. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 218–228, 2024.
  63. 63.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al. Least-to-most prompting enables complex reasoning in large language models. The International Conference on Learning Representations, 2023.

Citation

MLA
Chen, L., et al. “Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 37665–91, https://proceedings.neurips.cc/paper_files/paper/2024/file/4254e856d01a5e7b7ea050477c3ef9b9-Paper-Conference.pdf.
APA
Chen, L., Tong, P., Jin, Z., Sun, Y., Ye, J., & Xiong, H. (2024). Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs. Advances in Neural Information Processing Systems, 37, 37665–37691. https://proceedings.neurips.cc/paper_files/paper/2024/file/4254e856d01a5e7b7ea050477c3ef9b9-Paper-Conference.pdf
Chicago
Chen, L., P. Tong, Z. Jin, Y. Sun, J. Ye, and H. Xiong. 2024. “Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs”. Advances in Neural Information Processing Systems 37: 37665–91. https://proceedings.neurips.cc/paper_files/paper/2024/file/4254e856d01a5e7b7ea050477c3ef9b9-Paper-Conference.pdf.
Harvard
Chen, L. et al. (2024) “Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37665–37691. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/4254e856d01a5e7b7ea050477c3ef9b9-Paper-Conference.pdf.
Vancouver
1. Chen L, Tong P, Jin Z, Sun Y, Ye J, Xiong H (2024) Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37665–37691

BibTeX

@inproceedings{chen2024plan,
  title = {Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs},
  author = {Chen, Liyi and Tong, Panrong and Jin, Zhongming and Sun, Ying and Ye, Jieping and Xiong, Hui},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {37665-37691},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/4254e856d01a5e7b7ea050477c3ef9b9-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors