Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models

Linhao LuoZicheng ZhaoGholamreza HaffariYuan-Fang LiChen GongShirui Pan

article2025ICML84 citations

Proposes Graph-Constrained Reasoning, a framework that constrains language model decoding with a trie-based index of knowledge graph paths to eliminate reasoning hallucinations and achieve zero-shot generalization across question answering benchmarks.

Listen

Large language models show strong problem-solving skills but frequently generate inaccurate facts and ungrounded reasoning steps, commonly known as hallucinations. While structured knowledge graphs provide verified factual data to anchor these models, existing integration methods face severe trade-offs. Retrieval-based approaches often miss structural context or fail on novel queries, whereas interactive, agent-based approaches suffer from high latency, heavy computational overhead, and continued reasoning errors.

The article introduces and evaluates Graph-Constrained Reasoning, a framework designed to eliminate reasoning hallucinations by directly embedding structured knowledge graph constraints into the language model decoding process.

The evaluated approach indexes reachable knowledge graph paths into a prefix tree structure, called a knowledge graph trie, built either offline or on demand in roughly 0.28 seconds. During text generation, a lightweight, specialized language model is constrained by this index to generate only verified reasoning paths and preliminary answers. These candidate paths are then passed to a general large language model, which synthesizes multiple lines of evidence into a final answer in a single call. The authors evaluated this framework across standard question-answering benchmarks, including WebQuestionsSP, Complex WebQuestions, FreebaseQA, CommonsenseQA, and MedQA, against 22 baseline methods.

The evaluation yielded several key findings. First, the framework achieved state-of-the-art accuracy, reaching a 92.6% top-match rate on WebQuestionsSP and 75.8% on Complex WebQuestions, outperforming the best prior baseline methods by 2.1% and 9.1%, respectively. Second, the method achieved a 100% faithful reasoning rate, completely eliminating reasoning hallucinations on verified graph structures where baseline models exhibited hallucination rates of 33% to 52%. Third, the system maintained high operational efficiency, requiring only two model calls and 231 input tokens per query on average, compared to over 11 calls and 7,000 tokens for leading agent-based frameworks. Finally, the framework demonstrated zero-shot adaptability, improving accuracy on unseen knowledge graphs by up to 8.2% without extra training.

These results indicate that embedding structural graph constraints directly into model decoding offers a scalable, low-latency pathway to highly reliable artificial intelligence systems. Organizations deploying reasoning models in regulated or mission-critical settings can substantially reduce operational compute costs while enforcing verifiable compliance with trusted factual data sources.

Organizations seeking to deploy accurate reasoning systems should consider adopting constrained decoding architectures over costly multi-step agent frameworks. Engineering teams can implement dynamic caching for popular entities to keep query latency low. However, stakeholders should note that the system's faithfulness is inherently bounded by the completeness of the underlying knowledge base. If relevant facts are missing from the graph or if initial entity extraction selects an unrelated subgraph, the system may fail to produce the correct answer. Further work should explore combining structured graphs with external unstructured text documents to handle incomplete databases.

No sufficiently relevant recommendations were found.

Cover for Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models

Abstract

Large language models (LLMs) have demonstrated impressive reasoning abilities, but they still struggle with faithful reasoning due to knowledge gaps and hallucinations. To address these issues, knowledge graphs (KGs) have been utilized to enhance LLM reasoning through their structured knowledge. However, existing KG-enhanced methods, either retrieval-based or agent-based, encounter difficulties in accurately retrieving knowledge and efficiently traversing KGs at scale. In this work, we introduce graph-constrained reasoning (GCR), a novel framework that bridges structured knowledge in KGs with unstructured reasoning in LLMs. To eliminate hallucinations, GCR ensures faithful KG-grounded reasoning by integrating KG structure into the LLM decoding process through KG-Trie, a trie-based index that encodes KG reasoning paths. KG-Trie constrains the decoding process, allowing LLMs to directly reason on graphs and generate faithful reasoning paths grounded in KGs. Additionally, GCR leverages a lightweight KG-specialized LLM for graph-constrained reasoning alongside a powerful general LLM for inductive reasoning over multiple reasoning paths, resulting in accurate reasoning with zero reasoning hallucination. Extensive experiments on several KGQA benchmarks demonstrate that GCR achieves state-of-the-art performance and exhibits strong zero-shot generalizability to unseen KGs without additional training¹.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminary
  • 4. Approach
  • 4.1. From Chain-of-Thought Reasoning to Graph-constrained Reasoning
  • 4.2. Knowledge Graph Trie Construction
  • 4.3. Graph-constrained Decoding
  • 4.4. Graph Inductive Reasoning
  • 5. Experiment
  • 5.1. Experiment Setups
  • 5.2. RQ1: Reasoning Performance and Efficiency
  • 5.3. RQ2: Hallucination Elimination and Faithful Reasoning
  • 5.4. RQ3: Zero-shot Generalizability to Unseen KGs
  • 6. Conclusion
  • Acknowledgment
  • Impact Statement
  • References
  • Appendix
  • A. Detailed Related Work on KG-enhanced LLMs
  • B. KG-Trie Construction
  • B.1. Construction Strategies
  • B.2. Time and Space Complexity Analysis
  • B.2.1. THEORETICAL ANALYSIS
  • B.2.2. EMPIRICAL ANALYSIS
  • B.2.3. TIME CONSUMPTION BREAKDOWN
  • B.3. Strategies for Optimizing Efficiency
  • B.4. Real-World Applicability
  • C. Datasets
  • D. Baselines
  • E. Implementation Details and Experiment Settings
  • F. Additional Experiment Results
  • F.1. Performance on Different Hops of KG-Trie
  • F.2. Performance on Multi-path Reasoning
  • F.3. Performance on Multi-hop Reasoning
  • F.4. Logical Coherence in KG Reasoning Paths
  • F.5. Analysis of the Failure Cases
  • G. Limitations
  • H. Templates and Prompts

Knowls

  1. Knowl 1 — Graph-constrained reasoning combines KG path search with LLM induction

    model/method

    Graph-constrained reasoning (GCR) answers a question using two complementary language models and a knowledge graph (KG). A lightweight KG-specialized LLM generates KG-grounded reasoning paths and provisional answers; a more capable general LLM then reasons across multiple paths and provisional answers to produce the final answer. The KG structure constrains path generation during decoding, rather than being supplied only as retrieved text or explored through repeated agent–KG interactions.

  2. Knowl 2 — KG-Trie indexes question-relevant KG paths for constrained decoding

    model/method

    For a question, GCR identifies its linked KG entities and retrieves paths from those entities up to a chosen maximum hop count using breadth-first search. It formats each path as a sequence of entities and relations, tokenizes the formatted sequences with the LLM tokenizer, and stores them in a trie. Shared prefixes are represented by common trie prefixes, so the trie indicates which next tokens can continue a valid KG path. The index can be constructed for a question when needed or precomputed and loaded for inference. In the reported KGQA experiments, the maximum path length was 2 hops.

  3. Knowl 3 — Trie constraints restrict generated reasoning paths to KG paths

    model/method

    During graph-constrained decoding, the KG-specialized LLM generates a reasoning path one token at a time, but can select only tokens that continue a prefix represented in the question-specific KG-Trie. This prevents the path from continuing along a sequence absent from the indexed KG paths. Once a valid path has been generated, decoding switches back to ordinary LLM generation to produce a hypothesis answer conditioned on that path. GCR calls reasoning zero-hallucination when its generated paths are fully grounded in the KG; this operational criterion establishes KG traceability, not that every KG fact is complete or correct.

  4. Knowl 4 — KG-specialized LLM is fine-tuned to generate paths and hypothesis answers

    model/method

    GCR fine-tunes a KG-specialized LLM on question–reasoning-path–answer examples. For each training question and answer, the path supervision consists of shortest KG paths connecting the question entity to the answer entity; when multiple shortest paths exist, they produce multiple training examples. The model is trained to generate a relevant path and then a hypothesis answer conditioned on that path. The WebQSP and CWQ training splits yielded 28,307 and 181,602 examples, respectively, for 209,909 examples in total. The reported training setup used 3 epochs, batch size 4, learning rate 2e-5, and a cosine learning-rate schedule with warmup ratio 0.03; the main experiments used fine-tuned Llama-3.1-8B as the KG-specialized model.

  5. Knowl 5 — A general LLM synthesizes multiple path hypotheses

    model/method

    GCR uses beam search with graph-constrained decoding to produce the top K path–hypothesis-answer pairs in one call to the KG-specialized LLM. A general LLM receives those pairs together with the original question and produces final answers by reasoning across the alternatives. This lets the final model reconcile evidence and disregard noisy individual hypotheses, rather than simply returning the KG-specialized model’s provisional answer. The experiments used K = 10 and tested ChatGPT and GPT-4o-mini as the general LLM.

  6. Knowl 6 — GCR achieves leading WebQSP and CWQ results

    empirical result

    On the Freebase-based WebQSP and CWQ benchmarks, GCR was evaluated using Hit, which measures whether a prediction contains a correct answer, and F1, which balances answer precision and recall. With Llama-3.1-8B as the KG-specialized LLM and ChatGPT as the general LLM, GCR scored WebQSP Hit 92.6 and F1 73.2, and CWQ Hit 72.7 and F1 60.9. Replacing ChatGPT with GPT-4o-mini yielded WebQSP Hit 92.2 and F1 74.1, and CWQ Hit 75.8 and F1 61.7. For comparison, GNN-RAG+RA scored 90.7/73.5 on WebQSP and 68.7/60.4 on CWQ (Hit/F1). Thus, the GPT-4o-mini GCR configuration had higher WebQSP F1, and both GCR configurations exceeded the listed comparison on CWQ Hit and F1.

  7. Knowl 7 — KG constraints produce fully KG-grounded paths in the reported faithfulness evaluation

    empirical result

    The paper evaluates a reasoning path as faithful when it can be found in the KG, and reports the faithful-reasoning ratio among correctly answered questions. With graph constraints, GCR achieved a 100% faithful-reasoning ratio on both WebQSP and CWQ. Without the constraints, the reported ratios were 62.4% on WebQSP and 48.1% on CWQ. Removing constraints also substantially reduced answer performance on WebQSP, while answer Hit on CWQ remained nearly unchanged. These results support the narrower claim that the decoding constraint enforces grounding in the supplied KG; they do not rule out errors caused by incomplete or incorrect KG facts.

  8. Knowl 8 — GCR transfers to new KGQA datasets without additional fine-tuning

    empirical result

    For zero-shot transfer, the KG-specialized Llama-3.1-8B trained on Freebase was applied without additional fine-tuning to FreebaseQA, CommonsenseQA (CSQA, using ConceptNet), and MedQA (using a medical KG). The reported scores for ChatGPT were 85, 79, and 64, respectively; GCR with ChatGPT scored 92, 85, and 66. GPT-4o-mini scored 89, 91, and 75, while GCR with GPT-4o-mini scored 94, 94, and 79. The evaluation used 100 selected questions per dataset and constructed a KG-Trie from the corresponding KG. The results show transfer gains across these datasets, with smaller gains on MedQA than on the other two.

  9. Knowl 9 — GCR reduces LLM calls and input tokens relative to graph-search baselines

    empirical result

    In the WebQSP efficiency comparison, measured on a single A100-80G GPU, GCR achieved Hit 92.6 with an average runtime of 3.60 seconds, 2 LLM calls, and 231 input LLM tokens per question. ToG achieved Hit 75.1 with 16.14 seconds, 11.6 calls, and 7,069 tokens; RoG achieved Hit 85.7 with 2.60 seconds, 2 calls, and 521 tokens; GNN-RAG achieved Hit 85.7 with 1.52 seconds, 1 call, and 414 tokens. Dense S-Bert retrieval was faster at 0.87 seconds and used 1 call and 293 tokens, but achieved Hit 66.9. GCR was therefore not the fastest method in this comparison, but paired the highest reported Hit with substantially fewer calls and input tokens than the iterative agent baseline.

  10. Knowl 10 — Ablations show both LLM roles contribute to answer F1

    empirical result

    On WebQSP and CWQ, the full GCR system achieved F1 scores of 73.2 and 60.9. When the KG-specialized LLM was removed and all 2-hop paths were given to the general LLM, F1 fell to 52.9 and 37.5. When the general LLM was removed and the KG-specialized LLM’s hypothesis answers were used directly, F1 fell to 57.0 and 39.4. These ablations support the division of work in GCR: the specialized model finds question-relevant paths in the KG, while the general model improves final answering by reasoning across multiple path hypotheses.

Coverage note — The appendix’s detailed KG-Trie construction-time and storage breakdowns, caching and parallelization options, and finer-grained sensitivity analyses are omitted as supporting engineering analyses rather than separate central contributions; the main path-length and beam-size settings are included.

References

  1. 1.Agrawal, G., Kumarage, T., Alghamdi, Z., and Liu, H. Mindful-rag: A study of points of failure in retrieval augmented generation. arXiv preprint arXiv:2407.12216, 2024.
  2. 2.Bollacker, K., Evans, C., Paritosh, P., Sturge, T., and Taylor, J. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data, pp. 1247–1250, 2008.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020.
  4. 4.Chen, C., Wang, Y., Li, B., and Lam, K.-Y. Knowledge is flat: A seq2seq generative framework for various knowledge graph completion. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 4005–4017, 2022.
  5. 5.Chen, L., Tong, P., Jin, Z., Sun, Y., Ye, J., and Xiong, H. Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  6. 6.De Cao, N., Izacard, G., Riedel, S., and Petroni, F. Autoregressive entity retrieval. In International Conference on Learning Representations, 2022.
  7. 7.Dehghan, M., Alomrani, M., Bagga, S., Alfonso-Hermelo, D., Bibi, K., Ghaddar, A., Zhang, Y., Li, X., Hao, J., Liu, Q., Lin, J., Chen, B., Parthasarathi, P., Biparva, M., and Rezagholizadeh, M. EWEK-QA : Enhanced web and efficient knowledge graph retrieval for citation-based question answering systems. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14169–14187, Bangkok, Thailand, August 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.acl-long.764.
  8. 8.Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., and Weston, J. E. Chain-of-verification reduces hallucination in large language models. In ICLR 2024 Workshop on Reliable and Responsible Foundation Models.
  9. 9.Dong, Z., Peng, B., Wang, Y., Fu, J., Wang, X., Shan, Y., and Zhou, X. Effiqa: Efficient question-answering with strategic multi-model collaboration on knowledge graphs. arXiv preprint arXiv:2406.01238, 2024.
  10. 10.Erling, O. and Mikhailov, I. Rdf support in the virtuoso dbms. In Networked Knowledge-Networked Media: Integrating Knowledge Management, New Media Technologies and Semantic Systems, pp. 7–24. Springer, 2009.
  11. 11.Evans, J. S. B. Intuition and reasoning: A dual-process perspective. Psychological Inquiry, 21(4):313–326, 2010.
  12. 12.Federico, M., Cettolo, M., Brugnara, F., and Antoniol, G. Language modelling for efficient beam-search. Computer Speech and Language, 9(4):353–380, 1995.
  13. 13.Feng, Y., Chen, X., Lin, B. Y., Wang, P., Yan, J., and Ren, X. Scalable multi-hop relational reasoning for knowledge-aware question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1295–1309, 2020.
  14. 14.Fredkin, E. Trie memory. Communications of the ACM, 3(9):490–499, 1960.
  15. 15.He, G., Lan, Y., Jiang, J., Zhao, W. X., and Wen, J.-R. Improving multi-hop knowledge base question answering by learning intermediate supervision signals. In Proceedings of the 14th ACM international conference on web search and data mining, pp. 553–561, 2021.
  16. 16.Hoffman, M. D., Phan, D., Dohan, D., Douglas, S., Le, T. A., Parisi, A., Sountsov, P., Sutton, C., Vikram, S., and A Saurous, R. Training chain-of-thought via latent-variable inference. Advances in Neural Information Processing Systems, 36, 2024.
  17. 17.Huang, J. and Chang, K. C.-C. Towards reasoning in large language models: A survey. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 1049–1065, 2023.
  18. 18.Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D. Large language models cannot self-correct reasoning yet. In The Twelfth International Conference on Learning Representations, 2024.
  19. 19.Izacard, G. and Grave, E. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 874–880, 2021.
  20. 20.Jiang, J., Zhou, K., Zhao, X., and Wen, J.-R. Unikgqa: Unified retrieval and reasoning for solving multi-hop question answering over knowledge graph. In The Eleventh International Conference on Learning Representations, 2022.
  21. 21.Jiang, J., Zhou, K., Dong, Z., Ye, K., Zhao, W. X., and Wen, J.-R. Structgpt: A general framework for large language model to reason over structured data. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 9237–9251, 2023.
  22. 22.Jiang, J., Zhou, K., Zhao, W. X., Song, Y., Zhu, C., Zhu, H., and Wen, J.-R. Kg-agent: An efficient autonomous agent framework for complex reasoning over knowledge graph. arXiv preprint arXiv:2402.11163, 2024.
  23. 23.Jiang, K., Wu, D., and Jiang, H. Freebaseqa: A new factoid qa data set matching trivia-style question-answer pairs with freebase. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 318–323, 2019.
  24. 24.Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021.
  25. 25.Li, S., Gao, Y., Jiang, H., Yin, Q., Li, Z., Yan, X., Zhang, C., and Yin, B. Graph reasoning for question answering with triplet retrieval. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 3366–3375, 2023.
  26. 26.Li, Y., Song, D., Zhou, C., Tian, Y., Wang, H., Yang, Z., and Zhang, S. A framework of knowledge graph-enhanced large language model based on question decomposition and atomic retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 11472–11485, 2024.
  27. 27.Li, Y., Zhang, X., Luo, L., Chang, H., Ren, Y., King, I., and Li, J. G-refer: Graph retrieval-augmented large language model for explainable recommendation. In Proceedings of the ACM on Web Conference 2025, pp. 240–251, 2025.
  28. 28.Liang, K., Meng, L., Liu, M., Liu, Y., Tu, W., Wang, S., Zhou, S., and Liu, X. Learn from relational correlations and periodic events for temporal knowledge graph reasoning. In Proceedings of the 46th international ACM SIGIR conference on research and development in information retrieval, pp. 1559–1568, 2023.
  29. 29.Liang, K., Meng, L., Liu, Y., Liu, M., Wei, W., Liu, S., Tu, W., Wang, S., Zhou, S., and Liu, X. Simple yet effective: Structure guided pre-trained transformer for multi-modal knowledge graph reasoning. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 1554–1563, 2024.
  30. 30.Liu, B., Zhang, J., Lin, F., Yang, C., Peng, M., and Yin, W. Symagent: A neural-symbolic self-learning agent framework for complex reasoning over knowledge graphs. In Proceedings of the ACM on Web Conference 2025, pp. 98–108, 2025.
  31. 31.Luo, L., Li, Y.-F., Haffari, G., and Pan, S. Reasoning on graphs: Faithful and interpretable large language model reasoning. In International Conference on Learning Representations, 2024.
  32. 32.Luo, L., Zhao, Z., Haffari, G., Phung, D., Gong, C., and Pan, S. Gfm-rag: Graph foundation model for retrieval augmented generation. arXiv preprint arXiv:2502.01113, 2025.
  33. 33.Lv, Q., Wang, J., Chen, H., Li, B., Zhang, Y., and Wu, F. Coarse-to-fine highlighting: Reducing knowledge hallucination in large language models. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=JCG0KTPVYy.
  34. 34.Ma, J., Gao, Z., Chai, Q., Sun, W., Wang, P., Pei, H., Tao, J., Song, L., Liu, J., Zhang, C., et al. Debate on graph: a flexible and reliable reasoning framework for large language models. arXiv preprint arXiv:2409.03155, 2024.
  35. 35.Mavromatis, C. and Karypis, G. Rearev: Adaptive reasoning for question answering over knowledge graphs. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 2447–2458, 2022.
  36. 36.Mavromatis, C. and Karypis, G. Gnn-rag: Graph neural retrieval for large language model reasoning. arXiv preprint arXiv:2405.20139, 2024.
  37. 37.Meta. Build the future of ai with meta llama 3, 2024. URL https://llama.meta.com/llama3/.
  38. 38.Nguyen, T., Luo, L., Shiri, F., Phung, D., Li, Y.-F., Vu, T.-T., and Haffari, G. Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics ACL 2024, pp. 2862–2883, Bangkok, Thailand and virtual meeting, August 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.findings-acl.168.
  39. 39.OpenAI. Introducing chatgpt, 2022. URL https://openai.com/index/chatgpt/.
  40. 40.OpenAI. Hello gpt-4o, 2024a. URL https://openai.com/index/hello-gpt-4o/.
  41. 41.OpenAI. New embedding models and api updates, 2024b. URL https://openai.com/index/new-embedding-models-and-api-updates/.
  42. 42.OpenAI. Learning to reason with llms, 2024c. URL https://openai.com/index/learning-to-reason-with-llms/.
  43. 43.Pan, S., Luo, L., Wang, Y., Chen, C., Wang, J., and Wu, X. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering (TKDE), 2024.
  44. 44.Qiao, S., Ou, Y., Zhang, N., Chen, X., Yao, Y., Deng, S., Tan, C., Huang, F., and Chen, H. Reasoning with language model prompting: A survey. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5368–5393, 2023.
  45. 45.Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019. URL https://arxiv.org/abs/1908.10084.
  46. 46.Singh, D., Reddy, S., Hamilton, W., Dyer, C., and Yogatama, D. End-to-end training of multi-document reader and retriever for open-domain question answering. Advances in Neural Information Processing Systems, 34:25968–25981, 2021.
  47. 47.Speer, R., Chin, J., and Havasi, C. Conceptnet 5.5: An open multilingual graph of general knowledge. In Proceedings of the AAAI conference on artificial intelligence, volume 31, 2017.
  48. 48.Stanovich, K., West, R., and Hertwig, R. Individual differences in reasoning: Implications for the rationality debate?-open peer commentary-the questionable utility of cognitive ability in explaining cognitive illusions. 2000.
  49. 49.Sui, Y., He, Y., Liu, N., He, X., Wang, K., and Hooi, B. Fidelis: Faithful reasoning in large language model for knowledge graph question answering. arXiv preprint arXiv:2405.13873, 2024.
  50. 50.Sun, H., Dhingra, B., Zaheer, M., Mazaitis, K., Salakhutdinov, R., and Cohen, W. Open domain question answering using early fusion of knowledge bases and text. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4231–4242, 2018.
  51. 51.Sun, J., Xu, C., Tang, L., Wang, S., Lin, C., Gong, Y., Ni, L., Shum, H.-Y., and Guo, J. Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph. In The Twelfth International Conference on Learning Representations, 2024.
  52. 52.Talmor, A. and Berant, J. The web as a knowledge-base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 641–651, 2018.
  53. 53.Talmor, A., Herzig, J., Lourie, N., and Berant, J. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4149–4158, 2019.
  54. 54.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  55. 55.Wang, J., Sun, K., Luo, L., Wei, W., Hu, Y., Liew, A. W.-C., Pan, S., and Yin, B. Large language models-guided dynamic adaptation for temporal knowledge graph reasoning. Thirty-Eighth Annual Conference on Neural Information Processing Systems, 2024a.
  56. 56.Wang, K., Duan, F., Wang, S., Li, P., Xian, Y., Yin, C., Rong, W., and Xiong, Z. Knowledge-driven cot: Exploring faithful reasoning in llms for knowledge-intensive question answering. arXiv preprint arXiv:2308.13259, 2023.
  57. 57.Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2024b.
  58. 58.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022.
  59. 59.Wikipedia contributors. Trie. https://en.wikipedia.org/wiki/Trie, 2024. Accessed: 2024-09-11.
  60. 60.Wu, Z., Pan, S., Chen, F., Long, G., Zhang, C., and Philip, S. Y. A comprehensive survey on graph neural networks. IEEE transactions on neural networks and learning systems, 32(1):4–24, 2020.
  61. 61.Xia, F., Liu, J., Nie, H., Fu, Y., Wan, L., and Kong, X. Random walks: A review of algorithms and applications. IEEE Transactions on Emerging Topics in Computational Intelligence, 4(2):95–107, 2019.
  62. 62.Xie, X., Zhang, N., Li, Z., Deng, S., Chen, H., Xiong, F., Chen, M., and Chen, H. From discrimination to generation: Knowledge graph completion with generative transformer. In Companion Proceedings of the Web Conference 2022, pp. 162–165, 2022.
  63. 63.Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., and Fan, Z. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024a.
  64. 64.Yang, R., Liu, H., Zeng, Q., Ke, Y. H., Li, W., Cheng, L., Chen, Q., Caverlee, J., Matsuo, Y., and Li, I. Kg-rank: Enhancing large language models for medical qa with knowledge graphs and ranking techniques. arXiv preprint arXiv:2403.05881, 2024b.
  65. 65.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 2024.
  66. 66.Yasunaga, M., Ren, H., Bosselut, A., Liang, P., and Leskovec, J. Qa-gnn: Reasoning with language models and knowledge graphs for question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 535–546, 2021.
  67. 67.Yih, W.-t., Richardson, M., Meek, C., Chang, M.-W., and Suh, J. The value of semantic parse labeling for knowledge base question answering. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 201–206, 2016.
  68. 68.Yu, P., Wang, T., Golovneva, O., AlKhamissi, B., Verma, S., Jin, Z., Ghosh, G., Diab, M., and Celikyilmaz, A. Alert: Adapting language models to reasoning tasks. arXiv preprint arXiv:2212.08286, 2022.
  69. 69.Zhang, J., Zhang, X., Yu, J., Tang, J., Tang, J., Li, C., and Chen, H. Subgraph retrieval enhanced model for multi-hop knowledge base question answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5773–5784, 2022.
  70. 70.Zhang, P., Xiao, S., Liu, Z., Dou, Z., and Nie, J.-Y. Retrieve anything to augment large language models. arXiv preprint arXiv:2310.07554, 2023.
  71. 71.Zhu, Y., Qiao, S., Ou, Y., Deng, S., Zhang, N., Lyu, S., Shen, Y., Liang, L., Gu, J., and Chen, H. Knowagent: Knowledge-augmented planning for llm-based agents. arXiv preprint arXiv:2403.03101, 2024.

Citation

MLA
Luo, L., et al. “Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2410.13080v2.
APA
Luo, L., Zhao, Z., Haffari, G., Li, Y.-F., Gong, C., & Pan, S. (2024). Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models. arXiv. http://arxiv.org/abs/2410.13080v2
Chicago
Luo, L., Z. Zhao, G. Haffari, Y.-F. Li, C. Gong, and S. Pan. 2024. “Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models”. arXiv. http://arxiv.org/abs/2410.13080v2.
Harvard
Luo, L. et al. (2024) “Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.13080v2.
Vancouver
1. Luo L, Zhao Z, Haffari G, Li Y-F, Gong C, Pan S (2024) Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models. arXiv

BibTeX

@article{luo2024graph,
  title = {Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models},
  author = {Luo, Linhao and Zhao, Zicheng and Haffari, Gholamreza and Li, Yuan-Fang and Gong, Chen and Pan, Shirui},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.13080v2},
  eprint = {2410.13080}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/