MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models

Yilin WenZifeng WangJimeng Sun

article2024ACL180 citations

Proposes a plug-and-play prompting framework that combines external knowledge graphs with large language models to construct transparent reasoning pathways and reduce hallucinations in complex question answering.

Listen

Large language models often struggle in high-stakes fields like healthcare because they can produce inaccurate facts, rely on outdated internal information, and obscure their underlying reasoning. While external structured databases known as knowledge graphs provide factual, auditable connections between concepts, existing retrieval methods usually flatten these graphs into plain text, discarding structural relationships and failing when retrieved data is irrelevant or imperfect.

The article demonstrates and evaluates MindMap, a prompting framework designed to enable fixed large language models to ingest structured knowledge graphs, reason jointly across external facts and internal knowledge, and generate transparent visual reasoning paths.

The researchers developed a three-stage prompting pipeline: extracting key entities from user queries to retrieve path- and neighbor-based subgraphs from an external knowledge base, prompting the model to aggregate these paths into natural-language reasoning graphs, and instructing the model to synthesize the final answer along with an explicit decision tree that tracks evidence sources. The framework was evaluated across three medical question-answering benchmarks featuring clinical consultations, multi-turn dialogues, and pharmacist licensing examination questions, comparing against standard base models and several text- and graph-retrieval techniques.

First, the approach significantly improved factual reliability, securing the top average ranking from automated expert evaluation and reducing factual hallucination across clinical dialogue benchmarks. Second, when tested on examination questions with intentionally mismatched or noisy graph data, the framework achieved 61.7% accuracy, outperforming the base model at 52.2% and retrieval baselines that scored between 42.0% and 54.2%. Third, pairwise evaluations showed the framework consistently won over baseline methods across disease diagnosis and treatment recommendations, maintaining overall win rates between 78% and 88% on clinical consultations. Finally, ablation analysis revealed that combining both multi-hop path exploration and local neighbor exploration was essential to reduce reasoning errors.

These findings indicate that structured graph prompting allows organizations to deploy artificial intelligence in sensitive domains without fine-tuning model parameters, reducing deployment costs while mitigating the safety and compliance risks associated with false model outputs. Unlike standard retrieval systems that blindly trust external passages, this approach effectively blends internal model reasoning with external facts, remaining robust even when retrieved data contains errors.

Organizations evaluating this approach should consider piloting graph-based prompting pipelines in complex reasoning workflows rather than relying solely on unstructured document retrieval. Practitioners should ensure that both path-tracing and local entity exploration are implemented in the query engine, while instructing models to verify retrieved evidence against their pre-trained knowledge base.

The primary limitations involve the risk of propagating outdated data present within source knowledge bases and the potential for visual reasoning structures to become overly complex for end-users to interpret. Confidence in these results is high across medical question-answering tasks, though production deployments in clinical environments require continued caution and human expert oversight.

No sufficiently relevant recommendations were found.

Cover for MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models

Abstract

Large language models (LLMs) have achieved remarkable performance in natural language understanding and generation tasks. However, they often suffer from limitations such as difficulty in incorporating new knowledge, generating hallucinations, and explaining their reasoning process. To address these challenges, we propose a novel prompting pipeline, named MindMap, that leverages knowledge graphs (KGs) to enhance LLMs' inference and transparency. Our method enables LLMs to comprehend KG inputs and infer with a combination of implicit and external knowledge. Moreover, our method elicits the mind map of LLMs, which reveals their reasoning pathways based on the ontology of knowledge. We evaluate our method on diverse question & answering tasks, especially in medical domains, and show significant improvements over baselines. We also introduce a new hallucination evaluation benchmark and analyze the effects of different components of our method. Our results demonstrate the effectiveness and robustness of our method in merging knowledge from LLMs and KGs for combined inference. To reproduce our results and extend the framework further, we make our codebase available at https://github.com/wyl-willing/MindMap.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Step I: Evidence Graph Mining
  • 3.1.1 Entity Recognition
  • 3.1.2 Evidence Sub-graphs Exploration
  • 3.2 Step II: Evidence Graph Aggregation
  • 3.3 Step III: LLM Reasoning with Mind Map
  • 3.3.1 Prompting for Graph Reasoning
  • 3.3.2 Synergistic Inference with LLM and KG Knowledge
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Medical Question Answering
  • 4.2.1 Evaluation Metrics
  • 4.2.2 Results
  • 4.3 Long Dialogue Question Answering
  • 4.4 Generate with Mismatch Knowledge from KG
  • 4.4.1 Evaluation Metrics
  • 4.4.2 Results
  • 4.5 Ablation Study
  • 4.6 In-depth Analysis
  • 4.6.1 How does MindMap perform without correct KG knowledge?
  • 4.6.2 How robust is MindMap to unmatched fact queries?
  • 4.6.3 How does MindMap aggregate evidence graphs considering entity semantics?
  • 4.6.4 How does MindMap visualize the inference process and evidence sources?
  • 4.6.5 How does MindMap leverage LLM knowledge for various tasks?
  • 5 Conclusion
  • References
  • A Construction of Datasets
  • B Implementation of Knowledge Graph
  • C Implementation of Baselines
  • D Prompt Engine
  • E Evidence Subgraphs Exploration
  • F In-depth Analysis
  • G Pairwise Ranking Evaluation
  • H Limitations and Potential Risks

Knowls

  1. Knowl 1 — MindMap combines KG evidence mining, graph aggregation, and LLM reasoning

    model/method

    MindMap is a prompting pipeline for answering questions with both a fixed large language model’s (LLM’s) implicit knowledge and explicit knowledge retrieved from a knowledge graph (KG). Given a question, it identifies relevant entities, mines path-based and neighbor-based evidence subgraphs from a source KG, and asks an LLM to consolidate those subgraphs into indexed natural-language evidence. A final LLM prompt uses the question and the evidence to produce an answer, an inference trace, and a mind map that connects reasoning steps to their evidence sources. The workflow illustration on page 13 depicts this flow from question entities through evidence mining and aggregation to the final answer and provenance-bearing mind map.

  2. Knowl 2 — Question entities are identified by LLM extraction and BERT-based KG matching

    model/method

    For a question, MindMap first prompts an LLM to extract key medical entities. The extraction prompt includes the question, a cue phrase requesting extra entities, and two examples. It then encodes the extracted entity phrases and the entities in the source knowledge graph into dense representations and computes BERT-based cosine similarities between them. The highest-similarity KG entities are selected as the query entity set used to initiate evidence search. The paper does not specify a similarity threshold or a fixed number of selected entities.

  3. Knowl 3 — MindMap mines both connecting paths and local KG neighborhoods

    model/method

    MindMap represents source-KG facts as triples of the form (head entity, relation, tail entity), and constructs two complementary types of evidence subgraphs around the question entities. Path-based exploration searches for connecting pathways between query entities, allowing at most kk hops when seeking the next candidate entity; successful connections update the current start entity and remove the reached candidate, while unconnected path segments are stored and joined into evidence subgraphs. Neighbor-based exploration adds one-hop triples around each query entity and further expands a neighbor when it is judged semantically relevant to the question. Newly introduced intermediate entities are added to the query entity set. To limit evidence volume while retaining diversity, the resulting path and neighbor subgraphs are clustered and sampled by head entity. The paper does not give a numerical value for kk or detailed clustering and sampling parameters.

  4. Knowl 4 — LLMs aggregate indexed evidence and generate a provenance-linked mind map

    model/method

    For evidence aggregation, MindMap selects at least kk path-based and kk neighbor-based evidence subgraphs, formats each as an entity-relation chain, and assigns it a source identifier such as P1 or N1. An LLM converts each chain into a concise natural-language description, producing separate path-based and neighbor-based reasoning-graph inputs. In the final prompt, the LLM receives the question and these descriptions, along with instructions and examples to combine external evidence with its own knowledge, reason over the information, and track sources. The requested output has three parts: a summary answer, an inference process that identifies the evidence chains used, and a decision-tree-style mind map. Mind-map nodes can come from the evidence or from the LLM’s own knowledge-based expansions, with evidence-source labels attached to support traceability.

  5. Knowl 5 — Evaluation uses three medical QA datasets and two constructed knowledge graphs

    experimental setup

    MindMap was evaluated on three test sets: English patient-clinician question answering (GenMedGPT-5k), Chinese long-dialogue question answering (CMCQA), and five-option Chinese pharmacist-examination questions with explanations (ExplainCPE). The experiments used an English medical KG, EMCKG, for GenMedGPT-5k and a Chinese medical KG, CMCKG, for CMCQA and ExplainCPE; the latter pairing was used to test performance when retrieved knowledge could be mismatched. Retrieval-based methods used GPT-3.5-turbo-0613 as their LLM backbone. Comparisons included vanilla GPT-3.5 and GPT-4, tree-of-thought prompting, BM25 document retrieval, embedding-based document retrieval, and KG retrieval.

    Dataset Test questions KG nodes KG triples KG relations
    GenMedGPT-5k 714 1122 5802 6
    CMCQA 468 62282 506490 12
    ExplainCPE 400 62282 506490 12
  6. Knowl 6 — MindMap improves medical QA scores on GenMedGPT-5k

    empirical result

    On GenMedGPT-5k, which evaluates disease diagnosis and recommendations for tests and medication, MindMap achieved the best reported BERTScore F1 and GPT-4 average ranking among the compared prompting and retrieval methods. For GPT-4 ranking, lower values indicate a more preferred answer. The paper also reports a keyword-based hallucination quantification score: it extracts keywords from answers and labels with an NER-MT5 model, joins them into keyword strings, and compares the strings using TF-IDF similarity; the authors state that a lower score indicates more hallucination. MindMap’s score was higher than those of the listed baselines.

    Method BERT precision BERT recall BERT F1 GPT-4 ranking (avg.) Hallucination quantification
    MindMap 0.7936 0.7977 0.7954 1.8725 0.6070
    GPT-3.5 0.7612 0.8003 0.7800 4.8571 0.5563
    Tree-of-thought (TOT) 0.7202 0.7949 0.7554 - 0.5483
    GPT-4 0.7689 0.7893 0.7786 4.1764 0.5577
    BM25 Retriever 0.7693 0.7981 0.7831 3.5546 0.5834
    Embedding Retriever 0.7690 0.8038 0.7857 3.1232 0.5886
    KG Retriever 0.7717 0.8030 0.7868 3.4159 0.5871
  7. Knowl 7 — CMCQA results show favorable rankings but a smaller advantage

    empirical result

    On CMCQA, a Chinese long-dialogue QA test set, MindMap’s BERTScore F1 was close to the baselines, and its GPT-4 ranking tied the KG Retriever at 2.3 (lower is better). In GPT-4 pairwise judgments of disease-diagnosis and medication recommendations, MindMap had more wins than losses against each listed baseline, though ties were common. The authors attribute the narrower gap than on GenMedGPT-5k in part to CMCKG not covering all facts needed for CMCQA questions; they report that combining KG evidence with the LLM’s implicit knowledge still outperformed retrieval-only approaches in these comparisons.

    Method BERT precision BERT recall BERT F1 GPT-4 ranking (avg.)
    MindMap 0.9415 0.9321 0.9367 2.3
    GPT-3.5 0.9385 0.9361 0.9372 3.4
    GPT-4 0.9355 0.9358 0.9356 3.6
    BM25 Retriever 0.9365 0.9348 0.9356 3.7
    Embedding Retriever 0.9357 0.9334 0.9345 5.4
    KG Retriever 0.9318 0.9348 0.9332 2.3

    The GPT-4 pairwise win/tie/loss percentages below are averages over disease diagnosis and drug recommendation, comparing MindMap with each baseline.

    Baseline Win (%) Tie (%) Loss (%)
    GPT-3.5 41.5 35.29 23.21
    BM25 Retriever 39.045 39.775 21.175
    Embedding Retriever 41.075 37.43 21.495
    KG Retriever 39.365 38.385 22.25
    GPT-4 36.05 38.49 25.455
  8. Knowl 8 — MindMap is more accurate than retrieval baselines with mismatched KG knowledge

    empirical result

    ExplainCPE evaluates five-way examination questions using CMCKG, where retrieved information can be redundant or mismatched to the question. MindMap reached 61.7% accuracy, exceeding the retrieval baselines, including KG Retriever at 42.0%, but below GPT-4 at 72.0%. The result supports the paper’s qualified claim that MindMap is more robust than direct retrieval prompting when external facts are inaccurate, not that it eliminates errors: 37.7% of MindMap responses were wrong. Accuracy counts correct, wrong, and failed responses as reported below.

    Method Correct (%) Wrong (%) Failed (%)
    GPT-3.5 52.2 47.2 0.5
    BM25 Retriever 50 44.2 5.7
    Embedding Retriever 54.2 45.2 0.5
    KG Retriever 42 44 14
    GPT-4 72 27.7 0.2
    MindMap 61.7 37.7 0.5
  9. Knowl 9 — Using both evidence types improves results over single-type evidence

    empirical result

    On GenMedGPT-5k, the combined MindMap system outperformed path-only and neighbor-only variants on BERTScore F1 and the reported hallucination-quantification score, despite using more tokens. The authors characterize the neighbor-only variant as somewhat better than path-only for factual accuracy, while path-based evidence is useful for finding relevant external information but can be inadequate for multi-hop recommendations. Separately, on ExplainCPE, removing the instruction to “combine with the knowledge you already have” lowered accuracy from 61.7% to 53.5%, an 8.2 percentage-point difference. The measurements below report average tokens, BERTScore components, and hallucination quantification for the GenMedGPT-5k variants.

    Variant Tokens (avg.) BERT precision BERT recall BERT F1 Hallucination quantification
    Path-only 1028 0.6310 0.7885 0.7002 0.3854
    Neighbor-only 1236 0.6393 0.7930 0.7072 0.3894
    MindMap 1431 0.7938 0.7987 0.7960 0.5890
  10. Knowl 10 — KG errors, integration failures, and opaque mind maps remain risks

    limitation

    The paper identifies several risks rather than claiming that KG prompting resolves them. A KG can contain biased, outdated, incomplete, or erroneous facts that affect generated answers. Combining KG content and LLM reasoning can also introduce logical inconsistencies, particularly for complex or vague queries. Dependence on a KG may hurt performance when that graph is unavailable or lacks relevant information. Finally, a mind map may fail to improve practical interpretability if its structure is too complex or obscure for users to understand.

Coverage note — Detailed case-study transcripts and individual example questions were omitted because they illustrate, rather than add to, the general method and quantitative findings.

References

  1. 1.Rohaid Ali, Oliver Y Tang, Ian D Connolly, Jared S Fridley, John H Shin, Patricia L Zadnik Sullivan, Deus Cielo, Adetokunbo A Oyelese, Curtis E Doberstein, Albert E Telfeian, et al. 2022. Performance of chatgpt, gpt-4, and google bard on a neurosurgery oral boards preparation question bank. Neurosurgery, pages 10–1227.
  2. 2.Samy Ateia and Udo Kruschwitz. 2023. Is chatgpt a biomedical expert?–exploring the zero-shot performance of current gpt models in biomedical tasks. arXiv preprint arXiv:2306.16108.
  3. 3.Jinheon Baek, Alham Fikri Aji, and Amir Saffari. 2023. Knowledge-augmented language model prompting for zero-shot knowledge graph question answering. arXiv preprint arXiv:2306.04136.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901.
  5. 5.Yihan Cao, Yanbin Kang, and Lichao Sun. 2023. Instruction mining: High-quality instruction data selection for large language models. arXiv preprint arXiv:2307.06290.
  6. 6.Zhikai Chen, Haitao Mao, Hang Li, Wei Jin, Hongzhi Wen, Xiaochi Wei, Shuaiqiang Wang, Dawei Yin, Wenqi Fan, Hui Liu, et al. 2023. Exploring the potential of large language models (llms) in learning on graphs. arXiv preprint arXiv:2307.03393.
  7. 7.Nurendra Choudhary and Chandan K Reddy. 2023. Complex logical reasoning over knowledge graphs using large language models. arXiv preprint arXiv:2305.01157.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Falcon Z. Dai. 2020. Word2vec conjecture and a limitative result.
  10. 10.Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020. A survey of the state of explainable ai for natural language processing. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 447–459.
  11. 11.Jiayan Guo, Lun Du, and Hengyu Liu. 2023. GPT4Graph: Can large language models understand graph structured data? an empirical evaluation and benchmarking. arXiv preprint arXiv:2305.15066.
  12. 12.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  13. 13.Zhen Jia, Soumajit Pramanik, Rishiraj Saha Roy, and Gerhard Weikum. 2021. Complex temporal question answering on knowledge graphs. In Proceedings of the 30th ACM international conference on information & knowledge management, pages 792–802.
  14. 14.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  15. 15.Xun Liang, Shichao Song, Simin Niu, Zhiyu Li, Feiyu Xiong, Bo Tang, Zhaohui Wy, Dawei He, Peng Cheng, Zhonghao Wang, and Haiying Deng. 2023. Uhgeval: Benchmarking the hallucination of chinese large language models via unconstrained generation.
  16. 16.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023a. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  17. 17.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35.
  18. 18.Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2020. K-BERT: Enabling language representation with knowledge graph. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 2901–2908.
  19. 19.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  20. 20.Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2023. Unifying large language models and knowledge graphs: A roadmap. arXiv preprint arXiv:2306.08302.
  21. 21.Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813.
  22. 22.Anastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa, Mike Lewis, and Amjad Almahairi. 2022. Progressive prompts: Continual learning for language models. In The Eleventh International Conference on Learning Representations.
  23. 23.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910.
  24. 24.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  25. 25.Anil Sharma and Suresh Kumar. 2023. Ontology-based semantic retrieval of documents using word2vec model. Data & Knowledge Engineering, 144:102110.
  26. 26.Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge. Nature, pages 1–9.
  27. 27.Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Heung-Yeung Shum, and Jian Guo. 2023. Think-on-Graph: Deep and responsible reasoning of large language model with knowledge graph. arXiv preprint arXiv:2307.07697.
  28. 28.Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuan-Jing Huang, and Zheng Zhang. 2020. CoLAKE: Contextualized language and knowledge embedding. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3660–3670.
  29. 29.Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. 2021. ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137.
  30. 30.Cunxiang Wang, Sirui Cheng, Zhikun Xu, Bowen Ding, Yidong Wang, and Yue Zhang. 2023. Evaluating open question answering evaluation. arXiv preprint arXiv:2305.12421.
  31. 31.Xiaoyan Wang, Pavan Kapanipathi, Ryan Musa, Mo Yu, Kartik Talamadupula, Ibrahim Abdelaziz, Maria Chang, Achille Fokoue, Bassem Makni, Nicholas Mattei, et al. 2019. Improving natural language inference using external knowledge in the science questions domain. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 7208–7215.
  32. 32.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  33. 33.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models.
  34. 34.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022b. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  35. 35.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023a. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  36. 36.Yao Yao, Zuchao Li, and Hai Zhao. 2023b. Beyond chain-of-thought, effective graph-of-thought reasoning in large language models. arXiv preprint arXiv:2305.16582.
  37. 37.Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D Manning, Percy S Liang, and Jure Leskovec. 2022. Deep bidirectional language-knowledge graph pretraining. Advances in Neural Information Processing Systems, 35:37309–37323.
  38. 38.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: Reasoning with language models and knowledge graphs for question answering. In North American Chapter of the Association for Computational Linguistics (NAACL).
  39. 39.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019a. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  40. 40.Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D Manning, and Jure Leskovec. 2022. GreaseLM: Graph reasoning enhanced language models. In International Conference on Learning Representations.
  41. 41.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019b. ERNIE: Enhanced language representation with informative entities. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1441–1451.

Citation

MLA
Wen, Y., et al. “MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10370–88, https://doi.org/10.18653/v1/2024.acl-long.558.
APA
Wen, Y., Wang, Z., & Sun, J. (2024). MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10370–10388. https://doi.org/10.18653/v1/2024.acl-long.558
Chicago
Wen, Y., Z. Wang, and J. Sun. 2024. “MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10370–88. https://doi.org/10.18653/v1/2024.acl-long.558.
Harvard
Wen, Y., Wang, Z. and Sun, J. (2024) “MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10370–10388. Available at: https://doi.org/10.18653/v1/2024.acl-long.558.
Vancouver
1. Wen Y, Wang Z, Sun J (2024) MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10370–10388

BibTeX

@inproceedings{wen-etal-2024-mindmap,
    title = "{M}ind{M}ap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models",
    author = "Wen, Yilin  and
      Wang, Zifeng  and
      Sun, Jimeng",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.558/",
    doi = "10.18653/v1/2024.acl-long.558",
    pages = "10370--10388"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/