MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
Yilin WenZifeng WangJimeng Sun
Proposes a plug-and-play prompting framework that combines external knowledge graphs with large language models to construct transparent reasoning pathways and reduce hallucinations in complex question answering.
Large language models often struggle in high-stakes fields like healthcare because they can produce inaccurate facts, rely on outdated internal information, and obscure their underlying reasoning. While external structured databases known as knowledge graphs provide factual, auditable connections between concepts, existing retrieval methods usually flatten these graphs into plain text, discarding structural relationships and failing when retrieved data is irrelevant or imperfect.
The article demonstrates and evaluates MindMap, a prompting framework designed to enable fixed large language models to ingest structured knowledge graphs, reason jointly across external facts and internal knowledge, and generate transparent visual reasoning paths.
The researchers developed a three-stage prompting pipeline: extracting key entities from user queries to retrieve path- and neighbor-based subgraphs from an external knowledge base, prompting the model to aggregate these paths into natural-language reasoning graphs, and instructing the model to synthesize the final answer along with an explicit decision tree that tracks evidence sources. The framework was evaluated across three medical question-answering benchmarks featuring clinical consultations, multi-turn dialogues, and pharmacist licensing examination questions, comparing against standard base models and several text- and graph-retrieval techniques.
First, the approach significantly improved factual reliability, securing the top average ranking from automated expert evaluation and reducing factual hallucination across clinical dialogue benchmarks. Second, when tested on examination questions with intentionally mismatched or noisy graph data, the framework achieved 61.7% accuracy, outperforming the base model at 52.2% and retrieval baselines that scored between 42.0% and 54.2%. Third, pairwise evaluations showed the framework consistently won over baseline methods across disease diagnosis and treatment recommendations, maintaining overall win rates between 78% and 88% on clinical consultations. Finally, ablation analysis revealed that combining both multi-hop path exploration and local neighbor exploration was essential to reduce reasoning errors.
These findings indicate that structured graph prompting allows organizations to deploy artificial intelligence in sensitive domains without fine-tuning model parameters, reducing deployment costs while mitigating the safety and compliance risks associated with false model outputs. Unlike standard retrieval systems that blindly trust external passages, this approach effectively blends internal model reasoning with external facts, remaining robust even when retrieved data contains errors.
Organizations evaluating this approach should consider piloting graph-based prompting pipelines in complex reasoning workflows rather than relying solely on unstructured document retrieval. Practitioners should ensure that both path-tracing and local entity exploration are implemented in the query engine, while instructing models to verify retrieved evidence against their pre-trained knowledge base.
The primary limitations involve the risk of propagating outdated data present within source knowledge bases and the potential for visual reasoning structures to become overly complex for end-users to interpret. Confidence in these results is high across medical question-answering tasks, though production deployments in clinical environments require continued caution and human expert oversight.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). Graph of Thoughts introduces the graph-structured reasoning framework that MindMap adapts by grounding LLM thought pathways in knowledge-graph structure.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This roadmap lays out the knowledge-graph and language-model integration strategies that provide essential context for MindMap’s graph-augmented prompting approach.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Chain-of-Thought Prompting establishes the intermediate-reasoning prompting paradigm that MindMap extends into ontology-guided graph reasoning.
- Paper: JointLK: Joint Reasoning with Language Models and Knowledge Graphs for Commonsense Question Answering, Yueqing Sun et al. (2022). JointLK demonstrates how language models and knowledge graphs can jointly support reasoning, giving useful methodological precedent for MindMap’s integration.
No sufficiently relevant recommendations were found.
