Generate-on-Graph: Treat LLM as both Agent and KG for Incomplete Knowledge Graph Question Answering
Yao XuShizhu HeJiabei ChenZihao WangYangqiu SongHanghang TongGuang LiuJun ZhaoKang Liu
Proposes Generate-on-Graph, a training-free framework that enables large language models to answer complex questions over incomplete knowledge graphs by alternating between exploring existing graph triples and generating missing factual connections using internal model knowledge.
Large Language Models often struggle with factual inaccuracies, hallucinations, and gaps in specialized knowledge. While connecting language models to external knowledge graphs can ground responses in structured facts, traditional integration benchmarks assume these knowledge bases are completely comprehensive. In real-world enterprise environments, external knowledge graphs are frequently incomplete, causing standard methods—such as semantic parsing and path retrieval—to fail or produce ungrounded guesses when critical factual connections are missing.
The main objective of the article is to establish a realistic benchmark for question answering over incomplete knowledge graphs and to propose and evaluate Generate-on-Graph, a training-free framework that enables language models to act both as exploring agents and as dynamic knowledge generators.
To simulate real-world conditions, the authors created incomplete benchmark datasets from standard question-answering collections by randomly deleting key factual links at missingness rates between 20% and 80%. They evaluated the Generate-on-Graph framework against standard prompting, semantic parsing, and retrieval-augmented baselines using multiple language model backbones, including GPT-3.5, GPT-4, and open-source models. The framework operates through an iterative cycle where the model plans the next reasoning step, searches the graph using structured tools, and actively generates missing factual triples using internal knowledge whenever the graph lacks necessary links.
The evaluation yielded several key findings. First, Generate-on-Graph consistently outperformed all baseline methods across both complete and incomplete graph settings, achieving top accuracy scores of 84.4% on WebQSP and 75.2% on Complex WebQuestions when powered by GPT-4. Second, conventional methods experienced severe performance drops under incomplete graphs; for example, semantic parsing accuracy fell from 76.5% to 39.3% when 40% of key facts were missing. Third, Generate-on-Graph maintained robust performance across varying degrees of missing information, delivering an average accuracy improvement of 5.0% over competing methods on complex questions. Finally, ablation studies showed that dynamically generating missing facts with contextual graph cues significantly enhanced accuracy, though retrieving excessive, irrelevant graph context introduced noise that degraded performance.
These findings indicate that treating language models as both navigators and supplemental knowledge sources bridges critical data gaps in automated reasoning systems. For organizations deploying knowledge-assisted artificial intelligence, this approach reduces the cost and burden of maintaining exhaustive databases while mitigating the risk of query failures caused by incomplete records. However, because generating missing facts relies on model memory, systems remain susceptible to hallucinations and phrasing mismatches.
Organizations seeking to implement knowledge graph question answering should adopt hybrid agent-generator architectures rather than rigid semantic parsers. Before full deployment, teams should run controlled pilots to tune context retrieval thresholds and implement automated validation routines to check model-generated facts against known organizational constraints.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). StructGPT establishes the iterative tool-use approach to LLM question answering over structured data that Generate-on-Graph adapts to the harder case of missing graph facts.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). ReAct introduces the interleaving of language-model reasoning with external actions that helps situate Generate-on-Graph’s agent-driven graph exploration.
- Paper: Modeling Relational Data with Graph Convolutional Networks, Michael Schlichtkrull et al. (2018). R-GCN provides a foundational account of inferring missing knowledge-graph links, clarifying the graph incompleteness problem that Generate-on-Graph addresses through language-model generation.
No sufficiently relevant recommendations were found.
