ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph
Jinhao JiangKun ZhouWayne Xin ZhaoYaliang LiJi-Rong Wen
Proposes a unified pre-trained language model that performs structural subgraph reasoning directly via a specialized self-attention mechanism, outperforming traditional two-module graph neural network approaches for knowledge graph question answering while updating fewer parameters.
Organizations increasingly rely on automated question answering over knowledge graphs to extract factual insights from massive structured databases. Traditional systems split this challenge across two separate components: a pre-trained language model that interprets the user's natural language question, and a distinct graph neural network that traverses graph connections to locate answers. However, this dual-architecture design prevents seamless knowledge sharing and fine-grained interaction between question text and graph data, creating engineering complexity and sub-optimal reasoning performance on complex, multi-hop queries.
The article introduces and evaluates ReasoningLM, a unified framework that empowers a standard pre-trained language model to directly execute structured graph reasoning without relying on external graph neural network modules.
To evaluate this framework, the authors used a multi-stage approach. First, graph data was converted into sequential inputs using a breadth-first search method and processed using a novel subgraph-aware self-attention mechanism. This mechanism constrains attention masks so the language model can replicate graph message-passing while integrating textual context. Second, the authors created an adaptation training set of 20,000 subgraphs paired with synthetic questions generated using ChatGPT at a nominal cost of fifteen dollars. Finally, the system was evaluated across three widely recognized benchmark datasets (WebQuestionsSP, Complex WebQuestions 1.1, and MetaQA), updating only about one million parameters via lightweight adapter modules during downstream tuning.
The findings show that ReasoningLM significantly outperforms existing methods. On the challenging Complex WebQuestions dataset, the model achieved a 69.0% top-1 accuracy (Hits@1), representing an approximate 36.1% relative improvement over the previous best-performing baseline (UniKGQA at 50.7%). On the WebQuestionsSP dataset, it reached 78.5% top-1 accuracy, exceeding top baselines by 4.5% relative margin. Standalone large language models without graph grounding (such as ChatGPT and Davinci-003) struggled considerably on complex multi-hop queries, scoring below 45% on multi-hop benchmarks. Furthermore, ablation analyses confirmed that both the structural attention masking and the synthetic adaptation tuning are essential to performance, and the framework maintained consistent gains when implemented across multiple base language model architectures, including RoBERTa, BERT, and DeBERTa.
These results demonstrate that language models can be directly adapted for structured topological reasoning without altering their core architectures. For practical operations, this unified approach reduces operational complexity by replacing dual-model pipelines with a single model. It also significantly lowers training costs and data requirements: ReasoningLM matches or exceeds state-of-the-art benchmarks using as few as 5,000 adaptation samples and only a fraction of target task training data, making it highly suitable for low-resource or domain-specific applications.
Decision-makers should consider adopting single-model structured reasoning architectures to simplify enterprise question answering infrastructure and improve answer accuracy on complex factual queries. For next steps, engineering teams should evaluate lightweight adapter tuning on existing internal language model deployments before investing in complex multi-component graph network pipelines.
Decision-makers should note certain operational boundaries. Standard language model context windows (such as 512 tokens) restrict the maximum size of the retrieved graph that can be processed at one time, necessitating efficient initial graph retrieval. Additionally, while confidence in question answering benchmarks is high, the approach has not yet been evaluated on other structured graph tasks, such as knowledge graph completion, or scaled to billion-parameter language models due to computational resource constraints.
- Paper: Modeling Relational Data with Graph Convolutional Networks, Michael Schlichtkrull et al. (2018). Its relational graph convolution and message-passing formulation clarify the structural operations that ReasoningLM’s graph-aware attention is designed to reproduce.
- Paper: Deep Bidirectional Language-Knowledge Graph Pretraining, Michihiro Yasunaga et al. (2022). DRAGON establishes the deep language-model/knowledge-graph fusion context that makes ReasoningLM’s unified structural reasoning approach easier to understand.
- Paper: Knowledge Base Question Answering by Case-based Reasoning over Subgraphs, Rajarshi Das et al. (2022). CBR-SUBG provides a useful earlier example of subgraph-based reasoning for WebQuestionsSP and MetaQA, benchmarks central to ReasoningLM’s evaluation.
- Paper: TIARA: Multi-grained Retrieval for Robust Question Answering over Large Knowledge Base, Yiheng Shu et al. (2022). TIARA’s retrieval-and-query-decoding pipeline supplies the prior knowledge-base question-answering approach that ReasoningLM replaces with direct structural reasoning in a language model.
- Paper: RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering, Xi Ye et al. (2022). RnG-KBQA illustrates the earlier ranking-and-generation strategy for multi-hop knowledge-base questions that provides context for ReasoningLM’s unified alternative.
- Paper: Generate-on-Graph: Treat LLM as both Agent and KG for Incomplete Knowledge Graph Question Answering, Yao Xu et al. (2024). Generate-on-Graph extends language-model question answering over knowledge graphs to incomplete graphs by letting the model actively supply missing factual links.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Plan-on-Graph continues knowledge-graph question answering with adaptive exploration and self-correction rather than ReasoningLM’s subgraph-encoded structural attention.
- Paper: Multimodal Reasoning with Multimodal Knowledge Graph, Junlin Lee et al. (2024). MR-MKG generalizes language-model reasoning over graph substructures from text-only knowledge graphs to multimodal graphs containing visual context.
- Paper: LLaGA: Large Language and Graph Assistant, Runjin Chen et al. (2024). LLaGA carries structural graph representations into language models beyond question answering, applying them to graph classification, link prediction, and summarization.
