ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph

Jinhao JiangKun ZhouWayne Xin ZhaoYaliang LiJi-Rong Wen

article2023EMNLP66 citations

Proposes a unified pre-trained language model that performs structural subgraph reasoning directly via a specialized self-attention mechanism, outperforming traditional two-module graph neural network approaches for knowledge graph question answering while updating fewer parameters.

Listen

Organizations increasingly rely on automated question answering over knowledge graphs to extract factual insights from massive structured databases. Traditional systems split this challenge across two separate components: a pre-trained language model that interprets the user's natural language question, and a distinct graph neural network that traverses graph connections to locate answers. However, this dual-architecture design prevents seamless knowledge sharing and fine-grained interaction between question text and graph data, creating engineering complexity and sub-optimal reasoning performance on complex, multi-hop queries.

The article introduces and evaluates ReasoningLM, a unified framework that empowers a standard pre-trained language model to directly execute structured graph reasoning without relying on external graph neural network modules.

To evaluate this framework, the authors used a multi-stage approach. First, graph data was converted into sequential inputs using a breadth-first search method and processed using a novel subgraph-aware self-attention mechanism. This mechanism constrains attention masks so the language model can replicate graph message-passing while integrating textual context. Second, the authors created an adaptation training set of 20,000 subgraphs paired with synthetic questions generated using ChatGPT at a nominal cost of fifteen dollars. Finally, the system was evaluated across three widely recognized benchmark datasets (WebQuestionsSP, Complex WebQuestions 1.1, and MetaQA), updating only about one million parameters via lightweight adapter modules during downstream tuning.

The findings show that ReasoningLM significantly outperforms existing methods. On the challenging Complex WebQuestions dataset, the model achieved a 69.0% top-1 accuracy (Hits@1), representing an approximate 36.1% relative improvement over the previous best-performing baseline (UniKGQA at 50.7%). On the WebQuestionsSP dataset, it reached 78.5% top-1 accuracy, exceeding top baselines by 4.5% relative margin. Standalone large language models without graph grounding (such as ChatGPT and Davinci-003) struggled considerably on complex multi-hop queries, scoring below 45% on multi-hop benchmarks. Furthermore, ablation analyses confirmed that both the structural attention masking and the synthetic adaptation tuning are essential to performance, and the framework maintained consistent gains when implemented across multiple base language model architectures, including RoBERTa, BERT, and DeBERTa.

These results demonstrate that language models can be directly adapted for structured topological reasoning without altering their core architectures. For practical operations, this unified approach reduces operational complexity by replacing dual-model pipelines with a single model. It also significantly lowers training costs and data requirements: ReasoningLM matches or exceeds state-of-the-art benchmarks using as few as 5,000 adaptation samples and only a fraction of target task training data, making it highly suitable for low-resource or domain-specific applications.

Decision-makers should consider adopting single-model structured reasoning architectures to simplify enterprise question answering infrastructure and improve answer accuracy on complex factual queries. For next steps, engineering teams should evaluate lightweight adapter tuning on existing internal language model deployments before investing in complex multi-component graph network pipelines.

Decision-makers should note certain operational boundaries. Standard language model context windows (such as 512 tokens) restrict the maximum size of the retrieved graph that can be processed at one time, necessitating efficient initial graph retrieval. Additionally, while confidence in question answering benchmarks is high, the approach has not yet been evaluated on other structured graph tasks, such as knowledge graph completion, or scaled to billion-parameter language models due to computational resource constraints.

arXiv: 2401.00158RUCAIBox/ReasoningLM
Cover for ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph

Abstract

Question Answering over Knowledge Graph (KGQA) aims to seek answer entities for the natural language question from a large-scale Knowledge Graph (KG). To better perform reasoning on KG, recent work typically adopts a pre-trained language model (PLM) to model the question, and a graph neural network (GNN) based module to perform multi-hop reasoning on the KG. Despite the effectiveness, due to the divergence in model architecture, the PLM and GNN are not closely integrated, limiting the knowledge sharing and fine-grained feature interactions. To solve it, we aim to simplify the above two-module approach, and develop a more capable PLM that can directly support subgraph reasoning for KGQA, namely ReasoningLM. In our approach, we propose a subgraph-aware self-attention mechanism to imitate the GNN for performing structured reasoning, and also adopt an adaptation tuning strategy to adapt the model parameters with 20,000 subgraphs with synthesized questions. After adaptation, the PLM can be parameter-efficient fine-tuned on downstream tasks. Experiments show that ReasoningLM surpasses state-of-the-art models by a large margin, even with fewer updated parameters and less training data. Our codes and data are publicly available at https://github.com/RUCAIBox/ReasoningLM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 4 Approach
  • 4.1 Overview
  • 4.2 Adapting PLM for Subgraph Reasoning
  • 4.2.1 BFS-based Subgraph Serialization
  • 4.2.2 Subgraph-aware Self-Attention
  • 4.3 Adaptation Tuning
  • 4.3.1 Tuning Data Construction
  • 4.3.2 Answer Entity Prediction
  • 4.4 Efficient Fine-tuning
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Implementation Details
  • 5.3 Main Results
  • 5.4 Further Analysis
  • 6 Conclusion
  • 7 Limitations
  • Acknowledgments
  • References
  • A Datasets
  • B Baselines
  • C Ablation Study of Retrieval Subgraphs
  • D Prompt for ChatGPT

Knowls

  1. Knowl 1 — Subgraph-aware attention makes a Transformer perform graph reasoning

    model/method

    ReasoningLM encodes a natural-language question and its retrieved knowledge-graph subgraph in one Transformer. Its additive attention mask permits four information flows: question tokens attend to other question tokens; each entity or relation token attends to question tokens; graph tokens attend to tokens adjacent to them in a knowledge-graph triple; and all other flows are blocked. In particular, question tokens cannot attend to subgraph tokens, and graph tokens cannot freely attend to all other graph tokens. An entity can therefore aggregate information from the relations and entities in its incident triples, while a relation can aggregate information from its head and tail entities. Stacking Transformer layers propagates information across multiple graph hops while allowing each graph token to use the question representation.

    For a sequence of LL question and subgraph tokens, let A∈RL×LA\in\mathbb{R}^{L\times L} be the self-attention logits, M∈RL×LM\in\mathbb{R}^{L\times L} the additive mask, and Q,K,VQ,K,V the query, key, and value matrices. The mask has value 00 for permitted interactions and −∞-\infty for blocked ones; attention is computed as

    Attn⁡(Q,K,V)=softmax⁡(A+M)V.\operatorname{Attn}(Q,K,V)=\operatorname{softmax}(A+M)V.

    The softmax is taken over keys for each query token, so blocked interactions receive zero attention weight.

  2. Knowl 2 — BFS serialization preserves subgraph structure in the PLM input

    model/method

    ReasoningLM turns a retrieved subgraph into a sequence by breadth-first traversal beginning at a topic entity. It visits triples whose head is the current entity, then proceeds to triples whose head entities have already been visited. Entities and relations are included only on their first visit, producing a compact sequence in traversal order rather than repeating the full text of every triple. This lets the input length depend on the number of distinct visited entities and relations, while the attention mask separately encodes which graph tokens are adjacent in a triple.

    To represent the sequence in the language model's embedding space, the model tokenizes each entity and relation name into subwords and sums the subword embeddings into one vector for that graph item. It concatenates these graph-item vectors with the question-token embeddings, then adds positional embeddings before the Transformer processes the combined input.

  3. Knowl 3 — Synthetic subgraphs and questions adapt the PLM to graph reasoning

    model/method

    ReasoningLM uses an automatically constructed adaptation corpus to familiarize a pretrained language model with the combined question–subgraph input and its constrained attention. The corpus contains 20,000 question, subgraph, and answer examples drawn from Wikidata; 1,000 examples are reserved for validation. Sampling starts from 2,000 popular entities selected from Wikidata5M. A random walk of at most four hops supplies a reasoning path, whose endpoint is the answer; randomly sampled surrounding entities and relations form the subgraph, with the path guaranteed to remain included.

    Questions are synthesized from the topic entity and reasoning path. The authors use ChatGPT to generate varied, fluent questions that express the path constraints while omitting the answer, at a reported cost of approximately US$15 for the 20,000 questions. They also construct a rule-based question variant from hand-written templates, which provides a cheaper but less varied alternative.

  4. Knowl 4 — Answer-entity prediction trains the adapted model

    equation

    For adaptation, ReasoningLM receives a question and its subgraph and predicts which subgraph entity answers the question. Let LL be the input sequence length, dd the hidden dimension, and H∈RL×dH\in\mathbb{R}^{L\times d} the final Transformer hidden states. A linear prediction layer followed by a softmax produces position scores s∈RLs\in\mathbb{R}^{L}. The training loss is the KL divergence between predicted scores and the ground-truth answer scores s⋆s^\star:

    s=softmax⁡(Linear⁡(H)),Lat=DKL(s,s⋆).s=\operatorname{softmax}(\operatorname{Linear}(H)),\qquad \mathcal{L}_{\mathrm{at}}=D_{\mathrm{KL}}(s,s^\star).

    The ground-truth score marks labeled answer entities, and the loss is computed only at entity positions; question words and relation positions cannot be answers. The same answer-entity prediction task is used when fine-tuning the reasoning stage on downstream KGQA data.

  5. Knowl 5 — Downstream retrieval and reasoning use parameter-efficient adapters

    model/method

    After adaptation, ReasoningLM is applied in two stages. For subgraph retrieval, a retrieval adapter is trained on question–relation pairs drawn from shortest paths between topic and answer entities. At inference, starting from topic entities, the model scores neighboring relations for relevance to the question and iteratively adds the top-scoring relations and their triples to the retrieved subgraph. The number of relations added per step is 15 for WebQSP and CWQ and 3 for MetaQA. For answer reasoning, a reasoning adapter scores entities in the retrieved subgraph and the highest-scoring entity is returned as the answer.

    The reported downstream setup freezes the pretrained model and updates adapter parameters rather than the full PLM; the main comparison reports 1M updated parameters for ReasoningLM. Retrieval training uses AdamW with batch size 10 and learning rate 5×10−55\times10^{-5}. Reasoning training uses AdamW with learning rate 10−410^{-4} and batch sizes 60 for WebQSP, 300 for CWQ, and 4 for MetaQA.

  6. Knowl 6 — ReasoningLM benchmark results on five KGQA settings

    data/table

    The main evaluation uses RoBERTa-base and reports Hits@1 and F1 as percentages on WebQSP and CWQ, and Hits@1 on the one-, two-, and three-hop MetaQA settings. MetaQA uses a one-shot training subset with 161, 210, and 150 examples, respectively. The table compares UniKGQA, a strong PLM-based baseline, with the full-parameter adaptation (FPT) and parameter-efficient adaptation (PET) variants of ReasoningLM. ReasoningLM with FPT and LLM-generated questions achieves the strongest listed WebQSP and CWQ scores, and improves on UniKGQA for MetaQA two-hop and three-hop; its MetaQA one-hop Hits@1 is lower than UniKGQA's.

    Method Updated params WebQSP CWQ MQA-1H MQA-2H MQA-3H
    Hits@1 F1 Hits@1 F1 Hits@1 Hits@1 Hits@1
    UniKGQA 12M 75.1 70.2 50.7 48.0 97.1 98.2 92.6
    ReasoningLM, FPT, LLM-SYN 1M 78.5 71.0 69.0 64.0 96.5 98.3 92.7
    ReasoningLM, FPT, Rule-SYN 1M 78.0 70.5 62.8 55.4 96.1 96.9 91.0
    ReasoningLM, PET, LLM-SYN 1M 76.7 69.1 68.3 62.4 95.7 97.0 90.9

    FPT denotes full-parameter adaptation tuning; PET denotes parameter-efficient adaptation tuning; LLM-SYN and Rule-SYN denote LLM-generated and rule-generated adaptation questions. The result table supports the effectiveness of the full ReasoningLM setup on WebQSP and CWQ while also showing that synthesis and adaptation choices affect results.

  7. Knowl 7 — Both structural attention and adaptation tuning are important

    empirical result

    An ablation on WebQSP and CWQ removes either subgraph-aware self-attention (SA) or adaptation tuning (AT) from ReasoningLM. Both removals reduce Hits@1 and F1, with particularly large CWQ declines when structural attention is removed. The ablation table reports the following percentages:

    Model WebQSP CWQ
    Hits@1 F1 Hits@1 F1
    ReasoningLM 78.5 70.1 69.0 64.0
    Without SA 68.5 63.2 40.5 38.2
    Without AT 67.5 60.4 55.2 43.3

    These comparisons indicate that the graph-constrained attention mechanism and the adaptation procedure each contribute to downstream performance; neither alone accounts for the complete model's results.

  8. Knowl 8 — ReasoningLM retrieval improves other reasoners when they use its subgraphs

    data/table

    On CWQ, the authors evaluate whether ReasoningLM's retrieval stage supplies better subgraphs independently of its own answer reasoner by giving those subgraphs to NSM and UniKGQA. Both baselines improve relative to their own reported scores when using ReasoningLM retrieval: UniKGQA rises from 50.7 to 63.43 Hits@1 and from 48.0 to 57.65 F1; NSM rises from 47.6 to 61.9 Hits@1 and from 42.4 to 50.1 F1. The remaining gap between these models with ReasoningLM subgraphs and ReasoningLM itself also indicates that retrieval alone does not explain the full result.

    CWQ model and subgraph source Hits@1 F1
    ReasoningLM 69.0 64.0
    NSM 47.6 42.4
    NSM using ReasoningLM retrieval 61.9 50.1
    UniKGQA 50.7 48.0
    UniKGQA using ReasoningLM retrieval 63.43 57.65
  9. Knowl 9 — Performance transfers across PLM backbones and improves with more data

    empirical result

    On CWQ, ReasoningLM performs comparably with several PLM backbones, and RoBERTa-large exceeds RoBERTa-base on both reported metrics. These results support the authors' claim that the adaptation method can be used with different PLMs, although the tested models do not establish performance for very large language models.

    PLM backbone Hits@1 F1
    RoBERTa-base 69.0 64.0
    RoBERTa-large 70.0 65.2
    DeBERTa-base 68.1 63.1
    BERT-base 67.4 63.0

    The adaptation-sample analysis on WebQSP and CWQ shows that ReasoningLM reaches performance competitive with UniKGQA after adaptation on 5,000 examples; adding more adaptation examples further improves results before performance stabilizes. In a separate CWQ analysis using the same retrieval model, ReasoningLM outperforms NSM and UniKGQA across the tested amounts of downstream fine-tuning data. The plotted sample-efficiency analyses report trends rather than exact numerical values for each point.

  10. Knowl 10 — Input length and evaluation scope constrain the demonstrated method

    limitation

    ReasoningLM retains the original PLM architecture, so its input-length limit—typically 512 tokens for the models discussed—prevents it from processing arbitrarily large retrieved subgraphs in one input. The authors suggest relative positional embeddings or a better retrieval model that selects a suitably sized subgraph as possible ways to ease this constraint; these remedies are not established by the reported experiments. Evaluation is limited to KGQA datasets, so the method is not empirically validated here for other knowledge-graph reasoning tasks such as commonsense question answering or knowledge-graph completion. The experiments also do not cover PLMs larger than one billion parameters, owing to computational-resource constraints.

Coverage note — No other substantial contributed method or result was omitted; prompt examples and baseline descriptions were excluded because they are implementation detail or comparison background rather than separate contributions.

References

  1. 1.Kurt D. Bollacker, Colin Evans, Praveen K. Paritosh, Tim Sturge, and Jamie Taylor. 2008. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD 2008, Vancouver, BC, Canada, June 10-12, 2008, pages 1247–1250. ACM.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  3. 3.Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou, Juanzi Li, Bin He, and Hanwang Zhang. 2022. KQA pro: A dataset with explicit compositional programs for complex question answering over knowledge base. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 6101–6119. Association for Computational Linguistics.
  4. 4.Zhikai Chen, Haitao Mao, Hongzhi Wen, Haoyu Han, Wei Jin, Haiyang Zhang, Hui Liu, and Jiliang Tang. 2023. Label-free node classification on graphs with large language models (llms). arXiv preprint arXiv:2310.04668.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  6. 6.Vijay Prakash Dwivedi and Xavier Bresson. 2020. A generalization of transformer networks to graphs. CoRR, abs/2012.09699.
  7. 7.Gaole He, Yunshi Lan, Jing Jiang, Wayne Xin Zhao, and Ji-Rong Wen. 2021. Improving multi-hop knowledge base question answering by learning intermediate supervision signals. In WSDM ’21, The Fourteenth ACM International Conference on Web Search and Data Mining, Virtual Event, Israel, March 8-12, 2021, pages 553–561. ACM.
  8. 8.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 2790–2799. PMLR.
  9. 9.Drew A. Hudson and Christopher D. Manning. 2019. Learning by abstraction: The neural state machine. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 5901–5914.
  10. 10.Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Structgpt: A general framework for large language model to reason over structured data. CoRR, abs/2305.09645.
  11. 11.Jinhao Jiang, Kun Zhou, Ji-Rong Wen, and Xin Zhao. 2022a. greattruthsarealwayssimple:great truths are always simple: A rather simple knowledge encoder for enhancing the commonsense reasoning capacity of pre-trained models. In Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 1730–1741. Association for Computational Linguistics.
  12. 12.Jinhao Jiang, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2022b. Unikgqa: Unified retrieval and reasoning for solving multi-hop question answering over knowledge graph. CoRR, abs/2212.00959.
  13. 13.Alexander H. Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston. 2016a. Key-value memory networks for directly reading documents. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 1400–1409. The Association for Computational Linguistics.
  14. 14.Alexander H. Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston. 2016b. Key-value memory networks for directly reading documents. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 1400–1409. The Association for Computational Linguistics.
  15. 15.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  16. 16.Apoorv Saxena, Adrian Kochsiek, and Rainer Gemulla. 2022. Sequence-to-sequence knowledge graph completion and question answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 2814–2828. Association for Computational Linguistics.
  17. 17.Apoorv Saxena, Aditay Tripathi, and Partha P. Talukdar. 2020. Improving multi-hop question answering over knowledge graphs using knowledge base embeddings. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 4498–4507. Association for Computational Linguistics.
  18. 18.Jiaxin Shi, Shulin Cao, Lei Hou, Juanzi Li, and Hanwang Zhang. 2021. Transfernet: An effective and transparent framework for multi-hop question answering over relation graph. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 4149–4158. Association for Computational Linguistics.
  19. 19.Haitian Sun, Tania Bedrax-Weiss, and William W. Cohen. 2019. Pullnet: Open domain question answering with iterative retrieval on knowledge bases and text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2380–2390. Association for Computational Linguistics.
  20. 20.Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, and William W. Cohen. 2018. Open domain question answering using early fusion of knowledge bases and text. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 4231–4242. Association for Computational Linguistics.
  21. 21.Alon Talmor and Jonathan Berant. 2018. The web as a knowledge-base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 641–651. Association for Computational Linguistics.
  22. 22.Thomas Pellissier Tanon, Denny Vrandecic, Sebastian Schaffert, Thomas Steiner, and Lydia Pintscher. 2016. From freebase to wikidata: The great migration. In Proceedings of the 25th International Conference on World Wide Web, WWW 2016, Montreal, Canada, April 11 - 15, 2016, pages 1419–1428. ACM.
  23. 23.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  24. 24.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio’, and Yoshua Bengio. 2017. Graph attention networks. ArXiv, abs/1710.10903.
  25. 25.Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. KEPLER: A unified model for knowledge embedding and pre-trained language representation. Trans. Assoc. Comput. Linguistics, 9:176–194.
  26. 26.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 602–631. Association for Computational Linguistics.
  27. 27.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: reasoning with language models and knowledge graphs for question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 535–546. Association for Computational Linguistics.
  28. 28.Wen-tau Yih, Ming-Wei Chang, Xiaodong He, and Jianfeng Gao. 2015. Semantic parsing via staged query graph generation: Question answering with knowledge base. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26-31, 2015, Beijing, China, Volume 1: Long Papers, pages 1321–1331. The Association for Computer Linguistics.
  29. 29.Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation? In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 28877–28888.
  30. 30.Mohamad Zamini, Hassan Reza, and Minou Rabiei. 2022. A review of knowledge graph completion. Inf., 13(8):396.
  31. 31.Jing Zhang, Xiaokang Zhang, Jifan Yu, Jian Tang, Jie Tang, Cuiping Li, and Hong Chen. 2022. Subgraph retrieval enhanced model for multi-hop knowledge base question answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 5773–5784. Association for Computational Linguistics.
  32. 32.Yuyu Zhang, Hanjun Dai, Zornitsa Kozareva, Alexander J. Smola, and Le Song. 2018. Variational reasoning for question answering with knowledge graph. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 6069–6076. AAAI Press.
  33. 33.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR.

Citation

MLA
Jiang, J., et al. “ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 3721–35, https://doi.org/10.18653/v1/2023.emnlp-main.228.
APA
Jiang, J., Zhou, K., Zhao, X., Li, Y., & Wen, J.-R. (2023). ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3721–3735. https://doi.org/10.18653/v1/2023.emnlp-main.228
Chicago
Jiang, J., K. Zhou, X. Zhao, Y. Li, and J.-R. Wen. 2023. “ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3721–35. https://doi.org/10.18653/v1/2023.emnlp-main.228.
Harvard
Jiang, J. et al. (2023) “ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3721–3735. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.228.
Vancouver
1. Jiang J, Zhou K, Zhao X, Li Y, Wen J-R (2023) ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3721–3735

BibTeX

@inproceedings{jiang-etal-2023-reasoninglm,
    title = "{R}easoning{LM}: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph",
    author = "Jiang, Jinhao  and
      Zhou, Kun  and
      Zhao, Xin  and
      Li, Yaliang  and
      Wen, Ji-Rong",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.228/",
    doi = "10.18653/v1/2023.emnlp-main.228",
    pages = "3721--3735"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/