Deep Bidirectional Language-Knowledge Graph Pretraining

Michihiro YasunagaAntoine BosselutHongyu RenXikun ZhangChristopher D. ManningPercy LiangJure Leskovec

article2022NeurIPS291 citations

Presents DRAGON, a self-supervised framework that fuses text and knowledge graph subgraphs bidirectionally through joint masked language modeling and link prediction, significantly boosting accuracy on multi-step reasoning and biomedical question answering.

Listen

Modern artificial intelligence models frequently struggle with complex reasoning, subtle linguistic nuance, and structured factual grounding when relying solely on unstructured text. While structured knowledge graphs offer curated factual relationships and clear multi-step reasoning pathways, prior efforts to integrate them into language models have suffered from shallow, one-way information exchange or have limited their integration to task-specific fine-tuning on small datasets. As organizations increasingly depend on language systems for high-stakes decision-making and specialized domain analysis, there is a critical need for foundation models that deeply unify text understanding with structured knowledge representations at scale.

The article evaluates DRAGON (Deep Bidirectional Language-Knowledge Graph Pretraining), a self-supervised framework designed to deeply fuse text segments and knowledge graph subgraphs during pretraining. The primary objective is to demonstrate that bidirectional, large-scale pretraining across text and knowledge graphs establishes stronger, more generalizable reasoning capabilities than conventional text-only models or existing graph-augmented fine-tuning methods.

To achieve this, the approach pairs text excerpts with relevant local subgraphs extracted via entity linking. The architecture uses a cross-modal encoder that exchanges information across text tokens and graph nodes across multiple layers. The model is pretrained using a unified joint objective that combines masked language modeling (predicting hidden words using context and graph connections) and knowledge graph link prediction (predicting missing graph connections using network structure and textual context). The framework was evaluated in both a general commonsense domain (using the BookCorpus text dataset and the ConceptNet graph) and a biomedical domain (using PubMed abstracts and the Unified Medical Language System graph) across twelve downstream benchmark tasks.

The key findings show substantial performance improvements across all evaluation benchmarks. First, the model outperformed standard language models and graph-augmented fine-tuning baselines with an average absolute gain of about 5% across downstream tasks, including a 7% gain on the OpenBookQA benchmark. Second, the model excelled at complex reasoning tasks, achieving up to a 10% gain on questions requiring multi-step deduction, handling negation (showing a 14% improvement), and processing long contexts. Third, in low-resource settings with limited training data, the model achieved an approximate 8% performance boost and maintained a 5% advantage when fine-tuning data was restricted to just 10%. Finally, within the biomedical domain, the model set new state-of-the-art benchmarks on medical question-answering tasks, achieving a 3% gain on MedQA and outperforming top specialized biomedical language models.

These findings indicate that unifying structured knowledge with unstructured text during early-stage pretraining fundamentally improves reasoning robustness and data efficiency. For decision-makers and technical leaders, this approach lowers the risk of logical errors in automated analysis and reduces the financial and operational costs associated with collecting large volumes of task-specific labeled data. Furthermore, increasing the model capacity in this framework yields continuous performance gains, whereas capacity increases in fine-tuning-only models deliver diminishing returns.

Technical leaders and practitioners working on knowledge-intensive applications should consider adopting pretraining architectures that deeply fuse knowledge graphs with text, particularly in domains with dense factual structures such as healthcare, compliance, and scientific research. When deploying language models in data-constrained or highly complex analytical settings, incorporating structured graph data early in the training lifecycle is recommended over relying solely on post-hoc fine-tuning.

The article notes that the current implementation is strictly an encoder model designed for classification and question-answering tasks, meaning it does not currently support open-ended text generation. Stakeholders can have high confidence in the reported classification and reasoning improvements given the extensive ablation studies and multi-domain benchmarks, but should exercise caution if attempting to apply the architecture to generative natural language workflows without further architectural development.

Cover for Deep Bidirectional Language-Knowledge Graph Pretraining

Abstract

Pretraining a language model (LM) on text has been shown to help various downstream NLP tasks. Recent works show that a knowledge graph (KG) can complement text data, offering structured background knowledge that provides a useful scaffold for reasoning. However, these works are not pretrained to learn a deep fusion of the two modalities at scale, limiting the potential to acquire fully joint representations of text and KG. Here we propose DRAGON (Deep Bidirectional Language-Knowledge Graph Pretraining), a self-supervised method to pretrain a deeply joint language-knowledge foundation model from text and KG at scale. Specifically, our model takes pairs of text segments and relevant KG subgraphs as input and bidirectionally fuses information from both modalities. We pretrain this model by unifying two self-supervised reasoning tasks, masked language modeling and KG link prediction. DRAGON outperforms existing LM and LM+KG models on diverse downstream tasks including question answering across general and biomedical domains, with +5% absolute gain on average. In particular, DRAGON achieves strong performance on complex reasoning about language and knowledge (+10% on questions involving long contexts or multi-step reasoning) and low-resource QA (+8% on OBQA and RiddleSense), and new state-of-the-art results on various BioNLP tasks. Our code and trained models are available at https://github.com/michiyasunaga/dragon.

Table of Contents

  • 1 Introduction
  • 1.1 Related work
  • 2 Deep Bidirectional Language-Knowledge Graph Pretraining (DRAGON)
  • 2.1 Input representation
  • 2.2 Cross-modal encoder
  • 2.3 Pretraining objective
  • 2.4 Finetuning
  • 3 Experiments: General domain
  • 3.1 Pretraining setup
  • 3.2 Downstream evaluation tasks
  • 3.3 Baselines
  • 3.4 Results
  • 3.4.1 Analysis: Effect of knowledge graph
  • 3.4.2 Analysis: Effect of pretraining
  • 3.4.3 Analysis: Design choices of DRAGON
  • 4 Experiments: Biomedical domain
  • 5 Conclusion
  • Reproducibility
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — DRAGON Cross-Modal Architecture

    model/method

    DRAGON (Deep Bidirectional Language-Knowledge Graph Pretraining) employs a sequence-graph cross-modal encoder that fuses a text sequence and a knowledge graph (KG) subgraph across multiple bidirectional interaction layers.

    Given an input text token sequence (wint,w1,…,wI)(w_{\text{int}}, w_1, \dots, w_I) prepended with a special modality interaction token wintw_{\text{int}} and a set of local KG nodes (vint,v1,…,vJ)(v_{\text{int}}, v_1, \dots, v_J) including a special interaction node vintv_{\text{int}}, the model first computes initial representations using NN Transformer language model (LM) layers and node embedding lookups: (Hint(0),H1(0),…,HI(0))=LM-Layers(wint,w1,…,wI)(H_{\text{int}}^{(0)}, H_1^{(0)}, \dots, H_I^{(0)}) = \text{LM-Layers}(w_{\text{int}}, w_1, \dots, w_I) (Vint(0),V1(0),…,VJ(0))=Node-Embedding(vint,v1,…,vJ)(V_{\text{int}}^{(0)}, V_1^{(0)}, \dots, V_J^{(0)}) = \text{Node-Embedding}(v_{\text{int}}, v_1, \dots, v_J)

    The initial token representations H(0)H^{(0)} and node representations V(0)V^{(0)} are then jointly encoded through MM text-KG fusion layers. At each fusion layer ℓ∈{1,…,M}\ell \in \{1, \dots, M\}:

    1. Text representations are updated by a Transformer LM layer: (H~int(ℓ),H1(ℓ),…,HI(ℓ))=LM-Layer(Hint(ℓ−1),H1(ℓ−1),…,HI(ℓ−1))(\tilde{H}_{\text{int}}^{(\ell)}, H_1^{(\ell)}, \dots, H_I^{(\ell)}) = \text{LM-Layer}(H_{\text{int}}^{(\ell-1)}, H_1^{(\ell-1)}, \dots, H_I^{(\ell-1)})
    2. KG node representations are updated by a Graph Neural Network (GNN) layer: (V~int(ℓ),V1(ℓ),…,VJ(ℓ))=GNN-Layer(Vint(ℓ−1),V1(ℓ−1),…,VJ(ℓ−1))(\tilde{V}_{\text{int}}^{(\ell)}, V_1^{(\ell)}, \dots, V_J^{(\ell)}) = \text{GNN-Layer}(V_{\text{int}}^{(\ell-1)}, V_1^{(\ell-1)}, \dots, V_J^{(\ell-1)})
    3. Modality interaction is performed between the interaction token and interaction node representations via a multilayer perceptron MInt\text{MInt}: [Hint(ℓ);Vint(ℓ)]=MInt([H~int(ℓ);V~int(ℓ)])[H_{\text{int}}^{(\ell)}; V_{\text{int}}^{(\ell)}] = \text{MInt}([\tilde{H}_{\text{int}}^{(\ell)}; \tilde{V}_{\text{int}}^{(\ell)}]) where [⋅;⋅][\cdot ; \cdot] denotes concatenation.

    The final layer yields contextualized token vectors (Hint,H1,…,HI)(H_{\text{int}}, H_1, \dots, H_I) and contextualized node vectors (Vint,V1,…,VJ)(V_{\text{int}}, V_1, \dots, V_J) where information flows bidirectionally between text and KG.

  2. Knowl 2 — Unified Pretraining Objective for Joint Text-KG Reasoning

    model/method

    DRAGON is pretrained via a joint self-supervised objective L=LMLM+LLinkPred\mathcal{L} = \mathcal{L}_{\text{MLM}} + \mathcal{L}_{\text{LinkPred}} combining Masked Language Modeling (MLM) on text with Link Prediction (LinkPred) on the knowledge graph:

    1. Masked Language Modeling (LMLM\mathcal{L}_{\text{MLM}}): A subset of tokens M⊂W\mathcal{M} \subset W in the input text segment W=(w1,…,wI)W = (w_1, \dots, w_I) is replaced with the token [MASK]\text{[MASK]}. A linear prediction head over the contextualized token representations HiH_i optimizes the cross-entropy loss: LMLM=−∑i∈Mlog⁡p(wi∣Hi)\mathcal{L}_{\text{MLM}} = - \sum_{i \in \mathcal{M}} \log p(w_i \mid H_i)

    2. KG Link Prediction (LLinkPred\mathcal{L}_{\text{LinkPred}}): A subset of edge triplets S={(h,r,t)}⊆ES = \{(h, r, t)\} \subseteq \mathcal{E} is held out from the input KG G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E}). For each positive triplet (h,r,t)∈S(h, r, t) \in S, nn negative corruptions (h′,r,t′)(h', r, t') are sampled. The objective uses a margin γ\gamma and the sigmoid function σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}: LLinkPred=∑(h,r,t)∈S(−log⁡σ(ϕr(h,t)+γ)+1n∑(h′,r,t′)log⁡σ(ϕr(h′,t′)+γ))\mathcal{L}_{\text{LinkPred}} = \sum_{(h,r,t) \in S} \left( -\log \sigma(\phi_r(h, t) + \gamma) + \frac{1}{n} \sum_{(h',r,t')} \log \sigma(\phi_r(h', t') + \gamma) \right) where the scoring function ϕr(h,t)\phi_r(h, t) evaluates the plausibility of the triplet using contextualized node embeddings h=Vh,t=Vt∈Rdh = V_h, t = V_t \in \mathbb{R}^d and a learnable relation embedding r=Rr∈Rdr = R_r \in \mathbb{R}^d. Scoring functions include:

    • DistMult: ϕr(h,t)=⟨h,r,t⟩=∑khkrktk\phi_r(h, t) = \langle h, r, t \rangle = \sum_k h_k r_k t_k
    • TransE: ϕr(h,t)=−∥h+r−t∥\phi_r(h, t) = - \| h + r - t \|
    • RotatE: ϕr(h,t)=−∥h∘r−t∥\phi_r(h, t) = - \| h \circ r - t \|

    MLM encourages the model to use structured knowledge in the KG to resolve masked text tokens, while LinkPred encourages the model to use textual context to predict missing KG links.

  3. Knowl 3 — Aligned Text-KG Input Construction and Retrieval

    model/method

    To construct aligned multimodal input instances X=(W,G)X = (W, \mathcal{G}) from a text corpus W\mathcal{W} and a knowledge graph G=(Vall,Eall)G = (\mathcal{V}_{\text{all}}, \mathcal{E}_{\text{all}}), DRAGON performs subgraph retrieval and graph-text grounding:

    1. Text Sampling: A text segment W=(w1,…,wI)W = (w_1, \dots, w_I) is sampled from the corpus (up to 512 tokens).
    2. Entity Linking and Subgraph Extraction: Entity mentions in WW are identified and linked to entity nodes in the KG to form an initial seed node set Vel⊆Vall\mathcal{V}_{\text{el}} \subseteq \mathcal{V}_{\text{all}}. Then, 2-hop bridge nodes connecting nodes in Vel\mathcal{V}_{\text{el}} are retrieved from GG to form the node set V⊆Vall\mathcal{V} \subseteq \mathcal{V}_{\text{all}} (restricted to a maximum size of 200 nodes). All KG edges spanning V\mathcal{V} are retained to form edge set E⊆Eall\mathcal{E} \subseteq \mathcal{E}_{\text{all}}, yielding the local KG G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E}).
    3. Modality Interaction Anchors: A special interaction token wintw_{\text{int}} is prepended to the text sequence: (wint,w1,…,wI)(w_{\text{int}}, w_1, \dots, w_I). A special interaction node vintv_{\text{int}} is added to V\mathcal{V} and connected to every entity-linked seed node in Vel\mathcal{V}_{\text{el}} via a dedicated relation type relr_{\text{el}}. The interaction token and node serve as global pooling hubs and cross-modal interaction interfaces during encoding.
  4. Knowl 4 — Task Adaptation and Downstream Finetuning in DRAGON

    model/method

    To finetune DRAGON on downstream tasks (such as multiple-choice question answering or text classification), the input text (e.g., question concatenated with a candidate answer) and a retrieved local knowledge graph subgraph G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E}) are encoded by the pretrained cross-modal encoder.

    The encoder outputs contextualized token representations (Hint,H1,…,HI)(H_{\text{int}}, H_1, \dots, H_I) and node representations (Vint,V1,…,VJ)(V_{\text{int}}, V_1, \dots, V_J), where HintH_{\text{int}} and VintV_{\text{int}} are the modality interaction token and node vectors, respectively. A global pooled representation XX of the multimodal input is computed as: X=MLP(Hint,Vint,Gˉ)X = \text{MLP}(H_{\text{int}}, V_{\text{int}}, \bar{G}) where Gˉ=∑j=1JαjVj\bar{G} = \sum_{j=1}^J \alpha_j V_j is an attention-pooled representation of all KG entity nodes {V1,…,VJ}\{V_1, \dots, V_J\} using the interaction token vector HintH_{\text{int}} as the attention query. The vector XX is fed into a classification head to output task-specific predictions.

  5. Knowl 5 — General Domain Commonsense Reasoning Performance

    data/table

    In the general domain, DRAGON was pretrained using BookCorpus text and the ConceptNet knowledge graph (initialized from RoBERTa-Large with 19 LM layers and 5 text-KG fusion layers, totaling 360M parameters). It was evaluated across nine commonsense reasoning and natural language inference benchmarks: CommonsenseQA (CSQA), OpenBookQA (OBQA), RiddleSense (Riddle), AI2 Reasoning Challenge Challenge Set (ARC), CosmosQA, HellaSwag, PIQA, SIQA, and Abductive NLI (aNLI).

    Baselines compared:

    • RoBERTa: RoBERTa-Large with continued MLM pretraining on BookCorpus for the same number of steps without KG.
    • QA-GNN: RoBERTa-Large augmented with a GNN finetuned on the downstream task without joint pretraining.
    • GreaseLM: RoBERTa-Large with 5 text-KG fusion layers finetuned on the downstream task without joint pretraining.
    Method CSQA OBQA Riddle ARC CosmosQA HellaSwag PIQA SIQA aNLI
    RoBERTa 68.7 64.9 60.7 43.0 80.5 82.3 79.4 75.9 82.7
    QA-GNN 73.4 67.8 67.0 44.4 80.7 82.6 79.6 75.7 83.0
    GreaseLM 74.2 66.9 67.2 44.7 80.6 82.8 79.6 75.5 83.3
    DRAGON (Ours) 76.0 72.0 71.3 48.6 82.3 85.2 81.1 76.8 84.0

    DRAGON outperforms RoBERTa by +8% accuracy on average and achieves substantial gains over GreaseLM on low-resource datasets (e.g., +5.1% on OBQA, +4.1% on Riddle, +3.9% on ARC) and complex reasoning tasks (e.g., +1.7% on CosmosQA, +2.4% on HellaSwag).

  6. Knowl 6 — Biomedical Domain NLP and Reasoning Performance

    data/table

    In the biomedical domain, DRAGON was pretrained on PubMed abstracts (21GB text) and the Unified Medical Language System (UMLS) KG (300K nodes, 1M edges), initializing the LM component with BioLinkBERT-Large. DRAGON was evaluated on three biomedical question answering benchmarks: MedQA-USMLE (MedQA), PubMedQA, and BioASQ (accuracy metric).

    Method MedQA PubMedQA BioASQ
    BioBERT 36.7 60.2 84.1
    PubmedBERT 38.1 55.8 87.5
    BioLinkBERT 44.6 72.2 94.8
    + QAGNN 45.0 72.1 95.0
    + GreaseLM 45.1 72.4 94.9
    DRAGON (Ours) 47.5 73.4 96.4

    DRAGON achieved a +2.9% accuracy gain over BioLinkBERT and +2.4% over GreaseLM on MedQA, as well as +1.2% on PubMedQA and +1.6% on BioASQ over BioLinkBERT, establishing a new state of the art on all three biomedical benchmarks.

  7. Knowl 7 — Performance on Complex Reasoning Categories

    data/table

    To evaluate model robustness under complex reasoning constraints, accuracy was measured on subsets of the combined CSQA and OBQA development sets categorized by five linguistic and structural proxies: negation terms, conjunction terms, hedge terms, number of prepositional phrases (0, 1, 2, 3), and number of entity mentions (>10>10).

    Method Negation Conjunction Hedge # Prepositional Phrases # Entities
    0 1 2 3 >>10
    RoBERTa 61.7 70.9 68.6 67.6 71.0 71.1 73.1 74.5
    QA-GNN 65.1 74.5 74.2 72.1 71.6 75.6 71.3 78.6
    GreaseLM 65.1 74.9 76.6 75.6 73.8 74.7 73.6 79.4
    DRAGON (Ours) 75.2 79.6 77.5 79.1 78.2 77.8 80.9 83.5

    DRAGON outperformed RoBERTa by +13.5% and GreaseLM by +10.1% on negation questions, while achieving consistent gains on multi-constraint questions (+7.3%+7.3\% on 3 prepositional phrases and +4.1%+4.1\% on >10>10 entities over GreaseLM).

  8. Knowl 8 — Ablation Analysis of DRAGON Pretraining and Architecture Design Choices

    data/table

    An ablation study on CSQA and OBQA evaluated the contributions of DRAGON's pretraining objectives, link prediction scoring functions, modality fusion mechanisms, and graph representation choices:

    Ablation Type Ablation CSQA OBQA
    Pretraining objective MLM + LinkPred (final) 76.0 72.0
    MLM only 74.3 67.2
    LinkPred only 73.8 66.4
    LinkPred head DistMult (final) 76.0 72.0
    TransE 75.7 71.4
    RotatE 75.8 71.7
    Cross-modal model Bidirectional interaction (final) 76.0 72.0
    Concatenate at end 74.5 68.0
    KG structure Use graph (final) 76.0 72.0
    Convert to sentence 74.7 70.1

    Key observations:

    • Joint Objective: Joint pretraining with MLM + LinkPred outperformed MLM only (+4.8% on OBQA) and LinkPred only (+5.6% on OBQA).
    • Scoring Function: DistMult, TransE, and RotatE all performed well and improved upon the MLM-only baseline.
    • Modality Interaction: Bidirectional interaction across intermediate layers substantially outperformed shallow concatenation at the final layer (+4.0% on OBQA).
    • Graph Structure: Preserving graph topology via GNN layers outperformed linearizing KG triplets into natural language sentences (+1.9% on OBQA).
  9. Knowl 9 — Pretraining Benefits for Low-Resource Scenarios and Model Scaling

    empirical result

    DRAGON exhibits marked advantages in data efficiency and architectural scaling behavior compared to models that only finetune on knowledge graphs:

    1. Low-Resource Finetuning: When evaluated in a low-resource setting using only 10% of downstream training data, DRAGON achieved 77.9% accuracy on CosmosQA (vs. 72.2% for RoBERTa and 73.0% for GreaseLM) and 72.3% accuracy on PIQA (vs. 66.4% for RoBERTa and 67.0% for GreaseLM), demonstrating improved sample efficiency.
    2. Scaling Fusion Capacity: Increasing the cross-modal capacity by adding text-KG fusion layers from 5 to 7 (denoted "-Ex") improved DRAGON's accuracy on CSQA (76.0% →\to 76.3%) and OBQA (72.0% →\to 72.8%). In contrast, for the finetuning-only baseline GreaseLM, increasing the number of fusion layers decreased accuracy on CSQA (74.2% →\to 73.9%) and OBQA (66.9% →\to 66.2%).
  10. Knowl 10 — Encoder-Only Architecture Limitation

    limitation

    DRAGON is designed strictly as a bidirectional sequence-graph encoder (analogous to BERT and RoBERTa) and is trained for discriminative tasks and classification/multiple-choice question answering. Consequently, the model does not possess generative capabilities for free-form language generation.

Coverage note — No substantial contributed material was omitted; qualitative attention weight visualizations from the case studies were summarized within the complex reasoning and architecture analyses.

References

  1. 1.Rishi Bommasani et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  2. 2.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics (NAACL), 2019.
  3. 3.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  4. 4.Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In North American Chapter of the Association for Computational Linguistics (NAACL), 2018.
  5. 5.Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. Freebase: a collaboratively created graph database for structuring human knowledge. In SIGMOD, 2008.
  6. 6.Denny Vrandeˇci´c and Markus Krötzsch. Wikidata: A free collaborative knowledgebase. Communications of the ACM, 2014.
  7. 7.Robyn Speer, Joshua Chin, and Catherine Havasi. Conceptnet 5.5: An open multilingual graph of general knowledge. In Proceedings of the AAAI Conference on Artificial Intelligence, 2017.
  8. 8.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. QA-GNN: Reasoning with language models and knowledge graphs for question answering. In North American Chapter of the Association for Computational Linguistics (NAACL), 2021.
  9. 9.Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D Manning, and Jure Leskovec. Greaselm: Graph reasoning enhanced language models for question answering. In International Conference on Learning Representations (ICLR), 2022.
  10. 10.Hongyu Ren, Weihua Hu, and Jure Leskovec. Query2box: Reasoning over knowledge graphs in vector space using box embeddings. In International Conference on Learning Representations (ICLR), 2020.
  11. 11.Hongyu Ren, Hanjun Dai, Bo Dai, Xinyun Chen, Michihiro Yasunaga, Haitian Sun, Dale Schuurmans, Jure Leskovec, and Denny Zhou. Lego: Latent execution-guided reasoning for multi-hop question answering on knowledge graphs. In International Conference on Machine Learning (ICML), 2021.
  12. 12.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. Ernie: Enhanced language representation with informative entities. In Association for Computational Linguistics (ACL), 2019.
  13. 13.Wenhan Xiong, Jingfei Du, William Yang Wang, and Veselin Stoyanov. Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model. In International Conference on Learning Representations (ICLR), 2020.
  14. 14.Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. Kepler: A unified model for knowledge embedding and pre-trained language representation. Transactions of the Association for Computational Linguistics (TACL), 2021.
  15. 15.Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training. In North American Chapter of the Association for Computational Linguistics (NAACL), 2021.
  16. 16.Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137, 2021.
  17. 17.Olivier Bodenreider. The unified medical language system (UMLS): Integrating biomedical terminology. Nucleic acids research, 2004.
  18. 18.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  19. 19.Michihiro Yasunaga, Jure Leskovec, and Percy Liang. LinkBERT: Pretraining language models with document links. In Association for Computational Linguistics (ACL), 2022.
  20. 20.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. Realm: Retrieval-augmented language model pre-training. In International Conference on Machine Learning (ICML), 2020.
  21. 21.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  22. 22.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. arXiv preprint arXiv:2112.04426, 2021.
  23. 23.Matthew E. Peters, Mark Neumann, IV RobertLLogan, Roy Schwartz, V. Joshi, Sameer Singh, and Noah A. Smith. Knowledge enhanced contextual word representations. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  24. 24.Corby Rosset, Chenyan Xiong, Minh Phan, Xia Song, Paul Bennett, and Saurabh Tiwary. Knowledge-aware language model pretraining. arXiv preprint arXiv:2007.00655, 2020.
  25. 25.Tao Shen, Yi Mao, Pengcheng He, Guodong Long, Adam Trischler, and Weizhu Chen. Exploiting structured knowledge in text via graph-guided representation learning. In Empirical Methods in Natural Language Processing (EMNLP), 2020.
  26. 26.Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Collier. Self-alignment pretraining for biomedical entity representations. In North American Chapter of the Association for Computational Linguistics (NAACL), 2021.
  27. 27.Donghan Yu, Chenguang Zhu, Yiming Yang, and Michael Zeng. Jaket: Joint pre-training of knowledge graph and language understanding. In AAAI Conference on Artificial Intelligence, 2022.
  28. 28.Pei Ke, Haozhe Ji, Yu Ran, Xin Cui, Liwei Wang, Linfeng Song, Xiaoyan Zhu, and Minlie Huang. Jointgt: Graph-text joint representation learning for text generation from knowledge graphs. In Findings of ACL, 2021.
  29. 29.Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and P. Wang. K-bert: Enabling language representation with knowledge graph. In AAAI Conference on Artificial Intelligence, 2020.
  30. 30.Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuan-Jing Huang, and Zheng Zhang. Colake: Contextualized language and knowledge embedding. In International Conference on Computational Linguistics (COLING), 2020.
  31. 31.Bin He, Di Zhou, Jinghui Xiao, Xin Jiang, Qun Liu, Nicholas Jing Yuan, and Tong Xu. Integrating graph contextualized knowledge into pre-trained language models. In Findings of EMNLP, 2020.
  32. 32.Bill Yuchen Lin, Xinyue Chen, Jamin Chen, and Xiang Ren. Kagnet: Knowledge-aware graph networks for commonsense reasoning. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  33. 33.Yanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang, Jun Yan, and Xiang Ren. Scalable multi-hop relational reasoning for knowledge-aware question answering. In Empirical Methods in Natural Language Processing (EMNLP), 2020.
  34. 34.Shangwen Lv, Daya Guo, Jingjing Xu, Duyu Tang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, and Songlin Hu. Graph-based reasoning over heterogeneous external knowledge for commonsense question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  35. 35.Kuan Wang, Yuyu Zhang, Diyi Yang, Le Song, and Tao Qin. Gnn is a counter? revisiting gnn for question answering. In International Conference on Learning Representations (ICLR), 2022.
  36. 36.Todor Mihaylov and Anette Frank. Knowledgeable reader: Enhancing cloze-style reading comprehension with external commonsense knowledge. In Association for Computational Linguistics (ACL), 2018.
  37. 37.An Yang, Quan Wang, Jing Liu, Kai Liu, Yajuan Lyu, Hua Wu, Qiaoqiao She, and Sujian Li. Enhancing pre-trained language representations with rich knowledge for machine reading comprehension. In Association for Computational Linguistics (ACL), 2019.
  38. 38.Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, and William W Cohen. Open domain question answering using early fusion of knowledge bases and text. In Empirical Methods in Natural Language Processing (EMNLP), 2018.
  39. 39.Haitian Sun, Tania Bedrax-Weiss, and William W Cohen. Pullnet: Open domain question answering with iterative retrieval on knowledge bases and text. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  40. 40.Jun Yan, Mrigank Raman, Aaron Chan, Tianyu Zhang, Ryan Rossi, Handong Zhao, Sungchul Kim, Nedim Lipka, and Xiang Ren. Learning contextualized knowledge structures for commonsense reasoning. In Findings of ACL, 2021.
  41. 41.Yueqing Sun, Qi Shi, Le Qi, and Yu Zhang. Jointlk: Joint reasoning with language models and knowledge graphs for commonsense question answering. In North American Chapter of the Association for Computational Linguistics (NAACL), 2022.
  42. 42.Yichong Xu, Chenguang Zhu, Shuohang Wang, Siqi Sun, Hao Cheng, Xiaodong Liu, Jianfeng Gao, Pengcheng He, Michael Zeng, and Xuedong Huang. Human parity on commonsenseqa: Augmenting self-attention with external attention. In Association for Computational Linguistics (ACL), 2022.
  43. 43.Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. Complex embeddings for simple link prediction. In International conference on machine learning (ICML), 2016.
  44. 44.Seyed Mehran Kazemi and David Poole. Simple embedding for link prediction in knowledge graphs. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  45. 45.Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. Translating embeddings for modeling multi-relational data. In Advances in Neural Information Processing Systems (NeurIPS), 2013.
  46. 46.Bishan Yang, Wen-tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. Embedding entities and relations for learning and inference in knowledge bases. In International Conference on Learning Representations (ICLR), 2015.
  47. 47.Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. Rotate: Knowledge graph embedding by relational rotation in complex space. In International Conference on Learning Representations (ICLR), 2019.
  48. 48.Sebastian Riedel, Limin Yao, Andrew McCallum, and Benjamin M Marlin. Relation extraction with matrix factorization and universal schemas. In North American Chapter of the Association for Computational Linguistics (NAACL), 2013.
  49. 49.Kristina Toutanova, Danqi Chen, Patrick Pantel, Hoifung Poon, Pallavi Choudhury, and Michael Gamon. Representing text for joint embedding of text and knowledge bases. In Empirical Methods in Natural Language Processing (EMNLP), 2015.
  50. 50.Ruobing Xie, Zhiyuan Liu, Jia Jia, Huanbo Luan, and Maosong Sun. Representation learning of knowledge graphs with entity descriptions. In Proceedings of the AAAI Conference on Artificial Intelligence, 2016.
  51. 51.Liang Yao, Chengsheng Mao, and Yuan Luo. Kg-bert: Bert for knowledge graph completion. arXiv preprint arXiv:1909.03193, 2019.
  52. 52.Bosung Kim, Taesuk Hong, Youngjoong Ko, and Jungyun Seo. Multi-task learning for knowledge graph completion with pre-trained language models. In International Conference on Computational Linguistics (COLING), 2020.
  53. 53.Da Li, Sen Yang, Kele Xu, Ming Yi, Yukai He, and Huaimin Wang. Multi-task pre-training language model for semantic network completion. arXiv preprint arXiv:2201.04843, 2022.
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems (NeurIPS), 2017.
  55. 55.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In International Conference on Computer Vision (ICCV), 2015.
  56. 56.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In North American Chapter of the Association for Computational Linguistics (NAACL), 2019.
  57. 57.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Empirical Methods in Natural Language Processing (EMNLP), 2018.
  58. 58.Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee, and Xiang Ren. Riddlesense: Reasoning about riddle questions featuring linguistic creativity and commonsense knowledge. In Findings of ACL, 2021.
  59. 59.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  60. 60.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Cosmos qa: Machine reading comprehension with contextual commonsense reasoning. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  61. 61.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Association for Computational Linguistics (ACL), 2019.
  62. 62.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In AAAI Conference on Artificial Intelligence, 2020.
  63. 63.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  64. 64.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. Abductive commonsense reasoning. In International Conference on Learning Representations (ICLR), 2020.
  65. 65.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  66. 66.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  67. 67.Elliot G Brown, Louise Wood, and Sue Wood. The medical dictionary for regulatory activities (meddra). Drug safety, 20(2):109–117, 1999.
  68. 68.Carolyn E Lipscomb. Medical subject headings (mesh). Bulletin of the Medical Library Association, 88(3):265, 2000.
  69. 69.Marinka Zitnik, Monica Agrawal, and Jure Leskovec. Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics, 34(13):i457–i466, 2018.
  70. 70.Michael Ashburner, Catherine A Ball, Judith A Blake, David Botstein, Heather Butler, J Michael Cherry, Allan P Davis, Kara Dolinski, Selina S Dwight, Janan T Eppig, et al. Gene ontology: tool for the unification of biology. Nature genetics, 25(1):25–29, 2000.
  71. 71.David S Wishart, Yannick D Feunang, An C Guo, Elvis J Lo, Ana Marcu, Jason R Grant, Tanvir Sajed, Daniel Johnson, Carin Li, Zinat Sayeeda, et al. Drugbank 5.0: a major update to the drugbank database for 2018. Nucleic acids research, 2018.
  72. 72.Camilo Ruiz, Marinka Zitnik, and Jure Leskovec. Identification of disease treatment mechanisms through the multiscale interactome. Nature communications, 12(1):1–15, 2021.
  73. 73.PubMed. https://pubmed.ncbi.nlm.nih.gov/.
  74. 74.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 2020.
  75. 75.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pretraining for biomedical natural language processing. arXiv preprint arXiv:2007.15779, 2020.
  76. 76.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 2021.
  77. 77.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  78. 78.Anastasios Nentidis, Konstantinos Bougiatiotis, Anastasia Krithara, and Georgios Paliouras. Results of the seventh edition of the bioasq challenge. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2019.
  79. 79.Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. A survey of knowledge-enhanced text generation. ACM Computing Surveys (CSUR), 2022.
  80. 80.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. Towards controllable biases in language generation. In the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)-Findings, long, 2020.
  81. 81.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359, 2021.
  82. 82.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Findings of EMNLP, 2020.
  83. 83.Ninareh Mehrabi, Pei Zhou, Fred Morstatter, Jay Pujara, Xiang Ren, and A. G. Galstyan. Lawyers are dishonest? quantifying representational harms in commonsense knowledge resources. ArXiv, abs/2103.11320, 2021.

Citation

MLA
Yasunaga, M., et al. “Deep Bidirectional Language-Knowledge Graph Pretraining”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 37309–23, https://proceedings.neurips.cc/paper_files/paper/2022/file/f224f056694bcfe465c5d84579785761-Paper-Conference.pdf.
APA
Yasunaga, M., Bosselut, A., Ren, H., Zhang, X., Manning, C. D., Liang, P. S., & Leskovec, J. (2022). Deep Bidirectional Language-Knowledge Graph Pretraining. Advances in Neural Information Processing Systems, 35, 37309–37323. https://proceedings.neurips.cc/paper_files/paper/2022/file/f224f056694bcfe465c5d84579785761-Paper-Conference.pdf
Chicago
Yasunaga, M., A. Bosselut, H. Ren, et al. 2022. “Deep Bidirectional Language-Knowledge Graph Pretraining”. Advances in Neural Information Processing Systems 35: 37309–23. https://proceedings.neurips.cc/paper_files/paper/2022/file/f224f056694bcfe465c5d84579785761-Paper-Conference.pdf.
Harvard
Yasunaga, M. et al. (2022) “Deep Bidirectional Language-Knowledge Graph Pretraining”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37309–37323. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/f224f056694bcfe465c5d84579785761-Paper-Conference.pdf.
Vancouver
1. Yasunaga M, Bosselut A, Ren H, Zhang X, Manning CD, Liang PS, Leskovec J (2022) Deep Bidirectional Language-Knowledge Graph Pretraining. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37309–37323

BibTeX

@inproceedings{yasunaga2022deep,
  title = {Deep Bidirectional Language-Knowledge Graph Pretraining},
  author = {Yasunaga, Michihiro and Bosselut, Antoine and Ren, Hongyu and Zhang, Xikun and Manning, Christopher D and Liang, Percy S. and Leskovec, Jure},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {37309-37323},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/f224f056694bcfe465c5d84579785761-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors