Deep Bidirectional Language-Knowledge Graph Pretraining
Michihiro YasunagaAntoine BosselutHongyu RenXikun ZhangChristopher D. ManningPercy LiangJure Leskovec
Presents DRAGON, a self-supervised framework that fuses text and knowledge graph subgraphs bidirectionally through joint masked language modeling and link prediction, significantly boosting accuracy on multi-step reasoning and biomedical question answering.
Modern artificial intelligence models frequently struggle with complex reasoning, subtle linguistic nuance, and structured factual grounding when relying solely on unstructured text. While structured knowledge graphs offer curated factual relationships and clear multi-step reasoning pathways, prior efforts to integrate them into language models have suffered from shallow, one-way information exchange or have limited their integration to task-specific fine-tuning on small datasets. As organizations increasingly depend on language systems for high-stakes decision-making and specialized domain analysis, there is a critical need for foundation models that deeply unify text understanding with structured knowledge representations at scale.
The article evaluates DRAGON (Deep Bidirectional Language-Knowledge Graph Pretraining), a self-supervised framework designed to deeply fuse text segments and knowledge graph subgraphs during pretraining. The primary objective is to demonstrate that bidirectional, large-scale pretraining across text and knowledge graphs establishes stronger, more generalizable reasoning capabilities than conventional text-only models or existing graph-augmented fine-tuning methods.
To achieve this, the approach pairs text excerpts with relevant local subgraphs extracted via entity linking. The architecture uses a cross-modal encoder that exchanges information across text tokens and graph nodes across multiple layers. The model is pretrained using a unified joint objective that combines masked language modeling (predicting hidden words using context and graph connections) and knowledge graph link prediction (predicting missing graph connections using network structure and textual context). The framework was evaluated in both a general commonsense domain (using the BookCorpus text dataset and the ConceptNet graph) and a biomedical domain (using PubMed abstracts and the Unified Medical Language System graph) across twelve downstream benchmark tasks.
The key findings show substantial performance improvements across all evaluation benchmarks. First, the model outperformed standard language models and graph-augmented fine-tuning baselines with an average absolute gain of about 5% across downstream tasks, including a 7% gain on the OpenBookQA benchmark. Second, the model excelled at complex reasoning tasks, achieving up to a 10% gain on questions requiring multi-step deduction, handling negation (showing a 14% improvement), and processing long contexts. Third, in low-resource settings with limited training data, the model achieved an approximate 8% performance boost and maintained a 5% advantage when fine-tuning data was restricted to just 10%. Finally, within the biomedical domain, the model set new state-of-the-art benchmarks on medical question-answering tasks, achieving a 3% gain on MedQA and outperforming top specialized biomedical language models.
These findings indicate that unifying structured knowledge with unstructured text during early-stage pretraining fundamentally improves reasoning robustness and data efficiency. For decision-makers and technical leaders, this approach lowers the risk of logical errors in automated analysis and reduces the financial and operational costs associated with collecting large volumes of task-specific labeled data. Furthermore, increasing the model capacity in this framework yields continuous performance gains, whereas capacity increases in fine-tuning-only models deliver diminishing returns.
Technical leaders and practitioners working on knowledge-intensive applications should consider adopting pretraining architectures that deeply fuse knowledge graphs with text, particularly in domains with dense factual structures such as healthcare, compliance, and scientific research. When deploying language models in data-constrained or highly complex analytical settings, incorporating structured graph data early in the training lifecycle is recommended over relying solely on post-hoc fine-tuning.
The article notes that the current implementation is strictly an encoder model designed for classification and question-answering tasks, meaning it does not currently support open-ended text generation. Stakeholders can have high confidence in the reported classification and reasoning improvements given the extensive ablation studies and multi-domain benchmarks, but should exercise caution if attempting to apply the architecture to generative natural language workflows without further architectural development.
- Paper: ERNIE: Enhanced Language Representation with Informative Entities, Zhengyan Zhang et al. (2019). ERNIE establishes the fundamental paradigm of injecting knowledge graph entity representations directly into Transformer language pretraining, serving as a primary conceptual predecessor to DRAGON's bidirectional fusion framework.
- Paper: Modeling Relational Data with Graph Convolutional Networks, Michael Schlichtkrull et al. (2018). This paper introduces Relational Graph Convolutional Networks for link prediction and entity modeling, providing the foundational graph-relational encoding and objective formulations adapted in DRAGON's knowledge graph pretraining.
- Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). This work formulates core self-supervised pretraining strategies for graph neural networks, which directly underpin DRAGON's joint graph link prediction and node masking objectives.
- Paper: LXMERT: Learning Cross-Modality Encoder Representations from Transformers, Hao Tan et al. (2019). LXMERT introduces multi-layer cross-modality encoder representations for co-attending across distinct domains, providing the architectural foundation for DRAGON's bidirectional text-and-graph information exchange layers.
- Paper: A Survey on Knowledge Graphs: Representation, Acquisition, and Applications, Shaoxiong Ji et al. (2020). This survey provides an essential taxonomy of knowledge graph representations, embeddings, and reasoning frameworks necessary to understand the structured graph components utilized by DRAGON.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Retrieval-Augmented Generation defines the standard knowledge-augmented baseline that DRAGON improves upon by moving structured knowledge integration from post-hoc retrieval to early pretraining.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This roadmap synthesizes modern paradigms for unifying language models and knowledge graphs, contextualizing DRAGON within the broader progression of synergized and bidirectional reasoning systems.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). StructGPT builds upon graph-text integration principles by introducing iterative reading-and-reasoning interfaces that allow language models to reason over external structured knowledge without retraining.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Plan-on-Graph extends static text-graph reasoning frameworks into dynamic, self-correcting graph navigation and multi-hop planning powered by language models.
- Paper: Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting, Xinyan Guan et al. (2024). This work applies knowledge graph subgraphs downstream to iteratively detect and retrofit factual hallucinations in model-generated intermediate reasoning steps.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). GraphRAG extends structured graph-language integration by constructing hierarchical community knowledge graphs from text to perform global query-focused summarization.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). This comprehensive survey categorizes scientific large language models, surveying how cross-modal and graph-augmented architectures like DRAGON are applied to specialized scientific and biomedical discovery.
