Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora
Xisen JinDejiao ZhangHenghui ZhuWei XiaoShang-Wen LiXiaokai WeiAndrew O. ArnoldXiang Ren
Establishes a continual pretraining framework for language models across evolving domains and temporal data streams, demonstrating that distillation-based methods effectively prevent catastrophic forgetting on earlier tasks while improving adaptation to new data.
Organizations deploying natural language processing models face significant performance degradation over time as language, topics, and real-world conditions shift. Pretrained language models are traditionally trained on massive, static text corpora and then adapted to specific tasks. When new domains or temporal shifts emerge, maintaining these systems typically requires either expensive complete retraining from scratch or naive continuous updates that cause catastrophic forgetting, where the model loses its ability to handle earlier domains.
The article establishes a lifelong language model pretraining framework to evaluate how continuously updating a single model on incoming, unlabeled text streams impacts downstream performance. It specifically tests the model's ability to retain knowledge from past domains, adapt to the latest incoming data, and generalize across temporal gaps where downstream training data is outdated relative to evaluation data.
The researchers designed two real-world data streams using transformer models: a domain-incremental stream composed of millions of academic papers across four disciplines (biomedical, computer science, materials science, and physics) and a chronological stream comprising 100 million tweets across four separate years (2014, 2016, 2018, and 2020). Using these streams, they evaluated baseline approaches against multiple continual learning techniques, including parameter expansion (adapters), episodic memory replay, parameter regularization, and several knowledge distillation methods that penalize output or representation discrepancies between previous and updated model checkpoints.
The analysis revealed several critical findings. First, distillation-based continual learning methods—particularly output logit distillation—proved the most effective at preserving knowledge from earlier domains, improving downstream F1 scores by at least 1.0% to over 1.5% over sequential updating on earlier tasks and matching or exceeding offline multi-task retraining. Second, standard episodic memory replay largely failed to prevent forgetting and led to overfitting on cached examples, even when memory size was increased a hundredfold from 100,000 to 10 million instances. Third, continual pretraining improved performance on the latest temporal data and boosted temporal generalization across time gaps (such as applying models trained on 2016 data to 2020 test sets), outperforming models trained exclusively on current-year data. Finally, lifelong pretraining demonstrated its greatest relative performance gains in low-resource downstream fine-tuning scenarios where labeled target data is scarce.
These findings indicate that organizations do not need to choose between the prohibitive compute and storage costs of recurring full-corpus retraining and the performance degradation of naive continuous updating. Continuous pretraining using knowledge distillation delivers a single, robust model capable of maintaining legacy capabilities while adapting to new domains and temporal evolution, which significantly reduces compute overhead, storage requirements, and privacy risks associated with retaining historic data. However, distillation introduces a trade-off: strongly penalizing representational drift can slightly constrain the model's adaptability when absorbing radically distinct new domains.
Decision-makers should consider adopting distillation-based continual pretraining pipelines over simple replay buffers or parameter-regularization techniques when managing evolving text corpora. Teams should prioritize logit-distillation methods as the baseline continual learning strategy for production pipelines. Before large-scale deployment, organizations should conduct domain-similarity assessments; when domain shifts involve large vocabulary divergence (as in distinct scientific fields rather than social media), engineering teams should calibrate distillation regularization weights to avoid overly rigid model behavior on new tasks.
The primary limitations of the article involve task-dependent variations among advanced distillation methods, where no single contrastive or representation-based variant uniformly outperformed standard logit distillation across all settings. Furthermore, continuous pretraining risks compounding historical biases present in earlier streams unless targeted bias-mitigation techniques are applied. The core findings are supported by consistent evidence across distinct base model sizes (including base and large architectures) and diverse domain and temporal datasets, providing high confidence in distillation as a viable foundation for lifelong language model maintenance.
- Paper: Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks, Suchin Gururangan et al. (2020). This paper establishes domain-adaptive and task-adaptive pretraining for language models, providing the baseline pretraining paradigms that lifelong pretraining directly adapts to continual data streams.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This foundational work introduces distillation-based regularization for continual learning without past data, which serves as the core mechanism evaluated in the source for retaining downstream performance across domains.
- Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). This paper introduces gradient-based continual learning protocols and memory constraints that underpin standard evaluation setups in lifelong pretraining.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). This paper proposes Dark Experience Replay to align output distributions over streaming data, representing a key distillation and replay baseline adapted in continuous pretraining setups.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This comprehensive taxonomy outlines the stability-plasticity trade-offs and core algorithmic families that structure the lifelong pretraining methods evaluated in the source.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). This foundational work develops path-integral parameter regularization to prevent catastrophic forgetting, establishing the parameter-regularization baselines benchmarked against distillation.
- Paper: Time Waits for No One! Analysis and Challenges of Temporal Misalignment, Kelvin Luu et al. (2022). This work deepens the investigation of temporal misalignment in evolving text streams and quantifies performance degradation across longitudinal NLP benchmarks.
- Paper: An Empirical Investigation of the Role of Pre-training in Lifelong Learning, Sanket Vaibhav Mehta et al. (2023). This empirical study investigates the exact impact of scale and pre-training representations on mitigating catastrophic forgetting across diverse NLP tasks.
- Paper: Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time, Michael Y. Hu et al. (2026). This research extends continual midtraining by introducing dynamic adapter-based data mixing to balance old and new domains systematically during continuous pretraining.
- Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). This paper builds on continual adaptation by embedding self-updatable latent memory architectures directly inside language models to continuously integrate streaming knowledge.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). This study extends lifelong model updates to long horizons by evaluating compositions of distillation, replay, and adapter merging across 100 sequential tasks.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). This comprehensive survey provides an overarching theoretical and methodological framework synthesizing representation-based lifelong learning and foundation model adaptation.
