Lifelong Language Pretraining with Distribution-Specialized Experts
Wuyang ChenYanqi ZhouNan DuYanping HuangJames LaudonZhifeng ChenClaire Cui
Proposes Lifelong-MoE, an extensible mixture-of-experts pretraining framework that continually adapts large language models to streaming data distributions and prevents catastrophic forgetting by progressively expanding and freezing specialized experts without increasing inference computation.
Large language models typically rely on static, high-quality pretraining datasets, but real-world textual data arrives in continuous, evolving streams such as new web pages, multilingual content, and social media discussions. Sequentially updating models on changing text distributions usually causes catastrophic forgetting, where the model overwrites previously acquired capabilities to fit new data. Completely retraining large models from scratch whenever new data appears is computationally prohibitive, while traditional continual learning approaches focus primarily on adapting models across specific downstream tasks rather than handling continuous, task-agnostic pretraining.
The article demonstrates an extensible framework, termed Lifelong-MoE, that enables large language models to continually pretrain on streaming data distributions without forgetting prior knowledge. It evaluates how dynamically expanding modular expert networks and applying targeted regularizations preserves past learning while maintaining constant computational costs during inference and training.
The authors designed an experimental evaluation using a sequential stream of diverse data distributions comprising hundreds of billions of tokens: English web text and Wikipedia, non-English multilingual text, and public conversational data. Building upon sparsely activated mixture-of-experts architectures—where only the top two expert modules are active per token—the approach progressively adds new specialized experts and gating units for new data distributions, freezes previously trained experts and their gating mechanisms, and applies output distillation regularization to shared dense layers. The evaluated models ranged up to 1.878 billion activated parameters and were benchmarked against standard dense architectures, memory replay strategies, and parameter regularization baselines across nineteen downstream language understanding and generation benchmarks.
The investigation produced several key findings. First, Lifelong-MoE significantly reduced catastrophic forgetting; after completing sequential training across all distributions, its performance dropped by only 39.9% on question answering and 15.3% on translation, compared to drops of 48.5% and 72.7% respectively for standard parameter regularization. Second, the proposed method achieved superior or competitive overall downstream performance, reaching a translation score of 19.16 compared to 11.14 for an oracle dense model trained on all data simultaneously. Third, ablation experiments revealed that explicit protection requires freezing both old experts and their corresponding gating dimensions together, as freezing either component alone degraded performance. Finally, a controlled partial expansion strategy (expanding experts from 4 to 7 to 10) outperformed naive duplication of experts while avoiding exponential growth in total parameter storage.
These findings indicate that modular, sparsely activated neural network architectures offer a practical path for maintaining up-to-date language models. Organizations can incorporate new domain data continuously rather than incurring the substantial financial, energy, and hardware costs of full retraining cycles. Because the architecture keeps the number of activated parameters per token constant, deployment inference latency and per-token compute costs remain unaffected even as total model capacity grows.
Engineering and research teams maintaining production language models should adopt modular expert expansion and selective parameter freezing strategies when integrating streaming, out-of-domain text corpora. When designing expert growth paths, teams should implement partial expert allocation rather than doubling modules to prevent excessive memory footprints, and calibrate output regularization scaling to maintain training stability. Future work should investigate automated criteria for deciding when a distribution shift warrants adding new experts, evaluate scaling behavior on models beyond the two-billion parameter regime, and test the architecture across longer sequences of data distributions.
The article provides strong empirical confidence through systematic ablations and extensive multi-task evaluations across diverse benchmarks. However, the evaluation assumes clear, distinct phase boundaries between training distributions rather than gradual, mixed distribution shifts. Practitioners should exercise caution when deploying the framework in settings with highly uncurated, continuous data mixtures where distinct distribution boundaries are difficult to identify.
- Paper: Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora, Xisen Jin et al. (2022). This earlier study establishes the continual language-pretraining setting and compares replay, regularization, and distillation methods that Lifelong-MoE builds on.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). Its sparsely gated mixture-of-experts design supplies the architectural foundation for Lifelong-MoE’s selectively activated experts.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). Its work on stable, transferable sparse experts clarifies the MoE training and routing challenges that Lifelong-MoE must manage.
- Paper: DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale, Samyam Rajbhandari et al. (2022). Its MoE training and inference techniques provide practical context for Lifelong-MoE’s sparse expert architecture and compute-efficiency goals.
No sufficiently relevant recommendations were found.
