An Empirical Investigation of the Role of Pre-training in Lifelong Learning
Sanket Vaibhav MehtaDarshan PatilSarath ChandarEmma Strubell
Reveals that pre-trained initializations naturally reduce catastrophic forgetting in sequential task learning by converging to wider loss basins, and introduces a sharpness-aware optimization method to explicitly promote flat minima during fine-tuning.
Modern artificial intelligence systems frequently suffer from catastrophic forgetting, where learning new sequential tasks overwrites and degrades previously acquired knowledge. Deploying models capable of lifelong or continual learning is critical for reducing the substantial energy, computational costs, and infrastructure overhead associated with repeatedly retraining large models from scratch. While transfer learning using large pre-trained foundation models has become standard practice across artificial intelligence, prior continual learning research largely focused on models initialized from scratch. The article evaluates how pre-trained weight initializations influence catastrophic forgetting across sequential task learning and demonstrates how to leverage these optimization properties to design more resilient lifelong learning algorithms.
To conduct this evaluation, the authors performed extensive empirical experiments across computer vision and natural language processing benchmarks. The investigation evaluated vision models like ResNet-18 across datasets including Split CIFAR and a diverse five-dataset vision benchmark, while strictly removing class overlaps between pre-training and downstream data. For language tasks, the authors evaluated transformer models such as DistilBERT, BERT, RoBERTa, and T5 across standard datasets as well as a newly introduced 15-dataset natural language processing benchmark. The authors compared pre-trained and randomly initialized models across prominent continual learning algorithms, analyzed loss landscapes through geometric contour mapping and sharpness metrics, and tested an optimization approach known as Sharpness-Aware Minimization during sequential fine-tuning.
The investigation produced four central findings. First, generic pre-trained model initializations implicitly and substantially alleviate catastrophic forgetting across both vision and language domains. On a diverse five-dataset computer vision benchmark, pre-trained ResNet-18 reduced forgetting from 51.5% to 38.3% compared to random initialization, and standard fine-tuning of a pre-trained model frequently outperformed specialized continual learning algorithms applied to randomly initialized models. Second, model capacity and training data diversity directly strengthen retention; for instance, RoBERTa-base outperformed the larger BERT-Large on sequential tasks due to its broader pre-training data. Third, geometric loss landscape analysis revealed that pre-training places models into significantly wider, flatter loss basins, causing future weight updates to induce much smaller increases in earlier task loss. Fourth, explicitly seeking flat minima during sequential learning using Sharpness-Aware Minimization further reduced forgetting across benchmarks, improving final accuracy by approximately 3% to 13% across various baseline configurations.
These findings demonstrate that retaining prior knowledge during sequential learning is heavily governed by the geometric flatness of the optimization basin rather than memory storage alone. For organizations deploying machine learning systems, shifting to pre-trained architectures combined with flatness-seeking optimization lowers operational risks and computing costs by reducing reliance on massive data replay buffers. Rather than exclusively building complex external memory safeguards, machine learning practitioners should prioritize starting with highly diverse pre-trained representations and integrating sharpness-aware optimization routines into their model update pipelines.
Decision-makers should note certain boundary conditions regarding these conclusions. Although pre-training mitigates forgetting, sequential training across highly diverse tasks still exhibits noticeable performance degradation over long task sequences. Additionally, the experimental scope primarily focused on classification tasks within vision and language. Overall confidence in the empirical conclusions is high, supported by rigorous cross-domain evaluations across multiple seeds and task orderings, though teams deploying models for complex generative tasks should validate these techniques on domain-specific pilot pipelines before broad implementation.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Provides the foundational framework of Elastic Weight Consolidation and Fisher-information-based parameter regularization for mitigating catastrophic forgetting in sequential task learning.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Introduces Gradient Episodic Memory, establishing the primary replay- and gradient-projection-based continual learning baseline benchmarked in the source study.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Introduces the Learning without Forgetting paradigm that uses distillation losses to adapt networks sequentially without storing historical task datasets.
- Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). Presents Averaged Gradient Episodic Memory (A-GEM), a core task-sequential benchmark algorithm compared directly against the source's sharpness-aware optimization method.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Formulates Synaptic Intelligence for trajectory-based parameter importance weighting, serving as a standard baseline for continual learning without raw memory replay.
- Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). Develops Memory Aware Synapses, establishing output-sensitivity regularization as a key baseline for retaining capabilities across task sequences.
- Paper: Why Does Unsupervised Pre-training Help Deep Learning?, Dumitru Erhan et al. (2010). Offers foundational empirical and theoretical analysis on how pre-training alters optimization landscapes and guides neural networks toward better basins.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). Surveys the taxonomy and evaluation methodology of continual learning classification algorithms, providing context for the baselines investigated in the study.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Introduces Dark Experience Replay, representing modern experience replay baselines evaluated in sequential fine-tuning setups.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). Demonstrates how pre-trained representation backbones can be adapted for continual learning without full retraining, motivating the source's investigation into pre-training dynamics.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). Extends the study of pre-training and optimization-based loss landscapes into a broader, unified taxonomy of continual learning theory and methodology.
- Paper: Mitigating the Alignment Tax of RLHF, Yong Lin et al. (2024). Investigates the alignment tax and catastrophic forgetting specifically during the sequential post-pre-training alignment phase of large language models.
- Paper: Machine Unlearning of Pre-trained Large Language Models, Jin Yao et al. (2024). Applies the dynamics of forgetting and parameter preservation in pre-trained large language models to the deliberate task of machine unlearning.
- Paper: LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, Bo Liu et al. (2023). Extends continual learning evaluation of pre-trained models from standard vision and NLP benchmarks to complex sequential manipulation tasks in robotics.
