An Empirical Investigation of the Role of Pre-training in Lifelong Learning

Sanket Vaibhav MehtaDarshan PatilSarath ChandarEmma Strubell

article2023JMLR184 citations

Reveals that pre-trained initializations naturally reduce catastrophic forgetting in sequential task learning by converging to wider loss basins, and introduces a sharpness-aware optimization method to explicitly promote flat minima during fine-tuning.

Listen

Modern artificial intelligence systems frequently suffer from catastrophic forgetting, where learning new sequential tasks overwrites and degrades previously acquired knowledge. Deploying models capable of lifelong or continual learning is critical for reducing the substantial energy, computational costs, and infrastructure overhead associated with repeatedly retraining large models from scratch. While transfer learning using large pre-trained foundation models has become standard practice across artificial intelligence, prior continual learning research largely focused on models initialized from scratch. The article evaluates how pre-trained weight initializations influence catastrophic forgetting across sequential task learning and demonstrates how to leverage these optimization properties to design more resilient lifelong learning algorithms.

To conduct this evaluation, the authors performed extensive empirical experiments across computer vision and natural language processing benchmarks. The investigation evaluated vision models like ResNet-18 across datasets including Split CIFAR and a diverse five-dataset vision benchmark, while strictly removing class overlaps between pre-training and downstream data. For language tasks, the authors evaluated transformer models such as DistilBERT, BERT, RoBERTa, and T5 across standard datasets as well as a newly introduced 15-dataset natural language processing benchmark. The authors compared pre-trained and randomly initialized models across prominent continual learning algorithms, analyzed loss landscapes through geometric contour mapping and sharpness metrics, and tested an optimization approach known as Sharpness-Aware Minimization during sequential fine-tuning.

The investigation produced four central findings. First, generic pre-trained model initializations implicitly and substantially alleviate catastrophic forgetting across both vision and language domains. On a diverse five-dataset computer vision benchmark, pre-trained ResNet-18 reduced forgetting from 51.5% to 38.3% compared to random initialization, and standard fine-tuning of a pre-trained model frequently outperformed specialized continual learning algorithms applied to randomly initialized models. Second, model capacity and training data diversity directly strengthen retention; for instance, RoBERTa-base outperformed the larger BERT-Large on sequential tasks due to its broader pre-training data. Third, geometric loss landscape analysis revealed that pre-training places models into significantly wider, flatter loss basins, causing future weight updates to induce much smaller increases in earlier task loss. Fourth, explicitly seeking flat minima during sequential learning using Sharpness-Aware Minimization further reduced forgetting across benchmarks, improving final accuracy by approximately 3% to 13% across various baseline configurations.

These findings demonstrate that retaining prior knowledge during sequential learning is heavily governed by the geometric flatness of the optimization basin rather than memory storage alone. For organizations deploying machine learning systems, shifting to pre-trained architectures combined with flatness-seeking optimization lowers operational risks and computing costs by reducing reliance on massive data replay buffers. Rather than exclusively building complex external memory safeguards, machine learning practitioners should prioritize starting with highly diverse pre-trained representations and integrating sharpness-aware optimization routines into their model update pipelines.

Decision-makers should note certain boundary conditions regarding these conclusions. Although pre-training mitigates forgetting, sequential training across highly diverse tasks still exhibits noticeable performance degradation over long task sequences. Additionally, the experimental scope primarily focused on classification tasks within vision and language. Overall confidence in the empirical conclusions is high, supported by rigorous cross-domain evaluations across multiple seeds and task orderings, though teams deploying models for complex generative tasks should validate these techniques on domain-specific pilot pipelines before broad implementation.

arXiv: 2112.09153
Cover for An Empirical Investigation of the Role of Pre-training in Lifelong Learning

Abstract

The lifelong learning paradigm in machine learning is an attractive alternative to the more prominent isolated learning scheme not only due to its resemblance to biological learning but also its potential to reduce energy waste by obviating excessive model re-training. A key challenge to this paradigm is the phenomenon of catastrophic forgetting. With the increasing popularity and success of pre-trained models in machine learning, we pose the question: What role does pre-training play in lifelong learning, specifically with respect to catastrophic forgetting? We investigate existing methods in the context of large, pre-trained models and evaluate their performance on a variety of text and image classification tasks, including a large-scale study using a novel data set of 15 diverse NLP tasks. Across all settings, we observe that generic pre-training implicitly alleviates the effects of catastrophic forgetting when learning multiple tasks sequentially compared to randomly initialized models. We then further investigate why pre-training alleviates forgetting in this setting. We study this phenomenon by analyzing the loss landscape, finding that pre-trained weights appear to ease forgetting by leading to wider minima. Based on this insight, we propose jointly optimizing for current task loss and loss basin sharpness to explicitly encourage wider basins during sequential fine-tuning. We show that this optimization approach outperforms several state-of-the-art task-sequential continual learning algorithms across multiple settings, occasionally even without retaining a memory that scales in size with the number of tasks.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Problem Setup: Task Incremental Learning
  • 2.2 Benchmarks
  • 2.2.1 CV Benchmarks
  • 2.2.2 NLP Benchmarks
  • 2.2.3 15-Dataset-NLP Benchmark
  • 2.3 Task sequences
  • 2.4 Evaluation
  • 2.5 Methods
  • 3 Does pre-training implicitly alleviate forgetting?
  • 3.1 How much does pre-training help in alleviating forgetting?
  • 3.2 Do pre-trained models undergo similar forgetting on diverse and homogeneous tasks?
  • 3.3 How do different pre-trained initialization affect forgetting?
  • 4 Exploring the Loss Landscape
  • 4.1 Loss Contour
  • 4.2 Linear Model Interpolation
  • 4.3 Sharpness
  • 5 Lifelong Learning with Sharpness Aware Minimization (SAM)
  • 5.1 Loss Contours and Sharpness with SAM
  • 5.2 Analyzing the influence of pre-training task minima curvature on forgetting
  • 5.3 Analyzing the influence of task-agnostic favorable initializations on forgetting
  • 6 Related Work
  • 7 Discussion
  • A Implementation Details
  • A.1 CV Experiments
  • A.2 NLP Experiments
  • A.3 Sharpness metric
  • B Task-specific results
  • B.1 5-dataset-NLP
  • B.2 5-dataset-CV
  • C Loss Landscape
  • C.1 Loss Contours
  • References

Knowls

  1. Knowl 1 — No knowls could be extracted because the paper's content was not accessible

    limitation

    The source document for this extraction task was not readable, so no methods, results, or other contributed content from the paper could be identified or verified. This entry is not a claim about the paper's scientific content; it records only that the extraction could not be performed from the material provided.

Coverage note — The attached PDF's text could not be read, so no knowls could be extracted from the paper's contribution; nothing was deliberately omitted because no contributed material was accessible.

Citation

MLA
Mehta, S. V., et al. “An Empirical Investigation of the Role of Pre-training in Lifelong Learning”. Journal of Machine Learning Research, vol. 24, no. 214, 2023, pp. 1–0, https://www.jmlr.org/papers/v24/22-0496.html.
APA
Mehta, S. V., Patil, D., Chandar, S., & Strubell, E. (2023). An Empirical Investigation of the Role of Pre-training in Lifelong Learning. Journal of Machine Learning Research, 24(214), 1–50. https://www.jmlr.org/papers/v24/22-0496.html
Chicago
Mehta, S. V., D. Patil, S. Chandar, and E. Strubell. 2023. “An Empirical Investigation of the Role of Pre-training in Lifelong Learning”. Journal of Machine Learning Research 24 (214): 1–50. https://www.jmlr.org/papers/v24/22-0496.html.
Harvard
Mehta, S.V. et al. (2023) “An Empirical Investigation of the Role of Pre-training in Lifelong Learning”, Journal of Machine Learning Research, 24(214), pp. 1–50. Available at: https://www.jmlr.org/papers/v24/22-0496.html.
Vancouver
1. Mehta SV, Patil D, Chandar S, Strubell E (2023) An Empirical Investigation of the Role of Pre-training in Lifelong Learning. Journal of Machine Learning Research 24:1–50

BibTeX

@article{JMLR:v24:22-0496,
  author  = {Sanket Vaibhav Mehta and Darshan Patil and Sarath Chandar and Emma Strubell},
  title   = {An Empirical Investigation of the Role of Pre-training in Lifelong Learning},
  journal = {Journal of Machine Learning Research},
  year    = {2023},
  volume  = {24},
  number  = {214},
  pages   = {1--50},
  url     = {http://jmlr.org/papers/v24/22-0496.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/