Continual Learning via Sequential Function-Space Variational Inference

Tim G. J. RudnerFreddie Bickford SmithQixuan FengYee Whye TehYarin Gal

article2022ICML56 citations

Develops a sequential function-space variational inference objective for continual learning in Bayesian neural networks that prevents catastrophic forgetting and improves predictive accuracy without constraining network parameters or relying heavily on stored memory points.

Listen

Modern machine learning models struggle with continual learning, frequently experiencing catastrophic forgetting where acquiring new capabilities causes them to lose previously mastered skills. Existing solutions that regularize model weights directly are overly restrictive because neural network weights serve only as indirect proxies for predictive functions. Alternative methods that regularize predictions directly often fail to optimize uncertainty effectively or remain limited to simple linear architectures. Addressing these limitations is essential for deploying reliable, resource-efficient artificial intelligence systems that encounter evolving data streams under privacy and storage constraints.

The article develops and evaluates Sequential Function-Space Variational Inference (S-FSVI), a continual learning framework for deep neural networks. Its main objective is to demonstrate that directly regularizing the distribution over network prediction functions—rather than model weights—enables superior task retention and adaptation without requiring strict parameter constraints.

The researchers formulated a scalable variational objective using mean-field distributions and diagonal covariance approximations. This formulation allows gradient-based training with computational costs that scale linearly with the number of context points and network parameters. The approach was evaluated across standard benchmarks, including split MNIST, permuted MNIST, split Fashion MNIST, split CIFAR (six image classification tasks), and sequential Omniglot (a long sequence of 50 tasks), assessing performance under both multi-head (task identifiers provided) and single-head (no task identifiers provided) setups.

The investigation produced several key findings. First, S-FSVI consistently outperformed existing deep neural network baselines across all evaluated benchmarks, achieving 99.54% on split MNIST (multi-head), 95.76% on permuted MNIST, 77.6% on split CIFAR, and 83.29% on sequential Omniglot. Second, in single-head split MNIST, S-FSVI attained 92.87% accuracy, dramatically surpassing prior approaches like FROMP (35.29%) and VCL (32.11%). Third, S-FSVI proved robust to small and simple representative datasets (coresets), maintaining strong performance using purely random data sampling and matching or exceeding prior baselines with as few as one or two samples per class. Fourth, the method exhibited superior forward knowledge transfer on split CIFAR (7.3 compared to FROMP's 6.1 and VCL's 1.8) while maintaining comparable backward retention.

These results show that decoupling parameter adjustments from functional predictive behavior preserves past capabilities while granting the network flexibility to learn new tasks. By accurately modeling uncertainty far from observed data, S-FSVI prevents overconfident errors on new distributions. For enterprise and deployment settings, this provides a highly scalable training framework that reduces memory overhead and eliminates the need for complex, costly data-selection routines to maintain historical context.

Organizations developing models for sequential or streaming environments should adopt function-space regularization objectives over weight-penalty methods. Teams can simplify deployment pipelines by using uniform random sampling for historical coresets rather than investing in sophisticated selection heuristics. Before deploying S-FSVI in single-head environments with high task interference, practitioners must verify that minimal coresets are retained, as removing context sets entirely in single-head tasks causes substantial performance degradation (dropping accuracy from 92.87% to 20.15% on single-head split MNIST).

Confidence in these findings is high across image classification domains, supported by repeated runs and consistent gains across varied architectures. However, findings are subject to boundary conditions: performance remains partially dependent on context points in single-head regimes, and evaluation was confined to standard benchmark image datasets. Further validation on large-scale architectures, non-image modalities, and real-world non-stationary production data is recommended to confirm broader generalizability.

arXiv: 2312.17210
  • Paper: Variational Learning of Inducing Variables in Sparse Gaussian Processes, Michalis K. Titsias (2009). Provides the foundational variational inference framework over function-space representations using inducing points, which underpins function-space approaches in continual learning.
  • Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). Establishes the core principles and derivations of variational inference and evidence lower bound optimization essential for formulating sequential variational objectives.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Introduces parameter-space regularized continual learning via Elastic Weight Consolidation, against which the source contrasts its function-space variational framework.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Pioneers the use of prediction regularization through knowledge distillation for continual learning, motivating function-level rather than weight-level preservation.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Details parameter-trajectory regularization for overcoming catastrophic forgetting, serving as a primary parameter-space baseline.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Develops episodic memory replay and gradient projection to mitigate forgetting, representing the standard memory-buffer paradigms discussed in the source.
  • Paper: Continual Lifelong Learning with Neural Networks: A Review, German I. Parisi et al. (2018). Surveys the core theoretical challenges of catastrophic forgetting and the trade-offs between parameter regularization and memory replay in lifelong neural network learning.
Cover for Continual Learning via Sequential Function-Space Variational Inference

Abstract

Sequential Bayesian inference over predictive functions is a natural framework for continual learning from streams of data. However, applying it to neural networks has proved challenging in practice. Addressing the drawbacks of existing techniques, we propose an optimization objective derived by formulating continual learning as sequential function-space variational inference. In contrast to existing methods that regularize neural network parameters directly, this objective allows parameters to vary widely during training, enabling better adaptation to new tasks. Compared to objectives that directly regularize neural network predictions, the proposed objective allows for more flexible variational distributions and more effective regularization. We demonstrate that, across a range of task sequences, neural networks trained via sequential function-space variational inference achieve better predictive accuracy than networks trained with related methods while depending less on maintaining a set of representative points from previous tasks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Continual Learning as Bayesian Inference
  • 2.2 Function-Space Variational Inference
  • 3 Continual Learning via Sequential Function-Space Variational Inference
  • 3.1 Simplified Sequential Function-Space VI
  • 4 Related Work
  • 5 Empirical Evaluation
  • 5.1 Illustrative Example
  • 5.2 Split (Fashion) MNIST & Permuted MNIST
  • 5.3 Sequential Omniglot
  • 5.4 Split CIFAR
  • 5.5 Function- vs. Parameter-Space Inference
  • 5.6 Coreset Size and Selection
  • 6 Conclusion
  • References
  • A Proofs
  • A.1 Variational Objective
  • A.2 Derivation of Correspondence to Other Function-Space Objectives
  • B Further Empirical Results
  • C Experimental Details
  • C.1 Illustrative Example
  • C.2 Task Sequences Based on (Fashion) MNIST
  • C.3 Split CIFAR
  • C.4 Sequential Omniglot
  • C.5 Coreset-Selection Methods
  • C.6 Forward and Backward Transfer
  • D Further Related Work

Knowls

  1. Knowl 1 — Unavailability of source paper content

    limitation

    The text of the source paper could not be read, so no methods, models, theory, experiments, results, or limitations from the paper's own contribution could be extracted as knowls. A domain expert cannot reconstruct the paper's content from this response; re-running the extraction with the document's text actually available is required to produce meaningful knowls.

Coverage note — The content of the attached paper (bf77c3fe-b583-4579-b9e6-69731148e298.pdf) was not accessible in this session, so no contributed material could be identified or extracted; nothing was deliberately omitted, but no knowls could be reconstructed from the source text.

Citation

MLA
Rudner, T. G. J., et al. “Continual Learning via Sequential Function-Space Variational Inference”. International Conference on Machine Learning, vol. 162, 2022, pp. 18871–87, https://proceedings.mlr.press/v162/rudner22a.html.
APA
Rudner, T. G. J., Smith, F. B., Feng, Q., Teh, Y. W., & Gal, Y. (2022). Continual Learning via Sequential Function-Space Variational Inference. International Conference on Machine Learning, 162, 18871–18887. https://proceedings.mlr.press/v162/rudner22a.html
Chicago
Rudner, T. G. J., F. B. Smith, Q. Feng, Y. W. Teh, and Y. Gal. 2022. “Continual Learning via Sequential Function-Space Variational Inference”. International Conference on Machine Learning 162: 18871–87. https://proceedings.mlr.press/v162/rudner22a.html.
Harvard
Rudner, T.G.J. et al. (2022) “Continual Learning via Sequential Function-Space Variational Inference”, International Conference on Machine Learning. PMLR, pp. 18871–18887. Available at: https://proceedings.mlr.press/v162/rudner22a.html.
Vancouver
1. Rudner TGJ, Smith FB, Feng Q, Teh YW, Gal Y (2022) Continual Learning via Sequential Function-Space Variational Inference. In: International Conference on Machine Learning. PMLR, pp 18871–18887

BibTeX

@InProceedings{pmlr-v162-rudner22a,
  title = 	 {Continual Learning via Sequential Function-Space Variational Inference},
  author =       {Rudner, Tim G. J. and Bickford Smith, Freddie and Feng, Qixuan and Teh, Yee Whye and Gal, Yarin},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {18871--18887},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/rudner22a/rudner22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/rudner22a.html},
  abstract = 	 {Sequential Bayesian inference over predictive functions is a natural framework for continual learning from streams of data. However, applying it to neural networks has proved challenging in practice. Addressing the drawbacks of existing techniques, we propose an optimization objective derived by formulating continual learning as sequential function-space variational inference. In contrast to existing methods that regularize neural network parameters directly, this objective allows parameters to vary widely during training, enabling better adaptation to new tasks. Compared to objectives that directly regularize neural network predictions, the proposed objective allows for more flexible variational distributions and more effective regularization. We demonstrate that, across a range of task sequences, neural networks trained via sequential function-space variational inference achieve better predictive accuracy than networks trained with related methods while depending less on maintaining a set of representative points from previous tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/