Continual Learning via Sequential Function-Space Variational Inference
Tim G. J. RudnerFreddie Bickford SmithQixuan FengYee Whye TehYarin Gal
Develops a sequential function-space variational inference objective for continual learning in Bayesian neural networks that prevents catastrophic forgetting and improves predictive accuracy without constraining network parameters or relying heavily on stored memory points.
Modern machine learning models struggle with continual learning, frequently experiencing catastrophic forgetting where acquiring new capabilities causes them to lose previously mastered skills. Existing solutions that regularize model weights directly are overly restrictive because neural network weights serve only as indirect proxies for predictive functions. Alternative methods that regularize predictions directly often fail to optimize uncertainty effectively or remain limited to simple linear architectures. Addressing these limitations is essential for deploying reliable, resource-efficient artificial intelligence systems that encounter evolving data streams under privacy and storage constraints.
The article develops and evaluates Sequential Function-Space Variational Inference (S-FSVI), a continual learning framework for deep neural networks. Its main objective is to demonstrate that directly regularizing the distribution over network prediction functions—rather than model weights—enables superior task retention and adaptation without requiring strict parameter constraints.
The researchers formulated a scalable variational objective using mean-field distributions and diagonal covariance approximations. This formulation allows gradient-based training with computational costs that scale linearly with the number of context points and network parameters. The approach was evaluated across standard benchmarks, including split MNIST, permuted MNIST, split Fashion MNIST, split CIFAR (six image classification tasks), and sequential Omniglot (a long sequence of 50 tasks), assessing performance under both multi-head (task identifiers provided) and single-head (no task identifiers provided) setups.
The investigation produced several key findings. First, S-FSVI consistently outperformed existing deep neural network baselines across all evaluated benchmarks, achieving 99.54% on split MNIST (multi-head), 95.76% on permuted MNIST, 77.6% on split CIFAR, and 83.29% on sequential Omniglot. Second, in single-head split MNIST, S-FSVI attained 92.87% accuracy, dramatically surpassing prior approaches like FROMP (35.29%) and VCL (32.11%). Third, S-FSVI proved robust to small and simple representative datasets (coresets), maintaining strong performance using purely random data sampling and matching or exceeding prior baselines with as few as one or two samples per class. Fourth, the method exhibited superior forward knowledge transfer on split CIFAR (7.3 compared to FROMP's 6.1 and VCL's 1.8) while maintaining comparable backward retention.
These results show that decoupling parameter adjustments from functional predictive behavior preserves past capabilities while granting the network flexibility to learn new tasks. By accurately modeling uncertainty far from observed data, S-FSVI prevents overconfident errors on new distributions. For enterprise and deployment settings, this provides a highly scalable training framework that reduces memory overhead and eliminates the need for complex, costly data-selection routines to maintain historical context.
Organizations developing models for sequential or streaming environments should adopt function-space regularization objectives over weight-penalty methods. Teams can simplify deployment pipelines by using uniform random sampling for historical coresets rather than investing in sophisticated selection heuristics. Before deploying S-FSVI in single-head environments with high task interference, practitioners must verify that minimal coresets are retained, as removing context sets entirely in single-head tasks causes substantial performance degradation (dropping accuracy from 92.87% to 20.15% on single-head split MNIST).
Confidence in these findings is high across image classification domains, supported by repeated runs and consistent gains across varied architectures. However, findings are subject to boundary conditions: performance remains partially dependent on context points in single-head regimes, and evaluation was confined to standard benchmark image datasets. Further validation on large-scale architectures, non-image modalities, and real-world non-stationary production data is recommended to confirm broader generalizability.
- Paper: Variational Learning of Inducing Variables in Sparse Gaussian Processes, Michalis K. Titsias (2009). Provides the foundational variational inference framework over function-space representations using inducing points, which underpins function-space approaches in continual learning.
- Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). Establishes the core principles and derivations of variational inference and evidence lower bound optimization essential for formulating sequential variational objectives.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Introduces parameter-space regularized continual learning via Elastic Weight Consolidation, against which the source contrasts its function-space variational framework.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Pioneers the use of prediction regularization through knowledge distillation for continual learning, motivating function-level rather than weight-level preservation.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Details parameter-trajectory regularization for overcoming catastrophic forgetting, serving as a primary parameter-space baseline.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Develops episodic memory replay and gradient projection to mitigate forgetting, representing the standard memory-buffer paradigms discussed in the source.
- Paper: Continual Lifelong Learning with Neural Networks: A Review, German I. Parisi et al. (2018). Surveys the core theoretical challenges of catastrophic forgetting and the trade-offs between parameter regularization and memory replay in lifelong neural network learning.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). Contextualizes function-space variational objectives within the broader modern theoretical taxonomy and categorization of continual learning methods.
- Paper: FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning, Dipam Goswami et al. (2023). Extends the goal of mitigating forgetting without raw exemplar replay by modeling distributional feature representations in class-incremental scenarios.
- Paper: An Empirical Investigation of the Role of Pre-training in Lifelong Learning, Sanket Vaibhav Mehta et al. (2023). Investigates how large pre-trained representations interact with catastrophic forgetting and downstream continual learning objectives across diverse domains.
