Temporal-Difference Variational Continual Learning

Luckeciano Carvalho MeloAlessandro AbateYarin Gal

article2025NeurIPS1 citations

Introduces a temporal-difference-inspired variational continual learning objective that regularizes model updates using multiple past posterior estimates to prevent compounding approximation errors and reduce catastrophic forgetting.

Listen

Deployed machine learning models must continually learn from sequential streams of new data while retaining knowledge acquired from previous tasks. In practice, models frequently suffer from catastrophic forgetting, where new information rapidly overwrites existing capabilities, degrading system reliability in production environments. While Bayesian Continual Learning methods use variational inference to balance new learning with memory stability by regularizing updates against the most recent estimated model state, these conventional methods remain vulnerable. Over successive learning steps, single-step approximation errors propagate and compound, leading to severe performance loss over extended task sequences.

The article develops and evaluates a new optimization framework, called Temporal-Difference Variational Continual Learning (TD-VCL), which regularizes updates across a sequence of multiple past posterior estimates rather than relying solely on the single most recent one. By bridging continual learning with temporal-difference principles from reinforcement learning, the objective aims to dilute individual approximation errors and prevent compounding degradation.

The authors validated this approach through extensive empirical benchmarks comparing standard variational methods, replay-augmented techniques, non-variational baselines, and variants enhanced by the new framework. The evaluation utilized five distinct image classification suites, including newly designed, highly constrained single-head benchmarks (PermutedMNIST-Hard, SplitMNIST-Hard, SplitNotMNIST-Hard) with strictly limited replay memory, as well as complex vision benchmarks (CIFAR100-10 and TinyImageNet-10) using deep convolutional architectures trained from scratch without pre-trained features.

The experimental findings demonstrate substantial performance advantages over existing methods. First, TD-VCL consistently outperformed standard Variational Continual Learning across all benchmarks, achieving approximately 89–90% final accuracy on PermutedMNIST-Hard compared to 78–81% for standard baselines and 40–51% for non-variational approaches. Second, the method exhibited marked resilience on earlier tasks; on the first task in PermutedMNIST-Hard after observing ten sequential tasks, TD-VCL maintained 80–85% accuracy, whereas standard variational approaches dropped to 50–60% and non-variational baselines collapsed to 20%. Third, TD-VCL demonstrated scalability on harder datasets with deeper networks, maintaining superior performance on CIFAR100-10 (71% final accuracy versus 65–66% for baseline variational methods) and TinyImageNet-10. Finally, incorporating the temporal-difference objective into other Bayesian continual learning algorithms (such as UCL and UCB) consistently boosted their accuracy, confirming its general applicability across base architectures.

These results indicate that accounting for historical estimates effectively stabilizes long-term model memory without compromising plasticity or requiring unbounded data replay. For decision-makers and system architects, this capability reduces operational risks and performance degradation in continually adapting systems, providing uncertainty-aware predictions vital for safety-critical settings.

Organizations deploying sequential machine learning systems should consider adopting multi-step temporal-difference regularization frameworks. Technical teams implementing the approach should start with the flexible TD(λ)-VCL formulation and tune the history window parameter n and the weighting factor λ, as sensitivity analysis shows performance gains saturate beyond an optimal lookback window. When deploying on resource-constrained embedded systems, teams can mitigate memory overhead by offloading historical model checkpoints to central storage or computing regularization terms asynchronously on general processors.

Confidence in these findings is supported by rigorous multi-seed evaluation, explicit confidence intervals, and consistent gains across varied architectures. However, decision-makers should note that computational and memory requirements depend directly on the number of historical snapshots retained, and optimal hyperparameter configurations vary depending on specific task characteristics.

No sufficiently relevant recommendations were found.

Cover for Temporal-Difference Variational Continual Learning

Abstract

Machine Learning models in real-world applications must continuously learn new tasks to adapt to shifts in the data-generating distribution. Yet, for Continual Learning (CL), models often struggle to balance learning new tasks (plasticity) with retaining previous knowledge (memory stability). Consequently, they are susceptible to Catastrophic Forgetting, which degrades performance and undermines the reliability of deployed systems. In the Bayesian CL literature, variational methods tackle this challenge by employing a learning objective that recursively updates the posterior distribution while constraining it to stay close to its previous estimate. Nonetheless, we argue that these methods may be ineffective due to compounding approximation errors over successive recursions. To mitigate this, we propose new learning objectives that integrate the regularization effects of multiple previous posterior estimations, preventing individual errors from dominating future posterior updates and compounding over time. We reveal insightful connections between these objectives and Temporal-Difference methods, a popular learning mechanism in Reinforcement Learning and Neuroscience. Experiments on challenging CL benchmarks show that our approach effectively mitigates Catastrophic Forgetting, outperforming strong Variational CL methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 Temporal-Difference Variational Continual Learning
  • 4.1 Variational Continual Learning with n-Step KL Regularization
  • 4.2 From n-Step KL to Temporal-Difference Targets
  • 5 Experiments and Discussion
  • 5.1 Experiments
  • 6 Closing Remarks
  • References
  • A Derivation of the n-Step KL Regularization Objective
  • B Derivation of the Temporal-Difference VCL Objective
  • C The connection of TD Targets in TD-VCL and Reinforcement Learning
  • D TD(λ\lambda)-VCL is a discounted sum of n-Step TD targets
  • E TD-VCL: A spectrum of Continual Learning algorithms
  • F Impact Statement
  • G Implementation Details and Reproducibility
  • H Hyperparameters
  • I PermutedMNIST-Hard, SplitMNIST-Hard, and SplitNotMNIST-Hard: Introducing Higher Standards for MNIST/NotMNIST-based Continual Learning Benchmarks
  • J Benchmarks Description
  • K Per Task Performance: Additional Results
  • K.1 SplitMNIST-Hard
  • K.2 SplitNotMNIST-Hard
  • K.3 CIFAR100-10
  • K.4 TinyImageNet-10
  • L Hyperparameters Robustness Analysis
  • L.1 n-Step KL Regularization
  • L.2 TD(λ\lambda)-VCL
  • M Full Table Results
  • N Does TD-VCL Assume Knowledge of Task Boundaries?
  • O Further Questions
  • O.1 What is the computational cost associated with TD-VCL?
  • O.2 What is the memory cost of TD-VCL?
  • O.3 Is maintaining previous posteriors a major bottleneck? Can we optimize this cost?
  • O.4 When should one use nn-Step TD-VCL or TD(λ\lambda)-VCL?
  • O.5 What is the impact of Early Stopping in the presented methods?

Knowls

  1. Knowl 1 — n-Step KL regularization preserves the VCL update while using multiple past posteriors

    model/method

    For sequential tasks with datasets D_1,[?] replaced below by D1,…,DtD_1,\ldots,D_t, assume the task likelihood factorizes across tasks given parameters θ\theta, so that p(θ∣D1:t)∝p(θ∣D1:t−1)p(Dt∣θ)p(\theta\mid D_{1:t})\propto p(\theta\mid D_{1:t-1})p(D_t\mid\theta). Let Q\mathcal Q be a family of variational distributions, and let qs(θ)q_s(\theta) denote the variational posterior after task ss. For an integer horizon nn with 1≤n≤t1\leq n\leq t, the standard VCL update has the equivalent objective

    qt=arg⁡max⁡q∈Q{Eθ∼q[∑i=0n−1n−inlog⁡p(Dt−i∣θ)]−1n∑i=0n−1DKL ⁣(q(θ) ∥ qt−i−1(θ))}.q_t=\arg\max_{q\in\mathcal Q}\left\{\mathbb E_{\theta\sim q}\left[\sum_{i=0}^{n-1}\frac{n-i}{n}\log p(D_{t-i}\mid\theta)\right]-\frac{1}{n}\sum_{i=0}^{n-1}D_{\mathrm{KL}}\!\left(q(\theta)\,\|\,q_{t-i-1}(\theta)\right)\right\}.

    Here Dt−iD_{t-i} is the dataset from task t−it-i, and DKLD_{\mathrm{KL}} is Kullback–Leibler divergence. Thus one update can distribute its regularization across several stored posterior estimates rather than relying only on qt−1q_{t-1}. The task likelihood terms are also reweighted across the nn recent datasets, with more recent tasks receiving larger coefficients. Setting n=1n=1 recovers vanilla VCL.

  2. Knowl 2 — TD(λ)-VCL geometrically weights the influence of past posterior estimates

    model/method

    Let qs(θ)q_s(\theta) be the variational posterior after task ss, let DsD_s be that task's dataset, and let Q\mathcal Q be the variational family. For an integer horizon 1≤n≤t1\leq n\leq t and 0≤λ<10\leq\lambda<1, TD(λ\lambda)-VCL uses the objective

    qt=arg⁡max⁡q∈Q{Eθ∼q[∑i=0n−1λi(1−λn−i)1−λnlog⁡p(Dt−i∣θ)]−∑i=0n−1λi(1−λ)1−λnDKL ⁣(q(θ) ∥ qt−i−1(θ))}.q_t=\arg\max_{q\in\mathcal Q}\left\{\mathbb E_{\theta\sim q}\left[\sum_{i=0}^{n-1}\frac{\lambda^i(1-\lambda^{n-i})}{1-\lambda^n}\log p(D_{t-i}\mid\theta)\right]-\sum_{i=0}^{n-1}\frac{\lambda^i(1-\lambda)}{1-\lambda^n}D_{\mathrm{KL}}\!\left(q(\theta)\,\|\,q_{t-i-1}(\theta)\right)\right\}.

    The likelihood and KL terms therefore combine information from multiple recent tasks and posterior approximations, while the factor λi\lambda^i geometrically reduces the weight assigned farther back in time. The likelihood terms are expected under the candidate distribution qq; the previous qt−i−1q_{t-i-1} are fixed estimates from earlier updates. This objective is an equivalent representation of the standard VCL target under the task-wise Bayesian recursion.

  3. Knowl 3 — TD-VCL objectives are combinations of temporal-difference targets

    theoretical result

    For a candidate variational distribution qq at task tt, define the mm-step target, for 1≤m≤t1\leq m\leq t, by

    TDt(m;q)=Eθ∼q[∑i=0m−1log⁡p(Dt−i∣θ)]−DKL ⁣(q(θ) ∥ qt−m(θ)),\mathrm{TD}_t(m;q)=\mathbb E_{\theta\sim q}\left[\sum_{i=0}^{m-1}\log p(D_{t-i}\mid\theta)\right]-D_{\mathrm{KL}}\!\left(q(\theta)\,\|\,q_{t-m}(\theta)\right),

    where DsD_s is the dataset from task ss and qsq_s is its earlier variational posterior estimate. Maximizing this target over q∈Qq\in\mathcal Q gives the same optimizer as the standard VCL update. TD(λ\lambda)-VCL is equivalently the normalized discounted sum

    qt=arg⁡max⁡q∈Q1−λ1−λn∑k=0n−1λkTDt(k+1;q).q_t=\arg\max_{q\in\mathcal Q}\frac{1-\lambda}{1-\lambda^n}\sum_{k=0}^{n-1}\lambda^k\mathrm{TD}_t(k+1;q).

    The nn-step KL objective is the simple average of the first nn such targets. In this construction, likelihood expectations are estimated by Monte Carlo, while the KL term bootstraps from an earlier posterior estimate. The paper connects this structure to temporal-difference learning: the continual-learning recursion moves backward through past tasks, whereas the reinforcement-learning recursion uses future steps.

  4. Knowl 4 — The TD(λ)-VCL family interpolates between vanilla VCL and n-Step KL

    theoretical result

    In the TD(λ\lambda)-VCL objective, the horizon nn determines how many past task likelihoods and posterior estimates can contribute, while λ\lambda controls their relative weighting. At λ=0\lambda=0, only the most recent task likelihood and posterior estimate contribute, yielding the vanilla VCL objective for any admissible nn. In the limit λ→1\lambda\to1, the likelihood coefficient for task t−it-i converges to (n−i)/n(n-i)/n, and each KL term's coefficient converges to 1/n1/n; the objective therefore becomes the nn-Step KL regularization objective. Intermediate values of λ\lambda give geometrically greater emphasis to more recent estimates. The paper interprets this family as combining Monte Carlo likelihood estimates with posterior-based bootstrapping.

  5. Knowl 5 — Evaluation uses restricted-memory, single-head continual-learning benchmarks

    experimental setup

    The evaluation covers five sequential classification benchmarks. The three MNIST/NotMNIST “Hard” benchmarks enforce single-head classifiers and restrict replay: PermutedMNIST-Hard has 10 pixel-permutation tasks and permits 200 examples from each of the two most recent past tasks; SplitMNIST-Hard has five digit-pair tasks (0/1 through 8/9) and permits 40 examples from only the most recent past task; SplitNotMNIST-Hard uses five letter-pair tasks (A/F through E/J) with the same replay restriction. CIFAR100-10 and TinyImageNet-10 each have 10 tasks; replay is limited to 200 examples per task, with TinyImageNet-10 restricted to the three most recent tasks. The Hard benchmarks are designed to make forgetting more pronounced than in unrestricted versions.

    For the first three benchmarks, the reported fully connected architectures have hidden-layer widths [100, 100], [256, 256], and [150, 150, 150, 150], respectively. CIFAR100-10 and TinyImageNet-10 use Bayesian AlexNet without pretrained representations. The shared training settings are batch size 256, learning rate 10−310^{-3}, maximum 100 epochs, Adam optimization, and early stopping with patience five. Variational methods use a Gaussian mean-field posterior and Gaussian prior, calculate KL terms analytically, and estimate expected likelihoods with Monte Carlo; posterior predictive accuracy is also estimated by Monte Carlo. Results compare Online MLE, Batch MLE, VCL, VCL CoreSet, n-Step TD-VCL, and TD(λ\lambda)-VCL. Reported benchmark averages use 10 seeds for the first three benchmarks and 5 for the image benchmarks; table uncertainties are reported as two standard deviations.

  6. Knowl 6 — TD-VCL improves final average accuracy across the five benchmarks

    data/table

    The table compares average accuracy across all tasks observed at the final evaluation point for each benchmark. Each value is the reported mean ± two standard deviations; higher accuracy indicates better retention and acquisition across the task stream. Both proposed objectives exceed standard VCL on every benchmark, and achieve the top or tied-top result in each column. On PermutedMNIST-Hard, the gap between TD-VCL and VCL grows from 0.78 to 0.88–0.89; on the harder image benchmarks, the proposed methods also improve on the replay-based VCL CoreSet baseline.

    Method PermutedMNIST-Hard (t=10t=10) SplitMNIST-Hard (t=5t=5) SplitNotMNIST-Hard (t=5t=5) CIFAR100-10 (t=10t=10) TinyImageNet-10 (t=10t=10)
    Online MLE 0.40±0.080.40\pm0.08 0.57±0.060.57\pm0.06 0.51±0.040.51\pm0.04 0.52±0.040.52\pm0.04 0.44±0.030.44\pm0.03
    Batch MLE 0.51±0.060.51\pm0.06 0.59±0.030.59\pm0.03 0.50±0.060.50\pm0.06 0.54±0.070.54\pm0.07 0.51±0.030.51\pm0.03
    VCL 0.78±0.040.78\pm0.04 0.64±0.110.64\pm0.11 0.51±0.060.51\pm0.06 0.66±0.010.66\pm0.01 0.51±0.020.51\pm0.02
    VCL CoreSet 0.81±0.030.81\pm0.03 0.62±0.030.62\pm0.03 0.51±0.070.51\pm0.07 0.65±0.020.65\pm0.02 0.54±0.020.54\pm0.02
    n-Step TD-VCL 0.88±0.020.88\pm0.02 0.67±0.040.67\pm0.04 0.58±0.080.58\pm0.08 0.69±0.020.69\pm0.02 0.56±0.020.56\pm0.02
    TD(λ\lambda)-VCL 0.89±0.020.89\pm0.02 0.66±0.020.66\pm0.02 0.58±0.090.58\pm0.09 0.71±0.010.71\pm0.01 0.56±0.020.56\pm0.02
  7. Knowl 7 — The TD objective also improves UCL and UCB in most tested settings

    empirical result

    The authors incorporated the TD(λ\lambda) objective into UCL and UCB, which respectively alter VCL regularization and learning-rate adaptation. The table reports average accuracy over all observed tasks at each benchmark's final timestep, as mean ± two standard deviations. TD enhancement improves the corresponding base method in 14 of the 15 comparisons shown; the exception is UCL on SplitNotMNIST-Hard, where its score changes from 0.52±0.040.52\pm0.04 to 0.51±0.060.51\pm0.06. This supports compatibility with different Bayesian continual-learning mechanisms, while showing that improvements are not universal for every method–benchmark pair.

    Base method Variant PermutedMNIST-Hard (t=10t=10) SplitMNIST-Hard (t=5t=5) SplitNotMNIST-Hard (t=5t=5) CIFAR100-10 (t=10t=10) TinyImageNet-10 (t=10t=10)
    VCL VCL 0.78±0.040.78\pm0.04 0.64±0.110.64\pm0.11 0.51±0.060.51\pm0.06 0.66±0.010.66\pm0.01 0.51±0.020.51\pm0.02
    VCL TD(λ\lambda)-VCL 0.89±0.020.89\pm0.02 0.67±0.040.67\pm0.04 0.58±0.090.58\pm0.09 0.71±0.010.71\pm0.01 0.56±0.020.56\pm0.02
    UCL UCL 0.73±0.120.73\pm0.12 0.66±0.060.66\pm0.06 0.52±0.040.52\pm0.04 0.62±0.020.62\pm0.02 0.50±0.030.50\pm0.03
    UCL TD(λ\lambda)-UCL 0.84±0.040.84\pm0.04 0.70±0.040.70\pm0.04 0.51±0.060.51\pm0.06 0.67±0.030.67\pm0.03 0.56±0.010.56\pm0.01
    UCB UCB 0.83±0.020.83\pm0.02 0.75±0.100.75\pm0.10 0.61±0.050.61\pm0.05 0.66±0.010.66\pm0.01 0.42±0.030.42\pm0.03
    UCB TD(λ\lambda)-UCB 0.88±0.020.88\pm0.02 0.80±0.030.80\pm0.03 0.63±0.030.63\pm0.03 0.70±0.010.70\pm0.01 0.47±0.020.47\pm0.02
  8. Knowl 8 — TD-VCL does not require task boundaries

    theoretical result

    Under the paper's task-wise likelihood factorization, a Bayesian update may group any contiguous chunk of a data stream into one update: the chunk likelihood is the product of its constituent task likelihoods. Consequently, the variational objective can sum likelihood contributions within a chunk, and the update rule does not require the learner to know where task boundaries lie.

    The authors tested this claim on StreamingPermutedMNIST-Hard, which provides a data stream with randomly placed boundaries rather than explicit task boundaries. After the full stream, average accuracy was 0.89±0.020.89\pm0.02 for both n-Step TD-VCL and TD(λ\lambda)-VCL, compared with 0.88±0.020.88\pm0.02 and 0.89±0.020.89\pm0.02, respectively, on PermutedMNIST-Hard. The streaming results therefore showed no observed performance penalty for removing boundary information. The reported streaming values were averaged over 10 seeds.

  9. Knowl 9 — The objectives have an optimal useful horizon and modest λ sensitivity

    empirical result

    Hyperparameter sensitivity was evaluated on PermutedMNIST-Hard. For n-Step KL regularization, increasing the horizon nn improved performance up to about n=5n=5; beyond that, gains saturated and could become slightly detrimental. Performance was sensitive to likelihood tempering β\beta: values that were too high impaired fitting new tasks, while values that were too low weakened retention of previous knowledge; the tested moderate values 0.0010.001, 0.0050.005, and 0.010.01 balanced these effects well. TD(λ\lambda)-VCL showed mild sensitivity to λ\lambda, with differences becoming more visible as more tasks were observed and some horizons favoring lower values. The authors report that the useful choice of λ\lambda depends partly on whether the newest posterior estimates are more informative than older ones.

  10. Knowl 10 — TD-VCL trades additional posterior storage and tuning for reduced forgetting

    limitation

    TD-VCL requires retaining multiple earlier posterior estimates, so memory use grows with the horizon nn and with the size of the variational parameterization; its hyperparameters nn and λ\lambda may require tuning for a new setting. The paper characterizes posterior storage as O(n)O(n) and notes that this can become important for larger networks. The authors also report that the KL terms are computed analytically without data-dependent network forward passes, making training cost nearly the same as VCL; the principal additional costs are posterior storage and hyperparameter search. They suggest controlling nn to meet memory constraints and note that earlier posteriors need not remain in GPU memory.

Coverage note — Proof derivations, complete intermediate-timestep tables, and the per-task plots are omitted because the standalone objective statements and final-timestep comparisons retain the main theoretical and empirical contributions without reproducing supporting detail.

References

  1. 1.Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. An empirical investigation of catastrophic forgetting in gradient-based neural networks. In International Conference on Learning Representations, pages 1–10, 2015.
  2. 2.Michael McCloskey and Neal J. Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of Learning and Motivation, 24:109–165, 1989. URL https://api.semanticscholar.org/CorpusID:61019113.
  3. 3.Jeffrey C. Schlimmer and Douglas Fisher. A case study of incremental concept induction. In Proceedings of the Fifth AAAI National Conference on Artificial Intelligence, AAAI’86, page 496–501. AAAI Press, 1986.
  4. 4.Wickliffe C. Abraham and Anthony Robins. Memory retention – the synaptic stability versus plasticity dilemma. Trends in Neurosciences, 28(2):73–78, 2005. ISSN 0166-2236. doi: https://doi.org/10.1016/j.tins.2004.12.003. URL https://www.sciencedirect.com/science/article/pii/S0166223604003704.
  5. 5.Richard S. Sutton and Steven D. Whitehead. Online learning with random representations. In Proceedings of the Tenth International Conference on International Conference on Machine Learning, ICML’93, page 314–321, San Francisco, CA, USA, 1993. Morgan Kaufmann Publishers Inc. ISBN 1558603077.
  6. 6.Robert M. French. Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4):128–135, 1999. ISSN 1364-6613. doi: https://doi.org/10.1016/S1364-6613(99)01294-2. URL https://www.sciencedirect.com/science/article/pii/S1364661399012942.
  7. 7.Raia Hadsell, Dushyant Rao, Andrei A. Rusu, and Razvan Pascanu. Embracing change: Continual learning in deep neural networks. Trends in Cognitive Sciences, 24(12):1028–1040, 2020. ISSN 1364-6613. doi: https://doi.org/10.1016/j.tics.2020.09.004. URL https://www.sciencedirect.com/science/article/pii/S1364661320302199.
  8. 8.Joan Serra, Didac Suris, Marius Miron, and Alexandros Karatzoglou. Overcoming catastrophic forgetting with hard attention to the task. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 4548–4557. PMLR, 10–15 Jul 2018. URL https://proceedings.mlr.press/v80/serra18a.html.
  9. 9.Tom Rainforth, Adam Foster, Desi R. Ivanova, and Freddie Bickford Smith. Modern Bayesian Experimental Design. Statistical Science, 39(1):100 – 114, 2024. doi: 10.1214/23-STS915. URL https://doi.org/10.1214/23-STS915.
  10. 10.Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 5580–5590, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN 9781510860964.
  11. 11.Yarin Gal, Riashat Islam, and Zoubin Ghahramani. Deep bayesian active learning with image data. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 1183–1192. JMLR.org, 2017.
  12. 12.Luckeciano C. Melo, Panagiotis Tigas, Alessandro Abate, and Yarin Gal. Deep bayesian active learning for preference modeling in large language models, 2024. URL https://arxiv.org/abs/2406.10023.
  13. 13.Cuong V. Nguyen, Yingzhen Li, Thang D. Bui, and Richard E. Turner. Variational continual learning. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=BkQqq0gRb.
  14. 14.Richard S. Sutton. Learning to predict by the methods of temporal differences. Mach. Learn., 3(1):9–44, August 1988. ISSN 0885-6125. doi: 10.1023/A:1022633531479. URL https://doi.org/10.1023/A:1022633531479.
  15. 15.Wolfram Schultz, Peter Dayan, and P. Read Montague. A neural substrate of prediction and reward. Science, 275(5306):1593–1599, 1997. doi: 10.1126/science.275.5306.1593. URL https://www.science.org/doi/abs/10.1126/science.275.5306.1593.
  16. 16.Mark B. Ring. Child: A first step towards continual learning. Mach. Learn., 28(1):77–104, jul 1997. ISSN 0885-6125. doi: 10.1023/A:1007331723572. URL https://doi.org/10.1023/A:1007331723572.
  17. 17.Timo Flesch, Andrew Saxe, and Christopher Summerfield. Continual task learning in natural and artificial agents. Trends in Neurosciences, 46(3):199–210, 2023. ISSN 0166-2236. doi: https://doi.org/10.1016/j.tins.2022.12.006. URL https://www.sciencedirect.com/science/article/pii/S0166223622002600.
  18. 18.Tameem Adel, Han Zhao, and Richard E. Turner. Continual learning with adaptive weights (CLAW). In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=Hklso24Kwr.
  19. 19.James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114:3521 – 3526, 2016. URL https://api.semanticscholar.org/CorpusID:4704285.
  20. 20.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 3987–3995. JMLR.org, 2017.
  21. 21.Arslan Chaudhry, Puneet K. Dokania, Thalaiyasingam Ajanthan, and Philip H. S. Torr. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss, editors, Computer Vision – ECCV 2018, pages 556–572, Cham, 2018. Springer International Publishing. ISBN 978-3-030-01252-6.
  22. 22.David Lopez-Paz and Marc' Aurelio Ranzato. Gradient episodic memory for continual learning. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/f87522788a2be2d171666752f97ddebb-Paper.pdf.
  23. 23.Jihwan Bang, Heesu Kim, YoungJoon Yoo, Jung-Woo Ha, and Jonghyun Choi. Rainbow memory: Continual learning with a memory of diverse samples. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8214–8223, 2021. doi: 10.1109/CVPR46437.2021.00812.
  24. 24.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, G. Sperl, and Christoph H. Lampert. icarl: Incremental classifier and representation learning. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5533–5542, 2016. URL https://api.semanticscholar.org/CorpusID:206596260.
  25. 25.Guanxiong Zeng, Yang Chen, Bo Cui, and Shan Yu. Continual learning of context-dependent processing in neural networks. Nature Machine Intelligence, 1:364 – 372, 2018. URL https://api.semanticscholar.org/CorpusID:52908642.
  26. 26.Khurram Javed and Martha White. Meta-learning representations for continual learning. Curran Associates Inc., Red Hook, NY, USA, 2019.
  27. 27.Hao Liu and Huaping Liu. Continual learning with recursive gradient optimization. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=7YDLgf9_zgm.
  28. 28.Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. A comprehensive survey of continual learning: Theory, method and application. IEEE transactions on pattern analysis and machine intelligence, PP, February 2024. ISSN 0162-8828. doi: 10.1109/tpami.2024.3367329. URL https://arxiv.org/pdf/2302.00487.
  29. 29.Hippolyt Ritter, Aleksandar Botev, and David Barber. Online structured laplace approximations for overcoming catastrophic forgetting. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, page 3742–3752, Red Hook, NY, USA, 2018. Curran Associates Inc.
  30. 30.Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. Progress & compress: A scalable framework for continual learning. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 4528–4537. PMLR, 10–15 Jul 2018. URL https://proceedings.mlr.press/v80/schwarz18a.html.
  31. 31.Michalis K. Titsias, Jonathan Schwarz, Alexander G. de G. Matthews, Razvan Pascanu, and Yee Whye Teh. Functional regularisation for continual learning with gaussian processes. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=HkxCzeHFDB.
  32. 32.Pingbo Pan, Siddharth Swaroop, Alexander Immer, Runa Eschenhagen, Richard E. Turner, and Mohammad Emtiyaz Khan. Continual deep learning by functional regularisation of memorable past. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546.
  33. 33.Noel Loo, Siddharth Swaroop, and Richard E Turner. Generalized variational continual learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=_IM-AfFhna9.
  34. 34.Tim G. J. Rudner, Freddie Bickford Smith, Qixuan Feng, Yee Whye Teh, and Yarin Gal. Continual learning via sequential function-space variational inference. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 18871–18887. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/rudner22a.html.
  35. 35.Noel Loo, Siddharth Swaroop, and Richard E Turner. Combining variational continual learning with fiLM layers. In 4th Lifelong Machine Learning Workshop at ICML 2020, 2020. URL https://openreview.net/forum?id=fZBEGA1d-4Y.
  36. 36.Liu Guimeng, Guo Yang, Cheryl Wong Sze Yin, Ponnuthurai Nagartnam Suganathan, and Ramasamy Savitha. Unsupervised generative variational continual learning. In 2022 IEEE International Conference on Image Processing (ICIP), pages 4028–4032, 2022. doi: 10.1109/ICIP46576.2022.9897538.
  37. 37.Hanna Tseran. Natural variational continual learning. 2018. URL https://api.semanticscholar.org/CorpusID:155098533.
  38. 38.Sayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, and Marcus Rohrbach. Uncertainty-guided continual learning with bayesian neural networks. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=HklUCCVKDB.
  39. 39.Jeevan Thapa and Rui Li. Bayesian adaptation of network depth and width for continual learning. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2025.
  40. 40.Djohan Bonnet, Kellian Cottart, Tifenn Hirtzlin, Tarcisius Januel, Thomas Dalgaty, Elisa Vianello, and Damien Querlioz. Bayesian continual learning and forgetting in neural networks, 2025. URL https://arxiv.org/abs/2504.13569.
  41. 41.Sayantan Auddy, Jakob Hollenstein, and Matteo Saveriano. Can expressive posterior approximations improve variational continual learning? Workshop on Lifelong Learning for Long-term Human-Robot Interaction of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020.
  42. 42.Yang Yang, Bo Chen, and Hongwei Liu. Memorized variational continual learning for dirichlet process mixtures. IEEE Access, 7:150851–150862, 2019. doi: 10.1109/ACCESS.2019.2947722.
  43. 43.Hongjoon Ahn, Sungmin Cha, Donggyu Lee, and Taesup Moon. Uncertainty-based continual learning with adaptive regularization. Curran Associates Inc., Red Hook, NY, USA, 2019.
  44. 44.Chen Zeno, Itay Golan, Elad Hoffer, and Daniel Soudry. Task-agnostic continual learning using online variational bayes with fixed-point updates. Neural Computation, 33(11):3139–3177, 10 2021. ISSN 0899-7667. doi: 10.1162/neco_a_01430. URL https://doi.org/10.1162/neco_a_01430.
  45. 45.Zoubin Ghahramani and H. Attias. Online variational bayesian learning. In NeurIPS Workshop on Online Learning, NeurIPS, 2000.
  46. 46.Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA, 2018. ISBN 0262039249.
  47. 47.Diederik P. Kingma and Max Welling. Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014.
  48. 48.Brian Trippe and Richard Turner. Overpruning in variational bayesian neural networks, 2018.
  49. 49.Abhishek Kumar, Sunabha Chatterjee, and Piyush Rai. Bayesian structural adaptation for continual learning. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5850–5860. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/kumar21a.html.
  50. 50.Tatsuya Konishi, Mori Kurokawa, Chihiro Ono, Zixuan Ke, Gyuhak Kim, and Bing Liu. Parameter-level soft-masking for continual learning. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  51. 51.Alex Krizhevsky. Learning multiple layers of features from tiny images. In Technical Report, University of Toronto, 2009. URL http://www.cs.toronto.edu/~kriz/learning-features-2009-TR.pdf.
  52. 52.Adam X. Yang, Maxime Robeyns, Xi Wang, and Laurence Aitchison. Bayesian low-rank adaptation for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=FJiUyzOF1m.
  53. 53.Vikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, and Benjamin Van Roy. Efficient exploration for LLMs. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=PpPZ6W7rxy.
  54. 54.Chelsea Finn, Kelvin Xu, and Sergey Levine. Probabilistic model-agnostic meta-learning. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, page 9537–9548, Red Hook, NY, USA, 2018. Curran Associates Inc.
  55. 55.Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson. Varibad: A very good method for bayes-adaptive deep rl via meta-learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=Hkl9JlBYvr.
  56. 56.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Commun. ACM, 60(6):84–90, May 2017. ISSN 0001-0782. doi: 10.1145/3065386. URL https://doi.org/10.1145/3065386.
  57. 57.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), San Diego, CA, USA, 2015.
  58. 58.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. doi: 10.1109/CVPR.2009.5206848.
  59. 59.Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams. Scalable bayesian optimization using deep neural networks. In Francis Bach and David Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 2171–2180, Lille, France, 07–09 Jul 2015. PMLR. URL https://proceedings.mlr.press/v37/snoek15.html.

Citation

MLA
Melo, L. C., et al. “Temporal-Difference Variational Continual Learning”. arXiv, 2024, http://arxiv.org/abs/2410.07812v4.
APA
Melo, L. C., Abate, A., & Gal, Y. (2024). Temporal-Difference Variational Continual Learning. arXiv. http://arxiv.org/abs/2410.07812v4
Chicago
Melo, L. C., A. Abate, and Y. Gal. 2024. “Temporal-Difference Variational Continual Learning”. arXiv. http://arxiv.org/abs/2410.07812v4.
Harvard
Melo, L.C., Abate, A. and Gal, Y. (2024) “Temporal-Difference Variational Continual Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.07812v4.
Vancouver
1. Melo LC, Abate A, Gal Y (2024) Temporal-Difference Variational Continual Learning. arXiv

BibTeX

@article{melo2024temporal,
  title = {Temporal-Difference Variational Continual Learning},
  author = {Melo, Luckeciano C. and Abate, Alessandro and Gal, Yarin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.07812v4},
  eprint = {2410.07812}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/