keyword
online continual learning
Online continual learning is a machine learning paradigm in which a model continuously learns from a non-stationary stream of incoming data in real time, typically processing each data sample only once under strict computational and memory constraints. Unlike standard offline machine learning that trains over multiple iterations across a static, fully accessible dataset, online continual learning requires an agent to immediately adapt to shifting distributions and new classes as they appear sequentially. The primary challenge of this setting is balancing plasticity, which is the ability to integrate new information, with stability, which is the retention of previously acquired knowledge without experiencing catastrophic forgetting. To achieve this balance efficiently, approaches in online continual learning frequently employ mechanisms such as experience replay buffers, architectural growth, or regularization techniques designed to preserve learned representations during streaming parameter updates.
2 items

New Insights on Reducing Abrupt Representation Change in Online Continual Learning
Lucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars, Joelle Pineau, Eugene Belilovsky
Why you should read this
Demonstrates how Experience Replay causes disruptive representation shifts when new classes appear in online continual learning, and resolves this with an asymmetric update rule that forces incoming data to adapt to established representations.
In the online continual learning paradigm, agents must learn from a changing distribution while respecting memory and compute constraints. Experience Replay (ER), where a small subset of past data is stored and replayed alongside new data, has emerged as a simple and effective learning strategy. In this work, we focus on the change in representations of observed data that arises when previously unobserved classes appear in the incoming data stream, and new classes must be distinguished from previous ones. We shed new light on this question by showing that applying ER causes the newly added classes' representations to overlap significantly with the previous classes, leading to highly disruptive parameter updates. Based on this empirical analysis, we propose a new method which mitigates this issue by shielding the learned representations from drastic adaptation to accommodate new classes. We show that using an asymmetric update rule pushes new classes to adapt to the older ones (rather than the reverse), which is more effective especially at task boundaries, where much of the forgetting typically occurs. Empirical results show significant gains over strong baselines on standard continual learning benchmarks.
Added
2026-09-26

Self-Supervised Models are Continual Learners
Enrico Fini, Victor G. Turrisi da Costa, Xavier Alameda-Pineda, Elisa Ricci, Karteek Alahari, Julien Mairal
Why you should read this
Proposes CaSSLe, a general framework that mitigates catastrophic forgetting in continual self-supervised learning by using a predictor network to convert standard self-supervised loss functions into representation distillation mechanisms without requiring extra hyperparameter tuning.
Self-supervised models have been shown to produce comparable or better visual representations than their supervised counterparts when trained offline on unlabeled data at scale. However, their efficacy is catastrophically reduced in a Continual Learning (CL) scenario where data is presented to the model sequentially. In this paper, we show that self-supervised loss functions can be seamlessly converted into distillation mechanisms for CL by adding a predictor network that maps the current state of the representations to their past state. This enables us to devise a framework for Continual self-supervised visual representation Learning that (i) significantly improves the quality of the learned representations, (ii) is compatible with several state-of-the-art self-supervised objectives, and (iii) needs little to no hyperparameter tuning. We demonstrate the effectiveness of our approach empirically by training six popular self-supervised models in various CL settings. Code: github.com/DonkeyShot21/cassle.
Added
2026-09-26
