keyword
replay-based methods
Replay-based methods are continual learning techniques designed to prevent catastrophic forgetting by reintroducing past data to a machine learning model while it trains on new information. In sequential training environments where data distributions shift over time, these approaches preserve knowledge by maintaining a limited memory buffer of stored historical exemplars or by synthesizing approximations of previous data through generative models. During the training process on new tasks or incoming data streams, the stored or generated past examples are interleaved with current data samples, allowing the model to optimize its parameters across both historical and new information simultaneously to maintain stable representations and decision boundaries under bounded memory constraints.
3 items

New Insights on Reducing Abrupt Representation Change in Online Continual Learning
Lucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars, Joelle Pineau, Eugene Belilovsky
Why you should read this
Demonstrates how Experience Replay causes disruptive representation shifts when new classes appear in online continual learning, and resolves this with an asymmetric update rule that forces incoming data to adapt to established representations.
In the online continual learning paradigm, agents must learn from a changing distribution while respecting memory and compute constraints. Experience Replay (ER), where a small subset of past data is stored and replayed alongside new data, has emerged as a simple and effective learning strategy. In this work, we focus on the change in representations of observed data that arises when previously unobserved classes appear in the incoming data stream, and new classes must be distinguished from previous ones. We shed new light on this question by showing that applying ER causes the newly added classes' representations to overlap significantly with the previous classes, leading to highly disruptive parameter updates. Based on this empirical analysis, we propose a new method which mitigates this issue by shielding the learned representations from drastic adaptation to accommodate new classes. We show that using an asymmetric update rule pushes new classes to adapt to the older ones (rather than the reverse), which is more effective especially at task boundaries, where much of the forgetting typically occurs. Empirical results show significant gains over strong baselines on standard continual learning benchmarks.
Added
2026-09-26

Self-Supervised Models are Continual Learners
Enrico Fini, Victor G. Turrisi da Costa, Xavier Alameda-Pineda, Elisa Ricci, Karteek Alahari, Julien Mairal
Why you should read this
Proposes CaSSLe, a general framework that mitigates catastrophic forgetting in continual self-supervised learning by using a predictor network to convert standard self-supervised loss functions into representation distillation mechanisms without requiring extra hyperparameter tuning.
Self-supervised models have been shown to produce comparable or better visual representations than their supervised counterparts when trained offline on unlabeled data at scale. However, their efficacy is catastrophically reduced in a Continual Learning (CL) scenario where data is presented to the model sequentially. In this paper, we show that self-supervised loss functions can be seamlessly converted into distillation mechanisms for CL by adding a predictor network that maps the current state of the representations to their past state. This enables us to devise a framework for Continual self-supervised visual representation Learning that (i) significantly improves the quality of the learned representations, (ii) is compatible with several state-of-the-art self-supervised objectives, and (iii) needs little to no hyperparameter tuning. We demonstrate the effectiveness of our approach empirically by training six popular self-supervised models in various CL settings. Code: github.com/DonkeyShot21/cassle.
Added
2026-09-26

Class-Incremental Exemplar Compression for Class-Incremental Learning
Zilin Luo, Yaoyao Liu, Bernt Schiele, Qianru Sun
Why you should read this
Proposes an adaptive masking strategy optimized through bilevel learning that selectively compresses non-discriminative background pixels in exemplars, enabling memory-constrained class-incremental learning systems to store substantially more training samples and improve classification accuracy across benchmarks.
Exemplar-based class-incremental learning (CIL) [36] finetunes the model with all samples of new classes but few-shot exemplars of old classes in each incremental phase, where the “few-shot” abides by the limited memory budget. In this paper, we break this “few-shot” limit based on a simple yet surprisingly effective idea: compressing exemplars by downsampling non-discriminative pixels and saving “many-shot” compressed exemplars in the memory. Without needing any manual annotation, we achieve this compression by generating 0-1 masks on discriminative pixels from class activation maps (CAM) [49]. We propose an adaptive mask generation model called class-incremental masking (CIM) to explicitly resolve two difficulties of using CAM: 1) transforming the heatmaps of CAM to 0-1 masks with an arbitrary threshold leads to a trade-off between the coverage on discriminative pixels and the quantity of exemplars, as the total memory is fixed; and 2) optimal thresholds vary for different object classes, which is particularly obvious in the dynamic environment of CIL. We optimize the CIM model alternatively with the conventional CIL model through a bilevel optimization problem [40]. We conduct extensive experiments on high-resolution CIL benchmarks including Food-101, ImageNet-100, and ImageNet-1000, and show that using the compressed exemplars by CIM can achieve a new state-of-the-art CIL accuracy, e.g., 4.8 percentage points higher than FOSTER [42] on 10-phase ImageNet-1000. Our code is available at https://github.com/xflz/CIM-CIL.
Added
2026-09-26
