keyword
representation forgetting
Representation forgetting is the loss or degradation of the quality of intermediate feature representations learned by a neural network for previous tasks as it is sequentially trained on new data distributions. In continual learning, while catastrophic forgetting typically refers to the overall drop in performance on earlier tasks at the model output layer, representation forgetting specifically focuses on changes occurring within the latent embedding space. It assesses whether the underlying feature extractor retains sufficient information about past tasks, often evaluated by probing whether a linear classifier can still discriminate past classes from the updated representations. By separating the degradation of internal representations from the misalignment of final classification layers, the concept helps identify whether a model has genuinely lost the capacity to represent earlier knowledge or merely needs recalibrated decision boundaries.
2 items

Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning
Nader Asadi, MohammadReza Davari, Sudhir P. Mudur, Rahaf Aljundi, Eugene Belilovsky
Why you should read this
Proposes a replay-free continual learning framework that prevents catastrophic forgetting by distilling the relative similarities between incoming samples and past class prototypes, outperforming memory-buffer methods without storing historical data.
In Continual learning (CL) balancing effective adaptation while combating catastrophic forgetting is a central challenge. Many of the recent best-performing methods utilize various forms of prior task data, e.g. a replay buffer, to tackle the catastrophic forgetting problem. Having access to previous task data can be restrictive in many real-world scenarios, for example when task data is sensitive or proprietary. To overcome the necessity of using previous tasks' data, in this work, we start with strong representation learning methods that have been shown to be less prone to forgetting. We propose a holistic approach to jointly learn the representation and class prototypes while maintaining the relevance of old class prototypes and their embedded similarities. Specifically, samples are mapped to an embedding space where the representations are learned using a supervised contrastive loss. Class prototypes are evolved continually in the same latent space, enabling learning and prediction at any point. To continually adapt the prototypes without keeping any prior task data, we propose a novel distillation loss that constrains class prototypes to maintain relative similarities as compared to new task data. This method yields state-of-the-art performance in the task-incremental setting, outperforming methods relying on large amounts of data, and provides strong performance in the class-incremental setting without using any stored data points.
Added
2026-10-03

Probing Representation Forgetting in Supervised and Unsupervised Continual Learning
MohammadReza Davari, Nader Asadi, Sudhir P. Mudur, Rahaf Aljundi, Eugene Belilovsky
Why you should read this
Reveals through linear probing that neural networks retain substantially more past-task information during continual learning than standard accuracy metrics suggest, enabling a competitive rehearsal-free method based on supervised contrastive learning and class prototypes.
Continual Learning (CL) research typically focuses on tackling the phenomenon of catastrophic forgetting in neural networks. Catastrophic forgetting is associated with an abrupt loss of knowledge previously learned by a model when the task, or more broadly the data distribution, being trained on changes. In supervised learning problems this forgetting, resulting from a change in the model’s representation, is typically measured or observed by evaluating the decrease in old task performance. However, a model’s representation can change without losing knowledge about prior tasks. In this work we consider the concept of representation forgetting, observed by using the difference in performance of an optimal linear classifier before and after a new task is introduced. Using this tool we revisit a number of standard continual learning benchmarks and observe that, through this lens, model representations trained without any explicit control for forgetting often experience small representation forgetting and can sometimes be comparable to methods which explicitly control for forgetting, especially in longer task sequences. We also show that representation forgetting can lead to new insights on the effect of model capacity and loss function used in continual learning. Based on our results, we show that a simple yet competitive approach is to learn representations continually with standard supervised contrastive learning while constructing prototypes of class samples when queried on old samples.
Added
2026-09-26
