keyword
class-incremental learning
Class-incremental learning is a machine learning paradigm in which an artificial intelligence model sequentially learns to recognize new object or concept categories over time while retaining its ability to classify previously learned categories. Unlike traditional training where all training data is available simultaneously, or task-incremental settings where a task identifier is provided during evaluation, a class-incremental system must classify inputs across all observed classes within a single unified prediction space without knowing which learning phase an instance belongs to. The primary challenge in this setting is mitigating catastrophic forgetting, the phenomenon where learning novel classes severely degrades performance on older ones. To balance adapting to new categories with preserving prior knowledge, approaches typically employ strategies such as storing compressed exemplar samples, regularizing parameter updates, utilizing knowledge distillation, or dynamically expanding network architectures.
8 items

Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models
Didi Zhu, Zhongyi Sun, Zexi Li, Tao Shen, Ke Yan, Shouhong Ding, Chao Wu, Kun Kuang
Why you should read this
Proposes Model Tailor, a parameter-efficient post-training method that updates fewer than ten percent of model parameters through sparse masking and Hessian-based compensation, preventing catastrophic forgetting in multi-modal large language models while maintaining performance on both original and target tasks.
Catastrophic forgetting emerges as a critical challenge when fine-tuning multi-modal large language models (MLLMs), where improving performance on target tasks often leads to a significant performance drop on the original tasks. This paper presents a comprehensive analysis of catastrophic forgetting in MLLMs and introduces a post-training adjustment method called Model Tailor. Our method primarily preserves the pre-trained parameters while replacing a small number (≤ 10%) of fine-tuned parameters, maintaining ~ 99% effectiveness on original tasks versus pre-training, and achieving ~ 97% on new tasks compared to standard fine-tuning. Specifically, we derive a sparse mask to identify the “model patch”, based on a fusion strategy that integrates salience and sensitivity analysis. Subsequently, a compensation mechanism is introduced to “decorate the patch”, enhancing the model’s performance on both target and original tasks. Additionally, our method is adaptable to multi-task scenarios. Through extensive experiments on Instruct-BLIP and LLaVA-1.5 in both image captioning and visual question answering tasks, our approach demonstrates significant task adaptability while preserving inherent pre-trained capabilities.
Added
2026-10-03

Dense Network Expansion for Class Incremental Learning
Zhiyuan Hu, Yunsheng Li, Jiancheng Lyu, Dashan Gao, Nuno Vasconcelos
Why you should read this
Proposes a dense network expansion method for class incremental learning that transfers knowledge across frozen task experts using a cross-task attention block, achieving superior accuracy while significantly curbing parameter growth.
The problem of class incremental learning (CIL) is considered. State-of-the-art approaches use a dynamic architecture based on network expansion (NE), in which a task expert is added per task. While effective from a computational standpoint, these methods lead to models that grow quickly with the number of tasks. A new NE method, dense network expansion (DNE), is proposed to achieve a better trade-off between accuracy and model complexity. This is accomplished by the introduction of dense connections between the intermediate layers of the task expert networks, that enable the transfer of knowledge from old to new tasks via feature sharing and reusing. This sharing is implemented with a cross-task attention mechanism, based on a new task attention block (TAB), that fuses information across tasks. Unlike traditional attention mechanisms, TAB operates at the level of the feature mixing and is decoupled with spatial attentions. This is shown more effective than a joint spatial-and-task attention for CIL. The proposed DNE approach can strictly maintain the feature space of old classes while growing the network and feature scale at a much slower rate than previous methods. In result, it outperforms the previous SOTA methods by a margin of 4% in terms of accuracy, with similar or even smaller model scale.
Added
2026-09-26

Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning
Na Dong, Yongqiang Zhang, Mingli Ding, Gim Hee Lee
Why you should read this
Presents a DETR-based framework for incremental few-shot object detection that prevents catastrophic forgetting and overfitting by combining self-supervised fine-tuning on pseudo-labeled proposals with selective knowledge distillation on class-specific layers.
Incremental few-shot object detection aims at detecting novel classes without forgetting knowledge of the base classes with only a few labeled training data from the novel classes. Most related prior works are on incremental object detection that rely on the availability of abundant training samples per novel class that substantially limits the scalability to real-world setting where novel data can be scarce. In this paper, we propose the Incremental-DETR that does incremental few-shot object detection via fine-tuning and self-supervised learning on the DETR object detector. To alleviate severe over-fitting with few novel class data, we first fine-tune the class-specific components of DETR with self-supervision from additional object proposals generated using Selective Search as pseudo labels. We further introduce an incremental few-shot fine-tuning strategy with knowledge distillation on the class-specific components of DETR to encourage the network in detecting novel classes without forgetting the base classes. Extensive experiments conducted on standard incremental object detection and incremental few-shot object detection settings show that our approach significantly outperforms state-of-the-art methods by a large margin. Our source code is available at https://github.com/dongnana777/Incremental-DETR.
Added
2026-09-26

Class-Incremental Exemplar Compression for Class-Incremental Learning
Zilin Luo, Yaoyao Liu, Bernt Schiele, Qianru Sun
Why you should read this
Proposes an adaptive masking strategy optimized through bilevel learning that selectively compresses non-discriminative background pixels in exemplars, enabling memory-constrained class-incremental learning systems to store substantially more training samples and improve classification accuracy across benchmarks.
Exemplar-based class-incremental learning (CIL) [36] finetunes the model with all samples of new classes but few-shot exemplars of old classes in each incremental phase, where the “few-shot” abides by the limited memory budget. In this paper, we break this “few-shot” limit based on a simple yet surprisingly effective idea: compressing exemplars by downsampling non-discriminative pixels and saving “many-shot” compressed exemplars in the memory. Without needing any manual annotation, we achieve this compression by generating 0-1 masks on discriminative pixels from class activation maps (CAM) [49]. We propose an adaptive mask generation model called class-incremental masking (CIM) to explicitly resolve two difficulties of using CAM: 1) transforming the heatmaps of CAM to 0-1 masks with an arbitrary threshold leads to a trade-off between the coverage on discriminative pixels and the quantity of exemplars, as the total memory is fixed; and 2) optimal thresholds vary for different object classes, which is particularly obvious in the dynamic environment of CIL. We optimize the CIM model alternatively with the conventional CIL model through a bilevel optimization problem [40]. We conduct extensive experiments on high-resolution CIL benchmarks including Food-101, ImageNet-100, and ImageNet-1000, and show that using the compressed exemplars by CIM can achieve a new state-of-the-art CIL accuracy, e.g., 4.8 percentage points higher than FOSTER [42] on 10-phase ImageNet-1000. Our code is available at https://github.com/xflz/CIM-CIL.
Added
2026-09-26

Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task
Stan Weixian Lei, Difei Gao, Jay Zhangjie Wu, Yuxuan Wang, Wei Liu, Mengmi Zhang, Mike Zheng Shou
Why you should read this
Proposes a real-data-free continual learning framework that replays scene graphs instead of images to prevent catastrophic forgetting in visual question answering models across evolving visual environments and question types.
Added
2026-09-26

Endpoints Weight Fusion for Class Incremental Semantic Segmentation
Jia-Wen Xiao, Chang-Bin Zhang, Jiekang Feng, Xialei Liu, Joost van de Weijer, Ming-Ming Cheng
Why you should read this
Proposes a dynamic parameter-averaging strategy that fuses starting and ending network weights from each incremental step to mitigate catastrophic forgetting in semantic segmentation without increasing model size or requiring extra training.
Class incremental semantic segmentation (CISS) focuses on alleviating catastrophic forgetting to improve discrimination. Previous work mainly exploits regularization (e.g., knowledge distillation) to maintain previous knowledge in the current model. However, distillation alone often yields limited gain to the model since only the representations of old and new models are restricted to be consistent. In this paper, we propose a simple yet effective method to obtain a model with a strong memory of old knowledge, named Endpoints Weight Fusion (EWF). In our method, the model containing old knowledge is fused with the model retaining new knowledge in a dynamic fusion manner, strengthening the memory of old classes in ever-changing distributions. In addition, we analyze the relationship between our fusion strategy and a popular moving average technique EMA, which reveals why our method is more suitable for class-incremental learning. To facilitate parameter fusion with closer distance in the parameter space, we use distillation to enhance the optimization process. Furthermore, we conduct experiments on two widely used datasets, achieving state-of-the-art performance.
Added
2026-09-26

A Continual Learning Survey: Defying Forgetting in Classification Tasks
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, A. Leonardis, G. Slabaugh, T. Tuytelaars
Why you should read this
Presents a structured taxonomy and a dynamic hyperparameter selection framework to benchmark eleven continual learning methods across diverse classification datasets, showing how model capacity, regularization, and parameter isolation directly affect catastrophic forgetting.
Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods and 4 baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time and storage.
Added
2026-09-14

iCaRL: Incremental Classifier and Representation Learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, Christoph H. Lampert
Why you should read this
Proposes iCaRL, a class-incremental learning method that simultaneously trains deep feature representations and classifiers on streaming data to prevent catastrophic forgetting over consecutive image classification tasks.
A major open problem on the road to artificial intelligence is the development of incrementally learning systems that learn about more and more concepts over time from a stream of data. In this work, we introduce a new training strategy, iCaRL, that allows learning in such a class-incremental way: only the training data for a small number of classes has to be present at the same time and new classes can be added progressively. iCaRL learns strong classifiers and a data representation simultaneously. This distinguishes it from earlier works that were fundamentally limited to fixed data representations and therefore incompatible with deep learning architectures. We show by experiments on CIFAR-100 and ImageNet ILSVRC 2012 data that iCaRL can learn many classes incrementally over a long period of time where other strategies quickly fail.
Added
2026-09-10
