End-to-End Incremental Learning
Francisco M. CastroManuel J. Marín-JiménezNicolás GuilCordelia SchmidKarteek Alahari
Proposes a fully end-to-end incremental learning framework that mitigates catastrophic forgetting by combining knowledge distillation with a small exemplar set to jointly optimize feature representations and classifiers on sequentially added image classes.
Practical visual recognition systems frequently need to learn new categories over time without retraining from scratch on entire historical datasets. However, standard deep learning models suffer from catastrophic forgetting, where integrating new information causes a steep drop in accuracy on previously learned categories. Storing all historical data to retrain models is computationally expensive and quickly becomes unsustainable as the number of classes grows.
The main objective of the article is to demonstrate an end-to-end deep learning framework that learns new image classes incrementally while preserving past knowledge using only a minimal set of preserved reference images.
To accomplish this, the authors developed a convolutional neural network architecture trained using a combined objective called cross-distilled loss. This combines a cross-entropy loss to learn incoming categories with a distillation loss that transfers and locks in knowledge from older categories. The system retains a small representative memory of past categories selected via a mean-distance ranking method (herding). The process incorporates data augmentation and a balanced fine-tuning stage to correct for the numerical imbalance between new data and limited historical exemplars. The method was evaluated using ResNet architectures on standard image classification benchmarks (CIFAR-100 and ImageNet ILSVRC 2012) across varying incremental step sizes and memory configurations.
The framework achieved state-of-the-art results across major benchmarks. On CIFAR-100 with a fixed 2,000-exemplar memory, it reached an average incremental accuracy of roughly 63.4% to 63.8% across 2-to-20 class steps, significantly outperforming prior methods like iCaRL (54.1% to 62.0%) and LwF.MC (9.6% to 47.6%). On large-scale ImageNet benchmarks, the method attained top-5 accuracies of 90.4% (for 10-class steps) and 69.4% (for 100-class steps), beating previous techniques by over 5 percentage points. Furthermore, on fine-grained recognition tests with highly similar classes (such as distinguishing vehicle types or dog breeds), the end-to-end framework outperformed competing baselines by roughly 10 to 25 percentage points.
These findings demonstrate that jointly learning feature representations and classifiers avoids the performance bottlenecks of decoupled, prototype-based classifiers. By retaining only a small representative memory and adjusting for data imbalances during fine-tuning, organizations can continually expand visual recognition systems at stable memory footprints and reduced training overhead, making lifelong machine learning practical in real-world deployment.
Teams implementing incremental computer vision systems should adopt joint end-to-end training paired with distillation loss and balanced fine-tuning rather than freezing network layers or relying on external nearest-mean classifiers. Systems should utilize herding selection to manage exemplar storage. As next steps, the authors recommend exploring dynamic exemplar allocation to further optimize memory efficiency across classes.
The primary limitation of the approach arises when the number of retained past samples is severely restricted relative to large batches of new classes (e.g., adding 50 classes at once with a small fixed memory), which can degrade performance due to extreme data imbalance. Confidence in these results remains high across standard and fine-grained classification tasks, though practitioners should remain cautious in extreme class-imbalance scenarios without adequate fine-tuning.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). It introduces the exemplar-based class-incremental learning framework with distillation and classification losses that this work directly builds upon and turns into an end-to-end architecture.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). It establishes the foundational distillation loss strategy for preserving prior-task capabilities in neural networks without retraining on the full original dataset.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). It formalizes the replay-based continual learning paradigm and benchmark setups on CIFAR-100 that contextualize memory-constrained incremental training.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). It provides the essential foundational baseline for parameter-regularization approaches to catastrophic forgetting in sequential neural network learning.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). It introduces online synapse-importance estimation to protect prior knowledge during sequential task training, serving as key background in continual learning.
- Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). It establishes generative replay mechanisms as an alternative replay paradigm for mitigating catastrophic forgetting in sequential image classification.
- Paper: Large Scale Incremental Learning, Yue Wu et al. (2019). It directly tackles the output classification bias toward new classes that arises in exemplar-based incremental learning methods like this paper's end-to-end framework.
- Paper: Learning a Unified Classifier Incrementally via Rebalancing, Saihui Hou et al. (2019). It resolves the severe data imbalance between new and old classes in multi-phase class-incremental learning via cosine normalization and geometric constraints.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). It extends experience replay and distillation principles to boundary-free continual learning by matching continuous logit outputs alongside buffer replay.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). It provides a unified comparative benchmark and taxonomy evaluating class-incremental and task-incremental methods across standard vision datasets.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). It advances class-incremental learning to exemplar-free continual learning by learning prompt pools on top of frozen vision transformers.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It provides a broad theoretical and methodological survey synthesizing distillation, replay, and representation techniques across the continual learning landscape.
