Large Scale Incremental Learning
Yue WuYinpeng ChenLijuan WangYuancheng YeZicheng LiuYandong GuoYun Fu
Reveals that final-layer classification bias causes catastrophic forgetting in large-scale incremental learning and introduces a two-parameter linear correction that outperforms state-of-the-art methods by over 11% on ImageNet and MS-Celeb-1M.
Real-world artificial intelligence applications, such as large-scale facial recognition and visual classification systems, must continuously learn new categories over time without forgetting previously acquired knowledge. Deep neural networks typically suffer from catastrophic forgetting when past training data cannot be retained in full. While storing a small set of representative sample images from earlier categories can mitigate this issue in small setups, existing approaches experience severe accuracy degradation when scaled up to thousands of categories with visually similar classes.
The article evaluates the root causes of this scalability bottleneck and demonstrates a novel two-stage method, named Bias Correction (BiC), to correct classification errors caused by extreme data imbalance in large-scale incremental learning.
The researchers diagnosed that the primary failure in existing incremental models stems from the final classification layer, which develops a heavy systematic bias favoring newly introduced categories over data-constrained older categories. To address this, the article introduces a two-stage training strategy: the main neural network is first trained on the combined old and new image data using knowledge distillation, after which the feature representations are frozen. A simple linear correction layer with only two parameters is then optimized on a small, balanced validation subset of old and new samples to adjust the outputs of the new categories.
The empirical findings demonstrate that this lightweight correction delivers substantial performance gains, particularly at large scales. On the 1,000-class ImageNet benchmark across 10 incremental steps, the proposed method surpassed existing state-of-the-art techniques by an average of 11.1%, beating the next best method by 18.5% at the final stage. On a 10,000-class facial recognition benchmark (MS-Celeb-1M), the method outperformed prior leading approaches by an average of 13.2% and achieved an 87.98% final accuracy compared to 65.56% for the previous baseline. Furthermore, ablation experiments revealed that a 9:1 training-to-validation split on stored exemplars is optimal, and that the method remains robust regardless of whether older exemplars are selected randomly or using complex representative selection algorithms.
These results show that large-scale incremental learning can be made viable without requiring extensive retraining on historical data archives, directly reducing the computational expense, memory footprint, and storage overhead of maintaining continuous learning systems. Correcting output layer bias offers a high-impact, low-complexity solution that substantially bridges the performance gap between continuously updated models and ideal models trained on all historical data at once.
Organizations deploying large-scale visual recognition models should consider incorporating a two-stage bias correction step into their continual update pipelines. When adopting this method, engineering teams should allocate roughly 10% of their retained exemplar quota specifically for bias validation rather than model feature learning. Additionally, combining this method with standard data augmentation techniques could yield further gains during early deployment stages when class imbalances are mild.
Confidence in these findings is high given the rigorous evaluation across standard large-scale benchmarks. However, a minor performance gap remains between the proposed approach and an idealized model retrained entirely from scratch, as the method primarily targets classifier layer bias rather than residual drift within deeper feature extraction layers. Further analysis on non-visual domains and real-time streaming data represents an appropriate next step before full enterprise deployment across non-vision workloads.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). Introduces the exemplar replay and distillation framework for class-incremental learning that the source identifies as having scaling and class imbalance limitations.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Establishes distillation loss as a foundational mechanism to prevent catastrophic forgetting when incrementally learning new vision tasks.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Provides fundamental analysis and benchmark formulations for class-incremental learning using episodic memory of old tasks.
- Paper: A systematic study of the class imbalance problem in convolutional neural networks, Mateusz Buda et al. (2017). Provides a systematic empirical study on how extreme class imbalance degrades convolutional neural network representations and classifier decision boundaries.
- Paper: Decoupling Representation and Classifier for Long-Tailed Recognition, Bingyi Kang et al. (2019). Extends the principle of correcting classifier bias under data imbalance by decoupling representation learning from classifier adjustment across large-scale benchmarks.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). Provides a comprehensive retrospective survey synthesizing replay-based methods and classifier bias correction strategies across the broader continual learning landscape.
- Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). Builds on knowledge distillation under class distributions by decoupling target and non-target distillation terms to improve knowledge transfer.
