Learning a Unified Classifier Incrementally via Rebalancing
Saihui HouXinyu PanChen Change LoyZilei WangDahua Lin
Proposes a unified multi-class incremental learning framework that mitigates catastrophic forgetting by combining cosine normalization, a less-forget feature constraint, and margin-based inter-class separation to correct the data imbalance between old and new classes.
Modern computer vision and artificial intelligence systems frequently encounter continuous streams of new information. Updating these models incrementally is essential to avoid retraining from scratch, which is computationally expensive and impractical. However, existing systems suffer from severe forgetting when exposed to new categories. The article identifies that this performance drop is primarily caused by an extreme data imbalance: models are trained on abundant new data while retaining only a tiny fraction of previous data. This imbalance inflates the influence of new categories, degrades previously acquired knowledge, and causes confusion between old and new classes.
The article develops and evaluates a rebalanced machine learning framework designed to train a single, unified classifier across multiple incremental phases without forgetting past categories.
The researchers evaluated their framework against standard industry benchmarks across two major image datasets: CIFAR-100 and ImageNet (both a 100-class subset and the full 1,000-class collection). The models started with half of the total categories and incrementally incorporated the remaining categories across sequences of one, two, five, and ten phases while storing only twenty representative samples per past category. The proposed architecture integrated three primary techniques: cosine normalization to equalize prediction scales, a geometric feature constraint to preserve past representations, and an inter-class margin loss to enforce clear boundaries between old and new categories.
The evaluation produced several clear findings. First, the proposed framework substantially outperformed existing state-of-the-art methods, reducing classification errors by over 6% on CIFAR-100 and by over 13% on full ImageNet over ten incremental phases. Second, error analysis confirmed that predictions remained balanced across both past and recent classes, eliminating the bias toward newer data that plagued earlier models. Third, standard neural network predictions directly matched or exceeded prototype-based nearest-mean classification, simplifying the overall system design. Fourth, ablation studies demonstrated that all three mathematical components collectively drove these performance gains, with an adaptive weighting mechanism playing a critical role in multi-phase stability.
These results demonstrate that catastrophic forgetting in unified classifiers is largely a data-imbalance issue that can be resolved geometrically rather than a structural flaw in neural networks. For organizations deploying machine learning systems, this approach lowers operational costs and update latency by removing the requirement to store legacy training data or retrain entire models from the ground up.
Organizations operating continuous-learning systems should consider adopting cosine normalization and margin-based separation into their updating pipelines. Before full deployment, teams should conduct internal pilot tests tailored to their operational constraints, noting trade-offs between memory budgets (the number of stored legacy examples) and accuracy requirements. Further development should explore optimal methods to automatically tune distillation weights across extended update cycles.
Confidence in these findings is strong given the rigorous multi-dataset validation and substantial performance margins over baselines. Nevertheless, stakeholders should note that the evaluation was bounded by image classification tasks with a fixed number of stored samples per class, meaning performance should be verified on non-vision modalities and variable-memory environments.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL establishes the foundational exemplar-based class-incremental learning framework that the source paper directly builds upon and rebalances to avoid catastrophic forgetting.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Learning without Forgetting introduces the distillation-loss paradigm used across incremental learning models, providing the baseline mechanism modified by the source's less-forget constraint.
- Paper: Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence, Arslan Chaudhry et al. (2018). RWalk formalizes the trade-offs of catastrophic forgetting and intransigence in sequential single-head classification, framing the problem space targeted by unified incremental classifiers.
- Paper: A systematic study of the class imbalance problem in convolutional neural networks, Mateusz Buda et al. (2017). This systematic study analyzes how class imbalance impairs deep convolutional networks, establishing the theoretical rationale for rebalancing data distributions during incremental phases.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Elastic Weight Consolidation is a seminal baseline for mitigating catastrophic forgetting in neural networks by regularizing parameter drift on subsequent tasks.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Synaptic Intelligence provides foundational insight into online parameter-importance regularization for sequential task learning.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Gradient Episodic Memory defines key continual learning benchmarks and gradient-based exemplar constraints for managing historical data in sequential class learning.
- Paper: Large Scale Incremental Learning, Yue Wu et al. (2019). This work directly complements the source by diagnosing the classifier layer bias caused by data imbalance and proposing a post-hoc linear correction layer.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). DER++ extends replay-based continual learning by retaining intermediate logit trajectories under continuous domain and class shifts without task boundary knowledge.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). This survey provides a comprehensive taxonomy covering modern regularization, replay, and rebalancing strategies developed after the initial class-incremental learning breakthroughs.
- Paper: Decoupling Representation and Classifier for Long-Tailed Recognition, Bingyi Kang et al. (2019). This paper generalizes the decoupling of representation learning and classifier balancing investigated in class-incremental rebalancing methods to broad long-tailed recognition tasks.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This comprehensive survey categorizes and benchmarks subsequent continual learning methodologies across balanced and unbalanced data streams.
- Paper: Self-Distillation Enables Continual Learning, Idan Shenfeld et al. (2026). This work explores self-distillation fine-tuning to prevent catastrophic forgetting when incrementally adapting modern foundation models.
