Class-Incremental Exemplar Compression for Class-Incremental Learning
Zilin LuoYaoyao LiuBernt SchieleQianru Sun
Proposes an adaptive masking strategy optimized through bilevel learning that selectively compresses non-discriminative background pixels in exemplars, enabling memory-constrained class-incremental learning systems to store substantially more training samples and improve classification accuracy across benchmarks.
Artificial intelligence systems deployed in dynamic real-world environments must continually learn new object classes over time while retaining previously learned knowledge. A standard approach to prevent the forgetting of earlier classes is to store a small budget of past representative images, known as exemplars, to replay during future training. However, strict memory budgets typically restrict models to keeping only a handful of exemplars per old class, creating severe data imbalance between massive new data and scarce old data, which leads to substantial performance degradation.
The article evaluates and demonstrates a novel compression approach called class-incremental masking. The main objective is to overcome the severe sample shortage of old classes by downsampling non-essential background pixels while preserving critical object features at original resolutions, enabling systems to store substantially more exemplars within the exact same memory footprint without requiring manual localization annotations.
The researchers designed an adaptive mask generation framework that leverages the model's own visual attention mechanisms, known as class activation maps, to locate discriminative features. Because optimal visual cues shift as new classes are introduced, the method embeds learnable activation functions and uses a nested, two-level optimization routine to dynamically adjust masks across incremental learning stages. The framework was evaluated across standard high-resolution benchmark datasets—including Food-101, ImageNet-100, and large-scale ImageNet-1000—integrated as a plug-in component alongside top-performing continuous learning architectures under both fixed and expanding memory conditions.
The evaluation yielded several key findings. First, the proposed approach achieved state-of-the-art accuracy across all benchmark settings, consistently outperforming leading baselines like FOSTER and DER. Second, performance gains widened significantly under tighter memory constraints; on the 1,000-class benchmark with a restrictive 5,000-sample memory budget, the method boosted average accuracy by 4.8 percentage points in a 10-phase setting. Third, the framework delivered outsized benefits over longer training horizons with more incremental phases and achieved the highest relative accuracy gains on smaller target objects, where background downsampling yields greater memory savings and allows more exemplars to be preserved.
These findings indicate that strategic, selective image downsampling effectively mitigates the trade-off between sample quality and sample diversity in memory-constrained artificial intelligence systems. Rather than uniformly degrading image resolution or relying on fragile pixel synthesis, preserving fine-grained foreground cues while discarding irrelevant background details provides a cost-effective, high-performing pathway to continuous model updating without increasing physical memory costs or infrastructure footprint.
Engineering and research teams managing continuous machine learning pipelines should evaluate this masking approach as a modular plug-in to alleviate memory bottlenecks and catastrophic forgetting. Organizations should prioritize its application in scenarios characterized by high-resolution visual inputs, long operational deployment horizons, and tight edge-device storage constraints.
The methodology possesses clear boundary conditions. It is specifically tailored for high-resolution visual data and is ineffective on very low-resolution images, where compression parameter overhead outweighs the memory saved from downsampling. Additionally, the framework cannot retroactively modify previously saved exemplars once their original training data has been discarded. Nevertheless, given the thorough empirical validation and consistent improvements across standardized benchmarks, confidence in the reported performance gains remains high for high-resolution vision applications.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). It establishes the foundational exemplar-based class-incremental learning setup that stores historical samples under a fixed memory budget using distillation and herding.
- Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). It introduces an end-to-end rehearsal framework with distillation and balanced fine-tuning for exemplar-based incremental learning.
- Paper: Large Scale Incremental Learning, Yue Wu et al. (2019). It analyzes the severe class imbalance and classifier bias inherent in exemplar-based class-incremental learning benchmarks.
- Paper: Learning a Unified Classifier Incrementally via Rebalancing, Saihui Hou et al. (2019). It provides crucial background on resolving data imbalance and classifier drift when retaining limited exemplars per category across incremental phases.
- Paper: Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks, Zhiwei Deng et al. (2022). It explores compressing past knowledge into compact memory budgets via bi-level optimization, conceptually underpinning exemplar compression in continual learning.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). It provides a foundational continual learning replay baseline leveraging past experience buffers under tight storage constraints.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). It provides a systematic taxonomy and evaluation of catastrophic forgetting mitigation strategies, contextualizing memory-constrained rehearsal.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). It introduces distillation-based preservation of past knowledge in incremental neural network training.
No sufficiently relevant recommendations were found.
