Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning
Kai ZhuWei ZhaiYang CaoJiebo LuoZhengjun Zha
Develops a self-sustaining representation expansion framework for non-exemplar class-incremental learning that prevents catastrophic forgetting and parameter explosion by dynamically reorganizing network structures and selectively distilling knowledge using new class prototypes.
Modern computer vision models deployed in continuous, real-world environments must regularly learn new visual categories without forgetting previously acquired knowledge. While traditional incremental learning retains a memory bank of past images to prevent catastrophic forgetting, this practice frequently violates strict privacy regulations or exceeds edge-device storage capacities. Non-exemplar class-incremental learning addresses this challenge by assuming zero storage of past images, but existing systems suffer from severe forgetting and knowledge decay because training updates are guided entirely by new visual classes.
The article develops and evaluates a self-sustaining representation expansion framework designed to learn new classes without storing past data. Its primary objective is to maintain past knowledge and retain discriminative boundaries across expanding class sets while keeping the underlying model parameter size strictly constant after each training phase.
The authors designed a dual-mechanism approach evaluated across three standard benchmark datasets: CIFAR-100, TinyImageNet, and ImageNet-Subset. During incremental training phases, temporary residual adapter branches are inserted into a ResNet-18 architecture to absorb novel features while freezing the core feature extractor. Once training is complete, mathematical re-parameterization integrates these side-branch parameters back into the primary network without increasing overall parameter counts. Simultaneously, a prototype selection mechanism compares incoming new samples against stored class prototypes using cosine similarity: samples highly similar to old categories are routed to distillation loss to preserve historical knowledge, while dissimilar samples drive cross-entropy loss to learn novel classes.
The empirical findings demonstrate strong advantages over prior state-of-the-art methods across all tested benchmarks. The proposed framework outperforms existing non-exemplar techniques by average incremental accuracy margins of approximately 3% on CIFAR-100, 3% on TinyImageNet, and 6% on ImageNet-Subset (reaching 67.69% on the ImageNet subset). In multi-phase scenarios with 20 incremental steps on TinyImageNet, the method achieved a 6.10% absolute gain over previous approaches. Furthermore, the framework reduced average forgetting substantially—down to 9.17%–14.20% on TinyImageNet compared to 18.04%–30.55% for the best previous non-exemplar method—matching the performance of traditional systems that store 20 exemplar images per class.
These results establish that organizations can deploy continually updating vision models on resource-constrained edge hardware and in privacy-sensitive environments without sacrificing accuracy or incurring parameter growth. The prototype selection threshold of 0.8 achieved the optimal balance between historical retention and new class plasticity, confirming that selective routing of incoming data mitigates error accumulation across extended operational lifespans.
Organizations developing edge-based computer vision systems should adopt temporary residual adapters combined with structural re-parameterization to avoid parameter explosion. When deploying this architecture, teams must tune the prototype similarity threshold using validation data to maintain balanced routing between distillation and novel feature optimization. Further validation across broader enterprise-scale datasets and non-visual domains is recommended prior to full production deployment.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Introduces the foundational knowledge distillation mechanism for adapting neural networks to new classification tasks without storing historical exemplar data.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). Establishes the standard benchmark setting for class-incremental learning combining distillation and nearest-mean prototype classification.
- Paper: Lifelong Learning with Dynamically Expandable Networks, Jaehong Yoon et al. (2017). Pioneers dynamically expandable network architectures to accommodate sequential tasks without catastrophic forgetting.
- Paper: Learning a Unified Classifier Incrementally via Rebalancing, Saihui Hou et al. (2019). Analyzes the class imbalance between old and new categories in incremental learning and introduces cosine normalization to preserve unified decision boundaries.
- Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). Develops end-to-end incremental learning using combined cross-entropy and distillation objectives under memory constraints.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). Provides a comprehensive taxonomy and evaluation framework of continual learning strategies, parameter isolation, and stability-plasticity trade-offs.
- Paper: FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning, Dipam Goswami et al. (2023). Extends exemplar-free class-incremental learning by addressing heterogeneous class distributions in frozen feature extractor prototype representations.
- Paper: Dense Network Expansion for Class Incremental Learning, Zhiyuan Hu et al. (2023). Advances network expansion paradigms for class-incremental learning by incorporating lightweight task-attention blocks into vision transformers while curbing parameter growth.
- Paper: GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task, Huiping Zhuang et al. (2023). Builds upon non-iterative classifier expansion by introducing closed-form analytic learning with Gaussian kernels for few-shot class incremental tasks.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). Provides a comprehensive theoretical and methodological synthesis of subsequent developments across continual learning, including structural and representation expansion.
