Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

Kai ZhuWei ZhaiYang CaoJiebo LuoZhengjun Zha

article2022CVPR184 citations

Develops a self-sustaining representation expansion framework for non-exemplar class-incremental learning that prevents catastrophic forgetting and parameter explosion by dynamically reorganizing network structures and selectively distilling knowledge using new class prototypes.

Listen

Modern computer vision models deployed in continuous, real-world environments must regularly learn new visual categories without forgetting previously acquired knowledge. While traditional incremental learning retains a memory bank of past images to prevent catastrophic forgetting, this practice frequently violates strict privacy regulations or exceeds edge-device storage capacities. Non-exemplar class-incremental learning addresses this challenge by assuming zero storage of past images, but existing systems suffer from severe forgetting and knowledge decay because training updates are guided entirely by new visual classes.

The article develops and evaluates a self-sustaining representation expansion framework designed to learn new classes without storing past data. Its primary objective is to maintain past knowledge and retain discriminative boundaries across expanding class sets while keeping the underlying model parameter size strictly constant after each training phase.

The authors designed a dual-mechanism approach evaluated across three standard benchmark datasets: CIFAR-100, TinyImageNet, and ImageNet-Subset. During incremental training phases, temporary residual adapter branches are inserted into a ResNet-18 architecture to absorb novel features while freezing the core feature extractor. Once training is complete, mathematical re-parameterization integrates these side-branch parameters back into the primary network without increasing overall parameter counts. Simultaneously, a prototype selection mechanism compares incoming new samples against stored class prototypes using cosine similarity: samples highly similar to old categories are routed to distillation loss to preserve historical knowledge, while dissimilar samples drive cross-entropy loss to learn novel classes.

The empirical findings demonstrate strong advantages over prior state-of-the-art methods across all tested benchmarks. The proposed framework outperforms existing non-exemplar techniques by average incremental accuracy margins of approximately 3% on CIFAR-100, 3% on TinyImageNet, and 6% on ImageNet-Subset (reaching 67.69% on the ImageNet subset). In multi-phase scenarios with 20 incremental steps on TinyImageNet, the method achieved a 6.10% absolute gain over previous approaches. Furthermore, the framework reduced average forgetting substantially—down to 9.17%–14.20% on TinyImageNet compared to 18.04%–30.55% for the best previous non-exemplar method—matching the performance of traditional systems that store 20 exemplar images per class.

These results establish that organizations can deploy continually updating vision models on resource-constrained edge hardware and in privacy-sensitive environments without sacrificing accuracy or incurring parameter growth. The prototype selection threshold of 0.8 achieved the optimal balance between historical retention and new class plasticity, confirming that selective routing of incoming data mitigates error accumulation across extended operational lifespans.

Organizations developing edge-based computer vision systems should adopt temporary residual adapters combined with structural re-parameterization to avoid parameter explosion. When deploying this architecture, teams must tune the prototype similarity threshold using validation data to maintain balanced routing between distillation and novel feature optimization. Further validation across broader enterprise-scale datasets and non-visual domains is recommended prior to full production deployment.

Cover for Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

Abstract

Non-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a novel self-sustaining representation expansion scheme. Our scheme consists of a structure reorganization strategy that fuses main-branch expansion and side-branch updating to maintain the old features, and a main-branch distillation scheme to transfer the invariant knowledge. Furthermore, a prototype selection mechanism is proposed to enhance the discrimination between the old and new classes by selectively incorporating new samples into the distillation process. Extensive experiments on three benchmarks demonstrate significant incremental performance, outperforming the state-of-the-art methods by a margin of 3%, 3% and 6%, respectively.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Incremental Learning
  • 2.2. Residual Block
  • 3. Problem Description
  • 4. Methodology
  • 4.1. Standard NECIL Paradigm
  • 4.2. Optimization
  • 4.3. Self-Sustaining Representation Expansion
  • 5. Experiments
  • 5.1. Dataset and Settings
  • 5.2. Ablation Study
  • 5.3. Analysis
  • 5.4. Visualization
  • 5.5. Comparison with SOTA
  • 6. Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — Dynamic Structure Reorganization for Non-Exemplar Class-Incremental Learning

    model/method

    In non-exemplar class-incremental learning (NECIL), storing past data is forbidden, which typically causes parameter optimization to rapidly overwrite representations of past classes. Dynamic Structure Reorganization (DSR) addresses this by decoupling the optimization paths for old and new representations across incremental phases while maintaining a constant parameter budget during inference.

    At phase n>1n > 1, given an input image batch QQ from the current task training set XnX^n, the backbone feature extractor fen−1f_e^{n-1} parameterized by fixed weights θ^en−1\hat{\theta}_e^{n-1} is expanded by inserting a residual adapter Δθen\Delta \theta_e^n into each convolutional block:

    fen(Q;θen)=Ftransform(fen−1(Q;θ^en−1))=fen−1(Q;θ^en−1⊕Δθen)f_e^n(Q; \theta_e^n) = F_{\text{transform}}(f_e^{n-1}(Q; \hat{\theta}_e^{n-1})) = f_e^{n-1}(Q; \hat{\theta}_e^{n-1} \oplus \Delta \theta_e^n)

    where ⋅^\hat{\cdot} denotes frozen parameters, and ⊕\oplus denotes structural expansion. Gradient flow updates only the adapter parameters Δθen\Delta \theta_e^n, isolating the main backbone from destructive gradient updates.

    Once training for phase nn concludes, structural re-parameterization integrates the side-branch adapter weights into the main branch convolutional kernels and Batch Normalization layers losslessly via zero-padding and linear transformation:

    fen−1(Q;θ^en−1⊕Δθen)=fen(Q;θen′⊕0)=fen(Q;θen′)f_e^{n-1}(Q; \hat{\theta}_e^{n-1} \oplus \Delta \theta_e^n) = f_e^n(Q; {\theta_e^n}' \oplus 0) = f_e^n(Q; {\theta_e^n}')

    where θen′{\theta_e^n}' denotes the fused parameters. The adapter structures are then discarded, reverting the network to its standard topology and preventing linear parameter growth over incremental phases.

  2. Knowl 2 — Prototype Selection Mechanism for Selective Distillation

    model/method

    In non-exemplar class-incremental learning, knowledge distillation using only incoming new class samples often transfers imprecise old-task distribution information, confusing new class representations with similar old classes. The Prototype Selection Mechanism (PSM) selectively routes incremental samples to either distillation or task adaptation based on their feature-space proximity to stored class prototypes.

    For a sample representation rqn=fen(Q;θen)r_q^n = f_e^n(Q; \theta_e^n) in the current phase nn, the normalized cosine similarity score SiS_i against stored prototypes of all observed classes is computed as:

    Si=Cosine(Nor(rqn),Nor(Prototype))S_i = \text{Cosine}(\text{Nor}(r_q^n), \text{Nor}(\text{Prototype}))

    where Nor(⋅)\text{Nor}(\cdot) denotes L2L_2 normalization.

    Given a similarity threshold σ\sigma (set empirically to σ=0.8\sigma = 0.8):

    • Samples with Si≥σS_i \ge \sigma are considered semantically close to previously learned classes; their distillation mask Maskkd\text{Mask}_{kd} is activated (11) and classification mask Maskce\text{Mask}_{ce} is set to 00. These samples participate in knowledge distillation to preserve fine-grained old class statistics via soft labels.
    • Samples with Si<σS_i < \sigma are considered distinct from old classes; their classification mask Maskce\text{Mask}_{ce} is activated (11) and distillation mask Maskkd\text{Mask}_{kd} is set to 00. These samples drive the adaptation of the residual adapter to learn discriminative novel representations.
  3. Knowl 3 — Training Loss Function with Prototype Oversampling and Masked Objectives

    equation

    In the self-sustaining representation expansion framework for non-exemplar class-incremental learning at phase nn, the network is optimized using a combination of masked cross-entropy, masked knowledge distillation, and balanced prototype cross-entropy:

    L=Maskce(Lce)+λMaskkd(Lkd)+γLproto\mathcal{L} = \text{Mask}_{ce}(\mathcal{L}_{ce}) + \lambda \text{Mask}_{kd}(\mathcal{L}_{kd}) + \gamma \mathcal{L}_{proto}

    where:

    • Lce=Fce(sqn,yqn)\mathcal{L}_{ce} = F_{ce}(s_q^n, y_q^n) is the standard cross-entropy loss over the predicted logits sqn=gcn(rqn;θcn)s_q^n = g_c^n(r_q^n; \theta_c^n) and true novel labels yqn∈Yny_q^n \in Y^n.
    • Lkd=Fkd(rqn,rqn−1)=∥rqn−rqn−1∥22\mathcal{L}_{kd} = F_{kd}(r_q^n, r_q^{n-1}) = \|r_q^n - r_q^{n-1}\|^2_2 is the Euclidean distance distillation loss aligning the current representation rqn=fen(Q;θen)r_q^n = f_e^n(Q; \theta_e^n) with the predecessor feature extractor representation rqn−1=fen−1(Q;θen−1)r_q^{n-1} = f_e^{n-1}(Q; \theta_e^{n-1}).
    • Lproto=Fce(pB,yB)\mathcal{L}_{proto} = F_{ce}(p_B, y_B) is the prototype balance classification loss, where pB=UpB(Prototype)p_B = Up_B(\text{Prototype}) is generated by oversampling the stored deep feature prototypes (one per class) up to the batch size BB, and yBy_B denotes the corresponding oversampled prototype labels.
    • Maskce\text{Mask}_{ce} and Maskkd\text{Mask}_{kd} are binary selection masks determined by the Prototype Selection Mechanism based on cosine similarity threshold σ\sigma.
    • λ\lambda and γ\gamma are loss weighting hyperparameters, both set to 1010.
  4. Knowl 4 — Error Accumulation Lower Bound in Non-Exemplar Distillation

    theoretical result

    In class-incremental learning without exemplars (NECIL), old samples are completely absent during training phases n>1n > 1. Consequently, cross-entropy supervision Lce\mathcal{L}_{ce} optimizes solely for new-task discrimination, while distillation loss Lkd\mathcal{L}_{kd} relies entirely on new class inputs Q∈XnQ \in X^n to preserve old representations.

    Let α\alpha represent the representation forgetting rate of the distillation process during the initial transition phase (n=2n=2). Because distillation in each subsequent phase nn is supervised only by representations from the immediately preceding phase n−1n-1, approximation errors and bias drift accumulate across successive phases, yielding an accumulated forgetting rate FrnFr_n that satisfies:

    Frn≥αn−1Fr_n \ge \alpha^{n-1}

    This exponential compounding of error demonstrates why standard rehearsal-free distillation degenerates across multi-phase settings unless feature representations are structurally isolated and selectively distilled.

  5. Knowl 5 — Average Incremental Accuracy Across Standard Continual Learning Benchmarks

    data/table

    The table below compares the average incremental accuracy (%) of the self-sustaining representation expansion method against non-exemplar class-incremental learning (NECIL, E=0E=0) methods and exemplar-based (E=20E=20) methods across CIFAR-100, TinyImageNet, and ImageNet-Subset under 5, 10, and 20 incremental phases (PP).

    Methods CIFAR-100 TinyImageNet ImageNet-Subset
    P=5P=5 P=10P=10 P=20P=20 P=5P=5 P=10P=10 P=20P=20 P=10P=10
    Exemplar-based (E=20E=20)
    iCaRL-CNN 51.07 48.66 44.43 34.64 31.15 27.90 50.53
    iCaRL-NCM 58.56 54.19 50.51 45.86 43.29 38.04 60.79
    EEIL 60.37 56.05 52.34 47.12 45.01 40.50 63.34
    UCIR 63.78 62.39 59.07 49.15 48.52 42.83 66.16
    Non-exemplar (E=0E=0)
    EWC 24.48 21.20 15.89 18.80 15.77 12.39 20.40
    LwF_MC 45.93 27.43 20.07 29.12 23.10 17.43 31.18
    MUC 49.42 30.19 21.27 32.58 26.61 21.95 35.07
    SDC 56.77 57.00 58.90 - - - 61.12
    PASS 63.47 61.84 58.09 49.55 47.29 42.07 61.80
    Ours 65.88 65.04 61.70 50.39 48.93 48.17 67.69

    The proposed method outperforms the previous non-exemplar state-of-the-art (PASS) by 2.41%2.41\%, 3.20%3.20\%, and 2.80%2.80\% on CIFAR-100 (P=5,10,20P=5, 10, 20); by 0.84%0.84\%, 1.64%1.64\%, and 6.10%6.10\% on TinyImageNet; and by 5.89%5.89\% on ImageNet-Subset (P=10P=10). It also surpasses exemplar-based methods storing 20 exemplars per class (such as UCIR).

  6. Knowl 6 — Resistance to Catastrophic Forgetting in NECIL

    data/table

    The table below reports average forgetting (%) across different incremental phases (P=5,10,20P=5, 10, 20) on CIFAR-100 and TinyImageNet, measuring the performance degradation on past tasks throughout the incremental process.

    Method CIFAR-100 TinyImageNet
    P=5P=5 P=10P=10 P=20P=20 P=5P=5 P=10P=10 P=20P=20
    iCaRL-CNN 42.13 45.69 43.54 36.89 36.70 45.12
    iCaRL-NCM 24.90 28.32 35.53 27.15 28.89 37.40
    EEIL 23.36 26.65 32.40 25.56 25.91 35.04
    UCIR 21.00 25.12 28.65 20.61 22.25 33.74
    LwF_MC 44.23 50.47 55.46 54.26 54.37 63.54
    MUC 40.28 47.56 52.65 51.46 50.21 58.00
    PASS 25.20 30.25 30.61 18.04 23.11 30.55
    Ours 18.37 19.48 19.00 9.17 14.06 14.20

    The proposed method reduces average forgetting by over 11%11\% on CIFAR-100 (20 phases) and over 16%16\% on TinyImageNet (20 phases) compared to PASS, and maintains lower forgetting rates than exemplar-based approaches (E=20E=20) such as UCIR.

  7. Knowl 7 — Ablation of Dynamic Structure Reorganization, Distillation, and Prototype Selection

    data/table

    The individual and combined contributions of Dynamic Structure Reorganization (DSR), Main-Branch Distillation (MBD), and the Prototype Selection Mechanism (PSM) were evaluated on CIFAR-100 across 5, 10, and 20 incremental phases (PP).

    DSR MBD PSM CIFAR-100 Accuracy (%)
    P=5P=5 phases P=10P=10 phases P=20P=20 phases
    61.11 57.08 51.04
    ✓ 64.86 63.25 54.09
    ✓ 62.70 62.60 58.57
    ✓ ✓ 65.10 63.87 60.60
    ✓ ✓ ✓ 65.88 64.69 61.61

    Key findings include:

    • Adding DSR alone yields a 3.75%3.75\% improvement in the 5-phase setting (64.86%64.86\% vs. 61.11%61.11\%) by shielding old representations from gradient updates, but its standalone advantage decreases in longer phase sequences (P=20P=20).
    • Adding MBD alone provides larger gains in longer sequences (+7.53%+7.53\% for P=20P=20, reaching 58.57%58.57\%) by continuously enforcing feature invariance.
    • Combining all three components achieves the best performance across all phase divisions (65.88%65.88\%, 64.69%64.69\%, and 61.61%61.61\%).
  8. Knowl 8 — Adapter Convolution Architecture and Selection Threshold Sensitivity

    empirical result

    Evaluating different structural choices for the residual adapter in Dynamic Structure Reorganization and similarity thresholds in the Prototype Selection Mechanism on CIFAR-100 shows:

    1. Adapter Architecture: Comparing three residual branch designs:

      • 1×11 \times 1 convolution: achieves 65.87%65.87\% (P=5P=5), 65.12%65.12\% (P=10P=10), and 61.60%61.60\% (P=20P=20).
      • 1×11 \times 1 convolution + Batch Normalization: achieves 65.88%65.88\% (P=5P=5), 64.84%64.84\% (P=10P=10), and 60.72%60.72\% (P=20P=20).
      • 3×33 \times 3 convolution: achieves 64.28%64.28\% (P=5P=5), 63.47%63.47\% (P=10P=10), and 60.81%60.81\% (P=20P=20). A simple 1×11 \times 1 convolution provides optimal performance without requiring the parameter overhead of 3×33 \times 3 kernels.
    2. Prototype Selection Threshold σ\sigma: Evaluating cosine similarity thresholds σ∈[0.5,0.85]\sigma \in [0.5, 0.85] across 5, 10, and 20 phases shows that accuracy consistently peaks at σ=0.8\sigma = 0.8 (reaching 65.88%65.88\%, 64.69%64.69\%, and 61.61%61.61\% respectively). Lower thresholds misroute dissimilar samples into distillation, while higher thresholds starve the distillation process of necessary near-class samples.

Coverage note — None was omitted; all key architectural components (DSR, MBD, PSM, structural re-parameterization), analytical derivations, empirical results, and ablation studies were converted into standalone knowls.

References

  1. 1.Francisco M Castro, Manuel J Marín-Jiménez, Nicolás Guil, Cordelia Schmid, and Karteek Alahari. End-to-end incremental learning. In Proceedings of the European conference on computer vision (ECCV), pages 233–248, 2018.
  2. 2.Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. Smote: synthetic minority oversampling technique. Journal of artificial intelligence research, 16:321–357, 2002.
  3. 3.Zhen Cheng, Zhiwei Xiong, Chang Chen, Dong Liu, and Zheng-Jun Zha. Light field super-resolution with zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10010–10019, 2021.
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  5. 5.Xiaohan Ding, Yuchen Guo, Guiguang Ding, and J. Han. Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1911–1920, 2019.
  6. 6.Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun. Repvgg: Making vgg-style convnets great again. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13733–13742, 2021.
  7. 7.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16, pages 86–102. Springer, 2020.
  8. 8.Robert M French. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135, 1999.
  9. 9.Shuxuan Guo, José Manuel Álvarez, and Mathieu Salzmann. Expandnets: Linear over-parameterization to train compact convolutional networks. arXiv: Computer Vision and Pattern Recognition, 2020.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  11. 11.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 831–839, 2019.
  12. 12.Xinting Hu, Kaihua Tang, Chunyan Miao, Xiansheng Hua, and Hanwang Zhang. Distilling causal effect of data in classincremental learning. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3956–3965, 2021.
  13. 13.Steven C. Y. Hung, Cheng-Hao Tu, Cheng-En Wu, ChienHung Chen, Yi-Ming Chan, and Chu-Song Chen. Compacting, picking and growing for unforgetting continual learning. ArXiv, abs/1910.06562, 2019.
  14. 14.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015.
  15. 15.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka GrabskaBarwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  16. 16.A. Krizhevsky. Learning multiple layers of features from tiny images. 2009.
  17. 17.Ya Le and Xuan Yang. Tiny imagenet visual recognition challenge. CS 231N, 7(7):3, 2015.
  18. 18.Wei-Hong Li, Xialei Liu, and Hakan Bilen. Improving task adaptation for cross-domain few-shot learning. ArXiv, abs/2107.00358, 2021.
  19. 19.Yunsheng Li, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Ye Yu, Lu Yuan, Zicheng Liu, Mei Chen, and Nuno Vasconcelos. Revisiting dynamic convolution via matrix decomposition. arXiv preprint arXiv:2103.08756, 2021.
  20. 20.Jiawei Liu, Zheng-Jun Zha, Di Chen, Richang Hong, and Meng Wang. Adaptive transfer network for cross-domain person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7202–7211, 2019.
  21. 21.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Adaptive aggregation networks for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2544–2553, 2021.
  22. 22.Yaoyao Liu, Yuting Su, An-An Liu, Bernt Schiele, and Qianru Sun. Mnemonics training: Multi-class incremental learning without forgetting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12245–12254, 2020.
  23. 23.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(Nov):2579–2605, 2008.
  24. 24.Jathushan Rajasegaran, Munawar Hayat, Salman Khan, Fahad Shahbaz Khan, and Ling Shao. Random path selection for incremental learning. Advances in Neural Information Processing Systems, 2019.
  25. 25.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. Learning multiple visual domains with residual adapters. arXiv preprint arXiv:1705.08045, 2017.
  26. 26.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. Efficient parametrization of multi-domain deep neural networks. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8119–8127, 2018.
  27. 27.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2001–2010, 2017.
  28. 28.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  29. 29.Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. ArXiv, abs/1606.04671, 2016.
  30. 30.Christian Simon, Piotr Koniusz, and Mehrtash Harandi. On learning the geodesic path for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1591–1600, 2021.
  31. 31.K. Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2015.
  32. 32.Gido M. van de Ven and A. Tolias. Three scenarios for continual learning. ArXiv, abs/1904.07734, 2019.
  33. 33.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 374–382, 2019.
  34. 34.Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. Vitae: Vision transformer advanced by exploring intrinsic inductive bias. Advances in Neural Information Processing Systems, 34, 2021.
  35. 35.Shipeng Yan, Jiangwei Xie, and Xuming He. Der: Dynamically expandable representation for class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3014–3023, 2021.
  36. 36.Hongxu Yin, Pavlo Molchanov, Jose M Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K Jha, and Jan Kautz. Dreaming to distill: Data-free knowledge transfer via deepinversion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8715–8724, 2020.
  37. 37.L. Yu, S. Parisot, G. Slabaugh, J. Xu, and T. Tuytelaars. More classifiers, less forgetting: A generic multi-classifier paradigm for incremental learning. European Conference on Computer Vision, 2020.
  38. 38.Lu Yu, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang, Yongmei Cheng, Shangling Jui, and Joost van de Weijer. Semantic drift compensation for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6982–6991, 2020.
  39. 39.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In International Conference on Machine Learning, pages 3987–3995. PMLR, 2017.
  40. 40.Hanwang Zhang, Zheng-Jun Zha, Yang Yang, Shuicheng Yan, and Tat-Seng Chua. Robust (semi) nonnegative graph embedding. IEEE transactions on image processing, 23(7):2996–3012, 2014.
  41. 41.Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao. Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond. arXiv preprint arXiv:2202.10108, 2022.
  42. 42.Heliang Zheng, Jianlong Fu, Zheng-Jun Zha, and Jiebo Luo. Learning deep bilinear transformation for fine-grained image representation. Advances in Neural Information Processing Systems, 32, 2019.
  43. 43.Kecheng Zheng, Wu Liu, Lingxiao He, Tao Mei, Jiebo Luo, and Zheng-Jun Zha. Group-aware label transfer for domain adaptive person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5310–5319, 2021.
  44. 44.Fei Zhu, Xu-Yao Zhang, Chuang Wang, Fei Yin, and ChengLin Liu. Prototype augmentation and self-supervision for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5871–5880, 2021.
  45. 45.Kai Zhu, Yang Cao, Wei Zhai, Jie Cheng, and Zheng-Jun Zha. Self-promoted prototype refinement for few-shot classincremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6801–6810, 2021.

Citation

MLA
Zhu, K., et al. “Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning”. arXiv, 2022, http://arxiv.org/abs/2203.06359v2.
APA
Zhu, K., Zhai, W., Cao, Y., Luo, J., & Zha, Z.-J. (2022). Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning. arXiv. http://arxiv.org/abs/2203.06359v2
Chicago
Zhu, K., W. Zhai, Y. Cao, J. Luo, and Z.-J. Zha. 2022. “Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning”. arXiv. http://arxiv.org/abs/2203.06359v2.
Harvard
Zhu, K. et al. (2022) “Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.06359v2.
Vancouver
1. Zhu K, Zhai W, Cao Y, Luo J, Zha Z-J (2022) Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning. arXiv

BibTeX

@article{zhu2022self,
  title = {Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning},
  author = {Zhu, Kai and Zhai, Wei and Cao, Yang and Luo, Jiebo and Zha, Zheng-Jun},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.06359v2},
  eprint = {2203.06359}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE