Dense Network Expansion for Class Incremental Learning

Zhiyuan HuYunsheng LiJiancheng LyuDashan GaoNuno Vasconcelos

article2023CVPR89 citations

Proposes a dense network expansion method for class incremental learning that transfers knowledge across frozen task experts using a cross-task attention block, achieving superior accuracy while significantly curbing parameter growth.

Listen

Modern computer vision models face catastrophic forgetting, an issue where sequential training on new visual classes causes the system to erase previously acquired knowledge. While network expansion methods prevent forgetting by adding specialized sub-networks for each new task, they trigger rapid, unsustainable growth in computational footprint and model size. The article designs and evaluates Dense Network Expansion, an architecture that maintains past recognition accuracy while substantially curbing resource growth by sharing and reusing intermediate features across tasks.

Dense Network Expansion decouples spatial image analysis from cross-task knowledge transfer using Vision Transformers. Instead of duplicating entire large networks or entangling spatial and task tokens—which causes attention dilution—the framework freezes previous components and incorporates a lightweight task attention block directly into the feature-mixing stage. The authors benchmarked this design against leading continuous learning methods on standardized vision datasets, assessing accuracy, computational operations, and scalability across multi-task schedules up to 26 sequential stages.

The findings show that Dense Network Expansion outperforms existing class-incremental learning methods. When adding single-head task experts, the system achieved a 68.04% final accuracy on CIFAR-100, surpassing the previous top performer by nearly 4%, while outperforming the leading baseline on ImageNet-100 by 1.9%. Crucially, the approach narrows the performance gap relative to an ideal, fully retrained model by up to 50%. In extended sequences spanning 26 incremental steps, the architecture maintained high accuracy while consuming a fraction of the computational operations required by traditional network expansion baselines.

These results demonstrate that continuous learning does not require choosing between catastrophic forgetting and runaway infrastructure costs. By enabling compact sub-networks to selectively query and integrate older representations, organizations can deploy edge and cloud vision systems that continually scale to new classes with predictable computational overhead. Teams implementing incremental learning pipelines should adopt modular cross-task feature sharing rather than full-backbone expansion or joint spatial-task attention mechanisms.

The reported advantages rely on benchmarks where early tasks establish strong initial representations, and tasks with larger class additions may still require increasing the capacity of task branches. Nonetheless, the experimental evidence provides strong confidence that dense, decoupled attention provides a viable, resource-efficient foundation for sequential vision learning.

  • Paper: Lifelong Learning with Dynamically Expandable Networks, Jaehong Yoon et al. (2017). This foundational work on dynamically expandable networks establishes the architectural paradigm of adding and freezing sub-networks during sequential task learning that Dense Network Expansion directly optimizes and refines.
  • Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). Progressive Neural Networks introduced the core mechanism of freezing previous columns and transferring intermediate features to new task networks via lateral connections, a foundational precursor to Dense Network Expansion.
  • Paper: Densely Connected Convolutional Networks, Gao Huang et al. (2017). DenseNets pioneered the dense feature-reuse connectivity pattern that Dense Network Expansion adapts across temporal task boundaries to curb parameter growth.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This paper establishes the standard Learning without Forgetting formulation and knowledge distillation framework for multi-task and incremental vision adaptation.
  • Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). L2P provides essential background on leveraging frozen transformer backbones and modular dynamic parameter tuning for class-incremental vision tasks without buffer storage.
  • Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). This paper defines the standard class-incremental learning benchmark protocols and cross-distillation evaluation paradigms that modern incremental vision models are measured against.
  • Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey provides a comprehensive taxonomy of parameter-isolation and expansion techniques versus regularized continuous learning approaches for classification.
Cover for Dense Network Expansion for Class Incremental Learning

Abstract

The problem of class incremental learning (CIL) is considered. State-of-the-art approaches use a dynamic architecture based on network expansion (NE), in which a task expert is added per task. While effective from a computational standpoint, these methods lead to models that grow quickly with the number of tasks. A new NE method, dense network expansion (DNE), is proposed to achieve a better trade-off between accuracy and model complexity. This is accomplished by the introduction of dense connections between the intermediate layers of the task expert networks, that enable the transfer of knowledge from old to new tasks via feature sharing and reusing. This sharing is implemented with a cross-task attention mechanism, based on a new task attention block (TAB), that fuses information across tasks. Unlike traditional attention mechanisms, TAB operates at the level of the feature mixing and is decoupled with spatial attentions. This is shown more effective than a joint spatial-and-task attention for CIL. The proposed DNE approach can strictly maintain the feature space of old classes while growing the network and feature scale at a much slower rate than previous methods. In result, it outperforms the previous SOTA methods by a margin of 4% in terms of accuracy, with similar or even smaller model scale.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Dense Network Expansion
  • 4. Experiments
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Dense Network Expansion Framework for Class Incremental Learning

    model/method

    Dense Network Expansion (DNE) is a class incremental learning (CIL) architecture designed to balance classification accuracy and computational scale as new tasks arrive sequentially.

    In CIL, a model learns a sequence of MM classification tasks T={T1,T2,…,TM}\mathcal{T} = \{T_1, T_2, \dots, T_M\}, where task TtT_t has a dataset DtD_t over a disjoint class set YtY_t (Yi∩Yj=∅Y_i \cap Y_j = \emptyset for i≠ji \neq j). At step tt, the model must classify inputs across all observed classes ⋃k=1tYk\bigcup_{k=1}^t Y_k, typically using current task data DtD_t and an exemplar memory buffer M\mathcal{M} containing a small number of samples from previous tasks.

    Rather than adding an entire independent multi-head backbone network per task (standard Network Expansion) or compressing all knowledge into a fixed single network (single-model distillation), DNE introduces dense cross-task connections across intermediate layers of lightweight task-specific subnetwork branches (experts). When learning task tt:

    1. All previously learned task experts f1,…,ft−1f^1, \dots, f^{t-1} are frozen, preserving the representations of prior classes without catastrophic forgetting.
    2. A compact task expert ftf^t with LL layers is added. At block l∈{1,…,L}l \in \{1, \dots, L\}, expert tt generates its representations rltr_l^t by taking as input the intermediate feature representations from all task experts at block l−1l-1: rlt=flt(rl−11,…,rl−1t−1,rl−1t;θlt)r_l^t = f_l^t(r_{l-1}^1, \dots, r_{l-1}^{t-1}, r_{l-1}^t; \theta_l^t) where θlt\theta_l^t denotes the trainable parameters of block ll in expert tt.
    3. Feature representations from all experts at the final layer LL are concatenated to form rL=rL1⊕rL2⊕⋯⊕rLtr_L = r_L^1 \oplus r_L^2 \oplus \dots \oplus r_L^t, which is classified via a unified classifier h(rL;ϕ)h(r_L; \phi) with parameters ϕ={ϕ1,…,ϕt}\phi = \{\phi^1, \dots, \phi^t\}.

    By reusing and recombining features from prior frozen experts, DNE enables each added task expert to use significantly fewer parameters (e.g., only 1 to 4 attention heads per task rather than a full 12-head transformer), preventing unsustainable model growth.

  2. Knowl 2 — Task Attention Block for Disentangled Cross-Task Attention

    model/method

    In Vision Transformer (ViT) based architectures for class incremental learning, Dense Network Expansion (DNE) decouples spatial attention from cross-task attention. At layer ll, spatial attention is performed independently within each task expert using standard Multi-Head Self-Attention (MHSA) over image patch tokens: slt=rl−1t+MHSAl(LN(rl−1t))s_l^t = r_{l-1}^t + \text{MHSA}_l(\text{LN}(r_{l-1}^t)) where LN\text{LN} denotes layer normalization and rl−1tr_{l-1}^t is the output of the (l−1)(l-1)-th block for task tt.

    To share and fuse features across tasks without fragmenting spatial attention, DNE replaces the standard multi-layer perceptron (MLP) with a Task Attention Block (TAB). For patch index p∈{0,…,P−1}p \in \{0, \dots, P-1\} and spatial attention feature spt=spt,1⊕⋯⊕spt,Hts_p^t = s_p^{t,1} \oplus \dots \oplus s_p^{t,H_t} across HtH_t attention heads of expert tt:

    1. First Attention Stage (Cross-Task Feature Mixing): The query vector for head ii of patch pp in task expert tt is: Qpt,i=WqLN(spt,i)∈RD,Qp=[Qpt,1,…,Qpt,Ht]∈RD×HtQ_p^{t,i} = W_q \text{LN}(s_p^{t,i}) \in \mathbb{R}^D, \quad Q_p = [Q_p^{t,1}, \dots, Q_p^{t,H_t}] \in \mathbb{R}^{D \times H_t} where Wq∈RD×DW_q \in \mathbb{R}^{D \times D} is a learned projection matrix shared across heads, and DD is the token feature dimension per head.

    Keys are formed from all H=∑k=1tHkH = \sum_{k=1}^t H_k heads across all current and frozen previous task experts: Kpi,j=WkLN(spi,j)∈RD,Kp=[Kp1,1,…,Kpt,Ht]∈RD×HK_p^{i,j} = W_k \text{LN}(s_p^{i,j}) \in \mathbb{R}^D, \quad K_p = [K_p^{1,1}, \dots, K_p^{t,H_t}] \in \mathbb{R}^{D \times H} with shared projection matrix Wk∈RD×DW_k \in \mathbb{R}^{D \times D}.

    Attention weights Ap∈RHt×HA_p \in \mathbb{R}^{H_t \times H} are calculated by normalizing the similarity matrix Cp=QpTKp∈RHt×HC_p = Q_p^T K_p \in \mathbb{R}^{H_t \times H} via row-wise softmax: Api=SoftMax(CpiD)∈RHA_p^i = \text{SoftMax}\left(\frac{C_p^i}{\sqrt{D}}\right) \in \mathbb{R}^H

    Values use head-specific projection matrices Wvi,j∈RD′×DW_v^{i,j} \in \mathbb{R}^{D' \times D} (where D′=γDD' = \gamma D with expansion factor γ\gamma): Vpi,j=Wvi,jLN(spi,j)∈RD′,Vp=[Vp1,1,…,Vpt,Ht]∈RD′×HV_p^{i,j} = W_v^{i,j} \text{LN}(s_p^{i,j}) \in \mathbb{R}^{D'}, \quad V_p = [V_p^{1,1}, \dots, V_p^{t,H_t}] \in \mathbb{R}^{D' \times H}

    The intermediate response for head ii of task tt is: opt,i=GELU(λi∑j=1HApi,jVpj)∈RD′o_p^{t,i} = \text{GELU}\left(\lambda_i \sum_{j=1}^H A_p^{i,j} V_p^j\right) \in \mathbb{R}^{D'} where λi\lambda_i is a learned scalar multiplier and GELU\text{GELU} is the Gaussian Error Linear Unit.

    1. Second Attention Stage: A second task attention operation with identical structure projects the intermediate representations {op1,…,opt}\{o_p^1, \dots, o_p^t\} from dimension D′D' back to DD, producing the layer output for task tt: rpt=spt+TA(LN(op1,…,opt))r_p^t = s_p^t + \text{TA}(\text{LN}(o_p^1, \dots, o_p^t))

    The outputs and activations of previous task experts {rp1,…,rpt−1}\{r_p^1, \dots, r_p^{t-1}\} and {op1,…,opt−1}\{o_p^1, \dots, o_p^{t-1}\} remain frozen.

  3. Knowl 3 — Attention Fragmentation in Joint Spatial-and-Task Attention

    theoretical result

    When extending vision transformers to multi-branch class incremental learning (CIL), a joint Spatial-and-Task Attention (STA) mechanism calculates attention across all image patch tokens p∈{0,…,P−1}p \in \{0, \dots, P-1\} and all task heads t∈{1,…,T}t \in \{1, \dots, T\}.

    Because every task expert in image CIL processes the identical input image, four categories of pairwise attention dot-products arise between query and key tokens:

    1. SPSH: Same patch, same head (self-attention within the patch).
    2. SPDH: Same patch, different heads (attention between representations of the identical spatial image region across different task experts).
    3. DPSH: Different patches, same head (intra-task spatial context).
    4. DPDH: Different patches, different heads (cross-task spatial context).

    Because SPDH tokens correspond to identical image patches, their dot-products dominate the attention matrix:

    • In STA, same-patch connections (SPSH + SPDH) consume 25.8%25.8\% of total attention weight (compared to 12.67%12.67\% in Independent Attention models).
    • The remaining 74.2%74.2\% of attention weight is dispersed across all cross-head spatial pairs (DPDH), producing a large number of very small-valued matrix entries. This fragments spatial attention and severely dilutes spatial structure.

    As a consequence, STA attains only 41.07%41.07\% final accuracy on CIFAR100 with 6 tasks (50 base + 5 tasks of 10 classes), substantially below both Independent Attention (58.22%58.22\%) and Dense Network Expansion (68.04%68.04\%).

  4. Knowl 4 — Multi-Loss Training Objective in Dense Network Expansion

    equation

    During incremental task step tt, the trainable parameters θt\theta^t of task expert tt and classifier parameters ϕ\phi in Dense Network Expansion are optimized using a composite objective consisting of three loss terms: L=Lce+Lte+Ldis\mathcal{L} = \mathcal{L}_{ce} + \mathcal{L}_{te} + \mathcal{L}_{dis}

    1. Unified Cross-Entropy Loss (Lce\mathcal{L}_{ce}): Evaluated over the union of all observed classes ⋃i=1tYi\bigcup_{i=1}^t Y_i using the output probabilities of the concatenated classifier g(x;θ)=h(rL;ϕ)g(x; \theta) = h(r_L; \phi) on both current task data DtD_t and exemplar buffer M\mathcal{M}.

    2. Task Expertise Loss (Lte\mathcal{L}_{te}): Forces task expert tt to learn features discriminative for task tt. It collapses all previously learned classes into a single aggregate class Y′=⋃i=1t−1Yi\mathcal{Y}' = \bigcup_{i=1}^{t-1} Y_i and optimizes a (∣Yt∣+1)(|Y_t| + 1)-way classification objective: Lte=−∑(x,y)∈Dt∪Mlog⁡P(y′∣x)\mathcal{L}_{te} = -\sum_{(x, y) \in D_t \cup \mathcal{M}} \log P(y' \mid x) where y′=yy' = y if y∈Yty \in Y_t, and y′=classoldy' = \text{class}_{\text{old}} if y∈Y′y \in \mathcal{Y}'.

    3. Distillation Loss (Ldis\mathcal{L}_{dis}): Enforces consistency on historical classes between the updated model gtg^t and the frozen model gt−1g^{t-1} from the prior step. Softmax is applied to logits corresponding to past classes Y′=⋃i=1t−1Yi\mathcal{Y}' = \bigcup_{i=1}^{t-1} Y_i for both gt−1g^{t-1} and gtg^t, yielding distributions p^t−1\hat{p}^{t-1} and p^t\hat{p}^t, which are penalized via Kullback-Leibler (KL) divergence: Ldis=KL(p^t−1∥p^t)\mathcal{L}_{dis} = \text{KL}(\hat{p}^{t-1} \parallel \hat{p}^t)

  5. Knowl 5 — Computational Complexity and Efficiency Bound of Dense Network Expansion

    theoretical result

    For an incremental vision transformer processing PP patch tokens of feature dimension DD per head across TT sequential tasks, the computational complexity per layer is:

    1. Independent Attention Network Expansion (IA-NE): With TT task experts each possessing HH spatial attention heads, spatial self-attention requires O(THP2D2)\mathcal{O}(T H P^2 D^2) operations and standard MLP feature mixing requires O(TPH2D2)\mathcal{O}(T P H^2 D^2) operations, giving total complexity: OIA=O(THPD2(P+H))\mathcal{O}_{\text{IA}} = \mathcal{O}(T H P D^2(P + H))

    2. Dense Network Expansion (DNE): Spatial attention remains within-task with complexity O(THP2D2)\mathcal{O}(T H P^2 D^2). The Task Attention Block queries representations from all prior tasks per patch, adding O(T2H2PD2)\mathcal{O}(T^2 H^2 P D^2) operations. The total complexity is: ODNE=O(THPD2(P+TH))\mathcal{O}_{\text{DNE}} = \mathcal{O}(T H P D^2(P + T H))

    3. Efficiency Regime with Single-Head Experts: Because DNE reuses prior features across tasks, setting H=1H = 1 head per added task expert suffices for competitive accuracy, whereas IA-NE requires H>1H > 1 (e.g., H=12H = 12). With H=1H=1, DNE complexity is O(TPD2(P+T))\mathcal{O}(T P D^2(P + T)).

    The compute ratio between DNE (with 1 head per task) and IA-NE (with HH heads per task) is: Ratio=P+TH(P+H)\text{Ratio} = \frac{P + T}{H(P + H)} DNE is more computationally efficient than IA-NE whenever: T<H2+(H−1)PT < H^2 + (H - 1)P For standard Vision Transformer dimensions (H=12H = 12 heads, P=196P = 196 patches), this condition holds for up to T<122+11×196=2300T < 12^2 + 11 \times 196 = 2300 tasks.

  6. Knowl 6 — Benchmark Comparison on CIFAR100 and ImageNet100 for Step Size 10

    data/table

    Performance of class incremental learning methods evaluated on CIFAR100 and ImageNet100. The protocol starts with an initial task of 50 classes, followed by 5 incremental tasks of Ns=10N_s = 10 classes each (6 tasks total), with an exemplar buffer memory M\mathcal{M} of 2,000 images.

    Evaluation metrics:

    • Last Accuracy (LA, %): Average classification accuracy over all 100 classes after the final task step (AMA_M).
    • Average Incremental Accuracy (AA, %): Mean accuracy over all task steps (1M∑i=1MAi\frac{1}{M} \sum_{i=1}^M A_i).
    • Joint-to-CIL Difference (DD, %): Difference between the upper-bound joint training accuracy and the CIL model's LA (AM,joint−AM,modelA_{M,\text{joint}} - A_{M,\text{model}}).
    • FLOPs (FF, GigaFLOPs): Total computation for 100-class inference at the final step.
    Method CIFAR100, Ns=10N_s=10 ImageNet100, Ns=10N_s=10
    LA ↑\uparrow AA ↑\uparrow D↓D \downarrow F↓F \downarrow LA ↑\uparrow AA ↑\uparrow D↓D \downarrow F↓F \downarrow
    Joint (Transformer) 76.12 - 0 1.38G 79.12 - 0 1.38G
    Joint (ResNet18) 80.41 - 0 1.12G 81.20 - 0 1.12G
    iCaRL 44.72 59.32 31.40 1.12G 42.84 55.65 38.36 1.12G
    PODNet 52.46 66.41 27.95 1.12G 63.46 73.57 17.74 1.12G
    DER 63.78 71.69 16.63 6.68G 70.40 76.90 10.80 6.68G
    FOSTER 63.31 72.20 17.10 1.12G 67.68 75.85 13.52 1.12G
    Dytox 64.06 71.55 12.06 1.38G 68.84 75.54 10.28 1.38G
    DNE-1head 68.04 73.68 8.08 2.68G 72.30 78.09 6.82 2.68G
    DNE-2heads 69.73 74.61 6.39 3.10G 73.64 78.88 5.48 3.10G
    DNE-4heads 70.04 74.86 6.08 4.02G 73.58 78.56 5.54 4.02G

    DNE-1head achieves 68.04%68.04\% LA on CIFAR100 (+3.98% over Dytox) and 72.30%72.30\% LA on ImageNet100 (+1.90% over DER) while requiring less than half the FLOPs of DER (2.68G vs 6.68G). In terms of gap DD to upper-bound joint training, DNE reduces the degradation by up to 50%50\% relative to Dytox.

  7. Knowl 7 — CIFAR100 Incremental Performance Across Varying Task Step Sizes

    data/table

    Performance comparison on CIFAR100 with step sizes Ns=5N_s = 5 (11 tasks total) and Ns=25N_s = 25 (3 tasks total), using an initial task of 50 classes and a memory buffer of 2,000 images.

    Method CIFAR100, Ns=5N_s=5 CIFAR100, Ns=25N_s=25
    LA ↑\uparrow AA ↑\uparrow D↓D \downarrow F↓F \downarrow LA ↑\uparrow AA ↑\uparrow D↓D \downarrow F↓F \downarrow
    Joint (Transformer) 76.12 - 0 1.38G 76.12 - 0 1.38G
    Joint (ResNet18) 80.41 - 0 1.12G 80.41 - 0 1.12G
    iCaRL 42.89 55.23 33.23 1.12G 53.79 66.80 26.62 1.12G
    PODNet 48.18 60.69 27.94 1.12G 61.54 71.45 18.87 1.12G
    DER 60.73 70.42 15.39 12.19G 69.08 74.82 11.33 3.32G
    FOSTER 48.18 60.69 27.94 1.12G 70.14 76.33 10.27 1.12G
    Dytox 58.59 68.31 17.53 1.38G 69.29 74.10 6.83 1.38G
    DNE-1head 69.10 74.03 7.02 5.71G 67.61 73.27 8.51 1.19G
    DNE-2heads 69.72 74.27 6.40 7.39G 68.99 73.91 7.13 1.27G
    DNE-4heads 69.43 74.20 6.69 10.75G 71.47 75.70 4.65 1.45G

    The margin of DNE over existing SOTA methods is largest for smaller step sizes (Ns=5N_s = 5), where DNE-2heads outperforms DER by 8.99%8.99\% in LA while using roughly 40%40\% fewer FLOPs (7.39G vs 12.19G). For large step sizes (Ns=25N_s = 25), adding 4 heads provides the capacity needed for 25 new classes per task, outperforming FOSTER by 1.33%1.33\% in LA.

  8. Knowl 8 — Scalability and Stability Over Long Task Sequences

    empirical result

    When evaluated on a long incremental task sequence on CIFAR100 with step size Ns=2N_s = 2 (50 initial classes followed by 25 incremental tasks of 2 classes each, for a total of 26 tasks, with buffer size S=2000S = 2000):

    1. Distillation Methods (FOSTER): Maintain approximately constant computational cost (approx1.12G\\approx 1.12\text{G} FLOPs) across all tasks, but suffer from catastrophic forgetting as the task count increases, showing marked accuracy degradation toward step 26.
    2. Standard Network Expansion (DER): Avoids catastrophic forgetting by adding a full subnetwork backbone per task, but incurs linear FLOP scaling with task count, consuming over 26 times the baseline FLOPs (approx29G\\approx 29\text{G} FLOPs by step 26).
    3. Dense Network Expansion (DNE): Employs dense cross-task connections and a single attention head (H=1H = 1) per task expert. Because each expert is compact, DNE maintains a slow rate of FLOP growth that remains substantially below DER throughout all 26 tasks, while achieving consistently higher accuracy than both DER and FOSTER at every incremental task step.

Coverage note — Deliberately omitted general background on prior CIL architectures (such as standard descriptions of iCaRL, PODNet, DER, FOSTER, and Dytox) and standard Vision Transformer preliminaries as they represent prior work.

References

  1. 1.Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars. Expert gate: Lifelong learning with a network of experts. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3366–3375, 2017.
  2. 2.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738, 2021.
  3. 3.Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. Dark experience for general continual learning: a strong, simple baseline. Advances in neural information processing systems, 33:15920–15930, 2020.
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  5. 5.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  8. 8.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In European Conference on Computer Vision, pages 86–102. Springer, 2020.
  9. 9.Arthur Douillard, Alexandre Ramé, Guillaume Couairon, and Matthieu Cord. Dytox: Transformers for continual learning with dynamic token expansion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9285–9295, 2022.
  10. 10.Siavash Golkar, Michael Kagan, and Kyunghyun Cho. Continual learning via neural pruning. arXiv preprint arXiv:1903.04476, 2019.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  12. 12.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 831–839, 2019.
  13. 13.Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, Jiashi Feng, and Yannis Kalantidis. Decoupling representation and classifier for long-tailed recognition. arXiv preprint arXiv:1910.09217, 2019.
  14. 14.Minsoo Kang, Jaeyoo Park, and Bohyung Han. Class-incremental learning by knowledge distillation with adaptive feature consolidation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16071–16080, 2022.
  15. 15.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  16. 16.Yajing Kong, Liu Liu, Zhen Wang, and Dacheng Tao. Balancing stability and plasticity through advanced null space in continual learning. arXiv preprint arXiv:2207.12061, 2022.
  17. 17.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  18. 18.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017.
  19. 19.David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. Advances in neural information processing systems, 30, 2017.
  20. 20.Arun Mallya and Svetlana Lazebnik. Packnet: Adding multiple tasks to a single network by iterative pruning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 7765–7773, 2018.
  21. 21.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  22. 22.Jathushan Rajasegaran, Munawar Hayat, Salman H Khan, Fahad Shahbaz Khan, and Ling Shao. Random path selection for continual learning. Advances in Neural Information Processing Systems, 32, 2019.
  23. 23.Amal Rannen, Rahaf Aljundi, Matthew B Blaschko, and Tinne Tuytelaars. Encoder based lifelong learning. In Proceedings of the IEEE International Conference on Computer Vision, pages 1320–1328, 2017.
  24. 24.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2001–2010, 2017.
  25. 25.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  26. 26.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  27. 27.Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. Foster: Feature boosting and compression for class-incremental learning. arXiv preprint arXiv:2204.04662, 2022.
  28. 28.Shipeng Wang, Xiaorong Li, Jian Sun, and Zongben Xu. Training networks in null space of feature covariance for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 184–193, 2021.
  29. 29.Zhen Wang, Liu Liu, Yiqun Duan, Yajing Kong, and Dacheng Tao. Continual learning with lifelong vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 171–181, 2022.
  30. 30.Tz-Ying Wu, Gurumurthy Swaminathan, Zhizhong Li, Avinash Ravichandran, Nuno Vasconcelos, Rahul Bhotika, and Stefano Soatto. Class-incremental learning with strong pre-trained models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9601–9610, 2022.
  31. 31.Shipeng Yan, Jiangwei Xie, and Xuming He. Der: Dynamically expandable representation for class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3014–3023, 2021.
  32. 32.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In International Conference on Machine Learning, pages 3987–3995. PMLR, 2017.

Citation

MLA
Hu, Z., et al. “Dense Network Expansion for Class Incremental Learning”. arXiv, 2023, http://arxiv.org/abs/2303.12696v1.
APA
Hu, Z., Li, Y., Lyu, J., Gao, D., & Vasconcelos, N. (2023). Dense Network Expansion for Class Incremental Learning. arXiv. http://arxiv.org/abs/2303.12696v1
Chicago
Hu, Z., Y. Li, J. Lyu, D. Gao, and N. Vasconcelos. 2023. “Dense Network Expansion for Class Incremental Learning”. arXiv. http://arxiv.org/abs/2303.12696v1.
Harvard
Hu, Z. et al. (2023) “Dense Network Expansion for Class Incremental Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.12696v1.
Vancouver
1. Hu Z, Li Y, Lyu J, Gao D, Vasconcelos N (2023) Dense Network Expansion for Class Incremental Learning. arXiv

BibTeX

@article{hu2023dense,
  title = {Dense Network Expansion for Class Incremental Learning},
  author = {Hu, Zhiyuan and Li, Yunsheng and Lyu, Jiancheng and Gao, Dashan and Vasconcelos, Nuno},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.12696v1},
  eprint = {2303.12696}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE