Dense Network Expansion for Class Incremental Learning
Zhiyuan HuYunsheng LiJiancheng LyuDashan GaoNuno Vasconcelos
Proposes a dense network expansion method for class incremental learning that transfers knowledge across frozen task experts using a cross-task attention block, achieving superior accuracy while significantly curbing parameter growth.
Modern computer vision models face catastrophic forgetting, an issue where sequential training on new visual classes causes the system to erase previously acquired knowledge. While network expansion methods prevent forgetting by adding specialized sub-networks for each new task, they trigger rapid, unsustainable growth in computational footprint and model size. The article designs and evaluates Dense Network Expansion, an architecture that maintains past recognition accuracy while substantially curbing resource growth by sharing and reusing intermediate features across tasks.
Dense Network Expansion decouples spatial image analysis from cross-task knowledge transfer using Vision Transformers. Instead of duplicating entire large networks or entangling spatial and task tokens—which causes attention dilution—the framework freezes previous components and incorporates a lightweight task attention block directly into the feature-mixing stage. The authors benchmarked this design against leading continuous learning methods on standardized vision datasets, assessing accuracy, computational operations, and scalability across multi-task schedules up to 26 sequential stages.
The findings show that Dense Network Expansion outperforms existing class-incremental learning methods. When adding single-head task experts, the system achieved a 68.04% final accuracy on CIFAR-100, surpassing the previous top performer by nearly 4%, while outperforming the leading baseline on ImageNet-100 by 1.9%. Crucially, the approach narrows the performance gap relative to an ideal, fully retrained model by up to 50%. In extended sequences spanning 26 incremental steps, the architecture maintained high accuracy while consuming a fraction of the computational operations required by traditional network expansion baselines.
These results demonstrate that continuous learning does not require choosing between catastrophic forgetting and runaway infrastructure costs. By enabling compact sub-networks to selectively query and integrate older representations, organizations can deploy edge and cloud vision systems that continually scale to new classes with predictable computational overhead. Teams implementing incremental learning pipelines should adopt modular cross-task feature sharing rather than full-backbone expansion or joint spatial-task attention mechanisms.
The reported advantages rely on benchmarks where early tasks establish strong initial representations, and tasks with larger class additions may still require increasing the capacity of task branches. Nonetheless, the experimental evidence provides strong confidence that dense, decoupled attention provides a viable, resource-efficient foundation for sequential vision learning.
- Paper: Lifelong Learning with Dynamically Expandable Networks, Jaehong Yoon et al. (2017). This foundational work on dynamically expandable networks establishes the architectural paradigm of adding and freezing sub-networks during sequential task learning that Dense Network Expansion directly optimizes and refines.
- Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). Progressive Neural Networks introduced the core mechanism of freezing previous columns and transferring intermediate features to new task networks via lateral connections, a foundational precursor to Dense Network Expansion.
- Paper: Densely Connected Convolutional Networks, Gao Huang et al. (2017). DenseNets pioneered the dense feature-reuse connectivity pattern that Dense Network Expansion adapts across temporal task boundaries to curb parameter growth.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This paper establishes the standard Learning without Forgetting formulation and knowledge distillation framework for multi-task and incremental vision adaptation.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). L2P provides essential background on leveraging frozen transformer backbones and modular dynamic parameter tuning for class-incremental vision tasks without buffer storage.
- Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). This paper defines the standard class-incremental learning benchmark protocols and cross-distillation evaluation paradigms that modern incremental vision models are measured against.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey provides a comprehensive taxonomy of parameter-isolation and expansion techniques versus regularized continuous learning approaches for classification.
- Paper: VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding, Yi Xin et al. (2024). VMT-Adapter advances parameter-efficient vision transformer sharing by employing a unified cross-task adapter framework for multi-task dense scene understanding over a single frozen backbone pass.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). This comprehensive survey contextualizes modular architectural expansion and representation-based continual learning methods within the broader theoretical landscape of lifelong learning.
