Forget-free Continual Learning with Winning Subnetworks
Haeyong KangRusty John Lloyd MinaSultan Rizky Hikmawan MadjidJaehong YoonMark Hasegawa-JohnsonSung Ju HwangChang D. Yoo
Proposes a continual learning framework that completely prevents catastrophic forgetting by isolating and reusing task-specific sparse subnetworks within a single model, compressing the resulting binary masks with Huffman coding to achieve sub-linear memory growth across sequential tasks.
Continual learning requires artificial intelligence systems to learn a sequence of new tasks without degrading performance on previously learned tasks. In conventional deep neural networks, updating network parameters on new data typically causes catastrophic forgetting, where earlier knowledge is overwritten and lost. Existing remedies often expand network architectures as tasks accumulate or store historical data in memory buffers to replay past experiences. However, both approaches lead to substantial memory and computation overhead, making them impractical for resource-constrained or long-term operational environments.
The article aims to introduce and validate a continual learning framework called Winning SubNetworks. This method seeks to prevent catastrophic forgetting completely while maintaining high task accuracy and enforcing sub-linear growth in network memory capacity.
The proposed framework builds on the concept that dense neural networks contain compact, highly performant subnetworks. The approach jointly learns the model weights alongside separate task-adaptive binary masks, which identify and activate the most critical subnetwork connections for each incoming task. Once a subnetwork is selected for a task, its weights are frozen to prevent future tasks from altering them, while new tasks selectively reuse these frozen connections alongside a small set of unallocated weights. To store subnetwork configurations efficiently across long task sequences without ballooning storage, the article applies Huffman coding to compress the accumulated binary masks. The authors validated the method through multi-task classification experiments across six standard benchmark datasets—ranging up to 100 sequential tasks—using diverse network architectures.
The experimental evaluations yielded four principal findings. First, the method achieves complete immunity to catastrophic forgetting, maintaining zero performance loss on prior tasks across all tested datasets. Second, the framework achieved higher overall accuracy than competing continual learning techniques, including top average accuracies of 87.28% on Omniglot Rotation and 71.96% on TinyImageNet. Third, the selective reuse of weights and Huffman coding reduced capacity requirements significantly; for example, on TinyImageNet, the model operated within 48.65% capacity compared to baselines exceeding 100% to 200%. Finally, selective knowledge reuse substantially improved computational efficiency, leading to faster training convergence across all evaluated benchmarks.
These findings demonstrate that artificial intelligence systems can continually adapt to new operational requirements without requiring linear infrastructure expansion or expensive retraining cycles. For organizations deploying machine learning at scale, this method provides a way to reduce compute costs, cut storage footprints, and guarantee that existing capabilities remain stable. The strategy shifts continual learning from resource-heavy memory buffers toward structural subnetwork isolation.
Organizations developing edge computing devices or sequential machine learning pipelines should consider evaluating modular subnetwork masking architectures as a viable strategy to manage capacity constraints. In systems where past tasks share structural commonalities, tuning target capacity ratios—such as reserving roughly 10% capacity per task—provides a balanced trade-off between task accuracy and memory growth. Before wide production deployment, teams should conduct pilot evaluations to assess layer-specific sensitivity and verify performance on complex domain shifts, as the current experiments rely on supervised classification setups where task identities are explicitly known during inference.
- Paper: PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning, Arun Mallya et al. (2017). PackNet’s iterative pruning and task-specific parameter masks provide the closest earlier blueprint for WSN’s reuse of sparse subnetworks to prevent forgetting.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). EWC establishes the continual-learning problem and contrasts weight-importance regularization with WSN’s strategy of isolating task-specific weights.
- Paper: Overcoming catastrophic forgetting with hard attention to the task, Joan Serrà et al. (2018). HAT shows how learned task-specific masks protect neural pathways, making its gating approach a useful precursor to WSN’s binary subnetwork masks.
- Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). Progressive Neural Networks introduce isolated task-specific network components and selective reuse of prior features, clarifying the architectural trade-offs WSN addresses.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey organizes continual-learning methods, including parameter-isolation approaches, and supplies the taxonomy needed to situate WSN’s winning subnetworks.
- Paper: Task-Specific Skill Localization in Fine-tuned Language Models, Abhishek Panigrahi et al. (2023). Task-Specific Skill Localization carries binary-mask isolation into fine-tuned language models, extending WSN’s subnetwork idea to pinpointing compact task-specific capabilities.
- Paper: Localizing Task Information for Improved Model Merging and Compression, Ke Wang et al. (2024). This work uses task-specific binary masks to localize information in merged models, extending the mask-based separation central to WSN toward model merging and compression.
- Paper: Dense Network Expansion for Class Incremental Learning, Zhiyuan Hu et al. (2023). Dense Network Expansion continues the search for scalable task-specific capacity by reusing features across sequential tasks instead of allowing model growth to run unchecked.
- Paper: Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning, Kai Zhu et al. (2022). This approach extends architectural adaptation for continual learning with temporary branches that absorb new classes and are then consolidated without increasing model size.
