Lifelong Learning with Dynamically Expandable Networks
Jaehong YoonEunho YangJeongtae LeeSung Ju Hwang
Proposes Dynamically Expandable Networks, an online lifelong learning method that dynamically scales capacity and selectively retrains or duplicates units to prevent catastrophic forgetting while achieving batch-level performance with significantly fewer parameters.
Real-world artificial intelligence applications, such as autonomous vehicles and robotics, require systems to learn sequentially from incoming streams of new tasks. Traditional deep neural networks struggle in this continuous learning setting due to catastrophic forgetting, where learning new information degrades performance on previously learned tasks. Existing approaches either rigidly restrict model updates to preserve old knowledge—which hurts learning on new tasks—or continuously expand the model with fixed capacity, resulting in bloated, computationally expensive architectures.
The article introduces and evaluates the Dynamically Expandable Network, a deep learning architecture designed to learn sequentially arriving tasks efficiently. The objective of the article is to demonstrate that this framework can dynamically adjust its size, retain previously acquired knowledge without degradation, and match or exceed the performance of models trained on all tasks simultaneously while using substantially fewer parameters.
The researchers evaluated the framework across feedforward and convolutional neural networks on multiple benchmark image datasets, including MNIST-Variation, CIFAR-100, and Animals with Attributes (spanning sequences of up to 50 tasks). The framework operates through three non-technical steps: first, it selectively retrains only the subnetwork of existing parameters relevant to the new task; second, if the current capacity fails to adequately represent the task, it dynamically expands the network and removes redundant units; and third, it detects units whose internal representations drift too much and splits or duplicates them, assigning timestamps to ensure earlier tasks only rely on representations that existed when they were learned.
The evaluation produced four key findings. First, the framework substantially outperformed existing continual learning baselines across all benchmark datasets in classification accuracy. Second, it achieved comparable performance to models trained independently for each task while utilizing only 11.9% to 60.3% of their total parameter capacity, demonstrating exceptional efficiency. Third, selective retraining significantly reduced computational training time compared to full-network retraining while increasing accuracy by about two percentage points over base models. Fourth, after completing sequential learning, fine-tuning the resulting network structure across all tasks surpassed traditional batch multi-task models by 0.05 to 4.8 percentage points in accuracy.
These findings demonstrate that artificial intelligence systems do not need to choose between severe memory loss and ballooning computational costs. For organizations deploying machine learning at scale, this approach reduces hardware and energy overhead, shortens retraining timelines, and prevents catastrophic service degradation when introducing new operational capabilities. It also provides a practical method for discovering compact, high-performing network architectures automatically, even when all training data is available upfront.
Decision-makers can consider implementing dynamic network expansion for continuous data streaming pipelines and resource-constrained edge devices. However, leaders should note that the framework's performance depends on setting appropriate mathematical thresholds for selective expansion and unit splitting. Before deploying to production environments, organizations should conduct pilot evaluations to tune these hyperparameters on domain-specific data and ensure computational gains hold across specialized workflows.
- Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). This work introduces Progressive Neural Networks, establishing the foundational architecture-expanding baseline for continual learning that Dynamically Expandable Networks build upon and optimize for parameter efficiency.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This seminal paper introduces Learning without Forgetting (LwF), demonstrating how to adapt neural networks to sequential tasks using regularization and knowledge distillation, serving as a direct point of comparison for DEN.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). This foundational paper presents Elastic Weight Consolidation (EWC) to protect vital task parameters, providing the standard fixed-capacity benchmark against which dynamic network expansion is motivated.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). This work develops Synaptic Intelligence to compute online parameter importance trajectories, representing a key regularization baseline for mitigating catastrophic forgetting in sequential learning.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). This paper establishes iCaRL for incremental class learning, highlighting the challenges of feature drift in sequential classification that dynamic network architectures aim to resolve.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). This paper formalizes continuous lifelong learning protocols and introduces Gradient Episodic Memory, providing foundational metrics and replay-based baselines for continual learning.
- Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). This paper introduces cross-stitch units to learn shared versus task-specific representation structures, providing key conceptual background on selective multi-task feature sharing.
- Paper: Continual Lifelong Learning with Neural Networks: A Review, German I. Parisi et al. (2018). This comprehensive review organizes continual learning techniques across regularization, rehearsal, and dynamic architecture families, placing dynamic expansion methods like DEN into broader taxonomic context.
- Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). This work extends efficient lifelong learning under streaming single-pass constraints with A-GEM, contrasting memory-based gradient projection against architectural expansion strategies.
- Paper: PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning, Arun Mallya et al. (2017). PackNet develops an alternative architectural strategy for continual learning by using iterative pruning and parameter allocation to pack sequential tasks into a fixed network capacity.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey evaluates diverse continual learning approaches, including dynamic and parameter-isolation architectures, on standardized task-incremental vision benchmarks.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). This survey provides modern theoretical and structural categorizations of continual learning, systematically surveying architecture-expansion methods alongside replay and regularization paradigms.
- Paper: LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, Bo Liu et al. (2023). LIBERO extends continual learning benchmarks to sequential robotic manipulation tasks, empirically assessing how dynamic architecture techniques scale to complex procedural knowledge transfer.
