Lifelong Learning with Dynamically Expandable Networks

Jaehong YoonEunho YangJeongtae LeeSung Ju Hwang

article2017ICLR1,482 citations

Proposes Dynamically Expandable Networks, an online lifelong learning method that dynamically scales capacity and selectively retrains or duplicates units to prevent catastrophic forgetting while achieving batch-level performance with significantly fewer parameters.

Listen

Real-world artificial intelligence applications, such as autonomous vehicles and robotics, require systems to learn sequentially from incoming streams of new tasks. Traditional deep neural networks struggle in this continuous learning setting due to catastrophic forgetting, where learning new information degrades performance on previously learned tasks. Existing approaches either rigidly restrict model updates to preserve old knowledge—which hurts learning on new tasks—or continuously expand the model with fixed capacity, resulting in bloated, computationally expensive architectures.

The article introduces and evaluates the Dynamically Expandable Network, a deep learning architecture designed to learn sequentially arriving tasks efficiently. The objective of the article is to demonstrate that this framework can dynamically adjust its size, retain previously acquired knowledge without degradation, and match or exceed the performance of models trained on all tasks simultaneously while using substantially fewer parameters.

The researchers evaluated the framework across feedforward and convolutional neural networks on multiple benchmark image datasets, including MNIST-Variation, CIFAR-100, and Animals with Attributes (spanning sequences of up to 50 tasks). The framework operates through three non-technical steps: first, it selectively retrains only the subnetwork of existing parameters relevant to the new task; second, if the current capacity fails to adequately represent the task, it dynamically expands the network and removes redundant units; and third, it detects units whose internal representations drift too much and splits or duplicates them, assigning timestamps to ensure earlier tasks only rely on representations that existed when they were learned.

The evaluation produced four key findings. First, the framework substantially outperformed existing continual learning baselines across all benchmark datasets in classification accuracy. Second, it achieved comparable performance to models trained independently for each task while utilizing only 11.9% to 60.3% of their total parameter capacity, demonstrating exceptional efficiency. Third, selective retraining significantly reduced computational training time compared to full-network retraining while increasing accuracy by about two percentage points over base models. Fourth, after completing sequential learning, fine-tuning the resulting network structure across all tasks surpassed traditional batch multi-task models by 0.05 to 4.8 percentage points in accuracy.

These findings demonstrate that artificial intelligence systems do not need to choose between severe memory loss and ballooning computational costs. For organizations deploying machine learning at scale, this approach reduces hardware and energy overhead, shortens retraining timelines, and prevents catastrophic service degradation when introducing new operational capabilities. It also provides a practical method for discovering compact, high-performing network architectures automatically, even when all training data is available upfront.

Decision-makers can consider implementing dynamic network expansion for continuous data streaming pipelines and resource-constrained edge devices. However, leaders should note that the framework's performance depends on setting appropriate mathematical thresholds for selective expansion and unit splitting. Before deploying to production environments, organizations should conduct pilot evaluations to tune these hyperparameters on domain-specific data and ensure computational gains hold across specialized workflows.

arXiv: 1708.01547
  • Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). This work introduces Progressive Neural Networks, establishing the foundational architecture-expanding baseline for continual learning that Dynamically Expandable Networks build upon and optimize for parameter efficiency.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This seminal paper introduces Learning without Forgetting (LwF), demonstrating how to adapt neural networks to sequential tasks using regularization and knowledge distillation, serving as a direct point of comparison for DEN.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). This foundational paper presents Elastic Weight Consolidation (EWC) to protect vital task parameters, providing the standard fixed-capacity benchmark against which dynamic network expansion is motivated.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). This work develops Synaptic Intelligence to compute online parameter importance trajectories, representing a key regularization baseline for mitigating catastrophic forgetting in sequential learning.
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). This paper establishes iCaRL for incremental class learning, highlighting the challenges of feature drift in sequential classification that dynamic network architectures aim to resolve.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). This paper formalizes continuous lifelong learning protocols and introduces Gradient Episodic Memory, providing foundational metrics and replay-based baselines for continual learning.
  • Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). This paper introduces cross-stitch units to learn shared versus task-specific representation structures, providing key conceptual background on selective multi-task feature sharing.
Cover for Lifelong Learning with Dynamically Expandable Networks

Abstract

We propose a novel deep network architecture for lifelong learning which we refer to as Dynamically Expandable Network (DEN), that can dynamically decide its network capacity as it trains on a sequence of tasks, to learn a compact overlapping knowledge sharing structure among tasks. DEN is efficiently trained in an online manner by performing selective retraining, dynamically expands network capacity upon arrival of each task with only the necessary number of units, and effectively prevents semantic drift by splitting/duplicating units and timestamping them. We validate DEN on multiple public datasets under lifelong learning scenarios, on which it not only significantly outperforms existing lifelong learning methods for deep networks, but also achieves the same level of performance as the batch counterparts with substantially fewer number of parameters. Further, the obtained network fine-tuned on all tasks obtained significantly better performance over the batch models, which shows that it can be used to estimate the optimal network structure even when all tasks are available in the first place.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Incremental Learning of a Dynamically Expandable Network
  • 4 Experiment
  • 4.1 Quantitative evaluation
  • 5 Conclusion
  • References
  • A
  • A.1 Results on Permuted MNIST

Knowls

  1. Knowl 1 — Dynamically Expandable Network (DEN) Framework for Lifelong Learning

    model/method

    The Dynamically Expandable Network (DEN) is a deep neural network framework for lifelong learning over a sequence of tasks t=1,…,Tt = 1, \dots, T with corresponding task datasets Dt={(xi,yi)}i=1Nt\mathcal{D}_t = \{(x_i, y_i)\}_{i=1}^{N_t}, where prior training datasets D1,…,Dt−1\mathcal{D}_1, \dots, \mathcal{D}_{t-1} are inaccessible at step tt. Rather than fixing network capacity or fully freezing/retraining weights, DEN dynamically determines parameter sharing and layer capacity through three core operations:

    1. Selective Retraining: Fits a sparse connection from the top layer to identify the task-relevant subnetwork, retraining only those selected weights while keeping non-selected parameters fixed.
    2. Dynamic Network Expansion: If the loss on task tt after selective retraining exceeds a preset threshold τ\tau, adds a candidate pool of kk hidden units at each layer and applies group-sparsity regularization to prune unnecessary units.
    3. Network Split and Duplication: If incoming weights to an existing neuron change by an ℓ2\ell_2 distance exceeding a drift threshold σ\sigma during adaptation to task tt, the neuron is duplicated to preserve its legacy representation for earlier tasks while providing a decoupled copy for the new task.
    4. Timestamped Inference: Each unit is tagged with the task index tt at which it was introduced, ensuring that evaluating an earlier task t′t' uses only units and weights created up to stage t′t' (t′≤tt' \le t).
  2. Knowl 2 — Incremental Learning Algorithm for Dynamically Expandable Networks

    algorithm

    The sequential training procedure of DEN trains the initial task using ℓ1\ell_1 regularization to enforce sparsity, and subsequently alternates between selective retraining, threshold-driven capacity expansion, and drift-driven unit splitting for each subsequent task.

    Input: Task datasets (D1,…,DT)(\mathcal{D}_1, \dots, \mathcal{D}_T), expansion threshold τ\tau, split threshold σ\sigma, sparsity coefficient μ\mu
    Output: Sequence of model parameters W1,…,WTW^1, \dots, W^T
    for t=1,…,Tt = 1, \dots, T do
        if t=1t = 1 then
            Initialize network and optimize W1←arg⁡min⁡W1L(W1;D1)+μ∑l=1L∥Wl1∥1W^1 \leftarrow \arg\min_{W^1} \mathcal{L}(W^1; \mathcal{D}_1) + \mu \sum_{l=1}^L \|W_l^1\|_1
        else
            Wt←SelectiveRetraining(Wt−1,Dt)W^t \leftarrow \text{SelectiveRetraining}(W^{t-1}, \mathcal{D}_t)
            Compute task loss Lt=L(Wt;Dt)\mathcal{L}_t = \mathcal{L}(W^t; \mathcal{D}_t)
            if Lt>τ\mathcal{L}_t > \tau then
                Wt←DynamicExpansion(Wt,Dt,τ)W^t \leftarrow \text{DynamicExpansion}(W^t, \mathcal{D}_t, \tau)
            Wt←Split(Wt,Dt,σ)W^t \leftarrow \text{Split}(W^t, \mathcal{D}_t, \sigma)

    Here, LL is the number of layers, L(⋅)\mathcal{L}(\cdot) denotes the task-specific loss function, and Wt={Wlt}l=1LW^t = \{W_l^t\}_{l=1}^L represents the set of all layer weight matrices at step tt.

  3. Knowl 3 — Selective Retraining via Breadth-First Subnetwork Identification

    algorithm

    Selective retraining prevents negative transfer and saves computation by identifying and updating only the existing parameters that have active computational paths to the output of the newly arrived task tt.

    Input: Dataset Dt\mathcal{D}_t, previous parameters Wt−1W^{t-1}, regularizer coefficient μ\mu
    Output: Updated parameter tensor WtW^t
    Initialize active neuron set S={ot}S = \{o_t\} where oto_t is the output neuron for task tt
    Solve for topmost task connection weights: WL,tt←arg⁡min⁡WL,tL(WL,t;W1:L−1t−1,Dt)+μ∥WL,t∥1W_{L,t}^t \leftarrow \arg\min_{W_{L,t}} \mathcal{L}(W_{L,t}; W_{1:L-1}^{t-1}, \mathcal{D}_t) + \mu \|W_{L,t}\|_1
    for each hidden unit ii at layer L−1L-1 do
        if connection (i,ot)(i, o_t) in WL,tt≠0W_{L,t}^t \neq 0 then
            Add unit ii to SS
    for layer l=L−1,…,1l = L-1, \dots, 1 do
        for each neuron ii at layer l−1l-1 do
            if exists neuron j∈Sj \in S at layer ll such that Wl,ijt−1≠0W_{l,ij}^{t-1} \neq 0 then
                Add unit ii to SS
    Let WStW_S^t be all weights between units in SS, and WSct−1W_{S^c}^{t-1} be all remaining frozen weights
    Optimize active subnetwork: WSt←arg⁡min⁡WSL(WS;WSct−1,Dt)+μ∥WS∥2W_S^t \leftarrow \arg\min_{W_S} \mathcal{L}(W_S; W_{S^c}^{t-1}, \mathcal{D}_t) + \mu \|W_S\|_2

    By optimizing only WStW_S^t under ℓ2\ell_2 regularization, parameters unaffected by task tt remain exactly unchanged, preventing catastrophic forgetting on unshared subcomponents.

  4. Knowl 4 — Dynamic Network Expansion with Group Sparsity Regularization

    algorithm

    When selective retraining yields a loss Lt>τ\mathcal{L}_t > \tau, the network lacks the capacity to represent the new task. DEN introduces kk candidate units per layer and selects the optimal subset via group sparsity.

    Input: Dataset Dt\mathcal{D}_t, current parameter tensor WtW^t, expansion size kk, regularizers μ,γ\mu, \gamma, threshold τ\tau
    Output: Dynamically expanded parameter tensor WtW^t
    if L(Wt;Dt)>τ\mathcal{L}(W^t; \mathcal{D}_t) > \tau then
        for each layer l∈{1,…,L−1}l \in \{1, \dots, L-1\} do
            Add kk candidate units hlNh_l^{\mathcal{N}} to layer ll, expanding weight matrix to Wl=[Wlt−1;WlN]W_l = [W_l^{t-1}; W_l^{\mathcal{N}}]
        Solve group-sparse optimization over added weights WNW^{\mathcal{N}}:
            min⁡WNL(WN;Wt−1,Dt)+μ∥WN∥1+γ∑l=1L−1∑g∈G∥Wl,gN∥2\min_{W^{\mathcal{N}}} \mathcal{L}(W^{\mathcal{N}}; W^{t-1}, \mathcal{D}_t) + \mu \|W^{\mathcal{N}}\|_1 + \gamma \sum_{l=1}^{L-1} \sum_{g \in \mathcal{G}} \|W_{l,g}^{\mathcal{N}}\|_2
        for layer l=L−1,…,1l = L-1, \dots, 1 do
            Identify added units in hlNh_l^{\mathcal{N}} whose incoming group weights ∥Wl,gN∥2=0\|W_{l,g}^{\mathcal{N}}\|_2 = 0
            Remove those unselected units from the network

    For fully connected layers, each group g∈Gg \in \mathcal{G} corresponds to all incoming connection weights to a candidate neuron. For convolutional layers, each group corresponds to the weights of a convolutional filter (activation map).

  5. Knowl 5 — Network Unit Splitting and Duplication to Prevent Semantic Drift

    algorithm

    When weights shared with older tasks undergo large parameter shifts to accommodate task tt, semantic drift occurs. DEN detects this drift via ℓ2\ell_2 distance and duplicates affected neurons.

    Input: Parameter tensor Wt−1W^{t-1}, dataset Dt\mathcal{D}_t, drift threshold σ\sigma, regularization weight λ\lambda
    Output: Updated parameter tensor WtW^t
    Compute preliminary weights: W~t←arg⁡min⁡WL(W;Dt)+λ∥W−Wt−1∥22\widetilde{W}^t \leftarrow \arg\min_W \mathcal{L}(W; \mathcal{D}_t) + \lambda \|W - W^{t-1}\|_2^2
    for each hidden unit ii in the network do
        Compute parameter drift ρit=∥wit−wit−1∥2\rho_i^t = \|w_i^t - w_i^{t-1}\|_2 where wiw_i is the incoming weight vector to unit ii
        if ρit>σ\rho_i^t > \sigma then
            Duplicate unit ii into a new unit i′i'
            Assign legacy weights wit−1w_i^{t-1} to ii and modified weights wi′tw_{i'}^t to i′i'
    Re-optimize the network parameters initialized from W~t\widetilde{W}^t using min⁡WL(W;Dt)+λ∥W−Wt−1∥22\min_W \mathcal{L}(W; \mathcal{D}_t) + \lambda \|W - W^{t-1}\|_2^2 to obtain WtW^t

    Splitting unit ii creates two specialized features: one dedicated to older tasks requiring the original representation, and a distinct unit i′i' that can adapt to new tasks without degrading old task accuracy.

  6. Knowl 6 — Timestamped Inference Mechanism in DEN

    model/method

    In DEN, each hidden unit jj receives a timestamp index {z}j=t\{z\}_j = t, recording the exact task learning stage tt during which it was added (either via dynamic expansion or unit duplication).

    During inference for task tt, the forward pass only computes activations for units whose timestamp satisfies {z}j≤t\{z\}_j \le t. Any unit introduced at a later task stage t′>tt' > t is masked out. Unlike methods that permanently freeze all weights of older subnetworks, timestamped inference allows early tasks to utilize representations that were further updated in subsequent stages (provided those units were not split due to excessive drift), while strictly insulating early tasks from newly added network units.

  7. Knowl 7 — Experimental Configurations and Benchmark Datasets for DEN

    experimental setup

    DEN was evaluated across four lifelong learning benchmark datasets with feedforward and convolutional architectures:

    • MNIST-Variation: 62,000 images of handwritten digits rotated at arbitrary angles with background noise (1,000 train, 200 validation, 5,000 test images per class), configured as 10 one-vs-rest binary tasks. Base network: 2-layer MLP with 312-128 ReLU units.
    • CIFAR-100: 60,000 images across 100 classes (500 train, 100 test per class), grouped into 10 sequential tasks of 10 binary classification subtasks each. Base network: AlexNet variant with 5 conv layers (64–128–256–256–12864\text{--}128\text{--}256\text{--}256\text{--}128 filters of size 5×55 \times 5) and 3 fully-connected layers (384–192–100384\text{--}192\text{--}100 neurons).
    • AWA (Animals with Attributes): 30,475 images across 50 animal classes (50 binary tasks) using 500-dimensional PCA-reduced DeCAF features with 30/30/30 train/val/test splits. Base network: 2-layer MLP (312–128312\text{--}128 neurons).
    • Permuted MNIST: 70,000 images across 10 sequential 10-class tasks, where each task applies a distinct random pixel permutation.

    Baselines included single-task learning (DNN-STL), joint multi-task learning (DNN-MTL), continuous fine-tuning with ℓ2\ell_2 penalty (DNN-L2), Elastic Weight Consolidation (DNN-EWC), and Progressive Neural Networks (DNN-Progressive).

  8. Knowl 8 — Lifelong Learning Performance and Capacity Comparisons across Benchmarks

    empirical result

    Across benchmark datasets, DEN matched or outperformed lifelong learning baselines while maintaining compact parameter capacity:

    • MNIST-Variation (T=10T=10): DEN achieved an average per-task AUROC of 0.8131 at t=Tt=T, outperforming DNN-STL (0.7963), DNN-MTL (0.8047), DNN-Progressive (0.7817), DNN-EWC (0.7480), and DNN-L2 (0.7440). DEN matched STL performance using only 18.0%18.0\% of STL capacity.
    • CIFAR-100 (T=10T=10): DEN achieved 0.9225 AUROC, outperforming DNN-MTL (0.8890), DNN-Progressive (0.8819), DNN-EWC (0.8134), and DNN-L2 (0.7830), and approaching separate CNN-STL models (0.9345) using 60.3%60.3\% of its capacity.
    • AWA (T=50T=50): DEN achieved 0.8913 AUROC, significantly outperforming DNN-MTL (0.7222), DNN-Progressive (0.6465), DNN-EWC (0.5604), and DNN-L2 (0.5454), while matching DNN-STL (0.9064) using only 11.9%11.9\% of STL capacity.
    • Permuted MNIST (T=10T=10): DEN achieved 0.9965 AUROC, outperforming DNN-Progressive (0.9912), DNN-EWC (0.9732), and DNN-L2 (0.9636) with only 1.39×1.39\times the base network capacity.
    • DEN-Finetune: When the structure identified by DEN was fine-tuned on all tasks simultaneously, it outperformed batch DNN-MTL across all datasets by 0.05%p0.05\%\text{p} to 4.8%p4.8\%\text{p} AUROC.
  9. Knowl 9 — Ablation Analysis of Selective Retraining and Dynamic Network Expansion

    data/table

    An ablation study on the MNIST-Variation dataset evaluated the individual contributions of selective retraining and dynamic group-sparse expansion against fixed-size expansion baselines:

    Model AUROC Capacity (Relative to MTL)
    DNN-L2 0.7440±0.010.7440 \pm 0.01 1.001.00
    DNN-L2-Selective 0.7499±0.010.7499 \pm 0.01 1.001.00
    DEN-Constant (k=13k=13) 0.7611±0.010.7611 \pm 0.01 1.551.55
    DEN-Constant (k=20k=20) 0.7580±0.010.7580 \pm 0.01 1.891.89
    DEN-Dynamic 0.7684±0.010.7684 \pm 0.01 1.56±0.011.56 \pm 0.01

    Key takeaways:

    1. Selective retraining (DNN-L2-Selective) improved AUROC over standard ℓ2\ell_2 retraining (DNN-L2) while requiring significantly less GPU computation time than full network retraining, because it selectively preserved unaffected parameters and focused on generic lower-layer features.
    2. Dynamic group-sparse expansion (DEN-Dynamic) achieved the highest AUROC (0.76840.7684) while adding fewer parameters (1.56×1.56\times capacity) than uniform expansion with k=20k=20 (1.89×1.89\times capacity, 0.75800.7580 AUROC), demonstrating that unconstrained uniform expansion causes overfitting.
  10. Knowl 10 — Ablation of Unit Splitting and Timestamping on Semantic Drift

    empirical result

    Tracking the AUROC trajectory of early tasks (t=1,4,7t=1, 4, 7) on MNIST-Variation as subsequent tasks were trained revealed the roles of splitting and timestamping in preventing catastrophic forgetting:

    • DNN-L2 prevented semantic drift on t=1t=1 by heavily penalizing parameter changes, but suffered severely degraded accuracy on later tasks (t=4,7t=4, 7) due to constrained plasticity.
    • DNN-EWC showed less forgetting on later tasks than DNN-L2, but attained lower overall AUROC due to fixed capacity limitations.
    • DNN-Progressive completely avoided forgetting by freezing earlier subnetwork columns, but underperformed in raw task accuracy due to an inability to dynamically adjust per-layer unit counts.
    • DEN without timestamping (DEN-No-Stamp) achieved strong later-task accuracy but experienced slight performance degradation on earlier tasks as new units established unconstrained forward paths.
    • Full DEN (with unit splitting and timestamped inference) maintained strictly stable AUROC on early tasks across all subsequent training stages, exhibiting zero observable catastrophic forgetting while achieving the highest accuracy on newly introduced tasks.

Coverage note — None was omitted; all key contributions, algorithms, mathematical formulations, experimental settings, empirical results, and ablation studies are captured.

References

  1. 1.Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al. Tensorflow: Large-scale Machine Learning on Heterogeneous Distributed Systems. arXiv:1603.04467, 2016.
  2. 2.Jose M Alvarez and Mathieu Salzmann. Learning the number of neurons in deep networks. In Advances in Neural Information Processing Systems, pp. 2262–2270, 2016.
  3. 3.Corinna Cortes, Xavi Gonzalvo, Vitaly Kuznetsov, Mehryar Mohri, and Scott Yang. Adanet: Adaptive structural learning of artificial neural networks. arXiv preprint arXiv:1607.01097, 2016.
  4. 4.Eric Eaton and Paul L. Ruvolo. ELLA: An efficient lifelong learning algorithm. In Sanjoy Dasgupta and David Mcallester (eds.), ICML, volume 28, pp. 507–515. JMLR Workshop and Conference Proceedings, 2013.
  5. 5.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, pp. 201611835, 2017.
  6. 6.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. 2009.
  7. 7.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks. In NIPS, 2012.
  8. 8.Abhishek Kumar and Hal Daume III. Learning task grouping and overlap in multi-task learning. In ICML, 2012.
  9. 9.Christoph Lampert, Hannes Nickisch, and Stefan Harmeling. Learning to Detect Unseen Object Classes by Between-Class Attribute Transfer. In CVPR, 2009.
  10. 10.Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based Learning Applied to Document Recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  11. 11.Sang-Woo Lee, Jin-Hwa Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Overcoming catastrophic forgetting by incremental moment matching. arXiv preprint arXiv:1703.08475, 2017.
  12. 12.George Philipp and Jaime G. Carbonell. Nonparametric neural networks. In ICLR, 2017.
  13. 13.Andrei Rusu, Neil Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  14. 14.S. Thrun. A lifelong learning perspective for mobile robot control. In V. Graefe (ed.), Intelligent Robots and Systems. Elsevier, 1995.
  15. 15.Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. Learning structured sparsity in deep neural networks. In NIPS, pp. 2074–2082, 2016.
  16. 16.Tianjun Xiao, Jiaxing Zhang, Kuiyuan Yang, Yuxin Peng, and Zheng Zhang. Error-driven incremental learning in deep convolutional neural network for large-scale image classification. In Proceedings of the 22nd ACM international conference on Multimedia, pp. 177–186. ACM, 2014.
  17. 17.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In ICML, pp. 3987–3995, 2017.
  18. 18.Guanyu Zhou, Kihyuk Sohn, and Honglak Lee. Online incremental feature learning with denoising autoencoders. In International Conference on Artificial Intelligence and Statistics, pp. 1453–1461, 2012.

Citation

MLA
Yoon, J., et al. “Lifelong Learning with Dynamically Expandable Networks”. arXiv, 2017, http://arxiv.org/abs/1708.01547v11.
APA
Yoon, J., Yang, E., Lee, J., & Hwang, S. J. (2017). Lifelong Learning with Dynamically Expandable Networks. arXiv. http://arxiv.org/abs/1708.01547v11
Chicago
Yoon, J., E. Yang, J. Lee, and S. J. Hwang. 2017. “Lifelong Learning with Dynamically Expandable Networks”. arXiv. http://arxiv.org/abs/1708.01547v11.
Harvard
Yoon, J. et al. (2017) “Lifelong Learning with Dynamically Expandable Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1708.01547v11.
Vancouver
1. Yoon J, Yang E, Lee J, Hwang SJ (2017) Lifelong Learning with Dynamically Expandable Networks. arXiv

BibTeX

@article{yoon2017lifelong,
  title = {Lifelong Learning with Dynamically Expandable Networks},
  author = {Yoon, Jaehong and Yang, Eunho and Lee, Jeongtae and Hwang, Sung Ju},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1708.01547v11},
  eprint = {1708.01547}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors