Progressive Neural Networks

Andrei A. RusuNeil C. RabinowitzGuillaume DesjardinsHubert SoyerJames KirkpatrickKoray KavukcuogluRazvan PascanuRaia Hadsell

article2016arXiv3,025 citations

Introduces an architecture that eliminates catastrophic forgetting in sequential reinforcement learning tasks by allocating new network columns and leveraging lateral connections to transfer previously learned features.

Listen

Developing artificial intelligence systems capable of learning multiple sequential tasks without degrading previously acquired knowledge is a critical step toward human-level artificial intelligence. Traditional deep learning approaches typically rely on fine-tuning pretrained models on new domains. However, fine-tuning is inherently destructive: it overwrites prior capabilities—a failure mode known as catastrophic forgetting—and struggles when deciding which past model to build upon when facing sequences of diverse or incompatible tasks.

The article evaluates progressive neural networks, a modular architecture designed to accumulate knowledge across sequential tasks, prevent catastrophic forgetting entirely, and accelerate learning through knowledge transfer. The authors set out to demonstrate whether this architecture can reliably outperform standard fine-tuning approaches across diverse, sequential reinforcement learning environments.

To evaluate the system, the authors conducted reinforcement learning experiments across three distinct domains: synthetic variants of the Atari game Pong with modified visuals and mechanics, sequences of distinct Atari games, and three-dimensional navigation tasks within the Labyrinth maze environment. The architecture allocates a separate neural network column for each new task while freezing all prior columns to protect existing capabilities. Lateral connections equipped with adapter layers allow new columns to selectively reuse or ignore features learned by earlier columns. The authors benchmarked this design against training from scratch, standard full fine-tuning, and output-only fine-tuning, while using a sensitivity analysis based on Fisher Information to measure where and how feature reuse occurred.

The findings show that progressive neural networks consistently outperform conventional methods in learning speed and cumulative reward. In synthetic Pong variants, progressive networks achieved mean transfer scores of 209% to 222% relative to single-task baselines, outperforming full fine-tuning at 181%. In complex 3D maze tasks, two-column progressive networks reached an average transfer score of 491%, compared to 235% for full fine-tuning. In diverse Atari game sequences, progressive networks achieved positive transfer in 8 out of 12 target games—compared to only 5 out of 12 for standard fine-tuning—with transfer performance scaling positively up to four-column configurations. The analysis revealed that optimal transfer occurs when models balance the reuse of prior visual features with the acquisition of new, task-specific representations, while completely avoiding negative interference on unrelated tasks.

These results demonstrate that retaining prior models and expanding network capacity enables safer, more effective sequential learning than altering existing model parameters. For decision-makers and system designers, this approach significantly reduces the operational risk of performance degradation across deployed capabilities while lowering training time for adjacent tasks. Because prior tasks remain untouched, organizations can deploy multi-task systems into production environments without needing to maintain persistent, centralized training datasets for continuous retraining.

Organizations developing sequential machine learning workflows should consider modular, column-based architectures when learning sequences of tasks where past competency must be strictly preserved. However, engineering teams should plan for post-training compression, column pruning, or distillation pipelines before deployment, as network parameters scale quadratically with the number of tasks added. Further development is also recommended to automate task-identification routing during inference, removing the current operational requirement for explicit task labels.

The findings are supported by consistent results across diverse reinforcement learning benchmarks. However, the study evaluated up to four sequential tasks per run and relied on known task boundaries. Leaders should view the architecture as a proven strategy for controlled multi-task scaling, while exercising caution when deploying in unconstrained settings that require autonomous task switching or have strict parameter footprint limits.

arXiv: 1606.04671
  • Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). It establishes the foundational deep reinforcement learning framework on Atari benchmarks that Progressive Neural Networks directly adapt and evaluate on sequential task streams.
  • Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). It systematically investigates layer-wise feature transferability in deep neural networks, providing the empirical rationale for reusing frozen representations across related tasks.
  • Paper: Transfer Learning for Reinforcement Learning Domains: A Survey, Matthew E. Taylor et al. (2009). It surveys foundational transfer learning methods in reinforcement learning domains, establishing the core problem formulation of leveraging prior agent experience.
  • Paper: Why Does Unsupervised Pre-training Help Deep Learning?, Dumitru Erhan et al. (2010). It analyzes the mechanics of feature transfer, pre-training, and fine-tuning in deep architectures, which serve as standard baseline approaches contrasted in Progressive Networks.
  • Paper: Curriculum learning, Yoshua Bengio et al. (2009). It formalizes sequential training across tasks of varying difficulty, motivating architectural solutions for transferring representations sequentially without catastrophic forgetting.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). It introduces Elastic Weight Consolidation as an alternative parameter-regularization approach to overcoming catastrophic forgetting in sequential reinforcement learning tasks.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). It develops Gradient Episodic Memory to mitigate catastrophic forgetting and enable positive backward transfer without adding task-specific neural columns.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). It presents a knowledge distillation technique for sequential task adaptation that preserves prior knowledge without retaining historical training data or expanding the network architecture.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). It proposes Synaptic Intelligence to compute online path-integral weight importance, offering a fixed-capacity alternative to progressive column growth for continual learning.
  • Paper: Experience Replay for Continual Learning, David Rolnick et al. (2018). It uses experience replay combined with behavioral cloning to prevent catastrophic forgetting in sequential deep reinforcement learning without explicit task boundary metadata.
  • Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). It benchmarks continual learning methods against expanding architectural models and introduces an efficient episodic memory projection algorithm for streaming tasks.
  • Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). It builds upon continual learning principles by computing unsupervised parameter sensitivity to preserve critical pathways while adapting to incoming tasks.
  • Paper: Continual Lifelong Learning with Neural Networks: A Review, German I. Parisi et al. (2018). It provides a comprehensive review of lifelong and continual learning, systematically categorizing dynamic architectures alongside regularization and replay mechanisms.
  • Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). It benchmarks parameter isolation and dynamic architecture strategies against other continual learning paradigms across standardized task sequences.
  • Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). It tackles task interference and transfer during multi-task learning by projecting conflicting gradients, addressing transfer efficiency within a shared model architecture.
Cover for Progressive Neural Networks

Abstract

Learning to solve complex sequences of tasks--while both leveraging transfer and avoiding catastrophic forgetting--remains a key obstacle to achieving human-level intelligence. The progressive networks approach represents a step forward in this direction: they are immune to forgetting and can leverage prior knowledge via lateral connections to previously learned features. We evaluate this architecture extensively on a wide variety of reinforcement learning tasks (Atari and 3D maze games), and show that it outperforms common baselines based on pretraining and finetuning. Using a novel sensitivity measure, we demonstrate that transfer occurs at both low-level sensory and high-level control layers of the learned policy.

Table of Contents

  • 1 Introduction
  • 2 Progressive Networks
  • 3 Transfer Analysis
  • 4 Related Literature
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Pong Soup
  • 5.3 Atari Games
  • 5.4 Labyrinth
  • 6 Conclusion
  • References
  • A Perturbation Analysis
  • B Compressibility of Progressive Networks
  • C Setup Details
  • D Learning curves
  • E Labyrinth

Knowls

  1. Knowl 1 — Progressive Neural Network Architecture

    model/method

    A progressive neural network is a continual learning architecture designed to learn a sequence of KK distinct tasks without catastrophic forgetting while facilitating transfer through lateral connections.

    When training begins on the first task (k=1k=1), a deep neural network column consisting of LL layers is initialized and trained to convergence. When moving to task k>1k > 1, the parameters of all previous columns {Θ(j)}j=1k−1\{\Theta^{(j)}\}_{j=1}^{k-1} are frozen (treated as constants by the optimizer), and a new column with parameters Θ(k)\Theta^{(k)} is instantiated with random initialization. To enable transfer from previously acquired representations, hidden layer ii of column kk receives inputs from layer i−1i-1 of column kk as well as from layer i−1i-1 of all preceding columns j<kj < k via lateral connections.

    For a progressive network of KK columns, the feedforward activation hi(k)∈Rnih_i^{(k)} \in \mathbb{R}^{n_i} at layer i≤Li \le L of column kk is given by:

    hi(k)=f(Wi(k)hi−1(k)+∑j<kUi(k:j)hi−1(j))h_i^{(k)} = f\left(W_i^{(k)} h_{i-1}^{(k)} + \sum_{j < k} U_i^{(k:j)} h_{i-1}^{(j)}\right)

    where Wi(k)∈Rni×ni−1W_i^{(k)} \in \mathbb{R}^{n_i \times n_{i-1}} denotes the intra-column weight matrix of layer ii of column kk, Ui(k:j)∈Rni×ni−1(j)U_i^{(k:j)} \in \mathbb{R}^{n_i \times n_{i-1}^{(j)}} denotes the lateral weight matrix routing features from layer i−1i-1 of column jj to layer ii of column kk, h0h_0 is the network input, and ff is an element-wise non-linear activation function (such as f(x)=max⁡(0,x)f(x) = \max(0, x)). Because lateral connections are strictly directed from earlier columns (j<kj < k) to later columns and earlier parameters are frozen, task learning on column kk causes zero interference with policies learned on columns 1,…,k−11, \dots, k-1.

  2. Knowl 2 — Non-Linear Lateral Adapters for Progressive Networks

    model/method

    To prevent parameter count explosion when connecting to multiple prior columns and to improve conditioning across representations with differing scales, progressive networks employ non-linear adapter modules for lateral connections.

    Let hi−1(<k)=[hi−1(1),hi−1(2),…,hi−1(k−1)]∈Rni−1(<k)h_{i-1}^{(<k)} = \left[ h_{i-1}^{(1)}, h_{i-1}^{(2)}, \dots, h_{i-1}^{(k-1)} \right] \in \mathbb{R}^{n_{i-1}^{(<k)}} denote the concatenated anterior feature vector containing the layer i−1i-1 hidden activations across all previous k−1k-1 columns. The linear lateral connection is replaced by a single-hidden-layer multi-layer perceptron (MLP) adapter:

    hi(k)=σ(Wi(k)hi−1(k)+Ui(k:j)σ(Vi(k:j)αi−1(<k)hi−1(<k)))h_i^{(k)} = \sigma\left( W_i^{(k)} h_{i-1}^{(k)} + U_i^{(k:j)} \sigma\left( V_i^{(k:j)} \alpha_{i-1}^{(<k)} h_{i-1}^{(<k)} \right) \right)

    where:

    • αi−1(<k)\alpha_{i-1}^{(<k)} is a learnable scalar multiplier initialized to a small random value, serving to normalize differing activation scales across prior columns;
    • Vi(k:j)∈Rni−1×ni−1(<k)V_i^{(k:j)} \in \mathbb{R}^{n_{i-1} \times n_{i-1}^{(<k)}} is a projection matrix that projects the anterior feature vector onto an ni−1n_{i-1}-dimensional subspace;
    • Ui(k:j)U_i^{(k:j)} denotes the second linear layer of the adapter module;
    • σ\sigma denotes the element-wise activation function (e.g., ReLU).

    For convolutional layers, dimensionality reduction in the lateral adapter is implemented using 1×11 \times 1 convolutions. Projecting the concatenated anterior features into an ni−1n_{i-1}-dimensional subspace bounds the parameter growth of the lateral connections for each new column to the same order of magnitude as the original single column Θ(1)\Theta^{(1)}.

  3. Knowl 3 — Average Fisher Sensitivity

    equation

    Average Fisher Sensitivity (AFS) is an analytical metric used to quantify the relative reliance of a target policy π\pi on the representations learned across different columns and layers of a progressive neural network.

    Let h^i(k)\hat{h}_i^{(k)} denote the normalized hidden activations at layer ii of column kk, and let ρ(s,a)\rho(s, a) denote the stationary state-action distribution induced by the progressive network on the target task. The diagonal Fisher information matrix F^i(k)\hat{F}_i^{(k)} with respect to the layer activations is defined as:

    F^i(k)=Eρ(s,a)[(∂log⁡π(a∣s)∂h^i(k))(∂log⁡π(a∣s)∂h^i(k))T]\hat{F}_i^{(k)} = \mathbb{E}_{\rho(s, a)}\left[ \left( \frac{\partial \log \pi(a \mid s)}{\partial \hat{h}_i^{(k)}} \right) \left( \frac{\partial \log \pi(a \mid s)}{\partial \hat{h}_i^{(k)}} \right)^T \right]

    For convolutional layers, the diagonal elements F^i(k)(m,m)\hat{F}_i^{(k)}(m, m) implicitly sum over spatial pixel locations for feature map mm. The relative sensitivity of feature mm in layer ii of column kk across all KK columns is given by the Average Fisher Sensitivity (AFS):

    AFS(i,k,m)=F^i(k)(m,m)∑j=1KF^i(j)(m,m)\text{AFS}(i, k, m) = \frac{\hat{F}_i^{(k)}(m, m)}{\sum_{j=1}^K \hat{F}_i^{(j)}(m, m)}

    To compute column reliance at an entire layer ii, the score is summed over all MiM_i features in layer ii:

    AFS(i,k)=∑m=1MiAFS(i,k,m)\text{AFS}(i, k) = \sum_{m=1}^{M_i} \text{AFS}(i, k, m)

    Using normalized representations h^\hat{h} ensures that sensitivity values are scale-invariant and directly comparable across different layers and columns.

  4. Knowl 4 — Average Perturbation Sensitivity

    equation

    Average Perturbation Sensitivity (APS) is an empirical measure of the causal contribution of individual columns and layers to the overall target task performance in a progressive neural network.

    Gaussian noise is injected into the post-activation representations of layer ii in column kk during each forward pass. The noise variance is scaled proportionally to the empirical variance of activations in that layer to achieve scale invariance. Let σi2(k)\sigma_i^{2(k)} be the critical noise variance required to induce a 50% drop in target game performance over 10 evaluation episodes. The noise precision is defined as:

    Λi(k)=1σi2(k)\Lambda_i^{(k)} = \frac{1}{\sigma_i^{2(k)}}

    The Average Perturbation Sensitivity of layer ii in column kk, normalized across all KK columns, is:

    APS(i,k)=Λi(k)∑j=1KΛi(j)\text{APS}(i, k) = \frac{\Lambda_i^{(k)}}{\sum_{j=1}^K \Lambda_i^{(j)}}

    A high APS score indicates that the target task policy is highly sensitive to perturbations in column kk at layer ii, demonstrating that the network critically depends on representations transferred from that specific column.

  5. Knowl 5 — Reinforcement Learning Transfer Benchmark Results

    data/table

    The transfer capabilities of progressive neural networks were evaluated against standard transfer and single-task baselines across three reinforcement learning domains: Pong Soup (synthetic visual and control variants of Atari Pong), standard Atari games, and Labyrinth (3D first-person maze foraging tasks). Transfer score is measured as the area under the learning curve (AUC) of episode reward during training, normalized such that Baseline 1 (a single column trained exclusively on the target task) equals 100%.

    Pong Soup Atari Labyrinth
    Model Mean (%) Median (%) Mean (%) Median (%) Mean (%) Median (%)
    Baseline 1 (Target only) 100 100 100 100 100 100
    Baseline 2 (Finetune output only) 35 7 41 21 88 85
    Baseline 3 (Full finetuning) 181 160 133 110 235 112
    Baseline 4 (Random frozen column) 134 131 96 95 185 108
    Progressive (2 columns) 209 169 132 112 491 115
    Progressive (3 columns) 222 183 140 111 — —
    Progressive (4 columns) — — 141 116 — —

    Baseline 2 severely underperforms in deep RL (negative transfer across all domains), indicating that frozen source representations alone without target adaptation are insufficient. Full finetuning (Baseline 3) achieves strong transfer but destroys source task policies. Progressive networks consistently achieve higher mean and median transfer scores than all baselines, with performance monotonically increasing as more source columns are added.

  6. Knowl 6 — Progressive Reinforcement Learning Experimental Protocol

    experimental setup

    Progressive networks are trained using the Asynchronous Advantage Actor-Critic (A3C) framework across multi-threaded CPU workers (16 workers per experiment). The policy π(k)(a∣s)=hL(k)(s)\pi^{(k)}(a \mid s) = h_L^{(k)}(s) and value function V(k)(s)V^{(k)}(s) are simultaneously predicted from the top hidden layer.

    Key experimental parameters include:

    • Optimization: Shared RMSProp without momentum or variance centering.
    • Hyperparameter sampling: Learning rate sampled from {10−3,5⋅10−4,10−4}\{10^{-3}, 5 \cdot 10^{-4}, 10^{-4}\}; entropy regularization weight from {10−2,10−3,10−4}\{10^{-2}, 10^{-3}, 10^{-4}\}; gradient norm clipping threshold from {20,40}\{20, 40\}; lateral scalar multiplier α\alpha initialized from {1,10−1,10−2}\{1, 10^{-1}, 10^{-2}\}.
    • Network architecture for Atari: 3 convolutional layers (all with 12 feature maps: conv 1 has 8×88 \times 8 kernel with stride 4; conv 2 has 4×44 \times 4 kernel with stride 2; conv 3 has 3×43 \times 4 kernel with stride 1), followed by a fully connected layer with 256 hidden units.
    • Execution: Action repeat of 4 frames; training runs execute for up to 1.6×1081.6 \times 10^8 environment steps (4×1074 \times 10^7 agent-perceived steps).
    • Aggregation: Reported performance averages the top 3 of 25 runs (different random seeds and sampled hyperparameters) evaluated by the area under the learning curve (AUC) averaged in chunks of 2.5×1052.5 \times 10^5 environment steps.
  7. Knowl 7 — Hierarchical Visual and Control Feature Transfer Dynamics

    empirical result

    Sensitivity analysis (using Average Fisher Sensitivity and Average Perturbation Sensitivity) reveals distinct layer-specific transfer mechanisms depending on the nature of domain shifts between tasks:

    1. Visual Invariance with Policy Shift (Pong →\to H-Flip): When visual inputs are horizontally flipped, the agent fully reuses low-level and mid-level convolutional representations from the source column but learns a completely new fully connected policy layer to adjust for the inverted paddle and ball coordinate dynamics.
    2. Scale Shift (Pong →\to Zoom): When input dimensions are scaled and translated, the target column reuses a single low-level spatio-temporal filter with a strong temporal DC component from the source net (which detects ball and paddle motion), while mid-level and high-level convolutional layers are relearned.
    3. Noise Asymmetry (Pong ↔\leftrightarrow Noisy Pong): Transfer from clean Pong to Gaussian-noise Pong fails at low-level vision because clean filters lack noise robustness, requiring new conv1 filters. Conversely, transferring from Noisy Pong to clean Pong transfers all convolutional layers without requiring new visual features, demonstrating that noise-robust representations generalize unidirectionally to clean environments.
  8. Knowl 8 — Optimal Feature Transfer Sweet Spot in Multi-Task Atari Sequences

    empirical result

    Across 72 evaluated three-column progressive networks on diverse Atari game sequences (source games: Seaquest, River Raid, Pong; target games: Alien, Asterix, Boxing, Centipede, Gopher, Hero, James Bond, Krull, Robotank, Road Runner, Star Gunner, Wizard of Wor):

    • Progressive networks achieved positive transfer in 8 out of 12 target games (compared to 5 of 12 for full finetuning), with only 2 instances of negative transfer.
    • Transfer score strongly correlates with the Average Fisher Sensitivity distribution between source and target columns. The highest positive transfer occurs at a 'sweet spot' where the target column heavily leverages source convolutional features while simultaneously dedicating newly allocated convolutional capacity to learn target-specific visual features.
    • Severe negative transfer occurs when the network exhibits complete reliance on the convolutional layers of prior columns without learning any new visual features in the target column. This failure mode arises because transferred representations can act as a sub-optimal local minimum that restricts exploration or imposes an incompatible inductive bias.
  9. Knowl 9 — Sparsity and Compressibility of Lateral Capacity

    empirical result

    Analysis of Average Fisher Sensitivity (AFS) spectra across progressive networks with 2, 3, and 4 columns reveals that network capacity is utilized with increasing sparsity as more tasks are learned:

    1. Decreasing Source Utilization: Concatenating and sorting per-feature-map AFS values from all source columns shows that as the total number of columns increases (from 2 to 4 columns), the resulting AFS spectrum becomes significantly sparser. The network utilizes a progressively smaller fraction of the available source feature maps.
    2. Declining New Column Demand: The total capacity utilized by the final added column (measured by the area under its AFS spectrum) decreases as the number of prior columns grows, and its own feature spectrum becomes sparser.

    These findings indicate that full dense connectivity between columns is redundant, demonstrating that progressive networks can be substantially pruned or compressed online during continual learning to prevent parameter explosion.

  10. Knowl 10 — Scalability and Inference Limitations of Progressive Networks

    limitation

    While progressive networks eliminate catastrophic forgetting by design and provide positive transfer across sequential reinforcement learning tasks, they present two primary architectural limitations:

    1. Parameter and Capacity Growth: In their unpruned form, the number of hidden units and feature maps grows linearly with the number of tasks KK, and the number of parameters grows quadratically (O(K2) \mathcal{O}(K^2)) due to all-to-all lateral connections across columns at each layer.
    2. Task ID Requirement at Inference: Because each column encapsulates the policy for a specific task and columns remain distinct, the progressive network requires an explicit external task identifier at inference time to select which column's output head to execute.

Coverage note — None was omitted. All key models, equations, sensitivity analyses, empirical transfer results, and limitations were fully extracted into self-contained knowls.

References

  1. 1.Forest Agostinelli, Michael R Anderson, and Honglak Lee. Adaptive multi-column deep neural networks with application to robust image denoising. In Advances in Neural Information Processing Systems, 2013.
  2. 2.Shun-ichi Amari. Natural gradient works efficiently in learning. Neural Computation, 1998.
  3. 3.M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research (JAIR), 47:253–279, 2013.
  4. 4.Yoshua Bengio. Deep learning of representations for unsupervised and transfer learning. In JMLR: Workshop on Unsupervised and Transfer Learning, 2012.
  5. 5.Dan C. Ciresan, Ueli Meier, and Jürgen Schmidhuber. Multi-column deep neural networks for image classification. In Conf. on Computer Vision and Pattern Recognition, 2012.
  6. 6.Scott E. Fahlman and Christian Lebiere. The cascade-correlation learning architecture. In Advances in Neural Information Processing Systems, 1990.
  7. 7.G. E. Hinton and R. R. Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313(5786):504–507, July 2006.
  8. 8.Goeff Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. CoRR, abs/1503.02531, 2015.
  9. 9.Yann LeCun, John S. Denker, and Sara A. Solla. Optimal brain damage. In Advances in Neural Information Processing Systems, 1990.
  10. 10.Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. In Proc. of Int’l Conference on Learning Representations (ICLR), 2013.
  11. 11.G. Mesnil, Y. Dauphin, X. Glorot, S. Rifai, Y. Bengio, I. Goodfellow, E. Lavoie, X. Muller, G. Desjardins, D. Warde-Farley, P. Vincent, A. Courville, and J. Bergstra. Unsupervised and transfer learning challenge: a deep learning approach. In JMLR W& CP: Proc. of the Unsupervised and Transfer Learning challenge and workshop, volume 27, 2012.
  12. 12.V. Mnih, Kk Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  13. 13.Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In Int’l Conf. on Machine Learning (ICML), 2016.
  14. 14.Emilio Parisotto, Lei Jimmy Ba, and Ruslan Salakhutdinov. Actor-mimic: Deep multitask and transfer reinforcement learning. In Proc. of Int’l Conference on Learning Representations (ICLR), 2016.
  15. 15.Mark B. Ring. Continual Learning in Reinforcement Environments. R. Oldenbourg Verlag, 1995.
  16. 16.Artem Rozantsev, Mathieu Salzmann, and Pascal Fua. Beyond sharing weights for deep domain adaptation. CoRR, abs/1603.06432, 2016.
  17. 17.A. Rusu, S. Colmenarejo, Ç. Gülçehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell. Policy distillation. abs/1511.06295, 2016.
  18. 18.Paul Ruvolo and Eric Eaton. Ella: An efficient lifelong learning algorithm. In Proceedings of the 30th International Conference on Machine Learning (ICML-13), June 2013.
  19. 19.Daniel L. Silver, Qiang Yang, and Lianghao Li. Lifelong machine learning systems: Beyond learning algorithms. In AAAI Spring Symposium: Lifelong Machine Learning, 2013.
  20. 20.Matthew E. Taylor and Peter Stone. An introduction to inter-task transfer for reinforcement learning. AI Magazine, 32(1):15–34, 2011.
  21. 21.Alexander V. Terekhov, Guglielmo Montone, and J. Kevin O’Regan. Knowledge Transfer in Deep Block-Modular Neural Networks, pages 268–279. Springer International Publishing, Cham, 2015.
  22. 22.C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor. A Deep Hierarchical Approach to Lifelong Learning in Minecraft. ArXiv e-prints, 2016.
  23. 23.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? In Advances in Neural Information Processing Systems, pages 3320–3328, 2014.
  24. 24.Guanyu Zhou, Kihyuk Sohn, and Honglak Lee. Online incremental feature learning with denoising autoencoders. In Proc. of Int’l Conf. on Artificial Intelligence and Statistics (AISTATS), pages 1453–1461, 2012.

Citation

MLA
Rusu, A. A., et al. “Progressive Neural Networks”. arXiv, 2016, http://arxiv.org/abs/1606.04671v4.
APA
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., & Hadsell, R. (2016). Progressive Neural Networks. arXiv. http://arxiv.org/abs/1606.04671v4
Chicago
Rusu, A. A., N. C. Rabinowitz, G. Desjardins, et al. 2016. “Progressive Neural Networks”. arXiv. http://arxiv.org/abs/1606.04671v4.
Harvard
Rusu, A.A. et al. (2016) “Progressive Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1606.04671v4.
Vancouver
1. Rusu AA, Rabinowitz NC, Desjardins G, Soyer H, Kirkpatrick J, Kavukcuoglu K, Pascanu R, Hadsell R (2016) Progressive Neural Networks. arXiv

BibTeX

@article{rusu2016progressive,
  title = {Progressive Neural Networks},
  author = {Rusu, Andrei A. and Rabinowitz, Neil C. and Desjardins, Guillaume and Soyer, Hubert and Kirkpatrick, James and Kavukcuoglu, Koray and Pascanu, Razvan and Hadsell, Raia},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1606.04671v4},
  eprint = {1606.04671}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/