Understanding Plasticity in Neural Networks

Clare LyleZeyu ZhengEvgenii NikishinBernardo Ávila PiresRazvan PascanuWill Dabney

article2023ICML216 citations

Reveals that neural network plasticity loss stems primarily from unfavorable changes in loss curvature rather than unit saturation, providing practical architectural and optimization techniques like layer normalization to maintain continual learning capacity in deep reinforcement learning.

Listen

Deep reinforcement learning systems must continuously update their predictions as tasks, goals, and training data evolve over time. However, neural networks frequently suffer from plasticity loss—a phenomenon where models progressively lose their ability to adapt and learn new information as training continues. While this issue directly limits the long-term reliability and adaptability of artificial intelligence systems, its underlying mechanisms have remained poorly understood.

The article systematically analyzes why neural networks lose plasticity during non-stationary learning and evaluates targeted architectural and algorithmic interventions to preserve adaptability over extended training.

The researchers designed controlled empirical experiments using image-based reinforcement learning environments (MNIST and CIFAR-10) and full-scale benchmarks across 57 Atari games. They systematically evaluated how optimization dynamics, loss landscape geometry, and model scale affect plasticity. The study compared several previously proposed diagnostic metrics (such as weight norms, matrix rank, and inactive units) against landscape properties, and benchmarked practical interventions including normalization layers, output representations, parameter resets, and weight decay.

The investigation produced four primary findings. First, commonly cited explanations for plasticity loss—such as inactive network units, parameter norms, or feature rank—fail as universal root causes; their correlations with plasticity reverse depending on task structure and dataset. Second, plasticity loss is primarily driven by optimization dynamics making the loss landscape increasingly sharp and difficult to navigate, causing optimization to slow down rather than simply getting stuck in local plateaus. Third, simply increasing model width reduces plasticity loss but cannot eliminate it entirely even at hardware capacity limits. Fourth, stabilizing the loss landscape using layer normalization yielded the most consistent improvements, robustly outperforming the baseline architecture across the 57 Atari benchmark games without additional hyperparameter tuning, with performance gains exceeding 100% in several environments.

These findings indicate that maintaining a smooth loss landscape and stable gradient updates is essential for preventing learning degradation in dynamic environments. Rather than relying on heuristic parameter resets or regularization methods that can disrupt current task performance, engineering efforts should prioritize architectural choices and optimizer configurations that smooth optimization curvature. This shift in design focus can substantially improve model robustness, lower retraining costs, and prevent sudden performance failures in deployed adaptive systems.

Organizations developing adaptive machine learning and reinforcement learning systems should adopt layer normalization as a standard component in deep network architectures. Engineering teams should also tune optimizer stability parameters (such as increasing numerical stability factors and updating moment estimates more rapidly) when deploying models in non-stationary settings. While categorical output encodings showed strong plasticity preservation, their stability trade-offs warrant cautious implementation. Further pilot testing in large-scale, continuous-learning production pipelines is recommended to validate these architectural strategies across diverse operational domains.

arXiv: 2303.01486
Cover for Understanding Plasticity in Neural Networks

Abstract

Plasticity, the ability of a neural network to quickly change its predictions in response to new information, is essential for the adaptability and robustness of deep reinforcement learning systems. Deep neural networks are known to lose plasticity over the course of training even in relatively simple learning problems, but the mechanisms driving this phenomenon are still poorly understood. This paper conducts a systematic empirical analysis into plasticity loss, with the goal of understanding the phenomenon mechanistically in order to guide the future development of targeted solutions. We find that loss of plasticity is deeply connected to changes in the curvature of the loss landscape, but that it often occurs in the absence of saturated units. Based on this insight, we identify a number of parameterization and optimization design choices which enable networks to better preserve plasticity over the course of training. We validate the utility of these findings on larger-scale RL benchmarks in the Arcade Learning Environment.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Preliminaries
  • 2.2. Defining plasticity
  • 3. Methodology and Motivating Questions
  • 3.1. Measuring plasticity
  • 3.2. Environments
  • 3.3. Outline of experimental results
  • 4. Two Simple Studies on Plasticity
  • 4.1. Optimizer instability and non-stationarity
  • 4.2. Loss landscape evolution under non-stationarity
  • 5. Explaining Plasticity Loss
  • 5.1. Experimental setting
  • 5.2. Falsification of prior hypotheses
  • 5.3. Loss landscape evolution during training
  • 6. Solutions
  • 6.1. The role of scaling on plasticity
  • 6.2. Interventions in toy problems
  • 6.3. Application to larger benchmarks
  • 7. Related Work
  • 8. Conclusions
  • Acknowledgements
  • References
  • A. Experiment details
  • A.1. Case studies
  • A.2. Toy RL environments
  • A.3. Double DQN
  • B. Additional analysis
  • B.1. Detailed intervention analysis
  • B.2. Learning curves for classification MDPs
  • B.2.1. TRAINING ACCURACY
  • B.2.2. PROBE TASKS
  • B.3. Qualitative findings in DDQN

Knowls

  1. Knowl 1 — Operational definition of plasticity and plasticity loss

    definition

    Let O(θ,ℓ)\mathcal{O}(\theta,\ell) be an optimization procedure that starts from neural-network parameters θ∈Θ\theta\in\Theta and optimizes a loss function ℓ:Θ→R\ell:\Theta\to\mathbb{R}, producing parameters θ∗=O(θ,ℓ)\theta^*=\mathcal{O}(\theta,\ell). Let L\mathcal{L} be a distribution over possible probe losses and let bb be the loss achieved by a fixed baseline predictor. The plasticity of parameters θt\theta_t is defined as

    P(θt)=b−Eℓ∼L[ℓ(θt∗)],θt∗=O(θt,ℓ).P(\theta_t)=b-\mathbb{E}_{\ell\sim\mathcal{L}}[\ell(\theta_t^*)],\qquad \theta_t^*=\mathcal{O}(\theta_t,\ell).

    Thus, higher plasticity means that the network obtains a lower expected loss after being adapted to randomly sampled new learning objectives. For a training trajectory beginning at θ0\theta_0, the paper measures the change in plasticity as P(θt)−P(θ0)P(\theta_t)-P(\theta_0); a decrease in PP represents plasticity loss. Because the baseline bb cancels when comparing checkpoints, plasticity loss measures the relative change in adaptability independently of the absolute difficulty of the probe tasks.

  2. Knowl 2 — Randomized probe-task procedure for measuring plasticity

    model/method

    The paper probes a trained reinforcement-learning network ff using regression targets designed to perturb its predictions in effectively random directions. For an input distribution XX consisting of transitions stored in the agent's replay buffer, it samples an independently initialized network f(⋅;ω0)f(\cdot;\omega_0), with ω0\omega_0 drawn from the same initialization distribution as the trained network, and defines

    g(x)=a+sin⁡ ⁣(105f(x;ω0)),g(x)=a+\sin\!\left(10^5 f(x;\omega_0)\right),

    where aa is set to the current network's mean prediction over XX. The offset prevents later checkpoints, whose value predictions may have drifted away from zero, from being unfairly compared with randomly initialized networks. The resulting regression task is optimized with the same optimizer used for the primary reinforcement-learning task, for a budget of 2,000 optimizer steps. In the main plasticity experiments, the authors pause reinforcement-learning training every 5,000 optimizer steps, sample 10 independent target functions, optimize each probe from the saved checkpoint, and record the final probe loss. Lower final probe loss indicates greater plasticity.

  3. Knowl 3 — Loss-landscape and gradient-interference diagnostics

    equation

    For a loss ℓ(θ)\ell(\theta) at parameters θ∈Rd\theta\in\mathbb{R}^d, the paper analyzes the Hessian

    Hℓ(θ)=∇θ2ℓ(θ)∈Rd×d,H_\ell(\theta)=\nabla_\theta^2\ell(\theta)\in\mathbb{R}^{d\times d},

    whose largest eigenvalue is used as a measure of local loss-landscape sharpness. It also estimates the alignment of per-example gradients by sampling kk inputs x1,…,xkx_1,\ldots,x_k and constructing

    Ck[i,j]=⟨∇θℓ(θ,xi),∇θℓ(θ,xj)⟩∥∇θℓ(θ,xi)∥ ∥∇θℓ(θ,xj)∥,i,j∈{1,…,k}.C_k[i,j]= \frac{\left\langle\nabla_\theta\ell(\theta,x_i),\nabla_\theta\ell(\theta,x_j)\right\rangle} {\left\|\nabla_\theta\ell(\theta,x_i)\right\|\,\left\|\nabla_\theta\ell(\theta,x_j)\right\|}, \qquad i,j\in\{1,\ldots,k\}.

    Here ℓ(θ,xi)\ell(\theta,x_i) is the loss on input xix_i, and Ck[i,j]C_k[i,j] is the cosine similarity between the two parameter gradients. Negative off-diagonal values indicate interference: improving one group of inputs tends to worsen another. A low-rank or strongly block-structured matrix indicates approximately colinear gradients; positive colinearity can reflect shared progress, whereas negative colinearity reflects interference. The experiments use k=512k=512 for the toy reinforcement-learning analyses.

  4. Knowl 4 — Classification-inspired MDP testbed

    experimental setup

    The authors construct finite Markov decision processes with state space {0,…,9}\{0,\ldots,9\} and ten actions, while using either MNIST or CIFAR-10 images as observations. In the true-label environment, state ss emits an image from class ss, action aa receives reward δa=s\delta_{a=s}, and the process transitions randomly to a new state. The random-label environment has the same transition structure but assigns each image a randomized class label before associating observations with states. In the sparse-reward environment, observations follow the true-label mapping, the reward is δa=s=9\delta_{a=s=9}, and the process transitions to a random state when a≠sa\neq s but advances to s+1s+1 when a=sa=s.

    The true-label and random-label environments isolate non-stationarity caused by changing bootstrap targets without making state visitation depend on the policy. The sparse-reward environment additionally introduces policy-dependent visitation. These variants compare visually rich tasks that either align with the network's image-classification inductive bias or require fitting randomized labels, and that differ in reward density.

  5. Knowl 5 — Adaptive-optimizer instability after abrupt target changes

    empirical result

    A two-hidden-layer MLP with width 1,024 per hidden layer was trained to memorize randomly permuted labels for 5,000 fixed MNIST images. Labels were repeatedly re-randomized while the input set remained fixed. With Adam's default settings—learning rate 0.0010.001, first-moment decay β1=0.9\beta_1=0.9, second-moment decay β2=0.999\beta_2=0.999, numerical constant ϵ=10−9\epsilon=10^{-9}, and ϵˉ=0\bar\epsilon=0—abrupt target changes caused the optimizer to diverge, saturating most ReLU units and reducing accuracy to near-trivial levels.

    The instability is explained by Adam's update scaling,

    ut=αm^tv^t+ϵˉ+ϵ,u_t=\alpha\frac{\hat m_t}{\sqrt{\hat v_t+\bar\epsilon}+\epsilon},

    where α\alpha is the learning rate and m^t\hat m_t and v^t\hat v_t are exponentially averaged first- and second-moment estimates of the gradient. After a sudden loss increase, the first-moment estimate adapts more quickly than the second-moment estimate under the default settings, so the update can be divided by an underestimated second moment. Increasing the numerical stabilizer to ϵˉ=10−3\bar\epsilon=10^{-3} and using β2=0.9\beta_2=0.9 prevented the catastrophic divergence. The result identifies optimizer stability under non-stationary targets as one direct route by which plasticity can be destroyed.

  6. Knowl 6 — Gradient descent sharpens new-task landscapes more than equal-size random perturbations

    empirical result

    The authors compared two parameter trajectories that started from the same random initialization and received updates of equal norm: one trajectory used gradient-based optimization on a non-stationary Q-learning objective, while the other added Gaussian parameter perturbations with the same norm as the corresponding optimizer update. The experiment used the easy classification MDP with MNIST observations, stochastic gradient descent with learning rate 0.0010.001, batch size 512512, and target-network updates every 5,000 steps.

    To assess trainability of arbitrary new objectives rather than the primary Q-learning objective, the authors evaluated the Hessian and gradient structure of a probe loss of the form

    ℓ(θ)=[fθ(X)−stopgrad⁡(fθ(X))+ε]2,\ell(\theta)=\left[f_\theta(X)-\operatorname{stopgrad}(f_\theta(X))+\varepsilon\right]^2,

    where XX is a batch of inputs, fθ(X)f_\theta(X) is the current output vector, stopgrad⁡\operatorname{stopgrad} prevents the reference output from receiving gradients, and ε∼N(0,1)\varepsilon\sim\mathcal{N}(0,1) is a random output perturbation. The Hessian spectral norm increased for both trajectories, but its outlier eigenvalues grew substantially faster under gradient descent. Gradient descent also produced increasingly negative gradient similarities between inputs, whereas the random-perturbation trajectory did not. Thus, even when optimization avoids unit saturation, the directional bias of gradient descent can move parameters into regions that are sharper and more interference-prone for future learning objectives than regions reached by equally sized random parameter perturbations.

  7. Knowl 7 — Simple network statistics fail the causal falsification test

    theoretical result

    The paper tests whether four commonly proposed explanations of plasticity loss—parameter norm, weight-matrix rank, feature rank, and the number of inactive units—have stable explanatory power across learning problems. The falsification criterion is that a genuinely causal diagnostic should maintain a consistent association with plasticity when the observation space, reward structure, architecture, optimizer, or random seed is changed.

    Across 128 DQN agents trained under varied tasks, observation spaces, optimizers, and seeds, each candidate statistic exhibited a positive association with plasticity or plasticity loss in at least one setting and a negative association in another; many settings showed little association at all. For example, weight norm was positively correlated with plasticity loss for CIFAR-10 observations but slightly negatively correlated for MNIST observations. Feature rank and the number of inactive units likewise changed correlation sign when the reward structure or architecture changed. The correlations were generally weak, and their reversals rule out these statistics as standalone, robust causal explanations of plasticity loss.

  8. Knowl 8 — Plasticity loss manifests as slower probe optimization rather than a higher final plateau

    empirical result

    When probe-task learning curves were initialized from checkpoints taken progressively later during reinforcement-learning training, early checkpoints reduced the new target loss rapidly and reached low losses, whereas later checkpoints had progressively shallower learning curves. The later networks were therefore slower to navigate the new loss landscape rather than simply becoming trapped at higher-loss local minima.

    The later-checkpoint curves also showed increasing variance and non-monotonicity. In full-batch optimization, this behavior accompanied increasing loss-landscape sharpness; in mini-batch optimization, the authors additionally observed increasing interference between mini-batches and non-monotonic loss even on the batch used to compute the gradient. These observations connect the measured loss of plasticity to worsening optimization geometry for new objectives.

  9. Knowl 9 — Architectural and optimization interventions that preserve plasticity

    data/table

    The authors trained MLP, convolutional, ResNet-18, and small Vision Transformer agents for 100 iterations of 1,000 optimizer steps and measured the change in probe-task loss. The tested interventions were two-hot categorical output encoding, resetting the last layer at each target-network update, weight decay with coefficient 10−510^{-5}, spectral normalization of the initial linear layer, layer normalization after convolutional and fully connected layers, and Shrink-and-Perturb. Lower values indicate less plasticity loss; entries are reported mean ±\pm standard deviation.

    Could not parse LaTeX table

    The strongest and most consistent improvements came from parameterizations intended to smooth or stabilize the loss landscape, especially layer normalization and the two-hot categorical representation. Resetting the final layer helped in several architectures, whereas weight decay and Shrink-and-Perturb were less reliable. Two-hot encoding reduced probe plasticity loss but sometimes destabilized the primary learned policy and required different optimizer hyperparameters, so it was not a universally plug-in intervention. Increasing network width also consistently reduced plasticity loss across target-update periods of 1, 100, and 1,000 steps, but scaling a CNN up to the limit of a single GPU did not eliminate plasticity loss.

  10. Knowl 10 — Layer normalization improves Double-DQN performance across Atari games

    empirical result

    The paper inserted layer normalization after every hidden layer of a standard Double-DQN network and compared it with the unmodified network on all 57 games in the Arcade Learning Environment. Each condition used three random seeds, RMSProp, ϵ\epsilon-greedy exploration, frame stacking, a replay buffer of 100,000 transitions, training for 200 million frames, and one optimizer update every four environment steps. No additional hyperparameter tuning was performed for the layer-normalized agent.

    Layer normalization produced performance gains in many games, although the effect was not universal: the reported human-normalized improvement ranged from −52.4%-52.4\% on Double Dunk to +144.2%+144.2\% on Enduro, with large gains also reported for Space Invaders and Wizard of Wor. Games receiving substantial performance improvements often had degenerate gradient-covariance structure or ill-conditioned Hessians in the default architecture. Layer normalization tended to produce weaker gradient correlations in these cases. The authors treat this as evidence that regularizing optimization geometry is a promising route to more robust reinforcement learning, while explicitly noting that the benchmark gains cannot be definitively attributed to reduced plasticity loss alone.

Coverage note — Supplementary per-game Hessian plots, secondary TD-loss curves, and detailed initial-versus-final probe-loss breakdowns were omitted because they reinforce the included findings without constituting distinct central contributions.

References

  1. 1.Abbott, L. F. and Nelson, S. B. Synaptic plasticity: taming the beast. Nature neuroscience, 3(11):1178–1183, 2000.
  2. 2.Ash, J. and Adams, R. P. On warm-starting neural network training. Advances in Neural Information Processing Systems, 33:3884–3894, 2020.
  3. 3.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  4. 4.Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B. The shattered gradients problem: If resnets are the answer, then what is the question? In International Conference on Machine Learning, pp. 342–350. PMLR, 2017.
  5. 5.Bartlett, P. L. and Mendelson, S. Rademacher and gaussian complexities: Risk bounds and structural results. Journal of Machine Learning Research, 3(Nov):463–482, 2002.
  6. 6.Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, 2013.
  7. 7.Berariu, T., Czarnecki, W., De, S., Bornschein, J., Smith, S., Pascanu, R., and Clopath, C. A study on the plasticity of neural networks. arXiv preprint arXiv:2106.00042, 2021.
  8. 8.Bubeck, S. and Sellke, M. A universal law of robustness via isoperimetry. Journal of the ACM, 70(2):1–18, 2023.
  9. 9.Buhlmann, P. Invariance, causality and robustness. Statistical Science, 35(3):404–426, 2020.
  10. 10.Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A. Gradient descent on neural networks typically occurs at the edge of stability. arXiv preprint arXiv:2103.00065, 2021.
  11. 11.Delfosse, Q., Schramowski, P., Mundt, M., Molina, A., and Kersting, K. Adaptive rational activations to boost deep reinforcement learning. arXiv preprint arXiv:2102.09407, 2021.
  12. 12.Deng, L. The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine, 29(6):141–142, 2012.
  13. 13.Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. Sharp minima can generalize for deep nets. In International Conference on Machine Learning, pp. 1019–1028. PMLR, 2017.
  14. 14.Dohare, S., Mahmood, A. R., and Sutton, R. S. Continual backprop: Stochastic gradient descent with persistent randomness. arXiv preprint arXiv:2108.06325, 2021.
  15. 15.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  16. 16.Dziugaite, G. K., Drouin, A., Neal, B., Rajkumar, N., Caballero, E., Wang, L., Mitliagkas, I., and Roy, D. M. In search of robust measures of generalization. Advances in Neural Information Processing Systems, 33:11723–11733, 2020.
  17. 17.Fedus, W., Ghosh, D., Martin, J. D., Bellemare, M. G., Bengio, Y., and Larochelle, H. On catastrophic interference in atari 2600 games. arXiv preprint arXiv:2002.12499, 2020.
  18. 18.Fort, S., Nowak, P. K., and Narayanan, S. Stiffness: A new perspective on generalization in neural networks. CoRR, abs/1901.09491, 2019. URL http://arxiv.org/abs/1901.09491.
  19. 19.Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning, pp. 3259–3269. PMLR, 2020.
  20. 20.French, R. M. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135, 1999.
  21. 21.Ghorbani, B., Krishnan, S., and Xiao, Y. An investigation into neural net optimization via hessian eigenvalue density. In International Conference on Machine Learning, pp. 2232–2241. PMLR, 2019.
  22. 22.Gilmer, J., Ghorbani, B., Garg, A., Kudugunta, S. R., Neyshabur, B., Cardoze, D., Dahl, G. E., Nado, Z., and Firat, O. A loss curvature perspective on training instability in deep learning. In ICLR, 2022. URL https://arxiv.org/abs/2110.04369.
  23. 23.Glorot, X. and Bengio, Y. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 249–256. JMLR Workshop and Conference Proceedings, 2010.
  24. 24.Gogianu, F., Berariu, T., Rosca, M. C., Clopath, C., Busoniu, L., and Pascanu, R. Spectral normalisation for deep reinforcement learning: an optimisation perspective. In International Conference on Machine Learning, pp. 3734–3744. PMLR, 2021.
  25. 25.Gulcehre, C., Srinivasan, S., Sygnowski, J., Ostrovski, G., Farajtabar, M., Hoffman, M., Pascanu, R., and Doucet, A. An empirical study of implicit regularization in deep offline rl. Transactions of Machine Learning Research, 2022.
  26. 26.Hadsell, R., Rao, D., Rusu, A. A., and Pascanu, R. Embracing change: Continual learning in deep neural networks. Trends in Cognitive Sciences, 24(12):1028–1040, 2020. ISSN 1364-6613.
  27. 27.He, K., Zhang, X., Ren, S., and Sun, J. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pp. 1026–1034, 2015.
  28. 28.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  29. 29.Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S. Transient non-stationarity and generalisation in deep reinforcement learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Qun8fv4qSby.
  30. 30.Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho*, K., and Geras*, K. The break-even point on optimization trajectories of deep neural networks. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1g87C4KwB.
  31. 31.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  32. 32.Kemker, R., McClure, M., Abitino, A., Hayes, T., and Kanan, C. Measuring catastrophic forgetting in neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018.
  33. 33.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In ICLR (Poster), 2015.
  34. 34.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  35. 35.Kumar, A., Agarwal, R., Ghosh, D., and Levine, S. Implicit under-parameterization inhibits data-efficient deep reinforcement learning. In International Conference on Learning Representations, 2020.
  36. 36.Lee, S.-W., Kim, J.-H., Jun, J., Ha, J.-W., and Zhang, B.-T. Overcoming catastrophic forgetting by incremental moment matching. Advances in neural information processing systems, 30, 2017.
  37. 37.Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. The large learning rate phase of deep learning: the catapult mechanism. arXiv preprint arXiv:2003.02218, 2020.
  38. 38.Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. Visualizing the loss landscape of neural nets. Advances in neural information processing systems, 31, 2018.
  39. 39.Lyle, C., Rowland, M., and Dabney, W. Understanding and preventing capacity loss in reinforcement learning. In International Conference on Learning Representations, 2021.
  40. 40.Lyle, C., Rowland, M., Dabney, W., Kwiatkowska, M., and Gal, Y. Learning dynamics and generalization in deep reinforcement learning. In International Conference on Machine Learning, pp. 14560–14581. PMLR, 2022.
  41. 41.Martens, J., Ballard, A., Desjardins, G., Swirszcz, G., Dalibard, V., Sohl-Dickstein, J., and Schoenholz, S. S. Rapid training of deep neural networks without skip connections or normalization layers using deep kernel shaping. arXiv preprint arXiv:2110.01765, 2021.
  42. 42.Mermillod, M., Bugaiska, A., and Bonin, P. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013.
  43. 43.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015.
  44. 44.Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A. The primacy bias in deep reinforcement learning. In International Conference on Machine Learning, pp. 16828–16847. PMLR, 2022.
  45. 45.Park, N. and Kim, S. How do vision transformers work? In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=D78Go4hVcxO.
  46. 46.Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S. Exponential expressivity in deep neural networks through transient chaos. In Lee, D., Sugiyama, M., Luxburg, U., Guyon, I., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL https://proceedings.neurips.cc/paper/2016/file/148510031349642de5ca0c544f31b2ef-Paper.pdf.
  47. 47.Quan, J. and Ostrovski, G. DQN Zoo: Reference implementations of DQN-based agents, 2020. URL http://github.com/deepmind/dqn_zoo.
  48. 48.Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T., and Wayne, G. Experience replay for continual learning. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/fa7cdfad1a5aaf8370ebeda47a1ff1c3-Paper.pdf.
  49. 49.Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A. How does batch normalization help optimization? In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper/2018/file/905056c1ac1dad141560467e0a99e1cf-Paper.pdf.
  50. 50.Schmitt, S., Hudson, J. J., Zidek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Kuttler, H., Zisserman, A., Simonyan, K., et al. Kickstarting deep reinforcement learning. arXiv preprint arXiv:1803.03835, 2018.
  51. 51.Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J. Deep information propagation. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=H1W1UN9gg.
  52. 52.Sutskever, I., Martens, J., Dahl, G., and Hinton, G. On the importance of initialization and momentum in deep learning. In International conference on machine learning, pp. 1139–1147. PMLR, 2013.
  53. 53.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  54. 54.Van Hasselt, H., Guez, A., and Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI conference on artificial intelligence, volume 30, 2016.
  55. 55.Vapnik, V. On the uniform convergence of frequencies of occurrence of events to their probabilities. Theory of Probability and its Applications, 16, 1968.
  56. 56.Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. Dueling network architectures for deep reinforcement learning. In International conference on machine learning, pp. 1995–2003. PMLR, 2016.
  57. 57.Yang, G. and Schoenholz, S. Mean field residual networks: On the edge of chaos. Advances in neural information processing systems, 30, 2017.
  58. 58.Yang, G., Pennington, J., Rao, V., Sohl-Dickstein, J., and Schoenholz, S. S. A mean field theory of batch normalization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SyMDXnCcF7.
  59. 59.Zhang, C., Bengio, S., and Singer, Y. Are all layers created equal? Journal of Machine Learning Research, 2021a.
  60. 60.Zhang, G., Botev, A., and Martens, J. Deep learning without shortcuts: Shaping the kernel with tailored rectifiers. In International Conference on Learning Representations, 2021b.

Citation

MLA
Lyle, C., et al. “Understanding Plasticity in Neural Networks”. International Conference on Machine Learning, vol. 202, 2023, pp. 23190–211, https://proceedings.mlr.press/v202/lyle23b.html.
APA
Lyle, C., Zheng, Z., Nikishin, E., Pires, B. A., Pascanu, R., & Dabney, W. (2023). Understanding Plasticity in Neural Networks. International Conference on Machine Learning, 202, 23190–23211. https://proceedings.mlr.press/v202/lyle23b.html
Chicago
Lyle, C., Z. Zheng, E. Nikishin, B. A. Pires, R. Pascanu, and W. Dabney. 2023. “Understanding Plasticity in Neural Networks”. International Conference on Machine Learning 202: 23190–211. https://proceedings.mlr.press/v202/lyle23b.html.
Harvard
Lyle, C. et al. (2023) “Understanding Plasticity in Neural Networks”, International Conference on Machine Learning. PMLR, pp. 23190–23211. Available at: https://proceedings.mlr.press/v202/lyle23b.html.
Vancouver
1. Lyle C, Zheng Z, Nikishin E, Pires BA, Pascanu R, Dabney W (2023) Understanding Plasticity in Neural Networks. In: International Conference on Machine Learning. PMLR, pp 23190–23211

BibTeX

@InProceedings{pmlr-v202-lyle23b,
  title = 	 {Understanding Plasticity in Neural Networks},
  author =       {Lyle, Clare and Zheng, Zeyu and Nikishin, Evgenii and Avila Pires, Bernardo and Pascanu, Razvan and Dabney, Will},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {23190--23211},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/lyle23b/lyle23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/lyle23b.html},
  abstract = 	 {Plasticity, the ability of a neural network to quickly change its predictions in response to new information, is essential for the adaptability and robustness of deep reinforcement learning systems. Deep neural networks are known to lose plasticity over the course of training even in relatively simple learning problems, but the mechanisms driving this phenomenon are still poorly understood. This paper conducts a systematic empirical analysis into plasticity loss, with the goal of understanding the phenomenon mechanistically in order to guide the future development of targeted solutions. We find that loss of plasticity is deeply connected to changes in the curvature of the loss landscape, but that it often occurs in the absence of saturated units. Based on this insight, we identify a number of parameterization and optimization design choices which enable networks to better preserve plasticity over the course of training. We validate the utility of these findings on larger-scale RL benchmarks in the Arcade Learning Environment.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/