Wide Neural Networks Forget Less Catastrophically

Seyed-Iman MirzadehArslan ChaudhryDong YinHuiyi HuRazvan PascanuDilan GörürMehrdad Farajtabar

article2022ICML90 citations

Demonstrates that increasing network width significantly reduces catastrophic forgetting in continual learning—even matching the benefits of replay buffers—and explains this phenomenon through gradient orthogonality, activation sparsity, and the lazy training regime.

Listen

Modern artificial intelligence increasingly relies on continuous data streams where models must absorb new information over time. A central barrier in these continual learning settings is catastrophic forgetting, where a network abruptly erases previously learned capabilities upon training on new tasks. While prior research focused almost exclusively on specialized algorithms or memory replay buffers to mitigate forgetting, relatively little is known about how the underlying network architecture itself influences memory retention.

The article demonstrates that increasing network width substantially and consistently reduces catastrophic forgetting in continual learning. Specifically, it evaluates how structural over-parametrization through width versus depth impacts task retention, overall accuracy, and underlying optimization dynamics.

To evaluate this relationship, the authors conducted empirical experiments across standard continual learning benchmarks, including Rotated MNIST (domain-incremental image classification across five tasks) using multi-layer perceptrons and Split CIFAR-100 (task-incremental classification across twenty tasks) using Wide Residual Networks. The study tested varying widths and depths across extensive hyperparameter sweeps and multiple random seeds. Beyond standard accuracy and forgetting metrics, the evaluation tracked optimization properties such as gradient angles across tasks, gradient sparsity, parameter displacement, and layer-wise gradient norms.

The investigation produced several key findings. First, expanding network width significantly reduces forgetting while improving overall accuracy: expanding a two-layer perceptron's width from 32 to 2048 reduced first-task forgetting from 62% to 48%, while scaling WideResNet width eightfold reduced forgetting from 42% to 31%. Second, wider models match or exceed the retention benefits achieved by dedicated replay buffers without sacrificing the model's plasticity on new tasks. Third, over-parametrization via depth produces no such benefit and frequently worsens forgetting due to exploding gradient norms in earlier layers. Fourth, wider networks achieve higher retention because their task gradients are more orthogonal, gradient updates are sparser, and the weights remain closer to their initialization throughout training. Finally, under a fixed parameter budget, shallow and wide networks consistently outperform deep and thin networks, and width benefits are fully additive when combined with dedicated continual learning algorithms.

These findings indicate that catastrophic forgetting is heavily governed by model architecture, not just training algorithms. Engineering teams facing sequential learning challenges can improve retention by deploying wider architectures rather than simply increasing network depth. However, relying solely on extreme width creates substantial trade-offs in computational cost, memory footprint, and energy consumption. Therefore, scaling width is most effectively viewed as a foundational architectural complement to existing algorithmic techniques rather than a standalone replacement.

For practical application, practitioners should prioritize wider, shallower designs when configuring neural backbones for streaming or continual learning tasks. System architects should also account for network capacity when benchmarking continual learning methods to avoid misattributing width-driven improvements to algorithmic designs. Future research should evaluate how other architectural mechanisms—such as activation functions, normalization layers, and vision transformers—interact with continual learning dynamics across larger-scale enterprise applications.

The conclusions are supported by consistent results across multi-seed runs, diverse architectures, and multiple benchmarks. However, the study focuses primarily on controlled classification benchmarks and highlights that full mathematical bounds on deep network gradients remain loose when singular values exceed one. Leaders should account for the increased training compute and inference latency of wide architectures before executing large-scale production deployments.

Mirzadeh et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Wide Neural Networks Forget Less Catastrophically

Abstract

A primary focus area in continual learning research is alleviating the “catastrophic forgetting” problem in neural networks by designing new algorithms that are more robust to the distribution shifts. While the recent progress in continual learning literature is encouraging, our understanding of what properties of neural networks contribute to catastrophic forgetting is still limited. To address this, instead of focusing on continual learning algorithms, in this work, we focus on the model itself and study the impact of “width” of the neural network architecture on catastrophic forgetting, and show that width has a surprisingly significant effect on forgetting. To explain this effect, we study the learning dynamics of the network from various perspectives such as gradient orthogonality, sparsity, and lazy training regime. We provide potential explanations that are consistent with the empirical results across different architectures and continual learning benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Main Experiments
  • 3.1. Experimental Setup
  • 3.2. Main Results
  • 4. Analysis
  • 4.1. Increased Gradient Orthogonality
  • 4.2. Increased Gradient Sparsity
  • 4.3. Lazy Training Regime
  • 4.4. The Case for Depth
  • 5. Additional Experiments
  • 5.1. Width versus Depth
  • 5.2. Interacting with CL Algorithms
  • 6. Discussion and Conclusion
  • Acknowledgments
  • References
  • A. Experimental Setup Details
  • A.1. Design Choices
  • A.2. Hyperparameters
  • B. Additional Results
  • B.1. Evolution of the Average Accuracy Throughout the Learning Experience
  • B.2. Varying Width
  • B.3. Lazy Training Regime in Deeper Models
  • B.4. Detailed Results for Other Methods
  • B.5. Orthogonality Between Task 2 and Subsequent Tasks
  • B.6. Singular Values of Deeper Models
  • B.7. Interacting With Additional Algorithms
  • C. Task Construction in Claim 4.1

Knowls

  1. Knowl 1 — Increasing width improves continual-learning accuracy and reduces forgetting

    data/table

    The study compared two-layer ReLU MLPs on five-task Rotated MNIST and depth-10 WideResNets on 20-task Split CIFAR-100. Each task was trained for five epochs. Wider networks improved average accuracy and reduced average forgetting at every tested width; learning accuracy also increased. Results are means ± standard deviations across runs.

    Could not parse LaTeX table

    Here “avg. acc.” is average final accuracy over tasks, “avg. forg.” is average forgetting, “learn. acc.” is accuracy immediately after each task is learned, and “joint acc.” is accuracy after training on the combined data of all tasks. The results indicate that lower forgetting did not come at the expense of learning the current task: both average and learning accuracy rose with width. The authors emphasize that widening is a costly architectural change, not a claim of a compute-efficient state-of-the-art forgetting solution.

  2. Knowl 2 — Continual-learning benchmarks and experimental protocol

    experimental setup

    The experiments evaluated width and depth under two continual-learning settings. Rotated MNIST is a five-task domain-incremental benchmark, with MNIST images rotated by 0∘0^\circ, 22.5∘22.5^\circ, 45∘45^\circ, 67.5∘67.5^\circ, and 90∘90^\circ; its models were two-layer ReLU MLPs with varied widths. Split CIFAR-100 is a 20-task task-incremental benchmark, with five randomly selected classes per task and no class reuse; its models were WideResNets (WRNs) whose width factors scaled convolutional channel counts by 1,2,4,1,2,4, or 88. Models trained for five epochs on each task using SGD.

    For each architecture, the authors selected the best result from a hyperparameter grid: learning rates [0.001,0.01,0.05,0.1][0.001,0.01,0.05,0.1] (with 0.010.01 designated for MNIST and 0.050.05 for CIFAR in the reported grid), batch sizes 1616, 6464 for MLPs, and 3232 for CIFAR, momentum 00 for MNIST and 0.80.8 for CIFAR, and weight decay 00 for MNIST and 0.00010.0001 for CIFAR. Results were averaged over five network-initialization seeds; Split CIFAR-100 additionally used three task-order seeds.

  3. Knowl 3 — Definitions of the continual-learning evaluation metrics

    definition

    Let TT be the number of tasks, and let at,ia_{t,i} be validation accuracy on task ii after training through task tt. The final average accuracy is

    AT=1T∑i=1TaT,i.A_T=\frac{1}{T}\sum_{i=1}^{T}a_{T,i}.

    Average forgetting excludes the last task and averages, over the other tasks, the loss from each task’s best accuracy during the continual-learning sequence to its final accuracy:

    F=1T−1∑i=1T−1(max⁡t∈{1,…,T−1}at,i−aT,i).F=\frac{1}{T-1}\sum_{i=1}^{T-1}\left(\max_{t\in\{1,\ldots,T-1\}}a_{t,i}-a_{T,i}\right).

    Learning accuracy measures performance immediately after each task is learned:

    LAT=1T∑i=1Tai,i.LA_T=\frac{1}{T}\sum_{i=1}^{T}a_{i,i}.

    Joint accuracy is the model’s accuracy after training on the combined data of all tasks. Accuracies and forgetting are reported in percentage points in the experiments.

  4. Knowl 4 — Equal-capacity models can forget differently because of their parameterization

    theoretical result

    For inputs x∈Rdx\in\mathbb{R}^d and scalar outputs, consider a linear model fw(x)=x⊤wf_w(x)=x^\top w, with w∈Rdw\in\mathbb{R}^d, and a two-layer linear model fU,v(x)=v⊤Uxf_{U,v}(x)=v^\top Ux, with U∈Rh×dU\in\mathbb{R}^{h\times d} and v∈Rhv\in\mathbb{R}^h. These model classes represent the same set of functions, since the two-layer model has no nonlinear activation. Nevertheless, for two tasks whose input subspaces are orthogonal, after fitting task 1 with zero training loss and then using gradient descent on task 2, the linear parameterization preserves task-1 predictions, whereas the factored two-layer parameterization can incur positive task-1 forgetting.

    In the paper’s construction, let X1,X2∈Rn×dX_1,X_2\in\mathbb{R}^{n\times d} be the task input matrices and assume X1X2⊤=0X_1X_2^\top=0. For the two-layer model, positive forgetting is guaranteed when h≤nh\le n, X1(U0)⊤X_1(U^0)^\top has rank hh, and the task-2 training changes the output-layer parameter, vT≠v0v^T\ne v^0. Here (U0,v0)(U^0,v^0) and (UT,vT)(U^T,v^T) are the two-layer parameters at the start and end of task-2 training. The example shows that equal function-class capacity alone does not determine forgetting under gradient-based continual learning.

  5. Knowl 5 — Wider networks exhibit more orthogonal task gradients

    empirical result

    For a differentiable two-task loss, let w∈Rpw\in\mathbb{R}^p be the current parameter vector, let η>0\eta>0 be a gradient-descent step size, and let L1,L2L_1,L_2 be the task losses. After one step on task 2, w′=w−η∇L2(w)w'=w-\eta\nabla L_2(w). There is a ξ∈[0,1]\xi\in[0,1] such that

    L1(w′)−L1(w)=−η⟨∇L1(w−ξη∇L2(w)),∇L2(w)⟩.L_1(w')-L_1(w)=-\eta\left\langle\nabla L_1\bigl(w-\xi\eta\nabla L_2(w)\bigr),\nabla L_2(w)\right\rangle.

    Thus the change in task-1 loss depends on an inner product between the task gradients at nearby parameter values; small inner products, as occur for near-orthogonal gradients, limit one-step forgetting. Empirically, the authors concatenated gradients across all layers and measured angles between gradients at the task-1 optimum and subsequent task optima. For both Rotated MNIST MLPs and Split CIFAR-100 WRNs, those angles generally moved closer to 90∘90^\circ as width increased. The trend also held when gradients from task 2 were used as the reference, supporting—but not proving—the proposed explanation that width reduces forgetting through gradient orthogonalization.

  6. Knowl 6 — Wider networks have sparser gradients, which can limit one-step forgetting

    theoretical result

    Let ww be the current parameter vector, L1L_1 and L2L_2 differentiable task losses, and w′=w−η∇L2(w)w'=w-\eta\nabla L_2(w) a gradient-descent update on task 2 with step size η>0\eta>0. Suppose ∥∇L2(w)∥2≤B\|\nabla L_2(w)\|_2\le B throughout the parameter domain, and define

    b(w)=sup⁡∥w~−w∥2≤ηB∥∇L1(w~)∥∞.b(w)=\sup_{\|\widetilde w-w\|_2\le\eta B}\|\nabla L_1(\widetilde w)\|_\infty.

    Then the task-1 loss increase satisfies

    L1(w′)−L1(w)≤ηb(w)∥∇L2(w)∥0,L_1(w')-L_1(w)\le\eta b(w)\|\nabla L_2(w)\|_0,

    where ∥⋅∥0\|\cdot\|_0 counts nonzero vector entries. Under these conditions, an update with fewer nonzero gradient coordinates admits a smaller upper bound on forgetting. In the experiments, gradient magnitudes were measured across all layers on task 2 after task 1 had been learned and before task-2 training began. Wider MLPs and WRNs had distributions more concentrated near zero, which the authors interpret as evidence that fewer parameters need substantial changes when adapting to a new task.

  7. Knowl 7 — Wider networks move less in parameter space during continual training

    theoretical result

    Let wt∗w_t^* be the parameter vector at the end of training task tt, and let Lt−1L_{t-1} be the loss on the preceding task. Define the parameter displacement during task tt as Dt=∥wt∗−wt−1∗∥2D_t=\|w_t^*-w_{t-1}^*\|_2. For differentiable Lt−1L_{t-1}, the increase in that loss is bounded by

    Lt−1(wt∗)−Lt−1(wt−1∗)≤Dtsup⁡{w:∥w−wt−1∗∥2≤Dt}∥∇Lt−1(w)∥2.L_{t-1}(w_t^*)-L_{t-1}(w_{t-1}^*)\le D_t\sup_{\{w:\|w-w_{t-1}^*\|_2\le D_t\}}\|\nabla L_{t-1}(w)\|_2.

    The bound relates forgetting to how far training on the new task moves parameters away from the previous solution. Across the Rotated MNIST MLP and Split CIFAR-100 WRN experiments, wider models stayed substantially closer to the task-1 solution over subsequent training than narrower models. The authors interpret this as wider networks behaving more like the lazy-training regime, in which parameters move relatively little; this is offered as a potential explanation for the empirical reduction in forgetting.

  8. Knowl 8 — Increasing depth can amplify early-layer gradients and forgetting

    theoretical result

    Consider a feed-forward network with KK layers, weight matrices W1,…,WKW_1,\ldots,W_K, hidden representations h1,…,hKh_1,\ldots,h_K, and loss LL. Ignoring biases and normalization layers, backpropagation gives an early-layer gradient involving the product of Jacobians of later layers. With ReLU activations, whose diagonal Jacobians have spectral norm at most 11, the gradient norm obeys

    ∥∂L∂hl∥2≤∥∂L∂hK∥2∏i=lK−1∥Wi+1∥2,\left\|\frac{\partial L}{\partial h_l}\right\|_2\le \left\|\frac{\partial L}{\partial h_K}\right\|_2 \prod_{i=l}^{K-1}\|W_{i+1}\|_2,

    where ∥⋅∥2\|\cdot\|_2 on matrices denotes the spectral norm. If the relevant weight matrices have singular values greater than 11, this bound indicates a possible amplification of gradients in earlier layers as depth grows. The authors caution that it can become very loose and is not itself a general explanation or guarantee of exploding gradients. In their experiments, increasing depth did not produce the monotonic reduction in forgetting seen with width: forgetting rose in the tested MLP and WRN depth sweeps. For the Rotated MNIST MLP, the first-layer gradient norm at task 5 was almost three times larger for depth 8 than for depth 2, while increasing width had minimal or decreasing effect on that norm.

  9. Knowl 9 — At roughly matched parameter counts, wider and shallower networks perform better

    data/table

    The authors compared architectures with approximately matched parameter counts on both benchmarks. In each comparison, the shallower, wider model achieved higher average accuracy and lower average forgetting than the deeper, thinner alternative. The data also show that this width advantage is not merely the consequence of comparing models with very different parameter budgets.

    Could not parse LaTeX table

    Each benchmark comparison is grouped by roughly similar parameter budgets. Average accuracy and forgetting are the continual-learning metrics; joint accuracy is performance after training on all task data together. The results favor width over depth for continual retention at comparable model size.

  10. Knowl 10 — Width gains remain when replay or gradient-based continual-learning methods are used

    empirical result

    The width benefit was tested alongside Experience Replay (ER) and A-GEM, using replay buffers of 125 examples on Rotated MNIST and 100 examples on Split CIFAR-100. Average accuracy increased and average forgetting decreased with every tested width for both methods. The following endpoint results give average accuracy / average forgetting:

    Could not parse LaTeX table

    These results support the authors’ conclusion that widening can complement, rather than replace, continual-learning algorithms. Additional Rotated MNIST experiments reported the same direction of change with MC-SGD, LwF, and EWC, although the size of the gain varied by method.

Coverage note — The supplementary three-layer MLP ablation locating a stronger effect in the first layer was omitted because it is secondary to the paper’s main width, depth, and learning-dynamics findings; detailed per-layer ablation values are not needed to reconstruct those central contributions.

References

  1. 1.Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M., and Tuytelaars, T. (2018). Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision (ECCV), pages 139–154.
  2. 2.Aljundi, R., Chakravarty, P., and Tuytelaars, T. (2017). Expert gate: Lifelong learning with a network of experts. In CVPR, pages 7120–7129.
  3. 3.Allen-Zhu, Z., Li, Y., and Liang, Y. (2019). Learning and generalization in overparameterized neural networks, going beyond two layers. Advances in neural information processing systems.
  4. 4.Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R. (2019). On exact computation with an infinitely wide neural net. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pages 8141–8150.
  5. 5.Balaji, Y., Farajtabar, M., Yin, D., Mott, A., and Li, A. (2020). The effectiveness of memory replay in large scale continual learning. arXiv preprint arXiv:2010.02418.
  6. 6.Beaulieu, S., Frati, L., Miconi, T., Lehman, J., Stanley, K. O., Clune, J., and Cheney, N. (2020). Learning to continually learn. In ECAI 2020 - 24th European Conference on Artificial Intelligence.
  7. 7.Bengio, Y., Simard, P., and Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. IEEE transactions on neural networks, 5(2):157–166.
  8. 8.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020.
  9. 9.Chang, M., Gupta, A., Levine, S., and Griffiths, T. L. (2018). Automatically composing representation transformations as a means for generalization. In ICML workshop Neural Abstract Machines and Program Induction v2.
  10. 10.Chaudhry, A., Gordo, A., Dokania, P. K., Torr, P. H. S., and Lopez-Paz, D. (2021). Using hindsight to anchor past knowledge in continual learning. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021.
  11. 11.Chaudhry, A., Khan, N., Dokania, P. K., and Torr, P. H. (2020). Continual learning in low-rank orthogonal subspaces. In Advances in Neural Information Processing Systems.
  12. 12.Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M. (2018). Efficient lifelong learning with a-gem. In International Conference on Learning Representations.
  13. 13.Chaudhry, A., Rohrbach, M., Elhoseiny, M., Ajanthan, T., Dokania, P. K., Torr, P. H. S., and Ranzato, M. (2019). On tiny episodic memories in continual learning. arXiv preprint arXiv:1902.10486.
  14. 14.Chizat, L., Oyallon, E., and Bach, F. (2019). On lazy training in differentiable programming. In NeurIPS 2019-33rd Conference on Neural Information Processing Systems, pages 2937–2947.
  15. 15.Doan, T., Bennani, M. A., Mazoure, B., Rabusseau, G., and Alquier, P. (2021). A theoretical analysis of catastrophic forgetting through the ntk overlap matrix. In International Conference on Artificial Intelligence and Statistics.
  16. 16.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR, 2021.
  17. 17.Du, S., Lee, J., Li, H., Wang, L., and Zhai, X. (2019). Gradient descent finds global minima of deep neural networks. In International Conference on Machine Learning, pages 1675–1685. PMLR.
  18. 18.Farajtabar, M., Azizan, N., Mott, A., and Li, A. (2020). Orthogonal gradient descent for continual learning. In International Conference on Artificial Intelligence and Statistics, pages 3762–3773. PMLR.
  19. 19.Fernando, C., Banarse, D., Blundell, C., Zwols, Y., Ha, D., Rusu, A. A., Pritzel, A., and Wierstra, D. (2017). Pathnet: Evolution channels gradient descent in super neural networks. arXiv preprint arXiv:1701.08734.
  20. 20.Ferran Alet, Tomas Lozano-Perez, L. P. K. (2018). Modular meta-learning. arXiv preprint arXiv:1806.10166v1.
  21. 21.Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep learning. MIT press.
  22. 22.Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y. (2013). An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211.
  23. 23.He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  24. 24.Hombaiah, S. A., Chen, T., Zhang, M., Bendersky, M., and Najork, M. (2021). Dynamic language models for continuously evolving content. In KDD ’21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining.
  25. 25.Hsu, Y.-C., Liu, Y.-C., and Kira, Z. (2018). Re-evaluating continual learning scenarios: A categorization and case for strong baselines. arXiv preprint arXiv:1810.12488.
  26. 26.Jacot, A., Gabriel, F., and Hongler, C. (2018). Neural tangent kernel: convergence and generalization in neural networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pages 8580–8589.
  27. 27.Javed, K. and White, M. (2019). Meta-learning representations for continual learning. In Advances in Neural Information Processing Systems, pages 1820–1830.
  28. 28.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  29. 29.Kirichenko, P., Farajtabar, M., Rao, D., Lakshminarayanan, B., Levine, N., Li, A., Hu, H., Wilson, A. G., and Pascanu, R. (2021). Task-agnostic continual learning with hybrid probabilistic models. In ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models.
  30. 30.Kirkpatrick, J. N., Pascanu, R., Rabinowitz, N. C., Veness, J., and et. al. (2017). Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences of the United States of America, 114 13:3521–3526.
  31. 31.Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25:1097–1105.
  32. 32.Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., d’Autume, C. d. M., Ruder, S., Yogatama, D., et al. (2021). Pitfalls of static language modelling. arXiv preprint arXiv:2102.01951.
  33. 33.Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2019). Wide neural networks of any depth evolve as linear models under gradient descent. Advances in neural information processing systems, 32:8572–8583.
  34. 34.Li, X., Zhou, Y., Wu, T., Socher, R., and Xiong, C. (2019). Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting. In Proceedings of the 36th International Conference on Machine Learning, ICML, Proceedings of Machine Learning Research.
  35. 35.Li, Z. and Hoiem, D. (2018). Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40:2935–2947.
  36. 36.Lopez-Paz, D. and Ranzato, M. (2017). Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems, pages 6467–6476.
  37. 37.McClelland, J. L., McNaughton, B. L., and O’Reilly, R. C. (1995). Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory. Psychological review, 102(3):419.
  38. 38.McCloskey, M. and Cohen, N. J. (1989). Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of Learning and Motivation, 24:109–165.
  39. 39.Mehta, S. V., Patil, D., Chandar, S., and Strubell, E. (2021). An empirical investigation of the role of pre-training in lifelong learning. *ICML CL Workshop,. *
  40. 40.Mirzadeh, S. I., Chaudhry, A., Yin, D., Nguyen, T., Pascanu, R., Gorur, D., and Farajtabar, M. (2022). Architecture matters in continual learning. ArXiv, abs/2202.00275.
  41. 41.Mirzadeh, S. I., Farajtabar, M., and Ghasemzadeh, H. (2020a). Dropout as an implicit gating mechanism for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 232–233.
  42. 42.Mirzadeh, S. I., Farajtabar, M., Gorur, D., Pascanu, R., and Ghasemzadeh, H. (2021). Linear mode connectivity in multitask and continual learning. In International Conference on Learning Representations.
  43. 43.Mirzadeh, S. I., Farajtabar, M., Pascanu, R., and Ghasemzadeh, H. (2020b). Understanding the role of training regimes in continual learning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020.
  44. 44.Nguyen, C. V., Achille, A., Lam, M., Hassner, T., Mahadevan, V., and Soatto, S. (2019). Toward understanding catastrophic forgetting in continual learning. arXiv preprint arXiv:1908.01091.
  45. 45.Nguyen, C. V., Li, Y., Bui, T. D., and Turner, R. E. (2018). Variational continual learning. In 6th International Conference on Learning Representations, ICLR 2018.
  46. 46.Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., and Wermter, S. (2019). Continual lifelong learning with neural networks: A review. Neural Networks, 113:54–71.
  47. 47.Pascanu, R., Mikolov, T., and Bengio, Y. (2013). On the difficulty of training recurrent neural networks. In International conference on machine learning, pages 1310–1318. PMLR.
  48. 48.Ramasesh, V. V., Lewkowycz, A., and Dyer, E. (2022). Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations.
  49. 49.Rebuffi, S.-A., Kolesnikov, A. I., Sperl, G., and Lampert, C. H. (2016). icarl: Incremental classifier and representation learning. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5533–5542.
  50. 50.Riemer, M., Cases, I., Ajemian, R., Liu, M., Rish, I., Tu, Y., and Tesauro, G. (2018). Learning to learn without forgetting by maximizing transfer and minimizing interference. In International Conference on Learning Representations.
  51. 51.Ring, M. B. (1995). Continual learning in reinforcement environments. PhD thesis, University of Texas at Austin, TX, USA.
  52. 52.Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. P., and Wayne, G. (2019). Experience replay for continual learning. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019.
  53. 53.Rosenbaum, C., Klinger, T., and Riemer, M. (2018). Routing networks: Adaptive selection of non-linear functions for multi-task learning. In International Conference on Learning Representations.
  54. 54.Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. (2016). Progressive neural networks. arXiv preprint arXiv:1606.04671.
  55. 55.Shin, H., Lee, J. K., Kim, J., and Kim, J. (2017). Continual learning with deep generative replay. In Advances in Neural Information Processing Systems, pages 2990–2999.
  56. 56.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016). Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826.
  57. 57.Thrun, S. (1995). A lifelong learning perspective for mobile robot control. In Intelligent robots and systems, pages 201–214. Elsevier.
  58. 58.Titsias, M. K., Schwarz, J., Matthews, A. G. d. G., Pascanu, R., and Teh, Y. W. (2019). Functional regularisation for continual learning using gaussian processes. arXiv preprint arXiv:1901.11356.
  59. 59.Toneva, M., Sordoni, A., Combes, R. T. d., Trischler, A., Bengio, Y., and Gordon, G. J. (2019). An empirical study of example forgetting during deep neural network learning. In 7th International Conference on Learning Representations, ICLR 2019.
  60. 60.Veniat, T., Denoyer, L., and Ranzato, M. (2021). Efficient continual learning with modular networks and task-driven priors. In ICLR.
  61. 61.Wallingford, M., Kusupati, A., Alizadeh-Vahid, K., Walsman, A., Kembhavi, A., and Farhadi, A. (2020). In the wild: From ml models to pragmatic ml systems. ArXiv, abs/2007.02519.
  62. 62.Wortsman, M., Ramanujan, V., Liu, R., Kembhavi, A., Rastegari, M., Yosinski, J., and Farhadi, A. (2020). Supermasks in superposition. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020.
  63. 63.Xu, J. and Zhu, Z. (2018). Reinforced continual learning. In Advances in Neural Informatio Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018.
  64. 64.Yin, D., Farajtabar, M., Li, A., Levine, N., and Mott, A. (2020). Optimization and generalization of regularization-based continual learning: a loss approximation viewpoint. arXiv preprint arXiv:2006.10974.
  65. 65.Yoon, J., Yang, E., Lee, J., and Hwang, S. J. (2018). Lifelong learning with dynamically expandable networks. In Sixth International Conference on Learning Representations. ICLR.
  66. 66.Zagoruyko, S. and Komodakis, N. (2016). Wide residual networks. In British Machine Vision Conference 2016. British Machine Vision Association.
  67. 67.Zenke, F., Poole, B., and Ganguli, S. (2017). Continual learning through synaptic intelligence. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 3987–3995. JMLR.
  68. 68.Zou, D., Cao, Y., Zhou, D., and Gu, Q. (2020). Gradient descent optimizes over-parameterized deep relu networks. Machine Learning, 109(3):467–492.

Citation

MLA
Mirzadeh, S. I., et al. “Wide Neural Networks Forget Less Catastrophically”. International Conference on Machine Learning, vol. 162, 2022, pp. 15699–717, https://proceedings.mlr.press/v162/mirzadeh22a.html.
APA
Mirzadeh, S. I., Chaudhry, A., Yin, D., Hu, H., Pascanu, R., Gorur, D., & Farajtabar, M. (2022). Wide Neural Networks Forget Less Catastrophically. International Conference on Machine Learning, 162, 15699–15717. https://proceedings.mlr.press/v162/mirzadeh22a.html
Chicago
Mirzadeh, S. I., A. Chaudhry, D. Yin, et al. 2022. “Wide Neural Networks Forget Less Catastrophically”. International Conference on Machine Learning 162: 15699–717. https://proceedings.mlr.press/v162/mirzadeh22a.html.
Harvard
Mirzadeh, S.I. et al. (2022) “Wide Neural Networks Forget Less Catastrophically”, International Conference on Machine Learning. PMLR, pp. 15699–15717. Available at: https://proceedings.mlr.press/v162/mirzadeh22a.html.
Vancouver
1. Mirzadeh SI, Chaudhry A, Yin D, Hu H, Pascanu R, Gorur D, Farajtabar M (2022) Wide Neural Networks Forget Less Catastrophically. In: International Conference on Machine Learning. PMLR, pp 15699–15717

BibTeX

@InProceedings{pmlr-v162-mirzadeh22a,
  title = 	 {Wide Neural Networks Forget Less Catastrophically},
  author =       {Mirzadeh, Seyed Iman and Chaudhry, Arslan and Yin, Dong and Hu, Huiyi and Pascanu, Razvan and Gorur, Dilan and Farajtabar, Mehrdad},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {15699--15717},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/mirzadeh22a/mirzadeh22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/mirzadeh22a.html},
  abstract = 	 {A primary focus area in continual learning research is alleviating the "catastrophic forgetting" problem in neural networks by designing new algorithms that are more robust to the distribution shifts. While the recent progress in continual learning literature is encouraging, our understanding of what properties of neural networks contribute to catastrophic forgetting is still limited. To address this, instead of focusing on continual learning algorithms, in this work, we focus on the model itself and study the impact of "width" of the neural network architecture on catastrophic forgetting, and show that width has a surprisingly significant effect on forgetting. To explain this effect, we study the learning dynamics of the network from various perspectives such as gradient orthogonality, sparsity, and lazy training regime. We provide potential explanations that are consistent with the empirical results across different architectures and continual learning benchmarks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/