A Continual Learning Survey: Defying Forgetting in Classification Tasks

Matthias De LangeRahaf AljundiMarc MasanaSarah ParisotXu JiaAleš LeonardisGreg SlabaughTinne Tuytelaars

article2019TPAMI2,524 citations

Presents a structured taxonomy and a dynamic hyperparameter selection framework to benchmark eleven continual learning methods across diverse classification datasets, showing how model capacity, regularization, and parameter isolation directly affect catastrophic forgetting.

Listen

The article addresses the challenge of catastrophic forgetting in artificial neural networks, where models lose performance on earlier tasks after learning new ones. This problem arises in dynamic real-world settings with streaming data, where full retraining is often infeasible due to storage limits or privacy rules.

The study evaluates approaches for task incremental classification, seeking to let networks accumulate knowledge sequentially. It organizes existing methods into a taxonomy, proposes a framework that sets the stability-plasticity trade-off using only current-task data, and compares 11 representative techniques plus baselines on multiple benchmarks.

Experiments used the balanced Tiny Imagenet dataset along with unbalanced sequences from iNaturalist and RecogSeq. Tests examined effects of model size, dropout, weight decay, and task order. Parameter isolation methods such as PackNet reached the highest final accuracy with zero forgetting on balanced data, while iCaRL performed nearly as well when given larger replay buffers. Regularization methods showed mixed results, with MAS proving most robust; data-focused methods like LwF degraded sharply on dissimilar tasks.

These outcomes indicate that isolation techniques suit multi-head task-incremental settings but cannot scale easily to shared-head or class-incremental cases. Replay and regularization approaches allow greater flexibility yet demand careful hyperparameter selection to prevent inflated performance claims. Task ordering had negligible impact across all methods.

Practitioners should apply the proposed continual hyperparameter framework and favor isolation methods when task identity is known at test time. Further work is needed to handle unbounded data streams without explicit task boundaries and to reduce reliance on stored raw samples. The study is limited to classification problems with known boundaries, so results should be applied cautiously outside this scope.

  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). This work introduces Elastic Weight Consolidation (EWC), a foundational parameter regularization baseline evaluated directly in the survey.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This paper establishes Learning without Forgetting (LwF), one of the primary data-focused distillation approaches benchmarked and analyzed in the survey.
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). This paper proposes iCaRL, providing the core exemplar replay and distillation framework that the survey evaluates against regularization and isolation baselines.
  • Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). This paper introduces Memory Aware Synapses (MAS), a key unsupervised parameter-importance method analyzed and highlighted for its robustness in the survey.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). This paper defines Synaptic Intelligence (SI), presenting the path-integral-based parameter consolidation method evaluated in the survey's comparative study.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). This paper introduces Gradient Episodic Memory (GEM) and formalizes the standard metrics for backward and forward transfer in task-incremental continual learning.
  • Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). This paper introduces deep generative replay for continual learning, forming the basis for memory replay strategies without raw sample storage evaluated in the survey.
  • Paper: Continual Lifelong Learning with Neural Networks: A Review, German I. Parisi et al. (2018). This earlier review outlines the overarching continual learning taxonomy and biological foundations that the survey builds upon and empirically evaluates.
Cover for A Continual Learning Survey: Defying Forgetting in Classification Tasks

Abstract

Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods and 4 baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time and storage.

Table of Contents

  • I Introduction
  • II The Task Incremental Learning Setting
  • III Continual Learning Approaches
  • III-A Replay methods
  • III-B Regularization-based methods
  • III-B1 Data-focused methods
  • III-B2 Prior-focused methods
  • III-C Parameter isolation methods
  • IV Continual Hyperparameter Framework
  • V Compared Methods
  • V-A Replay methods
  • V-B Regularization-based methods
  • V-C Parameter isolation methods
  • VI Experiments
  • VI-A Experimental Setup
  • VI-B Comparing Methods on the Base Network
  • VI-B1 Discussion
  • VI-C Effects of Model Capacity
  • VI-C1 Discussion
  • VI-D Effects of Regularization
  • VI-D1 Discussion
  • VI-E Effects of a Real-world Setting
  • VI-E1 Discussion
  • VI-F Effects of Task Order
  • VI-F1 Discussion
  • VI-G Qualitative Comparison
  • VI-H Experiments Summary
  • VII Looking ahead
  • VIII Related machine learning fields
  • IX Conclusion
  • References
  • A Implementation Details
  • A-A Model Setup
  • A-B Framework Setup
  • A-C Methods Setup
  • A-D Qualitative Comparison Setup
  • B Supplemental Results
  • B-A Complete Mean-IMM Results
  • B-B Synaptic Intelligence (SI): Overfitting and Regularization
  • B-C Extra Replay Experiments: Epoch Sensitivity
  • B-D Reproduced RecogSeq Results
  • B-E HAT Analysis: Asymmetric Capacity Allocation
  • B-F PackNet Long Sequence Capacity Analysis
  • B-F1 Decreased Model Capacity
  • B-F2 Long Task Sequence Experiment

Knowls

  1. Knowl 1 — Continual Hyperparameter Selection Framework

    algorithm

    Continual learning methods require hyperparameters that balance the stability-plasticity trade-off. Standard grid search over held-out data from all tasks violates the continual learning constraint of having access only to current task data. The continual hyperparameter framework dynamically configures learning rate and stability parameters on each task t+1t+1 using only its local data Dt+1\mathcal{D}^{t+1}:

    1. Maximal Plasticity Search: A copy of previous model parameters θt\theta^t is finetuned on Dt+1\mathcal{D}^{t+1} over a coarse grid of learning rates Ψ\Psi. The learning rate η\eta^* yielding the maximum validation accuracy AA^* on the new task is selected, establishing an unconstrained plasticity upper bound.
    2. Stability Decay: The continual learning method is executed with η\eta^* initialized with maximum stability hyperparameter values H\mathcal{H} (e.g., maximum penalty on parameter drift or distillation weight). If validation accuracy AA is below (1p)A(1-p)A^* for a tolerated drop margin p[0,1]p \in [0, 1], hyperparameters in H\mathcal{H} are decayed by factor α[0,1]\alpha \in [0, 1] via HαH\mathcal{H} \leftarrow \alpha \cdot \mathcal{H} to incrementally grant plasticity until the threshold is met. Decayed hyperparameters are propagated to future tasks.
    Input: hyperparameter set H\mathcal{H}, decay factor α[0,1]\alpha \in [0, 1], accuracy drop margin p[0,1]p \in [0, 1], new task data Dt+1\mathcal{D}^{t+1}, learning rate grid Ψ\Psi
    Input: previous task model parameters θt\theta^t, continual learning method CLM\text{CLM}
    Output: updated model parameters θt+1\theta^{t+1}, adapted hyperparameter set H\mathcal{H}
    // Maximal Plasticity Search
    A0A^* \leftarrow 0
    η0\eta^* \leftarrow 0
    for ηΨ\eta \in \Psi do
        AFinetune(Dt+1,η;θt)A \leftarrow \text{Finetune}(\mathcal{D}^{t+1}, \eta; \theta^t)
        if A>AA > A^* then
            AAA^* \leftarrow A
            ηη\eta^* \leftarrow \eta
    // Stability Decay
    do
        A,θt+1CLM(Dt+1,η,H;θt)A, \theta^{t+1} \leftarrow \text{CLM}(\mathcal{D}^{t+1}, \eta^*, \mathcal{H}; \theta^t)
        if A<(1p)AA < (1 - p)A^* then
            HαH\mathcal{H} \leftarrow \alpha \cdot \mathcal{H}
    while A<(1p)AA < (1 - p)A^*
    return θt+1,H\theta^{t+1}, \mathcal{H}
  2. Knowl 2 — Taxonomy of Continual Learning Methods

    definition

    Continual learning methods for preventing catastrophic forgetting in neural networks are structured into three principal families based on how task-specific knowledge is retained and updated:

    1. Replay Methods: Alleviate forgetting by retaining past task information via data.
      • Rehearsal: Explicitly re-trains on a stored subset of past samples (exemplars) alongside new task data.
      • Pseudo-Rehearsal: Generates synthetic samples of past distributions using a generative model and replays them during new task training.
      • Constrained Optimization: Uses stored historical exemplars to constrain gradient updates for the new task, projecting the update vector onto a feasible half-space that does not increase historical task loss.
    2. Regularization-Based Methods: Augment the loss function with consolidation penalties to avoid storing raw input data.
      • Data-Focused: Distills prediction logits or latent feature representations from previously trained models on current task samples.
      • Prior-Focused: Estimates parameter importance distributions as Bayesian priors, adding quadratic penalties to restrict shifts in parameters deemed critical for previous tasks.
    3. Parameter Isolation Methods: Dedicates distinct network parameters to each individual task to eliminate parameter interference.
      • Fixed Network: Applies masks or pruning strategies within a static architecture to assign dedicated sub-networks per task.
      • Dynamic Architectures: Allocates new structural branches, modules, or network copies for each incoming task while freezing existing parameters.
  3. Knowl 3 — Empirical Performance Comparison of Continual Learning Methods Across Benchmarks

    data/table

    The table below compares the best average classification accuracy (%) achieved by representative continual learning methods across three distinct benchmarks: Tiny ImageNet (10 balanced tasks, 20 classes each, trained from scratch), iNaturalist (10 fine-grained, highly unbalanced super-category tasks, using pretrained AlexNet), and RecogSeq (8 diverse image recognition datasets spanning flowers, scenes, birds, cars, aircrafts, actions, letters, and digits, using pretrained AlexNet).

    Method Tiny ImageNet iNaturalist RecogSeq
    Finetuning (Lower Bound) 21.30 45.59 23.11
    Joint Training (Upper Bound) 55.70 63.90 65.82
    Replay Methods
    iCaRL (4.5k buffer) 48.55
    iCaRL (9k buffer) 49.94
    GEM (4.5k buffer) 45.27
    GEM (9k buffer) 44.23
    Regularization-Based
    LwF 48.11 48.02 30.59
    EBLL 48.17 53.30 33.82
    SI 43.74 51.77 43.40
    EWC 45.13 54.02 42.01
    MAS 48.98 54.59 45.72
    mean-IMM 32.42 49.82 31.43
    mode-IMM 42.41 55.61 34.45
    Parameter Isolation
    PackNet 55.96 60.61 64.88
    HAT 44.19

    PackNet attains top accuracy across all setups by pruning and freezing parameter subsets, guaranteeing 0% forgetting. Among regularization-based approaches, MAS displays the highest robustness and consistent performance across diverse task shifts. Data-focused methods (LwF and EBLL) perform competitively on homogeneous tasks (Tiny ImageNet) but experience catastrophic degradation under severe domain distribution shifts (RecogSeq). Mode-IMM surpasses mean-IMM across all datasets and benefits from positive backward transfer in unbalanced settings.

  4. Knowl 4 — Model Capacity and Architecture Effects in Continual Learning

    empirical result

    Evaluating continual learning across different convolutional architectures (SMALL with 613k parameters, BASE with 3.51M parameters, WIDE with 8.95M parameters, and DEEP with 6.64M parameters and 20 conv layers) demonstrates clear architectural preferences:

    • Deep Architectures: Deep networks (DEEP) consistently yield inferior continual learning and joint training accuracy across all method families compared to shallower models. The additional layers of abstraction aggravate overfitting on early tasks and hinder parameter accumulation across sequential task shifts.
    • Wide Architectures: For an equivalent number of feature extraction parameters, WIDE architectures achieve on average 11% higher average accuracy than DEEP architectures across all continual learning methods.
    • Small Architectures: Under-parameterized models (SMALL) suffer from accelerated catastrophic forgetting (on average 0.89% higher forgetting than BASE) due to lack of spare capacity for new task representations. While EWC achieves its peak performance on SMALL networks, methods like Finetuning and Synaptic Intelligence suffer severe degradation (>20% forgetting).
  5. Knowl 5 — Interactions Between Regularization (Dropout, Weight Decay) and Continual Learning

    empirical result

    Applying standard neural network regularization during continual learning reveals contrasting interactions across method families:

    • Dropout (p=0.5p = 0.5): Substantially improves baseline finetuning, joint training, PackNet, and Synaptic Intelligence (SI). SI benefits by up to 10%\sim 10\% in average accuracy because dropout prevents overfitting during online trajectory calculation, stabilizing parameter importance estimation. However, for prior-focused methods (EWC, MAS), dropout increases forgetting over long task sequences because co-adaptation suppression reduces parameter redundancy, leaving fewer unimportant parameters free to acquire future tasks.
    • Weight Decay (λ=104\lambda = 10^{-4}): Enhances performance predominantly in WIDE network architectures. For prior-focused and parameter isolation methods on standard or small architectures, weight decay accelerates forgetting because the continuous weight decay penalty actively attenuates consolidated weights that encode prior task knowledge.
  6. Knowl 6 — Insignificance of Task Ordering in Continual Learning

    empirical result

    Testing curriculum-based hypotheses across diverse task orderings—such as easy-to-hard versus hard-to-easy orderings on Tiny ImageNet, and relatedness-based orderings (related, unrelated, and random) on iNaturalist—demonstrates that task presentation order has a negligible impact on overall final average accuracy and catastrophic forgetting.

    While isolated small tasks with few classes (e.g., the 5-class 'Fungi' super-category in iNaturalist) induce temporary performance dips in prior-focused regularization methods due to local overfitting on small sample sizes, the relative rankings and long-term convergence characteristics of all 11 evaluated continual learning methods remain invariant to task sequence ordering.

  7. Knowl 7 — Resource Complexity and Qualitative Trade-Offs in Continual Learning

    data/table

    The qualitative comparison below outlines GPU memory overhead (relative to the minimum column value), relative computation time, extra storage footprint, support for task-agnostic inference (single-head evaluation without task ID oracle), and sample privacy risks:

    Method Memory Compute Task-Agnostic Privacy Additional Storage
    Train Test Train Test Possible Issues
    iCaRL 1.24 1.00 5.63 45.61 Yes Yes M+RM + R
    GEM 1.07 1.29 10.66 3.64 Yes Yes TM+RT \cdot M + R
    LwF 1.07 1.10 1.29 1.86 Yes No MM
    EBLL 1.53 1.08 2.24 1.34 Yes No M+TAM + T \cdot A
    SI 1.09 1.05 1.13 1.61 Yes No 3M3 \cdot M
    EWC 1.09 1.05 1.11 1.88 Yes No 2M2 \cdot M
    MAS 1.09 1.05 1.16 1.88 Yes No 2M2 \cdot M
    mean-IMM 1.01 1.03 1.09 1.18 Yes No TMT \cdot M
    mode-IMM 1.01 1.03 1.24 1.00 Yes No 2TM2 \cdot T \cdot M
    PackNet 1.00 1.94 2.66 2.40 No No TM[bit]T \cdot M[\text{bit}]
    HAT 1.21 1.17 1.00 2.06 No No TUT \cdot U

    where TT is the number of tasks, MM is model parameter size, RR is exemplar buffer size, AA is task autoencoder size, M[bit]M[\text{bit}] is a binary mask per parameter, and UU is unit-based embedding size.

    Replay methods present data privacy concerns by retaining raw samples RR and incur high compute costs (e.g., nearest-mean classification in iCaRL requires 45.6×45.6\times test compute; GEM requires quadratic programming optimization). Parameter isolation methods require minimal storage (M[bit]M[\text{bit}]) but strictly prohibit task-agnostic inference due to requiring a task oracle to select masks.

  8. Knowl 8 — Task-Incremental Learning Objective and Forgetting Metric

    equation

    In task-incremental learning over a sequence of TT distinct tasks, each task t{1,,T}t \in \{1, \dots, T\} provides data (X(t),Y(t))D(t)(X^{(t)}, Y^{(t)}) \sim \mathcal{D}^{(t)}. Training minimizes the overall statistical risk across all seen tasks without access to prior task data:

    t=1TE(X(t),Y(t))D(t)[(ft(X(t);θ),Y(t))]\sum_{t=1}^T \mathbb{E}_{(X^{(t)}, Y^{(t)}) \sim \mathcal{D}^{(t)}} \left[ \ell(f_t(X^{(t)}; \theta), Y^{(t)}) \right]

    where \ell is the loss function, θ\theta denotes network parameters, and ftf_t is the network output function specific to task tt. For the current task TT with NTN_T samples, the empirical risk approximation is:

    1NTi=1NT(fT(xi(T);θ),yi(T))\frac{1}{N_T} \sum_{i=1}^{N_T} \ell(f_T(x_i^{(T)}; \theta), y_i^{(T)})

    Catastrophic forgetting on task tt after training up to task TT (t<Tt < T) is quantified as the difference between the initial accuracy attained on task tt immediately after its training phase (at,ta_{t,t}) and the accuracy on task tt evaluated after training task TT (aT,ta_{T,t}):

    Forgetting(t,T)=at,taT,t\text{Forgetting}(t, T) = a_{t,t} - a_{T,t}

    Average forgetting across all historical tasks upon reaching task TT is given by:

    FT=1T1t=1T1(at,taT,t)F_T = \frac{1}{T-1} \sum_{t=1}^{T-1} (a_{t,t} - a_{T,t})

    where FT<0F_T < 0 reflects positive backward transfer.

  9. Knowl 9 — Parameter Importance Formulations in Prior-Focused Regularization (EWC, SI, MAS)

    equation

    Prior-focused regularization methods penalize shifts to important parameters via a quadratic surrogate loss kΩkT(θkθk)2\sum_k \Omega_k^T (\theta_k - \theta_k^{*})^2. The parameter importance weights ΩkT\Omega_k^T for parameter index kk on task TT are calculated differently across methods:

    • Elastic Weight Consolidation (EWC) computes the diagonal empirical Fisher Information Matrix post-training on supervised data DT\mathcal{D}^T: ΩkT=E(x,y)DT[(L(f(x;θ),y)θkT)2]\Omega_k^T = \mathbb{E}_{(x, y) \sim \mathcal{D}^T} \left[ \left( \frac{\partial \mathcal{L}(f(x; \theta), y)}{\partial \theta_k^T} \right)^2 \right] where L\mathcal{L} is the task loss function.

    • Synaptic Intelligence (SI) accumulates importance online during training along the optimization trajectory via a path integral: ΩkT=t=1Tωkt(Δθkt)2+ξ\Omega_k^T = \sum_{t=1}^T \frac{\omega_k^t}{(\Delta \theta_k^t)^2 + \xi} where ωkt\omega_k^t is the running path integral of parameter kk over task tt, Δθkt=θktθkt1\Delta \theta_k^t = \theta_k^t - \theta_k^{t-1} is the total parameter displacement during task tt, and ξ>0\xi > 0 is a damping parameter preventing division by zero.

    • Memory Aware Synapses (MAS) computes parameter sensitivity in an unsupervised setting using the squared L2L_2-norm of the learned model output function fTf_T: ΩkT=ExDT[fT(x;θ)22θkT]\Omega_k^T = \mathbb{E}_{x \sim \mathcal{D}^T} \left[ \frac{\partial \|f_T(x; \theta)\|_2^2}{\partial \theta_k^T} \right] enabling importance estimation without ground-truth labels.

  10. Knowl 10 — Parameter Merging in Incremental Moment Matching

    equation

    Incremental Moment Matching (IMM) resolves catastrophic forgetting by merging task-specific models θ1,,θT\theta^1, \dots, \theta^T into a single consolidated model parameter vector θ1:T\theta^{1:T} via posterior distribution matching:

    • Mean-IMM computes a weighted average of individual task model weights: θk1:T=t=1Tαktθkt,subject tot=1Tαkt=1\theta_k^{1:T} = \sum_{t=1}^T \alpha_k^t \theta_k^t, \quad \text{subject to} \quad \sum_{t=1}^T \alpha_k^t = 1 where αkt\alpha_k^t is the mixing ratio for parameter kk on task tt.

    • Mode-IMM merges models by locating the mode of the Gaussian mixture approximation, weighting task parameters by their Fisher information importance weights Ωkt\Omega_k^t: θk1:T=1Ωk1:Tt=1TαktΩktθkt,whereΩk1:T=t=1TαktΩkt\theta_k^{1:T} = \frac{1}{\Omega_k^{1:T}} \sum_{t=1}^T \alpha_k^t \Omega_k^t \theta_k^t, \quad \text{where} \quad \Omega_k^{1:T} = \sum_{t=1}^T \alpha_k^t \Omega_k^t

    Mode-IMM consistently outperforms Mean-IMM because it accounts for the curvature and certainty of each parameter's local minimum.

Coverage note — Omitted detailed sub-analyses from the appendices regarding specific layer-wise pruning fractions in PackNet, ternary feature mask variants, and individual per-dataset hyperparameter grids.

References

  1. 1.D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel et al., “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science, vol. 362, no. 6419, pp. 1140–1144, 2018.
  2. 2.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” IJCV, vol. 115, no. 3, pp. 211–252, 2015.
  3. 3.R. M. French, “Catastrophic forgetting in connectionist networks,” Trends in cognitive sciences, vol. 3, no. 4, pp. 128–135, 1999.
  4. 4.Z. Chen and B. Liu, “Lifelong machine learning,” Synthesis Lectures on Artificial Intelligence and Machine Learning, vol. 12, no. 3, pp. 1–207, 2018.
  5. 5.R. Aljundi, P. Chakravarty, and T. Tuytelaars, “Expert gate: Lifelong learning with a network of experts,” in CVPR, 2017, pp. 3366–3375.
  6. 6.A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny, “Efficient lifelong learning with A-GEM,” in ICLR, 2018.
  7. 7.G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks, 2019.
  8. 8.D. L. Silver and R. E. Mercer, “The task rehearsal method of life-long learning: Overcoming impoverished data,” in Conference of the Canadian Society for Computational Studies of Intelligence. Springer, 2002, pp. 90–101.
  9. 9.A. Rannen, R. Aljundi, M. B. Blaschko, and T. Tuytelaars, “Encoder based lifelong learning,” in ICCV, 2017, pp. 1320–1328.
  10. 10.R. Aljundi, M. Rohrbach, and T. Tuytelaars, “Selfless sequential learning,” arXiv preprint arXiv:1806.05421, 2018.
  11. 11.M. McCloskey and N. J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,” in Psychology of learning and motivation. Elsevier, 1989, vol. 24, pp. 109–165.
  12. 12.H. Shin, J. K. Lee, J. Kim, and J. Kim, “Continual learning with deep generative replay,” in NeurIPS, 2017, pp. 2990–2999.
  13. 13.R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, and T. Tuytelaars, “Memory aware synapses: Learning what (not) to forget,” in ECCV, 2018, pp. 139–154.
  14. 14.A. Chaudhry, P. K. Dokania, T. Ajanthan, and P. H. Torr, “Riemannian walk for incremental learning: Understanding forgetting and intransigence,” in ECCV, 2018, pp. 532–547.
  15. 15.A. Gepperth and C. Karaoguz, “A bio-inspired incremental learning architecture for applied perceptual problems,” Cognitive Computation, vol. 8, no. 5, pp. 924–934, 2016.
  16. 16.S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert, “icarl: Incremental classifier and representation learning,” in CVPR, 2017, pp. 2001–2010.
  17. 17.A. Rosenfeld and J. K. Tsotsos, “Incremental learning through deep adaptation,” TPAMI, 2018.
  18. 18.K. Shmelkov, C. Schmid, and K. Alahari, “Incremental learning of object detectors without catastrophic forgetting,” in ICCV, 2017, pp. 3400–3409.
  19. 19.Y. Zhang and Q. Yang, “A survey on multi-task learning,” arXiv preprint arXiv:1707.08114, 2017.
  20. 20.S. Grossberg, Studies of mind and brain : neural principles of learning, perception, development, cognition, and motor control, ser. Boston studies in the philosophy of science 70. Dordrecht: Reidel, 1982.
  21. 21.M. D. Lange, X. Jia, S. Parisot, A. Leonardis, G. Slabaugh, and T. Tuytelaars, “Unsupervised model personalization while preserving privacy and scalability: An open problem,” in CVPR, 2020, pp. 14 463–14 472.
  22. 22.C. S. Lee and A. Y. Lee, “Clinical applications of continual learning machine learning,” The Lancet Digital Health, vol. 2, no. 6, pp. e279–e281, 2020.
  23. 23.P. Gupta, Y. Chaudhary, T. Runkler, and H. Schuetze, “Neural topic modeling with continual lifelong learning,” in ICML. PMLR, 2020, pp. 3907–3917.
  24. 24.T. Lesort, V. Lomonaco, A. Stoian, D. Maltoni, D. Filliat, and N. D´ıaz-Rodr´ıguez, “Continual learning for robotics,” arXiv preprint arXiv:1907.00182, 2019.
  25. 25.B. Pfulb and A. Gepperth, “A comprehensive, application-oriented study of catastrophic forgetting in dnns,” arXiv preprint arXiv:1905.08101, 2019.
  26. 26.S. Farquhar and Y. Gal, “Towards robust evaluations of continual learning,” arXiv preprint arXiv:1805.09733, 2018.
  27. 27.J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al., “Overcoming catastrophic forgetting in neural networks,” PNAS, p. 201611835, 2017.
  28. 28.S.-W. Lee, J.-H. Kim, J. Jun, J.-W. Ha, and B.-T. Zhang, “Overcoming catastrophic forgetting by incremental moment matching,” in NeurIPS, 2017, pp. 4652–4662.
  29. 29.R. Kemker, M. McClure, A. Abitino, T. L. Hayes, and C. Kanan, “Measuring catastrophic forgetting in neural networks,” in AAAI, 2018.
  30. 30.C. Fernando, D. Banarse, C. Blundell, Y. Zwols, D. Ha, A. A. Rusu, A. Pritzel, and D. Wierstra, “Pathnet: Evolution channels gradient descent in super neural networks,” arXiv preprint arXiv:1701.08734, 2017.
  31. 31.I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio, “An empirical investigation of catastrophic forgetting in gradient-based neural networks,” arXiv preprint arXiv:1312.6211, 2013.
  32. 32.Y.-C. Hsu, Y.-C. Liu, and Z. Kira, “Re-evaluating continual learning scenarios: A categorization and case for strong baselines,” arXiv preprint arXiv:1810.12488, 2018.
  33. 33.M. D. Lange and T. Tuytelaars, “Continual prototype evolution: Learning online from non-stationary data streams,” 2020.
  34. 34.R. M. French, “Semi-distributed representations and catastrophic forgetting in connectionist networks,” Connection Science, vol. 4, no. 3-4, pp. 365–377, 1992.
  35. 35.——, “Dynamically constraining connectionist networks to produce distributed, orthogonal representations to reduce catastrophic interference,” network, vol. 1111, p. 00001, 1994.
  36. 36.J. K. Kruschke, “Alcove: an exemplar-based connectionist model of category learning.” Psychological review, vol. 99, no. 1, p. 22, 1992.
  37. 37.——, “Human category learning: Implications for backpropagation models,” Connection Science, vol. 5, no. 1, pp. 3–36, 1993.
  38. 38.S. A. Sloman and D. E. Rumelhart, “Reducing interference in distributed memories through episodic gating,” Essays in honor of WK Estes, vol. 1, pp. 227–248, 1992.
  39. 39.A. Robins, “Catastrophic forgetting, rehearsal and pseudorehearsal,” Connection Science, vol. 7, no. 2, pp. 123–146, 1995.
  40. 40.R. M. French, “Pseudo-recurrent connectionist networks: An approach to the’sensitivity-stability’dilemma,” Connection Science, vol. 9, no. 4, pp. 353–380, 1997.
  41. 41.J. Rueckl, “Jumpnet: A multiple-memory connectionist architecture,” CogSci, no. 24, pp. 866–871, 1993.
  42. 42.B. Ans and S. Rousset, “Avoiding catastrophic forgetting by coupling two reverberating neural networks,” Comptes Rendus de l’Acad´emie des Sciences-Series III-Sciences de la Vie, vol. 320, no. 12, pp. 989–997, 1997.
  43. 43.N. Srivastava, G. Hinton, A. Krizdhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” JMLR, vol. 15, no. 1, pp. 1929–1958, 2014.
  44. 44.A. Pentina and C. H. Lampert, “A pac-bayesian bound for lifelong learning.” in ICML, 2014, pp. 991–999.
  45. 45.P. Alquier, T. T. Mai, and M. Pontil, “Regret bounds for lifelong learning,” arXiv preprint arXiv:1610.08628, 2016.
  46. 46.A. Pentina and C. H. Lampert, “Lifelong learning with non-iid tasks,” in NeurIPS, 2015, pp. 1540–1548.
  47. 47.M.-F. Balcan, A. Blum, and S. Vempala, “Efficient representations for lifelong learning and autoencoding,” in Conference on Learning Theory, 2015, pp. 191–210.
  48. 48.R. Aljundi, M. Lin, B. Goujaud, and Y. Bengio, “Online continual learning with no task boundaries,” arXiv preprint arXiv:1903.08671, 2019.
  49. 49.D. Rolnick, A. Ahuja, J. Schwarz, T. Lillicrap, and G. Wayne, “Experience replay for continual learning,” in Advances in Neural Information Processing Systems, 2019, pp. 350–360.
  50. 50.D. Isele and A. Cosgun, “Selective experience replay for lifelong learning,” in AAAI, 2018.
  51. 51.A. Chaudhry, M. Rohrbach, M. Elhoseiny, T. Ajanthan, P. K. Dokania, P. H. Torr, and M. Ranzato, “Continual learning with tiny episodic memories,” arXiv preprint arXiv:1902.10486, 2019.
  52. 52.C. Atkinson, B. McCane, L. Szymanski, and A. Robins, “Pseudorecursal: Solving the catastrophic forgetting problem in deep neural networks,” arXiv preprint arXiv:1802.03875, 2018.
  53. 53.F. Lavda, J. Ramapuram, M. Gregorova, and A. Kalousis, “Continual classification learning using generative models,” arXiv preprint arXiv:1810.10612, 2018.
  54. 54.J. Ramapuram, M. Gregorova, and A. Kalousis, “Lifelong generative modeling,” arXiv preprint arXiv:1705.09847, 2017.
  55. 55.D. Lopez-Paz and M. Ranzato, “Gradient episodic memory for continual learning,” in Advances in neural information processing systems, 2017, pp. 6467–6476.
  56. 56.F. Zenke, B. Poole, and S. Ganguli, “Continual learning through synaptic intelligence,” in ICML. JMLR. org, 2017, pp. 3987–3995.
  57. 57.X. Liu, M. Masana, L. Herranz, J. Van de Weijer, A. M. Lopez, and A. D. Bagdanov, “Rotate your networks: Better weight consolidation and less catastrophic forgetting,” in ICPR. IEEE, 2018, pp. 2262–2268.
  58. 58.Z. Li and D. Hoiem, “Learning without forgetting,” in ECCV. Springer, 2016, pp. 614–629.
  59. 59.H. Jung, J. Ju, M. Jung, and J. Kim, “Less-forgetting learning in deep neural networks,” arXiv preprint arXiv:1607.00122, 2016.
  60. 60.J. Zhang, J. Zhang, S. Ghosh, D. Li, S. Tasci, L. Heck, H. Zhang, and C.-C. J. Kuo, “Class-incremental learning via deep model consolidation,” in WACV, 2020, pp. 1131–1140.
  61. 61.A. Mallya and S. Lazebnik, “Packnet: Adding multiple tasks to a single network by iterative pruning,” in CVPR, 2018, pp. 7765–7773.
  62. 62.A. Mallya, D. Davis, and S. Lazebnik, “Piggyback: Adapting a single network to multiple tasks by learning to mask weights,” in ECCV, 2018, pp. 67–82.
  63. 63.J. Serr`a, D. Sur´ıs, M. Miron, and A. Karatzoglou, “Overcoming catastrophic forgetting with hard attention to the task,” arXiv preprint arXiv:1801.01423, 2018.
  64. 64.A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell, “Progressive neural networks,” arXiv preprint arXiv:1606.04671, 2016.
  65. 65.J. Xu and Z. Zhu, “Reinforced Continual Learning,” in NeurIPS, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 899–908. [Online]. Available: http://papers.nips.cc/paper/7369-reinforced-continual-learning.pdf
  66. 66.M. Masana, X. Liu, B. Twardowski, M. Menta, A. D. Bagdanov, and J. van de Weijer, “Class-incremental learning: survey and performance evaluation,” arXiv preprint arXiv:2010.15277, 2020.
  67. 67.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NeurIPS, 2014, pp. 2672–2680.
  68. 68.Y. Bengio and Y. LeCun, Eds., ICLR, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014. [Online]. Available: https://openreview.net/group?id=ICLR.cc/2014
  69. 69.C. V. Nguyen, Y. Li, T. D. Bui, and R. E. Turner, “Variational continual learning,” arXiv preprint arXiv:1710.10628, 2017.
  70. 70.H. Ahn, S. Cha, D. Lee, and T. Moon, “Uncertainty-based continual learning with adaptive regularization,” in NeurIPS, 2019, pp. 4394–4404.
  71. 71.C. Zeno, I. Golan, E. Hoffer, and D. Soudry, “Task agnostic continual learning using online variational bayes,” arXiv preprint arXiv:1803.10123, 2018.
  72. 72.R. Aljundi, K. Kelchtermans, and T. Tuytelaars, “Task-free continual learning,” in CVPR, 2019, pp. 11 254–11 263.
  73. 73.G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
  74. 74.D. J. MacKay, “A practical bayesian framework for backpropagation networks,” Neural computation, vol. 4, no. 3, pp. 448–472, 1992.
  75. 75.R. Pascanu and Y. Bengio, “Revisiting natural gradient for deep networks,” arXiv preprint arXiv:1301.3584, 2013.
  76. 76.J. Martens, “New insights and perspectives on the natural gradient method,” arXiv preprint arXiv:1412.1193, 2014.
  77. 77.F. Husz´ar, “Note on the quadratic penalties in elastic weight consolidation,” PNAS, p. 201717042, 2018.
  78. 78.J. Schwarz, J. Luketina, W. M. Czarnecki, A. Grabska-Barwinska, Y. W. Teh, R. Pascanu, and R. Hadsell, “Progress & compress: A scalable framework for continual learning,” arXiv preprint arXiv:1805.06370, 2018.
  79. 79.I. J. Goodfellow, O. Vinyals, and A. M. Saxe, “Qualitatively characterizing neural network optimization problems,” arXiv preprint arXiv:1412.6544, 2014.
  80. 80.G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, “Improving neural networks by preventing co-adaptation of feature detectors,” arXiv preprint arXiv:1207.0580, 2012.
  81. 81.M. Masana, T. Tuytelaars, and J. van de Weijer, “Ternary feature masks: continual learning without any forgetting,” arXiv preprint arXiv:2001.08714, 2020.
  82. 82.Stanford, “Tiny ImageNet Challenge, CS231N Course.” [Online]. Available: https://tiny-imagenet.herokuapp.com/
  83. 83.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR. Ieee, 2009, pp. 248–255.
  84. 84.“iNaturalist dataset: FGVC5 workshop at CVPR 2018.” [Online]. Available: https://www.kaggle.com/c/inaturalist-2018
  85. 85.M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” IJCV, vol. 111, no. 1, pp. 98–136, 2015.
  86. 86.M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in ICVGIP. IEEE, 2008, pp. 722–729.
  87. 87.A. Quattoni and A. Torralba, “Recognizing indoor scenes,” in CVPR. IEEE, 2009, pp. 413–420.
  88. 88.P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona, “Caltech-ucsd birds 200,” 2010.
  89. 89.J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in ICCV workshop, 2013, pp. 554–561.
  90. 90.S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi, “Fine-grained visual classification of aircraft,” arXiv preprint arXiv:1306.5151, 2013.
  91. 91.T. E. De Campos, B. R. Babu, M. Varma et al., “Character recognition in natural images.” VISAPP (2), vol. 7, 2009.
  92. 92.Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” 2011.
  93. 93.C. Chung, S. Patel, R. Lee, L. Fu, S. Reilly, T. Ho, J. Lionetti, M. D. George, and P. Taylor, “Implementation of an integrated computerized prescriber order-entry system for chemotherapy in a multisite safety-net health system,” AJHP, 2018.
  94. 94.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NeurIPS, 2012, pp. 1097–1105.
  95. 95.Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in ICML. ACM, 2009, pp. 41–48.
  96. 96.C. V. Nguyen, A. Achille, M. Lam, T. Hassner, V. Mahadevan, and S. Soatto, “Toward understanding catastrophic forgetting in continual learning,” arXiv preprint arXiv:1908.01091, 2019.
  97. 97.S. J. Pan and Q. Yang, “A survey on transfer learning,” TKDE, vol. 22, no. 10, pp. 1345–1359, 2009.
  98. 98.G. Csurka, Domain adaptation in computer vision applications. Springer, 2017.
  99. 99.S. Shalev-Shwartz et al., “Online learning and online convex optimization,” Foundations and Trends® in Machine Learning, vol. 4, no. 2, pp. 107–194, 2012.
  100. 100.L. Bottou, “Online learning and stochastic approximations,” Online learning in neural networks, vol. 17, no. 9, p. 142, 1998.
  101. 101.H. Xu, B. Liu, L. Shu, and P. S. Yu, “Learning to accept new classes without training,” arXiv preprint arXiv:1809.06004, 2018.
  102. 102.A. Bendale and T. Boult, “Towards open world recognition,” in CVPR, 2015, pp. 1893–1902.

Citation

MLA
Delange, M., et al. “A Continual Learning Survey: Defying Forgetting in Classification Tasks”. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, pp. 1–1, https://doi.org/10.1109/TPAMI.2021.3057446.
APA
Delange, M., Aljundi, R., Masana, M., Parisot, S., Jia, X., Leonardis, A., Slabaugh, G., & Tuytelaars, T. (2021). A continual learning survey: Defying forgetting in classification tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1–1. https://doi.org/10.1109/TPAMI.2021.3057446
Chicago
Delange, M., R. Aljundi, M. Masana, et al. 2021. “A Continual Learning Survey: Defying Forgetting in Classification Tasks”. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1–1. https://doi.org/10.1109/TPAMI.2021.3057446.
Harvard
Delange, M. et al. (2021) “A continual learning survey: Defying forgetting in classification tasks”, IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1. Available at: https://doi.org/10.1109/TPAMI.2021.3057446.
Vancouver
1. Delange M, Aljundi R, Masana M, Parisot S, Jia X, Leonardis A, Slabaugh G, Tuytelaars T (2021) A continual learning survey: Defying forgetting in classification tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence 1–1

BibTeX

@article{Delange_2021, title={A continual learning survey: Defying forgetting in classification tasks}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2021.3057446}, DOI={10.1109/tpami.2021.3057446}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Delange, Matthias and Aljundi, Rahaf and Masana, Marc and Parisot, Sarah and Jia, Xu and Leonardis, Ales and Slabaugh, Greg and Tuytelaars, Tinne}, year={2021}, pages={1–1} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF