Overcoming catastrophic forgetting with hard attention to the task

Joan SerràDídac SurísMarius MironAlexandros Karatzoglou

article2018ICML1,326 citations

Proposes a task-based hard attention mechanism that learns parameter masks to protect prior knowledge during continual training, cutting catastrophic forgetting rates by 45 to 80%.

Listen

Modern artificial intelligence systems struggle when learning multiple tasks in sequence because neural networks tend to overwrite previously learned information upon acquiring new skills. This issue, known as catastrophic forgetting, presents a major obstacle to developing lifelong learning systems such as autonomous robotics or continually adapting software. Retraining models from scratch or constantly reprocessing historical data becomes computationally prohibitive and inefficient at scale.

The article demonstrates that a task-based gating technique called Hard Attention to the Task (HAT) effectively preserves previously learned knowledge without hindering the learning of new tasks. By learning unit-level attention masks concurrently with standard training, HAT shields critical network pathways from future updates.

The authors evaluated the approach across rigorous image classification benchmarks, comparing it against eleven baseline and state-of-the-art methods. The primary evaluation involved ten runs across randomized sequences of eight diverse image datasets using a standard 7.1-million-parameter network architecture. Supplementary tests evaluated incremental class learning and standard digit classification tasks, measuring performance via a normalized forgetting ratio.

The key findings reveal that HAT consistently outperforms existing continual learning approaches. First, HAT reduced forgetting rates by 45% to 80% across all sequential tasks, achieving a final eight-task forgetting ratio of -0.06 compared to -0.11 for the strongest baseline. Second, in incremental class learning across CIFAR benchmarks, HAT reduced the forgetting ratio by 55% compared to the best alternative. Third, on split and permuted benchmark tasks, HAT reduced error rates by 52% to 80%, achieving top accuracies of 98.6% and 99.0%. Finally, the mechanism enables significant model compression during training, reducing network size by 79% to 99% depending on the task.

These results imply that organizations can deploy continually updating AI models with minimal computational overhead and low operational risk. Because HAT relies on only two intuitive hyperparameters controlling stability and compactness, it avoids the hyperparameter sensitivity and performance instability typical of competing methods. Furthermore, the ability to monitor capacity usage and weight reuse across tasks provides actionable visibility into model lifecycle management.

Organizations developing adaptive AI or edge deployments should consider adopting hard attention mechanisms to retain cumulative knowledge and compress model footprints. Future development should evaluate HAT across non-vision modalities, such as natural language processing and reinforcement learning, and explore settings where task identifiers are not explicitly provided.

The study's primary limitation is its assumption that task identities are known during both training and evaluation, alongside its primary focus on image classification benchmarks. Nevertheless, the low variance across multiple random sequences and strong outperformance across standard setups provide high confidence in the method's effectiveness for structured sequential learning tasks.

  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Introduces Elastic Weight Consolidation (EWC) to protect critical parameters via quadratic penalties, providing the primary baseline and foundational regularized continual learning paradigm that the source contrasts with hard task-masking.
  • Paper: PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning, Arun Mallya et al. (2017). Introduces parameter isolation and network pruning for sequential task learning, serving as key conceptual prior work for using dedicated subnetworks to eliminate catastrophic forgetting.
  • Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). Establishes architectural parameter-allocation mechanisms for sequential task learning, establishing the foundation for modular parameter preservation in continual learning.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Demonstrates path-integral parameter importance regularization (Synaptic Intelligence), representing a key prior continual learning baseline for preventing catastrophic forgetting.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Presents Learning without Forgetting (LwF), establishing the classic task-incremental learning baseline using distillation to mitigate forgetting without historical raw data.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Formulates standard benchmark settings, evaluation metrics, and optimization constraints for continual learning against which task-based attention approaches are evaluated.
  • Paper: Lifelong Learning with Dynamically Expandable Networks, Jaehong Yoon et al. (2017). Proposes dynamically expanding subnetworks for sequential learning, motivating task-specific capacity allocation strategies.
  • Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). Introduces output-sensitivity parameter importance weighting to mitigate forgetting under fixed network capacity, complementing regularized continual learning formulations.
Cover for Overcoming catastrophic forgetting with hard attention to the task

Abstract

Catastrophic forgetting occurs when a neural network loses the information learned in a previous task after training on subsequent tasks. This problem remains a hurdle for artificial intelligence systems with sequential learning capabilities. In this paper, we propose a task-based hard attention mechanism that preserves previous tasks' information without affecting the current task's learning. A hard attention mask is learned concurrently to every task, through stochastic gradient descent, and previous masks are exploited to condition such learning. We show that the proposed mechanism is effective for reducing catastrophic forgetting, cutting current rates by 45 to 80%. We also show that it is robust to different hyperparameter choices, and that it offers a number of monitoring capabilities. The approach features the possibility to control both the stability and compactness of the learned knowledge, which we believe makes it also attractive for online learning or network compression applications.

Table of Contents

  • 1 Introduction
  • 2 Putting Hard Attention to the Task
  • 2.1 Motivation
  • 2.2 Architecture
  • 2.3 Network Training
  • 2.4 Hard Attention Training
  • 2.5 Embedding Gradient Compensation
  • 2.6 Promoting Low Capacity Usage
  • 3 Related Work
  • 4 Experiments
  • 4.1 Results
  • 4.2 Additional Results
  • 4.3 Hyperparameters
  • 4.4 Monitoring and Network Pruning
  • 5 Conclusion
  • References
  • A Data
  • B Raw Results
  • B.1 Task Mixture
  • B.2 Layer Use
  • B.3 Network Compression
  • B.4 Training Time
  • C Additional Results
  • C.1 Incremental CIFAR
  • C.2 Permuted MNIST
  • C.3 Split MNIST
  • D Variations to the Proposed Approach
  • D.1 Embedding Learning
  • D.2 Annealing
  • D.3 Gate
  • D.4 Cumulative Attention
  • D.5 Embedding Initialization
  • D.6 Attention Regularization
  • D.7 Hard Attention to the Input
  • E A Note on Binary Masks

Knowls

  1. Knowl 1 — Hard Attention to the Task Architecture and Unit Gating

    model/method

    In Hard Attention to the Task (HAT), task-specific conditioning is applied layer-wise across a feedforward neural network with LL layers. For a given task tt and layer l∈{1,…,L−1}l \in \{1, \dots, L-1\}, a learned single-layer task embedding vector elte_l^t is converted to a unit-level hard attention vector alta_l^t via a scaled sigmoid gate:

    alt=σ(selt),a_l^t = \sigma\left(s e_l^t\right),

    where σ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}} is the element-wise sigmoid function and s>0s > 0 is a positive scaling parameter controlling the sharpness of the gate. For the final classification layer LL, the attention vector aLta_L^t is a hard-coded binary vector that selects the output heads corresponding to task tt.

    Given the unconditioned activation vector hlh_l of layer ll (representing fully-connected linear units or convolutional filters), the conditioned activation vector hl′h_l' is computed via element-wise multiplication:

    hl′=alt⊙hl.h_l' = a_l^t \odot h_l.

    As s→∞s \to \infty, the attention values polarize to al,it∈{0,1}a_{l,i}^t \in \{0, 1\}, acting as hard inhibitory gates that dynamically activate or deactivate individual network units for task tt during standard forward propagation.

  2. Knowl 2 — Backward Gradient Conditioning via Cumulative Attention

    model/method

    To prevent interference with parameters used by previously learned tasks when training on a new task t+1t+1, HAT tracks the maximum attention allocated to each unit across all past tasks. For each layer ll, the cumulative attention vector al≤ta_l^{\le t} after learning task tt is recursively defined by:

    al≤t=max⁡(alt,al≤t−1),a_l^{\le t} = \max\left(a_l^t, a_l^{\le t-1}\right),

    initialized with al≤0=0a_l^{\le 0} = \mathbf{0}.

    When optimizing the network on task t+1t+1, the parameter gradient gl,ij=∂L∂Wl,ijg_{l,ij} = \frac{\partial \mathcal{L}}{\partial W_{l,ij}} for the weight connecting input unit jj from layer l−1l-1 to output unit ii in layer ll is modified to gl,ij′g'_{l,ij} according to the cumulative attention of both connected units:

    gl,ij′=[1−min⁡(al,i≤t,al−1,j≤t)]gl,ij.g'_{l,ij} = \left[ 1 - \min\left(a_{l,i}^{\le t}, a_{l-1,j}^{\le t}\right) \right] g_{l,ij}.

    When both connected units were heavily attended in previous tasks (al,i≤t→1a_{l,i}^{\le t} \to 1 and al−1,j≤t→1a_{l-1,j}^{\le t} \to 1), the gradient update is blocked (gl,ij′→0g'_{l,ij} \to 0), leaving the weight frozen. If either unit was inactive in all past tasks (a≈0a \approx 0), the weight remains fully plastic (gl,ij′≈gl,ijg'_{l,ij} \approx g_{l,ij}) to acquire representations for the new task. No attention mask is applied to raw image or audio input signals.

  3. Knowl 3 — Gating Temperature Annealing Scheme

    model/method

    To permit gradient flow through the task embeddings during training while achieving near-binary attention masks at test time, HAT anneals the scaling parameter ss linearly across the mini-batches of each training epoch:

    s=1smax⁡+(smax⁡−1smax⁡)b−1B−1,s = \frac{1}{s_{\max}} + \left( s_{\max} - \frac{1}{s_{\max}} \right) \frac{b-1}{B-1},

    where b∈{1,…,B}b \in \{1, \dots, B\} is the current mini-batch index, BB is the total number of mini-batches in the epoch, and smax⁡≥1s_{\max} \ge 1 is a hyperparameter setting the maximum gate sharpness.

    At the start of each epoch (b=1b=1), s=1smax⁡s = \frac{1}{s_{\max}}, which pushes the gate values al,it=σ(sel,it)≈0.5a_{l,i}^t = \sigma(s e_{l,i}^t) \approx 0.5, allowing gradients to distribute broadly across all units. As training progresses through the epoch, ss increases up to smax⁡≫1s_{\max} \gg 1, polarizing al,ita_{l,i}^t toward {0,1}\{0, 1\}. During inference and testing, ss is fixed to smax⁡s_{\max}.

  4. Knowl 4 — Embedding Gradient Compensation

    equation

    Because the annealing of the scaling factor ss dampens and narrows the gradients backpropagated to the task embeddings elte_{l}^t, HAT multiplies the raw embedding gradient ql,i=∂L∂el,itq_{l,i} = \frac{\partial \mathcal{L}}{\partial e_{l,i}^t} by an analytic compensation factor prior to updating el,ite_{l,i}^t:

    ql,i′=smax⁡[cosh⁡(sel,it)+1]s[cosh⁡(el,it)+1]ql,i.q'_{l,i} = \frac{s_{\max} \left[ \cosh\left(s e_{l,i}^t\right) + 1 \right]}{s \left[ \cosh\left(e_{l,i}^t\right) + 1 \right]} q_{l,i}.

    To preserve numerical stability, the scaled embedding values are clamped such that ∣sel,it∣≤50|s e_{l,i}^t| \le 50, and the task embedding parameters are bounded within the active sigmoid range el,it∈[−6,6]e_{l,i}^t \in [-6, 6]. This compensation restores the cumulative gradient magnitude achieved at s=smax⁡s = s_{\max} across the broader functional range of s=1s = 1.

  5. Knowl 5 — Attention-Weighted Capacity Regularization

    model/method

    To preserve unused network capacity for subsequent tasks without penalizing the reuse of units already engaged by earlier tasks, HAT trains task tt with an attention-weighted, normalized L1L_1 penalty on the current attention vectors At={a1t,…,aL−1t}A^t = \{a_1^t, \dots, a_{L-1}^t\}:

    L′(y,y^,At,A<t)=L(y,y^)+c R(At,A<t),\mathcal{L}'\left(y, \hat{y}, A^t, A^{<t}\right) = \mathcal{L}(y, \hat{y}) + c \, R\left(A^t, A^{<t}\right),

    where c≥0c \ge 0 is a compressibility hyperparameter, A<t={a1≤t−1,…,aL−1≤t−1}A^{<t} = \{a_1^{\le t-1}, \dots, a_{L-1}^{\le t-1}\} represents the set of cumulative attention vectors up to task t−1t-1, and the regularization function R(At,A<t)R(A^t, A^{<t}) is defined as:

    R(At,A<t)=∑l=1L−1∑i=1Nlal,it(1−al,i≤t−1)∑l=1L−1∑i=1Nl(1−al,i≤t−1),R\left(A^t, A^{<t}\right) = \frac{\sum_{l=1}^{L-1} \sum_{i=1}^{N_l} a_{l,i}^t \left(1 - a_{l,i}^{\le t-1}\right)}{\sum_{l=1}^{L-1} \sum_{i=1}^{N_l} \left(1 - a_{l,i}^{\le t-1}\right)},

    where NlN_l is the total number of units in layer ll.

    The weighting factor (1−al,i≤t−1)(1 - a_{l,i}^{\le t-1}) eliminates the regularization penalty on units for which al,i≤t−1→1a_{l,i}^{\le t-1} \to 1, enabling unpenalized reuse of previously allocated features, while applying full regularization pressure to previously unallocated units where al,i≤t−1→0a_{l,i}^{\le t-1} \to 0.

  6. Knowl 6 — Forgetting Ratio Metric

    definition

    The forgetting ratio ρτ≤t\rho_{\tau \le t} quantifies the retention of performance on task τ\tau after the sequential training of tasks 1,…,t1, \dots, t (with τ≤t\tau \le t):

    ρτ≤t=Aτ≤t−AτRAτ≤tJ−AτR−1,\rho_{\tau \le t} = \frac{A_{\tau \le t} - A_{\tau}^R}{A_{\tau \le t}^J - A_{\tau}^R} - 1,

    where Aτ≤tA_{\tau \le t} is the test classification accuracy on task τ\tau measured after completing training on task tt, AτRA_\tau^R is the accuracy of a random stratified classifier on task τ\tau, and Aτ≤tJA_{\tau \le t}^J is the test accuracy on task τ\tau obtained by a model trained jointly on all tt tasks in a multitask configuration.

    A value of ρτ≤t≈0\rho_{\tau \le t} \approx 0 indicates performance equivalent to joint multitask training (zero catastrophic forgetting), whereas ρτ≤t≈−1\rho_{\tau \le t} \approx -1 signifies complete degradation to random guessing. The overall retention across all tt learned tasks is evaluated via the average forgetting ratio:

    ρ≤t=1t∑τ=1tρτ≤t.\rho^{\le t} = \frac{1}{t} \sum_{\tau=1}^t \rho_{\tau \le t}.

  7. Knowl 7 — Sequential Image Classification Benchmark Setup

    experimental setup

    The continual learning experimental setup consists of random sequences of 8 image classification datasets: CIFAR-10, CIFAR-100, FaceScrub, FashionMNIST, NotMNIST, MNIST, SVHN, and TrafficSigns. All inputs are resized to 32×32×332 \times 32 \times 3 pixels. Dataset class numbers range from 10 to 100, training set sizes from 16,853 to 73,257 (with 15% reserved for validation), and test set sizes from 1,873 to 26,032.

    The base neural architecture is an AlexNet-style network with approximately 7.1 million parameters across all methods:

    • 3 convolutional layers: 64 filters (4×44 \times 4), 128 filters (3×33 \times 3), and 256 filters (2×22 \times 2), with ReLU activations and 2×22 \times 2 max-pooling.
    • 2 fully connected layers of 2048 units each, followed by a task-specific softmax output head.
    • Dropout: 0.2 for the first two conv layers, 0.5 for subsequent layers.
    • Initializations: Xavier uniform for network weights, N(0,1)\mathcal{N}(0, 1) for task embeddings.
    • Optimization: SGD with batch size 64, initial learning rate 0.05, decayed by a factor of 3 upon validation plateau (patience 5 epochs), terminating below 10−410^{-4} or at 200 epochs. Default HAT hyperparameters are smax⁡=400s_{\max} = 400 and c=0.75c = 0.75. Results are averaged over 10 randomized task sequence orderings.
  8. Knowl 8 — Comparative Forgetting Performance on 8-Dataset Sequence

    data/table

    The table compares the average forgetting ratio after the second task (ρ≤2\rho^{\le 2}) and after all eight tasks (ρ≤8\rho^{\le 8}) on the 8-dataset sequential image classification benchmark across 10 random runs (mean with standard deviation in parentheses):

    Approach ρ≤2\rho^{\le 2} ρ≤8\rho^{\le 8}
    LFL -0.73 (0.29) -0.92 (0.08)
    LWF -0.14 (0.13) -0.80 (0.06)
    SGD -0.20 (0.08) -0.66 (0.03)
    IMM-MODE -0.11 (0.08) -0.49 (0.05)
    SGD-F -0.20 (0.15) -0.44 (0.06)
    IMM-MEAN -0.12 (0.10) -0.42 (0.04)
    EWC -0.08 (0.06) -0.25 (0.03)
    PATHNET -0.09 (0.16) -0.17 (0.23)
    PNN -0.11 (0.10) -0.11 (0.01)
    HAT -0.02 (0.03) -0.06 (0.01)

    HAT achieves the lowest forgetting ratio across both horizons, cutting the forgetting of the best baseline by 75% at t=2t=2 (relative to EWC) and by 45% at t=8t=8 (relative to PNN), while exhibiting minimal variance across different task sequences.

  9. Knowl 9 — Performance on Incremental CIFAR, Permuted MNIST, and Split MNIST

    empirical result

    HAT was evaluated on three standard continual learning configurations:

    1. Incremental Class Learning (CIFAR-10 / CIFAR-100): In a 10-task sequential class-addition scenario, the strongest baseline was Elastic Weight Consolidation (EWC) with ρ≤10=−0.18\rho^{\le 10} = -0.18. HAT reached ρ≤10=−0.09\rho^{\le 10} = -0.09, yielding a 55% reduction in forgetting.
    2. Permuted MNIST (10 tasks): Across 10 sequential permuted MNIST tasks, the prior literature top result was Synaptic Intelligence (SI) with an average accuracy of A≤10=97.1%A^{\le 10} = 97.1\%. HAT attained A≤10=98.6%A^{\le 10} = 98.6\%, representing a 52% error rate reduction.
    3. Split MNIST (2 tasks): On the disjoint MNIST split setup, the prior best published result was Conceptor-Aided Backpropagation at A≤2=94.9%A^{\le 2} = 94.9\%. HAT achieved A≤2=99.0%A^{\le 2} = 99.0\%, corresponding to an 80% error rate reduction.
  10. Knowl 10 — Single-Task Network Compression and Feature Reuse Monitoring

    model/method

    The hard attention mechanism provides internal monitoring of unit usage and allows single-task network compression during SGD training. By initializing the task embedding with positive values el∼U(0,2)e_l \sim \mathcal{U}(0, 2) (allocating full capacity at epoch zero) and setting a high compressibility parameter c=1.5c = 1.5, the learned attention vectors alta_l^t prune unneeded units end-to-end.

    Pruning all units with al,it=0a_{l,i}^t = 0 compresses the AlexNet-like network to between 1%1\% (on MNIST) and 21%21\% (on CIFAR-100) of its original parameter count with negligible drop in validation accuracy compared to uncompressed vanilla SGD. Furthermore, computing pairwise inner products between cumulative task masks a≤tia^{\le t_i} and a≤tja^{\le t_j} allows direct monitoring of the percentage of weights from task tit_i reused in task tjt_j.

Coverage note — None was omitted; the extracted knowls fully cover the HAT architecture, gradient modification rule, annealing schedule, gradient compensation formula, capacity regularization, evaluation metric, experimental configuration, quantitative results across all benchmarks, and network pruning/monitoring capabilities.

References

  1. 1.Bakker, B. and Heskes, T. Task clustering and gating for bayesian multitask learning. Journal of Machine Learning Research, 4:83–99, 2003.
  2. 2.Bulatov, Y. NotMNIST dataset. Technical report, 2011. URL http://yaroslavvb.blogspot.it/2011/09/notmnist-dataset.html.
  3. 3.Clegg, B. A., DiGirolamo, G. J., and Keele, S. W. Sequence learning. Trends in Cognitive Sciences, 2(8):275–281, 1998.
  4. 4.Fernando, C., Banarse, D., Blundell, C., Zwols, Y., Ha, D., Rusu, A. A., Pritzel, A., and Wierstra, D. PathNet: evolution channels gradient descent in super neural networks. ArXiv, 1701.08734, 2017.
  5. 5.French, R. M. Using semi-distributed representations to overcome catastrophic forgetting in connectionist networks. In Proc. of the Annual Conf. of the Cognitive Science Society (CogSci), pp. 173–178, 1991.
  6. 6.Glorot, X. and Bengio, Y. Understanding the difficulty of training deep feedforward neural networks. In Proc. of the Int. Conf. on Artificial Intelligence and Statistics (AISTATS), pp. 249–256, 2010.
  7. 7.Goodfellow, I., Mizra, M., Da, X., Courville, A., and Bengio, Y. An empirical investigation of catastrophic forgetting in gradient-based neural networks. In Proc. of the Int. Conf. on Learning Representations (ICLR), 2014.
  8. 8.Gutsein, S. and Stump, E. Reduction of catastrophic forgetting with transfer learning and ternary output codes. In Proc. of the Int. Joint Conf. on Neural Networks (IJCNN), pp. 1–8, 2015.
  9. 9.Han, S., Mao, H., and Dally, W. J. Deep compression: compressing deep neural networks with pruning, trained quantization and Huffman coding. In Proc. of the Int. Conf. on Learning Representations (ICLR), 2016.
  10. 10.He, X. and Jaeger, H. Overcoming catastrophic interference using conceptor-aided backpropagation. In Proc. of the Int. Conf. on Learning Representations (ICLR), 2018.
  11. 11.Jung, H., Ju, J., Jung, M., and Kim, J. Less-forgetting learning in deep neural networks. ArXiv, 1607.00122, 2016.
  12. 12.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R. Overcoming catastrophic forgetting in neural networks. Proc. of the National Academy of Sciences of the USA, 114(13):3521–3526, 2017.
  13. 13.Krizhevsky, A. Learning multiple layers of features from tiny images. Msc thesis, University of Toronto, Toronto, Canada, 2009.
  14. 14.Krizhevsky, A., Sutskever, I., and Hinton, G. ImageNet classification with deep convolutional neural networks. In Pereira, F., Burges, C. J. C., Bottou, L., and Weinberger, K. Q. (eds.), Advances in Neural Information Processing Systems (NIPS), volume 25, pp. 1097–1105. 2012.
  15. 15.LeCun, Y., Denker, J. S., and Solla, S. A. Optimal brain damage. In Touretzky, D. S. (ed.), Advances in Neural Information Processing Systems (NIPS), volume 2, pp. 598–605. Morgan Kaufmann, 1990.
  16. 16.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradientbased learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  17. 17.Lee, S.-W., Kim, J.-H., Jun, J., Ha, J.-W., and Zhang, B.-T. Overcoming catastrophic forgetting by incremental moment matching. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems (NIPS), volume 30, pp. 4655–4665. Curran Associates Inc., 2017.
  18. 18.Legg, S. and Hutter, M. Universal intelligence: a definition of machine intelligence. Minds and Machines, 17(4): 391–444, 2007.
  19. 19.Li, Z. and Hoiem, D. Learning without forgetting. IEEE Trans. on Pattern Analysis and Machine Intelligence, PP (99):1–1, 2017.
  20. 20.Lopez-Paz, D. and Ranzato, M. A. Gradient episodic memory for continuum learning. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems (NIPS), volume 30, pp. 6449–6458. Curran Associates Inc., 2017.
  21. 21.Mallya, A. and Lazebnik, S. PackNet: adding multiple tasks to a single network by iterative pruning. ArXiv, 1711.05769, 2017.
  22. 22.McCloskey, M. and Cohen, N. Catastrophic interference in connectionist networks: the sequential learning problem. Psychology of Learning and Motivation, 24:109– 165, 1989.
  23. 23.McCulloch, W. S. and Pitts, W. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5(4):115–133, 1943.
  24. 24.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning (NIPSDeepLearning), 2011.
  25. 25.Ng, H.-W. and Winkler, S. A data-driven approach to cleaning large face datasets. In Proc. of the IEEE Int. Conf. on Image Processing (ICIP), pp. 343–347, 2014.
  26. 26.Nguyen, C., Li, Y., Bui, T. D., and Turner, R. E. Variational continual learning. ArXiv, 1710.10628, 2017.
  27. 27.Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. Automatic differentiation in PyTorch. In NIPS Workshop on The Future of Gradient-based Machine Learning Software & Techniques (NIPS-Autodiff), 2017.
  28. 28.Ratcliff, R. Connectionist models of recognition memory: constraints imposed by learning and forgetting functions. Psychological Review, 97:285–308, 1990.
  29. 29.Rebuffi, S., Kolesnikov, A., Sperl, G., and Lampert, C. iCaRL: incremental classifier and representation learning. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 2001–2010, 2017.
  30. 30.Robins, A. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science, 7:123–146, 1995.
  31. 31.Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. Progressive neural networks. ArXiv, 1606.04671, 2016.
  32. 32.Shin, H., Lee, J. K., Kim, J., and Kim, J. Continual learning with deep generative replay. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems (NIPS), volume 30, pp. 2993–3002. Curran Associates Inc., 2017.
  33. 33.Srivastava, R. K., Masci, J., Kazerounian, S., Gomez, F., and Schmidhuber, J. Compete to compute. In Burges, C. J. C., Bottou, L., Welling, M., Ghahramani, Z., and Weinberger, K. (eds.), Advances in Neural Information Processing Systems (NIPS), volume 26, pp. 2310–2318. Curran Associates Inc., 2013.
  34. 34.Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C. The German traffic sign recognition benchmark: a multi-class classification competition. In Proc. of the Int. Joint Conf. on Neural Networks (IJCNN), pp. 1453–1460, 2011.
  35. 35.Thrun, S. and Mitchell, T. Lifelong robot learning. Robotics and Autonomous Systems, 15:25–46, 1995.
  36. 36.Venkatesan, R., Venkateswara, H., Panchanathan, S., and Li, B. A strategy for an uncompromising incremental learner. ArXiv, 1705.00744, 2017.
  37. 37.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. ArXiv, 1708.07747, 2017.
  38. 38.Yoon, J., Yang, E., Lee, J., and Hwang, S. J. Lifelong learning with dynamically expandable networks. In Proc. of the Int. Conf. on Learning Representations (ICLR), 2018.
  39. 39.Zenke, F., Poole, B., and Ganguli, S. Improved multitask learning through synaptic intelligence. In Proc. of the Int. Conf. on Machine Learning (ICML), pp. 3987–3995, 2017.

Citation

MLA
Serra, J., et al. “Overcoming Catastrophic Forgetting with Hard Attention to the Task”. International Conference on Machine Learning, vol. 80, 2018, pp. 4548–57, https://proceedings.mlr.press/v80/serra18a.html.
APA
Serra, J., Suris, D., Miron, M., & Karatzoglou, A. (2018). Overcoming Catastrophic Forgetting with Hard Attention to the Task. International Conference on Machine Learning, 80, 4548–4557. https://proceedings.mlr.press/v80/serra18a.html
Chicago
Serra, J., D. Suris, M. Miron, and A. Karatzoglou. 2018. “Overcoming Catastrophic Forgetting with Hard Attention to the Task”. International Conference on Machine Learning 80: 4548–57. https://proceedings.mlr.press/v80/serra18a.html.
Harvard
Serra, J. et al. (2018) “Overcoming Catastrophic Forgetting with Hard Attention to the Task”, International Conference on Machine Learning. PMLR, pp. 4548–4557. Available at: https://proceedings.mlr.press/v80/serra18a.html.
Vancouver
1. Serra J, Suris D, Miron M, Karatzoglou A (2018) Overcoming Catastrophic Forgetting with Hard Attention to the Task. In: International Conference on Machine Learning. PMLR, pp 4548–4557

BibTeX

@InProceedings{pmlr-v80-serra18a,
  title = 	 {Overcoming Catastrophic Forgetting with Hard Attention to the Task},
  author =       {Serra, Joan and Suris, Didac and Miron, Marius and Karatzoglou, Alexandros},
  booktitle = 	 {Proceedings of the 35th International Conference on Machine Learning},
  pages = 	 {4548--4557},
  year = 	 {2018},
  editor = 	 {Dy, Jennifer and Krause, Andreas},
  volume = 	 {80},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {10--15 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v80/serra18a/serra18a.pdf},
  url = 	 {https://proceedings.mlr.press/v80/serra18a.html},
  abstract = 	 {Catastrophic forgetting occurs when a neural network loses the information learned in a previous task after training on subsequent tasks. This problem remains a hurdle for artificial intelligence systems with sequential learning capabilities. In this paper, we propose a task-based hard attention mechanism that preserves previous tasks’ information without affecting the current task’s learning. A hard attention mask is learned concurrently to every task, through stochastic gradient descent, and previous masks are exploited to condition such learning. We show that the proposed mechanism is effective for reducing catastrophic forgetting, cutting current rates by 45 to 80%. We also show that it is robust to different hyperparameter choices, and that it offers a number of monitoring capabilities. The approach features the possibility to control both the stability and compactness of the learned knowledge, which we believe makes it also attractive for online learning or network compression applications.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/