Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence

Arslan ChaudhryPuneet K. DokaniaThalaiyasingam AjanthanPhilip H. S. Torr

article2018ECCV1,438 citations

Proposes Riemannian Walk (RWalk), a continual learning method based on KL-divergence that optimizes the trade-off between catastrophic forgetting and model intransigence, while establishing formal metrics to quantify both behaviors.

Listen

Modern artificial intelligence systems struggle when required to learn new tasks sequentially over time without forgetting previously acquired knowledge. While standard machine learning models train on a complete, fixed dataset, real-world deployment requires continuous learning from sequential streams of data. Existing continuous learning methods often focus solely on mitigating catastrophic forgetting (losing past knowledge), but they frequently suffer from intransigence, which is the model's inability to adapt and learn new information effectively. This trade-off between remembering past tasks and acquiring new ones poses a major bottleneck for practical, long-term deployment.

The article establishes a formal framework to quantify both forgetting and intransigence in incremental learning settings and introduces RWalk, an efficient algorithm designed to balance these competing demands. The approach evaluates sequential image classification on standard visual benchmarks, including split MNIST and split CIFAR-100 datasets, under realistic evaluation setups where task identities are not provided at test time.

The researchers developed two formal metrics to independently measure forgetting and intransigence alongside standard classification accuracy. They then formulated RWalk by combining a streamlined, computationally efficient parameter-regularization technique with a parameter-importance metric that tracks learning progress over the optimization trajectory. To address the severe intransigence that occurs when distinguishing between classes from different tasks, the method integrates small subsets of stored exemplar samples (typically 0.2% to 5% of past data) selected through strategies such as Mean-of-Features and uniform sampling.

Key findings show that RWalk consistently outperforms existing baselines across varied benchmarks, achieving an average accuracy of 82.5% on MNIST and 34.0% on CIFAR-100 in the realistic single-head evaluation setting with exemplar samples, compared to 79.7% and 33.6% for enhanced Elastic Weight Consolidation. In tests on CIFAR-100 using deep residual architectures, RWalk reached 70.1% accuracy, outperforming competitive approaches like Gradient Episodic Memory (65.4%) and iCaRL (50.8%). Storing just a tiny fraction of historical data significantly reduced intransigence, dropping the intransigence score from 0.8 to 0.05 on MNIST and turning standard regularized methods into viable continuous learners. Furthermore, RWalk showed significantly lower sensitivity to changes in regularization hyperparameter values compared to previous methods, maintaining stable performance across broad hyperparameter ranges.

These results demonstrate that catastrophic forgetting cannot be addressed in isolation; machine learning systems must actively manage the trade-off between retaining past knowledge and learning new capabilities. The efficiency of the proposed method allows models to operate with a constant memory footprint regardless of the number of sequential tasks, reducing operational costs and computational overhead for edge and continual learning applications. The findings also highlight that simplified evaluation protocols (where the task identity is known at test time) create a false sense of security, as real-world scenarios require models to distinguish among all learned classes simultaneously.

Organizations implementing continuous learning systems should adopt comprehensive evaluation metrics covering both accuracy and learning rigidity, rather than evaluating memory retention alone. Engineering teams should pair parameter-regularization algorithms with small memory buffers of past data to maintain task discrimination, using uniform or feature-mean sampling to balance performance with low selection overhead. Future research and development should extend these continuous learning frameworks to complex, high-dimensional computer vision tasks such as semantic segmentation and evaluate them on broader, more diverse real-world data streams.

arXiv: 1801.10112
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Elastic Weight Consolidation introduces the Fisher information metric to constrain parameter changes, serving as the foundational regularization framework that Riemannian Walk directly generalizes.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Synaptic Intelligence establishes path-integral importance measures computed along the optimization trajectory, providing the other core mechanism fused and generalized by Riemannian Walk.
  • Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). Memory Aware Synapses develops online sensitivity-based parameter regularization, providing essential background on gradient-based importance estimation for continual learning.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Gradient Episodic Memory defines key transfer and retention metrics in continual learning while establishing gradient-projection optimization that motivates better stability-plasticity trade-offs.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Learning without Forgetting introduces knowledge distillation as a baseline mechanism for preserving past task performance in deep incremental learning.
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL formalizes the class-incremental learning setup and benchmarks on CIFAR-100 that Riemannian Walk adopts and evaluates against.
Cover for Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence

Abstract

Incremental learning (IL) has received a lot of attention recently, however, the literature lacks a precise problem definition, proper evaluation settings, and metrics tailored specifically for the IL problem. One of the main objectives of this work is to fill these gaps so as to provide a common ground for better understanding of IL. The main challenge for an IL algorithm is to update the classifier whilst preserving existing knowledge. We observe that, in addition to forgetting, a known issue while preserving knowledge, IL also suffers from a problem we call intransigence, inability of a model to update its knowledge. We introduce two metrics to quantify forgetting and intransigence that allow us to understand, analyse, and gain better insights into the behaviour of IL algorithms. We present RWalk, a generalization of EWC++ (our efficient version of EWC [Kirkpatrick2016EWC]) and Path Integral [Zenke2017Continual] with a theoretically grounded KL-divergence based perspective. We provide a thorough analysis of various IL algorithms on MNIST and CIFAR-100 datasets. In these experiments, RWalk obtains superior results in terms of accuracy, and also provides a better trade-off between forgetting and intransigence.

Table of Contents

  • 1 Introduction
  • 2 Problem Set-up and Preliminaries
  • 2.1 Single-head vs Multi-head Evaluations
  • Why single-head evaluation for il?
  • 2.2 Probabilistic Interpretation of Neural Network Output
  • 2.3 KL-divergence as the Distance in the Riemannian Manifold
  • 3 Forgetting and Intransigence
  • Average Accuracy (AA)
  • Forgetting Measure (FF)
  • Intransigence Measure (II)
  • 4 Riemannian Walk for Incremental Learning
  • 4.1 Avoiding Catastrophic Forgetting
  • KL-divergence based Regularization (ewc++)
  • Optimization-path based Parameter Importance
  • Final Objective Function (RWalk)
  • 4.2 Handling Intransigence
  • Uniform Sampling
  • Plane Distance-based Sampling
  • Entropy-based Sampling
  • Mean of Features (MoF)
  • 5 Related Work
  • 6 Experiments
  • Datasets
  • Architectures
  • Baselines
  • 6.1 Results
  • Interplay of Forgetting and Intransigence
  • Effect of Increasing the Number of Samples
  • Comparison of Different Sampling Strategies
  • 7 Discussion
  • 0.A Approximate KL divergence using Fisher Information Matrix
  • 0.A.1 Proof of Approximate KL divergence
  • 0.A.2 Empirical vs True Fisher
  • Empirical Fisher
  • True Fisher
  • 0.B Additional Experiments and Analysis
  • 0.B.1 Comparison with GEM [15] on ResNets
  • 0.B.2 Effect of Regularization Hyperparameter (λ\lambda)
  • 0.B.3 CIFAR Architecture and Task-Level Analysis
  • References

Knowls

  1. Knowl 1 — Riemannian Walk (RWalk) Objective

    model/method

    Riemannian Walk (RWalk) is a parameter-regularization framework for incremental classification that combines Fisher Information Matrix curvature regularization with an optimization-path-based parameter importance score evaluated on the Riemannian manifold.

    When training incrementally on task kk with dataset Dk\mathcal{D}_k, given parameter vector θk−1∈RP\theta^{k-1} \in \mathbb{R}^P converged after task k−1k-1, RWalk minimizes the regularized loss function:

    L~k(θ)=Lk(θ)+λ∑i=1P(Fˉθik−1+sˉt0tk−1(θi))(θi−θik−1)2\tilde{L}^k(\theta) = L^k(\theta) + \lambda \sum_{i=1}^P \left(\bar{F}_{\theta_i}^{k-1} + \bar{s}_{t_0}^{t_{k-1}}(\theta_i)\right) (\theta_i - \theta_i^{k-1})^2

    where:

    • Lk(θ)L^k(\theta) is the task-specific empirical cross-entropy loss on Dk\mathcal{D}_k.
    • λ>0\lambda > 0 is the regularization trade-off hyperparameter.
    • Fˉθk−1∈[0,1]P\bar{F}_{\theta}^{k-1} \in [0, 1]^P is the normalized diagonal empirical Fisher Information Matrix accumulated up to task k−1k-1.
    • sˉt0tk−1(θ)∈[0,1]P\bar{s}_{t_0}^{t_{k-1}}(\theta) \in [0, 1]^P is the normalized optimization-path parameter importance score accumulated from training iteration t0t_0 to tk−1t_{k-1}.
    • Continual score averaging is applied across tasks to prevent the regularization from becoming overly rigid: st0tk−1(θi)=12(st0tk−2(θi)+stk−2tk−1(θi))s_{t_0}^{t_{k-1}}(\theta_i) = \frac{1}{2}\left(s_{t_0}^{t_{k-2}}(\theta_i) + s_{t_{k-2}}^{t_{k-1}}(\theta_i)\right), ensuring older tasks decay in relative influence.
    • The total training space complexity of RWalk is O(P)O(P), remaining constant with respect to the number of tasks.
  2. Knowl 2 — EWC++: Online and Memory-Efficient Fisher Information Estimation

    model/method

    Standard Elastic Weight Consolidation (EWC) regularizes parameters around past optima using the diagonal of the empirical Fisher Information Matrix (FIM), requiring O(kP)O(kP) memory to store separate FIMs for kk tasks and an expensive offline forward-backward pass over each task's dataset at the end of training.

    EWC++ maintains a single running diagonal empirical Fisher matrix Fθt∈RPF_\theta^t \in \mathbb{R}^P across training iterations tt via an exponential moving average:

    Fθt=αF^θt+(1−α)Fθt−1F_\theta^t = \alpha \hat{F}_\theta^t + (1 - \alpha) F_\theta^{t-1}

    where α∈[0,1]\alpha \in [0, 1] is a smoothing factor (set to α=0.9\alpha = 0.9 in experiments) and F^θt\hat{F}_\theta^t is the empirical Fisher matrix estimated on the current mini-batch BtB_t:

    F^θt=1∣Bt∣∑(x,y)∈Bt(∂log⁡pθ(y∣x)∂θ)2\hat{F}_\theta^t = \frac{1}{|B_t|} \sum_{(x, y) \in B_t} \left( \frac{\partial \log p_\theta(y|x)}{\partial \theta} \right)^2

    At the completion of task k−1k-1, the running Fisher Fθtk−1F_\theta^{t_{k-1}} is stored as Fθk−1F_\theta^{k-1} and used to regularize task kk. This reduces memory storage to two parameter vectors (O(P)O(P)) at all times and eliminates the post-training dataset pass.

  3. Knowl 3 — Optimization-Path Parameter Importance on the Riemannian Manifold

    model/method

    Because the Fisher Information Matrix only reflects the local curvature at a local minimum of the current loss, it is blind to parameter influence along the optimization trajectory. RWalk computes parameter importance as the ratio of loss reduction to Riemannian manifold distance (approximated by the second-order KL-divergence of the model likelihoods).

    For parameter θi\theta_i over training steps from iteration t1t_1 to t2t_2, evaluated at discrete intervals Δt≥1\Delta t \ge 1, the importance score is:

    st1t2(θi)=∑t=t1t2ΔLtt+Δt(θi)12Fθit(Δθi(t))2+ϵs_{t_1}^{t_2}(\theta_i) = \sum_{t=t_1}^{t_2} \frac{\Delta L_t^{t+\Delta t}(\theta_i)}{\frac{1}{2} F_{\theta_i}^t (\Delta \theta_i(t))^2 + \epsilon}

    where:

    • Δθi(t)=θi(t+Δt)−θi(t)\Delta \theta_i(t) = \theta_i(t + \Delta t) - \theta_i(t) is the parameter displacement.
    • ϵ>0\epsilon > 0 is a small positive constant preventing numerical instability.
    • FθitF_{\theta_i}^t is the online diagonal empirical Fisher information at step tt.
    • ΔLtt+Δt(θi)≈−∑τ=tt+Δt−1∂L∂θi(τ)⋅(θi(τ+1)−θi(τ))\Delta L_t^{t+\Delta t}(\theta_i) \approx - \sum_{\tau=t}^{t+\Delta t - 1} \frac{\partial L}{\partial \theta_i}(\tau) \cdot (\theta_i(\tau + 1) - \theta_i(\tau)) is the first-order approximation of the accumulated loss reduction caused by parameter θi\theta_i.

    Negative values of s(θi)s(\theta_i) are clipped to zero so that only positive contributions to loss reduction are accumulated.

  4. Knowl 4 — Forgetting Measure in Incremental Learning

    definition

    Let ak,j∈[0,1]a_{k,j} \in [0, 1] denote the classification accuracy on the held-out test set of task j≤kj \le k after the model has been trained sequentially from task 11 through task kk.

    The task-level forgetting for task j<kj < k after completing task kk is defined as the maximum drop from the peak historical accuracy achieved on task jj during previous incremental training steps:

    fjk=max⁡l∈{1,…,k−1}al,j−ak,j,∀j<kf_j^k = \max_{l \in \{1, \dots, k-1\}} a_{l, j} - a_{k, j}, \quad \forall j < k

    The average forgetting at task kk, denoted Fk∈[−1,1]F_k \in [-1, 1], is the mean forgetting across all previously learned tasks:

    Fk=1k−1∑j=1k−1fjkF_k = \frac{1}{k-1} \sum_{j=1}^{k-1} f_j^k

    A lower value of FkF_k indicates less forgetting. A negative score fjk<0f_j^k < 0 corresponds to positive backward transfer (PBT), indicating that learning task kk improved performance on task jj, while fjk>0f_j^k > 0 indicates negative backward transfer (NBT).

  5. Knowl 5 — Intransigence Measure in Incremental Learning

    definition

    Intransigence quantifies an incremental learning algorithm's inability to learn new tasks due to parameter over-regularization or restricted capacity.

    Let ak,k∈[0,1]a_{k,k} \in [0, 1] be the test accuracy on task kk immediately after training the incremental model on task kk. Let ak∗∈[0,1]a_k^* \in [0, 1] be the reference test accuracy on task kk obtained by a joint/reference model trained with simultaneous access to all training datasets up to task kk, ⋃l=1kDl\bigcup_{l=1}^k \mathcal{D}_l.

    The intransigence metric on task kk, Ik∈[−1,1]I_k \in [-1, 1], is defined as:

    Ik=ak∗−ak,kI_k = a_k^* - a_{k,k}

    Lower values of IkI_k indicate better model plasticity. Ik<0I_k < 0 corresponds to positive forward transfer (PFT), where sequential learning enhances learning of task kk compared to the reference joint model; Ik>0I_k > 0 indicates negative forward transfer (NFT).

  6. Knowl 6 — Single-Head vs Multi-Head Incremental Learning Evaluation Protocols

    definition

    In incremental classification across a sequence of tasks with disjoint class label sets Yk\mathcal{Y}^k:

    • Multi-Head Evaluation: The task identifier kk is provided at test time. The network's output space for task kk is restricted to the task-specific label set Yk\mathcal{Y}^k. As a result, the model only performs intra-task classification and is never evaluated on its ability to discriminate classes across different tasks.

    • Single-Head Evaluation: The task identifier is unknown at test time. The network must predict labels from the unified output space comprising all classes seen so far, Y1:k=⋃j=1kYj\mathcal{Y}_{1:k} = \bigcup_{j=1}^k \mathcal{Y}^j. The model must simultaneously perform intra-task and inter-task discrimination with limited or no access to data from previous tasks.

    While multi-head evaluation often exhibits near-zero forgetting and low intransigence across algorithms, single-head evaluation reveals severe performance degradation, exposing catastrophic forgetting and high intransigence.

  7. Knowl 7 — Exemplar Sampling Strategies for Mitigating Intransigence

    model/method

    In single-head incremental learning, models suffer from high intransigence because the absence of past-task training samples prevents the network from learning inter-task decision boundaries. Storing a small subset of exemplars (≤5%\le 5\% of dataset size) per class alleviates this issue. Four sampling strategies are evaluated:

    1. Uniform Sampling: Selects mm exemplars uniformly at random from each class dataset.

    2. Plane Distance-based Sampling (PD): Computes the pseudo-distance to the decision boundary d(xi)=ϕ(xi)⊤wyid(x_i) = \phi(x_i)^\top w_{y_i}, where ϕ(xi)\phi(x_i) is the feature representation and wyiw_{y_i} are the final layer weights for class yiy_i. Exemplars are sampled with probability q(xi)∝1/d(xi)q(x_i) \propto 1 / d(x_i) to retain boundary-defining samples.

    3. Entropy-based Sampling: Selects samples with probability proportional to the predictive Shannon entropy of the softmax output distribution pθ(y∣xi)p_\theta(y|x_i), targeting uncertain instances.

    4. Mean of Features (MoF): Selects mm exemplars per class such that their average feature embedding 1m∑ϕ(xi)\frac{1}{m}\sum \phi(x_i) best approximates the empirical class feature mean, with subset selection complexity O(n⋅f⋅m)O(n \cdot f \cdot m) where nn is class size and ff is feature dimension.

  8. Knowl 8 — Empirical Benchmark on Split MNIST and Split CIFAR-100

    data/table

    Evaluation of incremental learning methods on 5-task Split MNIST (MLP with two 256-unit hidden layers) and 10-task Split CIFAR-100 (4-layer CNN) under multi-head and single-head evaluation protocols. Methods with exemplar replay are denoted by suffix '-S' (10 samples/class [0.2%] on MNIST; 25 samples/class [5%] on CIFAR-100, selected via Mean of Features):

    Methods MNIST CIFAR
    A5(%)A_5(\%) F5F_5 I5I_5 A10(%)A_{10}(\%) F10F_{10} I10I_{10}
    Multi-head Evaluation
    Vanilla 90.3 0.12 6.6×10−46.6 \times 10^{-4} 44.4 0.36 0.02
    EWC 99.3 0.001 0.01 72.8 0.001 0.07
    PI 99.3 0.002 0.01 73.2 0 0.06
    RWalk (Ours) 99.3 0.003 0.01 74.2 0.004 0.04
    Single-head Evaluation
    Vanilla 38.0 0.62 0.29 10.2 0.36 -0.06
    EWC 55.8 0.08 0.77 23.1 0.03 0.17
    PI 57.6 0.11 0.80 22.8 0.04 0.20
    iCaRL-hb1 36.6 0.68 -0.01 7.4 0.40 0.06
    iCaRL 55.8 0.19 0.46 9.5 0.11 0.35
    Vanilla-S 73.7 0.30 0.03 12.9 0.64 -0.30
    EWC-S 79.7 0.14 0.22 33.6 0.27 -0.05
    PI-S 78.7 0.24 0.05 33.6 0.27 -0.03
    RWalk (Ours) 82.5 0.15 0.14 34.0 0.28 -0.06

    In single-head evaluation without exemplars, parameter-regularized methods (EWC, PI) suffer from severe intransigence (I5≥0.77I_5 \ge 0.77 on MNIST), resulting in low average accuracy. Incorporating a small set of exemplars dramatically reduces intransigence (e.g., PI intransigence drops from 0.80 to 0.05 on MNIST). RWalk consistently achieves the highest average accuracy across settings (82.5% on MNIST-S, 34.0% on CIFAR-S) while maintaining a balanced trade-off between forgetting and intransigence.

  9. Knowl 9 — Comparison with Gradient Episodic Memory (GEM) on ResNet18

    data/table

    Comparison on Split CIFAR-100 divided into 20 tasks of 5 classes each using a ResNet18 backbone under the multi-head evaluation setting:

    Methods Total Number of Samples AkA_k (%)
    iCaRL 5120 50.8
    GEM 5120 65.4
    RWalk (Ours) 5000 70.1

    RWalk outperforms both iCaRL (50.8%) and GEM (65.4%) by achieving an average multi-head accuracy of 70.1% across the 20 tasks while utilizing a slightly smaller total memory buffer (5000 vs 5120 samples).

  10. Knowl 10 — Robustness to Regularization Hyperparameter Sensitivity

    empirical result

    In incremental learning algorithms regularized by parameter importance, sensitivity to the regularization coefficient λ\lambda determines practical applicability across varying numbers of tasks.

    When λ\lambda is varied by a factor of 1×1051 \times 10^5 on Split MNIST (from λ=0.1\lambda = 0.1 to λ=10000\lambda = 10000):

    • EWC exhibits a change of ΔF5=−0.06\Delta F_5 = -0.06 in forgetting and ΔI5=+0.14\Delta I_5 = +0.14 in intransigence (average accuracy varying between 80.3%80.3\% and 79.1%79.1\%).
    • Path Integral (PI) exhibits a change of ΔF5=−0.07\Delta F_5 = -0.07 and ΔI5=+0.13\Delta I_5 = +0.13 (average accuracy varying between 78.5%78.5\% and 80.3%80.3\%).
    • RWalk demonstrates zero change (ΔF5=0.0,ΔI5=0.0\Delta F_5 = 0.0, \Delta I_5 = 0.0) across the entire 10510^5 span of λ\lambda, with forgetting remaining constant at F5=0.16F_5 = 0.16, intransigence remaining at I5=0.12I_5 = 0.12, and average accuracy remaining stable at 81.6%−82.6%81.6\% - 82.6\%.

    A consistent pattern of insensitivity to λ\lambda is observed on CIFAR-100, which is attributed to RWalk's individual [0,1][0, 1] normalization and continual score averaging of Fisher and path-importance weights.

  11. Knowl 11 — Sampling Strategy Sensitivity and Replay Regime Behavior

    empirical result

    Analysis of exemplar sampling methods and buffer sizes indicates:

    • Mean of Features (MoF) achieves the highest incremental classification accuracy, but uniform random sampling is nearly as effective despite requiring no feature extraction or iterative optimization.
    • Regularized models (RWalk, EWC, PI) are largely insensitive to the choice of sampling strategy, whereas unregularized Vanilla training fluctuates substantially across strategies because its output layer weights for previous tasks undergo unconstrained drift.
    • In the small exemplar memory regime (e.g., ≤20\le 20 samples per class on MNIST, ≤50\le 50 samples per class on CIFAR-100), regularized methods strongly outperform Vanilla. In the high-sample regime (e.g., 200 samples per class or 40% of CIFAR-100), Vanilla training matches regularized methods as the network can effectively relearn previous tasks directly from memory.

Coverage note — The formal Taylor expansion proof of Lemma 1 (relating KL-divergence to the empirical Fisher Information Matrix) in the supplementary material was deliberately omitted in accordance with the rule prohibiting pure mathematical derivations.

References

  1. 1.Amari, S.I.: Natural gradient works efficiently in learning. Neural Computation (1998) 2, 4, 15, 16, 17
  2. 2.Grosse, R., Martens, J.: A kronecker-factored approximate fisher matrix for convolution layers. In: ICML (2016) 4
  3. 3.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016) 17
  4. 4.Hecht-Nielsen, R., et al.: Theory of the backpropagation neural network. Neural Networks 1(Supplement-1), 445–448 (1988) 9
  5. 5.Hinton, G., Vinyals, O., Dean, J.: Distilling the knowledge in a neural network. In: NIPS (2014) 1
  6. 6.Kingma, D., Ba, J.: Adam: A method for stochastic optimization. In: ICLR (2015) 10
  7. 7.Kirkpatrick, J., Pascanu, R., Rabinowitz, N.C., Veness, J., Desjardins, G., Rusu, A.A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., Hadsell, R.: Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences of the United States of America (PNAS) (2016) 1, 2, 4, 6, 7, 8, 9, 10, 18
  8. 8.Krizhevsky, A., Hinton, G.: Learning multiple layers of features from tiny images. https://www.cs.toronto.edu/ kriz/cifar.html (2009) 2
  9. 9.Kullback, S., Leibler, R.A.: On information and sufficiency. The Annals of Mathematical Statistics (1951) 4
  10. 10.Le Roux, N., Pierre-Antoine, M., Bengio, Y.: Topmoumoute online natural gradient algorithm. In: NIPS (2007) 4
  11. 11.LeCun, Y.: The mnist database of handwritten digits. http://yann. lecun. com/exdb/mnist/ (1998) 2
  12. 12.Lee, J.M.: Riemannian manifolds: an introduction to curvature, vol. 176. Springer Science & Business Media (2006) 4
  13. 13.Lee, S.W., Kim, J.H., Ha, J.W., Zhang, B.T.: Overcoming catastrophic forgetting by incremental moment matching. In: NIPS (2017) 3, 10
  14. 14.Li, Z., Hoiem, D.: Learning without forgetting. In: ECCV (2016) 9
  15. 15.Lopez-Paz, D., Ranzato, M.: Gradient episodic memory for continuum learning. In: NIPS (2017) 1, 5, 6, 10, 11, 15, 17, 18
  16. 16.Martens, J., Grosse, R.: Optimizing neural networks with kronecker-factored approximate curvature. In: ICML (2015) 4, 7
  17. 17.Nguyen, C.V., Li, Y., Bui, T.D., Turner, R.E.: Variational continual learning. ICLR (2018) 10
  18. 18.Pascanu, R., Bengio, Y.: Revisiting natural gradient for deep networks. In: ICLR (2014) 2, 4, 15, 16, 17
  19. 19.Rebuffi, S.A., Bilen, H., Vedaldi, A.: Learning multiple visual domains with residual adapters. In: NIPS (2017) 9
  20. 20.Rebuffi, S.V., Kolesnikov, A., Lampert, C.H.: iCaRL: Incremental classifier and representation learning. In: CVPR (2017) 1, 3, 9, 10, 12, 13, 17
  21. 21.Rusu, A.A., Rabinowitz, N.C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., Hadsell, R.: Progressive neural networks. arXiv preprint arXiv:1606.04671 (2016) 9
  22. 22.Schwarz, J., Luketina, J., Czarnecki, W.M., Grabska-Barwinska, A., Teh, Y.W., Pascanu, R., Hadsell, R.: Progress & compress: A scalable framework for continual learning. In: ICML (2018) 2, 7, 10
  23. 23.Shin, H., Lee, J.K., Kim, J., Kim, J.: Continual learning with deep generative replay. In: NIPS (2017) 10
  24. 24.Terekhov, A.V., Montone, G., ORegan, J.K.: Knowledge transfer in deep block-modular neural networks. In: Conference on Biomimetic and Biohybrid Systems. pp. 268–279 (2015) 9
  25. 25.Yoon, J., Yang, E., Lee, J., Hwang, S.J.: Lifelong learning with dynamically expandable networks. In: ICLR (2018) 9
  26. 26.Zenke, F., Poole, B., Ganguli, S.: Continual learning through synaptic intelligence. In: ICML (2017) 1, 2, 6, 7, 8, 9, 10, 11, 18, 19

Citation

MLA
Chaudhry, A., et al. “Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence”. Lecture Notes in Computer Science, Springer International Publishing, 2018, pp. 556–72, https://doi.org/10.1007/978-3-030-01252-6_33.
APA
Chaudhry, A., Dokania, P. K., Ajanthan, T., & Torr, P. H. S. (2018). Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence. In Lecture Notes in Computer Science (pp. 556–572). Springer International Publishing. https://doi.org/10.1007/978-3-030-01252-6_33
Chicago
Chaudhry, A., P. K. Dokania, T. Ajanthan, and P. H. S. Torr. 2018. “Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence”. In Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-030-01252-6_33.
Harvard
Chaudhry, A. et al. (2018) “Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence”, Lecture Notes in Computer Science. Springer International Publishing, pp. 556–572. Available at: https://doi.org/10.1007/978-3-030-01252-6_33.
Vancouver
1. Chaudhry A, Dokania PK, Ajanthan T, Torr PHS (2018) Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence. In: Lecture Notes in Computer Science. Springer International Publishing, pp 556–572

BibTeX

@inbook{Chaudhry_2018, title={Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence}, ISBN={9783030012526}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-030-01252-6_33}, DOI={10.1007/978-3-030-01252-6_33}, booktitle={Computer Vision – ECCV 2018}, publisher={Springer International Publishing}, author={Chaudhry, Arslan and Dokania, Puneet K. and Ajanthan, Thalaiyasingam and Torr, Philip H. S.}, year={2018}, pages={556–572} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF