Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
Arslan ChaudhryPuneet K. DokaniaThalaiyasingam AjanthanPhilip H. S. Torr
Proposes Riemannian Walk (RWalk), a continual learning method based on KL-divergence that optimizes the trade-off between catastrophic forgetting and model intransigence, while establishing formal metrics to quantify both behaviors.
Modern artificial intelligence systems struggle when required to learn new tasks sequentially over time without forgetting previously acquired knowledge. While standard machine learning models train on a complete, fixed dataset, real-world deployment requires continuous learning from sequential streams of data. Existing continuous learning methods often focus solely on mitigating catastrophic forgetting (losing past knowledge), but they frequently suffer from intransigence, which is the model's inability to adapt and learn new information effectively. This trade-off between remembering past tasks and acquiring new ones poses a major bottleneck for practical, long-term deployment.
The article establishes a formal framework to quantify both forgetting and intransigence in incremental learning settings and introduces RWalk, an efficient algorithm designed to balance these competing demands. The approach evaluates sequential image classification on standard visual benchmarks, including split MNIST and split CIFAR-100 datasets, under realistic evaluation setups where task identities are not provided at test time.
The researchers developed two formal metrics to independently measure forgetting and intransigence alongside standard classification accuracy. They then formulated RWalk by combining a streamlined, computationally efficient parameter-regularization technique with a parameter-importance metric that tracks learning progress over the optimization trajectory. To address the severe intransigence that occurs when distinguishing between classes from different tasks, the method integrates small subsets of stored exemplar samples (typically 0.2% to 5% of past data) selected through strategies such as Mean-of-Features and uniform sampling.
Key findings show that RWalk consistently outperforms existing baselines across varied benchmarks, achieving an average accuracy of 82.5% on MNIST and 34.0% on CIFAR-100 in the realistic single-head evaluation setting with exemplar samples, compared to 79.7% and 33.6% for enhanced Elastic Weight Consolidation. In tests on CIFAR-100 using deep residual architectures, RWalk reached 70.1% accuracy, outperforming competitive approaches like Gradient Episodic Memory (65.4%) and iCaRL (50.8%). Storing just a tiny fraction of historical data significantly reduced intransigence, dropping the intransigence score from 0.8 to 0.05 on MNIST and turning standard regularized methods into viable continuous learners. Furthermore, RWalk showed significantly lower sensitivity to changes in regularization hyperparameter values compared to previous methods, maintaining stable performance across broad hyperparameter ranges.
These results demonstrate that catastrophic forgetting cannot be addressed in isolation; machine learning systems must actively manage the trade-off between retaining past knowledge and learning new capabilities. The efficiency of the proposed method allows models to operate with a constant memory footprint regardless of the number of sequential tasks, reducing operational costs and computational overhead for edge and continual learning applications. The findings also highlight that simplified evaluation protocols (where the task identity is known at test time) create a false sense of security, as real-world scenarios require models to distinguish among all learned classes simultaneously.
Organizations implementing continuous learning systems should adopt comprehensive evaluation metrics covering both accuracy and learning rigidity, rather than evaluating memory retention alone. Engineering teams should pair parameter-regularization algorithms with small memory buffers of past data to maintain task discrimination, using uniform or feature-mean sampling to balance performance with low selection overhead. Future research and development should extend these continuous learning frameworks to complex, high-dimensional computer vision tasks such as semantic segmentation and evaluate them on broader, more diverse real-world data streams.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Elastic Weight Consolidation introduces the Fisher information metric to constrain parameter changes, serving as the foundational regularization framework that Riemannian Walk directly generalizes.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). Synaptic Intelligence establishes path-integral importance measures computed along the optimization trajectory, providing the other core mechanism fused and generalized by Riemannian Walk.
- Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). Memory Aware Synapses develops online sensitivity-based parameter regularization, providing essential background on gradient-based importance estimation for continual learning.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Gradient Episodic Memory defines key transfer and retention metrics in continual learning while establishing gradient-projection optimization that motivates better stability-plasticity trade-offs.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Learning without Forgetting introduces knowledge distillation as a baseline mechanism for preserving past task performance in deep incremental learning.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL formalizes the class-incremental learning setup and benchmarks on CIFAR-100 that Riemannian Walk adopts and evaluates against.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey systematically benchmarks and categorizes continual learning techniques, incorporating Riemannian Walk's stability-plasticity and intransigence insights into a broader empirical framework.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). This comprehensive survey contextualizes Riemannian Walk within the theoretical taxonomy of regularization and optimization methods balancing the stability-plasticity dilemma.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Dark Experience Replay extends the preservation of optimization geometry to general continual learning by matching continuous logit distributions across tasks.
- Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). Averaged GEM builds on single-pass streaming evaluation protocols and efficiency considerations in continual learning, contrasting with parameter-regularization methods like Riemannian Walk.
- Paper: Large Scale Incremental Learning, Yue Wu et al. (2019). Bias Correction investigates the failure modes of incremental learning classifiers at scale, addressing representation bias that complements intransigence analysis.
