New Insights on Reducing Abrupt Representation Change in Online Continual Learning
Lucas CacciaRahaf AljundiNader AsadiTinne TuytelaarsJoelle PineauEugene Belilovsky
Demonstrates how Experience Replay causes disruptive representation shifts when new classes appear in online continual learning, and resolves this with an asymmetric update rule that forces incoming data to adapt to established representations.
Real-world machine learning systems frequently need to learn continuously from an incoming stream of data where new categories appear over time. A major failure mode in this setup is catastrophic forgetting, where learning new information abruptly overwrites previously acquired knowledge. While standard approaches store and replay a small memory buffer of past examples alongside new data, models still experience severe performance drops whenever new classes are introduced, especially under strict memory and computation limits.
The article investigates the root cause of this performance drop and proposes a practical training strategy to prevent disruption when new categories arrive. Specifically, it demonstrates how standard replay methods cause learned internal representations of past classes to drift drastically, and it develops asymmetric loss functions that isolate incoming classes from past classes during initial learning updates.
The authors conducted empirical evaluations across standard continual learning vision benchmarks, including Split CIFAR-10, Split CIFAR-100, and Split Mini-ImageNet, processing streams under realistic constraints without task indicators. They compared standard experience replay and several state-of-the-art baselines against two proposed variants: Experience Replay with Asymmetric Metric Learning (ER-AML) and Experience Replay with Asymmetric Cross-Entropy (ER-ACE). The evaluation tracked accuracy over time, total computational floating-point operations, memory consumption, and behavior across blurry, non-discrete task transitions.
The key findings demonstrate significant stability and accuracy improvements. First, the article identifies that standard cross-entropy loss causes unlearned new class samples to severely displace established older class representations at transition points. Second, the proposed ER-ACE method resolves this drift by restricting the loss on incoming data to only current classes while allowing replayed data to consolidate all classes, achieving an average relative accuracy gain of 36% over standard experience replay. Third, these gains are most pronounced in constrained, small memory buffer regimes where baseline accuracy drops heavily. Fourth, the asymmetric cross-entropy technique achieves top-tier performance without adding computational overhead, whereas competing baselines incur substantial training or prototype-recalculation costs.
These results show that preventing abrupt representation drift is more critical than previously assumed class-imbalance corrections in streaming environments. For technical leadership and deployment planning, adopting an asymmetric loss provides a highly cost-effective upgrade: it delivers higher model accuracy, lower forgetting, and consistent operational uptime without requiring additional hardware memory or compute resources.
Organizations deploying streaming machine learning systems should implement asymmetric loss masking on incoming batches to stabilize performance during distribution shifts. When selecting methods, teams should audit full computational and latency costs rather than relying solely on final accuracy benchmarks. Before broad deployment in non-vision domains, engineering teams should conduct pilot studies, as the current evaluation focuses primarily on image classification architectures using fixed-size convolutional neural networks.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). This paper establishes Dark Experience Replay as a key general continual learning baseline, providing direct context for comparing replay-based representation preservation techniques.
- Paper: Experience Replay for Continual Learning, David Rolnick et al. (2018). This work introduces foundational principles of experience replay in boundary-free continual learning settings that motivate representation drift mitigation.
- Paper: Large Scale Incremental Learning, Yue Wu et al. (2019). This study analyzes class imbalance and classifier bias during incremental updates, which the source paper directly investigates and contrasts against representation change.
- Paper: Learning a Unified Classifier Incrementally via Rebalancing, Saihui Hou et al. (2019). This paper examines feature drift and class imbalance in incremental learning, offering foundational approaches to loss design that precede asymmetric cross-entropy.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). This foundational paper introduces exemplar-based representation learning and distillation for incremental classification that modern online continual learning builds upon.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). This work formalizes online continual learning from streaming tasks and evaluates episodic memory updates using gradient constraints.
- Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). This paper establishes realistic single-pass evaluation protocols and efficient gradient replay mechanisms essential for streaming continual learning setups.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This comprehensive survey provides standard classification benchmarks and taxonomies of catastrophic forgetting that frame online continual learning research.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). This survey provides an updated theoretical and methodological synthesis of continual learning, contextualizing representation-based and replay-based solutions.
- Paper: Probing Representation Forgetting in Supervised and Unsupervised Continual Learning, MohammadReza Davari et al. (2022). This study extends the investigation into representation forgetting by utilizing linear probes to measure internal feature degradation across varying loss functions.
- Paper: Improving Task-free Continual Learning by Distributionally Robust Memory Evolution, Zhenyi Wang et al. (2022). This paper builds on task-free continual learning by actively evolving the replay memory distribution to counter representation drift and buffer overfitting.
- Paper: Class-Incremental Exemplar Compression for Class-Incremental Learning, Zilin Luo et al. (2023). This work addresses memory budget constraints in class-incremental learning by compressing exemplars to improve replay efficiency.
- Paper: Energy-based Latent Aligner for Incremental Learning, K. J. Joseph et al. (2022). This research introduces an energy-based latent alignment method to measure and reverse internal representation shifts caused by incremental updates.
- Paper: An Empirical Investigation of the Role of Pre-training in Lifelong Learning, Sanket Vaibhav Mehta et al. (2023). This work evaluates how pre-trained feature initializations impact representation stability and catastrophic forgetting across sequential task domains.
- Paper: Understanding Plasticity in Neural Networks, Clare Lyle et al. (2023). This paper investigates the underlying optimization dynamics and loss landscape properties responsible for plasticity loss and representational failure in continual learning.
