New Insights on Reducing Abrupt Representation Change in Online Continual Learning

Lucas CacciaRahaf AljundiNader AsadiTinne TuytelaarsJoelle PineauEugene Belilovsky

article2022ICLR308 citationsBest Paper Award

Demonstrates how Experience Replay causes disruptive representation shifts when new classes appear in online continual learning, and resolves this with an asymmetric update rule that forces incoming data to adapt to established representations.

Listen

Real-world machine learning systems frequently need to learn continuously from an incoming stream of data where new categories appear over time. A major failure mode in this setup is catastrophic forgetting, where learning new information abruptly overwrites previously acquired knowledge. While standard approaches store and replay a small memory buffer of past examples alongside new data, models still experience severe performance drops whenever new classes are introduced, especially under strict memory and computation limits.

The article investigates the root cause of this performance drop and proposes a practical training strategy to prevent disruption when new categories arrive. Specifically, it demonstrates how standard replay methods cause learned internal representations of past classes to drift drastically, and it develops asymmetric loss functions that isolate incoming classes from past classes during initial learning updates.

The authors conducted empirical evaluations across standard continual learning vision benchmarks, including Split CIFAR-10, Split CIFAR-100, and Split Mini-ImageNet, processing streams under realistic constraints without task indicators. They compared standard experience replay and several state-of-the-art baselines against two proposed variants: Experience Replay with Asymmetric Metric Learning (ER-AML) and Experience Replay with Asymmetric Cross-Entropy (ER-ACE). The evaluation tracked accuracy over time, total computational floating-point operations, memory consumption, and behavior across blurry, non-discrete task transitions.

The key findings demonstrate significant stability and accuracy improvements. First, the article identifies that standard cross-entropy loss causes unlearned new class samples to severely displace established older class representations at transition points. Second, the proposed ER-ACE method resolves this drift by restricting the loss on incoming data to only current classes while allowing replayed data to consolidate all classes, achieving an average relative accuracy gain of 36% over standard experience replay. Third, these gains are most pronounced in constrained, small memory buffer regimes where baseline accuracy drops heavily. Fourth, the asymmetric cross-entropy technique achieves top-tier performance without adding computational overhead, whereas competing baselines incur substantial training or prototype-recalculation costs.

These results show that preventing abrupt representation drift is more critical than previously assumed class-imbalance corrections in streaming environments. For technical leadership and deployment planning, adopting an asymmetric loss provides a highly cost-effective upgrade: it delivers higher model accuracy, lower forgetting, and consistent operational uptime without requiring additional hardware memory or compute resources.

Organizations deploying streaming machine learning systems should implement asymmetric loss masking on incoming batches to stabilize performance during distribution shifts. When selecting methods, teams should audit full computational and latency costs rather than relying solely on final accuracy benchmarks. Before broad deployment in non-vision domains, engineering teams should conduct pilot studies, as the current evaluation focuses primarily on image classification architectures using fixed-size convolutional neural networks.

Cover for New Insights on Reducing Abrupt Representation Change in Online Continual Learning

Abstract

In the online continual learning paradigm, agents must learn from a changing distribution while respecting memory and compute constraints. Experience Replay (ER), where a small subset of past data is stored and replayed alongside new data, has emerged as a simple and effective learning strategy. In this work, we focus on the change in representations of observed data that arises when previously unobserved classes appear in the incoming data stream, and new classes must be distinguished from previous ones. We shed new light on this question by showing that applying ER causes the newly added classes' representations to overlap significantly with the previous classes, leading to highly disruptive parameter updates. Based on this empirical analysis, we propose a new method which mitigates this issue by shielding the learned representations from drastic adaptation to accommodate new classes. We show that using an asymmetric update rule pushes new classes to adapt to the older ones (rather than the reverse), which is more effective especially at task boundaries, where much of the forgetting typically occurs. Empirical results show significant gains over strong baselines on standard continual learning benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Learning Setting and Notation
  • 4 Methods
  • 4.1 A Distance Metric Learning Approach for Reducing Drift (ER-AML)
  • 4.2 Negative Selection Affects Representation Drift
  • 4.3 Cross-entropy Based Alternative (ER-ACE)
  • 5 Experiments
  • 5.1 Datasets
  • 5.2 Baselines
  • 5.3 Evaluation Metrics and Considerations
  • 5.4 Standard Online Continual Learning Settings
  • 5.5 Blurry Task Boundaries
  • 6 Conclusion
  • 7 Reproducibility Statement
  • 8 Acknowledgements
  • References
  • A Experimental Setup
  • A.1 Hyperparameters
  • A.2 Blurry Task Boundaries Experiment
  • B An in-depth analysis of SS-IL in the online setting
  • C Overfitting on buffered samples
  • D combining ER-ACE with DER++
  • E Gradient Norm
  • F ER-AML with Triplet loss
  • G Ablations Negative Selection
  • H Additional Drift Results
  • I Analysis of the Representations During the Second Task
  • J Additional Blurry Task Boundaries Experiments
  • K Experiments with limited training data available
  • L Additional Results
  • L.1 Anytime Evaluation without Data Augmentation
  • L.2 Anytime Evaluation with Data Augmentation

Knowls

  1. Knowl 1 — Experience Replay with Asymmetric Cross-Entropy

    model/method

    Experience Replay with Asymmetric Cross-Entropy (ER-ACE) is an online continual learning method designed to prevent abrupt representation drift of previously learned classes when new classes are introduced in the input stream.

    Let M\mathcal{M} denote a fixed-size rehearsal buffer. At each training step, the learner receives an incoming batch XinX^{in} from the data stream and samples a rehearsal batch XbfX^{bf} from M\mathcal{M}. Let CcurrC_{curr} denote the set of classes present in the incoming batch XinX^{in}, ColdC_{old} denote previously observed classes that are not present in XinX^{in}, and Call=Cold∪CcurrC_{all} = C_{old} \cup C_{curr} denote the set of all classes observed up to the current timestep.

    For any batch of inputs XX and subset of target classes C⊆CallC \subseteq C_{all}, the restricted cross-entropy loss is defined as:

    Lce(X,C)=−∑x∈Xlog⁡exp⁡(logitc(x)(x))∑c∈Cexp⁡(logitc(x))\mathcal{L}_{ce}(X, C) = - \sum_{x \in X} \log \frac{\exp(\text{logit}_{c(x)}(x))}{\sum_{c \in C} \exp(\text{logit}_c(x))}

    where c(x)c(x) is the ground-truth class label of sample xx, and logitc(x)=wcTfθ(x)\text{logit}_c(x) = w_c^T f_\theta(x) is the unnormalized output for class cc computed from the feature representation fθ(x)f_\theta(x) and class prototype vector wcw_c.

    The total training loss applied to the combined batch Xbf∪XinX^{bf} \cup X^{in} at each step is:

    Lace(Xbf∪Xin)=Lce(Xbf,Cold∪Ccurr)+Lce(Xin,Ccurr)\mathcal{L}_{ace}(X^{bf} \cup X^{in}) = \mathcal{L}_{ce}(X^{bf}, C_{old} \cup C_{curr}) + \mathcal{L}_{ce}(X^{in}, C_{curr})

    In implementation, the restricted cross-entropy Lce(Xin,Ccurr)\mathcal{L}_{ce}(X^{in}, C_{curr}) is computed by masking logits of classes outside CcurrC_{curr} with a large negative value (e.g., −109-10^9) before applying standard softmax cross-entropy. This allows representations of new classes to be formed in isolation without exerting negative gradient pressure on old class prototypes, while the unmasked rehearsal loss Lce(Xbf,Cold∪Ccurr)\mathcal{L}_{ce}(X^{bf}, C_{old} \cup C_{curr}) learns discrimination across all observed classes without requiring task identifiers.

  2. Knowl 2 — Experience Replay with Asymmetric Metric Learning

    model/method

    Experience Replay with Asymmetric Metric Learning (ER-AML) isolates incoming class representation learning using a Supervised Contrastive (SupCon) loss with restricted negative selection, combined with a prototype-based cross-entropy loss on replayed memory samples.

    Let fθ(x)f_\theta(x) denote the normalized hidden feature representation of input xx, and let sim(a,b)=exp⁡(aTbτ∥a∥∥b∥)\text{sim}(a, b) = \exp\left(\frac{a^T b}{\tau \|a\| \|b\|}\right) compute the exponential cosine similarity between two vectors with temperature scaling factor τ>0\tau > 0.

    For an incoming batch XinX^{in}, positive samples P(xi)⊂Xin∪MP(x_i) \subset X^{in} \cup \mathcal{M} are selected from samples sharing class label c(xi)c(x_i). Negative samples N(xi)N(x_i) are restricted strictly to samples whose labels belong to classes present in the current incoming mini-batch XinX^{in} (excluding previous classes not in XinX^{in}):

    L1(Xin)=−∑xi∈Xin1∣P(xi)∣∑xp∈P(xi)log⁡sim(fθ(xp),fθ(xi))∑xn∈N(xi)∪P(xi)sim(fθ(xn),fθ(xi))\mathcal{L}_1(X^{in}) = - \sum_{x_i \in X^{in}} \frac{1}{|P(x_i)|} \sum_{x_p \in P(x_i)} \log \frac{\text{sim}(f_\theta(x_p), f_\theta(x_i))}{\sum_{x_n \in N(x_i) \cup P(x_i)} \text{sim}(f_\theta(x_n), f_\theta(x_i))}

    For replayed buffer data Xbf∼MX^{bf} \sim \mathcal{M}, a modified cross-entropy loss is applied over all observed classes CallC_{all} using learnable class weight prototypes {wc}c∈Call\{w_c\}_{c \in C_{all}}:

    L2(Xbf)=−∑x∈Xbflog⁡sim(wc(x),fθ(x))∑c∈Callsim(wc,fθ(x))\mathcal{L}_2(X^{bf}) = - \sum_{x \in X^{bf}} \log \frac{\text{sim}(w_{c(x)}, f_\theta(x))}{\sum_{c \in C_{all}} \text{sim}(w_c, f_\theta(x))}

    The total training objective combines both terms:

    L(Xin∪Xbf)=γL1(Xin)+L2(Xbf)\mathcal{L}(X^{in} \cup X^{bf}) = \gamma \mathcal{L}_1(X^{in}) + \mathcal{L}_2(X^{bf})

    where γ>0\gamma > 0 is a scalar weighting hyperparameter. Restricting N(xi)N(x_i) prevents incoming unlearned samples from displacing learned representations of past classes, while prototype-based classification avoids the high inference cost of nearest-class-mean search.

  3. Knowl 3 — Gradient Dynamics and Root Cause of Representation Drift at Task Boundaries

    theoretical result

    In standard Experience Replay (ER), applying symmetric cross-entropy across all seen classes to both incoming and replayed data produces severe representation drift and catastrophic forgetting at task boundaries.

    Let fθ(x)f_\theta(x) be the penultimate layer representation of input xx, W=[w1,…,w∣Call∣]TW = [w_1, \dots, w_{|C_{all}|}]^T be the matrix of class prototype weights for all observed classes CallC_{all}, p\mathbf{p} be the softmax prediction vector over CallC_{all}, y\mathbf{y} be the one-hot ground-truth label vector, and 1y∈C\mathbf{1}_{y \in C} be a binary mask vector indicating membership in a class subset C⊆CallC \subseteq C_{all}. The gradient of the class-restricted cross-entropy loss Lce(x,C)\mathcal{L}_{ce}(x, C) with respect to fθ(x)f_\theta(x) is:

    ∂Lce(x,C)∂fθ(x)=WT((p−y)⊙1y∈C)\frac{\partial \mathcal{L}_{ce}(x, C)}{\partial f_\theta(x)} = W^T \left( (\mathbf{p} - \mathbf{y}) \odot \mathbf{1}_{y \in C} \right)

    When a new task begins, incoming samples from previously unobserved classes have representations that are initially unorganized and frequently lie near existing clusters of older classes. Under symmetric loss (C=CallC = C_{all}), these poorly embedded incoming samples act as negative anchors against all previous classes ColdC_{old}. The resulting negative gradient components on older class prototypes overwhelm the positive gradient signals, producing a sharp spike in the L1L_1-norm of the gradient with respect to previous class representations. This causes abrupt displacement of older class prototypes and features.

    By restricting the incoming loss denominator to classes present in the current batch (C=CcurrC = C_{curr}), 1y∈C\mathbf{1}_{y \in C} sets the gradient contribution for all c∈Coldc \in C_{old} to zero on incoming samples. This eliminates gradient spikes on past class representations and forces new class representations to adapt around established clusters rather than displacing them.

  4. Knowl 4 — ER-AML Continual Training Procedure

    algorithm

    The ER-AML training procedure processes an incoming stream of labeled data mini-batches while maintaining a rehearsal buffer M\mathcal{M} updated via reservoir sampling.

    Input: Learning rate α\alpha, loss scaling factor γ\gamma, temperature τ\tau
    Initialize: Memory buffer M←∅\mathcal{M} \leftarrow \emptyset, Model parameters θ\theta
    while stream has data do
        Receive mini-batch XinX^{in} from stream
        Xpos,Xneg←FetchPosNeg(Xin,M)X_{pos}, X_{neg} \leftarrow \text{FetchPosNeg}(X^{in}, \mathcal{M})
        Xbf←Sample(M)X^{bf} \leftarrow \text{Sample}(\mathcal{M})
        L←γL1(Xin,Xpos,Xneg)+L2(Xbf)\mathcal{L} \leftarrow \gamma \mathcal{L}_1(X^{in}, X_{pos}, X_{neg}) + \mathcal{L}_2(X^{bf})
        θ←θ−α∇θL\theta \leftarrow \theta - \alpha \nabla_\theta \mathcal{L}
        M←ReservoirUpdate(M,Xin)\mathcal{M} \leftarrow \text{ReservoirUpdate}(\mathcal{M}, X^{in})
    end while

    The subroutine FetchPosNeg(Xin,M)\text{FetchPosNeg}(X^{in}, \mathcal{M}) selects, for each anchor xi∈Xinx_i \in X^{in}, positive samples from Xin∪MX^{in} \cup \mathcal{M} belonging to c(xi)c(x_i), and negative samples drawn exclusively from classes present in XinX^{in}. Replay batch XbfX^{bf} is sampled uniformly from M\mathcal{M}. Model parameters θ\theta are updated via stochastic gradient descent, and buffer M\mathcal{M} is updated using standard reservoir sampling to ensure uniform representation of the stream without requiring task boundary signals.

  5. Knowl 5 — Online Continual Learning Formulation and Anytime Evaluation Metrics

    definition

    In the online continual learning single-head classification framework, an agent observes a sequence of data batches (Xtin,Ytin)(X^{in}_t, Y^{in}_t) sampled from a non-stationary data distribution Dt\mathcal{D}_t. The data distribution changes sequentially across TT tasks at unknown change points without task descriptors provided at training or evaluation time (single shared output head performing NN-way classification over all observed classes). The learner processes data in a single pass under constant buffer memory M\mathcal{M} and bounded compute budgets.

    To ensure fair evaluation across the entire stream rather than relying solely on post-training accuracy, the framework defines two primary metrics:

    1. Anytime Accuracy (AAkAA_k): The classification accuracy evaluated at time step kk across the combined test sets of all distributions observed up to step kk.

    2. Averaged Anytime Accuracy (AAAAAA): The mean of the Anytime Accuracies measured periodically across all TT evaluation checkpoints:

    AAA=1T∑t=1T(AA)tAAA = \frac{1}{T} \sum_{t=1}^T (AA)_t

    1. Cumulative Computation Cost: The total floating-point operations (FLOPs) expended across training and inference recomputations (such as Nearest Class Mean prototype updates) over TT steps:

    Comp=∑t=1TO(m(⋅;θt))\text{Comp} = \sum_{t=1}^T \mathcal{O}(m(\cdot; \theta_t))

  6. Knowl 6 — Performance Comparison on Split CIFAR-10 Across Buffer Budgets

    data/table

    The performance of online continual learning methods evaluated on Split CIFAR-10 (5 disjoint 2-class tasks, single pass, reduced ResNet-18, batch size 10) across rehearsal buffer capacities M∈{5,20,100}M \in \{5, 20, 100\} per class is summarized below. Averaged Anytime Accuracy (AAAAAA), final accuracy (AccAcc), total training TFLOPs, and memory usage (Mb) are reported over 10 runs.

    Method Aug. M=5M = 5 M=20M = 20 M=100M = 100 Train Mem.
    AAA Acc AAA Acc AAA Acc TFLOPs (Mb)
    iid No - 62.7±0.762.7\pm0.7 - 62.7±0.762.7\pm0.7 - 62.7±0.762.7\pm0.7 8 4
    iid++ No - 72.9±0.772.9\pm0.7 - 72.9±0.772.9\pm0.7 - 72.9±0.772.9\pm0.7 16 4
    DER++ Yes 50.7±1.150.7\pm1.1 31.8±0.931.8\pm0.9 55.6±1.255.6\pm1.2 39.3±1.039.3\pm1.0 60.1±1.360.1\pm1.3 52.3±1.152.3\pm1.1 24 (4, 7)
    ER No 40.0±0.840.0\pm0.8 19.7±0.319.7\pm0.3 45.2±1.345.2\pm1.3 26.7±1.026.7\pm1.0 55.4±1.455.4\pm1.4 38.7±0.838.7\pm0.8 17 (4, 7)
    ER Yes 45.6±1.145.6\pm1.1 28.4±1.028.4\pm1.0 55.9±1.255.9\pm1.2 40.3±0.640.3\pm0.6 60.3±1.360.3\pm1.3 49.4±1.349.4\pm1.3 17 (4, 7)
    iCaRL No 47.0±0.847.0\pm0.8 30.6±0.830.6\pm0.8 55.1±0.755.1\pm0.7 41.7±0.641.7\pm0.6 59.3±0.659.3\pm0.6 45.1±0.645.1\pm0.6 (21, 47) (8, 11)
    iCaRL Yes 49.1±1.049.1\pm1.0 33.4±1.033.4\pm1.0 54.4±0.754.4\pm0.7 39.2±0.839.2\pm0.8 56.9±0.756.9\pm0.7 42.3±0.842.3\pm0.8 (21, 47) (8, 11)
    MIR No 39.3±1.039.3\pm1.0 19.7±0.519.7\pm0.5 44.7±1.144.7\pm1.1 29.7±0.629.7\pm0.6 53.8±1.753.8\pm1.7 43.3±1.043.3\pm1.0 41 (4, 7)
    MIR Yes 44.9±0.944.9\pm0.9 29.8±0.829.8\pm0.8 49.7±1.049.7\pm1.0 41.8±0.641.8\pm0.6 54.6±1.454.6\pm1.4 49.3±0.649.3\pm0.6 41 (4, 7)
    SS-IL No 42.6±1.742.6\pm1.7 29.6±0.429.6\pm0.4 44.8±1.844.8\pm1.8 35.1±0.935.1\pm0.9 48.1±2.248.1\pm2.2 41.1±0.441.1\pm0.4 19 (8, 11)
    SS-IL Yes 41.1±1.641.1\pm1.6 31.6±0.531.6\pm0.5 47.0±1.247.0\pm1.2 38.3±0.438.3\pm0.4 48.1±1.748.1\pm1.7 47.5±0.747.5\pm0.7 19 (8, 11)
    ER-ACE No 53.1±1.0\mathbf{53.1\pm1.0} 35.6±1.0\mathbf{35.6\pm1.0} 58.0±0.7\mathbf{58.0\pm0.7} 42.6±0.742.6\pm0.7 61.9±0.961.9\pm0.9 52.2±0.752.2\pm0.7 17 (4, 7)
    ER-ACE Yes 52.6±0.952.6\pm0.9 35.1±0.835.1\pm0.8 56.4±1.056.4\pm1.0 43.4±1.643.4\pm1.6 61.7±0.961.7\pm0.9 53.7±1.153.7\pm1.1 17 (4, 7)
    ER-AML No 49.4±1.049.4\pm1.0 30.9±0.830.9\pm0.8 57.0±1.057.0\pm1.0 39.2±1.039.2\pm1.0 63.3±1.0\mathbf{63.3\pm1.0} 52.2±1.152.2\pm1.1 17 (4, 7)
    ER-AML Yes 50.4±1.350.4\pm1.3 36.4±1.4\mathbf{36.4\pm1.4} 56.8±1.056.8\pm1.0 47.7±0.7\mathbf{47.7\pm0.7} 62.0±0.962.0\pm0.9 55.7±1.3\mathbf{55.7\pm1.3} 17 (4, 7)
    GDUMB Yes 0±0.00\pm0.0 35.0±0.635.0\pm0.6 0±0.00\pm0.0 45.8±0.945.8\pm0.9 0±0.00\pm0.0 61.3±1.761.3\pm1.7 (43, 853) (11, 14)

    ER-ACE and ER-AML outperform standard Experience Replay by up to 36% relative gain, achieving top performance particularly in the small buffer regime (M=5M=5 and M=20M=20) while incurring no additional FLOPs over standard ER.

  7. Knowl 7 — Performance on Long Task Sequences: Split CIFAR-100 and Split MiniImageNet

    data/table

    Performance comparison on long task sequences: Split CIFAR-100 (20 disjoint tasks of 5 classes each, 32×3232 \times 32 images) and Split MiniImageNet (20 disjoint tasks of 5 classes each, 84×8484 \times 84 images) with memory buffer size M=100M = 100 per class. Results report the best configuration with or without data augmentation across 10 runs.

    Method Split CIFAR-100 Split MiniImageNet
    AAA Acc Train TFLOPs Mem (Mb) AAA Acc Train TFLOPs Mem (Mb)
    iid - 19.8±0.319.8\pm0.3 9 4 - 16.7±0.516.7\pm0.5 59 4
    iid++ - 28.3±0.328.3\pm0.3 17 4 - 25.0±0.825.0\pm0.8 118 4
    DER++ 23.3±0.523.3\pm0.5 15.1±0.415.1\pm0.4 25 36 21.7±0.621.7\pm0.6 12.9±0.312.9\pm0.3 176 217
    ER 24.2±0.624.2\pm0.6 19.8±0.419.8\pm0.4 17 35 26.2±0.826.2\pm0.8 18.2±0.518.2\pm0.5 118 216
    iCaRL 26.3±0.326.3\pm0.3 17.3±0.217.3\pm0.2 294 39 24.4±0.424.4\pm0.4 17.1±0.117.1\pm0.1 2097 220
    MIR 23.6±0.823.6\pm0.8 20.6±0.520.6\pm0.5 41 35 27.2±0.727.2\pm0.7 20.2±0.820.2\pm0.8 294 216
    SS-IL 31.5±0.531.5\pm0.5 25.0±0.325.0\pm0.3 19 39 29.7±0.629.7\pm0.6 23.5±0.5\mathbf{23.5\pm0.5} 137 220
    ER-ACE 32.7±0.5\mathbf{32.7\pm0.5} 25.8±0.4\mathbf{25.8\pm0.4} 17 35 30.2±0.6\mathbf{30.2\pm0.6} 22.7±0.622.7\pm0.6 118 216
    ER-AML 30.2±0.630.2\pm0.6 24.3±0.424.3\pm0.4 28 35 27.0±0.727.0\pm0.7 19.3±0.619.3\pm0.6 200 216

    ER-ACE achieves a ~35% relative gain in final accuracy over baseline ER on CIFAR-100 without increasing compute, matches SS-IL in accuracy while maintaining a lower computational and memory profile, and achieves near equal-compute iid++ performance on MiniImageNet.

  8. Knowl 8 — Blurry Task Boundaries Protocol and Continuous Distribution Shift Performance

    experimental setup

    In realistic online streams, task boundaries are rarely discrete. The blurry task boundary protocol models continuous distribution shifts by smoothly blending classes over time.

    For a stream lasting Tsteps=5000T_{steps} = 5000 mini-batches of size 10 on Split CIFAR-10, the unnormalized probability of observing class cc at step t∈{1,…,5000}t \in \{1, \dots, 5000\} is modeled as a Gaussian function:

    pc(t)∼N(μc−t,Nc4),where μc=(2c−1)Nc2p_c(t) \sim \mathcal{N}\left(\mu_c - t, \frac{N_c}{4}\right), \quad \text{where } \mu_c = \frac{(2c - 1)N_c}{2}

    and NcN_c is the total number of samples for class cc. Probabilities are normalized into a Categorical distribution across all classes at every step, yielding an average of 2 unique labels per incoming mini-batch.

    Methods relying on task identifiers (e.g., MIR, SS-IL) cannot execute in this setting. The final classification accuracy (averaged over 5 runs) on CIFAR-10 under blurry boundaries is:

    Method M=20M = 20 M=100M = 100
    ER 32.1±1.532.1 \pm 1.5 42.7±2.242.7 \pm 2.2
    DER++ 31.0±1.431.0 \pm 1.4 41.7±1.441.7 \pm 1.4
    ER-AML 45.6±1.2\mathbf{45.6 \pm 1.2} 55.2±1.1\mathbf{55.2 \pm 1.1}
    ER-ACE 44.5±0.544.5 \pm 0.5 50.2±1.150.2 \pm 1.1

    When varying the level of class overlap (from 1 to 5 average unique classes per mini-batch), ER-AML and ER-ACE consistently outperform ER and DER++ across all degrees of stream blurriness.

  9. Knowl 9 — Analytical and Empirical Evaluation of Separated Softmax in the Online Setting

    empirical result

    Separated Softmax (SS-IL) applies masked cross-entropy to both the incoming stream batch and the rehearsal buffer batch in isolation. In the online continual learning setting, this design exhibits two distinct properties:

    1. Failure to Learn the Current Task: Because both incoming and rehearsal losses mask out cross-task logits, the network never receives an objective that trains it to separate classes in the current task from classes in previous tasks. Consequently, during online streaming, SS-IL achieves below-random classification accuracy on the active task in a single-head setting until those classes enter the buffer in subsequent tasks. In contrast, ER-ACE leaves the rehearsal loss unmasked, successfully acquiring current-task discrimination while preserving past representations.

    2. Representation Drift Mitigation Beyond Class Imbalance: When evaluated on an artificially balanced stream (sampling 0, 10, 20, 30, 40 rehearsal points per batch across tasks 1 through 5 to eliminate class frequency imbalance entirely), SS-IL continues to substantially outperform ER on CIFAR-10:

    Method M=20M = 20 M=50M = 50 M=100M = 100
    ER 21.0±1.221.0 \pm 1.2 25.7±1.125.7 \pm 1.1 37.8±0.737.8 \pm 0.7
    SS-IL 30.3±1.030.3 \pm 1.0 34.6±0.834.6 \pm 0.8 39.1±0.639.1 \pm 0.6

    This demonstrates that the primary mechanism underlying logit-masking benefits in replay architectures is the mitigation of representation drift at distribution shifts rather than the correction of class imbalance.

  10. Knowl 10 — Mitigation of Buffer Overfitting and Orthogonal Integration with Dark Experience Replay

    empirical result

    ER-ACE reduces overfitting on buffered exemplars and provides orthogonal benefits when combined with Dark Experience Replay (DER++).

    1. Buffer Overfitting Reduction: Representation alignment between rehearsal samples xm∈Mx_m \in \mathcal{M} and unseen validation samples xv∈Vx_v \in \mathcal{V} of the same class is evaluated by computing the maximum cosine similarity:

    max⁡xv∈V,c(xv)=c(xm)fθ(xm)Tfθ(xv)∥fθ(xm)∥∥fθ(xv)∥\max_{x_v \in \mathcal{V}, c(x_v) = c(x_m)} \frac{f_\theta(x_m)^T f_\theta(x_v)}{\|f_\theta(x_m)\| \|f_\theta(x_v)\|}

    ER-ACE maintains significantly higher feature alignment between stored and held-out validation samples for earlier tasks (e.g., Tasks 1-3 on Split CIFAR-10) compared to ER, indicating superior generalization and reduced exemplar overfitting.

    1. Combination with DER++: Integrating ER-ACE's asymmetric masking into DER++ (DER++ACE) improves performance across all buffer sizes on CIFAR-10 and substantially reduces catastrophic forgetting:
    • At M=20M=20: Forgetting drops from ∼38%\sim 38\% (DER++) to ∼10%\sim 10\% (DER++ACE), while final accuracy increases from 31.8%31.8\% to ∼35.5%\sim 35.5\%.
    • At M=50M=50: Forgetting drops from ∼28%\sim 28\% to ∼12%\sim 12\%.
    • At M=100M=100: Forgetting drops from ∼22%\sim 22\% to ∼8%\sim 8\%.

Coverage note — None was omitted; the extracted knowls cover all core methodological, analytical, algorithmic, and empirical contributions of the paper.

References

  1. 1.Hongjoon Ahn, Jihwan Kwak, Subin Lim, Hyeonsu Bang, Hyojun Kim, and Taesup Moon. Ss-il: Separated softmax for incremental learning. arXiv preprint arXiv:2003.13947, 2020.
  2. 2.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. Memory aware synapses: Learning what (not) to forget. arXiv preprint arXiv:1711.09601, 2017.
  3. 3.Rahaf Aljundi, Klaas Kelchtermans, and Tinne Tuytelaars. Task-free continual learning. In CVPR 2019, 2018.
  4. 4.Rahaf Aljundi, Lucas Caccia, Eugene Belilovsky, Massimo Caccia, Laurent Charlin, and Tinne Tuytelaars. Online continual learning with maximally interfered retrieval. In Advances in Neural Information Processing (NeurIPS), 2019a.
  5. 5.Rahaf Aljundi, Min Lin, Baptiste Goujaud, and Yoshua Bengio. Gradient based sample selection for online continual learning. arXiv preprint arXiv:1903.08671, 2019b.
  6. 6.Zalán Borsos, Mojmír Mutnỳ, and Andreas Krause. Coresets via bilevel optimization for continual learning and streaming. arXiv preprint arXiv:2006.03875, 2020.
  7. 7.Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. Dark experience for general continual learning: a strong, simple baseline. arXiv preprint arXiv:2004.07211, 2020.
  8. 8.Massimo Caccia, Pau Rodriguez, Oleksiy Ostapenko, Fabrice Normandin, Min Lin, Lucas Caccia, Issam Laradji, Irina Rish, Alexandre Lacoste, David Vazquez, et al. Online fast adaptation and knowledge accumulation: a new approach to continual learning. arXiv preprint arXiv:2003.05856, 2020.
  9. 9.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. Efficient lifelong learning with a-gem. In ICLR 2019.
  10. 10.Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. Continual learning with tiny episodic memories. arXiv preprint arXiv:1902.10486, 2019.
  11. 11.Hung-Jen Chen, An-Chieh Cheng, Da-Cheng Juan, Wei Wei, and Min Sun. Mitigating forgetting in online continual learning via instance-aware parameterization. Advances in Neural Information Processing Systems, 33, 2020a.
  12. 12.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020b.
  13. 13.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. Continual learning: A comparative study on how to defy forgetting in classification tasks. arXiv preprint arXiv:1909.08383, 2019.
  14. 14.Sebastian Farquhar and Yarin Gal. Towards robust evaluations of continual learning. arXiv preprint arXiv:1805.09733, 2018.
  15. 15.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729–9738, 2020.
  16. 16.Xu He, Jakub Sygnowski, Alexandre Galashov, Andrei A. Rusu, Yee Whye Teh, and Razvan Pascanu. Task agnostic continual learning via meta learning. ArXiv, abs/1906.05201, 2019. URL https://arxiv.org/abs/1906.05201.
  17. 17.Elad Hoffer and Nir Ailon. Deep metric learning using triplet network. In International workshop on similarity-based pattern recognition, pp. 84–92. Springer, 2015.
  18. 18.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 831–839, 2019.
  19. 19.Xu Ji, Joao Henriques, Tinne Tuytelaars, and Andrea Vedaldi. Automatic recall machines: Internal replay, continual learning and the brain. arXiv preprint arXiv:2006.12323, 2020.
  20. 20.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 18661–18673. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/d89a66c7c80a29b1bdbab0f2a1a94af8-Paper.pdf.
  21. 21.Timothée Lesort, Andrei Stoian, and David Filliat. Regularization shortcomings for continual learning. arXiv preprint arXiv:1912.03049, 2019.
  22. 22.Timothée Lesort, Massimo Caccia, and Irina Rish. Understanding continual learning settings with data distribution drift analysis. arXiv preprint arXiv:2104.01678, 2021.
  23. 23.Zhizhong Li and Derek Hoiem. Learning without forgetting. In European Conference on Computer Vision, pp. 614–629. Springer, 2016.
  24. 24.David Lopez-Paz et al. Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems, pp. 6467–6476, 2017.
  25. 25.Zheda Mai, Ruiwen Li, Hyunwoo Kim, and Scott Sanner. Supervised contrastive replay: Revisiting the nearest class mean classifier in online class-incremental continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3589–3599, 2021.
  26. 26.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of learning and motivation, 24:109–165, 1989.
  27. 27.Fabrice Normandin, Florian Golemo, Oleksiy Ostapenko, Matthew Riemer, Pau Rodriguez, Julio Hurtado, Khimya Khetarpal, Timothée Lesort, Laurent Charlin, Irina Rish, and Massimo Caccia. Sequoia - towards a systematic organization of continual learning research. https://github.com/lebrice/Sequoia, 2021. URL https://github.com/lebrice/Sequoia.
  28. 28.Oleksiy Ostapenko, Pau Rodriguez, Massimo Caccia, and Laurent Charlin. Continual learning via local module composition. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/fe5e7cb609bdbe6d62449d61849c38b0-Abstract.html.
  29. 29.Ameya Prabhu, Philip HS Torr, and Puneet K Dokania. Gdumb: A simple approach that questions our progress in continual learning. In European Conference on Computer Vision, pp. 524–540. Springer, 2020.
  30. 30.Hang Qi, Matthew Brown, and David G Lowe. Low-shot learning with imprinted weights. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5822–5830, 2018.
  31. 31.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proc. CVPR, 2017.
  32. 32.David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P Lillicrap, and Greg Wayne. Experience replay for continual learning. arXiv preprint arXiv:1811.11682, 2018.
  33. 33.Joan Serrà, Dídac Surís, Marius Miron, and Alexandros Karatzoglou. Overcoming catastrophic forgetting with hard attention to the task. arXiv preprint arXiv:1801.01423, 2018.
  34. 34.Dongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner, Hyunwoo Kim, and Jongseong Jang. Online class-incremental continual learning with adversarial shapley value. arXiv e-prints, pp. arXiv–2009, 2020.
  35. 35.Dongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner, Hyunwoo Kim, and Jongseong Jang. Online class-incremental continual learning with adversarial shapley value. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 9630–9638, 2021.
  36. 36.Binh Tang and David S Matteson. Graph-based continual learning. arXiv preprint arXiv:2007.04813, 2020.
  37. 37.Gido M van de Ven and Andreas S Tolias. Three scenarios for continual learning. arXiv preprint arXiv:1904.07734, 2019. URL https://arxiv.org/abs/1904.07734.
  38. 38.Jeffrey S Vitter. Random sampling with a reservoir. ACM Transactions on Mathematical Software (TOMS), 11(1):37–57, 1985.
  39. 39.Johannes Von Oswald, Dominic Zhao, Seijin Kobayashi, Simon Schug, Massimo Caccia, Nicolas Zucchet, and Joao Sacramento. Learning where to learn: Gradient sparsity in meta and continual learning. Advances in Neural Information Processing Systems, 34, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/2a10665525774fa2501c2c8c4985ce61-Abstract.html.
  40. 40.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 374–382, 2019.
  41. 41.Chen Zeno, Itay Golan, Elad Hoffer, and Daniel Soudry. Task agnostic continual learning using online variational bayes. arXiv preprint arXiv:1803.10123, 2018.
  42. 42.Bowen Zhao, Xi Xiao, Guojun Gan, Bin Zhang, and Shutao Xia. Maintaining discrimination and fairness in class incremental learning. arXiv preprint arXiv:1911.07053, 2019.

Citation

MLA
Caccia, L., et al. “New Insights on Reducing Abrupt Representation Change in Online Continual Learning”. arXiv, 2021, http://arxiv.org/abs/2104.05025v3.
APA
Caccia, L., Aljundi, R., Asadi, N., Tuytelaars, T., Pineau, J., & Belilovsky, E. (2021). New Insights on Reducing Abrupt Representation Change in Online Continual Learning. arXiv. http://arxiv.org/abs/2104.05025v3
Chicago
Caccia, L., R. Aljundi, N. Asadi, T. Tuytelaars, J. Pineau, and E. Belilovsky. 2021. “New Insights on Reducing Abrupt Representation Change in Online Continual Learning”. arXiv. http://arxiv.org/abs/2104.05025v3.
Harvard
Caccia, L. et al. (2021) “New Insights on Reducing Abrupt Representation Change in Online Continual Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.05025v3.
Vancouver
1. Caccia L, Aljundi R, Asadi N, Tuytelaars T, Pineau J, Belilovsky E (2021) New Insights on Reducing Abrupt Representation Change in Online Continual Learning. arXiv

BibTeX

@article{caccia2021new,
  title = {New Insights on Reducing Abrupt Representation Change in Online Continual Learning},
  author = {Caccia, Lucas and Aljundi, Rahaf and Asadi, Nader and Tuytelaars, Tinne and Pineau, Joelle and Belilovsky, Eugene},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.05025v3},
  eprint = {2104.05025}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors