Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Zhenyi WangLi ShenLe FangQiuling SuoTiehang DuanMingchen Gao

article2022ICML51 citations

Proposes a distributionally robust memory evolution framework that uses Wasserstein gradient flows to dynamically update replay buffers, preventing overfitting to stored samples and mitigating catastrophic forgetting in task-free continual learning.

Listen

Real-world machine learning systems frequently encounter non-stationary streams of data where distributions shift over time without explicit notifications or defined task boundaries. In these task-free continuous learning environments, models must continuously assimilate new information while retaining previously acquired knowledge. Standard approaches rely on memory replay, which preserves a small subset of historical data in a storage buffer. However, models repeatedly train on these limited static samples, leading to severe overfitting and catastrophic forgetting of older data. Furthermore, small memory buffers introduce significant distribution uncertainty because they cannot adequately represent the full historical data stream.

The article aims to solve this limitation by introducing a principled memory evolution framework based on distributionally robust optimization. This approach actively and dynamically evolves the stored memory data distribution to make the buffer progressively harder to memorize and better representative of the entire data stream.

To achieve this, the article reinterprets the optimization problem as continuous dynamics, utilizing Wasserstein gradient flows to evolve the data distribution while standard gradient updates adjust the model parameters. The authors develop three practical evolution mechanisms: a diffusion-based stochastic method (Langevin Dynamics), a deterministic kernel-based approach (Stein Variational Gradient Descent), and a physics-inspired Hamiltonian method. The framework was evaluated across standard image benchmarks (CIFAR-10, CIFAR-100, and MiniImageNet) by integrating the proposed evolution techniques into existing replay baselines such as standard Experience Replay, Maximally Interfering Retrieval, and Gradient-based Memory Editing.

The evaluation yielded several key findings. First, integrating memory evolution consistently boosted test accuracy across all baselines, achieving absolute accuracy improvements of 3.6% to 4.5% on CIFAR-10, 0.9% to 1.6% on CIFAR-100, and 1.4% to 2.8% on MiniImageNet. Second, the framework maintained performance advantages across various memory buffer constraints (e.g., buffers ranging from 2,000 to 10,000 samples). Third, optimizing against worst-case distribution shifts inherently conferred substantial adversarial robustness; under strong attacks, baseline models degraded to near-zero accuracy, whereas the proposed method maintained noticeable resilience, outperforming naive baselines by 4% to 12% under projected gradient descent perturbations.

These findings demonstrate that actively diversifying and hardening memory distributions is far more effective than replaying static historical examples. The framework effectively narrows the representational gap between stored samples and past data streams, mitigating memory overfitting without requiring complex model architecture expansions. Importantly, because it is modular, the technique can be directly integrated into existing continuous learning pipelines, simultaneously improving model longevity and security against adversarial threats.

Organizations deploying continuous learning models on non-stationary data streams should adopt dynamic memory evolution strategies instead of static buffer replays. When implementing these methods, practitioners must weigh the trade-off between model robustness and compute efficiency, as evolving the memory buffer increases training runtime by approximately 3.4 to 4.1 times compared to naive replay baselines. Future development should focus on optimizing this computational overhead and exploring domain-specific geometry constraints.

Confidence in the reported improvements is high across the evaluated image classification benchmarks and hyperparameters. However, practitioners should exercise caution regarding boundary conditions: the empirical validation is restricted to continuous vision data domains, and while adaptations for discrete data (such as language embeddings) are theoretically feasible, they were not experimentally validated in the article.

arXiv: 2207.07256
Cover for Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Abstract

Task-free continual learning (CL) aims to learn a non-stationary data stream without explicit task definitions and not forget previous knowledge. The widely adopted memory replay approach could gradually become less effective for long data streams, as the model may memorize the stored examples and overfit the memory buffer. Second, existing methods overlook the high uncertainty in the memory data distribution since there is a big gap between the memory data distribution and the distribution of all the previous data examples. To address these problems, for the first time, we propose a principled memory evolution framework to dynamically evolve the memory data distribution by making the memory buffer gradually harder to be memorized with distributionally robust optimization (DRO). We then derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). The proposed DRO is w.r.t the worst-case evolved memory data distribution, thus guarantees the model performance and learns significantly more robust features than existing memory-replay-based methods. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than existing task-free CL methods.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem Setup
  • 3.2. DRO for Task-free CL
  • 3.3. Task-free DRO: A Continuous Dynamics View
  • 3.4. Training Algorithm for Dynamic DRO
  • 4. Experiments
  • 4.1. Experiment Setup
  • 4.2. Comparison to Continual Learning
  • 4.3. Robustness to Adversarial Perturbations
  • 4.4. Ablation Study
  • 5. Conclusion
  • 6. Acknowledgement
  • References
  • A. Baselines
  • B. Lagrangian Duality Derivation
  • C. More Experiments Results
  • Hyperparameter Sensitivity Analysis
  • D. Foundations of Calculus of Variations
  • E. Derivations of the memory evolution methods

Knowls

  1. Knowl 1 — Task-Free Distributionally Robust Optimization Objective for Memory Replay

    model/method

    In task-free continual learning (CL), a model f(x,θ)f(x, \theta) with parameters θ∈Θ\theta \in \Theta observes a stream of mini-batches (xk,yk)(x_k, y_k) without explicit task indicators or boundary markers while maintaining an episodic memory buffer M\mathcal{M} sampled from an unknown stationary data distribution. Standard memory replay minimizes empirical loss over raw memory distribution μ0\mu_0: min⁡θ∈Θ[L(θ,xk,yk)+Ex∼μ0L(θ,x,y)]\min_{\theta \in \Theta} [L(\theta, x_k, y_k) + \mathbb{E}_{x \sim \mu_0} L(\theta, x, y)], which leads to memory overfitting and cannot account for distributional uncertainty in small buffers.

    To address this, task-free Distributionally Robust Optimization (DRO) models the memory distribution μ\mu as lying within an ambiguity set P={μ:D(μ∥π)≤D(μ0∥π)≤ϵ}\mathcal{P} = \{\mu : D(\mu \parallel \pi) \le D(\mu_0 \parallel \pi) \le \epsilon\}, where D(⋅∥⋅)D(\cdot \parallel \cdot) is the Kullback-Leibler (KL) divergence and π\pi is the worst-case evolved memory distribution that maximizes loss, bounded by threshold ϵ\epsilon. To prevent the evolved data distribution from diverging excessively from μ0\mu_0, an additional constraint enforces directional gradient alignment: Ex∼μ,x′∼μ0[∇θL(θ,x,y)⋅∇θL(θ,x′,y)]≥λ\mathbb{E}_{x \sim \mu, x' \sim \mu_0} [\nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y)] \ge \lambda.

    Applying Lagrangian duality to the constrained supremum converts the task-free DRO formulation into an unconstrained saddle-point optimization problem: min⁡θ∈Θsup⁡μ[Eμ[L(θ,x,y)]−γD(μ∥π)+βEx∼μ,x′∼μ0[∇θL(θ,x,y)⋅∇θL(θ,x′,y)]]\min_{\theta \in \Theta} \sup_{\mu} \left[ \mathbb{E}_{\mu} [L(\theta, x, y)] - \gamma D(\mu \parallel \pi) + \beta \mathbb{E}_{x \sim \mu, x' \sim \mu_0} [\nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y)] \right] where γ>0\gamma > 0 and β>0\beta > 0 are regularization hyperparameters (with γ=1\gamma = 1 under gradient flow formulations). This objective forces the memory buffer to evolve toward hard-to-memorize distributions that narrow the gap to the ideal multi-task stationary data distribution.

  2. Knowl 2 — Dynamic DRO via Coupled Wasserstein and Euclidean Gradient Flows

    model/method

    To solve the task-free distributionally robust optimization (DRO) saddle-point problem over an infinite-dimensional probability measure space, the system is cast as a continuous gradient flow system termed Dynamic DRO. Let P2(Rd)\mathcal{P}_2(\mathbb{R}^d) denote the Wasserstein space of probability measures on Rd\mathbb{R}^d with finite second moments. An energy functional F(μ)\mathcal{F}(\mu) over probability measure μ\mu is defined as: F(μ)=V(μ)+D(μ∥π)\mathcal{F}(\mu) = \mathcal{V}(\mu) + D(\mu \parallel \pi) V(μ)=−Eμ[L(θ,x,y)]−βEx∼μ,x′∼μ0[∇θL(θ,x,y)⋅∇θL(θ,x′,y)]\mathcal{V}(\mu) = -\mathbb{E}_\mu [L(\theta, x, y)] - \beta \mathbb{E}_{x \sim \mu, x' \sim \mu_0} [\nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y)] where L(θ,x,y)L(\theta, x, y) is the task loss function, μ0\mu_0 is the raw memory probability measure, π\pi is the target worst-case probability density π(x)∝e−U(x,θ)\pi(x) \propto e^{-U(x, \theta)}, and β\beta controls gradient dot product regularization.

    The potential function U(x,θ)U(x, \theta) corresponding to the first variation δVδμ\frac{\delta \mathcal{V}}{\delta \mu} is given by: U(x,θ)=−L(θ,x,y)−β∇θL(θ,x,y)⋅∇θL(θ,x′,y)U(x, \theta) = -L(\theta, x, y) - \beta \nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y)

    The Dynamic DRO optimizes the inner supremum and outer minimization simultaneously via a coupled gradient flow system: {∂tμt=div(μt∇δFδμ(μt))dθdt=−∇θEμt[L(θ,x,y)]\begin{cases} \partial_t \mu_t = \text{div}\left( \mu_t \nabla \frac{\delta \mathcal{F}}{\delta \mu}(\mu_t) \right) \\ \frac{d\theta}{dt} = -\nabla_\theta \mathbb{E}_{\mu_t} [L(\theta, x, y)] \end{cases} Here, ∂tμt=div(μt∇δFδμ(μt))\partial_t \mu_t = \text{div}(\mu_t \nabla \frac{\delta \mathcal{F}}{\delta \mu}(\mu_t)) represents a Wasserstein Gradient Flow (WGF) driving the memory probability measure μt\mu_t along the steepest descent curve of F(μ)\mathcal{F}(\mu) in Wasserstein space, while θ\theta evolves according to standard gradient flow in Euclidean parameter space.

  3. Knowl 3 — Memory Evolution via Langevin Dynamics (WGF-LD)

    model/method

    Under the Dynamic DRO framework, when the Wasserstein gradient flow ∂tμt=div(μt∇δFδμ(μt))\partial_t \mu_t = \text{div}(\mu_t \nabla \frac{\delta \mathcal{F}}{\delta \mu}(\mu_t)) is directly implemented in continuous Euclidean space, it defines a diffusion process described by the Langevin stochastic differential equation (SDE): dXt=−∇XU(Xt,θ)dt+2dWtdX_t = -\nabla_X U(X_t, \theta) dt + \sqrt{2} dW_t where Xt∼μtX_t \sim \mu_t represents the evolved memory data random variable at continuous time tt, WtW_t is standard Brownian motion in Rd\mathbb{R}^d, and U(x,θ)=−L(θ,x,y)−β∇θL(θ,x,y)⋅∇θL(θ,x′,y)U(x, \theta) = -L(\theta, x, y) - \beta \nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y).

    Discretizing this continuous SDE with step size (evolution rate) α>0\alpha > 0 yields the particle update rule for each memory sample xtix^i_t (i∈{1,…,N}i \in \{1, \dots, N\}): xt+1i=xti−α∇xU(xti,θ)+2αξtx^i_{t+1} = x^i_t - \alpha \nabla_x U(x^i_t, \theta) + \sqrt{2\alpha} \xi_t where ξt∼N(0,I)\xi_t \sim \mathcal{N}(0, I) is standard Gaussian noise. The drift term −α∇xU(xti,θ)-\alpha \nabla_x U(x^i_t, \theta) pushes the memory samples toward worst-case configurations that maximize classification difficulty, while the Brownian diffusion term 2αξt\sqrt{2\alpha} \xi_t acts as a random force to diversify the evolved memory distribution.

  4. Knowl 4 — Deterministic Memory Evolution via Stein Variational Gradient Descent (WGF-SVGD)

    model/method

    To perform deterministic memory distribution evolution without adding stochastic noise, the Wasserstein gradient flow is mapped into a Reproducing Kernel Hilbert Space (RKHS) H\mathcal{H} associated with a positive-definite kernel k(x,x′)k(x, x'). The kernelized Wasserstein gradient flow is governed by: ∂tμt=div(μtKμt∇δFδμ(μt))\partial_t \mu_t = \text{div}\left( \mu_t \mathcal{K}_{\mu_t} \nabla \frac{\delta \mathcal{F}}{\delta \mu}(\mu_t) \right) where Kμf(x)=∫k(x,x′)f(x′)dμ(x′)\mathcal{K}_{\mu} f(x) = \int k(x, x') f(x') d\mu(x').

    Discretizing the corresponding deterministic ODE dXdt=−[Kμ∇δFδμ(μt)](X)\frac{dX}{dt} = -[\mathcal{K}_\mu \nabla \frac{\delta \mathcal{F}}{\delta \mu}(\mu_t)](X) over a batch of NN memory particles {xti}i=1N\{x_t^i\}_{i=1}^N produces the WGF-SVGD update equation: xt+1i=xti−αN∑j=1N[k(xti,xtj)∇xtjU(xtj,θ)+∇xtjk(xti,xtj)]x_{t+1}^i = x_t^i - \frac{\alpha}{N} \sum_{j=1}^N \left[ k(x_t^i, x_t^j) \nabla_{x_t^j} U(x_t^j, \theta) + \nabla_{x_t^j} k(x_t^i, x_t^j) \right] where α\alpha is the evolution rate, and U(x,θ)=−L(θ,x,y)−β∇θL(θ,x,y)⋅∇θL(θ,x′,y)U(x, \theta) = -L(\theta, x, y) - \beta \nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y) is the potential function. A Gaussian RBF kernel k(xi,xj)=exp⁡(−∥xi−xj∥22σ2)k(x_i, x_j) = \exp\left(-\frac{\|x_i - x_j\|^2}{2\sigma^2}\right) is used.

    In this update, the smoothed gradient term ∑jk(xti,xtj)∇xtjU(xtj,θ)\sum_j k(x_t^i, x_t^j) \nabla_{x_t^j} U(x_t^j, \theta) aggregates gradient information across memory points to drive samples toward the worst-case distribution π\pi, while the repulsive term ∑j∇xtjk(xti,xtj)\sum_j \nabla_{x_t^j} k(x_t^i, x_t^j) prevents particles from collapsing into a single mode.

  5. Knowl 5 — General Memory Evolution via Hamiltonian Dynamics (WGF-HMC)

    model/method

    Continuous Markov samplers targeting π(x)∝e−U(x,θ)\pi(x) \propto e^{-U(x, \theta)} can be expressed as a generalized Wasserstein gradient flow: ∂tμt=div(μt(D+Q)∇δFδμ(μt))\partial_t \mu_t = \text{div}\left( \mu_t (D + Q) \nabla \frac{\delta \mathcal{F}}{\delta \mu}(\mu_t) \right) where DD is a positive semidefinite diffusion matrix and QQ is a skew-symmetric curl matrix. Setting D=(000C),Q=(0−II0)D = \begin{pmatrix} 0 & 0 \\ 0 & C \end{pmatrix}, \quad Q = \begin{pmatrix} 0 & -I \\ I & 0 \end{pmatrix} with friction parameter CC and identity matrix II corresponds to continuous Hamiltonian dynamics with friction.

    Discretizing this system introduces an auxiliary momentum variable vtv_t with momentum friction coefficient τ\tau and step size α\alpha, yielding the WGF-HMC memory evolution update: {xt+1=xt+vtvt+1=(1−τ)vt−α∇xU(xt,θ)+2ταξt\begin{cases} x_{t+1} = x_t + v_t \\ v_{t+1} = (1 - \tau) v_t - \alpha \nabla_x U(x_t, \theta) + \sqrt{2\tau\alpha} \xi_t \end{cases} where ξt∼N(0,I)\xi_t \sim \mathcal{N}(0, I) is standard Gaussian noise and U(x,θ)=−L(θ,x,y)−β∇θL(θ,x,y)⋅∇θL(θ,x′,y)U(x, \theta) = -L(\theta, x, y) - \beta \nabla_\theta L(\theta, x, y) \cdot \nabla_\theta L(\theta, x', y). This formulation allows incorporating geometry constraints and momentum to traverse complex loss landscapes during memory evolution.

  6. Knowl 6 — Distributionally Robust Memory Evolution Algorithm for Task-Free Continual Learning

    algorithm

    The task-free continual learning procedure alternates between evolving a mini-batch of memory samples for TT steps via WGF and updating model parameters θ\theta.

    Input: Initial model parameters θ\theta, learning rate η\eta, evolution step size α\alpha, number of inner evolution steps TT, regularization weight β\beta, memory buffer M\mathcal{M}, data stream of mini-batches {(xk,yk)}k=1K\{(x_k, y_k)\}_{k=1}^K
    Output: Updated model parameters θ\theta
    for k=1k = 1 to KK do
        Receive incoming mini-batch (xk,yk)(x_k, y_k)
        Sample replay mini-batch (x,y)∼M(x, y) \sim \mathcal{M}
        Initialize (x0,y)=(x,y)(x_0, y) = (x, y)
        for t=0t = 0 to T−1T-1 do
            Compute potential gradient ∇xU(xt,θ)=−∇xL(θ,xt,y)−β∇x[∇θL(θ,xt,y)⋅∇θL(θ,x0,y)]\nabla_x U(x_t, \theta) = -\nabla_x L(\theta, x_t, y) - \beta \nabla_x [\nabla_\theta L(\theta, x_t, y) \cdot \nabla_\theta L(\theta, x_0, y)]
            Update particles (xt+1,y)(x_{t+1}, y) using WGF-LD, WGF-SVGD, or WGF-HMC
        end for
        Set evolved memory batch (x~,y)=(xT,y)(\tilde{x}, y) = (x_T, y)
        Update model parameters: θ←θ−η∇θ[L(θ,x~,y)+L(θ,xk,yk)]\theta \leftarrow \theta - \eta \nabla_\theta [L(\theta, \tilde{x}, y) + L(\theta, x_k, y_k)]
        Update memory buffer: M←reservoir_sampling(M,(xk,yk))\mathcal{M} \leftarrow \text{reservoir\_sampling}(\mathcal{M}, (x_k, y_k))
    end for
    return θ\theta

    Default hyperparameters used in experiments: ResNet-18 backbone, T=5T = 5 evolution steps, β=0.003\beta = 0.003, momentum τ=0.1\tau = 0.1, evolution rate α=0.01\alpha = 0.01 (CIFAR-10), α=0.05\alpha = 0.05 (CIFAR-100), α=0.001\alpha = 0.001 (MiniImageNet). Raw memory data in M\mathcal{M} is kept un-overwritten across stream steps for stability, with evolved copies generated per iteration.

  7. Knowl 7 — Continual Learning Classification Accuracy under Task-Free Memory Evolution

    empirical result

    Experiments evaluated on split CIFAR-10 (5 disjoint tasks), CIFAR-100 (20 disjoint tasks), and MiniImageNet (20 disjoint tasks) using a ResNet-18 model demonstrate that integrating WGF memory evolution (WGF-LD, WGF-SVGD, WGF-HMC) into baseline replay methods (Experience Replay (ER), Maximally Interfering Retrieval (MIR), GMED, and ER with data augmentation (ERaug)) consistently improves average test accuracy (mean ±\pm standard deviation over 10 runs):

    Algorithm CIFAR-10 CIFAR-100 MiniImagenet
    fine-tuning 18.9 ±\pm 0.1 3.1 ±\pm 0.2 2.9 ±\pm 0.5
    A-GEM 19.0 ±\pm 0.3 2.4 ±\pm 0.2 3.0 ±\pm 0.4
    GSS-Greedy 29.9 ±\pm 1.5 19.5 ±\pm 1.3 17.4 ±\pm 0.9
    ER 33.3 ±\pm 2.8 20.1 ±\pm 1.2 24.8 ±\pm 1.0
    ER + WGF-LD 37.6 ±\pm 1.5 21.5 ±\pm 1.3 27.3 ±\pm 1.0
    ER + WGF-SVGD 36.5 ±\pm 1.4 21.3 ±\pm 1.5 27.6 ±\pm 1.3
    ER + WGF-HMC 37.8 ±\pm 1.3 21.2 ±\pm 1.4 27.2 ±\pm 1.1
    MIR 34.4 ±\pm 2.5 20.0 ±\pm 1.7 25.3 ±\pm 1.7
    MIR + WGF-LD 38.2 ±\pm 1.2 21.6 ±\pm 1.2 26.9 ±\pm 1.0
    MIR + WGF-SVGD 37.0 ±\pm 1.4 21.2 ±\pm 1.5 27.4 ±\pm 1.2
    MIR + WGF-HMC 37.9 ±\pm 1.5 21.3 ±\pm 1.4 27.1 ±\pm 1.3
    GMED (ER) 34.8 ±\pm 2.2 20.9 ±\pm 1.6 27.3 ±\pm 1.8
    GMED + WGF-LD 38.4 ±\pm 1.6 21.7 ±\pm 1.7 28.3 ±\pm 1.9
    GMED + WGF-SVGD 37.6 ±\pm 1.7 21.8 ±\pm 1.5 28.7 ±\pm 1.5
    GMED + WGF-HMC 37.8 ±\pm 1.2 21.5 ±\pm 1.9 28.4 ±\pm 1.3
    ERaug + ER 46.3 ±\pm 2.7 18.3 ±\pm 1.9 30.8 ±\pm 2.2
    ERaug + WGF-LD 47.6 ±\pm 2.4 19.8 ±\pm 2.2 31.9 ±\pm 1.8
    ERaug + WGF-SVGD 47.9 ±\pm 2.5 19.9 ±\pm 2.3 32.2 ±\pm 1.5
    ERaug + WGF-HMC 47.8 ±\pm 2.6 20.3 ±\pm 2.1 31.7 ±\pm 2.0
    iid online 60.3 ±\pm 1.4 18.7 ±\pm 1.2 17.7 ±\pm 1.5
    iid offline 78.7 ±\pm 1.1 44.9 ±\pm 1.5 39.8 ±\pm 1.4

    Combining ER with WGF yields improvements of up to 4.5% on CIFAR-10, 1.4% on CIFAR-100, and 2.8% on MiniImageNet over standard ER. For MIR and GMED, WGF integration similarly yields consistent gains of 1.4% to 3.8%.

  8. Knowl 8 — Adversarial Robustness Improvements under Task-Free DRO Memory Evolution

    empirical result

    Because task-free DRO optimizes model parameters against the worst-case evolved memory data distribution, the resulting network exhibits increased robustness against adversarial input perturbations without explicit adversarial training during rehearsal.

    Under a Carlini & Wagner ℓ2\ell_2 norm attack, models trained with baseline replay methods degrade to 0.0% accuracy on CIFAR-100 and MiniImageNet, whereas ER combined with WGF retains non-zero classification performance:

    Algorithm CIFAR-10 CIFAR-100 Mini-Imagenet
    ER 2.0 ±\pm 0.1 0.0 0.0
    GMED 2.1 ±\pm 0.1 0.0 0.0
    ER + WGF-LD 8.0 ±\pm 0.2 3.0 ±\pm 0.2 3.1 ±\pm 0.1
    ER + WGF-SVGD 4.2 ±\pm 0.1 0.0 2.2 ±\pm 0.2
    ER + WGF-HMC 8.2 ±\pm 0.3 2.5 ±\pm 0.2 3.0 ±\pm 0.1

    Under 20-step PGD ℓ∞\ell_\infty attacks with perturbation magnitudes ϵ∈[1/255,10/255]\epsilon \in [1/255, 10/255] and step size 2/2552/255, ER+WGF-HMC and ER+WGF-LD maintain a 4% to 12% absolute accuracy advantage over standard ER across perturbation levels on CIFAR-100 and MiniImageNet. Stochastic evolution methods (WGF-LD and WGF-HMC) achieve higher adversarial robustness than deterministic WGF-SVGD because the noise term facilitates broader exploration of the input space to create harder adversarial samples.

  9. Knowl 9 — Ablation on Memory Buffer Size and Inner Evolution Steps

    empirical result

    Ablation experiments evaluate the impact of varying memory buffer sizes (NN) and inner evolution steps (TT) on continual learning performance:

    1. Memory Buffer Size Variation (ResNet-18 accuracy, mean ±\pm std):
    CIFAR-100 Mini-ImageNet
    Memory Size 2000 3000 5000 3000 5000 10000
    ER 11.2 ±\pm 1.0 15.0 ±\pm 0.9 20.1 ±\pm 1.2 13.4 ±\pm 1.4 17.9 ±\pm 1.6 24.8 ±\pm 0.9
    ER + WGF-LD 12.9 ±\pm 1.2 17.0 ±\pm 1.1 21.5 ±\pm 1.3 16.2 ±\pm 1.2 20.8 ±\pm 1.2 27.3 ±\pm 1.0
    ER + WGF-SVGD 12.3 ±\pm 1.1 16.0 ±\pm 1.2 21.3 ±\pm 1.5 15.7 ±\pm 1.2 21.3 ±\pm 1.0 27.6 ±\pm 1.3
    ER + WGF-HMC 12.7 ±\pm 1.0 17.2 ±\pm 1.0 21.2 ±\pm 1.4 15.9 ±\pm 1.5 20.6 ±\pm 1.4 27.2 ±\pm 1.1
    MIR 11.6 ±\pm 0.8 15.6 ±\pm 1.0 20.0 ±\pm 1.7 12.6 ±\pm 1.5 17.4 ±\pm 1.2 25.3 ±\pm 1.7
    MIR + WGF-LD 13.1 ±\pm 0.9 17.3 ±\pm 1.2 21.6 ±\pm 1.2 15.5 ±\pm 1.4 20.5 ±\pm 1.1 26.9 ±\pm 1.0
    MIR + WGF-SVGD 12.7 ±\pm 1.0 16.5 ±\pm 1.3 21.2 ±\pm 1.5 15.3 ±\pm 1.2 20.7 ±\pm 1.6 27.4 ±\pm 1.2
    MIR + WGF-HMC 13.2 ±\pm 1.2 17.5 ±\pm 1.1 21.3 ±\pm 1.4 15.8 ±\pm 1.7 20.3 ±\pm 1.5 27.1 ±\pm 1.3

    WGF variants maintain a consistent performance advantage over raw ER and MIR across all buffer sizes, with larger relative gains at smaller buffer capacities where distributional uncertainty is highest.

    1. Number of Inner Evolution Steps (TT) on Mini-ImageNet:
    • T=3T=3: ER+WGF-LD = 27.0±0.927.0 \pm 0.9, ER+WGF-SVGD = 27.2±1.227.2 \pm 1.2, ER+WGF-HMC = 27.1±1.327.1 \pm 1.3.
    • T=5T=5: ER+WGF-LD = 27.3±1.027.3 \pm 1.0, ER+WGF-SVGD = 27.6±1.327.6 \pm 1.3, ER+WGF-HMC = 27.2±1.127.2 \pm 1.1.
    • T=7T=7: ER+WGF-LD = 27.5±1.427.5 \pm 1.4, ER+WGF-SVGD = 27.2±1.227.2 \pm 1.2, ER+WGF-HMC = 27.6±1.027.6 \pm 1.0. Performance increases marginally with larger TT; T=5T=5 provides an optimal trade-off between sample hardness exploration and computational overhead.
  10. Knowl 10 — Computational Overhead of Dynamic DRO Memory Evolution

    empirical result

    The inner loop optimization of Dynamic DRO requires computing input space gradients and kernel/diffusion interactions for TT steps per mini-batch. The relative training time of ER combined with WGF methods compared to baseline Experience Replay (ER, normalized to 1.0) on standard benchmarks is:

    Algorithm Relative Training Time
    ER 1.0
    ER + WGF-LD 3.4
    ER + WGF-SVGD 4.1
    ER + WGF-HMC 3.5

    WGF-LD and WGF-HMC incur approximately 3.4×3.4\times to 3.5×3.5\times the computation of standard ER due to T=5T=5 gradient updates per rehearsal batch. WGF-SVGD incurs 4.1×4.1\times computation due to the additional O(N2)\mathcal{O}(N^2) pairwise kernel computations and kernel gradient calculations.

Coverage note — None was omitted; all contributed models, optimization formulations, gradient flow derivations, algorithms, benchmark evaluations, adversarial analyses, ablations, and runtime overhead measurements are covered.

References

  1. 1.Aljundi, R., Belilovsky, E., Tuytelaars, T., Charlin, L., Caccia, M., Lin, M., and Page-Caccia, L. Online continual learning with maximal interfered retrieval. Advances in Neural Information Processing Systems 32, pp. 11849–11860, 2019a.
  2. 2.Aljundi, R., Kelchtermans, K., and Tuytelaars, T. Task-free continual learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019b.
  3. 3.Aljundi, R., Lin, M., Goujaud, B., and Bengio, Y. Gradient based sample selection for online continual learning. In Advances in Neural Information Processing Systems 30, 2019c.
  4. 4.Ambrosio, L., Gigli, N., and Savare, G. Gradient flows: In metric spaces and in the space of probability measures. (Lectures in Mathematics. ETH), 2008.
  5. 5.Bass, R. F. Stochastic Processes. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2011. doi: 10.1017/CBO9780511997044.
  6. 6.Boyd, S. and Vandenberghe, L. Convex optimization. Cambridge university press, 2004.
  7. 7.Carlini, N. and Wagner, D. Towards evaluating the robustness of neural networks. 2017 IEEE Symposium on Security and Privacy (SP), 2017.
  8. 8.Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M. Efficient lifelong learning with a-gem. Proceedings of the International Conference on Learning Representations, 2019a.
  9. 9.Chaudhry, A., Rohrbach, M., Elhoseiny, M., Ajanthan, T., Dokania, P. K., Torr, P. H. S., and Ranzato, M. Continual learning with tiny episodic memories. https://arxiv.org/abs/1902.10486, 2019b.
  10. 10.Chen, T., Fox, E., and Guestrin, C. Stochastic gradient hamiltonian monte carlo. In Proceedings of the 31st International Conference on Machine Learning, 2014.
  11. 11.Chewi, S., Le Gouic, T., Lu, C., Maunu, T., and Rigollet, P. Svgd as a kernelized wasserstein gradient flow of the chi-squared divergence. In Advances in Neural Information Processing Systems, 2020.
  12. 12.Chrysakis, A. and Moens, M.-F. Online continual learning from imbalanced data. Proceedings of the 37th International Conference on Machine Learning, 119:1952–1961, 2020.
  13. 13.He, X., Sygnowski, J., Galashov, A., Rusu, A. A., Teh, Y. W., and Pascanu, R. Task agnostic continual learning via meta learning. https://arxiv.org/abs/1906.05201, 2019.
  14. 14.Jin, X., Sadhu, A., Du, J., and Ren, X. Gradient-based editing of memory examples for online task-free continual learning. Advances in Neural Information Processing Systems, 2021.
  15. 15.Jordan, R., Kinderlehrer, D., , and Otto., F. The variational formulation of the fokker–planck equation. SIAM Journal on Mathematical Analysis, 1998.
  16. 16.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 2017.
  17. 17.Krizhevsky, A. Learning multiple layers of features from tiny images. Technical report, 2009.
  18. 18.Lee, S., Ha, J., Zhang, D., and Kim, G. A neural dirichlet process mixture model for task-free continual learning. In Proceedings of the 17th International Conference on Machine Learning, 2020.
  19. 19.Liu, C., Zhuo, J., and Zhu, J. Understanding mcmc dynamics as flows on the wasserstein spac. Proceedings of the International Conference on Machine Learning, 2019.
  20. 20.Liu, Q. Stein variational gradient descent as gradient flow. Advances in Neural Information Processing Systems, 2017.
  21. 21.Liu, Q. and Wang, D. Stein variational gradient descent: A general purpose bayesian inference algorithm. Advances in Neural Information Processing Systems, 2016.
  22. 22.Lopez-Paz, D. and Ranzato, M. Gradient episodic memory for continual learning. Advances in Neural Information Processing Systems, 2017.
  23. 23.Ma, Y.-A., Chen, T., and Fox, E. B. A complete recipe for stochastic gradient mcmc. Advances in Neural Information Processing Systems, 2015.
  24. 24.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. Towards deep learning models resistant to adversarial attacks. Proceedings of the International Conference on Learning Representations, 2018.
  25. 25.Miersemann, E. Calculus of Variations, Lecture Notes. Leipzig University, 2012.
  26. 26.Nguyen, C. V., Li, Y., Bui, T. D., and Turner, R. E. Variational continual learning. Proceedings of the International Conference on Learning Representations, 2018.
  27. 27.Pham, Q., Liu, C., Sahoo, D., and HOI, S. Contextual transformation networks for online continual learning. Proceedings of the International Conference on Learning Representations, 2021.
  28. 28.Rahimian, H. and Mehrotra, S. Distributionally robust optimization: A review. 2019.
  29. 29.Riemer, M., Cases, I., Ajemian, R., Liu, M., Rish, I., Tu, Y., and Tesauro, G. Learning to learn without forgetting by maximizing transfer and minimizing interference. International Conference on Learning Representations, 2019.
  30. 30.Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. International Conference on Learning Representations, 2020.
  31. 31.Saha, G., Garg, I., and Roy, K. Gradient projection memory for continual learning. Proceedings of the International Conference on Learning Representations, 2021.
  32. 32.Vinyals, O., Blundell, C., Lillicrap, T., kavukcuoglu, k., and Wierstra, D. Matching networks for one shot learning. 29, 2016.
  33. 33.von Oswald, J., Henning, C., Sacramento, J., and Grewe, B. F. Continual learning with hypernetworks. https://arxiv.org/abs/1906.00695, 2019.
  34. 34.Wang, Z., Duan, T., Fang, L., Suo, Q., and Gao, M. Meta learning on a sequence of imbalanced domains with difficulty awareness. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  35. 35.Wang, Z., Shen, L., Duan, T., Zhan, D., Fang, L., and Gao, M. Learning to learn and remember super long multidomain task sequence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  36. 36.Welling, M. and Teh, Y. W. Bayesian learning via stochastic gradient langevin dynamics. Proceedings of the International Conference on Machine Learning, 2011.
  37. 37.Xu, Z., Dan, C., Khim, J., and Ravikumar, P. Class-weighted classification: Trade-offs and robust approaches. Proceedings of the International Conference on Machine Learning, 2020.
  38. 38.Zenke, F., Poole, B., and Ganguli, S. Continual learning through synaptic intelligence. https://arxiv.org/abs/1703.04200, 2017.
  39. 39.Zeno, C., Golan, I., Hoffer, E., and Soudry, D. Task agnostic continual learning using online variational bayes. https://arxiv.org/abs/1803.10123, 2019.
  40. 40.Zhai, R., Dan, C., Kolter, J. Z., and Ravikumar, P. Doro: Distributional and outlier robust optimization. Proceedings of the International Conference on Machine Learning, 2021.

Citation

MLA
Wang, Z., et al. “Improving Task-free Continual Learning by Distributionally Robust Memory Evolution”. International Conference on Machine Learning, vol. 162, 2022, pp. 22985–98, https://proceedings.mlr.press/v162/wang22v.html.
APA
Wang, Z., Shen, L., Fang, L., Suo, Q., Duan, T., & Gao, M. (2022). Improving Task-free Continual Learning by Distributionally Robust Memory Evolution. International Conference on Machine Learning, 162, 22985–22998. https://proceedings.mlr.press/v162/wang22v.html
Chicago
Wang, Z., L. Shen, L. Fang, Q. Suo, T. Duan, and M. Gao. 2022. “Improving Task-free Continual Learning by Distributionally Robust Memory Evolution”. International Conference on Machine Learning 162: 22985–98. https://proceedings.mlr.press/v162/wang22v.html.
Harvard
Wang, Z. et al. (2022) “Improving Task-free Continual Learning by Distributionally Robust Memory Evolution”, International Conference on Machine Learning. PMLR, pp. 22985–22998. Available at: https://proceedings.mlr.press/v162/wang22v.html.
Vancouver
1. Wang Z, Shen L, Fang L, Suo Q, Duan T, Gao M (2022) Improving Task-free Continual Learning by Distributionally Robust Memory Evolution. In: International Conference on Machine Learning. PMLR, pp 22985–22998

BibTeX

@InProceedings{pmlr-v162-wang22v,
  title = 	 {Improving Task-free Continual Learning by Distributionally Robust Memory Evolution},
  author =       {Wang, Zhenyi and Shen, Li and Fang, Le and Suo, Qiuling and Duan, Tiehang and Gao, Mingchen},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {22985--22998},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/wang22v/wang22v.pdf},
  url = 	 {https://proceedings.mlr.press/v162/wang22v.html},
  abstract = 	 {Task-free continual learning (CL) aims to learn a non-stationary data stream without explicit task definitions and not forget previous knowledge. The widely adopted memory replay approach could gradually become less effective for long data streams, as the model may memorize the stored examples and overfit the memory buffer. Second, existing methods overlook the high uncertainty in the memory data distribution since there is a big gap between the memory data distribution and the distribution of all the previous data examples. To address these problems, for the first time, we propose a principled memory evolution framework to dynamically evolve the memory data distribution by making the memory buffer gradually harder to be memorized with distributionally robust optimization (DRO). We then derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). The proposed DRO is w.r.t the worst-case evolved memory data distribution, thus guarantees the model performance and learns significantly more robust features than existing memory-replay-based methods. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than existing task-free CL methods.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/