Built independently by an author, for readers. Read the story and support ChapterPal

keyword

non-stationary data stream

A non-stationary data stream is a continuous, potentially unbounded sequence of incoming data whose underlying probability distribution and statistical properties change over time. In contrast to stationary streams where the data-generating process remains constant, non-stationary streams experience shifts in feature distributions or target relationships, commonly referred to as concept drift. Because this dynamic behavior violates traditional independent and identically distributed assumptions, systems that process non-stationary data streams must continually adapt to new information, track evolving trends, and mitigate the risk of performance degradation or the forgetting of previously learned patterns.

1 item

Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Zhenyi Wang, Li Shen, Le Fang, Qiuling Suo, Tiehang Duan, Mingchen Gao

OrganizationsJD.comMetaUniversity at Buffalo

Why you should read this

Proposes a distributionally robust memory evolution framework that uses Wasserstein gradient flows to dynamically update replay buffers, preventing overfitting to stored samples and mitigating catastrophic forgetting in task-free continual learning.

Task-free continual learning (CL) aims to learn a non-stationary data stream without explicit task definitions and not forget previous knowledge. The widely adopted memory replay approach could gradually become less effective for long data streams, as the model may memorize the stored examples and overfit the memory buffer. Second, existing methods overlook the high uncertainty in the memory data distribution since there is a big gap between the memory data distribution and the distribution of all the previous data examples. To address these problems, for the first time, we propose a principled memory evolution framework to dynamically evolve the memory data distribution by making the memory buffer gradually harder to be memorized with distributionally robust optimization (DRO). We then derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). The proposed DRO is w.r.t the worst-case evolved memory data distribution, thus guarantees the model performance and learns significantly more robust features than existing memory-replay-based methods. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than existing task-free CL methods.

Added

2026-09-26