Built independently by an author, for readers. Read the story and support ChapterPal

keyword

reservoir sampling

Reservoir sampling is a family of randomized algorithms designed to select a representative sample of a fixed size from a data stream of unknown or unbounded length in a single pass. The algorithm initially fills a fixed-capacity buffer with the first elements of the stream and subsequently processes each new element by deciding whether to retain it and replace an existing element according to a decreasing probability based on the total number of items observed. This process guarantees that every item observed in the stream has an equal chance of being present in the sample at any point in time, using minimal memory and processing time. In machine learning and streaming data applications, such as continual learning, reservoir sampling is commonly employed to maintain bounded experience replay buffers that preserve an unbiased distribution of past training examples over non-stationary data streams without requiring prior knowledge of the dataset size.

2 items

Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Zhenyi Wang, Li Shen, Le Fang, Qiuling Suo, Tiehang Duan, Mingchen Gao

OrganizationsJD.comMetaUniversity at Buffalo

Why you should read this

Proposes a distributionally robust memory evolution framework that uses Wasserstein gradient flows to dynamically update replay buffers, preventing overfitting to stored samples and mitigating catastrophic forgetting in task-free continual learning.

Task-free continual learning (CL) aims to learn a non-stationary data stream without explicit task definitions and not forget previous knowledge. The widely adopted memory replay approach could gradually become less effective for long data streams, as the model may memorize the stored examples and overfit the memory buffer. Second, existing methods overlook the high uncertainty in the memory data distribution since there is a big gap between the memory data distribution and the distribution of all the previous data examples. To address these problems, for the first time, we propose a principled memory evolution framework to dynamically evolve the memory data distribution by making the memory buffer gradually harder to be memorized with distributionally robust optimization (DRO). We then derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). The proposed DRO is w.r.t the worst-case evolved memory data distribution, thus guarantees the model performance and learns significantly more robust features than existing memory-replay-based methods. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than existing task-free CL methods.

Added

2026-09-26

Experience Replay for Continual Learning

Experience Replay for Continual Learning

David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, Greg Wayne

OrganizationsGoogleUniversity of Pennsylvania

Why you should read this

Demonstrates that combining experience replay with behavioral cloning effectively prevents catastrophic forgetting in continual reinforcement learning without requiring task boundary labels, matching the performance of specialized task-aware methods across Atari and DeepMind Lab.

Continual learning is the problem of learning new tasks or knowledge while protecting old knowledge and ideally generalizing from old experience to learn new tasks faster. Neural networks trained by stochastic gradient descent often degrade on old tasks when trained successively on new tasks with different data distributions. This phenomenon, referred to as catastrophic forgetting, is considered a major hurdle to learning with non-stationary data or sequences of new tasks, and prevents networks from continually accumulating knowledge and skills. We examine this issue in the context of reinforcement learning, in a setting where an agent is exposed to tasks in a sequence. Unlike most other work, we do not provide an explicit indication to the model of task boundaries, which is the most general circumstance for a learning agent exposed to continuous experience. While various methods to counteract catastrophic forgetting have recently been proposed, we explore a straightforward, general, and seemingly overlooked solution - that of using experience replay buffers for all past events - with a mixture of on- and off-policy learning, leveraging behavioral cloning. We show that this strategy can still learn new tasks quickly yet can substantially reduce catastrophic forgetting in both Atari and DMLab domains, even matching the performance of methods that require task identities. When buffer storage is constrained, we confirm that a simple mechanism for randomly discarding data allows a limited size buffer to perform almost as well as an unbounded one.

Added

2026-09-18