Built independently by an author, for readers. Read the story and support ChapterPal

keyword

replay buffers

A replay buffer is a memory storage mechanism in machine learning that saves past experiences, data samples, or environment transitions encountered during training so they can be reused at later stages. Commonly utilized in reinforcement learning and continual learning frameworks, this buffer stores records such as state transitions, actions, rewards, and historical task data. By sampling and replaying these stored experiences alongside or in place of immediate observations, learning algorithms can break temporal correlations between consecutive data points, improve sample efficiency, and mitigate catastrophic forgetting, thereby enabling neural networks to retain previously acquired skills and knowledge when exposed to non-stationary data streams or sequences of new tasks.

2 items

Experience Replay for Continual Learning

Experience Replay for Continual Learning

David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, Greg Wayne

OrganizationsGoogleUniversity of Pennsylvania

Why you should read this

Demonstrates that combining experience replay with behavioral cloning effectively prevents catastrophic forgetting in continual reinforcement learning without requiring task boundary labels, matching the performance of specialized task-aware methods across Atari and DeepMind Lab.

Continual learning is the problem of learning new tasks or knowledge while protecting old knowledge and ideally generalizing from old experience to learn new tasks faster. Neural networks trained by stochastic gradient descent often degrade on old tasks when trained successively on new tasks with different data distributions. This phenomenon, referred to as catastrophic forgetting, is considered a major hurdle to learning with non-stationary data or sequences of new tasks, and prevents networks from continually accumulating knowledge and skills. We examine this issue in the context of reinforcement learning, in a setting where an agent is exposed to tasks in a sequence. Unlike most other work, we do not provide an explicit indication to the model of task boundaries, which is the most general circumstance for a learning agent exposed to continuous experience. While various methods to counteract catastrophic forgetting have recently been proposed, we explore a straightforward, general, and seemingly overlooked solution - that of using experience replay buffers for all past events - with a mixture of on- and off-policy learning, leveraging behavioral cloning. We show that this strategy can still learn new tasks quickly yet can substantially reduce catastrophic forgetting in both Atari and DMLab domains, even matching the performance of methods that require task identities. When buffer storage is constrained, we confirm that a simple mechanism for randomly discarding data allows a limited size buffer to perform almost as well as an unbounded one.

Added

2026-09-18