Experience Replay for Continual Learning
David RolnickArun AhujaJonathan SchwarzTimothy P. LillicrapGreg Wayne
Demonstrates that combining experience replay with behavioral cloning effectively prevents catastrophic forgetting in continual reinforcement learning without requiring task boundary labels, matching the performance of specialized task-aware methods across Atari and DeepMind Lab.
Artificial intelligence systems deployed in real-world environments face the ongoing challenge of continual learning, where new skills must be mastered sequentially without erasing previously acquired capabilities. In deep reinforcement learning, sequential training frequently suffers from catastrophic forgetting, a failure where learning a new task overwrites earlier knowledge. Conventional training setups often bypass this by training all tasks simultaneously on massive compute clusters, but real-world settings such as robotics make simultaneous data collection costly and unfeasible. Moreover, practical applications rarely provide clear task labels or defined task boundaries, leaving standard systems vulnerable to severe performance degradation.
The article evaluates whether experience replay, combined with behavioral cloning, can mitigate catastrophic forgetting in deep reinforcement learning when tasks are presented in sequence without task boundary metadata.
To demonstrate this, the authors introduced Continual Learning with Experience And Replay (CLEAR), an approach implemented within a scalable, distributed actor-critic framework. CLEAR blends on-policy updates from new experiences to maintain learning plasticity with off-policy replay updates to preserve stability. It applies behavioral cloning loss terms during replay to prevent policy and value estimates from drifting away from historical behavior. The authors validated the method using multi-task suites in DeepMind Lab and sequential Atari benchmarks, testing various buffer capacities with reservoir sampling, evaluating new-to-replay data ratios, and benchmarking results against standard sequential training as well as state-of-the-art methods like Elastic Weight Consolidation and Progress & Compress.
The analysis yielded several key findings. First, CLEAR virtually eliminated catastrophic forgetting across sequential 3D navigation tasks, matching the performance of simultaneous multi-task training upper bounds. Second, a 50-50 mix of novel and replay data provided an optimal balance, enabling rapid acquisition of newly introduced probe tasks without slowing down adaptation as the buffer filled. Third, the approach proved highly memory-efficient, maintaining strong retention and performance even when the replay buffer was constrained by reservoir sampling to hold as little as 1 in 200 total experiences. Finally, on standard sequential Atari benchmarks, CLEAR matched or outperformed more complex weight-consolidation techniques, achieving top cumulative scores across multiple tasks while operating completely agnostically to task boundaries.
These findings indicate that complex parameter-isolation and weight-consolidation techniques may be unnecessary for many continual learning applications. Instead, leveraging modern storage to maintain a modest replay buffer offers a simpler, more direct solution. Because CLEAR does not require explicit signals when a task changes, it can be deployed in complex, continuous environments where tasks evolve dynamically. This reduces engineering complexity, system fragility, and operational overhead in lifelong learning pipelines.
For engineering and research teams addressing continual reinforcement learning, CLEAR is recommended as an effective, practical first line of defense against catastrophic forgetting. In constrained environments, teams should deploy reservoir sampling with a 50-50 new-to-replay ratio. If higher performance gains are required, future development can explore combining replay mechanisms with orthogonal parameter-consolidation approaches or more sophisticated off-policy value estimators.
The conclusions are supported by consistent multi-seed empirical evaluations across diverse reinforcement learning environments. However, decision-makers should note certain boundary conditions. Replay methods inherently assume a shared and stable action space; if the fundamental mechanics or action definitions change between tasks, historical replay and behavioral cloning may lead to negative interference, requiring supplemental selective forgetting mechanisms.
- Paper: Prioritized Experience Replay, Tom Schaul et al. (2016). It introduces foundational experience replay mechanisms and prioritization strategies in deep reinforcement learning that the source adapts for continual learning settings.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). It establishes Elastic Weight Consolidation as a seminal baseline for mitigating catastrophic forgetting in sequential task learning across classification and reinforcement learning.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). It introduces Gradient Episodic Memory, providing essential context on using stored episodic memories to prevent forgetting during sequential gradient updates.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). It provides the foundational framework of knowledge preservation and distillation under task distribution shifts without retraining from scratch.
- Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). It details Synaptic Intelligence as a key regularization-based alternative to replay for preserving performance across sequential neural network training.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). It formulates bounded-memory exemplar retention and distillation for incremental learning, directly motivating replay buffer management in continual learning.
- Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). It demonstrates generative replay as a mechanism to alleviate catastrophic forgetting in sequential tasks without storing original data.
- Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). It builds on replay-based continual learning by proposing Averaged GEM to dramatically reduce memory and computational overhead in single-pass streaming settings.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). It provides a comprehensive survey and empirical benchmark comparing replay-based, regularization, and parameter-isolation methods for defying forgetting.
- Paper: Recurrent Experience Replay in Distributed Reinforcement Learning, Steven Kapturowski et al. (2019). It extends experience replay in deep reinforcement learning to large-scale distributed recurrent architectures across Atari and DMLab environments.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). It provides a comprehensive tutorial and taxonomy of offline reinforcement learning methods, contextualizing the challenges of learning from fixed replay datasets and behavioral cloning.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). It standardizes benchmarks and evaluation protocols for offline, replay-driven reinforcement learning across complex multi-task and continuous control domains.
