Experience Replay for Continual Learning

David RolnickArun AhujaJonathan SchwarzTimothy P. LillicrapGreg Wayne

article2018NeurIPS1,830 citations

Demonstrates that combining experience replay with behavioral cloning effectively prevents catastrophic forgetting in continual reinforcement learning without requiring task boundary labels, matching the performance of specialized task-aware methods across Atari and DeepMind Lab.

Listen

Artificial intelligence systems deployed in real-world environments face the ongoing challenge of continual learning, where new skills must be mastered sequentially without erasing previously acquired capabilities. In deep reinforcement learning, sequential training frequently suffers from catastrophic forgetting, a failure where learning a new task overwrites earlier knowledge. Conventional training setups often bypass this by training all tasks simultaneously on massive compute clusters, but real-world settings such as robotics make simultaneous data collection costly and unfeasible. Moreover, practical applications rarely provide clear task labels or defined task boundaries, leaving standard systems vulnerable to severe performance degradation.

The article evaluates whether experience replay, combined with behavioral cloning, can mitigate catastrophic forgetting in deep reinforcement learning when tasks are presented in sequence without task boundary metadata.

To demonstrate this, the authors introduced Continual Learning with Experience And Replay (CLEAR), an approach implemented within a scalable, distributed actor-critic framework. CLEAR blends on-policy updates from new experiences to maintain learning plasticity with off-policy replay updates to preserve stability. It applies behavioral cloning loss terms during replay to prevent policy and value estimates from drifting away from historical behavior. The authors validated the method using multi-task suites in DeepMind Lab and sequential Atari benchmarks, testing various buffer capacities with reservoir sampling, evaluating new-to-replay data ratios, and benchmarking results against standard sequential training as well as state-of-the-art methods like Elastic Weight Consolidation and Progress & Compress.

The analysis yielded several key findings. First, CLEAR virtually eliminated catastrophic forgetting across sequential 3D navigation tasks, matching the performance of simultaneous multi-task training upper bounds. Second, a 50-50 mix of novel and replay data provided an optimal balance, enabling rapid acquisition of newly introduced probe tasks without slowing down adaptation as the buffer filled. Third, the approach proved highly memory-efficient, maintaining strong retention and performance even when the replay buffer was constrained by reservoir sampling to hold as little as 1 in 200 total experiences. Finally, on standard sequential Atari benchmarks, CLEAR matched or outperformed more complex weight-consolidation techniques, achieving top cumulative scores across multiple tasks while operating completely agnostically to task boundaries.

These findings indicate that complex parameter-isolation and weight-consolidation techniques may be unnecessary for many continual learning applications. Instead, leveraging modern storage to maintain a modest replay buffer offers a simpler, more direct solution. Because CLEAR does not require explicit signals when a task changes, it can be deployed in complex, continuous environments where tasks evolve dynamically. This reduces engineering complexity, system fragility, and operational overhead in lifelong learning pipelines.

For engineering and research teams addressing continual reinforcement learning, CLEAR is recommended as an effective, practical first line of defense against catastrophic forgetting. In constrained environments, teams should deploy reservoir sampling with a 50-50 new-to-replay ratio. If higher performance gains are required, future development can explore combining replay mechanisms with orthogonal parameter-consolidation approaches or more sophisticated off-policy value estimators.

The conclusions are supported by consistent multi-seed empirical evaluations across diverse reinforcement learning environments. However, decision-makers should note certain boundary conditions. Replay methods inherently assume a shared and stable action space; if the fundamental mechanics or action definitions change between tasks, historical replay and behavioral cloning may lead to negative interference, requiring supplemental selective forgetting mechanisms.

  • Paper: Prioritized Experience Replay, Tom Schaul et al. (2016). It introduces foundational experience replay mechanisms and prioritization strategies in deep reinforcement learning that the source adapts for continual learning settings.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). It establishes Elastic Weight Consolidation as a seminal baseline for mitigating catastrophic forgetting in sequential task learning across classification and reinforcement learning.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). It introduces Gradient Episodic Memory, providing essential context on using stored episodic memories to prevent forgetting during sequential gradient updates.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). It provides the foundational framework of knowledge preservation and distillation under task distribution shifts without retraining from scratch.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). It details Synaptic Intelligence as a key regularization-based alternative to replay for preserving performance across sequential neural network training.
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). It formulates bounded-memory exemplar retention and distillation for incremental learning, directly motivating replay buffer management in continual learning.
  • Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). It demonstrates generative replay as a mechanism to alleviate catastrophic forgetting in sequential tasks without storing original data.
Cover for Experience Replay for Continual Learning

Abstract

Continual learning is the problem of learning new tasks or knowledge while protecting old knowledge and ideally generalizing from old experience to learn new tasks faster. Neural networks trained by stochastic gradient descent often degrade on old tasks when trained successively on new tasks with different data distributions. This phenomenon, referred to as catastrophic forgetting, is considered a major hurdle to learning with non-stationary data or sequences of new tasks, and prevents networks from continually accumulating knowledge and skills. We examine this issue in the context of reinforcement learning, in a setting where an agent is exposed to tasks in a sequence. Unlike most other work, we do not provide an explicit indication to the model of task boundaries, which is the most general circumstance for a learning agent exposed to continuous experience. While various methods to counteract catastrophic forgetting have recently been proposed, we explore a straightforward, general, and seemingly overlooked solution - that of using experience replay buffers for all past events - with a mixture of on- and off-policy learning, leveraging behavioral cloning. We show that this strategy can still learn new tasks quickly yet can substantially reduce catastrophic forgetting in both Atari and DMLab domains, even matching the performance of methods that require task identities. When buffer storage is constrained, we confirm that a simple mechanism for randomly discarding data allows a limited size buffer to perform almost as well as an unbounded one.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 The CLEAR Method
  • 4 Results
  • 4.1 Catastrophic forgetting vs. interference
  • 4.2 Stability
  • 4.3 Plasticity
  • 4.4 Balance of on- and off-policy learning
  • 4.5 Limited-size buffers
  • 4.6 Comparison to P&C and EWC
  • 5 Discussion
  • References
  • A Implementation details
  • A.1 Distributed setup
  • A.2 Network
  • A.3 Buffers
  • A.4 Training
  • A.5 Evaluation
  • A.6 Experiments
  • B Figures replotted according to cumulative sum

Knowls

  1. Knowl 1 — CLEAR Framework for Continual Reinforcement Learning

    model/method

    Continual Learning with Experience And Replay (CLEAR) is an actor-critic deep reinforcement learning framework designed to mitigate catastrophic forgetting in sequential multi-task environments without requiring task labels or knowledge of task boundaries.

    CLEAR integrates two complementary mechanisms:

    1. Mixed On-Policy and Off-Policy Learning: Batches fed to the learner network consist of a mixture (typically 50–50 or 75–25) of novel on-policy experiences generated by actors executing the current policy πθ\pi_\theta and historical off-policy replay experiences sampled from a replay buffer. Novel experiences preserve plasticity (the ability to rapidly acquire new skills), while replayed experiences provide stability (retention of past skills). Off-policy distribution shifts are corrected using truncated importance sampling via the V-Trace algorithm.
    2. Behavioral Cloning on Replay Data: For replayed experiences only, auxiliary behavioral cloning loss terms penalize drift between the learner's current policy/value outputs and the historical policy/value outputs stored at the time the replayed trajectory was collected. This regularizes the network against policy drift on older tasks during new task acquisition.
  2. Knowl 2 — CLEAR Optimization Objectives and Behavioral Cloning Losses

    equation

    Let θ\theta denote the parameters of the neural network, πθ(a∣hs)\pi_\theta(a|h_s) the policy over actions aa given hidden state hsh_s at time step ss, Vθ(hs)V_\theta(h_s) the predicted state value, μ(a∣hs)\mu(a|h_s) the behavior policy that generated the experience trajectory, rsr_s the reward, and γ∈[0,1)\gamma \in [0, 1) the discount factor.

    The nn-step V-Trace value target vsv_s is defined as: vs:=V(hs)+∑t=ss+n−1γt−s(∏i=st−1ci)δtVv_s := V(h_s) + \sum_{t=s}^{s+n-1} \gamma^{t-s} \left( \prod_{i=s}^{t-1} c_i \right) \delta_t V where the temporal difference is: δtV:=ρt(rt+γV(ht+1)−V(ht))\delta_t V := \rho_t \left( r_t + \gamma V(h_{t+1}) - V(h_t) \right) with truncated importance sampling weights ci:=min⁡(cˉ,πθ(ai∣hi)μ(ai∣hi))c_i := \min\left(\bar{c}, \frac{\pi_\theta(a_i|h_i)}{\mu(a_i|h_i)}\right) and ρt:=min⁡(ρˉ,πθ(at∣ht)μ(at∣ht))\rho_t := \min\left(\bar{\rho}, \frac{\pi_\theta(a_t|h_t)}{\mu(a_t|h_t)}\right) for truncation constants cˉ\bar{c} and ρˉ\bar{\rho}.

    The standard actor-critic loss terms applied to both new and replay unrolls are: Lpolicy-gradient:=−ρslog⁡πθ(as∣hs)(rs+γvs+1−Vθ(hs))L_{\text{policy-gradient}} := -\rho_s \log \pi_\theta(a_s|h_s) \left( r_s + \gamma v_{s+1} - V_\theta(h_s) \right) Lvalue:=(Vθ(hs)−vs)2L_{\text{value}} := \left( V_\theta(h_s) - v_s \right)^2 Lentropy:=∑aπθ(a∣hs)log⁡πθ(a∣hs)L_{\text{entropy}} := \sum_a \pi_\theta(a|h_s) \log \pi_\theta(a|h_s)

    For replay experiences only, two behavioral cloning losses are added to prevent output drift: Lpolicy-cloning:=DKL(μ(⋅∣hs)∥πθ(⋅∣hs))=∑aμ(a∣hs)log⁡μ(a∣hs)πθ(a∣hs)L_{\text{policy-cloning}} := D_{\text{KL}}(\mu(\cdot|h_s) \parallel \pi_\theta(\cdot|h_s)) = \sum_a \mu(a|h_s) \log \frac{\mu(a|h_s)}{\pi_\theta(a|h_s)} Lvalue-cloning:=∥Vθ(hs)−Vreplay(hs)∥22L_{\text{value-cloning}} := \|V_\theta(h_s) - V_{\text{replay}}(h_s)\|_2^2

    Evaluating DKL(μ∥πθ)D_{\text{KL}}(\mu \parallel \pi_\theta) rather than DKL(πθ∥μ)D_{\text{KL}}(\pi_\theta \parallel \mu) guarantees that πθ(a∣hs)\pi_\theta(a|h_s) retains non-zero support across all actions selected by the historical policy μ\mu.

    In standard implementations, these loss functions are weighted with coefficients wpg=1.0w_{\text{pg}} = 1.0, wval=0.5w_{\text{val}} = 0.5, went≈0.005w_{\text{ent}} \approx 0.005, wpol-clone=0.01w_{\text{pol-clone}} = 0.01, and wval-clone=0.005w_{\text{val-clone}} = 0.005.

  3. Knowl 3 — Distributed Actor-Learner Architecture and Reservoir Replay in CLEAR

    algorithm

    CLEAR operates in an Importance Weighted Actor-Learner Architecture (IMPALA) distributed setting. The replay buffer is partitioned equally across all parallel CPU actors. When actor buffers reach capacity, reservoir sampling is used to maintain a uniform random sample of all historical experience without knowing task boundaries.

    Input: Actor pool AA, Learner network θ\theta, buffer capacity per actor MM, replay mixing ratio α∈[0,1]\alpha \in [0, 1]
    Initialize actor replay buffers Bk←∅B_k \leftarrow \emptyset, reservoir thresholds τk←0\tau_k \leftarrow 0 for each actor k∈Ak \in A
    Actor Process kk:
      while training is active do
        Update local network weights from learner θ\theta
        Collect trajectory unroll U=(st,at,rt,μ(at∣st),Vμ(st))t=0n−1U = (s_t, a_t, r_t, \mu(a_t|s_t), V_\mu(s_t))_{t=0}^{n-1}
        Draw random priority key u∼Uniform(0,1)u \sim \text{Uniform}(0, 1)
        if ∣Bk∣<M|B_k| < M then
          Insert (U,u)(U, u) into BkB_k
        else if u>τku > \tau_k then
          Replace element with lowest key in BkB_k with (U,u)(U, u)
          Update threshold τk←min⁡(U′,u′)∈Bku′\tau_k \leftarrow \min_{(U', u') \in B_k} u'
        end if
        Sample a replay unroll UreplayU_{\text{replay}} uniformly from BkB_k
        Enqueue candidate pair (U,Ureplay)(U, U_{\text{replay}}) to learner queue QQ
        Wait until the learner reads the pair
      end while
    Learner Process:
      while training is active do
        Dequeue batch of pairs from QQ
        Construct batch BB by selecting UU with probability 1−α1-\alpha and UreplayU_{\text{replay}} with probability α\alpha from each pair
        Compute Lpolicy-gradientL_{\text{policy-gradient}}, LvalueL_{\text{value}}, and LentropyL_{\text{entropy}} on all items in BB
        Compute Lpolicy-cloningL_{\text{policy-cloning}} and Lvalue-cloningL_{\text{value-cloning}} exclusively on replay items in BB
        Compute total weighted loss LtotalL_{\text{total}}
        Update learner weights θ←θ−η∇θLtotal\theta \leftarrow \theta - \eta \nabla_\theta L_{\text{total}}
        Broadcast updated weights θ\theta asynchronously to actors
      end while
  4. Knowl 4 — Dissociation Between Catastrophic Forgetting and Multi-Task Interference

    empirical result

    In deep multi-task reinforcement learning on DeepMind Lab (DMLab) navigation environments (explore_object_locations_small, rooms_collect_good_objects, and rooms_keys_doors_puzzle), performance degradation in sequential task training is shown to be driven almost entirely by catastrophic forgetting rather than destructive task interference.

    When comparing three training paradigms holding total environmental frames per task constant:

    1. Separate Training: Individual networks trained independently on each task.
    2. Simultaneous Training: A single network trained concurrently on a balanced mixture of all tasks.
    3. Sequential Training: A single network trained on tasks in repeating sequential cycles without replay.

    Simultaneous multi-task training matches or slightly outperforms separate single-task training due to modest constructive interference (shared visual processing and exploratory representations). In contrast, sequential training exhibits immediate, severe performance collapse on previously learned tasks as soon as training shifts to another task.

  5. Knowl 5 — Cumulative Performance Comparison on Sequential DMLab Tasks

    data/table

    The final cumulative average reward, defined at time tt by 1t∑s<trs\frac{1}{t}\sum_{s < t} r_s, was evaluated across three DMLab tasks presented in cyclically repeating sequences for a total of 900×106900 \times 10^6 environment frames. CLEAR variants are compared against separate, simultaneous, and naive sequential baselines.

    Training Method explore_object... rooms_collect... rooms_keys...
    Separate 29.24 8.79 19.91
    Simultaneous 32.35 8.81 20.56
    Sequential (no CLEAR) 17.99 5.01 10.87
    CLEAR (50-50 new-replay) 31.40 8.00 18.13
    CLEAR w/o behavioral cloning 28.66 7.79 16.63
    CLEAR, 75-25 new-replay 30.28 7.83 17.86
    CLEAR, 100% replay 31.09 7.48 13.39
    CLEAR, buffer 5M 30.33 8.00 18.07
    CLEAR, buffer 50M 30.82 7.99 18.21

    Separate and Simultaneous settings serve as upper bounds where catastrophic forgetting is absent. Standard sequential training loses over 40–50% of cumulative reward due to forgetting. CLEAR (50-50) virtually eliminates catastrophic forgetting, performing within 5–10% of the simultaneous upper bound. Ablating behavioral cloning reduces cumulative scores across all tasks, confirming the regularizing utility of cloning historical outputs.

  6. Knowl 6 — Preservation of Learning Plasticity with Experience Mixing

    empirical result

    A critical concern in replay-based continual learning is the loss of plasticity: as the replay buffer grows, novel task samples make up an increasingly small fraction of the buffer, which could impede rapid acquisition of new skills.

    To test plasticity, a novel probe task (natlab_varying_map_randomized) was introduced at varying positions (Position 1, Position 4, Position 7) within a repeating multi-task sequence of three DMLab tasks:

    • In standard CLEAR (50–50 mixture of novel and replay data), learning speed and asymptotic score on the probe task remain invariant to the insertion point in the sequence and the accumulation of past experiences in the replay buffer.
    • When training exclusively on replay data (100% replay), probe task learning speed deteriorates markedly at later positions (Position 7) because the probe task constitutes a diminishing fraction of the total buffer.

    This demonstrates that dedicated on-policy learning from novel unrolls is essential to maintain plasticity in continual RL.

  7. Knowl 7 — Trade-off Between Plasticity and Stability Across Replay Ratios

    empirical result

    Varying the proportion of novel to replayed unrolls in CLEAR highlights the trade-off between stability (resisting forgetting) and plasticity (learning speed):

    • 100% Novel (0% Replay): Standard training; exhibits rapid task learning but complete catastrophic forgetting when task switching occurs.
    • 75% Novel / 25% Replay: Substantially reduces catastrophic forgetting while preserving rapid acquisition of new skills, though minor performance dips persist upon task switches.
    • 50% Novel / 50% Replay: Provides an optimal balance, virtually eliminating catastrophic forgetting on historical tasks without degrading the acquisition rate of novel tasks.
    • 0% Novel / 100% Replay: Prevents catastrophic forgetting and can improve performance via off-policy learning, but exhibits noticeably slower learning curves and reduced early performance when introduced to challenging tasks (e.g., rooms_keys_doors_puzzle).
  8. Knowl 8 — Continual Learning Robustness Under Constrained Replay Buffer Sizes

    empirical result

    CLEAR retains its ability to prevent catastrophic forgetting even when replay buffer storage is severely constrained relative to the total training volume.

    In a 900×106900 \times 10^6 environment frame sequential DMLab training run, buffer capacities of 450×106450 \times 10^6 frames (storing 50% of all generated data), 50×10650 \times 10^6 frames (storing 1 in 18 frames), and 5×1065 \times 10^6 frames (storing 1 in 180 frames) were evaluated using reservoir sampling:

    • The 50M50\text{M} buffer achieves cumulative performance (30.8230.82, 7.997.99, 18.2118.21) identical to the large 450M450\text{M} buffer (30.8230.82, 8.008.00, 18.1318.13).
    • The 5M5\text{M} buffer exhibits only a minor decrease in stability (30.3330.33, 8.008.00, 18.0718.07), attributed to slight overfitting to the small set of stored trajectory unrolls repeatedly replayed by the learner.
  9. Knowl 9 — Comparison of CLEAR with EWC and Progress & Compress on Atari Benchmarks

    data/table

    CLEAR (75–25 new-replay ratio) was evaluated against Elastic Weight Consolidation (EWC), Progress & Compress (P&C), and an on-policy baseline on a sequential benchmark of six Atari 2600 games (space_invaders, krull, beam_rider, hero, star_gunner, ms_pacman) trained under identical architectures and frame budgets.

    Method space_invaders krull beam_rider hero star_gunner ms_pacman
    Baseline 346.47 5512.36 952.72 2737.32 1065.20 753.83
    CLEAR 426.72 8845.12 585.05 8106.22 991.74 1222.82
    EWC 549.22 5352.92 432.78 4499.35 704.20 639.75
    PC 455.34 7025.40 743.38 3433.10 1261.86 924.52

    CLEAR achieves cumulative performance comparable to or exceeding P&C and substantially outperforms EWC and the baseline across the majority of games (e.g., scoring 8845.128845.12 on Krull and 8106.228106.22 on Hero vs 7025.407025.40 and 3433.103433.10 for P&C). Unlike EWC and P&C, which explicitly require task IDs and notifications of task boundaries to compute parameter importance matrices or freeze subnetworks, CLEAR operates completely agnostic to task boundaries and identities.

  10. Knowl 10 — Limitations and Failure Modes of CLEAR

    limitation

    CLEAR exhibits specific structural and operational limitations in continual learning settings:

    1. Memory Storage Overhead: While reservoir sampling mitigates memory scaling, storing replay unrolls requires higher memory capacity than synaptic consolidation methods (e.g., EWC or P&C) that store only network parameters and Fisher information diagonals.
    2. Vulnerability to Action Space Shifts: If the action space or the semantic meaning of actions changes between sequential tasks, behavioral cloning and off-policy replay regularize the current policy toward obsolete or conflicting historical action distributions, causing negative transfer.
    3. Off-Policy Correction Limits: The importance sampling truncation in V-Trace (ar{\rho}, \bar{c}) is designed for moderate distribution shifts. For extreme policy divergences over long temporal horizons, pure V-Trace may suffer from variance or bias unless combined with action-value methods (e.g., Retrace or Q-learning).

Coverage note — None was omitted; all key theoretical formulations, algorithmic implementations, experimental comparisons across DMLab and Atari benchmarks, and ablation analyses from the paper are fully covered.

References

  1. 1.Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al. DeepMind Lab. Preprint arXiv:1612.03801, 2016.
  2. 2.Marcus K Benna and Stefano Fusi. Computational principles of synaptic memory consolidation. Nature neuroscience, 19(12):1697, 2016.
  3. 3.Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al. IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures. In ICML, 2018.
  4. 4.Robert M French. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135, 1999.
  5. 5.Tommaso Furlanello, Zachary C Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. Born again neural networks. ICML, 2018.
  6. 6.Stephen Grossberg. How does a brain build a cognitive code? In Studies of mind and brain, pages 1–52. Springer, 1982.
  7. 7.Shixiang Gu, Tim Lillicrap, Richard E Turner, Zoubin Ghahramani, Bernhard Schölkopf, and Sergey Levine. Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning. In NeurIPS, 2017.
  8. 8.Tyler L Hayes, Nathan D Cahill, and Christopher Kanan. Memory efficient experience replay for streaming learning. Preprint arXiv:1809.05922, 2018.
  9. 9.Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, et al. Deep Q-learning from demonstrations. In AAAI, 2018.
  10. 10.David Isele and Akansel Cosgun. Selective experience replay for lifelong learning. In AAAI, 2018.
  11. 11.Christos Kaplanis, Murray Shanahan, and Claudia Clopath. Continual reinforcement learning with complex synapses. In ICML, 2018.
  12. 12.Christos Kaplanis, Murray Shanahan, and Claudia Clopath. Policy consolidation for continual reinforcement learning. In ICML, 2019.
  13. 13.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. In Proceedings of the National Academy of Sciences, 2017.
  14. 14.Pascal Leimer, Michael Herzog, and Walter Senn. Synaptic weight decay with selective consolidation enables fast learning without catastrophic forgetting. BioRxiv, page 613265, 2019.
  15. 15.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.
  16. 16.Long-Ji Lin. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine learning, 8(3-4):293–321, 1992.
  17. 17.David Lopez-Paz et al. Gradient episodic memory for continual learning. In NeurIPS, pages 6467–6476, 2017.
  18. 18.James L McClelland. Complementary learning systems in the brain: A connectionist approach to explicit and implicit cognition and memory. Annals of the New York Academy of Sciences, 843(1):153–169, 1998.
  19. 19.Kieran Milan, Joel Veness, James Kirkpatrick, Michael Bowling, Anna Koop, and Demis Hassabis. The forget-me-not process. In NeurIPS, 2016.
  20. 20.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529, 2015.
  21. 21.Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare. Safe and efficient off-policy reinforcement learning. In NeurIPS, 2016.
  22. 22.Brendan O’Donoghue, Remi Munos, Koray Kavukcuoglu, and Volodymyr Mnih. Combining policy gradient and Q-learning. In ICLR, 2017.
  23. 23.Junhyuk Oh, Yijie Guo, Satinder Singh, and Honglak Lee. Self-imitation learning. ICML, 2018.
  24. 24.German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. Continual lifelong learning with neural networks: A review. Neural Networks, 2019.
  25. 25.Mark B Ring. Child: A first step towards continual learning. Machine Learning, 28(1):77–104, 1997.
  26. 26.Amanda Rios and Laurent Itti. Closed-loop GAN for continual learning. Preprint arXiv:1811.0114, 2018.
  27. 27.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. Preprint arXiv:1606.04671, 2016.
  28. 28.Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. Prioritized experience replay. In ICLR, 2016.
  29. 29.Jonathan Schwarz, Jelena Luketina, Wojciech M Czarnecki, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. Progress & compress: A scalable framework for continual learning. In ICML, 2018.
  30. 30.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. Continual learning with deep generative replay. In NeurIPS, 2017.
  31. 31.Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas. Sample efficient actor-critic with experience replay. In ICLR, 2017.
  32. 32.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In ICML, 2017.

Citation

MLA
Rolnick, D., et al. “Experience Replay for Continual Learning”. arXiv, 2018, http://arxiv.org/abs/1811.11682v2.
APA
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. P., & Wayne, G. (2018). Experience Replay for Continual Learning. arXiv. http://arxiv.org/abs/1811.11682v2
Chicago
Rolnick, D., A. Ahuja, J. Schwarz, T. P. Lillicrap, and G. Wayne. 2018. “Experience Replay for Continual Learning”. arXiv. http://arxiv.org/abs/1811.11682v2.
Harvard
Rolnick, D. et al. (2018) “Experience Replay for Continual Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1811.11682v2.
Vancouver
1. Rolnick D, Ahuja A, Schwarz J, Lillicrap TP, Wayne G (2018) Experience Replay for Continual Learning. arXiv

BibTeX

@article{rolnick2018experience,
  title = {Experience Replay for Continual Learning},
  author = {Rolnick, David and Ahuja, Arun and Schwarz, Jonathan and Lillicrap, Timothy P. and Wayne, Greg},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1811.11682v2},
  eprint = {1811.11682}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission