The Primacy Bias in Deep Reinforcement Learning

Evgenii NikishinMax SchwarzerPierluca D'OroPierre-Luc BaconAaron C. Courville

article2022ICML264 citations

Identifies how deep reinforcement learning agents overfit to early experience, proposing a simple periodic parameter-resetting mechanism that overcomes this primacy bias and consistently improves performance across discrete and continuous benchmarks.

Listen

Modern deep reinforcement learning algorithms frequently suffer from a failure mode where artificial agents overfit to their earliest interactions with an environment and subsequently fail to learn from new, higher-quality data. This dynamic, termed the primacy bias, creates a negative feedback loop: an overfitted agent executes poor actions, collects low-quality experience, and becomes permanently unable to reach optimal performance even after extensive additional training. As deep reinforcement learning is increasingly deployed in data-intensive and complex environments, overcoming this learning bottleneck is critical for improving data efficiency and final agent capabilities.

The article evaluates the causes and mechanics of the primacy bias across deep reinforcement learning algorithms and demonstrates a simple, lightweight remediation strategy. To test this, the authors evaluated standard algorithms across diverse benchmarks, including 26 discrete-action Atari games and 19 continuous-control robotic tasks from the DeepMind Control Suite, spanning both raw pixel inputs and dense physical states across multiple random seeds.

The proposed solution, termed resetting, periodically re-initializes the parameters of the agent's neural networks (or their final layers) while strictly preserving all historical interactions stored in the data replay buffer. The analysis reveals three primary findings. First, periodic resets consistently improve agent performance without adding computational overhead, increasing the aggregate interquartile mean score on Atari benchmarks by over 25% and lifting continuous-control baseline performance by 30% to 34%. Second, resets enable agents to thrive under aggressive training regimes: at high data reuse rates, adding resets doubled performance where standard algorithms collapsed. Third, traditional regularization techniques such as weight decay (L2) and dropout fail to prevent primacy bias, whereas resetting effectively resolves optimization failures like value estimation collapse and divergence.

These findings indicate that the primary barrier in many underperforming reinforcement learning systems is an optimization failure within the neural network rather than inadequate data collection. Retaining the non-parametric memory of the environment in the replay buffer while clearing out entrenched, overfitted network weights allows the agent to quickly recover and leverage accumulated experience. This permits organizations to achieve higher sample efficiency and utilize more aggressive training configurations without risking irreversible learning plateaus.

Teams developing deep reinforcement learning systems should adopt periodic resets, targeting 3 to 10 reset cycles per training run and re-initializing the last 1 to 3 network layers or full networks depending on the task's representation complexity. However, practitioners should note that resets cause brief transient dips in real-time performance immediately following a reset, meaning safety-critical or regret-minimizing operational deployments will require buffering mechanisms, such as offline post-training or policy blending, before full live deployment.

arXiv: 2205.07802
  • Paper: Understanding Plasticity in Neural Networks, Clare Lyle et al. (2023). Investigates the broader phenomenon of plasticity loss and optimization landscape curvature in deep RL, systematically analyzing mechanisms like parameter resets that directly build on the source's findings.
  • Paper: Actor Prioritized Experience Replay, Baturay Saglam et al. (2023). Extends the study of sampling and early-learning instabilities in continuous actor-critic setups by introducing tailored replay distributions for policy and value networks.
  • Paper: Empirical Design in Reinforcement Learning, Andrew Patterson et al. (2024). Provides a comprehensive empirical methodology guide that builds upon experimental design considerations highlighted in low-sample deep RL studies.
Cover for The Primacy Bias in Deep Reinforcement Learning

Abstract

This work identifies a common flaw of deep reinforcement learning (RL) algorithms: a tendency to rely on early interactions and ignore useful evidence encountered later. Because of training on progressively growing datasets, deep RL agents incur a risk of overfitting to earlier experiences, negatively affecting the rest of the learning process. Inspired by cognitive science, we refer to this effect as the primacy bias. Through a series of experiments, we dissect the algorithmic aspects of deep RL that exacerbate this bias. We then propose a simple yet generally-applicable mechanism that tackles the primacy bias by periodically resetting a part of the agent. We apply this mechanism to algorithms in both discrete (Atari 100k) and continuous action (DeepMind Control Suite) domains, consistently improving their performance.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 The Primacy Bias
  • 3.1 Heavy Priming Causes Unrecoverable Overfitting
  • 3.2 Experiences of Primed Agents are Sufficient
  • 4 Have You Tried Resetting It?
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Resets Consistently Improve Performance
  • 5.3 Learning Dynamics of Agents with Resets
  • 5.4 Elements Behind the Success of Resets
  • 5.5 Summary
  • 6 Related Work
  • 7 Future Work and Limitations
  • 8 Conclusion
  • References
  • A Experimental Details
  • B Ablations
  • C Per-Environment and Additional Results

Knowls

  1. Knowl 1 — Primacy Bias in Deep Reinforcement Learning

    definition

    Primacy bias in deep reinforcement learning (RL) is defined as the tendency of an agent parameterized by neural networks to overfit its early interactions with the environment, thereby impairing its ability to learn from subsequent experiences.

    In off-policy deep RL algorithms using an experience replay buffer, early transitions are resampled and trained upon repeatedly before newer data is collected. This can induce representations and value estimates that over-specialize to initial data distributions. The bias generates a compounding feedback loop: overfitted policy and value networks produce poor actions, resulting in the collection of low-quality or non-exploratory data, which further impedes recovery from the suboptimal solution.

  2. Knowl 2 — Periodic Network Resetting Mechanism

    model/method

    Periodic network resetting is a regularization and forgetting mechanism designed to mitigate the primacy bias in deep reinforcement learning without modifying the environment interaction loop or discarding collected experience.

    Let an agent's policy πθ\pi_\theta and/or action-value function QθQ_\theta be parameterized by a neural network with parameter set θ\theta. At regular intervals of TresetT_{\text{reset}} environment steps, a specified subset of the network parameters θreset⊆θ\theta_{\text{reset}} \subseteq \theta (such as the final linear layers or all fully connected layers) is re-initialized by sampling new values from the network's initial parameter distribution:

    θreset←θ0∼Dinit\theta_{\text{reset}} \leftarrow \theta_0 \sim \mathcal{D}_{\text{init}}

    Crucially, the experience replay buffer B\mathcal{B} is preserved entirely across resets (B←B\mathcal{B} \leftarrow \mathcal{B}), allowing the agent to retain its non-parametric record of transitions while discarding overfitted representations or policies. The optimizer's internal state (such as Adam's first and second moment vectors) corresponding to θreset\theta_{\text{reset}} may also be reset to zero. Immediately following a reset, standard training resumes directly on batches sampled from B\mathcal{B} without pre-training or offline burn-in steps.

  3. Knowl 3 — Heavy Priming Experiment Demonstrating Unrecoverable Overfitting

    empirical result

    Excessive training on initial environment interactions causes deep reinforcement learning agents to suffer irreversible performance degradation, demonstrating that the primacy bias alone can permanently stall learning.

    In an experimental test using Soft Actor-Critic (SAC) on the DeepMind Control Suite quadruped-run task:

    1. Standard SAC collects transitions and performs 1 gradient update per environment step, successfully learning to run and reaching undiscounted episode returns near 700700 within 10610^6 steps.
    2. In the "heavy priming" condition, SAC collects an initial set of only 100100 transitions, updates both the policy and value networks 10510^5 times on this initial buffer, and then resumes standard 1-update-per-step training.

    Despite collecting and training on nearly 10610^6 subsequent transitions in the environment, the heavily primed SAC agent completely fails to learn, with returns remaining close to 00. This demonstrates that deep RL agents cannot automatically escape parameter configurations overfitted to early data, even when presented with large amounts of subsequent data.

  4. Knowl 4 — Data Reusability and Experience Sufficiency of Overfitted Agents

    empirical result

    Data collected by a reinforcement learning agent failing due to the primacy bias contains sufficient information to learn an effective policy; the failure to improve stems from optimization entrapment in the agent's parameters rather than deficient exploration data.

    When Soft Actor-Critic (SAC) is trained on DMC quadruped-run with a high replay ratio of 9 gradient updates per environment step, the agent overfits to early data and fails, achieving flat episode returns around 250250. However, when a freshly initialized SAC agent is trained from scratch using the failing agent's full replay buffer as its initial offline data, the new agent rapidly learns and achieves near-optimal episode returns approaching 800–900800\text{--}900. This confirms that the accumulated replay buffer retains high-quality transitions, and the primary bottleneck of primacy bias is the overfitted network's inability to distill knowledge from newer experiences.

  5. Knowl 5 — Network Reset Configurations for SPR, SAC, and DrQ

    model/method

    The layer depth and frequency of periodic parameter resets depend on the baseline algorithm and the complexity of representation learning in the domain:

    1. Self-Predictive Representations (SPR) on Atari 100k:

      • Reset target: Only the final linear layer (layer 5 of 5) of the Q-network.
      • Periodicity: Every 2×1042 \times 10^4 environment steps.
      • Buffer: Prioritized experience replay (PER) storing all prior transitions.
    2. Soft Actor-Critic (SAC) on DeepMind Control Suite (dense states):

      • Reset target: Entire actor, critic, and target critic networks (all 3 fully connected layers).
      • Periodicity: Every 2×1052 \times 10^5 environment steps.
      • Buffer: Uniform replay buffer storing all prior transitions.
    3. Data-Regularized Q-learning (DrQ) on DeepMind Control Suite (pixel observations):

      • Reset target: The last 3 layers out of 7 (the multi-layer Q-head / actor head) across actor, critic, and target critic networks, leaving convolutional encoder weights intact.
      • Periodicity: Every 2×1052 \times 10^5 environment steps (or 4×105/R4 \times 10^5 / R steps, where RR is the task action repeat).
      • Buffer: Uniform replay buffer storing the most recent 10510^5 transitions.

    For all methods, optimizer moment statistics for the reinitialized layers are reset, and training continues immediately without pre-training.

  6. Knowl 6 — Aggregate Benchmark Performance with Network Resets on Atari 100k and DMC

    data/table

    Periodic resetting consistently improves sample-efficient deep RL algorithms across discrete and continuous action benchmarks, measured via Interquartile Mean (IQM), Median, and Mean across runs with 95% bootstrap confidence intervals.

    Method IQM Median Mean
    Atari 100k (26 discrete tasks, 20 seeds)
    SPR + resets 0.478 (0.46, 0.51) 0.512 (0.42, 0.57) 0.911 (0.84, 1.00)
    SPR 0.380 (0.36, 0.39) 0.433 (0.38, 0.48) 0.578 (0.56, 0.60)
    DrQ(ϵ)\text{DrQ}(\epsilon) 0.280 (0.27, 0.29) 0.304 (0.28, 0.33) 0.465 (0.46, 0.48)
    DER 0.183 (0.18, 0.19) 0.191 (0.18, 0.21) 0.351 (0.34, 0.36)
    CURL 0.113 (0.11, 0.12) 0.102 (0.09, 0.12) 0.261 (0.25, 0.27)
    DeepMind Control Suite (dense states, 10 seeds)
    SAC + resets 656 (549, 753) 617 (538, 681) 607 (547, 667)
    SAC 501 (389, 609) 475 (407, 563) 484 (420, 548)
    DeepMind Control Suite (pixels, 10 seeds)
    DrQ + resets 762 (704, 815) 680 (625, 731) 677 (632, 720)
    DrQ 569 (475, 662) 521 (470, 600) 535 (481, 589)

    On Atari 100k, SPR + resets achieves super-human performance on 7 games (compared to 4 for baseline SPR) and raises the human-normalized mean score from 0.579 to 0.901. On DMC, resets provide relative IQM gains of +30.9% for SAC and +33.9% for DrQ without incurring additional computational overhead.

  7. Knowl 7 — Interaction of Network Resets with Replay Ratio and n-Step TD Targets

    empirical result

    Network resets unlock higher replay ratios and longer nn-step returns, regimes where baseline deep RL agents fail due to exacerbated primacy bias.

    1. Replay Ratio (RR): The replay ratio defines the number of gradient updates per environment step. Higher replay ratios increase data reuse and sample efficiency but heighten the risk of overfitting to early transitions.
    • For SPR on Atari 100k, adding resets yields a >40%>40\% performance increase at RR=4\text{RR}=4.
    • For SAC on DMC, standard SAC performance drops precipitously at high replay ratios (from IQM≈500\text{IQM} \approx 500 at RR=1\text{RR}=1 to 301301 at RR=32\text{RR}=32, 149149 at RR=128\text{RR}=128, and 3939 at RR=256\text{RR}=256). Adding resets allows SAC performance to improve as replay ratio increases, peaking at IQM=651\text{IQM} = 651 at RR=32\text{RR}=32 and maintaining robust scores (IQM=584\text{IQM} = 584 at RR=128\text{RR}=128, IQM=520\text{IQM} = 520 at RR=256\text{RR}=256).
    1. nn-Step Temporal-Difference Targets: Larger multi-step horizons nn reduce target bias but increase target variance:

    Gt(n)=∑k=0n−1γkrt+k+γnQ(st+n,at+n)G_t^{(n)} = \sum_{k=0}^{n-1} \gamma^k r_{t+k} + \gamma^n Q(s_{t+n}, a_{t+n})

    Higher target variance increases susceptibility to early data overfitting. In SPR, resets provide no gain at n=3n=3, but improve performance by up to 40%40\% at n=20n=20. In SAC (RR=9\text{RR}=9), resets yield an improvement of 50–60%50\text{--}60\% for n=3n=3 and n=5n=5, compared to 40%40\% at n=1n=1.

  8. Knowl 8 — Comparison Between Network Resets and Standard Regularization Techniques

    empirical result

    Standard deep learning regularizers such as L2L_2 weight decay and dropout fail to alleviate the primacy bias in deep RL and reduce baseline agent performance compared to periodic resets.

    On DMC continuous control tasks, Soft Actor-Critic (SAC) and Data-Regularized Q-learning (DrQ) were evaluated against versions augmented with L2L_2 regularization (swept over penalty coefficients λ∈{10−4,5⋅10−4,10−3,5⋅10−3}\lambda \in \{10^{-4}, 5\cdot 10^{-4}, 10^{-3}, 5\cdot 10^{-3}\}) and dropout (rates p∈{0.1,0.5}p \in \{0.1, 0.5\}):

    Method IQM Median Mean
    SAC 501 (389, 609) 475 (407, 563) 484 (420, 548)
    SAC + resets 656 (549, 753) 617 (538, 681) 607 (547, 667)
    SAC + dropout 219 (160, 285) 254 (204, 307) 258 (216, 300)
    SAC + L2L_2 412 (299, 524) 415 (337, 495) 416 (351, 481)
    DrQ 569 (475, 662) 521 (470, 600) 535 (481, 589)
    DrQ + resets 762 (704, 815) 680 (625, 731) 677 (632, 720)
    DrQ + dropout 492 (414, 567) 480 (420, 541) 479 (431, 527)
    DrQ + L2L_2 463 (362, 566) 473 (403, 541) 472 (415, 529)

    Furthermore, applying L2L_2 regularization across eleven coefficients ranging from 10−510^{-5} to 1.01.0 completely failed to rescue SAC from the heavy priming collapse on quadruped-run. While resets implicitly control parameter norms by periodically restoring weights to initial scale, standard parameter shrinkages do not resolve the structural representational entrapment caused by primacy bias.

  9. Knowl 9 — Ablation Analysis on Replay Buffer Retention and Parameter Initialization

    empirical result

    Ablation studies on Data-Regularized Q-learning (DrQ) and Self-Predictive Representations (SPR) reveal the architectural and algorithmic factors governing recovery after network resets:

    1. Replay Buffer Retention: Clearing the replay buffer alongside network weights when resetting DrQ on DMC tasks (walker-stand, walker-run, hopper-stand, hopper-hop) completely destroys learning progress, yielding near-zero returns. Retaining collected transitions in the buffer is essential for rapid recovery after resets.
    2. Initialization Seed: Reinitializing network weights using the identical random seed used at step 0 yields performance identical to reinitializing with fresh random seeds. This confirms primacy bias is not caused by pathological initializations, but by non-stationary optimization dynamics over growing datasets.
    3. Optimizer Statistics: Resetting only optimizer moment estimates (Adam β1,β2\beta_1, \beta_2) without resetting network weights produces no effect on training curves, as momentum statistics adapt within ≈10–1000\approx 10\text{--}1000 updates.
    4. Target Subnetworks: Resetting the critic yields the predominant performance gain in DrQ compared to resetting only the actor.
    5. Reset Frequency: While a single reset after initial training helps, periodic continual resetting (yielding 3 to 10 resets across training) achieves the highest final performance.
  10. Knowl 10 — Resolution of Temporal-Difference Learning Divergence and Collapse via Resets

    empirical result

    Periodic network resets resolve two common failure modes of off-policy temporal-difference (TD) learning with neural network function approximation:

    1. Degenerate Critic Collapse in Sparse Rewards: In sparse-reward tasks such as cartpole-swingup_sparse, standard DrQ agents can bootstrap predominantly on their own near-zero value predictions and collapse into trivial critic outputs with vanishing gradient norms, failing to learn even when successful goal transitions are present in ≈2%\approx 2\% of the replay buffer trajectories. Resetting the critic and actor allows the optimization process to restart and successfully learn a near-optimal policy from the existing rewarding trajectories.

    2. Critic Value Divergence and Overestimation: In continuous control environments such as walker-stand, TD learning can experience severe value divergence where predicted Q-values explode to unphysically large magnitudes (>103>10^3) despite the use of double Q-learning and target networks. Once diverged, standard gradient descent cannot recover normal value magnitudes. A network reset restores the critic parameters to initial scales, enabling the agent to re-estimate values stably and reach optimal returns.

  11. Knowl 11 — Practical Limitations of Periodic Network Resets

    limitation

    While periodic network resetting improves final performance and sample efficiency in off-policy deep RL, it presents two primary practical limitations:

    1. Hyperparameter Selection: The reset periodicity TresetT_{\text{reset}} and the specific network depth to reset (e.g., last linear layer vs. entire head vs. full network) are problem- and architecture-dependent hyperparameters that must be set empirically. Resetting too deeply in representation-heavy domains (such as resetting convolutional visual representations in Atari) impairs performance, while resetting too shallowly in low-dimensional state spaces may not fully eliminate overfitted features.
    2. Transient Performance Collapses: Immediately following a parameter reset, policy returns experience brief temporary drops while the newly initialized layers learn from the replay buffer. While performance typically recovers within a few thousand gradient steps, these transient performance collapses lead to temporary drops in online return, which can be undesirable in settings where cumulative online regret minimization or safety constraints are required during training.

Coverage note — All primary contributions, empirical analyses, ablation studies, and limitations from the paper are covered. Detailed per-game raw scores for all 26 Atari 100k tasks from Appendix C (Table 7) were omitted in favor of the aggregate benchmark table to prevent dilution.

References

  1. 1.Achille, A., Rovere, M., and Soatto, S. Critical learning periods in deep networks. In International Conference on Learning Representations, 2018.
  2. 2.Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G. Deep reinforcement learning at the edge of the statistical precipice. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  3. 3.Alabdulmohsin, I., Maennel, H., and Keysers, D. The impact of reinitialization on generalization in convolutional neural networks. arXiv preprint arXiv:2109.00267, 2021.
  4. 4.Anderson, C. W. et al. Q-learning with hidden-unit restarting. Advances in Neural Information Processing Systems, pp. 81–81, 1993.
  5. 5.Asch, S. E. Forming impressions of personality. University of California Press, 1961.
  6. 6.Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Fantacci, C., Godwin, J., Jones, C., Hennigan, T., Hessel, M., Kapturowski, S., Keck, T., Kemaev, I., King, M., Martens, L., Mikulik, V., Norman, T., Quan, J., Papamakarios, G., Ring, R., Ruiz, F., Sanchez, A., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Stokowiec, W., and Viola, F. The DeepMind JAX Ecosystem, 2020. URL http://github.com/deepmind.
  7. 7.Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, 2013.
  8. 8.Bengio, E., Pineau, J., and Precup, D. Interference and generalization in temporal difference learning. In International Conference on Machine Learning, pp. 767–777. PMLR, 2020.
  9. 9.Besson, L. and Kaufmann, E. What doubling tricks can and can’t do for multi-armed bandits. arXiv preprint arXiv:1803.06971, 2018.
  10. 10.Bjorck, J., Gomes, C. P., and Weinberger, K. Q. Is high variance unavoidable in rl? a case study in continuous control. arXiv preprint arXiv:2110.11222, 2021.
  11. 11.Bjorck, J., Gomes, C. P., and Weinberger, K. Q. Is high variance unavoidable in RL? a case study in continuous control. In International Conference on Learning Representations, 2022.
  12. 12.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  13. 13.Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D. The value-improvement path: Towards better representations for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 7160–7168, 2021.
  14. 14.Dohare, S., Mahmood, A. R., and Sutton, R. S. Continual backprop: Stochastic gradient descent with persistent randomness. arXiv preprint arXiv:2108.06325, 2021.
  15. 15.D’Oro, P. and Jaśkowski, W. How to learn a useful critic? model-based action-gradient-estimator policy optimization. Advances in Neural Information Processing Systems, 33, 2020.
  16. 16.Erhan, D., Courville, A., Bengio, Y., and Vincent, P. Why does unsupervised pre-training help deep learning? In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 201–208. JMLR Workshop and Conference Proceedings, 2010.
  17. 17.Farahmand, A., Ghavamzadeh, M., Mannor, S., and Szepesvári, C. Regularized policy iteration. Advances in Neural Information Processing Systems, 21, 2008.
  18. 18.Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W. Revisiting fundamentals of experience replay. In International Conference on Machine Learning, pp. 3061–3071. PMLR, 2020.
  19. 19.French, R. M. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135, 1999.
  20. 20.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pp. 2052–2062. PMLR, 2019.
  21. 21.Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905, 2018.
  22. 22.Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. Learning latent dynamics for planning from pixels. In International Conference on Machine Learning, pp. 2555–2565. PMLR, 2019.
  23. 23.Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. Rainbow: Combining improvements in deep reinforcement learning. In Thirty-second AAAI conference on artificial intelligence, 2018.
  24. 24.Hunter, J. D. Matplotlib: A 2d graphics environment. IEEE Annals of the History of Computing, 9(03):90–95, 2007.
  25. 25.Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S. Transient non-stationarity and generalisation in deep reinforcement learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Qun8fv4qSby.
  26. 26.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pp. 448–456. PMLR, 2015.
  27. 27.Johnson, J. S. and Newport, E. L. Critical period effects in second language learning: The influence of maturational state on the acquisition of english as a second language. Cognitive psychology, 21(1):60–99, 1989.
  28. 28.Jones, E., Oliphant, T., and Peterson, P. SciPy: Open source scientific tools for Python. 2014.
  29. 29.Kaiser, Ł., Babaeizadeh, M., Miłos, P., Osiński, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. Model based reinforcement learning for atari. In ICLR, 2019.
  30. 30.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In ICLR (Poster), 2015.
  31. 31.Kirby, S. Spontaneous evolution of linguistic structure-an iterated learning model of the emergence of regularity and irregularity. IEEE Transactions on Evolutionary Computation, 5(2):102–110, 2001.
  32. 32.Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T. A survey of generalisation in deep reinforcement learning. arXiv preprint arXiv:2111.09794, 2021.
  33. 33.Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B. E., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J. B., Grout, J., Corlay, S., et al. Jupyter Notebooks-a publishing format for reproducible computational workflows., volume 2016. 2016.
  34. 34.Kostrikov, I. JAXRL: Implementations of Reinforcement Learning algorithms in JAX, 10 2021. URL https://github.com/ikostrikov/jaxrl.
  35. 35.Kostrikov, I., Yarats, D., and Fergus, R. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In ICLR, 2021.
  36. 36.Kumar, A., Agarwal, R., Ghosh, D., and Levine, S. Implicit under-parameterization inhibits data-efficient deep reinforcement learning. In International Conference on Learning Representations, 2020.
  37. 37.Li, C. J., Yu, Y., Loizou, N., Gidel, G., Ma, Y., Roux, N. L., and Jordan, M. I. On the convergence of stochastic extragradient for bilinear games with restarted iteration averaging. arXiv preprint arXiv:2107.00464, 2021.
  38. 38.Lin, L. J. Self-improving reactive agents based on reinforcement learning, planning and teaching. Mach. Learn., 8: 293–321, 1992.
  39. 39.Liu, Z., Li, X., Kang, B., and Darrell, T. Regularization matters in policy optimization - an empirical study on continuous control. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=yr1mzrH3IC.
  40. 40.Lyle, C., Rowland, M., and Dabney, W. Understanding and preventing capacity loss in reinforcement learning. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=ZkC8wKoLbQ7.
  41. 41.Marshall, P. H. and Werder, P. R. The effects of the elimination of rehearsal on primacy and recency. Journal of Verbal Learning and Verbal Behavior, 11(5):649–653, 1972.
  42. 42.McKinney, W. Python for data analysis: Data wrangling with Pandas, NumPy, and IPython. ” O’Reilly Media, Inc.”, 2012.
  43. 43.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. nature, 518(7540): 529–533, 2015.
  44. 44.Oliphant, T. E. A guide to NumPy, volume 1. Trelgol Publishing USA, 2006.
  45. 45.Oliphant, T. E. Python for scientific computing. Computing in Science & Engineering, 9(3):10–20, 2007.
  46. 46.Ren, Y., Guo, S., Labeau, M., Cohen, S. B., and Kirby, S. Compositional languages emerge in a neural iterated learning model. In International Conference on Learning Representations, 2019.
  47. 47.Robins, A. Catastrophic forgetting in neural networks: the role of rehearsal mechanisms. In Proceedings 1993 The First New Zealand International Two-Stream Conference on Artificial Neural Networks and Expert Systems, pp. 65–68. IEEE, 1993.
  48. 48.Schaul, T., Quan, J., Antonoglou, I., and Silver, D. Prioritized experience replay. In ICLR (Poster), 2016.
  49. 49.Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P. Data-efficient reinforcement learning with self-predictive representations. In International Conference on Learning Representations, 2020.
  50. 50.Sharkey, N. E. and Sharkey, A. J. An analysis of catastrophic interference. Connection Science, 1995.
  51. 51.Shteingart, H., Neiman, T., and Loewenstein, Y. The role of first impression in operant learning. Journal of Experimental Psychology: General, 142(2):476, 2013.
  52. 52.Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B. Observational overfitting in reinforcement learning. In International Conference on Learning Representations, 2019.
  53. 53.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  54. 54.Sutton, R. S. Learning to predict by the methods of temporal differences. Machine learning, 3(1):9–44, 1988.
  55. 55.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  56. 56.Taha, A., Shrivastava, A., and Davis, L. S. Knowledge evolution in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12843–12852, 2021.
  57. 57.Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N. dm control: Software and tasks for continuous control, 2020.
  58. 58.Teh, Y. W., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. Distral: Robust multitask reinforcement learning. In NIPS, 2017.
  59. 59.Van Der Walt, S., Colbert, S. C., and Varoquaux, G. The numpy array: a structure for efficient numerical computation. Computing in science & engineering, 13(2):22–30, 2011.
  60. 60.Van Hasselt, H., Guez, A., and Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI conference on artificial intelligence, volume 30, 2016.
  61. 61.van Hasselt, H. P., Hessel, M., and Aslanides, J. When to use parametric models in reinforcement learning? In NeurIPS, 2019a.
  62. 62.van Hasselt, H. P., Hessel, M., and Aslanides, J. When to use parametric models in reinforcement learning? Advances in Neural Information Processing Systems, 32:14322–14333, 2019b.
  63. 63.Van Rossum, G. and Drake Jr, F. L. Python tutorial, volume 620. Centrum voor Wiskunde en Informatica Amsterdam, 1995.
  64. 64.Vanseijen, H. and Sutton, R. A deeper look at planning as learning from replay. In Bach, F. and Blei, D. (eds.), Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pp. 2314–2322, Lille, France, 07–09 Jul 2015. PMLR. URL https://proceedings.mlr.press/v37/vanseijen15.html.
  65. 65.Wang, C., Wu, Y., Vuong, Q., and Ross, K. Striving for simplicity and performance in off-policy drl: Output normalization and non-uniform sampling. In International Conference on Machine Learning, pp. 10070–10080. PMLR, 2020.
  66. 66.Yalnizyan-Carson, A. and Richards, B. A. Forgetting enhances episodic control with structured memories. bioRxiv, 2021.
  67. 67.Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R. Improving sample efficiency in model-free reinforcement learning from images. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 10674–10681, 2021.
  68. 68.Zhang, C., Vinyals, O., Munos, R., and Bengio, S. A study on overfitting in deep reinforcement learning. arXiv preprint arXiv:1804.06893, 2018.
  69. 69.Zhang, C., Bengio, S., and Singer, Y. Are all layers created equal? arXiv preprint arXiv:1902.01996, 2019.
  70. 70.Zhou, H., Vani, A., Larochelle, H., and Courville, A. Fortuitous forgetting in connectionist networks. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=ei3SY1_zYsE.

Citation

MLA
Nikishin, E., et al. “The Primacy Bias in Deep Reinforcement Learning”. arXiv, 2022, http://arxiv.org/abs/2205.07802v1.
APA
Nikishin, E., Schwarzer, M., D'Oro, P., Bacon, P.-L., & Courville, A. (2022). The Primacy Bias in Deep Reinforcement Learning. arXiv. http://arxiv.org/abs/2205.07802v1
Chicago
Nikishin, E., M. Schwarzer, P. D'Oro, P.-L. Bacon, and A. Courville. 2022. “The Primacy Bias in Deep Reinforcement Learning”. arXiv. http://arxiv.org/abs/2205.07802v1.
Harvard
Nikishin, E. et al. (2022) “The Primacy Bias in Deep Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.07802v1.
Vancouver
1. Nikishin E, Schwarzer M, D'Oro P, Bacon P-L, Courville A (2022) The Primacy Bias in Deep Reinforcement Learning. arXiv

BibTeX

@article{nikishin2022the,
  title = {The Primacy Bias in Deep Reinforcement Learning},
  author = {Nikishin, Evgenii and Schwarzer, Max and D'Oro, Pierluca and Bacon, Pierre-Luc and Courville, Aaron},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.07802v1},
  eprint = {2205.07802}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/