The Primacy Bias in Deep Reinforcement Learning
Evgenii NikishinMax SchwarzerPierluca D'OroPierre-Luc BaconAaron C. Courville
Identifies how deep reinforcement learning agents overfit to early experience, proposing a simple periodic parameter-resetting mechanism that overcomes this primacy bias and consistently improves performance across discrete and continuous benchmarks.
Modern deep reinforcement learning algorithms frequently suffer from a failure mode where artificial agents overfit to their earliest interactions with an environment and subsequently fail to learn from new, higher-quality data. This dynamic, termed the primacy bias, creates a negative feedback loop: an overfitted agent executes poor actions, collects low-quality experience, and becomes permanently unable to reach optimal performance even after extensive additional training. As deep reinforcement learning is increasingly deployed in data-intensive and complex environments, overcoming this learning bottleneck is critical for improving data efficiency and final agent capabilities.
The article evaluates the causes and mechanics of the primacy bias across deep reinforcement learning algorithms and demonstrates a simple, lightweight remediation strategy. To test this, the authors evaluated standard algorithms across diverse benchmarks, including 26 discrete-action Atari games and 19 continuous-control robotic tasks from the DeepMind Control Suite, spanning both raw pixel inputs and dense physical states across multiple random seeds.
The proposed solution, termed resetting, periodically re-initializes the parameters of the agent's neural networks (or their final layers) while strictly preserving all historical interactions stored in the data replay buffer. The analysis reveals three primary findings. First, periodic resets consistently improve agent performance without adding computational overhead, increasing the aggregate interquartile mean score on Atari benchmarks by over 25% and lifting continuous-control baseline performance by 30% to 34%. Second, resets enable agents to thrive under aggressive training regimes: at high data reuse rates, adding resets doubled performance where standard algorithms collapsed. Third, traditional regularization techniques such as weight decay (L2) and dropout fail to prevent primacy bias, whereas resetting effectively resolves optimization failures like value estimation collapse and divergence.
These findings indicate that the primary barrier in many underperforming reinforcement learning systems is an optimization failure within the neural network rather than inadequate data collection. Retaining the non-parametric memory of the environment in the replay buffer while clearing out entrenched, overfitted network weights allows the agent to quickly recover and leverage accumulated experience. This permits organizations to achieve higher sample efficiency and utilize more aggressive training configurations without risking irreversible learning plateaus.
Teams developing deep reinforcement learning systems should adopt periodic resets, targeting 3 to 10 reset cycles per training run and re-initializing the last 1 to 3 network layers or full networks depending on the task's representation complexity. However, practitioners should note that resets cause brief transient dips in real-time performance immediately following a reset, meaning safety-critical or regret-minimizing operational deployments will require buffering mechanisms, such as offline post-training or policy blending, before full live deployment.
- Paper: Prioritized Experience Replay, Tom Schaul et al. (2016). Introduces prioritized experience replay, establishing the core mechanics of replay buffer sampling that the source analyzes when dissecting early-data overfitting and replay-induced bias.
- Paper: Deep Reinforcement Learning at the Edge of the Statistical Precipice, Rishabh Agarwal et al. (2021). Establishes rigorous statistical evaluation methodologies and standard protocols on the Atari 100k and DeepMind Control Suite benchmarks that the source directly utilizes.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). Presents Soft Actor-Critic, the foundational continuous-control reinforcement learning algorithm evaluated in the source's empirical analysis.
- Paper: Deep Reinforcement Learning with Double Q-learning, Hado van Hasselt et al. (2016). Analyzes value overestimation and optimization dynamics in Q-learning, providing essential context for how algorithmic choices exacerbate bias during RL agent training.
- Paper: Rainbow: Combining Improvements in Deep Reinforcement Learning, Matteo Hessel et al. (2017). Synthesizes major value-based deep RL algorithmic components, illustrating the baseline architectures whose susceptibility to early interactions is examined in the source.
- Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). Establishes standard deep Q-learning with experience replay, providing the foundational training paradigm whose vulnerability to primacy bias is investigated.
- Paper: Understanding Plasticity in Neural Networks, Clare Lyle et al. (2023). Investigates the broader phenomenon of plasticity loss and optimization landscape curvature in deep RL, systematically analyzing mechanisms like parameter resets that directly build on the source's findings.
- Paper: Actor Prioritized Experience Replay, Baturay Saglam et al. (2023). Extends the study of sampling and early-learning instabilities in continuous actor-critic setups by introducing tailored replay distributions for policy and value networks.
- Paper: Empirical Design in Reinforcement Learning, Andrew Patterson et al. (2024). Provides a comprehensive empirical methodology guide that builds upon experimental design considerations highlighted in low-sample deep RL studies.
