Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
Michal NaumanMichal BortkiewiczPiotr MilosTomasz TrzcinskiMateusz OstaszewskiMarek Cygan
Reveals through an extensive evaluation of over 60 actor-critic variants that general neural network regularizers outperform domain-specific reinforcement learning modifications, enabling basic Soft Actor-Critic agents to achieve state-of-the-art sample efficiency across diverse continuous control tasks.
Off-policy reinforcement learning has become a central method for training autonomous control systems in robotics and complex industrial automation. To make these systems learn efficiently from limited interactions, practitioners frequently increase the frequency of learning updates relative to gathered data. However, intensive updating often introduces severe training instabilities, including value overestimation, model overfitting, and loss of learning plasticity (the capacity of a neural network to acquire new information). While researchers have proposed numerous specialized techniques to address these issues, prior studies typically evaluated interventions in narrow settings or within single simulation environments, leaving engineering teams uncertain about which design choices truly deliver robust, generalizable performance.
The main objective of the article is to evaluate how diverse regularization techniques interact and affect off-policy reinforcement learning performance. Specifically, it demonstrates whether broad, standard neural network regularizers can outperform domain-specific reinforcement learning methods across different task types and update frequencies.
To establish these insights, the researchers implemented 64 distinct agent configurations within the standard Soft Actor-Critic framework. They tested these agents across 14 continuous control tasks spanning two established benchmark suites: the DeepMind Control Suite (locomotion tasks) and MetaWorld (robotic manipulation tasks). The experimental matrix systematically combined three categories of interventions: critic regularizations designed to prevent value overestimation, network regularizations to control overfitting, and plasticity regularizations to maintain adaptability. The configurations were evaluated across both low and high update-frequency regimes using multiple random trials to ensure statistical reliability.
The study reveals four key findings. First, generic neural network regularizations (such as layer normalization and spectral normalization) and plasticity interventions (such as periodic network resets) substantially outperform reinforcement learning-specific critic modifications. Second, commonly used critic regularizations like Clipped Double Q-learning often impair performance when general network regularizations are present, causing severe degradation in complex manipulation tasks. Third, applying standard layer normalization and periodic resets enables a basic model-free agent to master challenging quadruped locomotion tasks (such as the Dog domain) that previously required complex model-based architectures. Fourth, the primary statistical predictors of failure are severe value overestimation and exploding gradient magnitudes, both of which are effectively mitigated by generic normalization techniques rather than specialized reinforcement learning loss functions.
These findings suggest that complex reinforcement learning systems can be substantially simplified while achieving higher stability and performance. Engineering teams can eliminate complicated, domain-specific value-clipping heuristics in favor of standard deep learning practices like layer normalization and weight resets. This shift reduces system complexity, mitigates deployment risks, and shortens development timelines for continuous control applications. Moreover, the divergent results observed between locomotion and manipulation benchmarks underscore that algorithms tuned solely on single benchmark suites risk significant failure when deployed to new operational domains.
Practitioners should prioritize incorporating layer normalization or spectral normalization alongside periodic network resets into their continuous control pipelines. Teams should also re-evaluate or remove specialized critic clipping mechanisms unless empirical testing confirms their necessity for a given task. Before committing to widespread architectural changes, organizations should establish broad evaluation suites containing both locomotion and manipulation benchmarks rather than relying on isolated test environments.
The conclusions of the article are subject to certain limitations. The empirical evaluations were confined to continuous control tasks relying on direct physical state measurements within the Soft Actor-Critic architecture. Preliminary tests on image-based inputs showed that these benefits do not directly transfer to vision-based models under low update frequencies. While confidence in the reported state-based continuous control results is high, practitioners should exercise caution and conduct targeted pilot tests before applying these recommendations to vision-driven agents or non-actor-critic frameworks.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). Read the original Soft Actor-Critic paper first to understand the off-policy, maximum-entropy actor-critic framework and twin-critic design that this study directly evaluates.
- Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). TD3 establishes the critic overestimation problem and the clipped double-Q remedy that the source directly tests against generic regularization.
- Paper: Understanding Plasticity in Neural Networks, Clare Lyle et al. (2023). This study develops the plasticity-loss concepts and interventions—especially normalization and resets—that the source evaluates in continuous-control actor-critic learning.
- Paper: The Primacy Bias in Deep Reinforcement Learning, Evgenii Nikishin et al. (2022). The primacy-bias paper motivates periodic network resets as a remedy for overfitting to early experience, a central intervention in the source.
- Paper: Deep Reinforcement Learning with Double Q-learning, Hado van Hasselt et al. (2016). Double Q-learning explains the value-overestimation correction underlying clipped double-Q methods, which the source finds can hurt performance.
- Paper: Dropout Q-Functions for Doubly Efficient Reinforcement Learning, Takuya Hiraoka et al. (2022). DroQ combines layer normalization and dropout in high-update off-policy control, providing a direct precursor to the source’s comparison of generic and specialized regularization.
No sufficiently relevant recommendations were found.
