Dropout Q-Functions for Doubly Efficient Reinforcement Learning
Takuya HiraokaTakahisa ImagawaTaisei HashimotoTakashi OnishiYoshimasa Tsuruoka
Proposes DroQ, a continuous reinforcement learning algorithm that replaces large Q-ensembles with dropout-regularized networks to achieve the state-of-the-art sample efficiency of REDQ at the computational speed of standard Soft Actor-Critic.
Applying reinforcement learning to complex real-world control systems has long been hindered by high data requirements, often requiring millions of interactions to achieve adequate performance. Recent high-update methods address this by updating learning models multiple times per interaction, but they suffer from severe overestimation bias. To counteract this bias, state-of-the-art frameworks employ large ensembles of ten neural network models, which causes significant computational latency, high memory consumption, and slow development cycles on resource-constrained hardware.
The article develops and evaluates DroQ (Dropout Q-functions), an algorithm designed to achieve both high sample efficiency and high computational efficiency. The method replaces large model ensembles with a compact ensemble of just two models enhanced with dropout regularization and layer normalization to capture model uncertainty without massive computational overhead.
The researchers benchmarked DroQ against standard continuous-control tasks in the MuJoCo simulation suite (Hopper, Walker2d, Ant, and Humanoid). They systematically evaluated performance against standard baselines—including Soft Actor-Critic (SAC), Randomized Ensembled Double Q-learning (REDQ), and Double Uncertainty Value Networks (DUVN)—measuring learning speed, estimation bias, processing runtime, parameter counts, and peak memory usage.
The findings show that DroQ matches the sample efficiency and estimation bias reduction of large-ensemble methods like REDQ while running more than twice as fast per update loop (approximately 930–990 milliseconds compared to 2300–2400 milliseconds). Furthermore, DroQ reduces total model parameters by roughly 80% (about 140,000–166,000 parameters versus 700,000–820,000) and decreases peak memory consumption by nearly 67% (from roughly 241 MB down to 73 MB). Ablation analyses confirmed that combining dropout with layer normalization is essential, as layer normalization stabilizes the gradient oscillations introduced by dropout, whereas omitting either component severely degrades performance.
These results demonstrate that engineering small models with dropout and normalization can replicate the uncertainty estimation benefits of large ensembles. For engineering teams and stakeholders, this offers a direct path to deploying sample-efficient control algorithms on mobile, robotics, and edge hardware without the prohibitive computational costs and memory footprints of previous architectures.
Organizations developing resource-sensitive control systems should adopt small-ensemble dropout architectures like DroQ as a drop-in replacement for larger ensemble baselines. Implementation is straightforward, requiring only standard layer additions to existing network definitions. Development teams should use layer normalization rather than batch normalization and conduct hyperparameter tuning on task-specific dropout rates.
Confidence in these findings is high for standard simulation-based continuous control tasks. However, the evaluation remains bounded by synthetic physics benchmarks in simulated environments. Further validation through physical hardware pilots is recommended before full-scale operational deployment.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). It introduces Soft Actor-Critic (SAC), the continuous-action actor-critic baseline whose computational efficiency and dual Q-network setup DroQ builds upon and benchmarks against.
- Paper: Soft Actor-Critic Algorithms and Applications, Tuomas Haarnoja et al. (2018). It details the standard practical implementation and temperature-tuning mechanisms of Soft Actor-Critic that provide the continuous control foundation for DroQ.
- Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). It establishes the clipped double Q-learning and variance-reduction framework in continuous actor-critic methods that underlies ensembled Q-learning algorithms.
- Paper: Dropout: a simple way to prevent neural networks from overfitting, Nitish Srivastava et al. (2014). It introduces dropout as an implicit ensemble and regularization technique, which DroQ incorporates into Q-function architectures to replace large explicit ensembles.
- Paper: Double Q-learning, Hado van Hasselt (2010). It formalizes the fundamental maximization bias in Q-learning and the principle of decoupling value estimation from action selection using multiple estimators.
- Paper: Deep Reinforcement Learning with Double Q-learning, Hado van Hasselt et al. (2016). It demonstrates how to effectively mitigate overestimation bias in deep neural network value function approximation.
No sufficiently relevant recommendations were found.
