Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Takuya HiraokaTakahisa ImagawaTaisei HashimotoTakashi OnishiYoshimasa Tsuruoka

article2022ICLR175 citations

Proposes DroQ, a continuous reinforcement learning algorithm that replaces large Q-ensembles with dropout-regularized networks to achieve the state-of-the-art sample efficiency of REDQ at the computational speed of standard Soft Actor-Critic.

Listen

Applying reinforcement learning to complex real-world control systems has long been hindered by high data requirements, often requiring millions of interactions to achieve adequate performance. Recent high-update methods address this by updating learning models multiple times per interaction, but they suffer from severe overestimation bias. To counteract this bias, state-of-the-art frameworks employ large ensembles of ten neural network models, which causes significant computational latency, high memory consumption, and slow development cycles on resource-constrained hardware.

The article develops and evaluates DroQ (Dropout Q-functions), an algorithm designed to achieve both high sample efficiency and high computational efficiency. The method replaces large model ensembles with a compact ensemble of just two models enhanced with dropout regularization and layer normalization to capture model uncertainty without massive computational overhead.

The researchers benchmarked DroQ against standard continuous-control tasks in the MuJoCo simulation suite (Hopper, Walker2d, Ant, and Humanoid). They systematically evaluated performance against standard baselines—including Soft Actor-Critic (SAC), Randomized Ensembled Double Q-learning (REDQ), and Double Uncertainty Value Networks (DUVN)—measuring learning speed, estimation bias, processing runtime, parameter counts, and peak memory usage.

The findings show that DroQ matches the sample efficiency and estimation bias reduction of large-ensemble methods like REDQ while running more than twice as fast per update loop (approximately 930–990 milliseconds compared to 2300–2400 milliseconds). Furthermore, DroQ reduces total model parameters by roughly 80% (about 140,000–166,000 parameters versus 700,000–820,000) and decreases peak memory consumption by nearly 67% (from roughly 241 MB down to 73 MB). Ablation analyses confirmed that combining dropout with layer normalization is essential, as layer normalization stabilizes the gradient oscillations introduced by dropout, whereas omitting either component severely degrades performance.

These results demonstrate that engineering small models with dropout and normalization can replicate the uncertainty estimation benefits of large ensembles. For engineering teams and stakeholders, this offers a direct path to deploying sample-efficient control algorithms on mobile, robotics, and edge hardware without the prohibitive computational costs and memory footprints of previous architectures.

Organizations developing resource-sensitive control systems should adopt small-ensemble dropout architectures like DroQ as a drop-in replacement for larger ensemble baselines. Implementation is straightforward, requiring only standard layer additions to existing network definitions. Development teams should use layer normalization rather than batch normalization and conduct hyperparameter tuning on task-specific dropout rates.

Confidence in these findings is high for standard simulation-based continuous control tasks. However, the evaluation remains bounded by synthetic physics benchmarks in simulated environments. Further validation through physical hardware pilots is recommended before full-scale operational deployment.

No sufficiently relevant recommendations were found.

Cover for Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Abstract

Randomized ensembled double Q-learning (REDQ) (Chen et al., 2021b) has recently achieved state-of-the-art sample efficiency on continuous-action reinforcement learning benchmarks. This superior sample efficiency is made possible by using a large Q-function ensemble. However, REDQ is much less computationally efficient than non-ensemble counterparts such as Soft Actor-Critic (SAC) (Haarnoja et al., 2018a). To make REDQ more computationally efficient, we propose a method of improving computational efficiency called DroQ, which is a variant of REDQ that uses a small ensemble of dropout Q-functions. Our dropout Q-functions are simple Q-functions equipped with dropout connection and layer normalization. Despite its simplicity of implementation, our experimental results indicate that DroQ is doubly (sample and computationally) efficient. It achieved comparable sample efficiency with REDQ, much better computational efficiency than REDQ, and comparable computational efficiency with that of SAC.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Maximum Entropy Reinforcement Learning (maximum entropy RL)
  • 2.2 Randomized Ensembled Double Q-Learning (REDQ)
  • 3 Injecting Model Uncertainty into Target with Dropout Q-functions
  • 4 Experiments
  • 4.1 Sample efficiency and bias-reduction ability of DroQ
  • 4.2 Computational efficiency of DroQ
  • 4.3 Ablation study
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Effect of dropout rate on DroQ and its variants
  • A.1 DroQ
  • A.2 DroQ without layer normalization
  • A.3 Sin-DroQ: DroQ variant using A single dropout Q-function
  • B REDQ with different ensemble size N
  • C Additional ablation study of DroQ
  • D Why is the combination of dropout and layer normalization important?
  • E Effects of other normalization methods
  • F Relation between (1) ensemble size and (2) the effect of layer normalization and dropout on overall performance
  • G Effect of dropout Q-functions on SAC
  • H Hyperparameter settings
  • I Experiments on the original REDQ codebase
  • J Our source code

Citation

MLA
Hiraoka, T., et al. “Dropout Q-Functions for Doubly Efficient Reinforcement Learning”. arXiv, 2021, http://arxiv.org/abs/2110.02034v2.
APA
Hiraoka, T., Imagawa, T., Hashimoto, T., Onishi, T., & Tsuruoka, Y. (2021). Dropout Q-Functions for Doubly Efficient Reinforcement Learning. arXiv. http://arxiv.org/abs/2110.02034v2
Chicago
Hiraoka, T., T. Imagawa, T. Hashimoto, T. Onishi, and Y. Tsuruoka. 2021. “Dropout Q-Functions for Doubly Efficient Reinforcement Learning”. arXiv. http://arxiv.org/abs/2110.02034v2.
Harvard
Hiraoka, T. et al. (2021) “Dropout Q-Functions for Doubly Efficient Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.02034v2.
Vancouver
1. Hiraoka T, Imagawa T, Hashimoto T, Onishi T, Tsuruoka Y (2021) Dropout Q-Functions for Doubly Efficient Reinforcement Learning. arXiv

BibTeX

@article{hiraoka2021dropout,
  title = {Dropout Q-Functions for Doubly Efficient Reinforcement Learning},
  author = {Hiraoka, Takuya and Imagawa, Takahisa and Hashimoto, Taisei and Onishi, Takashi and Tsuruoka, Yoshimasa},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.02034v2},
  eprint = {2110.02034}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors