Hyperparameters in Reinforcement Learning and How To Tune Them

Theresa EimerMarius LindauerRoberta Raileanu

article2023ICML91 citations

Demonstrates that automated hyperparameter optimization outperforms manual tuning in reinforcement learning while using less compute, and provides concrete best practices alongside ready-to-use implementations to improve experimental reproducibility.

Listen

Deep reinforcement learning (RL) holds immense potential for solving complex autonomous decision-making problems, yet it faces severe reproducibility and evaluation challenges. In standard practice, practitioners tune internal algorithm settings—known as hyperparameters—manually or via exhaustive grid searches over only a few variables. These tuning processes, as well as the random seeds used during training and evaluation, are rarely reported transparently. This lack of standardization leads to costly compute waste, unfair algorithm comparisons, and brittle models that perform poorly in production.

The article evaluates how hyperparameter choices and random noise impact RL algorithm performance and explores whether automated hyperparameter optimization (HPO) methods can replace expensive manual tuning. To demonstrate this, the authors conduct extensive empirical analyses across major RL algorithm classes (PPO, SAC, DQN, and IDAAC) and diverse standard benchmark environments ranging from classic control tasks to complex robotics and procedurally generated games (such as Brax and Procgen). The study benchmarks several black-box automated tuning methods, including Random Search, evolutionary multi-fidelity methods like DEHB, and population-based training techniques like PB2 and BGT.

The findings establish that hyperparameter settings critically drive algorithm success, yet the optimization landscape is relatively smooth and predictable. Automated HPO tools regularly match or surpass the performance of meticulously hand-tuned models while consuming a fraction of the computational budget. For instance, on complex tasks, tools like DEHB achieve superior performance using less than one-twelfth of the tuning budget required by standard baseline sweeps. Furthermore, the analysis reveals that RL performance is highly sensitive to random seeds: models tuned on a single seed often overfit drastically, sometimes performing more than four times worse on unseen test seeds. Tuning across 3 to 5 seeds significantly improves generalizability, whereas evaluating beyond that can sharply increase search complexity.

These insights demonstrate that standard manual tuning practices inflate computational costs and obscure true algorithmic progress. Because hyperparameter landscapes are smooth, practitioners do not need complex, algorithm-specific custom tuners; accessible, general-purpose AutoML methods are fully sufficient and highly effective. Adopting these automated approaches reduces the cloud compute overhead required to deploy RL systems, accelerates development timelines, and mitigates the operational risk of deploying overfitted policies.

The article recommends that organizations integrate automated HPO natively into their machine learning pipelines. Teams should tune across broad search spaces rather than cherry-picking two or three parameters, strictly separate tuning seeds from final testing seeds to prevent overfitting, and report exact tuning budgets and protocols to ensure fair comparisons. To facilitate this transition, the authors provide open-source, plug-and-play tuning sweepers compatible with standard RL workflows.

While these conclusions are supported by robust experiments across diverse benchmarks, readers should note that population-based methods exhibited variability and may require larger budgets or specialized restart mechanics to prevent stagnation. Overall, confidence is high that adopting disciplined AutoML best practices and automated tuning frameworks will substantially improve RL reliability and resource efficiency.

  • Paper: Empirical Design in Reinforcement Learning, Andrew Patterson et al. (2024). It builds upon hyperparameter tuning and reproducibility practices to formalize comprehensive experimental design, statistical testing, and reporting protocols across reinforcement learning.
Cover for Hyperparameters in Reinforcement Learning and How To Tune Them

Abstract

In order to improve reproducibility, deep reinforcement learning (RL) has been adopting better scientific practices such as standardized evaluation metrics and reporting. However, the process of hyperparameter optimization still varies widely across papers, which makes it challenging to compare RL algorithms fairly. In this paper, we show that hyperparameter choices in RL can significantly affect the agent’s final performance and sample efficiency, and that the hyperparameter landscape can strongly depend on the tuning seed which may lead to overfitting. We therefore propose adopting established best practices from AutoML, such as the separation of tuning and testing seeds, as well as principled hyperparameter optimization (HPO) across a broad search space. We support this by comparing multiple state-of-the-art HPO tools on a range of RL algorithms and environments to their hand-tuned counterparts, demonstrating that HPO approaches often have higher performance and lower compute overhead. As a result of our findings, we recommend a set of best practices for the RL community, which should result in stronger empirical results with fewer computational costs, better reproducibility, and thus faster progress. In order to encourage the adoption of these practices, we provide plug-and-play implementations of the tuning algorithms used in this paper at https://github.com/facebookresearch/how-to-autorl.

Table of Contents

  • 1. Introduction
  • 2. The Hyperparameter Optimization Problem
  • 3. Related Work
  • 4. The Hyperparameter Landscape of RL
  • 4.1. Which RL Hyperparameters Should Be Tuned?
  • 4.2. Are Hyperparameters in RL Well Behaved?
  • 4.3. How Do We Account for Noise?
  • 5. Tradeoffs for Hyperparameter Optimization in Practice
  • 6. Recommendations & Best Practices
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Reproducibility Checklist for Tuning Hyperparameters in RL
  • B. Our AutoRL Hydra Sweepers
  • C. Additional Background on Tuning Methods Used
  • C.1. Random Search
  • C.2. DEHB
  • C.3. PBT Variants
  • D. An Overview of Hyperparameter Configurations & Search Spaces
  • D.1. Stable Baselines Default Configurations
  • D.2. Stable Baseline Sweep Values
  • D.3. Stable Baselines Search Spaces
  • D.4. Brax Experiment Settings
  • D.5. Procgen Experiment Settings
  • D.6. Hardware
  • E. Details on the Tuned Configurations
  • F. Tuning Results on Brax & Procgen in Tabular Form
  • G. Hyperparameter Sweeps for PPO, DQN and SAC
  • G.1. PPO Sweeps
  • G.2. DQN Sweeps
  • G.3. SAC Sweeps
  • H. Full Performance Pointplots
  • H.1. SAC Pointplots
  • H.2. DQN Pointplots
  • H.3. PPO Pointplots
  • I. Hyperparameter Importances using fANOVA
  • J. Partial Dependency Plots
  • J.1. SAC on Pendulum
  • J.2. PPO on Acrobot

Knowls

  1. Knowl 1 — RL hyperparameters broadly determine performance

    empirical result

    Across PPO, DQN, and SAC on discrete and continuous tasks, nearly every swept hyperparameter had a substantial effect on final return. The study evaluated 126 algorithm–environment–hyperparameter settings; in only 7 cases was the worst hyperparameter value within one standard deviation of the best value, and in only 13 cases did the median performance decrease by less than 20% from the best value. Even a hyperparameter that is rarely tuned, such as PPO’s clip range, could determine whether an agent succeeded or failed on a task such as Brax Ant.

    fANOVA nevertheless found that one or two hyperparameters accounted for most of the measured importance in each environment, although the dominant hyperparameters changed by environment: PPO learning rate was dominant on Acrobot, clip range on Pendulum, and GAE lambda on MiniGrid. Partial-dependence analyses on Pendulum and Acrobot revealed few complex interaction patterns. Because many hyperparameters matter, their relative importance is environment-dependent, and strong interference effects were uncommon in these analyses, the study recommends tuning as many potentially relevant hyperparameters as the optimization budget permits rather than selecting only two or three by hand.

  2. Knowl 2 — Smooth average landscapes coexist with severe seed variability

    empirical result

    Average RL performance usually changed smoothly as a hyperparameter value moved through its search range: nearby values tended to have nearby performance, and performance generally deteriorated gradually away from good values rather than changing abruptly. Configuration rankings were also often preserved during training, with good configurations learning quickly and poor configurations degrading early. This supports using partial training runs in multi-fidelity and population-based hyperparameter optimization.

    The same conclusion does not hold reliably for individual random seeds. For a fixed hyperparameter configuration, some seeds crashed, learned inconsistently, or performed exceptionally well despite the configuration having poor average performance. The largest spread often occurred for medium-quality configurations, although unusually strong seeds could also arise from unstable configurations. Thus, the principal obstacle in tuning these RL agents was seed sensitivity rather than an intrinsically irregular average hyperparameter landscape.

  3. Knowl 3 — Complex-benchmark tuning protocol

    experimental setup

    The larger comparison tuned PPO on Brax Ant, Halfcheetah, and Humanoid, and IDAAC on Procgen Bigfish, Climber, and Plunder. The compared methods were hand-tuned baselines, random search (RS), DEHB, and population-based methods including PB2 and BGT. Each HPO setting was repeated three times; every tuning repetition used five training seeds, and the selected incumbent was evaluated on ten previously unseen test seeds. Tuning budgets were at most 16 or 64 full algorithm runs, whereas the published IDAAC baseline used 810 runs.

    For Brax, the optimized cost was evaluation reward over one episode of the batched environment; for Procgen, it was evaluation reward averaged over ten episodes of the training environment. Tuning used seeds 0–4 and testing used seeds 5–14. For each environment, methods were ranked by test performance: the best mean and methods within its standard deviation received rank 1, and subsequent statistically overlapping groups received successive ranks. Domain-level scores were obtained by averaging these environment ranks.

  4. Knowl 4 — HPO improves or matches hand tuning on difficult domains

    data/table

    On the larger Brax and Procgen benchmarks, DEHB was the most consistent method and generally improved on hand-tuned baselines while using far fewer runs. With 16 runs, DEHB achieved mean ranks of 1.3 on Brax versus 1.7 for the baseline and 1.7 on Procgen versus 2.0 for the baseline. With 64 runs, its mean ranks improved to 1.0 on both domains, compared with 1.3 for the baselines. On Brax, RS ranked 2.3 with 16 runs and 3.3 with 64 runs; BGT often required more budget and performed poorly under the restricted budgets. On Procgen, RS ranked 2.7 and 3.0, while BGT ranked 3.8 and 2.7 for 16 and 64 runs, respectively. PB2 could produce strong incumbents but sometimes overfit severely; on Bigfish, its test score was about five times worse than its incumbent score.

    The detailed test and incumbent scores were:

    Could not parse LaTeX table
    Could not parse LaTeX table

    The incumbent–test gaps and the instability of some population-based methods show why test performance, rather than the score used during tuning, must determine comparisons. Despite these differences, the HPO overhead was small relative to RL training: BGT needed on average under two minutes to generate configurations at the 16-run budget and under two hours at the 64-run budget, while the other methods stayed below five minutes for either budget.

  5. Knowl 5 — DEHB and population-based optimization configurations

    model/method

    The study compared three black-box optimization families. Random search samples complete configurations from the search space and selects the best observed incumbent. DEHB combines differential evolution with HyperBand: it evaluates many configurations at progressively larger training budgets, discards low-performing configurations at lower fidelities, and evolves configurations from the survivors. Population-based training maintains agents with individual configurations, periodically evaluates and checkpoints them, replaces a fraction of poor agents with copies of good agents, and perturbs or model-selects their hyperparameters so that the final agent can have a schedule rather than one fixed configuration.

    For the smaller experiments, DEHB used three iterations with η=5\eta=5, giving three budget levels and a minimum budget equal to 1/1001/100 of a full training run. PB2 used a population of 8 and 20 configuration changes. For the larger experiments, DEHB used two iterations with η=1.9\eta=1.9, giving eight budget levels with the same 1/1001/100 minimum budget. PB2 used populations of 16 or 64 according to the available budget. BGT used 8 initial runs and a population of 8 at the smaller budget, and 48 initial runs and a population of 16 at the larger budget. PB2 and BGT replaced the worst 12.5% of agents with the best 12.5% at each replacement step. BGT additionally uses periodic Gaussian-process kernel restarts and full-budget initial runs to warm-start its model.

  6. Knowl 6 — Large search spaces can be tuned with very small budgets

    data/table

    The study compared tuning only the learning rate, a three-hyperparameter space, and the full space using PPO on Acrobot and SAC on Pendulum. Each HPO method received a budget of 10 full RL runs, whereas the reference sweep used 125 runs per environment. The metric is negative evaluation reward, so lower values are better; each incumbent value is the mean and standard deviation across five independent tuning runs.

    Could not parse LaTeX table

    Random search was strong on Acrobot but unreliable on Pendulum, where its tuning runs varied widely. PB2 became worse as the search space grew on Acrobot but improved with search-space size on Pendulum; its incumbents were often nearly static despite being able to learn schedules. DEHB was the most stable across tuning seeds. Overall, all three methods found reasonable configurations in the full spaces despite only 10 full runs, and the methods matched or exceeded the best single-seed results obtained by the much larger sweeps.

  7. Knowl 7 — Separate tuning and testing seeds are essential

    data/table

    The study evaluated whether averaging a configuration’s performance over more tuning seeds improves generalization. PPO on Acrobot and SAC on Pendulum were tuned over the full search spaces with 1, 3, 5, or 10 tuning seeds; each selected incumbent was then evaluated on 10 separate test seeds. The metric is negative evaluation reward, so lower values are better. Values are means and standard deviations across tuning repetitions.

    Could not parse LaTeX table

    Increasing the number of tuning seeds sometimes improved test performance, particularly for RS on Acrobot and Pendulum and for DEHB and PB2 on Pendulum, but the benefit was not monotonic. Using more than three seeds often reduced test performance and could sharply increase variance, as with five-seed RS or ten-seed PB2 on Pendulum. The selected incumbent can therefore overfit the tuning seeds even when its tuning score is excellent; for example, DEHB’s best Acrobot incumbents had test scores more than four times worse. Fair comparisons require the same test seeds for all methods and complete separation between tuning and test seeds.

  8. Knowl 8 — Recommended end-to-end RL tuning protocol

    algorithm

    The study recommends the following procedure for selecting RL hyperparameters:

    1. Define separate training and test settings. These may include environment variations, environment random seeds, initial-state seeds, agent and network-initialization seeds, and random seeds for the HPO tool.
    2. Define a configuration space containing all hyperparameters that could plausibly affect training success. The authors successfully tuned spaces containing up to 14 hyperparameters and recommend avoiding unnecessary pruning for spaces of roughly this size unless prior importance information is reliable.
    3. Select an HPO method and specify its budget or a self-termination rule.
    4. Choose a cost metric based on evaluation reward over enough episodes to obtain a reliable estimate, rather than relying only on a noisy training return.
    5. Run the HPO method on the training settings using multiple tuning seeds; the reproducibility checklist recommends at least five tuning seeds when feasible.
    6. Evaluate the resulting incumbent configurations on separate, unseen test settings and test seeds, and report the test results.

    For a fair algorithm comparison, the proposed method and every baseline should use the same configuration-space definition, tuning budget, cost metric, comparable hardware when the budget is measured in time, and test seeds. Reporting should include the HPO package and optimization method, all search ranges and hyperparameter types, tuning and test seeds, the selection protocol, final configurations, software versions, hardware, and the code implementing the tuning process. Existing hyperparameters should be reused only with their source, original tuning protocol, budget, and search space disclosed; otherwise, the algorithm should be retuned.

  9. Knowl 9 — Plug-and-play Hydra sweepers for AutoRL

    model/method

    The authors provide Hydra-based sweepers implementing DEHB and several population-based variants, including standard PBT, PB2, and BGT, at https://github.com/facebookresearch/how-to-autorl. The sweepers launch configurations locally or as parallel cluster jobs, can resume interrupted optimizations, and expose tuning seeds directly. They are black-box interfaces, so they can be applied to arbitrary RL algorithms and environments without changing the learning algorithm itself.

    To integrate a sweeper, an RL implementation must return a scalar cost or success metric after training; population-based methods additionally require checkpointing and loading the agent’s training state. A Hydra configuration specifies the search space and optimizer, after which the complete tuning process can be run from one configuration file. The implementation also adds initial full-budget runs to the original PBT and PB2 variants to stabilize them and supports multi-fidelity tuning directly, avoiding a separate orchestration script.

  10. Knowl 10 — Scope limitations and unresolved seed-generalization problem

    limitation

    The experiments focus on black-box hyperparameter optimization and do not establish how architecture search should be performed for RL; neural architecture search is treated as a separate problem. The conclusions are also budget-dependent: BGT was likely disadvantaged by the small budgets used in the comparison, and population-based methods often changed configurations only once or not at all instead of learning flexible schedules.

    The remaining central limitation is strong dependence on the random seed for a fixed configuration. Separate tuning and testing seeds reduce misleading evaluations but do not remove the underlying variance, and increasing the number of tuning seeds makes the optimization problem substantially harder. The authors therefore identify seed-responsive dynamic hyperparameter policies, gradient-based methods, and other RL-specific approaches as promising future directions, while noting that gradient-based approaches require access to algorithm gradients and can incur additional computational overhead.

Coverage note — The exhaustive appendix plots, default-configuration tables, complete search-range listings, and formal AC/DAC background definitions were omitted because they expand or formalize the reported findings without adding separate load-bearing contributions.

References

  1. 1.Adriaensen, S., Biedenkapp, A., Shala, G., Awad, N., Eimer, T., Lindauer, M., and Hutter, F. Automated dynamic algorithm configuration. Journal of Artificial Intelligence Research, 2022.
  2. 2.Agarwal, R., Schwarzer, M., Castro, P., Courville, A., and Bellemare, M. Deep reinforcement learning at the edge of the statistical precipice. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS, pp. 29304–29320, 2021.
  3. 3.Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2019.
  4. 4.Andrychowicz, M., Raichuk, A., Stanczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O. What matters for on-policy deep actor-critic methods? A large-scale study. In 9th International Conference on Learning Representations, ICLR. OpenReview.net, 2021.
  5. 5.Awad, N., Mallik, N., and Hutter, F. DEHB: evolutionary hyberband for scalable, robust and efficient hyperparameter optimization. In Zhou, Z. (ed.), Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI, pp. 2147–2153. ijcai.org, 2021.
  6. 6.Badia, A., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z., and Blundell, C. Agent57: Outperforming the atari human benchmark. In Proceedings of the 37th International Conference on Machine Learning, ICML, volume 119 of Proceedings of Machine Learning Research, pp. 507–517. PMLR, 2020.
  7. 7.Bakshy, E., Dworkin, L., Karrer, B., Kashin, K., Letham, B., Murthy, A., and Singh, S. Ae: A domain-agnostic platform for adaptive experimentation. 2018.
  8. 8.Bechtle, S., Molchanov, A., Chebotar, Y., Grefenstette, E., Righetti, L., Sukhatme, G., and Meier, F. Meta learning via learned loss. In 25th International Conference on Pattern Recognition, ICPR, pp. 4161–4168. IEEE, 2020.
  9. 9.Bergstra, J. and Bengio, Y. Random search for hyperparameter optimization. 13:281–305, 2012.
  10. 10.Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Jozefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S. Dota 2 with large scale deep reinforcement learning. CoRR, abs/1912.06680, 2019.
  11. 11.Biedenkapp, A., Bozkurt, H. F., Eimer, T., Hutter, F., and Lindauer, M. Dynamic Algorithm Configuration: Foundation of a New Meta-Algorithmic Framework. In Lang, J., Giacomo, G. D., Dilkina, B., and Milano, M. (eds.), Proceedings of the Twenty-fourth European Conference on Artificial Intelligence (ECAI’20), pp. 427–434, June 2020.
  12. 12.Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. Openai gym, 2016.
  13. 13.Chevalier-Boisvert, M., Willems, L., and Pal, S. Minimalistic gridworld environment for gymnasium, 2018. URL https://github.com/Farama-Foundation/Minigrid.
  14. 14.Co-Reyes, J., Miao, Y., Peng, D., Real, E., Le, Q., Levine, S., Lee, H., and Faust, A. Evolving reinforcement learning algorithms. In 9th International Conference on Learning Representations, ICLR. OpenReview.net, 2021. URL https://openreview.net/forum?id=0XXpJ4OtjW.
  15. 15.Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. Leveraging procedural generation to benchmark reinforcement learning. In Proceedings of the 37th International Conference on Machine Learning, ICML, volume 119 of Proceedings of Machine Learning Research, pp. 2048–2056. PMLR, 2020.
  16. 16.Duan, Y., Schulman, J., Chen, X., Bartlett, P., Sutskever, I., and Abbeel, P. Rlˆ2ˆ2: Fast reinforcement learning via slow reinforcement learning. CoRR, abs/1611.02779, 2016.
  17. 17.Eggensperger, K., Lindauer, M., and Hutter, F. Neural networks for predicting algorithm runtime distributions. In Lang, J. (ed.), Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI), pp. 1442–1448. ijcai.org, 2018.
  18. 18.Eggensperger, K., Lindauer, M., and Hutter, F. Pitfalls and best practices in algorithm configuration. pp. 861–893, 2019.
  19. 19.Eggensperger, K., Muller, P., Mallik, N., Feurer, M., Sass, R., Klein, A., Awad, N., Lindauer, M., and Hutter, F. Hpobench: A collection of reproducible multi-fidelity benchmark problems for HPO. In Vanschoren, J. and Yeung, S. (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks, 2021.
  20. 20.Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. Implementation matters in deep RL: A case study on PPO and TRPO. In 8th International Conference on Learning Representations, ICLR. OpenReview.net, 2020.
  21. 21.Flennerhag, S., Schroecker, Y., Zahavy, T., van Hasselt, H., Silver, D., and Singh, S. Bootstrapped meta-learning. In The Tenth International Conference on Learning Representations, ICLR. OpenReview.net, 2022.
  22. 22.Franke, J., Kohler, G., Biedenkapp, A., and Hutter, F. Sample-efficient automated deep reinforcement learning. In 9th International Conference on Learning Representations, ICLR. OpenReview.net, 2021.
  23. 23.Franke, J. K., Kohler, G., Biedenkapp, A., and Hutter, F. Sample-efficient automated deep reinforcement learning. arXiv:2009.01555 [cs.LG], 2020.
  24. 24.Freeman, C., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O. Brax - A differentiable physics engine for large scale rigid body simulation. In Vanschoren, J. and Yeung, S. (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, 2021.
  25. 25.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Dy, J. G. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML, volume 80 of Proceedings of Machine Learning Research, pp. 1856–1865. PMLR, 2018.
  26. 26.Hambro, E., Raileanu, R., Rothermel, D., Mella, V., Rocktaschel, T., Kuttler, H., and Murray, N. Dungeons and data: A large-scale nethack dataset. 2022.
  27. 27.Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. Deep reinforcement learning that matters. In McIlraith, S. and Weinberger, K. (eds.), Proceedings of the Conference on Artificial Intelligence (AAAI’18). AAAI Press, 2018.
  28. 28.Hsu, C., Mendler-Dunner, C., and Hardt, M. Revisiting design choices in proximal policy optimization. CoRR, abs/2009.10897, 2020.
  29. 29.Huang, S., Dossa, R., Ye, C., and Braga, J. Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms. CoRR, abs/2111.08819, 2021.
  30. 30.Hutter, F., Hoos, H., and Leyton-Brown, K. An efficient approach for assessing hyperparameter importance. In Xing, E. and Jebara, T. (eds.), Proceedings of the 31th International Conference on Machine Learning, (ICML’14), pp. 754–762. Omnipress, 2014.
  31. 31.Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K. Population based training of neural networks. arXiv:1711.09846 [cs.LG], 2017.
  32. 32.Kiran, M. and Ozyildirim, B. Hyperparameter tuning for deep reinforcement learning applications. CoRR, abs/2201.11182, 2022. URL https://arxiv.org/abs/2201.11182.
  33. 33.Li, A., Spyra, O., Perel, S., Dalibard, V., Jaderberg, M., Gu, C., Budden, D., Harley, T., and Gupta, P. A generalized framework for population based training. In Teredesai, A., Kumar, V., Li, Y., Rosales, R., Terzi, E., and Karypis, G. (eds.), Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD, pp. 1791–1799. ACM, 2019.
  34. 34.Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A. Hyperband: A novel bandit-based approach to hyperparameter optimization. 18(185):1–52, 2018.
  35. 35.Liaw, R., Liang, E., Nishihara, R., Moritz, P., Gonzalez, J., and Stoica, I. Tune: A research platform for distributed model selection and training. CoRR, abs/1807.05118, 2018.
  36. 36.Lindauer, M. and Hutter, F. Best practices for scientific research on neural architecture search. Journal of Machine Learning Research, 21:1–18, 2020.
  37. 37.Lindauer, M., Eggensperger, K., Feurer, M., Biedenkapp, A., Deng, D., Benjamins, C., Ruhkopf, T., Sass, R., and Hutter, F. SMAC3: A versatile bayesian optimization package for hyperparameter optimization. J. Mach. Learn. Res., 23:54:1–54:9, 2022.
  38. 38.Lu, C., Kuba, J., Letcher, A., Metz, L., de Witt, C., and Foerster, J. Discovered policy optimisation. CoRR, abs/2210.05639, 2022.
  39. 39.Makarova, A., Shen, H., Perrone, V., Klein, A., Faddoul, J., Krause, A., Seeger, M., and Archambeau, C. Automatic termination for hyperparameter optimization. In Guyon, I., Lindauer, M., van der Schaar, M., Hutter, F., and Garnett, R. (eds.), International Conference on Automated Machine Learning, AutoML, volume 188 of Proceedings of Machine Learning Research, pp. 7/1–21. PMLR, 2022.
  40. 40.Metz, L., Harrison, J., Freeman, C., Merchant, A., Beyer, L., Bradbury, J., Agrawal, N., Poole, B., Mordatch, I., Roberts, A., and Sohl-Dickstein, J. Velo: Training versatile learned optimizers by scaling up. CoRR, abs/2211.09760, 2022.
  41. 41.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  42. 42.Obando-Ceron, J. and Castro, P. Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML, volume 139 of Proceedings of Machine Learning Research, pp. 1373–1383. PMLR, 2021.
  43. 43.Parker-Holder, J., Nguyen, V., and Roberts, S. Provably efficient online hyperparameter optimization with population-based bandits. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS, 2020.
  44. 44.Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., and Lindauer, M. Automated reinforcement learning (autorl): A survey and open problems. J. Artif. Intell. Res., 74:517–568, 2022.
  45. 45.Paul, S., Kurin, V., and Whiteson, S. Fast efficient hyperparameter tuning for policy gradient methods. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS, pp. 4618–4628, 2019.
  46. 46.Pushak, Y. and Hoos, H. H. Automl loss landscapes. ACM Trans. Evol. Learn. Optim., 2(3):10:1–10:30, 2022.
  47. 47.Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N. Stable-baselines3: Reliable reinforcement learning implementations. J. Mach. Learn. Res., 22:268:1–268:8, 2021.
  48. 48.Raileanu, R. and Fergus, R. Decoupling value and policy for generalization in reinforcement learning. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML, volume 139 of Proceedings of Machine Learning Research, pp. 8787–8798. PMLR, 2021.
  49. 49.Raileanu, R. and Rocktaschel, T. RIDE: rewarding impact-driven exploration for procedurally-generated environments. In 8th International Conference on Learning Representations, ICLR. OpenReview.net, 2020.
  50. 50.Sass, R., Bergman, E., Biedenkapp, A., Hutter, F., and Lindauer, M. Deepcave: An interactive analysis tool for automated machine learning. CoRR, abs/2206.03493, 2022.
  51. 51.Schede, E., Brandt, J., Tornede, A., Wever, M., Bengs, V., Hullermeier, E., and Tierney, K. A survey of methods for automated algorithm configuration. J. Artif. Intell. Res., 75:425–487, 2022.
  52. 52.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv:1707.06347 [cs.LG], 2017.
  53. 53.Shala, G., Arango, S., Biedenkapp, A., Hutter, F., and Grabocka, J. Autorl-bench 1.0. In Workshop on Meta-Learning (MetaLearn@NeurIPS’22), 2022.
  54. 54.Storn, R. and Price, K. Differential evolution - A simple and efficient heuristic for global optimization over continuous spaces. J. Glob. Optim., 11(4):341–359, 1997.
  55. 55.Turner, R., Eriksson, D., McCourt, M., Kiili, J., Laaksonen, E., Xu, Z., and Guyon, I. Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020. CoRR, abs/2104.10201, 2021.
  56. 56.Wan, X., Lu, C., Parker-Holder, J., Ball, P., Nguyen, V., Ru, B., and Osborne, M. Bayesian generational population-based training. In Guyon, I., Lindauer, M., van der Schaar, M., Hutter, F., and Garnett, R. (eds.), International Conference on Automated Machine Learning, AutoML, volume 188 of Proceedings of Machine Learning Research, pp. 14/1–27. PMLR, 2022.
  57. 57.Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. Dueling network architectures for deep reinforcement learning. In Balcan, M. and Weinberger, K. (eds.), Proceedings of the 33rd International Conference on Machine Learning (ICML’17), volume 48, pp. 1995–2003. Proceedings of Machine Learning Research, 2016.
  58. 58.Xu, Z., van Hasselt, H., and Silver, D. Meta-gradient reinforcement learning. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS, pp. 2402–2413, 2018.
  59. 59.Xu, Z., van Hasselt, H., Hessel, M., Oh, J., Singh, S., and Silver, D. Meta-gradient reinforcement learning with an objective discovered online. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS, 2020.
  60. 60.Yadan, O. Hydra - a framework for elegantly configuring complex applications. Github, 2019. URL https://github.com/facebookresearch/hydra.
  61. 61.Zahavy, T., Xu, Z., Veeriah, V., Hessel, M., Oh, J., van Hasselt, H., Silver, D., and Singh, S. A self-tuning actor-critic algorithm. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS, 2020.
  62. 62.Zhang, B., Rajan, R., Pineda, L., Lambert, N., Biedenkapp, A., Chua, K., Hutter, F., and Calandra, R. On the importance of hyperparameter optimization for model-based reinforcement learning. In Banerjee, A. and Fukumizu, K. (eds.), The 24th International Conference on Artificial Intelligence and Statistics, AISTATS, volume 130 of Proceedings of Machine Learning Research, pp. 4015–4023. PMLR, 2021a.
  63. 63.Zhang, S. and Jiang, N. Towards hyperparameter-free policy selection for offline reinforcement learning. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS, pp. 12864–12875, 2021.
  64. 64.Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J., and Tian, Y. Noveld: A simple yet effective exploration criterion. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS, pp. 25217–25230, 2021b.
  65. 65.Zimmer, L., Lindauer, M., and Hutter, F. Auto-pytorch: Multi-fidelity metalearning for efficient and robust autodl. IEEE Trans. Pattern Anal. Mach. Intell., 43(9):3079–3090, 2021.

Citation

MLA
Eimer, T., et al. “Hyperparameters in Reinforcement Learning and How To Tune Them”. International Conference on Machine Learning, vol. 202, 2023, pp. 9104–49, https://proceedings.mlr.press/v202/eimer23a.html.
APA
Eimer, T., Lindauer, M., & Raileanu, R. (2023). Hyperparameters in Reinforcement Learning and How To Tune Them. International Conference on Machine Learning, 202, 9104–9149. https://proceedings.mlr.press/v202/eimer23a.html
Chicago
Eimer, T., M. Lindauer, and R. Raileanu. 2023. “Hyperparameters in Reinforcement Learning and How To Tune Them”. International Conference on Machine Learning 202: 9104–49. https://proceedings.mlr.press/v202/eimer23a.html.
Harvard
Eimer, T., Lindauer, M. and Raileanu, R. (2023) “Hyperparameters in Reinforcement Learning and How To Tune Them”, International Conference on Machine Learning. PMLR, pp. 9104–9149. Available at: https://proceedings.mlr.press/v202/eimer23a.html.
Vancouver
1. Eimer T, Lindauer M, Raileanu R (2023) Hyperparameters in Reinforcement Learning and How To Tune Them. In: International Conference on Machine Learning. PMLR, pp 9104–9149

BibTeX

@InProceedings{pmlr-v202-eimer23a,
  title = 	 {Hyperparameters in Reinforcement Learning and How To Tune Them},
  author =       {Eimer, Theresa and Lindauer, Marius and Raileanu, Roberta},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {9104--9149},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/eimer23a/eimer23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/eimer23a.html},
  abstract = 	 {In order to improve reproducibility, deep reinforcement learning (RL) has been adopting better scientific practices such as standardized evaluation metrics and reporting. However, the process of hyperparameter optimization still varies widely across papers, which makes it challenging to compare RL algorithms fairly. In this paper, we show that hyperparameter choices in RL can significantly affect the agent’s final performance and sample efficiency, and that the hyperparameter landscape can strongly depend on the tuning seed which may lead to overfitting. We therefore propose adopting established best practices from AutoML, such as the separation of tuning and testing seeds, as well as principled hyperparameter optimization (HPO) across a broad search space. We support this by comparing multiple state-of-the-art HPO tools on a range of RL algorithms and environments to their hand-tuned counterparts, demonstrating that HPO approaches often have higher performance and lower compute overhead. As a result of our findings, we recommend a set of best practices for the RL community, which should result in stronger empirical results with fewer computational costs, better reproducibility, and thus faster progress. In order to encourage the adoption of these practices, we provide plug-and-play implementations of the tuning algorithms used in this paper at https://github.com/facebookresearch/how-to-autorl.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/