Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Michal NaumanMichal BortkiewiczPiotr MilosTomasz TrzcinskiMateusz OstaszewskiMarek Cygan

article2024ICML57 citations

Reveals through an extensive evaluation of over 60 actor-critic variants that general neural network regularizers outperform domain-specific reinforcement learning modifications, enabling basic Soft Actor-Critic agents to achieve state-of-the-art sample efficiency across diverse continuous control tasks.

Listen

Off-policy reinforcement learning has become a central method for training autonomous control systems in robotics and complex industrial automation. To make these systems learn efficiently from limited interactions, practitioners frequently increase the frequency of learning updates relative to gathered data. However, intensive updating often introduces severe training instabilities, including value overestimation, model overfitting, and loss of learning plasticity (the capacity of a neural network to acquire new information). While researchers have proposed numerous specialized techniques to address these issues, prior studies typically evaluated interventions in narrow settings or within single simulation environments, leaving engineering teams uncertain about which design choices truly deliver robust, generalizable performance.

The main objective of the article is to evaluate how diverse regularization techniques interact and affect off-policy reinforcement learning performance. Specifically, it demonstrates whether broad, standard neural network regularizers can outperform domain-specific reinforcement learning methods across different task types and update frequencies.

To establish these insights, the researchers implemented 64 distinct agent configurations within the standard Soft Actor-Critic framework. They tested these agents across 14 continuous control tasks spanning two established benchmark suites: the DeepMind Control Suite (locomotion tasks) and MetaWorld (robotic manipulation tasks). The experimental matrix systematically combined three categories of interventions: critic regularizations designed to prevent value overestimation, network regularizations to control overfitting, and plasticity regularizations to maintain adaptability. The configurations were evaluated across both low and high update-frequency regimes using multiple random trials to ensure statistical reliability.

The study reveals four key findings. First, generic neural network regularizations (such as layer normalization and spectral normalization) and plasticity interventions (such as periodic network resets) substantially outperform reinforcement learning-specific critic modifications. Second, commonly used critic regularizations like Clipped Double Q-learning often impair performance when general network regularizations are present, causing severe degradation in complex manipulation tasks. Third, applying standard layer normalization and periodic resets enables a basic model-free agent to master challenging quadruped locomotion tasks (such as the Dog domain) that previously required complex model-based architectures. Fourth, the primary statistical predictors of failure are severe value overestimation and exploding gradient magnitudes, both of which are effectively mitigated by generic normalization techniques rather than specialized reinforcement learning loss functions.

These findings suggest that complex reinforcement learning systems can be substantially simplified while achieving higher stability and performance. Engineering teams can eliminate complicated, domain-specific value-clipping heuristics in favor of standard deep learning practices like layer normalization and weight resets. This shift reduces system complexity, mitigates deployment risks, and shortens development timelines for continuous control applications. Moreover, the divergent results observed between locomotion and manipulation benchmarks underscore that algorithms tuned solely on single benchmark suites risk significant failure when deployed to new operational domains.

Practitioners should prioritize incorporating layer normalization or spectral normalization alongside periodic network resets into their continuous control pipelines. Teams should also re-evaluate or remove specialized critic clipping mechanisms unless empirical testing confirms their necessity for a given task. Before committing to widespread architectural changes, organizations should establish broad evaluation suites containing both locomotion and manipulation benchmarks rather than relying on isolated test environments.

The conclusions of the article are subject to certain limitations. The empirical evaluations were confined to continuous control tasks relying on direct physical state measurements within the Soft Actor-Critic architecture. Preliminary tests on image-based inputs showed that these benefits do not directly transfer to vision-based models under low update frequencies. While confidence in the reported state-based continuous control results is high, practitioners should exercise caution and conduct targeted pilot tests before applying these recommendations to vision-driven agents or non-actor-critic frameworks.

arXiv: 2403.00514

No sufficiently relevant recommendations were found.

Cover for Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Abstract

Recent advancements in off-policy Reinforcement Learning (RL) have significantly improved sample efficiency, primarily due to the incorporation of various forms of regularization that enable more gradient update steps than traditional agents. However, many of these techniques have been tested in limited settings, often on tasks from single simulation benchmarks and against well-known algorithms rather than a range of regularization approaches. This limits our understanding of the specific mechanisms driving RL improvements. To address this, we implemented over 60 different off-policy agents, each integrating established regularization techniques from recent state-of-the-art algorithms. We tested these agents across 14 diverse tasks from 2 simulation benchmarks, measuring training metrics related to overestimation, overfitting, and plasticity loss — issues that motivate the examined regularization techniques. Our findings reveal that while the effectiveness of a specific regularization setup varies with the task, certain combinations consistently demonstrate robust and superior performance. Notably, a simple Soft Actor-Critic agent, appropriately regularized, reliably finds a better-performing policy within the training regime, which previously was achieved mainly through model-based approaches.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Overestimation
  • 2.2 Overfitting
  • 2.3 Plasticity
  • 3 Study Design
  • 4 Experiments
  • 4.1 Combination of interventions – First-order marginalization
  • 4.2 Combination of interventions – Second-order marginalization
  • 4.3 A closer look on Dog environment performance
  • 4.4 Correlation of Overestimation, Overfitting and Plasticity metrics with Performance
  • 5 Related Works
  • 6 Limitations
  • 7 Conclusions
  • References
  • A Details of experiments
  • B Architecture details
  • B.1 Hyperparameters
  • C Further Experiments
  • C.1 Third-order marginalization
  • C.2 Gradient Norm Analysis
  • C.3 Comparison of ReDO to other plasticity-inducing methods
  • C.4 Closer look on CDQL performance on Hopper and Quadruped environments
  • C.5 Regression and Spearman correlation analysis
  • C.6 Image-based DeepMind Control
  • C.7 Best combinations of intervention performance plots
  • C.8 Other

Knowls

  1. Knowl 1 — Factorial evaluation of SAC regularization across tasks and replay ratios

    experimental setup

    The study evaluated Soft Actor-Critic (SAC) on 14 proprioceptive continuous-control tasks: seven DeepMind Control Suite tasks (acrobot-swingup, hopper-hop, humanoid-walk, humanoid-run, dog-trot, dog-run, and quadruped-run) and seven MetaWorld tasks (Hammer, Push, Sweep, Coffee-Push, Stick-Pull, Reach, and Hand-Insert). Experiments used replay ratios (gradient updates per environment step) of 2 and 16, with 10 seeds per configuration. The design crossed three intervention families, allowing either no intervention or one of three choices from each family: critic regularization (Clipped Double Q-learning, Tactical Optimism and Pessimism, or Generalized Pessimism Learning), network regularization (layer normalization, spectral normalization, or weight decay), and plasticity regularization (full-parameter resets, Concatenated ReLU, or Sharpness-Aware Minimization). This yields 43=644^3=64 configurations and excludes combinations of multiple methods within the same family. Actor and critic were three-layer MLPs with 256-unit hidden layers and ReLU activations; network regularizers were applied only to the critic. Layer normalization was applied at each hidden layer, spectral normalization at the final hidden layer, and weight decay across layers. Final performance was the mean of the last 10 policy evaluations; reported interquartile means used 500 bootstrap samples.

  2. Knowl 2 — General network regularization is more robust than critic-specific regularization

    empirical result

    Across the tested SAC configurations, network and plasticity regularization were generally more robust and beneficial than critic regularization methods designed specifically for off-policy value estimation. In particular, adding critic regularization often reduced performance when a network or plasticity regularizer was already present. Tactical Optimism and Pessimism was an exception, performing comparatively well on DeepMind Control tasks, whereas Generalized Pessimism Learning was among the least robust choices on both benchmarks. These are aggregate empirical tendencies rather than claims that every critic regularizer fails on every task.

  3. Knowl 3 — The best network regularizer depends on the benchmark

    empirical result

    Layer normalization was especially effective on DeepMind Control Suite tasks: it ranked near the top across non-dog tasks and was particularly important on Dog-Trot and Dog-Run. It performed poorly on MetaWorld manipulation tasks. Spectral normalization was more effective on MetaWorld and offered more consistent performance across the two benchmarks when tasks were considered together. Weight decay generally performed poorly as a standalone addition to SAC, although pairing it with full-parameter resets could produce strong results.

  4. Knowl 4 — Full-parameter resets are a strong high-replay intervention

    empirical result

    At replay ratio 16, periodic full-parameter resets were among the most robust plasticity interventions and generally outperformed Concatenated ReLU and Sharpness-Aware Minimization. On DeepMind Control tasks, combining resets with layer normalization produced exceptional performance across critic-regularization choices. On MetaWorld, weight decay could become beneficial when paired with resets, despite its weak standalone performance. A comparison with ReDo found that full resets performed better overall, although both approaches reduced critic gradient norm, value overestimation, and dormant-neuron prevalence on MetaWorld and non-dog DeepMind Control tasks. ReDo was unstable on dog tasks in runs without layer or spectral normalization.

  5. Knowl 5 — Regularized model-free SAC achieves strong performance on dog control tasks

    empirical result

    On the challenging proprioceptive Dog-Trot and Dog-Run tasks, combinations featuring layer normalization were prominent among the best-performing agents; at replay ratio 16, layer normalization paired with full-parameter resets was especially effective. The authors report that this simple model-free SAC configuration found a policy that outperformed previously reported model-free results within the training regime, and came close to a model-based method using proprioceptive states, which performed slightly better. Their longer dog-task comparison trained for 4 million environment steps. The result is specific to the tested tasks, state inputs, and training budget.

  6. Knowl 6 — Operational measures of value overestimation, overfitting, and plasticity

    definition

    The study measured critic overestimation as the critic's state-action value minus an estimate of the on-policy value: b(s,a)=Qϕ(s,a)−Qπ(s,a)b(s,a)=Q_\phi(s,a)-Q^\pi(s,a), where ss is a state, aa an action, QϕQ_\phi the learned critic with parameters ϕ\phi, and QπQ^\pi the on-policy value estimated using five Monte Carlo rollouts. Overfitting was measured as the ratio of average temporal-difference (TD) error on held-out evaluation trajectories to average TD error on training data: oϕ=TDvalidation/TDtrainingo_\phi=\mathrm{TD}_{\mathrm{validation}}/\mathrm{TD}_{\mathrm{training}}. Plasticity loss was not represented by one direct measure; the study tracked the percentage of dormant neurons, the rank of critic penultimate-layer representations, critic weight L2 norm, and critic gradient norm as proxies.

  7. Knowl 7 — Different proxy metrics explain performance in different environments

    empirical result

    In the pooled analysis of runs, overestimation had the strongest reported relationship with return on MetaWorld and on DeepMind Control tasks excluding the dog environments. In dog environments, overestimation was not the best performance predictor despite being particularly high; critic gradient norm had the strongest monotonic relationship with return there. Dormant-neuron prevalence was negatively associated with performance on DeepMind Control tasks and was closely related to the rank of critic representations, although the pattern was not universal across tasks. Overfitting showed the greatest variation in its relationship with return across environments. The authors also observed that interventions aimed at one issue could shift proxies for other issues: full-parameter resets, for example, affected overestimation and overfitting measures as well as plasticity-related measures. They hypothesize that overestimation becomes more predictive once more fundamental plasticity problems are mitigated.

  8. Knowl 8 — High replay ratios expose large critic gradients

    empirical result

    Critic gradient norms were much larger in the replay-ratio-16 MetaWorld experiments than in DeepMind Control experiments, reaching differences of orders of magnitude. Dog environments also showed very large gradient norms, including at replay ratio 2. The study found a negative association between critic gradient norm and return that was most apparent in difficult settings, particularly at high replay ratios; the sign of the association was not uniform across all DeepMind Control tasks. Layer normalization was more robust than spectral normalization at mitigating gradient growth on DeepMind Control, especially in the dog tasks. These observations support treating gradient behavior as environment-dependent rather than as a universal performance proxy.

  9. Knowl 9 — Clipped Double Q-learning has task-dependent performance effects

    empirical result

    Clipped Double Q-learning (CDQ), which uses the minimum of two critic estimates, did not have a uniformly beneficial effect. It could substantially reduce performance on MetaWorld tasks and on DeepMind Control Hopper-Hop; in Hopper-Hop, adding resets or layer normalization did not remove the adverse effect. CDQ also had a negative effect on Quadruped-Run, although additional regularization could mitigate it there. By contrast, it was effective in some DeepMind Control locomotion settings. Thus, its usefulness in the tested actor-critic agents depended on the task and benchmark.

  10. Knowl 10 — A small image-based DrQ study did not establish transfer of state-based findings

    limitation

    The study also tested four DrQ variants on image observations: baseline DrQ, DrQ with critic layer normalization, DrQ with full-parameter resets every 200,000 environment steps, and DrQ with both interventions. The replay ratio was 1, with three seeds per task; non-humanoid tasks ran for 1 million frames and humanoid tasks for 3 million frames. The reported final returns were:

Coverage note — The image-based DrQ knowl’s table is omitted here because the source’s table reports only four of the six tested tasks and the small pilot does not support a general transfer conclusion. Detailed per-task regression plots and individual training curves are also omitted because they do not add a distinct finding beyond the environment-dependent relationships summarized above.

References

  1. 1.Abbas, Z., Zhao, R., Modayil, J., White, A., and Machado, M. C. Loss of plasticity in continual deep reinforcement learning. arXiv preprint arXiv: 2303.07507, 2023.
  2. 2.Achille, A., Rovere, M., and Soatto, S. Critical learning periods in deep neural networks. arXiv preprint arXiv: 1711.08856, 2017.
  3. 3.Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M. Deep reinforcement learning at the edge of the statistical precipice. Advances in neural information processing systems, 34:29304–29320, 2021.
  4. 4.Andrychowicz, M., Raichuk, A., Stanczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., et al. What matters in on-policy reinforcement learning? a large-scale empirical study. In ICLR 2021-Ninth International Conference on Learning Representations, 2021.
  5. 5.Ash, J. and Adams, R. P. On warm-starting neural network training. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 3884–3894. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/288cd2567953f06e460a33951f55daaf-Paper.pdf.
  6. 6.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  7. 7.Ball, P. J., Smith, L., Kostrikov, I., and Levine, S. Efficient online reinforcement learning with offline data. arXiv preprint arXiv: 2302.02948, 2023.
  8. 8.Bjorck, J., Gomes, C. P., and Weinberger, K. Q. Towards deeper deep reinforcement learning with spectral normalization. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 8242–8255, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/4588e674d3f0faf985047d4c3f13ed0d-Abstract.html.
  9. 9.Brock, A., Donahue, J., and Simonyan, K. Large scale gan training for high fidelity natural image synthesis. International Conference on Learning Representations, 2018.
  10. 10.Cetin, E. and Celiktutan, O. Learning pessimism for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 6971–6979, 2023.
  11. 11.Chen, X., Wang, C., Zhou, Z., and Ross, K. W. Randomized ensembled double q-learning: Learning fast without a model. In International Conference on Learning Representations, 2020.
  12. 12.Ciosek, K., Vuong, Q., Loftin, R., and Hofmann, K. Better exploration with optimistic actor critic. Advances in Neural Information Processing Systems, 32, 2019.
  13. 13.Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., Donner, C., Fritz, L., Galperti, C., Huber, A., Keeling, J., Tsimpoukelli, M., Kay, J., Merle, A., Moret, J.-M., Noury, S., Pesamosca, F., Pfau, D., Sauter, O., Sommariva, C., Coda, S., Duval, B., Fasoli, A., Kohli, P., Kavukcuoglu, K., Hassabis, D., and Riedmiller, M. Magnetic control of tokamak plasmas through deep reinforcement learning. Nature, 602 (7897):414—419, February 2022. ISSN 0028-0836. doi: 10.1038/s41586-021-04301-9.
  14. 14.Dohare, S., Sutton, R. S., and Mahmood, A. R. Continual backprop: Stochastic gradient descent with persistent randomness. arXiv preprint arXiv: 2108.06325, 2021.
  15. 15.D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A. Sample-efficient reinforcement learning by breaking the replay ratio barrier. In The Eleventh International Conference on Learning Representations, 2022.
  16. 16.Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-aware minimization for efficiently improving generalization. arXiv preprint arXiv: 2010.01412, 2020.
  17. 17.Fujimoto, S., Hoof, H., and Meger, D. Addressing function approximation error in actor-critic methods. In International conference on machine learning, pp. 1587–1596. PMLR, 2018.
  18. 18.Gogianu, F., Berariu, T., Rosca, M. C., Clopath, C., Busoniu, L., and Pascanu, R. Spectral normalisation for deep reinforcement learning: an optimisation perspective. In International Conference on Machine Learning, pp. 3734–3744. PMLR, 2021.
  19. 19.Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. Reinforcement learning with deep energy-based policies. In International conference on machine learning, pp. 1352–1361. PMLR, 2017.
  20. 20.Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905, 2018.
  21. 21.Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023.
  22. 22.Hansen, N., Wang, X., and Su, H. Temporal difference learning for model predictive control. In International Conference on Machine Learning, PMLR, 2022.
  23. 23.Hiraoka, T., Imagawa, T., Hashimoto, T., Onishi, T., and Tsuruoka, Y. Dropout q-functions for doubly efficient reinforcement learning. In International Conference on Learning Representations, 2021.
  24. 24.Hussing, M., Voelcker, C., Gilitschenski, I., Farahmand, A.-m., and Eaton, E. Dissecting deep rl with high update ratios: Combatting value overestimation and divergence. arXiv preprint arXiv:2403.05996, 2024.
  25. 25.Janner, M., Fu, J., Zhang, M., and Levine, S. When to trust your model: Model-based policy optimization. Advances in Neural Information Processing Systems, 32, 2019.
  26. 26.Ji, T., Luo, Y., Sun, F., Zhan, X., Zhang, J., and Xu, H. Seizing serendipity: Exploiting the value of past success in off-policy actor-critic, 2024.
  27. 27.Kumar, A., Agarwal, R., Ghosh, D., and Levine, S. Implicit under-parameterization inhibits data-efficient deep reinforcement learning. arXiv preprint arXiv: 2010.14498, 2020.
  28. 28.Kumar, S., Marklund, H., and Roy, B. V. Maintaining plasticity in continual learning via regenerative regularization. arXiv preprint arXiv: 2308.11958, 2023.
  29. 29.Lee, H., Cho, H., Kim, H., Gwak, D., Kim, J., Choo, J., Yun, S.-Y., and Yun, C. Plastic: Improving input and label plasticity for sample efficient reinforcement learning. NEURIPS, 2023.
  30. 30.Li, Q., Kumar, A., Kostrikov, I., and Levine, S. Efficient deep reinforcement learning requires regulating overfitting. In The Eleventh International Conference on Learning Representations, 2022.
  31. 31.Liu, Z., Li, X., Kang, B., and Darrell, T. Regularization matters in policy optimization-an empirical study on continuous control. In International Conference on Learning Representations, 2020.
  32. 32.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. International Conference on Learning Representations, 2017.
  33. 33.Lyle, C., Rowland, M., and Dabney, W. Understanding and preventing capacity loss in reinforcement learning. International Conference on Learning Representations, 2022. doi: 10.48550/arXiv.2204.09560.
  34. 34.Lyle, C., Zheng, Z., Nikishin, E., Avila Pires, B., Pascanu, R., and Dabney, W. Understanding plasticity in neural networks. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 23190–23211. PMLR, 23-29 Jul 2023.
  35. 35.Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y. Spectral normalization for generative adversarial networks. International Conference on Learning Representations, 2018.
  36. 36.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015.
  37. 37.Moskovitz, T., Parker-Holder, J., Pacchiano, A., Arbel, M., and Jordan, M. Tactical optimism and pessimism for deep reinforcement learning. Advances in Neural Information Processing Systems, 34:12849–12863, 2021.
  38. 38.Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A. The primacy bias in deep reinforcement learning. In International conference on machine learning, pp. 16828–16847. PMLR, 2022.
  39. 39.Nikishin, E., Oh, J., Ostrovski, G., Lyle, C., Pascanu, R., Dabney, W., and Barreto, A. Deep reinforcement learning with plasticity injection. In Workshop on Reincarnating Reinforcement Learning at ICLR 2023, 2023. URL https://openreview.net/forum?id=O9cJADBZT1.
  40. 40.OpenAI, :, Berner, C., et al. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv: 1912.06680, 2019.
  41. 41.Puterman, M. L. Markov decision processes: discrete stochastic dynamic programming. John Wiley & Sons, 2014.
  42. 42.Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S. Bigger, better, faster: Human-level atari with human-level efficiency. In International Conference on Machine Learning, pp. 30365–30380. PMLR, 2023.
  43. 43.Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. Deterministic policy gradient algorithms. In International conference on machine learning, pp. 387–395. PMLR, 2014.
  44. 44.Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. Mastering the game of go without human knowledge. nature, 550(7676):354–359, 2017.
  45. 45.Smith, L., Kostrikov, I., and Levine, S. A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning. arXiv preprint arXiv:2208.07860, 2022.
  46. 46.Sokar, G., Agarwal, R., Castro, P. S., and Evci, U. The dormant neuron phenomenon in deep reinforcement learning. International Conference on Machine Learning, 2023. doi: 10.48550/arXiv.2302.12902.
  47. 47.Sutton, R. The bitter lesson. Incomplete Ideas (blog), 13(1), 2019.
  48. 48.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  49. 49.Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al. Deepmind control suite. arXiv preprint arXiv:1801.00690, 2018.
  50. 50.Thrun, S. and Schwartz, A. Issues in using function approximation for reinforcement learning. In Proceedings of the 1993 connectionist models summer school, pp. 255–263. Psychology Press, 2014.
  51. 51.Yarats, D., Kostrikov, I., and Fergus, R. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In International conference on learning representations, 2020.
  52. 52.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Kaelbling, L. P., Kragic, D., and Sugiura, K. (eds.), 3rd Annual Conference on Robot Learning, CoRL 2019, Osaka, Japan, October 30 - November 1, 2019, Proceedings, volume 100 of Proceedings of Machine Learning Research, pp. 1094–1100. PMLR, 2019. URL http://proceedings.mlr.press/v100/yu20a.html.
  53. 53.Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A. Self-attention generative adversarial networks. arXiv preprint arXiv: 1805.08318, 2018.
  54. 54.Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al. Maximum entropy inverse reinforcement learning. In Aaai, volume 8, pp. 1433–1438. Chicago, IL, USA, 2008.

Citation

MLA
Nauman, M., et al. “Overestimation, Overfitting, and Plasticity in Actor-Critic: The Bitter Lesson of Reinforcement Learning”. arXiv, 2024, http://arxiv.org/abs/2403.00514v2.
APA
Nauman, M., Bortkiewicz, M., Miłoś, P., Trzciński, T., Ostaszewski, M., & Cygan, M. (2024). Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning. arXiv. http://arxiv.org/abs/2403.00514v2
Chicago
Nauman, M., M. Bortkiewicz, P. Miłoś, T. Trzciński, M. Ostaszewski, and M. Cygan. 2024. “Overestimation, Overfitting, and Plasticity in Actor-Critic: The Bitter Lesson of Reinforcement Learning”. arXiv. http://arxiv.org/abs/2403.00514v2.
Harvard
Nauman, M. et al. (2024) “Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.00514v2.
Vancouver
1. Nauman M, Bortkiewicz M, Miłoś P, Trzciński T, Ostaszewski M, Cygan M (2024) Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning. arXiv

BibTeX

@article{nauman2024overestimation,
  title = {Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning},
  author = {Nauman, Michal and Bortkiewicz, Michał and Miłoś, Piotr and Trzciński, Tomasz and Ostaszewski, Mateusz and Cygan, Marek},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.00514v2},
  eprint = {2403.00514}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/