Is Value Learning Really the Main Bottleneck in Offline RL?

Seohong ParkKevin FransSergey LevineAviral Kumar

article2024NeurIPS60 citations

Reveals that offline reinforcement learning performance is primarily limited by policy extraction objectives and test-time policy generalization rather than imperfect value learning, providing practical test-time improvement techniques to bridge this gap.

Listen

Data-driven control algorithms often struggle to perform reliably when trained entirely on historical datasets. While offline reinforcement learning theoretically holds an advantage over standard imitation learning by learning from suboptimal data through value functions, it frequently underperforms in practical settings. Historically, researchers have attributed this shortcoming to inaccurate value estimation. The article addresses the root causes of this underperformance to establish whether learning the value function is truly the primary bottleneck holding back offline reinforcement learning.

The main objective of the article is to systematically evaluate how value estimation, policy extraction, and test-time generalization limit algorithm performance and scalability. To do this, the authors conducted an extensive empirical study spanning over 15,000 experimental runs across eight diverse robotics, manipulation, and continuous-control environments. By decoupling value estimation from policy training, they analyzed how varying the quality, quantity, and coverage of training data affected different algorithmic components.

The findings reveal that policy extraction and test-time generalization—rather than value function estimation—are the dominant bottlenecks in offline reinforcement learning. First, the method chosen to extract a policy from a learned value function significantly dictates success; behavior-constrained policy gradient methods outperformed or matched popular value-weighted behavioral cloning methods in 15 out of 16 tested settings. In complex environments, gradient-based extraction achieved performance scores near 191 to 193 compared to under 100 for weighted cloning. Second, existing algorithms already optimize well on in-distribution training data, but they fail during deployment when encountering novel, out-of-distribution states. Third, increasing dataset state coverage with exploratory, noisy actions produced substantially better deployment policies than using smaller, highly optimal datasets.

These insights challenge the prevailing belief that improving value function estimation is the most important path forward. Instead, conventional policy extraction methods often fail to fully exploit learned value functions due to severe sample inefficiency and overfitting. Furthermore, because failures occur primarily when agents encounter unseen states at deployment time, relying solely on conservative training objectives is insufficient to ensure dependable real-world performance.

To improve offline reinforcement learning in practice, practitioners should prioritize collecting datasets with high state coverage rather than focusing solely on clean, expert demonstrations. Decision-makers should also deploy behavior-constrained gradient methods instead of weighted imitation objectives and integrate test-time policy improvement techniques—such as on-the-fly action adjustment or test-time updates—to steer actions effectively during execution. The conclusions are well-supported across continuous-control environments, though caution is warranted for discrete-action tasks or settings with multiple equally optimal actions, where the study's proxy metrics may have limited applicability.

arXiv: 2406.09329

No sufficiently relevant recommendations were found.

Cover for Is Value Learning Really the Main Bottleneck in Offline RL?

Abstract

While imitation learning requires access to high-quality data, offline reinforcement learning (RL) should, in principle, perform similarly or better with substantially lower data quality by using a value function. However, current results indicate that offline RL often performs worse than imitation learning, and it is often unclear what holds back the performance of offline RL. Motivated by this observation, we aim to understand the bottlenecks in current offline RL algorithms. While poor performance of offline RL is typically attributed to an imperfect value function, we ask: is the main bottleneck of offline RL indeed in learning the value function, or something else? To answer this question, we perform a systematic empirical study of (1) value learning, (2) policy extraction, and (3) policy generalization in offline RL problems, analyzing how these components affect performance. We make two surprising observations. First, we find that the choice of a policy extraction algorithm significantly affects the performance and scalability of offline RL, often more so than the value learning objective. For instance, we show that common value-weighted behavioral cloning objectives (e.g., AWR) do not fully leverage the learned value function, and switching to behavior-constrained policy gradient objectives (e.g., DDPG+BC) often leads to substantial improvements in performance and scalability. Second, we find that a big barrier to improving offline RL performance is often imperfect policy generalization on test-time states out of the support of the training data, rather than policy learning on in-distribution states. We then show that the use of suboptimal but high-coverage data or test-time policy training techniques can address this generalization issue in practice. Specifically, we propose two simple test-time policy improvement methods and show that these methods lead to better performance.

Knowls

  1. Knowl 1 — Three bottlenecks in offline reinforcement learning

    definition

    The paper decomposes offline reinforcement learning performance into three potential bottlenecks: (B1) inaccurate value-function estimation from the offline dataset, (B2) imperfect extraction of a policy from the learned value function, and (B3) poor policy generalization to states encountered during evaluation but insufficiently represented in the training data. The paper’s central empirical hypothesis is that B2 and B3 often constrain performance more strongly than B1, contrary to the common emphasis on value-function quality.

  2. Knowl 2 — Policy extraction objectives compared in the study

    model/method

    The study compares three ways to extract a continuous-action policy from a learned value function QQ using an offline dataset DD of state-action pairs (s,a)(s,a). Here, π(a∣s)\pi(a\mid s) is the learned policy density, V(s)V(s) is a state-value baseline, πβ(a∣s)\pi_\beta(a\mid s) is a behavior-cloning policy, α\alpha controls the strength or temperature of behavior regularization, and NN is the number of sampled candidate actions.

    Weighted behavioral cloning, instantiated as advantage-weighted regression (AWR), maximizes

    JAWR(π)=E(s,a)∼D[exp⁡(α(Q(s,a)−V(s)))log⁡π(a∣s)].J_{\mathrm{AWR}}(\pi)=\mathbb{E}_{(s,a)\sim D}\left[\exp\left(\alpha\left(Q(s,a)-V(s)\right)\right)\log \pi(a\mid s)\right].

    Behavior-constrained policy gradient, instantiated as DDPG+BC, maximizes

    JDDPG+BC(π)=E(s,a)∼D[Q(s,μπ(s))+αlog⁡π(a∣s)],J_{\mathrm{DDPG+BC}}(\pi)=\mathbb{E}_{(s,a)\sim D}\left[Q\left(s,\mu_\pi(s)\right)+\alpha\log \pi(a\mid s)\right],

    where μπ(s)=Ea∼π(⋅∣s)[a]\mu_\pi(s)=\mathbb{E}_{a\sim\pi(\cdot\mid s)}[a] is the policy’s mean action.

    Sampling-based action selection, instantiated as SfBC, samples a1,…,aNa_1,\ldots,a_N independently from πβ(⋅∣s)\pi_\beta(\cdot\mid s) and selects

    π(s)=arg max⁡a∈{a1,…,aN}  Q(s,a).\pi(s)=\underset{a\in\{a_1,\ldots,a_N\}}{\operatorname{arg\,max}}\;Q(s,a).

    The value-learning methods are trained separately from policy extraction: IQL estimates an optimal value function using in-sample expectile regression, SARSA estimates a behavior-policy value function through temporal-difference evaluation, and CRL estimates a behavior-policy goal-conditioned value function through contrastive learning.

  3. Knowl 3 — Data-scaling matrices isolate value learning from policy extraction

    experimental setup

    To distinguish value-learning limitations from policy-extraction limitations, the paper uses decoupled offline RL procedures: a value function is trained first without a policy, and multiple policies are then extracted from that same fixed value function. The amount of data used for value learning and the amount used for policy extraction are varied independently. Each resulting matrix has value-data size on the horizontal axis and policy-data size on the vertical axis. Vertical color gradients indicate that performance is mainly policy-data-limited, horizontal gradients indicate value-data limitation, and diagonal gradients indicate that both matter.

    The study combines IQL with AWR, DDPG+BC, or SfBC; combines SARSA with those extractors on non-goal-conditioned tasks; and uses CRL instead of SARSA on goal-conditioned tasks. It evaluates eight environments spanning locomotion, manipulation, goal-conditioned control, and pixel-based control: gc-antmaze-large, antmaze-large, d4rl-hopper, d4rl-walker2d, exorl-walker, exorl-cheetah, kitchen, and gc-roboverse. The experiments aggregate 15,488 runs, using eight random seeds per matrix cell. Agents are trained for 10610^6 gradient steps, except for gc-roboverse, which uses 5×1055\times10^5 steps; performance is evaluated every 10510^5 steps with 50 rollouts and averaged over the final three evaluations. Policy-extraction hyperparameters are individually tuned for each matrix cell.

  4. Knowl 4 — Policy extraction often matters more than the value-learning objective

    data/table

    Across the data-scaling experiments, the choice of policy extractor frequently changes performance and scaling more than the choice of value-learning method, even when all extractors use the same learned value function. DDPG+BC is better than or at least as good as AWR in 15 of the 16 settings according to the paper’s aggregate comparison. The scores below are aggregated over each full value-data/policy-data matrix and eight random seeds; the reported uncertainty is the standard deviation.

    Task and value learner AWR DDPG+BC SfBC
    gc-antmaze-large (IQL) 51±251\pm2 58±258\pm2 58±158\pm1
    gc-antmaze-large (CRL) 37±237\pm2 58±258\pm2 51±251\pm2
    antmaze-large (IQL) 12±212\pm2 17±417\pm4 24±324\pm3
    antmaze-large (SARSA) 0±00\pm0 0±00\pm0 0±00\pm0
    kitchen (IQL) 80±180\pm1 86±186\pm1 75±175\pm1
    kitchen (SARSA) 79±179\pm1 83±183\pm1 73±173\pm1
    exorl-walker (IQL) 99±199\pm1 191±6191\pm6 140±1140\pm1
    exorl-walker (SARSA) 94±094\pm0 193±5193\pm5 125±1125\pm1
    exorl-cheetah (IQL) 71±171\pm1 101±2101\pm2 77±277\pm2
    exorl-cheetah (SARSA) 78±178\pm1 131±3131\pm3 89±189\pm1
    d4rl-hopper (IQL) 53±153\pm1 52±352\pm3 43±143\pm1
    d4rl-hopper (SARSA) 56±156\pm1 61±361\pm3 50±250\pm2
    d4rl-walker2d (IQL) 73±173\pm1 81±181\pm1 68±168\pm1
    d4rl-walker2d (SARSA) 79±079\pm0 84±084\pm0 81±181\pm1
    gc-roboverse (IQL) 23±223\pm2 20±220\pm2 14±214\pm2
    gc-roboverse (CRL) 13±113\pm1 16±216\pm2 15±115\pm1

    The strongest gaps occur on the highly suboptimal, diverse ExORL datasets: DDPG+BC reaches 191±6191\pm6 versus 99±199\pm1 for AWR with IQL on exorl-walker, and 101±2101\pm2 versus 71±171\pm1 on exorl-cheetah. AWR’s data-scaling matrices are generally policy-limited or jointly limited, indicating that weighted cloning does not fully exploit additional value-learning data. DDPG+BC more often exhibits value-limited scaling, meaning that improvements in the value function translate more directly into policy performance.

  5. Knowl 5 — Why behavior-constrained policy gradients outperform AWR

    empirical result

    The paper identifies three mechanisms that explain the superior scaling of DDPG+BC over AWR in many continuous-control experiments.

    First, AWR is purely mode-covering: it reweights dataset actions and therefore keeps learned actions within the convex hull of observed actions. DDPG+BC combines mode-covering behavioral cloning with mode-seeking first-order maximization of QQ, allowing the policy to hill-climb the learned value function and extrapolate moderately beyond the dataset while remaining behavior-regularized. Policies learned by DDPG+BC consequently produce a wider action range than AWR on exorl-walker.

    Second, AWR’s exponential advantage weights can make only a small fraction of the dataset contribute substantially to learning. Large temperatures allow a few high-advantage transitions to dominate the objective, reducing the effective sample size and causing overfitting. In the paper’s low-data gc-antmaze-large experiment, high-temperature AWR shows a widening training/validation policy-loss gap, whereas DDPG+BC does not show the same overfitting pattern.

    Third, every AWR coefficient multiplying log⁡π(a∣s)\log\pi(a\mid s) is positive. With limited action coverage, an expressive policy can increase the probability of every observed state-action pair, including poor actions, and can memorize the dataset; this can make AWR behave similarly to unweighted behavioral cloning. The paper therefore recommends behavior-constrained policy gradients rather than value-weighted behavioral cloning when the goal is to make policy performance scale with additional value data.

    The dependence on the behavior-regularization coefficient also differs. On gc-antmaze-large, AWR remains policy-bounded across tested coefficients, while DDPG+BC has a policy-bounded regime for strong regularization and a value-bounded regime for weak regularization. An intermediate DDPG+BC coefficient of α=1.0\alpha=1.0 performs strongly across both regimes.

  6. Knowl 6 — Training, validation, and evaluation policy-error metrics

    model/method

    To measure policy generalization, the paper compares the learned policy π\pi with an approximately optimal reference policy π∗\pi^*. Let DtrainD_{\mathrm{train}} be the offline training-state distribution, DvalD_{\mathrm{val}} an held-out validation-state distribution, and pπp_\pi the state distribution induced by rolling out π\pi during evaluation. For vector-valued actions, the paper uses the squared Euclidean error between the learned and reference actions:

    MSEtrain=Es∼Dtrain[∥π(s)−π∗(s)∥2].\mathrm{MSE}_{\mathrm{train}}=\mathbb{E}_{s\sim D_{\mathrm{train}}}\left[\left\|\pi(s)-\pi^*(s)\right\|^2\right]. MSEval=Es∼Dval[∥π(s)−π∗(s)∥2].\mathrm{MSE}_{\mathrm{val}}=\mathbb{E}_{s\sim D_{\mathrm{val}}}\left[\left\|\pi(s)-\pi^*(s)\right\|^2\right]. MSEeval=Es∼pπ[∥π(s)−π∗(s)∥2].\mathrm{MSE}_{\mathrm{eval}}=\mathbb{E}_{s\sim p_\pi}\left[\left\|\pi(s)-\pi^*(s)\right\|^2\right].

    The reference policy is used only for analysis and visualization. Training and validation MSEs measure accuracy on states drawn from the offline-data distribution, whereas evaluation MSE measures accuracy on the states that the learned policy actually visits at test time. The paper treats these quantities as practical proxies for policy accuracy rather than universally valid performance measures.

  7. Knowl 7 — Test-time policy generalization is often the dominant performance bottleneck

    empirical result

    The paper studies offline-to-online IQL on six tasks—antmaze-medium, antmaze-large, kitchen, adroit-pen, adroit-hammer, and adroit-door—using eight random seeds and tracking the three policy-error metrics during additional online interaction. In many tasks, online training improves evaluation MSE while training MSE and validation MSE remain nearly flat. Evaluation MSE is also the most predictive of return among the three metrics.

    These results indicate that the tested offline RL policies are often already close to their best achievable accuracy on states represented by the offline dataset. Their returns are instead limited by errors on the policy-induced test-time state distribution, where the policy encounters novel states. Thus, additional online data helps primarily by correcting behavior on states reached during deployment, not by improving policy accuracy on the original training or validation distributions.

  8. Knowl 8 — High state coverage can be more valuable than high data optimality

    empirical result

    The paper compares datasets collected from expert policies with different action-noise levels, thereby trading off state coverage against action optimality. In experiments with IQL on gc-antmaze-large and adroit-pen, datasets with more exploratory, lower-optimality actions generally produce better offline RL performance than higher-optimality datasets with narrower coverage.

    The result is consistent with the policy-generalization analysis: the main difficulty in these experiments is often not extracting a useful policy from suboptimal actions, but ensuring that the learned policy is accurate on the states it will visit at evaluation. The paper therefore recommends prioritizing coverage—especially around states likely to be visited by an optimal policy—when collecting additional offline data. It also finds that value-gradient-based extraction such as DDPG+BC is important in the high-coverage regime, because AWR can fail to exploit the improved value function in low-data settings.

  9. Knowl 9 — On-the-fly policy extraction with value gradients

    algorithm

    On-the-fly policy extraction (OPEX) improves a fixed offline policy during evaluation without changing its training procedure. Its inputs are a frozen learned value function QQ, an offline policy π\pi, a test-time state ss, and a step-size hyperparameter β\beta. At each evaluation state, sample an action a∼π(⋅∣s)a\sim\pi(\cdot\mid s) and make one value-gradient update:

    a←a+β∇aQ(s,a).a\leftarrow a+\beta\nabla_a Q(s,a).

    The updated action is executed in the environment. The operation is repeated independently at every test-time state during the rollout; the value function and policy parameters remain unchanged. The gradient step moves the sampled action in the direction that locally increases the learned value. In the experiments, the tested β\beta values were 0.30.3 for antmaze, 0.00030.0003 for kitchen, 0.030.03 for d4rl-hopper, 0.10.1 for d4rl-walker2d, and 11 for exorl-walker and exorl-cheetah. OPEX requires only this additional action-update line at evaluation time and does not require offline retraining.

  10. Knowl 10 — Test-time training distills the value function on visited states

    algorithm

    Test-time training (TTT) updates policy parameters during evaluation while keeping both the learned offline value function QQ and the original offline policy πoff\pi_{\mathrm{off}} fixed. Let pπp_\pi denote the state distribution generated by the current evaluation policy and let D∪pπ(⋅)D\cup p_\pi(\cdot) denote a mixture of offline-data states and evaluation states. TTT repeatedly optimizes the current policy π\pi using

    JTTT(π)=Es,a∼D∪pπ(⋅)[Q(s,μπ(s))−βDKL(πoff ∥ π)],J_{\mathrm{TTT}}(\pi)=\mathbb{E}_{s,a\sim D\cup p_\pi(\cdot)}\left[Q\left(s,\mu_\pi(s)\right)-\beta D_{\mathrm{KL}}\left(\pi_{\mathrm{off}}\,\|\,\pi\right)\right],

    where μπ(s)=Ea∼π(⋅∣s)[a]\mu_\pi(s)=\mathbb{E}_{a\sim\pi(\cdot\mid s)}[a], DKLD_{\mathrm{KL}} is the Kullback–Leibler divergence, and β\beta controls how strongly the current policy is kept near the offline policy. The procedure collects test-time states, performs policy-gradient updates on the objective, and continues this process as more rollout data arrive. In the experiments, TTT uses a learning rate of 0.000030.00003 and up to 2×1062\times10^6 total policy-gradient steps; tested β\beta values are 0.30.3 for antmaze, 55 for kitchen, 0.50.5 for d4rl-hopper, 0.30.3 for d4rl-walker2d, and 0.010.01 for exorl-walker and exorl-cheetah. TTT and OPEX improve over vanilla IQL and non-gradient SfBC on many of the eight evaluated tasks, often by substantial margins, by directly improving policy accuracy on test-time states.

Coverage note — The appendix’s alternative state-representation experiment and detailed value-learning losses were omitted because they are secondary analyses rather than load-bearing parts of the paper’s main bottleneck study.

References

  1. 1.Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song. Uncertainty-based offline reinforcement learning with diversified q-ensemble. In Neural Information Processing Systems (NeurIPS), 2021.
  2. 2.Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. In Neural Information Processing Systems (NeurIPS), 2017.
  3. 3.Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. ArXiv, abs/1607.06450, 2016.
  4. 4.Suneel Belkhale, Yuchen Cui, and Dorsa Sadigh. Data quality in imitation learning. In Neural Information Processing Systems (NeurIPS), 2023.
  5. 5.David Brandfonbrener, William F. Whitney, Rajesh Ranganath, and Joan Bruna. Offline rl without off-policy evaluation. In Neural Information Processing Systems (NeurIPS), 2021.
  6. 6.Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. Exploration by random network distillation. In International Conference on Learning Representations (ICLR), 2019.
  7. 7.Huayu Chen, Cheng Lu, Chengyang Ying, Hang Su, and Jun Zhu. Offline reinforcement learning via high-fidelity generative behavior modeling. In International Conference on Learning Representations (ICLR), 2023.
  8. 8.Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal. Adversarially trained actor critic for offline reinforcement learning. In International Conference on Machine Learning (ICML), 2022.
  9. 9.Open X-Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu, Charlotte Le, Chelsea Finn, Chen Wang, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Driess, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Foster, Fangchen Liu, Federico Ceola, Fei Xia, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su, Hao-Shu Fang, Haochen Shi, Henghui Bao, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim, Jaimyn Drake, Jan Peters, Jan Schneider, Jasmine Hsu, Jeannette Bohg, Jeffrey Bingham, Jeffrey Wu, Jensen Gao, Jiaheng Hu, Jiajun Wu, Jialin Wu, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jingyun Yang, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kaiyuan Wang, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Lin, Kevin Zhang, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Lawrence Yunliang Chen, Lerrel Pinto, Li Fei-Fei, Liam Tan, Linxi "Jim" Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Ning Liu, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R Sanketi, Patrick "Tree" Miller, Patrick Yin, Paul Wohlhart, Peng Xu, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Mart’in-Mart’in, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Ruohan Zhang, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Shan Lin, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiaolong Wang, Xinghao Zhu, Xinyang Geng, Xiyuan Liu, Xu Liangwei, Xuanlin Li, Yao Lu, Yecheng Jason Ma, Yejin Kim, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Yilin Wu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yue Cao, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zhuo Xu, Zichen Jeff Cui, Zichen Zhang, and Zipeng Lin. Open x-embodiment: Robotic learning datasets and rt-x models. In IEEE International Conference on Robotics and Automation (ICRA), 2024.
  10. 10.Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine. Rvs: What is essential for offline rl via supervised learning? In International Conference on Learning Representations (ICLR), 2022.
  11. 11.Benjamin Eysenbach, Tianjun Zhang, Ruslan Salakhutdinov, and Sergey Levine. Contrastive learning as goal-conditioned reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2022.
  12. 12.Justin Fu, Aviral Kumar, Ofir Nachum, G. Tucker, and Sergey Levine. D4rl: Datasets for deep data-driven reinforcement learning. ArXiv, abs/2004.07219, 2020.
  13. 13.Yuwei Fu, Di Wu, and Benoît Boulet. A closer look at offline rl agents. In Neural Information Processing Systems (NeurIPS), 2022.
  14. 14.Scott Fujimoto and Shixiang Shane Gu. A minimalist approach to offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2021.
  15. 15.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning (ICML), 2019.
  16. 16.Scott Fujimoto, David Meger, Doina Precup, Ofir Nachum, and Shixiang Shane Gu. Why should i trust you, bellman? the bellman error is a poor replacement for value error. In International Conference on Machine Learning (ICML), 2022.
  17. 17.Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon. Extreme q-learning: Maxent rl without entropy. In International Conference on Learning Representations (ICLR), 2023.
  18. 18.Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu. Emaq: Expected-max q-learning operator for simple yet effective offline and online rl. In International Conference on Machine Learning (ICML), 2021.
  19. 19.Seyed Kamyar Seyed Ghasemipour, Shixiang Shane Gu, and Ofir Nachum. Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters. In Neural Information Processing Systems (NeurIPS), 2022.
  20. 20.Dibya Ghosh. dibyaghosh/jaxrl_m, 2023. URL https://github.com/dibyaghosh/jaxrl_m.
  21. 21.Philippe Hansen-Estruch, Ilya Kostrikov, Michael Janner, Jakub Grudzien Kuba, and Sergey Levine. Idql: Implicit q-learning as an actor-critic method with diffusion policies. ArXiv, abs/2304.10573, 2023.
  22. 22.Leslie Pack Kaelbling. Learning to achieve goals. In International Joint Conference on Artificial Intelligence (IJCAI), 1993.
  23. 23.Bingyi Kang, Xiao Ma, Yi-Ren Wang, Yang Yue, and Shuicheng Yan. Improving and benchmarking offline reinforcement learning algorithms. ArXiv, abs/2306.00972, 2023.
  24. 24.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  25. 25.Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations (ICLR), 2022.
  26. 26.Aviral Kumar, Aurick Zhou, G. Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2020.
  27. 27.Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, and Sergey Levine. Implicit under-parameterization inhibits data-efficient deep reinforcement learning. In International Conference on Learning Representations (ICLR), 2021.
  28. 28.Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine. Should i run offline reinforcement learning or behavioral cloning? In International Conference on Learning Representations (ICLR), 2021.
  29. 29.Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville, G. Tucker, and Sergey Levine. Dr3: Value-based deep reinforcement learning requires explicit regularization. In International Conference on Learning Representations (ICLR), 2022.
  30. 30.Cassidy Laidlaw, Stuart J. Russell, and Anca D. Dragan. Bridging rl theory and practice with the effective horizon. In Neural Information Processing Systems (NeurIPS), 2023.
  31. 31.Sascha Lange, Thomas Gabel, and Martin Riedmiller. Batch reinforcement learning. In Reinforcement learning: State-of-the-art, pages 45–73. Springer, 2012.
  32. 32.Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joëlle Pineau, and Kee-Eung Kim. Optidice: Offline policy optimization via stationary distribution correction estimation. In International Conference on Machine Learning (ICML), 2021.
  33. 33.Sergey Levine, Aviral Kumar, G. Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. ArXiv, abs/2005.01643, 2020.
  34. 34.Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. In International Conference on Learning Representations (ICLR), 2016.
  35. 35.Cong Lu, Philip J. Ball, Tim G. J. Rudner, Jack Parker-Holder, Michael A. Osborne, and Yee Whye Teh. Challenges and opportunities in offline reinforcement learning from visual observations. Transactions on Machine Learning Research (TMLR), 2023.
  36. 36.Ajay Mandlekar, Danfei Xu, J. Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Mart’in-Mart’in. What matters in learning from offline human demonstrations for robot manipulation. In Conference on Robot Learning (CoRL), 2021.
  37. 37.Bogdan Mazoure, Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson. Improving zero-shot generalization in offline reinforcement learning using generalized similarity functions. In Neural Information Processing Systems (NeurIPS), 2022.
  38. 38.Ishita Mediratta, Qingfei You, Minqi Jiang, and Roberta Raileanu. The generalization gap in offline reinforcement learning. In International Conference on Learning Representations (ICLR), 2024.
  39. 39.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. Playing atari with deep reinforcement learning. ArXiv, abs/1312.5602, 2013.
  40. 40.Rémi Munos. Error bounds for approximate policy iteration. In International Conference on Machine Learning (ICML), 2003.
  41. 41.Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans. Algaedice: Policy gradient from arbitrary experience. ArXiv, abs/1912.02074, 2019.
  42. 42.Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine. Accelerating online reinforcement learning with offline datasets. ArXiv, abs/2006.09359, 2020.
  43. 43.Whitney Newey and James L. Powell. Asymmetric least squares estimation and testing. Econometrica, 55: 819–847, 1987.
  44. 44.Seohong Park, Dibya Ghosh, Benjamin Eysenbach, and Sergey Levine. Hiql: Offline goal-conditioned rl with latent states as actions. In Neural Information Processing Systems (NeurIPS), 2023.
  45. 45.Seohong Park, Tobias Kreiman, and Sergey Levine. Foundation policies with hilbert representations. In International Conference on Machine Learning (ICML), 2024.
  46. 46.Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. ArXiv, abs/1910.00177, 2019.
  47. 47.Jan Peters and Stefan Schaal. Reinforcement learning by reward-weighted regression for operational space control. In International Conference on Machine Learning (ICML), 2007.
  48. 48.Rafael Rafailov, Kyle Beltran Hatch, Anikait Singh, Aviral Kumar, Laura Smith, Ilya Kostrikov, Philippe Hansen-Estruch, Victor Kolev, Philip J Ball, Jiajun Wu, et al. D5rl: Diverse datasets for data-driven deep reinforcement learning. In Reinforcement Learning Conference (RLC), 2024.
  49. 49.Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-maron, Mai Giménez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent. Transactions on Machine Learning Research (TMLR), 2022.
  50. 50.Harshit S. Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum. Dual rl: Unification and new methods for reinforcement and imitation learning. In International Conference on Learning Representations (ICLR), 2024.
  51. 51.Jost Tobias Springenberg, Abbas Abdolmaleki, Jingwei Zhang, Oliver Groth, Michael Bloesch, Thomas Lampe, Philemon Brakel, Sarah Bechtle, Steven Kapturowski, Roland Hafner, Nicolas Manfred Otto Heess, and Martin A. Riedmiller. Offline actor-critic reinforcement learning scales to large models. In International Conference on Machine Learning (ICML), 2024.
  52. 52.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research (JMLR), 15(1):1929–1958, 2014.
  53. 53.Denis Tarasov, Vladislav Kurenkov, Alexander Nikulin, and Sergey Kolesnikov. Revisiting the minimalist approach to offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2023.
  54. 54.Denis Tarasov, Alexander Nikulin, Dmitry Akimov, Vladislav Kurenkov, and Sergey Kolesnikov. Corl: Research-oriented deep offline reinforcement learning library. In Neural Information Processing Systems (NeurIPS), 2023.
  55. 55.Ruosong Wang, Dean Phillips Foster, and Sham M. Kakade. What are the statistical limits of offline rl with linear function approximation? In International Conference on Learning Representations (ICLR), 2021.
  56. 56.Ruosong Wang, Yifan Wu, Ruslan Salakhutdinov, and Sham M. Kakade. Instabilities of offline rl with pre-trained neural representation. In International Conference on Machine Learning (ICML), 2021.
  57. 57.Tongzhou Wang, Antonio Torralba, Phillip Isola, and Amy Zhang. Optimal goal-reaching reinforcement learning via quasimetric learning. In International Conference on Machine Learning (ICML), 2023.
  58. 58.Ziyun Wang, Alexander Novikov, Konrad Zolna, Jost Tobias Springenberg, Scott E. Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Manfred Otto Heess, and Nando de Freitas. Critic regularized regression. In Neural Information Processing Systems (NeurIPS), 2020.
  59. 59.Yifan Wu, G. Tucker, and Ofir Nachum. Behavior regularized offline reinforcement learning. ArXiv, abs/1911.11361, 2019.
  60. 60.Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh. Uncertainty weighted actor-critic for offline reinforcement learning. In International Conference on Machine Learning (ICML), 2021.
  61. 61.Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Victor Chan, and Xianyuan Zhan. Offline rl with no ood actions: In-sample learning via implicit value regularization. In International Conference on Learning Representations (ICLR), 2023.
  62. 62.Mengjiao Yang and Ofir Nachum. Representation matters: Offline pretraining for sequential decision making. In International Conference on Machine Learning (ICML), 2021.
  63. 63.Rui Yang, Yong Lin, Xiaoteng Ma, Haotian Hu, Chongjie Zhang, and T. Zhang. What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning (ICML), 2023.
  64. 64.Denis Yarats, David Brandfonbrener, Hao Liu, Michael Laskin, P. Abbeel, Alessandro Lazaric, and Lerrel Pinto. Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning. ArXiv, abs/2201.13425, 2022.
  65. 65.Chongyi Zheng, Benjamin Eysenbach, Homer Walke, Patrick Yin, Kuan Fang, Ruslan Salakhutdinov, and Sergey Levine. Stabilizing contrastive rl: Techniques for offline goal reaching. ArXiv, abs/2306.03346, 2023.

Citation

MLA
Park, S., et al. “Is Value Learning Really the Main Bottleneck in Offline RL?”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 79029–56, https://proceedings.neurips.cc/paper_files/paper/2024/file/8ffb4e3118280a66b192b6f06e0e2596-Paper-Conference.pdf.
APA
Park, S., Frans, K., Levine, S., & Kumar, A. (2024). Is Value Learning Really the Main Bottleneck in Offline RL?. Advances in Neural Information Processing Systems, 37, 79029–79056. https://proceedings.neurips.cc/paper_files/paper/2024/file/8ffb4e3118280a66b192b6f06e0e2596-Paper-Conference.pdf
Chicago
Park, S., K. Frans, S. Levine, and A. Kumar. 2024. “Is Value Learning Really the Main Bottleneck in Offline RL?”. Advances in Neural Information Processing Systems 37: 79029–56. https://proceedings.neurips.cc/paper_files/paper/2024/file/8ffb4e3118280a66b192b6f06e0e2596-Paper-Conference.pdf.
Harvard
Park, S. et al. (2024) “Is Value Learning Really the Main Bottleneck in Offline RL?”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 79029–79056. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/8ffb4e3118280a66b192b6f06e0e2596-Paper-Conference.pdf.
Vancouver
1. Park S, Frans K, Levine S, Kumar A (2024) Is Value Learning Really the Main Bottleneck in Offline RL?. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 79029–79056

BibTeX

@inproceedings{park2024value,
  title = {Is Value Learning Really the Main Bottleneck in Offline RL?},
  author = {Park, Seohong and Frans, Kevin and Levine, Sergey and Kumar, Aviral},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {79029-79056},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/8ffb4e3118280a66b192b6f06e0e2596-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors