Is Value Learning Really the Main Bottleneck in Offline RL?
Seohong ParkKevin FransSergey LevineAviral Kumar
Reveals that offline reinforcement learning performance is primarily limited by policy extraction objectives and test-time policy generalization rather than imperfect value learning, providing practical test-time improvement techniques to bridge this gap.
Data-driven control algorithms often struggle to perform reliably when trained entirely on historical datasets. While offline reinforcement learning theoretically holds an advantage over standard imitation learning by learning from suboptimal data through value functions, it frequently underperforms in practical settings. Historically, researchers have attributed this shortcoming to inaccurate value estimation. The article addresses the root causes of this underperformance to establish whether learning the value function is truly the primary bottleneck holding back offline reinforcement learning.
The main objective of the article is to systematically evaluate how value estimation, policy extraction, and test-time generalization limit algorithm performance and scalability. To do this, the authors conducted an extensive empirical study spanning over 15,000 experimental runs across eight diverse robotics, manipulation, and continuous-control environments. By decoupling value estimation from policy training, they analyzed how varying the quality, quantity, and coverage of training data affected different algorithmic components.
The findings reveal that policy extraction and test-time generalization—rather than value function estimation—are the dominant bottlenecks in offline reinforcement learning. First, the method chosen to extract a policy from a learned value function significantly dictates success; behavior-constrained policy gradient methods outperformed or matched popular value-weighted behavioral cloning methods in 15 out of 16 tested settings. In complex environments, gradient-based extraction achieved performance scores near 191 to 193 compared to under 100 for weighted cloning. Second, existing algorithms already optimize well on in-distribution training data, but they fail during deployment when encountering novel, out-of-distribution states. Third, increasing dataset state coverage with exploratory, noisy actions produced substantially better deployment policies than using smaller, highly optimal datasets.
These insights challenge the prevailing belief that improving value function estimation is the most important path forward. Instead, conventional policy extraction methods often fail to fully exploit learned value functions due to severe sample inefficiency and overfitting. Furthermore, because failures occur primarily when agents encounter unseen states at deployment time, relying solely on conservative training objectives is insufficient to ensure dependable real-world performance.
To improve offline reinforcement learning in practice, practitioners should prioritize collecting datasets with high state coverage rather than focusing solely on clean, expert demonstrations. Decision-makers should also deploy behavior-constrained gradient methods instead of weighted imitation objectives and integrate test-time policy improvement techniques—such as on-the-fly action adjustment or test-time updates—to steer actions effectively during execution. The conclusions are well-supported across continuous-control environments, though caution is warranted for discrete-action tasks or settings with multiple equally optimal actions, where the study's proxy metrics may have limited applicability.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). This comprehensive tutorial establishes the core formulation of offline reinforcement learning, foundational value-learning challenges, and distributional shift concepts that the source paper dissects.
- Paper: Offline Reinforcement Learning with Implicit Q-Learning, Ilya Kostrikov et al. (2021). It introduces Implicit Q-Learning and advantage-weighted policy extraction, providing the primary baseline paradigm whose policy extraction and value learning bottlenecks are directly evaluated in the source paper.
- Paper: Conservative Q-Learning for Offline Reinforcement Learning, Aviral Kumar et al. (2020). It details Conservative Q-Learning, establishing the canonical value-regularization framework that represents the imperfect value learning hypothesis re-examined by the source paper.
- Paper: Off-Policy Deep Reinforcement Learning without Exploration, Scott Fujimoto et al. (2018). This seminal paper introduces batch-constrained policy learning to mitigate extrapolation error, forming the basis of behavior-constrained policy extraction analyzed in the source.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). It establishes the standardized D4RL offline benchmarks and data distribution regimes upon which the empirical analyses of value learning and policy extraction in the source rely.
- Paper: Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction, Aviral Kumar et al. (2019). It defines support-constrained offline policy optimization mechanisms that the source investigates when evaluating behavior-constrained policy gradients.
- Paper: Supported Policy Optimization for Offline Reinforcement Learning, Jialong Wu et al. (2022). It develops explicit density-supported policy extraction on top of TD3, exemplifying the behavior-constrained actor-critic frameworks compared in the source study.
- Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). It derives the deterministic policy gradient framework (DDPG) whose behavior-constrained variants (e.g., DDPG+BC) are highlighted by the source as superior policy extraction objectives.
No sufficiently relevant recommendations were found.
