Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning

Yihao SunJiaji ZhangChengxing JiaHaoxin LinJunyin YeYang Yu

article2023ICML59 citations

Proposes MOBILE, a model-based offline reinforcement learning algorithm that measures uncertainty using Bellman estimation inconsistency across a dynamics ensemble to approximate true Bellman errors and achieve state-of-the-art performance on D4RL and NeoRL benchmarks.

Listen

Offline reinforcement learning enables decision-making systems to learn purely from previously collected static data, avoiding the high costs and physical risks associated with live trial-and-error exploration. To expand limited datasets, model-based methods build predictive dynamics models to simulate additional scenarios. However, because predictive models inevitably contain errors in regions with sparse data, decision policies often exploit these inaccuracies by overestimating potential rewards, ultimately leading to severe operational failures.

The article develops and evaluates a new offline reinforcement learning framework called MOBILE (Model-Bellman Inconsistency Penalized Offline Policy Optimization). The primary objective is to demonstrate that directly estimating decision uncertainty across an ensemble of models—rather than relying solely on transition prediction errors—provides a tighter, more reliable penalty against risky, out-of-distribution actions.

The researchers designed an uncertainty metric, termed Model-Bellman Inconsistency, which measures the variance in expected future values across an ensemble of learned dynamics models. This penalty directly discounts value estimates on synthetic data generated in uncertain territory. The authors evaluated MOBILE across standard benchmarks, including the D4RL suite (Gym and robotic Adroit domains) and the near-real-world NeoRL benchmark, comparing performance, runtime efficiency, and memory footprint against leading model-free and model-based baselines.

Empirical findings demonstrate that the proposed uncertainty metric correlates substantially higher with the true Bellman estimation error than existing transition-focused quantifiers. Across 27 benchmark datasets, MOBILE achieved state-of-the-art results on 20 tasks. On the standard D4RL Gym suite, MOBILE achieved an average score of 80.0, outperforming model-free methods like EDAC (76.0) and model-based baselines such as MOPO (70.3). Furthermore, on the conservative NeoRL dataset designed to mirror real-world logging conditions, MOBILE achieved an average score of 60.7 compared to CQL (56.1) and MOPO (28.5). Computational analysis showed that MOBILE matches the training runtime of standard model-based methods while requiring only 2.2 million parameters, compared to 13.7 million for high-performing ensemble critics.

These results indicate that combining environment dynamics with policy value functions enables agents to explore productive simulated pathways without straying into dangerous, poorly modeled scenarios. For practitioners, this translates to improved policy reliability and reduced operational risk when deploying autonomous systems trained exclusively on legacy data, all without imposing additional computational overhead.

Organizations aiming to implement offline reinforcement learning in data-constrained domains should adopt Model-Bellman Inconsistency penalization in place of traditional dynamics-only reward penalties. When resources permit, expanding the model ensemble size can offer further performance gains. However, caution is advised in extreme low-data regimes with complex hand manipulation, where purely model-free methods may still prove complementary. Future efforts should focus on validating the framework in physical industrial pilots and refining automatic hyperparameter tuning across varying data qualities.

No sufficiently relevant recommendations were found.

Cover for Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning

Abstract

For offline reinforcement learning (RL), model-based methods are expected to be data-efficient as they incorporate dynamics models to generate more data. However, due to inevitable model errors, straightforwardly learning a policy in the model typically fails in the offline setting. Previous studies have incorporated conservatism to prevent out-of-distribution exploration. For example, MOPO penalizes rewards through uncertainty measures from predicting the next states, which we have discovered are loose bounds of the ideal uncertainty, i.e., the Bellman error. In this work, we propose MOdel-Bellman Inconsistency penalized OffLinE Policy Optimization (MOBILE), a novel uncertainty-driven offline RL algorithm. MOBILE conducts uncertainty quantification through the inconsistency of Bellman estimations under an ensemble of learned dynamics models, which can be a better approximator to the true Bellman error, and penalizes the Bellman estimation based on this uncertainty. Empirically we have verified that our proposed uncertainty quantification can be significantly closer to the true Bellman error than the compared methods. Consequently, MOBILE outperforms prior offline RL approaches on most tasks of D4RL and NeoRL benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. MDPs and Offline RL
  • 2.2. Model-based Offline RL Algorithms
  • 3. Model-Bellman Inconsistency Penalized Offline Policy Optimization
  • 3.1. Pessimistic Value Estimation via Model-Bellman Inconsistency
  • 3.2. Theoretical Connections to PEVI
  • 3.3. Algorithm
  • 4. Experiments
  • 4.1. Benchmark Results
  • 4.1.1. D4RL
  • 4.1.2. NEORL
  • 4.2. Uncertainty Quantification
  • 4.3. Bellman Error Estimation
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Introduction to Pessimistic Value Iteration
  • A.1. Background of Pessimistic Value Iteration
  • A.2. Practical Implementation of PEVI
  • B. Implementation Details
  • B.1. Dynamics Model Training
  • B.2. Policy Optimization
  • C. Experimental Details
  • C.1. Benchmarks
  • C.2. Hyperparameters
  • C.3. Tuning for MOPO
  • C.4. Computational Cost Comparison
  • D. Omitted Experiments
  • D.1. Experiments in Adroit Domain
  • D.2. Hyperparameters for TD3+BC and EDAC in NeoRL
  • D.3. MOBILE with Larger Ensemble
  • D.4. Omitted Uncertainty Quantification in Section 4.2
  • D.5. Omitted Bellman Error Estimation in Section 4.3

Knowls

  1. Knowl 1 — Model-Bellman Inconsistency as a Bellman-error penalty

    equation

    Let ss be a state, aa an action, r(s,a)r(s,a) the known reward, and gamma\in(0,1) the discount factor. Let T∗T^* be the unknown true transition law and let {T^i}i=1N\{\widehat T_i\}_{i=1}^N be an ensemble of learned transition models, with mean model T^=1N∑i=1NT^i\widehat T=\frac{1}{N}\sum_{i=1}^N\widehat T_i. For a policy π\pi and a target action-value function Q−Q^-, define the Bellman estimate under model ii by

    T^iπQ−(s,a)=r(s,a)+γ Es′∼T^i(⋅∣s,a), a′∼π(⋅∣s′)[Q−(s′,a′)].\widehat T_i^\pi Q^-(s,a)=r(s,a)+\gamma\,\mathbb{E}_{s'\sim\widehat T_i(\cdot\mid s,a),\,a'\sim\pi(\cdot\mid s')}[Q^-(s',a')].

    The mean ensemble estimate is T^πQ−=1N∑iT^iπQ−\widehat T^\pi Q^- = \frac{1}{N}\sum_i\widehat T_i^\pi Q^-. MOBILE defines Model-Bellman Inconsistency as the population standard deviation of the ensemble’s Bellman estimates:

    U(s,a)=1N∑i=1N(T^iπQ−(s,a)−T^πQ−(s,a))2.U(s,a)=\sqrt{\frac{1}{N}\sum_{i=1}^N\left(\widehat T_i^\pi Q^-(s,a)-\widehat T^\pi Q^-(s,a)\right)^2}.

    The corresponding pessimistic target, called the Model-Bellman Inconsistency Penalized operator, is T^MOBIPQ−(s,a)=T^πQ−(s,a)−βU(s,a)\widehat T^{\mathrm{MOBIP}}Q^-(s,a)=\widehat T^\pi Q^-(s,a)-\beta U(s,a), where β≥0\beta\geq 0 is a penalty coefficient. Unlike a penalty based only on predicted transition uncertainty, UU measures disagreement in value-relevant Bellman estimates, incorporating both model predictions and the value function.

  2. Knowl 2 — MOBILE training procedure and practical configuration

    model/method

    MOBILE combines synthetic-data generation with a soft actor-critic (SAC) learner and applies the Model-Bellman Inconsistency penalty only to synthetic transitions. It assumes an offline dataset DD and a known reward function. A probabilistic dynamics ensemble is trained on DD; each network outputs a Gaussian distribution over the next state and reward. In the reported experiments, seven models were trained independently by maximum likelihood, each as a four-layer feedforward network with 200 hidden units; five models with the lowest validation prediction error on a held-out set of 1,000 transitions were retained. A model was randomly selected from those five for simulated rollouts.

    Before agent updates, MOBILE generates short rollouts of length hh from states in DD and adds the resulting transitions to a synthetic replay set DmodelD_{\mathrm{model}}. For each critic QψkQ_{\psi_k}, k∈{1,2}k\in\{1,2\}, a real transition (s,a,r,s′)∈D(s,a,r,s')\in D uses target

    yreal=r+γ[min⁡k∈{1,2}Qψk−(s′,a′)−αlog⁡πϕ(a′∣s′)],y_{\mathrm{real}}=r+\gamma\left[\min_{k\in\{1,2\}}Q^-_{\psi_k}(s',a')-\alpha\log\pi_\phi(a'\mid s')\right],

    where a′∼πϕ(⋅∣s′)a'\sim\pi_\phi(\cdot\mid s'), Qψk−Q^-_{\psi_k} is a target critic, and α\alpha is the SAC entropy coefficient. A synthetic transition uses the same target minus βU(s,a)\beta U(s,a). The inconsistency estimate is the ensemble standard deviation of the model-specific discounted expectations of min⁡kQψk−(s′,a′)\min_k Q^-_{\psi_k}(s',a'); the implementation omits reward disagreement from this calculation because the authors report it is relatively small. The two critics are fitted by squared-error regression to their targets. The actor maximizes the expected minimum critic value minus the entropy term, using states from both replay sets.

    For Gym and NeoRL, the experiments used two critics, actor and critic networks with two 256-unit fully connected layers and ReLU activations, discount 0.990.99, Adam, actor learning rate 10−410^{-4}, critic learning rate 3×10−43\times10^{-4}, target-network smoothing coefficient 5×10−35\times10^{-3}, batch size 256, and three million gradient steps. Each batch contained 5% real and 95% synthetic transitions. Rollout lengths were selected per task from h∈{1,5}h\in\{1,5\} for Gym and NeoRL; Adroit used h∈{1,3}h\in\{1,3\}. The penalty coefficient β\beta was tuned per task.

  3. Knowl 3 — Conditional validity of the ensemble inconsistency estimate

    theoretical result

    For a finite-horizon problem, let T^Vh+1(s,a)\widehat T V_{h+1}(s,a) be the Bellman estimate produced by an ensemble-mean learned dynamics model for a value function Vh+1V_{h+1}, and let TVh+1(s,a)T V_{h+1}(s,a) be the estimate under the true dynamics. Suppose the ensemble standard deviation of the model-specific Bellman estimates is an admissible error estimator: for a selected coefficient βh\beta_h, it satisfies, simultaneously for all states ss and actions aa,

    ∣T^Vh+1(s,a)−TVh+1(s,a)∣≤βh Stdi=1,…,N ⁣(T^iVh+1(s,a)).\left|\widehat T V_{h+1}(s,a)-T V_{h+1}(s,a)\right| \leq \beta_h\,\mathrm{Std}_{i=1,\ldots,N}\!\left(\widehat T_i V_{h+1}(s,a)\right).

    Under this assumption, the penalty Γh(s,a)=βh Stdi(T^iVh+1(s,a))\Gamma_h(s,a)=\beta_h\,\mathrm{Std}_{i}(\widehat T_i V_{h+1}(s,a)) is a valid uncertainty quantifier: with the assumed coverage probability, it upper-bounds the Bellman estimation error uniformly over state-action pairs. The paper’s result is conditional on the admissibility assumption; it does not prove that an ensemble standard deviation always bounds the true error.

  4. Knowl 4 — PEVI relates small uncertainty penalties to policy quality

    theoretical result

    In pessimistic value iteration (PEVI), suppose the penalty functions {Γh}h=1H\{\Gamma_h\}_{h=1}^H satisfy ∣T^Vh+1(s,a)−TVh+1(s,a)∣≤Γh(s,a)|\widehat T V_{h+1}(s,a)-T V_{h+1}(s,a)|\leq\Gamma_h(s,a) for every state-action pair at each step, with probability at least 1−ξ1-\xi. Then the policy π^\widehat\pi returned by PEVI has, simultaneously for every initial state ss, the following suboptimality bound with probability at least 1−ξ1-\xi:

    ∣Vπ∗(s)−Vπ^(s)∣≤2∑h=1HEπ∗ ⁣[Γh(sh,ah)∣s1=s].\left|V^{\pi^*}(s)-V^{\widehat\pi}(s)\right| \leq 2\sum_{h=1}^H\mathbb{E}_{\pi^*}\!\left[\Gamma_h(s_h,a_h)\mid s_1=s\right].

    Here HH is the episode horizon, π∗\pi^* is an optimal policy in the true MDP, and the expectation is over its induced trajectory. This result motivates MOBILE’s design: a useful pessimism penalty must cover Bellman error, while a tighter penalty can yield a tighter policy-performance bound.

  5. Knowl 5 — Model-Bellman Inconsistency targets a tighter error than MOPO’s model penalty

    theoretical result

    The paper contrasts direct estimation of Bellman error with MOPO’s transition-model uncertainty penalty. Let T^θ(⋅∣s,a)\widehat T_\theta(\cdot\mid s,a) be a learned transition distribution, T∗(⋅∣s,a)T^*(\cdot\mid s,a) the true distribution, and UMOPO(s,a)U^{\mathrm{MOPO}}(s,a) an uncertainty estimate satisfying the assumed bound DTV(T^θ,T∗)≤UMOPO(s,a)D_{\mathrm{TV}}(\widehat T_\theta,T^*)\leq U^{\mathrm{MOPO}}(s,a). If rewards are bounded in magnitude by rmax⁡r_{\max} and the discount is γ∈(0,1)\gamma\in(0,1), the paper’s argument yields a Bellman-error upper bound proportional to

    γrmax⁡1−γ UMOPO(s,a).\frac{\gamma r_{\max}}{1-\gamma}\,U^{\mathrm{MOPO}}(s,a).

    Thus, subject to the admissibility assumption, a suitably scaled MOPO uncertainty is a valid uncertainty quantifier, but its bound carries a factor of order 1/(1−γ)1/(1-\gamma). MOBILE instead estimates disagreement in Bellman values directly; under its own admissibility assumption, the paper characterizes its uncertainty scaling as O(1)O(1) with respect to the horizon, rather than first bounding transition error and converting it into Bellman error. The tighter scaling is a theoretical comparison, not an unconditional guarantee that the ensemble estimate is accurate.

  6. Knowl 6 — Bellman-error estimation is more correlated with MOBILE’s inconsistency measure

    empirical result

    The authors tested whether Model-Bellman Inconsistency tracks true Bellman error better than uncertainty measures based only on dynamics predictions. They evaluated state-action pairs generated by learned models, used the real environment to compute the corresponding Bellman error, and measured the correlation between that error and each uncertainty estimate. The comparisons were max-aleatoric uncertainty, max-pairwise difference between ensemble predictions, ensemble standard deviation combining epistemic and aleatoric uncertainty, and Model-Bellman Inconsistency. Across the reported Gym-task curves, Model-Bellman Inconsistency showed higher correlation with the true Bellman error than the other three measures. In an additional comparison, the alternative penalties were rescaled to have comparable uncertainty magnitudes; MOBILE still achieved better policy performance than the three MOPO-style variants. The paper presents these findings graphically and does not report exact correlation coefficients in the text.

  7. Knowl 7 — D4RL Gym benchmark performance

    data/table

    The table reports normalized returns on 12 D4RL Gym v2 datasets, averaged over four random seeds. MOBILE has the highest reported score on eight individual datasets and the highest overall average, 80.0. MOPO and MOPO* are distinct: MOPO scores are from the original paper’s v0 data, whereas MOPO* scores use the authors’ implementation on v2 data with retuned hyperparameters. Scores are reproduced as reported; MOBILE entries include the reported ±\pm values.

    Task BC CQL TD3+BC EDAC MOPO MOPO* COMBO TT RAMBO MOBILE
    halfcheetah-random 2.2 31.3 11.0 28.4 35.4 38.5 38.8 6.1 39.5 39.3±\pm3.0
    hopper-random 3.7 5.3 8.5 25.3 11.7 31.7 17.9 6.9 25.4 31.9±\pm0.6
    walker2d-random 1.3 5.4 1.6 16.6 13.6 7.4 7.0 5.9 0.0 17.9±\pm6.6
    halfcheetah-medium 43.2 46.9 48.3 65.9 42.3 73.0 54.2 46.9 77.9 74.6±\pm1.2
    hopper-medium 54.1 61.9 59.3 101.6 28.0 62.8 97.2 67.4 87.0 106.6±\pm0.6
    walker2d-medium 70.9 79.5 83.7 92.5 17.8 84.1 81.9 81.3 84.9 87.7±\pm1.1
    halfcheetah-medium-replay 37.6 45.3 44.6 61.3 53.1 72.1 55.1 44.1 68.7 71.7±\pm1.2
    hopper-medium-replay 16.6 86.3 60.9 101.0 67.5 103.5 89.5 99.4 99.5 103.9±\pm1.0
    walker2d-medium-replay 20.3 76.8 81.8 87.1 39.0 85.6 56.0 82.6 89.2 89.9±\pm1.5
    halfcheetah-medium-expert 44.0 95.0 90.7 106.3 63.3 90.8 90.0 95.0 95.4 108.2±\pm2.5
    hopper-medium-expert 53.9 96.9 98.0 110.7 23.7 81.6 111.1 110.0 88.2 112.6±\pm0.2
    walker2d-medium-expert 90.1 109.1 110.1 114.7 44.6 112.9 103.3 101.9 56.7 115.2±\pm0.7
    Average 36.5 61.6 58.2 76.0 36.7 70.3 66.8 62.3 67.7 80.0
  8. Knowl 8 — NeoRL benchmark performance

    data/table

    The table gives normalized returns on nine NeoRL tasks using 1,000 training trajectories per task; each value is averaged over four random seeds. The tasks combine HalfCheetah, Hopper, and Walker2d with low (L), medium (M), and high (H) dataset quality. MOBILE reaches the highest overall average, 60.7, compared with 56.1 for CQL, and leads on five of the nine tasks. Its gains are especially large on HalfCheetah-L, HalfCheetah-M, and Hopper-H, while it is not the best method on every task.

    Task BC CQL TD3+BC EDAC MOPO MOBILE
    HalfCheetah-L 29.1 38.2 30.0 31.3 40.1 54.7±\pm3.0
    Hopper-L 15.1 16.0 15.8 18.3 6.2 17.4±\pm3.9
    Walker2d-L 28.5 44.7 43.0 40.2 11.6 37.6±\pm2.0
    HalfCheetah-M 49.0 54.6 52.3 54.9 62.3 77.8±\pm1.4
    Hopper-M 51.3 64.5 70.3 44.9 1.0 51.1±\pm13.3
    Walker2d-M 48.7 57.3 58.5 57.6 39.9 62.2±\pm1.6
    HalfCheetah-H 71.3 77.4 75.3 81.4 65.9 83.0±\pm4.6
    Hopper-H 43.1 76.6 75.3 52.5 11.5 87.8±\pm26.0
    Walker2d-H 72.6 75.3 69.6 75.5 18.0 74.9±\pm3.4
    Average 45.4 56.1 54.5 50.7 28.5 60.7
  9. Knowl 9 — Adroit results depend on the amount and type of offline data

    empirical result

    On D4RL Adroit, the paper evaluated human-demonstration datasets containing 25 trajectories and cloned datasets mixing demonstrations with behavior-cloned data. MOBILE performed best among the listed methods on all three cloned tasks, but did not match EDAC or CQL on the human datasets. The authors hypothesize that the human datasets’ small size—about 5,000 transitions—makes it difficult to learn a dynamics model that generalizes well. This is a stated explanation, not a separately tested causal result. Scores are normalized returns averaged over four seeds; MOBILE’s reported values include the accompanying ±\pm terms.

    Task BC CQL TD3+BC EDAC MOPO MOBILE
    pen-human 25.8 35.2 -1.0 52.1 10.7 30.1±\pm14.6
    door-human 2.8 9.1 -0.2 10.7 -0.2 -0.2±\pm0.1
    hammer-human 3.1 0.6 0.2 0.8 0.3 0.4±\pm0.2
    pen-cloned 38.3 27.2 -2.1 68.2 54.6 69.0±\pm9.3
    door-cloned 0.0 3.5 0.0 9.6 15.3 24.0±\pm22.8
    hammer-cloned 0.7 1.4 -0.1 0.3 0.5 1.5±\pm0.4
  10. Knowl 10 — MOBILE has near-MOPO computational cost in the reported comparison

    data/table

    The authors measured runtime per epoch of 1,000 gradient steps and parameter count on the hopper-medium-v2 task, using one GeForce GTX 3070 GPU and an AMD Ryzen 5900X CPU at 4.8 GHz. MOBILE used 8 seconds per epoch and 2.2 million parameters, close to MOPO at 7 seconds and 2.2 million parameters. It was faster than the reported CQL and EDAC runs; EDAC had the largest parameter count because it uses an ensemble of Q-networks.

    Algorithm Runtime (s/epoch) Parameters
    CQL 12 0.7M
    EDAC 10 13.7M
    MOPO 7 2.2M
    MOBILE 8 2.2M

Coverage note — The secondary OOD-action uncertainty plots, per-task hyperparameter listings, MOPO-variant tuning results, and larger-ensemble ablation are omitted because they provide supporting diagnostics or implementation detail rather than additional central findings.

References

  1. 1.An, G., Moon, S., Kim, J., and Song, H. O. Uncertainty-based offline reinforcement learning with diversified q-ensemble. In Advances in Neural Information Processing Systems 34 (NeurIPS’21), virtual event, 2021.
  2. 2.Bai, C., Wang, L., Yang, Z., Deng, Z., Garg, A., Liu, P., and Wang, Z. Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning. In The Tenth International Conference on Learning Representations (ICLR’22), virtual event, 2022.
  3. 3.Chen, X., Yu, Y., Li, Q., Luo, F., Qin, Z. T., Shang, W., and Ye, J. Offline model-based adaptable policy learning. In Advances in Neural Information Processing Systems 34 (NeurIPS’21), virtual event, 2021.
  4. 4.Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., González, J. E., and Levine, S. Model-based value estimation for efficient model-free reinforcement learning. CoRR, abs/1803.00101, 2018. URL http://arxiv.org/abs/1803.00101.
  5. 5.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4RL: datasets for deep data-driven reinforcement learning. CoRR, abs/2004.07219, 2020. URL https://arxiv.org/abs/2004.07219.
  6. 6.Fujimoto, S. and Gu, S. S. A minimalist approach to offline reinforcement learning. In Advances in Neural Information Processing Systems 34 (NeurIPS’21), virtual event, 2021.
  7. 7.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. In Proceedings of the 36th International Conference on Machine Learning (ICML’19), Long Beach, USA, 2019.
  8. 8.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning (ICML’18), Stockholmsmässan, Sweden, 2018.
  9. 9.Janner, M., Fu, J., Zhang, M., and Levine, S. When to trust your model: Model-based policy optimization. In Advances in Neural Information Processing Systems 32 (NeurIPS’19), Vancouver, Canada, 2019.
  10. 10.Janner, M., Li, Q., and Levine, S. Offline reinforcement learning as one big sequence modeling problem. In Advances in Neural Information Processing Systems 34 (NeurIPS’21), virtual event, 2021.
  11. 11.Jeong, J., Wang, X., Gimelfarb, M., Kim, H., Abdulhai, B., and Sanner, S. Conservative bayesian model-based value expansion for offline policy optimization. CoRR, abs/2210.03802, 2022. doi: 10.48550/arXiv.2210.03802. URL https://doi.org/10.48550/arXiv.2210.03802.
  12. 12.Jin, Y., Yang, Z., and Wang, Z. Is pessimism provably efficient for offline rl? In Proceedings of the 38th International Conference on Machine Learning (ICML’21), virtual event, 2021.
  13. 13.Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. Morel: Model-based offline reinforcement learning. In Advances in Neural Information Processing Systems 33 (NeurIPS’20), virtual event, 2020.
  14. 14.Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. Stabilizing off-policy q-learning via bootstrapping error reduction. In Advances in Neural Information Processing Systems 32 (NeurIPS’19), Vancouver, Canada, 2019.
  15. 15.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems 33 (NeurIPS’20), virtual event, 2020.
  16. 16.Lange, S., Gabel, T., and Riedmiller, M. A. Batch reinforcement learning. In Reinforcement Learning, volume 12, pp. 45–73. 2012.
  17. 17.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. CoRR, abs/2005.01643, 2020. URL https://arxiv.org/abs/2005.01643.
  18. 18.Lin, H., Sun, Y., Zhang, J., and Yu, Y. Model-based reinforcement learning with multi-step plan value estimation. CoRR, abs/2209.05530, 2022. URL https://doi.org/10.48550/arXiv.2209.05530.
  19. 19.Lu, C., Ball, P. J., Parker-Holder, J., Osborne, M. A., and Roberts, S. J. Revisiting design choices in offline model based reinforcement learning. In The Tenth International Conference on Learning Representations (ICLR’22), virtual event, 2022.
  20. 20.Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T. Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees. In The Seventh International Conference on Learning Representations (ICLR’19), New Orleans, USA, 2019.
  21. 21.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. Human-level control through deep reinforcement learning. Nat., 518(7540):529–533, 2015.
  22. 22.Müller, A. Integral probability metrics and their generating classes of functions. Advances in Applied Probability, 29 (2):429–443, 1997.
  23. 23.Qin, R.-J., Zhang, X., Gao, S., Chen, X.-H., Li, Z., Zhang, W., and Yu, Y. NeoRL: A near real-world benchmark for offline reinforcement learning. In Advances in Neural Information Processing Systems 35 (NeurIPS’22, Datasets and Benchmarks), New Orleans, USA, 2022.
  24. 24.Rigter, M., Lacerda, B., and Hawes, N. RAMBO-RL: robust adversarial model-based offline reinforcement learning. CoRR, abs/2204.12581, 2022. doi: 10.48550/arXiv.2204.12581. URL https://doi.org/10.48550/arXiv.2204.12581.
  25. 25.Xu, T., Li, Z., and Yu, Y. Error bounds of imitating policies and environments. In Advances in Neural Information Processing Systems 33 (NeurIPS’20), virtual event, 2020.
  26. 26.Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. MOPO: model-based offline policy optimization. In Advances in Neural Information Processing Systems 33 (NeurIPS’20), virtual event, 2020.
  27. 27.Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. COMBO: conservative offline model-based policy optimization. In Advances in Neural Information Processing Systems 34 (NeurIPS’21), virtual event, 2021.

Citation

MLA
Sun, Y., et al. “Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 33177–94, https://proceedings.mlr.press/v202/sun23q.html.
APA
Sun, Y., Zhang, J., Jia, C., Lin, H., Ye, J., & Yu, Y. (2023). Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning. International Conference on Machine Learning, 202, 33177–33194. https://proceedings.mlr.press/v202/sun23q.html
Chicago
Sun, Y., J. Zhang, C. Jia, H. Lin, J. Ye, and Y. Yu. 2023. “Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning”. International Conference on Machine Learning 202: 33177–94. https://proceedings.mlr.press/v202/sun23q.html.
Harvard
Sun, Y. et al. (2023) “Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning”, International Conference on Machine Learning. PMLR, pp. 33177–33194. Available at: https://proceedings.mlr.press/v202/sun23q.html.
Vancouver
1. Sun Y, Zhang J, Jia C, Lin H, Ye J, Yu Y (2023) Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning. In: International Conference on Machine Learning. PMLR, pp 33177–33194

BibTeX

@InProceedings{pmlr-v202-sun23q,
  title = 	 {Model-{B}ellman Inconsistency for Model-based Offline Reinforcement Learning},
  author =       {Sun, Yihao and Zhang, Jiaji and Jia, Chengxing and Lin, Haoxin and Ye, Junyin and Yu, Yang},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {33177--33194},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/sun23q/sun23q.pdf},
  url = 	 {https://proceedings.mlr.press/v202/sun23q.html},
  abstract = 	 {For offline reinforcement learning (RL), model-based methods are expected to be data-efficient as they incorporate dynamics models to generate more data. However, due to inevitable model errors, straightforwardly learning a policy in the model typically fails in the offline setting. Previous studies have incorporated conservatism to prevent out-of-distribution exploration. For example, MOPO penalizes rewards through uncertainty measures from predicting the next states, which we have discovered are loose bounds of the ideal uncertainty, i.e., the Bellman error. In this work, we propose MOdel-Bellman Inconsistency penalized offLinE Policy Optimization (MOBILE), a novel uncertainty-driven offline RL algorithm. MOBILE conducts uncertainty quantification through the inconsistency of Bellman estimations under an ensemble of learned dynamics models, which can be a better approximator to the true Bellman error, and penalizes the Bellman estimation based on this uncertainty. Empirically we have verified that our proposed uncertainty quantification can be significantly closer to the true Bellman error than the compared methods. Consequently, MOBILE outperforms prior offline RL approaches on most tasks of D4RL and NeoRL benchmarks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/