When to Trust Your Model: Model-Based Policy Optimization

Michael JannerJustin FuMarvin ZhangSergey Levine

article2019NeurIPS1,324 citations

Proposes Model-Based Policy Optimization, a framework that avoids compounding prediction errors by branching short model rollouts from real experience to match the asymptotic performance of model-free reinforcement learning with drastically superior sample efficiency.

Listen

Reinforcement learning offers powerful techniques for automating complex control tasks, yet deploying these methods in real-world physical systems is often hindered by high data collection costs. Model-free approaches achieve high final performance but demand millions of interactions, whereas traditional model-based methods learn rapidly by simulating environments but suffer from compounding predictive errors that degrade asymptotic performance. The article addresses this fundamental efficiency–accuracy trade-off by introducing and evaluating Model-Based Policy Optimization, an algorithmic framework designed to maximize learning speed while preserving optimal final performance.

The approach combines theoretical performance bounds with empirical modeling of generalization error. Rather than generating long, unreliable simulated trajectories from initial states, the method initiates short-horizon rollouts branched directly from real state data stored in a replay buffer. An ensemble of probabilistic neural networks captures both data noise and model uncertainty, while a standard off-policy optimization algorithm trains on the synthetic rollouts. The authors evaluated this framework across standard continuous-control robotic simulation benchmarks without modifying task horizons or granting access to privileged environmental information.

The evaluation yielded several key findings. First, the proposed method achieved learning speeds up to an order of magnitude faster than leading model-free baselines while matching their final peak performance; for instance, on complex tasks, it achieved comparable results using only one-tenth of the environment samples. Second, the method scaled successfully to high-dimensional control tasks where competing model-based algorithms failed entirely. Third, the analysis demonstrated that even single-step simulated rollouts provide substantial training stability and efficiency gains, effectively preventing the policy from exploiting inaccuracies in the predictive model.

These findings indicate that organizations can significantly cut the time, energy, and financial expenses associated with physical hardware experimentation. Restricting synthetic rollouts to short horizons provides a reliable safeguard against model errors, eliminating the historical performance penalties associated with model-based learning. For engineering and data science teams deploying reinforcement learning in physical control or robotics, the recommended next step is to adopt short-horizon branched rollouts rather than pursuing complex, long-horizon dynamics planning. While these results show high reliability across standard benchmarks, practitioners should conduct pilot evaluations in target physical settings to assess model generalization before deploying policies on physical hardware.

  • Paper: Approximately Optimal Approximate Reinforcement Learning, S. Kakade et al. (2002). This paper establishes the foundational conservative policy iteration framework and monotonic improvement bounds that MBPO directly builds upon and adapts for model-based policy evaluation.
  • Paper: Trust Region Policy Optimization, John Schulman et al. (2015). Trust Region Policy Optimization develops practical monotonic improvement bounds via policy divergence constraints, providing the core theoretical mechanics extended by MBPO.
  • Paper: Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models, Kurtland Chua et al. (2018). PETS introduces ensembles of deep probabilistic dynamics models to combat model bias and epistemic uncertainty, which serves as the dynamics modeling backbone employed in MBPO.
  • Paper: PILCO: A Model-Based and Data-Efficient Approach to Policy Search, Marc Peter Deisenroth et al. (2011). PILCO introduces principled uncertainty-aware dynamics modeling to mitigate model exploitation in model-based reinforcement learning, framing the key problem MBPO addresses.
  • Paper: Constrained Policy Optimization, Joshua Achiam et al. (2017). Constrained Policy Optimization extends conservative policy iteration bounds to continuous state-action spaces, directly informing MBPO's bounding techniques under model error.
Cover for When to Trust Your Model: Model-Based Policy Optimization

Abstract

Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate and analyze a model-based reinforcement learning algorithm with a guarantee of monotonic improvement at each step. In practice, this analysis is overly pessimistic and suggests that real off-policy data is always preferable to model-generated on-policy data, but we show that an empirical estimate of model generalization can be incorporated into such analysis to justify model usage. Motivated by this analysis, we then demonstrate that a simple procedure of using short model-generated rollouts branched from real data has the benefits of more complicated model-based algorithms without the usual pitfalls. In particular, this approach surpasses the sample efficiency of prior model-based methods, matches the asymptotic performance of the best model-free algorithms, and scales to horizons that cause other model-based methods to fail entirely.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Background
  • 4 Monotonic improvement with model bias
  • 4.1 Monotonic model-based improvement
  • 4.2 Interpolating model-based and model-free updates
  • 4.3 Model generalization in practice
  • 5 Model-based policy optimization with deep reinforcement learning
  • 6 Experiments
  • 6.1 Comparative evaluation
  • 6.2 Design evaluation
  • 7 Discussion
  • References
  • A Model-based Policy Optimization with Performance Guarantees
  • B Useful Lemmas
  • C Hyperparameter Settings

Knowls

  1. Knowl 1 — Model-Based Policy Optimization (MBPO) Algorithm

    algorithm

    Model-Based Policy Optimization (MBPO) is a reinforcement learning algorithm that trains an off-policy model-free agent (Soft Actor-Critic) using short, branched rollouts generated from an ensemble of learned dynamics models.

    Input: Policy πϕ\pi_\phi, predictive dynamics model pθp_\theta, real environment dataset Denv\mathcal{D}_{\text{env}}, model dataset Dmodel\mathcal{D}_{\text{model}}
    Initialize policy parameters ϕ\phi and model parameters θ\theta
    Initialize Denv←∅\mathcal{D}_{\text{env}} \leftarrow \emptyset, Dmodel←∅\mathcal{D}_{\text{model}} \leftarrow \emptyset
    for epoch =1,2,…,N= 1, 2, \dots, N do
        Train predictive model pθp_\theta on Denv\mathcal{D}_{\text{env}} via maximum likelihood
        for step =1,2,…,E= 1, 2, \dots, E do
            Take action in environment according to πϕ(a∣s)\pi_\phi(a|s); add transition (s,a,s′,r)(s, a, s', r) to Denv\mathcal{D}_{\text{env}}
            for rollout =1,2,…,M= 1, 2, \dots, M do
                Sample state sts_t uniformly at random from Denv\mathcal{D}_{\text{env}}
                Perform kk-step rollout starting from sts_t under policy πϕ\pi_\phi and dynamics pθp_\theta
                Add all generated transitions to Dmodel\mathcal{D}_{\text{model}}
            end for
            for update =1,2,…,G= 1, 2, \dots, G do
                Sample a batch of transitions from Dmodel\mathcal{D}_{\text{model}}
                Update Q-function and policy parameters ϕ\phi via Soft Actor-Critic gradient steps
            end for
        end for
    end for

    The algorithm leverages short model rollouts of length kk rooted at real states sampled from Denv\mathcal{D}_{\text{env}}. This generates a large volume of synthetic transitions ({st,at,rt,st+1}\{s_t, a_t, r_t, s_{t+1}\}), allowing the agent to perform G∈[20,40]G \in [20, 40] policy optimization gradient updates per single environment interaction step without suffering from catastrophic model divergence.

  2. Knowl 2 — Return Lower Bound for k-Branched Rollouts under Generalizing Dynamics Models

    theoretical result

    Let an MDP have discount factor γ∈(0,1)\gamma \in (0, 1) and reward bounded by rmax⁡r_{\max}, i.e., max⁡s,a∣r(s,a)∣≤rmax⁡\max_{s, a} |r(s, a)| \le r_{\max}. In a kk-branched rollout, trajectories are generated by executing a data-collecting policy πD\pi_D under the true dynamics pp up to a branch point sampled with probability proportional to γt\gamma^t, and subsequently executing the candidate policy π\pi for kk steps under the learned dynamics model p^\hat{p}.

    Let the model error under the updated policy π\pi be bounded across all timesteps tt by: ϵm′≥max⁡tEs∼πt[DTV(p(s′∣s,a)∥p^(s′∣s,a))]\epsilon_{m'} \ge \max_t \mathbb{E}_{s \sim \pi_t} [D_{TV}(p(s'|s, a) \parallel \hat{p}(s'|s, a))] and the policy divergence be bounded by max⁡sDTV(πD(a∣s)∥π(a∣s))≤ϵπ\max_s D_{TV}(\pi_D(a|s) \parallel \pi(a|s)) \le \epsilon_\pi.

    The true expected discounted return η[π]=Eπ[∑t=0∞γtr(st,at)]\eta[\pi] = \mathbb{E}_\pi \left[ \sum_{t=0}^\infty \gamma^t r(s_t, a_t) \right] is bounded below by the expected return under kk-branched rollouts ηbranch[π]\eta^{\text{branch}}[\pi] according to: η[π]≥ηbranch[π]−2rmax⁡[γk+1ϵπ(1−γ)2+γkϵπ1−γ+k1−γϵm′]\eta[\pi] \ge \eta^{\text{branch}}[\pi] - 2 r_{\max} \left[ \frac{\gamma^{k+1} \epsilon_\pi}{(1 - \gamma)^2} + \frac{\gamma^k \epsilon_\pi}{1 - \gamma} + \frac{k}{1 - \gamma} \epsilon_{m'} \right]

    Unlike bounds based purely on static dataset error, this bound justifies non-zero model rollouts: when the generalization error ϵm′\epsilon_{m'} is sufficiently low, the optimal model rollout horizon: k∗=arg⁡min⁡k[γk+1ϵπ(1−γ)2+γkϵπ1−γ+k1−γϵm′]k^* = \arg\min_k \left[ \frac{\gamma^{k+1} \epsilon_\pi}{(1 - \gamma)^2} + \frac{\gamma^k \epsilon_\pi}{1 - \gamma} + \frac{k}{1 - \gamma} \epsilon_{m'} \right] satisfies k∗>0k^* > 0, analytically validating the use of short, truncated model rollouts.

  3. Knowl 3 — Return Discrepancy Bound for Full Model Rollouts

    theoretical result

    In a Markov decision process with discount factor γ∈(0,1)\gamma \in (0, 1) and maximum reward magnitude rmax⁡=max⁡s,a∣r(s,a)∣r_{\max} = \max_{s, a} |r(s, a)|, let η[π]\eta[\pi] denote the expected discounted return of policy π\pi under the true transition dynamics p(s′∣s,a)p(s'|s, a), and let η^[π]\hat{\eta}[\pi] denote its expected discounted return under a learned transition model p^(s′∣s,a)\hat{p}(s'|s, a).

    If the expected total variation (TV) distance between the true and learned dynamics distributions under the state distribution of a data-collecting policy πD\pi_D is bounded at each timestep by: max⁡tEs∼πD,t[DTV(p(s′,r∣s,a)∥p^(s′,r∣s,a))]≤ϵm\max_t \mathbb{E}_{s \sim \pi_{D, t}} [D_{TV}(p(s', r|s, a) \parallel \hat{p}(s', r|s, a))] \le \epsilon_m and the policy divergence is bounded by max⁡sDTV(πD(a∣s)∥π(a∣s))≤ϵπ\max_s D_{TV}(\pi_D(a|s) \parallel \pi(a|s)) \le \epsilon_\pi, then the true return satisfies: η[π]≥η^[π]−[2γrmax⁡(ϵm+2ϵπ)(1−γ)2+4rmax⁡ϵπ1−γ]\eta[\pi] \ge \hat{\eta}[\pi] - \left[ \frac{2\gamma r_{\max}(\epsilon_m + 2\epsilon_\pi)}{(1 - \gamma)^2} + \frac{4 r_{\max} \epsilon_\pi}{1 - \gamma} \right]

    This provides a monotonic improvement guarantee: optimizing η^[π]\hat{\eta}[\pi] under the learned model guarantees improvement on the true environment as long as the model return improvement exceeds the error penalty term C(ϵm,ϵπ)=2γrmax⁡(ϵm+2ϵπ)(1−γ)2+4rmax⁡ϵπ1−γC(\epsilon_m, \epsilon_\pi) = \frac{2\gamma r_{\max}(\epsilon_m + 2\epsilon_\pi)}{(1 - \gamma)^2} + \frac{4 r_{\max} \epsilon_\pi}{1 - \gamma}.

  4. Knowl 4 — Pessimistic Return Lower Bound for k-Branched Rollouts

    theoretical result

    Let ηbranch[π]\eta^{\text{branch}}[\pi] denote the expected return of policy π\pi under a kk-branched rollout procedure branching from real data generated by πD\pi_D. If model error is bounded exclusively on the training distribution of πD\pi_D such that max⁡tEs∼πD,t[DTV(p(s′∣s,a)∥p^(s′∣s,a))]≤ϵm\max_t \mathbb{E}_{s \sim \pi_{D, t}} [D_{TV}(p(s'|s, a) \parallel \hat{p}(s'|s, a))] \le \epsilon_m, and policy divergence is bounded by max⁡sDTV(πD(a∣s)∥π(a∣s))≤ϵπ\max_s D_{TV}(\pi_D(a|s) \parallel \pi(a|s)) \le \epsilon_\pi, then the true return η[π]\eta[\pi] is lower-bounded by: η[π]≥ηbranch[π]−2rmax⁡[γk+1ϵπ(1−γ)2+γk+21−γϵπ+k1−γ(ϵm+2ϵπ)]\eta[\pi] \ge \eta^{\text{branch}}[\pi] - 2 r_{\max} \left[ \frac{\gamma^{k+1} \epsilon_\pi}{(1 - \gamma)^2} + \frac{\gamma^k + 2}{1 - \gamma} \epsilon_\pi + \frac{k}{1 - \gamma} (\epsilon_m + 2\epsilon_\pi) \right]

    Because the linear penalty in kk involves the pessimistic distribution shift estimate ϵm+2ϵπ\epsilon_m + 2\epsilon_\pi, the lower bound is mathematically maximized at k=0k = 0, reflecting the standard pessimistic conclusion that pure real off-policy data is preferable to model data in the absence of model generalization assumptions.

  5. Knowl 5 — Bootstrap Ensemble of Probabilistic Dynamics Models

    model/method

    Dynamics modeling in MBPO uses an ensemble of BB probabilistic neural networks {pθ1,…,pθB}\{p_\theta^1, \dots, p_\theta^B\}. Each model member outputs parameters of a Gaussian distribution with diagonal covariance: pθi(st+1,r∣st,at)=N(μθi(st,at),Σθi(st,at))p_\theta^i(s_{t+1}, r|s_t, a_t) = \mathcal{N}\left(\mu_\theta^i(s_t, a_t), \Sigma_\theta^i(s_t, a_t)\right)

    This architecture accounts for two sources of uncertainty:

    1. Aleatoric uncertainty (stochasticity and noise in state transitions), captured by the predicted Gaussian variances Σθi(st,at)\Sigma_\theta^i(s_t, a_t).
    2. Epistemic uncertainty (model parameter uncertainty due to finite data), captured via the bootstrap ensemble of independently initialized and trained networks.

    To generate synthetic rollouts without compounding correlation errors, transitions are sampled along the rollout by selecting a model uniformly at random from the ensemble at each step.

  6. Knowl 6 — Sample Efficiency and Performance of MBPO on Continuous Control Benchmarks

    empirical result

    MBPO was evaluated on standard 1000-step horizon MuJoCo benchmarks (InvertedPendulum, Hopper, Walker2d, Ant, HalfCheetah, Humanoid) against model-free methods (SAC, PPO) and model-based algorithms (PETS, STEVE, SLBO).

    Key empirical findings include:

    1. MBPO achieves the asymptotic performance of model-free algorithms (matching SAC) while attaining a 10-fold to 20-fold improvement in sample efficiency over SAC.
    2. On the Ant benchmark, MBPO at 300,000 environment steps reaches a return of ~5,500, matching the performance of SAC at 3,000,000 steps.
    3. On Hopper and Walker2d, MBPO learns near-optimal policies with 14 minutes and 40 minutes of simulated real-time interaction, respectively.
    4. Unlike planning-based model-based baselines (e.g., PETS) that degrade or fail on high-dimensional systems like Ant and Humanoid, MBPO scales effectively to these complex state spaces.
  7. Knowl 7 — Empirical Impact of Rollout Horizon on Model-Based Policy Optimization

    empirical result

    Ablations on the rollout length kk within MBPO reveal clear trade-offs between sample efficiency and model error compounding:

    1. Setting a fixed single-step model rollout (k=1k = 1) provides a strong baseline that retains most of the benefit of model-based acceleration and outperforms methods relying on long rollouts from the initial state distribution.
    2. Linearly increasing the rollout length over training epochs (e.g., from k=1k = 1 to k=15k = 15 or k=25k = 25) achieves the best overall performance and final return.
    3. Extended model rollouts (k=200k = 200) degrade policy performance due to accumulated transition error, and rollouts of length k=500k = 500 fail entirely to learn effective policies.
    4. High policy gradient update ratios (G=20G = 20 to 4040 updates per environment step) are stable and effective only when trained with model-augmented transitions; running pure SAC with high GG fails to yield comparable sample efficiency.
  8. Knowl 8 — Mitigation of Model Exploitation via Short Branched Rollouts

    empirical result

    Evaluating multi-step rollouts and policy return distributions demonstrates how short rollouts prevent model exploitation:

    1. Long action sequences (e.g., 450-step hopping sequences) rolled out under learned dynamics exhibit rapidly growing prediction variance and kinematic deterioration due to compounding error.
    2. When policies are trained using short branched rollouts (small kk), the cumulative returns under the model correlate strongly with returns in the true environment, indicating that the policy optimization does not exploit model inaccuracies.
    3. Across MuJoCo benchmark environments, returns evaluated under the learned model tend to underestimate, rather than overestimate, real environment returns during MBPO training.
  9. Knowl 9 — Hyperparameter Configurations for MBPO Across Continuous Control Benchmarks

    data/table

    The hyperparameters used by MBPO across continuous control benchmark environments are specified below. Rollout schedules x→yx \to y over epochs a→ba \to b denote the thresholded linear function f(e)=min⁡(max⁡(x+e−ab−a(y−x),x),y)f(e) = \min\left(\max\left(x + \frac{e-a}{b-a}(y-x), x\right), y\right) at epoch ee.

    Hyperparameter HalfCheetah InvertedPendulum Walker2d Ant Hopper Humanoid
    Epochs (NN) 400 15 300 300 125 300
    Env steps / epoch (EE) 1000 1000 1000 1000 1000 1000
    Model rollouts / step (MM) 400 400 400 400 400 400
    Ensemble size (BB) 7 7 7 7 7 7
    Network architecture MLP 4×2004\times 200 MLP 4×2004\times 200 MLP 4×2004\times 200 MLP 4×2004\times 200 MLP 4×2004\times 200 MLP 4×4004\times 400
    Policy updates / step (GG) 40 20 20 20 20 20
    Model horizon (kk) 1 1 1→251 \to 25 (20→10020 \to 100) 1→251 \to 25 (20→10020 \to 100) 1→151 \to 15 (20→10020 \to 100) 1→251 \to 25 (20→30020 \to 300)

    All tasks employ an ensemble of 7 probabilistic neural networks and sample 400 short model rollouts per environment step. The policy update ratio GG is set to 20 for most tasks and 40 for HalfCheetah, while the model horizon kk increases dynamically from 1 up to 15 or 25 as model fidelity improves with collected data.

  10. Knowl 10 — Empirical Characterization of Model Generalization Under Policy Shift

    empirical result

    Empirical measurements of dynamics model generalization show that model error under a new policy π\pi, denoted ϵm′\epsilon_{m'}, scales with policy divergence ϵπ=DKL(π∥πD)\epsilon_\pi = D_{KL}(\pi \parallel \pi_D) from the data-collecting policy πD\pi_D.

    Key empirical findings:

    1. Model error increases approximately linearly with policy shift: ϵ^m′(ϵπ)≈ϵm+ϵπdϵm′dϵπ\hat{\epsilon}_{m'}(\epsilon_\pi) \approx \epsilon_m + \epsilon_\pi \frac{d\epsilon_{m'}}{d\epsilon_\pi}.
    2. The local sensitivity of model error to policy divergence, measured by the derivative dϵm′dϵπ∣ϵπ=0\frac{d\epsilon_{m'}}{d\epsilon_\pi}\Big|_{\epsilon_\pi = 0}, decreases monotonically by orders of magnitude as the training dataset size increases (from 5k5\text{k} to 55k55\text{k} transitions on Hopper and Walker2d).
    3. Models trained on larger datasets not only achieve lower error on their training distribution, but also exhibit substantially improved out-of-distribution generalization to state distributions induced by nearby updated policies.

Coverage note — None was omitted; all major theoretical bounds, algorithmic designs, empirical evaluations, and ablation studies from the paper have been extracted.

References

  1. 1.Asadi, K., Misra, D., and Littman, M. Lipschitz continuity in model-based reinforcement learning. In International Conference on Machine Learning, 2018.
  2. 2.Atkeson, C. G. and Schaal, S. Learning tasks from a single demonstration. In International Conference on Robotics and Automation, 1997.
  3. 3.Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H. Sample-efficient reinforcement learning with stochastic ensemble value expansion. In Advances in Neural Information Processing Systems, 2018.
  4. 4.Chua, K., Calandra, R., McAllister, R., and Levine, S. Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In Advances in Neural Information Processing Systems. 2018.
  5. 5.Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P. Model-based reinforcement learning via meta-policy optimization. In Conference on Robot Learning, 2018.
  6. 6.Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S. On the sample complexity of the linear quadratic regulator. arXiv preprint arXiv:1710.01688, 2017.
  7. 7.Deisenroth, M. and Rasmussen, C. E. PILCO: A model-based and data-efficient approach to policy search. In International Conference on Machine Learning, 2011.
  8. 8.Depeweg, S., Hernández-Lobato, J. M., Doshi-Velez, F., and Udluft, S. Learning and policy search in stochastic dynamical systems with bayesian neural networks. In International Conference on Learning Representations, 2016.
  9. 9.Draeger, A., Engell, S., and Ranke, H. Model predictive control using neural networks. IEEE Control Systems Magazine, 1995.
  10. 10.Du, Y. and Narasimhan, K. Task-agnostic dynamics priors for deep reinforcement learning. In International Conference on Machine Learning, 2019.
  11. 11.Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A. X., and Levine, S. Visual foresight: Model-based deep reinforcement learning for vision-based robotic control. arXiv preprint arXiv:1812.00568, 2018.
  12. 12.Farahmand, A.-M., Barreto, A., and Nikovski, D. Value-aware loss function for model-based reinforcement learning. In International Conference on Artificial Intelligence and Statistics, 2017.
  13. 13.Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S. Model-based value estimation for efficient model-free reinforcement learning. In International Conference on Machine Learning, 2018.
  14. 14.Gal, Y., McAllister, R., and Rasmussen, C. E. Improving PILCO with Bayesian neural network dynamics models. In ICML Workshop on Data-Efficient Machine Learning Workshop, 2016.
  15. 15.Gu, S., Lillicrap, T., Sutskever, I., and Levine, S. Continuous deep Q-learning with model-based acceleration. In International Conference on Machine Learning, 2016.
  16. 16.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, 2018.
  17. 17.Heess, N., Wayne, G., Silver, D., Lillicrap, T., Tassa, Y., and Erez, T. Learning continuous control policies by stochastic value gradients. In Advances in Neural Information Processing Systems, 2015.
  18. 18.Holland, G. Z., Talvitie, E. J., and Bowling, M. The effect of planning shape on dyna-style planning in high-dimensional state spaces. arXiv preprint arXiv:1806.01825, 2018.
  19. 19.Kaelbling, L. P., Littman, M. L., and Moore, A. P. Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4:237–285, 1996.
  20. 20.Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowsi, P., Levine, S., Sepassi, R., Tucker, G., and Michalewski, H. Model-based reinforcement learning for Atari. arXiv preprint arXiv:1903.00374, 2019.
  21. 21.Kalweit, G. and Boedecker, J. Uncertainty-driven imagination for continuous deep reinforcement learning. In Conference on Robot Learning, 2017.
  22. 22.Kumar, V., Todorov, E., and Levine, S. Optimal control with learned local models: Application to dexterous manipulation. In International Conference on Robotics and Automation, 2016.
  23. 23.Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P. Model-ensemble trust-region policy optimization. In International Conference on Learning Representations, 2018.
  24. 24.Levine, S. and Koltun, V. Guided policy search. In International Conference on Machine Learning, 2013.
  25. 25.Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. Continuous control with deep reinforcement learning. In International Conference on Learning Representations, 2016.
  26. 26.Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T. Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees. In International Conference on Learning Representations, 2019.
  27. 27.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. Human-level control through deep reinforcement learning. Nature, 2015.
  28. 28.Nagabandi, A., Kahn, G., S. Fearing, R., and Levine, S. Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning. In International Conference on Robotics and Automation, 2018.
  29. 29.Oh, J., Guo, X., Lee, H., Lewis, R., and Singh, S. Action-conditional video prediction using deep networks in Atari games. In Advances in Neural Information Processing Systems, 2015.
  30. 30.Oh, J., Singh, S., and Lee, H. Value prediction network. In Advances in Neural Information Processing Systems, 2017.
  31. 31.Piche, A., Thomas, V., Ibrahim, C., Bengio, Y., and Pal, C. Probabilistic planning with sequential Monte Carlo methods. In International Conference on Learning Representations, 2019.
  32. 32.Racaniere, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Jimenez Rezende, D., Puigdomenech Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Hassabis, D., Silver, D., and Wierstra, D. Imagination-augmented agents for deep reinforcement learning. In Advances in Neural Information Processing Systems. 2017.
  33. 33.Rajeswaran, A., Ghotra, S., Levine, S., and Ravindran, B. EPOpt: Learning robust neural network policies using model ensembles. In International Conference on Learning Representations, 2017.
  34. 34.Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. Trust region policy optimization. In International Conference on Machine Learning, 2015.
  35. 35.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  36. 36.Shalev-Shwartz, S. and Ben-David, S. Understanding machine learning: From theory to algorithms. Cambridge university press, 2014.
  37. 37.Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., and Degris, T. The predictron: End-to-end learning and planning. In International Conference on Machine Learning, 2017.
  38. 38.Sun, W., Gordon, G. J., Boots, B., and Bagnell, J. Dual policy iteration. In Advances in Neural Information Processing Systems, 2018.
  39. 39.Sutton, R. S. Integrated architectures for learning, planning, and reacting based on approximating dynamic programming. In International Conference on Machine Learning, 1990.
  40. 40.Szita, I. and Szepesvari, C. Model-based reinforcement learning with nearly tight exploration complexity bounds. In International Conference on Machine Learning, 2010.
  41. 41.Talvitie, E. Model regularization for stable sample rollouts. In Conference on Uncertainty in Artificial Intelligence, 2014.
  42. 42.Talvitie, E. Self-correcting models for model-based reinforcement learning. In AAAI Conference on Artificial Intelligence, 2016.
  43. 43.Tamar, A., WU, Y., Thomas, G., Levine, S., and Abbeel, P. Value iteration networks. In Advances in Neural Information Processing Systems. 2016.
  44. 44.Todorov, E., Erez, T., and Tassa, Y. MuJoCo: A physics engine for model-based control. In International Conference on Intelligent Robots and Systems, 2012.
  45. 45.Whitney, W. and Fergus, R. Understanding the asymptotic performance of model-based RL methods. 2018.

Citation

MLA
Janner, M., et al. “When to Trust Your Model: Model-Based Policy Optimization”. Advances in Neural Information Processing Systems, vol. 32, 2019, https://proceedings.neurips.cc/paper_files/paper/2019/file/5faf461eff3099671ad63c6f3f094f7f-Paper.pdf.
APA
Janner, M., Fu, J., Zhang, M., & Levine, S. (2019). When to Trust Your Model: Model-Based Policy Optimization. Advances in Neural Information Processing Systems, 32. https://proceedings.neurips.cc/paper_files/paper/2019/file/5faf461eff3099671ad63c6f3f094f7f-Paper.pdf
Chicago
Janner, M., J. Fu, M. Zhang, and S. Levine. 2019. “When to Trust Your Model: Model-Based Policy Optimization”. Advances in Neural Information Processing Systems 32. https://proceedings.neurips.cc/paper_files/paper/2019/file/5faf461eff3099671ad63c6f3f094f7f-Paper.pdf.
Harvard
Janner, M. et al. (2019) “When to Trust Your Model: Model-Based Policy Optimization”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2019/file/5faf461eff3099671ad63c6f3f094f7f-Paper.pdf.
Vancouver
1. Janner M, Fu J, Zhang M, Levine S (2019) When to Trust Your Model: Model-Based Policy Optimization. Advances in Neural Information Processing Systems 32:

BibTeX

@inproceedings{janner2019when,
  title = {When to Trust Your Model: Model-Based Policy Optimization},
  author = {Janner, Michael and Fu, Justin and Zhang, Marvin and Levine, Sergey},
  year = {2019},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {32},
  url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/5faf461eff3099671ad63c6f3f094f7f-Paper.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors