When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning

Haoyi NiuShubham SharmaYiwen QiuMing LiGuyue ZhouJianming HuXianyuan Zhan

article2022NeurIPS80 citations

Proposes the H2O framework, a hybrid reinforcement learning method that bridges sim-to-real dynamics gaps by adaptively penalizing value function estimates during simulation rollouts while learning directly from limited real-world datasets.

Listen

Deploying reinforcement learning algorithms to control physical, real-world systems remains difficult due to fundamental data constraints. Traditional online learning requires millions of trial-and-error interactions that are unsafe or impractical on real hardware, while low-cost computer simulators introduce dynamic discrepancies that cause policies to fail when deployed in reality. Conversely, pure offline learning trains solely on pre-collected historical datasets without physical risk, but its effectiveness is heavily bottlenecked by limited state-action coverage and overly conservative constraints. Combining limited real data with unrestricted simulation exploration is therefore critical to making reinforcement learning practical for high-stakes industrial and robotic applications.

The article develops and evaluates the Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning (H2O) framework. The primary objective is to demonstrate that an agent can combine pre-collected real-world data with online simulation interactions by adaptively penalizing simulated transitions with large physical discrepancies, bridging the gap between simulation and reality without requiring an exact digital twin.

To evaluate this framework, the authors conducted theoretical analyses and extensive benchmark testing across both simulated physics environments and physical hardware. The simulated benchmarks altered standard MuJoCo locomotion tasks with substantial physical distortions, such as doubling gravity, reducing friction to 30%, and injecting control noise, while using standard offline datasets of varying quality. For real-world validation, the authors tested the approach on a physical two-wheeled balancing robot using 100,000 human-controlled transitions alongside a simulator with significant motor dead zones and unmodeled friction. The framework was benchmarked against leading purely online, purely offline, and hybrid domain-adaptation algorithms.

The findings show that H2O systematically outperformed all baseline methods across nearly all benchmark scenarios. In simulated tasks, H2O achieved the highest average returns, frequently exceeding purely offline algorithms by 15% to 25% and outperforming purely online simulation baselines by more than 50%. In physical robot deployments, H2O maintained stable balance and smooth tracking at the target speed of 0.2 meters per second, whereas purely offline baselines exceeded target speeds by nearly 100% before losing stability, and online baselines failed immediately. Ablation tests confirmed that both adaptive value regularization and importance-weighted updates are essential, as removing the dynamics ratio correction caused performance to drop by more than 30%.

These results demonstrate that organizations do not need perfectly calibrated, expensive simulators to train capable autonomous policies, provided real-world data is used to dynamically identify and discount inaccurate simulated dynamics. Crucially, the findings reveal that high performance within an uncalibrated simulator does not correlate with real-world success, indicating that simulation-only verification poses serious operational and safety risks for physical deployments. Incorporating hybrid offline-and-online training mitigates the risk of policy failure while reducing the cost of extensive physical data collection.

Organizations developing autonomous robotics and industrial control systems should consider adopting hybrid training architectures rather than relying strictly on domain randomization or purely offline datasets. Stakeholders should also re-examine safety verification protocols to ensure policies are not evaluated solely against simulated metrics. Before deploying this approach to mission-critical systems, teams should conduct real-world pilot tests to confirm that baseline data collection sufficiently covers operational boundaries.

Confidence in these findings is strong across the tested mechanical control and locomotion regimes, supported by both formal proofs and physical robot validations. However, readers should note certain limitations: the method relies on statistical discriminators and Gaussian approximations to estimate physical discrepancies, and its underlying algorithm inherits some conservative constraints. Further validation is required for higher-dimensional tasks with severe visual shifts or unobservable environmental states.

Cover for When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning

Abstract

Learning effective reinforcement learning (RL) policies to solve real-world complex tasks can be quite challenging without a high-fidelity simulation environment. In most cases, we are only given imperfect simulators with simplified dynamics, which inevitably lead to severe sim-to-real gaps in RL policy learning. The recently emerged field of offline RL provides another possibility to learn policies directly from pre-collected historical data. However, to achieve reasonable performance, existing offline RL algorithms need impractically large offline data with sufficient state-action space coverage for training. This brings up a new question: is it possible to combine learning from limited real data in offline RL and unrestricted exploration through imperfect simulators in online RL to address the drawbacks of both approaches? In this study, we propose the Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning (H2O) framework to provide an affirmative answer to this question. H2O introduces a dynamics-aware policy evaluation scheme, which adaptively penalizes the Q-function learning on simulated state-action pairs with large dynamics gaps, while also simultaneously allowing learning from a fixed real-world dataset. Through extensive simulation and real-world tasks, as well as theoretical analysis, we demonstrate the superior performance of H2O against other cross-domain online and offline RL algorithms. H2O provides a brand new hybrid offline-and-online RL paradigm, which can potentially shed light on future RL algorithm design for solving practical real-world tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Reinforcement Learning Using Simulators with Dynamics Gap
  • 2.2 Offline Reinforcement Learning
  • 3 Background
  • 3.1 Reinforcement Learning
  • 3.2 Offline RL via Value Regularization
  • 4 Hybrid Offline-and-Online Reinforcement Learning
  • 4.1 Incorporating Offline Data in Online Learning
  • 4.2 Adaptive Value Regularization on High Dynamics-Gap Samples
  • 4.3 Fixing Bellman Error due to Dynamics Gap
  • 4.4 Practical Implementation
  • 5 Interpretation of Dynamics-Aware Policy Evaluation
  • 6 Experiments
  • 6.1 Experimental Environment Setups
  • 6.2 Baselines
  • 6.3 Comparative Evaluation of H2O in Simulation and Real-World Experiments
  • 6.4 Ablation Study
  • 7 Conclusion and Perspectives
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Dynamics-aware hybrid policy evaluation

    model/method

    H2O is a hybrid reinforcement-learning framework that trains a policy using both a fixed real-world dataset D\mathcal D and online rollouts from an imperfect simulator with transition kernel PM^P_{\hat M}, while the real system has transition kernel PMP_M. Its critic is trained by minimizing

    LH2O(Q)=β(log⁡∑s,aω(s,a)exp⁡(Q(s,a))−E(s,a)∼D[Q(s,a)])+E~(Q,B^πQ^),\mathcal L_{\mathrm{H2O}}(Q)=\beta\left(\log\sum_{s,a}\omega(s,a)\exp(Q(s,a))-\mathbb E_{(s,a)\sim\mathcal D}[Q(s,a)]\right)+\widetilde{\mathcal E}(Q,\widehat{\mathcal B}^{\pi}\widehat Q),

    where QQ is the learned action-value function, Q^\widehat Q is a target critic, π\pi is the learned policy, β>0\beta>0 controls value regularization, ω(s,a)\omega(s,a) is a normalized distribution representing the dynamics gap at state-action pair (s,a)(s,a), and E~\widetilde{\mathcal E} is a Bellman-error loss corrected for the simulator’s transition bias. The log-sum-exp term penalizes high values at state-action pairs with large ω(s,a)\omega(s,a), whereas the real-data term offsets this penalty on state-action pairs supported by the trustworthy dataset. Thus, H2O supplements the limited coverage of D\mathcal D with broad simulator exploration without treating all simulated transitions as equally reliable.

  2. Knowl 2 — Estimating dynamics gaps with transition-density ratios

    equation

    For a state-action pair (s,a)(s,a), H2O measures the simulator-to-real dynamics gap using the Kullback–Leibler divergence

    u(s,a)=DKL ⁣(PM^(s′∣s,a) ∥ PM(s′∣s,a))=Es′∼PM^(⋅∣s,a)[log⁡PM^(s′∣s,a)PM(s′∣s,a)],u(s,a)=D_{\mathrm{KL}}\!\left(P_{\hat M}(s'\mid s,a)\,\middle\|\,P_M(s'\mid s,a)\right)=\mathbb E_{s'\sim P_{\hat M}(\cdot\mid s,a)}\left[\log\frac{P_{\hat M}(s'\mid s,a)}{P_M(s'\mid s,a)}\right],

    where s′s' is the next state, PMP_M is the real transition distribution, and PM^P_{\hat M} is the simulator transition distribution. The regularization distribution is the normalized gap measure ω(s,a)=u(s,a)/∑s~,a~u(s~,a~)\omega(s,a)=u(s,a)/\sum_{\tilde s,\tilde a}u(\tilde s,\tilde a).

    H2O estimates the transition-density ratio using two binary discriminators. Let pR(s,a,s′)=p(real∣s,a,s′)p_R(s,a,s')=p(\mathrm{real}\mid s,a,s') and pR(s,a)=p(real∣s,a)p_R(s,a)=p(\mathrm{real}\mid s,a) be the probabilities that a transition is classified as real rather than simulated. Bayes’ rule gives

    PM^(s′∣s,a)PM(s′∣s,a)=1−pR(s,a,s′)pR(s,a,s′)1−pR(s,a)pR(s,a).\frac{P_{\hat M}(s'\mid s,a)}{P_M(s'\mid s,a)}= \frac{\frac{1-p_R(s,a,s')}{p_R(s,a,s')}}{\frac{1-p_R(s,a)}{p_R(s,a)}}.

    The probabilities are approximated by discriminators trained with cross-entropy on real offline transitions and simulated transitions: one discriminator receives (s,a,s′)(s,a,s'), and the other receives (s,a)(s,a). The resulting ratio is used both to construct ω(s,a)\omega(s,a) and to correct Bellman errors on simulated data.

  3. Knowl 3 — Importance-weighted Bellman error for simulated transitions

    equation

    Because a simulated transition (s,a,s′)(s,a,s') uses s′∼PM^(⋅∣s,a)s'\sim P_{\hat M}(\cdot\mid s,a) rather than the real next-state distribution, directly applying the Bellman loss to simulator data can train the critic toward incorrect targets. H2O instead uses

    E~(Q,B^πQ^)=12E(s,a,s′)∼D[δQ(s,a,s′)2]+12E(s,a,s′)∼B[PM(s′∣s,a)PM^(s′∣s,a) δQ(s,a,s′)2],\widetilde{\mathcal E}(Q,\widehat{\mathcal B}^{\pi}\widehat Q) =\frac{1}{2}\mathbb E_{(s,a,s')\sim\mathcal D}\left[\delta_Q(s,a,s')^2\right] +\frac{1}{2}\mathbb E_{(s,a,s')\sim\mathcal B}\left[\frac{P_M(s'\mid s,a)}{P_{\hat M}(s'\mid s,a)}\,\delta_Q(s,a,s')^2\right],

    where B\mathcal B is the simulated replay buffer and

    δQ(s,a,s′)=Q(s,a)−[r(s,a)+γ Ea′∼π(⋅∣s′)Q^(s′,a′)].\delta_Q(s,a,s')=Q(s,a)-\left[r(s,a)+\gamma\,\mathbb E_{a'\sim\pi(\cdot\mid s')}\widehat Q(s',a')\right].

    Here r(s,a)r(s,a) is the reward, γ∈(0,1)\gamma\in(0,1) is the discount factor, and the ratio PM/PM^P_M/P_{\hat M} is estimated from the discriminators. The real-data term uses the observed real transition without correction; the simulated-data term reweights each squared Bellman residual so that its expectation corresponds to sampling next states from the real dynamics.

  4. Knowl 4 — H2O training procedure with actor-critic learning

    algorithm

    H2O can be instantiated with Soft Actor-Critic. Its inputs are a fixed real-world transition dataset D\mathcal D, an imperfect simulator M^\hat M, and training length TT. It maintains a critic QθQ_\theta, target critic QθˉQ_{\bar\theta}, stochastic actor πϕ\pi_\phi, simulated replay buffer B\mathcal B, and the two transition discriminators used to estimate the dynamics ratio.

    Input: real offline dataset D\mathcal D, imperfect simulator M^\hat M, training steps TT
    Initialize critic QθQ_\theta, target critic QθˉQ_{\bar\theta}, actor πϕ\pi_\phi, and empty simulated buffer B\mathcal B
    Initialize discriminators DsasD_{sas} and DsaD_{sa}
    for step t=1,…,Tt=1,\ldots,T do
        Collect a rollout with πϕ\pi_\phi in M^\hat M and append its transitions to B\mathcal B
        Update DsasD_{sas} using cross-entropy on real transitions from D\mathcal D and simulated transitions from B\mathcal B
        Update DsaD_{sa} using cross-entropy on real state-action pairs from D\mathcal D and simulated pairs from B\mathcal B
        Estimate the transition ratio, dynamics-gap weights ω\omega, and corrected Bellman loss E~\widetilde{\mathcal E}
        Update θ\theta by minimizing the H2O critic objective
        Update ϕ\phi by ascending the gradient of E(s,a)∼D∪B[Qθ(s,a)−λlog⁡πϕ(a∣s)]\mathbb E_{(s,a)\sim\mathcal D\cup\mathcal B}[Q_\theta(s,a)-\lambda\log\pi_\phi(a\mid s)]
        Every target-update period, set θˉ←(1−τ)θˉ+τθ\bar\theta\leftarrow(1-\tau)\bar\theta+\tau\theta
    end for
    Output: actor πϕ\pi_\phi

    The temperature λ\lambda in the entropy-regularized actor objective is automatically tuned. In implementation, the state-action sum in the log-sum-exp critic term is approximated with simulated replay-buffer minibatches. Because a black-box simulator may not provide samples from its transition distribution for arbitrary (s,a)(s,a), the expectation in the KL gap estimate is approximated using NN random next states drawn from a Gaussian centered at the observed simulated next state, with covariance equal to the state covariance computed from the real offline dataset.

  5. Knowl 5 — Adaptive reward interpretation and safety property

    theoretical result

    Under the approximation that the weighted log-sum-exp term can be replaced by the weighted mean of QQ, H2O’s critic objective becomes

    min⁡Q  β(E(s,a)∼ω[Q(s,a)]−E(s,a)∼D[Q(s,a)])+E~(Q,B^πQ^).\min_Q\;\beta\left(\mathbb E_{(s,a)\sim\omega}[Q(s,a)]-\mathbb E_{(s,a)\sim\mathcal D}[Q(s,a)]\right)+\widetilde{\mathcal E}(Q,\widehat{\mathcal B}^{\pi}\widehat Q).

    The approximation is supported by the bound

    Eω[Q(s,a)]≤log⁡Eω[exp⁡(Q(s,a))]≤Eω[Q(s,a)]+Var⁡ω[exp⁡(Q(s,a))]2exp⁡(2Qmin⁡),\mathbb E_{\omega}[Q(s,a)]\leq \log\mathbb E_{\omega}[\exp(Q(s,a))]\leq \mathbb E_{\omega}[Q(s,a)]+\frac{\operatorname{Var}_{\omega}[\exp(Q(s,a))]}{2\exp(2Q_{\min})},

    where the learned Q-values lie in [Qmin⁡,Qmax⁡][Q_{\min},Q_{\max}]. It is expected to be accurate when Qmin⁡Q_{\min} is sufficiently large and the variance of exp⁡(Q)\exp(Q) under ω\omega is not too large.

    For a tabular critic, the corresponding approximate dynamic-programming update is

    Q^k+1(s,a)=(B^πQ^k)(s,a)−βω(s,a)−dMπD(s,a)dMπD(s,a)+dM^π(s,a),\widehat Q^{k+1}(s,a)=\left(\widehat{\mathcal B}^{\pi}\widehat Q^k\right)(s,a)-\beta\frac{\omega(s,a)-d_M^{\pi_D}(s,a)}{d_M^{\pi_D}(s,a)+d_{\hat M}^{\pi}(s,a)},

    where dMπD(s,a)d_M^{\pi_D}(s,a) is the real state-action marginal induced by the behavior policy πD\pi_D that generated the offline dataset, and dM^π(s,a)d_{\hat M}^{\pi}(s,a) is the simulated state-action marginal induced by the learned policy π\pi. The correction acts like an adaptive reward adjustment. If ω(s,a)>dMπD(s,a)\omega(s,a)>d_M^{\pi_D}(s,a), it penalizes the Q-value, which occurs at high-gap or poorly covered state-action pairs. If ω(s,a)<dMπD(s,a)\omega(s,a)<d_M^{\pi_D}(s,a), it boosts the Q-value, favoring low-gap regions with substantial real-data support. The paper’s analysis further shows that this mechanism underestimates the value function in high-dynamics-gap regions, supporting safer policy learning under the hybrid data distribution.

  6. Knowl 6 — MuJoCo simulation evaluation under controlled dynamics gaps

    data/table

    The simulation study uses the original MuJoCo HalfCheetah environment as the real environment and creates three imperfect simulators: Gravity doubles gravitational acceleration, Friction uses 0.30.3 times the friction coefficient, and Joint Noise adds independent action noise sampled from N(0,1)\mathcal N(0,1) to every action dimension. Offline data are the corresponding D4RL Medium, Medium Replay, and Medium Expert datasets collected in the original environment. Policies are trained online in the modified simulator and evaluated by average return in the unchanged environment. Results are averaged over five random seeds.

    Could not parse LaTeX table

    H2O obtains the highest return in all Medium and Medium Replay conditions and remains competitive on Medium Expert, although DARC is better on the Medium Expert Friction condition and marginally better on the other two Medium Expert conditions. SAC trained only in the biased simulator is particularly weak in several settings, while naïvely adding real data through DARC+ does not consistently improve over DARC.

  7. Knowl 7 — Real wheel-legged robot transfer

    empirical result

    The real-world evaluation uses a two-wheeled wheel-legged robot whose state is S=(θ,θ˙,x,x˙)S=(\theta,\dot\theta,x,\dot x), with body tilt angle θ\theta, displacement xx, angular velocity θ˙\dot\theta, and linear velocity x˙\dot x; the action is the motor torque applied to the two wheels. Two tasks are tested: standing still, which requires maintaining balance without moving or falling, and moving straight, which requires tracking a target forward velocity while remaining balanced. The offline dataset contains 100,000 human-controlled transitions with task-specific rewards. Policies are trained in an Isaac Gym simulator containing unmodeled real-world effects such as motor dead zones, inaccurate sliding friction, and unmatchable wheel-ground friction and wear.

    H2O keeps the robot standing for at least 11 seconds, whereas CQL contacts the ground and loses control after 12 seconds and SAC, DARC, and DARC+ fail shortly after initialization. In the moving-straight task, H2O remains balanced and closely tracks the target velocity v=0.2 m/sv=0.2\,\mathrm{m/s}; CQL overshoots the target by nearly a factor of two, while H2O also exhibits smoother tilt-angle behavior. SAC, DARC, and DARC+ fail at the beginning. The comparison of simulator and real-robot returns also shows that high performance in a biased simulator does not guarantee real-world transfer: a policy with lower simulated return can transfer better than one that scores highly in simulation.

  8. Knowl 8 — Ablation of adaptive weighting and Bellman correction

    data/table

    The ablation study uses the HalfCheetah Medium Replay task with Gravity dynamics. The suffix −a-a removes adaptive dynamics-gap weights and uses a uniform weight, −dr-dr removes the transition-ratio correction from the Bellman error, and −reg-reg removes value regularization entirely. The average-return results are

    Could not parse LaTeX table

    Removing adaptive weighting lowers performance because high-gap simulated samples are no longer selectively regularized. Removing the dynamics-ratio correction causes a much larger degradation, showing that Bellman targets must also be corrected when real and simulated transitions are mixed. Removing value regularization produces a substantial drop, and the comparison between H2O-reg-dr and H2O-dr indicates that Bellman-error correction remains beneficial even without the full regularization term.

  9. Knowl 9 — Empirical behavior of the learned gap measure

    empirical result

    To test whether the learned gap measure identifies state-action pairs where simulation is less reliable, the authors modify HalfCheetah by adding action-noise offsets that increase with the robot’s X-velocity. Higher X-velocity also corresponds to higher task reward and therefore higher critic values, creating a setting where the algorithm must distinguish desirable but unreliable simulated samples.

    Across training snapshots at 1, 20, 200, and 500 epochs, the estimated u(s,a)u(s,a) increasingly concentrates on the higher-noise, high-X-velocity part of the data. H2O consequently applies stronger value penalties to the high-Q samples that have larger estimated dynamics gaps. This behavior supports the intended use of the discriminator-based gap estimate rather than a uniform penalty over all simulated transitions.

  10. Knowl 10 — Approximate gap quantification and remaining limitations

    limitation

    The practical H2O implementation does not evaluate the exact transition KL divergence. It approximates the expectation over next states with NN Gaussian samples centered at a simulated next state, using the covariance of states in the real offline dataset. It also approximates the full state-action log-sum-exp with simulated replay-buffer minibatches. These choices are computationally convenient and work well in the reported experiments, but the paper identifies more accurate, non-approximate dynamics-gap quantification as an open direction.

    The authors also identify the use of a conservative offline-RL/value-regularization backbone as a possible source of unnecessary conservatism. They propose investigating less conservative offline objectives and better gap estimators in future work; therefore, the reported results do not establish that H2O is optimal for every offline dataset, simulator, or choice of hybrid-RL backbone.

Coverage note — Standard MDP and offline-RL background, related work, acknowledgements, and proof derivations were omitted because they are not independent contributions; appendix-only implementation details were included only when they materially affect H2O’s practical behavior.

References

  1. 1.Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al. Learning dexterous in-hand manipulation. The International Journal of Robotics Research, 39(1):3–20, 2020. 3
  2. 2.Andrew G Barto, Richard S Sutton, and Charles W Anderson. Neuronlike adaptive elements that can solve difficult learning control problems. IEEE transactions on systems, man, and cybernetics, (5):834–846, 1983. 3
  3. 3.Dimitris Bertsimas, David B Brown, and Constantine Caramanis. Theory and applications of robust optimization. SIAM review, 53(3):464–501, 2011. 16
  4. 4.Alex Bewley, Jessica Rigley, Yuxuan Liu, Jeffrey Hawke, Richard Shen, Vinh-Dieu Lam, and Alex Kendall. Learning to drive from simulation without real world labels. In 2019 International conference on robotics and automation (ICRA), pages 4818–4824. IEEE, 2019. 2
  5. 5.Jacob Buckman, Carles Gelada, and Marc G Bellemare. The importance of pessimism in fixed-dataset policy optimization. In International Conference on Learning Representations, 2020. 3
  6. 6.Yevgen Chebotar, Ankur Handa, Viktor Makoviychuk, Miles Macklin, Jan Issac, Nathan Ratliff, and Dieter Fox. Closing the sim-to-real loop: Adapting simulation randomization with real world experience. In 2019 International Conference on Robotics and Automation (ICRA), pages 8973–8979. IEEE, 2019. 3
  7. 7.Benjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine, and Ruslan Salakhutdinov. Off-dynamics reinforcement learning: Training for transfer with domain classifiers. In International Conference on Learning Representations, 2020. 2, 3, 5, 8, 19, 24
  8. 8.Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020. 7
  9. 9.Scott Fujimoto and Shixiang Shane Gu. A minimalist approach to offline reinforcement learning. Advances in Neural Information Processing Systems, 34, 2021. 2, 3
  10. 10.Scott Fujimoto, Herke Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning, pages 1587–1596, 2018. 2
  11. 11.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pages 2052–2062. PMLR, 2019. 2, 3
  12. 12.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, pages 1861–1870, 2018. 5, 8
  13. 13.Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims. Morel: Model-based offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2020. 3
  14. 14.Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum. Offline reinforcement learning with fisher divergence critic regularization. In International Conference on Machine Learning, pages 5774–5783. PMLR, 2021. 2, 3
  15. 15.Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations, 2022. 2, 3
  16. 16.Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. Stabilizing off-policy q-learning via bootstrapping error reduction. In Advances in Neural Information Processing Systems, pages 11761–11771, 2019. 2, 3, 4
  17. 17.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2020. 2, 3, 4, 7, 8, 19, 20
  18. 18.Alex X Lee, Coline Manon Devin, Yuxiang Zhou, Thomas Lampe, Konstantinos Bousmalis, Jost Tobias Springenberg, Arunkumar Byravan, Abbas Abdolmaleki, Nimrod Gileadi, David Khosid, et al. Beyond pick-and-place: Tackling robotic stacking of diverse shapes. In 5th Annual Conference on Robot Learning, 2021. 2
  19. 19.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020. 2, 3
  20. 20.Jianxiong Li, Xianyuan Zhan, Haoran Xu, Xiangyu Zhu, Jingjing Liu, and Ya-Qin Zhang. Distance-sensitive offline reinforcement learning. arXiv preprint arXiv:2205.11027, 2022. 3
  21. 21.Jinning Li, Chen Tang, Masayoshi Tomizuka, and Wei Zhan. Dealing with the unknown: Pessimistic offline reinforcement learning. In Conference on Robot Learning, pages 1455–1464. PMLR, 2022. 3, 4, 19
  22. 22.JG Liao and Arthur Berg. Sharpening jensen’s inequality. The American Statistician, 2018. 6, 16
  23. 23.Jinxin Liu, Zhang Hongyin, and Donglin Wang. Dara: Dynamics-aware reward augmentation in offline reinforcement learning. In International Conference on Learning Representations, 2022. 2, 3
  24. 24.Lennart Ljung. System identification. In Signal analysis and prediction, pages 163–173. Springer, 1998. 3
  25. 25.Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv preprint arXiv:2108.10470, 2021. 8
  26. 26.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013. 1
  27. 27.Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel. Sim-to-real transfer of robotic control with dynamics randomization. In 2018 IEEE international conference on robotics and automation (ICRA), pages 3803–3810. IEEE, 2018. 2, 3
  28. 28.Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine. Epopt: Learning robust neural network policies using model ensembles. arXiv preprint arXiv:1610.01283, 2016. 3
  29. 29.Kanishka Rao, Chris Harris, Alex Irpan, Sergey Levine, Julian Ibarz, and Mohi Khansari. Rl-cyclegan: Reinforcement learning aware simulation-to-real. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11157–11166, 2020. 2
  30. 30.David Silver, Julian Schrittwieser, Koray Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. Mastering the game of go without human knowledge. nature, 550(7676):354–359, 2017. 1
  31. 31.Richard S Sutton, Andrew G Barto, et al. Introduction to reinforcement learning. 1998. 3
  32. 32.Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30. IEEE, 2017. 2
  33. 33.Emanuel Todorov, Tom Erez, and Yuval Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pages 5026–5033. IEEE, 2012. 7
  34. 34.Quan Vuong, Sharad Vikram, Hao Su, Sicun Gao, and Henrik I Christensen. How to pick the domain randomization parameters for sim-to-real transfer of reinforcement learning policies? arXiv preprint arXiv:1903.11774, 2019. 3
  35. 35.Guan Wang, Haoyi Niu, Desheng Zhu, Jianming Hu, Xianyuan Zhan, and Guyue Zhou. A versatile and efficient reinforcement learning approach for autonomous driving. In Machine Learning for Autonomous Driving Workshop NeurIPS 2022, 2022. 2
  36. 36.Yifan Wu, George Tucker, and Ofir Nachum. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019. 3
  37. 37.Teng Xiao and Donglin Wang. A general offline reinforcement learning framework for interactive recommendation. In The Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI, volume 2021, 2021. 2
  38. 38.Haoran Xu, Xianyuan Zhan, Jianxiong Li, and Honglei Yin. Offline reinforcement learning with soft behavior regularization. arXiv preprint arXiv:2110.07395, 2021. 3
  39. 39.Haoran Xu, Jiang Li, Jianxiong Li, and Xianyuan Zhan. A policy-guided imitation approach for offline reinforcement learning. In Advances in Neural Information Processing Systems, 2022. 2, 3
  40. 40.Haoran Xu, Xianyuan Zhan, and Xiangyu Zhu. Constraints penalized q-learning for safe offline reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022. 2, 3
  41. 41.Wenhao Yu, Jie Tan, C Karen Liu, and Greg Turk. Preparing for the unknown: Learning a universal policy with online system identification. arXiv preprint arXiv:1702.02453, 2017. 3
  42. 42.Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. Mopo: Model-based offline policy optimization. In Neural Information Processing Systems (NeurIPS), 2020. 3, 23
  43. 43.Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. Combo: Conservative offline model-based policy optimization. Advances in Neural Information Processing Systems, 34, 2021. 3, 4, 19, 23
  44. 44.Xianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu, Honglei Yin, and Yu Zheng. Deepthermal: Combustion optimization for thermal power generating units using offline reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022. 2, 3
  45. 45.Xianyuan Zhan, Xiangyu Zhu, and Haoran Xu. Model-based offline planning with trajectory pruning. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, pages 3716–3722, 2022. 3

Citation

MLA
Niu, H., et al. “When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 36599–612, https://proceedings.neurips.cc/paper_files/paper/2022/file/ed3cd2520148b577039adfade82a5566-Paper-Conference.pdf.
APA
Niu, H., sharma, . shubham ., Qiu, Y., Li, M., Zhou, G., HU, J., & Zhan, X. (2022). When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning. Advances in Neural Information Processing Systems, 35, 36599–36612. https://proceedings.neurips.cc/paper_files/paper/2022/file/ed3cd2520148b577039adfade82a5566-Paper-Conference.pdf
Chicago
Niu, H., . shubham . sharma, Y. Qiu, et al. 2022. “When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning”. Advances in Neural Information Processing Systems 35: 36599–612. https://proceedings.neurips.cc/paper_files/paper/2022/file/ed3cd2520148b577039adfade82a5566-Paper-Conference.pdf.
Harvard
Niu, H. et al. (2022) “When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 36599–36612. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/ed3cd2520148b577039adfade82a5566-Paper-Conference.pdf.
Vancouver
1. Niu H, sharma shubham, Qiu Y, Li M, Zhou G, HU J, Zhan X (2022) When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 36599–36612

BibTeX

@inproceedings{niu2022when,
  title = {When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning},
  author = {Niu, Haoyi and sharma, shubham and Qiu, Yiwen and Li, Ming and Zhou, Guyue and HU, Jianming and Zhan, Xianyuan},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {36599-36612},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/ed3cd2520148b577039adfade82a5566-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors