Do Differentiable Simulators Give Better Policy Gradients?
Hyung Ju Terry SuhMax SimchowitzKaiqing ZhangRuss Tedrake
Explains why physical discontinuities and stiffness introduce severe bias and variance into differentiable simulator gradients, while developing a hybrid alpha-order estimator that effectively blends zeroth- and first-order techniques for reliable policy optimization.
Modern robotics and continuous control increasingly rely on simulation to train policies via reinforcement learning. While differentiable simulators promise substantially faster computation by providing exact first-order derivatives, it has remained unclear whether these exact gradients actually improve optimization in complex physical tasks involving contact, friction, and stiff dynamics.
The article systematically evaluates the trade-offs between first-order gradient estimators (which leverage exact simulator derivatives) and zeroth-order estimators (which rely on function evaluations under randomized smoothing). To solve the pitfalls of both extremes, the article proposes and demonstrates an adaptive interpolation method called the alpha-order gradient estimator.
The evaluation combines mathematical analysis of bias and variance with simulated robotic benchmarks, including object pushing under varying spring stiffness, sliding trajectories governed by Coulomb friction, double-pendulum chaotic dynamics, and a paddle tennis policy optimization task. These experiments compare optimization stability, trajectory quality, and convergence speed across the different gradient estimators.
The investigation reveals four primary findings. First, first-order estimators become severely biased in the presence of physical discontinuities—such as collisions or geometric boundaries—causing optimization to stall in flat regions or fail entirely. Second, continuous approximations of contact do not eliminate this problem; under finite sample sizes, stiff continuous relaxations exhibit "empirical bias," appearing to have near-zero variance while producing entirely incorrect gradient directions. Third, in stiff contact regimes or chaotic dynamics, the variance of first-order estimators escalates dramatically over long horizons, whereas zeroth-order estimators remain unbiased and often exhibit lower variance. Fourth, the proposed alpha-order estimator reliably identifies when exact gradients are misleading and switches toward zeroth-order estimation, matching or exceeding the convergence performance of both individual estimators across all evaluated tasks.
These findings demonstrate that blindly relying on first-order gradients from differentiable physics engines can introduce silent optimization failures and poor policy performance in contact-rich settings. Because empirical variance alone cannot detect empirical bias, standard variance-reduction heuristics fail near physical boundaries. In contrast, integrating zeroth-order robustness constraints allows practitioners to harvest the speed benefits of differentiable engines in smooth regions without compromising stability near impact surfaces.
Engineering and research teams should avoid using unconstrained first-order policy gradients for contact-heavy control tasks. Instead, teams should implement hybrid interpolation strategies, such as the alpha-order estimator, and design physics engines with contact formulations that avoid excessive numerical stiffness. Further work is recommended to evaluate differentiable engines based on optimization-based implicit time-stepping and to test these methods on complex hardware deployments.
The primary analytical limitations involve the use of heuristic parameters for the statistical confidence bounds rather than fully rigorous concentration bounds, as well as the restriction of empirical validations to simulated benchmarks rather than physical robotic systems. Nonetheless, the theoretical proofs and consistent empirical evidence provide high confidence that unadjusted first-order gradients are fundamentally vulnerable in discontinuous physical domains.
- Paper: Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming, Saeed Ghadimi et al. (2013). Its convergence analysis of first- and zeroth-order stochastic estimators provides the bias, variance, and robustness framework this paper uses to compare gradient choices.
- Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). The policy gradient theorem establishes the reinforcement-learning gradient objective that this paper examines estimating through differentiable and stochastic simulators.
- Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). Its deterministic policy-gradient derivation grounds the continuous-control setting in which the paper evaluates simulator-based policy gradients.
- Paper: PILCO: A Model-Based and Data-Efficient Approach to Policy Search, Marc Peter Deisenroth et al. (2011). PILCO's analytic policy gradients through learned dynamics offer a concrete model-based control precedent for the paper's comparison with zeroth-order estimates.
No sufficiently relevant recommendations were found.
