Efficient Diffusion Policies For Offline Reinforcement Learning

Bingyi KangXiao MaChao DuTianyu PangShuicheng Yan

article2023NeurIPS145 citationsOutstanding Paper Award

Proposes an efficient diffusion policy framework that approximates clean actions from noise to drastically cut offline reinforcement learning training time from days to hours while enabling compatibility with likelihood-based algorithms and setting new state-of-the-art benchmarks on D4RL.

Listen

Offline reinforcement learning allows autonomous systems to learn decision-making policies from pre-existing datasets without risky or costly real-world interactions. While representing policies with expressive generative diffusion models significantly enhances performance over simple Gaussian distributions, existing diffusion policies face severe practical hurdles. Specifically, previous methods require slow, computationally prohibitive multi-step sampling chains during training and remain incompatible with popular likelihood-based reinforcement learning algorithms because diffusion models lack tractable likelihoods.

The article demonstrates an efficient diffusion policy framework that resolves both computational and algorithmic bottlenecks, evaluating its speed and decision-making performance across standard benchmark suites.

To overcome these challenges, the authors introduced an action approximation technique that estimates clean actions from noisy dataset samples in a single network pass rather than running an entire reverse sampling chain during training. They paired this with a fast differential equation-based solver for accelerated evaluation and introduced an approximated objective that enables integration with maximum likelihood-based algorithms. The framework was evaluated across continuous control domains—including locomotion, maze navigation, and robotic manipulation—integrated with multiple reinforcement learning base algorithms.

The evaluation produced several key findings. First, the proposed method reduces policy training time by a factor of 25, cutting locomotion benchmark runtimes from five days down to roughly five hours. Second, when paired with implicit Q-learning, the approach establishes new state-of-the-art benchmarks, improving average normalized performance over prior feed-forward baselines across kitchen tasks (from 53.3 to 59.4), complex hand manipulation tasks (from 54.4 to 71.3), and ant maze environments (from 63.0 to 73.4). Third, the efficiency gains permit training with 1,000 diffusion steps rather than the 5 to 100 steps used previously, yielding superior action quality without incurring performance degradation.

These findings indicate that expressive generative diffusion policies can be scaled to complex industrial and robotic control problems at computational costs comparable to standard neural networks. By removing algorithm-specific constraints, organizations can integrate diffusion-based policies into existing reinforcement learning pipelines with minimal infrastructure overhead, achieving higher task success rates in complex multi-modal environments.

Teams deploying offline reinforcement learning in continuous control settings should adopt single-step action approximation frameworks and modern differential equation solvers to accelerate training cycles. Furthermore, evaluation should employ running average metrics during training rather than peak checkpoint selection to avoid overestimating deployment stability on volatile tasks.

While the empirical results on simulated benchmarks provide high confidence in the methodology's efficiency and representational benefits, the article's evaluations are confined to simulated environments with standardized data distributions. Stakeholders should exercise appropriate validation through targeted real-world pilot tests, noting potential risks if high-performance autonomous control frameworks are adapted into safety-critical or dual-use physical systems.

Cover for Efficient Diffusion Policies For Offline Reinforcement Learning

Abstract

Offline reinforcement learning (RL) aims to learn optimal policies from offline datasets, where the parameterization of policies is crucial but often overlooked. Recently, Diffsuion-QL [37] significantly boosts the performance of offline RL by representing a policy with a diffusion model, whose success relies on a parametrized Markov Chain with hundreds of steps for sampling. However, Diffusion-QL suffers from two critical limitations. 1) It is computationally inefficient to forward and backward through the whole Markov chain during training. 2) It is incompatible with maximum likelihood-based RL algorithms (e.g., policy gradient methods) as the likelihood of diffusion models is intractable. Therefore, we propose efficient diffusion policy (EDP) to overcome these two challenges. EDP approximately constructs actions from corrupted ones at training to avoid running the sampling chain. We conduct extensive experiments on the D4RL benchmark. The results show that EDP can reduce the diffusion policy training time from 5 days to 5 hours on gym-locomotion tasks. Moreover, we show that EDP is compatible with various offline RL algorithms (TD3, CRR, and IQL) and achieves new state-of-the-art on D4RL by large margins over previous methods. Our code is available at https://github.com/sail-sg/edp.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Offline Reinforcement Learning
  • 3.2 Diffusion Models
  • 4 Efficient Diffusion Policy
  • 4.1 Diffusion Policy
  • 4.2 Reinforcement-Guided Diffusion Policy Learning
  • 4.3 Generalization to Various RL algorithms
  • 4.4 Comparison to Diffusion-QL
  • 4.5 Controlled Sampling from Diffusion Policies
  • 5 Experiments
  • 5.1 Efficiency and Reproducibility
  • 5.2 Generality and Overall Results
  • 5.3 Ablation Study
  • Energy-Based Action Selection
  • 6 Conclusion
  • References
  • Broader Impact
  • A Reinforcement Learning Algorithms
  • A.1 TD3+BC
  • A.2 Critic Regularized Regression
  • A.3 Implicit Q Learning
  • B Reinforcement Guided Diffusion Policy Details
  • C Environmental Details
  • C.1 Hyper-Parameters
  • C.2 More Results
  • C.3 More Results on EAS
  • C.4 More Experiments on Controlled Sampling
  • C.5 Computational Cost Comparison between EDP and Feed-Forward Policy networks
  • C.6 The effect of action approximation
  • D Negative Societal Impacts

Knowls

  1. Knowl 1 — Action Approximation for Efficient Diffusion Policy Training

    model/method

    In standard diffusion policies, computing policy gradients requires unrolling the full multi-step reverse diffusion chain, which is computationally expensive during training. Action approximation avoids this multi-step sampling by performing one-step denoising directly on corrupted dataset actions.

    Let (s,a0)(s, a^0) denote a state-action pair drawn from an offline dataset D\mathcal{D}. In a forward diffusion process governed by a variance schedule β1,…,βK\beta_1, \dots, \beta_K with αk=1−βk\alpha_k = 1 - \beta_k and αˉk=∏j=1kαj\bar{\alpha}_k = \prod_{j=1}^k \alpha_j, the noisy action aka^k at diffusion timestep k∼U({1,…,K})k \sim \mathcal{U}(\{1, \dots, K\}) is given in closed form by: ak=αˉka0+1−αˉkϵ,ϵ∼N(0,I)a^k = \sqrt{\bar{\alpha}_k} a^0 + \sqrt{1 - \bar{\alpha}_k} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)

    Using a parameterised noise-prediction network ϵθ(ak,k;s)\epsilon_\theta(a^k, k; s), the approximated clean action a^0\hat{a}^0 is reconstructed via one-step algebraic inversion: a^0=1αˉkak−1−αˉkαˉkϵθ(ak,k;s)\hat{a}^0 = \frac{1}{\sqrt{\bar{\alpha}_k}} a^k - \frac{\sqrt{1 - \bar{\alpha}_k}}{\sqrt{\bar{\alpha}_k}} \epsilon_\theta(a^k, k; s)

    For direct policy optimization against a learned critic Qϕ(s,a)Q_\phi(s, a), the policy improvement objective is computed using this approximated action: Lπ(θ)=−Es∼D,ak,k[Qϕ(s,a^0)]\mathcal{L}_\pi(\theta) = -\mathbb{E}_{s \sim \mathcal{D}, a^k, k} \left[ Q_\phi(s, \hat{a}^0) \right]

    This approximation requires only a single forward and backward pass through the noise-prediction network per training iteration, enabling diffusion policies to train efficiently with large diffusion horizons (e.g., K=1000K = 1000).

  2. Knowl 2 — Likelihood-Based Offline Policy Optimization for Diffusion Policies

    model/method

    Likelihood-based offline reinforcement learning algorithms (such as CRR, AWR, and IQL) update policies by maximising a weighted log-likelihood objective: max⁡θE(s,a)∼D[f(Qϕ(s,a))log⁡πθ(a∣s)]\max_\theta \mathbb{E}_{(s, a) \sim \mathcal{D}} \left[ f(Q_\phi(s, a)) \log \pi_\theta(a \mid s) \right] where f(Qϕ(s,a))f(Q_\phi(s, a)) is a monotonically increasing weighting function conditioned on state ss and action aa.

    Because exact log-likelihood evaluation log⁡πθ(a∣s)\log \pi_\theta(a \mid s) is intractable in diffusion probabilistic models, two tractable objectives are formulated:

    1. Evidence Lower Bound (ELBO) Weighting: Applying the variational lower bound of DDPM yields the weighted loss: LELBO(θ)=Ek,ϵ,(s,a)∼D[βkf(Qϕ(s,a))2αk(1−αˉk−1)∥ϵ−ϵθ(ak,k;s)∥2]\mathcal{L}_{\text{ELBO}}(\theta) = \mathbb{E}_{k, \epsilon, (s, a) \sim \mathcal{D}} \left[ \frac{\beta_k f(Q_\phi(s, a))}{2 \alpha_k (1 - \bar{\alpha}_{k-1})} \left\| \epsilon - \epsilon_\theta(a^k, k; s) \right\|^2 \right] where αk=1−βk\alpha_k = 1 - \beta_k, αˉk=∏j=1kαj\bar{\alpha}_k = \prod_{j=1}^k \alpha_j, and ak=αˉka+1−αˉkϵa^k = \sqrt{\bar{\alpha}_k} a + \sqrt{1 - \bar{\alpha}_k} \epsilon with ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I).

    2. Approximated Gaussian Policy Likelihood: Approximating the policy distribution as π^θ(a∣s)≜N(a^0,I)\hat{\pi}_\theta(a \mid s) \triangleq \mathcal{N}(\hat{a}^0, I), where a^0=1αˉkak−1−αˉkαˉkϵθ(ak,k;s)\hat{a}^0 = \frac{1}{\sqrt{\bar{\alpha}_k}} a^k - \frac{\sqrt{1 - \bar{\alpha}_k}}{\sqrt{\bar{\alpha}_k}} \epsilon_\theta(a^k, k; s), converts the weighted maximum likelihood into a weighted mean squared error regression: Lπ^(θ)=Ek,ϵ,(s,a)∼D[f(Qϕ(s,a))∥a−a^0∥2]\mathcal{L}_{\hat{\pi}}(\theta) = \mathbb{E}_{k, \epsilon, (s, a) \sim \mathcal{D}} \left[ f(Q_\phi(s, a)) \left\| a - \hat{a}^0 \right\|^2 \right]

    Both variants achieve comparable empirical performance, with the approximated Gaussian policy formulation providing simpler implementation.

  3. Knowl 3 — Reinforcement-Guided Diffusion Policy Learning Algorithm

    algorithm

    Reinforcement-Guided Diffusion Policy Learning (RGDPL) integrates action approximation during policy improvement with fast ODE-based sampling (DPM-Solver) during policy evaluation to train diffusion policies efficiently.

    Input: Total iterations II, policy loss coefficient λ\lambda, learning rate η\eta, offline dataset D\mathcal{D}, noise network ϵθ(ak,k;s)\epsilon_\theta(a^k, k; s), critic network Qϕ(s,a)Q_\phi(s, a), target critic Qϕ^(s,a)Q_{\hat{\phi}}(s, a)
    Output: Policy parameters θ\theta, critic parameters ϕ\phi
    for i=1i = 1 to II do
        Sample mini-batch transitions (s,a,s′,r)∼D(s, a, s', r) \sim \mathcal{D}
        Sample next actions a′∼πθ(⋅∣s′)a' \sim \pi_\theta(\cdot \mid s') using DPM-Solver (15 model calls)
        Compute temporal difference loss:
        LTD(ϕ)=1∣B∣∑(r+γQϕ^(s′,a′)−Qϕ(s,a))2\mathcal{L}_{\text{TD}}(\phi) = \frac{1}{|\mathcal{B}|} \sum (r + \gamma Q_{\hat{\phi}}(s', a') - Q_\phi(s, a))^2
        Update critic parameters: ϕ←ϕ−η∇ϕLTD(ϕ)\phi \leftarrow \phi - \eta \nabla_\phi \mathcal{L}_{\text{TD}}(\phi)
        Sample diffusion timesteps k∼U({1,…,K})k \sim \mathcal{U}(\{1, \dots, K\}) and noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I)
        Compute perturbed actions: ak=αˉka+1−αˉkϵa^k = \sqrt{\bar{\alpha}_k} a + \sqrt{1 - \bar{\alpha}_k} \epsilon
        Compute approximated clean actions: a^0=1αˉkak−1−αˉkαˉkϵθ(ak,k;s)\hat{a}^0 = \frac{1}{\sqrt{\bar{\alpha}_k}} a^k - \frac{\sqrt{1 - \bar{\alpha}_k}}{\sqrt{\bar{\alpha}_k}} \epsilon_\theta(a^k, k; s)
        Compute diffusion noise loss: Ldiff(θ)=∥ϵ−ϵθ(ak,k;s)∥2\mathcal{L}_{\text{diff}}(\theta) = \|\epsilon - \epsilon_\theta(a^k, k; s)\|^2
        Compute policy improvement loss Lπ(θ)\mathcal{L}_\pi(\theta) via −Qϕ(s,a^0)-Q_\phi(s, \hat{a}^0) or weighted regression
        Update policy parameters: θ←θ−η∇θ(Ldiff(θ)+λLπ(θ))\theta \leftarrow \theta - \eta \nabla_\theta (\mathcal{L}_{\text{diff}}(\theta) + \lambda \mathcal{L}_\pi(\theta))
    end for

    The algorithm uses batch size 256, Adam optimizer with gradient norm clipping threshold of 5, and K=1000K = 1000 diffusion training steps.

  4. Knowl 4 — Energy-Based Action Selection for Evaluation

    algorithm

    Because diffusion policies generate actions via stochastic sampling from an implicit distribution without closed-form statistics (such as the distribution mean), raw evaluation sampling exhibits high variance. Energy-Based Action Selection (EAS) reduces action variance at test time by filtering stochastic candidate actions through the learned critic Qϕ(s,a)Q_\phi(s, a).

    Input: Candidate pool size NN, state ss, diffusion policy πθ(a∣s)\pi_\theta(a \mid s), critic network Qϕ(s,a)Q_\phi(s, a)
    Output: Executed action aa
    for i=1i = 1 to NN do
        Sample candidate action ai∼πθ(⋅∣s)a_i \sim \pi_\theta(\cdot \mid s) using DPM-Solver
        Evaluate candidate critic value: qi=Qϕ(s,ai)q_i = Q_\phi(s, a_i)
    end for
    Compute selection probabilities: pi=exp⁡(qi)∑j=1Nexp⁡(qj)p_i = \frac{\exp(q_i)}{\sum_{j=1}^N \exp(q_j)} for each i∈{1,…,N}i \in \{1, \dots, N\}
    Sample index m∼Categorical(p1,…,pN)m \sim \text{Categorical}(p_1, \dots, p_N)
    Set executed action a=ama = a_m
    return aa

    This procedure samples actions from the energy-reweighted policy distribution p(a∣s)∝exp⁡(Qϕ(s,a))πθ(a∣s)p(a \mid s) \propto \exp(Q_\phi(s, a)) \pi_\theta(a \mid s). Setting N=10N = 10 provides an optimal trade-off between execution speed and task return.

  5. Knowl 5 — Running Average at Training Metric for Offline RL Evaluation

    definition

    Online Model Selection (OMS) evaluates an offline RL algorithm by taking the maximum evaluation score over all checkpoints saved during training. Because offline RL training trajectories can be volatile or suffer catastrophic late-stage performance drops, OMS yields an overly optimistic estimate of policy quality.

    The Running Average at Training (RAT) metric evaluates stability and final convergence by computing the arithmetic mean of evaluation performance across the final 10 consecutive training checkpoints: RAT=110∑m=M−9MRm\text{RAT} = \frac{1}{10} \sum_{m = M - 9}^{M} R_m where RmR_m is the evaluation return evaluated at checkpoint mm, and MM is the total number of checkpoints across training.

  6. Knowl 6 — Integration of EDP with TD3+BC, CRR, and IQL

    model/method

    Efficient Diffusion Policy (EDP) interfaces with different offline RL algorithms through specific critic updates and policy advantage weightings:

    • EDP + TD3+BC: Uses clipped double Q-learning with target Q^(s,a)=r(s,a)+γmin⁡(Q1(s′,a′),Q2(s′,a′))\hat{Q}(s, a) = r(s, a) + \gamma \min(Q_1(s', a'), Q_2(s', a')), where a′∼πθ(⋅∣s′)a' \sim \pi_\theta(\cdot \mid s') is sampled using DPM-Solver. Policy improvement minimises Ldiff(θ)−λE[Q(s,a^0)]\mathcal{L}_{\text{diff}}(\theta) - \lambda \mathbb{E}[Q(s, \hat{a}^0)], where QQ is randomly selected from {Q1,Q2}\{Q_1, Q_2\} with equal probability, and a^0\hat{a}^0 is generated via action approximation.
    • EDP + CRR: Evaluates policy improvement weights via an exponential advantage function fCRR(s,a)=exp⁡(1βA(s,a))f_{\text{CRR}}(s, a) = \exp\left(\frac{1}{\beta} A(s, a)\right), with advantage defined as: A(s,a)=min⁡(Q1(s,a),Q2(s,a))−1N∑i=1Nmin⁡(Q1(s,a^i),Q2(s,a^i))A(s, a) = \min(Q_1(s, a), Q_2(s, a)) - \frac{1}{N} \sum_{i=1}^N \min(Q_1(s, \hat{a}_i), Q_2(s, \hat{a}_i)) where a^i∼N(a^0,I)\hat{a}_i \sim \mathcal{N}(\hat{a}^0, I), β=1\beta = 1, and N=10N = 10.
    • EDP + IQL: Replaces out-of-distribution advantage sampling by learning an expectile value network Vψ(s)V_\psi(s) using asymmetric loss L2τ(u)=∣τ−1(u<0)∣u2L_2^\tau(u) = |\tau - \mathbf{1}(u < 0)| u^2 on u=min⁡(Q1(s,a),Q2(s,a))−Vψ(s)u = \min(Q_1(s, a), Q_2(s, a)) - V_\psi(s). The weighting function is: fIQL(s,a)=exp⁡(min⁡(Q1(s,a),Q2(s,a))−Vψ(s)τIQL)f_{\text{IQL}}(s, a) = \exp\left(\frac{\min(Q_1(s, a), Q_2(s, a)) - V_\psi(s)}{\tau_{\text{IQL}}}\right) Hyperparameters use τ=0.9\tau = 0.9 for AntMaze tasks, τ=0.7\tau = 0.7 for other tasks, and temperature τIQL=1.0\tau_{\text{IQL}} = 1.0.
  7. Knowl 7 — Benchmark Normalized Scores of EDP across D4RL Domains

    data/table

    The following table compares the average normalized scores on the D4RL benchmark under the Running Average at Training (RAT) metric across 5 random seeds. FF denotes standard feed-forward MLP policies (Gaussian parameterization), while EDP denotes the Efficient Diffusion Policy.

    Domain / Task BC 10%BC CQL TD3+BC CRR IQL
    FF EDP FF EDP FF EDP
    halfcheetah-medium-v2 42.6 42.5 44.0 48.3 52.1 44.0 49.2 47.4 48.1
    hopper-medium-v2 52.9 56.9 58.5 59.3 81.9 58.5 78.7 66.3 63.1
    walker2d-medium-v2 75.3 75.0 72.5 83.7 86.9 72.5 82.5 78.3 85.4
    halfcheetah-medium-replay-v2 36.6 40.6 45.5 44.6 49.4 45.5 43.5 44.2 43.8
    hopper-medium-replay-v2 18.1 75.9 95.0 60.9 101.0 95.0 99.0 94.7 99.1
    walker2d-medium-replay-v2 26.0 62.5 77.2 81.8 94.9 77.2 63.3 73.9 84.0
    halfcheetah-medium-expert-v2 55.2 92.9 91.6 90.7 95.5 91.6 85.6 86.7 86.7
    hopper-medium-expert-v2 52.5 110.9 105.4 98.0 97.4 105.4 92.9 91.5 99.6
    walker2d-medium-expert-v2 107.5 109.0 108.8 110.1 110.2 108.8 110.1 109.6 109.0
    Locomotion Average 51.9 74.0 77.6 75.3 85.5 77.6 78.3 77.0 79.9
    kitchen-complete-v0 65.0 7.2 43.8 2.2 61.5 43.8 73.9 62.5 75.5
    kitchen-partial-v0 38.0 66.8 49.8 0.7 52.8 49.8 40.0 46.3 46.3
    kitchen-mixed-v0 51.5 50.9 51.0 0.0 60.8 51.0 46.1 51.0 56.5
    Kitchen Average 51.5 41.6 48.2 1.0 58.4 48.2 53.3 53.3 59.4
    pen-human-v0 63.9 -2.0 37.5 5.9 48.2 37.5 70.2 71.5 72.7
    pen-cloned-v0 37.0 0.0 39.2 17.2 15.9 39.2 54.0 37.3 70.0
    Adroit Average 50.5 -1.0 38.4 11.6 32.1 38.4 62.1 54.4 71.3
    antmaze-umaze-v0 54.6 62.8 74.0 40.2 96.6 0.0 95.9 87.5 94.2
    antmaze-umaze-diverse-v0 45.6 50.2 84.0 58.0 69.5 41.9 15.9 62.2 79.0
    antmaze-medium-play-v0 0.0 5.4 61.2 0.2 0.0 0.0 33.5 71.2 81.8
    antmaze-medium-diverse-v0 0.0 9.8 53.7 0.0 6.4 0.0 32.7 70.0 82.3
    antmaze-large-play-v0 0.0 0.0 15.8 0.0 1.6 0.0 26.0 39.6 42.3
    antmaze-large-diverse-v0 0.0 6.0 14.9 0.0 4.4 0.0 58.5 47.5 60.6
    AntMaze Average 16.7 22.4 50.6 16.4 29.8 7.0 43.8 63.0 73.4

    EDP + TD3 achieves the highest average score (85.5) on Gym-locomotion, while EDP + IQL achieves the best performance across Kitchen (59.4), Adroit (71.3), and AntMaze (73.4).

  8. Knowl 8 — Training and Inference Speedups of EDP

    empirical result

    On the walker2d-medium-expert-v2 benchmark (evaluated over 10,000 policy update iterations and 10,000 environment interaction steps):

    • Diffusion-QL (PyTorch implementation): 4.66 training iterations per second (IPS), 18.67 sampling steps per second (SPS).
    • Diffusion-QL (JAX re-implementation): 22.30 training IPS, 123.70 sampling SPS.
    • EDP without DPM-Solver (Action Approximation + DDPM): 50.94 training IPS, 123.06 sampling SPS (a 2.3×2.3\times training speedup over the JAX baseline attributable to action approximation).
    • EDP without Action Approximation (DPM-Solver only): 38.40 training IPS, 411.00 sampling SPS (a 3.3×3.3\times sampling speedup over the JAX baseline attributable to DPM-Solver).
    • Full EDP (Action Approximation + DPM-Solver): 116.21 training IPS, 411.79 sampling SPS (a 5.2×5.2\times speedup over the JAX baseline and an overall ∼25×\sim 25\times speedup over Diffusion-QL PyTorch).

    On D4RL gym-locomotion benchmarks, total training time decreases from 5 days (Diffusion-QL) to 5 hours (EDP).

  9. Knowl 9 — Ablation on DPM-Solver Steps and EAS Action Pool Size

    empirical result

    Ablation experiments on D4RL gym-locomotion benchmarks examine the sensitivity of performance to DPM-Solver call steps and Energy-Based Action Selection (EAS) sample size:

    • DPM-Solver Steps: Varying the number of function evaluations in DPM-Solver from 3 to 30 shows that policy performance increases steadily up to 15 model calls, after which performance plateaus. 15 model calls is sufficient to match full DDPM sampling accuracy while minimizing latency.
    • EAS Sample Pool Size (NN): Varying NN from 1 to 200 shows that normalized returns increase monotonically with larger sample sizes on 8 out of 9 gym-locomotion environments. Setting N=10N = 10 captures the majority of the performance gain while keeping evaluation compute low.
    • EAS on Feed-Forward Policies: Applying EAS to standard TD3+BC with feed-forward policies results in no performance gain (69.7 normalized return without EAS vs. 69.2 with EAS across gym-locomotion-v2), confirming that EAS specifically addresses the stochasticity of diffusion-based policies.
  10. Knowl 10 — Comparison of Controlled Sampling Strategies for Diffusion Policies

    empirical result

    Two alternative controlled sampling strategies to reduce variance in diffusion policies were evaluated against Energy-Based Action Selection (EAS):

    • Policy Scaling: Scaling the score model outputs by temperature τ>1\tau > 1 models a sharpened policy πθτ(a∣s)∝(πθ(a∣s))τ\pi_\theta^\tau(a \mid s) \propto (\pi_\theta(a \mid s))^\tau. Varying τ∈[0.5,2.0]\tau \in [0.5, 2.0] on gym-locomotion tasks shows that peak performance occurs strictly at τ=1.0\tau = 1.0, indicating that artificial score scaling degrades action selection.
    • Deterministic Sampling (Initial Noise Scaling): Setting the initial Gaussian noise aK∼N(0,c2I)a^K \sim \mathcal{N}(0, c^2 I) with scale factor c∈[0.0,1.0]c \in [0.0, 1.0] demonstrates that zero initial noise (c=0.0c = 0.0) produces the highest score among noise scaling factors, and standard noise (c=1.0c = 1.0) performs the worst. However, deterministic sampling with c=0.0c = 0.0 remains inferior to EAS.

Coverage note — All primary contributed methods (Action Approximation, likelihood-based ELBO and Gaussian approximations, EAS algorithm, RGDPL algorithm), experimental frameworks (RAT metric), empirical benchmark tables, efficiency comparisons, and controlled sampling ablations are covered. Basic background derivations of standard DDPM and offline RL preliminaries were omitted.

References

  1. 1.David Brandfonbrener, Will Whitney, Rajesh Ranganath, and Joan Bruna. Offline rl without off-policy evaluation. Advances in Neural Information Processing Systems, 34:4933–4946, 2021.
  2. 2.David Brandfonbrener, Will Whitney, Rajesh Ranganath, and Joan Bruna. Offline rl without off-policy evaluation. Advances in Neural Information Processing Systems, 34:4933–4946, 2021.
  3. 3.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34:15084–15097, 2021.
  4. 4.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34:15084–15097, 2021.
  5. 5.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  6. 6.Scott Fujimoto and Shixiang Shane Gu. A minimalist approach to offline reinforcement learning. Advances in neural information processing systems, 34:20132–20145, 2021.
  7. 7.Scott Fujimoto, Herke Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In International conference on machine learning, pages 1587–1596. PMLR, 2018.
  8. 8.Scott Fujimoto, Herke Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In International conference on machine learning, pages 1587–1596. PMLR, 2018.
  9. 9.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International conference on machine learning, 2019.
  10. 10.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  11. 11.Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34:1273–1286, 2021.
  12. 12.Bingyi Kang, Xiao Ma, Yirui Wang, Yang Yue, and Shuicheng Yan. Improving and benchmarking offline reinforcement learning algorithms. arXiv preprint arXiv:2306.00972, 2023.
  13. 13.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  14. 14.Roger Koenker and Kevin F Hallock. Quantile regression. Journal of economic perspectives, 15(4):143–156, 2001.
  15. 15.Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. arXiv preprint arXiv:2110.06169, 2021.
  16. 16.Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. Stabilizing off-policy q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems, 32, 2019.
  17. 17.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. Advances in Neural Information Processing Systems, 33:1179–1191, 2020.
  18. 18.Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  19. 19.Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  20. 20.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022.
  21. 21.Xiao Ma, Bingyi Kang, Zhongwen Xu, Min Lin, and Shuicheng Yan. Mutual information regularized offline reinforcement learning. arXiv preprint arXiv:2210.07484, 2022.
  22. 22.Michael Mathieu, Sherjil Ozair, Srivatsan Srinivasan, Caglar Gulcehre, Shangtong Zhang, Ray Jiang, Tom Le Paine, Konrad Zolna, Richard Powell, Julian Schrittwieser, et al. Starcraft ii unplugged: Large scale offline reinforcement learning. In Deep RL Workshop NeurIPS 2021, 2021.
  23. 23.Diganta Misra. Mish: A self regularized non-monotonic neural activation function. arXiv preprint arXiv:1908.08681, 4(2):10–48550, 2019.
  24. 24.Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans. Bridging the gap between value and policy based reinforcement learning. Advances in neural information processing systems, 2017.
  25. 25.Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine. Awac: Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020.
  26. 26.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021.
  27. 27.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc., 2019.
  28. 28.Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
  29. 29.Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
  30. 30.John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. Trust region policy optimization. In International conference on machine learning, pages 1889–1897. PMLR, 2015.
  31. 31.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  32. 32.Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, and Lerrel Pinto. Behavior transformers: Cloning k modes with one stone. arXiv preprint arXiv:2206.11251, 2022.
  33. 33.Noah Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller. Keep doing what worked: Behavior modelling priors for offline reinforcement learning. In International Conference on Learning Representations, 2020.
  34. 34.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  35. 35.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  37. 37.Zhendong Wang, Jonathan J Hunt, and Mingyuan Zhou. Diffusion policies as an expressive policy class for offline reinforcement learning. arXiv preprint arXiv:2208.06193, 2022.
  38. 38.Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Springenberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, et al. Critic regularized regression. Advances in Neural Information Processing Systems, 33:7768–7778, 2020.
  39. 39.Yifan Wu, George Tucker, and Ofir Nachum. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  40. 40.Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. Combo: Conservative offline model-based policy optimization. Advances in neural information processing systems, 34:28954–28967, 2021.
  41. 41.Yang Yue, Bingyi Kang, Xiao Ma, Gao Huang, Shiji Song, and Shuicheng Yan. Offline prioritized experience replay. arXiv preprint arXiv:2306.05412, 2023.
  42. 42.Yang Yue, Bingyi Kang, Xiao Ma, Zhongwen Xu, Gao Huang, and Shuicheng Yan. Boosting offline reinforcement learning via data rebalancing. arXiv preprint arXiv:2210.09241, 2022.

Citation

MLA
Kang, B., et al. “Efficient Diffusion Policies For Offline Reinforcement Learning”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 67195–212, https://proceedings.neurips.cc/paper_files/paper/2023/file/d45e0bfb5a39477d56b55c0824200008-Paper-Conference.pdf.
APA
Kang, B., Ma, X., Du, C., Pang, T., & Yan, S. (2023). Efficient Diffusion Policies For Offline Reinforcement Learning. Advances in Neural Information Processing Systems, 36, 67195–67212. https://proceedings.neurips.cc/paper_files/paper/2023/file/d45e0bfb5a39477d56b55c0824200008-Paper-Conference.pdf
Chicago
Kang, B., X. Ma, C. Du, T. Pang, and S. Yan. 2023. “Efficient Diffusion Policies For Offline Reinforcement Learning”. Advances in Neural Information Processing Systems 36: 67195–212. https://proceedings.neurips.cc/paper_files/paper/2023/file/d45e0bfb5a39477d56b55c0824200008-Paper-Conference.pdf.
Harvard
Kang, B. et al. (2023) “Efficient Diffusion Policies For Offline Reinforcement Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 67195–67212. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/d45e0bfb5a39477d56b55c0824200008-Paper-Conference.pdf.
Vancouver
1. Kang B, Ma X, Du C, Pang T, Yan S (2023) Efficient Diffusion Policies For Offline Reinforcement Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 67195–67212

BibTeX

@inproceedings{kang2023efficient,
  title = {Efficient Diffusion Policies For Offline Reinforcement Learning},
  author = {Kang, Bingyi and Ma, Xiao and Du, Chao and Pang, Tianyu and Yan, Shuicheng},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {67195-67212},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/d45e0bfb5a39477d56b55c0824200008-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors