Diffusion Actor-Critic with Entropy Regulator

Yinuo WangLikun WangYuxuan JiangWenjun ZouTong LiuXujie SongWenxuan WangLiming XiaoJiang WuJingliang Duan

article2024NeurIPS101 citations

Proposes an online reinforcement learning framework that leverages diffusion models to express complex multimodal policies and uses Gaussian mixture models to estimate policy entropy for adaptive exploration, achieving state-of-the-art performance on continuous control benchmarks.

Listen

Modern autonomous systems, from industrial robotics to automated driving, rely heavily on reinforcement learning to solve complex control problems. However, standard reinforcement learning algorithms generally assume actions follow simple, single-peaked Gaussian distributions. This simplification restricts their ability to learn in complex environments where multiple distinct actions could achieve high rewards, often causing algorithms to average distinct good choices into a single, ineffective intermediate decision.

The article aims to resolve this limitation by introducing Diffusion Actor-Critic with Entropy Regulator (DACER), an online reinforcement learning framework that models complex, multimodal action distributions. Specifically, the article demonstrates how generative diffusion processes can be integrated into online reinforcement learning while adaptively regulating exploration to maximize policy performance.

To achieve this, the article repurposes the reverse denoising process of a diffusion model to act directly as a flexible policy network. Because diffusion policies lack an exact mathematical formula for entropy—the metric traditionally used to measure and guide exploration—the method fits actions with a Gaussian mixture model to approximate entropy. It then uses this estimate to dynamically adjust the exploration noise added during training. The researchers evaluated DACER against six prominent reinforcement learning baselines across eight standardized continuous control benchmarks in the MuJoCo simulation physics suite, as well as on a dedicated multimodal navigation task.

The evaluation produced several decisive findings. Across standard benchmarks, DACER matched or outperformed all baseline algorithms. In the challenging Humanoid control task, DACER achieved an average return of 11,888, outperforming top-tier baselines like Distributional Soft Actor-Critic (10,829) and Soft Actor-Critic (9,335) by roughly 10% to 27%, while exceeding earlier algorithms by over 100%. In multimodal navigation tests, DACER successfully discovered all symmetric optimal paths simultaneously, whereas baseline methods failed to capture distinct concurrent modes. Ablation experiments also showed that adaptively regulating noise via estimated entropy was essential; omitting entropy regulation or using fixed noise schedules caused substantial drops in performance. Furthermore, selecting twenty reverse diffusion steps balanced gradient stability and performance, whereas thirty steps triggered training instability.

These findings indicate that generative diffusion architectures can substantially elevate the control precision and learning efficiency of online autonomous systems facing complex decision spaces. By expanding representational flexibility without relying on offline imitation data, this approach reduces the risk of systems settling for suboptimal compromise actions. For enterprise implementations, this promises higher task completion rates in high-dimensional robotics and simulation-driven control tasks.

Decision-makers and engineering teams exploring advanced reinforcement learning should consider adopting generative diffusion policies when standard algorithms plateau on complex or multimodal control tasks. Future development should focus on optimizing the computational overhead of entropy estimation to support real-time training at high frequencies. Readers should note that current validations are confined to simulated physics benchmarks across five random seeds, meaning computational efficiency, sensor noise, and real-time inference latency must be validated in physical hardware pilot deployments before mission-critical execution.

arXiv: 2405.15177
Cover for Diffusion Actor-Critic with Entropy Regulator

Abstract

Reinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution with learned mean and variance, which constrains their capability to acquire complex policies. In response to this problem, we propose an online RL algorithm termed diffusion actor-critic with entropy regulator (DACER). This algorithm conceptualizes the reverse process of the diffusion model as a novel policy function and leverages the capability of the diffusion model to fit multimodal distributions, thereby enhancing the representational capacity of the policy. Since the distribution of the diffusion policy lacks an analytical expression, its entropy cannot be determined analytically. To mitigate this, we propose a method to estimate the entropy of the diffusion policy utilizing Gaussian mixture model. Building on the estimated entropy, we can learn a parameter α that modulates the degree of exploration and exploitation. Parameter α will be employed to adaptively regulate the variance of the added noise, which is applied to the action output by the diffusion model. Experimental trials on MuJoCo benchmarks and a multimodal task demonstrate that the DACER algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting a stronger representational capacity of the diffusion policy.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Online Reinforcement Learning
  • 3.2 Diffusion Models
  • 4 Method
  • 4.1 Diffusion Policy Representation
  • 4.2 Diffusion Policy Learning
  • 4.3 Diffusion Policy with Entropy
  • 5 Experiments
  • 5.1 Comparative Evaluation
  • 5.2 Policy Representation Experiment
  • 5.3 Ablation Study
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • A Environmental Details
  • A.1 Experimental Environment Introduction
  • A.2 Training Details on MuJoCo tasks
  • B Limitation and Future Work
  • C Positive and Negative Social Impact
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Diffusion Actor-Critic with Entropy Regulator Framework

    model/method

    Diffusion Actor-Critic with Entropy Regulator (DACER) is an online model-free reinforcement learning framework that parameterizes the policy using the reverse process of a conditional Denoising Diffusion Probabilistic Model (DDPM). Unlike offline diffusion reinforcement learning methods that rely on behavior cloning and dataset imitation, DACER operates in the online setting without behavior-cloning loss terms.

    In DACER, policy improvement is driven by directly maximizing the expected state-action value (QQ-value) evaluated on actions generated by the reverse diffusion process, backpropagating action gradients through the entire diffusion chain. Because the continuous probability density function of a multi-step diffusion policy lacks a closed-form analytical expression, DACER estimates the policy's differential entropy using a Gaussian Mixture Model (GMM) fitted to sampled actions. An adaptive entropy regulator parameter α\alpha is adjusted according to the difference between a target entropy and the estimated entropy, and this parameter scales zero-mean Gaussian noise added to generated actions during environment rollouts to dynamically balance exploration and exploitation.

  2. Knowl 2 — Diffusion Policy Parameterization and Sampling Mechanism

    model/method

    The policy πθ(a∣s)\pi_\theta(a|s) in DACER is modeled as the reverse chain of a conditional diffusion model over TT discrete reverse steps starting from standard Gaussian noise aT∼N(0,I)a_T \sim \mathcal{N}(0, I):

    πθ(a∣s)=pθ(a0:T∣s)=p(aT)∏t=1Tpθ(at−1∣at,s)\pi_\theta(a|s) = p_\theta(a_{0:T}|s) = p(a_T) \prod_{t=1}^T p_\theta(a_{t-1}|a_t, s)

    where each reverse transition is parameterized as a Gaussian distribution with fixed variance Σθ(at,s,t)=βtI\Sigma_\theta(a_t, s, t) = \beta_t I and mean μθ(at,s,t)\mu_\theta(a_t, s, t) defined by a parameterized noise prediction network ϵθ(at,s,t)\epsilon_\theta(a_t, s, t):

    μθ(at,s,t)=1αt(at−βt1−αˉtϵθ(at,s,t))\mu_\theta(a_t, s, t) = \frac{1}{\sqrt{\alpha_t}} \left( a_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon_\theta(a_t, s, t) \right)

    Here, βt∈(0,1)\beta_t \in (0, 1) is a predefined variance schedule, αt=1−βt\alpha_t = 1 - \beta_t, and αˉt=∏k=1tαk\bar{\alpha}_t = \prod_{k=1}^t \alpha_k. The reverse timestep tt is encoded into a 16-dimensional representation using sinusoidal positional embeddings and concatenated with the state ss and current latent action ata_t. Sampling an action a0a_0 is executed sequentially from t=Tt = T down to 11 via the reparameterization trick:

    at−1=1αt(at−βt1−αˉtϵθ(at,s,t))+βtϵ,ϵ∼N(0,I)a_{t-1} = \frac{1}{\sqrt{\alpha_t}} \left( a_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon_\theta(a_t, s, t) \right) + \sqrt{\beta_t} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)

    The terminal sample a0a_0 of the reverse chain serves as the action output by the diffusion policy.

  3. Knowl 3 — Direct Q-Value Optimization for Online Diffusion Policies

    model/method

    In DACER, policy parameters θ\theta are updated without behavioral cloning objectives by directly maximizing the expected state-action value Qϕ(s,a0)Q_\phi(s, a_0) over state batches drawn from the replay buffer B\mathcal{B}:

    max⁡θEs∼B,a0∼πθ(⋅∣s)[Qϕ(s,a0)]\max_\theta \mathbb{E}_{s \sim \mathcal{B}, a_0 \sim \pi_\theta(\cdot|s)} [Q_\phi(s, a_0)]

    The objective gradient ∇θE[Qϕ(s,a0)]\nabla_\theta \mathbb{E}[Q_\phi(s, a_0)] requires backpropagating the action-gradient ∇aQϕ(s,a)\nabla_a Q_\phi(s, a) through all TT steps of the unrolled stochastic reverse diffusion sampling chain to update the parameters of the noise prediction network ϵθ\epsilon_\theta.

    Policy evaluation uses double QQ-learning with two critic networks Qϕ1,Qϕ2Q_{\phi_1}, Q_{\phi_2} and target networks Qϕ1′,Qϕ2′Q_{\phi'_1}, Q_{\phi'_2}. The critic parameters ϕi\phi_i (i∈{1,2}i \in \{1, 2\}) are trained by minimizing the mean squared Bellman error:

    min⁡ϕiE(s,a,r,s′)∼B[(r(s,a)+γmin⁡j=1,2Qϕj′(s′,a′)−Qϕi(s,a))2]\min_{\phi_i} \mathbb{E}_{(s, a, r, s') \sim \mathcal{B}} \left[ \left( r(s, a) + \gamma \min_{j=1,2} Q_{\phi'_j}(s', a') - Q_{\phi_i}(s, a) \right)^2 \right]

    where γ∈(0,1)\gamma \in (0, 1) is the discount factor, and the next-state action a′a' is generated by passing s′s' through the diffusion policy πθ(⋅∣s′)\pi_\theta(\cdot|s').

  4. Knowl 4 — GMM-Based Differential Entropy Estimation for Diffusion Policies

    model/method

    Because the multi-step marginal action distribution generated by a reverse diffusion chain πθ(a∣s)\pi_\theta(a|s) lacks an analytical probability density function, its differential entropy cannot be computed in closed form. DACER estimates the entropy at state ss by sampling NN actions a1,a2,…,aN∼πθ(⋅∣s)a^1, a^2, \dots, a^N \sim \pi_\theta(\cdot|s) and fitting a KK-component Gaussian Mixture Model (GMM):

    f^(a)=∑k=1KwkN(a∣μk,Σk)\hat{f}(a) = \sum_{k=1}^K w_k \mathcal{N}(a | \mu_k, \Sigma_k)

    where wk≥0w_k \ge 0, ∑k=1Kwk=1\sum_{k=1}^K w_k = 1, and μk,Σk\mu_k, \Sigma_k denote the mixing weights, mean vectors, and covariance matrices estimated via the Expectation-Maximization (EM) algorithm with posterior responsibilities:

    γ(zki)=wkN(ai∣μk,Σk)∑j=1KwjN(ai∣μj,Σj)\gamma(z_k^i) = \frac{w_k \mathcal{N}(a^i | \mu_k, \Sigma_k)}{\sum_{j=1}^K w_j \mathcal{N}(a^i | \mu_j, \Sigma_j)}

    wk=1N∑i=1Nγ(zki),μk=∑i=1Nγ(zki)ai∑i=1Nγ(zki),Σk=∑i=1Nγ(zki)(ai−μk)(ai−μk)T∑i=1Nγ(zki)w_k = \frac{1}{N} \sum_{i=1}^N \gamma(z_k^i), \quad \mu_k = \frac{\sum_{i=1}^N \gamma(z_k^i) a^i}{\sum_{i=1}^N \gamma(z_k^i)}, \quad \Sigma_k = \frac{\sum_{i=1}^N \gamma(z_k^i) (a^i - \mu_k)(a^i - \mu_k)^T}{\sum_{i=1}^N \gamma(z_k^i)}

    The state-dependent entropy HsH_s is then approximated using the closed-form upper bound for GMM differential entropy:

    Hs≈−∑k=1Kwklog⁡wk+∑k=1Kwk12log⁡((2πe)d∣Σk∣)H_s \approx -\sum_{k=1}^K w_k \log w_k + \sum_{k=1}^K w_k \frac{1}{2} \log \left( (2\pi e)^d |\Sigma_k| \right)

    where d=dim(A)d = \text{dim}(\mathcal{A}) is the action dimension. The overall estimated entropy H^\hat{H} is the sample average H^=Es∼B[Hs]\hat{H} = \mathbb{E}_{s \sim \mathcal{B}}[H_s] across a batch of states from the replay buffer.

  5. Knowl 5 — Adaptive Exploration Noise Regulation via Estimated Entropy

    model/method

    To adaptively balance exploration and exploitation without an analytical policy entropy gradient, DACER learns a scalar regulator parameter α\alpha driven by the discrepancy between the target entropy Hˉ\bar{H} (set to −0.9⋅dim(A)-0.9 \cdot \text{dim}(\mathcal{A})) and the GMM-estimated entropy H^\hat{H}:

    α←α−βα[Hˉ−H^]\alpha \leftarrow \alpha - \beta_\alpha [\bar{H} - \hat{H}]

    where βα\beta_\alpha is the learning rate for α\alpha.

    During training and environment interaction, exploration is injected by perturbing the final action aa generated by the diffusion reverse process using zero-mean Gaussian noise scaled by α\alpha and a constant hyperparameter λ\lambda:

    anoisy=a+λα⋅N(0,I)a_{\text{noisy}} = a + \lambda \alpha \cdot \mathcal{N}(0, I)

    During policy evaluation, no noise is added, and the clean output aa of the diffusion model is executed directly.

  6. Knowl 6 — Diffusion Actor-Critic with Entropy Regulator (DACER) Algorithm

    algorithm

    The complete training procedure for DACER executes environment sampling with entropy-modulated noise injection, critic updates, policy updates via backpropagation through the reverse diffusion chain, and periodic GMM-based entropy estimation to tune α\alpha.

    Input: noise scale λ\lambda, initial policy parameters θ\theta, critic parameters ϕ1,ϕ2\phi_1, \phi_2, target critic parameters ϕ1′,ϕ2′\phi'_1, \phi'_2, initial entropy parameter α\alpha, critic learning rate βq\beta_q, alpha learning rate βα\beta_\alpha, actor learning rate βπ\beta_\pi, target smoothing rate ρ\rho, target entropy Hˉ\bar{H}
    for each iteration do
        for each sampling step do
            Sample action a∼πθ(⋅∣s)a \sim \pi_\theta(\cdot|s) via reverse diffusion
            Add exploration noise a←a+λα⋅N(0,I)a \leftarrow a + \lambda \alpha \cdot \mathcal{N}(0, I)
            Execute action aa, receive reward rr and next state s′s'
            Store transition (s,a,r,s′)(s, a, r, s') in replay buffer B\mathcal{B}
        end for
        for each update step do
            Sample mini-batch of transitions from B\mathcal{B}
            Update critic networks ϕi←ϕi−βq∇ϕiLq(ϕi)\phi_i \leftarrow \phi_i - \beta_q \nabla_{\phi_i} \mathcal{L}_q(\phi_i) for i∈{1,2}i \in \{1, 2\}
            Update diffusion policy θ←θ−βπ∇θLπ(θ)\theta \leftarrow \theta - \beta_\pi \nabla_\theta \mathcal{L}_\pi(\theta)
            if step  mod 10000==0\bmod 10000 == 0 then
                Estimate policy entropy H^=Es∼B[Hs]\hat{H} = \mathbb{E}_{s \sim \mathcal{B}} [H_s] using GMM fitting
                Update entropy parameter α←α−βα[Hˉ−H^]\alpha \leftarrow \alpha - \beta_\alpha [\bar{H} - \hat{H}]
                Update target critics ϕi′←ρϕi′+(1−ρ)ϕi\phi'_i \leftarrow \rho \phi'_i + (1 - \rho) \phi_i for i∈{1,2}i \in \{1, 2\}
            end if
        end for
    end for
  7. Knowl 7 — Comparative Performance of DACER on MuJoCo Benchmarks

    data/table

    The performance of DACER was evaluated against six model-free continuous control baselines (DSAC, SAC, TD3, DDPG, TRPO, PPO) across eight MuJoCo benchmark environments. Each algorithm was trained for 1.5 million steps over five random seeds. The evaluation metric is the mean and standard deviation of the highest returns observed during the final 10% of iterations, evaluated every 15,000 steps over 10 test episodes.

    Task DACER DSAC SAC TD3 DDPG TRPO PPO
    Humanoid-v3 11888 ±\pm 244 10829 ±\pm 243 9335 ±\pm 695 5631 ±\pm 435 5291 ±\pm 662 965 ±\pm 555 6869 ±\pm 1563
    Ant-v3 9108 ±\pm 103 7086 ±\pm 261 6427 ±\pm 804 6184 ±\pm 486 4549 ±\pm 788 6203 ±\pm 578 6156 ±\pm 185
    HalfCheetah-v3 17177 ±\pm 176 17025 ±\pm 157 16573 ±\pm 224 8632 ±\pm 4041 13970 ±\pm 2083 4785 ±\pm 967 5789 ±\pm 2200
    Walker2d-v3 6701 ±\pm 62 6424 ±\pm 147 6200 ±\pm 263 5237 ±\pm 335 4095 ±\pm 68 5502 ±\pm 593 4831 ±\pm 637
    InvertedDoublePendulum-v3 9360 ±\pm 0 9360 ±\pm 0 9360 ±\pm 0 9347 ±\pm 15 9183 ±\pm 9 6259 ±\pm 2065 9356 ±\pm 2
    Hopper-v3 4104 ±\pm 49 3660 ±\pm 533 2483 ±\pm 943 3569 ±\pm 455 2644 ±\pm 659 3474 ±\pm 400 2647 ±\pm 482
    Pusher-v2 -19 ±\pm 1 -19 ±\pm 1 -20 ±\pm 0 -21 ±\pm 1 -30 ±\pm 6 -23 ±\pm 2 -23 ±\pm 1
    Swimmer-v3 152 ±\pm 7 138 ±\pm 6 140 ±\pm 14 134 ±\pm 5 146 ±\pm 4 70 ±\pm 38 130 ±\pm 2

    DACER matches or outperforms all baselines across all eight tasks. On Humanoid-v3, DACER achieves performance improvements of 9.8% over DSAC, 27.3% over SAC, 111.1% over TD3, 124.7% over DDPG, 73.1% over PPO, and 1131.9% over TRPO.

  8. Knowl 8 — Multimodal Policy and Value Function Representation in Multi-Goal Navigation

    empirical result

    In a 2D multi-goal navigation environment where a 2D point mass on a 7×77 \times 7 plane must navigate towards one of four symmetrically placed goals located at (0,5)(0, 5), (0,−5)(0, -5), (5,0)(5, 0), and (−5,0)(-5, 0):

    1. Policy Distribution and Action Vectors: DACER generates action vectors that point directly towards the nearest goal from any state, whereas TD3 and PPO produce disordered, random action directions across the state space.
    2. Multimodal Trajectory Sampling: When initialized at points requiring multimodal decisions (such as (0,0)(0, 0), (0.5,0.5)(0.5, 0.5), (0.5,−0.5)(0.5, -0.5), (−0.5,−0.5)(-0.5, -0.5), and (−0.5,0.5)(-0.5, 0.5)) across 100 sampled trajectories per point, DACER samples paths branching evenly toward all four goals, demonstrating high expressive multimodal coverage. In contrast, Gaussian policy baselines like DSAC collapse onto unimodal trajectories.
    3. Learned Value Landscape: DACER's learned QQ-function accurately forms four symmetric, distinct peaks corresponding to the four goal locations across the entire state space, which baseline methods fail to reconstruct.
  9. Knowl 9 — Ablation Analysis on Entropy Regulation and Diffusion Timesteps

    empirical result

    Ablation experiments conducted on the Walker2d-v3 task examine the influence of entropy regulation, noise modulation mechanisms, and the reverse diffusion chain length TT:

    1. Entropy Regulator Necessity: A diffusion actor-critic trained solely with QQ-learning without the entropy regulator mechanism (DAC) generates overly deterministic actions during training, causing deficient exploration and achieving substantially lower returns than DACER.
    2. Noise Modulation Strategy: Adaptively updating the noise factor via GMM-estimated entropy outperforms both a constant noise scale (fixed at 0.1) and a heuristic linear decay schedule (decaying linearly from 0.27 to 0.1 over training iterations).
    3. Number of Reverse Steps: Comparing T∈{10,20,30}T \in \{10, 20, 30\}, setting T=20T = 20 yields superior asymptotic performance and stable learning curves. Increasing TT to 30 leads to training instability and gradient explosion during backpropagation through the diffusion chain, while T=10T = 10 yields suboptimal returns.
  10. Knowl 10 — Computational Bottleneck of GMM Entropy Estimation in Online Diffusion RL

    limitation

    Estimating policy entropy using a Gaussian Mixture Model requires sampling N=200N = 200 actions per state across batches of states and iteratively running the Expectation-Maximization algorithm until convergence, which takes approximately 40 ms per estimation batch. To maintain computational tractability and prevent excessive training slowdown, entropy estimation and parameter α\alpha updates must be performed intermittently (every 10,000 environment steps) rather than at every gradient step. This coarse update interval prevents seamless step-by-step integration with standard maximum-entropy RL objectives.

Coverage note — None was omitted; all contributed aspects of the DACER algorithm, its GMM-based entropy estimation, experimental results on MuJoCo benchmarks, multimodal multi-goal navigation, ablation studies, and limitations are fully covered.

References

  1. 1.Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. An optimistic perspective on offline reinforcement learning. In International Conference on Machine Learning, pages 104–114. PMLR, 2020.
  2. 2.Anurag Ajay, Yilun Du, Abhi Gupta, Joshua Tenenbaum, Tommi Jaakkola, and Pulkit Agrawal. Is conditional generative modeling all you need for decision-making? The Eleventh International Conference on Learning Representations, 2023.
  3. 3.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
  4. 4.David Brandfonbrener, Will Whitney, Rajesh Ranganath, and Joan Bruna. Offline rl without off-policy evaluation. Advances in Neural Information Processing Systems, 34:4933–4946, 2021.
  5. 5.Yuhui Chen, Haoran Li, and Dongbin Zhao. Boosting continuous control with consistency policy. arXiv preprint arXiv:2310.06343, 2023.
  6. 6.Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137, 2023.
  7. 7.Felipe Codevilla, Eder Santana, Antonio M López, and Adrien Gaidon. Exploring the limitations of behavior cloning for autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9329–9338, 2019.
  8. 8.Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  9. 9.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  10. 10.Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun, and Bo Cheng. Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors. IEEE Transactions on Neural Networks and Learning Systems, 33(11):6584–6598, 2021.
  11. 11.Jingliang Duan, Wenxuan Wang, Liming Xiao, Jiaxin Gao, and Shengbo Eben Li. Dsac-t: Distributional soft actor-critic with three refinements. arXiv preprint arXiv:2310.05858, 2023.
  12. 12.Scott Fujimoto, Herke Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning, pages 1587–1596. PMLR, 2018.
  13. 13.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pages 2052–2062. PMLR, 2019.
  14. 14.Yang Guan, Yangang Ren, Qi Sun, Shengbo Eben Li, Haitong Ma, Jingliang Duan, Yifan Dai, and Bo Cheng. Integrated decision and control: Toward interpretable and computationally efficient driving intelligence. IEEE Transactions on Cybernetics, 53(2):859–873, 2022.
  15. 15.Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. Reinforcement learning with deep energy-based policies. In International Conference on Machine Learning, pages 1352–1361. PMLR, 2017.
  16. 16.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, pages 1861–1870. PMLR, 2018.
  17. 17.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  18. 18.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  19. 19.Marco F Huber, Tim Bailey, Hugh Durrant-Whyte, and Uwe D Hanebeck. On entropy approximation for gaussian mixture random vectors. In 2008 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems, pages 181–188. IEEE, 2008.
  20. 20.Bingyi Kang, Xiao Ma, Chao Du, Tianyu Pang, and Shuicheng Yan. Efficient diffusion policies for offline reinforcement learning. Conference and Workshop on Neural Information Processing Systems, 2023.
  21. 21.Bingyi Kang, Xiao Ma, Chao Du, Tianyu Pang, and Shuicheng Yan. Efficient diffusion policies for offline reinforcement learning. Advances in Neural Information Processing Systems, 36, 2024.
  22. 22.Elia Kaufmann, Leonard Bauersfeld, Antonio Loquercio, Matthias Müller, Vladlen Koltun, and Davide Scaramuzza. Champion-level drone racing using deep reinforcement learning. Nature, 620(7976):982–987, 2023.
  23. 23.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  24. 24.Vijay Konda and John Tsitsiklis. Actor-critic algorithms. Advances in Neural Information Processing Systems, 12, 1999.
  25. 25.S Eben Li. Reinforcement Learning for Sequential Decision and Optimal Control. Springer Verlag, Singapore, 2023.
  26. 26.Abdoulaye O Ly and Moulay Akhloufi. Learning to drive by imitation: An overview of deep behavior cloning methods. IEEE Transactions on Intelligent Vehicles, 6(2):195–209, 2020.
  27. 27.Diganta Misra. Mish: A self regularized non-monotonic activation function. arXiv preprint arXiv:1908.08681, 2019.
  28. 28.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023.
  29. 29.Baiyu Peng, Qi Sun, Shengbo Eben Li, Dongsuk Kum, Yuming Yin, Junqing Wei, and Tianyu Gu. End-to-end autonomous driving through dueling double deep q-network. Automotive Innovation, 4:328–337, 2021.
  30. 30.Michael Psenka, Alejandro Escontrela, Pieter Abbeel, and Yi Ma. Learning a diffusion model policy from rewards via q-score matching. arXiv preprint arXiv:2312.11752, 2023.
  31. 31.John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. Trust region policy optimization. In International Conference on Machine Learning, pages 1889–1897. PMLR, 2015.
  32. 32.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  33. 33.David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. Deterministic policy gradient algorithms. In International Conference on Machine Learning, pages 387–395. PMLR, 2014.
  34. 34.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  35. 35.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021.
  36. 36.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019.
  37. 37.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  38. 38.Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction. MIT press, 2018.
  39. 39.Emanuel Todorov, Tom Erez, and Yuval Tassa. Mujoco: A physics engine for model-based control. In Intelligent Robots and Systems, 2012.
  40. 40.Hado Van Hasselt, Arthur Guez, and David Silver. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI conference on artificial intelligence, volume 30, 2016.
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017.
  42. 42.Wenxuan Wang, Yuhang Zhang, Jiaxin Gao, Yuxuan Jiang, Yujie Yang, Zhilong Zheng, Wenjun Zou, Jie Li, Congsheng Zhang, Wenhan Cao, et al. Gops: A general optimal control problem solver for autonomous driving and industrial control applications. Communications in Transportation Research, 3:100096, 2023.
  43. 43.Zhendong Wang, Jonathan J Hunt, and Mingyuan Zhou. Diffusion policies as an expressive policy class for offline reinforcement learning. The Eleventh International Conference on Learning Representations, 2023.
  44. 44.Saining Xie Xinlei Chen, Zhuang Liu and Kaiming He. Deconstructing denoising diffusion models for self-supervised learning. arXiv preprint arXiv:2401.14404, 2024.
  45. 45.Long Yang, Zhixiong Huang, Fenghao Lei, Yucun Zhong, Yiming Yang, Cong Fang, Shiting Wen, Binbin Zhou, and Zhouchen Lin. Policy representation via diffusion probability model for reinforcement learning. arXiv preprint arXiv:2305.13122, 2023.
  46. 46.Ruihan Yang, Prakhar Srivastava, and Stephan Mandt. Diffusion probabilistic modeling for video generation. Entropy, 25(10):1469, 2023.

Citation

MLA
Wang, Y., et al. “Diffusion Actor-Critic with Entropy Regulator”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 54183–204, https://proceedings.neurips.cc/paper_files/paper/2024/file/6174c67b136621f3f2e4a6b1d3286f6b-Paper-Conference.pdf.
APA
Wang, Y., Wang, L., Jiang, Y., Zou, W., Liu, T., Song, X., Wang, W., Xiao, L., Wu, J., Duan, J., & Li, S. E. (2024). Diffusion Actor-Critic with Entropy Regulator. Advances in Neural Information Processing Systems, 37, 54183–54204. https://proceedings.neurips.cc/paper_files/paper/2024/file/6174c67b136621f3f2e4a6b1d3286f6b-Paper-Conference.pdf
Chicago
Wang, Y., L. Wang, Y. Jiang, et al. 2024. “Diffusion Actor-Critic with Entropy Regulator”. Advances in Neural Information Processing Systems 37: 54183–204. https://proceedings.neurips.cc/paper_files/paper/2024/file/6174c67b136621f3f2e4a6b1d3286f6b-Paper-Conference.pdf.
Harvard
Wang, Y. et al. (2024) “Diffusion Actor-Critic with Entropy Regulator”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 54183–54204. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/6174c67b136621f3f2e4a6b1d3286f6b-Paper-Conference.pdf.
Vancouver
1. Wang Y, Wang L, Jiang Y, et al (2024) Diffusion Actor-Critic with Entropy Regulator. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 54183–54204

BibTeX

@inproceedings{wang2024diffusion,
  title = {Diffusion Actor-Critic with Entropy Regulator},
  author = {Wang, Yinuo and Wang, Likun and Jiang, Yuxuan and Zou, Wenjun and Liu, Tong and Song, Xujie and Wang, Wenxuan and Xiao, Liming and Wu, Jiang and Duan, Jingliang and Li, Shengbo E.},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {54183-54204},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/6174c67b136621f3f2e4a6b1d3286f6b-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors