Diffusion Actor-Critic with Entropy Regulator
Yinuo WangLikun WangYuxuan JiangWenjun ZouTong LiuXujie SongWenxuan WangLiming XiaoJiang WuJingliang Duan
Proposes an online reinforcement learning framework that leverages diffusion models to express complex multimodal policies and uses Gaussian mixture models to estimate policy entropy for adaptive exploration, achieving state-of-the-art performance on continuous control benchmarks.
Modern autonomous systems, from industrial robotics to automated driving, rely heavily on reinforcement learning to solve complex control problems. However, standard reinforcement learning algorithms generally assume actions follow simple, single-peaked Gaussian distributions. This simplification restricts their ability to learn in complex environments where multiple distinct actions could achieve high rewards, often causing algorithms to average distinct good choices into a single, ineffective intermediate decision.
The article aims to resolve this limitation by introducing Diffusion Actor-Critic with Entropy Regulator (DACER), an online reinforcement learning framework that models complex, multimodal action distributions. Specifically, the article demonstrates how generative diffusion processes can be integrated into online reinforcement learning while adaptively regulating exploration to maximize policy performance.
To achieve this, the article repurposes the reverse denoising process of a diffusion model to act directly as a flexible policy network. Because diffusion policies lack an exact mathematical formula for entropy—the metric traditionally used to measure and guide exploration—the method fits actions with a Gaussian mixture model to approximate entropy. It then uses this estimate to dynamically adjust the exploration noise added during training. The researchers evaluated DACER against six prominent reinforcement learning baselines across eight standardized continuous control benchmarks in the MuJoCo simulation physics suite, as well as on a dedicated multimodal navigation task.
The evaluation produced several decisive findings. Across standard benchmarks, DACER matched or outperformed all baseline algorithms. In the challenging Humanoid control task, DACER achieved an average return of 11,888, outperforming top-tier baselines like Distributional Soft Actor-Critic (10,829) and Soft Actor-Critic (9,335) by roughly 10% to 27%, while exceeding earlier algorithms by over 100%. In multimodal navigation tests, DACER successfully discovered all symmetric optimal paths simultaneously, whereas baseline methods failed to capture distinct concurrent modes. Ablation experiments also showed that adaptively regulating noise via estimated entropy was essential; omitting entropy regulation or using fixed noise schedules caused substantial drops in performance. Furthermore, selecting twenty reverse diffusion steps balanced gradient stability and performance, whereas thirty steps triggered training instability.
These findings indicate that generative diffusion architectures can substantially elevate the control precision and learning efficiency of online autonomous systems facing complex decision spaces. By expanding representational flexibility without relying on offline imitation data, this approach reduces the risk of systems settling for suboptimal compromise actions. For enterprise implementations, this promises higher task completion rates in high-dimensional robotics and simulation-driven control tasks.
Decision-makers and engineering teams exploring advanced reinforcement learning should consider adopting generative diffusion policies when standard algorithms plateau on complex or multimodal control tasks. Future development should focus on optimizing the computational overhead of entropy estimation to support real-time training at high frequencies. Readers should note that current validations are confined to simulated physics benchmarks across five random seeds, meaning computational efficiency, sensor noise, and real-time inference latency must be validated in physical hardware pilot deployments before mission-critical execution.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). It establishes the maximum entropy actor-critic framework that DACER builds upon and extends using diffusion-based action policies and entropy regulation.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). It introduces action generation formulated as conditional denoising diffusion processes to capture expressive multimodal action distributions.
- Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). It analyzes the computational challenges and approximation methods required for integrating diffusion-based policies into continuous reinforcement learning architectures.
- Paper: Reinforcement Learning with Deep Energy-Based Policies, Tuomas Haarnoja et al. (2017). It provides foundational principles for learning expressive, multimodal stochastic policies in continuous action spaces using maximum entropy objectives.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). It details how generative diffusion models can be repurposed for sequential decision-making and policy synthesis.
- Paper: Diffusion Models: A Comprehensive Survey of Methods and Applications, Ling Yang et al. (2022). It provides a comprehensive mathematical survey of diffusion formulations, reverse denoising trajectories, and sampling techniques underlying modern diffusion policies.
- Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). It establishes core continuous-action actor-critic stabilization mechanisms, including twin critics and target policy smoothing, utilized in deep reinforcement learning baselines.
- Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). It derives the fundamental continuous-control actor-critic gradient theorems that underpin modern policy optimization algorithms.
- Paper: DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving, Bencheng Liao et al. (2025). It extends diffusion-based policy representations to real-time continuous control in end-to-end autonomous driving using truncated denoising schedules.
- Paper: Flow Q-Learning, Seohong Park et al. (2025). It advances expressive generative modeling in continuous control by combining flow matching with single-step policy extraction for reinforcement learning.
- Paper: CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control, Guy Tevet et al. (2025). It pairs real-time generative diffusion planning in a continuous closed feedback loop with reinforcement learning controllers for physics-based character animation.
- Paper: Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control, Carles Domingo-Enrich et al. (2025). It builds on reward-guided diffusion and flow optimization by formulating generative policy fine-tuning via stochastic optimal control.
