RORL: Robust Offline Reinforcement Learning via Conservative Smoothing
Rui YangChenjia BaiXiaoteng MaZhaoran WangChongjie ZhangLei Han
Proposes a conservative smoothing technique for offline reinforcement learning that simultaneously regularizes policy and value functions near the dataset support, achieving state-of-the-art D4RL benchmark performance while defending against adversarial observation perturbations.
Offline reinforcement learning enables automated systems to learn decision-making strategies strictly from pre-collected historical datasets, eliminating the risks and expenses of live trial-and-error. However, conventional methods heavily prioritize conservatism to avoid unfamiliar actions, resulting in fragile decision models that degrade sharply when exposed to small observation errors, sensor noise, or adversarial inputs. The article introduces and evaluates Robust Offline Reinforcement Learning (RORL), an algorithm designed to balance conservative decision-making with robustness against state and observation disturbances.
To achieve this, RORL introduces a conservative smoothing framework. The method applies smoothness constraints to both the policy and its value functions near observed data while penalizing unfamiliar state-action pairs using uncertainty quantified across an ensemble of value networks. The authors evaluated the approach across standard continuous control benchmark tasks in the D4RL benchmark suite and subjected the system to multiple adversarial observation attack scenarios across different perturbation scales.
Key findings demonstrate that RORL achieves top-tier benchmark performance, yielding an average normalized return of 85.7 across standard continuous control tasks and outperforming strong ensemble-based baselines like EDAC (82.9) and SAC-10 (50.8). Crucially, RORL accomplishes this with only 10 ensemble networks, whereas previous competitive baselines require up to 50 networks in certain complex tasks. Under adversarial observation perturbations, RORL maintains significantly higher stability and performance than baseline methods across all tested attack scales. Theoretical analysis confirms that RORL provides a provably tighter suboptimality bound than earlier pessimistic offline learning methods. Furthermore, ablation experiments show that penalizing out-of-distribution values is the single most critical component for preserving robustness under observation attacks.
These results indicate that offline decision systems can achieve operational robustness without sacrificing baseline performance or requiring prohibitively large computational architectures. By mitigating vulnerability to sensory noise and adversarial shifts, the method reduces operational risk in safety-critical deployments such as robotics and industrial automation. While RORL introduces modest computational overhead due to adversarial state generation—running at roughly 29.6 seconds per training epoch compared to 17.9 seconds for EDAC—it runs substantially faster than previous uncertainty-based methods while maintaining moderate GPU memory requirements.
Organizations evaluating offline learning pipelines for physical or security-sensitive domains should consider incorporating conservative smoothing to protect against observation drift. Future work should focus on accelerating the adversarial state generation step and exploring smoothing directly within compact latent representations rather than raw observation spaces.
- Paper: Conservative Q-Learning for Offline Reinforcement Learning, Aviral Kumar et al. (2020). It introduces Conservative Q-Learning (CQL) to lower-bound values on out-of-distribution actions, providing the core conservative value estimation foundation that RORL builds on and refines with smoothing.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). It presents the standardized D4RL offline reinforcement learning benchmarks and continuous control evaluation protocols used directly by RORL.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). It provides a comprehensive foundational review of distributional shift, value overestimation, and pessimism principles in offline reinforcement learning.
- Paper: Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction, Aviral Kumar et al. (2019). It formalizes bootstrapping error propagation on out-of-distribution actions in offline policy evaluation, motivating RORL's pessimistic regularization.
- Paper: Off-Policy Deep Reinforcement Learning without Exploration, Scott Fujimoto et al. (2018). It demonstrates how extrapolation error degrades offline off-policy algorithms and introduces batch-constrained policy learning.
- Paper: Offline Reinforcement Learning with Implicit Q-Learning, Ilya Kostrikov et al. (2021). It establishes in-sample dynamic programming via expectile regression, serving as a primary baseline and conceptual contrast to RORL's out-of-distribution smoothing.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). It introduces Soft Actor-Critic, the core maximum entropy actor-critic backbone utilized and modified by RORL's ensemble framework.
- Paper: Is Value Learning Really the Main Bottleneck in Offline RL?, Seohong Park et al. (2024). It critically examines whether value learning or policy generalization and extraction are the primary failure points in offline reinforcement learning, directly extending the inquiry into value-level vs. policy-level conservatism.
- Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). It investigates expressive diffusion-based policy representations in offline reinforcement learning to bypass traditional conservative policy parameterization bottlenecks.
- Paper: On Training in Imagination, Nadav Timor et al. (2026). It analyzes robustness guarantees and Lipschitz regularity under model errors when training policies entirely within learned imaginary rollouts.
