Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Tim SalimansJonathan HoXi ChenIlya Sutskever
Demonstrates that evolution strategies can match standard reinforcement learning methods on Atari and MuJoCo benchmarks while scaling efficiently across thousands of parallel CPU workers to achieve dramatic training speedups.
Modern artificial intelligence heavily relies on reinforcement learning to teach autonomous agents how to solve complex control and decision-making tasks. However, standard reinforcement learning algorithms rely on backpropagating gradients and estimating future reward values, which creates substantial communication overhead and makes it difficult to scale training across modern distributed computer clusters. The article addresses this operational bottleneck by evaluating whether Evolution Strategies, an older class of black-box optimization algorithms, can serve as a highly parallel, scalable alternative to standard reinforcement learning.
The article set out to demonstrate that Evolution Strategies can reliably train deep neural network policies to competitive performance levels on standard continuous robotic control and pixel-based video game benchmarks while achieving dramatic speedups through distributed computing. To evaluate this approach, the authors tested Evolution Strategies on continuous control environments in the MuJoCo physics simulator—including complex tasks like 3D humanoid walking—and 51 standard Atari 2600 games. The implementation eliminated heavy network communication by synchronizing random seeds across workers, allowing over a thousand parallel central processing units on commercial cloud hardware to communicate solely via single scalar reward values per episode.
The core finding is that Evolution Strategies scales near-linearly across large computing clusters, reducing training turnaround time from nearly a full day to mere minutes. Using 1,440 processor cores, the system solved the challenging 3D humanoid walking task in under 10 minutes, compared to roughly 11 hours on a single machine. In terms of sample efficiency, Evolution Strategies required approximately 3 to 10 times more environment interaction data than baseline algorithms on complex tasks, but this gap was offset by requiring roughly 3 times less computation per step because it avoids gradient backpropagation and value function approximation entirely. On the Atari benchmark, one hour of parallel Evolution Strategies training matched the computational cost of standard 24-hour reinforcement learning baselines, outperforming the established asynchronous actor-critic baseline on 23 games while falling short on 28. Furthermore, the approach exhibited robust exploration behavior and was completely invariant to action frequency and delayed rewards.
These findings indicate that computing turnaround time and engineering simplicity can often outweigh pure data efficiency. Because Evolution Strategies only evaluates complete episodes and perturbs policy parameters directly, it eliminates common reinforcement learning failure modes such as exploding gradients and high sensitivity to time-discounting parameters. Organizations can leverage standard cloud compute instances or low-precision inference hardware without requiring specialized high-bandwidth networking or complex gradient synchronization. Decision-makers facing long training cycles or delayed-reward problems should consider Evolution Strategies as a practical option whenever high-volume simulation data can be generated in parallel.
Next steps include applying Evolution Strategies to meta-learning and tasks with long time horizons or non-differentiable network components. Although Evolution Strategies proves highly effective when massive parallel simulation is available, its reliance on higher overall sample counts means it remains less suitable for physical systems where data collection is slow, expensive, or limited to real time.
- Paper: Asynchronous Methods for Deep Reinforcement Learning, Volodymyr Mnih et al. (2016). Its asynchronous actor-critic results establish the distributed reinforcement-learning baseline that this paper directly compares against.
- Paper: Benchmarking Deep Reinforcement Learning for Continuous Control, Yan Duan et al. (2016). Its standardized MuJoCo control suite supplies the benchmark context for assessing the source’s evolutionary approach against gradient-based methods.
- Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). Its Atari DQN work establishes the pixel-based benchmark and reinforcement-learning baseline used to interpret the source’s game results.
- Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). Its DDPG method provides a key continuous-control comparison, clarifying the sample-efficiency tradeoff discussed in the source.
- Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). Its policy-gradient and value-function framework makes clear which gradient-based training machinery the source’s evolution-strategy alternative avoids.
- Paper: A Brief Survey of Deep Reinforcement Learning, Kai Arulkumaran et al. (2017). Its survey organizes the deep-RL algorithms and Atari and MuJoCo benchmarks that frame the source’s comparisons.
- Paper: Recurrent World Models Facilitate Policy Evolution, David Ha et al. (2018). It carries evolution strategies from direct policy optimization into a learned world model, showing how policy evolution can train agents in an imagined environment.
- Paper: Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning, Xin Qiu et al. (2025). It extends the source’s gradient-free, parallel evolution-strategy approach to fine-tuning billion-parameter language models on reasoning tasks.
