Deep Surrogate Assisted Generation of Environments
Varun BhattBryon TjanakaMatthew C. FontaineStefanos Nikolaidis
Proposes a sample-efficient quality diversity framework that uses deep surrogate models to predict agent behaviors, drastically reducing expensive environment simulations when generating diverse test levels for reinforcement learning and planning agents.
Deploying autonomous systems and reinforcement learning agents into real-world applications requires rigorous validation across a diverse range of environments. Traditionally, agents are evaluated on human-authored test suites, which are expensive, tedious to design, and often fail to expose the full spectrum of possible agent behaviors and failure modes. Quality diversity algorithms can automatically discover collections of challenging, diverse environments, but standard approaches require thousands of full agent simulations, making the generation process computationally prohibitive.
The article introduces and evaluates Deep Surrogate Assisted Generation of Environments (DSAGE), an algorithm designed to efficiently generate collections of high-quality, diverse environments by leveraging deep learning prediction models to minimize costly simulations.
The approach operates via an outer-loop framework alternating through three phases: exploiting a deep surrogate model to search for promising environments, selectively simulating agents in generated environments to collect ground truth performance data, and updating the surrogate model with self-supervised feedback. DSAGE incorporates two core mechanisms: predicting ancillary agent occupancy grids before estimating high-level performance metrics, and downsampling the surrogate archive to uniformly evaluate only a subset of candidate solutions. The authors validated DSAGE across two benchmark environments—a discrete Maze domain evaluated with a reinforcement learning agent and a continuous Mario platformer domain evaluated with a pathfinding agent—comparing performance against traditional quality diversity algorithms, random domain randomization, and ablated versions over five independent trials per configuration.
The findings demonstrate that DSAGE dramatically outperforms baseline methods in both discovery quality and sample efficiency. In the Maze domain, DSAGE achieved an average quality-diversity score of approximately 16,447 (40% archive coverage) compared to 10,481 (25% coverage) for standard algorithms. In the Mario domain, it scored 4,362 (30% coverage), whereas standard approaches achieved 1,840 (13% coverage) and domain randomization scored below 93. DSAGE reached baseline target performance benchmarks in roughly 34,000 evaluations in the Maze and 2,500 evaluations in Mario, whereas standard methods required 100,000 and 5,760 evaluations respectively—reducing required simulations by approximately 57% to 66%. Ablation analyses confirmed that downsampling enabled iterative self-correction across more training cycles, while ancillary occupancy prediction significantly reduced measurement error.
These results demonstrate that surrogate models make automated agent stress-testing economically feasible and practical. By rapidly revealing behavioral edge cases, catastrophic errors, and unexpected traps prior to physical deployment, the method reduces development timelines and mitigates operational safety risks for autonomous systems. The article notes that while surrogate assistance without enhancements offers limited gains, combining downsampling with behavior prediction consistently outperforms existing baselines.
Organizations developing autonomous systems should adopt surrogate-assisted diversity generation to automate safety verification, benchmark agent robustness, and uncover hidden vulnerabilities. Technical teams should integrate behavioral proxy metrics, like spatial occupancy tracking, when modeling complex agent policies. Furthermore, automated environment generation should be explored as an adaptive training curriculum to improve agent generalization.
The study's primary limitations include its reliance on two-dimensional grid domains and the omission of temporal sequence data in spatial occupancy grids, which makes specific dynamic actions, such as jumping sequences, harder to predict. While confidence in the experimental findings is high across the evaluated settings, practitioners should validate the methodology in higher-fidelity, three-dimensional simulations before relying entirely on surrogate predictions for safety-critical hardware deployments.
- Paper: Diversity is All You Need: Learning Skills without a Reward Function, Benjamin Eysenbach et al. (2018). This work establishes the foundation of learning diverse behavioral repertoires and skills in reinforcement learning without task rewards, which directly underlies the quality diversity principles used in DSAGE.
- Paper: Recurrent World Models Facilitate Policy Evolution, David Ha et al. (2018). This paper introduces the use of learned world models to simulate and evaluate agent behavior without costly physical rollouts, motivating the deep surrogate modeling approach adopted in DSAGE.
- Paper: Emergent Tool Use From Multi-Agent Autocurricula, Bowen Baker et al. (2020). This paper demonstrates how automated environment and curriculum evolution drive the discovery of diverse, complex agent behaviors, motivating the need for efficient environment generation methods.
- Paper: Evolution Strategies as a Scalable Alternative to Reinforcement Learning, Tim Salimans et al. (2017). This study introduces scalable black-box optimization methods for policy and parameter search that inform the evolutionary search mechanisms used in quality diversity algorithms.
- Paper: Robust Reinforcement Learning via Genetic Curriculum, Yeeho Song et al. (2022). This paper extends automated environment generation to robustness curricula by evolving failure-inducing scenarios for reinforcement learning agents.
- Paper: Mastering Diverse Domains through World Models, Danijar Hafner et al. (2023). This work advances general-purpose world modeling to train and evaluate reinforcement learning policies across diverse and complex domain distributions.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). This paper generalizes automated open-ended search frameworks from environment generation to the automated design and refinement of entire agentic systems.
