Deep Surrogate Assisted Generation of Environments

Varun BhattBryon TjanakaMatthew C. FontaineStefanos Nikolaidis

article2022NeurIPS49 citations

Proposes a sample-efficient quality diversity framework that uses deep surrogate models to predict agent behaviors, drastically reducing expensive environment simulations when generating diverse test levels for reinforcement learning and planning agents.

Listen

Deploying autonomous systems and reinforcement learning agents into real-world applications requires rigorous validation across a diverse range of environments. Traditionally, agents are evaluated on human-authored test suites, which are expensive, tedious to design, and often fail to expose the full spectrum of possible agent behaviors and failure modes. Quality diversity algorithms can automatically discover collections of challenging, diverse environments, but standard approaches require thousands of full agent simulations, making the generation process computationally prohibitive.

The article introduces and evaluates Deep Surrogate Assisted Generation of Environments (DSAGE), an algorithm designed to efficiently generate collections of high-quality, diverse environments by leveraging deep learning prediction models to minimize costly simulations.

The approach operates via an outer-loop framework alternating through three phases: exploiting a deep surrogate model to search for promising environments, selectively simulating agents in generated environments to collect ground truth performance data, and updating the surrogate model with self-supervised feedback. DSAGE incorporates two core mechanisms: predicting ancillary agent occupancy grids before estimating high-level performance metrics, and downsampling the surrogate archive to uniformly evaluate only a subset of candidate solutions. The authors validated DSAGE across two benchmark environments—a discrete Maze domain evaluated with a reinforcement learning agent and a continuous Mario platformer domain evaluated with a pathfinding agent—comparing performance against traditional quality diversity algorithms, random domain randomization, and ablated versions over five independent trials per configuration.

The findings demonstrate that DSAGE dramatically outperforms baseline methods in both discovery quality and sample efficiency. In the Maze domain, DSAGE achieved an average quality-diversity score of approximately 16,447 (40% archive coverage) compared to 10,481 (25% coverage) for standard algorithms. In the Mario domain, it scored 4,362 (30% coverage), whereas standard approaches achieved 1,840 (13% coverage) and domain randomization scored below 93. DSAGE reached baseline target performance benchmarks in roughly 34,000 evaluations in the Maze and 2,500 evaluations in Mario, whereas standard methods required 100,000 and 5,760 evaluations respectively—reducing required simulations by approximately 57% to 66%. Ablation analyses confirmed that downsampling enabled iterative self-correction across more training cycles, while ancillary occupancy prediction significantly reduced measurement error.

These results demonstrate that surrogate models make automated agent stress-testing economically feasible and practical. By rapidly revealing behavioral edge cases, catastrophic errors, and unexpected traps prior to physical deployment, the method reduces development timelines and mitigates operational safety risks for autonomous systems. The article notes that while surrogate assistance without enhancements offers limited gains, combining downsampling with behavior prediction consistently outperforms existing baselines.

Organizations developing autonomous systems should adopt surrogate-assisted diversity generation to automate safety verification, benchmark agent robustness, and uncover hidden vulnerabilities. Technical teams should integrate behavioral proxy metrics, like spatial occupancy tracking, when modeling complex agent policies. Furthermore, automated environment generation should be explored as an adaptive training curriculum to improve agent generalization.

The study's primary limitations include its reliance on two-dimensional grid domains and the omission of temporal sequence data in spatial occupancy grids, which makes specific dynamic actions, such as jumping sequences, harder to predict. While confidence in the experimental findings is high across the evaluated settings, practitioners should validate the methodology in higher-fidelity, three-dimensional simulations before relying entirely on surrogate predictions for safety-critical hardware deployments.

  • Paper: Robust Reinforcement Learning via Genetic Curriculum, Yeeho Song et al. (2022). This paper extends automated environment generation to robustness curricula by evolving failure-inducing scenarios for reinforcement learning agents.
  • Paper: Mastering Diverse Domains through World Models, Danijar Hafner et al. (2023). This work advances general-purpose world modeling to train and evaluate reinforcement learning policies across diverse and complex domain distributions.
  • Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). This paper generalizes automated open-ended search frameworks from environment generation to the automated design and refinement of entire agentic systems.
Cover for Deep Surrogate Assisted Generation of Environments

Abstract

Recent progress in reinforcement learning (RL) has started producing generally capable agents that can solve a distribution of complex environments. These agents are typically tested on fixed, human-authored environments. On the other hand, quality diversity (QD) optimization has been proven to be an effective component of environment generation algorithms, which can generate collections of high-quality environments that are diverse in the resulting agent behaviors. However, these algorithms require potentially expensive simulations of agents on newly generated environments. We propose Deep Surrogate Assisted Generation of Environments (DSAGE), a sample-efficient QD environment generation algorithm that maintains a deep surrogate model for predicting agent behaviors in new environments. Results in two benchmark domains show that DSAGE significantly outperforms existing QD environment generation algorithms in discovering collections of environments that elicit diverse behaviors of a state-of-the-art RL agent and a planning agent. Our source code and videos are available at https://dsagepaper.github.io/

Table of Contents

  • 1 Introduction
  • 2 Problem Definition
  • 3 Background and Related Work
  • 4 Deep Surrogate Assisted Generation of Environments (DSAGE)
  • 5 Domains
  • 6 Experiments
  • 6.1 Experiment Design
  • 6.2 Analysis
  • 6.3 Ablation Study
  • 6.4 Qualitative Results
  • 7 Societal Impacts
  • 8 Limitations and Future Work
  • Acknowledgments and Disclosure of Funding
  • References
  • Checklist

Knowls

  1. Knowl 1 — Deep Surrogate Assisted Generation of Environments Algorithm

    algorithm

    The Deep Surrogate Assisted Generation of Environments (DSAGE) algorithm generates diverse collections of environments that elicit varied agent behaviors. DSAGE maintains two archives: a ground-truth archive AgtA_{gt} containing environments evaluated via ground-truth agent simulations, and a surrogate archive AsurrogateA_{surrogate} generated during an inner optimization loop guided by a deep surrogate model.

    The algorithm operates in an outer loop with three phases:

    1. Model Exploitation: A quality diversity (QD) optimizer (e.g., MAP-Elites or CMA-ME) generates candidate environment parameters θ\theta. An environment generator g(θ)g(\theta) creates the environment layout. A deep surrogate model predicts ancillary agent behavior data y^\hat{y} (such as an occupancy grid) and then predicts the objective f^\hat{f} and diversity measures m^\hat{m}. The QD optimizer updates AsurrogateA_{surrogate} using these predictions.
    2. Agent Simulation: A subset of solutions from AsurrogateA_{surrogate} is selected via downsampling. The selected environments are evaluated via true agent simulation to obtain ground-truth objective ff, measures mm, and ancillary data yy. Valid improvements are added to AgtA_{gt}, and the evaluation data is appended to dataset D\mathcal{D}.
    3. Model Improvement: The deep surrogate model is trained in a self-supervised manner on dataset D\mathcal{D}.
    Input: N: Maximum evaluation budget, n_rand: Initial random evaluation count, N_exploit: Model exploitation iteration count, B: Batch size for QD optimizer
    Output: Final ground-truth archive A_gt
    Initialize ground-truth archive A_gt, dataset D, and deep surrogate model sm
    \Theta \leftarrow generate_random_solutions(n_rand)
    for \theta \in \Theta do
        env \leftarrow g(\theta)
        f, m, y \leftarrow evaluate(env)
        D \leftarrow D \cup {(\theta, f, m, y)}
        A_gt \leftarrow add_solution(A_gt, (\theta, f, m))
    end for
    evals \leftarrow n_rand
    while evals < N do
        Initialize QD optimizer qd with surrogate archive A_surrogate
        for itr \in {1, 2, ..., N_exploit} do
            \Theta \leftarrow qd.ask(B)
            for \theta \in \Theta do
                env \leftarrow g(\theta)
                \hat{y} \leftarrow sm.predict_ancillary(env)
                \hat{f}, \hat{m} \leftarrow sm.predict(env, \hat{y})
                qd.tell(\theta, \hat{f}, \hat{m})
            end for
        end for
        \Theta \leftarrow select_solutions(A_surrogate)
        for \theta \in \Theta do
            env \leftarrow g(\theta)
            f, m, y \leftarrow evaluate(env)
            D \leftarrow D \cup {(\theta, f, m, y)}
            A_gt \leftarrow add_solution(A_gt, (\theta, f, m))
            evals \leftarrow evals + 1
        end for
        sm.train(D)
    end while
  2. Knowl 2 — Two-Stage Self-Supervised Prediction of Ancillary Agent Behavior

    model/method

    To predict agent performance and behavioral measures from initial environment configurations without experiencing severe regression errors caused by sensitive trajectory shifts, DSAGE implements a two-stage self-supervised prediction pipeline:

    1. Ancillary Behavior Network: A convolutional neural network takes a one-hot representation of the initial environment state and predicts an occupancy grid y^∈RH×W\hat{y} \in \mathbb{R}^{H \times W}, which represents the expected visitation frequency of the agent across each spatial grid tile (h,w)(h, w).
    2. Performance and Measure Network: The predicted occupancy grid y^\hat{y} is concatenated channel-wise with the one-hot representation of the environment. This combined tensor is passed into a second convolutional neural network to predict the scalar task objective f^∈R\hat{f} \in \mathbb{R} and the multi-dimensional diversity measure vector m^∈Rm\hat{m} \in \mathbb{R}^m.

    Ground-truth occupancy grids yy are automatically logged during true agent simulations during the agent simulation phase, providing supervision without human annotation.

  3. Knowl 3 — Surrogate Archive Downsampling for Evaluation Selection

    model/method

    In surrogate-assisted Quality Diversity (QD), the surrogate archive AsurrogateA_{surrogate} contains many candidate environments predicted to be high-performing and diverse. Evaluating all candidates with ground-truth simulations in every outer iteration rapidly exhausts the evaluation budget NN, leading to few surrogate model training rounds.

    DSAGE addresses this by spatially downsampling AsurrogateA_{surrogate}:

    • The continuous or discrete measure space of AsurrogateA_{surrogate} is partitioned into uniform spatial sub-regions of cells.
    • Exactly one solution is sampled uniformly at random from each populated sub-region to form the evaluation batch Θ\Theta.

    This downsampling strategy provides three advantages: (1) it decreases the number of ground-truth simulations per outer iteration, enabling more frequent dataset updates and surrogate retraining iterations for iterative error correction; (2) it balances the distribution of training data across under-explored parts of the measure space; and (3) it prevents clusters of redundant solutions from collapsing into identical cells in the ground-truth archive AgtA_{gt}.

  4. Knowl 4 — Quality Diversity Formulation for Automated Environment Generation

    definition

    In Quality Diversity (QD) environment generation, the objective is to find a set of environment parameters θ∈Rn\theta \in \mathbb{R}^n (e.g., latent vectors passed to a generator g(θ)g(\theta) or direct grid cell layouts) that maximize an objective function while spanning a specified behavioral measure space.

    Let f:Rn→Rf: \mathbb{R}^n \to \mathbb{R} be an objective function evaluating environment validity or agent task performance, and let m:Rn→Rmm: \mathbb{R}^n \to \mathbb{R}^m be a joint measure function mapping an environment to an mm-dimensional behavior descriptor. The measure space S⊆RmS \subseteq \mathbb{R}^m is discretized into MM discrete cells. The archive retains the solution θi\theta_i that maximizes f(θi)f(\theta_i) for each cell i∈{1,…,M}i \in \{1, \dots, M\}.

    The quality and diversity of the generated archive are evaluated using two metrics:

    1. QD-Score: QD-score=∑i=1Mf(θi)\text{QD-score} = \sum_{i=1}^M f(\theta_i) where empty cells contribute an objective value of 00.

    2. Archive Coverage: Coverage=1M∑i=1M1θi\text{Coverage} = \frac{1}{M} \sum_{i=1}^M \mathbf{1}_{\theta_i} where 1θi=1\mathbf{1}_{\theta_i} = 1 if cell ii contains a solution and 00 otherwise.

  5. Knowl 5 — Benchmark Domains for Quality Diversity Environment Generation

    experimental setup

    DSAGE was evaluated on two 2D grid benchmark domains:

    1. Maze Domain:
    • Representation: Direct search over 16×1616 \times 16 grid configurations containing wall tiles, an start tile, and a goal tile.
    • Agent: A pre-trained ACCEL reinforcement learning agent.
    • Objective (ff): Binary solvability (11 if the agent reaches the goal, 00 otherwise).
    • Measures (mm): (1) Number of wall cells (range [0,256][0, 256]), and (2) Mean agent path length over 50 episodes (range [0,648][0, 648], where 648648 indicates navigation timeout/failure).
    • Optimizer: Discrete MAP-Elites.
    1. Mario Domain:
    • Representation: Latent vectors θ\theta fed into a pre-trained Wasserstein GAN to generate Super Mario Bros level segments.
    • Agent: An A* search planning agent.
    • Objective (ff): Level completion rate (proportion of the level segment completed before dying, range [0.0,1.0][0.0, 1.0]).
    • Measures (mm): (1) Number of sky tiles in the top half of the level grid (range [0,150][0, 150]), and (2) Number of jump actions executed by the agent (range [0,100][0, 100]), averaged over 5 episodes.
    • Optimizer: Continuous Covariance Matrix Adaptation MAP-Elites (CMA-ME).
  6. Knowl 6 — Comparative Performance on QD-Score and Archive Coverage

    empirical result

    DSAGE was evaluated over 5 independent trials against several baselines:

    • DSAGE-Only Anc: DSAGE with ancillary occupancy grid prediction and full surrogate archive selection (no downsampling).
    • DSAGE-Only Down: DSAGE with surrogate archive downsampling and direct prediction (no ancillary data).
    • DSAGE Basic: DSAGE with full selection and direct prediction.
    • Baseline QD: MAP-Elites (Maze) or CMA-ME (Mario) without surrogate assistance.
    • Domain Randomization (DR): Uniform random solution sampling.

    One-way ANOVA showed significant differences in QD-score across algorithms for Maze (F(5,24)=430.98,p<0.001F(5, 24) = 430.98, p < 0.001) and Mario (F(5,24)=238.09,p<0.001F(5, 24) = 238.09, p < 0.001). Pairwise Bonferroni-corrected comparisons confirmed DSAGE significantly outperformed DSAGE Basic, Baseline QD, and DR in both domains (p<0.001p < 0.001).

    Maze Mario
    Algorithm QD-score Archive Coverage QD-score Archive Coverage
    DSAGE 16,446.60±42.2716,446.60 \pm 42.27 0.40±0.000.40 \pm 0.00 4,362.29±72.544,362.29 \pm 72.54 0.30±0.000.30 \pm 0.00
    DSAGE-Only Anc 14,568.00±434.5614,568.00 \pm 434.56 0.35±0.010.35 \pm 0.01 2,045.28±201.642,045.28 \pm 201.64 0.16±0.010.16 \pm 0.01
    DSAGE-Only Down 14,205.20±40.8614,205.20 \pm 40.86 0.34±0.000.34 \pm 0.00 4,067.42±102.064,067.42 \pm 102.06 0.30±0.010.30 \pm 0.01
    DSAGE Basic 11,740.00±84.1311,740.00 \pm 84.13 0.28±0.000.28 \pm 0.00 1,306.11±50.901,306.11 \pm 50.90 0.11±0.010.11 \pm 0.01
    Baseline QD 10,480.80±150.1310,480.80 \pm 150.13 0.25±0.000.25 \pm 0.00 1,840.17±95.761,840.17 \pm 95.76 0.13±0.010.13 \pm 0.01
    DR 5,199.60±30.325,199.60 \pm 30.32 0.13±0.000.13 \pm 0.00 92.75±3.0192.75 \pm 3.01 0.01±0.000.01 \pm 0.00

    Values show mean ±\pm standard error of the mean over 5 runs.

  7. Knowl 7 — Sample Efficiency of DSAGE to Target QD-Score

    data/table

    The sample efficiency of DSAGE and its variants was quantified by measuring the total number of ground-truth environment evaluations required to reach fixed target QD-scores (10,480.8010,480.80 in Maze, matching final baseline MAP-Elites; and 1,306.111,306.11 in Mario, matching final DSAGE Basic).

    Algorithm Maze Evaluations (Target: 10,480.8010,480.80) Mario Evaluations (Target: 1,306.111,306.11)
    DSAGE 33,930.40±1,411.0433,930.40 \pm 1,411.04 2,464.40±356.362,464.40 \pm 356.36
    DSAGE-Only Anc 51,919.60±8,254.2451,919.60 \pm 8,254.24 7,727.40±1,433.337,727.40 \pm 1,433.33
    DSAGE-Only Down 42,816.60±691.3842,816.60 \pm 691.38 2,768.60±586.342,768.60 \pm 586.34
    DSAGE Basic 85,328.60±2,947.2485,328.60 \pm 2,947.24 10,00010,000
    Baseline QD 100,000100,000 5,760.00±516.145,760.00 \pm 516.14

    Values denote mean ±\pm standard error of the mean over 5 trials. In Maze, DSAGE reached the target score with approximately 34%34\% of the evaluations required by MAP-Elites (33,930.4033,930.40 vs. 100,000100,000). In Mario, DSAGE achieved the target with less than half the evaluations of CMA-ME (2,464.402,464.40 vs. 5,760.005,760.00).

  8. Knowl 8 — Mean Absolute Error of Deep Surrogate Model Predictions

    data/table

    Surrogate model prediction accuracy was evaluated across DSAGE variants on a unified test dataset compiled from separate experimental runs. The table reports the Mean Absolute Error (MAE) for objective and measure predictions.

    Maze Mario
    Algorithm Objective MAE Wall Cells MAE Path Length MAE Objective MAE Sky Tiles MAE Jumps MAE
    DSAGE 0.03 0.37 96.58 0.10 1.10 7.16
    DSAGE-Only Anc 0.04 0.96 95.14 0.20 1.11 9.97
    DSAGE-Only Down 0.10 0.95 151.50 0.11 0.87 6.52
    DSAGE Basic 0.18 5.48 157.69 0.20 2.16 10.71

    Predicting ancillary agent occupancy grids significantly reduced the error on agent-behavior-dependent measures (Mean Agent Path Length MAE dropped from 157.69157.69 in DSAGE Basic to 96.5896.58 in DSAGE). Downsampling further improved accuracy across objective and measure predictions by enabling iterative dataset updates across a higher number of outer iterations.

  9. Knowl 9 — Main Effects of Ancillary Behavior Modeling and Archive Downsampling

    empirical result

    A 2×22 \times 2 factorial ablation analysis evaluated the main effects of ancillary behavior prediction (occupancy grid prediction vs. direct prediction) and surrogate archive selection (downsampling vs. full selection):

    1. Ancillary Data Prediction: Algorithms predicting ancillary data (DSAGE, DSAGE-Only Anc) performed significantly better in QD-score than models without ancillary prediction (DSAGE-Only Down, DSAGE Basic) in both domains (p<0.001p < 0.001). In Maze, occupancy grid prediction improved path length estimation because path length is directly proportional to cumulative spatial occupancy frequency. In Mario, ancillary occupancy data produced less improvement for jump count predictions because jumping actions depend on sequential execution order rather than occupancy frequency alone.
    2. Surrogate Downsampling: Downsampling (DSAGE, DSAGE-Only Down) significantly outperformed full selection (DSAGE-Only Anc, DSAGE Basic) in both domains (p<0.001p < 0.001). Downsampling increased the number of outer iterations for a fixed evaluation budget (e.g., from 6–7 iterations to ~220 in Maze), allowing the surrogate model to iteratively correct prediction errors.
    3. Interaction: A two-way ANOVA found no significant interaction between ancillary prediction and downsampling in either domain, demonstrating that both techniques contribute independent and complementary performance improvements.
  10. Knowl 10 — Limitations of Spatial Ancillary Representations and 2D Benchmarks

    limitation

    The DSAGE framework has two primary documented limitations:

    1. Lack of Temporal Dynamics: The ancillary behavioral model predicts static 2D spatial occupancy grids. While this avoids compounding errors from auto-regressive multi-step rollouts, omitting temporal sequence information makes it difficult to predict dynamic behaviors that depend on action order and timing, such as jump execution in platformer games.
    2. Low-Dimensional 2D Environments: Experiments were limited to 2D grid environments (Maze and Mario) where individual ground-truth simulations require seconds to minutes. The scalability and surrogate model accuracy in 3D physics-based simulation environments remain untested.

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.R. Wang, J. Lehman, J. Clune, and K. O. Stanley, ‘‘Paired open-ended trailblazer (POET): endlessly generating increasingly complex and diverse learning environments and their solutions,’’ CoRR, vol. abs/1901.01753, 2019.
  2. 2.R. Wang, J. Lehman, A. Rawal, J. Zhi, Y. Li, J. Clune, and K. O. Stanley, ‘‘Enhanced POET: open-ended reinforcement learning through unbounded invention of learning challenges and their solutions,’’ in Proceedings of the 37th International Conference on Machine Learning, ICML, 2020.
  3. 3.M. Dennis, N. Jaques, E. Vinitsky, A. M. Bayen, S. Russell, A. Critch, and S. Levine, ‘‘Emergent complexity and zero-shot transfer via unsupervised environment design,’’ in Advances in Neural Information Processing Systems 33, 2020.
  4. 4.J. Parker-Holder, M. Jiang, M. Dennis, M. Samvelyan, J. N. Foerster, E. Grefenstette, and T. Rocktäschel, ‘‘Evolving curricula with regret-based environment design,’’ CoRR, vol. abs/2203.01302, 2022.
  5. 5.A. Dharna, A. K. Hoover, J. Togelius, and L. Soros, ‘‘Transfer dynamics in emergent evolutionary curricula,’’ IEEE Transactions on Games, 2022.
  6. 6.M. Jiang, M. Dennis, J. Parker-Holder, J. N. Foerster, E. Grefenstette, and T. Rocktäschel, ‘‘Replay-guided adversarial environment design,’’ in Advances in Neural Information Processing Systems 34, 2021.
  7. 7.D. P. Liebana, S. Samothrakis, J. Togelius, T. Schaul, and S. M. Lucas, ‘‘General video game AI: competition, challenges and opportunities,’’ in Proceedings of the 30th AAAI Conference on Artificial Intelligence, 2016.
  8. 8.E. Hambro, S. P. Mohanty, D. Babaev, M. Byeon, D. Chakraborty, E. Grefenstette, M. Jiang, et al., ‘‘Insights from the NeurIPS 2021 NetHack challenge,’’ CoRR, vol. abs/2203.11889, 2022.
  9. 9.D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al., ‘‘A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play,’’ Science, 2018.
  10. 10.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, et al., ‘‘Mastering the game of Go with deep neural networks and tree search,’’ Nature, 2016.
  11. 11.O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al., ‘‘Grandmaster level in StarCraft II using multi-agent reinforcement learning,’’ Nature, 2019.
  12. 12.M. Moravcík, M. Schmid, N. Burch, V. Lisý, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, and M. H. Bowling, ‘‘Deepstack: Expert-level artificial intelligence in no-limit poker,’’ Science, 2017.
  13. 13.N. Brown and T. Sandholm, ‘‘Superhuman AI for multiplayer poker,’’ Science, 2019.
  14. 14.M. C. Fontaine, J. Togelius, S. Nikolaidis, and A. K. Hoover, ‘‘Covariance matrix adaptation for the rapid illumination of behavior space,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2020.
  15. 15.M. C. Fontaine, Y. Hsu, Y. Zhang, B. Tjanaka, and S. Nikolaidis, ‘‘On the importance of environments in human-robot coordination,’’ in Robotics: Science and Systems, 2021.
  16. 16.M. C. Fontaine, R. Liu, A. Khalifa, J. Modi, J. Togelius, A. K. Hoover, and S. Nikolaidis, ‘‘Illuminating mario scenes in the latent space of a generative adversarial network,’’ in Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021.
  17. 17.M. C. Fontaine and S. Nikolaidis, ‘‘A quality diversity approach to automatically generating human-robot interaction scenarios in shared autonomy,’’ in Robotics: Science and Systems, 2021.
  18. 18.A. Gaier, A. Asteroth, and J.-B. Mouret, ‘‘Data-efficient design exploration through surrogate-assisted illumination,’’ Evolutionary Computation, 2018.
  19. 19.Y. Zhang, M. C. Fontaine, A. K. Hoover, and S. Nikolaidis, ‘‘Deep surrogate assisted MAP-Elites for automated hearthstone deckbuilding,’’ CoRR, vol. abs/2112.03534, 2021.
  20. 20.N. Sturtevant, N. Decroocq, A. Tripodi, and M. Guzdial, ‘‘The unexpected consequence of incremental design changes,’’ in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, 2020.
  21. 21.S. Karakovskiy and J. Togelius, ‘‘The mario AI benchmark and competitions,’’ IEEE Transactions on Computational Intelligence and AI in Games, 2012.
  22. 22.R. Baumgarten, ‘‘Infinite super mario AI,’’ 2009.
  23. 23.M. C. Fontaine and S. Nikolaidis, ‘‘Differentiable quality diversity,’’ Advances in Neural Information Processing Systems, vol. 34, 2021.
  24. 24.A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret, ‘‘Robots that can adapt like animals,’’ Nature, 2015.
  25. 25.J. Mouret and J. Clune, ‘‘Illuminating search spaces by mapping elites,’’ CoRR, vol. abs/1504.04909, 2015.
  26. 26.J. K. Pugh, L. B. Soros, and K. O. Stanley, ‘‘Quality diversity: A new frontier for evolutionary computation,’’ Frontiers in Robotics and AI, 2016.
  27. 27.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, ‘‘Generative adversarial nets,’’ in Advances in Neural Information Processing Systems, 2014.
  28. 28.J. Lehman and K. O. Stanley, ‘‘Abandoning objectives: Evolution through the search for novelty alone,’’ Evolutionary Computation, 2011.
  29. 29.J. Lehman and K. O. Stanley, ‘‘Evolving a diversity of virtual creatures through novelty search and local competition,’’ in Proceedings of the 13th Annual Conference on Genetic and Evolutionary Computation, 2011.
  30. 30.P. Kent and J. Branke, ‘‘Bop-elites, a bayesian optimisation algorithm for quality-diversity search,’’ CoRR, vol. abs/2005.04320, 2020.
  31. 31.T. J. Choi and J. Togelius, ‘‘Self-referential quality diversity through differential map-elites,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2021.
  32. 32.E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune, ‘‘Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents,’’ in Advances in Neural Information Processing Systems 31, 2018.
  33. 33.C. Colas, V. Madhavan, J. Huizinga, and J. Clune, ‘‘Scaling map-elites to deep neuroevolution,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2020.
  34. 34.B. Tjanaka, M. C. Fontaine, J. Togelius, and S. Nikolaidis, ‘‘Approximating gradients for differentiable quality diversity in reinforcement learning,’’ CoRR, vol. abs/2202.03666, 2022.
  35. 35.O. Nilsson and A. Cully, ‘‘Policy gradient assisted map-elites,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2021.
  36. 36.G. Cideron, T. Pierrot, N. Perrin, K. Beguir, and O. Sigaud, ‘‘QD-RL: efficient mixing of quality and diversity in reinforcement learning,’’ CoRR, vol. abs/2006.08505, 2020.
  37. 37.A. Hagg, S. Berns, A. Asteroth, S. Colton, and T. Bäck, ‘‘Expressivity of parameterized and data-driven representations in quality diversity search,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2021.
  38. 38.T. Bartz-Beielstein, ‘‘A survey of model-based methods for global optimization,’’ Bioinspired Optimization Methods and Their Applications, 2016.
  39. 39.T. M. Moerland, J. Broekens, and C. M. Jonker, ‘‘Model-based reinforcement learning: A survey,’’ CoRR, vol. abs/2006.16712, 2020.
  40. 40.A. Hagg, D. Wilde, A. Asteroth, and T. Bäck, ‘‘Designing air flow with surrogate-assisted phenotypic niching,’’ in Proceedings of the International Conference on Parallel Problem Solving from Nature, 2020.
  41. 41.L. Cazenille, N. Bredeche, and N. Aubert-Kato, ‘‘Exploring self-assembling behaviors in a swarm of bio-micro-robots using surrogate-assisted map-elites,’’ in IEEE Symposium Series on Computational Intelligence (SSCI), 2019.
  42. 42.A. Gaier, A. Asteroth, and J.-B. Mouret, ‘‘Discovering representations for black-box optimization,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2020.
  43. 43.N. Rakicevic, A. Cully, and P. Kormushev, ‘‘Policy manifold search: Exploring the manifold hypothesis for diversity-based neuroevolution,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2021.
  44. 44.L. Keller, D. Tanneberg, S. Stark, and J. Peters, ‘‘Model-based quality-diversity search for efficient robot learning,’’ in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS, 2020.
  45. 45.B. Lim, L. Grillotti, L. Bernasconi, and A. Cully, ‘‘Dynamics-aware quality-diversity for efficient learning of skill repertoires,’’ CoRR, vol. abs/2109.08522, 2021.
  46. 46.N. Shaker, J. Togelius, and M. J. Nelson, Procedural Content Generation in Games. Springer, 2016.
  47. 47.D. Gravina, A. Khalifa, A. Liapis, J. Togelius, and G. N. Yannakakis, ‘‘Procedural content generation through quality diversity,’’ in Proceedings of the IEEE Conference on Games (CoG), 2019.
  48. 48.S. Earle, J. Snider, M. C. Fontaine, S. Nikolaidis, and J. Togelius, ‘‘Illuminating diverse neural cellular automata for level generation,’’ CoRR, vol. abs/2109.05489, 2021.
  49. 49.A. Khalifa, S. Lee, A. Nealen, and J. Togelius, ‘‘Talakat: Bullet hell generation through constrained map-elites,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2018.
  50. 50.K. Steckel and J. Schrum, ‘‘Illuminating the space of beatable lode runner levels produced by various generative adversarial networks,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2021.
  51. 51.J. Schrum, V. Volz, and S. Risi, ‘‘CPPN2GAN: combining compositional pattern producing networks and gans for large-scale pattern generation,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2020.
  52. 52.A. Sarkar and S. Cooper, ‘‘Generating and blending game levels via quality-diversity in the latent space of a variational autoencoder,’’ in Proceedings of the 16th International Conference on the Foundations of Digital Games, 2021.
  53. 53.A. Summerville, S. Snodgrass, M. Guzdial, C. Holmgård, A. K. Hoover, A. Isaksen, A. Nealen, and J. Togelius, ‘‘Procedural content generation via machine learning (PCGML),’’ IEEE Transactions on Games, 2018.
  54. 54.J. Liu, S. Snodgrass, A. Khalifa, S. Risi, G. N. Yannakakis, and J. Togelius, ‘‘Deep learning for procedural content generation,’’ Neural Computing and Applications, 2021.
  55. 55.S. Snodgrass and S. Ontañón, ‘‘Experiments in map generation using markov chains,’’ in Proceedings of the 9th International Conference on the Foundations of Digital Games, FDG, 2014.
  56. 56.M. Guzdial and M. Riedl, ‘‘Game level generation from gameplay videos,’’ in Proceedings of the 12th Artificial Intelligence and Interactive Digital Entertainment Conference, 2016.
  57. 57.A. Summerville and M. Mateas, ‘‘Super mario as a string: Platformer level generation via LSTMs,’’ in Proceedings of the First Joint International Conference of Digital Games Research Association and Foundation of Digital Games, DiGRA/FDG, 2016.
  58. 58.V. Volz, J. Schrum, J. Liu, S. M. Lucas, A. Smith, and S. Risi, ‘‘Evolving mario levels in the latent space of a deep convolutional generative adversarial network,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2018.
  59. 59.E. Giacomello, P. L. Lanzi, and D. Loiacono, ‘‘DOOM level generation using generative adversarial networks,’’ in Proceedings of the IEEE Games, Entertainment, Media Conference (GEM), 2018.
  60. 60.R. R. Torrado, A. Khalifa, M. C. Green, N. Justesen, S. Risi, and J. Togelius, ‘‘Bootstrapping conditional gans for video game level generation,’’ in Proceedings of the IEEE Conference on Games, 2020.
  61. 61.A. Sarkar, Z. Yang, and S. Cooper, ‘‘Conditional level generation and game blending,’’ CoRR, vol. abs/2010.07735, 2020.
  62. 62.A. Khalifa, P. Bontrager, S. Earle, and J. Togelius, ‘‘PCGRL: procedural content generation via reinforcement learning,’’ in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, 2020.
  63. 63.S. Earle, M. Edwards, A. Khalifa, P. Bontrager, and J. Togelius, ‘‘Learning controllable content generators,’’ in Proceedings of the IEEE Conference on Games (CoG), 2021.
  64. 64.D. Karavolos, A. Liapis, and G. N. Yannakakis, ‘‘A multifaceted surrogate model for search-based procedural content generation,’’ IEEE Transactions on Games, 2021.
  65. 65.J. Togelius, G. N. Yannakakis, K. O. Stanley, and C. Browne, ‘‘Search-based procedural content generation: A taxonomy and survey,’’ IEEE Transactions on Computational Intelligence and AI in Games, 2011.
  66. 66.J. Arnold and R. Alexander, ‘‘Testing autonomous robot control software using procedural content generation,’’ in Proceedings of the 32nd International Conference on Computer Safety, Reliability, and Security, 2013.
  67. 67.G. E. Mullins, P. G. Stankiewicz, R. C. Hawthorne, and S. K. Gupta, ‘‘Adaptive generation of challenging scenarios for testing and evaluation of autonomous vehicles,’’ Journal of Systems and Software, 2018.
  68. 68.Y. Abeysirigoonawardena, F. Shkurti, and G. Dudek, ‘‘Generating adversarial driving scenarios in high-fidelity simulators,’’ in Proceedings of the International Conference on Robotics and Automation (ICRA), 2019.
  69. 69.E. Rocklage, H. Kraft, A. Karatas, and J. Seewig, ‘‘Automated scenario generation for regression testing of autonomous vehicles,’’ in Proceedings of the 20th IEEE International Conference on Intelligent Transportation Systems (ITSC), 2017.
  70. 70.A. Gambi, M. Mueller, and G. Fraser, ‘‘Automatically testing self-driving cars with search-based procedural content generation,’’ in Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis, 2019.
  71. 71.D. Sadigh, S. S. Sastry, and S. A. Seshia, ‘‘Verifying robustness of human-aware autonomous cars,’’ IFAC-PapersOnLine, 2019.
  72. 72.D. J. Fremont, T. Dreossi, S. Ghosh, X. Yue, A. L. Sangiovanni-Vincentelli, and S. A. Seshia, ‘‘Scenic: a language for scenario specification and scene generation,’’ in Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design and Implementation, 2019.
  73. 73.Y. Zhou, S. Booth, N. Figueroa, and J. Shah, ‘‘RoCUS: robot controller understanding via sampling,’’ in Proceedings of the Conference on Robot Learning, 2021.
  74. 74.S. Risi and J. Togelius, ‘‘Increasing generality in machine learning through procedural content generation,’’ Nature Machine Intelligence, 2020.
  75. 75.N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi, ‘‘Illuminating generalization in deep reinforcement learning through procedural level generation,’’ arXiv preprint arXiv:1806.10729, 2018.
  76. 76.K. Cobbe, C. Hesse, J. Hilton, and J. Schulman, ‘‘Leveraging procedural generation to benchmark reinforcement learning,’’ in Proceedings of the International Conference on Machine Learning, 2020.
  77. 77.T. Gabor, A. Sedlmeier, M. Kiermeier, T. Phan, M. Henrich, M. Pichlmair, B. Kempter, C. Klein, H. Sauer, R. S. AG, et al., ‘‘Scenario co-evolution for reinforcement learning on a grid world smart factory domain,’’ in Proceedings of the Genetic and Evolutionary Computation Conference, 2019.
  78. 78.D. M. Bossens and D. Tarapore, ‘‘QED: using quality-environment-diversity to evolve resilient robot swarms,’’ IEEE Transactions on Evolutionary Computation, 2020.
  79. 79.A. Dharna, J. Togelius, and L. B. Soros, ‘‘Co-generation of game levels and game-playing agents,’’ in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, 2020.
  80. 80.M. Jiang, E. Grefenstette, and T. Rocktäschel, ‘‘Prioritized level replay,’’ in Proceedings of the 38th International Conference on Machine Learning, ICML, 2021.
  81. 81.R. Kirk, A. Zhang, E. Grefenstette, and T. Rocktäschel, ‘‘A survey of generalisation in deep reinforcement learning,’’ CoRR, vol. abs/2111.09794, 2021.
  82. 82.M. Chevalier-Boisvert, L. Willems, and S. Pal, ‘‘Minimalistic gridworld environment for OpenAI gym.’’ https://github.com/maximecb/gym-minigrid, 2018.
  83. 83.J. Togelius, S. Karakovskiy, and R. Baumgarten, ‘‘The 2009 mario AI competition,’’ in Proceedings of the IEEE Congress on Evolutionary Computation, CEC, 2010.
  84. 84.N. Jakobi, ‘‘Evolutionary robotics and the radical envelope-of-noise hypothesis,’’ Adaptive Behavior, 1998.
  85. 85.F. Sadeghi and S. Levine, ‘‘CAD2RL: real single-image flight without a single real image,’’ in Proceedings of Robotics: Science and Systems XIII, 2017.
  86. 86.J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, ‘‘Domain randomization for transferring deep neural networks from simulation to the real world,’’ in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS, 2017.
  87. 87.K. O. Stanley, J. Lehman, and L. Soros, ‘‘Open-endedness: The last grand challenge you’ve never heard of,’’ 2017.
  88. 88.A. Ecoffet, J. Clune, and J. Lehman, ‘‘Open questions in creating safe open-ended AI: tensions between control and creativity,’’ CoRR, vol. abs/2006.07495, 2020.
  89. 89.D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt, ‘‘Unsolved problems in ML safety,’’ CoRR, vol. abs/2109.13916, 2021.
  90. 90.A. Roy, N. Memon, J. Togelius, and A. Ross, ‘‘Evolutionary methods for generating synthetic masterprint templates: Dictionary attack in fingerprint recognition,’’ in 2018 International Conference on Biometrics (ICB), pp. 39–46, IEEE, 2018.
  91. 91.C. Xiao, Y. Wu, C. Ma, D. Schuurmans, and M. Müller, ‘‘Learning to combat compounding-error in model-based reinforcement learning,’’ CoRR, vol. abs/1912.11206, 2019.
  92. 92.P. R. Wurman, S. Barrett, K. Kawamoto, J. MacGlashan, K. Subramanian, T. J. Walsh, R. Capobianco, A. Devlic, F. Eckert, F. Fuchs, et al., ‘‘Outracing champion gran turismo drivers with deep reinforcement learning,’’ Nature, 2022.
  93. 93.T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein, Introduction to algorithms. MIT press, 2022.
  94. 94.A. Khalifa, ‘‘Mario AI framework.’’ https://github.com/amidos2006/Mario-AI-Framework, 2019.
  95. 95.M. Arjovsky, S. Chintala, and L. Bottou, ‘‘Wasserstein generative adversarial networks,’’ in Proceedings of the 34th International Conference on Machine Learning, 2017.
  96. 96.I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. Courville, ‘‘Improved training of wasserstein GANs,’’ in Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017.
  97. 97.S. Ioffe and C. Szegedy, ‘‘Batch normalization: Accelerating deep network training by reducing internal covariate shift,’’ in Proceedings of the 32nd International Conference on Machine Learning, ICML, 2015.
  98. 98.K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Deep residual learning for image recognition,’’ in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  99. 99.B. Tjanaka, M. C. Fontaine, D. H. Lee, T. T. M. Vu, Y. Zhang, S. Sommerer, N. Dennler, and S. Nikolaidis, ‘‘pyribs: A bare-bones python library for quality diversity optimization.’’ https://github.com/icaros-usc/pyribs, 2021.
  100. 100.D. P. Kingma and J. Ba, ‘‘Adam: A method for stochastic optimization,’’ in Proceedings of the 3rd International Conference on Learning Representations, ICLR, 2015.
  101. 101.A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, et al., ‘‘PyTorch: an imperative style, high-performance deep learning library,’’ in Advances in Neural Information Processing Systems, 2019.
  102. 102.A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra, ‘‘Grokking: Generalization beyond overfitting on small algorithmic datasets,’’ CoRR, vol. abs/2201.02177, 2022.

Citation

MLA
Bhatt, V., et al. “Deep Surrogate Assisted Generation of Environments”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 37762–77, https://proceedings.neurips.cc/paper_files/paper/2022/file/f649556471416b35e60ae0de7c1e3619-Paper-Conference.pdf.
APA
Bhatt, V., Tjanaka, B., Fontaine, M., & Nikolaidis, S. (2022). Deep Surrogate Assisted Generation of Environments. Advances in Neural Information Processing Systems, 35, 37762–37777. https://proceedings.neurips.cc/paper_files/paper/2022/file/f649556471416b35e60ae0de7c1e3619-Paper-Conference.pdf
Chicago
Bhatt, V., B. Tjanaka, M. Fontaine, and S. Nikolaidis. 2022. “Deep Surrogate Assisted Generation of Environments”. Advances in Neural Information Processing Systems 35: 37762–77. https://proceedings.neurips.cc/paper_files/paper/2022/file/f649556471416b35e60ae0de7c1e3619-Paper-Conference.pdf.
Harvard
Bhatt, V. et al. (2022) “Deep Surrogate Assisted Generation of Environments”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37762–37777. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/f649556471416b35e60ae0de7c1e3619-Paper-Conference.pdf.
Vancouver
1. Bhatt V, Tjanaka B, Fontaine M, Nikolaidis S (2022) Deep Surrogate Assisted Generation of Environments. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37762–37777

BibTeX

@inproceedings{bhatt2022deep,
  title = {Deep Surrogate Assisted Generation of Environments},
  author = {Bhatt, Varun and Tjanaka, Bryon and Fontaine, Matthew and Nikolaidis, Stefanos},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {37762-37777},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/f649556471416b35e60ae0de7c1e3619-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors