A Survey of Zero-shot Generalisation in Deep Reinforcement Learning

Robert KirkAmy ZhangEdward GrefenstetteTim Rocktäschel

article2023JAIR305 citations

Presents a unifying mathematical formalism and taxonomy for zero-shot generalization in deep reinforcement learning, critically evaluating existing benchmarks and methods to guide the development of policies that successfully transfer to unseen environments.

Listen

Reinforcement learning offers strong potential for automation across critical applications such as autonomous vehicles, robotics, and healthcare systems. However, traditional algorithms are evaluated in the exact environments where they were trained, leading to severe performance drops when faced with new, unfamiliar conditions. Because direct trial-and-error training in real-world environments is often unsafe and expensive, systems must be able to perform reliably upon deployment without additional training. The article evaluates the current landscape of zero-shot generalisation in reinforcement learning, establishing a unified framework to categorize existing benchmarks and algorithmic solutions across diverse deployment challenges.

The authors conducted a comprehensive review and structural analysis covering 55 simulation environments and dozens of modern algorithms. Using an extended contextual decision-making model, the analysis classifies environments by four main operational variables: state layout, visual observation, physical dynamics, and reward functions. It also organizes evaluation protocols based on whether tests require interpolating within familiar ranges or extrapolating to unseen conditions, and it categorizes methods by how they modify training data, adjust network architectures, or alter optimization objectives.

The analysis reveals several critical findings across the field. First, existing research is heavily skewed toward visual and spatial changes—with spatial variation present in roughly 76% of surveyed environments and visual variation in 53%—while physical dynamics (35%) and goal or reward variation (36%) remain significantly underrepresented. Second, standard benchmarks that rely purely on randomized procedural generation hide the underlying factors of variation, preventing precise diagnosis of failure modes. Third, most existing solutions attempt to improve robustness solely through adjusted loss functions, leaving model architecture improvements and within-episode adaptation largely underutilized. Finally, out-of-distribution performance cannot be achieved generically; successful transfer strictly requires problem-specific structural assumptions and tailored inductive biases.

These findings mean that current benchmarks may provide misleading confidence for real-world readiness. High performance on standard visual benchmarks does not translate to resilience against changes in real-world physics, equipment wear, or shifting operational objectives. For leadership, relying on broad generalisation claims without verifying the exact type of variation creates safety, compliance, and financial risks, as unseen operational shifts can lead to sudden system failures.

Organizations developing these systems should adopt benchmark environments that combine procedural variation with explicit, controllable parameters rather than relying on pure random generation. Teams should prioritize evaluating algorithms across multiple distinct operational dimensions—such as context efficiency, offline data learning, and task variations—using multidimensional scorecards rather than single performance metrics. While the article provides high confidence regarding the structural limitations of current benchmarks and algorithms, it focuses primarily on empirical literature rather than theoretical mathematical bounds. Decision-makers should exercise caution and require targeted verification protocols before deploying autonomous policies into safety-critical operations.

Cover for A Survey of Zero-shot Generalisation in Deep Reinforcement Learning

Abstract

The study of zero-shot generalisation (ZSG) in deep Reinforcement Learning (RL) aims to produce RL algorithms whose policies generalise well to novel unseen situations at deployment time, avoiding overfitting to their training environments. Tackling this is vital if we are to deploy reinforcement learning algorithms in real world scenarios, where the environment will be diverse, dynamic and unpredictable. This survey is an overview of this nascent field. We rely on a unifying formalism and terminology for discussing different ZSG problems, building upon previous works. We go on to categorise existing benchmarks for ZSG, as well as current methods for tackling these problems. Finally, we provide a critical discussion of the current state of the field, including recommendations for future work. Among other conclusions, we argue that taking a purely procedural content generation approach to benchmark design is not conducive to progress in ZSG, we suggest fast online adaptation and tackling RL-specific problems as some areas for future work on methods for ZSG, and we recommend building benchmarks in underexplored problem settings such as offline RL ZSG and reward-function variation.

Table of Contents

  • 1. Introduction
  • 2. Related Work: Surveys In Reinforcement Learning Subfields
  • 3. Formalising Zero-shot Generalisation In Reinforcement Learning
  • 3.1 Background: Generalisation In Supervised Learning
  • 3.2 Background: Reinforcement Learning
  • 3.3 Contextual Markov Decision Processes
  • 3.4 Training And Testing Contexts
  • 3.5 Real World Examples of This Formalism
  • 3.6 Additional Assumptions For More Feasible Generalisation
  • 3.7 Remarks And Discussion
  • 4. Benchmarks For Zero-shot Generalisation In Reinforcement Learning
  • 4.1 Environments
  • 4.1.1 Trends In Environments
  • 4.2 Evaluation Protocols For Zsg
  • 4.3 Discussion
  • 5. Methods For Zero-shot Generalisation In Reinforcement Learning
  • 5.1 Increasing Similarity Between Training And Testing
  • 5.1.1 Data Augmentation and Domain Randomisation
  • 5.1.2 Environment Generation
  • 5.1.3 Optimisation Objectives
  • 5.2 Handling Differences Between Training And Testing
  • 5.2.1 Encoding Inductive Biases
  • 5.2.2 Regularisation and Simplicity
  • 5.2.3 Learning Invariances
  • 5.2.4 Adapting Online
  • 5.3 RL-Specific Problems And Improvements
  • 5.3.1 RL-specific Problems
  • 5.3.2 Better Optimisation without Overfitting
  • 5.4 Discussion
  • 6. Discussion And Future Work
  • 6.1 Generalisation Beyond Zero-Shot Policy Transfer
  • 6.2 Real World Reinforcement Learning Generalisation
  • 6.3 Multi-Dimensional Evaluation Of Generalisation
  • 6.4 Tackling Stronger Types Of Variation
  • 6.5 Understanding Generalisation In Reinforcement Learning
  • 6.6 Future Work On Methods For Zero-shot Generalisation
  • 7. Conclusion
  • Acknowledgements
  • Appendix A. Other Structural Assumptions on MDPs
  • References

Knowls

  1. Knowl 1 — Contextual Markov Decision Process Formalism for Zero-Shot Generalisation

    definition

    A Contextual Markov Decision Process (CMDP) is defined as a tuple M=(S′,A,O,R,T,C,ϕ:S′×C→O,p(s′∣c),p(c))\mathcal{M} = \left(\mathcal{S}', \mathcal{A}, \mathcal{O}, \mathcal{R}, \mathcal{T}, \mathcal{C}, \phi: \mathcal{S}' \times \mathcal{C} \to \mathcal{O}, p(s'|c), p(c)\right) where:

    • S′\mathcal{S}' is the underlying state space,
    • A\mathcal{A} is the action space,
    • O\mathcal{O} is the observation space,
    • C\mathcal{C} is the context space (the set of parameters, seeds, or IDs specifying distinct environment instances or tasks),
    • p(c)p(c) is the context distribution,
    • p(s′∣c)p(s'|c) is the initial state distribution over S′\mathcal{S}' conditioned on context cc,
    • R:S′×C×A×S′×C→R\mathcal{R}: \mathcal{S}' \times \mathcal{C} \times \mathcal{A} \times \mathcal{S}' \times \mathcal{C} \to \mathbb{R} is the reward function,
    • T\mathcal{T} is the transition probability distribution over next state given current state and action,
    • ϕ:S′×C→O\phi: \mathcal{S}' \times \mathcal{C} \to \mathcal{O} is the observation emission function.

    The CMDP is formally equivalent to a POMDP defined over the augmented state space S=S′×C\mathcal{S} = \mathcal{S}' \times \mathcal{C} with initial state distribution p((s′,c))=p(c)p(s′∣c)p((s', c)) = p(c)p(s'|c). The transition dynamics are strictly factored such that the context cc remains constant within an episode: T((s′,c),a)((s′′,c′))=0if c′≠c.\mathcal{T}((s', c), a)((s'', c')) = 0 \quad \text{if } c' \neq c.

    If O=O′×C\mathcal{O} = \mathcal{O}' \times \mathcal{C} and ϕ((s′,c))=(ϕ′(s′),c)\phi((s', c)) = (\phi'(s'), c), the context is observed; otherwise, it is unobserved. A single context c∗c^* defines a context-MDP Mc∗\mathcal{M}_{c^*}, which is the restriction of M\mathcal{M} where p(c∗)=1p(c^*) = 1.

  2. Knowl 2 — Zero-Shot Policy Transfer Problem Formulation

    definition

    Given a Contextual MDP M=(S′,A,O,R,T,C,ϕ,p(s′∣c),p(c))\mathcal{M} = (\mathcal{S}', \mathcal{A}, \mathcal{O}, \mathcal{R}, \mathcal{T}, \mathcal{C}, \phi, p(s'|c), p(c)), any context subset C′⊆C\mathcal{C}' \subseteq \mathcal{C} defines a restricted CMDP M∣C′\mathcal{M}|_{\mathcal{C}'} with context distribution p′(c)={p(c)Zif c∈C′,0otherwise,p'(c) = \begin{cases} \frac{p(c)}{Z} & \text{if } c \in \mathcal{C}', \\ 0 & \text{otherwise}, \end{cases} where Z=∑c∈C′p(c)Z = \sum_{c \in \mathcal{C}'} p(c) is the normalisation constant.

    A Zero-Shot Policy Transfer (ZSPT) problem is defined by a CMDP M\mathcal{M} together with training and testing context sets Ctrain,Ctest⊆C\mathcal{C}_{\text{train}}, \mathcal{C}_{\text{test}} \subseteq \mathcal{C}. The objective is to find a non-Markovian policy π:H[O,A]→A\pi: \mathcal{H}[\mathcal{O}, \mathcal{A}] \to \mathcal{A} mapping histories of observations and actions ht=(o1,a1,r1,…,ot)∈H[O,A]h_t = (o_1, a_1, r_1, \dots, o_t) \in \mathcal{H}[\mathcal{O}, \mathcal{A}] to actions that maximises the expected return on the testing CMDP: J(π):=R(π,M∣Ctest)=Ec∼p(c∣Ctest)[R(π,Mc)],J(\pi) := \mathcal{R}(\pi, \mathcal{M}|_{\mathcal{C}_{\text{test}}}) = \mathbb{E}_{c \sim p(c|\mathcal{C}_{\text{test}})}\left[\mathcal{R}(\pi, \mathcal{M}_c)\right], where π\pi is learned solely through interaction with the training CMDP M∣Ctrain\mathcal{M}|_{\mathcal{C}_{\text{train}}} within a budget of NsN_s environment steps and NeN_e training episodes, with zero updates or interaction allowed on Ctest\mathcal{C}_{\text{test}} during evaluation.

    In a controllable context ZSPT problem, the learning algorithm is allowed to actively adapt the training context sampling distribution ptrain(c)p_{\text{train}}(c) during training, provided ptrain(c)=0p_{\text{train}}(c) = 0 for all c∉Ctrainc \notin \mathcal{C}_{\text{train}}.

  3. Knowl 3 — Reinforcement Learning Generalisation Gap

    equation

    In the Contextual Markov Decision Process (CMDP) framework, the generalisation gap of a policy π\pi trained on context set Ctrain\mathcal{C}_{\text{train}} and evaluated on test context set Ctest\mathcal{C}_{\text{test}} is defined as: GenGap(π):=R(π,M∣Ctrain)−R(π,M∣Ctest)\text{GenGap}(\pi) := \mathcal{R}(\pi, \mathcal{M}|_{\mathcal{C}_{\text{train}}}) - \mathcal{R}(\pi, \mathcal{M}|_{\mathcal{C}_{\text{test}}}) where R(π,M∣C′)=Ec∼p(c∣C′)[R(π,Mc)]\mathcal{R}(\pi, \mathcal{M}|_{\mathcal{C}'}) = \mathbb{E}_{c \sim p(c|\mathcal{C}')}[\mathcal{R}(\pi, \mathcal{M}_c)] denotes the expected cumulative return of policy π\pi evaluated on the subset CMDP M∣C′\mathcal{M}|_{\mathcal{C}'}.

    A smaller generalisation gap indicates that performance does not degrade drastically from training to testing. However, the generalisation gap has two core limitations:

    1. A policy executing random actions on both train and test context sets yields GenGap(π)≈0\text{GenGap}(\pi) \approx 0 despite having near-zero task capability.
    2. When test contexts differ in intrinsic difficulty or reward scaling from training contexts, the numerical magnitude of GenGap(π)\text{GenGap}(\pi) is uninformative about true transfer capability.

    Consequently, the generalisation gap serves as an auxiliary diagnostic metric alongside expected test return R(π,M∣Ctest)\mathcal{R}(\pi, \mathcal{M}|_{\mathcal{C}_{\text{test}}}), useful for breaking ties between algorithms or evaluating deployment safety assurances.

  4. Knowl 4 — Taxonomy of Environment Variation Dimensions in Contextual MDPs

    definition

    In Contextual Markov Decision Processes M=(S′,A,O,R,T,C,ϕ,p(s′∣c),p(c))\mathcal{M} = (\mathcal{S}', \mathcal{A}, \mathcal{O}, \mathcal{R}, \mathcal{T}, \mathcal{C}, \phi, p(s'|c), p(c)), variations across context-MDPs Mc\mathcal{M}_c are categorized along four fundamental dimensions:

    1. State Variation (SS): The context cc alters the initial underlying state distribution p(s′∣c)p(s'|c) and reachable layout geometries or entity placements (e.g. maze layouts, object positions).
    2. Observation Variation (OO): The context cc alters the observation emission function ϕ(s′,c)\phi(s', c), altering visual appearance, textures, camera angles, or distractors while leaving the underlying state transition dynamics and reward structures unchanged.
    3. Dynamics Variation (DD): The context cc modulates the transition probability distribution T((s′,c),a)\mathcal{T}((s', c), a), altering physical parameters such as friction coefficients, gravity, mass distributions, or motor delays.
    4. Reward Variation (RR): The context cc changes the reward function R(s′,c,a)\mathcal{R}(s', c, a), corresponding to shifts in goals, preferences, or tasks. To remain solvable zero-shot by a single policy, reward variation requires supplying goal or task specifications in the observation space to maintain consistent optimality.
  5. Knowl 5 — Evaluation Protocol Taxonomy for Controllable Context Spaces

    model/method

    In controllable contextual environments where the context space C\mathcal{C} comprises structured parameter axes (continuous, discrete ordinal, or discrete cardinal), evaluation protocols define specific relations between the training context distribution ptrain(c)p_{\text{train}}(c) and testing distribution ptest(c)p_{\text{test}}(c):

    1. Interpolation: Testing context parameters lie strictly inside the support or convex hull of the training context set conv(Ctrain)\text{conv}(\mathcal{C}_{\text{train}}).
    2. Single-Factor or Multi-Factor Extrapolation: Testing context parameters take values strictly outside the training range along one or more ordinal or continuous factor axes (e.g. higher gravity or larger maze dimensions).
    3. Combinatorial Interpolation (Systematicity): The training distribution is non-convex across multiple factor dimensions (i.e. factors vary during training, but certain cross-combinations are held out). Test contexts evaluate combinations of factor values that were each observed independently during training but never experienced simultaneously. This tests systematic compositional generalisation.
  6. Knowl 6 — Taxonomy of Solution Methods for Zero-Shot Generalisation in Deep RL

    model/method

    Methods for tackling zero-shot generalisation (ZSG) in deep reinforcement learning are categorized into three major families based on their underlying approach:

    1. Increasing Similarity Between Training and Testing:
      • Data Augmentation and Domain Randomisation: Expanding training diversity through image transformations (e.g. DrQ, SODA, SVEA) or randomized simulator dynamics (e.g. ADR, CAD2RL).
      • Environment Generation and Unsupervised Environment Design (UED): Learning or evolving solvable, high-utility context-MDP curricula (e.g. POET, PAIRED, Prioritized Level Replay).
      • Optimisation Objectives: Modifying the RL objective to target worst-case perturbations across context sets (e.g. Robust RL, RARL, WR2L).
    2. Handling Differences Between Training and Testing:
      • Encoding Inductive Biases: Incorporating architectural structures such as relational reasoning, modular attention, or invariant representations (e.g. AttentionAgent, SchemaNetworks, IDAAC, DARLA).
      • Regularisation and Simplicity: Applying information bottlenecks, dropout, weight decay, or large residual network backbones to prevent overfitting.
      • Learning Invariances: Using bisimulation metrics, contrastive learning, or causal state abstractions to ignore task-irrelevant features (e.g. DBC, PSM, ICP, DRIBO).
      • Online Adaptation within an Episode: Inferring latent context online or deploying meta-learned recurrent policies capable of adapting within a single test episode (e.g. UP-OSI, RMA, VariBAD, BOReL, PAD).
    3. RL-Specific Improvements:
      • Countering non-stationarity via iterative policy distillation (e.g. ITER).
      • Decoupling policy and value optimization architectures to allow value heads to train longer without inducing policy overfitting (e.g. PPG, DAAC).
      • Model-based reinforcement learning with learned world models and tree search planning for procedural generalization (e.g. MuZero Reanalyse).
  7. Knowl 7 — Structured Contextual MDP Formulations: Factored and Block MDPs

    definition

    To enable formal generalisation guarantees, structured MDP frameworks impose explicit conditional independence or latent state decompositions on the environment:

    1. Factored MDP: The state space is represented by a set of discrete state variables S:={S1,S2,…,Sn}\mathcal{S} := \{S_1, S_2, \dots, S_n\}. The transition probability and expected reward decompose conditionally over parent subsets PA(Si)⊆S\text{PA}(S_i) \subseteq \mathcal{S}: T(s′∣s,a)=∏i=1nP(si′∣PA(si′),a)\mathcal{T}(s'|s, a) = \prod_{i=1}^n P(s'_i | \text{PA}(s'_i), a) E[R(s,a)]=∑i=1nE[Ri(si,a)]\mathbb{E}[\mathcal{R}(s, a)] = \sum_{i=1}^n \mathbb{E}[\mathcal{R}_i(s_i, a)] Contexts can be framed as assignments to specific factors, enabling formal compositional generalisation (systematicity) to unseen factor combinations.

    2. Block MDP: Defined as a tuple ⟨S,A,X,p,q,R⟩\langle \mathcal{S}, \mathcal{A}, \mathcal{X}, p, q, \mathcal{R} \rangle, where S\mathcal{S} is a compact, unobservable latent state space, X\mathcal{X} is a high-dimensional observable space, p(s′∣s,a)p(s'|s,a) is the latent transition distribution, q(x∣s)q(x|s) is the observation emission function, and R(s,a)\mathcal{R}(s,a) is the reward function. Disjoint observation sets correspond to distinct latent states, allowing agents to learn invariant representations that filter out irrelevant observation variations.

  8. Knowl 8 — Hierarchy of Generalisation Difficulty across Factor Types

    theoretical result

    In contextual reinforcement learning with parameterized factor spaces, generalisation difficulty follows a partial ordering based on whether factor axes are ordinal (possessing an ordering or metric structure, including continuous variables) or cardinal (unordered discrete categorical choices), and whether test contexts require interpolation or extrapolation:

    1. Ordinal Interpolation: Evaluating on unseen values situated strictly within observed training bounds along continuous or ordered discrete axes (lowest difficulty).
    2. Cardinal Interpolation (Combinatorial Interpolation): Evaluating on novel combinations of individually observed discrete cardinal values paired with other factor values.
    3. Ordinal Extrapolation: Evaluating on parameter values outside the minimum and maximum bounds observed during training along continuous or ordinal dimensions (e.g. higher velocities, larger gravity, or longer horizon lengths).
    4. Cardinal Extrapolation: Evaluating on entirely novel, unseen discrete entities, object categories, mechanics, or game modes not present in the training set (highest difficulty; intractable without external prior knowledge, transfer learning, or strong inductive biases).
  9. Knowl 9 — Limitations of Purely Black-Box Procedural Content Generation for Generalisation Benchmarking

    limitation

    Purely black-box Procedural Content Generation (PCG) benchmarks—where environment instances are generated solely from an unexposed scalar random seed (such as standard OpenAI Procgen or NetHack Learning Environment)—have critical shortcomings for zero-shot generalisation research:

    1. Lack of Factor Disentanglement: Because a single seed jointly controls layout structure, asset textures, entity placement, and dynamics, researchers cannot isolate which specific factor of variation (state, observation, or dynamics) causes policy failure.
    2. Restricted Evaluation Protocol Options: Evaluation protocols on purely PCG environments are limited exclusively to adjusting the number of training seeds NtrainN_{\text{train}} versus evaluating on the full seed distribution. They cannot systematically test structured interpolation, combinatorial interpolation, or extrapolation splits.
    3. Conflation of Optimization and Generalisation: Evaluating on a small set of random seeds tests whether an algorithm avoids memorization, but does not diagnose targeted generalisation failures or compositional capabilities.

    A hybrid benchmark design is recommended: using PCG to produce rich background variation while providing explicit, controllable parameters for specific factors of variation to enable precise scientific analysis.

  10. Knowl 10 — The Principle of Unchanged Optimality in Zero-Shot Generalisation

    assumption

    The Principle of Unchanged Optimality requires that across all context-MDPs Mc\mathcal{M}_c within a contextual MDP M\mathcal{M}, the optimal action choice for a given underlying environment state s′s' remains invariant, or that the observation provides sufficient goal/task conditioning to preserve a unique optimal mapping.

    Under observation variation (OO) or state variation (SS), the optimal policy mapped from underlying state features typically remains consistent across context instances. Under dynamics variation (DD) or reward variation (RR), this principle is easily violated:

    • Dynamics shifts alter state-transition probabilities, potentially rendering previously optimal trajectories sub-optimal or hazardous.
    • Reward shifts alter the task objective entirely; zero-shot transfer without extra information is ill-defined unless the agent receives explicit goal representations, task IDs, or language instructions g(c)g(c) that condition the policy π(a∣o,g(c))\pi(a | o, g(c)) on the active task.

Coverage note — Specific empirical benchmark comparison tables and individual literature citations for all 55 surveyed environments and algorithms were omitted to focus on the paper's core conceptual formalisms, taxonomies, and theoretical frameworks.

References

  1. 1.Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., & Riedmiller, M. A. (2018). Maximum a posteriori policy optimisation. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  2. 2.Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., & Wang, J. (2019). Wasserstein Robust Reinforcement Learning.. arXiv:1907.13196 [cs, stat]..
  3. 3.Ada, S. E., Ugur, E., & Akin, H. L. (2021). Generalization in Transfer Learning.. arXiv:1909.01331 [cs, stat]., Comment: 23 pages, 36 figures.
  4. 4.Agarwal, R., Machado, M. C., Castro, P. S., & Bellemare, M. G. (2021). Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning.. arXiv:2101.05265 [cs, stat]., Comment: ICLR 2021 (Spotlight). Website: https://agarwl.github.io/pse.
  5. 5.Ahmed, O., Träuble, F., Goyal, A., Neitz, A., Bengio, Y., Schölkopf, B., Wüthrich, M., & Bauer, S. (2020). CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning.. arXiv:2010.04296 [cs, stat]., Comment: The first two authors contributed equally, the last two authors avised jointly.
  6. 6.Albrecht, S. V., & Stone, P. (2018). Autonomous agents modelling other agents: A comprehensive survey and open problems.. Artificial Intelligence..
  7. 7.Amin, S., Gomrokchi, M., Satija, H., van Hoof, H., & Precup, D. (2021). A Survey of Exploration Methods in Reinforcement Learning.. arXiv:2109.00157 [cs]..
  8. 8.Anand, A., Walker, J., Li, Y., Vértes, E., Schrittwieser, J., Ozair, S., Weber, T., & Hamrick, J. B. (2021). Procedural Generalization by Planning with Self-Supervised World Models.. arXiv:2111.01587 [cs]..
  9. 9.Arjovsky, M., Bottou, L., Gulrajani, I., & Lopez-Paz, D. (2020). Invariant Risk Minimization.. arXiv:1907.02893 [cs, stat]..
  10. 10.Arora, S., & Doshi, P. (2020). A Survey of Inverse Reinforcement Learning: Challenges, Methods and Progress.. arXiv:1806.06877 [cs, stat]..
  11. 11.Ball, P. J., Lu, C., Parker-Holder, J., & Roberts, S. (2021). Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment.. arXiv:2104.05632 [cs]., Comment: Accepted @ ICML 2021; Spotlight @ ICLR 2021 "Self-Supervision for Reinforcement Learning Workshop".
  12. 12.Bapst, V., Sanchez-Gonzalez, A., Doersch, C., Stachenfeld, K. L., Kohli, P., Battaglia, P. W., & Hamrick, J. B. (2019). Structured agents for physical construction. In Chaudhuri, K., & Salakhutdinov, R. (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, pp. 464–474. PMLR.
  13. 13.Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G., Vaswani, A., Allen, K., Nash, C., Langston, V., Dyer, C., Heess, N., Wierstra, D., Kohli, P., Botvinick, M., Vinyals, O., Li, Y., & Pascanu, R. (2018). Relational inductive biases, deep learning, and graph networks.. arXiv:1806.01261 [cs, stat]..
  14. 14.Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The Arcade Learning Environment: An Evaluation Platform for General Agents.. Journal of Artificial Intelligence Research..
  15. 15.Bengio, E., Pineau, J., & Precup, D. (2020). Interference and generalization in temporal difference learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 767–777. PMLR.
  16. 16.Benjamins, C., Eimer, T., Schubert, F., Biedenkapp, A., Rosenhahn, B., Hutter, F., & Lindauer, M. (2021). CARL: A Benchmark for Contextual and Adaptive Reinforcement Learning..
  17. 17.Bertrán, M., Martínez, N., Phielipp, M., & Sapiro, G. (2020). Instance-based generalization in reinforcement learning. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., & Lin, H. (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  18. 18.Biedenkapp, A., Bozkurt, H. F., Eimer, T., Hutter, F., & Lindauer, M. (2020). Dynamic Algorithm Configuration: Foundation of a New Meta-Algorithmic Framework.. ECAI 2020..
  19. 19.Boutilier, C., Dearden, R., & Goldszmidt, M. (2000). Stochastic dynamic programming with factored representations.. Artificial Intelligence..
  20. 20.Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI Gym.. arXiv:1606.01540 [cs]..
  21. 21.Chen, A. S., Nair, S., & Finn, C. (2021). Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos.. arXiv:2103.16817 [cs]., Comment: https://sites.google.com/view/dvd-human-videos.
  22. 22.Chen, J. Z. (2020). Reinforcement Learning Generalization with Surprise Minimization.. arXiv:2004.12399 [cs]., Comment: Inductive biases, invariances and generalization in RL Workshop, ICML 2020.
  23. 23.Chen, S., & Li, Y. (2020). An Overview of Robust Reinforcement Learning. In 2020 IEEE International Conference on Networking, Sensing and Control (ICNSC), pp. 1–6.
  24. 24.Chen, V., Gupta, A., & Marino, K. (2021). Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning.. arXiv:2011.00517 [cs]..
  25. 25.Chevalier-Boisvert, M. (2021). Minimalistic Gridworld Environment (MiniGrid)..
  26. 26.Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., & Bengio, Y. (2019). Babyai: A platform to study the sample efficiency of grounded language learning. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  27. 27.Cobbe, K., Hesse, C., Hilton, J., & Schulman, J. (2020a). Leveraging procedural generation to benchmark reinforcement learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 2048–2056. PMLR.
  28. 28.Cobbe, K., Hilton, J., Klimov, O., & Schulman, J. (2020b). Phasic Policy Gradient.. arXiv:2009.04416 [cs, stat]..
  29. 29.Cobbe, K., Klimov, O., Hesse, C., Kim, T., & Schulman, J. (2019). Quantifying generalization in reinforcement learning. In Chaudhuri, K., & Salakhutdinov, R. (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, pp. 1282–1289. PMLR.
  30. 30.Côté, M.-A., Kádár, Á., Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Tao, R. Y., Hausknecht, M., Asri, L. E., Adada, M., Tay, W., & Trischler, A. (2019). TextWorld: A Learning Environment for Text-based Games.. arXiv:1806.11532 [cs, stat]., Comment: Presented at the Computer Games Workshop at IJCAI 2018, Stockholm.
  31. 31.Crosby, M., Beyret, B., Shanahan, M., Hernández-Orallo, J., Cheke, L., & Halina, M. (2020). The Animal-AI Testbed and Competition. In Proceedings of the NeurIPS 2019 Competition and Demonstration Track, pp. 164–176. PMLR.
  32. 32.Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., & Levine, S. (2020). Emergent complexity and zero-shot transfer via unsupervised environment design. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., & Lin, H. (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  33. 33.Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  34. 34.DeVries, T., & Taylor, G. W. (2017). Improved Regularization of Convolutional Neural Networks with Cutout.. arXiv:1708.04552 [cs]..
  35. 35.Diuk, C., Cohen, A., & Littman, M. L. (2008). An object-oriented representation for efficient reinforcement learning. In Cohen, W. W., McCallum, A., & Roweis, S. T. (Eds.), Machine Learning, Proceedings of the Twenty-Fifth International Conference (ICML 2008), Helsinki, Finland, June 5-9, 2008, Vol. 307 of ACM International Conference Proceeding Series, pp. 240–247. ACM.
  36. 36.Dorfman, R., Shenfeld, I., & Tamar, A. (2021). Offline Meta Learning of Exploration.. arXiv:2008.02598 [cs, stat]..
  37. 37.Doshi-Velez, F., & Konidaris, G. D. (2016). Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations. In Kambhampati, S. (Ed.), Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, IJCAI 2016, New York, NY, USA, 9-15 July 2016, pp. 1432–1440. IJCAI/AAAI Press.
  38. 38.Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., & Koltun, V. (2017). CARLA: An Open Urban Driving Simulator.. arXiv:1711.03938 [cs]., Comment: Published at the 1st Conference on Robot Learning (CoRL).
  39. 39.Du, S. S., Kakade, S. M., Wang, R., & Yang, L. F. (2020). Is a good representation sufficient for sample efficient reinforcement learning?. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  40. 40.Du, S. S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudík, M., & Langford, J. (2019). Provably efficient RL with rich observations via latent state decoding. In Chaudhuri, K., & Salakhutdinov, R. (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, pp. 1665–1674. PMLR.
  41. 41.Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., & Abbeel, P. (2016). Rl 22: Fast reinforcement learning via slow reinforcement learning.. arXiv preprint ˆ arXiv:1611.02779..
  42. 42.Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., & Hester, T. (2021). An empirical investigation of the challenges of real-world reinforcement learning.. arXiv:2003.11881 [cs]., Comment: arXiv admin note: text overlap with arXiv:1904.12901.
  43. 43.E. Todorov, T. Erez, & Y. Tassa (2012). MuJoCo: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026–5033.
  44. 44.Eimer, T., Biedenkapp, A., Reimer, M., Adriaensen, S., Hutter, F., & Lindauer, M. (2021). DACBench: A Benchmark Library for Dynamic Algorithm Configuration.. arXiv:2105.08541 [cs]., Comment: Accepted at IJCAI 2021.
  45. 45.Eysenbach, B., Gupta, A., Ibarz, J., & Levine, S. (2019). Diversity is all you need: Learning skills without a reward function. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  46. 46.Eysenbach, B., Salakhutdinov, R., & Levine, S. (2021). Robust Predictable Control.. arXiv:2109.03214 [cs]., Comment: Project site with videos and code: https://beneysenbach.github.io/rpc.
  47. 47.Fan, J., & Li, W. (2021). DRIBO: Robust Deep Reinforcement Learning via Multi-View Information Bottleneck.. arXiv:2102.13268 [cs]., Comment: 27 pages.
  48. 48.Fan, L., Wang, G., Huang, D.-A., Yu, Z., Fei-Fei, L., Zhu, Y., & Anandkumar, A. (2021). SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies.. arXiv:2106.09678 [cs]., Comment: ICML 2021. Website: https://linxifan.github.io/secant-site/.
  49. 49.Farebrother, J., Machado, M. C., & Bowling, M. (2020). Generalization and Regularization in DQN.. arXiv:1810.00123 [cs, stat]., Comment: Earlier versions of this work were presented both at the NeurIPS’18 Deep Reinforcement Learning Workshop and the 4th Multidisciplinary Conference on Reinforcement Learning and Decision Making (RLDM’19).
  50. 50.Filos, A., Tigkas, P., McAllister, R., Rhinehart, N., Levine, S., & Gal, Y. (2020). Can autonomous vehicles identify, recover from, and adapt to distribution shifts?. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 3145–3153. PMLR.
  51. 51.Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In Precup, D., & Teh, Y. W. (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, Vol. 70 of Proceedings of Machine Learning Research, pp. 1126–1135. PMLR.
  52. 52.Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Badia, A. P., Buttimore, G., Deck, C., Leibo, J. Z., & Blundell, C. (2019). Generalization of reinforcement learners with working and episodic memory. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., & Garnett, R. (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 12448–12457.
  53. 53.François-Lavet, V., Bengio, Y., Precup, D., & Pineau, J. (2019). Combined reinforcement learning via abstract representations. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pp. 3582–3589. AAAI Press.
  54. 54.Fu, J., Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2021). D4RL: Datasets for Deep Data-Driven Reinforcement Learning.. arXiv:2004.07219 [cs, stat]., Comment: Website available at https://sites.google.com/view/d4rl/home.
  55. 55.Gamrian, S., & Goldberg, Y. (2019). Transfer learning for related reinforcement learning tasks via image-to-image translation. In Chaudhuri, K., & Salakhutdinov, R. (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, pp. 2063–2072. PMLR.
  56. 56.Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., & Levine, S. (2021). Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability.. arXiv:2107.06277 [cs, stat]., Comment: First two authors contributed equally.
  57. 57.Goel, S., Tatiya, G., Scheutz, M., & Sinapov, J. (2021). NovelGridworlds: A benchmark environment for detecting and adapting to novelties in open worlds..
  58. 58.Grigsby, J., & Qi, Y. (2020). Measuring Visual Generalization in Continuous Control from Pixels.. arXiv:2010.06740 [cs]., Comment: A total of 20 pages, 8 pages as the main text.
  59. 59.Gulcehre, C., Wang, Z., Novikov, A., Paine, T. L., Colmenarejo, S. G., Zolna, K., Agarwal, R., Merel, J., Mankowitz, D., Paduraru, C., Dulac-Arnold, G., Li, J., Norouzi, M., Hoffman, M., Nachum, O., Tucker, G., Heess, N., & de Freitas, N. (2021). RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning.. arXiv:2006.13888 [cs, stat]., Comment: NeurIPS paper. 21 pages including supplementary material, the github link for the datasets: https://github.com/deepmind/deepmind-research/rl_unplugged.
  60. 60.Hafner, D. (2021). Benchmarking the Spectrum of Agent Capabilities.. arXiv:2109.06780 [cs]., Comment: Website: https://danijar.com/crafter.
  61. 61.Hallak, A., Di Castro, D., & Mannor, S. (2015). Contextual Markov Decision Processes.. arXiv:1502.02259 [cs, stat]..
  62. 62.Han, I., Park, D.-H., & Kim, K.-J. (2021). A New Open-Source Off-Road Environment for Benchmark Generalization of Autonomous Driving.. IEEE Access..
  63. 63.Hansen, N., Jangir, R., Sun, Y., Alenyà, G., Abbeel, P., Efros, A. A., Pinto, L., & Wang, X. (2021a). Self-Supervised Policy Adaptation during Deployment.. arXiv:2007.04309 [cs, stat]., Comment: Website: https://nicklashansen.github.io/PAD/ Code: https://github.com/nicklashansen/policy-adaptation-during-deployment ICLR 2021.
  64. 64.Hansen, N., Su, H., & Wang, X. (2021b). Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation.. arXiv:2107.00644 [cs]., Comment: Code and videos are available at https://nicklashansen.github.io/SVEA.
  65. 65.Hansen, N., & Wang, X. (2021). Generalization in Reinforcement Learning by Soft Data Augmentation.. arXiv:2011.13389 [cs]., Comment: Website: https://nicklashansen.github.io/SODA/ Code: https://github.com/nicklashansen/dmcontrol-generalization-benchmark. Presented at International Conference on Robotics and Automation (ICRA) 2021.
  66. 66.Hao, B., Lattimore, T., Szepesvári, C., & Wang, M. (2021). Online sparse reinforcement learning. In Banerjee, A., & Fukumizu, K. (Eds.), The 24th International Conference on Artificial Intelligence and Statistics, AISTATS 2021, April 13-15, 2021, Virtual Event, Vol. 130 of Proceedings of Machine Learning Research, pp. 316–324. PMLR.
  67. 67.Harries, L., Lee, S., Rzepecki, J., Hofmann, K., & Devlin, S. (2019). MazeExplorer: A Customisable 3D Benchmark for Assessing Generalisation in Reinforcement Learning. In 2019 IEEE Conference on Games (CoG), pp. 1–4.
  68. 68.Harrison, J., Garg, A., Ivanovic, B., Zhu, Y., Savarese, S., Fei-Fei, L., & Pavone, M. (2017). ADAPT: Zero-Shot Adaptive Policy Transfer for Stochastic Dynamical Systems.. arXiv:1707.04674 [cs]., Comment: International Symposium on Robotics Research (ISRR), 2017.
  69. 69.Hausknecht, M. J., Ammanabrolu, P., Côté, M., & Yuan, X. (2020). Interactive fiction games: A colossal adventure. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pp. 7903–7910. AAAI Press.
  70. 70.Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. In McIlraith, S. A., & Weinberger, K. Q. (Eds.), Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pp. 3207–3214. AAAI Press.
  71. 71.Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., & Lerchner, A. (2017a). beta-vae: Learning basic visual concepts with a constrained variational framework. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  72. 72.Higgins, I., Pal, A., Rusu, A. A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., & Lerchner, A. (2017b). DARLA: improving zero-shot transfer in reinforcement learning. In Precup, D., & Teh, Y. W. (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, Vol. 70 of Proceedings of Machine Learning Research, pp. 1480–1490. PMLR.
  73. 73.Hill, F., Lampinen, A. K., Schneider, R., Clark, S., Botvinick, M., McClelland, J. L., & Santoro, A. (2020a). Environmental drivers of systematicity and generalization in a situated agent. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  74. 74.Hill, F., Mokra, S., Wong, N., & Harley, T. (2020b). Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text.. arXiv:2005.09382 [cs]..
  75. 75.Hu, H., Lerer, A., Cui, B., Wu, D., Pineda, L., Brown, N., & Foerster, J. (2021). Off-Belief Learning.. arXiv:2103.04000 [cs]..
  76. 76.Hu, H., Lerer, A., Peysakhovich, A., & Foerster, J. N. (2020). "other-play" for zero-shot coordination. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 4399–4410. PMLR.
  77. 77.Huang, B., Feng, F., Lu, C., Magliacane, S., & Zhang, K. (2021). AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning.. arXiv:2107.02729 [cs, stat]..
  78. 78.Hupkes, D., Dankers, V., Mul, M., & Bruni, E. (2020). Compositionality decomposed: How do neural networks generalise?.. arXiv:1908.08351 [cs, stat]..
  79. 79.Igl, M., Ciosek, K., Li, Y., Tschiatschek, S., Zhang, C., Devlin, S., & Hofmann, K. (2019). Generalization in reinforcement learning with selective noise injection and information bottleneck. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., & Garnett, R. (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 13956–13968.
  80. 80.Igl, M., Farquhar, G., Luketina, J., Boehmer, W., & Whiteson, S. (2021). Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning.. arXiv:2006.05826 [cs, stat]..
  81. 81.Irpan, A., & Song, X. (2019). The Principle of Unchanged Optimality in Reinforcement Learning Generalization.. arXiv:1906.00336 [cs, stat]., Comment: Published at ICML 2019 Workshop "Understanding and Improving Generalization in Deep Learning".
  82. 82.Jain, A., Szot, A., & Lim, J. J. (2020). Generalization to new actions in reinforcement learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 4661–4672. PMLR.
  83. 83.James, S., Ma, Z., Arrojo, D. R., & Davison, A. J. (2019a). RLBench: The Robot Learning Benchmark & Learning Environment.. arXiv:1909.12271 [cs]., Comment: Videos and code: https://sites.google.com/view/rlbench.
  84. 84.James, S., Wohlhart, P., Kalakrishnan, M., Kalashnikov, D., Irpan, A., Ibarz, J., Levine, S., Hadsell, R., & Bousmalis, K. (2019b). Sim-to-real via sim-to-sim: Data-efficient robotic grasping via randomized-to-canonical adaptation networks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pp. 12627–12637. Computer Vision Foundation / IEEE.
  85. 85.Jiang, M., Dennis, M., Parker-Holder, J., Foerster, J., Grefenstette, E., & Rocktäschel, T. (2021a). Replay-Guided Adversarial Environment Design.. arXiv:2110.02439 [cs]., Comment: NeurIPS 2021.
  86. 86.Jiang, M., Grefenstette, E., & Rocktäschel, T. (2021b). Prioritized Level Replay.. arXiv:2010.03934 [cs]..
  87. 87.Jiang, M., Luketina, J., Nardelli, N., Minervini, P., Torr, P. H., Whiteson, S., & Rocktäschel, T. (2020). WordCraft: An environment for benchmarking commonsense agents. In Workshop on Language in Reinforcement Learning (LaRel).
  88. 88.Johnson, M., Hofmann, K., Hutton, T., & Bignell, D. (2016). The malmo platform for artificial intelligence experimentation. In Kambhampati, S. (Ed.), Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, IJCAI 2016, New York, NY, USA, 9-15 July 2016, pp. 4246–4247. IJCAI/AAAI Press.
  89. 89.Juliani, A., Khalifa, A., Berges, V., Harper, J., Teng, E., Henry, H., Crespi, A., Togelius, J., & Lange, D. (2019). Obstacle tower: A generalization challenge in vision, control, and planning. In Kraus, S. (Ed.), Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pp. 2684–2691. ijcai.org.
  90. 90.Kanagawa, Y., & Kaneko, T. (2019). Rogue-Gym: A New Challenge for Generalization in Reinforcement Learning.. arXiv:1904.08129 [cs, stat]., Comment: 8 pages, 14 figures, 4 tables, accepted to IEEE COG 2019.
  91. 91.Kansky, K., Silver, T., Mély, D. A., Eldawy, M., Lázaro-Gredilla, M., Lou, X., Dorfman, N., Sidor, S., Phoenix, D. S., & George, D. (2017). Schema networks: Zero-shot transfer with a generative causal model of intuitive physics. In Precup, D., & Teh, Y. W. (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, Vol. 70 of Proceedings of Machine Learning Research, pp. 1809–1818. PMLR.
  92. 92.Katz, G., Huang, D. A., Ibeling, D., Julian, K., Lazarus, C., Lim, R., Shah, P., Thakoor, S., Wu, H., Zeljić, A., Dill, D. L., Kochenderfer, M. J., & Barrett, C. (2019). The Marabou Framework for Verification and Analysis of Deep Neural Networks. In Dillig, I., & Tasiran, S. (Eds.), Computer Aided Verification, Lecture Notes in Computer Science, pp. 443–452, Cham. Springer International Publishing.
  93. 93.Ke, N. R., Dasgupta, I., Chiappa, S., Goyal, A., Weber, T., Mitrovic, J., Hill, F., Chan, S. C. Y., Mozer, M. C., Rezende, D. J., & Kohli, P. (2021). Parametric Generalization for Benchmarking Reinforcement Learning Algorithms..
  94. 94.Kemertas, M., & Aumentado-Armstrong, T. (2021). Towards Robust Bisimulation Metric Learning.. arXiv:2110.14096 [cs]., Comment: Accepted to NeurIPS 2021.
  95. 95.Kempka, M., Wydmuch, M., Runc, G., Toczek, J., & Jaśkowski, W. (2016). ViZDoom: A Doom-based AI research platform for visual reinforcement learning. In IEEE Conference on Computational Intelligence and Games, pp. 341–348, Santorini, Greece. IEEE. The best paper award.
  96. 96.Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., Tsarkov, D., Wang, X., van Zee, M., & Bousquet, O. (2020). Measuring compositional generalization: A comprehensive method on realistic data. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  97. 97.Khetarpal, K., Riemer, M., Rish, I., & Precup, D. (2020). Towards Continual Reinforcement Learning: A Review and Perspectives.. arXiv:2012.13490 [cs]., Comment: Preprint, 52 pages, 8 figures.
  98. 98.Ko, B., & Ok, J. (2021). Time Matters in Using Data Augmentation for Vision-based Deep Reinforcement Learning.. arXiv:2102.08581 [cs]..
  99. 99.Kostrikov, I., Yarats, D., & Fergus, R. (2021). Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels.. arXiv:2004.13649 [cs, eess, stat]..
  100. 100.Koutras, D. I., Kapoutsis, A. C., Amanatiadis, A. A., & Kosmatopoulos, E. B. (2021). MarsExplorer: Exploration of Unknown Terrains via Deep Reinforcement Learning and Procedurally Generated Environments.. arXiv:2107.09996 [cs]..
  101. 101.Kumar, A., Fu, Z., Pathak, D., & Malik, J. (2021). RMA: Rapid Motor Adaptation for Legged Robots.. arXiv:2107.04034 [cs]., Comment: RSS 2021. Webpage at https://ashishkmr.github.io/rma-legged-robots/.
  102. 102.Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative q-learning for offline reinforcement learning. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., & Lin, H. (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  103. 103.Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., & Rocktäschel, T. (2020). The nethack learning environment. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., & Lin, H. (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  104. 104.Lee, K., Lee, K., Shin, J., & Lee, H. (2020). Network randomization: A simple technique for generalization in deep reinforcement learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  105. 105.Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.. arXiv:2005.01643 [cs, stat]..
  106. 106.Li, B., François-Lavet, V., Doan, T., & Pineau, J. (2021a). Domain Adversarial Reinforcement Learning.. arXiv:2102.07097 [cs]..
  107. 107.Li, Q., Peng, Z., Xue, Z., Zhang, Q., & Zhou, B. (2021b). MetaDrive: Composing Diverse Driving Scenarios for Generalizable Reinforcement Learning..
  108. 108.Liu, G. T., Cheng, P.-J., & Lin, G. (2020). Cross-State Self-Constraint for Feature Generalization in Deep Reinforcement Learning..
  109. 109.Lomonaco, V., Desai, K., Culurciello, E., & Maltoni, D. (2020). Continual Reinforcement Learning in 3D Non-stationary Environments.. arXiv:1905.10112 [cs, stat]., Comment: Accepted in the CLVision Workshop at CVPR2020: 13 pages, 4 figures, 5 tables.
  110. 110.Lu, X., Lee, K., Abbeel, P., & Tiomkin, S. (2020). Dynamics Generalization via Information Bottleneck in Deep Reinforcement Learning.. arXiv:2008.00614 [cs, stat]., Comment: 16 pages.
  111. 111.Luketina, J., Nardelli, N., Farquhar, G., Foerster, J. N., Andreas, J., Grefenstette, E., Whiteson, S., & Rocktäschel, T. (2019). A survey of reinforcement learning informed by natural language. In Kraus, S. (Ed.), Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pp. 6309–6317. ijcai.org.
  112. 112.Lynch, C., & Sermanet, P. (2021). Language Conditioned Imitation Learning over Unstructured Data.. arXiv:2005.07648 [cs]., Comment: Published at RSS 2021.
  113. 113.Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., & Bowling, M. (2017). Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents.. arXiv:1709.06009 [cs]..
  114. 114.Malik, D., Li, Y., & Ravikumar, P. (2021). When Is Generalizable Reinforcement Learning Tractable?.. arXiv:2101.00300 [cs, stat]., Comment: v2 extends results to function approximation setting.
  115. 115.Mankowitz, D. J., Levine, N., Jeong, R., Abdolmaleki, A., Springenberg, J. T., Shi, Y., Kay, J., Hester, T., Mann, T. A., & Riedmiller, M. A. (2020). Robust reinforcement learning for continuous control with model misspecification. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  116. 116.Mausam, D. S. W. (2003). Solving relational MDPs with first-order machine learning. In IN PROC. ICAPS WORKSHOP ON PLANNING UNDER UNCERTAINTY AND INCOMPLETE INFORMATION. Citeseer.
  117. 117.Mazoure, B., Ahmed, A. M., MacAlpine, P., Hjelm, R. D., & Kolobov, A. (2021). Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL.. arXiv:2106.02193 [cs]..
  118. 118.Milani, S., Topin, N., Veloso, M., & Fang, F. (2022). A Survey of Explainable Reinforcement Learning..
  119. 119.Mishra, N., Rohaninejad, M., Chen, X., & Abbeel, P. (2018). A simple neural attentive meta-learner. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  120. 120.Morimoto, J., & Doya, K. (2000). Robust reinforcement learning. In Leen, T. K., Dietterich, T. G., & Tresp, V. (Eds.), Advances in Neural Information Processing Systems 13, Papers from Neural Information Processing Systems (NIPS) 2000, Denver, CO, USA, pp. 1061–1067. MIT Press.
  121. 121.Müller-Brockhausen, M., Preuss, M., & Plaat, A. (2021). Procedural Content Generation: Better Benchmarks for Transfer Reinforcement Learning.. arXiv:2105.14780 [cs]..
  122. 122.Nagabandi, A., Clavera, I., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., & Finn, C. (2019a). Learning to adapt in dynamic, real-world environments through meta-reinforcement learning. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  123. 123.Nagabandi, A., Finn, C., & Levine, S. (2019b). Deep online learning via meta-learning: Continual adaptation for model-based RL. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  124. 124.Nair, A., Gupta, A., Dalal, M., & Levine, S. (2021). AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.. arXiv:2006.09359 [cs, stat]., Comment: 17 pages. Website: https://awacrl.github.io/.
  125. 125.Narvekar, S., Peng, B., Leonetti, M., Sinapov, J., Taylor, M. E., & Stone, P. (2020). Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey.. arXiv:2003.04960 [cs, stat]..
  126. 126.Ng, A. Y., & Russell, S. J. (2000). Algorithms for inverse reinforcement learning. In Langley, P. (Ed.), Proceedings of the Seventeenth International Conference on Machine Learning (ICML 2000), Stanford University, Stanford, CA, USA, June 29 - July 2, 2000, pp. 663–670. Morgan Kaufmann.
  127. 127.Ni, T., Eysenbach, B., & Salakhutdinov, R. (2021). Recurrent Model-Free RL is a Strong Baseline for Many POMDPs.. arXiv:2110.05038 [cs]..
  128. 128.OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., & Zhang, L. (2019a). Solving Rubik’s Cube with a Robot Hand.. arXiv:1910.07113 [cs, stat]..
  129. 129.OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., Pinto, H. P. d. O., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., & Zhang, S. (2019b). Dota 2 with Large Scale Deep Reinforcement Learning.. arXiv:1912.06680 [cs, stat]..
  130. 130.OpenAI, O. (2016). OpenAI Gym: The DoomTakeCover-v0 environment. https://gym.openai.com/envs/DoomTakeCover-v0.
  131. 131.Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvári, C., Singh, S., Roy, B. V., Sutton, R. S., Silver, D., & van Hasselt, H. (2020). Behaviour suite for reinforcement learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  132. 132.Osband, I., & Roy, B. V. (2014). Near-optimal reinforcement learning in factored mdps. In Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N. D., & Weinberger, K. Q. (Eds.), Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, pp. 604–612.
  133. 133.Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., & Song, D. (2019). Assessing Generalization in Deep Reinforcement Learning.. arXiv:1810.12282 [cs, stat]., Comment: 17 pages, 6 figures.
  134. 134.Peng, X. B., Andrychowicz, M., Zaremba, W., & Abbeel, P. (2018). Sim-to-Real Transfer of Robotic Control with Dynamics Randomization.. 2018 IEEE International Conference on Robotics and Automation (ICRA)..
  135. 135.Perez, C., Such, F., & Karaletsos, T. (2020). Generalized Hidden Parameter MDPs:Transferable Model-Based RL in a Handful of Trials.. Proceedings of the AAAI Conference on Artificial Intelligence..
  136. 136.Perez-Liebana, D., Liu, J., Khalifa, A., Gaina, R. D., Togelius, J., & Lucas, S. M. (2019). General Video Game AI: A Multi-Track Framework for Evaluating Agents, Games and Content Generation Algorithms.. arXiv:1802.10363 [cs]., Comment: 20 pages, 1 figure, accepted by IEEE ToG.
  137. 137.Pinto, L., Davidson, J., Sukthankar, R., & Gupta, A. (2017). Robust adversarial reinforcement learning. In Precup, D., & Teh, Y. W. (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, Vol. 70 of Proceedings of Machine Learning Research, pp. 2817–2826. PMLR.
  138. 138.Portelas, R., Colas, C., Weng, L., Hofmann, K., & Oudeyer, P. (2020). Automatic curriculum learning for deep RL: A short survey. In Bessiere, C. (Ed.), Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020, pp. 4819–4825. ijcai.org.
  139. 139.Powers, S., Xing, E., Kolve, E., Mottaghi, R., & Gupta, A. (2021). CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents.. arXiv:2110.10067 [cs]., Comment: Repository available at https://github.com/AGILabs/continual_rl.
  140. 140.Racanière, S., Weber, T., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P. W., Hassabis, D., Silver, D., & Wierstra, D. (2017). Imagination-augmented agents for deep reinforcement learning. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., & Garnett, R. (Eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5690–5701.
  141. 141.Raileanu, R., & Fergus, R. (2021). Decoupling Value and Policy for Generalization in Reinforcement Learning.. arXiv:2102.10330 [cs]..
  142. 142.Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., & Fergus, R. (2021). Automatic Data Augmentation for Generalization in Deep Reinforcement Learning.. arXiv:2006.12862 [cs]..
  143. 143.Rajan, R., Diaz, J. L. B., Guttikonda, S., Ferreira, F., Biedenkapp, A., von Hartz, J. O., & Hutter, F. (2021). MDP Playground: A Design and Debug Testbed for Reinforcement Learning.. arXiv:1909.07750 [cs, stat]., Comment: NeurIPS 2021 Data and Benchmark Track submission (with slight formatting differences, most notably citation style).
  144. 144.Ren, Y., Duan, J., Li, S. E., Guan, Y., & Sun, Q. (2020). Improving Generalization of Reinforcement Learning with Minimax Distributional Soft Actor-Critic.. arXiv:2002.05502 [cs, stat]..
  145. 145.Risi, S., & Togelius, J. (2020). Increasing generality in machine learning through procedural content generation.. Nat Mach Intell..
  146. 146.Sadeghi, F., & Levine, S. (2017). CAD2RL: Real Single-Image Flight without a Single Real Image.. arXiv:1611.04201 [cs]., Comment: To appear at Robotics: Science and Systems Conference (R:SS), 2017. Supplementary video: https://www.youtube.com/watch?v=nXBWmzFrj5s.
  147. 147.Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Kuttler, H., Grefenstette, E., & Rocktäschel, T. (2021). MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research. In Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  148. 148.Schmeckpeper, K., Rybkin, O., Daniilidis, K., Levine, S., & Finn, C. (2020). Reinforcement Learning with Videos: Combining Offline Observations with Interaction.. arXiv:2011.06507 [cs]..
  149. 149.Schölkopf, B. (2019). Causality for Machine Learning.. arXiv:1911.10500 [cs, stat]..
  150. 150.Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., & Silver, D. (2020). Mastering Atari, Go, chess and shogi by planning with a learned model.. Nature..
  151. 151.Schrittwieser, J., Hubert, T., Mandhane, A., Barekatain, M., Antonoglou, I., & Silver, D. (2021). Online and Offline Reinforcement Learning by Planning with a Learned Model.. arXiv:2104.06294 [cs]..
  152. 152.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms.. arXiv:1707.06347 [cs]..
  153. 153.Seo, Y., Lee, K., Clavera, I., Kurutach, T., Shin, J., & Abbeel, P. (2020). Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement Learning.. arXiv:2010.13303 [cs]., Comment: Accepted in NeurIPS2020. First two authors contributed equally, website: https://sites.google.com/view/trajectory-mcl code: https://github.com/younggyoseo/trajectory_mcl.
  154. 154.Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., & Levine, S. (2018). Time-Contrastive Networks: Self-Supervised Learning from Video.. arXiv:1704.06888 [cs]..
  155. 155.Sestini, A., Kuhnle, A., & Bagdanov, A. D. (2020). Demonstration-efficient Inverse Reinforcement Learning in Procedurally Generated Environments.. arXiv:2012.02527 [cs]., Comment: Presented at the AAAI-21 Workshop on Reinforcement Learning in Games.
  156. 156.Shapley, L. S. (1953). Stochastic Games*.. Proceedings of the National Academy of Sciences..
  157. 157.Shen, B., Xia, F., Li, C., Martín-Martín, R., Fan, L., Wang, G., Pérez-D’Arpino, C., Buch, S., Srivastava, S., Tchapmi, L. P., Tchapmi, M. E., Vainio, K., Wong, J., Fei-Fei, L., & Savarese, S. (2021). iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes.. arXiv:2012.02924 [cs]..
  158. 158.Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on Image Data Augmentation for Deep Learning.. J Big Data..
  159. 159.Singh, J., & Zheng, L. (2021). Sparse Attention Guided Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning.. arXiv:2102.07266 [cs, stat]..
  160. 160.Sodhani, S., Meier, F., Pineau, J., & Zhang, A. (2022). Block Contextual MDPs for Continual Learning. In Proceedings of The 4th Annual Learning for Dynamics and Control Conference, pp. 608–623. PMLR.
  161. 161.Sonar, A., Pacelli, V., & Majumdar, A. (2020). Invariant Policy Optimization: Towards Stronger Generalization in Reinforcement Learning.. arXiv:2006.01096 [cs, stat]., Comment: 16 pages, 4 figures.
  162. 162.Song, X., Jiang, Y., Tu, S., Du, Y., & Neyshabur, B. (2020). Observational overfitting in reinforcement learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  163. 163.Stone, A., Ramirez, O., Konolige, K., & Jonschkowski, R. (2021). The Distracting Control Suite – A Challenging Benchmark for Reinforcement Learning from Pixels.. arXiv:2101.02722 [cs]., Comment: Code available at https://github.com/google-research/google-research/tree/master/distracting_control.
  164. 164.Strehl, A. L., Diuk, C., & Littman, M. L. (2007). Efficient structure learning in factored-state MDPs. In Proceedings of the 22nd National Conference on Artificial Intelligence - Volume 1, AAAI’07, pp. 645–650, Vancouver, British Columbia, Canada. AAAI Press.
  165. 165.Tachet, R., Bachman, P., & van Seijen, H. (2020). Learning Invariances for Policy Generalization.. arXiv:1809.02591 [cs, stat]., Comment: 7 pages, 1 figure.
  166. 166.Tang, Y., & Ha, D. (2021). The Sensory Neuron as a Transformer: Permutation-Invariant Neural Networks for Reinforcement Learning.. arXiv:2109.02869 [cs]..
  167. 167.Tang, Y., Nguyen, D., & Ha, D. (2020). Neuroevolution of Self-Interpretable Agents.. Proceedings of the 2020 Genetic and Evolutionary Computation Conference., Comment: To appear at the Genetic and Evolutionary Computation Conference (GECCO 2020) as a full paper.
  168. 168.Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., & Riedmiller, M. (2018). DeepMind Control Suite.. arXiv:1801.00690 [cs]., Comment: 24 pages, 7 figures, 2 tables.
  169. 169.Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., & Czarnecki, W. M. (2021). Open-Ended Learning Leads to Generally Capable Agents.. arXiv:2107.12808 [cs]..
  170. 170.Tishby, N., & Zaslavsky, N. (2015). Deep learning and the information bottleneck principle. In 2015 IEEE Information Theory Workshop (ITW), pp. 1–5.
  171. 171.Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., & Abbeel, P. (2017). Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World.. arXiv:1703.06907 [cs]., Comment: 8 pages, 7 figures. Submitted to 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2017).
  172. 172.Tosch, E., Clary, K., Foley, J., & Jensen, D. (2019). Toybox: A Suite of Environments for Experimental Evaluation of Deep Reinforcement Learning.. arXiv:1905.02825 [cs, stat]..
  173. 173.Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Apps, C., & Silver, D. (2019). Grandmaster level in StarCraft II using multi-agent reinforcement learning.. Nature..
  174. 174.Vithayathil Varghese, N., & Mahmoud, Q. H. (2020). A survey of multi-task deep reinforcement learning.. Electronics..
  175. 175.Vlastelica, M., Rolínek, M., & Martius, G. (2021). Neuro-algorithmic Policies enable Fast Combinatorial Generalization.. arXiv:2102.07456 [cs]., Comment: 15 pages.
  176. 176.Wang, J. X., King, M., Porcel, N., Kurth-Nelson, Z., Zhu, T., Deck, C., Choy, P., Cassin, M., Reynolds, M., Song, F., Buttimore, G., Reichert, D. P., Rabinowitz, N., Matthey, L., Hassabis, D., Lerchner, A., & Botvinick, M. (2021). Alchemy: A structured task distribution for meta-reinforcement learning.. arXiv:2102.02926 [cs]., Comment: 16 pages, 9 figures.
  177. 177.Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., & Botvinick, M. (2017). Learning to reinforcement learn.. arXiv:1611.05763 [cs, stat]., Comment: 17 pages, 7 figures, 1 table.
  178. 178.Wang, K., Kang, B., Shao, J., & Feng, J. (2020). Improving Generalization in Reinforcement Learning with Mixture Regularization.. arXiv:2010.10814 [cs, stat]., Comment: NeurIPS 2020.
  179. 179.Wang, R., Lehman, J., Clune, J., & Stanley, K. O. (2019). Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.. arXiv:1901.01753 [cs]., Comment: 28 pages, 9 figures.
  180. 180.Wang, R., Lehman, J., Rawal, A., Zhi, J., Li, Y., Clune, J., & Stanley, K. O. (2020). Enhanced POET: open-ended reinforcement learning through unbounded invention of learning challenges and their solutions. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 9940–9951. PMLR.
  181. 181.Wang, X., Lian, L., & Yu, S. X. (2021). Unsupervised Visual Attention and Invariance for Reinforcement Learning.. arXiv:2104.02921 [cs]., Comment: Accepted at CVPR 2021.
  182. 182.Wellmer, Z., & Kwok, J. T. (2021). Dropout’s Dream Land: Generalization from Learned Simulators to Reality. In Oliver, N., Pérez-Cruz, F., Kramer, S., Read, J., & Lozano, J. A. (Eds.), Machine Learning and Knowledge Discovery in Databases. Research Track, Lecture Notes in Computer Science, pp. 255–270, Cham. Springer International Publishing.
  183. 183.Wenke, S., Saunders, D., Qiu, M., & Fleming, J. (2019). Reasoning and Generalization in RL: A Tool Use Perspective.. arXiv:1907.02050 [cs]., Comment: 13 pages, 5 figures.
  184. 184.Whiteson, S., Tanner, B., Taylor, M. E., & Stone, P. (2011). Protecting against evaluation overfitting in empirical reinforcement learning. In 2011 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL), pp. 120–127.
  185. 185.Wolpert, D., & Macready, W. (1997). No free lunch theorems for optimization.. IEEE Transactions on Evolutionary Computation..
  186. 186.Xie, S., Ma, X., Yu, P., Zhu, Y., Wu, Y. N., & Zhu, S.-C. (2021). HALMA: Humanlike Abstraction Learning Meets Affordance in Rapid Problem Solving.. arXiv:2102.11344 [cs]..
  187. 187.Xing, E., Gupta, A., Powers, S., & Dean, V. (2021a). Evaluating Generalization of Policy Learning Under Domain Shifts..
  188. 188.Xing, E., Gupta, A., Powers, S., & Dean, V. (2021b). KitchenShift: Evaluating Zero-Shot Generalization of Imitation-Based Policy Learning Under Domain Shifts. In NeurIPS 2021 Workshop on Distribution Shifts: Connecting Methods and Applications.
  189. 189.Xue, C., Pinto, V., Gamage, C., Nikonova, E., Zhang, P., & Renz, J. (2021). Phy-Q: A Benchmark for Physical Reasoning.. arXiv:2108.13696 [cs]., Comment: For the associated website, see https://github.com/phy-q/benchmark.
  190. 190.Yang, R., Xu, H., Wu, Y., & Wang, X. (2020a). Multi-Task Reinforcement Learning with Soft Modularization.. arXiv:2003.13661 [cs, stat]., Comment: Our project page: https://rchalyang.github.io/SoftModule.
  191. 191.Yang, Y., Cer, D., Ahmad, A., Guo, M., Law, J., Constant, N., Hernandez Abrego, G., Yuan, S., Tar, C., Sung, Y.-h., Strope, B., & Kurzweil, R. (2020b). Multilingual universal sentence encoder for semantic retrieval. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 87–94, Online. Association for Computational Linguistics.
  192. 192.Yen-Chen, L., Bauza, M., & Isola, P. (2019). Experience-Embedded Visual Foresight.. arXiv:1911.05071 [cs]., Comment: CoRL 2019. Project website: http://yenchenlin.me/evf/.
  193. 193.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., & Levine, S. (2019). Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.. arXiv preprint arXiv:1910.10897..
  194. 194.Yu, W., Tan, J., Liu, C. K., & Turk, G. (2017). Preparing for the Unknown: Learning a Universal Policy with Online System Identification.. arXiv:1702.02453 [cs]., Comment: Accepted as a conference paper at RSS 2017.
  195. 195.Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., Lockhart, E., Shanahan, M., Langston, V., Pascanu, R., Botvinick, M., Vinyals, O., & Battaglia, P. (2018). Relational Deep Reinforcement Learning.. arXiv:1806.01830 [cs, stat]..
  196. 196.Zambaldi, V. F., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D. P., Lillicrap, T. P., Lockhart, E., Shanahan, M., Langston, V., Pascanu, R., Botvinick, M., Vinyals, O., & Battaglia, P. W. (2019). Deep reinforcement learning with relational inductive biases. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  197. 197.Zhang, A., Ballas, N., & Pineau, J. (2018). A Dissection of Overfitting and Generalization in Continuous Reinforcement Learning.. arXiv:1806.07937 [cs, stat]., Comment: 20 pages, 16 figures.
  198. 198.Zhang, A., Lyle, C., Sodhani, S., Filos, A., Kwiatkowska, M., Pineau, J., Gal, Y., & Precup, D. (2020). Invariant causal prediction for block mdps. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, pp. 11214–11224. PMLR.
  199. 199.Zhang, A., McAllister, R., Calandra, R., Gal, Y., & Levine, S. (2021). Learning Invariant Representations for Reinforcement Learning without Reconstruction.. arXiv:2006.10742 [cs, stat]., Comment: Accepted as an oral at ICLR 2021.
  200. 200.Zhang, A., Wu, Y., & Pineau, J. (2018a). Natural Environment Benchmarks for Reinforcement Learning.. arXiv:1811.06032 [cs, stat]., Comment: 12 figures.
  201. 201.Zhang, C., Vinyals, O., Munos, R., & Bengio, S. (2018b). A Study on Overfitting in Deep Reinforcement Learning.. arXiv:1804.06893 [cs, stat]..
  202. 202.Zhang, H., Cissé, M., Dauphin, Y. N., & Lopez-Paz, D. (2018). mixup: Beyond empirical risk minimization. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  203. 203.Zhang, H., & Guo, Y. (2021). Generalization of Reinforcement Learning with Policy-Aware Adversarial Data Augmentation.. arXiv:2106.15587 [cs]..
  204. 204.Zhao, C., & Hospedales, T. (2020). Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation.. arXiv:2012.04839 [cs]..
  205. 205.Zhao, C., Sigaud, O., Stulp, F., & Hospedales, T. M. (2019). Investigating Generalisation in Continuous Deep Reinforcement Learning.. arXiv:1902.07015 [cs, stat]..
  206. 206.Zhao, W., Queralta, J. P., & Westerlund, T. (2020). Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: A Survey. In 2020 IEEE Symposium Series on Computational Intelligence (SSCI), pp. 737–744.
  207. 207.Zhong, V., Rocktäschel, T., & Grefenstette, E. (2021). RTFM: Generalising to Novel Environment Dynamics via Reading.. arXiv:1910.08210 [cs]., Comment: ICLR 2020; 17 pages, 13 figures.
  208. 208.Zhou, K., Liu, Z., Qiao, Y., Xiang, T., & Loy, C. C. (2022). Domain Generalization: A Survey.. IEEE Trans. Pattern Anal. Mach. Intell.., Comment: IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2022.
  209. 209.Zhou, K., Yang, Y., Qiao, Y., & Xiang, T. (2021). Domain Generalization with MixStyle.. arXiv:2104.02008 [cs]., Comment: ICLR 2021; Code is available at https://github.com/KaiyangZhou/mixstyle-release.
  210. 210.Zhu, Y., Wong, J., Mandlekar, A., & Martín-Martín, R. (2020). Robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.. arXiv:2009.12293 [cs]., Comment: For more information, please visit https://robosuite.ai.
  211. 211.Zhu, Z., Lin, K., & Zhou, J. (2021). Transfer Learning in Deep Reinforcement Learning: A Survey.. arXiv:2009.07888 [cs, stat]..
  212. 212.Zintgraf, L., Feng, L., Lu, C., Igl, M., Hartikainen, K., Hofmann, K., & Whiteson, S. (2021). Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning.. arXiv:2010.01062 [cs, stat]., Comment: Published at the International Conference on Machine Learning (ICML) 2021.
  213. 213.Zintgraf, L. M., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., & Whiteson, S. (2020). Varibad: A very good method for bayes-adaptive deep RL via meta-learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.

Citation

MLA
Kirk, R., et al. “A Survey of Zero-shot Generalisation in Deep Reinforcement Learning”. Journal of Artificial Intelligence Research, vol. 76, 2023, pp. 201–64, https://doi.org/10.1613/jair.1.14174.
APA
Kirk, R., Zhang, A., Grefenstette, E., & Rocktäschel, T. (2023). A Survey of Zero-shot Generalisation in Deep Reinforcement Learning. Journal of Artificial Intelligence Research, 76, 201–264. https://doi.org/10.1613/jair.1.14174
Chicago
Kirk, R., A. Zhang, E. Grefenstette, and T. Rocktäschel. 2023. “A Survey of Zero-shot Generalisation in Deep Reinforcement Learning”. Journal of Artificial Intelligence Research 76: 201–64. https://doi.org/10.1613/jair.1.14174.
Harvard
Kirk, R. et al. (2023) “A Survey of Zero-shot Generalisation in Deep Reinforcement Learning”, Journal of Artificial Intelligence Research, 76, pp. 201–264. Available at: https://doi.org/10.1613/jair.1.14174.
Vancouver
1. Kirk R, Zhang A, Grefenstette E, Rocktäschel T (2023) A Survey of Zero-shot Generalisation in Deep Reinforcement Learning. Journal of Artificial Intelligence Research 76:201–264

BibTeX

@article{Kirk_2023, title={A Survey of Zero-shot Generalisation in Deep Reinforcement Learning}, volume={76}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/jair.1.14174}, DOI={10.1613/jair.1.14174}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Kirk, Robert and Zhang, Amy and Grefenstette, Edward and Rocktäschel, Tim}, year={2023}, month=Jan, pages={201–264} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/