A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
Robert KirkAmy ZhangEdward GrefenstetteTim Rocktäschel
Presents a unifying mathematical formalism and taxonomy for zero-shot generalization in deep reinforcement learning, critically evaluating existing benchmarks and methods to guide the development of policies that successfully transfer to unseen environments.
Reinforcement learning offers strong potential for automation across critical applications such as autonomous vehicles, robotics, and healthcare systems. However, traditional algorithms are evaluated in the exact environments where they were trained, leading to severe performance drops when faced with new, unfamiliar conditions. Because direct trial-and-error training in real-world environments is often unsafe and expensive, systems must be able to perform reliably upon deployment without additional training. The article evaluates the current landscape of zero-shot generalisation in reinforcement learning, establishing a unified framework to categorize existing benchmarks and algorithmic solutions across diverse deployment challenges.
The authors conducted a comprehensive review and structural analysis covering 55 simulation environments and dozens of modern algorithms. Using an extended contextual decision-making model, the analysis classifies environments by four main operational variables: state layout, visual observation, physical dynamics, and reward functions. It also organizes evaluation protocols based on whether tests require interpolating within familiar ranges or extrapolating to unseen conditions, and it categorizes methods by how they modify training data, adjust network architectures, or alter optimization objectives.
The analysis reveals several critical findings across the field. First, existing research is heavily skewed toward visual and spatial changes—with spatial variation present in roughly 76% of surveyed environments and visual variation in 53%—while physical dynamics (35%) and goal or reward variation (36%) remain significantly underrepresented. Second, standard benchmarks that rely purely on randomized procedural generation hide the underlying factors of variation, preventing precise diagnosis of failure modes. Third, most existing solutions attempt to improve robustness solely through adjusted loss functions, leaving model architecture improvements and within-episode adaptation largely underutilized. Finally, out-of-distribution performance cannot be achieved generically; successful transfer strictly requires problem-specific structural assumptions and tailored inductive biases.
These findings mean that current benchmarks may provide misleading confidence for real-world readiness. High performance on standard visual benchmarks does not translate to resilience against changes in real-world physics, equipment wear, or shifting operational objectives. For leadership, relying on broad generalisation claims without verifying the exact type of variation creates safety, compliance, and financial risks, as unseen operational shifts can lead to sudden system failures.
Organizations developing these systems should adopt benchmark environments that combine procedural variation with explicit, controllable parameters rather than relying on pure random generation. Teams should prioritize evaluating algorithms across multiple distinct operational dimensions—such as context efficiency, offline data learning, and task variations—using multidimensional scorecards rather than single performance metrics. While the article provides high confidence regarding the structural limitations of current benchmarks and algorithms, it focuses primarily on empirical literature rather than theoretical mathematical bounds. Decision-makers should exercise caution and require targeted verification protocols before deploying autonomous policies into safety-critical operations.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). This comprehensive survey outlines core concepts, taxonomies, and methods for generalizing to unseen domains that the survey builds upon for the reinforcement learning setting.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). It provides foundational taxonomies and evaluation methodologies for domain generalization across unseen environments that inform the framing of zero-shot generalization.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). It delivers critical benchmark insights and empirical standards for evaluating generalization to out-of-distribution test domains without overfitting.
- Paper: Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning, Tianhe Yu et al. (2019). It introduces standard multi-task and meta-RL benchmarks that underpin discussions on task generalization and procedural environment design.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). It establishes key principles and challenges in offline reinforcement learning, directly preparing readers for the survey's discussion on offline RL zero-shot generalization.
- Paper: Transfer Learning for Reinforcement Learning Domains: A Survey, Matthew E. Taylor et al. (2009). It formalizes foundational knowledge transfer and multi-task taxonomies that precede modern zero-shot generalization in reinforcement learning.
- Paper: Deep Surrogate Assisted Generation of Environments, Varun Bhatt et al. (2022). It explores automated generation and evaluation of diverse RL environments, contextualizing the survey's critiques of procedural content generation in benchmark design.
- Paper: An Introduction to Deep Reinforcement Learning, Vincent François-Lavet et al. (2018). It outlines core deep reinforcement learning algorithms and basic generalization strategies, establishing necessary technical background for the survey.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). It empirically investigates how reinforcement learning generalizes to novel rules and visual variants out-of-distribution compared to supervised learning in foundation models.
- Paper: RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation, Yufei Wang et al. (2024). It realizes automated generation of diverse simulation environments and tasks using generative models, advancing beyond traditional procedural generation.
- Paper: METRA: Scalable Unsupervised RL with Metric-Aware Abstraction, Seohong Park et al. (2024). It develops unsupervised representation and skill learning that enables zero-shot downstream goal-reaching and fast adaptation across environments.
- Paper: Is Value Learning Really the Main Bottleneck in Offline RL?, Seohong Park et al. (2024). It systematically investigates test-time generalization bottlenecks in offline reinforcement learning, directly addressing areas highlighted by the survey.
- Paper: LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, Bo Liu et al. (2023). It introduces a benchmark isolating distribution shifts and procedural knowledge transfer in robotic manipulation.
- Paper: Diffusion Actor-Critic with Entropy Regulator, Yinuo Wang et al. (2024). It presents a diffusion-based online policy framework that improves exploration and generalization in high-dimensional continuous control.
