Multi-Game Decision Transformers
Kuang-Huei LeeOfir NachumMengjiao YangLisa LeeDaniel FreemanSergio GuadarramaIan FischerWinnie XuEric JangHenryk Michalewski
Demonstrates that a single offline transformer model trained across diverse experience can master dozens of Atari games simultaneously while displaying predictable performance scaling and rapid fine-tuning to unseen tasks.
Developing generalist artificial intelligence agents that can successfully perform a wide variety of tasks across fundamentally different environments remains a major challenge. While natural language processing and computer vision have achieved general-purpose capabilities by training large transformer models on vast, diverse datasets, decision-making and control domains have traditionally relied on smaller, specialist models trained on single tasks.
The article evaluates whether the same scaling and pretraining strategies used in language and vision can produce a capable, generalist decision-making agent. Specifically, it demonstrates that a single transformer model with a single set of parameters can learn to play dozens of distinct Atari games from a diverse mix of previously collected expert and non-expert experience.
To conduct this evaluation, the researchers cast the problem of offline decision-making as a sequence prediction task. They trained an autoregressive transformer on an uncurated dataset spanning 41 Atari games, comprising 4.1 billion environment steps (nearly 160 billion tokens) recorded across beginner to expert skill levels. Image observations were divided into patches and tokenized alongside actions, clipped rewards, and target returns. Because the training data included suboptimal gameplay, the authors applied a guided generation technique during evaluation to sample high-return targets and select expert-level actions at runtime. They compared this Multi-Game Decision Transformer against traditional reinforcement learning algorithms, representation learning baselines, and behavioral cloning across various model sizes and transfer settings.
The findings show that the generalist model achieves 126% of human-level performance across all 41 games, outperforming the aggregate score of the training data (101%) and beating behavioral cloning in 31 out of 41 games. Furthermore, the model exhibited clear scaling trends: larger architectures consistently improved interactive gameplay performance and learned faster per token, whereas traditional reinforcement learning baselines suffered from instability or performance degradation when scaled. When adapted to five completely new, held-out games using only 1% of target game data, the pretrained model transferred rapidly and outperformed all baseline approaches.
These results indicate that sequence modeling transformers provide a scalable, highly stable framework for generalist agents, avoiding the instability often found in conventional reinforcement learning algorithms. Incorporating diverse non-expert data enhances world representation and performance over training solely on expert data, provided guided generation is used to steer action selection at inference time. This suggests that large-scale pretraining can dramatically lower the data and time costs required to deploy competent agents on new tasks.
Organizations developing decision-making agents should explore sequence modeling transformers over standard value-based reinforcement learning when working with large, diverse offline datasets. Before deploying these models in real-world or mission-critical settings, however, stakeholders should conduct targeted pilot tests, as the current evaluation relies on self-contained video games with shared visual tokenizations and discrete action spaces. While confidence is high that scaling laws apply to multi-game offline environments, further research is required to verify whether these capabilities transfer to continuous physical control, robotics, and complex real-world systems.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). This foundational paper introduces the Decision Transformer architecture that reframes offline reinforcement learning as autoregressive sequence modeling, providing the core algorithmic paradigm directly scaled to multi-game settings.
- Paper: The Arcade Learning Environment: An Evaluation Platform for General Agents, Marc G. Bellemare et al. (2013). This benchmark paper establishes the Arcade Learning Environment testbed used to evaluate multi-game generalist reinforcement learning capabilities across Atari 2600 environments.
- Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). This landmark study demonstrates human-level control from raw pixel inputs on Atari games using deep reinforcement learning, establishing the primary comparative baseline and evaluation suite.
- Paper: IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures, Lasse Espeholt et al. (2018). This work introduces scalable multi-task reinforcement learning across Atari suites, establishing key distributed learning principles and baselines for multi-game agent training.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). This paper analyzes gradient interference in multi-task reinforcement learning, addressing fundamental optimization challenges when training a single model across diverse task domains.
- Paper: Progressive Neural Networks, Andrei A. Rusu et al. (2016). This research provides foundational insights into sequential multi-task transfer and preventing catastrophic forgetting across distinct Atari environments.
- Paper: RT-1: Robotics Transformer for Real-World Control at Scale, Anthony Brohan et al. (2023). This paper extends large-scale transformer-based sequence modeling from multi-game virtual domains to real-world multi-task robotic control at scale.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). This work pushes transformer-based generalist agent policies further by translating vision-language pretraining directly into unified robotic action token generation.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). This study advances generalist multi-task agent modeling by integrating multi-modal perception and language models for embodied task planning and physical control.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). This project continues the pursuit of generalist transformer policies across diverse datasets and embodiments by combining open-source architectures with diffusion action decoders.
- Paper: Mastering Diverse Domains through World Models, Danijar Hafner et al. (2023). This work presents an alternative world-model approach to mastering diverse multi-task domains with a single set of fixed hyperparameters without domain-specific retuning.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). This research scales transformer-based action policy architectures to billion-parameter diffusion foundation models for multi-task physical manipulation.
- Paper: Code World Models for General Game Playing, Wolfgang Lehrach et al. (2026). This book explores code-based world modeling and planning to achieve general game-playing capabilities across diverse multi-game environments.
