Multi-Game Decision Transformers

Kuang-Huei LeeOfir NachumMengjiao YangLisa LeeDaniel FreemanSergio GuadarramaIan FischerWinnie XuEric JangHenryk Michalewski

article2022NeurIPS275 citations

Demonstrates that a single offline transformer model trained across diverse experience can master dozens of Atari games simultaneously while displaying predictable performance scaling and rapid fine-tuning to unseen tasks.

Listen

Developing generalist artificial intelligence agents that can successfully perform a wide variety of tasks across fundamentally different environments remains a major challenge. While natural language processing and computer vision have achieved general-purpose capabilities by training large transformer models on vast, diverse datasets, decision-making and control domains have traditionally relied on smaller, specialist models trained on single tasks.

The article evaluates whether the same scaling and pretraining strategies used in language and vision can produce a capable, generalist decision-making agent. Specifically, it demonstrates that a single transformer model with a single set of parameters can learn to play dozens of distinct Atari games from a diverse mix of previously collected expert and non-expert experience.

To conduct this evaluation, the researchers cast the problem of offline decision-making as a sequence prediction task. They trained an autoregressive transformer on an uncurated dataset spanning 41 Atari games, comprising 4.1 billion environment steps (nearly 160 billion tokens) recorded across beginner to expert skill levels. Image observations were divided into patches and tokenized alongside actions, clipped rewards, and target returns. Because the training data included suboptimal gameplay, the authors applied a guided generation technique during evaluation to sample high-return targets and select expert-level actions at runtime. They compared this Multi-Game Decision Transformer against traditional reinforcement learning algorithms, representation learning baselines, and behavioral cloning across various model sizes and transfer settings.

The findings show that the generalist model achieves 126% of human-level performance across all 41 games, outperforming the aggregate score of the training data (101%) and beating behavioral cloning in 31 out of 41 games. Furthermore, the model exhibited clear scaling trends: larger architectures consistently improved interactive gameplay performance and learned faster per token, whereas traditional reinforcement learning baselines suffered from instability or performance degradation when scaled. When adapted to five completely new, held-out games using only 1% of target game data, the pretrained model transferred rapidly and outperformed all baseline approaches.

These results indicate that sequence modeling transformers provide a scalable, highly stable framework for generalist agents, avoiding the instability often found in conventional reinforcement learning algorithms. Incorporating diverse non-expert data enhances world representation and performance over training solely on expert data, provided guided generation is used to steer action selection at inference time. This suggests that large-scale pretraining can dramatically lower the data and time costs required to deploy competent agents on new tasks.

Organizations developing decision-making agents should explore sequence modeling transformers over standard value-based reinforcement learning when working with large, diverse offline datasets. Before deploying these models in real-world or mission-critical settings, however, stakeholders should conduct targeted pilot tests, as the current evaluation relies on self-contained video games with shared visual tokenizations and discrete action spaces. While confidence is high that scaling laws apply to multi-game offline environments, further research is required to verify whether these capabilities transfer to continuous physical control, robotics, and complex real-world systems.

Cover for Multi-Game Decision Transformers

Abstract

A longstanding goal of the field of AI is a method for learning a highly capable, generalist agent from diverse experience. In the subfields of vision and language, this was largely achieved by scaling up transformer-based models and training them on large, diverse datasets. Motivated by this progress, we investigate whether the same strategy can be used to produce generalist reinforcement learning agents. Specifically, we show that a single transformer-based model – with a single set of weights – trained purely offline can play a suite of up to 46 Atari games simultaneously at close-to-human performance. When trained and evaluated appropriately, we find that the same trends observed in language and vision hold, including scaling of performance with model size and rapid adaptation to new games via fine-tuning. We compare several approaches in this multi-game setting, such as online and offline RL methods and behavioral cloning, and find that our Multi-Game Decision Transformer models offer the best scalability and performance. We release the pre-trained models and code to encourage further research in this direction.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Reinforcement Learning as Sequence Modeling
  • 3.2 Tokenization
  • 3.3 Training Dataset
  • 3.4 Expert Action Inference
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Baseline Methods
  • 4.3 How do different online and offline methods perform in the multi-game regime?
  • 4.4 How do different methods scale with model size?
  • 4.5 How effective are different methods at transfer to novel games?
  • 4.6 Does multi-game decision transformer improve upon training data?
  • 4.7 Does optimal action inference improve upon behavior cloning?
  • 4.8 Does training on expert and non-expert data bring benefits over expert-only training?
  • 5 Conclusion
  • Acknowledgements
  • References
  • Checklist

Knowls

  1. Knowl 1 — Multi-Game Decision Transformer Sequence Modeling Formulation

    model/method

    The Multi-Game Decision Transformer poses multi-environment offline reinforcement learning as an autoregressive next-token prediction problem over trajectory sequences. An episodic sequence is represented as:

    x=⟨…,o1t,…,oMt,R^t,at,rt,… ⟩x = \langle \dots, o_1^t, \dots, o_M^t, \hat{R}^t, a^t, r^t, \dots \rangle

    where tt denotes the environment time-step, o1t,…,oMto_1^t, \dots, o_M^t are M=36M = 36 visual patch tokens representing observation image oto^t (divided into a 6×66 \times 6 spatial grid of non-overlapping 14×1414 \times 14 pixel patches, each linearly projected and added to a learned spatial position embedding), R^t=∑k>trk\hat{R}^t = \sum_{k > t} r^k is the target future return from step tt onward, ata^t is the discrete action token, and rt∈{−1,0,+1}r^t \in \{-1, 0, +1\} is the clipped scalar reward token.

    Future returns R^t\hat{R}^t are discretized via uniform quantization into integer bins over the range {−20,…,100}\{-20, \dots, 100\} with bin width 1. The model is parametrized as a causal GPT-2 decoder trained with standard cross-entropy loss to predict the discrete return, action, and reward tokens given prior context. Observations are placed immediately before the return token R^t\hat{R}^t to enable the transformer to predict the return distribution Pθ(Rt∣o≤t,… )P_\theta(R^t \mid o_{\le t}, \dots) directly at inference time.

  2. Knowl 2 — Expert Action Inference via Classifier-Guided Return Sampling

    algorithm

    To generate expert-level behavior from a model trained on mixed-quality (non-expert and expert) offline data without modifying the training objective, the Multi-Game Decision Transformer applies a classifier-guided autoregressive inference procedure inspired by controlled language generation.

    Assuming a binary classifier P(expertt∣Rt,… )∝exp⁡(κ⋅Rt−RlowRhigh−Rlow)P(\text{expert}^t \mid R^t, \dots) \propto \exp\left(\kappa \cdot \frac{R^t - R_{\text{low}}}{R_{\text{high}} - R_{\text{low}}}\right) with inverse temperature κ=10\kappa = 10, return lower bound Rlow=−20R_{\text{low}} = -20, and upper bound Rhigh=100R_{\text{high}} = 100, the expert-conditioned return distribution is computed via Bayes' rule:

    log⁡P(Rt∣expertt,… )=log⁡Pθ(Rt∣… )+κ⋅Rt−RlowRhigh−Rlow+C\log P(R^t \mid \text{expert}^t, \dots) = \log P_\theta(R^t \mid \dots) + \kappa \cdot \frac{R^t - R_{\text{low}}}{R_{\text{high}} - R_{\text{low}}} + C

    where CC is a normalizing constant.

    Input: Model parameters θ\theta, current observation patches o1t,…,oMto_1^t, \dots, o_M^t, token history x<tx_{<t}, return bounds Rlow=−20,Rhigh=100R_{\text{low}} = -20, R_{\text{high}} = 100, inverse temperature κ=10\kappa = 10
    Output: Executed action ata^t
    1. Append observation patches to history: x≤ot=⟨x<t,o1t,…,oMt⟩x_{\le o^t} = \langle x_{<t}, o_1^t, \dots, o_M^t \rangle
    2. Compute unnormalized return logits from model: z(R)=logitsθ(R∣x≤ot)z(R) = \text{logits}_\theta(R \mid x_{\le o^t}) for each discrete return bin R∈{Rlow,…,Rhigh}R \in \{R_{\text{low}}, \dots, R_{\text{high}}\}
    3. Compute guidance bias: g(R)=κ⋅R−RlowRhigh−Rlowg(R) = \kappa \cdot \frac{R - R_{\text{low}}}{R_{\text{high}} - R_{\text{low}}}
    4. Compute guided return distribution: P(R∣expertt)=exp⁡(z(R)+g(R))∑R′exp⁡(z(R′)+g(R′))P(R \mid \text{expert}^t) = \frac{\exp(z(R) + g(R))}{\sum_{R'} \exp(z(R') + g(R'))}
    5. Sample target return: R^t∼P(R∣expertt)\hat{R}^t \sim P(R \mid \text{expert}^t)
    6. Append return token to context: x≤R^t=⟨x≤ot,R^t⟩x_{\le \hat{R}^t} = \langle x_{\le o^t}, \hat{R}^t \rangle
    7. Compute action probabilities: P(a∣x≤R^t)=softmax(logitsθ(a∣x≤R^t))P(a \mid x_{\le \hat{R}^t}) = \text{softmax}(\text{logits}_\theta(a \mid x_{\le \hat{R}^t}))
    8. Sample action: at∼P(a∣x≤R^t)a^t \sim P(a \mid x_{\le \hat{R}^t})
    9. return ata^t
  3. Knowl 3 — Multi-Game Offline Pretraining and Adaptation Setup

    experimental setup

    The Multi-Game Decision Transformer evaluation protocol assesses multi-task performance and transfer across Atari games in the Arcade Learning Environment (ALE):

    • Dataset: Fixed offline trajectories collected from the training progress of DQN agents across 46 Atari games. 41 games are used for pretraining (2 runs ×\times 50 checkpoints ×\times 1 million environment steps = 4.1 billion environment steps, totaling ~160 billion discrete tokens including beginner to expert data). 5 games are held out for evaluation of transfer: Alien, MsPacman, Pong, SpaceInvaders, and StarGunner.
    • Model Variants: GPT-2 causal transformer decoders scaled across three sizes: DT-10M (10 million parameters), DT-40M (40 million parameters), and DT-200M (200 million parameters). The sequence context window spans 4 environment frames (156 tokens).
    • Pretraining: Trained on TPUv4 using the LAMB optimizer, learning rate 3×10−43 \times 10^{-4}, linear warm-up of 4000 steps, no weight decay, gradient clipping at 1.0, β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, batch size 2048, for 10 million gradient steps with image augmentations.
    • Fine-Tuning: Pretrained models are fine-tuned per held-out game on 1% of the dataset (1 million steps uniformly sampled without filtering) for 100,000 steps using learning rate 10−410^{-4}, weight decay 10−210^{-2}, and batch size 256.
    • Evaluation Metrics: Human-Normalized Score HNS=score−scorerandomscorehuman−scorerandom\text{HNS} = \frac{\text{score} - \text{score}_{\text{random}}}{\text{score}_{\text{human}} - \text{score}_{\text{random}}}, DQN-normalized score, and aggregate Inter-Quartile Mean (IQM) across games.
  4. Knowl 4 — Aggregate Multi-Game Performance of Decision Transformers vs Baselines

    empirical result

    In interactive evaluation across 41 Atari training games, a single generalist Multi-Game Decision Transformer (DT-200M) achieves an aggregate Inter-Quartile Mean (IQM) human-normalized score of 126%, significantly outperforming all other multi-game generalist models:

    • Multi-Game Decision Transformer (Offline): 126% IQM
    • Multi-Game Behavioral Cloning Transformer (Offline): 68% IQM
    • Multi-Game C51 DQN with Impala CNN (Online): 66% IQM (68% median)
    • Multi-Game Conservative Q-Learning (CQL) with Impala CNN (Offline): 38% IQM

    For comparison, single-game specialist models achieve 144% IQM (Single-Game DQN, Online) and 135% IQM (Single-Game BCQ, Offline). The training dataset contains an average IQM human-normalized score of 101%, which Multi-Game DT surpasses by 25 percentage points.

  5. Knowl 5 — Model Parameter Scaling in Multi-Game Offline RL

    empirical result

    Multi-Game Decision Transformers display predictable performance scaling with parameter count on interactive in-game evaluations across 41 Atari games:

    • DT-10M: ~56% IQM human-normalized score
    • DT-40M: ~93% IQM human-normalized score
    • DT-200M: ~126% IQM human-normalized score

    Behavioral Cloning transformers also scale positively (from ~46% at 10M to ~68% at 200M), but plateau well below Decision Transformers.

    In contrast, offline Temporal Difference (TD) methods show inverse or flat scaling: Conservative Q-Learning (CQL) with Impala CNN architectures achieves ~38% IQM at 5M parameters and drops to ~35% IQM at 60M parameters on training games. When fine-tuned on novel games, CQL IQM drops sharply from ~46% at 5M parameters down to ~6% at 120M parameters, highlighting severe instability when scaling TD-based value networks in multi-environment offline RL.

  6. Knowl 6 — Few-Shot Adaptation to Novel Games via Pretrained Decision Transformers

    empirical result

    Pretraining on 41 Atari games enables rapid adaptation to 5 held-out games (Alien, MsPacman, Pong, SpaceInvaders, StarGunner) using only 1% (1M steps) of uncurated data per game fine-tuned for 100,000 steps:

    • DT Pretraining: Achieves the highest transfer performance across all 5 held-out games, scaling from ~72% DQN-normalized IQM at 10M parameters to ~91% IQM at 200M parameters.
    • Comparison to Non-Pretrained Baselines: DT pretraining and CQL pretraining significantly outperform CQL trained from scratch on 1% data and DT trained without pretraining.
    • Comparison to Representation Learning Baselines: State representation pretraining methods (Attentive Contrastive Learning [ACL], Masked BERT pretraining, Contrastive Predictive Coding [CPC]) fine-tuned with behavioral cloning underperform DT pretraining across all held-out tasks, indicating that learning task-agnostic state representations alone without joint action-return sequence modeling is insufficient for optimal transfer.
  7. Knowl 7 — Super-Expert Policy Synthesis from Suboptimal Offline Data

    empirical result

    When evaluating the top-3 interactive rollouts per game, the Multi-Game Decision Transformer with expert action inference exceeds the best individual demonstration present in the offline training dataset across multiple Atari games.

    Observed percentage improvements over the highest demonstration score in the dataset include:

    • Asterix: >300% improvement
    • Breakout: >150% improvement
    • Gopher: >150% improvement
    • Frostbite: >100% improvement
    • Seaquest: >100% improvement

    This demonstrates that sequence modeling combined with classifier-guided return sampling performs dynamic stitching and policy improvement across suboptimal trajectory fragments rather than simple imitation of the dataset's best trajectory.

  8. Knowl 8 — Effect of Mixed-Quality Offline Data on Decision Transformers vs Behavioral Cloning

    empirical result

    Comparing 40M-parameter models trained on either the full uncurated dataset (containing beginner, intermediate, and expert trajectories) or an expert-only dataset (filtered to the top 10% highest-return trajectories per game):

    • Behavioral Cloning (BC-40M): Training on expert-only data yields 79% IQM human-normalized score, whereas training on full mixed data drops BC performance to 54% IQM.
    • Decision Transformer (DT-40M): Training on full mixed data achieves 93% IQM, outperforming DT trained on expert-only data (89% IQM).
    • Cross-Model Comparison: DT-40M trained on the full mixed dataset (93% IQM) outperforms BC-40M trained on expert-only data (79% IQM), and DT outperforms BC on 31 out of 41 individual games.

    This confirms that non-expert trajectories provide diverse state coverage and transition structure that benefit Decision Transformer representation learning when paired with return conditioning, whereas standard behavioral cloning suffers from suboptimal trajectory pollution.

  9. Knowl 9 — Limitations of Multi-Game Decision Transformers

    limitation

    The methodology and empirical conclusions of Multi-Game Decision Transformers are subject to several explicit limitations:

    1. Environment Homogeneity: The evaluation is confined to the Atari suite, where observation spaces (standardized 2D pixel grids) and action spaces (discrete 18-directional joystick inputs) are structurally aligned across games. The approach does not address heterogeneous action spaces or disparate observation modalities found in diverse robotics or continuous control domains.
    2. Dataset Scale vs Web Pretraining: While Atari replay datasets span billions of steps, they remain substantially smaller and less diverse than web-scale text or vision corpora, making it uncertain whether reinforcement learning scaling laws continue indefinitely without saturating.
    3. Lack of Zero-Shot Generalization: The generalist agent cannot adapt zero-shot to unseen games without fine-tuning on domain-specific demonstration data.

Coverage note — All primary contributions from the main text are covered; qualitative attention analysis and minor network architecture variations detailed in the appendices are omitted.

References

  1. 1.Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. An optimistic perspective on offline reinforcement learning. In International Conference on Machine Learning, pages 104–114. PMLR, 2020.
  2. 2.Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare. Deep reinforcement learning at the edge of the statistical precipice. Advances in Neural Information Processing Systems, 34, 2021.
  3. 3.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022.
  4. 4.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022.
  5. 5.William H Alexander and Samuel J Gershman. Representation learning with reward prediction errors. arXiv preprint arXiv:2108.12402, 2021.
  6. 6.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6836–6846, 2021.
  7. 7.Igor Babuschkin, Kate Baumli, Alison Bell, Surya Bhupatiraju, Jake Bruce, Peter Buchlovsky, David Budden, Trevor Cai, Aidan Clark, Ivo Danihelka, Claudio Fantacci, Jonathan Godwin, Chris Jones, Tom Hennigan, Matteo Hessel, Steven Kapturowski, Thomas Keck, Iurii Kemaev, Michael King, Lena Martens, Vladimir Mikulik, Tamara Norman, John Quan, George Papamakarios, Roman Ring, Francisco Ruiz, Alvaro Sanchez, Rosalia Schneider, Eren Sezener, Stephen Spencer, Srivatsan Srinivasan, Wojciech Stokowiec, and Fabio Viola. The DeepMind JAX Ecosystem, 2020. URL http://github.com/deepmind.
  8. 8.Osbert Bastani, Yewen Pu, and Armando Solar-Lezama. Verifiable reinforcement learning via policy extraction. Advances in neural information processing systems, 31, 2018.
  9. 9.M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, jun 2013.
  10. 10.Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, 2013.
  11. 11.Marc G Bellemare, Will Dabney, and Rémi Munos. A distributional perspective on reinforcement learning. In International Conference on Machine Learning, pages 449–458. PMLR, 2017.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  13. 13.Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G Bellemare. Dopamine: A research framework for deep reinforcement learning. arXiv preprint arXiv:1812.06110, 2018.
  14. 14.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34, 2021.
  15. 15.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  16. 16.Giuseppe Cuccu, Julian Togelius, and Philippe Cudré-Mauroux. Playing atari with six neurons. arXiv preprint arXiv:1806.01363, 2018.
  17. 17.Rishabh Dabral, Ganesh Ramakrishnan, Preethi Jyothi, et al. Rudder: A cross lingual video and text retrieval dataset. arXiv preprint arXiv:2103.05457, 2021.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  19. 19.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  20. 20.Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al. Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. In International Conference on Machine Learning, pages 1407–1416. PMLR, 2018.
  21. 21.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pages 2052–2062. PMLR, 2019.
  22. 22.Hiroki Furuta, Tadashi Kozuno, Tatsuya Matsushima, Yutaka Matsuo, and Shixiang Shane Gu. Coadaptation of algorithmic and implementational innovations in inference-based deep reinforcement learning. Advances in neural information processing systems, 34:9828–9842, 2021.
  23. 23.Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu. Generalized decision transformer for offline hindsight information matching. arXiv preprint arXiv:2111.10364, 2021.
  24. 24.Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare. Deepmdp: Learning continuous latent space models for representation learning. In International Conference on Machine Learning, pages 2170–2179. PMLR, 2019.
  25. 25.Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al. Rl unplugged: A suite of benchmarks for offline reinforcement learning. Advances in Neural Information Processing Systems, 33:7248–7259, 2020.
  26. 26.Agrim Gupta, Silvio Savarese, Surya Ganguli, and Li Fei-Fei. Embodied intelligence via learning and evolution. Nature communications, 12(1):1–12, 2021.
  27. 27.Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. In International Conference on Learning Representations, 2019.
  28. 28.Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. In International conference on machine learning, pages 2555–2565. PMLR, 2019.
  29. 29.Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. In International Conference on Learning Representations, 2020.
  30. 30.Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver. Rainbow: Combining improvements in deep reinforcement learning. In Thirty-second AAAI conference on artificial intelligence, 2018.
  31. 31.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  32. 32.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019.
  33. 33.Wenlong Huang, Igor Mordatch, and Deepak Pathak. One policy to control them all: Shared modular policies for agent-agnostic control. In International Conference on Machine Learning, pages 4455–4464. PMLR, 2020.
  34. 34.Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, pages 991–1002. PMLR, 2022.
  35. 35.Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34, 2021.
  36. 36.Lukasz Kaiser, Aidan N Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit. One model to learn them all. arXiv preprint arXiv:1706.05137, 2017.
  37. 37.Dmitry Kalashnikov, Jacob Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman. Mt-opt: Continuous multi-task robotic reinforcement learning at scale. arXiv preprint arXiv:2104.08212, 2021.
  38. 38.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  39. 39.Hilbert J Kappen, Vicenç Gómez, and Manfred Opper. Optimal control as a graphical model inference problem. Machine learning, 87(2):159–182, 2012.
  40. 40.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367, 2020.
  41. 41.Aviral Kumar, Xue Bin Peng, and Sergey Levine. Reward-conditioned policies. arXiv preprint arXiv:1912.13465, 2019.
  42. 42.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. Advances in Neural Information Processing Systems, 33:1179–1191, 2020.
  43. 43.Vitaly Kurin, Maximilian Igl, Tim Rocktäschel, Wendelin Boehmer, and Shimon Whiteson. My body is a cage: the role of morphology in graph-based incompatible control. arXiv preprint arXiv:2010.01856, 2020.
  44. 44.Kuang-Huei Lee, Ian Fischer, Anthony Liu, Yijie Guo, Honglak Lee, John Canny, and Sergio Guadarrama. Predictive information accelerates learning in rl. Advances in Neural Information Processing Systems, 33: 11890–11901, 2020.
  45. 45.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  46. 46.Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. Pretrained transformers as universal computation engines. arXiv preprint arXiv:2103.05247, 2021.
  47. 47.Clare Lyle, Mark Rowland, Georg Ostrovski, and Will Dabney. On the effect of auxiliary tasks on representation dynamics. In International Conference on Artificial Intelligence and Statistics, pages 1–9. PMLR, 2021.
  48. 48.Corey Lynch and Pierre Sermanet. Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648, 2020.
  49. 49.Horia Mania, Aurelia Guy, and Benjamin Recht. Simple random search provides a competitive approach to reinforcement learning. arXiv preprint arXiv:1803.07055, 2018.
  50. 50.John McCarthy, Marvin L Minsky, Nathaniel Rochester, and Claude E Shannon. A proposal for the dartmouth summer research project on artificial intelligence, august 31, 1955. AI magazine, 27(4):12–12, 2006.
  51. 51.Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, and Deepak Pathak. Discovering and achieving goals via world models. Advances in Neural Information Processing Systems, 34, 2021.
  52. 52.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  53. 53.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015.
  54. 54.Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In International conference on machine learning, pages 1928–1937. PMLR, 2016.
  55. 55.Ofir Nachum and Mengjiao Yang. Provable representation learning for imitation with contrastive fourier features. Advances in Neural Information Processing Systems, 34, 2021.
  56. 56.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  57. 57.Pedro A Ortega, Markus Kunesch, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Joel Veness, Jonas Buchli, Jonas Degrave, Bilal Piot, Julien Perolat, et al. Shaking the foundations: delusions in sequence models for interaction and control. arXiv preprint arXiv:2110.10819, 2021.
  58. 58.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  59. 59.Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov. Actor-mimic: Deep multitask and transfer reinforcement learning. arXiv preprint arXiv:1511.06342, 2015.
  60. 60.Dean A Pomerleau. Efficient training of artificial neural networks for autonomous navigation. Neural computation, 3(1):88–97, 1991.
  61. 61.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  62. 62.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  63. 63.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  64. 64.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128, 2021.
  65. 65.Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas. A generalist agent, 2022. URL https://arxiv.org/abs/2205.06175.
  66. 66.Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. Can wikipedia help offline reinforcement learning? arXiv preprint arXiv:2201.12122, 2022.
  67. 67.Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell. Policy distillation. arXiv preprint arXiv:1511.06295, 2015.
  68. 68.Juergen Schmidhuber. Reinforcement learning upside down: Don’t predict rewards–just map them to actions. arXiv preprint arXiv:1912.02875, 2019.
  69. 69.Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al. Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588(7839):604–609, 2020.
  70. 70.Ross D Shachter. Probabilistic inference and influence diagrams. Operations research, 36(4):589–604, 1988.
  71. 71.Rupesh Kumar Srivastava, Pranav Shyam, Filipe Mutz, Wojciech Jaśkowski, and Jürgen Schmidhuber. Training agents using upside-down reinforcement learning. arXiv preprint arXiv:1912.02877, 2019.
  72. 72.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021.
  73. 73.Emanuel Todorov. Linearly-solvable markov decision problems. Advances in neural information processing systems, 19, 2006.
  74. 74.Marc Toussaint. Robot trajectory optimization using approximate inference. In Proceedings of the 26th annual international conference on machine learning, pages 1049–1056, 2009.
  75. 75.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34: 200–212, 2021.
  76. 76.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 6309–6318, 2017.
  77. 77.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  78. 78.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, 2021.
  79. 79.Kevin Yang and Dan Klein. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218, 2021.
  80. 80.Mengjiao Yang and Ofir Nachum. Representation matters: Offline pretraining for sequential decision making. In International Conference on Machine Learning, pages 11784–11794. PMLR, 2021.
  81. 81.Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training bert in 76 minutes. arXiv preprint arXiv:1904.00962, 2019.
  82. 82.Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, pages 1094–1100. PMLR, 2020.
  83. 83.Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine. Learning invariant representations for reinforcement learning without reconstruction. arXiv preprint arXiv:2006.10742, 2020.
  84. 84.Qinqing Zheng, Amy Zhang, and Aditya Grover. Online decision transformer. arXiv preprint arXiv:2202.05607, 2022.

Citation

MLA
Lee, K.-H., et al. “Multi-Game Decision Transformers”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 27921–36, https://proceedings.neurips.cc/paper_files/paper/2022/file/b2cac94f82928a85055987d9fd44753f-Paper-Conference.pdf.
APA
Lee, K.-H., Nachum, O., Yang, M. (Sherry) ., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., & Mordatch, I. (2022). Multi-Game Decision Transformers. Advances in Neural Information Processing Systems, 35, 27921–27936. https://proceedings.neurips.cc/paper_files/paper/2022/file/b2cac94f82928a85055987d9fd44753f-Paper-Conference.pdf
Chicago
Lee, K.-H., O. Nachum, M. (Sherry) . Yang, et al. 2022. “Multi-Game Decision Transformers”. Advances in Neural Information Processing Systems 35: 27921–36. https://proceedings.neurips.cc/paper_files/paper/2022/file/b2cac94f82928a85055987d9fd44753f-Paper-Conference.pdf.
Harvard
Lee, K.-H. et al. (2022) “Multi-Game Decision Transformers”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 27921–27936. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/b2cac94f82928a85055987d9fd44753f-Paper-Conference.pdf.
Vancouver
1. Lee K-H, Nachum O, Yang M (Sherry), et al (2022) Multi-Game Decision Transformers. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 27921–27936

BibTeX

@inproceedings{lee2022multi,
  title = {Multi-Game Decision Transformers},
  author = {Lee, Kuang-Huei and Nachum, Ofir and Yang, Mengjiao (Sherry) and Lee, Lisa and Freeman, Daniel and Guadarrama, Sergio and Fischer, Ian and Xu, Winnie and Jang, Eric and Michalewski, Henryk and Mordatch, Igor},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {27921-27936},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/b2cac94f82928a85055987d9fd44753f-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission