Supervised Pretraining Can Learn In-Context Reinforcement Learning

Jonathan LeeAnnie XieAldo PacchianoYash ChandakChelsea FinnOfir NachumEmma Brunskill

article2023NeurIPS170 citations

Demonstrates that supervised pretraining to predict optimal actions enables transformers to perform sample-efficient in-context reinforcement learning with provable regret guarantees by effectively executing Bayesian posterior sampling.

Listen

Modern decision-making systems in robotics, recommendation engines, and autonomous operations require artificial intelligence agents that can quickly adapt to new environments without costly retraining. While transformer architectures excel at in-context learning for language tasks, adapting them to sequential decision-making and reinforcement learning has remained challenging. Standard reinforcement learning methods often require extensive data collection or computationally prohibitive Bayesian updates, while existing decision-focused transformers struggle to generalize beyond their training demonstrations or properly balance risk and exploration.

The article introduces and evaluates the Decision-Pretrained Transformer, a supervised pretraining framework designed to give transformer models in-context reinforcement learning capabilities. The primary objective is to demonstrate that standard supervised pretraining on optimal action predictions allows a transformer to execute efficient decision-making strategies—both online and offline—across unseen tasks and environments without any test-time model weight updates.

To evaluate this framework, the authors conducted extensive computational simulations across multi-armed bandits, structured linear bandits, discrete grid navigation tasks, and high-dimensional 3D vision-based environments. The model was trained using a standard causal GPT-2 backbone on datasets of past interactions paired with optimal action labels across diverse pretraining tasks. The model was then evaluated on entirely new tasks in two modes: offline, where it selects actions given fixed historical datasets, and online, where it collects its own data sequentially to solve an unknown task from scratch. The authors compared this approach against established algorithms, including Upper Confidence Bound, Thompson Sampling, Algorithm Distillation, and Proximal Policy Optimization, while deriving theoretical guarantees relating the framework to Bayesian posterior sampling.

The analysis reveals several central findings. First, purely training the model to predict optimal actions produces sophisticated, emergent decision strategies: the model matches classical algorithms in online exploration without explicit exploration incentives and demonstrates conservative hedging when handling noisy offline data. Second, the model successfully uncovers latent problem structure; when trained on linear bandit data gathered by basic algorithms, it achieves regret performance comparable to specialized linear algorithms, outperforming the suboptimal algorithms that generated its training data. Third, in complex spatial and visual domains, the model generalizes robustly to unseen goals, permuted dynamics, and out-of-distribution noise levels, achieving high average returns (such as 61.5 return in spatial tasks from random offline data where baseline return was 1.1) and successfully stitching separate demonstrations into optimal paths. Finally, the authors prove theoretically that the model implements an efficient form of Bayesian posterior sampling, establishing bounded cumulative regret guarantees across finite decision processes.

These findings indicate that supervised pretraining offers a scalable alternative to hand-engineered reinforcement learning algorithms. By bypassing the computational bottleneck of maintaining explicit Bayesian posterior distributions, this approach allows organizations to deploy a single foundation-style model that rapidly adapts to novel decision environments in real time. Furthermore, the demonstrated ability to surpass the performance of pretraining data sources reduces the risk and expense associated with curating perfect operational demonstrations.

Organizations developing autonomous systems should explore supervised action pretraining pipelines for multi-task adaptation, prioritizing diverse task collections over complex algorithm design. When deploying these models, practitioners can safely use standard algorithm rollouts or policy approximations for data generation. Future technical work should focus on scaling the framework to broader continuous-control domains, testing integration with existing large language models, and developing automated methods for labeling near-optimal actions in complex real-world settings where optimal ground truth is unavailable.

Confidence in the reported results is high across the evaluated simulation benchmarks and supported by rigorous theoretical proofs. However, practical application carries minor uncertainties, as the empirical validation remains restricted to controlled simulated domains, and theoretical guarantees assume bounded statistical model complexity and compliant pretraining data collection.

arXiv: 2306.14892
Cover for Supervised Pretraining Can Learn In-Context Reinforcement Learning

Abstract

Large transformer models trained on diverse datasets have shown a remarkable ability to learn in-context, achieving high few-shot performance on tasks they were not explicitly trained to solve. In this paper, we study the in-context learning capabilities of transformers in decision-making problems, i.e., reinforcement learning (RL) for bandits and Markov decision processes. To do so, we introduce and study the Decision-Pretrained Transformer (DPT), a supervised pretraining method where a transformer predicts an optimal action given a query state and an in-context dataset of interactions from a diverse set of tasks. While simple, this procedure produces a model with several surprising capabilities. We find that the trained transformer can solve a range of RL problems in-context, exhibiting both exploration online and conservatism offline, despite not being explicitly trained to do so. The model also generalizes beyond the pretraining distribution to new tasks and automatically adapts its decision-making strategies to unknown structure. Theoretically, we show DPT can be viewed as an efficient implementation of Bayesian posterior sampling, a provably sample-efficient RL algorithm. We further leverage this connection to provide guarantees on the regret of the in-context algorithm yielded by DPT, and prove that it can learn faster than algorithms used to generate the pretraining data. These results suggest a promising yet simple path towards instilling strong in-context decision-making abilities in transformers.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 In-Context Learning Model
  • 4 Learning in Bandits
  • 5 Learning in Markov Decision Processes
  • 5.1 Experimental Setup
  • 5.2 Main Results
  • 5.3 Learning from Algorithm-Generated Policies and Rollouts
  • 6 Theory
  • 6.1 History-Dependent Pretraining and Assumptions
  • 6.2 Main Results
  • 7 Discussion
  • Acknowledgments and Disclosure of Funding
  • References
  • Additional Related Work
  • A Implementation and Experiment Details
  • A.1 DPT Architecture: Formal Description
  • A.2 Implementation Details
  • A.2.1 Bandit algorithms
  • A.2.2 RL Algorithms
  • A.3 Bandit Pretraining and Testing
  • A.4 MDP Environment Details
  • A.5 MDP Pretraining Datasets
  • B Additional Experimental Results
  • B.1 Bandits
  • B.2 Markov Decision Processes
  • B.3 Sensitivity Analysis
  • C Additional Theory and Omitted Proofs
  • C.1 Posterior Sampling
  • C.2 Proof of Theorem 1
  • C.3 Proof of Corollary 6.2
  • C.4 Proof of Corollary 6.3
  • C.5 Proof of Proposition 6.4

Knowls

  1. Knowl 1 — Decision-Pretrained Transformer objective

    model/method

    The Decision-Pretrained Transformer (DPT) is a causal GPT-2-style transformer trained across a distribution of reinforcement-learning tasks. A task is an MDP with an unknown reward and transition function. For each task, the pretraining procedure samples an in-context interaction dataset D={(si,ai,si′,ri)}i=1nD=\{(s_i,a_i,s'_i,r_i)\}_{i=1}^n, a query state sqs_q, and an optimal-action label a⋆a^\star drawn from the task’s optimal policy πτ⋆(⋅∣sq)\pi^\star_\tau(\cdot\mid s_q). The partial dataset Dj={(si,ai,si′,ri)}i=1jD_j=\{(s_i,a_i,s'_i,r_i)\}_{i=1}^j contains the first jj interactions.

    The model Mθ(⋅∣sq,Dj)M_\theta(\cdot\mid s_q,D_j) outputs a probability distribution over actions, with parameters θ\theta. DPT is trained by minimizing the expected negative log-likelihood of the same optimal action at every context prefix:

    min⁡θ  E(τ,D,sq,a⋆)∼Ppre[∑j=1n−log⁡Mθ(a⋆∣sq,Dj)].\min_\theta\;\mathbb E_{(\tau,D,s_q,a^\star)\sim P_{\mathrm{pre}}}\left[\sum_{j=1}^{n}-\log M_\theta(a^\star\mid s_q,D_j)\right].

    Here PpreP_{\mathrm{pre}} is the joint distribution induced by sampling the task, interaction dataset, query state, and optimal-action label during pretraining. For discrete actions, the model uses a softmax output; continuous actions can instead be predicted directly. The implementation embeds each transition tuple as one vector, prepends a query-state vector marked by zero padding, and applies causal self-attention so the prediction at prefix DjD_j cannot use later interactions. The model parameters are frozen after pretraining; adaptation occurs through the forward pass conditioned on the in-context dataset.

  2. Knowl 2 — Offline, online, and history-conditioned deployment

    algorithm

    DPT can be deployed without parameter updates on a previously collected dataset or while collecting data online. In offline deployment, an unknown task τ\tau supplies a static dataset DD; at each visited state shs_h, the agent selects the most probable action ah∈arg⁡max⁡a∈AMθ(a∣sh,D)a_h\in\arg\max_{a\in\mathcal A}M_\theta(a\mid s_h,D). In online deployment, the dataset is initially empty. During each episode, the agent samples ah∼Mθ(⋅∣sh,D)a_h\sim M_\theta(\cdot\mid s_h,D), records the resulting state-action-reward transitions, appends the completed episode to DD, and uses the enlarged dataset in the next episode. Sampling rather than taking an argmax is what allows uncertainty-driven exploration.

    For the theoretical analysis in MDPs, the model is given an additional sequence of previously visited state-action pairs, ξh−1=((s1,a1),…,(sh−1,ah−1))\xi_{h-1}=((s_1,a_1),\ldots,(s_{h-1},a_{h-1})). The training variant labels both the query action and auxiliary pairs with optimal-policy actions: ξh=((s1,a1⋆),…,(sh,ah⋆))\xi_h=((s_1,a_1^\star),\ldots,(s_h,a_h^\star)), where the states are sampled independently of the task and ai⋆∼πτ⋆(⋅∣si)a_i^\star\sim\pi_\tau^\star(\cdot\mid s_i). At test time, the model instead appends its own visited states and sampled actions to the history. This history-conditioned version is used for the posterior-sampling equivalence theorem; the experimentally used MDP DPT omits the explicit history and is treated as a practical approximation.

  3. Knowl 3 — DPT is equivalent to posterior sampling

    theoretical result

    Under the following conditions, the history-conditioned DPT implements posterior sampling in-context. The model must fit the pretraining conditional distribution exactly, meaning Mθ(a∣sq,D,ξ)=Ppre(a∣sq,D,ξ)M_\theta(a\mid s_q,D,\xi)=P_{\mathrm{pre}}(a\mid s_q,D,\xi) for every action aa, query state sqs_q, in-context dataset DD, and state-action history ξ\xi. The pretraining dataset-generation policy must be compliant: its action choices may depend on the observed state and partial dataset but not on privileged task information. The query-state, dataset, and auxiliary-history distributions must have sufficient support.

    Posterior sampling uses the pretraining task distribution TpreT_{\mathrm{pre}} as its prior. Given historical data DD in a fixed current MDP τc\tau_c, it samples a task from the posterior P(τ∣D)P(\tau\mid D), executes that sampled task’s optimal policy in τc\tau_c, and repeats this process as new episodes are collected. Let PPS(ξH∣D,τc)P_{\mathrm{PS}}(\xi_H\mid D,\tau_c) denote the distribution over length-HH state-action trajectories produced by this procedure, and let PMθ(ξH∣D,τc)P_{M_\theta}(\xi_H\mid D,\tau_c) denote the distribution produced by history-conditioned DPT. The paper proves

    PPS(ξH∣D,τc)=PMθ(ξH∣D,τc)P_{\mathrm{PS}}(\xi_H\mid D,\tau_c)=P_{M_\theta}(\xi_H\mid D,\tau_c)

    for every trajectory ξH\xi_H. Thus, supervised prediction of optimal actions can implement posterior sampling without explicitly computing, storing, or sampling from the posterior over task models.

  4. Knowl 4 — Regret guarantees inherited from posterior sampling

    theoretical result

    Because the history-conditioned DPT has the same trajectory distribution as posterior sampling, it inherits posterior-sampling regret guarantees under the corresponding assumptions. In a finite-horizon MDP with state-space size SS, action-space size AA, horizon HH, and KK online episodes, suppose the pretraining context is generated by uniform sampling, expected rewards lie in [0,1][0,1], query states and auxiliary histories have uniform support, and the test-task density is bounded relative to the pretraining-task density by Ttest(τ)/Tpre(τ)≤CT_{\mathrm{test}}(\tau)/T_{\mathrm{pre}}(\tau)\le C. The expected cumulative regret satisfies

    Eτ∼Ttest[Reg⁡τ(Mθ)]≤O~ ⁣(CH3/2SAK),\mathbb E_{\tau\sim T_{\mathrm{test}}}[\operatorname{Reg}_\tau(M_\theta)] \le \widetilde O\!\left(C H^{3/2}S\sqrt{AK}\right),

    where O~\widetilde O suppresses polylogarithmic factors.

    For the paper’s linear-bandit setting, the state is a singleton, each action aa has a fixed feature vector ϕ(a)∈Rd\phi(a)\in\mathbb R^d with sup⁡a∥ϕ(a)∥2≤1\sup_a\|\phi(a)\|_2\le 1, the task parameter is θτ∼N(0,Id/d)\theta_\tau\sim\mathcal N(0,I_d/d), and rewards satisfy r∼N(⟨θτ,ϕ(a)⟩,1)r\sim\mathcal N(\langle\theta_\tau,\phi(a)\rangle,1). When the pretraining and test distributions coincide and the context datasets are generated by Gaussian Thompson Sampling, DPT satisfies

    Eτ∼Ttest[Reg⁡τ(Mθ)]≤O~(dK),\mathbb E_{\tau\sim T_{\mathrm{test}}}[\operatorname{Reg}_\tau(M_\theta)]\le \widetilde O(d\sqrt K),

    which is the structure-aware rate rather than the generic O~(∣A∣K)\widetilde O(\sqrt{|\mathcal A|K}) rate.

  5. Knowl 5 — Compliance and invariance to the data-collection algorithm

    theoretical result

    An in-context dataset distribution Dpre(⋅;τ)D_{\mathrm{pre}}(\cdot;\tau) is compliant when, for every dataset position ii, its action-selection conditional distribution satisfies

    Dpre(ai∣si,Di−1;τ)=Dpre(ai∣si,Di−1),D_{\mathrm{pre}}(a_i\mid s_i,D_{i-1};\tau)=D_{\mathrm{pre}}(a_i\mid s_i,D_{i-1}),

    where sis_i is the current observed state, Di−1D_{i-1} is the preceding partial dataset, and τ\tau is the task. Compliance permits random policies and adaptive algorithms such as Thompson Sampling, PPO, or other policies that use only observed history; it excludes datasets whose actions use privileged knowledge of the task, such as deliberately selecting the known optimal action.

    Under exact fitting of the pretraining distribution, consider two pretraining distributions that have the same task, query-state, and history distributions but use different compliant dataset-generation policies with the same support. Their conditional optimal-action distributions are identical:

    Ppre(1)(a⋆∣sq,D,ξ)=Ppre(2)(a⋆∣sq,D,ξ).P^{(1)}_{\mathrm{pre}}(a^\star\mid s_q,D,\xi)=P^{(2)}_{\mathrm{pre}}(a^\star\mid s_q,D,\xi).

    Consequently, changing a compliant source algorithm does not change the ideal DPT policy. In contrast, expert-biased datasets generated using privileged knowledge of the optimal action violate compliance and can produce a qualitatively different DPT behavior.

  6. Knowl 6 — Bandit uncertainty handling and distributional robustness

    empirical result

    In the basic bandit experiments, DPT was pretrained on five-armed Gaussian bandits with independent means μa∼Unif⁡[0,1]\mu_a\sim\operatorname{Unif}[0,1], reward noise standard deviation 0.30.3, and diverse context datasets of up to 500 interactions. Offline performance was measured by suboptimality μa⋆−μa^\mu_{a^\star}-\mu_{\hat a}, and online performance by cumulative regret over 500 actions.

    On in-distribution offline datasets, the DPT policy substantially outperformed the empirical-mean and lower-confidence-bound policies and approximately matched Gaussian Thompson Sampling. This indicates that optimal-action prediction caused the transformer to account for uncertainty from noisy or undersampled arms rather than simply choosing the arm with the largest observed mean. When actions were sampled from DPT’s predicted distribution online, its cumulative regret approximately matched UCB and Thompson Sampling, despite DPT never being explicitly trained with an exploration objective. The curves and error-bar plots reported on page 6 show this online exploration behavior.

    The same model remained effective when the reward-noise standard deviation was changed at test time, including values not used during pretraining, and when Gaussian rewards were replaced by Bernoulli rewards. The additional Bernoulli and reward-shift plots on pages 22–23 show that the learned uncertainty handling generalized beyond the original reward family.

  7. Knowl 7 — DPT discovers latent linear-bandit structure

    empirical result

    The paper tested whether DPT could exploit structure that was not explicitly supplied to the transformer and was not used by the algorithm generating its context data. Tasks were linear bandits with expected reward E[r∣a,τ]=⟨θτ,ϕ(a)⟩\mathbb E[r\mid a,\tau]=\langle\theta_\tau,\phi(a)\rangle, where ϕ(a)\phi(a) is a fixed feature representation shared across tasks. DPT was not given ϕ\phi directly; its context datasets were generated by Thompson Sampling that treated the arms independently and therefore ignored the linear structure.

    Nevertheless, DPT learned an implicit surrogate for the shared representation. Offline, it achieved performance close to greedy linear regression and substantially better than the Thompson Sampling source policy. Online, its cumulative regret was nearly that of LinUCB, even though LinUCB was explicitly given ϕ\phi and DPT was not. The linear-bandit curves on page 7 therefore provide evidence that supervised pretraining can learn a more informative exploration and decision strategy than the policy used to generate the pretraining data.

  8. Knowl 8 — Pretraining controls conservatism on expert-biased offline data

    empirical result

    DPT’s offline behavior changes according to the distribution of datasets used during pretraining. The standard DPT was trained on bandit datasets with diverse action frequencies. A second model, DPT-Exp, was trained on datasets mixed with varying fractions of expert actions, where the expert action was the known optimal arm. At test time, both models were evaluated on datasets ranging from fully random to fully expert demonstrations.

    The standard DPT continued to behave like Thompson Sampling across these dataset compositions, whereas DPT-Exp became increasingly similar to a lower-confidence-bound policy and performed best when the test dataset contained the expert bias represented during its pretraining. Thus, supervised pretraining can teach the transformer an appropriate degree of offline conservatism from the data distribution itself, without manually specifying a pessimistic objective. These comparisons are shown in the expert-dataset plot on page 7.

  9. Knowl 9 — Generalization, exploration, and trajectory stitching in MDPs

    empirical result

    DPT was evaluated on two sparse-reward MDP families. In Dark Room, the agent navigates a 10×1010\times10 grid for 100 steps and receives reward 11 only at an unknown goal. In Miniworld, the agent navigates from 25×2525\times25 RGB observations toward one of four colored target boxes and receives reward 11 only near the correct box.

    DPT generalized to held-out Dark Room goals and to image-based Miniworld observations. Offline, it performed near-optimally with expert demonstrations and substantially improved over random datasets: in held-out Dark Room tasks, random datasets averaged total reward 1.11.1, while DPT obtained average return 61.561.5. Online from an empty context, DPT explored more efficiently than Algorithm Distillation and achieved higher final return than RL2; PPO trained from scratch made little progress in the limited interaction budget. The offline and online comparisons are summarized by the bar charts and learning curves on pages 7–8.

    DPT also stitched information from separate demonstrations into a new trajectory. In a three-task Dark Room variant, pretraining and test-time context contained expert demonstrations for two tasks, while the evaluated third goal was unseen. DPT inferred a direct path to the third goal by combining the useful state-action subsequences, demonstrating that it could construct a higher-return trajectory rather than merely replaying an observed one.

  10. Knowl 10 — Optimal labels can be replaced by algorithm-generated labels, but limitations remain

    limitation

    A principal limitation of DPT is that its original objective assumes access to optimal-action labels for every pretraining task, which may be impractical in general MDPs. The paper tested a relaxation in which PPO-trained actions supplied the labels and PPO replay buffers supplied the in-context datasets. PPO was trained on 80 Dark Room training tasks for 1,000 episodes per task, producing 80,000 rollouts. DPT trained with PPO datasets and PPO action labels performed comparably to the original DPT and better than Algorithm Distillation; a variant using random datasets with PPO labels was somewhat weaker but only marginally so in several settings.

    The authors identify three unresolved limitations: the best way to use imperfect or heterogeneous multi-task decision datasets is not yet understood; the experimentally convenient, history-free MDP implementation is not identical to the history-conditioned posterior-sampling construction used in the theory; and the implications for large pretrained language models deployed in decision-making remain open.

Coverage note — Detailed architecture hyperparameters, auxiliary sensitivity analyses, and implementation details were omitted because they support rather than define the paper’s main methodological, theoretical, and empirical contributions.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  2. 2.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837, 2022.
  3. 3.Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill. Data distributional properties drive emergent in-context learning in transformers. Advances in Neural Information Processing Systems, 35:18878–18891, 2022.
  4. 4.Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant. What can transformers learn in-context? a case study of simple function classes. Advances in Neural Information Processing Systems, 35:30583–30598, 2022.
  5. 5.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022.
  6. 6.Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. What learning algorithm is in-context learning? investigations with linear models. arXiv preprint arXiv:2211.15661, 2022.
  7. 7.Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, et al. In-context reinforcement learning with algorithm distillation. arXiv preprint arXiv:2210.14215, 2022.
  8. 8.Mengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu, Ding Zhao, Joshua Tenenbaum, and Chuang Gan. Prompting decision transformer for few-shot policy generalization. In International Conference on Machine Learning, pages 24631–24645. PMLR, 2022.
  9. 9.Mengdi Xu, Yuchen Lu, Yikang Shen, Shun Zhang, Ding Zhao, and Chuang Gan. Hyper-decision transformer for efficient online policy adaptation. arXiv preprint arXiv:2304.08487, 2023.
  10. 10.Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction. MIT press, 2018.
  11. 11.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  12. 12.Malcolm Strens. A bayesian framework for reinforcement learning. In ICML, volume 2000, pages 943–950, 2000.
  13. 13.Ian Osband, Daniel Russo, and Benjamin Van Roy. (more) efficient reinforcement learning via posterior sampling. Advances in Neural Information Processing Systems, 26, 2013.
  14. 14.Tom Schaul and Jürgen Schmidhuber. Metalearning. Scholarpedia, 5(6):4650, 2010.
  15. 15.Yoshua Bengio, Samy Bengio, and Jocelyn Cloutier. Learning a synaptic learning rule. Citeseer, 1990.
  16. 16.Justin Fu, Sergey Levine, and Pieter Abbeel. One-shot learning of manipulation skills with online dynamics adaptation and neural network priors. In 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4019–4026. IEEE, 2016.
  17. 17.Anusha Nagabandi, Ignasi Clavera, Simin Liu, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn. Learning to adapt in dynamic, real-world environments through meta-reinforcement learning. arXiv preprint arXiv:1803.11347, 2018.
  18. 18.Nicholas C Landolfi, Garrett Thomas, and Tengyu Ma. A model-based approach for sample-efficient multi-task reinforcement learning. arXiv preprint arXiv:1907.04964, 2019.
  19. 19.Max Simchowitz, Christopher Tosh, Akshay Krishnamurthy, Daniel J Hsu, Thodoris Lykouris, Miro Dudik, and Robert E Schapire. Bayesian decision-making under misspecified priors with applications to meta-learning. Advances in Neural Information Processing Systems, 34:26382–26394, 2021.
  20. 20.Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen. Efficient off-policy meta-reinforcement learning via probabilistic context variables. In International conference on machine learning, pages 5331–5340. PMLR, 2019.
  21. 21.Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess. Meta reinforcement learning as task inference. arXiv preprint arXiv:1905.06424, 2019.
  22. 22.Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson. Varibad: A very good method for bayes-adaptive deep rl via meta-learning. In International Conference on Learning Representation (ICLR), 2020.
  23. 23.Evan Z Liu, Aditi Raghunathan, Percy Liang, and Chelsea Finn. Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices. In International conference on machine learning, pages 6925–6935. PMLR, 2021.
  24. 24.Theodore J Perkins, Doina Precup, et al. Using options for knowledge transfer in reinforcement learning. Technical report, Citeseer, 1999.
  25. 25.Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine. Meta-reinforcement learning of structured exploration strategies. Advances in neural information processing systems, 31, 2018.
  26. 26.Yiding Jiang, Evan Liu, Benjamin Eysenbach, J Zico Kolter, and Chelsea Finn. Learning options via compression. Advances in Neural Information Processing Systems, 35:21184–21199, 2022.
  27. 27.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning, pages 1126–1135. PMLR, 2017.
  28. 28.Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel. Promp: Proximal meta-policy search. arXiv preprint arXiv:1810.06784, 2018.
  29. 29.Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel. Rl2: Fast reinforcement learning via slow reinforcement learning. arXiv preprint arXiv:1611.02779, 2016.
  30. 30.Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick. Learning to reinforcement learn. arXiv preprint arXiv:1611.05763, 2016.
  31. 31.Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. A simple neural attentive meta-learner. arXiv preprint arXiv:1707.03141, 2017.
  32. 32.Adaptive Agent Team, Jakob Bauer, Kate Baumli, Satinder Baveja, Feryal Behbahani, Avishkar Bhoopchand, Nathalie Bradley-Schmieg, Michael Chang, Natalie Clay, Adrian Collister, et al. Human-timescale adaptation in an open-ended task space. arXiv preprint arXiv:2301.07608, 2023.
  33. 33.Chris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto, Jakob Foerster, Satinder Singh, and Feryal Behbahani. Structured state space models for in-context reinforcement learning. arXiv preprint arXiv:2303.03982, 2023.
  34. 34.Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, Pieter Abbeel, and Dale Schuurmans. Foundation models for decision making: Problems, methods, and opportunities. arXiv preprint arXiv:2303.04129, 2023.
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  36. 36.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  37. 37.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34:15084–15097, 2021.
  38. 38.Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34:1273–1286, 2021.
  39. 39.Kuang-Huei Lee, Ofir Nachum, Mengjiao Sherry Yang, Lisa Lee, Daniel Freeman, Sergio Guadarrama, Ian Fischer, Winnie Xu, Eric Jang, Henryk Michalewski, et al. Multi-game decision transformers. Advances in Neural Information Processing Systems, 35:27921–27936, 2022.
  40. 40.Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent. arXiv preprint arXiv:2205.06175, 2022.
  41. 41.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022.
  42. 42.Nur Muhammad Shafiullah, Zichen Cui, Ariuntuya Arty Altanzaya, and Lerrel Pinto. Behavior transformers: Cloning k modes with one stone. Advances in Neural Information Processing systems, 35:22955–22968, 2022.
  43. 43.David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche, and Joan Bruna. When does return-conditioned supervised learning work for offline reinforcement learning? arXiv preprint arXiv:2206.01079, 2022.
  44. 44.Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum. Dichotomy of control: Separating what you can control from what you cannot. arXiv preprint arXiv:2210.13435, 2022.
  45. 45.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. Advances in Neural Information Processing Systems, 33:1179–1191, 2020.
  46. 46.Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. Combo: Conservative offline model-based policy optimization. Advances in Neural Information Processing Systems, 34:28954–28967, 2021.
  47. 47.Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill. Provably good batch off-policy reinforcement learning without great exploration. Advances in Neural Information Processing Systems, 33:1264–1274, 2020.
  48. 48.Kamyar Ghasemipour, Shixiang Shane Gu, and Ofir Nachum. Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters. Advances in Neural Information Processing Systems, 35:18267–18281, 2022.
  49. 49.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International conference on machine learning, pages 2052–2062. PMLR, 2019.
  50. 50.Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. Stabilizing off-policy q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems, 32, 2019.
  51. 51.Yifan Wu, George Tucker, and Ofir Nachum. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  52. 52.Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller. Keep doing what worked: Behavioral modelling priors for offline reinforcement learning. arXiv preprint arXiv:2002.08396, 2020.
  53. 53.Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill. Off-policy policy gradient with state distribution correction. UAI, 2019.
  54. 54.Lanqing Li, Rui Yang, and Dijun Luo. Focal: Efficient fully-offline meta-reinforcement learning via distance metric learning and behavior regularization. arXiv preprint arXiv:2010.01112, 2020.
  55. 55.Eric Mitchell, Rafael Rafailov, Xue Bin Peng, Sergey Levine, and Chelsea Finn. Offline meta-reinforcement learning with advantage weighting. In International Conference on Machine Learning, pages 7780–7791. PMLR, 2021.
  56. 56.Ron Dorfman, Idan Shenfeld, and Aviv Tamar. Offline meta reinforcement learning–identifiability challenges and effective data collection strategies. Advances in Neural Information Processing Systems, 34:4607–4618, 2021.
  57. 57.Vitchyr H Pong, Ashvin V Nair, Laura M Smith, Catherine Huang, and Sergey Levine. Offline meta-reinforcement learning with online self-supervision. In International Conference on Machine Learning, pages 17811–17829. PMLR, 2022.
  58. 58.Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al. A tutorial on thompson sampling. Foundations and Trends® in Machine Learning, 11(1):1–96, 2018.
  59. 59.William R Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25(3-4):285–294, 1933.
  60. 60.Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. Finite-time analysis of the multiarmed bandit problem. Machine learning, 47:235–256, 2002.
  61. 61.Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvari, and Dale Schuurmans. On the optimality of batch policy optimization algorithms. In International Conference on Machine Learning, pages 11362–11371. PMLR, 2021.
  62. 62.Ying Jin, Zhuoran Yang, and Zhaoran Wang. Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096. PMLR, 2021.
  63. 63.Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári. Improved algorithms for linear stochastic bandits. In Advances in Neural Information Processing Systems, pages 2312–2320, 2011.
  64. 64.Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell. Bridging offline reinforcement learning and imitation learning: A tale of pessimism. Advances in Neural Information Processing Systems, 34:11702–11716, 2021.
  65. 65.Maxime Chevalier-Boisvert. Miniworld: Minimalistic 3d environment for rl and robotics research, 2018.
  66. 66.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  67. 67.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. An explanation of in-context learning as implicit bayesian inference. arXiv preprint arXiv:2111.02080, 2021.
  68. 68.Daniel Russo and Benjamin Van Roy. Learning to optimize via posterior sampling. Mathematics of Operations Research, 39(4):1221–1243, 2014.
  69. 69.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023.
  70. 70.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Vladymyrov. Transformers learn in-context by gradient descent. arXiv preprint arXiv:2212.07677, 2022.
  71. 71.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895, 2022.
  72. 72.Louis Kirsch, James Harrison, Jascha Sohl-Dickstein, and Luke Metz. General-purpose in-context learning by meta-learning transformers. arXiv preprint arXiv:2212.04458, 2022.
  73. 73.Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woomyoung Park, Jung-Woo Ha, et al. On the effect of pretraining corpora on in-context learning by a large-scale language model. arXiv preprint arXiv:2204.13509, 2022.
  74. 74.Yingcong Li, M Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak. Transformers as algorithms: Generalization and implicit model selection in in-context learning. arXiv preprint arXiv:2301.07067, 2023.
  75. 75.Noam Wies, Yoav Levine, and Amnon Shashua. The learnability of in-context learning. arXiv preprint arXiv:2303.07895, 2023.
  76. 76.Jacob Abernethy, Alekh Agarwal, Teodor V Marinov, and Manfred K Warmuth. A mechanism for sample-efficient in-context learning for sparse retrieval tasks. arXiv preprint arXiv:2305.17040, 2023.
  77. 77.Nan Rosemary Ke, Silvia Chiappa, Jane Wang, Anirudh Goyal, Jorg Bornschein, Melanie Rey, Theophane Weber, Matthew Botvinic, Michael Mozer, and Danilo Jimenez Rezende. Learning to induce causal structure. arXiv preprint arXiv:2204.04875, 2022.
  78. 78.Anirudh Goyal, Abram Friesen, Andrea Banino, Theophane Weber, Nan Rosemary Ke, Adria Puigdomenech Badia, Arthur Guez, Mehdi Mirza, Peter C Humphreys, Ksenia Konyushova, et al. Retrieval-augmented reinforcement learning. In International Conference on Machine Learning, pages 7740–7765. PMLR, 2022.
  79. 79.Tung Nguyen and Aditya Grover. Transformer neural processes: Uncertainty-aware meta learning via sequence modeling. arXiv preprint arXiv:2207.04179, 2022.
  80. 80.Shipra Agrawal and Navin Goyal. Near-optimal regret bounds for thompson sampling. Journal of the ACM (JACM), 64(5):1–24, 2017.
  81. 81.Shipra Agrawal and Randy Jia. Optimistic posterior sampling for reinforcement learning: worst-case regret bounds. Advances in Neural Information Processing Systems, 30, 2017.
  82. 82.Xiuyuan Lu and Benjamin Van Roy. Ensemble sampling. Advances in neural information processing systems, 30, 2017.
  83. 83.Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep exploration via bootstrapped dqn. Advances in neural information processing systems, 29, 2016.
  84. 84.Ian Osband, John Aslanides, and Albin Cassirer. Randomized prior functions for deep reinforcement learning. Advances in Neural Information Processing Systems, 31, 2018.
  85. 85.Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy. Approximate thompson sampling via epistemic neural networks. arXiv preprint arXiv:2302.09205, 2023.
  86. 86.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  87. 87.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  88. 88.Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann. Stable-baselines3: Reliable reinforcement learning implementations. The Journal of Machine Learning Research, 22(1):12348–12355, 2021.
  89. 89.Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun. Flambe: Structural complexity and representation learning of low rank mdps. Advances in neural information processing systems, 33:20095–20107, 2020.

Citation

MLA
Lee, J., et al. “Supervised Pretraining Can Learn In-Context Reinforcement Learning”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 43057–83, https://proceedings.neurips.cc/paper_files/paper/2023/file/8644b61a9bc87bf7844750a015feb600-Paper-Conference.pdf.
APA
Lee, J., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., & Brunskill, E. (2023). Supervised Pretraining Can Learn In-Context Reinforcement Learning. Advances in Neural Information Processing Systems, 36, 43057–43083. https://proceedings.neurips.cc/paper_files/paper/2023/file/8644b61a9bc87bf7844750a015feb600-Paper-Conference.pdf
Chicago
Lee, J., A. Xie, A. Pacchiano, et al. 2023. “Supervised Pretraining Can Learn In-Context Reinforcement Learning”. Advances in Neural Information Processing Systems 36: 43057–83. https://proceedings.neurips.cc/paper_files/paper/2023/file/8644b61a9bc87bf7844750a015feb600-Paper-Conference.pdf.
Harvard
Lee, J. et al. (2023) “Supervised Pretraining Can Learn In-Context Reinforcement Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 43057–43083. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/8644b61a9bc87bf7844750a015feb600-Paper-Conference.pdf.
Vancouver
1. Lee J, Xie A, Pacchiano A, Chandak Y, Finn C, Nachum O, Brunskill E (2023) Supervised Pretraining Can Learn In-Context Reinforcement Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 43057–43083

BibTeX

@inproceedings{lee2023supervised,
  title = {Supervised Pretraining Can Learn In-Context Reinforcement Learning},
  author = {Lee, Jonathan and Xie, Annie and Pacchiano, Aldo and Chandak, Yash and Finn, Chelsea and Nachum, Ofir and Brunskill, Emma},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {43057-43083},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/8644b61a9bc87bf7844750a015feb600-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors