Supervised Pretraining Can Learn In-Context Reinforcement Learning
Jonathan LeeAnnie XieAldo PacchianoYash ChandakChelsea FinnOfir NachumEmma Brunskill
Demonstrates that supervised pretraining to predict optimal actions enables transformers to perform sample-efficient in-context reinforcement learning with provable regret guarantees by effectively executing Bayesian posterior sampling.
Modern decision-making systems in robotics, recommendation engines, and autonomous operations require artificial intelligence agents that can quickly adapt to new environments without costly retraining. While transformer architectures excel at in-context learning for language tasks, adapting them to sequential decision-making and reinforcement learning has remained challenging. Standard reinforcement learning methods often require extensive data collection or computationally prohibitive Bayesian updates, while existing decision-focused transformers struggle to generalize beyond their training demonstrations or properly balance risk and exploration.
The article introduces and evaluates the Decision-Pretrained Transformer, a supervised pretraining framework designed to give transformer models in-context reinforcement learning capabilities. The primary objective is to demonstrate that standard supervised pretraining on optimal action predictions allows a transformer to execute efficient decision-making strategies—both online and offline—across unseen tasks and environments without any test-time model weight updates.
To evaluate this framework, the authors conducted extensive computational simulations across multi-armed bandits, structured linear bandits, discrete grid navigation tasks, and high-dimensional 3D vision-based environments. The model was trained using a standard causal GPT-2 backbone on datasets of past interactions paired with optimal action labels across diverse pretraining tasks. The model was then evaluated on entirely new tasks in two modes: offline, where it selects actions given fixed historical datasets, and online, where it collects its own data sequentially to solve an unknown task from scratch. The authors compared this approach against established algorithms, including Upper Confidence Bound, Thompson Sampling, Algorithm Distillation, and Proximal Policy Optimization, while deriving theoretical guarantees relating the framework to Bayesian posterior sampling.
The analysis reveals several central findings. First, purely training the model to predict optimal actions produces sophisticated, emergent decision strategies: the model matches classical algorithms in online exploration without explicit exploration incentives and demonstrates conservative hedging when handling noisy offline data. Second, the model successfully uncovers latent problem structure; when trained on linear bandit data gathered by basic algorithms, it achieves regret performance comparable to specialized linear algorithms, outperforming the suboptimal algorithms that generated its training data. Third, in complex spatial and visual domains, the model generalizes robustly to unseen goals, permuted dynamics, and out-of-distribution noise levels, achieving high average returns (such as 61.5 return in spatial tasks from random offline data where baseline return was 1.1) and successfully stitching separate demonstrations into optimal paths. Finally, the authors prove theoretically that the model implements an efficient form of Bayesian posterior sampling, establishing bounded cumulative regret guarantees across finite decision processes.
These findings indicate that supervised pretraining offers a scalable alternative to hand-engineered reinforcement learning algorithms. By bypassing the computational bottleneck of maintaining explicit Bayesian posterior distributions, this approach allows organizations to deploy a single foundation-style model that rapidly adapts to novel decision environments in real time. Furthermore, the demonstrated ability to surpass the performance of pretraining data sources reduces the risk and expense associated with curating perfect operational demonstrations.
Organizations developing autonomous systems should explore supervised action pretraining pipelines for multi-task adaptation, prioritizing diverse task collections over complex algorithm design. When deploying these models, practitioners can safely use standard algorithm rollouts or policy approximations for data generation. Future technical work should focus on scaling the framework to broader continuous-control domains, testing integration with existing large language models, and developing automated methods for labeling near-optimal actions in complex real-world settings where optimal ground truth is unavailable.
Confidence in the reported results is high across the evaluated simulation benchmarks and supported by rigorous theoretical proofs. However, practical application carries minor uncertainties, as the empirical validation remains restricted to controlled simulated domains, and theoretical guarantees assume bounded statistical model complexity and compliant pretraining data collection.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). It introduces the Decision Transformer architecture, framing sequential decision-making as autoregressive sequence modeling which directly inspires the supervised pretraining formulation of Decision-Pretrained Transformers.
- Paper: Multi-Game Decision Transformers, Kuang-Huei Lee et al. (2022). It establishes multi-task pretraining of sequence-modeling transformers on diverse datasets, laying the foundational methodology for scaling decision-making models across varied environments.
- Paper: Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Sébastien Bubeck et al. (2012). It provides the foundational regret analysis and theoretical principles for multi-armed bandits, which the source builds upon to theoretically bound in-context decision regret.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). It reviews key principles and challenges of offline reinforcement learning, including distributional shift and conservatism, which the source model implicitly reproduces in-context.
- Paper: Conservative Q-Learning for Offline Reinforcement Learning, Aviral Kumar et al. (2020). It details how conservative regularization enables reliable offline decision-making from fixed historical datasets, clarifying the offline conservatism observed in Decision-Pretrained Transformers.
- Paper: In-context Convergence of Transformers, Yu Huang et al. (2024). It provides theoretical convergence and optimization dynamics for softmax attention in in-context learning, extending the theoretical understanding of transformer-based in-context adaptation.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). It builds on supervised fine-tuning and in-context conditioning to dynamically align models with multi-objective preferences without heavy reinforcement learning loops.
- Paper: Sample-Efficient Learning from Agent Experience, Chenhui Gou et al. (2026). It investigates distilling ephemeral in-context agent experiences directly into model parameters, extending the practical application of in-context decision-making.
- Paper: Next-Latent Prediction Transformers Learn Compact World Models, Jayden Teoh et al. (2025). It explores how next-latent prediction objectives help transformers learn compact internal world models, extending the mechanisms of sequence modeling for sequential reasoning and planning.
