Prompting Decision Transformer for Few-Shot Policy Generalization
Mengdi XuYikang ShenShun ZhangYuchen LuDing ZhaoJoshua B. TenenbaumChuang Gan
Presents Prompt-DT, a sequence-modeling approach that conditions decision transformers on short demonstration trajectory prompts to enable offline policy adaptation to unseen and out-of-distribution control tasks without fine-tuning.
Deploying reinforcement learning agents in high-stakes, real-world domains such as robotics, healthcare, and autonomous driving requires learning from pre-collected historical data rather than dangerous or costly live trial-and-error. However, conventional offline reinforcement learning methods struggle significantly when faced with new, unseen tasks, requiring complex algorithms and computationally heavy model updates to adapt.
The article demonstrates that conditioning an autoregressive Transformer architecture on short demonstration segments—termed trajectory prompts—enables an artificial intelligence agent to rapidly generalize to unseen control tasks without any test-time model parameter updates or fine-tuning.
The authors evaluate their proposed method, Prompt-based Decision Transformer (Prompt-DT), across five continuous control benchmarks in simulated robotic environments. The approach treats decision-making as a sequence-modeling problem, prepending very short demonstration trajectories (ranging from 2 to 15 timesteps) directly to the agent's recent history to specify the target task. Prompt-DT is compared against traditional multi-task models and established meta-learning algorithms, with performance measured by cumulative task rewards across multiple test environments.
The evaluation yields several critical findings. First, Prompt-DT consistently outperforms strong offline meta-reinforcement learning baselines such as MACAW by a wide margin, converging faster and achieving higher performance while requiring substantially less adaptation data (for instance, 5 timesteps versus 256 samples). Second, trajectory prompts provide far more effective task identification than historical reward signals alone, allowing the agent to adapt immediately without test-time fine-tuning. Third, the system exhibits strong robustness to prompt length—achieving near-peak performance with as few as two timesteps—but displays high sensitivity to prompt quality, performing best when conditioned on high-quality demonstrations. Finally, Prompt-DT successfully generalizes to out-of-distribution tasks that lie beyond the range of training goals, an evaluation setting where prior methods fail.
These findings suggest that inductive architecture design and prompting techniques can replace complex meta-optimization loops, reducing the computational cost, training time, and engineering overhead of deploying adaptable autonomous agents. By eliminating online fine-tuning, this framework minimizes the risk of catastrophic forgetting and catastrophic failures during field deployment.
Organizations developing autonomous systems should explore prompt-conditioned sequence architectures over gradient-based meta-learning when rapid task switching is required from pre-collected data. When implementing these systems, teams should prioritize the collection of high-quality demonstration prompts rather than large volumes of adaptation data. Before deploying in production, further research is recommended to address current limitations, particularly the model's struggle on highly complex compositional tasks (such as multi-stage tool manipulation in the ML10 benchmark) and its dependence on high-quality offline datasets.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). It introduces the Decision Transformer architecture that reformulates offline reinforcement learning as conditional sequence modeling, establishing the foundational policy backbone that Prompt-DT builds upon.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). It provides essential background on few-shot adaptation using demonstration prompts and prompt-based tuning that motivates Prompt-DT's trajectory prompt framework.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). It presents a comprehensive taxonomy and theoretical foundation of prompting paradigms across sequence models, clarifying the prompting mechanics adapted for policy generalization.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). It investigates parameter-efficient soft prompt conditioning in Transformers, providing architectural insight into how prompt tokens steer sequence models without full-parameter retraining.
- Paper: Offline Reinforcement Learning with Implicit Q-Learning, Ilya Kostrikov et al. (2021). It details key concepts and benchmark standards in offline reinforcement learning without online environmental interactions, framing the offline policy optimization paradigm used in Prompt-DT.
- Paper: Supervised Pretraining Can Learn In-Context Reinforcement Learning, Jonathan Lee et al. (2023). It generalizes in-context reinforcement learning via supervised pretraining of Transformers across online and offline tasks, extending Prompt-DT's prompt-guided policy formulation.
- Paper: Multi-Game Decision Transformers, Kuang-Huei Lee et al. (2022). It scales the decision-transformer sequence modeling approach across dozens of multi-game environments, expanding on the trajectory-conditioned generalist agent paradigm.
- Paper: Human-Timescale Adaptation in an Open-Ended Task Space, Jakob Bauer et al. (2023). It explores large-scale fast adaptation and in-context policy improvement in complex 3D environments, continuing Prompt-DT's investigation into architecture-driven few-shot RL.
- Paper: PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs, Soroush Nasiriany et al. (2024). It applies iterative prompting to vision-language models for zero-shot and few-shot continuous robotic control, building on prompt-based decision-making concepts.
- Paper: A Survey on In-context Learning, Qingxiu Dong et al. (2024). It synthesizes emerging theories and empirical mechanisms of in-context learning and demonstration conditioning, contextualizing trajectory-based prompting methods.
