Guiding Pretraining in Reinforcement Learning with Large Language Models
Yuqing DuOlivia WatkinsZihan WangCédric ColasTrevor DarrellPieter AbbeelAbhishek GuptaJacob Andreas
Proposes ELLM, a framework that uses pretrained large language models to generate context-sensitive, common-sense exploration goals and intrinsic rewards, steering reinforcement learning agents toward meaningful behaviors without manual reward engineering or human intervention.
Reinforcement learning systems often struggle when operating in large, complex environments without frequent and carefully designed feedback. Standard intrinsic exploration techniques attempt to resolve this by rewarding agents whenever they encounter novel or unpredictable states. However, these novelty-seeking methods frequently fail in open-ended settings because they spend excessive time exploring irrelevant or nonsensical variations, such as tracking random environmental noise, rather than learning practical and meaningful behaviors.
The article introduces and evaluates Exploring with Large Language Models (ELLM), a framework designed to bias agent pretraining toward diverse, context-aware, and common-sense behaviors without requiring human intervention or task-specific manual reward engineering.
ELLM operates by using a state captioner to translate an agent's current environment observation and inventory into a text prompt. A pretrained large language model processes this prompt to suggest plausible, human-meaningful goals, such as chopping a tree or picking up a misplaced item. When the agent carries out an action, a transition captioner describes the resulting change, and the agent receives an intrinsic reward proportional to the semantic similarity between the transition description and the suggested goals. The authors evaluated ELLM across two simulated benchmarks: Crafter, an open-world survival game, and Housekeep, an embodied robotics simulator focused on household organization.
The experiments produced several critical findings. First, prompted language models generate high-quality exploratory targets; in Crafter, roughly 65% of generated goals were feasible, context-sensitive, and aligned with common sense. Second, ELLM significantly improved pretraining exploration, enabling agents to unlock an average of about 6 unique achievements per episode in Crafter, compared to fewer than 3 achievements for traditional novelty-seeking baselines like Random Network Distillation and Active Pre-Training. Third, ELLM was the only tested framework to achieve positive performance across all downstream transfer tasks in Crafter. In downstream deployment, using the pretrained ELLM policy to guide the exploratory actions of a newly initialized model proved more reliable than directly fine-tuning the pretrained weights, which often led to instability due to shifting reward scales.
These results demonstrate that large-scale linguistic pretraining can substitute for expensive, hand-engineered reward functions during autonomous exploration. Integrating common-sense background knowledge substantially reduces the sample inefficiency and wasted computational effort typical of unsupervised reinforcement learning. The findings also suggest that when transferring exploratory behaviors to specific target tasks, leveraging pretrained models as exploratory guides rather than direct policy initializations mitigates the risk of catastrophic unlearning.
Organizations developing autonomous agents for complex or open-ended environments should consider incorporating language model priors into their exploratory training pipelines to improve learning efficiency. Next development steps should focus on pairing language model rewards with standard novelty bonuses to prevent blind spots, implementing strict safety and bias filters on generated suggestions, and utilizing caching to manage API querying costs and latency.
Decision-makers should interpret these conclusions in light of several limitations. ELLM's performance relies heavily on effective prompt design and the presence of accurate observation captioners. Language models can exhibit domain gaps, such as omitting necessary crafting steps, which may prevent agents from discovering certain critical skills. Additionally, deploying language-guided exploration in real-world or physical robotics settings will require robust vision-to-language models and safeguards against unwanted or biased behaviors encoded within the language models.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Establishes how large language models act as zero-shot planners by extracting actionable common-sense knowledge into interactive environments, providing the core premise for ELLM's goal generation.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Introduces closing the loop between LLM planning and embodied environmental feedback via natural language captions, directly underlying ELLM's state and transition captioning setup.
- Paper: Exploration by Random Network Distillation, Yuri Burda et al. (2019). Defines Random Network Distillation, the foundational novelty-seeking intrinsic exploration baseline that ELLM explicitly compares against and seeks to improve upon.
- Paper: Curiosity-Driven Exploration by Self-Supervised Prediction, Deepak Pathak et al. (2017). Introduces prediction-error-based intrinsic curiosity rewards in sparse reinforcement learning, framing the classical novelty-driven exploration paradigm replaced by ELLM's language priors.
- Paper: Diversity is All You Need: Learning Skills without a Reward Function, Benjamin Eysenbach et al. (2018). Provides the foundational unsupervised reinforcement learning framework for discovering diverse, task-agnostic skills prior to downstream reward-driven fine-tuning.
- Paper: Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, Tejas D. Kulkarni et al. (2016). Demonstrates hierarchical reinforcement learning with intrinsic goal setting, establishing the architectural paradigm of using intrinsic meta-goals to guide lower-level policies.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Demonstrates the grounding of LLM outputs into executable embodied robot behaviors and control interfaces.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Surveys the broader ecosystem of LLM-based autonomous agents, contextualizing ELLM's exploration framework within unified agent architectures and capability acquisition paradigms.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). Synthesizes modern LLM agent controllers across perception, planning, and tool use, contextualizing language-guided reinforcement learning exploration in multi-environment settings.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). Extends LLM-guided decision-making and exploration by integrating tree search, external environmental feedback, and self-reflective evaluations.
- Paper: METRA: Scalable Unsupervised RL with Metric-Aware Abstraction, Seohong Park et al. (2024). Advances unsupervised pre-training and exploration in reinforcement learning by formalizing metric-aware abstractions based on temporal distance.
- Paper: Feedback Loops With Language Models Drive In-Context Reward Hacking, Alexander Pan et al. (2024). Examines the failure modes and safety concerns of deploying LLMs in closed feedback loops where proxy rewards lead to unintended behavioral shifts.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). Provides a comprehensive overview of agentic reasoning, categorizing post-training exploration, interactive learning, and multi-agent coordination.
