Guiding Pretraining in Reinforcement Learning with Large Language Models

Yuqing DuOlivia WatkinsZihan WangCédric ColasTrevor DarrellPieter AbbeelAbhishek GuptaJacob Andreas

article2023ICML287 citations

Proposes ELLM, a framework that uses pretrained large language models to generate context-sensitive, common-sense exploration goals and intrinsic rewards, steering reinforcement learning agents toward meaningful behaviors without manual reward engineering or human intervention.

Listen

Reinforcement learning systems often struggle when operating in large, complex environments without frequent and carefully designed feedback. Standard intrinsic exploration techniques attempt to resolve this by rewarding agents whenever they encounter novel or unpredictable states. However, these novelty-seeking methods frequently fail in open-ended settings because they spend excessive time exploring irrelevant or nonsensical variations, such as tracking random environmental noise, rather than learning practical and meaningful behaviors.

The article introduces and evaluates Exploring with Large Language Models (ELLM), a framework designed to bias agent pretraining toward diverse, context-aware, and common-sense behaviors without requiring human intervention or task-specific manual reward engineering.

ELLM operates by using a state captioner to translate an agent's current environment observation and inventory into a text prompt. A pretrained large language model processes this prompt to suggest plausible, human-meaningful goals, such as chopping a tree or picking up a misplaced item. When the agent carries out an action, a transition captioner describes the resulting change, and the agent receives an intrinsic reward proportional to the semantic similarity between the transition description and the suggested goals. The authors evaluated ELLM across two simulated benchmarks: Crafter, an open-world survival game, and Housekeep, an embodied robotics simulator focused on household organization.

The experiments produced several critical findings. First, prompted language models generate high-quality exploratory targets; in Crafter, roughly 65% of generated goals were feasible, context-sensitive, and aligned with common sense. Second, ELLM significantly improved pretraining exploration, enabling agents to unlock an average of about 6 unique achievements per episode in Crafter, compared to fewer than 3 achievements for traditional novelty-seeking baselines like Random Network Distillation and Active Pre-Training. Third, ELLM was the only tested framework to achieve positive performance across all downstream transfer tasks in Crafter. In downstream deployment, using the pretrained ELLM policy to guide the exploratory actions of a newly initialized model proved more reliable than directly fine-tuning the pretrained weights, which often led to instability due to shifting reward scales.

These results demonstrate that large-scale linguistic pretraining can substitute for expensive, hand-engineered reward functions during autonomous exploration. Integrating common-sense background knowledge substantially reduces the sample inefficiency and wasted computational effort typical of unsupervised reinforcement learning. The findings also suggest that when transferring exploratory behaviors to specific target tasks, leveraging pretrained models as exploratory guides rather than direct policy initializations mitigates the risk of catastrophic unlearning.

Organizations developing autonomous agents for complex or open-ended environments should consider incorporating language model priors into their exploratory training pipelines to improve learning efficiency. Next development steps should focus on pairing language model rewards with standard novelty bonuses to prevent blind spots, implementing strict safety and bias filters on generated suggestions, and utilizing caching to manage API querying costs and latency.

Decision-makers should interpret these conclusions in light of several limitations. ELLM's performance relies heavily on effective prompt design and the presence of accurate observation captioners. Language models can exhibit domain gaps, such as omitting necessary crafting steps, which may prevent agents from discovering certain critical skills. Additionally, deploying language-guided exploration in real-world or physical robotics settings will require robust vision-to-language models and safeguards against unwanted or biased behaviors encoded within the language models.

arXiv: 2302.06692
Cover for Guiding Pretraining in Reinforcement Learning with Large Language Models

Abstract

Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivated exploration methods address this limitation by rewarding agents for visiting novel states or transitions, but these methods offer limited benefits in large environments where most discovered novelty is irrelevant for downstream tasks. We describe a method that uses background knowledge from text corpora to shape exploration. This method, called ELLM (Exploring with LLMs) rewards an agent for achieving goals suggested by a language model prompted with a description of the agent’s current state. By leveraging large-scale language model pretraining, ELLM guides agents toward human-meaningful and plausibly useful behaviors without requiring a human in the loop. We evaluate ELLM in the Crafter game environment and the Housekeep robotic simulator, showing that ELLM-trained agents have better coverage of common-sense behaviors during pretraining and usually match or improve performance on a range of downstream tasks.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Work
  • 3. Structuring Exploration with LLM Priors
  • 3.1. Problem Description
  • 3.2. Goal-based Exploration Desiderata
  • 3.3. Goal Generation with LLMs ( G )
  • 3.4. Rewarding LLM Goals ( R int )
  • 3.5. Implementation Details
  • 4. Experiments
  • 4.1. Crafter
  • 4.2. Housekeep
  • 5. Conclusions and Discussion
  • 6. Acknowledgements
  • References
  • A. Crafter Pretraining Ablation
  • B. Crafter Downstream Training
  • C. Crafter Env Modifications
  • D. Crafter Prompt
  • E. Crafter Action Space
  • F. Housekeep Tasks
  • G. Housekeep Prompt
  • H. Algorithmic Details
  • I. Hard-coded Captioner Details
  • J. Learned Crafter Captioner
  • K. Crafter LLM Analysis
  • L. Novelty Bonus Ablation
  • M. Analysis of Downstream Training Approaches
  • N. Additional Baselines
  • O. Code and Compute
  • P. Societal Impact

Knowls

  1. Knowl 1 — Exploring with Large Language Models (ELLM) Framework

    model/method

    Exploring with Large Language Models (ELLM) is an intrinsically motivated reinforcement learning (RL) framework that uses background knowledge from pretrained large language models (LLMs) to guide task-agnostic exploration towards human-meaningful and plausibly useful behaviors.

    In a partially observed Markov decision process (POMDP) defined by (S,A,O,Ω,T,γ,R)(\mathcal{S}, \mathcal{A}, \mathcal{O}, \Omega, \mathcal{T}, \gamma, \mathcal{R}) with observations o∈Ωo \in \Omega and actions a∈Aa \in \mathcal{A}, ELLM operates as a competence-based intrinsically motivated RL method. Instead of manually specifying a goal distribution G\mathcal{G} or goal-conditioned intrinsic reward Rint(o,a,o′∣g)R_{\text{int}}(o, a, o' \mid g), ELLM:

    1. Converts the current observation oto_t into a textual description via a state captioner Cobs(ot)C_{\text{obs}}(o_t).
    2. Prompts a pretrained autoregressive LLM with the text observation and available action vocabulary to generate a candidate set of kk context-sensitive, common-sense goals {gt1,…,gtk}\{g_t^1, \dots, g_t^k\}.
    3. Employs an episode-level novelty filter to remove any goal gtig_t^i that has already been achieved earlier in the current episode.
    4. Uses a transition captioner Ctransition(ot,at,ot+1)C_{\text{transition}}(o_t, a_t, o_{t+1}) and a sentence embedding model E(⋅)E(\cdot) to evaluate semantic similarity between the executed transition and the suggested goals, rewarding the agent if cosine similarity exceeds a threshold TT.
  2. Knowl 2 — ELLM Goal-Conditioned Intrinsic Reward Formulation

    equation

    ELLM defines the intrinsic reward Rint(ot,at,ot+1∣g)R_{\text{int}}(o_t, a_t, o_{t+1} \mid g) for achieving a suggested natural language goal gg by comparing the text description of the state transition against gg using an encoder E(⋅)E(\cdot) (specifically SentenceBERT):

    Rint(ot,at,ot+1∣g)={Δ(Ctransition(ot,at,ot+1),g)if Δ(Ctransition(ot,at,ot+1),g)>T0otherwiseR_{\text{int}}(o_t, a_t, o_{t+1} \mid g) = \begin{cases} \Delta(C_{\text{transition}}(o_t, a_t, o_{t+1}), g) & \text{if } \Delta(C_{\text{transition}}(o_t, a_t, o_{t+1}), g) > T \\ 0 & \text{otherwise} \end{cases}

    where Ctransition:Ω×A×Ω→Σ∗C_{\text{transition}}: \Omega \times \mathcal{A} \times \Omega \to \Sigma^* maps the transition (ot,at,ot+1)(o_t, a_t, o_{t+1}) to a text description, T∈[0,1]T \in [0, 1] is a similarity threshold hyperparameter, and Δ(u,v)\Delta(u, v) is the cosine similarity between sentence embeddings:

    Δ(u,v)=E(u)⋅E(v)∥E(u)∥∥E(v)∥\Delta(u, v) = \frac{E(u) \cdot E(v)}{\|E(u)\| \|E(v)\|}

    When the LLM suggests a set of kk goals {gt1,…,gtk}\{g_t^1, \dots, g_t^k\}, the overall intrinsic reward is the maximum similarity across all suggested goals:

    Δmax⁡=max⁡i=1,…,kΔ(Ctransition(ot,at,ot+1),gti)\Delta^{\max} = \max_{i=1,\dots,k} \Delta(C_{\text{transition}}(o_t, a_t, o_{t+1}), g_t^i)

    Rint(ot,at,ot+1)=Eg1:k∼LLM(⋅∣Cobs(ot))[Δmax⁡⋅I(Δmax⁡>T)]R_{\text{int}}(o_t, a_t, o_{t+1}) = \mathbb{E}_{g^{1:k} \sim \text{LLM}(\cdot \mid C_{\text{obs}}(o_t))}\left[ \Delta^{\max} \cdot \mathbb{I}(\Delta^{\max} > T) \right]

  3. Knowl 3 — ELLM Exploration and Pretraining Algorithm

    algorithm

    The ELLM pretraining procedure trains a policy π\pi (which may optionally be conditioned on goal embeddings E(gt1:k)E(g_t^{1:k}) and/or text state embeddings E(Cobs(ot))E(C_{\text{obs}}(o_t))) via deep Q-learning using LLM-suggested goals and semantic similarity intrinsic rewards.

    Input: Environment env, State captioner CobsC_{\text{obs}}, Transition captioner CtransitionC_{\text{transition}}, Language model LLM, Sentence encoder EE, Similarity threshold TT, Maximum environment steps NN
    Output: Pretrained policy π\pi
    Initialize untrained policy π\pi and replay buffer Buffer←∅\text{Buffer} \leftarrow \emptyset
    t←0t \leftarrow 0
    ot←env.RESET()o_t \leftarrow \text{env.RESET}()
    AchievedGoals←∅\text{AchievedGoals} \leftarrow \emptyset
    while t<Nt < N do
        ct←Cobs(ot)c_t \leftarrow C_{\text{obs}}(o_t)
        RawGoals←LLM(ct)\text{RawGoals} \leftarrow \text{LLM}(c_t)
        gt1:k←[g for g∈RawGoals if g∉AchievedGoals]g_t^{1:k} \leftarrow [g \text{ for } g \in \text{RawGoals if } g \notin \text{AchievedGoals}]
        
        at∼π(at∣ot,E(ct),E(gt1:k))a_t \sim \pi(a_t \mid o_t, E(c_t), E(g_t^{1:k}))
        ot+1,done←env.STEP(at)o_{t+1}, \text{done} \leftarrow \text{env.STEP}(a_t)
        
        ctrans←Ctransition(ot,at,ot+1)c_{\text{trans}} \leftarrow C_{\text{transition}}(o_t, a_t, o_{t+1})
        Δmax⁡←max⁡i=1,…,kΔ(ctrans,gti)\Delta^{\max} \leftarrow \max_{i=1,\dots,k} \Delta(c_{\text{trans}}, g_t^i)
        
        if Δmax⁡>T\Delta^{\max} > T then
            rt←Δmax⁡r_t \leftarrow \Delta^{\max}
            g∗←arg⁡max⁡g∈gt1:kΔ(ctrans,g)g^* \leftarrow \arg\max_{g \in g_t^{1:k}} \Delta(c_{\text{trans}}, g)
            AchievedGoals←AchievedGoals∪{g∗}\text{AchievedGoals} \leftarrow \text{AchievedGoals} \cup \{g^*\}
        else
            rt←0r_t \leftarrow 0
        end if
        
        Buffer←Buffer∪{(ot,at,gt1:k,rt,ot+1)}\text{Buffer} \leftarrow \text{Buffer} \cup \{(o_t, a_t, g_t^{1:k}, r_t, o_{t+1})\}
        π←UPDATE(π,Buffer)\pi \leftarrow \text{UPDATE}(\pi, \text{Buffer})
        
        if done then
            ot+1←env.RESET()o_{t+1} \leftarrow \text{env.RESET}()
            AchievedGoals←∅\text{AchievedGoals} \leftarrow \emptyset
        end if
        
        ot←ot+1o_t \leftarrow o_{t+1}
        t←t+1t \leftarrow t + 1
    end while
    return π\pi
  4. Knowl 4 — LLM Goal Generation Strategies: Open-Ended vs Closed-Form Prompting

    model/method

    ELLM extracts goal proposals from large language models using two distinct prompting strategies depending on the structure of the goal space:

    1. Open-Ended Generation: The LLM is prompted with a list of valid action verbs, a natural language caption of the current state Cobs(o)C_{\text{obs}}(o), and a question prompt (e.g., "What do you do?"). The text generated directly by the LLM (e.g., "Cut down the tree", "Attack the cow") is extracted as candidate goals. This mode is suited for open-ended environments with unconstrained goal spaces (such as Crafter), where goal descriptions cannot be easily enumerated in advance. Few-shot in-context examples are supplied in the prompt prefix to enforce consistent phrasing.
    2. Closed-Form Generation: In environments with large but delimitable goal spaces (such as object rearrangement in Housekeep), candidate goals are framed as individual queries to the LLM (e.g., "Should you store [object] in/on [receptacle]? (Yes/No)"). A goal proposal is accepted if and only if log⁡P("Yes")>log⁡P("No")\log P(\text{"Yes"}) > \log P(\text{"No"}). This format allows caching and reusing query outputs for repeated object-receptacle combinations across timesteps, substantially cutting API costs and runtime.
  5. Knowl 5 — Guided Exploration for Downstream Policy Transfer

    model/method

    When transferring an exploration-pretrained policy πpre\pi_{\text{pre}} to downstream tasks with extrinsic reward R\mathcal{R}, naive fine-tuning (directly continuing policy optimization from πpre\pi_{\text{pre}} on R\mathcal{R}) often fails due to catastrophic unlearning caused by the sudden shift in reward scale and density.

    Instead, downstream transfer is executed via guided exploration:

    1. A new downstream agent πdown\pi_{\text{down}} is initialized from scratch.
    2. The pretrained policy πpre\pi_{\text{pre}} is frozen.
    3. During reinforcement learning with ϵ\epsilon-greedy exploration on the target task, 50%50\% of the randomly sampled exploration actions are replaced with actions sampled from πpre\pi_{\text{pre}}:

    at={arg⁡max⁡aQdown(st,a)with probability 1−ϵa∼πpre(⋅∣st)with probability 0.5ϵa∼Uniform(A)with probability 0.5ϵa_t = \begin{cases} \arg\max_a Q_{\text{down}}(s_t, a) & \text{with probability } 1 - \epsilon \\ a \sim \pi_{\text{pre}}(\cdot \mid s_t) & \text{with probability } 0.5 \epsilon \\ a \sim \text{Uniform}(\mathcal{A}) & \text{with probability } 0.5 \epsilon \end{cases}

    This retains the downstream task optimization dynamics while biasing exploratory actions toward meaningful, contextually grounded behaviors learned during pretraining.

  6. Knowl 6 — Exploration Performance of ELLM in Crafter Pretraining

    empirical result

    In the Crafter survival environment with an expanded combinatorial action space of 260 actions (formed by verb + noun pairs like 'chop tree', 'eat plant', 'drink grass'), pretraining exploration quality is measured by the average number of unique achievements unlocked per episode without external task rewards:

    • Oracle (ground truth reward / context-sensitive valid goals): unlocks ≈9\approx 9 achievements per episode.
    • ELLM (goals) / ELLM (no goals): unlocks ≈6\approx 6 achievements per episode, discovering high-level prerequisites such as crafting tables and defending against enemies.
    • Novelty / Uniform baselines: achieve <3< 3 achievements per episode because exploration is wasted on nonsensical combinations (e.g., 'drink furnace', 'eat zombie').
    • Knowledge-based intrinsic motivation baselines (APT, RND, NovelD): achieve <3< 3 achievements per episode, demonstrating that optimizing solely for visual or state novelty in large combinatorial action spaces is insufficient to drive structured progress.
  7. Knowl 7 — Quality and Error Categorization of LLM-Generated Goals in Crafter

    data/table

    Evaluation of goal suggestions produced by Codex (few-shot prompted) in the Crafter environment categorized by context-sensitivity, common sense, and environmental affordances. Goals are broken down into total suggestions generated by the LLM versus those that were actually achieved and rewarded:

    Goal Category Suggested (%) Rewarded (%)
    Good (context-sensitive, common-sense, achievable) 64.9% 66.5%
    Context-Insensitive (e.g., make table without wood, mine stone without pickaxe) 13.6% 1.1%
    Common-Sense Insensitive (e.g., mine grass, make diamond, attack plant) 16.4% 32.4%
    Impossible under Crafter physics (e.g., make path, make wood, place lava) 5.0% 0.0%

    The majority (64.9%64.9\%) of suggested goals are valid, achievable, and context-appropriate. Because physical prerequisites prevent the completion of most impossible and context-insensitive actions, rewarded transitions are dominated by valid goals (66.5%66.5\%), filtering out invalid proposals during interaction.

  8. Knowl 8 — Downstream Task Performance on Crafter Benchmarks

    empirical result

    Downstream task evaluation in Crafter across seven sparse-reward tasks (Place Crafting Table, Attack Cow, Make Wood Sword, Mine Stone, Deforestation, Plant Row, Gardening) and the overall Crafter Game Score using guided exploration:

    • Performance Across Tasks: Goal-conditioned ELLM achieves the highest average success rate across tasks. Both goal-conditioned and goal-free ELLM are the only methods achieving non-zero success across all evaluated downstream tasks.
    • Comparison to Prior-Free Baselines: Novelty-only baselines (RND, APT) and training from scratch frequently fail to solve multi-step prerequisite tasks (e.g., Make Wood Sword, Mine Stone, Gardening), achieving near-zero success rates due to the vast combinatorial search space.
    • Goal Conditioning Advantage: While goal conditioning provides minimal benefit during pretraining exploration (since context alone reveals achievable actions), providing the target sequence of subgoals to goal-conditioned ELLM at test time significantly improves downstream transfer speed and convergence over unconditioned ELLM.
  9. Knowl 9 — Zero-Shot Goal Verification and Rearrangement in Housekeep

    empirical result

    In the Housekeep embodied robotics environment across 4 distinct multi-object rearrangement tasks (5 misplaced objects per room), LLM (InstructGPT text-davinci-002) classification accuracy and rearrangement pretraining performance demonstrate:

    Metric Task 1 Task 2 Task 3 Task 4
    Match Accuracy (True Positives) 85.7% 87.5% 50.0% 66.7%
    Mismatch Accuracy (True Negatives) 93.8% 90.1% 94.0% 87.6%

    InstructGPT accurately rejects misplaced object-receptacle combinations (>87%> 87\% mismatch accuracy across all tasks). In pretraining rearrangement success rate:

    • In Tasks 1 and 2 (where LLM match accuracy is >85%> 85\%), ELLM achieves pretraining rearrangement success rates substantially higher than RND, APT, and Novelty baselines.
    • In Tasks 3 and 4, lower true-positive accuracy reduces pretraining exploration efficiency, directly confirming that LLM goal precision governs pretraining coverage.
    • When fine-tuning or performing guided exploration on ground-truth human arrangements, ELLM matches or outperforms all baselines across all four tasks.
  10. Knowl 10 — Robustness to Imperfect Learned Transition Captioning

    empirical result

    When replacing the ground-truth programmatic transition captioner CtransitionC_{\text{transition}} in Crafter with a learned captioner trained via ClipCap (using a frozen CLIP ViT-B-32 visual prefix and a frozen GPT-2 language decoder trained on 847 human labels and 900 synthetic labels):

    • Error Characteristics of Learned Captioner: The captioner achieves an average false negative rate of 11%11\% (though reaching 100%100\% on 'chop grass' due to descriptive divergence like 'collect sapling'). However, the false positive rate is elevated when distinct achievements share descriptive words (e.g., 'wood', 'stone', 'collect').
    • Exploration Robustness: When using a cosine similarity threshold of T=0.5T = 0.5, ELLM pretraining performance with the learned ClipCap captioner experiences a moderate drop in achievements unlocked compared to ground-truth captioning, but continues to unlock ≈5\approx 5 achievements per episode, substantially outperforming prior-free intrinsic exploration baselines (RND, APT, Novelty <3< 3).
  11. Knowl 11 — Necessity of Episode-Level Novelty Filtering in ELLM

    empirical result

    Ablating the novelty bonus in ELLM—which removes already achieved goals within the current episode (the PREV_ACHIEVED\text{PREV\_ACHIEVED} filter)—causes catastrophic performance degradation during pretraining in both Crafter and Housekeep:

    • In Crafter, an ELLM agent trained without the intra-episode novelty filter repeatedly executes a small set of easily accessible, high-similarity goals (such as collecting saplings or chopping nearby grass) and fails to explore higher-order behaviors along the achievement hierarchy, plateauing at <3< 3 achievements per episode.
    • In Housekeep, removing the novelty filter causes the agent to repeatedly pick and place the same nearby object into the first valid receptacle rather than seeking out remaining misplaced objects.
  12. Knowl 12 — Limitations of ELLM

    limitation

    ELLM exhibits several operational and structural limitations:

    1. Prompt Sensitivity and False Negatives: LLM outputs are sensitive to prompt phrasing. Missing domain-specific common-sense knowledge can cause permanent false negatives where critical skills are never suggested (e.g., in Crafter, the LLM fails to suggest crafting a wooden pickaxe).
    2. Dependency on Captioning: ELLM requires state descriptions Cobs(o)C_{\text{obs}}(o) and transition descriptions Ctransition(o,a,o′)C_{\text{transition}}(o, a, o'). While achievable via ground truth in simulators, real-world deployment requires object detectors, instance segmenters, or vision-language models whose errors can introduce false-positive rewards.
    3. Applicability Scope: The method requires environments where exploratory goals and state information can be naturally articulated in natural language strings; it is less suited for continuous low-level motor control or fine-grained manipulation lacking semantic abstraction.
    4. API Latency and Financial Cost: Querying commercial LLM APIs repeatedly per environment step is compute- and cost-prohibitive without caching mechanisms.

Coverage note — None was omitted; all primary methods, algorithms, formulations, experimental results across Crafter and Housekeep, captioner analyses, ablations, and stated limitations are included.

References

  1. 1.Abid, A., Farooqi, M., and Zou, J. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. 298–306, 2021.
  2. 2.Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G. Deep reinforcement learning at the edge of the statistical precipice. Advances in Neural Information Processing Systems, 2021.
  3. 3.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., and Yan, M. Do as i can, not as i say: Grounding language in robotic affordances, 2022. URL https://arxiv.org/abs/2204.01691.
  4. 4.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022.
  5. 5.Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mane, D. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565, 2016.
  6. 6.Aubret, A., Matignon, L., and Hassas, S. A survey on intrinsic motivation in reinforcement learning. arXiv preprint arXiv:1908.06976, 2019.
  7. 7.Baranes, A. and Oudeyer, P.-Y. Active learning of inverse models with intrinsically motivated goal exploration in robots. Robotics and Autonomous Systems, 61(1):49–73, 2013.
  8. 8.Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. Unifying count-based exploration and intrinsic motivation. Advances in neural information processing systems, 29, 2016.
  9. 9.Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp. 610–623, 2021.
  10. 10.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020. URL https://arxiv.org/abs/2005.14165.
  11. 11.Burda, Y., Edwards, H., Storkey, A., and Klimov, O. Exploration by random network distillation. In Seventh International Conference on Learning Representations, pp. 1–17, 2019.
  12. 12.Chan, H., Wu, Y., Kiros, J., Fidler, S., and Ba, J. Actrce: Augmenting experience via teacher’s advice for multi-goal reinforcement learning. arXiv preprint arXiv:1902.04546, 2019.
  13. 13.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  14. 14.Choi, K., Cundy, C., Srivastava, S., and Ermon, S. LMPriors: Pre-trained language models as task-specific priors. arXiv preprint arXiv:2210.12530, 2022.
  15. 15.Colas, C., Sigaud, O., and Oudeyer, P.-Y. Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms. In International conference on machine learning, pp. 1039–1048. PMLR, 2018.
  16. 16.Colas, C., Karch, T., Lair, N., Dussoux, J.-M., Moulin-Frier, C., Dominey, P., and Oudeyer, P.-Y. Language as a cognitive tool to imagine goals in curiosity driven exploration. Advances in Neural Information Processing Systems, 33:3761–3774, 2020.
  17. 17.Colas, C., Karch, T., Sigaud, O., and Oudeyer, P.-Y. Autotelic agents with intrinsically motivated goal-conditioned reinforcement learning: a short survey. Journal of Artificial Intelligence Research, 74:1159–1199, 2022.
  18. 18.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  19. 19.Dubey, R., Agrawal, P., Pathak, D., Griffiths, T. L., and Efros, A. A. Investigating human priors for playing video games. arXiv preprint arXiv:1802.10217, 2018.
  20. 20.Hafner, D. Benchmarking the spectrum of agent capabilities. arXiv preprint arXiv:2109.06780, 2021.
  21. 21.Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., et al. Grounded language learning in a simulated 3d world. arXiv preprint arXiv:1706.06551, 2017.
  22. 22.Hill, F., Lampinen, A., Schneider, R., Clark, S., Botvinick, M., McClelland, J. L., and Santoro, A. Environmental drivers of systematicity and generalization in a situated agent. arXiv preprint arXiv:1910.00571, 2019.
  23. 23.Hill, F., Mokra, S., Wong, N., and Harley, T. Human instruction-following with deep reinforcement learning via transfer-learning from text. arXiv preprint arXiv:2005.09382, 2020.
  24. 24.Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. arXiv preprint arXiv:2201.07207, 2022a.
  25. 25.Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608, 2022b.
  26. 26.Kant, Y., Ramachandran, A., Yenamandra, S., Gilitschenski, I., Batra, D., Szot, A., and Agrawal, H. Housekeep: Tidying virtual households using commonsense reasoning. In Avidan, S., Brostow, G., Cisse, M., Farinella, G. M., and Hassner, T. (eds.), Computer Vision – ECCV 2022, pp. 355–373, Cham, 2022. Springer Nature Switzerland. ISBN 978-3-031-19842-7.
  27. 27.Kong, Y. and Fu, Y. Human action recognition and prediction: A survey. International Journal of Computer Vision, 130(5):1366–1401, 2022.
  28. 28.Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D. Reward design with language models. In International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=10uNUgI5Kl.
  29. 29.Ladosz, P., Weng, L., Kim, M., and Oh, H. Exploration in deep reinforcement learning: A survey. Information Fusion, 2022.
  30. 30.Lehman, J., Stanley, K. O., et al. Exploiting open-endedness to solve problems through the search for novelty. In ALIFE, pp. 329–336, 2008.
  31. 31.Lehman, J., Clune, J., Misevic, D., Adami, C., Altenberg, L., Beaulieu, J., Bentley, P. J., Bernard, S., Beslon, G., Bryson, D. M., Cheney, N., Chrabaszcz, P., Cully, A., Doncieux, S., Dyer, F. C., Ellefsen, K. O., Feldt, R., Fischer, S., Forrest, S., F´renoy, A., Gagne, C., Le Goff, L., Grabowski, L. M., Hodjat, B., Hutter, F., Keller, L., Knibbe, C., Krcah, P., Lenski, R. E., Lipson, H., MacCurdy, R., Maestre, C., Miikkulainen, R., Mitri, S., Moriarty, D. E., Mouret, J.-B., Nguyen, A., Ofria, C., Parizeau, M., Parsons, D., Pennock, R. T., Punch, W. F., Ray, T. S., Schoenauer, M., Schulte, E., Sims, K., Stanley, K. O., Taddei, F., Tarapore, D., Thibault, S., Watson, R., Weimer, W., and Yosinski, J. The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities. Artificial Life, 26(2):274–306, 05 2020. ISSN 1064-5462. doi: 10.1162/artl_a_00319. URL https://doi.org/10.1162/artl_a_00319.
  32. 32.Linke, C., Ady, N. M., White, M., Degris, T., and White, A. Adapting behavior via intrinsic reward: A survey and empirical study. Journal of Artificial Intelligence Research, 69:1287–1332, 2020.
  33. 33.Liu, H. and Abbeel, P. Behavior from the void: Unsupervised active pre-training. Advances in Neural Information Processing Systems, 34, 2021.
  34. 34.Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktaschel, T. A survey of reinforcement learning informed by natural language. arXiv preprint arXiv:1906.03926, 2019.
  35. 35.Lynch, C. and Sermanet, P. Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648, 2020.
  36. 36.Min, B., Ross, H., Sulem, E., Veyseh, A. P. B., Nguyen, T. H., Sainz, O., Agirre, E., Heinz, I., and Roth, D. Recent advances in natural language processing via large pre-trained language models: A survey. arXiv preprint arXiv:2111.01243, 2021.
  37. 37.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  38. 38.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. nature, 518(7540): 529–533, 2015.
  39. 39.Mokady, R., Hertz, A., and Bermano, A. H. Clipcap: Clip prefix for image captioning, 2021a. URL https://arxiv.org/abs/2111.09734.
  40. 40.Mokady, R., Hertz, A., and Bermano, A. H. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734, 2021b.
  41. 41.Mu, J., Zhong, V., Raileanu, R., Jiang, M., Goodman, N., Rocktaschel, T., and Grefenstette, E. Improving intrinsic exploration with language abstractions. arXiv preprint arXiv:2202.08938, 2022.
  42. 42.Nadeem, M., Bethke, A., and Reddy, S. Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456, 2020.
  43. 43.Oudeyer, P.-Y. and Kaplan, F. What is intrinsic motivation? a typology of computational approaches. Frontiers in neurorobotics, pp. 6, 2009.
  44. 44.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  45. 45.Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. Curiosity-driven exploration by self-supervised prediction. In International conference on machine learning, pp. 2778–2787. PMLR, 2017.
  46. 46.Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. Improving language understanding by generative pre-training. 2018.
  47. 47.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  48. 48.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  49. 49.Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. 11 2019. URL http://arxiv.org/abs/1908.10084.
  50. 50.Sharma, P., Torralba, A., and Andreas, J. Skill induction and planning with latent language. arXiv preprint arXiv:2110.01517, 2021.
  51. 51.Stanic, A., Tang, Y., Ha, D., and Schmidhuber, J. Learning to generalize with object-centric agents in the open world survival game crafter. arXiv preprint arXiv:2208.03374, 2022.
  52. 52.Stefanini, M., Cornia, M., Baraldi, L., Cascianelli, S., Fiameni, G., and Cucchiara, R. From show to tell: a survey on deep learning-based image captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  53. 53.Sutton, R. S., Barto, A. G., et al. Introduction to reinforcement learning, volume 135. MIT press Cambridge, 1998.
  54. 54.Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F. Semantic exploration from language abstractions and pretrained representations. arXiv preprint arXiv:2204.05080, 2022.
  55. 55.Ten, A., Oudeyer, P.-Y., and Moulin-Frier, C. Curiosity-driven exploration. The Drive for Knowledge: The Science of Human Information Seeking, pp. 53, 2022.
  56. 56.Van Hasselt, H., Guez, A., and Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI conference on artificial intelligence, volume 30, 2016.
  57. 57.Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. Dueling network architectures for deep reinforcement learning. In International conference on machine learning, pp. 1995–2003. PMLR, 2016.
  58. 58.Yao, S., Rao, R., Hausknecht, M., and Narasimhan, K. Keep calm and explore: Language models for action generation in text-based games. arXiv preprint arXiv:2010.02903, 2020.
  59. 59.Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. Reinforcement learning with prototypical representations. In International Conference on Machine Learning, pp. 11920–11931. PMLR, 2021.
  60. 60.Zaidi, S. S. A., Ansari, M. S., Aslam, A., Kanwal, N., Asghar, M., and Lee, B. A survey of modern deep learning based object detection models. Digital Signal Processing, pp. 103514, 2022.
  61. 61.Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J. E., and Tian, Y. Noveld: A simple yet effective exploration criterion. Advances in Neural Information Processing Systems, 34, 2021.

Citation

MLA
Du, Y., et al. “Guiding Pretraining in Reinforcement Learning with Large Language Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 8657–77, https://proceedings.mlr.press/v202/du23f.html.
APA
Du, Y., Watkins, O., Wang, Z., Colas, C., Darrell, T., Abbeel, P., Gupta, A., & Andreas, J. (2023). Guiding Pretraining in Reinforcement Learning with Large Language Models. International Conference on Machine Learning, 202, 8657–8677. https://proceedings.mlr.press/v202/du23f.html
Chicago
Du, Y., O. Watkins, Z. Wang, et al. 2023. “Guiding Pretraining in Reinforcement Learning with Large Language Models”. International Conference on Machine Learning 202: 8657–77. https://proceedings.mlr.press/v202/du23f.html.
Harvard
Du, Y. et al. (2023) “Guiding Pretraining in Reinforcement Learning with Large Language Models”, International Conference on Machine Learning. PMLR, pp. 8657–8677. Available at: https://proceedings.mlr.press/v202/du23f.html.
Vancouver
1. Du Y, Watkins O, Wang Z, Colas C, Darrell T, Abbeel P, Gupta A, Andreas J (2023) Guiding Pretraining in Reinforcement Learning with Large Language Models. In: International Conference on Machine Learning. PMLR, pp 8657–8677

BibTeX

@InProceedings{pmlr-v202-du23f,
  title = 	 {Guiding Pretraining in Reinforcement Learning with Large Language Models},
  author =       {Du, Yuqing and Watkins, Olivia and Wang, Zihan and Colas, C\'{e}dric and Darrell, Trevor and Abbeel, Pieter and Gupta, Abhishek and Andreas, Jacob},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {8657--8677},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/du23f/du23f.pdf},
  url = 	 {https://proceedings.mlr.press/v202/du23f.html},
  abstract = 	 {Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivated exploration methods address this limitation by rewarding agents for visiting novel states or transitions, but these methods offer limited benefits in large environments where most discovered novelty is irrelevant for downstream tasks. We describe a method that uses background knowledge from text corpora to shape exploration. This method, called ELLM (Exploring with LLMs) rewards an agent for achieving goals suggested by a language model prompted with a description of the agent’s current state. By leveraging large-scale language model pretraining, ELLM guides agents toward human-meaningful and plausibly useful behaviors without requiring a human in the loop. We evaluate ELLM in the Crafter game environment and the Housekeep robotic simulator, showing that ELLM-trained agents have better coverage of common-sense behaviors during pretraining and usually match or improve performance on a range of downstream tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/