Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning

Thomas CartaClément RomacThomas WolfSylvain LamprierOlivier SigaudPierre-Yves Oudeyer

article2023ICML255 citations

Introduces GLAM, a method that uses online reinforcement learning to functionally ground large language models in interactive environments, improving sample efficiency and policy generalization across decision-making tasks.

Listen

Large Language Models possess vast statistical knowledge about the world, yet they frequently fail in interactive, goal-oriented settings because their internal representations lack functional grounding—the alignment between language and external physical or spatial dynamics. In real-world applications such as robotics and virtual assistants, this misalignment leads to poor decision-making and unreliable execution. While standard techniques like human feedback align models with conversational preferences, they do not ground models in environments where actions alter physical states. The article introduces and evaluates Grounded Language Models (GLAM), a framework designed to functionally ground language models directly as interactive policies using online Reinforcement Learning.

The research evaluated how using online Reinforcement Learning (specifically Proximal Policy Optimization) enables language models to learn interactive tasks, adapt to unfamiliar objects, generalize to new instructions, and outperform offline imitation methods. The authors developed BabyAI-Text, a procedurally generated, text-based navigation and reasoning platform derived from the BabyAI benchmark. In this environment, the agent receives natural language observations detailing spatial relationships, evaluates candidate actions using the internal likelihood scores of a pretrained model (FLAN-T5, primarily the 780-million parameter version), and refines its policy using sparse reward signals collected through environmental interaction.

The findings demonstrate substantial operational advantages for this approach. First, the functionally grounded model showed high learning efficiency, achieving an 80% success rate within 250,000 steps and reaching 90% around 600,000 steps across mixed tasks. In contrast, standard reinforcement learning baselines and non-pretrained models failed to reach a 20% success rate even after 1.5 million steps. Second, the model proved robust to task complexity, maintaining steady performance when irrelevant distractor objects increased from 4 to 16 and when irrelevant choices were added to the action set. Third, the grounded model exhibited strong zero-shot generalization to novel and invented object names (retaining an 87% to 88% success rate), confirming it grounded spatial relationships and geometry rather than memorizing specific items. Finally, interactive online learning consistently outperformed offline Behavioral Cloning—even when imitation data came from a flawless automated bot—by allowing the model to recover from errors through trial-and-error intervention.

These results indicate that pre-existing model knowledge provides a valuable foundation that eliminates the need to train policies from scratch, significantly reducing the data and time required for autonomous agents to become proficient. However, the study also revealed clear boundaries: models showed limited generalization when action verbs were replaced with synonyms (dropping to a 12% success rate) and failed entirely when instructions were translated into another language (dropping to 2%). These failures show that functional grounding remains tied to the specific vocabulary experienced during interactive training. Additionally, smaller models (80 million parameters) failed to display these sample-efficiency benefits, indicating that a sufficient parameter scale is required for grounded behavior to emerge.

Before deploying these systems in high-stakes or physical applications, organizations should treat this approach as an experimental foundation rather than a production-ready solution. Practical scaling faces high computational costs during action evaluation, which the authors mitigated using a distributed infrastructure library called Lamorel. Future work must validate these techniques in complex multi-modal domains (such as vision and robotics), improve computational efficiency for large action spaces, and develop strategies to ensure grounded knowledge transfers across different languages and varied action phrasing.

arXiv: 2302.02662
Cover for Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning

Abstract

Recent works successfully leveraged Large Language Models' (LLM) abilities to capture abstract knowledge about world's physics to solve decision-making problems. Yet, the alignment between LLMs' knowledge and the environment can be wrong and limit functional competence due to lack of grounding. In this paper, we study an approach (named GLAM) to achieve this alignment through functional grounding: we consider an agent using an LLM as a policy that is progressively updated as the agent interacts with the environment, leveraging online Reinforcement Learning to improve its performance to solve goals. Using an interactive textual environment designed to study higher-level forms of functional grounding, and a set of spatial and navigation tasks, we study several scientific questions: 1) Can LLMs boost sample efficiency for online learning of various RL tasks? 2) How can it boost different forms of generalization? 3) What is the impact of online learning? We study these questions by functionally grounding several variants (size, architecture) of FLAN-T5.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 GLAM: Grounding LLMs with online RL
  • 3.1 Problem statement
  • 3.2 LLMs as policies in interactive environments
  • 3.3 PPO finetuning
  • 3.4 Distributed LLM policies using Lamorel
  • 4 Experiments
  • 4.1 How fast can an LLM adapt and learn to solve tasks? (Q1)
  • 4.1.1 Impact of the dimension of the action space
  • 4.1.2 Impact of the number of distractors
  • 4.2 Q2. Generalization to new objects
  • 4.3 Q3. Generalization to new tasks
  • 4.4 What is the impact of using RL vs Behavioral Cloning for grounding? (Q4)
  • 5 Conclusion
  • References
  • A Environments
  • A.1 BabyAI
  • A.2 BabyAI-Text
  • B Additional results
  • B.1 Per-task success rate
  • B.2 Averaging success rate over the tasks
  • B.3 Textual vs symbolic representation
  • B.4 Impact of pretraining
  • B.5 Impact of the size of the LLM
  • B.6 Impact of varying action space and distractors
  • B.6.1 Impact of the dimension of the action space
  • B.6.2 Impact of the number of distractors
  • B.7 Robustness to domain-specific vocabulary
  • C Evolution of actions distribution on evaluation prompts
  • D Generalization tests details
  • D.1 Recapitulating results table
  • D.2 Complementary tests for Q2
  • D.3 Complementary tests for Q3
  • D.4 LLM grounding of temporal symbols: "then" and "after"
  • E Distributed experimental setup
  • F finetuning details
  • F.1 PPO finetuning details
  • F.2 Behavioral Cloning
  • G Confidence interval
  • G.1 Confidence intervals for GFlan-T5, Flan-T5 and DRRN
  • G.2 Confidence intervals for random agents
  • H Word substitutions for generalization tests
  • H.1 Out of vocabulary
  • H.2 Invented words
  • H.3 Synonym actions
  • H.4 Translation to French

Knowls

  1. Knowl 1 — Grounded Language Models (GLAM) Framework

    model/method

    The Grounded Language Models (GLAM) framework uses a pretrained Large Language Model (LLM) directly as an agent policy in an interactive textual reinforcement learning setting. The environment is formalized as a goal-augmented Partially Observable Markov Decision Process (POMDP) M=(S,V,A,T,R,G,O,γ)\mathcal{M} = (\mathcal{S}, \mathcal{V}, \mathcal{A}, \mathcal{T}, \mathcal{R}, \mathcal{G}, \mathcal{O}, \gamma), where S\mathcal{S} is the state space, V\mathcal{V} is the language vocabulary, A⊂VN\mathcal{A} \subset \mathcal{V}^N is the textual action space, G⊂VN\mathcal{G} \subset \mathcal{V}^N is the goal space, T:S×A→S\mathcal{T}: \mathcal{S} \times \mathcal{A} \to \mathcal{S} is the state transition function, R:S×A×G→R\mathcal{R}: \mathcal{S} \times \mathcal{A} \times \mathcal{G} \to \mathbb{R} is the goal-conditioned reward function, O:S→VN\mathcal{O}: \mathcal{S} \to \mathcal{V}^N is the observation function mapping state to text, and γ∈[0,1)\gamma \in [0, 1) is the discount factor.

    At decision step tt, a textual prompt ptp_t is constructed containing the list of accessible actions, the target goal gg, and a short-term history buffer of the previous 3 observations and 2 actions. The LLM policy π^(a∣o,g)\hat{\pi}(a \mid o, g) evaluates candidate action strings ai∈Aa_i \in \mathcal{A} via autoregressive conditional likelihood. To approximate the state-value function V^(o∣g)=V^(p)≈Ea∼π^[R(s,g,a)+γV(T(s,a),g)]\hat{V}(o \mid g) = \hat{V}(p) \approx \mathbb{E}_{a \sim \hat{\pi}}[\mathcal{R}(s, g, a) + \gamma V(\mathcal{T}(s, a), g)], a multi-layer perceptron (MLP) value head is attached to the final layer of the first Decoder block. The entire network (LLM backbone and value head) is finetuned online using Proximal Policy Optimization (PPO) using sparse environment rewards.

  2. Knowl 2 — Action Probability Estimation via Pretrained Language Model Heads

    equation

    For an agent policy parameterized by a Large Language Model with prompt pp and a discrete textual action space A\mathcal{A}, the policy avoids adding separate task-specific classification heads by directly querying the pretrained language modeling heads. For an action ai∈Aa_i \in \mathcal{A} represented as a sequence of tokens ai={w0,…,w∣ai∣}a_i = \{w_0, \dots, w_{|a_i|}\}, the conditional log-likelihood under the LLM is computed as:

    LPLLM(ai∣p)=∑j=0∣ai∣log⁡PLLM(wj∣p,w<j)LP_{\text{LLM}}(a_i \mid p) = \sum_{j=0}^{|a_i|} \log P_{\text{LLM}}(w_j \mid p, w_{<j})

    where PLLM(wj∣p,w<j)P_{\text{LLM}}(w_j \mid p, w_{<j}) denotes the probability assigned to token wjw_j conditioned on prompt pp and preceding action tokens w<jw_{<j}. The action selection distribution over A\mathcal{A} is obtained via softmax normalization:

    P(ai∣p)=exp⁡(LPLLM(ai∣p))∑aj∈Aexp⁡(LPLLM(aj∣p))\mathbb{P}(a_i \mid p) = \frac{\exp(LP_{\text{LLM}}(a_i \mid p))}{\sum_{a_j \in \mathcal{A}} \exp(LP_{\text{LLM}}(a_j \mid p))}

    This formulation eliminates ad-hoc token-to-action string mapping, preserves the language model's pretrained semantic priors, and operates without structural modification across arbitrary discrete textual action spaces.

  3. Knowl 3 — BabyAI-Text Interactive Environment

    experimental setup

    BabyAI-Text is a text-only procedural reinforcement learning environment adapted from the MiniGrid-based BabyAI benchmark. An agent interacts in procedurally generated 8×88 \times 8 grid rooms containing distractor objects, keys, doors, balls, and boxes instantiated in 6 colors (red, blue, green, yellow, grey, purple).

    The agent receives a partial view of its 6×66 \times 6 forward visual field translated into template-based natural language descriptions:

    • "You see a <object> <location>" for objects and walls (with walls described at the closest coordinate).
    • "You see a(n) open/closed door <location>" for doors.
    • "You carry a <object>" when holding an entity.

    Relative coordinates <location> specify step counts along grid axes relative to agent orientation (e.g., "2 steps left and 1 step forward").

    The action space comprises 6 commands: turn left, turn right, go forward, pick up, drop, and toggle. When the agent fulfills the language goal in NN steps within horizon HH, it receives a sparse scalar reward rN=1−0.9NHr_N = 1 - 0.9 \frac{N}{H}, which is multiplied by a scaling factor of 20 during training (and 0 for non-terminal steps).

  4. Knowl 4 — Lamorel Distributed Framework for LLM Reinforcement Learning

    model/method

    In language-conditioned online RL with EE parallel environments and a discrete action space A\mathcal{A}, computing action distributions autoregressively requires E×∣A∣E \times |\mathcal{A}| forward passes per environment step. For large language models, this creates an execution bottleneck.

    The Lamorel framework resolves this via a distributed client-server architecture. The reinforcement learning training script acts as a client sending state-action requests to a master server, which dispatches inference over NN parallel LLM workers via PyTorch Distributed using the GLOO communication backend. Action scoring achieves quasi-linear speedup with respect to worker count NN. During policy updates, Lamorel dispatches forward and backward passes across workers in a Distributed Data Parallel (DDP) manner, aggregates gradients across workers, and updates both the LLM backbone and attached custom heads (e.g., MLP value heads).

  5. Knowl 5 — Sample Efficiency and Robustness of Online Grounded LLMs

    empirical result

    GFlan-T5 (Flan-T5-780M grounded via GLAM) was evaluated across multi-task BabyAI-Text goals (Go to <object>, Pick up <object>, Put <object A> next to <object B>, Pick up <object A> then go to <object B>, and Unlock <door>). GFlan-T5 achieves an average success rate of 0.8 in 250,000 environment steps and 0.9 in ~600,000 steps.

    In comparison, after 1,500,000 environment steps:

    • DRRN (~1M parameter baseline) achieves a success rate under 0.2.
    • NPAE-Flan-T5 (Flan-T5 with only pretrained embeddings and randomly initialized weights/action heads) remains under 0.2.
    • Symbolic-PPO (a PPO agent trained on BabyAI's native 3×6×63 \times 6 \times 6 symbolic matrix inputs) reaches ~0.4.

    In robustness analyses on Go To <object> using the sample efficiency metric SE=1T∑t=0TSRtSE = \frac{1}{T} \sum_{t=0}^T SR_t:

    • Action space scaling: Expanding the action space from 3 useful actions to 6 actions (3 useful, 3 useless) and 9 actions (3 useful, 6 useless: sleep, do nothing, think) causes no performance degradation for GFlan-T5, whereas baselines degrade substantially.
    • Distractor scaling: Increasing distractors from 4 to 16 reduces GFlan-T5 success rate by only 14%, compared to a 38% drop for Symbolic-PPO.
  6. Knowl 6 — Impact of Model Scale and Action Head Architecture on Grounding

    empirical result

    Ablation experiments on BabyAI-Text's Go To <object> task isolate the roles of parameter count, pretraining, and action scoring architecture:

    1. Model Scale: Pretrained prior knowledge provides emergent grounding efficiency as model scale increases. GFlan-T5-Small (80M parameters) exhibits low sample efficiency over 400,000 steps, while GFlan-T5-Large (780M) and GFlan-T5-XL (3B) rapidly achieve asymptotic success rates above 0.9 within 150,000 to 200,000 steps.

    2. Action Architecture: Utilizing the LLM's pretrained language modeling heads to score action sequences (GFlan-T5) substantially outperforms replacing the LM head with a separate MLP action head over the LLM trunk (AFlan-T5). AFlan-T5 requires ~250,000 environment steps to surpass non-pretrained baselines because the trunk representations must realign to the new head.

    3. Pretraining Dependency: Without pretraining the LLM trunk, LM head scoring (NPE-Flan-T5) fails to learn entirely, because gradients from the small set of active action tokens cannot effectively optimize the 32,000-token language head. In contrast, randomly initialized architectures with specialized action heads (NPAE-Flan-T5 and NPA-Flan-T5) learn, though at much lower sample efficiency than fully pretrained models.

  7. Knowl 7 — Zero-Shot Generalization to Out-of-Vocabulary and Invented Objects

    data/table

    Zero-shot generalization across 1,000 test episodes evaluated over 2 random seeds (reported as mean success rate ±\pm 99% confidence interval). Agents were trained on a mixture of 5 tasks and evaluated on environments where object and color tokens were replaced with out-of-vocabulary words or invented pseudo-words.

    Environment GFlan-T5 Flan-T5 (zero-shot) NPAE DRRN Random
    Mix - no change 0.89 0.05 0.11 0.03 0.17 0.04 0.14 0.02 0.15 0.05
    Mix - out-of-vocabulary nouns 0.87 0.05 0.09 0.02 0.16 0.05 0.15 0.00 0.15 0.05
    Mix - invented nouns and adjectives 0.88 0.06 0.11 0.03 0.16 0.03 0.16 0.00 0.15 0.05
    Mix - unseen in-vocabulary objects 0.87 0.03 0.12 0.03 0.16 0.08 0.17 0.09 0.15 0.05
    Mix - out-of-vocabulary adjectives 0.87 0.07 0.16 0.03 0.16 0.02 0.16 0.01 0.15 0.05

    GFlan-T5 maintains performance (0.87–0.88 vs 0.89 baseline), demonstrating that online functional grounding aligns the underlying relative spatial geometry and instructions (e.g., "steps", "left", "forward") rather than overfitting to specific object identity tokens.

  8. Knowl 8 — Generalization Bounds on Action Synonyms, Task Composition, and Cross-Lingual Transfer

    data/table

    Zero-shot transfer of GFlan-T5 was evaluated on novel compositions, action synonym substitutions, and language translation (1,000 test episodes, 99% confidence interval across 2 random seeds).

    Task / Modification GFlan-T5 Flan-T5 (zero-shot) NPAE DRRN Random
    Pick up then/after pick up (New composition) 0.12 0.06 0.02 0.00 0.06 0.01 0.06 0.03 0.05 0.05
    Mix - synonym actions 0.12 0.12 0.02 0.00 0.16 0.04 0.17 0.04 0.15 0.05
    Go To - English 0.99 0.01 0.27 0.03 0.31 0.04 0.31 0.03 0.30 0.05
    Go To - French 0.02 0.01 0.03 0.00 0.30 0.02 0.31 0.02 0.30 0.05
    Go To - English with actions in French 0.15 0.04 0.26 0.02 0.31 0.01 0.33 0.00 0.30 0.05

    Key behavioral findings include:

    • Action synonyms (e.g., replacing "go forward" with "move ahead") lead to an 87% relative performance drop, indicating action-vocabulary overfitting during finetuning.
    • Full French translation collapses performance to 0.02 (worse than random chance 0.30), showing functional grounding fails when all symbols in the prompt are shifted simultaneously.
    • Temporal connective grounding is asymmetric: tasks specifying chronological order with "then" achieve a 0.22 ±\pm 0.11 success rate, compared to 0.17 ±\pm 0.05 for inverted order with "after".
  9. Knowl 9 — Online RL Grounding vs. Offline Behavioral Cloning

    data/table

    A comparative evaluation on the Go To task of online RL grounding (GFlan-T5 trained over 400,000 steps) versus offline Behavioral Cloning (BC) trained on 400,000 transitions. BC-GFlan-T5 was trained on trajectories generated by a finetuned GFlan-T5, while BC-Bot was trained on expert trajectories from the BabyAI procedural bot.

    Environment GFlan-T5 BC-GFlan-T5 BC-Bot Random
    Go To task (no change) 0.82 0.02 0.69 0.08 0.73 0.07 0.30 0.05
    Go To task (invented words) 0.74 0.004 0.70 0.07 0.63 0.08 -

    Online RL grounding through trial-and-error interactions and interventions achieves higher absolute task performance (0.82 vs 0.73) and exhibits greater robustness against out-of-distribution lexical shifts (dropping by only 0.08 on invented words compared to a 0.10 drop for BC-Bot).

  10. Knowl 10 — Online Grounding Adaptability Under Inverted Action Semantics

    empirical result

    When the semantic rules of an environment conflict with natural language pretraining priors, GLAM adapts and overwrites its prior bias through online interaction. In an experimental condition where the physical transitions of the actions turn left and turn right were inverted in the BabyAI-Text Go To task (i.e., executing turn left turns the agent rightward in the grid world), GFlan-T5 initially exhibited lower success rates during early training episodes due to misleading language priors. However, through trial-and-error RL updates, GFlan-T5 learned the inverted dynamics and converged to an asymptotic success rate (~0.9) at a training speed comparable to training in the canonical environment.

Coverage note — None was omitted; all key contributions including the GLAM framework, action probability scoring formulation, BabyAI-Text environment, Lamorel distributed framework, sample efficiency results, ablations on architecture and model scale, out-of-vocabulary and task generalization tests, behavioral cloning comparisons, and semantic inversion tests are covered.

References

  1. 1.Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard. Can language models encode perceptual structure without grounding? a case study in color. In Proceedings of the 25th Conference on Computational Natural Language Learning (CoNLL),, 2021.
  2. 2.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alexander Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil Jayant Joshi, Ryan C. Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego M Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, and Mengyuan Yan. Do as i can, not as i say: Grounding language in robotic affordances. ArXiv, abs/2204.01691, 2022.
  3. 3.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. ArXiv, abs/2204.14198, 2022.
  4. 4.Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Seyedarian Hosseini, Pushmeet Kohli, and Edward Grefenstette. Learning to understand goal specifications by modelling reward. In International Conference on Learning Representations, 2018.
  5. 5.Emily M. Bender and Alexander Koller. Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.463.
  6. 6.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. Experience Grounds Language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8718–8735, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.703.
  7. 7.Gemma Boleda. Distributional semantics and linguistic theory. ArXiv, abs/1905.01896, 2019.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. ArXiv, abs/2005.14165, 2020.
  9. 9.Angelo Cangelosi and Francesca Stramandinoli. A review of abstract concept learning in embodied agents and robots. Philosophical Transactions of the Royal Society B: Biological Sciences, 373 (1752):20170131, June 2018. doi: 10.1098/rstb.2017.0131. Publisher: Royal Society.
  10. 10.Angelo Cangelosi, Giorgio Metta, Gerhard Sagerer, Stefano Nolfi, Chrystopher Nehaniv, Kerstin Fischer, Jun Tani, Tony Belpaeme, Giulio Sandini, Francesco Nori, et al. Integration of action and language knowledge: A roadmap for developmental robotics. IEEE Transactions on Autonomous Mental Development, 2(3):167–195, 2010.
  11. 11.Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio. Babyai: A platform to study the sample efficiency of grounded language learning. In International Conference on Learning Representations (ICLR), 2019.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. ArXiv, abs/2204.02311, 2022.
  13. 13.Cédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux, Clément Moulin-Frier, Peter Ford Dominey, and Pierre-Yves Oudeyer. Language as a cognitive tool to imagine goals in curiosity-driven exploration. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  14. 14.Marc-Alexandre Côté, Ákos Kádár, Xingdi Yuan, Ben A. Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew J. Hausknecht, Layla El Asri, Mahmoud Adada, Wendy Tay, and Adam Trischler. Textworld: A learning environment for text-based games. In CGW@IJCAI, 2018.
  15. 15.Ishita Dasgupta, Christine Kaeser Chen, Kenneth Marino, Arun Ahuja, Sheila Babayan, Felix Hill, and Rob Fergus. Collaborating with language models for embodied reasoning. In Advances in Neural Information Processing Systems (NeurIPS) LaReL workshop, 2022.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. ArXiv, abs/1810.04805, 2019.
  17. 17.Linxi (Jim) Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar. Minedojo: Building open-ended embodied agents with internet-scale knowledge. ArXiv, abs/2206.08853, 2022.
  18. 18.Tarun Gupta, Peter Karkus, Tong Che, Danfei Xu, and Marco Pavone. Foundation models for semantic novelty in reinforcement learning. ArXiv, abs/2211.04878, 2022.
  19. 19.Steve Harnad. The symbol grounding problem. Physica D, 42:335–346, 1990.
  20. 20.Zellig S. Harris. Distributional Structure. In Zellig S. Harris and Henry Hiz,˙ editors, Papers on Syntax, Synthese Language Library, pages 3–22. Springer Netherlands, Dordrecht, 1981. ISBN 978-94-009-8467-7. doi: 10.1007/978-94-009-8467-7_1.
  21. 21.Ji He, Mari Ostendorf, Xiaodong He, Jianshu Chen, Jianfeng Gao, Lihong Li, and Li Deng. Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads. arXiv:1606.03667 [cs], September 2016. URL http://arxiv.org/abs/1606.03667. arXiv: 1606.03667.
  22. 22.Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojciech M. Czarnecki, Max Jaderberg, Denis Teplyashin, Marcus Wainwright, Chris Apps, Demis Hassabis, and Phil Blunsom. Grounded language learning in a simulated 3d world. ArXiv, abs/1706.06551, 2017.
  23. 23.Felix Hill, Sona Mokra, Nathaniel Wong, and Tim Harley. Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text. arXiv:2005.09382 [cs], May 2020. arXiv: 2005.09382.
  24. 24.Wenlong Huang, P. Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. ArXiv, abs/2201.07207, 2022a.
  25. 25.Wenlong Huang, F. Xia, Ted Xiao, Harris Chan, Jacky Liang, Peter R. Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter. Inner monologue: Embodied reasoning through planning with language models. ArXiv, abs/2207.05608, 2022b.
  26. 26.Peter A. Jansen. A Systematic Survey of Text Worlds as Embodied Natural Language Environments. arXiv:2107.04132 [cs], July 2021. arXiv: 2107.04132.
  27. 27.Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi (Jim) Fan. Vima: General robot manipulation with multimodal prompts. ArXiv, abs/2210.03094, 2022.
  28. 28.Jared Kaplan, Sam McCandlish, T. J. Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. Scaling laws for neural language models. ArXiv, abs/2001.08361, 2020.
  29. 29.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014.
  30. 30.Shuang Li, Xavier Puig, Yilun Du, Clinton Jia Wang, Ekin Akyürek, Antonio Torralba, Jacob Andreas, and Igor Mordatch. Pre-trained language models for interactive decision-making. ArXiv, abs/2202.01771, 2022.
  31. 31.J. Liang, Wenlong Huang, F. Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R. Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. ArXiv, abs/2209.07753, 2022.
  32. 32.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal Tasks. February 2023. URL https://openreview.net/forum?id=E01k9048soZ.
  33. 33.Yao Lu, Karol Hausman, Yevgen Chebotar, Mengyuan Yan, Eric Jang, Alexander Herzog, Ted Xiao, Alex Irpan, Mohi Khansari, Dmitry Kalashnikov, and Sergey Levine. Aw-opt: Learning robotic skills with imitation and reinforcement at scale. In Conference on Robot Learning, 2021.
  34. 34.Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob N. Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel. A survey of reinforcement learning informed by natural language. In International Joint Conference on Artificial Intelligence, 2019.
  35. 35.Kyle Mahowald, Anna A. Ivanova, Idan Asher Blank, Nancy G. Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. Dissociating language and thought in large language models: a cognitive perspective. ArXiv, abs/2301.06627, 2023.
  36. 36.Suvir Mirchandani, Siddharth Karamcheti, and Dorsa Sadigh. Ella: Exploration through learned language abstraction. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  37. 37.Dipendra Kumar Misra, John Langford, and Yoav Artzi. Mapping instructions and visual observations to actions with reinforcement learning. In Conference on Empirical Methods in Natural Language Processing, 2017.
  38. 38.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022.
  39. 39.Roma Patel and Elizabeth-Jane Pavlick. Mapping language models to grounded conceptual spaces. In International Conference on Learning Representations, 2022.
  40. 40.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In arXiv preprint arXiv:2103.00020v1, 2021.
  41. 41.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John F. J. Mellor, Irina Higgins, Antonia Creswell, Nathan McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, L. Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, N. K. Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Tobias Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew G. Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem W. Ayoub, Jeff Stanway, L. L. Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training gopher. ArXiv, abs/2112.11446, 2021.
  42. 42.Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization. ArXiv, abs/2210.01241, 2022.
  43. 43.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. ArXiv, abs/2102.12092, 2021.
  44. 44.Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. Can wikipedia help offline reinforcement learning? ArXiv, abs/2201.12122, 2022.
  45. 45.Teven Le Scao, Angela Fan, Christopher Akiki, Elizabeth-Jane Pavlick, Suzana Ili’c, Daniel Hesslow, Roman Castagn’e, Alexandra Sasha Luccioni, Franccois Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Rose Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurenccon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa Etxabe, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris C. Emezue, Christopher Klamm, Colin Leong, Daniel Alexander van Strien, David Ifeoluwa Adelani, Dragomir R. Radev, Eduardo G. Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady ElSahar, Hamza Benyamina, Hieu Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jorg Frohberg, Josephine L. Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro von Werra, Leon Weber, Long Phan, Loubna Ben Allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, Mar’ia Grandury, Mario vSavsko, Max Huang, Maximin Coavoux, Mayank Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad Ali Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto L’opez, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, S. Longpre, Somaieh Nikpoor, Stanislav Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-shaibani, Matteo Manica, Nihal V. Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Févry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiang Tang, Zheng Xin Yong, Zhiqing Sun, Shaked Brody, Y Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre Franccois Lavall’ee, Rémi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, Stéphane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aur’elie N’ev’eol, Charles Lovering, Daniel H Garrette, Deepak R. Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Jordan Clive, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, S. Osher Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdenvek Kasner, Alice Rueda, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ananda Santa Rosa Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Olusola Ajibade, Bharat Kumar Saxena, Carlos Muñoz Ferrandis, Danish Contractor, David M. Lansky, Davis David, Douwe Kiela, Duong Anh Nguyen, Edward Tan, Emily Baylor, Ezinwanne Ozoani, Fatim T Mirza, Frankline Ononiwu, Habib Rezanejad, H.A. Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jan Passmore, Joshua Seltzer, Julio Bonis Sanz, Karen Fort, Lívia Macedo Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, M. K. K. Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nourhan Fahmy, Olanrewaju Modupe Samuel, Ran An, R. P. Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas L. Wang, Sourav Roy, Sylvain Viguier, Thanh-Cong Le, Tobi Oyebade, Trieu Nguyen Hai Le, Yoyo Yang, Zachary Kyle Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Kumar Singh, Benjamin Beilharz, Bo Wang, Caio Matheus Fonseca de Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Clémentine Fourrier, Daniel Le’on Perin’an, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully A. Burns, Helena U. Vrabec, Iman I.B. Bello, Isha Dash, Ji Soo Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthi Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc Pàmies, María Andrea Castillo, Marianna Nezhurina, Mario Sanger, Matthias Samwald, Michael Cullan, Michael Weinberg, M Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patricia Haller, R. Chandrasekhar, R. Eisenberg, Robert Martin, Rodrigo L. Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sang-aroonsiri, Srishti Kumar, Stefan Schweter, Sushil Pratap Bharati, T. A. Laud, Th’eo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yashasvi Bajaj, Y. Venkatraman, Yifan Xu, Ying Xu, Yun chao Xu, Zhee Xao Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, and Thomas Wolf. Bloom: A 176b-parameter open-access multilingual language model. ArXiv, abs/2211.05100, 2022.
  46. 46.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. ArXiv, abs/1707.06347, 2017.
  47. 47.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew J. Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. ArXiv, abs/2010.03768, 2020.
  48. 48.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan J. Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback. ArXiv, abs/2009.01325, 2020.
  49. 49.Shiro Takagi. On the effect of pre-training for transformer in different modality on offline reinforcement learning. ArXiv, abs/2211.09817, 2022.
  50. 50.Serge Thill, Sebastia Padó n, and Tom Ziemke. On the importance of a rich embodiment in the grounding of concepts: Perspectives from embodied cognitive science and computational linguistics. Topics in Cognitive Science, 6(3):545–558, 2014. doi: https://doi.org/10.1111/tops.12093.
  51. 51.Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. Large language models still can’t plan (a benchmark for llms on planning and reasoning about change). ArXiv, abs/2206.10498, 2022.
  52. 52.Ruoyao Wang, Peter Alexander Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. Scienceworld: Is your agent smarter than a 5th grader? ArXiv, abs/2203.07540, 2022.
  53. 53.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed Huai hsin Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. ArXiv, abs/2206.07682, 2022.
  54. 54.Peratham Wiriyathammabhum, Douglas Summers-Stay, Cornelia Fermüller, and Yiannis Aloimonos. Computer vision and natural language processing: Recent approaches in multimedia and robotics. ACM Comput. Surv., 49(4), dec 2016. ISSN 0360-0300. doi: 10.1145/3009906.
  55. 55.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. ArXiv, abs/2210.03629, 2022.
  56. 56.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel C. F. Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang. Florence: A new foundation model for computer vision. ArXiv, abs/2111.11432, 2021.
  57. 57.Qinqing Zheng, Amy Zhang, and Aditya Grover. Online Decision Transformer. In Proceedings of the 39th International Conference on Machine Learning, pages 27042–27059. PMLR, June 2022. URL https://proceedings.mlr.press/v162/zheng22c.html. ISSN: 2640-3498.

Citation

MLA
Carta, T., et al. “Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning”. PMLR 202 (2023):3676-3713, 2023, http://arxiv.org/abs/2302.02662v5.
APA
Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., & Oudeyer, P.-Y. (2023). Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning. PMLR 202 (2023):3676-3713. http://arxiv.org/abs/2302.02662v5
Chicago
Carta, T., C. Romac, T. Wolf, S. Lamprier, O. Sigaud, and P.-Y. Oudeyer. 2023. “Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning”. PMLR 202 (2023):3676-3713. http://arxiv.org/abs/2302.02662v5.
Harvard
Carta, T. et al. (2023) “Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning”, PMLR 202 (2023):3676-3713 [Preprint]. Available at: http://arxiv.org/abs/2302.02662v5.
Vancouver
1. Carta T, Romac C, Wolf T, Lamprier S, Sigaud O, Oudeyer P-Y (2023) Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning. PMLR 202 (2023):3676-3713

BibTeX

@article{carta2023grounding,
  title = {Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning},
  author = {Carta, Thomas and Romac, Clément and Wolf, Thomas and Lamprier, Sylvain and Sigaud, Olivier and Oudeyer, Pierre-Yves},
  year = {2023},
  journal = {PMLR 202 (2023):3676-3713},
  url = {http://arxiv.org/abs/2302.02662v5},
  eprint = {2302.02662}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/