Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
Thomas CartaClément RomacThomas WolfSylvain LamprierOlivier SigaudPierre-Yves Oudeyer
Introduces GLAM, a method that uses online reinforcement learning to functionally ground large language models in interactive environments, improving sample efficiency and policy generalization across decision-making tasks.
Large Language Models possess vast statistical knowledge about the world, yet they frequently fail in interactive, goal-oriented settings because their internal representations lack functional grounding—the alignment between language and external physical or spatial dynamics. In real-world applications such as robotics and virtual assistants, this misalignment leads to poor decision-making and unreliable execution. While standard techniques like human feedback align models with conversational preferences, they do not ground models in environments where actions alter physical states. The article introduces and evaluates Grounded Language Models (GLAM), a framework designed to functionally ground language models directly as interactive policies using online Reinforcement Learning.
The research evaluated how using online Reinforcement Learning (specifically Proximal Policy Optimization) enables language models to learn interactive tasks, adapt to unfamiliar objects, generalize to new instructions, and outperform offline imitation methods. The authors developed BabyAI-Text, a procedurally generated, text-based navigation and reasoning platform derived from the BabyAI benchmark. In this environment, the agent receives natural language observations detailing spatial relationships, evaluates candidate actions using the internal likelihood scores of a pretrained model (FLAN-T5, primarily the 780-million parameter version), and refines its policy using sparse reward signals collected through environmental interaction.
The findings demonstrate substantial operational advantages for this approach. First, the functionally grounded model showed high learning efficiency, achieving an 80% success rate within 250,000 steps and reaching 90% around 600,000 steps across mixed tasks. In contrast, standard reinforcement learning baselines and non-pretrained models failed to reach a 20% success rate even after 1.5 million steps. Second, the model proved robust to task complexity, maintaining steady performance when irrelevant distractor objects increased from 4 to 16 and when irrelevant choices were added to the action set. Third, the grounded model exhibited strong zero-shot generalization to novel and invented object names (retaining an 87% to 88% success rate), confirming it grounded spatial relationships and geometry rather than memorizing specific items. Finally, interactive online learning consistently outperformed offline Behavioral Cloning—even when imitation data came from a flawless automated bot—by allowing the model to recover from errors through trial-and-error intervention.
These results indicate that pre-existing model knowledge provides a valuable foundation that eliminates the need to train policies from scratch, significantly reducing the data and time required for autonomous agents to become proficient. However, the study also revealed clear boundaries: models showed limited generalization when action verbs were replaced with synonyms (dropping to a 12% success rate) and failed entirely when instructions were translated into another language (dropping to 2%). These failures show that functional grounding remains tied to the specific vocabulary experienced during interactive training. Additionally, smaller models (80 million parameters) failed to display these sample-efficiency benefits, indicating that a sufficient parameter scale is required for grounded behavior to emerge.
Before deploying these systems in high-stakes or physical applications, organizations should treat this approach as an experimental foundation rather than a production-ready solution. Practical scaling faces high computational costs during action evaluation, which the authors mitigated using a distributed infrastructure library called Lamorel. Future work must validate these techniques in complex multi-modal domains (such as vision and robotics), improve computational efficiency for large action spaces, and develop strategies to ensure grounded knowledge transfers across different languages and varied action phrasing.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). SayCan provides the foundational paradigm for grounding frozen LLM reasoning in physical affordances, motivating GLAM's transition to online RL fine-tuning of the LLM policy itself.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Inner Monologue introduces closed-loop natural language feedback mechanisms for LLM planning in interactive settings, which directly underpins interactive textual decision-making environments.
- Paper: The Symbol Grounding Problem, Stevan Harnad (1990). Harnad's seminal work establishes the foundational symbol grounding problem that GLAM explicitly aims to address through online reinforcement learning in interactive environments.
- Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). This text provides the theoretical foundation for policy gradient methods with function approximation that enable fine-tuning language model parameters via reinforcement learning.
- Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). Generalized Advantage Estimation introduces the variance reduction policy gradient framework used widely in modern online RL optimization for language-based policies.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E extends grounded LLM decision-making from interactive text environments into end-to-end multimodal embodied planning across physical robotic platforms.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). Reflexion offers an alternative, verbal reinforcement learning approach that adapts LLM agents interactively via episodic verbal memory instead of online gradient-based policy updates.
- Paper: An Embodied Generalist Agent in 3D World, Jiangyong Huang et al. (2024). LEO builds upon grounded embodied language agent principles by scaling unified autoregressive LLM architectures to full 3D perception and spatial navigation tasks.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA translates language-guided decision-making into open-source vision-language-action policies that map perceptual tokens directly to robotic actions.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This survey contextualizes GLAM within the broader landscape of LLM-based autonomous agents and interactive decision-making frameworks.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This comprehensive survey provides an expansive overview of autonomous agent architectures, capability acquisition, and reinforcement learning grounding techniques.
