Inner Monologue: Embodied Reasoning through Planning with Language Models
Wenlong HuangFei XiaTed XiaoHarris ChanJacky LiangPete FlorenceAndy ZengJonathan TompsonIgor MordatchYevgen Chebotar
Demonstrates how integrating real-time textual environment feedback into large language models enables closed-loop embodied reasoning and replanning without additional training, substantially improving the success rate of complex, long-horizon robotic tasks.
Deploying robotic systems in dynamic real-world environments requires agents to understand high-level goals, sequence complex behaviors, and dynamically adapt when actions fail. While pre-trained large language models (LLMs) possess extensive common-sense reasoning and planning capabilities, conventional methods use them in an open-loop manner by generating a plan upfront without updating it based on execution outcomes. Consequently, when low-level actions fail or unexpected disturbances occur, these systems cannot recover autonomously.
The article evaluates whether providing pre-trained LLMs with continuous, natural language feedback from their environment enables closed-loop reasoning and autonomous replanning without requiring any model fine-tuning or retraining.
The authors develop a framework termed "Inner Monologue," which continually injects real-time textual observations—such as success detection, passive scene descriptions, and active human responses—into the LLM planning prompt as the robot acts. The approach was tested across three distinct setups: a simulated tabletop rearrangement environment subjected to test-time noise, a real-world robotic arm performing pick-and-place sorting with visual occlusions, and a real-world mobile manipulator executing complex kitchen tasks across 120 evaluations with and without adversarial human disturbances.
The experimental findings demonstrate substantial performance gains from closed-loop feedback. First, in real-world mobile manipulation tasks under adversarial disturbances, Inner Monologue achieved an overall success rate of 60.4%, whereas the open-loop SayCan baseline fell to near 0% on several tasks and achieved only 30.8% overall. Second, combining multiple feedback sources yielded the highest reliability; on the physical tabletop platform, pairing object recognition with success detection achieved a 90% overall success rate, compared to 20% for the open-loop baseline. Third, the system generalized zero-shot to unseen long-horizon tasks, whereas specialized end-to-end models like CLIPort struggled completely (0% success on unseen multi-step tasks). Finally, the framework exhibited emergent interactive behaviors, including adjusting to mid-task human instruction changes, proposing alternative goals when encountering physical constraints, understanding multilingual instructions, and answering questions about the scene.
These results indicate that natural language provides an effective, flexible interface for connecting perception, high-level planning, and low-level control. By enabling robots to detect execution failures and replan dynamically, this approach significantly reduces the risk of mission failure and lowers deployment costs by avoiding task-specific model retraining.
Organizations developing embodied automation should adopt closed-loop textual feedback architectures rather than static planning pipelines. For near-term applications, engineering teams should implement complementary feedback channels (combining both state verification and semantic object recognition). Future efforts should focus on replacing human-in-the-loop observations with fully automated visual question answering and image captioning models, as well as integrating explicit safety and ethical verification layers into the planning prompt.
Readers should note certain limitations: the system remains fundamentally constrained by the physical capabilities of its low-level control policies and the accuracy of underlying perception models. Inaccurate success detections or object misidentifications can lead to unnecessary retries or execution errors. However, across both simulated and physical evaluations, the evidence consistently supports the conclusion that closed-loop language feedback markedly enhances the robustness of embodied robotic reasoning.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). SayCan provides the foundational framework for grounding high-level LLM task planning in robotic affordances that Inner Monologue directly builds upon and extends with closed-loop natural language feedback.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). This paper establishes the core paradigm of zero-shot step-by-step action planning with large language models in embodied environments, which Inner Monologue enhances via dynamic environmental feedback loops.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It introduces chain-of-thought prompting, the foundational reasoning technique that enables language models to generate sequential intermediate steps for planning.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). It presents a complementary foundational approach for translating language model outputs directly into executable robotic control policies in physical and simulated settings.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). It details how language models can systematically break complex challenges into simpler sequential subproblems, informing the task decomposition strategies used in robotic planning.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). ReAct formalizes and generalizes the interleaving of language-based reasoning traces and environment actions across interactive reasoning and decision-making benchmarks.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). Reflexion advances language-driven closed-loop adaptation by incorporating episodic verbal self-reflection and reinforcement across sequential decision-making trials.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E extends embodied reasoning from text-mediated interaction loops into an end-to-end multimodal model that directly incorporates continuous visual and sensor inputs into the language architecture.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 takes the integration of web-scale semantic reasoning and physical control a step further by directly generating low-level robot actions as tokens within a vision-language-action foundation model.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). Self-Refine abstracts the concept of multi-turn internal feedback and iterative plan correction into a general-purpose framework for non-embodied LLM generation tasks.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey provides a comprehensive architectural taxonomy of autonomous LLM agent frameworks, contextualizing feedback-driven planning systems like Inner Monologue.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This survey synthesizes the broader evolution of LLM-based agents across perception, memory, and embodied planning in physical and virtual environments.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA builds upon embodied reasoning principles by embedding structured, multi-step affordance reasoning directly into vision-language-action models for robotic manipulation.
