Inner Monologue: Embodied Reasoning through Planning with Language Models

Wenlong HuangFei XiaTed XiaoHarris ChanJacky LiangPete FlorenceAndy ZengJonathan TompsonIgor MordatchYevgen Chebotar

article2022CoRL1,566 citations

Demonstrates how integrating real-time textual environment feedback into large language models enables closed-loop embodied reasoning and replanning without additional training, substantially improving the success rate of complex, long-horizon robotic tasks.

Listen

Deploying robotic systems in dynamic real-world environments requires agents to understand high-level goals, sequence complex behaviors, and dynamically adapt when actions fail. While pre-trained large language models (LLMs) possess extensive common-sense reasoning and planning capabilities, conventional methods use them in an open-loop manner by generating a plan upfront without updating it based on execution outcomes. Consequently, when low-level actions fail or unexpected disturbances occur, these systems cannot recover autonomously.

The article evaluates whether providing pre-trained LLMs with continuous, natural language feedback from their environment enables closed-loop reasoning and autonomous replanning without requiring any model fine-tuning or retraining.

The authors develop a framework termed "Inner Monologue," which continually injects real-time textual observations—such as success detection, passive scene descriptions, and active human responses—into the LLM planning prompt as the robot acts. The approach was tested across three distinct setups: a simulated tabletop rearrangement environment subjected to test-time noise, a real-world robotic arm performing pick-and-place sorting with visual occlusions, and a real-world mobile manipulator executing complex kitchen tasks across 120 evaluations with and without adversarial human disturbances.

The experimental findings demonstrate substantial performance gains from closed-loop feedback. First, in real-world mobile manipulation tasks under adversarial disturbances, Inner Monologue achieved an overall success rate of 60.4%, whereas the open-loop SayCan baseline fell to near 0% on several tasks and achieved only 30.8% overall. Second, combining multiple feedback sources yielded the highest reliability; on the physical tabletop platform, pairing object recognition with success detection achieved a 90% overall success rate, compared to 20% for the open-loop baseline. Third, the system generalized zero-shot to unseen long-horizon tasks, whereas specialized end-to-end models like CLIPort struggled completely (0% success on unseen multi-step tasks). Finally, the framework exhibited emergent interactive behaviors, including adjusting to mid-task human instruction changes, proposing alternative goals when encountering physical constraints, understanding multilingual instructions, and answering questions about the scene.

These results indicate that natural language provides an effective, flexible interface for connecting perception, high-level planning, and low-level control. By enabling robots to detect execution failures and replan dynamically, this approach significantly reduces the risk of mission failure and lowers deployment costs by avoiding task-specific model retraining.

Organizations developing embodied automation should adopt closed-loop textual feedback architectures rather than static planning pipelines. For near-term applications, engineering teams should implement complementary feedback channels (combining both state verification and semantic object recognition). Future efforts should focus on replacing human-in-the-loop observations with fully automated visual question answering and image captioning models, as well as integrating explicit safety and ethical verification layers into the planning prompt.

Readers should note certain limitations: the system remains fundamentally constrained by the physical capabilities of its low-level control policies and the accuracy of underlying perception models. Inaccurate success detections or object misidentifications can lead to unnecessary retries or execution errors. However, across both simulated and physical evaluations, the evidence consistently supports the conclusion that closed-loop language feedback markedly enhances the robustness of embodied robotic reasoning.

arXiv: 2207.05608
Cover for Inner Monologue: Embodied Reasoning through Planning with Language Models

Abstract

Recent works have shown how the reasoning capabilities of Large Language Models (LLMs) can be applied to domains beyond natural language processing, such as planning and interaction for robots. These embodied problems require an agent to understand many semantic aspects of the world: the repertoire of skills available, how these skills influence the world, and how changes to the world map back to the language. LLMs planning in embodied environments need to consider not just what skills to do, but also how and when to do them - answers that change over time in response to the agent's own choices. In this work, we investigate to what extent LLMs used in such embodied contexts can reason over sources of feedback provided through natural language, without any additional training. We propose that by leveraging environment feedback, LLMs are able to form an inner monologue that allows them to more richly process and plan in robotic control scenarios. We investigate a variety of sources of feedback, such as success detection, scene description, and human interaction. We find that closed-loop language feedback significantly improves high-level instruction completion on three domains, including simulated and real table top rearrangement tasks and long-horizon mobile manipulation tasks in a kitchen environment in the real world.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Leveraging Embodied Language Feedback with Inner Monologue
  • 3.1 Problem Statement
  • 3.2 Inner Monologue
  • 3.3 Sources of Feedback
  • 4 Experimental Results
  • 4.1 Simulated Tabletop Rearrangement
  • 4.2 Real-World Tabletop Rearrangement
  • 4.3 Real-World Mobile Manipulator in a Kitchen Setting
  • 4.4 Emergent Capabilities
  • 5 Limitations
  • 6 Conclusion
  • References
  • A Inner Monologue Implementation Details
  • A.1 Inner Monologue for Simulated Tabletop Rearrangement
  • A.2 Inner Monologue for Real-World Tabletop Rearrangement
  • A.3 Inner Monologue for Real-World Mobile Manipulation in a Kitchen Setting
  • B Experiment Details
  • B.1 Simulated Tabletop Rearrangement Environment
  • B.2 Real Tabletop Rearrangement
  • B.3 Real Kitchen Mobile Manipulation
  • C Additional Results
  • D Prompts

Knowls

  1. Knowl 1 — Inner Monologue Closed-Loop Embodied Planning Formulation

    model/method

    Inner Monologue is a robotic task planning framework that couples a pre-trained, frozen Large Language Model (LLM) with multimodal environment feedback expressed as natural language text. An embodied agent receives a high-level natural language instruction ii and executes skills from a library of pre-trained policies πk∈Π\pi_k \in \Pi, each associated with a textual description lkl_k. Rather than generating an entire plan open-loop or relying solely on one-directional affordance grounding, the LLM functions as an interactive reasoner by maintaining an evolving prompt (the "inner monologue") that is continually updated with observations oo.

    The textual observations injected into the prompt fall into three main feedback categories:

    1. Success Detection (Success feedback): Binary task-completion feedback indicating whether the most recently executed low-level policy πk\pi_k succeeded or failed (e.g., "Success: False").
    2. Passive Scene Description (Object/Scene feedback): Structured semantic scene state information provided automatically after each action step without explicit querying by the LLM. This includes object recognition (lists of currently visible and occluded objects) and task-progress descriptions (lists of currently satisfied semantic sub-goals).
    3. Active Scene Description (Human feedback): Unstructured semantic information generated in response to active questions posed by the LLM planner during execution (e.g., asking a human or a visual question answering model for clarification on object locations or user preferences).

    By evaluating pre-trained LLMs solely via few-shot prompting, Inner Monologue enables robots to dynamically retry failed actions, replan when obstacles arise, adapt to mid-task goal changes, and propose alternative sub-goals.

  2. Knowl 2 — Dual-Model Success Detection Architecture for Mobile Manipulation

    model/method

    For real-world robotic mobile manipulation, task success detection is implemented using a combination of two neural models trained on offline teleoperated and autonomous robot rollouts:

    1. Foresight Success Detector: Given the initial visual observation image o0o_0, the final visual observation image ofo_f upon policy termination, and the text description of the attempted skill lkl_k (e.g., "Pick coke can"), the model uses CLIP image encoders to obtain embeddings for o0o_0 and ofo_f. These visual embeddings are concatenated and processed through a fusion multi-layer perceptron (MLP). The resulting image representation is concatenated with the CLIP text embedding of lkl_k and passed through a classification MLP to output a scalar probability P(success∣o0,of,lk)P(\text{success} \mid o_0, o_f, l_k). The model and CLIP backbone are trained using binary cross-entropy loss against ground-truth success labels. At inference, if P(success∣o0,of,lk)P(\text{success} \mid o_0, o_f, l_k) is below a threshold τ\tau, the feedback "[success: no]" is appended to the LLM prompt.

    2. Hindsight Success Predictor: To prevent false positives, a second model takes the image pair (o0,of)(o_0, o_f) and computes an image-fusion embedding, then takes the dot product between this embedding and the CLIP text embeddings of all candidate skill descriptions in the library. Applying a softmax with a learned temperature parameter yields a categorical probability distribution over all possible skills. The model is trained via symmetric contrastive loss.

    Inference Combination: The system first evaluates whether the foresight success probability exceeds τ\tau. If it does, the hindsight predictor is queried; execution is classified as successful if and only if arg⁡max⁡lPhindsight(l∣o0,of)\arg\max_{l} P_{\text{hindsight}}(l \mid o_0, o_f) matches the attempted skill lkl_k.

  3. Knowl 3 — Simulated Tabletop Rearrangement Evaluation under Disturbances

    data/table

    In a Ravens-based simulated tabletop rearrangement environment, a robotic arm rearranges up to 4 colored blocks and 3 colored bowls across 9 possible workspace locations based on natural language instructions. The planner uses InstructGPT-1.3B paired with a pre-trained CLIP-based Transporter Net pick-and-place primitive (trained on 20,000 demonstrations). To test robustness to dynamic failures, Gaussian noise is injected at test time into pixel observations (N(0,3)\mathcal{N}(0, 3)), pick-place heatmap predictions (N(0,2.5)\mathcal{N}(0, 2.5)), and placement positions (N(0,0.02 m)\mathcal{N}(0, 0.02\text{ m})). Methods are evaluated across 50 episodes per task on four seen tasks (used in few-shot prompts or baseline training) and four unseen tasks.

    Tasks CLIPort +oracle +LLM Object +IM Object+Success +IM Object+Scene
    Seen Tasks
    “Pick and place” 24.0% 74.0% 80.0% 90.0% 94.0%
    “Stack all the blocks” 2.0% 32.0% 4.0% 10.0% 26.0%
    “Put all the blocks on the [x] corner/side” 2.0% 32.0% 30.0% 28.0% 30.0%
    “Put all the blocks in the [x] bowl” 32.0% 94.0% 52.0% 46.0% 56.0%
    Unseen Tasks
    “Put all the blocks in different corners” 0.0% 0.0% 20.0% 20.0% 26.0%
    “Put the blocks in their matching bowls” 0.0% 0.0% 56.0% 70.0% 82.0%
    “Put the blocks on mismatched bowls” 0.0% 0.0% 62.0% 76.0% 86.0%
    “Stack all the blocks on the [x] corner/side” 0.0% 0.0% 0.0% 4.0% 6.0%

    A standalone multi-task CLIPort policy trained on the 4 seen tasks completely fails (0.0% success) on unseen long-horizon tasks. Providing LLM planning with object presence (+LLM Object) enables zero-shot generalization to unseen tasks. Adding closed-loop Inner Monologue feedback (+IM Object+Success and +IM Object+Scene) significantly outperforms open-loop LLM planning; +IM Object+Scene achieves the highest performance by tracking achieved sub-goals in the prompt, allowing the agent to recover when earlier achievements (such as unstable block towers) are knocked over by disturbance.

  4. Knowl 4 — Real-World Tabletop Pick-and-Place Experimental Setup

    experimental setup

    The real-world tabletop rearrangement platform comprises a 6-DoF Universal Robots UR5e arm equipped with a wrist-mounted Intel RealSense RGB-D camera and a pneumatic suction gripper. The workspace contains toy blocks, plastic food items, condiments, and plates.

    1. Planning: Multi-step planning is performed by InstructGPT-1.3B with few-shot prompts.
    2. Perception (Object Feedback): Object detection is performed zero-shot using pre-trained MDETR. The prompt receives a list of currently visible detected objects and previously seen objects that are currently occluded.
    3. Success Detection (Success Feedback): Bounding box centers detected by MDETR after an action are deprojected to 3D camera coordinates and transformed to the robot base frame. If the 2D Euclidean distance in the base frame between the detected object center and the intended target location is below a threshold (3 cm for block stacking, 10 cm for plate sorting), the action is classified as successful ("Successful action: True"), otherwise "Successful action: False".
    4. Low-Level Control: Target objects are parsed from LLM action strings into bounding box coordinates; the suction gripper moves 15 cm above the pick/place target and descends until a 5 N contact force is registered.
    5. Perturbation: Zero-mean Gaussian noise (standard deviation σ=1.5 cm\sigma = 1.5\text{ cm} for stacking, σ=0.7 cm\sigma = 0.7\text{ cm} for sorting, capped at 1.5σ1.5\sigma) is added to the planar pick coordinates to stress-test failure recovery.
  5. Knowl 5 — Performance of Inner Monologue in Real-World Tabletop Rearrangement

    data/table

    Real-world pick-and-place experiments on a UR5e manipulator evaluate Inner Monologue under injected execution noise across two task families (10 trials each): (1) finishing a 3-block stacking tower where the bottom block is initially occluded by a pre-stacked top block, and (2) sorting fruit and condiment bottles into separate plates based on LLM semantic categorization.

    Task Family Open-Loop LLM Object IM Object IM Success IM Object + Success
    Finish 3-block stacking 20% 40% 40% 100%
    Sort fruits from bottles 20% 50% 40% 80%
    Total 20% 45% 40% 90%

    The baseline open-loop LLM policy (which runs object detection only once at the beginning) achieves only 20% overall success because it fails to discover occluded objects and cannot recover from grasp slips. While adding closed-loop object feedback alone (45%) or success feedback alone (40%) provides moderate gains, combining both in Inner Monologue yields 90% total success. Closed-loop scene descriptions dynamically expose unoccluded objects to the planner, while success detections trigger prompt-driven retries.

  6. Knowl 6 — Real-World Mobile Manipulation Kitchen Setup

    experimental setup

    The real-world mobile manipulation domain utilizes an Everyday Robots mobile manipulator (7-DoF arm, parallel-jaw gripper, omnidirectional mobile base, and RGB cameras) operating in an office kitchen environment with 5 designated location zones (Table, Close Counter, Far Counter, Trash Can, Drawers) and 15 distinct household objects.

    • High-Level Planner: PaLM-540B language model evaluated with few-shot prompting.
    • Affordance Grounding: Following the SayCan framework, candidate action scores output by the LLM are multiplied by value functions V(s,a)V(s, a) from pre-trained Q-networks of reinforcement learning policies to ensure physical feasibility from the robot's current state.
    • Low-Level Control Policies: Behavior cloning policies for manipulation (counter picking, drawer opening/closing, drawer picking, placing) trained on 68,000 teleoperated demonstrations and 12,000 autonomous successes, paired with scripted navigation primitives.
    • Disturbance Protocol: In the adversarial disturbance condition, human operators actively perturb rollouts (e.g., knocking objects out of the gripper or displacing items) to force skill policy failure and test replanning capabilities.
    • Evaluation Set: 120 total real-world evaluations divided across three task families: Manipulation (4 instructions), Mobile Manipulation (2 long-horizon instructions), and Dexterous Drawer Manipulation (2 instructions).
  7. Knowl 7 — Performance of Inner Monologue in Real-World Mobile Kitchen Manipulation

    data/table

    Mobile manipulation in an office kitchen is evaluated over 120 trials across three task families (Manipulation, Mobile Manipulation, and Drawers), comparing SayCan against Inner Monologue variants with and without adversarial human physical disturbances during execution.

    Task Family SayCan IM Success IM Object + Success
    No Disturbances
    Manipulation 50.0% 62.5% 75.0%
    Mobile Manipulation 50.0% 50.0% 75.0%
    Drawers 83.3% 83.3% 100.0%
    With Disturbances
    Manipulation 12.5% 25.0% 33.3%
    Mobile Manipulation 0.0% 25.0% 75.0%
    Drawers 0.0% 44.4% 44.4%
    Total 30.8% 48.7% 60.4%

    In undisturbed conditions, Inner Monologue with Object and Success feedback improves total success over SayCan from 30.8% to 60.4% across all conditions by recovering from natural execution slips. Under external adversarial disturbances, SayCan fails completely (0.0% on Mobile Manipulation and Drawers, 12.5% on Manipulation) because value-function affordance grounding alone lacks a mechanism for prompting high-level retries or skill replanning. In contrast, Inner Monologue maintains 60.4% overall success under disturbances by reacting to [success: no] signals and updated scene contents in the LLM context prompt.

  8. Knowl 8 — Emergent Reasoning Behaviors in Grounded Inner Monologue

    model/method

    When pre-trained LLMs receive structured and unstructured textual environment feedback in an inner monologue prompt, several complex reasoning behaviors emerge without domain-specific training or explicit few-shot demonstrations of these behaviors:

    1. Dynamic Adaptation to Changing Instructions: When a human user injects mid-execution corrections or changes the goal (e.g., switching targets twice or issuing "nevermind, finish your previous task"), the LLM updates its goal representation and switches action sequences accordingly. It also interprets human requests like "please stop" as a prompt to predict the termination action done.
    2. Autonomous Goal Reformation under Infeasibility: When an attempted action fails due to physical constraints and the environment provides explanatory feedback (e.g., feedback stating "The purple block is too heavy to be picked up"), the LLM autonomously proposes a replacement sub-goal (such as "I need to find a lighter block") and selects a different valid object.
    3. Multilingual Instruction Grounding: The LLM seamlessly processes instructions and human interruptions provided in non-English languages (such as Chinese), translates them into English goal assertions, and continues plan execution.
    4. Interactive Post-Task Scene Reasoning: After plan completion, the dialogue history in the inner monologue allows the LLM to accurately answer factual visual and spatial questions about the final scene state (e.g., identifying which items remain inside a container versus on the table).
    5. Robustness to Feedback Permutation and Typos: The planner correctly processes unexpected prompt ordering (such as human interruptions appearing mid-routine) and typographical errors in human input.
  9. Knowl 9 — Limitations and Failure Modes of Inner Monologue

    limitation

    The Inner Monologue framework is subject to several practical limitations and failure modes:

    1. Perception and Detector Inaccuracies: The system depends heavily on accurate feedback models. False negatives from success detectors cause the robot to perform redundant retries of already-completed skills, while false positives leave failures uncorrected, causing subsequent steps to execute under partial observability. In open-vocabulary object detection evaluations in the kitchen domain, ViLD achieved 85.7% precision and 72.0% recall, while MDETR achieved 39.6% precision and 87.5% recall; detector omissions or bounding box inaccuracies in cluttered scenes frequently degraded planning quality.
    2. Low-Level Policy Bottlenecks: The scope of feasible tasks and overall success rates remain strictly bounded by the execution capabilities and robustness of the underlying low-level control policies, regardless of LLM reasoning capability.
    3. LLM Hallucination and Feedback Disregard: In certain failure cases, the language model planner ignores textual feedback indicating an object is missing and hallucinates plans involving objects not present in the workspace.
    4. Oracle Assumptions: In multiple evaluation settings, scene descriptions or object verifications rely on simulated ground truth or human observers in the loop rather than fully autonomous multimodal perception systems.

Coverage note — All substantial contributions—including the Inner Monologue formulation, the dual foresight/hindsight success detector, the experimental setups and evaluation results across simulated tabletop, real tabletop, and real kitchen domains, emergent capabilities, and limitations—are fully covered in the knowls.

References

  1. 1.L. P. Kaelbling and T. Lozano-Perez. Integrated task and motion planning in belief space. ´ The International Journal of Robotics Research, 32(9-10):1194–1227, 2013.
  2. 2.A. G. Barto and S. Mahadevan. Recent advances in hierarchical reinforcement learning. Discrete event dynamic systems, 13(1):41–77, 2003.
  3. 3.F. Petroni, T. Rocktaschel, P. Lewis, A. Bakhtin, Y. Wu, A. H. Miller, and S. Riedel. Language models ¨ as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  4. 4.Z. Jiang, F. F. Xu, J. Araki, and G. Neubig. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438, 2020.
  5. 5.J. Davison, J. Feldman, and A. M. Rush. Commonsense knowledge mining from pretrained models. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th in￾ternational joint conference on natural language processing (EMNLP-IJCNLP), pages 1173–1178, 2019.
  6. 6.A. Talmor, Y. Elazar, Y. Goldberg, and J. Berant. olmpics-on what language model pre-training captures. Transactions of the Association for Computational Linguistics, 8:743–758, 2020.
  7. 7.A. Roberts, C. Raffel, and N. Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020.
  8. 8.A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  9. 9.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  10. 10.J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  11. 11.T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  12. 12.A. K. Lampinen, I. Dasgupta, S. C. Chan, K. Matthewson, M. H. Tessler, A. Creswell, J. L. McClelland, J. X. Wang, and F. Hill. Can language models learn from explanations in context? arXiv preprint arXiv:2204.02329, 2022.
  13. 13.M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  14. 14.L. S. Vygotsky. Thought and language. MIT press, 2012.
  15. 15.P. Carruthers. Thinking in language?: evolution and a modularist possibility. Cambridge University Press, 1998.
  16. 16.L. Vygotsky. Tool and symbol in child development. The vygotsky reader, 1994.
  17. 17.L. S. Vygotsky. Play and its role in the mental development of the child. Soviet psychology, 5(3): 6–18, 1967.
  18. 18.C. Colas, T. Karch, C. Moulin-Frier, and P.-Y. Oudeyer. Vygotskian autotelic artificial intelligence: Language and culture internalization for human-like ai. arXiv preprint arXiv:2206.01134, 2022.
  19. 19.A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598, 2022.
  20. 20.W. Huang, P. Abbeel, D. Pathak, and I. Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning. PMLR, 2022.
  21. 21.M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, K. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, and M. Yan. Do as i can and not as i say: Grounding language in robotic affordances. In arXiv preprint arXiv:2204.01691, 2022.
  22. 22.L. P. Kaelbling and T. Lozano-Perez. Hierarchical planning in the now. In ´ Workshops at the Twenty-Fourth AAAI Conference on Artificial Intelligence, 2010.
  23. 23.S. Srivastava, E. Fang, L. Riano, R. Chitnis, S. Russell, and P. Abbeel. Combined task and motion planning through an extensible planner-independent interface layer. In 2014 IEEE international conference on robotics and automation (ICRA), 2014.
  24. 24.R. E. Fikes and N. J. Nilsson. Strips: A new approach to the application of theorem proving to problem solving. Artificial intelligence, 1971.
  25. 25.E. D. Sacerdoti. A structure for plans and behavior. Technical report, SRI International, Menlo Park California Artificial Intelligence Center, 1975.
  26. 26.D. Nau, Y. Cao, A. Lotem, and H. Munoz-Avila. Shop: Simple hierarchical ordered planner. In Proceedings of the 16th international joint conference on Artificial intelligence, 1999.
  27. 27.S. M. LaValle. Planning algorithms. Cambridge university press, 2006.
  28. 28.M. Toussaint. Logic-geometric programming: An optimization-based approach to combined task and motion planning. In Twenty-Fourth International Joint Conference on Artificial Intelligence, 2015.
  29. 29.M. A. Toussaint, K. R. Allen, K. A. Smith, and J. B. Tenenbaum. Differentiable physics and stable modes for tool-use and manipulation planning. Robotics: Science and Systems Foundation, 2018.
  30. 30.B. Eysenbach, R. R. Salakhutdinov, and S. Levine. Search on the replay buffer: Bridging planning and reinforcement learning. Advances in Neural Information Processing Systems, 2019.
  31. 31.D. Xu, S. Nair, Y. Zhu, J. Gao, A. Garg, L. Fei-Fei, and S. Savarese. Neural task programming: Learning to generalize across hierarchical tasks. In 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018.
  32. 32.D. Xu, R. Mart´ın-Mart´ın, D.-A. Huang, Y. Zhu, S. Savarese, and L. F. Fei-Fei. Regression planning networks. Advances in Neural Information Processing Systems, 32, 2019.
  33. 33.T. Silver, R. Chitnis, N. Kumar, W. McClinton, T. Lozano-Perez, L. P. Kaelbling, and J. Tenenbaum. Inventing relational state and action abstractions for effective and efficient bilevel planning. arXiv preprint arXiv:2203.09634, 2022.
  34. 34.D. Shah, P. Xu, Y. Lu, T. Xiao, A. Toshev, S. Levine, and B. Ichter. Value function spaces: Skill-centric state abstractions for long-horizon reasoning. ICLR, 2022. URL https://openreview.net/pdf?id=vgqS1vkkCbE.
  35. 35.A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn. Universal planning networks: Learning generalizable representations for visuomotor control. In International Conference on Machine Learning, pages 4732–4741. PMLR, 2018.
  36. 36.T. Kurutach, A. Tamar, G. Yang, S. J. Russell, and P. Abbeel. Learning plannable representations with causal infogan. Advances in Neural Information Processing Systems, 31, 2018.
  37. 37.A. Akakzia, C. Colas, P.-Y. Oudeyer, M. Chetouani, and O. Sigaud. Grounding language to autonomously-acquired skills via goal generation. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=chPj I5KMHG.
  38. 38.S. Pirk, K. Hausman, A. Toshev, and M. Khansari. Modeling long-horizon tasks as sequential interaction landscapes. arXiv preprint arXiv:2006.04843, 2020.
  39. 39.T. Kollar, S. Tellex, D. Roy, and N. Roy. Toward understanding natural language directions. In 2010 5th ACM/IEEE International Conference on Human-Robot Interaction (HRI), pages 259–266. IEEE, 2010.
  40. 40.S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy. Understanding natural language commands for robotic navigation and mobile manipulation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 25, pages 1507–1514, 2011.
  41. 41.M. Bollini, S. Tellex, T. Thompson, N. Roy, and D. Rus. Interpreting and executing recipes with a cooking robot. In Experimental Robotics, pages 481–495. Springer, 2013.
  42. 42.S. Tellex, R. Knepper, A. Li, D. Rus, and N. Roy. Asking for help using inverse semantics. 2014.
  43. 43.T. Kollar, S. Tellex, D. Roy, and N. Roy. Grounding verbs of motion in natural language commands to robots. In Experimental robotics, pages 31–47. Springer, 2014.
  44. 44.V. Blukis, Y. Terme, E. Niklasson, R. A. Knepper, and Y. Artzi. Learning to map natural language in￾structions to physical quadcopter control using simulated flight. arXiv preprint arXiv:1910.09664, 2019.
  45. 45.S. Nair and C. Finn. Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation. ArXiv, abs/1909.05829, 2020.
  46. 46.F. Xia, C. Li, R. Mart´ın-Mart´ın, O. Litany, A. Toshev, and S. Savarese. Relmogen: Integrating motion generation in reinforcement learning for mobile manipulation. In 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021.
  47. 47.C. Li, F. Xia, R. Martin-Martin, and S. Savarese. Hrl4in: Hierarchical reinforcement learning for interactive navigation with mobile manipulators. In Conference on Robot Learning, 2020.
  48. 48.Y. Jiang, S. Gu, K. Murphy, and C. Finn. Language as an abstraction for hierarchical deep reinforcement learning. In NeurIPS, 2019.
  49. 49.D. Hafner, K.-H. Lee, I. Fischer, and P. Abbeel. Deep hierarchical planning from pixels. arXiv preprint arXiv:2206.04114, 2022.
  50. 50.S. Mirchandani, S. Karamcheti, and D. Sadigh. Ella: Exploration through learned language abstraction. Advances in Neural Information Processing Systems, 34:29529–29540, 2021.
  51. 51.P. A. Jansen. Visually-grounded planning without vision: Language models infer detailed plans from high-level instructions. arXiv preprint arXiv:2009.14259, 2020.
  52. 52.P. Sharma, A. Torralba, and J. Andreas. Skill induction and planning with latent language. arXiv preprint arXiv:2110.01517, 2021.
  53. 53.S. Li, X. Puig, Y. Du, C. Wang, E. Akyurek, A. Torralba, J. Andreas, and I. Mordatch. Pre-trained language models for interactive decision-making. arXiv preprint arXiv:2202.01771, 2022.
  54. 54.M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  55. 55.Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  56. 56.N. Reimers and I. Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
  57. 57.J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  58. 58.C. Paxton, Y. Bisk, J. Thomason, A. Byravan, and D. Foxl. Prospection: Interpretable plans from language by predicting the future. In 2019 International Conference on Robotics and Automation (ICRA), pages 6942–6948. IEEE, 2019.
  59. 59.S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor. Language-conditioned imitation learning for robot manipulation tasks. Advances in Neural Information Processing Systems, 33:13139–13150, 2020.
  60. 60.V. Blukis, R. A. Knepper, and Y. Artzi. Few-shot object grounding and mapping for natural language robot instruction following. arXiv preprint arXiv:2011.07384, 2020.
  61. 61.C. Lynch and P. Sermanet. Language conditioned imitation learning over unstructured data. Robotics: Science and Systems, 2021. URL https://arxiv.org/abs/2005.07648.
  62. 62.Y. Chen, R. Xu, Y. Lin, and P. A. Vela. A joint network for grasp detection conditioned on natural language commands. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 4576–4582. IEEE, 2021.
  63. 63.O. Mees, L. Hermann, and W. Burgard. What matters in language conditioned robotic imitation learning. arXiv preprint arXiv:2204.06252, 2022.
  64. 64.C. Yan, F. Carnevale, P. Georgiev, A. Santoro, A. Guy, A. Muldal, C.-C. Hung, J. Abramson, T. Lillicrap, and G. Wayne. Intra-agent speech permits zero-shot task acquisition. arXiv preprint arXiv:2206.03139, 2022.
  65. 65.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  66. 66.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  67. 67.J. Lu, D. Batra, D. Parikh, and S. Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019.
  68. 68.Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904, 2021.
  69. 69.A. Suglia, Q. Gao, J. Thomason, G. Thattai, and G. Sukhatme. Embodied bert: A transformer model for embodied, language-guided visual task completion. arXiv preprint arXiv:2108.04927, 2021.
  70. 70.T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. E. Hinton. Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems, 33:22243–22255, 2020.
  71. 71.A. Jain, M. Guo, K. Srinivasan, T. Chen, S. Kudugunta, C. Jia, Y. Yang, and J. Baldridge. Mural: multimodal, multitask retrieval across languages. arXiv preprint arXiv:2109.05125, 2021.
  72. 72.J. Sun, D.-A. Huang, B. Lu, Y.-H. Liu, B. Zhou, and A. Garg. Plate: Visually-grounded planning with transformers in procedural tasks. IEEE Robotics and Automation Letters, 7(2):4924–4930, 2022.
  73. 73.F. Sener and A. Yao. Zero-shot anticipation for instructional activities. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 862–871, 2019.
  74. 74.A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi. Simple but effective: Clip embeddings for embodied ai. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14829–14838, 2022.
  75. 75.A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, and J. Lee. Transporter networks: Rearranging the visual world for robotic manipulation. Conference on Robot Learning (CoRL), 2020.
  76. 76.M. Shridhar, L. Manuelli, and D. Fox. Cliport: What and where pathways for robotic manipulation. In Conference on Robot Learning, pages 894–906. PMLR, 2022.
  77. 77.X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui. Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921, 2021.
  78. 78.I. Lenz, H. Lee, and A. Saxena. Deep learning for detecting robotic grasps. The International Journal of Robotics Research, 34(4-5):705–724, 2015.
  79. 79.F.-J. Chu, R. Xu, and P. A. Vela. Real-world multiobject, multigrasp detection. IEEE Robotics and Automation Letters, 3(4):3355–3362, 2018.
  80. 80.D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman. Mt-opt: Continuous multi-task robotic reinforcement learning at scale. arXiv preprint arXiv:2104.08212, 2021.
  81. 81.T. Migimatsu and J. Bohg. Grounding predicates through actions. arXiv preprint arXiv:2109.14718, 2021.
  82. 82.Y. Cui, S. Niekum, A. Gupta, V. Kumar, and A. Rajeswaran. Can foundation models perform zero-shot task specification for robot manipulation? In Learning for Dynamics and Control Conference, pages 893–905. PMLR, 2022.
  83. 83.M. Liang and X. Hu. Recurrent convolutional neural network for object recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3367–3375, 2015.
  84. 84.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015.
  85. 85.Z. Zou, Z. Shi, Y. Guo, and J. Ye. Object detection in 20 years: A survey. arXiv preprint arXiv:1905.05055, 2019.
  86. 86.A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020.
  87. 87.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433, 2015.
  88. 88.L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao. Unified vision-language pre-training for image captioning and vqa. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 13041–13049, 2020.
  89. 89.H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han. Once for all: Train one network and specialize it for efficient deployment. In International Conference on Learning Representations, 2020. URL https://arxiv.org/pdf/1908.09791.pdf.
  90. 90.J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, J. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022.
  91. 91.L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  92. 92.A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1780–1790, 2021.
  93. 93.T. Xiao, E. Jang, D. Kalashnikov, S. Levine, J. Ibarz, K. Hausman, and A. Herzog. Thinking while moving: Deep reinforcement learning with concurrent control. arXiv preprint arXiv:2004.06089, 2020.

Citation

MLA
Huang, W., et al. “Inner Monologue: Embodied Reasoning Through Planning with Language Models”. arXiv, 2022, http://arxiv.org/abs/2207.05608v1.
APA
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Brown, N., Jackson, T., Luu, L., Levine, S., Hausman, K., & Ichter, B. (2022). Inner Monologue: Embodied Reasoning through Planning with Language Models. arXiv. http://arxiv.org/abs/2207.05608v1
Chicago
Huang, W., F. Xia, T. Xiao, et al. 2022. “Inner Monologue: Embodied Reasoning Through Planning with Language Models”. arXiv. http://arxiv.org/abs/2207.05608v1.
Harvard
Huang, W. et al. (2022) “Inner Monologue: Embodied Reasoning through Planning with Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.05608v1.
Vancouver
1. Huang W, Xia F, Xiao T, et al (2022) Inner Monologue: Embodied Reasoning through Planning with Language Models. arXiv

BibTeX

@article{huang2022inner,
  title = {Inner Monologue: Embodied Reasoning through Planning with Language Models},
  author = {Huang, Wenlong and Xia, Fei and Xiao, Ted and Chan, Harris and Liang, Jacky and Florence, Pete and Zeng, Andy and Tompson, Jonathan and Mordatch, Igor and Chebotar, Yevgen and Sermanet, Pierre and Brown, Noah and Jackson, Tomas and Luu, Linda and Levine, Sergey and Hausman, Karol and Ichter, Brian},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.05608v1},
  eprint = {2207.05608}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/