Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
Wenlong HuangPieter AbbeelDeepak PathakIgor Mordatch
Demonstrates that large language models can act as zero-shot planners for embodied agents by decomposing high-level natural language instructions and semantically mapping them into executable, environment-admissible actions.
Deploying artificial intelligence agents to carry out daily, open-ended human activities in interactive environments represents a major frontier in robotics. While large language models internalize vast amounts of common-sense knowledge from text, traditional systems rely heavily on expensive, environment-specific demonstrations and supervised training to map natural language into robotic commands.
The article evaluates whether pre-trained large language models already possess the actionable knowledge required to break down high-level tasks into step-by-step plans without any additional training or model fine-tuning. It demonstrates a zero-shot planning framework that translates free-form language generation into valid, executable actions within an interactive simulator.
To test this capability, the researchers conducted experiments in the VirtualHome household simulation across 88 held-out everyday tasks evaluated over seven unique scenes. The proposed pipeline employs an autoregressive planning language model to generate steps, dynamically selects prompt examples based on semantic similarity, and uses a sentence embedding model to map generated text onto admissible environment actions while applying on-the-fly step corrections. Action quality was measured through mechanical executability across simulator constraints and human evaluations assessing semantic correctness.
The findings show that large models such as GPT-3 and Codex can spontaneously break down high-level tasks into remarkably realistic plans, achieving raw human-rated correctness of 77.86%, which surpasses human baseline plans scored at 70.05%. However, naively generated plans are rarely executable in the simulator, achieving an executability rate of only 7.79% for GPT-3. The proposed semantic translation and trajectory correction pipeline dramatically increases executability to 73.05% for GPT-3 and 78.57% for Codex, raising the proportion of plans that are simultaneously executable and correct from under 7% to over 35%.
These results indicate that pre-trained language models can serve as effective off-the-shelf high-level planners, significantly lowering development costs, engineering timelines, and data collection burdens for robotic systems. Because the framework operates entirely at inference time without modifying underlying model weights, it can be integrated directly into existing artificial intelligence serving pipelines without specialized retraining.
Organizations developing embodied agents should adopt modular architectures that decouple high-level language planning from mid-level grounding and low-level execution controllers. Before deploying such frameworks in physical operations, stakeholders must conduct pilot testing to resolve trade-offs between strict executability and task correctness, particularly when handling complex multi-step instructions.
Confidence in these findings is supported by consistent performance across diverse household tasks and human evaluator agreement. However, key limitations remain: the current system lacks real-time sensory perception of the environment, cannot distinguish between multiple objects of the same class, and relies on an assumed low-level execution controller to perform physical interactions.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Introduces the foundational paradigm of evaluating whether pretrained language models store actionable world and commonsense knowledge without fine-tuning, directly motivating using LLMs as zero-shot planners.
- Paper: PIQA: Reasoning about Physical Commonsense in Natural Language, Yonatan Bisk et al. (2019). Establishes benchmarks and evaluation methodology for physical commonsense reasoning in language models, which is essential for understanding how LLMs represent everyday physical tasks and steps.
- Paper: Language Models are Unsupervised Multitask Learners, Alec Radford et al. (2019). Demonstrates the zero-shot task decomposition and multitask execution capabilities of large language models that the source paper directly harnesses for embodied planning.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). Provides the foundational framework for grounding high-level natural language instructions into physical navigation and sequential embodied actions.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). Extends the concept of extracting zero-shot LLM plans to physical robotics by combining language task decomposition with learned robotic affordance value functions (SayCan).
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Builds on LLM-based embodied action decomposition by having language models generate executable code and hierarchical sub-functions for robotic control.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). Generalizes language-based embodied planning by directly ingesting raw multimodal sensor inputs into the language model architecture for end-to-end task planning and robot execution.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). Integrates iterative reasoning traces with environment feedback during action execution, improving upon static zero-shot plan generation.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Transitions from high-level symbolic action mapping to unified end-to-end vision-language-action policies that translate web-scale semantic knowledge into direct robotic control.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). Synthesizes the broader landscape of LLM-based autonomous agent architectures, placing zero-shot planning and grounding into comprehensive cognitive and multi-agent frameworks.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Presents a comprehensive review of architectural strategies for task planning, memory, and grounding across LLM-powered autonomous agents.
