Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Wenlong HuangPieter AbbeelDeepak PathakIgor Mordatch

article2022ICML1,510 citations

Demonstrates that large language models can act as zero-shot planners for embodied agents by decomposing high-level natural language instructions and semantically mapping them into executable, environment-admissible actions.

Listen

Deploying artificial intelligence agents to carry out daily, open-ended human activities in interactive environments represents a major frontier in robotics. While large language models internalize vast amounts of common-sense knowledge from text, traditional systems rely heavily on expensive, environment-specific demonstrations and supervised training to map natural language into robotic commands.

The article evaluates whether pre-trained large language models already possess the actionable knowledge required to break down high-level tasks into step-by-step plans without any additional training or model fine-tuning. It demonstrates a zero-shot planning framework that translates free-form language generation into valid, executable actions within an interactive simulator.

To test this capability, the researchers conducted experiments in the VirtualHome household simulation across 88 held-out everyday tasks evaluated over seven unique scenes. The proposed pipeline employs an autoregressive planning language model to generate steps, dynamically selects prompt examples based on semantic similarity, and uses a sentence embedding model to map generated text onto admissible environment actions while applying on-the-fly step corrections. Action quality was measured through mechanical executability across simulator constraints and human evaluations assessing semantic correctness.

The findings show that large models such as GPT-3 and Codex can spontaneously break down high-level tasks into remarkably realistic plans, achieving raw human-rated correctness of 77.86%, which surpasses human baseline plans scored at 70.05%. However, naively generated plans are rarely executable in the simulator, achieving an executability rate of only 7.79% for GPT-3. The proposed semantic translation and trajectory correction pipeline dramatically increases executability to 73.05% for GPT-3 and 78.57% for Codex, raising the proportion of plans that are simultaneously executable and correct from under 7% to over 35%.

These results indicate that pre-trained language models can serve as effective off-the-shelf high-level planners, significantly lowering development costs, engineering timelines, and data collection burdens for robotic systems. Because the framework operates entirely at inference time without modifying underlying model weights, it can be integrated directly into existing artificial intelligence serving pipelines without specialized retraining.

Organizations developing embodied agents should adopt modular architectures that decouple high-level language planning from mid-level grounding and low-level execution controllers. Before deploying such frameworks in physical operations, stakeholders must conduct pilot testing to resolve trade-offs between strict executability and task correctness, particularly when handling complex multi-step instructions.

Confidence in these findings is supported by consistent performance across diverse household tasks and human evaluator agreement. However, key limitations remain: the current system lacks real-time sensory perception of the environment, cannot distinguish between multiple objects of the same class, and relies on an assumed low-level execution controller to perform physical interactions.

Cover for Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Abstract

Can world knowledge learned by large language models (LLMs) be used to act in interactive environments? In this paper, we investigate the possibility of grounding high-level tasks, expressed in natural language (e.g. "make breakfast"), to a chosen set of actionable steps (e.g. "open fridge"). While prior work focused on learning from explicit step-by-step examples of how to act, we surprisingly find that if pre-trained LMs are large enough and prompted appropriately, they can effectively decompose high-level tasks into mid-level plans without any further training. However, the plans produced naively by LLMs often cannot map precisely to admissible actions. We propose a procedure that conditions on existing demonstrations and semantically translates the plans to admissible actions. Our evaluation in the recent VirtualHome environment shows that the resulting method substantially improves executability over the LLM baseline. The conducted human evaluation reveals a trade-off between executability and correctness but shows a promising sign towards extracting actionable knowledge from language models. Website at this https URL

Table of Contents

  • 1 Introduction
  • 2 Evaluation Framework
  • 2.1 Evaluated Environment: VirtualHome
  • 2.2 Metrics
  • 3 Method
  • 3.1 Querying LLMs for Action Plans
  • 3.2 Admissible Action Parsing by Semantic Translation
  • 3.3 Autoregressive Trajectory Correction
  • 3.4 Dynamic Example Selection for Improved Knowledge Extraction
  • 4 Results
  • 4.1 Do LLMs contain actionable knowledge for high-level tasks?
  • 4.2 How executable are the LLM action plans?
  • 4.3 Can LLM action plans be made executable by proposed procedure?
  • 5 Analysis and Discussions
  • 5.1 Ablation of design decisions
  • 5.2 Are the generated action plans grounded in the environment?
  • 5.3 Effect of Different Translation LMs
  • 5.4 Can LLMs generate actionable programs by following step-by-step instructions?
  • 5.5 Analysis of program length
  • 6 Related Works
  • 7 Conclusion, Limitations & Future Work
  • References
  • A Appendix
  • A.1 Hyperparameter Search
  • A.2 Details of Human Evaluations
  • A.3 All Evaluated Tasks
  • A.4 Natural Language Templates for All Atomic Actions
  • A.5 Random Samples of Action Plans

Knowls

  1. Knowl 1 — Algorithm for Zero-Shot Action Plan Generation with Pre-Trained Language Models

    algorithm

    To extract executable action sequences from frozen, pre-trained language models for high-level tasks without domain-specific fine-tuning, the planning process combines dynamic in-context example retrieval, step-by-step autoregressive generation via a causal planning model (LMPLM_P), and semantic action translation into admissible environment actions via a masked sentence embedding model (LMTLM_T).

    Input: Query task name QQ, demonstration set of pairs of task names and plans (Ti,Ei)i=1N{(T_i, E_i)}_{i=1}^N, planning language model LMPLM_P parameterized by θ\theta, translation language model LMTLM_T with sentence embedding function ff, set of all admissible environment actions Aenv\mathcal{A}_{\text{env}}, candidate sample count kk, scoring balance coefficient β\beta, out-of-distribution termination threshold ϵ\epsilon, maximum step limit MM.
    Output: Sequence of executable environment actions Π=(ae,1∗,ae,2∗,… )\Pi = (a_{e,1}^*, a_{e,2}^*, \dots).
    Find demonstration index i∗=arg⁡max⁡i∈{1,…,N}f(Ti)⋅f(Q)∥f(Ti)∥∥f(Q)∥i^* = \arg\max_{i \in \{1, \dots, N\}} \frac{f(T_i) \cdot f(Q)}{\|f(T_i)\| \|f(Q)\|}
    Initialize prompt text: S←Ti∗+Ei∗+QS \leftarrow T_{i^*} + E_{i^*} + Q
    Initialize generated plan: Π←()\Pi \leftarrow ()
    step ←1\leftarrow 1
    while step≤M\text{step} \le M do
        Sample kk single-step candidate action phrases {a^1,…,a^k}\{\hat{a}_1, \dots, \hat{a}_k\} from LMPLM_P conditioned on prompt SS
        if more than 50% of candidate samples are 0-length then
            break
        end if
        for each candidate sample a^∈{a^1,…,a^k}\hat{a} \in \{\hat{a}_1, \dots, \hat{a}_k\} and each action ae∈Aenva_e \in \mathcal{A}_{\text{env}} do
            Calculate score R(a^,ae)=f(a^)⋅f(ae)∥f(a^)∥∥f(ae)∥+β⋅Pθ(a^)R(\hat{a}, a_e) = \frac{f(\hat{a}) \cdot f(a_e)}{\|f(\hat{a})\| \|f(a_e)\|} + \beta \cdot P_\theta(\hat{a})
        end for
        Select best admissible action ae∗=arg⁡max⁡ae∈Aenvmax⁡a^R(a^,ae)a_e^* = \arg\max_{a_e \in \mathcal{A}_{\text{env}}} \max_{\hat{a}} R(\hat{a}, a_e)
        if max⁡ae,a^R(a^,ae)<ϵ\max_{a_e, \hat{a}} R(\hat{a}, a_e) < \epsilon then
            break
        end if
        Append ae∗a_e^* to prompt SS
        Append ae∗a_e^* to plan Π\Pi
        step ←step+1\leftarrow \text{step} + 1
    end while
    return Π\Pi

    In this procedure, Pθ(a^)=1n∑j=1nlog⁡pθ(xj∣x<j)P_\theta(\hat{a}) = \frac{1}{n} \sum_{j=1}^n \log p_\theta(x_j \mid x_{<j}) represents the mean token log-probability of generated candidate a^=(x1,…,xn)\hat{a} = (x_1, \dots, x_n) under LMPLM_P. Embeddings for all actions in Aenv\mathcal{A}_{\text{env}} can be pre-computed offline to eliminate overhead during runtime inference.

  2. Knowl 2 — Admissible Action Selection and Step-by-Step Trajectory Correction

    model/method

    When large language models generate free-form text plans, actions frequently fail to map to valid environment commands due to non-standard phrasing, unrecognized objects, or lexical ambiguity. To resolve this, generated candidate action phrases are mapped to the closest admissible environment action ae∈Aenva_e \in \mathcal{A}_{\text{env}} using an embedding similarity metric combined with generation likelihood.

    Let f(⋅)f(\cdot) denote a sentence embedding function produced by mean-pooling the final hidden layer of a Sentence-BERT/RoBERTa model. The cosine similarity between a predicted candidate phrase a^\hat{a} and an admissible environment action aea_e is defined as:

    C(f(a^),f(ae))=f(a^)⋅f(ae)∥f(a^)∥∥f(ae)∥C(f(\hat{a}), f(a_e)) = \frac{f(\hat{a}) \cdot f(a_e)}{\|f(\hat{a})\| \|f(a_e)\|}

    At each planning step, rather than translating an entire plan after complete generation, generation and translation are interleaved. The Planning LM samples kk candidates (a^1,…,a^k)(\hat{a}_1, \dots, \hat{a}_k), and the admissible action ae∗a_e^* is selected by maximizing a joint objective that balances semantic closeness and model confidence:

    ae∗=arg⁡max⁡ae∈Aenv[max⁡a^(C(f(a^),f(ae))+β⋅Pθ(a^))]a_e^* = \arg\max_{a_e \in \mathcal{A}_{\text{env}}} \left[ \max_{\hat{a}} \left( C(f(\hat{a}), f(a_e)) + \beta \cdot P_\theta(\hat{a}) \right) \right]

    where β\beta is a weighting hyperparameter (set to 0.30.3) and Pθ(a^)=1n∑j=1nlog⁡pθ(xj∣x<j)P_\theta(\hat{a}) = \frac{1}{n} \sum_{j=1}^n \log p_\theta(x_j \mid x_{<j}) is the mean log-probability under the causal planning model parameterized by θ\theta.

    The selected admissible action ae∗a_e^* is appended back into the prompt for the next step, ensuring that all subsequent generations are conditioned directly on valid, admissible steps. Early termination is triggered if max⁡ae,a^[C(f(a^),f(ae))+βPθ(a^)]<ϵ\max_{a_e, \hat{a}} [C(f(\hat{a}), f(a_e)) + \beta P_\theta(\hat{a})] < \epsilon (detecting out-of-distribution actions) or if over 50% of the sampled candidate action phrases are of length zero.

  3. Knowl 3 — Dynamic In-Context Demonstration Selection via Semantic Task Embedding

    model/method

    Fixed in-context prompting often biases large language models toward implicit environmental assumptions from the single prompt example that may not match the queried task (e.g., assuming an agent is seated at a computer). To provide weak supervision at inference time without updating model parameters, an in-context demonstration is retrieved dynamically from an offline demonstration set {(Ti,Ei)}i=1N\{(T_i, E_i)\}_{i=1}^N, where TiT_i is a task name and EiE_i is its annotated executable plan.

    Using the same frozen sentence embedding model f(⋅)f(\cdot) employed for action translation, the cosine similarity between the query task name QQ and each candidate task name TiT_i is evaluated. The pair (T∗,E∗)(T^*, E^*) is chosen via:

    T∗=arg⁡max⁡Tif(Ti)⋅f(Q)∥f(Ti)∥∥f(Q)∥T^* = \arg\max_{T_i} \frac{f(T_i) \cdot f(Q)}{\|f(T_i)\| \|f(Q)\|}

    The prompt provided to the causal Planning LM is then initialized with the formatted demonstration T∗+E∗+QT^* + E^* + Q prior to starting autoregressive plan generation.

  4. Knowl 4 — Executability, Sequence Similarity, and Correctness of Language Model Planners

    data/table

    In the VirtualHome simulation environment across 7 household scenes and 88 held-out test tasks, action plans generated by vanilla causal language models, supervised models, human annotators, and translated language model planners were evaluated. Executability measures the proportion of plans that parse syntactically and satisfy all precondition and postcondition constraints. LCS is the longest common subsequence match between generated and human-expert programs normalized by the maximum length of the two. Correctness is measured via Amazon Mechanical Turk, where 10 human raters per task evaluate whether the action sequence successfully completes the given task.

    Language Model Executability LCS Correctness
    Vanilla GPT-2 117M 18.66% 3.19% 15.81% (4.90%)
    Vanilla GPT-2 1.5B 39.40% 7.78% 29.25% (5.28%)
    Vanilla Codex 2.5B 17.62% 15.57% 63.08% (7.12%)
    Vanilla GPT-Neo 2.7B 29.92% 11.52% 65.29% (9.08%)
    Vanilla Codex 12B 18.07% 16.97% 64.87% (5.41%)
    Vanilla GPT-3 13B 25.87% 13.40% 49.44% (8.14%)
    Vanilla GPT-3 175B 7.79% 17.82% 77.86% (6.42%)
    Human 100.00% N/A 70.05% (5.44%)
    Fine-tuned GPT-3 13B 66.07% 34.08% 64.92% (5.96%)
    Translated Codex 12B (Ours) 78.57% 24.72% 54.88% (5.90%)
    Translated GPT-3 175B (Ours) 73.05% 24.09% 66.13% (8.38%)

    While vanilla GPT-3 175B achieves the highest raw human correctness score (77.86%, surpassing human-written plans restricted by environment syntax), its plans achieve only 7.79% executability in the environment. Applying the translation procedure elevates executability to 78.57% for Codex 12B and 73.05% for GPT-3 175B, while increasing LCS to ~24%, though accompanied by a modest reduction in human-perceived correctness.

  5. Knowl 5 — Joint Executability and Semantic Correctness Grounding Rate

    empirical result

    To measure true environment grounding, the percentage of generated plans that are simultaneously executable in the simulator and judged semantically correct by human annotators (defined as receiving ≥70%\ge 70\% 'Yes' votes across 10 human raters) was evaluated across 88 VirtualHome tasks:

    • Human-written plans: 65.91% correct and executable (100.00% executable, 70.05% mean correctness)
    • Translated GPT-3 175B: 35.23%
    • Translated Codex 12B: 27.27%
    • Vanilla GPT-3 175B: 6.82%
    • Vanilla Codex 12B: 4.55%
    • Vanilla GPT-3 13B: 2.27%
    • Vanilla GPT-2 1.5B: 1.14%
    • Vanilla GPT-2 0.1B: 0.00%

    Without semantic translation and trajectory correction, vanilla language models scale poorly in joint grounding rate despite increasing model parameter size. Semantic translation combined with autoregressive correction increases the proportion of valid, grounded plans by roughly 5×5\times to 6×6\times over vanilla baselines.

  6. Knowl 6 — Ablation of Translation, Dynamic Selection, and Trajectory Correction

    data/table

    An ablation study isolating the three components of the plan translation framework—action translation via sentence embeddings, dynamic demonstration selection, and step-by-step autoregressive trajectory correction—demonstrates the relative contribution of each mechanism in VirtualHome across 88 tasks.

    Methods Executability LCS
    Translated Codex 12B 78.57% 24.72%
    – w/o Action Translation 31.49% 22.53%
    – w/o Dynamic Example 50.86% 22.84%
    – w/o Trajectory Correction 55.19% 24.43%
    Translated GPT-3 175B 73.05% 24.09%
    – w/o Action Translation 36.04% 24.31%
    – w/o Dynamic Example 60.82% 22.92%
    – w/o Trajectory Correction 40.10% 24.98%

    Removing action translation produces the largest drop in plan executability (dropping from 78.57% to 31.49% on Codex 12B and from 73.05% to 36.04% on GPT-3 175B). Removing step-by-step trajectory correction reduces executability by 23.38% on Codex 12B and 32.95% on GPT-3 175B, confirming that conditioning future steps on translated admissible actions prevents compounding errors.

  7. Knowl 7 — Comparison of Embedding Representations for Action Translation

    data/table

    Evaluation of different sentence embedding architectures used as the Translation LM (LMTLM_T) to project free-form action phrases to the 47,522 admissible environment actions in VirtualHome reveals that contextual Transformer embeddings are critical, whereas model scale among BERT/RoBERTa variants has minimal impact.

    Translation LM Parameter Count Executability LCS
    Codex 12B as Planning LM
    Avg. GloVe embeddings - 46.92% 9.71%
    Sentence BERT (base) 110M 73.21% 24.10%
    Sentence BERT (large) 340M 75.16% 20.79%
    Sentence RoBERTa (base) 125M 74.35% 22.82%
    Sentence RoBERTa (large) 325M 78.57% 24.72%
    GPT-3 175B as Planning LM
    Avg. GloVe embeddings - 47.40% 12.16%
    Sentence BERT (base) 110M 77.60% 24.49%
    Sentence BERT (large) 340M 67.86% 21.24%
    Sentence RoBERTa (base) 125M 72.73% 23.64%
    Sentence RoBERTa (large) 325M 73.05% 24.09%

    Non-contextual averaged GloVe word embeddings perform poorly, achieving only ~47% executability and ~10-12% LCS. In contrast, all Sentence-BERT and Sentence-RoBERTa variants achieve 67%–78% executability and ~21%–25% LCS, indicating that pre-trained contextual embeddings sufficiently capture single-step semantic mappings.

  8. Knowl 8 — Zero-Shot Plan Synthesis from Step-by-Step Instructions vs Supervised Baselines

    data/table

    When pre-trained language models are prompted with both a high-level task name and step-by-step natural language instructions (how-to descriptions) in the prompt, zero-shot translated language models match the sequence similarity (LCS) of an in-domain supervised model trained on human-annotated data.

    Methods Executability LCS
    Translated Codex 12B 78.57% 32.87%
    Translated GPT-3 175B 74.15% 31.05%
    Supervised LSTM - 34.00%

    Without fine-tuning on any domain data from VirtualHome, Translated Codex 12B and Translated GPT-3 175B achieve 32.87% and 31.05% LCS respectively, approaching the 34.00% LCS achieved by a supervised LSTM trained directly on VirtualHome human annotations, while sustaining >74% plan executability.

  9. Knowl 9 — Plan Length and Failure Modes Across Language Model Scales

    data/table

    Evaluating the average program step length across 88 tasks illustrates the distinct failure modes exhibited by smaller versus larger language models.

    Methods Executability Average Length
    Vanilla GPT-2 1.5B 39.40% 4.24
    Vanilla Codex 12B 18.07% 7.22
    Vanilla GPT-3 175B 7.79% 9.716
    Translated Codex 12B 78.57% 7.13
    Translated GPT-3 175B 73.05% 7.36
    Human 100.00% 9.66

    Smaller models such as Vanilla GPT-2 1.5B exhibit artificially high executability (39.40%) by generating short programs (average length 4.24) that merely rephrase the goal in a single step (e.g., 'Go to sleep' →\rightarrow 'Go to bed') or repeat the prompt example. Larger models like Vanilla GPT-3 175B generate expressive, human-length plans (average length 9.72 vs Human 9.66) but fail on executability (7.79%). Translated models preserve plan expressiveness (length ~7.1–7.4) while attaining >73% executability.

  10. Knowl 10 — Limitations in Environmental Grounding and Sensorimotor Execution

    limitation

    The approach of extracting zero-shot action plans from pre-trained language models has several identified limitations:

    1. Executability vs. Correctness Trade-Off: Semantic action translation introduces errors when mapping compounded natural language actions (e.g., 'brush teeth with toothbrush and toothpaste') into atomic steps. Furthermore, when the environment lacks specific objects or actions required for a task, translation can map to inappropriate actions or trigger premature termination, resulting in lower human-rated correctness for translated models compared to unconstrained vanilla models.
    2. Mid-Level Planning Scope: The method addresses mid-level symbolic planning (generating actions such as 'grab cup' that satisfy precondition/postcondition logic) and assumes the presence of a low-level controller for sensorimotor execution (continuous navigation, visual interaction masks, trajectory control), which still requires domain-specific training.
    3. Lack of Perceptual Feedback: The planner operates open-loop without visual observation context or dynamic environment state feedback. It cannot resolve object instance ambiguities (e.g., distinguishing between multiple plates in tasks like 'stack two plates on the right side of a cup').

Coverage note — None was omitted; all key contributions, algorithms, equations, empirical findings, ablation studies, and limitations are fully covered.

References

  1. 1.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3674–3683, 2018.
  2. 2.Yoav Artzi and Luke Zettlemoyer. Weakly supervised learning of semantic parsers for mapping instructions to actions. Transactions of the Association for Computational Linguistics, 1:49–62, 2013.
  3. 3.BIG-bench collaboration. Beyond the imitation game: Measuring and extrapolating the capabilities of language models. In preparation, 2021. URL https://github.com/google/BIG-bench/.
  4. 4.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Kohd, Mark Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, Julian Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Rob Reich, Hongyu Ren, Frieda Rong, Yusuf Roohani, Camilo Ruiz, Jack Ryan, Christopher Ré, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishnan Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. On the opportunities and risks of foundation models, 2021.
  5. 5.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  6. 6.Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055, 2017.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harri Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  9. 9.Joe Davison, Joshua Feldman, and Alexander M Rush. Commonsense knowledge mining from pretrained models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1173–1178, 2019.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  11. 11.Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker-follower models for vision-and-language navigation. arXiv preprint arXiv:1806.02724, 2018.
  12. 12.Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama. From language to goals: Inverse reinforcement learning for vision-based instruction following. arXiv preprint arXiv:1902.07742, 2019.
  13. 13.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723, 2020.
  14. 14.Brent Harrison and Mark O Riedl. Learning from stories: using crowdsourced narratives to train virtual agents. In Twelfth Artificial Intelligence and Interactive Digital Entertainment Conference, 2016.
  15. 15.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al. Measuring coding challenge competence with apps. arXiv preprint arXiv:2105.09938, 2021.
  16. 16.Felix Hill, Sona Mokra, Nathaniel Wong, and Tim Harley. Human instruction-following with deep reinforcement learning via transfer-learning from text. arXiv preprint arXiv:2005.09382, 2020.
  17. 17.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8): 1735–1780, 1997.
  18. 18.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019.
  19. 19.Gabriel Ilharco, Rowan Zellers, Ali Farhadi, and Hannaneh Hajishirzi. Probing text models for common ground with visual representations. arXiv e-prints, pages arXiv–2005, 2020.
  20. 20.Peter A Jansen. Visually-grounded planning without vision: Language models infer detailed plans from high-level instructions. arXiv preprint arXiv:2009.14259, 2020.
  21. 21.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438, 2020.
  22. 22.Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017.
  23. 23.Belinda Z Li, Maxwell Nye, and Jacob Andreas. Implicit representations of meaning in neural language models. arXiv preprint arXiv:2106.00737, 2021.
  24. 24.Shuang Li, Xavier Puig, Yilun Du, Clinton Wang, Ekin Akyurek, Antonio Torralba, Jacob Andreas, and Igor Mordatch. Pre-trained language models for interactive decision-making. arXiv preprint arXiv:2202.01771, 2022.
  25. 25.Yuan-Hong Liao, Xavier Puig, Marko Boben, Antonio Torralba, and Sanja Fidler. Synthesizing environment-aware activities via activity sketches. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6291–6299, 2019.
  26. 26.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804, 2021.
  27. 27.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  28. 28.Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. Pretrained transformers as universal computation engines. arXiv preprint arXiv:2103.05247, 2021.
  29. 29.Corey Lynch and Pierre Sermanet. Grounding language in play. arXiv preprint arXiv:2005.07648, 2020.
  30. 30.Corey Lynch and Pierre Sermanet. Language conditioned imitation learning over unstructured data. Proceedings of Robotics: Science and Systems. doi, 10, 2021.
  31. 31.Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. Improving vision-and-language navigation with image-text pairs from the web. In European Conference on Computer Vision, pages 259–274. Springer, 2020.
  32. 32.Suvir Mirchandani, Siddharth Karamcheti, and Dorsa Sadigh. Ella: Exploration through learned language abstraction. arXiv preprint arXiv:2103.05825, 2021.
  33. 33.Dipendra Misra, Kejia Tao, Percy Liang, and Ashutosh Saxena. Environment-driven lexicon induction for high-level instructions. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 992–1002, 2015.
  34. 34.Dipendra K Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena. Tell me dave: Context-sensitive grounding of natural language to manipulation instructions. The International Journal of Robotics Research, 35(1-3):281–300, 2016.
  35. 35.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543, 2014.
  36. 36.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  37. 37.Gabriel Poesia, Oleksandr Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. Synchromesh: Reliable code generation from pre-trained language models. arXiv preprint arXiv:2201.11227, 2022.
  38. 38.Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. Virtualhome: Simulating household activities via programs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8494–8502, 2018.
  39. 39.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  41. 41.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
  42. 42.Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020.
  43. 43.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. Learning to retrieve prompts for in-context learning. arXiv preprint arXiv:2112.08633, 2021.
  44. 44.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al. Habitat: A platform for embodied ai research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9339–9347, 2019.
  45. 45.Pratyusha Sharma, Antonio Torralba, and Jacob Andreas. Skill induction and planning with latent language. arXiv preprint arXiv:2110.01517, 2021.
  46. 46.Jianhao Shen, Yichun Yin, Lin Li, Lifeng Shang, Xin Jiang, Ming Zhang, and Qun Liu. Generate & rank: A multi-task framework for math word problems. arXiv preprint arXiv:2109.03034, 2021.
  47. 47.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10740–10749, 2020.
  48. 48.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768, 2020.
  49. 49.Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, and Gaurav Sukhatme. Embodied bert: A transformer model for embodied, language-guided visual task completion. arXiv preprint arXiv:2108.04927, 2021.
  50. 50.Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. olmpics-on what language model pre-training captures. Transactions of the Association for Computational Linguistics, 8: 743–758, 2020.
  51. 51.Moritz Tenorth, Daniel Nyga, and Michael Beetz. Understanding and executing instructions for everyday manipulation tasks from the world wide web. In 2010 ieee international conference on robotics and automation, pages 1486–1491. IEEE, 2010.
  52. 52.Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. arXiv preprint arXiv:2106.13884, 2021.
  53. 53.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  54. 54.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6629–6638, 2019.
  55. 55.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.

Citation

MLA
Huang, W., et al. “Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents”. International Conference on Machine Learning, vol. 162, 2022, pp. 9118–47, https://proceedings.mlr.press/v162/huang22a.html.
APA
Huang, W., Abbeel, P., Pathak, D., & Mordatch, I. (2022). Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents. International Conference on Machine Learning, 162, 9118–9147. https://proceedings.mlr.press/v162/huang22a.html
Chicago
Huang, W., P. Abbeel, D. Pathak, and I. Mordatch. 2022. “Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents”. International Conference on Machine Learning 162: 9118–47. https://proceedings.mlr.press/v162/huang22a.html.
Harvard
Huang, W. et al. (2022) “Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents”, International Conference on Machine Learning. PMLR, pp. 9118–9147. Available at: https://proceedings.mlr.press/v162/huang22a.html.
Vancouver
1. Huang W, Abbeel P, Pathak D, Mordatch I (2022) Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents. In: International Conference on Machine Learning. PMLR, pp 9118–9147

BibTeX

@InProceedings{pmlr-v162-huang22a,
  title = 	 {Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents},
  author =       {Huang, Wenlong and Abbeel, Pieter and Pathak, Deepak and Mordatch, Igor},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {9118--9147},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/huang22a/huang22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/huang22a.html},
  abstract = 	 {Can world knowledge learned by large language models (LLMs) be used to act in interactive environments? In this paper, we investigate the possibility of grounding high-level tasks, expressed in natural language (e.g. “make breakfast”), to a chosen set of actionable steps (e.g. “open fridge”). While prior work focused on learning from explicit step-by-step examples of how to act, we surprisingly find that if pre-trained LMs are large enough and prompted appropriately, they can effectively decompose high-level tasks into mid-level plans without any further training. However, the plans produced naively by LLMs often cannot map precisely to admissible actions. We propose a procedure that conditions on existing demonstrations and semantically translates the plans to admissible actions. Our evaluation in the recent VirtualHome environment shows that the resulting method substantially improves executability over the LLM baseline. The conducted human evaluation reveals a trade-off between executability and correctness but shows a promising sign towards extracting actionable knowledge from language models.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/