ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Shibo HaoTianyang LiuZhen WangZhiting Hu

article2023NeurIPS311 citations

Proposes ToolkenGPT, a framework that represents external tools as learned token embeddings within the vocabulary of a frozen language model, enabling plug-and-play selection across massive tool sets without costly fine-tuning or prompt-length constraints.

Listen

Large language models often struggle with tasks requiring strict factual accuracy, complex numerical reasoning, and environment-grounded planning. Augmenting these models with external tools—such as calculators, database queries, and robotic actions—addresses these limitations, but existing paradigms have severe operational trade-offs. Fine-tuning models directly is computationally expensive, rigid, and hard to update as new tools emerge. Conversely, relying on few-shot demonstrations within prompts restricts the model to a small number of tools due to strict context window limits and often fails to provide deep operational understanding. To resolve these challenges, the article evaluates ToolkenGPT, a hybrid framework designed to let language models integrate large tool libraries efficiently without modifying base model parameters.

The framework represents external tools as specialized tokens, termed "toolkens," and optimizes only lightweight toolken embedding vectors appended to the model's prediction head while keeping the primary language model frozen. During text generation, the system predicts toolkens just as it would standard word tokens. When a toolken is triggered, the model pauses generation, temporarily shifts to a focused tool mode with specific demonstrations to complete necessary arguments, executes the tool call, and incorporates the output back into the primary generation sequence. The authors tested this method across three major domains: numerical problem solving on enhanced and synthetic arithmetic datasets, knowledge-based question answering using over two hundred database relations from Wikidata, and robotic task planning in the VirtualHome simulation environment.

The empirical results show substantial performance gains across all evaluated settings. In complex multi-hop numerical reasoning involving thirteen distinct operators, ToolkenGPT achieved an accuracy of 15%, substantially outperforming standard reasoning and tool-augmented baselines, which recorded 3% and 6% respectively. In knowledge-based queries with larger toolsets of up to 234 relation APIs, conventional prompt-based tool approaches collapsed due to token limits, whereas the proposed framework maintained superior accuracy whether trained on ground-truth demonstrations or synthetic data. In embodied task planning across 58 discrete robot actions and objects, ToolkenGPT achieved a 68% task success rate, nearly doubling the 38% rate of existing grounded decoding approaches by properly learning environmental constraints from training demonstrations.

From a resource and risk standpoint, these findings show that high tool proficiency can be achieved at minimal operational cost. The decoupled embedding architecture requires roughly two minutes of single-GPU training compared to forty minutes across eight specialized GPUs for low-rank fine-tuning, dramatically reducing development overhead while preventing destructive catastrophic forgetting in the base model. Organizations can dynamically plug in, update, or remove domain-specific tools without retraining underlying models. For deployment, stakeholders should consider adopting toolken-based embeddings for complex multi-tool architectures and can leverage synthetic training data generation when labeled real-world demonstrations are scarce.

Decision-makers should nevertheless account for key boundary conditions. Embedding effectiveness is bounded by the quality and domain coverage of available training demonstrations, as seen when synthetic data exhibited performance drops relative to fully supervised datasets. Additionally, evaluations were conducted within controlled mathematical, factual retrieval, and simulated household environments. Further testing in dynamic, non-deterministic production environments is necessary to confirm reliability before deploying this approach to mission-critical autonomous systems.

arXiv: 2305.11554
Cover for ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Abstract

Augmenting large language models (LLMs) with external tools has emerged as a promising approach to solving complex problems. However, traditional methods, which fine-tune LLMs with tool demonstration data, can be both costly and restricted to a predefined set of tools. Recent in-context learning paradigm alleviates these issues, but the limited context length only allows for a few shots of demonstrations, leading to suboptimal understandings of the tools. Moreover, when there are numerous tools to choose from, in-context learning could completely fail to work. In this paper, we propose an alternative approach, ToolkenGPT, which combines the benefits of both sides. Our approach represents each tool as a token (“toolken”) and learns an embedding for it, enabling tool calls in the same way as generating a regular word token. Once a toolken is triggered, the LLM is prompted to complete arguments for the tool to execute. ToolkenGPT offers the flexibility to plug in an arbitrary number of tools by expanding the set of toolkens on the fly. In addition, it improves tool use by allowing extensive demonstration data for learning the toolken embeddings. In diverse domains, including numerical reasoning, knowledge-based question answering, and embodied plan generation, our approach effectively augments LLMs with tools and substantially outperforms various latest baselines. ToolkenGPT demonstrates the promising ability to use relevant tools from a large tool set in complex scenarios.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 ToolkenGPT for Mastering Massive Tools
  • 3.1 Framework Overview
  • 3.2 Learning Toolken Embeddings
  • 4 Experiments
  • 4.1 Numerical Reasoning
  • 4.2 Knowledge-based Question Answering
  • 4.3 Embodied Plan Generation
  • 4.4 Analysis
  • 5 Conclusion
  • References
  • A Details of Numerical Reasoning
  • A.1 GSM8K-XL
  • A.1.1 Data Synthesis
  • A.1.2 Training Details
  • A.1.3 Prompt for GSM8K-XL Dataset
  • A.2 FuncQA
  • A.2.1 Training Details
  • A.2.2 Prompt for Synthetic Training Data
  • A.2.3 Prompt for FuncQA One-Hop
  • A.2.4 Prompt for FuncQA Multi-Hop
  • B Details of Knowledge-based QA
  • B.1 Getting Text Description
  • B.2 Synthetic Data
  • B.3 Training Details
  • C Details of Embodied Plan Generation
  • C.1 Preprocessing
  • C.2 Prompts
  • C.3 Training Details
  • D Computational Resources
  • E Safeguard Statement

Knowls

  1. Knowl 1 — ToolkenGPT Framework for Tool Use via Toolken Embeddings

    model/method

    ToolkenGPT augments a frozen large language model (LLM) with external tools by representing each tool as an independent token called a "toolken" in an expanded vocabulary. Each tool τ∈T\tau \in \mathcal{T} is assigned a continuous toolken embedding vector. The collection of toolken embeddings forms a matrix Wτ∈R∣T∣×dW_\tau \in \mathbb{R}^{|\mathcal{T}| \times d}, where ∣T∣|\mathcal{T}| is the number of tools and dd is the hidden dimension of the LLM.

    During text generation, the LLM alternates between two operating modes:

    1. Reasoning Mode: The language model generates tokens over the union of regular vocabulary words and tool tokens (V∪TV \cup \mathcal{T}). If a standard word token is predicted, generation continues in the reasoning mode.
    2. Tool Mode: When a toolken τ\tau is predicted, generation pauses and the LLM switches to the tool mode. In this mode, the LLM is provided with a prompt containing in-context demonstrations specific to the selected tool τ\tau. The LLM completes the arguments for the tool call (e.g., using syntax such as [tool](arguments)). The tool is then executed externally with the generated arguments, and the resulting output is returned and appended to the context in the reasoning mode.

    This design separates tool selection (handled globally via token prediction over WτW_\tau) from argument generation (handled locally via focused, single-tool in-context demonstrations), allowing arbitrary new tools to be plugged in or unplugged on the fly by expanding or modifying WτW_\tau.

  2. Knowl 2 — Toolken Next-Token Prediction Formulation

    equation

    In ToolkenGPT, the probability distribution over the next token tit_i given the preceding token sequence t<it_{<i} is computed by concatenating the frozen language model vocabulary embedding head Wv∈R∣V∣×dW_v \in \mathbb{R}^{|V| \times d} with the trainable toolken embedding matrix Wτ∈R∣T∣×dW_\tau \in \mathbb{R}^{|\mathcal{T}| \times d}:

    P(ti∣t<i)=softmax([Wv;Wτ]⋅hi−1)P(t_i \mid t_{<i}) = \text{softmax}\left([W_v; W_\tau] \cdot h_{i-1}\right)

    where:

    • VV is the regular vocabulary of the language model, and ∣V∣|V| is its vocabulary size.
    • T={τ1,τ2,… }\mathcal{T} = \{\tau_1, \tau_2, \dots\} is the set of available tools, with ∣T∣|\mathcal{T}| being the total number of tools.
    • ti∈V∪Tt_i \in V \cup \mathcal{T} is the predicted token, which can be either a regular vocabulary word or a toolken.
    • [Wv;Wτ]∈R(∣V∣+∣T∣)×d[W_v; W_\tau] \in \mathbb{R}^{(|V| + |\mathcal{T}|) \times d} denotes the row-wise concatenation of the word embedding matrix and the toolken embedding matrix.
    • hi−1∈Rdh_{i-1} \in \mathbb{R}^d is the last hidden state of the language model produced from the context t<it_{<i}.
    • softmax(⋅)\text{softmax}(\cdot) normalizes the combined logit vector across all ∣V∣+∣T∣|V| + |\mathcal{T}| tokens into a valid probability distribution.
  3. Knowl 3 — Training Toolken Embeddings via Masked Sequence Loss

    model/method

    ToolkenGPT trains toolken embeddings while keeping all parameters of the underlying large language model (LLM) completely frozen. The trainable parameter set is strictly the toolken embedding matrix Wτ∈R∣T∣×dW_\tau \in \mathbb{R}^{|\mathcal{T}| \times d}. Because gradients do not backpropagate through the transformer layers of the LLM and only require computing gradients with respect to WτW_\tau from precomputed final hidden states hi−1h_{i-1}, training maintains minimal GPU memory overhead, comparable to LLM inference.

    Training uses a dataset D={(s,s′)}\mathcal{D} = \{(s, s')\} of paired token sequences:

    • s=(t1,t2,…,tN)s = (t_1, t_2, \dots, t_N) is the ground-truth text sequence containing the natural language context and the returned results of tool executions.
    • s′=(t1′,t2′,…,tN′)s' = (t'_1, t'_2, \dots, t'_N) is a parallel target sequence where the first token of each tool output span in ss is replaced by the corresponding toolken token τ∈T\tau \in \mathcal{T}, and all subsequent tokens of that tool's output are replaced by a special token [N/A][\text{N/A}].

    The training objective optimizes WτW_\tau via the negative log-likelihood loss over valid tokens:

    L(Wτ)=∑(s,s′)∈D∑i=1N−log⁡P(ti′∣t<i)1ti′≠[N/A]\mathcal{L}(W_\tau) = \sum_{(s, s') \in \mathcal{D}} \sum_{i=1}^N -\log P(t'_i \mid t_{<i}) \mathbf{1}_{t'_i \neq [\text{N/A}]}

    where P(ti′∣t<i)P(t'_i \mid t_{<i}) is the next-token probability computed using the concatenated head [Wv;Wτ][W_v; W_\tau], and 1ti′≠[N/A]\mathbf{1}_{t'_i \neq [\text{N/A}]} is an indicator function that masks out loss calculation over positions marked with [N/A][\text{N/A}].

  4. Knowl 4 — Synthetic Generation of Tool Demonstration Data

    model/method

    When in-domain tool demonstration data is unavailable, ToolkenGPT creates training data by synthesizing tool demonstrations using an instruction-tuned LLM (such as ChatGPT) through self-instruct prompting.

    The synthesis pipeline operates as follows:

    1. The teacher LLM is prompted with the tool name, its functional documentation/description, and two few-shot demonstration examples demonstrating tool invocation using an explicit syntax (e.g., The capital of U.S. is <capital>("U.S.")="Washington D.C.").
    2. The model generates diverse single-step or multi-step questions and answers containing tool calls with placeholder arguments.
    3. Generated examples undergo syntactic filtering to remove malformed instances and deduplication to remove repeated queries.
    4. The synthesized texts are converted into paired training sequences (s,s′)(s, s'), where natural language spans representing tool outputs are substituted with the corresponding toolken and trailing [N/A][\text{N/A}] masks.

    This distillation process allows learning implicit tool calling semantics directly from documentation without requiring manual human annotations.

  5. Knowl 5 — Numerical Reasoning Performance on GSM8K-XL and FuncQA Benchmarks

    data/table

    ToolkenGPT was evaluated on two numerical reasoning benchmarks using LLaMA-33B as the base model (except for 0-shot ChatGPT):

    • GSM8K-XL: A 568-example test set derived from GSM8K by magnifying numbers into cubic magnitudes to test arithmetic tool use with 4 basic operators (+,−,×,÷+, -, \times, \div). Toolken embeddings were trained on 5,054 examples.
    • FuncQA: A synthetic benchmark requiring 13 arithmetic operators (e.g., power, sqrt, lcm, gcd, choose), split into a 68-question one-hop subset (FuncQA-one) and a 60-question multi-hop subset (FuncQA-multi). Toolken embeddings were trained purely on 611 one-hop synthetic examples.

    Accuracy was evaluated by exact match (rounded to two decimal places) for GSM8K-XL and FuncQA-one, and within a 0.1% error margin for FuncQA-multi.

    Method GSM8K-XL (4 tools) FuncQA (13 tools)
    One-Hop Multi-Hops
    0-shot ChatGPT 0.17 0.55 0.09
    Chain-of-Thought (CoT) 0.18 0.20 0.03
    ReAct 0.32 0.57 0.06
    ToolkenGPT 0.33 0.73 0.15

    When the toolset expands from 4 to 13 operators in FuncQA, ReAct degrades because in-context prompts cannot accommodate demonstrations for all tools. In contrast, ToolkenGPT achieves 0.73 on FuncQA-one and generalizes to 0.15 on FuncQA-multi despite training only on one-hop synthetic data without multi-hop Chain-of-Thought examples.

  6. Knowl 6 — Knowledge-Based Question Answering Scaling over Massive APIs

    empirical result

    On the KAMEL knowledge-base question answering benchmark, ToolkenGPT was evaluated across test subsets with increasing numbers of relational query tools (30, 60, 100, and 234 tools from Wikidata), using LLaMA-13B as the base model. Each subset contains 500 questions.

    Two ToolkenGPT variants were tested:

    • ToolkenGPT (sup): Toolken embeddings trained via supervised learning on 200 examples per relation sampled from KAMEL.
    • ToolkenGPT (syn): Toolken embeddings trained on synthetic data generated via ChatGPT from relation descriptions (averaging 40 synthetic examples per relation).

    Key findings:

    1. Zero-shot direct prompting of LLaMA-13B achieves ~20% accuracy across all tool subset sizes, indicating parametric memory limits.
    2. In-context learning (ICL) baselines fail to scale: standard few-shot ICL can only accommodate up to 30 tool demonstrations within the 2048-token context window, and zero-shot ICL with descriptions (ICL desc) suffers severe performance degradation as tools increase (falling below 20% accuracy).
    3. ToolkenGPT maintains robust performance across all tool scales: ToolkenGPT (sup) achieves near 1.0 (100%) accuracy across 30, 60, 100, and 234 tools. ToolkenGPT (syn) consistently outperforms all ICL baselines across all tool set sizes, achieving over 50% accuracy on all subsets without seeing any in-domain training data.
  7. Knowl 7 — Embodied Plan Generation in VirtualHome

    data/table

    ToolkenGPT was evaluated on embodied agent task planning using a 297-task dataset derived from the VirtualHome / ActivityPrograms environment (247 training tasks, 50 test tasks). The environment encompasses 58 toolkens: 25 executable action verbs, 32 environment objects, and 1 [END] token. The base model is LLaMA-13B.

    Plan quality was assessed using four metrics:

    • Grounding: Proportion of generated scripts where all actions and objects exist in the admissible environment candidate set.
    • Executable: Proportion of scripts that execute in the VirtualHome simulator without physical rule violations.
    • Success: Proportion of scripts whose final simulated state graph matches the ground-truth final state graph.
    • Success (R): Relaxed success rate, measuring whether the simulator reached the target final state at any step during execution.
    Method Grounding Executable Success Success (R)
    In-context Learning 0.74 0.42 0.20 0.30
    + Translation 1.00 0.52 0.24 0.32
    + Grounded Decoding 1.00 0.66 0.38 0.42
    ToolkenGPT 1.00 0.82 0.68 0.70

    While post-processing methods (Translation) and decoding constraints (Grounded Decoding) enforce 1.00 grounding, they fail to align with environment-specific affordances (e.g., mistranslating the prompt instruction "sit at desk" into [SIT] <desk>, which violates the simulator's non-sittable desk constraint). ToolkenGPT learns these implicit environment affordances directly from training demonstrations, predicting [SIT] <chair> and reaching 0.82 executability and 0.68 exact success.

  8. Knowl 8 — Computational Efficiency: ToolkenGPT vs. LoRA Fine-Tuning

    data/table

    The computational efficiency and performance of ToolkenGPT were compared against parameter-efficient fine-tuning via LoRA on the FuncQA dataset using LLaMA-7B as the base model.

    Method One-Hop Acc Multi-Hop Acc Computing Resource Training Time
    Prompting 0.10 0.00 - -
    ReAct 0.40 0.03 - -
    Fine-tune w/ LoRA 0.62 0.07 8 ×\times NVIDIA A100 (80GB) 40 min
    ToolkenGPT 0.55 0.06 1 ×\times NVIDIA RTX 3090 (24GB) 2 min

    While LoRA modifies internal attention projection matrices and requires backpropagation across the entire LLM, ToolkenGPT only optimizes the toolken embedding matrix Wτ∈R∣T∣×dW_\tau \in \mathbb{R}^{|\mathcal{T}| \times d}. This reduces hardware requirements from a multi-GPU cluster (8×A1008 \times \text{A100}) to a single commodity GPU (1×RTX 30901 \times \text{RTX 3090}) and shortens training time from 40 minutes to 2 minutes (20×20\times speedup), while achieving competitive accuracy (0.55 vs 0.62 one-hop, 0.06 vs 0.07 multi-hop).

  9. Knowl 9 — Ablation Analysis of Toolken Embeddings vs. Separate Argument Tool Mode

    data/table

    To determine whether ToolkenGPT's performance gains stem from improved tool selection via toolken embeddings or from the dedicated argument completion sub-routine (tool mode), an ablation study was conducted on FuncQA using LLaMA-30B.

    The baseline ReAct + Tool mode combines standard ReAct-style text prompting for tool selection with the isolated tool-mode prompting scheme for argument generation.

    Method One-Hop Acc Multi-Hop Acc
    ReAct 0.57 0.06
    ReAct + Tool mode 0.60 0.07
    ToolkenGPT 0.73 0.15

    Adding the specialized tool mode to ReAct yields modest gains (from 0.57 to 0.60 on one-hop, 0.06 to 0.07 on multi-hop) by providing cleaner argument completion context. However, full ToolkenGPT achieves significantly higher accuracy (0.73 and 0.15), demonstrating that embedding-based toolken prediction is the primary driver for accurate tool identification and triggering.

  10. Knowl 10 — Impact of Training Data Scale and Source on Toolken Learning

    data/table

    The effect of demonstration sample size and data origin on toolken embedding quality was evaluated on a 30-relation test subset of the KAMEL dataset using LLaMA-13B. Embeddings were trained with 10, 20, or 40 examples per tool using either in-domain supervised training data or synthetic data generated by ChatGPT.

    Examples per Tool Synthetic Data Acc Supervised Data Acc
    10 0.36 0.56
    20 0.46 0.90
    40 0.52 0.95

    Performance scales monotonically with the number of demonstration examples for both data sources. Supervised data yields superior performance at every sample size, reaching 0.95 accuracy with 40 examples per tool, compared to 0.52 for synthetic data. The gap is attributed to the distribution shift between ChatGPT-synthesized training expressions and the KAMEL test distribution.

Coverage note — Extensive prompt templates provided in Appendices A, B, and C were omitted from individual knowls and instead summarized conceptually in the methodology and experimental setup knowls.

References

  1. 1.Razvan Azamfirei, Sapna R Kudchadkar, and James Fackler. Large language models and the perils of their hallucinations. Critical Care, 27(1):1–2, 2023.
  2. 2.Michael Bommarito II and Daniel Martin Katz. Gpt takes the bar exam. arXiv preprint arXiv:2212.14402, 2022.
  3. 3.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR, 2022.
  4. 4.Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning, pages 287–318. PMLR, 2023.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  6. 6.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  7. 7.Harrison Chase. LangChain, 10 2022. URL https://github.com/hwchase17/langchain.
  8. 8.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588, 2022.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  11. 11.Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. arXiv preprint arXiv:2203.06904, 2022.
  12. 12.Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, et al. A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level. Proceedings of the National Academy of Sciences, 119(32):e2123433119, 2022.
  13. 13.Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. Successive prompting for decomposing complex questions. arXiv preprint arXiv:2212.04092, 2022.
  14. 14.Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. Gpts are gpts: An early look at the labor market impact potential of large language models. arXiv preprint arXiv:2303.10130, 2023.
  15. 15.Jacqueline Fagard, Lauriane Rat-Fischer, Rana Esseily, Eszter Somogyi, and JK O’Regan. What does it take for an infant to learn how to use a tool by observation? Frontiers in psychology, 7: 267, 2016.
  16. 16.Bin Fu, Yunqi Qiu, Chengguang Tang, Yang Li, Haiyang Yu, and Jian Sun. A survey on complex question answering over knowledge base: Recent advances and challenges. arXiv preprint arXiv:2007.13069, 2020.
  17. 17.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. arXiv preprint arXiv:2211.10435, 2022.
  18. 18.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR, 2020.
  19. 19.Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  20. 20.Shibo Hao, Bowen Tan, Kaiwen Tang, Bin Ni, Xiyan Shao, Hengzhe Zhang, Eric Xing, and Zhiting Hu. BertNet: Harvesting knowledge graphs with arbitrary relations from pretrained language models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 5000–5015, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.309. URL https://aclanthology.org/2023.findings-acl.309.
  21. 21.Joy He-Yueya, Gabriel Poesia, Rose E Wang, and Noah D Goodman. Solving math word problems by combining language models with symbolic solvers. arXiv preprint arXiv:2304.09102, 2023.
  22. 22.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR, 2019.
  23. 23.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2021.
  24. 24.Zhiting Hu and Eric P Xing. Toward a ’standard model’ of machine learning. Harvard Data Science Review, 2022.
  25. 25.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pages 9118–9147. PMLR, 2022.
  26. 26.Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608, 2022.
  27. 27.Wenlong Huang, Fei Xia, Dhruv Shah, Danny Driess, Andy Zeng, Yao Lu, Pete Florence, Igor Mordatch, Sergey Levine, Karol Hausman, et al. Grounded decoding: Guiding text generation with grounded models for robot control. arXiv preprint arXiv:2303.00855, 2023.
  28. 28.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
  29. 29.Qiao Jin, Yifan Yang, Qingyu Chen, and Zhiyong Lu. Genegpt: Augmenting large language models with domain tools for improved access to biomedical information. ArXiv, 2023.
  30. 30.Jan-Christoph Kalo and Leandra Fichtel. Kamel: Knowledge analysis with multitoken entities in language models. In Proceedings of the Conference on Automated Knowledge Base Construction, 2022.
  31. 31.Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. Compacter: Efficient low-rank hypercomplex adapter layers. Advances in Neural Information Processing Systems, 34:1022–1035, 2021.
  32. 32.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406, 2022.
  33. 33.Yann LeCun. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Review, 62, 2022.
  34. 34.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, 2021.
  35. 35.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474, 2020.
  36. 36.Minghao Li, Feifan Song, Bowen Yu, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. Api-bank: A benchmark for tool-augmented llms. arXiv preprint arXiv:2304.08244, 2023.
  37. 37.Shuang Li, Xavier Puig, Chris Paxton, Yilun Du, Clinton Wang, Linxi Fan, Tao Chen, De-An Huang, Ekin Akyürek, Anima Anandkumar, et al. Pre-trained language models for interactive decision-making. Advances in Neural Information Processing Systems, 35:31199–31212, 2022.
  38. 38.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, 2021.
  39. 39.Yaobo Liang, Chenfei Wu, Ting Song, Wenshan Wu, Yan Xia, Yu Liu, Yang Ou, Shuai Lu, Lei Ji, Shaoguang Mao, et al. Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis. arXiv preprint arXiv:2303.16434, 2023.
  40. 40.Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone. Llm+ p: Empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477, 2023.
  41. 41.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35, 2023.
  42. 42.Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021.
  43. 43.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. Chameleon: Plug-and-play compositional reasoning with large language models. arXiv preprint arXiv:2304.09842, 2023.
  44. 44.Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. Faithful chain-of-thought reasoning. arXiv preprint arXiv:2301.13379, 2023.
  45. 45.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al. Augmented language models: a survey. arXiv preprint arXiv:2302.07842, 2023.
  46. 46.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  47. 47.OpenAI. Gpt-4 technical report, 2023.
  48. 48.Batu Ozturkler, Nikolay Malkin, Zhen Wang, and Nebojsa Jojic. Thinksum: Probabilistic reasoning over sets using large language models. arXiv preprint arXiv:2210.01293, 2022.
  49. 49.Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014, 2023.
  50. 50.Aaron Parisi, Yao Zhao, and Noah Fiedel. Talm: Tool augmented language models. arXiv preprint arXiv:2205.12255, 2022.
  51. 51.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  52. 52.Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. Virtualhome: Simulating household activities via programs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8494–8502, 2018.
  53. 53.Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shi Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bo Li, Ziwei Tang, Jing Yi, Yu Zhu, Zhenning Dai, Lan Yan, Xin Cong, Ya-Ting Lu, Weilin Zhao, Yuxiang Huang, Jun-Han Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Zhiyuan Liu, and Maosong Sun. Tool learning with foundation models. ArXiv, abs/2304.08354, 2023.
  54. 54.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
  55. 55.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, et al. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, 2021.
  56. 56.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023.
  57. 57.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580, 2023.
  58. 58.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. Retrieval augmentation reduces hallucination in conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3784–3803, 2021.
  59. 59.Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models. arXiv preprint arXiv:2209.11302, 2022.
  60. 60.Alon Talmor and Jonathan Berant. The web as a knowledge-base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 641–651, 2018.
  61. 61.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  62. 62.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  63. 63.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022.
  64. 64.Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, and Yoon Kim. Multitask prompt tuning enables parameter-efficient transfer learning. In The Eleventh International Conference on Learning Representations, 2023.
  65. 65.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  66. 66.Jiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, and Zhiting Hu. Language Models Meet World Models: Embodied Experiences Enhance Language Models. NeurIPS, 2023.
  67. 67.Yaqi Xie, Chen Yu, Tongyao Zhu, Jinbin Bai, Ze Gong, and Harold Soh. Translating natural language to planning goals with large-language models. arXiv preprint arXiv:2302.05128, 2023.
  68. 68.Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, Pieter Abbeel, and Dale Schuurmans. Foundation models for decision making: Problems, methods, and opportunities. arXiv preprint arXiv:2303.04129, 2023.
  69. 69.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.
  70. 70.Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. A survey of knowledge-enhanced text generation. ACM Computing Surveys, 54(11s):1–38, 2022.
  71. 71.Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–9, 2022.
  72. 72.Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. AlignScore: Evaluating factual consistency with a unified alignment function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11328–11348, Toronto, Canada, July 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.acl-long.634.
  73. 73.Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. Text alignment is an efficient unified model for massive nlp tasks. NeurIPS, 2023.

Citation

MLA
Hao, S., et al. “ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 45870–94, https://proceedings.neurips.cc/paper_files/paper/2023/file/8fd1a81c882cd45f64958da6284f4a3f-Paper-Conference.pdf.
APA
Hao, S., Liu, T., Wang, Z., & Hu, Z. (2023). ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings. Advances in Neural Information Processing Systems, 36, 45870–45894. https://proceedings.neurips.cc/paper_files/paper/2023/file/8fd1a81c882cd45f64958da6284f4a3f-Paper-Conference.pdf
Chicago
Hao, S., T. Liu, Z. Wang, and Z. Hu. 2023. “ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings”. Advances in Neural Information Processing Systems 36: 45870–94. https://proceedings.neurips.cc/paper_files/paper/2023/file/8fd1a81c882cd45f64958da6284f4a3f-Paper-Conference.pdf.
Harvard
Hao, S. et al. (2023) “ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 45870–45894. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/8fd1a81c882cd45f64958da6284f4a3f-Paper-Conference.pdf.
Vancouver
1. Hao S, Liu T, Wang Z, Hu Z (2023) ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 45870–45894

BibTeX

@inproceedings{hao2023toolkengpt,
  title = {ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings},
  author = {Hao, Shibo and Liu, Tianyang and Wang, Zhen and Hu, Zhiting},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {45870-45894},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/8fd1a81c882cd45f64958da6284f4a3f-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors