API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Minghao LiYingxiu ZhaoBowen YuFeifan SongHangyu LiHaiyang YuZhoujun LiFei HuangYongbin Li

article2023EMNLP389 citations

Presents API-Bank, a runnable benchmark and large-scale training corpus of over two thousand APIs that standardizes how large language models plan, retrieve, and execute external tools while identifying key failure modes across leading models.

Listen

Large Language Models often struggle with outdated knowledge and limited domain coverage because they rely solely on static pre-training data. Enabling these models to interact with external tools and application programming interfaces (APIs)—such as databases, calculators, and search engines—is critical to expanding their real-world capabilities. However, prior to this work, researchers lacked realistic, standardized frameworks to accurately measure tool proficiency, train models efficiently, and identify underlying failure modes across complex, multi-step tasks.

The article introduces and evaluates API-Bank, a comprehensive benchmark and training framework specifically designed to evaluate and improve the tool-use capabilities of language models. It establishes concrete standards across three progressive levels of tool utilization: direct API calling, API retrieval combined with calling, and multi-step planning combined with retrieval and execution.

To build the framework, the authors implemented an executable testbed of 73 real-world APIs and manually annotated 314 diverse multi-turn dialogues containing 753 API calls across eight domains. For model training, they developed a collaborative five-agent automated synthesis pipeline using conversational models to generate 1,888 dialogues covering 2,138 APIs and 1,000 distinct domains. This automated multi-agent approach reduced data creation costs by 98% compared to human annotation while maintaining a 94% usability rate. The authors evaluated major public models, including GPT-3, GPT-3.5, and GPT-4, and fine-tuned an open-source 7-billion-parameter baseline (Alpaca-7B) to create a specialized tool-augmented model named Lynx.

The evaluation revealed several critical findings. First, raw model scale does not guarantee tool-use competence; base GPT-3 Davinci achieved less than 1% accuracy across tasks, demonstrating that instruction tuning is mandatory for tool integration. Second, commercial models showed varied capabilities across task difficulty: GPT-3.5 reached 59.4% accuracy on direct calling but dropped significantly to 22.0% on multi-step planning and retrieval, whereas GPT-4 excelled in planning tasks with 70.0% accuracy. Third, fine-tuning the open-source Alpaca-7B model using API-Bank data yielded Lynx, which improved overall accuracy from 15.2% to 39.6% (a 24 percentage-point gain), closely approaching GPT-3.5's overall score of 47.2%. Finally, error analysis showed that while smaller models primarily struggle with API name hallucinations (accounting for 61.4% of errors in Lynx) and invalid input parameters (32% combined), advanced models like GPT-4 fail predominantly during API retrieval (representing 67.9% of its errors).

These findings demonstrate that automated multi-agent data generation is an effective and cost-efficient strategy for training specialized, open-source tool-using models without relying exclusively on expensive commercial systems. However, deployment in production environments still carries operational risks due to frequent parameter formatting errors and tool hallucinations, which can cause system exceptions or failed transactions. In addition, prompt-based in-context learning alone remains insufficient for dependable tool retrieval in complex workflows.

Organizations developing tool-augmented models should adopt specialized fine-tuning pipelines and enforce strict decoding constraints or parameter-validation layers to prevent execution errors. Future efforts must focus on improving semantic tool-retrieval mechanisms and scaling diverse synthetic training datasets to further reduce error rates.

These results are derived from a controlled English-language environment and focused fine-tuning on a 7-billion-parameter architecture. Readers should exercise caution when extrapolating these benchmarks directly to multilingual settings or mission-critical enterprise systems without additional real-time testing and guardrails.

  • Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Toolformer pioneered self-supervised methods for teaching language models to invoke external tools via APIs, establishing the foundational paradigm evaluated and expanded in API-Bank.
  • Paper: PAL: Program-aided Language Models, Luyu Gao et al. (2023). PAL introduced the concept of augmenting language model reasoning with programmatic execution, providing critical motivation for tool-augmented dialogue benchmarks.
  • Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). WebGPT established early methods for combining search APIs and external tool interaction with human feedback, laying early groundwork for evaluating tool-augmented LLMs.
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This seminal paper demonstrated few-shot prompting capabilities in large language models that serve as the baseline execution mechanism tested across API-Bank.

Abstract

Recent research has demonstrated that Large Language Models (LLMs) can enhance their capabilities by utilizing external tools. However, three pivotal questions remain unanswered: (1) How effective are current LLMs in utilizing tools? (2) How can we enhance LLMs' ability to utilize tools? (3) What obstacles need to be overcome to leverage tools? To address these questions, we introduce API-Bank, a groundbreaking benchmark, specifically designed for tool-augmented LLMs. For the first question, we develop a runnable evaluation system consisting of 73 API tools. We annotate 314 tool-use dialogues with 753 API calls to assess the existing LLMs' capabilities in planning, retrieving, and calling APIs. For the second question, we construct a comprehensive training set containing 1,888 tool-use dialogues from 2,138 APIs spanning 1,000 distinct domains. Using this dataset, we train Lynx, a tool-augmented LLM initialized from Alpaca. Experimental results demonstrate that GPT-3.5 exhibits improved tool utilization compared to GPT-3, while GPT-4 excels in planning. However, there is still significant potential for further improvement. Moreover, Lynx surpasses Alpaca's tool utilization performance by more than 26 pts and approaches the effectiveness of GPT-3.5. Through error analysis, we highlight the key challenges for future research in this field to answer the third question.

Table of Contents

  • 1 Introduction
  • 2 Design Principles of API-Bank
  • 2.1 Ability Grading
  • 2.2 Data Standards
  • 3 Evaluation System of API-Bank
  • 3.1 System Implementation
  • 3.2 Dialogue Annotation
  • 3.3 Evaluation Metrics
  • 4 Training Set of API-Bank
  • 5 Benchmark Analysis
  • 6 Related Work
  • 7 Experiments
  • 7.1 Baselines
  • 7.2 Main Results
  • 7.3 Error Analysis
  • 7.4 Dataset Comparison
  • 8 Conclusion
  • 9 Limitations
  • 10 Ethical Statement
  • References
  • A Appendix
  • A.1 Error Definitions
  • A.2 Implement Details
  • A.3 Examples

Knowls

  1. Knowl 1 — Three levels of API-use ability

    definition

    API-Bank grades tool use along two dimensions: whether the model knows a small API pool or must search a large one, and whether the user’s request needs one API call or a sequence of calls. Because planning multiple calls is relatively straightforward when all APIs are already supplied, the benchmark merges the two few-API conditions into one ability and defines three levels:

    • Call: select and invoke an API when its description is provided.
    • Retrieval+Call: find and invoke one suitable API when the available APIs are not given in advance.
    • Plan+Retrieval+Call: plan a sequence of API calls, retrieving an API for each needed step, to fulfill one complex request.

    The conceptual diagrams on page 3 illustrate these abilities as the progression from a known API pool to searching an unknown pool, and from a single call to a multi-call plan.

  2. Knowl 2 — Runnable evaluation system and API retrieval

    model/method

    API-Bank implements a runnable system with 73 API tools, including a dedicated API Search tool. The tools cover everyday functions and services such as weather forecasts and text-to-image generation. To make execution reproducible, database APIs use initialized databases, and APIs that access external information return recorded results for test queries rather than changing live results. Implementing the APIs took 98 person-days of engineering work.

    For Retrieval+Call and Plan+Retrieval+Call, the model is not given the API pool in advance. It must first condense the user’s need into search keywords. API Search embeds those keywords and the APIs’ metadata—including names, descriptions, and parameters—then uses cosine similarity to return the highest-matching API’s metadata. A search is required before each API call in these settings.

  3. Knowl 3 — Manual evaluation dialogues and scoring

    experimental setup

    The evaluation dialogues were manually constructed for all three API-use abilities. For Call, annotators sampled APIs, created queries that those APIs could fulfill, specified calls, executed them, and wrote responses based on the outputs; dialogues could include multiple turns using the same API set. For Retrieval+Call, annotators began with a complex requirement and decomposed it into simpler queries, each fulfilled by one API. For Plan+Retrieve+Call, they kept the complex requirement intact and annotated the sequence of API calls and the response based on the final execution.

    Two annotators discussed each dialogue, and two additional annotators checked formatting, logical consistency, and whether calls were reasonable. Of 400 annotated dialogues, 314 were retained after discarding 21.5% for annotation problems; the retained set contains 753 API calls.

    API-call accuracy is the fraction of predictions judged correct. Correctness is based on whether the predicted calls perform the same database queries or modifications and return the same results as the annotated calls. Responses after API execution are evaluated with ROUGE-L.

  4. Knowl 4 — API-Bank corpus composition and benchmark coverage

    data/table

    API-Bank combines automatically generated training data with manually annotated evaluation data. The training split contains 1,000 domains, 2,138 APIs, 1,888 dialogues, and 5,221 turns; the evaluation split contains 8 domains, 73 APIs, 314 dialogues, and 914 turns. Together, the paper reports 1,008 domains, 2,211 APIs, 2,202 dialogues, and 6,135 turns. The training and evaluation sets average 2.76 and 2.91 turns per dialogue, respectively.

    The reported ability counts are 720 Call, 719 Retrieval+Call, and 449 Plan+Retrieve+Call dialogues in training, versus 214, 50, and 50 in evaluation. The paper also reports 3,147 single-call and 493 multiple-call instances in training, and 363 single-call and 122 multiple-call instances in evaluation.

    Compared with the six other benchmarks in the paper’s comparison, API-Bank covers the largest number of domains among those listed and is the only one marked as covering all of multi-turn dialogues, multi-call dialogues, API-call and response evaluation, and the three API-use abilities. ToolBench1 has more APIs (16,464 versus API-Bank’s reported 2,138 training APIs), so API-Bank’s distinction is breadth of combined coverage rather than the largest API count.

  5. Knowl 5 — Five-agent generation of tool-use training data

    algorithm

    API-Bank’s Multi-agent method generates training examples through five prompt-driven ChatGPT agents, with each stage conditioned on the preceding stages:

    1. Generate several domains.
    2. Given the domains, propose plausible APIs. To encourage API authenticity, provide examples from public API collections as part of the agent’s input.
    3. Select one or more proposed APIs and one of the benchmark’s three abilities; generate a user query that matches that ability and can be fulfilled by the selected APIs.
    4. Given the domain, APIs, ability, and query, generate the required API calls, simulate their execution, and produce an answer to the query.
    5. Test whether the resulting example follows the benchmark’s data standards and discard examples that fail.

    The page-5 workflow diagram depicts the four data-generation agents producing domains, APIs, queries, calls, and responses, followed by a separate tester agent. The method decomposes a complex generation instruction into smaller dependent tasks rather than asking one agent to satisfy all requirements at once.

  6. Knowl 6 — Generated-data quality and annotation cost

    empirical result

    In a human quality review of 100 randomly sampled Multi-agent training examples, the reported usable-data rate was 94%. The paper compares this with 5% availability when a single ChatGPT agent was given the complex generation instruction, and reports an 89% improvement. In a separate test of that single-agent instruction, GPT-4 raised availability to 25%, but many errors remained. The fifth, tester agent rejects 35% of generated instances; among automatically filtered instances examined by the authors, 78% failed to follow the intended data principles.

    Multi-agent generation costs about 0.1perdialogue,comparedwith0.1 per dialogue, compared with 8 per manually annotated dialogue, which the paper reports as a 98% cost reduction.

  7. Knowl 7 — Model performance on the API-Bank evaluation set

    empirical result

    The authors evaluated zero-shot Alpaca-7B, ChatGLM-6B, GPT-3 Davinci, GPT-3.5-turbo-0613, and GPT-4-0613, alongside Lynx-7B fine-tuned from Alpaca-7B initialization on API-Bank. Lynx was fine-tuned for three epochs with batch size 256 and learning rate 2×10−52\times10^{-5}. The table reports API-call correctness accuracy and response ROUGE-L for each ability and overall.

    The results show that performance generally falls as the task requires retrieval and planning in addition to calling. GPT-4 performs especially well on Plan+Retrieval+Call, with 70.00% correctness, compared with 22.00% for GPT-3.5 and 20.00% for Lynx. Lynx improves over Alpaca in Call correctness by 25.81 percentage points and reaches 39.58% overall correctness, between GPT-3.5 at 47.16% and GPT-4 at 60.24%.

    Model Call Acc. Call ROUGE-L Ret.+Call Acc. Ret.+Call ROUGE-L Plan+Ret.+Call Acc. Plan+Ret.+Call ROUGE-L Total Acc. Total ROUGE-L
    Alpaca-7B 24.06% 0.0204 5.19% 0.0019 0.00% 0.086 15.19% 0.0318
    ChatGLM-6B 23.62% 0.2451 13.33% 0.2173 0.00% 0.1522 16.42% 0.2191
    GPT-3 Davinci 0.50% 0.1035 1.48% 0.091 0.00% 0.0156 0.57% 0.0814
    GPT-3.5-turbo 59.40% 0.4598 38.52% 0.3758 22.00% 0.3809 47.16% 0.4267
    GPT-4 63.66% 0.3691 37.04% 0.351 70.00% 0.4808 60.24% 0.3910
    Lynx-7B 49.87% 0.4332 30.37% 0.2503 20.00% 0.3425 39.58% 0.3794
  8. Knowl 8 — Distinct failure patterns in Alpaca, Lynx, and GPT-4

    empirical result

    The paper classifies evaluation errors and reports each type as a share of that model’s errors. Alpaca most often made no API call; after fine-tuning, Lynx’s dominant problem was API-name mismatch or hallucination. GPT-4’s main difficulty was failed API retrieval. The authors also identify problematic parameters and unparsable call formatting as recurring issues.

    Error type Alpaca-7B Lynx-7B GPT-4
    No API Call 36.77% 5.29% –
    API Hallucination 15.93% 61.38% –
    Invalid Input Parameters 7.96% 8.47% 7.14%
    False API Call Format 23.65% 6.88% 17.86%
    Missing Input Parameters 1.17% 1.59% 7.14%
    Has Exception – 16.40% –
    Failed API Retrieval – – 67.86%

    A dash means the error type was not listed for that model in the paper’s distribution, not that its rate was established to be zero. The authors attribute Lynx’s hallucinations partly to generating APIs encountered during training even when they were not supplied or relevant. They also report that parameter-related failures—including exceptions, invalid parameters, and missing parameters—together account for about 32% of Lynx’s errors. GPT-4 sometimes emitted simultaneous calls that the evaluation system could not parse.

  9. Knowl 9 — Lynx compared with a ToolAlpaca-trained baseline

    empirical result

    To compare training data, the authors converted ToolAlpaca’s training set into API-Bank’s format, producing 10,366 training samples, and fine-tuned Alpaca on it. Lynx was trained on 6,184 samples. Since ToolAlpaca does not train API retrieval, this comparison evaluates only API calling. Despite using fewer samples, Lynx scored slightly higher than the ToolAlpaca-trained model on both reported metrics:

    Model Training samples Accuracy (Call) ROUGE (Call)
    Alpaca fine-tuned on ToolAlpaca 10,366 53.88 39.75
    Lynx 6,184 54.64 39.80

    The authors interpret this result as evidence that their generated training data is effective and of higher quality for the tested API-calling task.

  10. Knowl 10 — Stated limitations of API-Bank and Lynx

    limitation

    API-Bank covers English only; the authors leave construction and evaluation in other languages to future work. Their reported fine-tuning experiments use Lynx-7B, so they do not establish how the training approach scales to larger models. They also state that a commercially viable tool-augmented model based on a larger in-house LLM was trained, but its results could not be reported or analyzed because of anonymity constraints.

Coverage note — The paper’s example dialogues and ethical statement are omitted because they illustrate the benchmark or describe research practice rather than adding a distinct result or method beyond the extracted knowls.

References

  1. 1.Stanley H Ambrose. 2001. Paleolithic technology and human evolution. Science, 291(5509):1748–1753.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  4. 4.Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. 2023. Large language models as tool makers. arXiv preprint arXiv:2305.17126.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  6. 6.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  7. 7.Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings. arXiv preprint arXiv:2305.11554.
  8. 8.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  9. 9.Yaobo Liang, Chenfei Wu, Ting Song, Wenshan Wu, Yan Xia, Yu Liu, Yang Ou, Shuai Lu, Lei Ji, Shaoguang Mao, et al. 2023. Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis. arXiv preprint arXiv:2303.16434.
  10. 10.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al. 2023. Augmented language models: a survey. arXiv preprint arXiv:2302.07842.
  11. 11.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  12. 12.Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. 2023. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014.
  13. 13.Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E Gonzalez. 2023. Gorilla: Large language model connected with massive apis. arXiv preprint arXiv:2305.15334.
  14. 14.Cheng Qian, Chi Han, Yi R Fung, Yujia Qin, Zhiyuan Liu, and Heng Ji. 2023. Creator: Disentangling abstract and concrete reasonings of large language models through tool creation. arXiv preprint arXiv:2305.14318.
  15. 15.Shuofei Qiao, Honghao Gui, Huajun Chen, and Ningyu Zhang. 2023. Making language models better tool learners with execution feedback. arXiv preprint arXiv:2305.13068.
  16. 16.Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, et al. 2023a. Tool learning with foundation models. arXiv preprint arXiv:2304.08354.
  17. 17.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. 2023b. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789.
  18. 18.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  19. 19.Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023. Preference ranking optimization for human alignment. arXiv preprint arXiv:2306.17492.
  20. 20.Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, and Le Sun. 2023. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases. arXiv preprint arXiv:2306.05301.
  21. 21.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  22. 22.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  23. 23.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  24. 24.Qiantong Xu, Fenglu Hong, Bo Li, Changran Hu, Zhengyu Chen, and Jian Zhang. 2023. On the tool manipulation capability of open-source large language models.
  25. 25.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  26. 26.Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al. 2022. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598.
  27. 27.Yingxiu Zhao, Bowen Yu, Binyuan Hui, Haiyang Yu, Fei Huang, Yongbin Li, and Nevin L Zhang. 2023. A preliminary study of the intrinsic relationship between complexity and alignment. arXiv preprint arXiv:2308.05696.
  28. 28.Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2023. Toolqa: A dataset for llm question answering with external tools. arXiv preprint arXiv:2306.13304.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/