API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Minghao LiYingxiu ZhaoBowen YuFeifan SongHangyu LiHaiyang YuZhoujun LiFei HuangYongbin Li
Presents API-Bank, a runnable benchmark and large-scale training corpus of over two thousand APIs that standardizes how large language models plan, retrieve, and execute external tools while identifying key failure modes across leading models.
Large Language Models often struggle with outdated knowledge and limited domain coverage because they rely solely on static pre-training data. Enabling these models to interact with external tools and application programming interfaces (APIs)—such as databases, calculators, and search engines—is critical to expanding their real-world capabilities. However, prior to this work, researchers lacked realistic, standardized frameworks to accurately measure tool proficiency, train models efficiently, and identify underlying failure modes across complex, multi-step tasks.
The article introduces and evaluates API-Bank, a comprehensive benchmark and training framework specifically designed to evaluate and improve the tool-use capabilities of language models. It establishes concrete standards across three progressive levels of tool utilization: direct API calling, API retrieval combined with calling, and multi-step planning combined with retrieval and execution.
To build the framework, the authors implemented an executable testbed of 73 real-world APIs and manually annotated 314 diverse multi-turn dialogues containing 753 API calls across eight domains. For model training, they developed a collaborative five-agent automated synthesis pipeline using conversational models to generate 1,888 dialogues covering 2,138 APIs and 1,000 distinct domains. This automated multi-agent approach reduced data creation costs by 98% compared to human annotation while maintaining a 94% usability rate. The authors evaluated major public models, including GPT-3, GPT-3.5, and GPT-4, and fine-tuned an open-source 7-billion-parameter baseline (Alpaca-7B) to create a specialized tool-augmented model named Lynx.
The evaluation revealed several critical findings. First, raw model scale does not guarantee tool-use competence; base GPT-3 Davinci achieved less than 1% accuracy across tasks, demonstrating that instruction tuning is mandatory for tool integration. Second, commercial models showed varied capabilities across task difficulty: GPT-3.5 reached 59.4% accuracy on direct calling but dropped significantly to 22.0% on multi-step planning and retrieval, whereas GPT-4 excelled in planning tasks with 70.0% accuracy. Third, fine-tuning the open-source Alpaca-7B model using API-Bank data yielded Lynx, which improved overall accuracy from 15.2% to 39.6% (a 24 percentage-point gain), closely approaching GPT-3.5's overall score of 47.2%. Finally, error analysis showed that while smaller models primarily struggle with API name hallucinations (accounting for 61.4% of errors in Lynx) and invalid input parameters (32% combined), advanced models like GPT-4 fail predominantly during API retrieval (representing 67.9% of its errors).
These findings demonstrate that automated multi-agent data generation is an effective and cost-efficient strategy for training specialized, open-source tool-using models without relying exclusively on expensive commercial systems. However, deployment in production environments still carries operational risks due to frequent parameter formatting errors and tool hallucinations, which can cause system exceptions or failed transactions. In addition, prompt-based in-context learning alone remains insufficient for dependable tool retrieval in complex workflows.
Organizations developing tool-augmented models should adopt specialized fine-tuning pipelines and enforce strict decoding constraints or parameter-validation layers to prevent execution errors. Future efforts must focus on improving semantic tool-retrieval mechanisms and scaling diverse synthetic training datasets to further reduce error rates.
These results are derived from a controlled English-language environment and focused fine-tuning on a 7-billion-parameter architecture. Readers should exercise caution when extrapolating these benchmarks directly to multilingual settings or mission-critical enterprise systems without additional real-time testing and guardrails.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Toolformer pioneered self-supervised methods for teaching language models to invoke external tools via APIs, establishing the foundational paradigm evaluated and expanded in API-Bank.
- Paper: PAL: Program-aided Language Models, Luyu Gao et al. (2023). PAL introduced the concept of augmenting language model reasoning with programmatic execution, providing critical motivation for tool-augmented dialogue benchmarks.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). WebGPT established early methods for combining search APIs and external tool interaction with human feedback, laying early groundwork for evaluating tool-augmented LLMs.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This seminal paper demonstrated few-shot prompting capabilities in large language models that serve as the baseline execution mechanism tested across API-Bank.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). ToolLLM scales tool augmentation far beyond API-Bank by creating a benchmark and retrieval framework covering more than 16,000 real-world RESTful APIs.
- Paper: Gorilla: Large Language Model Connected with Massive APIs, Shishir G. Patil et al. (2023). Gorilla extends API-focused fine-tuning and evaluation by integrating dynamic documentation retrieval to handle massive, evolving API libraries.
- Paper: ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings, Shibo Hao et al. (2023). ToolkenGPT introduces an alternative embedding-based approach to tool use that overcomes the context-window and fine-tuning trade-offs highlighted in API benchmarks.
- Paper: EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction, Siyu Yuan et al. (2025). EASYTOOL builds upon tool-use benchmarks to optimize and standardize tool documentation for improved LLM agent performance and reduced context consumption.
- Paper: SciAgent: Tool-augmented Language Models for Scientific Reasoning, Yubo Ma et al. (2024). SciAgent applies tool-augmented planning and execution specifically to complex multi-step scientific and mathematical reasoning tasks.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). Language Agent Tree Search advances the linear planning and API execution studied in API-Bank by integrating deliberate tree search and self-reflection.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena expands API and tool evaluation into a fully interactive, realistic web environment for evaluating multi-step autonomous agent actions.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). BigCodeBench provides a specialized evaluation suite focusing on diverse function calling and complex software library utilization in executable code generation.