EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Siyu YuanKaitao SongJiangjie ChenXu TanYongliang ShenKan RenDongsheng LiDeqing Yang

article2025NAACL201 citations

Proposes EASYTOOL, a framework that converts lengthy, inconsistent tool documentation into standardized and concise instructions, substantially lowering prompt token costs while boosting tool-use accuracy across diverse agent tasks.

Listen

Autonomous agents powered by large language models increasingly rely on external software tools to complete complex tasks, such as accessing live web data and performing specialized computations. However, existing tool documentation across diverse sources suffers from severe formatting inconsistencies, high verbosity, and a lack of clear usage examples. These defects cause language models to exceed context limits, misinterpret tool purposes, and supply invalid parameters, resulting in frequent task failures.

The article introduces and evaluates EASYTOOL, a standardized framework that condenses raw, lengthy documentation into concise, structured tool instructions. The objective is to leverage the instruction-following capabilities of language models to improve tool retrieval, tool selection, and parameter accuracy while reducing computational overhead.

The authors tested EASYTOOL across three distinct benchmarks: ToolBench for real-world question answering across diverse web services, RestBench for multi-step service planning, and FuncQA for multi-step mathematical reasoning. The evaluation compared commercial and open-source models—including ChatGPT, GPT-4, GPT-4o, and Llama 3.1 variants—using raw documentation, standard prompt baselines, and EASYTOOL instructions. Independent human evaluators validated instruction quality and analyzed failure modes.

The investigation produced four primary findings. First, EASYTOOL drastically reduces token consumption, shrinking prompt sizes by 70.43% in ToolBench and 97.35% in RestBench. Second, standardized instructions significantly improve end-to-end task performance; on ToolBench, ChatGPT's success rate rose from 15.0% to 52.8%, and GPT-4o reached a 77.0% success rate. Third, tool selection and execution errors plummeted, virtually eliminating tool name errors and reducing parameter errors from 25% to 6% in ChatGPT and from 17% to 1% in GPT-4. Fourth, smaller open-source models benefited substantially; Llama-3.1-8B-Instruct improved its success rate from 5.5% to 48.5%, surpassing larger, specialized fine-tuned models.

These findings demonstrate that restructuring documentation into concise functional summaries and concrete parameter examples is far more effective than feeding raw documentation or using generic prompt compression techniques, which often destroy critical parameter syntax. For organizations building agentic systems, adopting standardized instruction layers substantially reduces token-based operating costs, increases system reliability, and allows teams to deploy smaller, lower-cost open-source models without sacrificing task accuracy.

Organizations developing tool-augmented language model workflows should implement structured pre-processing pipelines to convert complex API documentation into unified instructions with explicit scenario examples. Before broad deployment, development teams should run pilot evaluations on their specific API libraries, ensuring that tool descriptions capture multi-functional capabilities and exact parameter constraints.

The findings are supported by high inter-annotator agreement and consistent performance gains across diverse task types. Nevertheless, decision-makers should note key limitations: the framework requires source documentation to fit within the pre-processing model's context window, does not explicitly model inter-tool dependencies, and relies on base models possessing baseline instruction-following capabilities.

arXiv: 2401.06201main/easytool
Cover for EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Abstract

There has been a rising interest in utilizing tools in applications of autonomous agents based on large language models (LLMs) to address intricate real-world tasks. To develop LLM-based agents, it usually requires LLMs to understand many tool functions from different tool documentations. However, these documentations could be diverse, redundant, or incomplete, which immensely affects the capability of LLMs in using tools. Current LLMs exhibit satisfactory instruction-following capabilities based on instruction-following fine-tuning process. Motivated by this, in this paper, we introduce EASYTOOL, a framework transforming diverse and lengthy tool documentation into a unified and concise tool instruction to fully leverage instruction-following capabilities of LLMs for easier tool usage. EASYTOOL purifies essential information from extensive tool documentation of different sources, and elaborates a unified interface (i.e., tool instruction) to offer standardized tool descriptions and functionalities for LLM-based agents. Extensive experiments on multiple different tasks demonstrate that EASYTOOL can significantly reduce token consumption and improve the performance of LLM-based agents on tool utilization in real-world scenarios. Our code is available in supplemental materials. Our code is available at https://github.com/microsoft/JARVIS/tree/main/easytool.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 3.1 Task Formulation
  • 3.2 Analysis
  • 4 Method
  • 4.1 Tool Description Generation
  • 4.2 Tool Functionality Guidelines Construction
  • I: Tool Description Generation
  • II: Tool Function Guidelines Construction
  • 4.3 Evaluation
  • 5 Experiment
  • 5.1 Real-World Question Answering
  • 5.2 Real-World Web Services
  • 5.3 Numerical Reasoning
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgement
  • References
  • A Crowd-sourcing Details for Tool Instruction Evaluation and Error Analysis
  • B Details of ToolBench
  • B.1 Success Rate Evaluation
  • B.2 The Details about Baselines of ToolBench
  • B.3 Error Analysis about Example Number
  • C Tool Instruction Generation
  • D Robustness Evaluation
  • E Hallucination Evaluation
  • F Prompt Compression Method
  • G Examples of Tool Instruction
  • G.1 Data Examples of ToolBench
  • G.2 Data Examples of RestBench
  • G.3 Data Examples of FuncQA

Knowls

  1. Knowl 1 — EASYTOOL Framework for Concise Tool Instruction Generation

    model/method

    EASYTOOL is a framework that converts verbose, unstructured, and heterogeneous tool documentation into standardized, concise, and executable tool instructions for large language model (LLM) agents. The framework operates in two sequential stages:

    1. Tool Description Generation: An LLM (e.g., ChatGPT) parses the raw tool documentation to extract high-level tool identity and summarize its core functionality while stripping out extraneous metadata (such as internal URLs, product IDs, pricing tiers, and host endpoints). For multi-endpoint tools, it creates an enumerated summary listing each built-in function name and a brief statement of its specific purpose.

    2. Tool Functionality Guidelines Construction: The framework extracts parameter schemas (names, types, descriptions, and required/optional status) and uses an LLM to synthesize concrete execution examples formatted as JSON objects containing:

      • Scenario: A natural-language description specifying under what user condition or intent the function should be called.
      • Parameters: A structured key-value mapping demonstrating valid arguments conforming to the extracted schema.

    To ensure semantic fidelity and parameter validity, candidate parameter configurations generated during this step are executed against the actual underlying API tools before deployment.

  2. Knowl 2 — Deficiencies of Raw Tool Documentation for LLM Agents

    definition

    Raw tool documentations gathered across diverse web services, APIs, and model hubs suffer from three primary structural deficiencies that impede LLM-based agent performance:

    • Inconsistency: Tools originating from heterogeneous providers lack standardized formatting, schema conventions, and descriptions, forcing LLMs to parse erratic syntax across multiple domains.
    • Redundancy: Documentation frequently embeds extraneous technical metadata (e.g., hostnames, product IDs, pricing tiers, tracking URLs) that inflates prompt token consumption and distracts the LLM from identifying core functionalities.
    • Incompleteness: Real-world documentation often lacks contextual usage guidance and concrete invocation scenarios, providing only parameter types or isolated code snippets without explaining when and how specific arguments should be populated.
  3. Knowl 3 — Token Consumption Reduction via EASYTOOL

    data/table

    Converting raw documentation into EASYTOOL instructions dramatically compresses prompt length while retaining operational semantics across standard benchmarks.

    Dataset TokenDoc. TokenIns. Reduce (%)
    ToolBench 2,530 748 70.43%
    RestBench 3,881 103 97.35%

    Token counts are measured using OpenAI cl100k_base BPE encoding. TokenDoc. denotes the average token count of the raw documentation per tool, TokenIns. denotes the average token count after EASYTOOL restructuring, and Reduce denotes the percentage token savings. On ToolBench, EASYTOOL achieves a 70.43% reduction, and on RestBench, it achieves a 97.35% reduction.

  4. Knowl 4 — Agent Execution Performance on Multi-Tool ToolBench Tasks

    data/table

    Standardized instructions from EASYTOOL consistently improve the Pass Rate (proportion of requests completed within budget), Win Rate (evaluator preference relative to ChatGPT-ReACT baseline), and Success Rate (GPT-4 evaluation of answer correctness) on ToolBench subsets requiring multi-category (I2-Category) and multi-instruction (I3-Instruction) tool compositions.

    I2-Category I3-Instruction Average
    Model Method Pass Win Succ Pass Win Succ Pass Win Succ
    ChatGPT ReACT 39.0 - 18.0 23.0 - 1.0 31.0 - 9.5
    DFSDT 64.5 63.0 24.0 60.0 70.0 6.0 62.3 66.5 15.0
    +EASYTOOL 74.5 76.5 68.5 65.0 88.0 37.0 69.8 82.3 52.8
    +EASYTOOL +Re. 69.0 71.0 60.5 66.0 89.0 42.0 67.5 80.0 51.3
    ToolLLaMA-7B ReACT 30.0 45.5 9.5 22.0 49.0 3.0 26.0 47.3 6.3
    DFSDT 66.0 55.0 24.0 56.0 56.0 6.0 61.0 55.5 15.0
    +Re. 57.0 60.0 11.5 54.0 69.0 2.0 55.5 64.5 6.8
    Llama-3.1-8B-Ins ReACT 3.0 0.0 0.0 0.0 0.0 0.0 1.5 0.0 0.0
    DFSDT 12.0 32.0 10.0 8.0 3.0 1.0 10.0 17.5 5.5
    +EASYTOOL 75.0 75.0 60.0 69.0 88.0 37.0 72.0 81.5 48.5
    +EASYTOOL +Re. 69.0 69.0 57.5 68.0 89.0 42.0 68.5 79.0 49.8
    Llama-3.1-70B-Ins ReACT 24.0 10.0 11.0 20.0 13.0 10.0 22.0 11.5 10.5
    DFSDT 45.0 40.0 27.0 35.0 56.0 21.0 40.0 48.0 24.0
    +EASYTOOL 71.5 79.0 70.0 70.0 89.0 60.0 70.8 84.0 65.0
    +EASYTOOL +Re. 70.0 71.0 65.0 71.0 89.0 54.0 70.5 80.0 59.5
    GPT-4 ReACT 67.5 53.5 27.0 40.0 71.0 4.0 53.8 62.3 15.5
    DFSDT 69.5 57.0 42.0 59.0 73.0 50.0 64.3 65.0 46.0
    +EASYTOOL 76.5 78.5 76.0 69.0 89.0 64.0 72.8 83.8 70.0
    +EASYTOOL +Re. 72.5 72.0 73.5 69.0 90.0 53.0 70.8 81.0 63.3
    GPT-4o ReACT 66.0 56.5 30.0 42.0 71.0 5.0 54.0 63.8 17.5
    DFSDT 72.5 63.5 54.5 60.0 80.0 63.0 66.3 71.8 63.8
    +EASYTOOL 80.5 81.5 83.0 73.0 90.0 71.0 76.8 85.8 77.0
    +EASYTOOL +Re. 76.5 79.0 81.0 69.0 90.0 67.0 72.8 84.5 74.0

    All non-+Re. variants select candidate tools from the ground-truth toolset, whereas +Re. combines EASYTOOL with a dense retriever. With EASYTOOL, general open-weight models (e.g., Llama-3.1-8B-Instruct) surpass dedicated tool-fine-tuned models (ToolLLaMA-7B), boosting average Success Rate from 5.5% to 48.5%.

  5. Knowl 5 — Tool Retrieval Enhancement via EASYTOOL

    data/table

    Replacing raw tool documentation with concise tool descriptions generated by EASYTOOL substantially boosts the retrieval quality of embedding-based retrievers compared to raw documentation embeddings and specialized dense retrievers.

    I2-Category I3-Instruction Average
    Method @1 @5 @1 @5 @1 @5
    BERT Retriever 68.2 77.9 81.7 87.1 75.0 82.5
    Ada (text-embedding-ada-002) 36.8 30.7 54.6 46.8 45.7 38.8
    Ada + EASYTOOL 73.4 82.7 80.1 88.5 76.7 85.6

    Performance is evaluated using Normalized Discounted Cumulative Gain (NDCG@1 and NDCG@5). Using raw documentation with Ada yields an average NDCG@1 of 45.7% due to distracting metadata; replacing the documentation with EASYTOOL concise descriptions increases Ada's average NDCG@1 to 76.7% and NDCG@5 to 85.6%, outperforming the supervised BERT-base retriever.

  6. Knowl 6 — Tool Invocation Error Reduction and Example Number Ablation

    data/table

    Manual examination by three annotators (Fleiss's κ=0.91\kappa = 0.91) on 100 sampled tasks from ToolBench categorizes tool invocation failures into two classes:

    • Name Error: The model attempts to invoke a non-existent tool or endpoint name.
    • Parameter Error: The model calls a valid tool but supplies invalid, malformed, or missing parameters.
    Model Method Name Error (%) Parameter Error (%)
    ChatGPT Tool documentation 8 25
    EASYTOOL w/o example 0 21
    EASYTOOL w/ 1 example 0 6
    EASYTOOL w/ 3 examples 0 5
    GPT-4 Tool documentation 5 17
    EASYTOOL w/o example 0 14
    EASYTOOL w/ 1 example 0 1
    EASYTOOL w/ 3 examples 0 1

    The concise functional description in EASYTOOL completely eliminates Name Errors (0%0\%) across both ChatGPT and GPT-4. Supplying synthesized execution examples in the functionality guidelines resolves Parameter Errors, reducing ChatGPT parameter error rate from 25%25\% to 6%6\% with 1 example and 5%5\% with 3 examples, and GPT-4 parameter error rate from 17%17\% to 1%1\%.

  7. Knowl 7 — Web Service Planning Accuracy on RestBench

    empirical result

    On the TMDB subset of RestBench (comprising 55 RESTful API endpoints for movie database interactions), tool execution requires determining multi-step sequential calling paths. Accuracy is measured via Correct Path Rate (CP%), defined as the proportion of generated tool execution trajectories that contain the ground-truth sequence as a valid subsequence.

    • For Vicuna-13B-based RestGPT: Raw RestGPT achieves approximately 15%15\% CP%, ToolDec (constrained decoding) achieves ≈19%\approx 19\% CP%, whereas integrating EASYTOOL instructions improves CP% to ≈25%\approx 25\%.
    • For ChatGPT-based RestGPT: ReAct achieves ≈58%\approx 58\% CP%, RestGPT achieves ≈64%\approx 64\% CP%, whereas RestGPT with EASYTOOL instructions reaches ≈78%\approx 78\% CP%.

    This confirms that concise tool instructions assist LLMs in identifying correct multi-step API invocation paths without requiring specialized decoding-time constraints.

  8. Knowl 8 — Numerical Reasoning Enhancement on Incomplete Tool Documentation (FuncQA)

    data/table

    FuncQA evaluates LLM numerical reasoning across 13 arithmetic tools where original documentation is incomplete, containing only tool names and calling signatures without descriptions or examples. EASYTOOL uses the tool names and signatures to generate synthetic functionality descriptions and usage scenarios.

    Model / Method One-hop (%) ↑\uparrow Multi-hop (%) ↑\uparrow Error (%) ↓\downarrow
    Vicuna-30B (0-shot) 15.00 1.00 -
    + CoT 13.33 4.00 -
    + ReAct 45.00 7.35 20.31
    + EASYTOOL 65.00 11.76 10.15
    ChatGPT (0-shot) 55.00 9.00 -
    + CoT 48.33 17.64 -
    + ReAct 85.00 41.17 9.38
    + EASYTOOL 91.66 48.53 2.34

    One-hop includes 68 single-tool questions; Multi-hop includes 60 questions averaging 2.78 tool invocations. Error measures the percentage of tasks encountering at least one tool-related execution error. EASYTOOL increases ChatGPT one-hop accuracy from 85.00% (ReAct) to 91.66% and multi-hop accuracy from 41.17% to 48.53%, while reducing tool error rate from 9.38% to 2.34%.

  9. Knowl 9 — Hallucination Mitigation in Tool Instruction and Usage

    data/table

    LLM hallucinations in tool use are evaluated across three categories:

    • Input-conflicting: Outputs deviating from user input instructions.
    • Context-conflicting: Outputs contradicting the model's own prior generation.
    • Fact-conflicting: Outputs violating established world facts.

    During the generation of EASYTOOL instructions itself (evaluated on 500 ToolBench samples), ChatGPT hallucinations are rare: Human Eval reports 1.2%1.2\% input-conflicting, 0.8%0.8\% context-conflicting, and 1.0%1.0\% fact-conflicting errors; LLM Judge (GPT-4o) reports 0.6%0.6\%, 0.8%0.8\%, and 0.8%0.8\%, respectively.

    When GPT-4 subsequently uses tools to solve 200 ToolBench user requests, replacing raw tool documentation (w/ Doc.) with EASYTOOL instructions (w/ Inst.) substantially lowers downstream hallucinations:

    Method Metrics Input (%) Context (%) Fact (%)
    GPT-4 w/ Doc. Human Eval 17.0 6.0 5.0
    LLM Judge 14.0 8.0 3.8
    GPT-4 w/ Inst. Human Eval 4.2 1.6 0.0
    LLM Judge 3.5 0.0 0.0
  10. Knowl 10 — Failure of Generic Token-Pruning Prompt Compression for Tool Use

    empirical result

    Applying generic prompt compression algorithms such as LLMLingua—which drop tokens based on perplexity budgets without structural awareness—proves detrimental to LLM tool usage. Perplexity-based token pruning drops essential syntactic and semantic components from tool documentation, including required parameter keys, constraints, and method signatures (e.g., mangling parameter names and schema definitions into fragmented substrings). In contrast, EASYTOOL performs structured semantic restructuring, preserving parameter integrity and operational schema while removing extraneous narrative text.

  11. Knowl 11 — Limitations of EASYTOOL

    limitation

    The EASYTOOL framework has three principal limitations:

    1. Context Window Constraint on Documentation Processing: EASYTOOL relies on a teacher LLM (e.g., ChatGPT) to process the raw tool documentation. Consequently, individual tool documentation that exceeds the model's context window cannot be transformed without prior chunking or preprocessing.
    2. Absence of Inter-Tool Dependency Modeling: Tool instructions are created per tool in isolation, omitting relational dependencies, compositional protocols, and execution ordering constraints across different tools in an inventory.
    3. Dependence on Instruction-Following Capabilities: EASYTOOL relies on the base LLM's intrinsic instruction-following ability, making it ineffective on base models lacking instruction tuning.

Coverage note — Omitted qualitative UI screenshot figures (Figures 7-9) and extended raw JSON prompt/output listings from the appendices, as their substantive contributions are fully captured in the methodology, error analysis, and quantitative results knowls.

References

  1. 1.Zehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang, Dahua Lin, Kai Chen, et al. 2023. T-eval: Evaluating the tool utilization capability step by step. arXiv preprint arXiv:2312.14033.
  2. 2.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  3. 3.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  4. 4.Yu Du, Fangyun Wei, and Hongyang Zhang. 2024. Anytool: Self-reflective, hierarchical agents for large-scale api calls. arXiv preprint arXiv:2402.04253.
  5. 5.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  6. 6.Significant Gravitas. 2023. Auto-gpt: An autonomous gpt-4 experiment. https://github.com/Significant-Gravitas/Auto-GPT.
  7. 7.Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings. volume abs/2305.11554.
  8. 8.Cheng-Yu Hsieh, Si-An Chen, Chun-Liang Li, Yasuhisa Fujii, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. 2023. Tool documentation enables zero-shot tool-usage with large language models.
  9. 9.Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of ir techniques. ACM Trans. Inf. Syst., 20(4):422–446.
  10. 10.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023a. Mistral 7b.
  11. 11.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023b. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13358–13376, Singapore. Association for Computational Linguistics.
  12. 12.Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. Api-bank: A comprehensive benchmark for tool-augmented llms. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 3102–3116.
  13. 13.Xukun Liu, Zhiyuan Peng, Xiaoyuan Yi, Xing Xie, Lirong Xiang, Yuchen Liu, and Dongkuan Xu. 2024. Toolnet: Connecting large language models with massive tools via tool graph. arXiv preprint arXiv:2403.00839.
  14. 14.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. volume abs/2304.09842.
  15. 15.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076–12100, Singapore. Association for Computational Linguistics.
  16. 16.Jesse Mu, Xiang Lisa Li, and Noah D. Goodman. 2023. Learning to compress prompts with gist tokens. CoRR, abs/2304.08467.
  17. 17.Niels Mundler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2024. Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation. In The Twelfth International Conference on Learning Representations.
  18. 18.OpenAI. 2022. Chatgpt.
  19. 19.OpenAI. 2023. GPT-4 technical report.
  20. 20.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  21. 21.Aaron Parisi, Yao Zhao, and Noah Fiedel. 2022. TALM: tool augmented language models. CoRR, abs/2205.12255.
  22. 22.Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2023. Gorilla: Large language model connected with massive apis. CoRR, abs/2305.15334.
  23. 23.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. CoRR, abs/2307.16789.
  24. 24.Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. 2024. Tool learning with large language models: A survey. arXiv preprint arXiv:2405.17935.
  25. 25.Mengjie Ren, Boxi Cao, Hongyu Lin, Liu Cao, Xianpei Han, Ke Zeng, Guanglu Wan, Xunliang Cai, and Le Sun. 2024. Learning or self-aligning? rethinking instruction fine-tuning. arXiv preprint arXiv:2402.18243.
  26. 26.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. CoRR, abs/2302.04761.
  27. 27.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023a. Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface. volume abs/2303.17580.
  28. 28.Yongliang Shen, Kaitao Song, Xu Tan, Wenqi Zhang, Kan Ren, Siyu Yuan, Weiming Lu, Dongsheng Li, and Yueting Zhuang. 2023b. Taskbench: Benchmarking large language models for task automation. arXiv preprint arXiv:2311.18760.
  29. 29.Yifan Song, Weimin Xiong, Dawei Zhu, Wenhao Wu, Han Qian, Mingbo Song, Hailiang Huang, Cheng Li, Ke Wang, Rong Yao, Ye Tian, and Sujian Li. 2023. Restgpt: Connecting large language models with real-world restful apis.
  30. 30.Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, and Le Sun. 2023. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.
  31. 31.Gemini Team and Google. 2023. Gemini: A family of highly capable multimodal models.
  32. 32.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  34. 34.Boshi Wang, Hao Fang, Jason Eisner, Benjamin Van Durme, and Yu Su. 2024. Llms in the imaginarium: tool learning through simulated trial and error. arXiv preprint arXiv:2403.04746.
  35. 35.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13484–13508, Toronto, Canada. Association for Computational Linguistics.
  36. 36.Qiantong Xu, Fenglu Hong, Bo Li, Changran Hu, Zhengyu Chen, and Jian Zhang. 2023. On the tool manipulation capability of open-source large language models.
  37. 37.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
  38. 38.Kexun Zhang, Hongqiao Chen, Lei Li, and William Wang. 2023a. Syntax error-free and generalizable tool use for llms via finite-state decoding.
  39. 39.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023b. Siren’s song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219.

Citation

MLA
Yuan, S., et al. “EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 951–72, https://doi.org/10.18653/v1/2025.naacl-long.44.
APA
Yuan, S., Song, K., Chen, J., Tan, X., Shen, Y., Ren, K., Li, D., & Yang, D. (2025). EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 951–972. https://doi.org/10.18653/v1/2025.naacl-long.44
Chicago
Yuan, S., K. Song, J. Chen, et al. 2025. “EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 951–72. https://doi.org/10.18653/v1/2025.naacl-long.44.
Harvard
Yuan, S. et al. (2025) “EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 951–972. Available at: https://doi.org/10.18653/v1/2025.naacl-long.44.
Vancouver
1. Yuan S, Song K, Chen J, Tan X, Shen Y, Ren K, Li D, Yang D (2025) EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 951–972

BibTeX

@inproceedings{yuan-etal-2025-easytool,
    title = "{EASYTOOL}: Enhancing {LLM}-based Agents with Concise Tool Instruction",
    author = "Yuan, Siyu  and
      Song, Kaitao  and
      Chen, Jiangjie  and
      Tan, Xu  and
      Shen, Yongliang  and
      Ren, Kan  and
      Li, Dongsheng  and
      Yang, Deqing",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.44/",
    doi = "10.18653/v1/2025.naacl-long.44",
    pages = "951--972",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/