EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
Siyu YuanKaitao SongJiangjie ChenXu TanYongliang ShenKan RenDongsheng LiDeqing Yang
Proposes EASYTOOL, a framework that converts lengthy, inconsistent tool documentation into standardized and concise instructions, substantially lowering prompt token costs while boosting tool-use accuracy across diverse agent tasks.
Autonomous agents powered by large language models increasingly rely on external software tools to complete complex tasks, such as accessing live web data and performing specialized computations. However, existing tool documentation across diverse sources suffers from severe formatting inconsistencies, high verbosity, and a lack of clear usage examples. These defects cause language models to exceed context limits, misinterpret tool purposes, and supply invalid parameters, resulting in frequent task failures.
The article introduces and evaluates EASYTOOL, a standardized framework that condenses raw, lengthy documentation into concise, structured tool instructions. The objective is to leverage the instruction-following capabilities of language models to improve tool retrieval, tool selection, and parameter accuracy while reducing computational overhead.
The authors tested EASYTOOL across three distinct benchmarks: ToolBench for real-world question answering across diverse web services, RestBench for multi-step service planning, and FuncQA for multi-step mathematical reasoning. The evaluation compared commercial and open-source models—including ChatGPT, GPT-4, GPT-4o, and Llama 3.1 variants—using raw documentation, standard prompt baselines, and EASYTOOL instructions. Independent human evaluators validated instruction quality and analyzed failure modes.
The investigation produced four primary findings. First, EASYTOOL drastically reduces token consumption, shrinking prompt sizes by 70.43% in ToolBench and 97.35% in RestBench. Second, standardized instructions significantly improve end-to-end task performance; on ToolBench, ChatGPT's success rate rose from 15.0% to 52.8%, and GPT-4o reached a 77.0% success rate. Third, tool selection and execution errors plummeted, virtually eliminating tool name errors and reducing parameter errors from 25% to 6% in ChatGPT and from 17% to 1% in GPT-4. Fourth, smaller open-source models benefited substantially; Llama-3.1-8B-Instruct improved its success rate from 5.5% to 48.5%, surpassing larger, specialized fine-tuned models.
These findings demonstrate that restructuring documentation into concise functional summaries and concrete parameter examples is far more effective than feeding raw documentation or using generic prompt compression techniques, which often destroy critical parameter syntax. For organizations building agentic systems, adopting standardized instruction layers substantially reduces token-based operating costs, increases system reliability, and allows teams to deploy smaller, lower-cost open-source models without sacrificing task accuracy.
Organizations developing tool-augmented language model workflows should implement structured pre-processing pipelines to convert complex API documentation into unified instructions with explicit scenario examples. Before broad deployment, development teams should run pilot evaluations on their specific API libraries, ensuring that tool descriptions capture multi-functional capabilities and exact parameter constraints.
The findings are supported by high inter-annotator agreement and consistent performance gains across diverse task types. Nevertheless, decision-makers should note key limitations: the framework requires source documentation to fit within the pre-processing model's context window, does not explicitly model inter-tool dependencies, and relies on base models possessing baseline instruction-following capabilities.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). ToolLLM provides the foundational ToolBench dataset and evaluation paradigm for multi-API tool use that EASYTOOL directly adopts and optimizes.
- Paper: Gorilla: Large Language Model Connected with Massive APIs, Shishir G. Patil et al. (2023). Gorilla establishes the challenge of hallucinated parameters and verbose API documentation retrieval, defining the core problem EASYTOOL solves through structured instruction simplification.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Toolformer introduces the foundational paradigm of augmenting language models with external API call execution that modern agent architectures build upon.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). LongLLMLingua demonstrates the mechanics and limitations of generic prompt compression in long-context tasks, highlighting why task-tailored tool instruction formatting is necessary.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). SWE-agent demonstrates how designing concise, structured agent-computer interfaces substantially mitigates the failure modes of raw, verbose system documentation.
- Paper: ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings, Shibo Hao et al. (2023). ToolkenGPT highlights context-length bottlenecks and retrieval challenges when handling massive tool spaces, motivating EASYTOOL's standardized token-reduction approach.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey generalizes the principles of token reduction, API selection, and efficient tool execution evaluated in EASYTOOL into a unified framework for efficient agent design.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). Agentic Context Engineering extends the idea of static, concise tool instructions by dynamically evolving structured context playbooks based on execution experience.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). SWE-Pruner applies task-aware context compression and pruning principles specifically to coding agents navigating verbose software repositories.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Agent Workflow Memory builds on structured agent execution by inducing and reusing procedural workflows from past successful tool interactions.
