The Tool Illusion: Rethinking Tool Use in Web Agents
Renze LouBaolin PengWenlin YaoQianhui WuHao ChengSuman NathWenpeng YinJianfeng Gao
Challenges the assumption that tool use consistently benefits web agents by systematically analyzing their limitations across diverse models and benchmarks, while establishing practical design principles for building reliable agent tools.
Autonomous web agents powered by large artificial intelligence models are increasingly applied to interact with software interfaces and complete online tasks. Traditionally, these agents operate through low-level actions such as clicking and typing, which can be inefficient and error-prone. To solve this, developers have increasingly adopted tools—such as programming interfaces and synthesized scripts—that bundle multi-step operations into single commands. However, existing research has largely relied on limited evaluations, creating conflicting evidence around whether tool integration consistently improves performance, how tools should be designed, and what hidden operational overheads they introduce.
The article systematically evaluates the real-world utility, design principles, and efficiency trade-offs of tool-use web agents. Across controlled experiments spanning five underlying language models, three distinct agent frameworks, and two realistic web evaluation benchmarks, the analysis examines whether tools deliver consistent performance gains and under what conditions they succeed or fail.
The findings show that tool synthesis functions primarily as a one-way capability transfer from stronger models to weaker models. Tools automatically generated by artificial intelligence provide steady improvements only when the executing agent is less capable than the model that created the tools. When advanced models execute tools generated by equal or weaker models, the performance benefits become negligible or negative (for instance, dropping success rates by up to 2–4 percentage points). In contrast, human-curated programming interfaces consistently improved task success across all tested models, boosting success rates by roughly 12 to 20 percentage points on the primary benchmark. Furthermore, the analysis reveals that highly complex, end-to-end tools suffer from poor generalization: in one major framework, 79% of synthesized tools were never invoked during testing. The presence of large tool libraries also introduced substantial overhead, more than tripling token consumption in some frameworks and increasing navigation steps due to retrieval and selection delays.
These results challenge the common assumption that adding programmatic tools is an unambiguous improvement for web agents. Overly specialized tools crowd out the flexible reasoning necessary to navigate dynamic interfaces, while bloated tool repositories increase operational costs and execution latency. Converting procedural tools into transparent, natural-language guidance (termed semantic skills) proved more effective for high-capacity models, as it allows flexible adaptation rather than rigid execution. Additionally, visual grounding remains critical; while tools make agents somewhat more resilient when visual data is absent, combining visual inputs with tools consistently produces the strongest results.
Organizations developing or deploying web agents should decouple deterministic user-interface actions from dynamic decision-making. Tools should focus on reusable, composable, and low-to-medium complexity operations rather than attempting to automate entire workflows in a single command. If budget or engineering constraints prevent building high-quality tool suites with top-tier models or human developers, teams should consider providing semantic natural-language guidelines to capable models instead of rigid programmatic functions. Further exploration in live, multi-site production environments is recommended to validate these design principles under broader real-world network conditions and interface changes.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena establishes a realistic, reproducible benchmark for multi-step web tasks, helping frame the source’s controlled comparisons across evaluation settings.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web provides the broad real-website task and interaction-trace foundation that makes later claims about general web-agent capabilities easier to interpret.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). WebVoyager demonstrates an end-to-end multimodal web agent and benchmark, supplying a concrete baseline for understanding the source’s reassessment of agent tool use.
- Paper: AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, Ori Yoran et al. (2024). AssistantBench establishes realistic, time-consuming open-web tasks and agent comparisons that contextualize the source’s effort to broaden and standardize evaluation.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). ToolLLM develops large-scale API selection and multi-tool use, providing essential background for the source’s investigation of tool sources and tool-use frameworks.
No sufficiently relevant recommendations were found.
