The Tool Illusion: Rethinking Tool Use in Web Agents

Renze LouBaolin PengWenlin YaoQianhui WuHao ChengSuman NathWenpeng YinJianfeng Gao

article2026arXiv2 citations

Challenges the assumption that tool use consistently benefits web agents by systematically analyzing their limitations across diverse models and benchmarks, while establishing practical design principles for building reliable agent tools.

Listen

Autonomous web agents powered by large artificial intelligence models are increasingly applied to interact with software interfaces and complete online tasks. Traditionally, these agents operate through low-level actions such as clicking and typing, which can be inefficient and error-prone. To solve this, developers have increasingly adopted tools—such as programming interfaces and synthesized scripts—that bundle multi-step operations into single commands. However, existing research has largely relied on limited evaluations, creating conflicting evidence around whether tool integration consistently improves performance, how tools should be designed, and what hidden operational overheads they introduce.

The article systematically evaluates the real-world utility, design principles, and efficiency trade-offs of tool-use web agents. Across controlled experiments spanning five underlying language models, three distinct agent frameworks, and two realistic web evaluation benchmarks, the analysis examines whether tools deliver consistent performance gains and under what conditions they succeed or fail.

The findings show that tool synthesis functions primarily as a one-way capability transfer from stronger models to weaker models. Tools automatically generated by artificial intelligence provide steady improvements only when the executing agent is less capable than the model that created the tools. When advanced models execute tools generated by equal or weaker models, the performance benefits become negligible or negative (for instance, dropping success rates by up to 2–4 percentage points). In contrast, human-curated programming interfaces consistently improved task success across all tested models, boosting success rates by roughly 12 to 20 percentage points on the primary benchmark. Furthermore, the analysis reveals that highly complex, end-to-end tools suffer from poor generalization: in one major framework, 79% of synthesized tools were never invoked during testing. The presence of large tool libraries also introduced substantial overhead, more than tripling token consumption in some frameworks and increasing navigation steps due to retrieval and selection delays.

These results challenge the common assumption that adding programmatic tools is an unambiguous improvement for web agents. Overly specialized tools crowd out the flexible reasoning necessary to navigate dynamic interfaces, while bloated tool repositories increase operational costs and execution latency. Converting procedural tools into transparent, natural-language guidance (termed semantic skills) proved more effective for high-capacity models, as it allows flexible adaptation rather than rigid execution. Additionally, visual grounding remains critical; while tools make agents somewhat more resilient when visual data is absent, combining visual inputs with tools consistently produces the strongest results.

Organizations developing or deploying web agents should decouple deterministic user-interface actions from dynamic decision-making. Tools should focus on reusable, composable, and low-to-medium complexity operations rather than attempting to automate entire workflows in a single command. If budget or engineering constraints prevent building high-quality tool suites with top-tier models or human developers, teams should consider providing semantic natural-language guidelines to capable models instead of rigid programmatic functions. Further exploration in live, multi-site production environments is recommended to validate these design principles under broader real-world network conditions and interface changes.

No sufficiently relevant recommendations were found.

Cover for The Tool Illusion: Rethinking Tool Use in Web Agents

Abstract

As web agents rapidly evolve, an increasing body of work has moved beyond conventional atomic browser interactions and explored tool use as a higher-level action paradigm. Although prior studies have shown the promise of tools, their conclusions are often drawn from limited experimental scales and sometimes non-comparable settings. As a result, several fundamental questions remain unclear: i) whether tools provide consistent gains for web agents, ii) what practical design principles characterize effective tools, and iii) what side effects tool use may introduce. To establish a stronger empirical foundation for future research, we revisit tool use in web agents through an extensive and carefully controlled study across diverse tool sources, backbone models, tool-use frameworks, and evaluation benchmarks. Our findings both revise some prior conclusions and complement others with broader evidence. We hope this study provides a more reliable empirical basis and inspires future research on tool-use web agents.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background and Experimental Setup
  • 4 Experiments and Analyses
  • 4.1 Synthetic Tools Do Not Always Yield Consistent Gains
  • 4.2 Tool Synthesis as Capability Distillation: Strong-to-Weak Transfer
  • 4.3 Tools Should Not Be Overly Complex: Decoupling Agentic Reasoning from Tools
  • 4.4 The Essence of Tool Design: Functional Coverage and Composition
  • 4.5 The Tax of Tools: Token Cost and Action Overhead
  • 4.6 Skills as a Flexible Alternative to Tools
  • 4.7 Vision Remains Beneficial, While Tools Reduce Dependence on It
  • 5 Conclusion and Practical Takeaways
  • References
  • A Reproduction Details and Statement
  • A.1 WALT Reproduction
  • A.2 SkillWeaver Reproduction
  • A.3 Hybrid-Agent Reproduction
  • A.4 Other Implementation Details
  • B Additional Framework Details and Tool Examples
  • B.1 Framework Details Not Fully Documented in the Original Papers
  • B.2 Tool Examples
  • C Tool Complexity Levels
  • C.1 Prompt for Tool Complexity Classification
  • C.2 Heuristic Rules for Complexity Classification of Hybrid-Agent’s REST APIs
  • D Converting Tool Functions to Semantic Skills

Knowls

  1. Knowl 1 — Tool benefits depend on the source of the tools and the user model

    empirical result

    Across WEBARENA and VISUALWEBARENA, human-developed Hybrid-Agent tools improved success rates for all five tested backbone models, whereas the LLM-synthesized tools in WALT and SkillWeaver produced mixed results. The following average success rates are percentages, with values ordered as GPT-5, GPT-5-mini, GPT-5-nano, Grok-4.1-Fast-Reasoning, and Mistral-Large-3.1; each pair gives no tools → tools.

    • WALT, WEBARENA: 52.9→50.9, 46.2→46.0, 24.4→28.6, 44.5→46.6, 35.3→40.2. WALT, VISUALWEBARENA: 52.8→51.9, 46.0→47.6, 23.9→31.0, 43.9→44.7, 38.5→41.0.
    • SkillWeaver, WEBARENA: 39.2→37.4, 26.0→27.3, 11.0→13.1, 32.8→30.4, 23.6→22.9. SkillWeaver, VISUALWEBARENA: 37.9→34.4, 25.9→24.3, 11.2→13.2, 33.2→29.6, 22.3→20.6.
    • Hybrid-Agent, WEBARENA: 36.5→49.9, 19.3→38.9, 10.1→22.5, 26.2→42.7, 16.6→31.2. Hybrid-Agent, VISUALWEBARENA: 32.0→35.4, 16.0→24.2, 8.9→14.4, 25.5→29.5, 11.7→17.6.

    The pattern indicates that tool utility is conditional rather than universal: reliable, broad human-developed APIs helped across model strengths, while synthesized tools were most consistently useful to weaker tool users and could reduce performance for stronger users.

  2. Knowl 2 — LLM-synthesized tools transfer capability mainly from stronger constructors to weaker users

    empirical result

    The study varied the model constructing the tools while holding the tool-using model fixed. Tool scores were measured on CMS for SkillWeaver and Reddit for WALT; each triplet is ordered GPT-5, GPT-5-mini, GPT-5-nano, and all values are success-rate percentages.

    • SkillWeaver without tools: 43.4, 25.3, 9.9. Tools constructed by GPT-4o: 42.9, 28.6, 12.1; by GPT-5-mini: 41.8, 29.7, 16.5; by GPT-5: 44.5, 31.3, 17.6.
    • WALT without tools: 47.4, 38.6, 17.5. Tools constructed by GPT-4o: 37.7, 31.6, 24.6; by GPT-5-mini: 39.5, 33.3, 27.2; by GPT-5: 48.2, 43.0, 29.8.

    GPT-5-nano improved in every constructor/framework condition. GPT-5-mini improved reliably only when GPT-5 constructed the tools, while GPT-5 generally did not benefit except with GPT-5-constructed tools. This supports the paper's interpretation of synthesized tools as a form of mainly one-way capability distillation: knowledge encoded by a stronger constructor can help a weaker user, but the reverse transfer is not reliably beneficial.

  3. Knowl 3 — Complex, task-specific tools can lose generality and remain unused

    empirical result

    The study classified tools as high complexity when they attempted complex or end-to-end tasks and/or used many steps or state-dependent control logic; medium when they covered focused multi-step intentions; and low when they represented short, atomic actions. LLM prompting classified WALT and SkillWeaver functions, while documentation-based heuristics classified Hybrid-Agent APIs. The resulting high/medium/low shares were WALT 7% (3 tools)/34% (14)/59% (24), SkillWeaver 62% (272)/6% (27)/32% (139), and Hybrid-Agent 4% (73)/29% (566)/67% (1,307).

    On WEBARENA, the shares of tools invoked in zero, one, two, or at least three tasks were WALT 22% (9)/10% (4)/10% (4)/58% (23); SkillWeaver 79% (341)/8%/3%/10%; and Hybrid-Agent 20% (389)/18% (342)/11% (210)/51% (986). Thus, most SkillWeaver tools were never invoked, and its tool set had a much larger high-complexity share than the other two frameworks. The authors associate complex, UI-state-dependent synthesized procedures with task-specificity and reduced practical utility; adding more such tools did not translate into proportional use or generalizability.

  4. Knowl 4 — Composable functional coverage matters more than end-to-end tools or tool-set size

    empirical result

    For WEBARENA tasks, the distribution and success rates by number of distinct tools used were: WALT—0 tools, 357 tasks (45%), 56.3% success; 1 tool, 291 (37%), 43.6%; at least 2 tools, 140 (18%), 32.1%. SkillWeaver—0 tools, 159 (21%), 42.1%; 1 tool, 408 (54%), 39.2%; at least 2 tools, 194 (25%), 27.5%. Hybrid-Agent—0 tools, 82 (11%), 39.0%; 1 tool, 286 (37%), 50.3%; at least 2 tools, 396 (52%), 41.2%.

    Hybrid-Agent's tools were used in multi-tool combinations for a larger share of tasks than in the two synthesized-tool frameworks. Its smaller share of tasks completed with no tools also suggests broader coverage of benchmark intentions. The comparison motivates a design principle: tools need not solve a whole task individually; a set of focused tools can be useful when the tools compose reliably and collectively cover common user intentions. Tool count or individual comprehensiveness alone does not guarantee better performance.

  5. Knowl 5 — Tools impose token and action overhead that can outweigh their shortcuts

    empirical result

    On WEBARENA, average token cost per website increased when tools were enabled in all three frameworks. Prompt/completion/reasoning costs, in millions of tokens, were WALT 11.0/2.0/1.5 without tools versus 12.9/1.9/1.5 with tools (total 14.5M→16.3M); SkillWeaver 6.1/3.6/1.9 versus 11.9/4.6/3.0 (total 11.6M→19.5M); and Hybrid-Agent 5.8/2.7/2.0 versus 22.1/6.9/4.9 (total 10.5M→33.9M). The especially large Hybrid-Agent increase was associated with its very large API library.

    Average agent steps per task changed from 10.2 to 9.3 for WALT, 7.3 to 7.2 for SkillWeaver, and 7.1 to 8.3 for Hybrid-Agent, without tools versus with tools. Thus, only WALT reduced step count, SkillWeaver changed little, and Hybrid-Agent required more steps. The study attributes overhead to tool retrieval and selection, as well as error recovery and parameter adjustment for low-utility tools; tool construction itself also carries non-trivial but unquantified expense.

  6. Knowl 6 — Semantic skills can outperform executable tools for models able to use flexible guidance

    empirical result

    The researchers converted SkillWeaver Python tool functions into natural-language procedural descriptions using GPT-4.1, then supplied these descriptions as guidance rather than executable calls. On CMS and Reddit, respectively, success rates for GPT-5 were: no tools 43.4% and 31.1%, executable tools 42.9% and 30.2%, semantic skills 45.1% and 34.9%. For GPT-5-nano, the corresponding results were CMS 9.9%, 12.1%, and 9.3%; Reddit 7.5%, 11.3%, and 10.4%.

    Semantic skills outperformed both the baseline and tool condition for GPT-5, but not for GPT-5-nano, for which executable tools did better than semantic descriptions. The authors interpret this as a trade-off: natural-language skills expose procedures so a capable agent can inspect, adapt, partially follow, or ignore them, while requiring more reasoning than directly invoking a tool.

  7. Knowl 7 — Visual grounding remains useful with tools, while tools can reduce dependence on it

    empirical result

    An ablation compared success rates with and without webpage screenshots on CMS and Reddit in WEBARENA. Values below are with vision→without vision, shown separately for agents without and with tools.

    • WALT: CMS 52.2→46.2 without tools and 57.1→50.0 with tools; Reddit 47.4→37.7 without tools and 39.5→36.8 with tools.
    • SkillWeaver: CMS 43.4→37.4 without tools and 42.9→38.5 with tools; Reddit 31.1→26.4 without tools and 30.2→28.3 with tools.
    • Hybrid-Agent, with screenshots added to its prompts: CMS 46.7→39.6 without tools and 52.7→50.5 with tools; Reddit 42.5→38.7 without tools and 58.5→57.5 with tools.

    Removing vision lowered performance in every reported condition, so screenshots remained beneficial even when tools were available. In most comparisons, the vision/no-vision gap was smaller with tools than without them, suggesting that tools partially reduce dependence on direct visual grounding rather than replacing its value.

  8. Knowl 8 — Controlled evaluation spans three tool frameworks, five models, and two benchmarks

    experimental setup

    The study compared three open-source web-agent frameworks across GPT-5, GPT-5-mini, GPT-5-nano, Grok-4.1-Fast-Reasoning, and Mistral-Large-3.1. Hybrid-Agent uses human-developed REST APIs (about 437 tools per website on average); SkillWeaver uses LLM-synthesized Python functions (about 88 per website); and WALT uses LLM-synthesized UI-action flows (about 8 per website). The main experiments used released tool sets; the tool-constructor scaling analysis built additional tools with different constructors.

    Evaluation used binary task success on WEBARENA—Shopping (187 tasks), CMS/Shopping Admin (182), GitLab (180), Reddit (106), and Map (109)—and VISUALWEBARENA—Classifieds (234), Shopping (466), and Reddit (210). Multi-site WEBARENA tasks were excluded because SkillWeaver tools targeted single-site tasks. WEBARENA evaluates task completion with benchmark scripts using string matching or HTML parsing; VISUALWEBARENA adds visually grounded tasks. The broad model, framework, and benchmark comparison was designed to test how tool utility varies with tool source, model strength, and evaluation setting.

  9. Knowl 9 — The study's conclusions are scoped by focused constructor tests and benchmark adaptations

    limitation

    Because tool synthesis was expensive, the experiment varying tool constructors covered only CMS for SkillWeaver and Reddit for WALT, rather than the full set of websites. Multi-site WEBARENA tasks were also omitted because SkillWeaver's tools were designed for single-site tasks. Hybrid-Agent was originally text-only; its VISUALWEBARENA evaluation required adding screenshots to prompts and truncating web-state observations to control token costs. These design choices limit the direct scope of those particular comparisons, so the constructor-transfer findings are based on two representative sites and the Hybrid-Agent visual results use an adapted setup.

Coverage note — Appendix-level reproduction instructions, representative tool code, and the full complexity-classification prompts are omitted as implementation detail rather than distinct empirical contributions.

References

  1. 1.Anthropic. Public repository for Agent Skills. https://github.com/anthropics/skills, 2026. GitHub repository.
  2. 2.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 9313–9332, 2024.
  3. 3.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36:28091–28114, 2023.
  4. 4.Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur. Multimodal web navigation with instruction-finetuned foundation models. In The Twelfth International Conference on Learning Representations, 2024.
  5. 5.Yu Gu, Kai Zhang, Yuting Ning, Boyuan Zheng, Boyu Gou, Tianci Xue, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, et al. Is your llm secretly a world model of the internet? model-based planning for web agents. Transactions on Machine Learning Research, 2024.
  6. 6.Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. Webvoyager: Building an end-to-end web agent with large multimodal models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6864–6890, 2024.
  7. 7.Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, and Dong Yu. Openwebvoyager: Building multimodal web agents via iterative real-world exploration, feedback and optimization. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 27545–27564, 2025.
  8. 8.SU Hongjin, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan O Arik. Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments. In The Thirteenth International Conference on Learning Representations, 2025.
  9. 9.Linxi Jiang, Rui Xi, Zhijie Liu, Shuo Chen, Zhiqiang Lin, and Suman Nath. Web verbs: Typed abstractions for reliable task composition on the agentic web. arXiv preprint arXiv:2602.17245, 2026.
  10. 10.Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, and Wei-Lun Chao. Dual-view visual contextualization for web navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14445–14454, 2024.
  11. 11.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. arXiv preprint arXiv:2401.13649, 2024a.
  12. 12.Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov. Tree search for language model agents. arXiv preprint arXiv:2407.01476, 2024b.
  13. 13.Shambhavi Krishna, Zheng Chen, Yuan Ling, Xiaojiang Huang, Yingjie Li, Fan Yang, and Xiang Li. Paffa: Premeditated actions for fast agents. arXiv preprint arXiv:2412.07958, 2024.
  14. 14.Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, et al. Skillsbench: Benchmarking how well agent skills work across diverse tasks. arXiv preprint arXiv:2602.12670, 2026.
  15. 15.Junteng Liu, Yunji Li, Chi Zhang, Jingyang Li, Aili Chen, Ke Ji, Weiyu Cheng, Zijia Wu, Chengyu Du, Qidi Xu, et al. Webexplorer: Explore and evolve for training long-horizon web agents. arXiv preprint arXiv:2509.06501, 2025.
  16. 16.Mistral AI. Introducing mistral 3. https://mistral.ai/news/mistral-3, 2025. Official Mistral AI announcement.
  17. 17.Vardaan Pahuja, Yadong Lu, Corby Rosset, Boyu Gou, Arindam Mitra, Spencer Whitehead, Yu Su, and Ahmed Hassan. Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 6300–6323, 2025.
  18. 18.Viraj Prabhu, Yutong Dai, Matthew Fernandez, Jing Gu, Krithika Ramakrishnan, Yanqi Luo, Silvio Savarese, Caiming Xiong, Junnan Li, Zeyuan Chen, et al. Walt: Web agents that learn tools. arXiv preprint arXiv:2510.01524, 2025.
  19. 19.Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. Agent q: Advanced reasoning and learning for autonomous ai agents. arXiv preprint arXiv:2408.07199, 2024.
  20. 20.Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267, 2025.
  21. 21.Paloma Sodhi, SRK Branavan, Yoav Artzi, and Ryan McDonald. Step: Stacked llm policies for web actions. In First Conference on Language Modeling, 2024.
  22. 22.Yueqi Song, Frank F Xu, Shuyan Zhou, and Graham Neubig. Beyond browsing: Api-based web agents. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 11066–11085, 2025.
  23. 23.Brandon Walderman, Leo Lee, Andrew Nolan, David Bokan, Khushal Sagar, and Hannah Van Opstal. WebMCP. https://github.com/webmachinelearning/webmcp, 2025. GitHub repository.
  24. 24.Zhaoyang Wang, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao, et al. Adapting web agents with synthetic supervision. arXiv preprint arXiv:2511.06101, 2025a.
  25. 25.Zora Zhiruo Wang, Apurva Gandhi, Graham Neubig, and Daniel Fried. Inducing programmatic skills for agentic tasks. arXiv preprint arXiv:2504.06821, 2025b.
  26. 26.Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He, Deyu Zhou, Pengjun Xie, et al. Webwalker: Benchmarking llms in web traversal. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 10290–10305, 2025.
  27. 27.xAI. Grok 4.1. https://x.ai/news/grok-4-1, 2025. Official xAI announcement, November 17, 2025.
  28. 28.Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, et al. Skillrl: Evolving agents via recursive skill-augmented reinforcement learning. arXiv preprint arXiv:2602.08234, 2026.
  29. 29.Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. Travelplanner: a benchmark for real-world planning with language agents. In Proceedings of the 41st International Conference on Machine Learning, pp. 54590–54613, 2024a.
  30. 30.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2024b.
  31. 31.Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023.
  32. 32.Yuhao Yang, Zhen Yang, Zi-Yi Dou, Anh Nguyen, Keen You, Omar Attia, Andrew Szot, Michael Feng, Ram Ramrakhya, Alexander Toshev, et al. Ultracua: A foundation model for computer use agents with hybrid action. arXiv preprint arXiv:2510.17790, 2025.
  33. 33.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757, 2022.
  34. 34.Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, Xiyao Wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, et al. Agent learning via early experience. arXiv preprint arXiv:2510.08558, 2025.
  35. 35.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v (ision) is a generalist web agent, if grounded. In International Conference on Machine Learning, pp. 61349–61385. PMLR, 2024.
  36. 36.Boyuan Zheng, Michael Y Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, et al. Skillweaver: Web agents can self-improve by discovering and honing skills. arXiv preprint arXiv:2504.07079, 2025.
  37. 37.Hongbin Zhong, Fazle Faisal, Luis França, Tanakorn Leesatapornwongsa, Adriana Szekeres, Kexin Rong, and Suman Nath. Actionengine: From reactive to programmatic gui agents via state machine memory. arXiv preprint arXiv:2602.20502, 2026.
  38. 38.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. In The Twelfth International Conference on Learning Representations, 2024.

Citation

MLA
Lou, R., et al. “The Tool Illusion: Rethinking Tool Use in Web Agents”. arXiv, 2026, http://arxiv.org/abs/2604.03465v2.
APA
Lou, R., Peng, B., Yao, W., Wu, Q., Cheng, H., Nath, S., Yin, W., & Gao, J. (2026). The Tool Illusion: Rethinking Tool Use in Web Agents. arXiv. http://arxiv.org/abs/2604.03465v2
Chicago
Lou, R., B. Peng, W. Yao, et al. 2026. “The Tool Illusion: Rethinking Tool Use in Web Agents”. arXiv. http://arxiv.org/abs/2604.03465v2.
Harvard
Lou, R. et al. (2026) “The Tool Illusion: Rethinking Tool Use in Web Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.03465v2.
Vancouver
1. Lou R, Peng B, Yao W, Wu Q, Cheng H, Nath S, Yin W, Gao J (2026) The Tool Illusion: Rethinking Tool Use in Web Agents. arXiv

BibTeX

@article{lou2026the,
  title = {The Tool Illusion: Rethinking Tool Use in Web Agents},
  author = {Lou, Renze and Peng, Baolin and Yao, Wenlin and Wu, Qianhui and Cheng, Hao and Nath, Suman and Yin, Wenpeng and Gao, Jianfeng},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.03465v2},
  eprint = {2604.03465}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/