AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
Ori YoranSamuel Joseph AmouyalChaitanya MalaviyaBen BoginOfir PressJonathan Berant
Introduces AssistantBench, an automatically evaluated benchmark of realistic and time-consuming web tasks that exposes severe limitations in existing language models, alongside SeePlanAct, a new agent architecture designed to improve multi-step web execution.
Automated artificial intelligence assistants have significant potential to help users handle routine, time-consuming internet research, such as market monitoring or multi-step local searches. However, existing benchmarks primarily evaluate these tools on single websites or restricted sandbox environments rather than the open web. Consequently, current performance metrics fail to reflect whether automated systems can reliably navigate multiple live websites, plan dynamic browsing paths, and synthesize real-world information.
The article introduces ASSISTANTBENCH, a new evaluation benchmark designed to measure how effectively language models and autonomous web agents solve realistic, multi-step web tasks. The authors also present and evaluate a new web agent architecture, SEEPLANACT (SPA), created to improve browsing performance through explicit planning and memory mechanisms.
The benchmark comprises 214 realistic, automatically verifiable tasks spanning diverse domains collected from 53 human contributors, including 35 domain experts. Completing these tasks requires navigating more than 525 distinct webpages across 258 websites. The authors conducted comparative experiments using leading closed-book models, retrieval-augmented models with search engine access, existing state-of-the-art web agents, and their proposed SPA agent, using both GPT-4-Turbo and Claude-3.5-Sonnet engines.
The investigation produced several key findings. First, all evaluated systems struggled severely, with no standalone model exceeding 26% accuracy. Second, standalone autonomous web agents achieved poor overall results due to frequent navigation failures; the baseline SEEACT agent scored only 4.1% accuracy. Third, while the proposed SPA agent outperformed the baseline by 7 points in accuracy and reached 11.1% on its own, closed-book models achieved higher raw accuracy (up to 22.2%) simply because they abstained less often. However, closed-book models suffered from severe unreliability, hallucinating facts in 85% of their error cases. Fourth, an ensemble pairing the SPA agent with a closed-book fallback reached the top overall test accuracy of 25.2% to 26.4%. Finally, web agents exhibited high failure rates on both very short and very long interaction trajectories, with error rates peaking sharply on tasks requiring more than 15 browsing actions.
These findings indicate that relying on current AI models for open-web assistance introduces substantial operational and informational risks. Pure language models frequently present plausible but fabricated data, while retrieval-augmented systems routinely fail to retrieve complete context from complex web interfaces. Autonomous web agents currently suffer from navigation loops and interaction grounding errors, making them unreliable for unsupervised deployment in professional or high-stakes environments.
Organizations developing or deploying AI-driven web assistants should not rely on current standalone agents for end-to-end autonomous research without human verification. Developers should prioritize hybrid architectures that integrate explicit step-by-step planning and structured memory, while establishing robust fallback mechanisms when agents encounter navigation uncertainty. Further research must focus on improving open-web navigation training and developing reliable evaluation protocols for time-sensitive, dynamic web content.
Confidence in these findings is reinforced by realistic multi-domain task sourcing and consistent performance trends across multiple advanced language model backbones. Nevertheless, readers should account for certain limitations: the benchmark contains a relatively compact set of 214 tasks restricted to verifiable static outputs, and evaluation was bounded to proprietary commercial models due to the high computational costs of multi-turn web browsing.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). It provides the foundational framework and evaluation standard for realistic, multi-step autonomous web agents that AssistantBench expands upon for time-consuming open-web tasks.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). It establishes essential methodologies for building and evaluating generalist agents on raw, unmodified websites across diverse web domains.
- Paper: WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents, Shunyu Yao et al. (2022). It introduces grounding interactive language agents in real-world web environments, serving as a direct precursor to open-web task execution.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). It defines the core paradigm of interleaving reasoning traces and environment actions that powers modern web agents like SeePlanAct.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). It demonstrates early techniques for equipping language models with browser navigation and web-search capabilities to answer information-seeking questions.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). It provides a comprehensive taxonomy of architectures, planning modules, and evaluation benchmarks for LLM-based autonomous agents.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). It advances web-based agentic performance by training models via reinforcement learning to autonomously interleave reasoning with search engine interactions.
- Paper: Search-o1: Agentic Search-Enhanced Large Reasoning Models, Xiaoxi Li et al. (2025). It extends agentic web retrieval into large reasoning models by dynamically generating queries and synthesizing documents during multi-step problem solving.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). It extends autonomous agent interface design by tailoring specialized computer interfaces to tackle complex, long-horizon real-world tasks.
- Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). It addresses the memory overload and context limits encountered by agents executing deep web research and long-horizon tasks.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It synthesizes subsequent advancements in agentic reasoning, planning, and tool use across evolving real-world interactive environments.
