Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Jiwan ChungJiHyuk ByunVibhav VineetSeon Joo Kim
Introduces WebStep, an evaluation benchmark that uses automatic semantic state tracking to diagnose step-by-step agent failures, revealing critical skill-specific differences hidden behind identical overall success rates.
Web agents are increasingly deployed to automate complex interactive tasks such as online shopping, administrative scheduling, and software management. However, conventional benchmarks assess these systems primarily through terminal success, compressing extended interaction sequences into a single pass-or-fail outcome. This outcome-only perspective obscures whether a failure was caused by inadequate exploration, misinterpreting retrieved information, or making minor mechanical errors during the final action. The article introduces WEBSTEP, a diagnostic evaluation benchmark designed to provide automated, step-by-step visibility into web agent behavior and isolate the precise points where workflows break down.
The research team constructed ten self-hosted, deterministic websites spanning multiple realistic application categories, encompassing 1,800 distinct task instances. Each graphical interface is directly coupled to a structured background state machine, allowing the environment to automatically record semantic states and transitions without requiring human annotation. The authors evaluated six generalist and specialist agents across multiple difficulty levels, systematically controlling task complexity by planting plausible distractor items that require deep page inspection to differentiate from target items. The evaluation framework measures exploration reach, execution accuracy conditioned on finding the correct target, and the timing of specific interaction skills.
The analysis reveals several critical findings that terminal scores mask. First, agents with nearly identical terminal success rates display fundamentally different operational profiles; among smaller models clustering between 34% and 37% overall success, UI-TARS achieved 5.9% higher exploration success than GUI-Owl but underperformed it during final execution by 4.7%, while Fara suffered from limited information exploration. Second, capabilities vary significantly across specific interaction skills within identical tasks; on question-and-answer workflows, Claude CUA outperformed OpenAI CUA by 30.0% on navigation actions while trailing by 6.7% on detail inspection. Third, trajectory analysis indicates that agents diverge into errors through model-specific mechanisms, such as GUI-Owl failing through filter misapplication (31% of early divergences) while others faltered during search reformulations or coordinate-level button clicking. Finally, capability gaps widen sharply as task complexity increases: leading closed models retained approximately 90% of their exploration accuracy when distractors scaled from one to three, whereas smaller models experienced performance drops between 26% and 49%.
These results demonstrate that outcome-only benchmarks provide an incomplete and potentially misleading basis for selecting and improving autonomous agents. By separating exploration from execution and mapping failures to specific skills, process-level evaluation enables engineers to target precise training interventions—such as refining coordinate grounding for execution-weak models or reinforcing deeper evidence gathering before task commitment. This structural diagnosis reduces development costs, mitigates deployment risks, and provides clarity on operational failure modes before agents are released into high-stakes environments.
Organizations developing or deploying web agents should implement process-level state tracking within their evaluation pipelines to isolate skill-specific deficiencies. Key limitations of the study include its reliance on deterministic, simulated environments and synthetic data templates, which do not fully replicate asynchronous loading, dynamic live-web updates, or open-ended user requests. Nonetheless, the high degree of experimental control and exact semantic alignment provides strong confidence in the diagnostic conclusions regarding foundational agent behaviors.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena establishes the standard realistic web simulation environment and outcome-based evaluation framework that the source specifically critiques and improves through process-level semantic state tracking.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web provides foundational task formulations and multi-step web interaction traces across diverse real-world websites upon which web agent benchmarks build.
- Paper: Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement, Weimin Xiong et al. (2024). This paper introduces step-level process refinement and intermediate state monitoring for interactive agents, establishing the core rationale for moving beyond binary outcome rewards.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). WebVoyager demonstrates visual and multimodal interaction for autonomous web agents, providing the execution paradigm evaluated across browser interfaces in the source.
- Paper: AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, Ori Yoran et al. (2024). AssistantBench illustrates the limitations of evaluating agents solely on complex, multi-site open-ended tasks without fine-grained intermediate execution diagnostics.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). Agent-as-a-Judge highlights the necessity of inspecting entire trajectories and intermediate execution states rather than relying strictly on final outcome verifications.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This work lays the conceptual foundation for step-by-step process supervision versus outcome supervision in complex reasoning and action chains.
- Paper: Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning, Rodrigo Toro Icarte et al. (2022). Reward Machines introduces formal state-transition representations to expose intermediate environment dynamics, underlying the semantic MDP formulation used in WebStep.
- Paper: AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation, Priyam Sahoo et al. (2026). AgentLens extends process-level trajectory analysis to software engineering agents to diagnose misleading successful outcomes that conceal chaotic intermediate failures.
- Paper: TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents, Leitian Tao et al. (2026). TRACE applies turn-level state transition credit assignment to long-horizon agent trajectories, addressing the credit assignment challenges diagnosed by process evaluation.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). MemGym investigates intermediate memory management failures across long interactive trajectories, complementing fine-grained trajectory and skill decomposition in web environments.
- Paper: Covering Human Action Space for Computer Use: Data Synthesis and Benchmark, Miaosen Zhang et al. (2026). CUActSpot builds upon diagnostic failure analyses by broadening and benchmarking the fine-grained visual action grounding space required for robust computer use.
- Paper: Scaling Laws for Agent Harnesses via Effective Feedback Compute, Xuanliang Zhang et al. (2026). This work formulates scaling laws based on informative intermediate feedback compute across agent harnesses, leveraging trajectory-level evaluation principles.
