Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Jiwan ChungJiHyuk ByunVibhav VineetSeon Joo Kim

article2026arXiv3 citations

Introduces WebStep, an evaluation benchmark that uses automatic semantic state tracking to diagnose step-by-step agent failures, revealing critical skill-specific differences hidden behind identical overall success rates.

Listen

Web agents are increasingly deployed to automate complex interactive tasks such as online shopping, administrative scheduling, and software management. However, conventional benchmarks assess these systems primarily through terminal success, compressing extended interaction sequences into a single pass-or-fail outcome. This outcome-only perspective obscures whether a failure was caused by inadequate exploration, misinterpreting retrieved information, or making minor mechanical errors during the final action. The article introduces WEBSTEP, a diagnostic evaluation benchmark designed to provide automated, step-by-step visibility into web agent behavior and isolate the precise points where workflows break down.

The research team constructed ten self-hosted, deterministic websites spanning multiple realistic application categories, encompassing 1,800 distinct task instances. Each graphical interface is directly coupled to a structured background state machine, allowing the environment to automatically record semantic states and transitions without requiring human annotation. The authors evaluated six generalist and specialist agents across multiple difficulty levels, systematically controlling task complexity by planting plausible distractor items that require deep page inspection to differentiate from target items. The evaluation framework measures exploration reach, execution accuracy conditioned on finding the correct target, and the timing of specific interaction skills.

The analysis reveals several critical findings that terminal scores mask. First, agents with nearly identical terminal success rates display fundamentally different operational profiles; among smaller models clustering between 34% and 37% overall success, UI-TARS achieved 5.9% higher exploration success than GUI-Owl but underperformed it during final execution by 4.7%, while Fara suffered from limited information exploration. Second, capabilities vary significantly across specific interaction skills within identical tasks; on question-and-answer workflows, Claude CUA outperformed OpenAI CUA by 30.0% on navigation actions while trailing by 6.7% on detail inspection. Third, trajectory analysis indicates that agents diverge into errors through model-specific mechanisms, such as GUI-Owl failing through filter misapplication (31% of early divergences) while others faltered during search reformulations or coordinate-level button clicking. Finally, capability gaps widen sharply as task complexity increases: leading closed models retained approximately 90% of their exploration accuracy when distractors scaled from one to three, whereas smaller models experienced performance drops between 26% and 49%.

These results demonstrate that outcome-only benchmarks provide an incomplete and potentially misleading basis for selecting and improving autonomous agents. By separating exploration from execution and mapping failures to specific skills, process-level evaluation enables engineers to target precise training interventions—such as refining coordinate grounding for execution-weak models or reinforcing deeper evidence gathering before task commitment. This structural diagnosis reduces development costs, mitigates deployment risks, and provides clarity on operational failure modes before agents are released into high-stakes environments.

Organizations developing or deploying web agents should implement process-level state tracking within their evaluation pipelines to isolate skill-specific deficiencies. Key limitations of the study include its reliance on deterministic, simulated environments and synthetic data templates, which do not fully replicate asynchronous loading, dynamic live-web updates, or open-ended user requests. Nonetheless, the high degree of experimental control and exact semantic alignment provides strong confidence in the diagnostic conclusions regarding foundational agent behaviors.

arXiv: 2606.15673
Cover for Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Abstract

Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work, we conduct a process-level analysis of web agents. We introduce WebStep, a benchmark of 1,800 task instances with controlled difficulty and automatic semantic state tracking. Each website exposes a deterministic semantic MDP alongside the GUI: the agent operates on the interface, while the environment records high-level states and transitions in the background, enabling fine-grained analysis without manual annotation. Based on the semantic trajectory, we first show that process metrics reveal differences invisible to outcome evaluation: three agents whose success rates cluster within 34-37% diverge in exploration reach versus execution accuracy. Then, decomposing by skill characterizes the nature of these differences, exposing opposite per-skill rankings hidden within the same website: e.g., on Q&A, Claude CUA outperforms OpenAI CUA by 30% on navigation actions yet underperforms it by 6.7% on inspection, pinpointing a concrete skill to improve even within a domain. Bifurcation analysis further localizes the decisive error that loses the task and shows that this error is agent-specific rather than shared. Finally, these differences widen as tasks grow harder: success rate is similar on easy tasks but separates sharply as exploration becomes more demanding. Our process-level analysis opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved. Project page: this https URL

Table of Contents

  • 1 Introduction
  • 2 Process-level evaluation via semantic MDP duals
  • 2.1 Semantic MDP
  • 2.2 Task and world generation
  • 3 WebStep
  • 4 Experiments and Results
  • 4.1 Aggregate Results
  • 4.2 Skill-level diagnosis
  • 4.3 Trajectory bifurcations
  • 4.4 Exploration success by task complexity
  • 5 Conclusion
  • References
  • A Limitations
  • B Related Work
  • C Implementation Details
  • C.1 Agent Configurations
  • C.2 Compute and Runtime Statistics
  • C.3 Evaluation Protocol
  • C.4 Information Coverage
  • C.5 Reproducibility
  • C.6 LLM Usage
  • D Benchmark Specification
  • D.1 Task Distribution
  • D.2 Site Screenshots
  • D.3 Per-Site MDP Specifications
  • E Data Generation Pipeline
  • E.1 Site Construction Pipeline
  • E.2 World Generation
  • E.3 Oracle Trajectory Generation
  • F Extended Quantitative Results
  • F.1 Per-site performance decomposition
  • F.2 Per-site skill invocation
  • F.3 Aggregate skill invocation
  • F.4 Exploration SR by problem complexities
  • G Artifact Viewer
  • H Data Examples
  • H.1 Task Template Catalog
  • H.2 Example Trajectories: Agent vs. Oracle Trajectory
  • H.3 MDP Surface Transition Graphs
  • I Case Study
  • I.1 Case 1: Model Scale and Browsing Quality
  • I.2 Case 2: Premature Commitment to Hard Negative
  • I.3 Case 3: Thorough Exploration Before Commitment
  • I.4 Case 4: Exploration Success, Execution Failure
  • I.5 Case 5: Exploration Failure, Execution Success
  • I.6 Case 6: Redundant Process Despite Success
  • I.7 Case 7: Failure by Safeguard, Not by Competence

Knowls

  1. Knowl 1 — Semantic MDP Dual Architecture for Web Environments

    model/method

    To evaluate web agents at the process level without manual annotation, a web environment can be constructed as a semantic Markov Decision Process (MDP) dual. The environment is formalized as a tuple:

    M=(S,A,T,ρ0)\mathcal{M} = (\mathcal{S}, \mathcal{A}, \mathcal{T}, \rho_0)

    where:

    • S\mathcal{S} is a structured semantic state space. Each state is factored as s=(sp,si)s = (s^p, s^i), where sps^p is the positional interface state (e.g., current URL, search/filter settings, pagination, and open modal states) and sis^i is the informational state representing item attributes revealed to the agent (e.g., product prices, ratings, or email bodies rendered on screen).
    • A\mathcal{A} is a typed semantic action space abstracting raw coordinate-level operations (e.g., issuing queries, navigating surfaces, opening detail pages, or submitting commit actions).
    • T:S×A→S\mathcal{T}: \mathcal{S} \times \mathcal{A} \to \mathcal{S} is a deterministic transition function such that st+1=T(st,at)s_{t+1} = \mathcal{T}(s_t, a_t).
    • ρ0\rho_0 is the distribution of initial environment and task configurations.

    In this architecture, the graphical user interface (GUI) is rendered directly as a deterministic projection of the underlying semantic MDP state. When an agent executes low-level GUI actions (clicking, typing, scrolling), a dispatch layer maps task-relevant actions to typed semantic actions at∈Aa_t \in \mathcal{A}, computes the transition st+1=T(st,at)s_{t+1} = \mathcal{T}(s_t, a_t), and re-renders the DOM from the updated state. This guarantees exact recovery of the causal semantic trajectory τ=(s0,a0,s1,a1,…,sT)\tau = (s_0, a_0, s_1, a_1, \dots, s_T) directly from execution.

  2. Knowl 2 — Process-Level Web Agent Evaluation Metrics

    definition

    Given an agent's recorded semantic trajectory τ=(s0,a0,s1,…,sT)\tau = (s_0, a_0, s_1, \dots, s_T) in a deterministic semantic MDP, five process-level metrics quantify behavior beyond terminal binary success:

    1. Exploration Success Rate (Exploration SR): Measures whether the agent identifies the correct target entity e⋆e^\star before attempting a terminal action. If Vc=(v1,…,vk)V_c = (v_1, \dots, v_k) denotes the sequence of unique entities visited on detail pages prior to the first commit action, exploration succeeds if and only if vk=e⋆v_k = e^\star.

    2. Execution Success Rate (Execution SR): The proportion of tasks where the agent completes all terminal verifier conditions, conditioned strictly on successful exploration:

    Execution SR=P(Terminal Success∣Exploration Success)\text{Execution SR} = \mathbb{P}(\text{Terminal Success} \mid \text{Exploration Success})

    1. Informational Coverage (COVtCOV_{t}): The proportion of task-relevant evidence constraints C={c1,…,cK}C = \{c_1, \dots, c_K\} satisfied by information observed up to step tt:

    COVt=∣{c∈C:sat(c,t)}∣∣C∣COV_t = \frac{|\{c \in C : \text{sat}(c, t)\}|}{|C|}

    where sat(c,t)\text{sat}(c, t) is a boolean predicate indicating whether constraint cc has been revealed via on-screen rendering or disjunctive search matching up to step tt. Reported coverage is measured at the first commit step tcommitt_{\text{commit}}.

    1. Step Efficiency: The ratio of total low-level GUI actions to semantic MDP transitions (GUI Steps/Semantic Steps\text{GUI Steps} / \text{Semantic Steps}), measuring how many GUI actions produce meaningful state changes.

    2. Safe Pass Terminal Success Rate: Standard terminal success requires all task verifier conditions on the terminal semantic state to be satisfied. Safe Pass additionally credits trajectories where the agent correctly navigated to the target entity and gathered all necessary information but halted before the final commit action due to safety guardrails (e.g., policies forbidding submission of payment or personal credentials).

  3. Knowl 3 — Procedural Task and World Generation with Controlled Hard Negatives

    model/method

    To evaluate web agent discrimination and exploration under controlled difficulty, each task instance combines a natural language instruction with a synthetically generated world populated with concrete entities. The generation procedure consists of:

    1. Target Entity Creation: Generates a single target entity e⋆e^\star satisfying all constraints specified in a parameterized task template.
    2. Hard Negative Injection: Generates distractor entities that match the target on all surface-level attributes visible on list, card, or search-result pages (e.g., identical title, price range, or category) but differ on attributes visible only on item detail pages (e.g., seller, warranty terms, delivery policy, or body text). The hard negative count HN∈{0,1,2,3}HN \in \{0, 1, 2, 3\} sets a calibrated lower bound on required discrimination effort.
    3. Filler Entity Population: Populates remaining catalog slots with clearly distinguishable filler items to maintain realistic catalog density.
    4. Complexity Categorization: Tasks are structured along three axes:
      • Hard negative count: HN∈{0,1,2,3}HN \in \{0, 1, 2, 3\}.
      • Oracle trajectory length: The minimal number of semantic steps required to complete the task.
      • Information access level: Categorized into Card (solvable using list-level cards alone), Filter (requires search or filter actions but no detail views), and Detail (requires opening item detail views).
  4. Knowl 4 — Semantic MDP Specification Construction from Workflow Traces

    algorithm

    The construction of a deterministic semantic MDP from live website interaction traces follows a multi-stage specification and validation procedure:

    Input: User workflow traces WW containing screenshots and user actions
    Output: Verified executable semantic MDP specification M\mathcal{M}
    Identify domain entities with typed attributes and unique identifiers
    Define entity relations with explicit cardinalities
    Map interface surfaces Σ\Sigma, visible fields per surface, and available actions
    Construct the surface navigation graph and verify surface reachability
    Define state schema S=(sinterface,squery,ssession,srevealed)\mathcal{S} = (s_{\text{interface}}, s_{\text{query}}, s_{\text{session}}, s_{\text{revealed}})
    Specify observation function observe(s,world)→o\text{observe}(s, \text{world}) \to o mapping state to visible attributes
    Define semantic action set A\mathcal{A} with preconditions, state transitions T\mathcal{T}, and reveal rules
    Specify deterministic query functions for filtering, sorting, and pagination
    Define flow guards for sequential action dependencies and stage-locked fields
    Specify derived display values as deterministic functions of state and world data
    Define DOM rendering contract, data-test-id bindings, and UI layout templates
    Define seeded world generator with distribution parameters and consistency invariants
    Perform multi-level validation:
        Run MDP unit tests for valid transitions, invalid action rejection, and visibility
        Run Playwright GUI tests for coordinate clickability, element occlusion, and action binding
        Replay oracle trajectories in headless browser to verify task verifier satisfaction
        Conduct manual QA for residual verifier and world-generation ambiguities
    if any validation check fails then
        Revise the relevant specification stage and rerun validation
    end if
    return compiled executable MDP application
  5. Knowl 5 — Semantic Skill Taxonomy and Invocation Diagnosis

    definition

    Web agent actions are classified into five semantic skill categories:

    • Inspect: Actions that reveal additional information about a specific entity without altering application state (e.g., opening detail views, expanding email threads, viewing user profiles).
    • Navigate: Actions that transition between views or manage navigation history without narrowing candidate entities (e.g., pagination, tab switching, browser back navigation, temporal calendar shifting).
    • Search: Actions that formulate or clear text queries to change the visible candidate set.
    • Filter: Actions that narrow or reorder candidates within the current result set using structured UI controls (e.g., categorical filters, range sliders, sorting criteria).
    • Commit: The task-relative, concluding action that fulfills the goal on the target entity (e.g., placing an order, sending a message, starring a repository, or booking a listing).

    To diagnose agent capability, skill utilization is evaluated via Skill Invocation. Let S⋆S^\star be the set of skill categories present in the minimal oracle trajectory for a given task, and let SS be the set of skill categories executed in the agent's trajectory. For each required skill k∈S⋆k \in S^\star, skill invocation is evaluated as a binary indicator I(k∈S)\mathbb{I}(k \in S), measuring whether the agent recognizes the necessity of invoking skill kk during task execution.

  6. Knowl 6 — Trajectory Bifurcation Analysis at Shared Semantic States

    model/method

    Trajectory bifurcation analysis localizes the exact step where an unsuccessful agent trajectory deviates from task success. Because all agents interact within a deterministic, Markovian semantic environment, trajectories for the same task instance can be aligned at shared semantic states s∈Ss \in \mathcal{S}. The bifurcation point is defined as the last shared semantic state visited before the failing trajectory and a reference successful trajectory diverge.

    Bifurcation points are categorized into three distinct failure types:

    1. Wrong Branch: From the bifurcation state ss, the successful trajectory executes action asucca_{\text{succ}} and the failing trajectory executes action afaila_{\text{fail}}, where asucc≠afaila_{\text{succ}} \neq a_{\text{fail}} and neither action is Commit. The action afaila_{\text{fail}} identifies the exact decision that diverts the agent off the winning path.
    2. Delayed Commit: From state ss, the successful next action is Commit, but the failing agent continues taking non-commit exploration or navigation actions. The suffix of actions taken by the failing agent characterizes redundant or post-target behaviors.
    3. Premature Commit: From state ss, the failing agent executes Commit prematurely, whereas the successful trajectory executes further non-commit actions (such as inspecting or searching). The remaining successful suffix identifies the critical evidence gathering that the failing agent skipped.
  7. Knowl 7 — Aggregate Performance Breakdown Across Web Agent Architectures

    data/table

    Evaluation of six vision-based web agents on the 1,800-task WEBSTEP benchmark demonstrates that terminal success rate obscures substantial underlying differences in exploration reach and execution fidelity:

    Agent Terminal SR (%) Exploration SR (%) Execution SR (%) Coverage (%) GUI Steps GUI/Semantic
    Fara-7B 35.7 47.3 75.0 76.4 18.9 2.3
    GUI-Owl-1.5-8B 34.4 43.3 78.6 79.5 28.3 2.7
    UI-TARS-1.5-7B 37.1 49.2 73.9 80.3 35.0 2.5
    Qwen3.5-122B 58.2 66.3 86.7 89.1 22.1 2.3
    OpenAI CUA (GPT-5.4) 82.7 89.4 92.4 96.8 19.7 2.0
    Claude CUA (Sonnet 4.6) 85.3 91.0 93.7 96.3 14.3 1.5

    While the three smaller specialist models (Fara-7B, GUI-Owl-1.5-8B, and UI-TARS-1.5-7B) achieve clustered terminal success rates between 34.4% and 37.1%, process metrics disentangle distinct operational profiles:

    • UI-TARS achieves higher exploration success (49.2%) and information coverage (80.3%) via broader search and higher step counts, but has lower execution accuracy (73.9%).
    • GUI-Owl has lower exploration reach (43.3%) but higher conditional execution accuracy (78.6%).
    • Fara takes fewer steps (18.9 GUI steps) and achieves the lowest informational coverage (76.4%), indicating under-exploration.
    • Frontier models (OpenAI CUA and Claude CUA) achieve over 89% exploration SR and >92% execution SR, with Claude CUA exhibiting the highest step efficiency (1.5 GUI actions per semantic step).
  8. Knowl 8 — Within-Domain Skill Inversion Between Web Agents

    empirical result

    Decomposing agent performance into semantic skill categories reveals that overall domain scores can conceal opposing skill competencies within the same website environment:

    • On the Coding Q&A site, Claude CUA outperforms OpenAI CUA by 30.0% in Navigate skill invocation (73.0% vs. 43.0%), but underperforms OpenAI CUA by 6.7% in Inspect skill invocation (92.0% vs. 98.9%).
    • On the Housing platform, OpenAI CUA exhibits superior Navigate invocation relative to Claude CUA (+3.9%), but inferior Filter invocation (-16.2%).
    • Among open-weight specialist agents, GUI-Owl achieves an aggregate Filter invocation rate of 94.0% compared to UI-TARS's 77.3% (including a +48.3% gap on Code Repo and +32.0% on Q&A), whereas UI-TARS outperforms GUI-Owl on Navigate invocation (65.2% aggregate vs. 58.1%, with a +32.0% gap on Housing and +21.7% on Code Repo).

    These opposing intra-site skill profiles demonstrate that aggregate website success rates average out distinct behavioral weaknesses.

  9. Knowl 9 — Agent-Specific Divergence Patterns in Trajectory Bifurcations

    empirical result

    Empirical analysis of bifurcation points across agent trajectories shows distinct, model-specific error signatures:

    • First Wrong Branch: While Inspect is the most common initial divergent action across all models (accounting for 33% to 79% of initial mistakes due to opening non-target items), secondary divergence modes differ sharply: GUI-Owl enters failure branches through incorrect Filter actions in 31% of cases, whereas Fara and UI-TARS predominantly diverge by falling back to Search queries (30% of cases each).
    • Delayed Commit Behaviors: When models fail to commit after having gathered sufficient evidence, GUI-Owl and FARA spend 48% of subsequent steps executing redundant Navigate actions. In contrast, OpenAI CUA spends 64% of delayed steps on repeated Inspect actions and 32% on Search query reformulation.
    • Premature Commit Omissions: Across all models, premature commits predominantly omit Inspect (37–47%) and Navigate (38–47%) actions. However, smaller models (FARA, GUI-Owl, UI-TARS) omit required Search actions in 10–16% of premature commits, compared to only 1–3% for OpenAI CUA and Claude CUA, indicating that smaller models frequently fail to recognize when information retrieval beyond the current viewport is needed.
  10. Knowl 10 — Exploration Success Scaling Across Task Complexity Regimes

    data/table

    Evaluating exploration success across hard negative distractor counts (HNHN) and information access levels demonstrates how agent exploration degrades under controlled task complexity:

    Hard Negative Count (Exploration SR %) Information Access Level (Exploration SR %)
    Agent HN=0 HN=1 HN=2 HN=3 Card Filter Detail
    Claude CUA 90.6 98.7 90.3 89.8 95.1 98.2 88.3
    OpenAI CUA 91.7 93.6 87.9 83.5 93.2 96.7 85.8
    Qwen3.5-122B 76.4 78.2 65.2 51.4 70.4 90.1 59.4
    Fara-7B 60.5 64.1 41.9 32.8 58.6 72.8 38.9
    GUI-Owl-1.5-8B 59.4 43.6 34.5 32.3 60.5 65.0 35.2
    UI-TARS-1.5-7B 60.6 53.8 45.6 37.3 61.7 61.4 44.2

    Key trends across complexity axes:

    • As hard negative distractors increase from HN=1HN=1 to HN=3HN=3, exploration success declines monotonically for all agents. Claude CUA and OpenAI CUA retain approximately 90% of their initial exploration performance, whereas open models lose between 26% and 49% of their exploration success.
    • On tasks requiring detail-level inspection (Detail), performance drops sharply compared to list-only tasks (Card and Filter), widening the performance gap between frontier models (Claude CUA: 88.3%, OpenAI CUA: 85.8%) and open models (Fara: 38.9%, GUI-Owl: 35.2%, UI-TARS: 44.2%).
  11. Knowl 11 — Limitations of Deterministic Semantic MDP Web Benchmarking

    limitation

    The construction of web agent evaluation around deterministic semantic MDPs introduces several specific scope limitations:

    1. Semantic Abstraction Gap: Implementing websites as deterministic MDPs omits dynamic live-web behaviors such as asynchronous network loading, session timeouts, live multi-user concurrency, personalized content feeds, third-party authentication/captchas, and external API integrations.
    2. Synthetic Data Distributions: World entities and attribute catalogs are generated through seeded procedural methods, which do not fully replicate real-world phenomena such as long-tail distribution skews, organic attribute correlations, or temporal data drift.
    3. Structured Template Formulation: Task instructions are instantiated from formal parameterized templates with deterministic verifiers, which do not capture the natural language ambiguity, underspecification, or open-ended negotiation found in unconstrained human queries.
    4. Invocation vs. Success Granularity: Skill-level evaluation tracks whether required action types are invoked along the realized trajectory rather than measuring latent intent or per-action success, because partial or failed actions lack objective intent labels from raw GUI traces.
    5. Interaction Modality Scope: The evaluation environment is restricted to coordinate-based clicking, typing, and scrolling, omitting alternative interaction modalities such as drag-and-drop, native keyboard shortcuts, or direct URL navigation.

Coverage note — None was omitted; the knowls cover the benchmark architecture, specification algorithm, metrics, skill taxonomy, bifurcation analysis, experimental results, complexity scalings, and limitations.

References

  1. 1.Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. In Advances in neural information processing systems, 2017.
  2. 2.Anthropic. Claude opus 4.6 system card, 2026. URL https://www-cdn.anthropic.com/14e4fb01875d2a69f646fa5e574dea2b1c0ff7b5.pdf.
  3. 3.Ahmed Awadallah, Yash Lara, Raghav Magazine, Hussein Mozannar, Akshay Nambi, Yash Pandya, Aravind Rajeswaran, Corby Rosset, Alexey Taymanov, Vibhav Vineet, Spencer Whitehead, and Andrew Zhao. Fara-7b: An efficient agentic model for computer use. arXiv preprint arXiv:2511.19663, 2025.
  4. 4.Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sellier De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. WorkArena++: Towards compositional planning and reasoning-based common knowledge work tasks. arXiv preprint arXiv:2407.05291, 2024.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  6. 6.De Chezelles, Thibault Le Sellier, Sahar Omidi Shayegan, Lawrence Keunho Jang, Xing Han Lù, Ori Yoran, Dehan Kong, Frank F Xu, Siva Reddy, Quentin Cappart, et al. The browsergym ecosystem for web agent research. arXiv preprint arXiv:2412.05467, 2024.
  7. 7.Jiwan Chung, Neel Joshi, Pratyusha Sharma, Youngjae Yu, and Vibhav Vineet. What mllms learn about when they learn about multimodal reasoning: Perception, reasoning, or their integration? arXiv preprint arXiv:2510.01719, 2025.
  8. 8.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. In NeurIPS, 2023.
  9. 9.Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H Laradji, Manuel Del Verme, Tom Marber, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. Workarena: How capable are web agents at solving common knowledge work tasks? arXiv preprint arXiv:2403.07718, 2024.
  10. 10.Divyansh Garg, Shaun VanWeelden, Diego Caples, Andis Draguns, Nikil Ravi, Pranav Putta, Naman Garg, Tomas Abraham, Michael Lara, Federico Lopez, et al. Real: Benchmarking autonomous agents on deterministic simulations of real websites. arXiv preprint arXiv:2504.11543, 2025.
  11. 11.Milan Gritta, Debjit Paul, Xiaoguang Li, Lifeng Shang, Jun Wang, and Gerasimos Lampouras. Process evaluation for agentic systems. In Findings of the Association for Computational Linguistics: EACL 2026, pp. 2678–2692, 2026.
  12. 12.Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I Wang. Cruxeval: a benchmark for code reasoning, understanding and execution. In Proceedings of the 41st International Conference on Machine Learning, pp. 16568–16621, 2024.
  13. 13.Tanmay Gupta et al. MolmoWeb: Open visual web agent and open data for the open web. Technical report, Allen Institute for AI, 2026. URL https://allenai.org/papers/molmoweb.
  14. 14.Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. Webvoyager: Building an end-to-end web agent with large multimodal models. arXiv preprint arXiv:2401.13919, 2024.
  15. 15.Rodrigo Toro Icarte, Toryn Klassen, Richard Valenzano, and Sheila McIlraith. Using reward machines for high-level task specification and decomposition in reinforcement learning. In International Conference on Machine Learning, pp. 2107–2116. PMLR, 2018.
  16. 16.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. arXiv preprint arXiv:2401.13649, 2024.
  17. 17.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. ICLR, 2024.
  18. 18.Evan Zheran Liu, Kelvin Guo, Panupong Pasupat, Tianlin Shi, and Percy Liang. Reinforcement learning on web interfaces using workflow-guided exploration. arXiv preprint arXiv:1802.08802, 2018.
  19. 19.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. Agentbench: Evaluating llms as agents. In ICLR, 2024.
  20. 20.Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stanczak, Peter Shaw, Christopher Pal, and Siva Reddy. Agentrewardbench: Evaluating automatic evaluations of web agent trajectories. In Second Conference on Language Modeling, 2025.
  21. 21.Chang Ma, Junlei Zhang, Zhihao Liu, Jiawei Chen, Yitao Wang, et al. AgentBoard: An analytical evaluation board of multi-turn LLM agents. In NeurIPS, 2024.
  22. 22.Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. Gaia: A benchmark for general ai assistants. arXiv preprint arXiv:2311.12983, 2023.
  23. 23.OpenAI. Operator system card, 2025. URL https://cdn.openai.com/operator_system_card.pdf.
  24. 24.Yichen Pan, Dehan Kong, Sida Zhou, Cheng Cui, Yifei Leng, Bing Jiang, Hangyu Liu, Yanyi Shang, Shuyan Zhou, Tongshuang Wu, and Zhengyang Wu. WebCanvas: Benchmarking web agents in online environments. In ICML, 2024.
  25. 25.Eduardo Pignatelli, Johan Ferret, Matthieu Geist, Thomas Mesnard, Hado van Hasselt, and Laura Toni. A survey of temporal credit assignment in deep reinforcement learning. Transactions on Machine Learning Research, 2024.
  26. 26.Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian Li, Haihua Yang, Haifeng Liu, Feng Lin, Tao Peng, Xin Liu, and Guang Shi. UI-TARS: Pioneering automated GUI interaction with native agents. arXiv preprint arXiv:2501.12326, 2025.
  27. 27.Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy P Lillicrap. Androidinthewild: A large-scale dataset for android device control. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.
  28. 28.Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William E Bishop, Wei Li, Folawiyo Campbell-Ajala, et al. Androidworld: A dynamic benchmarking environment for autonomous agents. In The Thirteenth International Conference on Learning Representations, 2025.
  29. 29.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Cote, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. In International Conference on Learning Representations, 2021.
  30. 30.Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction, volume 1. MIT press Cambridge, 1998.
  31. 31.Daniel Toyama, Philippe Hamel, Anita Gergely, Gheorghe Comanici, Amelia Glaese, Zafarali Ahmed, Tyler Jackson, Shibl Mourad, and Doina Precup. Androidenv: A reinforcement learning platform for android. arXiv preprint arXiv:2105.13231, 2021.
  32. 32.Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. Scienceworld: Is your agent smarter than a 5th grader? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 11279–11298, 2022.
  33. 33.Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglong Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He, Deyu Zhou, Pengjun Xie, and Fei Huang. WebWalker: Benchmarking LLMs in web traversal. arXiv preprint arXiv:2501.07572, 2025.
  34. 34.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv preprint arXiv:2404.07972, 2024.
  35. 35.Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. An illusion of progress? assessing the current state of web agents. In COLM, 2025.
  36. 36.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  37. 37.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. NeurIPS, 2022.
  38. 38.Jiabo Ye, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Zhaoqing Zhu, Ziwei Zheng, Feiyu Gao, Junjie Cao, Zhengxi Lu, Jitong Liao, Qi Zheng, Fei Huang, Jingren Zhou, and Ming Yan. Mobile-agent-v3: Fundamental agents for GUI automation. arXiv preprint arXiv:2508.15144, 2025.
  39. 39.Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, and Jonathan Berant. AssistantBench: Can web agents solve realistic and time-consuming tasks? arXiv preprint arXiv:2407.15711, 2024.
  40. 40.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al. Webarena: A realistic web environment for building autonomous agents. In ICLR, 2024.

Citation

MLA
Chung, J., et al. “Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking”. COLM 2026, 2026, http://arxiv.org/abs/2606.15673v2.
APA
Chung, J., Byun, J., Vineet, V., & Kim, S. J. (2026). Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking. COLM 2026. http://arxiv.org/abs/2606.15673v2
Chicago
Chung, J., J. Byun, V. Vineet, and S. J. Kim. 2026. “Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking”. COLM 2026. http://arxiv.org/abs/2606.15673v2.
Harvard
Chung, J. et al. (2026) “Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking”, COLM 2026 [Preprint]. Available at: http://arxiv.org/abs/2606.15673v2.
Vancouver
1. Chung J, Byun J, Vineet V, Kim SJ (2026) Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking. COLM 2026

BibTeX

@article{chung2026where,
  title = {Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking},
  author = {Chung, Jiwan and Byun, JiHyuk and Vineet, Vibhav and Kim, Seon Joo},
  year = {2026},
  journal = {COLM 2026},
  url = {http://arxiv.org/abs/2606.15673v2},
  eprint = {2606.15673}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/