Built independently by an author, for readers. Read the story and support ChapterPal

keyword

semantic MDP traces

Semantic MDP traces are recorded sequences of high-level, structured states and transitions that capture an autonomous agent's interactions within an environment modeled as a Markov Decision Process. Unlike raw, low-level execution logs such as pixel inputs, interface coordinates, or basic input events, a semantic MDP trace reflects the underlying functional states of the system and the abstract actions executed throughout an interaction. By maintaining a background log of these meaningful state transitions, semantic MDP traces enable process-level evaluation, skill decomposition, and precise error localization across multi-step tasks without relying solely on binary terminal success or manual annotation.

1 item

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Jiwan Chung, JiHyuk Byun, Vibhav Vineet, Seon Joo Kim

OrganizationsMicrosoftYonsei University

Why you should read this

Introduces WebStep, an evaluation benchmark that uses automatic semantic state tracking to diagnose step-by-step agent failures, revealing critical skill-specific differences hidden behind identical overall success rates.

Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work, we conduct a process-level analysis of web agents. We introduce WebStep, a benchmark of 1,800 task instances with controlled difficulty and automatic semantic state tracking. Each website exposes a deterministic semantic MDP alongside the GUI: the agent operates on the interface, while the environment records high-level states and transitions in the background, enabling fine-grained analysis without manual annotation. Based on the semantic trajectory, we first show that process metrics reveal differences invisible to outcome evaluation: three agents whose success rates cluster within 34-37% diverge in exploration reach versus execution accuracy. Then, decomposing by skill characterizes the nature of these differences, exposing opposite per-skill rankings hidden within the same website: e.g., on Q&A, Claude CUA outperforms OpenAI CUA by 30% on navigation actions yet underperforms it by 6.7% on inspection, pinpointing a concrete skill to improve even within a domain. Bifurcation analysis further localizes the decisive error that loses the task and shows that this error is agent-specific rather than shared. Finally, these differences widen as tasks grow harder: success rate is similar on easy tasks but separates sharply as exploration becomes more demanding. Our process-level analysis opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved. Project page: this https URL

Added

2026-09-29