Skim: Speculative Execution for Fast and Efficient Web Agents
Mike WongKevin HsiehSuman NathRavi Netravali
Introduces a speculative execution framework for web agents that bypasses expensive multi-step LLM planning by synthesizing destination URLs from structural website patterns, reducing per-task cost by 1.9x and latency by 33.4% with zero accuracy loss.
Autonomous web agents that browse live websites to answer queries are increasingly used in deep research and enterprise automation, but their operational costs and response times remain major barriers to adoption. Standard architectures run heavy frontier artificial intelligence models and full browser rendering on every step of an interaction loop, incurring costs between twenty and fifty cents and delays between thirty and one hundred twenty seconds per task. The article demonstrates that these heavy overheads stem from treating all execution steps uniformly, even though the vast majority of steps are routine navigation across highly structured and predictable websites.
The article evaluates Skim, a speculative execution framework designed to accelerate web agents by bypassing full browser interaction and expensive frontier models whenever website structure permits. Skim works by creating offline profiles of websites to record common link structures, search behaviors, and expected data layouts. At runtime, when a user submits a query, the system attempts a fast path by predicting the direct destination link, fetching page content simply, and extracting the answer using a smaller, lower-cost model. A lightweight, two-stage checking mechanism verifies the speculative answer; if the check fails, the task smoothly escalates to the standard full-weight agent, resuming directly from the speculative link rather than restarting from scratch.
The researchers evaluated Skim across more than three hundred web tasks using standard benchmarks and three distinct agent backends. The findings demonstrate that Skim cuts median task latency by 33.4% and reduces median per-task operating costs by 1.9 times with no loss in overall accuracy. When configured in an aggregation mode that reinvests these cost savings into running multiple parallel fast-path attempts, overall task accuracy improved by 4.2 to 16.7 percentage points within the baseline cost budget. Furthermore, failed fast paths still delivered substantial value because their synthesized links provided effective starting points that significantly reduced recovery times for the fallback agent.
These results show that organizations can achieve major performance and cost improvements without sacrificing reliability or modifying their core agent logic. By replacing uniform execution loops with structure-aware routing and verification, systems can operate at significantly higher query volumes within existing infrastructure budgets. Organizations deploying web agents should implement speculative routing for read-dominant workflows, using lightweight verifiers biased toward quick escalation to prevent inaccurate data from reaching end users.
Decision-makers should note that the current evaluation focuses primarily on information retrieval, search, and data extraction tasks. The framework is not yet designed to handle state-altering actions, such as executing financial transactions or submitting binding forms, which carry irreversible side effects and require dedicated transactional safeguards. Confidence in the reported performance gains for read-heavy workloads is high given the consistent improvements demonstrated across multiple agent backends and diverse websites.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). Introduces WebVoyager, a multimodal agent and benchmark used directly as an evaluated backbone baseline in Skim's experiments.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Establishes how reusable procedural workflows can be extracted and cached from web navigation tasks, providing foundation for Skim's offline trajectory profiling.
- Paper: Fast Inference from Transformers via Speculative Decoding, Yaniv Leviathan et al. (2023). Defines the core principles of speculative drafting and parallel verification that Skim adapts from token generation to end-to-end web agent trajectories.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). Presents WebArena, the foundational realistic web environment and benchmark suite that grounds evaluation of web navigation agents.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Introduces Mind2Web, establishing standard task formulations and multi-step interaction paradigms across diverse websites that Skim targets for acceleration.
- Paper: SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, Xupeng Miao et al. (2024). Explores tree-based speculative inference and draft verification, underpinning Skim's architecture of fast-path generation gated by verification and fallback.
- Paper: SGLang: Efficient Execution of Structured Language Model Programs, Lianmin Zheng et al. (2023). Provides runtime optimization frameworks and speculative execution concepts for multi-step structured language model programs.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). Surveys system-wide efficiency strategies across agent memory, tool use, and planning, contextualizing Skim's trajectory-level speculative shortcuts within the broader efficiency taxonomy.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). Extends procedural trajectory extraction and execution reuse from single-agent web navigation to modular multi-agent workflow automation.
