Skim: Speculative Execution for Fast and Efficient Web Agents

Mike WongKevin HsiehSuman NathRavi Netravali

article2026arXiv1 citations

Introduces a speculative execution framework for web agents that bypasses expensive multi-step LLM planning by synthesizing destination URLs from structural website patterns, reducing per-task cost by 1.9x and latency by 33.4% with zero accuracy loss.

Listen

Autonomous web agents that browse live websites to answer queries are increasingly used in deep research and enterprise automation, but their operational costs and response times remain major barriers to adoption. Standard architectures run heavy frontier artificial intelligence models and full browser rendering on every step of an interaction loop, incurring costs between twenty and fifty cents and delays between thirty and one hundred twenty seconds per task. The article demonstrates that these heavy overheads stem from treating all execution steps uniformly, even though the vast majority of steps are routine navigation across highly structured and predictable websites.

The article evaluates Skim, a speculative execution framework designed to accelerate web agents by bypassing full browser interaction and expensive frontier models whenever website structure permits. Skim works by creating offline profiles of websites to record common link structures, search behaviors, and expected data layouts. At runtime, when a user submits a query, the system attempts a fast path by predicting the direct destination link, fetching page content simply, and extracting the answer using a smaller, lower-cost model. A lightweight, two-stage checking mechanism verifies the speculative answer; if the check fails, the task smoothly escalates to the standard full-weight agent, resuming directly from the speculative link rather than restarting from scratch.

The researchers evaluated Skim across more than three hundred web tasks using standard benchmarks and three distinct agent backends. The findings demonstrate that Skim cuts median task latency by 33.4% and reduces median per-task operating costs by 1.9 times with no loss in overall accuracy. When configured in an aggregation mode that reinvests these cost savings into running multiple parallel fast-path attempts, overall task accuracy improved by 4.2 to 16.7 percentage points within the baseline cost budget. Furthermore, failed fast paths still delivered substantial value because their synthesized links provided effective starting points that significantly reduced recovery times for the fallback agent.

These results show that organizations can achieve major performance and cost improvements without sacrificing reliability or modifying their core agent logic. By replacing uniform execution loops with structure-aware routing and verification, systems can operate at significantly higher query volumes within existing infrastructure budgets. Organizations deploying web agents should implement speculative routing for read-dominant workflows, using lightweight verifiers biased toward quick escalation to prevent inaccurate data from reaching end users.

Decision-makers should note that the current evaluation focuses primarily on information retrieval, search, and data extraction tasks. The framework is not yet designed to handle state-altering actions, such as executing financial transactions or submitting binding forms, which carry irreversible side effects and require dedicated transactional safeguards. Confidence in the reported performance gains for read-heavy workloads is high given the consistent improvements demonstrated across multiple agent backends and diverse websites.

arXiv: 2605.16565
Cover for Skim: Speculative Execution for Fast and Efficient Web Agents

Abstract

Skim is a speculative execution framework for web agents that exploits the predictable structure of purpose-built websites. Today's web-agent expense is not intrinsic to the tasks but a property of how agents are composed: frontier-model inference, browser rendering, and ReAct-style planning are applied to every step of every task regardless of complexity. Skim's key observation is that websites enforce stable URL patterns, answer formats, and task-to-trajectory mappings across queries of the same type, so most queries can bypass these heavyweight components entirely. An offline profiler captures these patterns once per site. At runtime, Skim matches each query to a template, synthesizes the destination URL, and extracts the answer with a small model. A lightweight verifier gates each fast-path output against the query and schema; rare misspeculations cascade to the full agent, warm-started by the fast path's final URL to preserve upstream trajectory progress. Across standard web-agent benchmarks paired with three backboneagents (WebVoyager, AgentOccam, BrowserUse), Skim reduces median per-task cost by 1.9x and latency by 33.4% with no accuracy loss.

Table of Contents

  • 1 Introduction
  • 2 Background and Motivation
  • 2.1 Overview of web agents
  • 2.2 Opportunities for specialization
  • 2.3 Challenges
  • 3 Design of Skim
  • 3.1 Overview
  • 3.2 Offline Profiling Pipeline
  • 3.3 Runtime Speculation
  • 3.4 Query Support and Deployment Discussion
  • 4 Implementation
  • 5 Evaluation
  • 5.1 Methodology
  • 5.2 Main Results
  • 5.3 Detailed Analysis
  • 6 Related Work
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Skim Speculative Execution Framework for Web Agents

    model/method

    Skim is a speculative execution framework that accelerates ReAct-style web agents by bypassing the standard loop of frontier LLM reasoning, headless browser rendering, and multi-step action planning for predictable tasks on purpose-built websites.

    Skim interposes between the user and an underlying web agent. Prior to runtime, an automated offline pipeline profiles reachable websites once to extract parameterized URL templates, search semantics, answer schemas, and capability metadata. When an incoming natural-language task arrives at runtime, Skim predicts the minimal execution resources required and synthesizes a direct destination URL using the site profile. It then executes a speculative fast path by fetching the target page via the cheapest viable method (such as a plain HTTP request), stripping irrelevant HTML based on the profile's answer schema, and extracting the candidate answer using a lightweight local model (e.g., Qwen2.5-14B). The candidate result is evaluated by a lightweight verifier; if accepted, it is returned immediately, whereas upon rejection, the task cascades to the full ReAct agent warm-started at the synthesized URL to preserve navigational progress.

  2. Knowl 2 — Offline Site Profiling and Structural Drift Handling

    model/method

    The offline profiling pipeline constructs a reusable structural site profile once per website to capture slow-evolving website scaffolding.

    The pipeline profiles a site by executing lightweight HTTP probes to evaluate response latency, visible content density, accessibility, and bot-detection behavior, escalating to a headless browser if HTTP responses return incomplete or client-side rendered content. It analyzes form schemas and common search URL conventions to identify search endpoints, navigates result lists to determine pagination behavior, and validates parameterized URL templates against real web interactions. A local language model (such as Qwen2.5-14B-Instruct) analyzes discovered endpoints and sample task queries to generate query-construction rules, representative templates, and expected answer schemas.

    To manage structural drift over time, Skim continuously monitors runtime signals, including schema verification failures, missing mandatory fields, unexpected page structures, and persistent HTTP 404 errors. When these failure signals exceed a defined threshold, Skim automatically re-triggers the offline probing pipeline to regenerate the affected profile components, amortizing the regeneration cost across future queries.

  3. Knowl 3 — Template-Driven URL Synthesis and Regex-Constrained Slot Filling

    model/method

    URL synthesis converts natural language task instructions into direct destination URLs using offline profile templates, collapsing multi-step navigation (e.g., search, filter, and sort) into a single retrieval.

    The synthesis procedure operates in two stages:

    1. Template Selection: A lightweight non-frontier language model matches the user query to a candidate template class (e.g., direct identifier lookup, filtered search, sorted retrieval, or paginated traversal).
    2. Slot Extraction and Validation: The model extracts query parameters (keywords, numeric ranges, sort orders) from the task text. Extracted parameters are validated and formatted using regular expressions and conversion rules defined in the profile (for example, parsing a price constraint such as "$100" into an integer cent value for a query parameter).

    Because the vocabulary of plausible user intents is tightly bounded by a website's specific domain and purpose, a lightweight non-frontier model constrained by typed regular expressions achieves sufficient parameterization accuracy without requiring a frontier model.

  4. Knowl 4 — Multi-Axis Resource Determination and Warm-Start Cascading

    model/method

    Skim configures runtime task execution along three orthogonal resource axes:

    • Page acquisition: Direct URL fetch synthesized from a profile template vs. multi-step ReAct browser navigation.
    • Page rendering: Lightweight HTTP-only fetch vs. headless browser execution (required for pages with dynamic client-side JavaScript or authenticated states).
    • Reasoning model: Non-frontier extraction model (e.g., Qwen2.5-14B) vs. frontier reasoning model (e.g., GPT-4o).

    Skim initially executes the task at the cheapest tier justified by the site profile. If the speculative result fails verification, Skim escalates along the specific axis indicated by the failure: an empty or blocked page escalates rendering from HTTP to a browser, a missing or malformed extraction field escalates the extraction model, and repeated failures escalate page acquisition to the full ReAct agent.

    During escalation to the full ReAct agent, the agent is initialized at the final URL reached by the speculative fast path rather than the homepage (a warm start). Because web tasks share common navigational prefixes (e.g., search, category navigation, sorting), the warm-start URL preserves partial navigational progress even when the fast path's extracted answer is rejected.

  5. Knowl 5 — Two-Stage Fast-Path Verification

    model/method

    Skim gates candidate outputs from speculative fast paths using a two-stage verification mechanism designed to prevent incorrect answers from committing while keeping verification overhead low:

    1. Schema Check: A rule-based filter that verifies whether the candidate answer is non-empty, type-compatible with the task requirements, and within expected value ranges specified in the site profile. This check eliminates empty results, bot-detection responses, and type mismatches without incurring model inference costs.
    2. Semantic Judge: A lightweight non-frontier model evaluates candidates that pass the schema check. The model inspects a compressed state summary containing the task description, the current URL, a summary of the schema-cleaned HTML, and the candidate answer to determine task consistency.

    The verifier is biased toward rejection: ambiguous, weakly supported, or borderline speculative outputs are escalated to higher execution tiers rather than committed.

  6. Knowl 6 — Deployment Policies: Accelerate Mode vs. Aggregate Mode

    model/method

    Skim supports two deployment modes based on its speculative execution primitive:

    • Accelerate Mode: Focuses on latency and cost reduction. When a speculative fast-path output passes two-stage verification, Skim immediately commits and returns the answer to the user, bypassing full ReAct execution.
    • Aggregate Mode: Focuses on task accuracy. The compute and cost savings achieved by speculative execution are reinvested into running multiple parallel speculative trials per task within the cost budget of a single full ReAct execution. The verifier scores and ranks candidate outputs across the diverse trajectories to select the highest-confidence answer.
  7. Knowl 7 — End-to-End Latency, Cost, and Accuracy of Skim on Web Benchmarks

    empirical result

    On an evaluation spanning over 300 tasks randomly sampled from the WebVoyager (15 live websites) and WebShop benchmarks using GPT-4o-backed agents (WebVoyager, AgentOccam, and BrowserUse), Skim operating in accelerate mode reduces median per-task execution cost by 1.9×1.9\times and median per-task latency by 33.4%33.4\%.

    Agent Backend Skim Accuracy Default Agent Accuracy
    WebVoyager 40.6% 37.6%
    AgentOccam 52.0% 49.6%
    BrowserUse 45.6% 45.0%

    The table compares the task success rate of Skim in accelerate mode against default ReAct execution across the three agent backends. Skim matches or slightly exceeds the accuracy of default agents (+3.0+3.0 percentage points on WebVoyager, +2.4+2.4 percentage points on AgentOccam, and +0.6+0.6 percentage points on BrowserUse) because template-grounded fast paths avoid compounding navigation errors, while verification failures fall back to full ReAct execution.

  8. Knowl 8 — Accuracy Improvement via Multi-Trial Speculation in Aggregate Mode

    empirical result

    In aggregate mode on the WebVoyager benchmark with AgentOccam, Skim reinvests its cost savings to execute an average of 4 additional speculative trials per task within the cost envelope of a single baseline ReAct execution.

    Aggregating the resulting multi-trial candidate answers produces the following accuracy changes relative to single-execution ReAct:

    • Majority voting across trials improves end-to-end task accuracy by 4.24.2 percentage points.
    • Oracle selection (upper bound selecting the best candidate among generated trials) improves end-to-end task accuracy by 16.716.7 percentage points.
  9. Knowl 9 — Latency Decomposition and Warm-Start Fallback Savings

    empirical result

    Stage-level latency breakdowns for tasks completing on Skim's speculative fast path show that systems-level retrieval operations are substantially faster than semantic reasoning operations:

    • HTTP fetch operations typically complete within 100–300 ms100\text{--}300\text{ ms}.
    • HTML cleaning executes in under 100 ms100\text{ ms}.
    • Two-stage verification adds approximately 1 s1\text{ s} of overhead.
    • Semantic routing and capability prediction require ∼3 s\sim 3\text{ s} per task.
    • URL synthesis requires ∼2–3 s\sim 2\text{--}3\text{ s} for most queries (with a tail reaching ∼9 s\sim 9\text{ s}).
    • Lightweight local extraction requires <1 s< 1\text{ s} for direct lookups and 8–14 s8\text{--}14\text{ s} for synthesis-heavy extraction.

    Across evaluated benchmarks, 12.6%–45.3%12.6\%\text{--}45.3\% of tasks complete entirely on the speculative fast path. For cascaded tasks, initializing ReAct execution from the warm-start URL reached during the speculative attempt significantly compresses the tail of execution latency compared to cold-start ReAct initialized from the homepage.

  10. Knowl 10 — Performance and Efficiency of Lightweight Verification vs. Full-DOM Frontier Judge

    empirical result

    Compared to a frontier-model verifier (GPT-4o) scoring candidate answers against full page DOM trees, Skim's two-stage lightweight verifier (rule-based schema check followed by Qwen2.5-14B evaluating schema-cleaned summaries) is 11.5×11.5\times cheaper per call while achieving:

    • Precision: 82.0%82.0\%
    • Recall: 86.2%86.2\%
    • F1F_1 Score: 0.840.84
    • Accuracy: 86.9%86.9\%

    Verifier reliability is highest on structured pages with consistent schemas. On pages with heavy client-side dynamic rendering or unstructured content, verifier uncertainty triggers execution cascades to the full ReAct agent rather than committing unverified speculative answers.

  11. Knowl 11 — Read-Dominant Scope and Stateful Action Limitation

    limitation

    Skim is designed for read-dominant web workloads, including search, catalog retrieval, information comparison, and structured extraction. It does not support end-to-end speculative execution for state-mutating actions such as online purchases, form submissions with external side effects, or account modifications, which require transactional safety, authentication management, and side-effect rollback mechanisms. For workflows that conclude with state mutations, Skim can accelerate the exploratory read-only navigational prefix before delegating the final stateful action to a standard browser agent.

Coverage note — None was omitted; all key architectural components, offline profiling mechanisms, URL synthesis and verification algorithms, experimental results across benchmarks and agent backends, and stated limitations are covered.

References

  1. 1.Song Bian, Minghao Yan, Anand Jayarajan, Gennady Pekhimenko, and Shivaram Venkataraman. What limits agentic systems efficiency? In SEA @ NeurIPS 2025 Workshop, 2025.
  2. 2.Islem Bouzenia and Michael Pradel. You name it, i run it: An llm agent to execute tests of arbitrary projects. ISSTA 2025, 2024.
  3. 3.browser-use. browser-use. https://github.com/browser-use/browser-use, 2026. Open-source browser agent framework. Accessed 2026-05-14.
  4. 4.Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023.
  5. 5.Lingjiao Chen, Matei Zaharia, and James Zou. Frugalgpt: How to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176, 2023.
  6. 6.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  7. 7.Masafumi Enomoto, Ryoma Obara, Haochen Zhang, and Masafumi Oyamada. Read more, think more: Revisiting observation reduction for web agents. arXiv preprint arXiv:2604.01535, 2026.
  8. 8.Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for gui agents. In International Conference on Learning Representations (ICLR), 2025.
  9. 9.Yilin Guan, Qingfeng Lan, Sun Fei, Dujian Ding, Devang Acharya, Chi Wang, William Yang Wang, and Wenyue Hua. Dynamic speculative agent planning. arXiv preprint arXiv:2509.01920, 2025.
  10. 10.Tanmay Gupta, Piper Wolters, Zixian Ma, Peter Sushko, Rock Yuren Pang, Diego Llanes, Yue Yang, Taira Anderson, Boyuan Zheng, Zhongzheng Ren, Harsh Trivedi, Taylor Blanton, Caleb Ouellette, Winson Han, Ali Farhadi, and Ranjay Krishna. Molmoweb: Open visual web agent and open data for the open web. arXiv preprint arXiv:2604.08516, 2026.
  11. 11.Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. A real-world webagent with planning, long context understanding, and program synthesis. In International Conference on Learning Representations (ICLR), 2024.
  12. 12.Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust. Understanding html with large language models. arXiv preprint arXiv:2210.03945, 2022.
  13. 13.Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. Webvoyager: Building an end-to-end web agent with large multimodal models. In arXiv preprint arXiv:2401.13919, 2024.
  14. 14.Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. Odysseys: Benchmarking web agents on realistic long horizon tasks. arXiv preprint arXiv:2604.24964, 2026.
  15. 15.Sheng Jia, Jamie Kiros, and Jimmy Ba. Dom-q-net: Grounded rl on structured language. In International Conference on Learning Representations (ICLR), 2019.
  16. 16.Nicholas Kushmerick. Wrapper induction: Efficiency and expressiveness. Artificial Intelligence, 118(1-2):15–68, 2000.
  17. 17.Donguk Kwon and Dongha Lee. Region4web: Rethinking observation space granularity for web agents. arXiv preprint arXiv:2605.07134, 2026.
  18. 18.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP ’23, page 611–626, New York, NY, USA, 2023. Association for Computing Machinery.
  19. 19.J. Liang. Caesar: Deep agentic web exploration for creative answer synthesis. arXiv, 2026.
  20. 20.Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Weixian Lei, Lijuan Wang, and Mike Zheng Shou. Showui: One vision-language-action model for gui visual agent. arXiv preprint arXiv:2411.17465, 2024.
  21. 21.Xing Han Lù, Zdeněk Kasner, and Siva Reddy. Weblinx: Real-world website navigation with multi-turn dialogue. arXiv preprint arXiv:2402.05930, 2024.
  22. 22.Zijian Lu, Yiping Zuo, Yupeng Nie, Xin He, Weibei Fan, Lianyong Qi, and Shi Jin. Contractskill: Repairable contract-based skills for multimodal web agents. arXiv preprint arXiv:2603.20340, 2026.
  23. 23.Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. Gaia: a benchmark for general ai assistants. arXiv preprint arXiv:2311.12983, 2023.
  24. 24.Niels Mündler, Mark Niklas Müller, Jingxuan He, and Martin Vechev. Swt-bench: Testing and validating real-world bug-fixes with code agents. NeurIPS, 2024.
  25. 25.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. Webgpt: Browser-assisted question-answering with human feedback. In arXiv preprint arXiv:2112.09332, 2021.
  26. 26.Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina Toutanova. From pixels to ui actions: Learning to follow instructions via graphical user interfaces. arXiv preprint arXiv:2306.00245, 2023.
  27. 27.Zhaoyang Wang, Qianhui Wu, Xuchao Zhang, Chaoyun Zhang, Wenlin Yao, Fazle Elahi Faisal, Baolin Peng, Si Qin, Suman Nath, Qingwei Lin, Chetan Bansal, Dongmei Zhang, Saravan Rajmohan, Jianfeng Gao, and Huaxiu Yao. Webxskill: Skill learning for autonomous web agents. arXiv preprint arXiv:2604.13318, 2026.
  28. 28.Yong Wu, Yanzhao Zheng, Tianze Xu, ZhenTao Zhang, YuanQiang Yu, JiHuai Zhu, Chao Ma, BinBin Lin, Baohua Dong, Hangcheng Zhu, Ruohui Huang, and Gang Yu. Contextbudget: Budget-aware context management for long-horizon search agents. arXiv preprint arXiv:2604.01664, 2026.
  29. 29.Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, and Huzefa Rangwala. Agentoccam: A simple yet strong baseline for llm-based web agents. arXiv preprint arXiv:2410.13825, 2024.
  30. 30.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  31. 31.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  32. 32.Naimeng Ye, Arnav Ahuja, Georgios Liargkovas, Yunan Lu, Kostis Kaffes, and Tianyi Peng. Speculative actions: A lossless framework for faster agentic systems. arXiv preprint arXiv:2510.04371, 2025.
  33. 33.Jiayuan Zhang, Kaiquan Chen, Zhihao Lu, Enshen Zhou, Qian Yu, and Jing Zhang. Prune4web: Dom tree pruning programming for web agent. In Proceedings of the AAAI Conference on Artificial Intelligence, 2026.
  34. 34.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded. arXiv preprint arXiv:2401.01614, 2024.
  35. 35.Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents. In arXiv preprint arXiv:2307.13854, 2023.

Citation

MLA
Wong, M., et al. “Skim: Speculative Execution for Fast and Efficient Web Agents”. arXiv, 2026, https://doi.org/10.48550/arxiv.2605.16565.
APA
Wong, M., Hsieh, K., Nath, S., & Netravali, R. (2026). Skim: Speculative Execution for Fast and Efficient Web Agents. arXiv. https://doi.org/10.48550/arxiv.2605.16565
Chicago
Wong, M., K. Hsieh, S. Nath, and R. Netravali. 2026. “Skim: Speculative Execution for Fast and Efficient Web Agents”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2605.16565.
Harvard
Wong, M. et al. (2026) “Skim: Speculative Execution for Fast and Efficient Web Agents”. arXiv. Available at: https://doi.org/10.48550/arxiv.2605.16565.
Vancouver
1. Wong M, Hsieh K, Nath S, Netravali R (2026) Skim: Speculative Execution for Fast and Efficient Web Agents. https://doi.org/10.48550/arxiv.2605.16565

BibTeX

@misc{https://doi.org/10.48550/arxiv.2605.16565,
  doi = {10.48550/ARXIV.2605.16565},
  url = {https://arxiv.org/abs/2605.16565},
  author = {Wong, Mike and Hsieh, Kevin and Nath, Suman and Netravali, Ravi},
  keywords = {Artificial Intelligence (cs.AI), Operating Systems (cs.OS), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Skim: Speculative Execution for Fast and Efficient Web Agents},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/