WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Hongliang HeWenlin YaoKaixin MaWenhao YuYong DaiHongming ZhangZhenzhong LanDong Yu

article2024ACL452 citations

Presents WebVoyager, an end-to-end multimodal web agent that interacts directly with live websites using visual and textual cues, outperforming text-only baselines alongside a benchmark of real-world tasks and an automated evaluation protocol.

Listen

Autonomous web agents have significant potential to automate complex online tasks, yet existing systems rely primarily on simplified simulators or raw website code. These text-only approaches ignore the visual layout and intuitive design of modern websites, making it difficult for automated tools to handle dynamic online interfaces. To address this limitation, the article introduces and evaluates WebVoyager, an end-to-end multimodal agent designed to navigate live, real-world websites autonomously by combining visual screenshots with textual interface elements.

The research evaluated the agent across an automated browsing environment using a new benchmark of 643 real-world tasks across 15 popular websites, including Amazon, Booking.com, and Google Flights. The agent identified interactive components by overlaying visual tags onto webpage screenshots, allowing it to reason through decisions and execute human-like actions such as clicking, typing, and scrolling. In addition to testing against standard baselines, the article introduced an automated evaluation method powered by a multimodal model to score navigation recordings, validating these scores against independent human evaluations.

The findings show that WebVoyager achieved a 59.1% task success rate, substantially outperforming both a text-only setup at 40.1% and a leading commercial integrated tool at 30.8%. Visual perception proved particularly decisive on complex interfaces; for instance, on booking and flight platforms that require calendar interactions, the visual agent succeeded where text-only agents failed. Conversely, text-only inputs performed slightly better on text-heavy websites where small text was harder to resolve from screenshots alone. Additionally, the automated scoring framework demonstrated an 85.3% agreement rate with human judges, confirming that vision-language models can reliably evaluate multi-step web navigation without constant manual oversight.

These results indicate that effective digital assistants require both visual and textual understanding to navigate modern web interfaces reliably. When the agent failed, the primary bottlenecks were getting stuck in repetitive navigation loops (44.4% of errors), visual misidentification of closely grouped elements (24.8%), and generating incomplete or hallucinated responses (21.8%). Organizations building or adopting web automation should focus future development on hybrid inputs that extract clean text alongside visual snapshots to handle dense content, while refining navigation prompts to prevent looping.

While confidence in the reported performance gains is supported by solid human validation, leaders should account for current operational boundaries. The agent does not support complex input actions like dragging, cannot parse video media, and was evaluated exclusively on public, non-login workflows without security verifications. Before deploying autonomous web agents into live operational settings, organizations must implement robust safety safeguards to prevent unintended actions, such as submitting unauthorized information or interacting with insecure sites.

Cover for WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Abstract

The rapid advancement of large language models (LLMs) has led to a new era marked by the development of autonomous applications in real-world scenarios, which drives innovation in creating advanced web agents. Existing web agents typically only handle one input modality and are evaluated only in simplified web simulators or static web snapshots, greatly limiting their applicability in real-world scenarios. To bridge this gap, we introduce WebVoyager, an innovative Large Multimodal Model (LMM) powered web agent that can complete user instructions end-to-end by interacting with real-world websites. Moreover, we establish a new benchmark by compiling real-world tasks from 15 popular websites and introduce an automatic evaluation protocol leveraging multimodal understanding abilities of GPT-4V to evaluate open-ended web agents. We show that WebVoyager achieves a 59.1% task success rate on our benchmark, significantly surpassing the performance of both GPT-4 (All Tools) and the WebVoyager (text-only) setups, underscoring the exceptional capability of WebVoyager. The proposed automatic evaluation metric achieves 85.3% agreement with human judgment, indicating its effectiveness in providing reliable and accurate assessments of web agents.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 WebVoyager
  • 3.1 Browsing Environment
  • 3.2 Interaction Formulation
  • 3.3 Observation Space
  • 3.4 Action Space
  • 4 Benchmark for WebVoyager
  • 4.1 Website Selection
  • 4.2 Data Construction
  • 4.3 Annotation Process
  • 5 Experiment
  • 5.1 Evaluation Methods
  • 5.2 Result
  • 5.3 Discussions
  • 5.4 Error Analysis
  • 6 Conclusion
  • References
  • A Prompt for WebVoyager
  • B Prompt for Auto Evaluation
  • C Action Space
  • D Additional Trajectories
  • E Additional Related Work
  • F Error Cases

Knowls

  1. Knowl 1 — WebVoyager Autonomous Multimodal Web Agent Architecture

    model/method

    WebVoyager is an autonomous, end-to-end multimodal web browsing agent powered by a Large Multimodal Model (LMM). Given a natural language user instruction II, the agent interacts with live websites via Selenium without intermediate human intervention.

    At time step tt, the context ctc_t is constructed from the user instruction II, historical actions aia_i, and observations oio_i: ct=(o1,a1,…,ot−1,at−1,ot,I)c_t = (o_1, a_1, \dots, o_{t-1}, a_{t-1}, o_t, I)

    Following the ReAct paradigm, the model generates a step output at=(st,a^t)=M(ct)a_t = (s_t, \hat{a}_t) = \mathcal{M}(c_t), where sts_t is a natural language thought process and a^t\hat{a}_t is an executable action code. The environment E\mathcal{E} executes a^t\hat{a}_t to return the next observation ot+1=E(ot,a^t)o_{t+1} = \mathcal{E}(o_t, \hat{a}_t). The cycle repeats until the agent outputs a termination action or exceeds the maximum budget of 15 steps.

    To prevent context overflow and visual confusion across long browsing episodes, WebVoyager implements context clipping: only the three most recent webpage observations (ot−2,ot−1,ot)(o_{t-2}, o_{t-1}, o_t) are retained in the multimodal prompt, while the complete textual history of thoughts and actions is preserved. If an action raises an execution exception in Selenium, the error message is appended to the prompt for a retry, which consumes one step from the exploration budget.

  2. Knowl 2 — Set-of-Mark Web Observation and Visual Grounding in WebVoyager

    model/method

    Rather than relying solely on verbose HTML DOM trees or accessibility trees, WebVoyager processes rendered webpage screenshots paired with auxiliary textual metadata.

    To ground agent decisions without training an external object detector, WebVoyager uses a rule-based JavaScript tool (GPT-4V-ACT) to identify interactive elements based on HTML tag types. The tool overlays bounding boxes with unique numerical index tags in the top-left corner of each element directly onto the screenshot. Black borders and black label backgrounds are used uniformly, which empirically yields higher action success rates than multi-colored labeling.

    At step tt, the observation oto_t consists of:

    1. A rendered screenshot of the viewport at a standardized resolution of 1024×7681024 \times 768 pixels with black numerical Set-of-Mark labels.
    2. Auxiliary text comprising the interactable element's HTML type, embedded inner text, and aria-label accessibility text.
    3. Text extracted from downloaded PDF files parsed automatically via the OpenAI Assistant API when clicking a document link.

    Interactions are restricted to a single browser tab to simplify the observation space.

  3. Knowl 3 — WebVoyager Browser Action Space

    definition

    WebVoyager operates on live web pages using a discrete set of keyboard and mouse action primitives formatted with numerical element IDs:

    1. Click [Numerical_Label]: Clicks on the specified webpage element (e.g., link or button). If the click initiates a PDF download, the PDF content is parsed and incorporated into the observation.
    2. Type [Numerical_Label]; [Content]: A composite action that focuses on the text box indexed by Numerical_Label, deletes any pre-existing text, inputs Content, and automatically executes an ENTER keystroke.
    3. Scroll [Numerical_Label or WINDOW]; [up or down]: Scrolls the entire browser window or a specific scrollable element container vertically up or down.
    4. Wait: Pauses execution to wait for dynamic webpage assets, AJAX requests, or rendering scripts to finish loading.
    5. GoBack: Navigates to the previous URL in the browser history.
    6. Google: Jumps directly to Google Search to start a new search trajectory when navigation is blocked or unproductive.
    7. ANSWER; [Content]: Concludes the interaction episode and returns the final response Content to the user.
  4. Knowl 4 — WebVoyager Benchmark for Real-World Web Navigation

    experimental setup

    The WebVoyager benchmark evaluates end-to-end autonomous web browsing on live, real-world websites. It comprises 643 tasks spanning 15 popular websites: Allrecipes, Amazon, Apple, ArXiv, BBC News, Booking, Cambridge Dictionary, Coursera, ESPN, GitHub, Google Flights, Google Map, Google Search, Huggingface, and Wolfram Alpha.

    The benchmark dataset was constructed via a three-step semi-automated process:

    1. Initial seed tasks were manually sampled and rewritten from Mind2Web for 5 websites.
    2. In-context self-instruct generation with GPT-4 Turbo produced candidate tasks across all 15 websites, followed by manual verification to confirm that each task was solvable on the target site.
    3. Iterative task pool expansion generated 40+ unique tasks per website. Task diversity was confirmed using the all-mpnet-base-v2 sentence embedding model: 99.68%99.68\% of all 206,403206,403 task pairs exhibited pairwise cosine similarity under 0.600.60, and only 4949 pairs exceeded 0.800.80.

    Tasks are categorized by answer evaluation type:

    • Golden Responses (22.3% of tasks): Tasks with fixed, stable, and exhaustively verifiable ground-truth answers.
    • Possible Responses (77.7% of tasks): Tasks involving open-ended summaries, multiple valid answers, or real-time information (e.g., live flight prices or hotel availability).
  5. Knowl 5 — GPT-4V Multimodal Trajectory Auto-Evaluation Protocol

    model/method

    To automatically evaluate open-ended web agent navigation trajectories without predefined golden action paths, WebVoyager introduces an LMM-based evaluation protocol using GPT-4V.

    The auto-evaluator receives three components:

    1. The natural language web task instruction.
    2. The sequence of kk screenshots from the interaction episode (where k∈{1,2,3}k \in \{1, 2, 3\} or k=Fullk = \text{Full} for the entire trajectory).
    3. The textual final response produced by the web agent.

    With decoding temperature set to 00, GPT-4V evaluates whether the visual states and final response satisfy all requirements in the instruction, outputting either SUCCESS or NOT SUCCESS. When supplied with full trajectory screenshots (k=Fullk = \text{Full}), the auto-evaluator achieves an 85.3%85.3\% agreement and a Cohen's kappa κ=0.70\kappa = 0.70 with consolidated human expert judgments, matching human-to-human inter-annotator agreement (Fleiss's κ=0.70\kappa = 0.70).

  6. Knowl 6 — Task Success Rate on the WebVoyager Benchmark

    data/table

    The table below compares the Task Success Rate across 15 websites on the 643-task WebVoyager benchmark. Evaluated systems include GPT-4 (All Tools), a text-only WebVoyager baseline using accessibility trees, and multimodal WebVoyager across multiple LMM backbones. Human evaluations and GPT-4V auto-evaluations (mean ±\pm std over 3 runs) are reported.

    Method Allrecipes Amazon Apple ArXiv GitHub Booking ESPN Coursera
    GPT-4 (All Tools) 11.1% 17.1% 44.2% 14.0% 48.8% 22.7% 31.8% 31.0%
    WebVoyagerText-only_{\text{Text-only}} 55.6% 31.7% 34.9% 32.6% 61.0% 2.3% 36.4% 23.8%
    WebVoyager (Human) 53.3% 58.5% 65.1% 51.2% 63.4% 43.2% 38.6% 73.8%
    WebVoyagerText-only∗_{\text{Text-only}}^* 57.8%±\pm0.0% 43.1%±\pm1.4% 36.4%±\pm3.5% 50.4%±\pm1.4% 63.4%±\pm2.5% 2.3%±\pm0.0% 38.6%±\pm2.3% 24.6%±\pm1.4%
    WebVoyager∗^* (GPT-4V) 51.1%±\pm2.2% 52.9%±\pm1.4% 62.8%±\pm2.3% 52.0%±\pm1.3% 59.3%±\pm3.7% 32.6%±\pm2.7% 47.0%±\pm1.3% 57.9%±\pm2.7%
    WebVoyagerClaude∗_{\text{Claude}}^* 45.9%±\pm3.4% 58.6%±\pm4.2% 58.1%±\pm4.0% 55.0%±\pm7.0% 56.9%±\pm1.4% 19.0%±\pm1.3% 46.2%±\pm1.3% 68.2%±\pm1.3%
    WebVoyagerGPT-4o∗_{\text{GPT-4o}}^* 56.3%±\pm1.3% 53.7%±\pm2.5% 56.6%±\pm1.3% 60.5%±\pm0.0% 57.7%±\pm3.7% 43.9%±\pm3.5% 44.0%±\pm2.7% 65.1%±\pm2.8%
    Method Cambridge BBC News Google Flights Google Map Google Search Huggingface Wolfram Overall
    GPT-4 (All Tools) 25.6% 9.5% 2.4% 53.7% 60.5% 37.2% 52.2% 30.8%
    WebVoyagerText-only_{\text{Text-only}} 62.8% 45.2% 7.1% 61.0% 67.4% 20.9% 58.7% 40.1%
    WebVoyager (Human) 65.1% 61.9% 59.5% 70.7% 76.7% 44.2% 63.0% 59.1%
    WebVoyagerText-only∗_{\text{Text-only}}^* 66.7%±\pm3.6% 45.2%±\pm2.4% 7.1%±\pm0.0% 62.6%±\pm2.8% 75.2%±\pm1.3% 31.0%±\pm1.4% 60.2%±\pm1.3% 44.3%±\pm0.6%
    WebVoyager∗^* (GPT-4V) 71.3%±\pm1.3% 60.3%±\pm2.8% 51.6%±\pm1.4% 64.3%±\pm2.8% 77.5%±\pm2.7% 55.8%±\pm2.3% 60.9%±\pm2.2% 57.1%±\pm0.2%
    WebVoyagerClaude∗_{\text{Claude}}^* 71.3%±\pm3.6% 66.7%±\pm4.8% 15.1%±\pm5.5% 55.3%±\pm1.4% 72.9%±\pm1.3% 53.5%±\pm4.7% 51.5%±\pm5.4% 52.8%±\pm1.4%
    WebVoyagerGPT-4o∗_{\text{GPT-4o}}^* 82.2%±\pm1.3% 54.8%±\pm2.4% 28.6%±\pm0.0% 56.9%±\pm2.8% 63.6%±\pm1.3% 42.6%±\pm3.6% 65.2%±\pm2.2% 55.5%±\pm0.8%

    Multimodal WebVoyager achieves a human-evaluated overall success rate of 59.1%59.1\%, substantially outperforming GPT-4 (All Tools) (30.8%30.8\%) and WebVoyagerText-only_{\text{Text-only}} (40.1%40.1\%). Text-only baselines drop markedly on visually rich websites such as Booking (2.3%2.3\%) and Google Flights (7.1%7.1\%) where accessibility trees are verbose and complex.

  7. Knowl 7 — Agreement and Cross-Evaluator Consistency of Multimodal Auto-Evaluators

    data/table

    The reliability of GPT-4V as an automatic trajectory evaluator was validated against consolidated human expert annotations on a subset of 300 tasks by varying the number kk of trajectory screenshots provided:

    Evaluation Context Overall Success Rate Human Agreement (%) Cohen's Kappa (κ\kappa)
    k=1k = 1 47.7% 75.3% 0.51
    k=2k = 2 55.3% 79.7% 0.59
    k=3k = 3 54.3% 81.3% 0.62
    Full trajectory 58.3% 85.3% 0.70

    Auto-evaluator consistency with human judgments increases with the number of trajectory frames provided. When comparing multiple multimodal backbones acting as both agent and evaluator across all tasks:

    Agent Backbone GPT-4V Evaluator Claude-3-Opus Evaluator GPT-4o Evaluator
    GPT-4V 57.1% 55.1% 63.0%
    Claude-3-Opus 52.8% 61.6% 55.4%
    GPT-4o 55.5% 54.9% 64.1%

    GPT-4o auto-evaluation is more lenient across all backbones (63.0%–64.1%63.0\%\text{--}64.1\%), while GPT-4V auto-evaluation is stricter. Claude-3-Opus auto-evaluation exhibits self-preference bias (61.6%61.6\% for Claude-3-Opus vs. ≈55%\approx 55\% for GPT-4V and GPT-4o). GPT-4o achieves κ=0.72\kappa = 0.72 with human judges, compared to κ=0.70\kappa = 0.70 for GPT-4V and κ=0.60\kappa = 0.60 for Claude-3-Opus.

  8. Knowl 8 — Performance on GAIA and SeeAct External Benchmarks

    empirical result

    WebVoyager was evaluated on external autonomous web agent benchmarks to test generalization:

    1. GAIA Benchmark (Web browsing subset, initialized via Google Search):
      • Level 1 tasks (n=26n=26): WebVoyager achieves a 38.5%38.5\% task success rate, outperforming GPT-4 (All Tools) (23.1%23.1\%) and WebVoyagerText-only_{\text{Text-only}} (19.2%19.2\%).
      • Level 2 tasks (n=64n=64): WebVoyager achieves a 15.6%15.6\% task success rate, outperforming GPT-4 (All Tools) (12.5%12.5\%) and WebVoyagerText-only_{\text{Text-only}} (12.5%12.5\%).
    2. SeeAct Online Test Set (50 interactive open-web tasks):
      • WebVoyager attains a 30%30\% task success rate without task-specific fine-tuning or specialized candidate selection modules, outperforming the best autonomous SeeAct agent (26%26\%), which relies on a fine-tuned cross-encoder model.
  9. Knowl 9 — Taxonomy and Distribution of WebVoyager Failure Modes

    data/table

    A manual error analysis of 300 failed navigation episodes from the WebVoyager benchmark identified four primary failure categories:

    Failure Category Proportion (%)
    Navigation Stuck 44.4%
    Visual Grounding Issue 24.8%
    Hallucination 21.8%
    Prompt Misalignment 9.0%

    The mechanisms causing these failures include:

    • Navigation Stuck (44.4%): Step exhaustion caused by imprecise search queries overwhelmed by irrelevant results, failure to locate small scrollable viewport elements, indecision between scrolling up or down, or repeating previously failed actions due to context clipping of older observations.
    • Visual Grounding Issue (24.8%): Misinterpretation of visual patterns (e.g., phonetic pronunciation marks or mathematical symbols), failing to recognize subtle UI state changes, or selecting incorrect elements due to visual proximity (such as confusing calendar day numbers with Set-of-Mark numerical tags).
    • Hallucination (21.8%): Overlooking constraints in the prompt (e.g., picking a visible cheap product without applying the required sorting filter) or entering data into an incorrect input field without raising an execution error.
    • Prompt Misalignment (9.0%): Outputting unparseable responses missing the required action code syntax, or executing ANSWER prematurely while stating in text that the task remains incomplete.
  10. Knowl 10 — Limitations and Safety Risks of WebVoyager

    limitation

    The design and deployment of WebVoyager have several documented constraints:

    1. Incomplete Action Space: WebVoyager supports discrete mouse and keyboard actions but lacks continuous operations like drag-and-drop (e.g., dragging sliders, reordering items, or panning map regions) due to the non-finite pixel coordinate space.
    2. File Format Constraints: The observation pipeline handles text and downloaded PDF documents (via the OpenAI Assistant API) but cannot parse dynamic media such as video streams or complex binary formats.
    3. Incompatibility with Open-Source LMMs: Open-source models (such as LLaVA) are hindered by low input image resolutions (224×224224 \times 224 or 336×336336 \times 336) that render fine web text illegible, and context limits of ≤4096\le 4096 tokens that cannot support 15-step multimodal trajectories requiring 7000+7000+ tokens.
    4. Autonomous Deployment Risks: Open-web execution carries risks of unintentionally submitting sensitive personal data, interacting with unauthorized malicious sites, triggering anti-bot CAPTCHAs, or sending high-frequency requests that burden web servers. The framework is restricted to non-login tasks to comply with website terms of service.

Coverage note — None was omitted; the top 10 knowls comprehensively capture WebVoyager's agent architecture, visual grounding approach, action space, benchmark dataset creation, auto-evaluation protocol, benchmark results, cross-evaluator consistency, external benchmark performance, error taxonomy, and limitations.

References

  1. 1.Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, et al. 2022. Cm3: A causal masked multimodal model of the internet. arXiv preprint arXiv:2201.07520.
  2. 2.AI Anthropic. 2024. Introducing the next generation of claude.
  3. 3.AutoGPT. 2022. AutoGPT.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  6. 6.Qi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou, Hsiang-Ting Chen, and Qi Wu. 2023. Webvln: Vision-and-language navigation on websites.
  7. 7.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024. Seeclick: Harnessing gui grounding for advanced visual gui agents.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113.
  9. 9.Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement, 20(1):37–46.
  10. 10.Yong Dai, Duyu Tang, Liangxin Liu, Minghuan Tan, Cong Zhou, Jingquan Wang, Zhangyin Feng, Fan Zhang, Xueyu Hu, and Shuming Shi. 2022. One model, multiple modalities: A sparsely activated approach for text, sound, image, video and code. arXiv preprint arXiv:2205.06126.
  11. 11.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2web: Towards a generalist agent for the web. arXiv preprint arXiv:2306.06070.
  12. 12.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey for in-context learning. arXiv preprint arXiv:2301.00234.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
  14. 14.Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378.
  15. 15.Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur. 2023. Multimodal web navigation with instruction-finetuned foundation models. arXiv preprint arXiv:2305.11854.
  16. 16.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913.
  17. 17.Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2023. A real-world webagent with planning, long context understanding, and program synthesis. arXiv preprint arXiv:2307.12856.
  18. 18.Jack Hessel, Jena D Hwang, Jae Sung Park, Rowan Zellers, Chandra Bhagavatula, Anna Rohrbach, Kate Saenko, and Yejin Choi. 2022. The abduction of sherlock holmes: A dataset for visual abductive reasoning. In European Conference on Computer Vision, pages 558–575. Springer.
  19. 19.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. 2024. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.
  20. 20.Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. Pix2struct: Screenshot parsing as pretraining for visual language understanding. In International Conference on Machine Learning, pages 18893–18912. PMLR.
  21. 21.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557.
  22. 22.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems, 36.
  23. 23.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521.
  24. 24.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. arXiv preprint arXiv:2304.09842.
  25. 25.Kaixin Ma, Hongming Zhang, Hongwei Wang, Xiaoman Pan, and Dong Yu. 2023. Laser: Llm agent with state-space exploration for web navigation. arXiv preprint arXiv:2309.08172.
  26. 26.Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. Gaia: a benchmark for general ai assistants. arXiv preprint arXiv:2311.12983.
  27. 27.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  28. 28.OpenAI. 2023. Gpt-4 technical report.
  29. 29.OpenAI. 2024. Hello gpt-4o.
  30. 30.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  31. 31.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  33. 33.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  34. 34.Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces. arXiv preprint arXiv:2306.00245.
  35. 35.Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. 2017. World of bits: An open-domain platform for web-based agents. In International Conference on Machine Learning, pages 3135–3144. PMLR.
  36. 36.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems.
  37. 37.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  38. 38.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  39. 39.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904.
  40. 40.Lilian Weng. 2023. Llm-powered autonomous agents. lilianweng.github.io.
  41. 41.An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, Zicheng Liu, and Lijuan Wang. 2023. Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation.
  42. 42.Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. 2023a. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441.
  43. 43.Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023b. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9(1).
  44. 44.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022a. Webshop: Towards scalable real-world web interaction with grounded language agents.
  45. 45.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022b. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  46. 46.Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. 2023. Lumos: Learning agents with unified data, modular design, and open-source llms. arXiv preprint arXiv:2311.05657.
  47. 47.Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6720–6731.
  48. 48.Chi Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023. Appagent: Multimodal agents as smartphone users.
  49. 49.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. Gpt-4v (ision) is a generalist web agent, if grounded. arXiv preprint arXiv:2401.01614.
  50. 50.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854.
  51. 51.Zhengxia Zou, Keyan Chen, Zhenwei Shi, Yuhong Guo, and Jieping Ye. 2023. Object detection in 20 years: A survey. Proceedings of the IEEE.

Citation

MLA
He, H., et al. “WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6864–90, https://doi.org/10.18653/v1/2024.acl-long.371.
APA
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., & (于东), D. Y. (2024). WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6864–6890. https://doi.org/10.18653/v1/2024.acl-long.371
Chicago
He, H., W. Yao, K. Ma, et al. 2024. “WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6864–90. https://doi.org/10.18653/v1/2024.acl-long.371.
Harvard
He, H. et al. (2024) “WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6864–6890. Available at: https://doi.org/10.18653/v1/2024.acl-long.371.
Vancouver
1. He H, Yao W, Ma K, Yu W, Dai Y, Zhang H, Lan Z, (于东) DY (2024) WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6864–6890

BibTeX

@inproceedings{he-etal-2024-webvoyager,
    title = "{W}eb{V}oyager: Building an End-to-End Web Agent with Large Multimodal Models",
    author = "He, Hongliang  and
      Yao, Wenlin  and
      Ma, Kaixin  and
      Yu, Wenhao  and
      Dai, Yong  and
      Zhang, Hongming  and
      Lan, Zhenzhong  and
      Yu, Dong",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.371/",
    doi = "10.18653/v1/2024.acl-long.371",
    pages = "6864--6890"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/