WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Hongliang HeWenlin YaoKaixin MaWenhao YuYong DaiHongming ZhangZhenzhong LanDong Yu
Presents WebVoyager, an end-to-end multimodal web agent that interacts directly with live websites using visual and textual cues, outperforming text-only baselines alongside a benchmark of real-world tasks and an automated evaluation protocol.
Autonomous web agents have significant potential to automate complex online tasks, yet existing systems rely primarily on simplified simulators or raw website code. These text-only approaches ignore the visual layout and intuitive design of modern websites, making it difficult for automated tools to handle dynamic online interfaces. To address this limitation, the article introduces and evaluates WebVoyager, an end-to-end multimodal agent designed to navigate live, real-world websites autonomously by combining visual screenshots with textual interface elements.
The research evaluated the agent across an automated browsing environment using a new benchmark of 643 real-world tasks across 15 popular websites, including Amazon, Booking.com, and Google Flights. The agent identified interactive components by overlaying visual tags onto webpage screenshots, allowing it to reason through decisions and execute human-like actions such as clicking, typing, and scrolling. In addition to testing against standard baselines, the article introduced an automated evaluation method powered by a multimodal model to score navigation recordings, validating these scores against independent human evaluations.
The findings show that WebVoyager achieved a 59.1% task success rate, substantially outperforming both a text-only setup at 40.1% and a leading commercial integrated tool at 30.8%. Visual perception proved particularly decisive on complex interfaces; for instance, on booking and flight platforms that require calendar interactions, the visual agent succeeded where text-only agents failed. Conversely, text-only inputs performed slightly better on text-heavy websites where small text was harder to resolve from screenshots alone. Additionally, the automated scoring framework demonstrated an 85.3% agreement rate with human judges, confirming that vision-language models can reliably evaluate multi-step web navigation without constant manual oversight.
These results indicate that effective digital assistants require both visual and textual understanding to navigate modern web interfaces reliably. When the agent failed, the primary bottlenecks were getting stuck in repetitive navigation loops (44.4% of errors), visual misidentification of closely grouped elements (24.8%), and generating incomplete or hallucinated responses (21.8%). Organizations building or adopting web automation should focus future development on hybrid inputs that extract clean text alongside visual snapshots to handle dense content, while refining navigation prompts to prevent looping.
While confidence in the reported performance gains is supported by solid human validation, leaders should account for current operational boundaries. The agent does not support complex input actions like dragging, cannot parse video media, and was evaluated exclusively on public, non-login workflows without security verifications. Before deploying autonomous web agents into live operational settings, organizations must implement robust safety safeguards to prevent unintended actions, such as submitting unauthorized information or interacting with insecure sites.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena established the benchmark environment and task formulation for autonomous web agents that WebVoyager directly builds upon and extends to live websites.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web introduced realistic, multi-website navigation tasks and element-interaction pipelines that serve as the direct conceptual foundation for WebVoyager's real-world agent design.
- Paper: Voyager: An Open-Ended Embodied Agent with Large Language Models, Guanzhi Wang et al. (2023). Voyager pioneered the paradigm of LLM-driven autonomous agents interacting iteratively with open-ended environments, directly inspiring WebVoyager's end-to-end framework.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). This work demonstrates how agents can transition from structured code to raw visual pixel inputs for GUI actions, an essential design prerequisite for multimodal web agents.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). WebGPT established foundational browser-assisted interaction and web navigation paradigms for language models.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). LLaVA provides the core visual instruction tuning architecture that powers modern large multimodal models used in WebVoyager.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). WebPilot extends autonomous web navigation by introducing strategic exploration and tree search to solve complex web tasks beyond single-trajectory agent models.
- Paper: AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, Ori Yoran et al. (2024). AssistantBench builds on web agent benchmarks to evaluate whether autonomous systems can execute complex, time-consuming multi-page research on the live web.
- Paper: SafeArena: Evaluating the Safety of Autonomous Web Agents, Ada Defne Tur et al. (2025). SafeArena analyzes the safety risks and adversarial vulnerabilities that arise when deploying autonomous multimodal web agents like WebVoyager across interactive platforms.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent advances visual web navigation by implementing dedicated high-resolution visual modules for fine-grained GUI element perception and grounding.
- Paper: FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation, Kaiyi Huang et al. (2026). AGUVIS generalizes multimodal web interaction into a purely vision-based GUI automation architecture featuring explicit planning and inner monologue.
