Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Yiheng XuZekun WangJunli WangDunjie LuTianbao XieAmrita SahaDoyen SahooTao YuCaiming Xiong
Presents a fully autonomous vision-based GUI interaction framework that standardizes cross-platform actions and incorporates structured inner monologue reasoning to eliminate dependence on text representations and proprietary language models.
Automating graphical user interface tasks across mobile devices, desktop applications, and websites is a crucial frontier for digital productivity. However, existing automated agents face severe bottlenecks. Most rely on lengthy, platform-specific text representations such as raw website code or accessibility trees, which demand heavy computing power and do not generalize across different operating systems. Furthermore, leading solutions depend heavily on proprietary, closed-source artificial intelligence models or generate purely reactive actions without deliberate reasoning, limiting their scalability, privacy, and adaptability in complex workflows.
The article demonstrates that an autonomous, pure-vision framework can effectively operate computer interfaces using only screen screenshots and standard keyboard-and-mouse commands without relying on closed-source models. The primary objective is to evaluate whether separating visual grounding from high-level reasoning and embedding structured self-reflection enables open-source vision-language models to achieve state-of-the-art automation across multiple digital platforms.
To achieve this, the authors constructed a large-scale training dataset combining more than one million visual grounding instances with 35,000 multi-step planning trajectories enriched with explicit thought processes. They implemented a two-stage training strategy: first training the model to locate and interact with interface elements efficiently, and second training it to engage in structured inner monologue to reason and plan before generating commands. The system standardizes interactions through a universal automation interface supplemented by modular extensions. The framework was evaluated across diverse offline benchmarks and live, interactive web and operating system environments against leading commercial and open-source models.
The evaluations yielded several major findings. First, the open-source system established state-of-the-art performance across diverse benchmarks, with its largest variant achieving an 89.2% grounding accuracy on cross-platform tasks and outperforming proprietary commercial models. Second, offline task planning improved dramatically, yielding an average step success rate increase of approximately 52% on complex website interactions over prior visual baselines. Third, the pure-vision approach reduced operational computational overhead significantly, cutting input token volume per interaction step by roughly 70% and lowering execution costs by approximately 93% compared to commercial text-parsing alternatives. Fourth, training across multiple operating environments enabled strong zero-shot transfer; models trained solely on mobile and web data successfully generalized to desktop operating system workflows, outperforming several commercial baselines on complex computer-use benchmarks. Finally, ablation analyses confirmed that structured inner monologue reasoning was essential, improving low-level execution accuracy by up to 11%.
These findings indicate that organizations can achieve robust, human-like digital task automation without locking themselves into expensive, opaque proprietary software pipelines. Operating purely on visual screen data lowers infrastructure costs, provides consistent performance across mobile, web, and desktop environments, and maintains predictable computational requirements regardless of how complex an underlying interface is structured. This significantly improves data privacy and deployment economics for enterprise-scale automation.
For practical implementation, organizations should adopt modular, vision-first architectures for cross-platform automation pipelines. Before deploying autonomous agents into live production or security-critical settings, decision-makers should invest in mechanisms that allow agents to express uncertainty and seek human clarification when instructions are ambiguous. Development teams should also introduce safety training to prevent unintended actions and implement dynamic reasoning controls to balance execution speed against complex planning needs.
Confidence in these findings is high across standard desktop, web, and mobile navigation tasks due to extensive testing across multiple recognized benchmarks. However, key operational limitations remain. The model currently lacks a mechanism to decline ambiguous instructions, which accounted for 40% of observed errors during testing. Additionally, real-world deployment faces practical obstacles from anti-automation defenses, such as security verification prompts and network blocks, which require further engineering before achieving unattended end-to-end reliability.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent establishes the visual language model paradigm for GUI grounding and navigation from high-resolution screen captures that Aguvis directly builds upon.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). Pix2Act provides the foundational formulation for operating autonomous agents directly from raw pixel observations without relying on underlying DOM or view hierarchies.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web provides the core web interaction benchmark and formulation that Aguvis adapts for its multimodal and live evaluation settings.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena defines the standard realistic benchmark environment and evaluation protocols for multi-step autonomous GUI and web agents.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). Spotlight introduces vision-only modeling and coordinate grounding for mobile user interfaces without structural metadata, directly informing Aguvis's pure vision grounding stage.
- Paper: Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration, Junyang Wang et al. (2024). Mobile-Agent-v2 formalizes visual state reflection and planning for device operation, which Aguvis unifies into a single-model inner monologue architecture.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision provides the open-source multimodal backbone architectures and staged instruction-tuning techniques utilized by modern open vision-action agents.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Visual CoT details the multi-modal chain-of-thought and visual grounding mechanisms that underpin Aguvis's structured reasoning and inner monologue framework.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). UI-R1 extends pure vision GUI action prediction by replacing heavy supervised fine-tuning with reinforcement learning over visual grounding tasks like ScreenSpot.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Agent Workflow Memory improves multi-step agent performance on benchmarks like Mind2Web and WebArena by inducing and reusing procedural workflows across tasks.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). WebPilot extends autonomous web interaction by introducing a dual-level optimization strategy with reflection-guided tree search on top of visual web environments.
- Paper: OpenHands: An Open Platform for AI Software Developers as Generalist Agents, Xingyao Wang et al. (2025). OpenHands builds a broader open platform for generalist agents interacting with code, bash environments, and web GUIs.
- Paper: ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots, Yu-Chung Hsiao et al. (2025). ScreenQA advances screen reading comprehension and question answering over app screenshots to complement visual agent execution.
