Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Yiheng XuZekun WangJunli WangDunjie LuTianbao XieAmrita SahaDoyen SahooTao YuCaiming Xiong

article2025ICML282 citations

Presents a fully autonomous vision-based GUI interaction framework that standardizes cross-platform actions and incorporates structured inner monologue reasoning to eliminate dependence on text representations and proprietary language models.

Listen

Automating graphical user interface tasks across mobile devices, desktop applications, and websites is a crucial frontier for digital productivity. However, existing automated agents face severe bottlenecks. Most rely on lengthy, platform-specific text representations such as raw website code or accessibility trees, which demand heavy computing power and do not generalize across different operating systems. Furthermore, leading solutions depend heavily on proprietary, closed-source artificial intelligence models or generate purely reactive actions without deliberate reasoning, limiting their scalability, privacy, and adaptability in complex workflows.

The article demonstrates that an autonomous, pure-vision framework can effectively operate computer interfaces using only screen screenshots and standard keyboard-and-mouse commands without relying on closed-source models. The primary objective is to evaluate whether separating visual grounding from high-level reasoning and embedding structured self-reflection enables open-source vision-language models to achieve state-of-the-art automation across multiple digital platforms.

To achieve this, the authors constructed a large-scale training dataset combining more than one million visual grounding instances with 35,000 multi-step planning trajectories enriched with explicit thought processes. They implemented a two-stage training strategy: first training the model to locate and interact with interface elements efficiently, and second training it to engage in structured inner monologue to reason and plan before generating commands. The system standardizes interactions through a universal automation interface supplemented by modular extensions. The framework was evaluated across diverse offline benchmarks and live, interactive web and operating system environments against leading commercial and open-source models.

The evaluations yielded several major findings. First, the open-source system established state-of-the-art performance across diverse benchmarks, with its largest variant achieving an 89.2% grounding accuracy on cross-platform tasks and outperforming proprietary commercial models. Second, offline task planning improved dramatically, yielding an average step success rate increase of approximately 52% on complex website interactions over prior visual baselines. Third, the pure-vision approach reduced operational computational overhead significantly, cutting input token volume per interaction step by roughly 70% and lowering execution costs by approximately 93% compared to commercial text-parsing alternatives. Fourth, training across multiple operating environments enabled strong zero-shot transfer; models trained solely on mobile and web data successfully generalized to desktop operating system workflows, outperforming several commercial baselines on complex computer-use benchmarks. Finally, ablation analyses confirmed that structured inner monologue reasoning was essential, improving low-level execution accuracy by up to 11%.

These findings indicate that organizations can achieve robust, human-like digital task automation without locking themselves into expensive, opaque proprietary software pipelines. Operating purely on visual screen data lowers infrastructure costs, provides consistent performance across mobile, web, and desktop environments, and maintains predictable computational requirements regardless of how complex an underlying interface is structured. This significantly improves data privacy and deployment economics for enterprise-scale automation.

For practical implementation, organizations should adopt modular, vision-first architectures for cross-platform automation pipelines. Before deploying autonomous agents into live production or security-critical settings, decision-makers should invest in mechanisms that allow agents to express uncertainty and seek human clarification when instructions are ambiguous. Development teams should also introduce safety training to prevent unintended actions and implement dynamic reasoning controls to balance execution speed against complex planning needs.

Confidence in these findings is high across standard desktop, web, and mobile navigation tasks due to extensive testing across multiple recognized benchmarks. However, key operational limitations remain. The model currently lacks a mechanism to decline ambiguous instructions, which accounted for 40% of observed errors during testing. Additionally, real-world deployment faces practical obstacles from anti-automation defenses, such as security verification prompts and network blocks, which require further engineering before achieving unattended end-to-end reliability.

arXiv: 2412.04454
Cover for Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Abstract

Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce AGUVIS, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct AGUVIS DATA COLLECTION, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that AGUVIS achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research.

Table of Contents

  • 1. Introduction
  • 2. AGUVIS
  • 2.1. Problem Formulation
  • 2.2. Unified GUI Interaction Framework
  • 2.3. AGUVIS DATA COLLECTION
  • 2.4. Model Architecture
  • 2.5. Training Paradigm
  • 3. Experiments
  • 3.1. GUI Grounding Evaluation
  • 3.2. Offline GUI Agent Evaluation
  • 3.3. Online GUI Agent Evaluation
  • 4. Analysis
  • 4.1. Impact of Training Stages
  • 4.2. Role of Inner Monologue
  • 4.3. Cross-Platform Benefits
  • 4.4. Efficiency Benifits from Pure Vision Perception
  • 4.5. Error Analysis and Future Work
  • 5. Related Work
  • 5.1. GUI Agent Benchmarks
  • 5.2. GUI Agent Models
  • 6. Conclusion
  • Acknowledgment
  • Impact Statement
  • References
  • A. AGUVIS Unified Design
  • A.1. Details of Action Space in AGUVIS
  • A.2. Pluggable Functions: Mobile Environments as An Example
  • B. Data Curation of AGUVIS DATA COLLECTION
  • B.1. Detailed Source Dataset Statistics
  • B.2. Prompt for Augmenting Planning & Reasoning Trajectories
  • B.3. Human Study on Augmented Data
  • B.3.1. QUALITATIVE HUMAN STUDY
  • B.3.2. FAILURE CASES UNDER NOISY TRAINING DATA
  • C. AGUVIS Training
  • C.1. Training Example Schema
  • C.2. Training Details
  • D. Evaluation Benchmarks
  • D.1. GUI Grounding Evaluation
  • D.2. Offline GUI Agent Evaluation
  • D.3. Online GUI Agent Evaluation
  • D.3.1. PROMPTS FOR USING GPT-4O AS PLANNING MODEL
  • E. Analysis
  • E.1. More Training Ablation
  • E.1.1. TRAINING STRATEGY ABLATION
  • E.1.2. DATA STRATEGY ABLATION
  • E.2. Planning Analysis
  • E.2.1. PROMPTS FOR SELF-PLANNING AND ENFORCED PLANNING MODE.
  • E.2.2. CASES OF INNER MONOLOGUE BONUS
  • E.3. AGUVIS Trajectories Examples on Online Evaluation
  • E.3.1. MIND2WEB-LIVE CASE: AGUVIS-72B AS PLANNER AND GROUNDER
  • E.3.2. MIND2WEB-LIVE CASE: GPT-4O AS PLANNER AND AGUVIS-7B AS GROUNDER
  • E.3.3. ANDROIDWORLD CASE: AGUVIS-72B AS PLANNER AND GROUNDER
  • E.3.4. ANDROIDWORLD CASE: GPT-4O AS PLANNER AND AGUVIS-7B AS GROUNDER
  • E.4. Case of AGUVIS Generalization in Real-World Scenarios

Knowls

  1. Knowl 1 — Unified Pure-Vision GUI Agent Framework and Action Space

    model/method

    The AGUVIS framework models Graphical User Interface (GUI) automation as a Partially Observable Markov Decision Process (POMDP) defined by the tuple (S,A,O,T,O)(S, A, O, T, \mathcal{O}), where SS denotes the set of environment states, AA represents executable actions, OO denotes observations, T:S×A×S→[0,1]T: S \times A \times S \to [0, 1] is the state transition probability distribution, and O:S×A×O→[0,1]\mathcal{O}: S \times A \times O \to [0, 1] is the observation probability distribution.

    Rather than parsing textual structures such as HTML Document Object Model (DOM) trees or Accessibility Trees (which typically consume 4,000 to 6,000 input tokens per step and scale with page complexity), AGUVIS operates exclusively on raw screenshot images ot∈Oo_t \in O. For standard 720p screenshots (1280×7201280 \times 720), this pure-vision formulation maintains a constant observation footprint of 1,196 tokens per step regardless of the underlying interface complexity.

    To standardize interactions across web, desktop, and mobile platforms, AGUVIS defines an action space anchored in Python's pyautogui library alongside modular pluggable functions:

    1. Core pyautogui Actions: Programmatic commands using normalized coordinates (x,y)∈[0,1]2(x, y) \in [0, 1]^2:
      • pyautogui.moveTo(x, y)
      • pyautogui.click(x, y)
      • pyautogui.write('text')
      • pyautogui.press('enter')
      • pyautogui.hotkey('ctrl', 'c')
      • pyautogui.scroll(pixels)
      • pyautogui.dragTo(x, y)
    2. Pluggable Extensions: Environment-specific operations that maintain consistent thought-to-action alignment, including:
      • Web: browser.select_option(x, y, value)
      • Mobile gestures: mobile.swipe(from, to), mobile.home(), mobile.back(), mobile.open_app(name), mobile.long_press(x, y)
      • Task termination and dialogue: terminate(status) (where status ∈\in ['success', 'infeasible']) and answer(text).
  2. Knowl 2 — Structured Inner Monologue Reasoning and Control Routing

    model/method

    To bridge visual observation with low-level execution without relying on closed-source external reasoning models, AGUVIS introduces a two-component inner monologue generated sequentially prior to action generation.

    At time step tt, given the task goal GG, current image observation oto_t, and prior low-level action history a1instr,…,at−1instra_1^{\text{instr}}, \dots, a_{t-1}^{\text{instr}}, the agent outputs:

    1. Explicit Reasoning (hth_t): A predictive thought analyzing the current interface state relative to the goal GG and previous thoughts ht−1h_{t-1}.
    2. Low-Level Instruction (atinstra_t^{\text{instr}}): A single-sentence natural language specification of the immediate next interface interaction.
    3. Executable Action (ata_t): The concrete pyautogui or pluggable function command.

    Control flow and message routing are managed through dedicated recipient tokens:

    • Natural language reasoning is routed via <|im_start|>assistant<|recipient|>all\nThought: {h_t}\nLow-level Instruction: {a_t^{\text{instr}}}<|im_end|>.
    • Executable programmatic actions are routed via <|im_start|>assistant<|recipient|>os\nAction: {a_t}<|diff_marker|>.

    During inference, the framework supports two execution regimes:

    • Enforced Planning Mode: The prompt is appended with <|im_start|>assistant<|recipient|>all\nThought:, forcing the model to generate high-level reasoning before committing to an action.
    • Self-Planning Mode: The prompt terminates at <|im_start|>assistant<|recipient|>, allowing the agent to autonomously decide between emitting an immediate reactive action (os) or initiating inner monologue (all) based on perceived task complexity.
  3. Knowl 3 — AGUVIS Data Collection and Trajectory Augmentation Pipeline

    model/method

    The AGUVIS Data Collection comprises two primary training splits covering website, mobile, and desktop environments:

    1. Grounding Split (1.036M single-step trajectories): Aggregates diverse UI datasets, converting varied annotation formats into standardized pyautogui commands: SeeClick (271K), GUIEnv (328K), GUIAct (67K), WebUI (57K), Widget Captioning (101K), RicoSCA (173K), UI RefExp (16K), RICO Icon (16K), and OmniACT (7K). For UI bounding-box datasets lacking action annotations (e.g., GUIEnv, WebUI, RICO Icon), synthetic instruction-action pairs are generated via deterministic templates covering multi-type interactions (clicks, drags, double-clicks, right-clicks).

    2. Planning & Reasoning Split (35K multi-step trajectories): Aggregates cross-platform sequences from MM-Mind2Web (1,009 trajectories, avg. 7.7 steps), GUIAct (2,482 trajectories, avg. 6.7 steps), MiniWoB++ (2,762 trajectories, avg. 3.6 steps), AitZ (1,987 trajectories, avg. 6.0 steps), AndroidControl (13,594 trajectories, avg. 5.5 steps), GUI Odyssey (7,735 trajectories, avg. 15.3 steps), AMEX (2,991 trajectories, avg. 11.9 steps), and AitW (2,346 trajectories, avg. 8.1 steps).

    VLM-Augmentation Process: To provide intermediate reasoning for datasets containing only raw action sequences, trajectories are augmented using GPT-4o. At step tt, the observation screenshot oto_t is marked with a red bounding box around the ground-truth target element. GPT-4o is provided with the overall goal GG, prior natural language actions a1instr,…,at−1instra_1^{\text{instr}}, \dots, a_{t-1}^{\text{instr}}, and the ground-truth action ata_t, and is prompted to synthesize forward-looking, predictive thoughts hth_t and a low-level instruction atinstra_t^{\text{instr}} from a first-person perspective without referencing hindsight markers. Human evaluation of 90 sampled augmented trajectories showed that 86.7% correctly generated predictive reasoning aligned with ground-truth goals, 7.8% failed due to dataset noise (redundant recorded user actions), and 5.5% exhibited semantic misinterpretations.

  4. Knowl 4 — Two-Stage Training Pipeline for Pure-Vision GUI Agents

    algorithm

    AGUVIS trains vision-language backbones (e.g., Qwen2-VL or LLaVA-OneVision) through a sequential two-stage pipeline that separates element grounding from multi-step reasoning.

    Input: Grounding dataset DgroundD_{\text{ground}}, Planning dataset DplanD_{\text{plan}}, Base vision-language model θ0\theta_0
    Output: Trained autonomous GUI agent θ2\theta_2
    Stage 1: Grounding Training
    Freeze Vision Transformer (ViT) image encoder parameters.
    Apply grounding packing to DgroundD_{\text{ground}}: bundle multiple instruction-action pairs per screenshot into single-image multi-turn dialogue sequences.
    Set peak learning rate to η1\eta_1 (10−510^{-5} for 7B, 5×10−65 \times 10^{-6} for 72B), sequence length to 8192, and maximum image resolution to 1280×7201280 \times 720.
    Train θ0\theta_0 on packed DgroundD_{\text{ground}} for 1 epoch using Adam optimizer with cosine decay and 3% warmup steps.
    Obtain grounding model θ1\theta_1 (AGUVIS-G).
    Stage 2: Planning & Reasoning Training
    Initialize from θ1\theta_1 with ViT parameters kept frozen.
    Construct training batches from DplanD_{\text{plan}} using a reasoning mixture strategy (combining direct actions with structured thoughts and low-level instructions).
    Set peak learning rate to η2\eta_2 (10−510^{-5} for 7B, 5×10−65 \times 10^{-6} for 72B) and batch size to 128.
    Train θ1\theta_1 on DplanD_{\text{plan}} for 1 epoch using Adam optimizer with cosine decay.
    Obtain full agent model θ2\theta_2 (AGUVIS).

    The grounding packing strategy in Stage 1 eliminates redundant image feature extractions across interactable objects on the same screen, reducing Stage 1 GPU training time on web grounding data from 6 hours to 1 hour while improving ScreenSpot web grounding accuracy from 73.3% to 76.8%.

    Training AGUVIS-7B requires 8 nodes of H100-80G GPUs (5 hours for Stage 1, 1 hour for Stage 2); AGUVIS-72B requires 16 nodes of H100-80G GPUs (30 hours for Stage 1, 6 hours for Stage 2).

  5. Knowl 5 — GUI Visual Grounding Performance on ScreenSpot

    data/table

    GUI visual grounding performance evaluated on the ScreenSpot benchmark (1.2K single-step instructions) across mobile, desktop, and web platforms, comparing text elements and icons/widgets under both direct (Original Instructions) and self-planned (Self-Plan) evaluation protocols.

    Planner Grounder Mobile Desktop Web Avg
    Text Icon/Widget Text Icon/Widget Text Icon/Widget
    Original Instructions Evaluation
    – GPT-4 22.6 24.5 20.2 11.8 9.2 8.8 16.2
    – GPT-4o 20.2 24.9 21.1 23.6 12.2 7.8 18.3
    – CogAgent 67.0 24.0 74.2 20.0 70.4 28.6 47.4
    – SeeClick 78.0 52.0 72.2 30.0 55.7 32.5 53.4
    – Qwen2-VL 75.5 60.7 76.3 54.3 35.2 25.7 55.3
    – UGround 82.8 60.3 82.5 63.6 80.4 70.4 73.3
    – AGUVIS-G-7B 88.3 78.2 88.1 70.7 85.7 74.8 81.8
    Self-Plan Evaluation
    GPT-4 SeeClick 76.6 55.5 68.0 28.6 40.9 23.3 48.8
    GPT-4 OmniParser 93.9 57.0 91.3 63.6 81.3 51.0 73.0
    GPT-4 UGround 90.1 70.3 87.1 55.7 85.7 64.6 75.6
    GPT-4o SeeClick 81.0 59.8 69.6 33.6 43.9 26.2 52.3
    GPT-4o UGround 93.4 76.9 92.8 67.9 88.7 68.9 81.4
    – AGUVIS-7B 95.6 77.7 93.8 67.1 88.3 75.2 84.4
    – AGUVIS-72B 94.5 85.2 95.4 77.9 91.3 85.9 89.2

    AGUVIS-G-7B (trained purely through Stage 1) achieves an average accuracy of 81.8% on original instructions, outperforming dedicated GUI grounders such as UGround (73.3%) and SeeClick (53.4%). Under the self-plan setting, the full model AGUVIS-7B attains 84.4% average accuracy, exceeding pipeline systems combining GPT-4o with external grounders (81.4%). Scaled AGUVIS-72B achieves 89.2% overall accuracy, establishing superior performance across all modalities and platforms.

  6. Knowl 6 — Offline GUI Interaction Performance on Multimodal-Mind2Web and AndroidControl

    data/table

    Performance of AGUVIS on offline website interaction (Multimodal-Mind2Web) and mobile device control (AndroidControl).

    Obs. Planner Grounder Cross-Task Cross-Website Cross-Domain
    Ele.Acc Op.F1 Step SR Ele.Acc Op.F1 Step SR Ele.Acc Op.F1 Step SR
    T GPT-3.5 Choice 19.4 59.2 16.8 14.9 56.5 14.1 25.2 57.9 24.1
    T GPT-4 Choice 40.8 63.1 32.3 30.2 61.0 27.0 35.4 61.9 29.7
    T+I GPT-4 Choice 46.4 73.4 40.2 38.0 67.8 32.4 42.4 69.3 36.8
    T+I GPT-4 SoM 29.6 – 20.3 20.1 – 13.9 27.0 – 23.7
    I GPT-4o SeeClick 32.1 – – 33.1 – – 33.5 – –
    I GPT-4V OmniParser 42.4 87.6 39.4 41.0 84.8 36.5 45.5 85.7 42.0
    I GPT-4o UGround 47.7 – – 46.0 – – 46.6 – –
    I SeeClick-9.6B – 28.3 87.0 25.5 21.4 80.6 16.4 23.2 84.8 20.8
    I AGUVIS-7B – 64.2 89.8 60.4 60.7 88.1 54.6 60.4 89.2 56.6
    I AGUVIS-72B – 69.5 90.8 64.0 62.6 88.6 56.5 63.5 88.5 58.2

    T indicates raw HTML text input, I indicates screenshot images, and T+I indicates both. On Multimodal-Mind2Web, AGUVIS operates purely on images, achieving Step Success Rates (Step SR) of 60.4% / 54.6% / 56.6% (7B) and 64.0% / 56.5% / 58.2% (72B), outperforming multimodal GPT-4 + Set-of-Marks (SoM) and GPT-4V + OmniParser.

    On the AndroidControl out-of-domain (OOD) test set (500 step-actions):

    Observation Planner Grounder Step Accuracy (%)
    High-Level Tasks Low-Level Tasks
    Accessibility Tree GPT-4-Turbo Choice 42.1 55.0
    Accessibility Tree PaLM 2S* Choice 58.5 77.5
    Image GPT-4-Turbo SeeClick 39.4 47.2
    Image GPT-4-Turbo UGround 46.2 58.0
    Image GPT-4o SeeClick 41.8 52.8
    Image GPT-4o UGround 48.4 62.4
    Image AGUVIS-7B – 61.5 80.5
    Image AGUVIS-72B – 66.4 84.4

    AGUVIS-72B outperforms closed-source models utilizing textual accessibility trees (e.g., PaLM 2S* at 58.5% / 77.5%) while relying strictly on visual screenshots.

  7. Knowl 7 — Online Agent Interaction, Inference Efficiency, and OSWorld Generalization

    empirical result

    AGUVIS was evaluated across dynamic online benchmarks (Mind2Web-Live, AndroidWorld, MobileMiniWob) and open-ended desktop operating systems (OSWorld):

    • Mind2Web-Live (104 online web tasks): Operating purely on screenshots via BrowserGym, AGUVIS-72B achieves a Task Success Rate (SR) of 27.1% with an average inference cost of $0.012 per successful step. It outperforms HTML-based GPT-4o Choice (22.1% SR, $0.142 cost), Llama-3.1-405B Choice (24.0% SR, $0.174 cost), and vision-based GPT-4o + AGUVIS-7B (24.0% SR, $0.106 cost). Compared to GPT-4o over HTML (~4,000 tokens/step), AGUVIS reduces input token consumption by 70% (1,196 tokens/step) and financial inference cost by 93%.

    • AndroidWorld (116 mobile tasks) & MobileMiniWob (92 tasks): On AndroidWorld, GPT-4o Planner + AGUVIS-7B Grounder achieves 37.1% Task SR, surpassing Accessibility Tree (AXTree) GPT-4-Turbo (30.6%) and Image+AXTree GPT-4-Turbo SoM (25.4%). Standalone AGUVIS-72B achieves 26.1% on AndroidWorld and 66.0% on MobileMiniWob, outperforming AXTree-based GPT-4-Turbo (59.7%) and Gemini 1.5 Pro (57.4%).

    • OSWorld Benchmark (369 desktop OS and workflow tasks): Despite being trained exclusively on web and mobile trajectories, AGUVIS exhibits strong zero-shot transfer to desktop environments:

      • GPT-4o Planner + AGUVIS-72B Grounder achieves a 17.04% Task SR (and AGUVIS-7B achieves 14.79%), outperforming GPT-4o + Set-of-Marks (4.59%) and Claude Computer-Use (14.9%).
      • Standalone AGUVIS-72B achieves a 10.26% Task SR, outperforming standalone proprietary models including GPT-4o (5.03%), GPT-4V (5.26%), and Gemini-Pro-1.5 (5.40%).
  8. Knowl 8 — Ablation Analysis of Training Stages, Inner Monologue, and Cross-Platform Transfer

    data/table

    Ablation results on AGUVIS-7B (Qwen2-VL backbone) and AGUVISLLAVA-OV (LLaVA-OneVision backbone) evaluating training stage configurations, inner monologue (IM) reasoning, and cross-platform data mixing.

    Configuration ScreenSpot Multimodal-Mind2Web (Step SR) AndroidControl (Step Acc.)
    Avg Cross-Task Cross-Website Cross-Domain High-Level Low-Level
    AGUVIS-7B (Qwen2-VL Backbone)
    Stage 1 →\to 2 84.4 58.5 55.4 54.8 61.5 80.5
    Stage 1 + 2 (Joint) 85.0 56.1 53.1 55.6 59.2 80.9
    w/o Stage 2 81.8 50.9 45.2 45.3 58.0 75.6
    w/o Stage 1 77.4 59.7 55.3 56.8 58.8 79.8
    w/o Stage 1 2 55.3 50.9 44.9 47.7 59.1 59.2
    w/o Inner Monologue 79.3 55.4 53.7 54.9 60.3 69.1
    AGUVISLLAVA-OV (LLaVA-OneVision Backbone)
    Stage 1 →\to 2 81.2 55.3 50.0 50.8 60.7 82.4
    w/o Stage 2 70.0 43.4 39.0 40.7 54.9 65.6
    w/o Stage 1 71.3 42.5 40.3 42.8 61.4 80.5
    w/o Stage 1 2 3.8 33.8 30.5 32.4 50.4 50.0

    Key Findings:

    1. Sequential vs Joint Training: Sequential training (Stage 1 →\to Stage 2) preserves multi-step planning performance better than Joint training (Stage 1 + 2), because Stage 1 grounding data is much larger (1.036M vs 35K trajectories) and dominates optimization during joint training.
    2. Backbone Agnosticism: On the weaker baseline LLaVA-OneVision, Stage 1 →\to Stage 2 training improves ScreenSpot accuracy from 3.8% to 81.2% and Multimodal-Mind2Web Cross-Task Step SR from 33.8% to 55.3%.
    3. Role of Inner Monologue: Removing inner monologue causes drops across all tasks, notably dropping ScreenSpot from 84.4% to 79.3% and AndroidControl Low-Level accuracy from 80.5% to 69.1%.
    4. Cross-Platform Transfer: Multi-domain training (Web + Mobile, 35K trajectories) improves Multimodal-Mind2Web performance over Web-only training (6K trajectories) and Mind2Web-only fine-tuning (1K trajectories), yielding Step SR of 58.5% vs 53.1% vs 50.9% on Cross-Task.
  9. Knowl 9 — Error Distribution and Ambiguity Failure Modes in GUI Grounding

    limitation

    Detailed failure analysis of 50 incorrect predictions on ScreenSpot under the self-planning setting reveals two primary error categories:

    1. Instruction Ambiguity (40% of errors): Natural language instructions that refer ambiguously to multiple plausible interface elements. The current agent architecture lacks mechanisms to estimate confidence, output uncertainty estimates, ask clarifying questions, or refuse execution in ambiguous scenarios.
    2. Grounding Errors (60% of errors): Coordinate prediction inaccuracies on unambiguous targets.

    Enforcing planning by compelling the agent to generate inner monologue before acting resolves 20% of grounding errors. However, a significant operational failure mode remains: for queries that are syntactically concise but require deeper semantic context or domain knowledge, the model frequently fails to identify the need for multi-step reasoning and inappropriately defaults to direct, reactive grounding.

Coverage note — None was omitted; all key theoretical formulations, dataset curation statistics, two-stage training designs, empirical evaluations across grounding/offline/online benchmarks, ablation studies, and error analyses were extracted into self-contained knowls.

References

  1. 1.Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., and y Arcas, B. A. Uibert: Learning generic multimodal representations for UI understanding. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, 2021.
  2. 2.Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J. Qwen2.5-vl technical report. CoRR, abs/2502.13923, 2025.
  3. 3.Bonatti, R., Zhao, D., Bonacci, F., Dupont, D., Abdali, S., Li, Y., Lu, Y., Wagle, J., Koishida, K., Bucker, A. F. C., Jang, L., and Hui, Z. Windows agent arena: Evaluating multi-modal os agents at scale. ArXiv preprint, 2024. URL https://api.semanticscholar.org/CorpusID:272600411.
  4. 4.Cao, R., Lei, F., Wu, H., Chen, J., Fu, Y., Gao, H., Xiong, X., Zhang, H., Hu, W., Mao, Y., Xie, T., Xu, H., Zhang, D., Wang, S. I., Sun, R., Yin, P., Xiong, C., Ni, A., Liu, Q., Zhong, V., Chen, L., Yu, K., and Yu, T. Spider2-v: How far are multimodal agents from automating data science and engineering workflows? In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, 2024.
  5. 5.Chai, Y., Huang, S., Niu, Y., Xiao, H., Liu, L., Zhang, D., Gao, P., Ren, S., and Li, H. Amex: Android multi-annotation expo dataset for mobile gui agents. ArXiv preprint, 2024. URL https://arxiv.org/abs/2407.17490.
  6. 6.Chen, W., Cui, J., Hu, J., Qin, Y., Fang, J., Zhao, Y., Wang, C., Liu, J., Chen, G., Huo, Y., et al. Guicourse: From general vision language models to versatile gui agents. ArXiv preprint, 2024a. URL https://arxiv.org/abs/2406.11317.
  7. 7.Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., Liu, Z., Gu, L., Wang, X., Li, Q., Ren, Y., Chen, Z., Luo, J., Wang, J., Jiang, T., Wang, B., He, C., Shi, B., Zhang, X., Lv, H., Wang, Y., Shao, W., Chu, P., Tu, Z., He, T., Wu, Z., Deng, H., Ge, J., Chen, K., Dou, M., Lu, L., Zhu, X., Lu, T., Lin, D., Qiao, Y., Dai, J., and Wang, W. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. CoRR, abs/2412.05271, 2024b. doi: 10.48550/ARXIV.2412.05271. URL https://doi.org/10.48550/arXiv.2412.05271.
  8. 8.Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z. Seeclick: Harnessing GUI grounding for advanced visual GUI agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, 2024.
  9. 9.Deitke, M., Clark, C., Lee, S., Tripathi, R., Yang, Y., Park, J. S., Salehi, M., Muennighoff, N., Lo, K., Soldaini, L., Lu, J., Anderson, T., Bransom, E., Ehsani, K., Ngo, H., Chen, Y., Patel, A., Yatskar, M., Callison-Burch, C., Head, A., Hendrix, R., Bastani, F., VanderBilt, E., Lambert, N., Chou, Y., Chheda, A., Sparks, J., Skjonsberg, S., Schmitz, M., Sarnat, A., Bischoff, B., Walsh, P., Newell, C., Wolters, P., Gupta, T., Zeng, K.-H., Borchardt, J., Groeneveld, D., Dumas, J., Nam, C., Lebrecht, S., Wittlif, C., Schoenick, C., Michel, O., Krishna, R., Weihs, L., Smith, N. A., Hajishirzi, H., Girshick, R., Farhadi, A., and Kembhavi, A. Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146, 2024.
  10. 10.Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., and Kumar, R. Rico: A mobile app dataset for building data-driven design applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology, 2017.
  11. 11.Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y. Mind2web: Towards a generalist agent for the web. In Advances in Neural Information Processing Systems, 2023.
  12. 12.Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Verme, M. D., Marty, T., Vazquez, D., Chapados, N., and Lacoste, A. Workarena: How capable are web agents at solving common knowledge work tasks? In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=BRfqYrikdo.
  13. 13.Gou, B., Wang, R., Zheng, B., Xie, Y., Chang, C., Shu, Y., Sun, H., and Su, Y. Navigating the digital world as humans do: Universal visual grounding for GUI agents. In The Thirteenth International Conference on Learning Representations, 2025.
  14. 14.Gur, I., Furuta, H., Huang, A. V., Safdari, M., Matsuo, Y., Eck, D., and Faust, A. A real-world webagent with planning, long context understanding, and program synthesis. In International Conference on Learning Representations, 2024.
  15. 15.Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., Ding, M., et al. Cogagent: A visual language model for gui agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
  16. 16.Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Jackson, T., Brown, N., Luu, L., Levine, S., Hausman, K., and Ichter, B. Inner monologue: Embodied reasoning through planning with language models. In Conference on Robot Learning, CoRL 2022, 2022.
  17. 17.Kapoor, R., Butala, Y. P., Russak, M., Koh, J. Y., Kamble, K., AlShikh, W., and Salakhutdinov, R. Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web. In Computer Vision - ECCV 2024 - 18th European Conference, 2024.
  18. 18.Kim, G., Baldi, P., and McAleer, S. Language models can solve computer tasks. In Advances in Neural Information Processing Systems, 2023.
  19. 19.Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024a.
  20. 20.Koh, J. Y., McAleer, S., Fried, D., and Salakhutdinov, R. Tree search for language model agents. ArXiv preprint, 2024b.
  21. 21.Lai, H., Liu, X., Iong, I. L., Yao, S., Chen, Y., Shen, P., Yu, H., Zhang, H., Zhang, X., Dong, Y., and Tang, J. Autowebglm: A large language model-based web navigating agent. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 5295—-5306, 2024.
  22. 22.Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., and Li, C. Llava-onevision: Easy visual task transfer. ArXiv preprint, 2024a. URL https://arxiv.org/abs/2408.03326.
  23. 23.Li, T., Li, G., Deng, Z., Wang, B., and Li, Y. A zero-shot language agent for computer control with structured reflection. In Findings of the Association for Computational Linguistics: EMNLP 2023, 2023.
  24. 24.Li, T., Li, G., Zheng, J., Wang, P., and Li, Y. MUG: interactive multimodal grounding on user interfaces. In Findings of the Association for Computational Linguistics: EACL 2024, 2024b.
  25. 25.Li, W., Bishop, W. E., Li, A., Rawles, C., Campbell-Ajala, F., Tyamagundlu, D., and Riva, O. On the effects of data scale on UI control agents. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, 2024c.
  26. 26.Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J. Mapping natural language instructions to mobile UI action sequences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2020a.
  27. 27.Li, Y., Li, G., He, L., Zheng, J., Li, H., and Guan, Z. Widget captioning: Generating natural language description for mobile user interface elements. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2020b. URL https://aclanthology.org/2020.emnlp-main.443.
  28. 28.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  29. 29.Lu, Q., Shao, W., Liu, Z., Meng, F., Li, B., Chen, B., Huang, S., Zhang, K., Qiao, Y., and Luo, P. Gui odyssey: A comprehensive dataset for cross-app gui navigation on mobile devices. ArXiv preprint, 2024. URL https://arxiv.org/abs/2406.08451.
  30. 30.Lu, X. H., Kasner, Z., and Reddy, S. Weblinx: Real-world website navigation with multi-turn dialogue. In Forty-first International Conference on Machine Learning, ICML 2024, 2024.
  31. 31.Lu, Y., Yang, J., Shen, Y., and Awadallah, A. Omniparser for pure vision based gui agent, 2024. URL https://arxiv.org/abs/2408.00203.
  32. 32.Microsoft. Playwright for python documentation. https://playwright.dev/python/, 2024.
  33. 33.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. Webgpt: Browser-assisted question-answering with human feedback. ArXiv preprint, 2021. URL https://arxiv.org/abs/2112.09332.
  34. 34.Niu, R., Li, J., Wang, S., Fu, Y., Hu, X., Leng, X., Kong, H., Chang, Y., and Wang, Q. Screenagent: A vision language model-driven computer control agent. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI 2024, 2024.
  35. 35.OpenAI. Hello gpt-4o, 2024. URL https://openai.com/index/hello-gpt-4o.
  36. 36.Pan, Y., Kong, D., Zhou, S., Cui, C., Leng, Y., Jiang, B., Liu, H., Shang, Y., Zhou, S., Wu, T., et al. Webcanvas: Benchmarking web agents in online environments. ArXiv preprint, 2024. URL https://arxiv.org/abs/2406.12373.
  37. 37.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 2019.
  38. 38.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, 2020.
  39. 39.Rawles, C., Clinckemaillie, S., Chang, Y., Waltz, J., Lau, G., Fair, M., Li, A., Bishop, W., Li, W., Campbell-Ajala, F., Toyama, D., Berry, R., Tyamagundlu, D., Lillicrap, T., and Riva, O. Androidworld: A dynamic benchmarking environment for autonomous agents, 2024a. URL https://arxiv.org/abs/2405.14573.
  40. 40.Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T. Androidinthewild: A large-scale dataset for android device control. Advances in Neural Information Processing Systems, 2024b.
  41. 41.Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. ArXiv preprint, abs/2403.05530, 2024. URL https://arxiv.org/abs/2403.05530.
  42. 42.Wang, H., Li, T., Deng, Z., Roth, D., and Li, Y. Devil’s advocate: Anticipatory reflection for LLM agents. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, 2024a.
  43. 43.Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. ArXiv preprint, 2024b. URL https://arxiv.org/abs/2409.12191.
  44. 44.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv preprint, 2019. URL https://arxiv.org/abs/1910.03771.
  45. 45.Wu, J., Wang, S., Shen, S., Peng, Y.-H., Nichols, J., and Bigham, J. Webui: A dataset for enhancing visual ui understanding with web semantics. ACM Conference on Human Factors in Computing Systems (CHI), 2023.
  46. 46.Wu, Z., Wu, Z., Xu, F., Wang, Y., Sun, Q., Jia, C., Cheng, K., Ding, Z., Chen, L., Liang, P. P., and Qiao, Y. OS-ATLAS: A foundation action model for generalist GUI agents. In The Thirteenth International Conference on Learning Representations, 2025.
  47. 47.Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, 2024.
  48. 48.Xu, T., Chen, L., Wu, D.-J., Chen, Y., Zhang, Z., Yao, X., Xie, Z., Chen, Y., Liu, S., Qian, B., Torr, P. H. S., Ghanem, B., and Li, G. Crab: Cross-environment agent benchmark for multimodal language model agents. ArXiv preprint, 2024a. URL https://arxiv.org/abs/2407.01511.
  49. 49.Xu, Y., Su, H., Xing, C., Mi, B., Liu, Q., Shi, W., Hui, B., Zhou, F., Liu, Y., Xie, T., Cheng, Z., Zhao, S., Kong, L., Wang, B., Xiong, C., and Yu, T. Lemur: Harmonizing natural language and code for language agents. In International Conference on Learning Representations, 2024b.
  50. 50.Yin, D., Brahman, F., Ravichander, A., Chandu, K. R., Chang, K., Choi, Y., and Lin, B. Y. Agent lumos: Unified and modular training for open-source language agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024.
  51. 51.Zhang, C., Yang, Z., Liu, J., Li, Y., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G. Appagent: Multimodal agents as smartphone users. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, 2025.
  52. 52.Zhang, J., Lan, T., Zhu, M., Liu, Z., Hoang, T., Kokane, S., Yao, W., Tan, J., Prabhakar, A., Chen, H., Liu, Z., Feng, Y., Awalgaonkar, T., Murthy, R., Hu, E., Chen, Z., Xu, R., Niebles, J. C., Heinecke, S., Wang, H., Savarese, S., and Xiong, C. xlam: A family of large action models to empower ai agent systems. ArXiv preprint, 2024a.
  53. 53.Zhang, J., Wu, J., Teng, Y., Liao, M., Xu, N., Xiao, X., Wei, Z., and Tang, D. Android in the zoo: Chain-of-action-thought for GUI agents. In Findings of the Association for Computational Linguistics: EMNLP 2024, 2024b.
  54. 54.Zhang, Z. and Zhang, A. You only look at screens: Multimodal chain-of-action agents. In Findings of the Association for Computational Linguistics, 2024.
  55. 55.Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y. Gpt-4v(ision) is a generalist web agent, if grounded. In Forty-first International Conference on Machine Learning, ICML 2024, 2024a.
  56. 56.Zheng, L., Wang, R., Wang, X., and An, B. Synapse: Trajectory-as-exemplar prompting with memory for computer control. In The Twelfth International Conference on Learning Representations, ICLR 2024. OpenReview.net, 2024b.
  57. 57.Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., and Neubig, G. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, 2024.

Citation

MLA
Xu, Y., et al. “Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction”. arXiv, 2024, http://arxiv.org/abs/2412.04454v2.
APA
Xu, Y., Wang, Z., Wang, J., Lu, D., Xie, T., Saha, A., Sahoo, D., Yu, T., & Xiong, C. (2024). Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction. arXiv. http://arxiv.org/abs/2412.04454v2
Chicago
Xu, Y., Z. Wang, J. Wang, et al. 2024. “Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction”. arXiv. http://arxiv.org/abs/2412.04454v2.
Harvard
Xu, Y. et al. (2024) “Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2412.04454v2.
Vancouver
1. Xu Y, Wang Z, Wang J, Lu D, Xie T, Saha A, Sahoo D, Yu T, Xiong C (2024) Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction. arXiv

BibTeX

@article{xu2024aguvis,
  title = {Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction},
  author = {Xu, Yiheng and Wang, Zekun and Wang, Junli and Lu, Dunjie and Xie, Tianbao and Saha, Amrita and Sahoo, Doyen and Yu, Tao and Xiong, Caiming},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2412.04454v2},
  eprint = {2412.04454}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/