GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
Qianhui WuKanzhi ChengRui YangChaoyun ZhangJianwei YangHuiqiang JiangJian MuBaolin PengBo QiaoReuben Tan
Presents GUI-Actor, a coordinate-free visual grounding framework for GUI agents that replaces text coordinate generation with an attention-based action head and verifier, enabling a 7B vision-language model to outperform 72B baselines on interface benchmarks while training only 100M parameters.
Autonomous graphical user interface (GUI) agents powered by vision-language models (VLMs) have gained significant traction for automating complex tasks across mobile, desktop, and web environments. A core capability of these agents is visual grounding—the ability to map natural language instructions directly to actionable screen regions. Prevailing approaches formulate grounding as text-based coordinate generation, where the model outputs numeric pixel coordinates. However, this text-based paradigm suffers from weak spatial-semantic alignment, overly penalizes valid interaction variations via single-point supervision, and creates a mismatch between coarse visual patch representations and dense coordinate targets.
The article demonstrates GUI-Actor, a coordinate-free framework designed to achieve accurate, human-like visual grounding. The system introduces an attention-based action head anchored by a dedicated context token, which attends directly over screenshot image patches rather than predicting textual numbers. The model is trained using spatial-aware multi-patch supervision, which treats all patches covered by ground-truth bounding boxes as positive targets to accommodate natural interaction ambiguity. In addition, the framework incorporates a lightweight grounding verifier that evaluates candidate regions generated in a single model forward pass to select the most plausible action before execution. The approach was evaluated across major benchmarks including ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, and a 49-task live Windows desktop testbed (OS-World-W) using roughly one million training screenshots.
The experimental findings show substantial improvements in performance and efficiency over existing methods. On the high-resolution, out-of-domain ScreenSpot-Pro benchmark, the 7-billion-parameter GUI-Actor achieves an accuracy of 40.7 to 44.6 (depending on the vision-language backbone), outperforming much larger 72-billion-parameter coordinate-based systems such as UI-TARS-72B (38.1). The smaller 2-billion-parameter GUI-Actor (36.7) also surpasses several competing 7-billion models. Adding the grounding verifier boosts performance further, reaching 44.2 on ScreenSpot-Pro for GUI-Actor-7B. Crucially, the system achieves peak benchmark accuracy with only about 60% of the training data required by baseline models. In live Windows environment testing, GUI-Actor-7B achieved a 12.2% task success rate, outperforming previous visual grounding baselines such as Aguvis-7B (4.0%), NAVI (10.2%), and OmniAgent (10.2%).
These results demonstrate that direct spatial attention is structurally superior to text-coordinate generation for computer control tasks. By grounding actions at the native resolution of the vision backbone, the method exhibits strong generalization to unseen screen layouts and resolutions while eliminating costly multi-turn sampling during inference. Furthermore, experiments confirm that freezing the underlying vision-language model and training only the lightweight action head (approximately 100 million parameters for a 7-billion model) achieves competitive accuracy. This lightweight training option provides a cost-effective pathway to equip general-purpose models with interface-control capabilities without degrading their core language and reasoning strengths.
Organizations developing or deploying automated computer agents should adopt coordinate-free, attention-based action heads and patch-level verifiers to reduce training compute, lower inference latency, and enhance execution reliability across varied screen resolutions. A noted limitation of the framework involves very small interface elements (such as icons under 10 by 10 pixels), which can fall between coarse 28-by-28-pixel visual patches. While patch-clustering mitigates this, future work should focus on higher perceptual visual resolutions and fine-grained spatial offsets to support precision-demanding applications like industrial design and computer-aided engineering software.
- Paper: ShowUI: One Vision-Language-Action Model for GUI Visual Agent, Kevin Qinghong Lin et al. (2025). ShowUI establishes the vision-language-action GUI setting and patch-based visual processing that GUI-Actor refines with coordinate-free grounding.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent shows how high-resolution VLMs perceive and ground GUI elements, providing essential context for GUI-Actor’s patch-token action head.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). PIX2ACT frames screenshot-driven interaction as visual grounding and action prediction, clarifying the coordinate-generation problem GUI-Actor seeks to overcome.
- Paper: Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction, Yiheng Xu et al. (2025). AGUVIS develops visual grounding for autonomous GUI agents and benchmarks, situating GUI-Actor’s grounding mechanism within the same vision-only agent setting.
- Paper: UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis, Xinyi Liu et al. (2025). UI-E2I-Synth examines instruction-to-element grounding and its data challenges, helping explain the supervision problem that GUI-Actor addresses.
- Paper: Covering Human Action Space for Computer Use: Data Synthesis and Benchmark, Miaosen Zhang et al. (2026). CUActSpot extends GUI grounding beyond clicks to ordered, multi-point, and continuous actions, testing whether coordinate-free grounding scales to richer computer-use behavior.
- Paper: FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation, Kaiyi Huang et al. (2026). The 2026 AGUVIS work continues visual GUI-agent research with cross-platform grounding and planning, offering a broader agent-level direction after GUI-Actor’s grounding advances.
- Paper: Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models, Marcel Gropl et al. (2026). Entropy-gradient grounding extends patch-level localization by extracting decision-relevant visual regions from VLMs without task-specific training.
