GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Qianhui WuKanzhi ChengRui YangChaoyun ZhangJianwei YangHuiqiang JiangJian MuBaolin PengBo QiaoReuben Tan

article2025NeurIPS96 citations

Presents GUI-Actor, a coordinate-free visual grounding framework for GUI agents that replaces text coordinate generation with an attention-based action head and verifier, enabling a 7B vision-language model to outperform 72B baselines on interface benchmarks while training only 100M parameters.

Listen

Autonomous graphical user interface (GUI) agents powered by vision-language models (VLMs) have gained significant traction for automating complex tasks across mobile, desktop, and web environments. A core capability of these agents is visual grounding—the ability to map natural language instructions directly to actionable screen regions. Prevailing approaches formulate grounding as text-based coordinate generation, where the model outputs numeric pixel coordinates. However, this text-based paradigm suffers from weak spatial-semantic alignment, overly penalizes valid interaction variations via single-point supervision, and creates a mismatch between coarse visual patch representations and dense coordinate targets.

The article demonstrates GUI-Actor, a coordinate-free framework designed to achieve accurate, human-like visual grounding. The system introduces an attention-based action head anchored by a dedicated context token, which attends directly over screenshot image patches rather than predicting textual numbers. The model is trained using spatial-aware multi-patch supervision, which treats all patches covered by ground-truth bounding boxes as positive targets to accommodate natural interaction ambiguity. In addition, the framework incorporates a lightweight grounding verifier that evaluates candidate regions generated in a single model forward pass to select the most plausible action before execution. The approach was evaluated across major benchmarks including ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, and a 49-task live Windows desktop testbed (OS-World-W) using roughly one million training screenshots.

The experimental findings show substantial improvements in performance and efficiency over existing methods. On the high-resolution, out-of-domain ScreenSpot-Pro benchmark, the 7-billion-parameter GUI-Actor achieves an accuracy of 40.7 to 44.6 (depending on the vision-language backbone), outperforming much larger 72-billion-parameter coordinate-based systems such as UI-TARS-72B (38.1). The smaller 2-billion-parameter GUI-Actor (36.7) also surpasses several competing 7-billion models. Adding the grounding verifier boosts performance further, reaching 44.2 on ScreenSpot-Pro for GUI-Actor-7B. Crucially, the system achieves peak benchmark accuracy with only about 60% of the training data required by baseline models. In live Windows environment testing, GUI-Actor-7B achieved a 12.2% task success rate, outperforming previous visual grounding baselines such as Aguvis-7B (4.0%), NAVI (10.2%), and OmniAgent (10.2%).

These results demonstrate that direct spatial attention is structurally superior to text-coordinate generation for computer control tasks. By grounding actions at the native resolution of the vision backbone, the method exhibits strong generalization to unseen screen layouts and resolutions while eliminating costly multi-turn sampling during inference. Furthermore, experiments confirm that freezing the underlying vision-language model and training only the lightweight action head (approximately 100 million parameters for a 7-billion model) achieves competitive accuracy. This lightweight training option provides a cost-effective pathway to equip general-purpose models with interface-control capabilities without degrading their core language and reasoning strengths.

Organizations developing or deploying automated computer agents should adopt coordinate-free, attention-based action heads and patch-level verifiers to reduce training compute, lower inference latency, and enhance execution reliability across varied screen resolutions. A noted limitation of the framework involves very small interface elements (such as icons under 10 by 10 pixels), which can fall between coarse 28-by-28-pixel visual patches. While patch-clustering mitigates this, future work should focus on higher perceptual visual resolutions and fine-grained spatial offsets to support precision-demanding applications like industrial design and computer-aided engineering software.

Cover for GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Abstract

One of the principal challenges in building VLM-powered GUI agents is visual grounding, i.e., localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these approaches suffer from several limitations: weak spatial-semantic alignment, inability to handle ambiguous supervision targets, and a mismatch between the dense nature of screen coordinates and the coarse, patch-level granularity of visual features extracted by models like Vision Transformers. In this paper, we propose GUI-Actor, a VLM-based method for coordinate-free GUI grounding. At its core, GUI-Actor introduces an attention-based action head that learns to align a dedicated <ACTOR> token with all relevant visual patch tokens, enabling the model to propose one or more action regions in a single forward pass. In line with this, we further design a grounding verifier to evaluate and select the most plausible action region from the candidates proposed for action execution. Extensive experiments show that GUI-Actor outperforms prior state-of-the-art methods on multiple GUI action grounding benchmarks, with improved generalization to unseen screen resolutions and layouts. Notably, GUI-Actor-7B even surpasses UI-TARS-72B (38.1) on ScreenSpot-Pro, achieving scores of 40.7 with Qwen2-VL and 44.6 with Qwen2.5-VL as backbones. Furthermore, by incorporating the verifier, we find that fine-tuning only the newly introduced action head (~100M parameters for 7B model) while keeping the VLM backbone frozen is sufficient to achieve performance comparable to previous state-of-the-art models, highlighting that GUI-Actor can endow the underlying VLM with effective grounding capabilities without compromising its general-purpose strengths.

Citation

MLA
Wu, Q., et al. “GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents”. arXiv, 2025, http://arxiv.org/abs/2506.03143v1.
APA
Wu, Q., Cheng, K., Yang, R., Zhang, C., Yang, J., Jiang, H., Mu, J., Peng, B., Qiao, B., Tan, R., Qin, S., Liden, L., Lin, Q., Zhang, H., Zhang, T., Zhang, J., Zhang, D., & Gao, J. (2025). GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents. arXiv. http://arxiv.org/abs/2506.03143v1
Chicago
Wu, Q., K. Cheng, R. Yang, et al. 2025. “GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents”. arXiv. http://arxiv.org/abs/2506.03143v1.
Harvard
Wu, Q. et al. (2025) “GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2506.03143v1.
Vancouver
1. Wu Q, Cheng K, Yang R, et al (2025) GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents. arXiv

BibTeX

@article{wu2025gui,
  title = {GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents},
  author = {Wu, Qianhui and Cheng, Kanzhi and Yang, Rui and Zhang, Chaoyun and Yang, Jianwei and Jiang, Huiqiang and Mu, Jian and Peng, Baolin and Qiao, Bo and Tan, Reuben and Qin, Si and Liden, Lars and Lin, Qingwei and Zhang, Huan and Zhang, Tong and Zhang, Jianbing and Zhang, Dongmei and Gao, Jianfeng},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2506.03143v1},
  eprint = {2506.03143}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/