Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
Miaosen ZhangXiaohan ZhaoZhihong TanHuoshen ZhouYijia FanYifan YangKai QiuBei LiuJustin WagleChenzhong Yin
Presents an automated data-synthesis pipeline and the multimodal CUActSpot benchmark to tackle long-tail interaction failures in computer-use agents, producing a 4B model that outperforms open-source alternatives with fewer than 32B parameters.
Digital automation through computer-use agents promises to transform workplace productivity, yet these systems frequently fail when performing complex on-screen actions. In professional applications such as spreadsheets, document editors, and graphic design software, agent failures disproportionately stem from action grounding—the core visual capability of identifying exact screen coordinates to execute commands. Existing industry benchmarks and datasets remain narrowly focused on simple, single-click operations on standard interface buttons. The article investigates the practical bottlenecks that hinder GUI-based computer-use agents and demonstrates how broadening action modeling and utilizing diverse synthetic data can resolve this reliability gap.
To address this challenge, the article introduces CUActSpot, a diagnostic benchmark spanning five interaction modalities: standard interfaces, text documents, spreadsheets, design canvases, and natural images. Unlike click-centric assessments, CUActSpot evaluates multi-point, ordered, and continuous operations, such as highlighting text spans, dragging table borders, and tracing image boundaries. Alongside the benchmark, the authors developed a code-based rendering pipeline that procedurally creates diverse digital workspaces, extracts precise coordinate metadata, and uses advanced language models to synthesize complex natural language instructions and mouse action traces, producing a 50-million-sample training corpus. Using this corpus, the authors trained Phi-Ground-Any-4B, a four-billion-parameter visual model tailored for general computer use.
The investigation yields four key findings. First, scaling data volume within a single modality produced diminishing returns, whereas expanding task and modality diversity substantially improved general performance across all interactions—a principle termed variety scaling. Second, Phi-Ground-Any-4B achieved an overall score of 44.4 percent on the CUActSpot benchmark, outperforming all open-source models with fewer than 32 billion parameters. Third, the model exhibited cross-task generalization, successfully completing 27 detailed task types on the benchmark despite being trained on only 20 synthetic categories. Finally, high performance on legacy click benchmarks failed to predict success in end-to-end realistic environments, whereas grounding accuracy on CUActSpot closely aligned with real-world computer task execution.
These results indicate that enterprise deployment of computer-use agents has been hindered by a mismatch between benchmark design and actual operational requirements. For leaders developing or deploying automation agents, expanding training data across diverse interaction types offers a cost-effective pathway to enhance agent reliability without inflating model size. Organizations should update their evaluation metrics to include multi-step, non-widget operations and leverage procedural data generation to train agents on complex software interactions. Future work should focus on closing the residual gap between synthetic and native software distributions and extending evaluation to long-horizon, stateful workflows.
Confidence in these findings is reinforced by rigorous ablation studies and consistent performance across simulated environments. However, decision-makers should recognize that CUActSpot consists of a curated sample of 206 diagnostic tasks and does not capture every long-term state change encountered in live enterprise systems. Prudent next steps involve conducting controlled pilot tests in target workplace applications to validate agent execution before full-scale autonomous deployment.
- Paper: Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction, Yiheng Xu et al. (2025). It provides a foundational pure-vision framework and grounding paradigm for autonomous GUI interaction that CUActSpot builds upon to address complex, non-click action spaces.
- Paper: ShowUI: One Vision-Language-Action Model for GUI Visual Agent, Kevin Qinghong Lin et al. (2025). It establishes a lightweight vision-language-action grounding architecture for screenshot-based GUI automation, motivating the need for richer synthetic training data across diverse interaction types.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). It introduces high-resolution visual language modeling for visual grounding in GUI agents, setting the standard for visual perception upon which CUActSpot and Phi-Ground-Any-4B build.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). It pioneers learning directly from raw visual pixels to low-level UI mouse and keyboard actions, establishing the direct action-space formulation that the source paper expands beyond simple clicks.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). It introduces synthetic trajectory and task generation pipelines for training autonomous agents in realistic digital environments, preceding the renderer-based synthesis proposed in the source.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). It develops pixel-only coordinate grounding on raw UI screenshots without view hierarchies, providing prerequisite techniques for visual element spotting.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). It demonstrates end-to-end multimodal agent web navigation based on visual screenshots and action execution, contextualizing the benchmark and modeling challenges addressed by CUActSpot.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). It introduces the standard benchmark and multi-turn action framework for generalist web agents, highlighting the early focus on click-centric widget navigation.
- Paper: ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots, Yu-Chung Hsiao et al. (2025). It establishes visual question answering and bounding-box grounding directly over raw screenshot pixels, serving as essential background for visual element localization.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). It applies reinforcement learning with concise action-location reward functions to improve GUI action prediction efficiency, extending the supervised visual grounding models introduced in the source.
- Paper: Neural Computers, Mingchen Zhuge et al. (2026). It generalizes screen-action interaction traces into unified video-generative neural computers that simulate runtime interfaces directly from user actions and visual frames.
