Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Miaosen ZhangXiaohan ZhaoZhihong TanHuoshen ZhouYijia FanYifan YangKai QiuBei LiuJustin WagleChenzhong Yin

article2026arXiv0 citations

Presents an automated data-synthesis pipeline and the multimodal CUActSpot benchmark to tackle long-tail interaction failures in computer-use agents, producing a 4B model that outperforms open-source alternatives with fewer than 32B parameters.

Listen

Digital automation through computer-use agents promises to transform workplace productivity, yet these systems frequently fail when performing complex on-screen actions. In professional applications such as spreadsheets, document editors, and graphic design software, agent failures disproportionately stem from action grounding—the core visual capability of identifying exact screen coordinates to execute commands. Existing industry benchmarks and datasets remain narrowly focused on simple, single-click operations on standard interface buttons. The article investigates the practical bottlenecks that hinder GUI-based computer-use agents and demonstrates how broadening action modeling and utilizing diverse synthetic data can resolve this reliability gap.

To address this challenge, the article introduces CUActSpot, a diagnostic benchmark spanning five interaction modalities: standard interfaces, text documents, spreadsheets, design canvases, and natural images. Unlike click-centric assessments, CUActSpot evaluates multi-point, ordered, and continuous operations, such as highlighting text spans, dragging table borders, and tracing image boundaries. Alongside the benchmark, the authors developed a code-based rendering pipeline that procedurally creates diverse digital workspaces, extracts precise coordinate metadata, and uses advanced language models to synthesize complex natural language instructions and mouse action traces, producing a 50-million-sample training corpus. Using this corpus, the authors trained Phi-Ground-Any-4B, a four-billion-parameter visual model tailored for general computer use.

The investigation yields four key findings. First, scaling data volume within a single modality produced diminishing returns, whereas expanding task and modality diversity substantially improved general performance across all interactions—a principle termed variety scaling. Second, Phi-Ground-Any-4B achieved an overall score of 44.4 percent on the CUActSpot benchmark, outperforming all open-source models with fewer than 32 billion parameters. Third, the model exhibited cross-task generalization, successfully completing 27 detailed task types on the benchmark despite being trained on only 20 synthetic categories. Finally, high performance on legacy click benchmarks failed to predict success in end-to-end realistic environments, whereas grounding accuracy on CUActSpot closely aligned with real-world computer task execution.

These results indicate that enterprise deployment of computer-use agents has been hindered by a mismatch between benchmark design and actual operational requirements. For leaders developing or deploying automation agents, expanding training data across diverse interaction types offers a cost-effective pathway to enhance agent reliability without inflating model size. Organizations should update their evaluation metrics to include multi-step, non-widget operations and leverage procedural data generation to train agents on complex software interactions. Future work should focus on closing the residual gap between synthetic and native software distributions and extending evaluation to long-horizon, stateful workflows.

Confidence in these findings is reinforced by rigorous ablation studies and consistent performance across simulated environments. However, decision-makers should recognize that CUActSpot consists of a curated sample of 206 diagnostic tasks and does not capture every long-term state change encountered in live enterprise systems. Prudent next steps involve conducting controlled pilot tests in target workplace applications to validate agent execution before full-scale autonomous deployment.

  • Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). It applies reinforcement learning with concise action-location reward functions to improve GUI action prediction efficiency, extending the supervised visual grounding models introduced in the source.
  • Paper: Neural Computers, Mingchen Zhuge et al. (2026). It generalizes screen-action interaction traces into unified video-generative neural computers that simulate runtime interfaces directly from user actions and visual frames.
Cover for Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Abstract

Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions is still poor, limiting user trust. Our analysis of failure cases from advanced models suggests a long-tail pattern in GUI operations, where a relatively small fraction of complex and diverse interactions accounts for a disproportionate share of task failures. We hypothesize that this issue largely stems from the scarcity of data for complex interactions. To address this problem, we propose a new benchmark CUActSpot for evaluating models' capabilities on complex interactions across five modalities: GUI, text, table, canvas, and natural image, as well as a variety of actions (click, drag, draw, etc.), covering a broader range of interaction types than prior click-centric benchmarks that focus mainly on GUI widgets. We also design a renderer-based data-synthesis pipeline: scenes are automatically generated for each modality, screenshots and element coordinates are recorded, and an LLM produces matching instructions and action traces. After training on this corpus, our Phi-Ground-Any-4B outperforms open-source models with fewer than 32B parameters. We will release our benchmark, data, code, and models at this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 CUActSpot Benchmark
  • 3.1 Evaluation Rules and Metrics
  • 3.2 Benchmark Statistics
  • 4 General Action Grounding Data Synthetic Pipeline
  • 4.1 General Synthetic Pipeline
  • 4.2 GUI Element and Table Modal
  • 4.3 Text and Canvas Modal
  • 4.4 Natural Image Modal
  • 5 Experiments and Evaluations
  • 5.1 Training Details
  • 5.2 Benchmarks Studies
  • 5.3 Empirical Studies and Ablations
  • 6 Conclusions and Limitations
  • References
  • A CUActSpot Details
  • A.1 Detailed Tasks Breakdown
  • A.2 Benchmark examples
  • B More Training Details
  • B.1 Data sampling
  • B.2 Data ablation results
  • C Data Synthesis Details
  • C.1 GUI
  • C.2 Text
  • C.3 Table
  • C.4 Canvas
  • C.5 Natural Image
  • D Case study

Knowls

  1. Knowl 1 — CUActSpot Benchmark Structure and Modality Taxonomy

    definition

    The CUActSpot benchmark is an evaluation suite designed to assess computer-use agents (CUAs) on complex on-screen mouse interactions across diverse modalities beyond standard GUI widgets. The benchmark comprises 206 manually verified diagnostic samples covering 12 high-level task types and 33 fine-grained interaction types across five distinct visual modalities:

    • GUI: Standard interface widgets (e.g., buttons, slide bars, tabs, file icons, and window borders), evaluated on single clicks and drag-and-drop operations.
    • Text: Document and text editor operations (e.g., in word processors or code editors), including placing the text insertion cursor between characters, dragging to select short or long text spans, and dragging text snippets.
    • Table: Spreadsheet manipulation (e.g., in Microsoft Excel or LibreOffice Calc), including selecting blank or content cells, dragging cell ranges, dragging column or row border headers to resize, and dragging cell fill-handles (bottom-right corners) to auto-fill formulas.
    • Canvas: Vector and graphic editing (e.g., in Microsoft PowerPoint), including selecting underlying shape layers, dragging shapes, connecting nodes with arrows, dragging bounding-box control points or rotation handles, and drawing multi-point closed polygons.
    • Natural Image: Bitmap image editing (e.g., in Adobe Photoshop or GIMP), including single-click foreground/background selection, dragging image layers, cropping bounding regions, face/body contour warping, and tracing multi-point polygonal cutout boundaries or zigzag scribble paths for erasure masks.

    Task complexity is categorized by the number of key points involved: 1-point (clicks), 2-point (drags), and NN-point (multi-point drawing/tracing), with explicit distinctions between order-sensitive and order-insensitive operations.

  2. Knowl 2 — CUActSpot Evaluation Metric and Correctness Rules

    algorithm

    The CUActSpot benchmark evaluates predicted mouse actions against annotated regions on screen. Each evaluation sample specifies two classes of ground-truth regions:

    • Correct Region: One or more valid target areas where predicted coordinates (e.g., click points, drag endpoints, or polygon vertices) must land. Correct regions can optionally carry an integer rank attribute k∈Nk \in \mathbb{N} to define order-sensitive action sequences (e.g., dragging along a directional arrow or tracing polygon vertices in order). If no rank is provided, the action is order-insensitive.
    • Banned Region: Screen regions where no predicted coordinate is permitted to fall, preventing artificial score inflation through unconstrained random clicking.

    For a sample with predicted coordinate sequence P=(p1,p2,…,pM)\mathcal{P} = (p_1, p_2, \dots, p_M), correctness is determined by applying the following hierarchical decision rules in order of priority:

    Input: Predicted key points P=(p1,…,pM)\mathcal{P} = (p_1, \dots, p_M), Correct regions Rcorrect\mathcal{R}_{\text{correct}}, Banned regions Rbanned\mathcal{R}_{\text{banned}}
    Output: Boolean sample success S∈{True,False}S \in \{\text{True}, \text{False}\}
    if Rbanned≠∅\mathcal{R}_{\text{banned}} \neq \emptyset then
        for each point p∈Pp \in \mathcal{P} do
            for each region B∈RbannedB \in \mathcal{R}_{\text{banned}} do
                if p∈Bp \in B then
                    return False
    if all R∈RcorrectR \in \mathcal{R}_{\text{correct}} have rank attributes then
        Let KK be the number of distinct ranks {1,2,…,K}\{1, 2, \dots, K\}
        if M≠KM \neq K then
            return False
        for k=1k = 1 to KK do
            Let Rk={R∈Rcorrect∣rank(R)=k}\mathcal{R}_k = \{R \in \mathcal{R}_{\text{correct}} \mid \text{rank}(R) = k\}
            if not (∃R∈Rk such that pk∈R)(\exists R \in \mathcal{R}_k \text{ such that } p_k \in R) then
                return False
        return True
    else
        for each region R∈RcorrectR \in \mathcal{R}_{\text{correct}} do
            if not (∃p∈P such that p∈R)(\exists p \in \mathcal{P} \text{ such that } p \in R) then
                return False
        return True

    The benchmark metric is the overall sample success rate, defined as the proportion of test instances satisfying all applicable evaluation rules.

  3. Knowl 3 — Procedural Rendering and LLM-Driven Coordinate Synthesis Pipeline

    model/method

    To synthesize large-scale training data for complex computer-use operations without manual annotation, a code-based rendering and LLM task design pipeline is used:

    1. Procedural Rendering & Ground-Truth Extraction: Synthetic user interface scenes are rendered programmatically using modality-specific rendering engines (Playwright, Selenium, PyQt5, Matplotlib, and SAM segmentation masks). Because visual elements (e.g., buttons, table cells, characters, shapes, control points, contours) are rendered via code, exact pixel coordinates, bounding boxes [x1,y1,x2,y2][x_1, y_1, x_2, y_2], and control handles are extracted directly from the rendering tree without visual annotation error.
    2. LLM Task and Trajectory Synthesis: The screenshot and the structured dictionary of extracted element metadata are supplied to an advanced reasoning LLM (OpenAI o3). The model is prompted with modality-specific guidelines to formulate natural language instructions and executable PyAutoGUI action scripts (such as pyautogui.click, pyautogui.moveTo, pyautogui.dragTo, pyautogui.mouseDown, and pyautogui.mouseUp).
    3. Symbolic Coordinate Mapping & Intermediate Calculation: The LLM constructs actions using symbolic coordinate tokens (e.g., x1, y1, x2, y2) paired with a separate coordinate_map dictionary. The reasoning LLM is prompted to perform intermediate algebraic calculations over geometric metadata—for example, computing midpoints, resizing bounding boxes, aligning shape anchors (e.g., moving the center of an arrow (x1,y1)(x_1, y_1) with tip (x1,yc)(x_1, y_c) to align with the top (x2,yt)(x_2, y_t) of an ellipse by setting the new center to (x2,yt+y1−yc)(x_2, y_t + y_1 - y_c)), or ordering polygon boundary vertices.

    Across all modalities, this pipeline generated 50 million training samples (30M GUI, 5M Text, 5M Table, 5M Canvas, 5M Natural Image).

  4. Knowl 4 — Modality-Specific Synthesis Pipelines for GUI, Text, Table, Canvas, and Images

    model/method

    The data synthesis framework implements five modality-specific procedural rendering pipelines:

    • GUI Pipeline: Extracts 2.6 billion URLs from CommonCrawl (CC-MAIN-2024-46), deduplicates by English language and HTTP status, and caps pages per domain at 50 (yielding 475.45M URLs). Pages are rendered with Selenium and Chrome Driver across 1080p, 2K, and 4K resolutions and aspect ratios between 2:1 and 1:2. Role-based filtering preserves interactive widgets (buttons, input fields) across 73.5M pages. Elements are spatially discretized and sampled (prioritizing icon elements), labeled with GPT-4o captions, and transformed into multi-step actions using OpenAI o3.
    • Text Pipeline: Employs PyQt5 to render text from Wikipedia and GitHub across 2,500 open-source English fonts on ~200 real background images (e.g., Microsoft Word documents, Notepad windows) with randomized font sizes, colors, and weights. The pipeline generates six scenarios across code and natural language: dragging to select short text spans, dragging to select long spans, and clicking to position insertion carets.
    • Table Pipeline: Begins with 16,000 TableVQA seed tables, which are modified by an LLM to alter topics and introduce topological variations (such as merged cells and multi-column headers), creating 160,000 unique HTML tables. An LLM generates 1,000 CSS templates with randomized colors, borders, and fonts (yielding 10,000 CSS instances). Half of the rendered tables undergo cell masking (replacing cell contents with empty cells). Coordinate bounds, row/column indices, and headers are extracted via JavaScript.
    • Canvas Pipeline: A procedural PowerPoint-style simulator renders 76 primitive shape types across 9 categories (rectangles, ellipses, triangles, quadrilaterals, polygons, stars, arrows, lines/connectors, callouts/decorations) on canvases of width W∈[800,2560]W \in [800, 2560] and height H∈[600,1440]H \in [600, 1440]. Colors are sampled in HSV space enforcing redmean perceptual Euclidean distances ≥100\ge 100 against the background and ≥60\ge 60 between fill and outline. Overlap between shapes is constrained to ≤0.25\le 0.25. Rendered elements receive slide-editor selection chrome: 8 bounding box control points, vertex markers, and rotation handles. Referring expressions are uniquely generated using a 44-color palette, relative size qualifiers, and 3×33 \times 3 to 5×55 \times 5 spatial grid locators.
    • Natural Image Pipeline: Uses images and masks from Segment Anything (SAM). GPT-4o generates fine-grained region captions. The Suzuki-Abe border following algorithm extracts 20-point polygonal contours from masks. OpenAI o3 generates multi-point drawing, lasso selection, and zigzag scribble erasing trajectories.
  5. Knowl 5 — Phi-Ground-Any-4B Architecture and Training Configuration

    experimental setup

    Phi-Ground-Any-4B is a vision-language action grounding model built upon the Phi-3.5-VL backbone (4 billion parameters), chosen because it has not undergone prior GUI-specific pretraining.

    Model Inputs and Optimization Hyperparameters:

    • Visual inputs are partitioned into a maximum of 16 image crops with data augmentations applied.
    • Effective batch size: 5,120.
    • Learning rate: 8×10−58 \times 10^{-5}.
    • Weight decay: 0.01.
    • Gradient clipping: 0.1.
    • Total training budget: approximately 100 billion tokens, executed over 30 hours on an 80 ×\times NVIDIA H100 GPU cluster.
    • Checkpoint selection: models are checkpointed every 100 training steps, and the best-performing checkpoint is reported.

    Training Data Composition and Weighting:

    Dataset Modality Available Samples Used Samples Epochs Sampling Weight
    GUI 30,432,242 6,800,000 0.223 0.34
    Text 6,083,400 5,000,000 0.822 0.25
    Table 5,242,630 2,000,000 0.381 0.10
    Canvas 4,323,253 2,000,000 0.463 0.10
    Natural Image 4,743,675 3,000,000 0.632 0.15
    OpenCUA 340,665 1,200,000 3.523 0.06

    OpenCUA human-annotated data is upweighted through multi-epoch repetition (3.52 epochs) due to its high quality.

  6. Knowl 6 — Grounding Performance Across Modalities on CUActSpot and Baseline Benchmarks

    data/table

    The table below compares grounding models on ScreenSpot-Pro (SS-pro), UI-Vision (UI-V), their absolute performance gap (Δ=SS-pro−UI-V\Delta = \text{SS-pro} - \text{UI-V}), and the five modality splits of CUActSpot (GUI, Text, Table, Canvas, Image, and Overall success rate in %).

    Model Date SS-pro UI-V Δ\Delta CUActSpot
    GUI Text Table Canvas Image Overall
    Phi-Ground-4B-16C 2025-07 38.0 24.5 13.5 5.3 6.2 6.2 4.7 2.4 5.0
    Uground-V1-2B 2024-10 27.1 12.8 14.3 10.5 0.0 9.4 6.2 0.0 5.2
    Uground-V1-7B 2024-10 31.1 12.9 18.2 18.4 0.0 3.1 9.4 2.4 6.7
    OS-Atlas-Base-7B 2024-10 18.9 9.0 9.9 15.8 0.0 12.5 10.9 0.0 7.8
    InfiGUI-R1-3B 2025-04 45.2 22.0 23.2 23.7 3.1 9.4 7.8 0.0 8.8
    UI-Venus-Ground-7B 2025-08 50.8 26.5 24.3 23.7 3.1 18.8 9.4 0.0 11.0
    GUI-G2^2-7B 2025-07 47.5 26.4 21.1 23.7 6.2 15.6 7.8 4.8 11.6
    MAI-UI-2B 2025-12 57.4 30.3 27.1 18.4 3.1 18.8 12.5 9.5 12.5
    GUI-Owl-1.5-8B-Think 2026-02 57.6 33.2 24.4 23.7 9.4 18.8 10.9 7.1 14.0
    MAI-UI-8B 2025-12 65.8 40.7 25.1 26.3 18.8 18.8 7.8 4.8 15.3
    GUI-Owl-1.5-8B-Instruct 2026-02 71.1 37.4 33.7 23.7 15.6 18.8 9.4 9.5 15.4
    UI-Venus-Ground-72B 2025-08 61.9 36.8 25.1 28.9 18.8 18.8 10.9 9.5 17.4
    InfiGUI-G1-7B 2025-08 51.9 26.1 25.8 44.7 18.8 37.5 9.4 4.8 23.0
    EvoCUA-8B 2026-01 45.4 15.6 29.8 18.4 40.6 34.4 9.4 16.7 23.9
    UI-TARS-1.5-7B 2025-04 42.6 22.3 20.3 42.1 28.1 34.4 14.1 23.8 28.5
    EvoCUA-32B 2026-01 49.8 20.9 28.9 28.9 31.2 40.6 25.0 16.7 28.5
    OpenCUA-7B 2025-08 50.0 25.5 24.5 42.1 37.5 53.1 28.1 38.1 39.8
    Phi-Ground-Any-4B (ours) 2026-05 26.3 15.8 10.5 44.7 34.4 68.8 40.6 33.3 44.4
    + APP data finetuned 2026-05 41.5 29.7 11.8 52.6 18.8 59.4 32.8 19.0 36.5
    OpenCUA-32B 2025-08 55.3 26.3 29.0 55.3 46.9 68.8 39.1 52.4 52.5
    GPT-5.4 (Azure) 2026-03 44.5 37.9 6.6 73.7 43.8 87.5 65.6 47.6 63.6

    Phi-Ground-Any-4B achieves 44.4% overall on CUActSpot, outperforming all open-source models with fewer than 32B parameters (including UI-TARS-1.5-7B at 28.5% and OpenCUA-7B at 39.8%). Traditional GUI models (e.g., GUI-Owl-1.5-8B-Instruct and MAI-UI-8B) score high on ScreenSpot-Pro (71.1% and 65.8%) but drop sharply on CUActSpot (15.4% and 15.3%), demonstrating their inability to generalize to non-widget modalities and multi-point actions. When fine-tuned on application-specific click data (+ APP data finetuned), ScreenSpot-Pro performance increases from 26.3% to 41.5% while CUActSpot overall performance decreases from 44.4% to 36.5%.

  7. Knowl 7 — Controlled Grounding Evaluation on OSWorld vs. Click Benchmarks

    data/table

    To isolate action grounding performance in end-to-end computer-use tasks, OSWorld agentic tasks were evaluated under a controlled setup: GPT-5.4 was used as the high-level task planner to generate single-step natural language instructions, while different dedicated grounding models predicted the exact action parameters. The maximum allowed actions per task was capped at 30.

    Planner Grounder ScreenSpot-Pro (%) OSWorld Success (%)
    GPT-5.4 GUI-Owl-1.5-8B-Instruct 71.1 37.7
    GPT-5.4 MAI-UI-8B 65.8 38.2
    GPT-5.4 GPT-5.4 44.5 44.1
    GPT-5.4 Phi-Ground-Any-4B 26.3 42.4

    Models that score highest on single-click GUI benchmarks like ScreenSpot-Pro (GUI-Owl-1.5-8B-Instruct at 71.1% and MAI-UI-8B at 65.8%) achieve lower agentic success on OSWorld (37.7% and 38.2%). In contrast, Phi-Ground-Any-4B, despite obtaining a low ScreenSpot-Pro score of 26.3% due to the lack of specialized application icon memorization, attains 42.4% on OSWorld, closely matching GPT-5.4 (44.1%). This demonstrates that complex, multi-modal grounding capability (as measured by CUActSpot) correlates more closely with real-world agent execution than widget-only click benchmarks.

  8. Knowl 8 — Variety Scaling vs. Single-Modality Data Scaling in Action Grounding

    empirical result

    Ablation experiments demonstrate that increasing data diversity across modalities and interaction types (variety scaling) is significantly more effective than scaling dataset size within a single modality for training computer-use grounding models:

    1. Single-Modality Saturation: Training solely on increasing volumes of GUI web data (from 500K up to 3,000K samples) results in plateauing overall benchmark performance with minimal gains across non-GUI categories.
    2. Cross-Modality Transfer: Sequentially expanding the training distribution by adding distinct modalities not only boosts performance on the newly added modality but also yields positive transfer to previously added modalities:
      • Starting with GUI (2M samples): Overall CUActSpot = 14.8%, ScreenSpot-Pro = 16.4%.
      • Adding Text (1M samples): Text increases by +24.9% (to 31.3%), GUI increases by +2.6% (to 34.2%), Table by +6.2% (to 28.1%); Overall = 21.5%.
      • Adding Table (1M samples): Table increases by +12.5% (to 40.6%), Canvas by +1.5%; Overall = 22.5%.
      • Adding Canvas (1M samples): Canvas increases by +14.1% (to 25.0%), Image by +7.2% (to 16.7%), GUI by +5.2%; Overall = 28.5%.
      • Adding Natural Image (1M samples): Image increases by +7.1% (to 23.8%), Canvas by +4.7% (to 29.7%), Table by +3.1% (to 50.0%); Overall = 31.6%.
      • Adding OpenCUA (0.5M samples): GUI increases by +7.9%, Canvas by +7.8%, Table by +6.3%, Image by +2.4%; Overall reaches 37.1%, and ScreenSpot-Pro reaches 24.6%.

    This indicates that task and modality diversity is a primary driver of generalizable visual interaction capabilities in computer-use models.

  9. Knowl 9 — Compositional Cross-Task Generalization Across Interaction Types

    empirical result

    When evaluated on the 33 detailed interaction tasks defined in the CUActSpot benchmark, Phi-Ground-Any-4B successfully completes at least one sample in 27 detailed tasks, despite the synthetic training corpus explicitly containing only 20 detailed task types:

    Task Source Number of Solved Detailed Tasks
    CUActSpot Total 33
    Synthetic Training Data Explicit Types 20
    Phi-Ground-Any-4B Solved 27

    This demonstrates that the model exhibits zero-shot compositional generalization. By learning visual grounding primitives across separate modalities (e.g., text manipulation and vector shape selection), the model generalizes to unobserved task compositions, such as editing text embedded inside presentation canvas shapes or selecting text segments within natural images.

  10. Knowl 10 — Limitations of CUActSpot Benchmark and Synthetic Training Data

    limitation

    The benchmark and data synthesis methodology present two primary limitations:

    • Diagnostic and Static Nature of CUActSpot: CUActSpot contains 206 manually curated isolated interaction instances designed for fine-grained diagnostic evaluation. It does not evaluate long-horizon, multi-turn stateful task workflows or interactive feedback loops where environment states change dynamically across actions.
    • Domain Gap and Alignment in Synthetic Data: While procedural rendering enables large-scale, controllable annotation of complex geometries, control handles, and drag trajectories, synthetic interface environments differ from real-world software desktop distributions. Additional alignment and fine-tuning on diverse proprietary software environments is required to close the remaining gap with human operators.

Coverage note — Verbatim prompt templates from Appendix C were omitted as raw strings, but their structural guidelines, coordinate calculation mechanics, and output formats were fully integrated into the synthetic pipeline knowls.

References

  1. 1.Anthropic. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. Technical report, Anthropic, October 2024. URL https://www.anthropic.com/news/3-5-models-and-computer-use.
  2. 2.OpenAI. Computer-Using Agent. Technical report, OpenAI, January 2025. URL https://openai.com/index/computer-using-agent/.
  3. 3.John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37:50528–50652, 2024.
  4. 4.Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. Openhands: An open platform for ai software developers as generalist agents. arXiv preprint arXiv:2407.16741, 2024.
  5. 5.Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s2: A compositional generalist-specialist framework for computer use agents. arXiv preprint arXiv:2504.00906, 2025.
  6. 6.OpenAI. Introducing GPT-5.4. Technical report, OpenAI, March 2026. URL https://openai.com/index/introducing-gpt-5-4/.
  7. 7.Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. An illusion of progress? assessing the current state of web agents. arXiv preprint arXiv:2504.01382, 2025.
  8. 8.Miaosen Zhang, Qi Dai, Yifan Yang, Jianmin Bao, Dongdong Chen, Kai Qiu, Chong Luo, Xin Geng, and Baining Guo. Magebench: Bridging large multimodal models to agents. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1415–1427, 2026.
  9. 9.Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, et al. Visualagentbench: Towards large multimodal models as visual foundation agents. arXiv preprint arXiv:2408.06327, 2024.
  10. 10.Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for gui agents. arXiv preprint arXiv:2410.05243, 2024.
  11. 11.Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, et al. Os-atlas: A foundation action model for generalist gui agents. arXiv preprint arXiv:2410.23218, 2024.
  12. 12.Miaosen Zhang, Ziqiang Xu, Jialiang Zhu, Qi Dai, Kai Qiu, Yifan Yang, Chong Luo, Tianyi Chen, Justin Wagle, Tim Franklin, et al. Phi-ground tech report: Advancing perception in gui grounding. arXiv preprint arXiv:2507.23779, 2025.
  13. 13.Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al. Ui-tars: Pioneering automated gui interaction with native agents. arXiv preprint arXiv:2501.12326, 2025.
  14. 14.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9313–9332, 2024.
  15. 15.Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua. Screenspot-pro: Gui grounding for professional high-resolution computer use. In Proceedings of the 33rd ACM International Conference on Multimedia, pages 8778–8786, 2025.
  16. 16.Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A Rodriguez, Montek Kalsi, Rabiul Awal, Nicolas Chapados, M Tamer Özsu, Aishwarya Agrawal, David Vazquez, et al. Ui-vision: A desktop-centric gui benchmark for visual perception and interaction. arXiv preprint arXiv:2503.15661, 2025.
  17. 17.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2024.
  18. 18.Yuhang Liu, Pengxiang Li, Congkai Xie, Xavier Hu, Xiaotian Han, Shengyu Zhang, Hongxia Yang, and Fei Wu. Infigui-r1: Advancing multimodal gui agents from reactive actors to deliberative reasoners. arXiv preprint arXiv:2504.14239, 2025.
  19. 19.Zhangxuan Gu, Zhengwen Zeng, Zhenyu Xu, Xingran Zhou, Shuheng Shen, Yunfei Liu, Beitong Zhou, Changhua Meng, Tianyu Xia, Weizhi Chen, et al. Ui-venus technical report: Building high-performance ui agents with rft. arXiv preprint arXiv:2508.10833, 2025.
  20. 20.Fei Tang, Zhangxuan Gu, Zhengxi Lu, Xuyang Liu, Shuheng Shen, Changhua Meng, Wen Wang, Wenqi Zhang, Yongliang Shen, Weiming Lu, et al. GUI-G2^2: Gaussian reward modeling for gui grounding. arXiv preprint arXiv:2507.15846, 2025.
  21. 21.Yuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li, Congkai Xie, Jiasheng Wang, Xueyu Hu, Xiaotian Han, Jianbo Yuan, Xinyao Wang, et al. Infigui-g1: Advancing gui grounding with adaptive exploration policy optimization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 32267–32275, 2026.
  22. 22.Hanzhang Zhou, Xu Zhang, Panrong Tong, Jianan Zhang, Liangyu Chen, Quyu Kong, Chenglin Cai, Chen Liu, Yue Wang, Jingren Zhou, et al. Mai-ui technical report: Real-world centric foundation gui agents. arXiv preprint arXiv:2512.22047, 2025.
  23. 23.Haiyang Xu, Xi Zhang, Haowei Liu, Junyang Wang, Zhaozai Zhu, Shengjie Zhou, Xuhao Hu, Feiyu Gao, Junjie Cao, Zihua Wang, et al. Mobile-agent-v3. 5: Multi-platform fundamental gui agents. arXiv preprint arXiv:2602.16855, 2026.
  24. 24.Tianbao Xie, Jiaqi Deng, Xiaochuan Li, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang, Yuhui Xu, Zekun Wang, et al. Scaling computer-use grounding via user interface decomposition and synthesis. arXiv preprint arXiv:2505.13227, 2025.
  25. 25.Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, et al. Opencua: Open foundations for computer-use agents. arXiv preprint arXiv:2508.09123, 2025.
  26. 26.Taofeng Xue, Chong Peng, Mianqiu Huang, Linsen Guo, Tiancheng Han, Haozhe Wang, Jianing Wang, Xiaocheng Zhang, Xin Yang, Dengchang Zhao, et al. Evocua: Evolving computer use agents via learning from scalable synthetic experience. arXiv preprint arXiv:2601.15876, 2026.
  27. 27.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v (ision) is a generalist web agent, if grounded. arXiv preprint arXiv:2401.01614, 2024.
  28. 28.Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023.
  29. 29.Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al. Cogagent: A visual language model for gui agents. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14281–14290, 2024.
  30. 30.Rui Qian, Xin Yin, Chuanhang Deng, Zhiyuan Peng, Jian Xiong, Wei Zhai, and Dejing Dou. Uground: Towards unified visual grounding with unrolled transformers. arXiv preprint arXiv:2510.03853, 2025.
  31. 31.Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. Aguvis: Unified pure vision agents for autonomous gui interaction. arXiv preprint arXiv:2412.04454, 2024.
  32. 32.Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Zechen Bai, Weixian Lei, Lijuan Wang, and Mike Zheng Shou. Showui: One vision-language-action model for generalist gui agent. In NeurIPS 2024 Workshop on Open-World Agents, 2024.
  33. 33.Yuhao Yang, Yue Wang, Dongxu Li, Ziyang Luo, Bei Chen, Chao Huang, and Junnan Li. Aria-ui: Visual grounding for gui instructions. In Findings of the Association for Computational Linguistics: ACL 2025, pages 22418–22433, 2025.
  34. 34.Haoming Wang, Haoyang Zou, Huatong Song, Jiazhan Feng, Junjie Fang, Junting Lu, Longxiang Liu, Qinyu Luo, Shihao Liang, Shijue Huang, et al. Ui-tars-2 technical report: Advancing gui agent with multi-turn reinforcement learning. arXiv preprint arXiv:2509.02544, 2025.
  35. 35.Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, et al. Windows agent arena: Evaluating multi-modal os agents at scale. arXiv preprint arXiv:2409.08264, 2024.
  36. 36.Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, et al. Androidworld: A dynamic benchmarking environment for autonomous agents. arXiv preprint arXiv:2405.14573, 2024.
  37. 37.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854, 2023.
  38. 38.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36:28091–28114, 2023.
  39. 39.OpenAI. Introducing OpenAI o3 and o4-mini. Technical report, OpenAI, April 2025. URL https://openai.com/index/introducing-o3-and-o4-mini/.
  40. 40.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023.
  41. 41.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  42. 42.Satoshi Suzuki et al. Topological structural analysis of digitized binary images by border following. Computer vision, graphics, and image processing, 30(1):32–46, 1985.
  43. 43.Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  44. 44.Yan Yang, Dongxu Li, Yutong Dai, Yuhao Yang, Ziyang Luo, Zirui Zhao, Zhiyuan Hu, Junzhe Huang, Amrita Saha, Zeyuan Chen, et al. Gta1: Gui test-time scaling agent. arXiv preprint arXiv:2507.05791, 2025.

Citation

MLA
Zhang, M., et al. “Covering Human Action Space for Computer Use: Data Synthesis and Benchmark”. arXiv, 2026, http://arxiv.org/abs/2605.12501v1.
APA
Zhang, M., Zhao, X., Tan, Z., Huoshen, Z., Fan, Y., Yang, Y., Qiu, K., Liu, B., Wagle, J., Yin, C., Cheng, M., Li, J., Dai, Q., Luo, C., Yang, X., Geng, X., & Guo, B. (2026). Covering Human Action Space for Computer Use: Data Synthesis and Benchmark. arXiv. http://arxiv.org/abs/2605.12501v1
Chicago
Zhang, M., X. Zhao, Z. Tan, et al. 2026. “Covering Human Action Space for Computer Use: Data Synthesis and Benchmark”. arXiv. http://arxiv.org/abs/2605.12501v1.
Harvard
Zhang, M. et al. (2026) “Covering Human Action Space for Computer Use: Data Synthesis and Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.12501v1.
Vancouver
1. Zhang M, Zhao X, Tan Z, et al (2026) Covering Human Action Space for Computer Use: Data Synthesis and Benchmark. arXiv

BibTeX

@article{zhang2026covering,
  title = {Covering Human Action Space for Computer Use: Data Synthesis and Benchmark},
  author = {Zhang, Miaosen and Zhao, Xiaohan and Tan, Zhihong and Huoshen, Zhou and Fan, Yijia and Yang, Yifan and Qiu, Kai and Liu, Bei and Wagle, Justin and Yin, Chenzhong and Cheng, Mingxi and Li, Ji and Dai, Qi and Luo, Chong and Yang, Xu and Geng, Xin and Guo, Baining},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.12501v1},
  eprint = {2605.12501}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/publicdomain/zero/1.0/