UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
Xinyi LiuXiaoyi ZhangZiyun ZhangYan Lu
Presents an automated instruction synthesis pipeline and benchmark that resolve data scarcity challenges in training vision-based agents to accurately ground complex graphical user interface elements.
Autonomous digital agents rely heavily on vision-based graphical user interface grounding to map natural language instructions directly to screen coordinates without depending on fragile application metadata. However, building reliable vision-language models for interface automation has been constrained by expensive manual data collection and existing benchmarks that do not reflect realistic operational complexities. Previous evaluations often overlook disproportionately large element-to-screen ratios, ignore long-tailed interface components like toggles and dropdowns, and depend almost entirely on direct, explicit commands rather than realistic implicit user requests.
The article aims to resolve these data and evaluation bottlenecks by developing an automated, large-scale instruction synthesis pipeline and introducing a realistic evaluation benchmark. It demonstrates that models trained on systematically synthesized data can achieve state-of-the-art interface grounding performance across web, desktop, and mobile operating environments.
To achieve this, the authors created UI-E2I-Synth, a multi-step data synthesis framework. The pipeline gathers interface screenshots across web, desktop, and mobile sources, applies heuristic parsing to extract reliable element attributes, and uses advanced vision models to generate both explicit and implicit referring expressions. It then parameterizes user actions to formulate realistic first-person instructions. This process produced a massive training dataset comprising approximately 1.6 million screenshots and 9.9 million instructions. In parallel, the authors established UI-I2E-Bench, an expert-validated benchmark featuring 1,477 multi-platform instructions designed with realistic element-to-screen proportions and high proportions of implicit commands.
The experimental findings show significant improvements across key benchmarks. Models fine-tuned on the synthetic data achieved a 9.7% relative improvement in overall grounding accuracy compared to prior state-of-the-art models, despite using roughly 28% less training data. On the new, more challenging UI-I2E-Bench, the seven-billion parameter model achieved an average accuracy of 69.5%, outperforming the previous leading baseline by 12.1 percentage points on implicit instructions. Evaluation on complex desktop applications in the ScreenSpot-Pro benchmark revealed a grounding accuracy of 23.6%, compared to 18.9% for the closest competitor. Furthermore, integrating the resulting vision model into an end-to-end task execution agent on the OSWorld benchmark improved the operational success rate from 3.6% to 12.0%.
These findings indicate that existing benchmarks have significantly overestimated the readiness of interface agents by evaluating them on overly simple, high-ratio text elements. Synthetic data generation offers a viable, highly cost-effective path to train robust visual agents without human labeling bottlenecks. In practice, enhancing grounding precision translates directly to higher reliability and fewer catastrophic execution errors in automated business workflows.
Organizations developing autonomous interface automation should transition away from brittle metadata scrapers toward robust visual grounding models, while deliberately rebalancing training distributions to emphasize non-text and long-tailed interface elements. Future development should focus on expanding synthetic instruction pipelines beyond English to multilingual settings, scaling base model capacities, and implementing chain-of-thought reasoning to resolve spatial hierarchies and icon recognition in specialized software environments.
Confidence in these findings is supported by consistent performance gains across multiple external benchmarks and thorough ablation studies. However, decision-makers should note that autonomous task execution rates remain modest overall, and models still exhibit vulnerabilities when encountering highly specialized application icons, deep interface hierarchies, or tasks requiring strict spatial counting.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). PIX2ACT establishes that agents can follow instructions and act from screenshots alone, providing a direct foundation for UI-E2I-Synth’s visual grounding task.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent develops visual perception and element grounding for GUI agents, clarifying the grounding challenges that UI-E2I-Synth targets with synthesized training data.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). Spotlight’s screenshot-based command grounding provides an earlier mobile-UI precedent for the visual localization task studied in UI-E2I-Synth.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Self-Instruct shows how language models can generate instruction-following training data, a key methodological precursor to UI-E2I-Synth’s synthetic instruction pipeline.
- Paper: Covering Human Action Space for Computer Use: Data Synthesis and Benchmark, Miaosen Zhang et al. (2026). CUActSpot extends synthetic GUI grounding data to diverse workspaces and complex actions, carrying UI-E2I-Synth’s data-scaling approach beyond element-level clicks.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). UI-R1 explores reinforcement learning as a complementary way to improve GUI action prediction and grounding beyond the synthesized supervised training studied here.
