UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis

Xinyi LiuXiaoyi ZhangZiyun ZhangYan Lu

article2025ACL22 citationsACL 2025 Best Theme Paper Award

Presents an automated instruction synthesis pipeline and benchmark that resolve data scarcity challenges in training vision-based agents to accurately ground complex graphical user interface elements.

Listen

Autonomous digital agents rely heavily on vision-based graphical user interface grounding to map natural language instructions directly to screen coordinates without depending on fragile application metadata. However, building reliable vision-language models for interface automation has been constrained by expensive manual data collection and existing benchmarks that do not reflect realistic operational complexities. Previous evaluations often overlook disproportionately large element-to-screen ratios, ignore long-tailed interface components like toggles and dropdowns, and depend almost entirely on direct, explicit commands rather than realistic implicit user requests.

The article aims to resolve these data and evaluation bottlenecks by developing an automated, large-scale instruction synthesis pipeline and introducing a realistic evaluation benchmark. It demonstrates that models trained on systematically synthesized data can achieve state-of-the-art interface grounding performance across web, desktop, and mobile operating environments.

To achieve this, the authors created UI-E2I-Synth, a multi-step data synthesis framework. The pipeline gathers interface screenshots across web, desktop, and mobile sources, applies heuristic parsing to extract reliable element attributes, and uses advanced vision models to generate both explicit and implicit referring expressions. It then parameterizes user actions to formulate realistic first-person instructions. This process produced a massive training dataset comprising approximately 1.6 million screenshots and 9.9 million instructions. In parallel, the authors established UI-I2E-Bench, an expert-validated benchmark featuring 1,477 multi-platform instructions designed with realistic element-to-screen proportions and high proportions of implicit commands.

The experimental findings show significant improvements across key benchmarks. Models fine-tuned on the synthetic data achieved a 9.7% relative improvement in overall grounding accuracy compared to prior state-of-the-art models, despite using roughly 28% less training data. On the new, more challenging UI-I2E-Bench, the seven-billion parameter model achieved an average accuracy of 69.5%, outperforming the previous leading baseline by 12.1 percentage points on implicit instructions. Evaluation on complex desktop applications in the ScreenSpot-Pro benchmark revealed a grounding accuracy of 23.6%, compared to 18.9% for the closest competitor. Furthermore, integrating the resulting vision model into an end-to-end task execution agent on the OSWorld benchmark improved the operational success rate from 3.6% to 12.0%.

These findings indicate that existing benchmarks have significantly overestimated the readiness of interface agents by evaluating them on overly simple, high-ratio text elements. Synthetic data generation offers a viable, highly cost-effective path to train robust visual agents without human labeling bottlenecks. In practice, enhancing grounding precision translates directly to higher reliability and fewer catastrophic execution errors in automated business workflows.

Organizations developing autonomous interface automation should transition away from brittle metadata scrapers toward robust visual grounding models, while deliberately rebalancing training distributions to emphasize non-text and long-tailed interface elements. Future development should focus on expanding synthetic instruction pipelines beyond English to multilingual settings, scaling base model capacities, and implementing chain-of-thought reasoning to resolve spatial hierarchies and icon recognition in specialized software environments.

Confidence in these findings is supported by consistent performance gains across multiple external benchmarks and thorough ablation studies. However, decision-makers should note that autonomous task execution rates remain modest overall, and models still exhibit vulnerabilities when encountering highly specialized application icons, deep interface hierarchies, or tasks requiring strict spatial counting.

arXiv: 2504.11257
Cover for UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis

Abstract

Recent advancements in Large Vision-Language Models are accelerating the development of Graphical User Interface (GUI) agents that utilize human-like vision perception capabilities to enhance productivity on digital devices. Compared to approaches predicated on GUI metadata, which are platform-dependent and vulnerable to implementation variations, vision-based approaches offer broader applicability. In this vision-based paradigm, the GUI instruction grounding, which maps user instruction to the location of corresponding element on the given screenshot, remains a critical challenge, particularly due to limited public training dataset and resource-intensive manual instruction data annotation. In this paper, we delve into unexplored challenges in this task including element-to-screen ratio, unbalanced element type, and implicit instruction. To address these challenges, we introduce a large-scale data synthesis pipeline UI-E2I-Synth for generating varying complex instruction datasets using GPT-4o instead of human annotators. Furthermore, we propose a new GUI instruction grounding benchmark UI-I2E-Bench, which is designed to address the limitations of existing benchmarks by incorporating diverse annotation aspects. Our model, trained on the synthesized data, achieves superior performance in GUI instruction grounding, demonstrating the advancements of proposed data synthesis pipeline. The proposed benchmark, accompanied by extensive analyses, provides practical insights for future research in GUI grounding. We will release corresponding artifacts at this https URL .

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Approach
  • 3.1 Raw Data Collection and Parsing
  • 3.2 Referring Expression Generation
  • 3.3 Instruction Synthesis
  • 4 UI-I2E-Bench: A Comprehensive Grounding Benchmark
  • 5 Experiments
  • 5.1 Training Details
  • 5.2 GUI Instruction Grounding Evaluation
  • 5.3 Agent Task Online Evaluation
  • 6 Conclusion
  • References
  • A Data Statistics
  • B Training Details
  • C Details of UI-I2E-Bench Annotation
  • D Ablation on Non-Text Element Proportion
  • E Related literature discussion
  • F Prompt templates
  • G Insights on failure cases in UI-I2E-Bench
  • H GPT-4o response example in Instruction Synthesis

Knowls

  1. Knowl 1 — UI-E2I-Synth generates user instructions through staged element and action synthesis

    model/method

    UI-E2I-Synth converts screenshot–metadata pairs into examples of the form screenshot, user instruction, and target-element coordinates. First, platform-specific metadata parsers extract an element’s type, content, and bounding box; the parsers favor high precision over recall, and element types are resampled to form a more balanced candidate pool. The five element types are Text (a button with visible text), Inputfield (an element for entering or editing content), Dropdown (a control for selecting among options), Icon (a graphical button), and Toggle (a two-state control such as a checkbox, radio button, or switch). Second, GPT-4o receives the extracted attributes and a Set-of-Marks screenshot as context, generates descriptions of element function and interaction outcome, and produces referring expressions. Explicit expressions identify visible or obvious features; implicit expressions identify an element by its function or its relationship to nearby elements without relying on those obvious features. Third, GPT-4o simulates user actions by generating action type and action content, then uses those parameters and an element expression to write a concise, first-person user instruction. This separates describing the target from expressing a plausible user intention, and yields varied instructions paired with the target coordinates.

  2. Knowl 2 — UI-I2E-Bench evaluates grounding across target size, element type, and instruction implicitness

    experimental setup

    UI-I2E-Bench is a semi-automatically constructed GUI instruction-grounding benchmark containing 1,477 screenshot–instruction–element examples from Web, Windows, and Android interfaces. The synthesis pipeline generates candidate examples; sampling is balanced by element type, and human reviewers check the target boxes and instructions, remove serious errors, and correct retained slight errors. Each example is annotated for element type and whether its instruction is explicit or implicit. Compared with ScreenSpot’s 1,272 examples, UI-I2E-Bench has a lower landscape element-to-screen ratio (0.042 versus 0.088), a higher minimum proportion for any element type (12.34% versus 3.21%), and an annotated implicit-instruction proportion of 63.03%; ScreenSpot lacks fine-grained element-type and implicitness annotations. The element-to-screen ratio is defined as the square root of target-box area divided by the square root of screenshot area. These dimensions support evaluation on smaller targets, less frequent element types, and instructions that require inference beyond direct visible-text matching.

  3. Knowl 3 — The grounding training corpus contains nearly 9.9 million instructions across platforms

    data/table

    The final training corpus contains 1,635,594 screenshots and 9,899,581 grounding instructions. Its components are UI-E2I-Synth-Web (1,536,200 screenshots; 9,097,736 instructions), UI-E2I-Synth-Desktop (14,087; 334,397), UI-E2I-Synth-AndroidControl (40,199; 109,126), MOTIF (30,699; 320,219), and WidgetCaption (14,409; 38,103). Web examples were rendered from selected Common Crawl pages; desktop data came from traversing Windows applications; AndroidControl supplied mobile screenshot–metadata pairs, while MOTIF and WidgetCaption added existing mobile grounding data. The authors report that non-text elements make up 23.0% of their synthetic data, compared with 8.7% in a manually labeled random sample of 100 SeeClick web elements. In separate random samples of 1,000 landscape elements, the proportions with element-to-screen ratios 0.00–0.02, 0.02–0.04, and 0.04–1.00 were respectively 36.92%, 40.43%, and 22.65% for the authors’ web data, versus 11.49%, 43.30%, and 45.21% for SeeClick web data. Thus, the curated data includes a larger share of non-text elements and small targets.

  4. Knowl 4 — UI-I2E-VLM uses full-parameter fine-tuning with model-specific high-resolution input handling

    experimental setup

    The authors fine-tuned two vision-language models on the 9,899,581-instruction grounding corpus: InternVL2-4B, which was not pretrained on GUI data, and Qwen2-VL-7B, whose pretraining included GUI screenshots. Both models were fully fine-tuned. For UI-I2E-VLM-7B, bounding boxes were normalized to the range [0, 1000), examples for the same image were packed into one conversation with at most 15 examples, and the maximum image size was set to 1500 × 1500 pixels; training took about 60 hours on 16 A100 GPUs. For UI-I2E-VLM-4B, boxes used original image coordinates, data packing was not used because it reduced performance, and each image was represented using 12 tiles of 448 × 448 pixels; training took about 90 hours on 64 A100 GPUs. These settings were intended to support grounding on high-resolution interfaces.

  5. Knowl 5 — UI-I2E-VLM-7B leads the reported aggregate grounding comparison with fewer training instructions than OS-Atlas-7B

    empirical result

    Grounding accuracy was measured as the percentage of examples for which the predicted point fell inside the target bounding box. Across ScreenSpot, UI-I2E-Bench, and ScreenSpot-Pro, UI-I2E-VLM-7B scored 82.5% on ScreenSpot, 69.5% on UI-I2E-Bench, and 23.6% on ScreenSpot-Pro; the arithmetic mean of these three benchmark averages was 58.5%. OS-Atlas-7B scored 82.5%, 58.6%, and 18.9%, respectively, for a mean of 53.3%. The paper reports this as a 9.7% relative improvement in average performance. UI-I2E-VLM-7B used 9.9 million training instructions, compared with 13.6 million for OS-Atlas-7B. UI-I2E-VLM-4B, trained on the same 9.9 million instructions, scored 70.4% on ScreenSpot, 53.4% on UI-I2E-Bench, 12.2% on ScreenSpot-Pro, and 45.3% on the three-benchmark mean. ScreenSpot’s comparatively strong results did not ensure similar performance on the more challenging UI-I2E-Bench and ScreenSpot-Pro: for example, ShowUI averaged 76.8% on ScreenSpot but 41.5% on UI-I2E-Bench and 7.7% on ScreenSpot-Pro.

  6. Knowl 6 — UI-I2E-VLM-7B gains most clearly on implicit instructions and less common element types

    empirical result

    On UI-I2E-Bench, UI-I2E-VLM-7B achieved 72.0% accuracy on explicit instructions and 67.9% on implicit instructions, for an overall benchmark accuracy of 69.5%. OS-Atlas-7B scored 63.2%, 55.8%, and 58.6% on those measures, respectively, so the difference on implicit instructions was 12.1 percentage points. UI-I2E-VLM-7B also exceeded OS-Atlas-7B across all five reported element types: Button 77.0% versus 69.1%, Icon 68.2% versus 58.7%, Dropdown 84.8% versus 80.3%, Input 86.2% versus 70.1%, and Toggle 44.4% versus 32.3%. Its platform accuracies were 62.1% for Web, 64.0% for desktop, and 76.2% for mobile, compared with OS-Atlas-7B’s 52.2%, 48.9%, and 68.1%. The particularly large differences for icons, input elements, and implicit instructions are consistent with the benchmark’s emphasis on annotation dimensions that earlier evaluations did not expose.

  7. Knowl 7 — Ablations support both user-action instruction synthesis and GPT-4o-generated expressions

    empirical result

    To test data quality and synthesis components, the authors sampled 500,000 web instructions for each setting, fine-tuned the same InternVL2-4B model, and evaluated it on ScreenSpot. The full UI-E2I-Synth data achieved an average accuracy of 65.8%, compared with 44.9% for OS-Atlas-Web. Removing the final instruction-synthesis step and using referring expressions as instructions reduced accuracy to 54.3%, an 11.5-percentage-point drop from the full pipeline. Using raw element attributes as instructions, thereby also removing GPT-4o-generated referring expressions, reduced accuracy to 46.9%, an 18.9-point drop. The full pipeline’s mobile, desktop, and web Text/Icon accuracies were 89.0/34.1%, 93.3/42.9%, and 79.6/44.7%. The corresponding averages for the no-instruction-synthesis setting were 85.4/20.5%, 84.0/18.6%, and 69.6/30.1%; for the raw-attributes setting they were 58.6/37.1%, 52.1/37.9%, and 46.1/44.2%. These controlled comparisons indicate that both natural-language referring expressions and the later conversion into user-oriented instructions contribute to the tested model’s grounding performance.

  8. Knowl 8 — Grounding accuracy decreases for smaller targets, while UI-I2E-VLM performs better on small targets than UGround

    empirical result

    In an analysis on UI-I2E-Bench that grouped examples by target-element size relative to screenshot size, accuracy declined as the element-to-screen ratio became smaller. UI-I2E-VLM-4B performed better than UGround in the smaller-target ranges. The authors attribute this advantage to their training data’s greater coverage of small elements and to the larger number of input image tokens used by their model. This result supports evaluating GUI grounding with small targets and high-resolution interfaces rather than relying only on examples with comparatively large target regions.

  9. Knowl 9 — UI-I2E-VLM improves GPT-4o-planned task success on OSWorld

    empirical result

    The authors evaluated grounding models in OSWorld, a live computer-use benchmark with 369 open-ended tasks across Ubuntu, Windows, and macOS. GPT-4o served as the planner and received the user task and current screenshot at each step; a grounding model supplied coordinates for the planner’s predicted action. The screenshot-only baseline used GPT-4o to produce coordinates directly. Overall task success was 3.6% for the screenshot-only baseline, 8.5% with OS-Atlas, and 12.0% with UI-I2E-VLM. Across the reported OS, Office, Daily, Professional, and Workflow categories, the respective success rates were 8.3%, 2.6%, 6.5%, 0%, and 2.1% for the screenshot-only baseline; 30.4%, 2.7%, 11.7%, 16.3%, and 2.2% with OS-Atlas; and 26.1%, 5.2%, 23.4%, 18.8%, and 2.8% with UI-I2E-VLM. UI-I2E-VLM therefore improved overall success in this setup, although OS-Atlas had higher success in the OS category.

  10. Knowl 10 — Remaining errors include spatial, hierarchical, icon-knowledge, and element-type failures; the study is limited to English

    limitation

    The authors identify recurring UI-I2E-VLM errors: failure to recognize icons without text, incorrect counting or positioning among elements in rows and columns, misunderstanding spatial relations, failure to reason through interface hierarchy (for example, locating a login module before its email field), and confusing element types such as a checkbox and adjacent text. They suggest, but do not establish, that quantity- and spatial-relation training examples, explanatory documentation for specialized icons, and chain-of-thought reasoning for hierarchical tasks could help. The study also covers English instructions only. The authors identify scaling the dataset and model and extending synthesis to other languages as future work.

Coverage note — The full per-baseline benchmark tables, individual benchmark examples, and prompt templates are omitted because they provide supporting detail rather than distinct contributions beyond the summarized comparisons, pipeline, and analyses.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Anthropic. 2024. Claude 3 model card.
  3. 3.Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer. 2022. A dataset for interactive vision language navigation with unknown command feasibility. In European Conference on Computer Vision (ECCV).
  4. 4.Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv preprint arXiv:2312.14238.
  5. 5.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024a. Seeclick: Harnessing gui grounding for advanced visual gui agents. arXiv preprint arXiv:2401.10935.
  6. 6.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024b. Seeclick: Harnessing gui grounding for advanced visual gui agents. arXiv preprint arXiv:2401.10935.
  7. 7.Common Crawl. 2024. Common crawl - open repository of web crawl data.
  8. 8.Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuolin Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. Nvlm: Open frontier-class multimodal llms. arXiv preprint arXiv:2409.11402.
  9. 9.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2024. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36.
  10. 10.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233.
  11. 11.Jingwen Fu, Xiaoyi Zhang, Yuwang Wang, Wenjun Zeng, and Nanning Zheng. 2024. Understanding mobile gui: From pixel-words to screen-sentences. Neurocomputing, 601:128200.
  12. 12.Google. 2024. Profile your layout with hierarchy viewer.
  13. 13.Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2024. Navigating the digital world as humans do: Universal visual grounding for gui agents. arXiv preprint arXiv:2410.05243.
  14. 14.Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2023. A real-world webagent with planning, long context understanding, and program synthesis. arXiv preprint arXiv:2307.12856.
  15. 15.Wei Li, William Bishop, Alice Li, Chris Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, and Oriana Riva. 2024. On the effects of data scale on computer control agents. arXiv preprint arXiv:2406.03679.
  16. 16.Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan. 2020. Widget captioning: Generating natural language description for mobile user interface elements. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5495–5510.
  17. 17.Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Zechen Bai, Weixian Lei, Lijuan Wang, and Mike Zheng Shou. Showui: One vision-language-action model for generalist gui agent. In NeurIPS 2024 Workshop on Open-World Agents.
  18. 18.Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Awadallah. 2024. Omniparser for pure vision based gui agent. arXiv preprint arXiv:2408.00203.
  19. 19.Microsoft. 2021. Ui automation overview - .net framework.
  20. 20.Mozilla. 2024. Document object model (dom) - web apis - mdn web docs.
  21. 21.OpenAI. 2025. Operator system card.
  22. 22.Guilherme Penedo, Hynek Kydlícek, Loubna Ben al-ˇlal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. The fineweb datasets: Decanting the web for the finest text data at scale. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  23. 23.Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces. Advances in Neural Information Processing Systems, 36:34354–34370.
  24. 24.Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths. Cognitive architectures for language agents. Transactions on Machine Learning Research.
  25. 25.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. Preprint, arXiv:2409.12191.
  26. 26.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560.
  27. 27.WebAIM. 2024. Webaim: The webaim million - the 2024 report on the accessibility of the top 1,000,000 home pages. Accessed on September 30, 2024.
  28. 28.Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, et al. 2024. Os-atlas: A foundation action model for generalist gui agents. arXiv preprint arXiv:2410.23218.
  29. 29.Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2025. The rise and potential of large language model based agents: A survey. Science China Information Sciences, 68(2):121101.
  30. 30.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv preprint arXiv:2404.07972.
  31. 31.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  32. 32.Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. 2023. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. Preprint, arXiv:2310.11441.
  33. 33.John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik R Narasimhan, and Ofir Press. 2024. Swe-agent: Agent-computer interfaces enable automated software engineering. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  34. 34.Zhizheng Zhang, Wenxuan Xie, Xiaoyi Zhang, and Yan Lu. 2023a. Reinforced ui instruction grounding: Towards a generic ui task automation api. arXiv preprint arXiv:2310.04716.
  35. 35.Zhizheng Zhang, Xiaoyi Zhang, Wenxuan Xie, and Yan Lu. 2023b. Responsible task automation: Empowering large language models as responsible task automators. arXiv preprint arXiv:2306.01242.
  36. 36.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. Gpt-4v(ision) is a generalist web agent, if grounded. In Forty-first International Conference on Machine Learning.
  37. 37.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854.

Citation

MLA
Liu, X., et al. “UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis”. Findings of the Association for Computational Linguistics: ACL 2025, 2025, pp. 15668–84, https://doi.org/10.18653/v1/2025.findings-acl.809.
APA
Liu, X., Zhang, X., Zhang, Z., & Lu, Y. (2025). UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis. Findings of the Association for Computational Linguistics: ACL 2025, 15668–15684. https://doi.org/10.18653/v1/2025.findings-acl.809
Chicago
Liu, X., X. Zhang, Z. Zhang, and Y. Lu. 2025. “UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis”. Findings of the Association for Computational Linguistics: ACL 2025, 15668–84. https://doi.org/10.18653/v1/2025.findings-acl.809.
Harvard
Liu, X. et al. (2025) “UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis”, Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, pp. 15668–15684. Available at: https://doi.org/10.18653/v1/2025.findings-acl.809.
Vancouver
1. Liu X, Zhang X, Zhang Z, Lu Y (2025) UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis. In: Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, pp 15668–15684

BibTeX

@inproceedings{liu-etal-2025-ui,
    title = "{UI}-{E}2{I}-Synth: Advancing {GUI} Grounding with Large-Scale Instruction Synthesis",
    author = "Liu, Xinyi  and
      Zhang, Xiaoyi  and
      Zhang, Ziyun  and
      Lu, Yan",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.809/",
    doi = "10.18653/v1/2025.findings-acl.809",
    pages = "15668--15684",
    ISBN = "979-8-89176-256-5"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/