Large Language Model-Brained GUI Agents: A Survey

Chaoyun ZhangShilin HeJiaxu QianBowen LiLiqun LiSi QinYu KangMing-Jie MaQingwei Lin 林庆维Saravan Rajmohan

article2025arXiv212 citations

Systematizes the development of multimodal language model agents for graphical user interfaces by detailing agent architectures, training datasets, action modeling techniques, and evaluation benchmarks across web, mobile, and desktop environments.

Listen

Modern digital workflows increasingly rely on graphical user interfaces (GUIs), which are inherently designed for human visual interaction rather than machine automation. Historically, automating tasks across software required rigid, script-based or rule-based methods that easily broke when layouts shifted or workflows changed. Meanwhile, application programming interface (API) automation remains constrained because many applications lack exposed, unified, or accessible APIs. The rapid rise of multimodal Large Language Models (LLMs) and Visual Language Models (VLMs) offers a solution by enabling "LLM-brained" GUI agents that interpret visual interfaces, follow natural language instructions, and interact with software just like human users.

This article provides a comprehensive survey that traces the evolution, architectures, training pipelines, benchmarks, and real-world applications of LLM-powered GUI agents. It systematically evaluates how these agents function across web, mobile, and desktop environments, and outlines the major technical challenges and strategic steps required to achieve robust, autonomous computer interaction.

To establish a holistic overview, the article synthesizes findings from over 500 research studies and industry milestones. It analyzes the end-to-end agent architecture, which comprises environmental perception (via screenshots, accessibility trees, and computer vision), prompt engineering, model inference, action execution (mouse clicks, keystrokes, gestures, and API calls), and short- and long-term memory systems. The analysis highlights key technical enhancements, such as multi-agent collaboration, vision-based UI grounding, self-reflection frameworks, self-evolution mechanisms, and reinforcement learning.

Several key findings emerge from the synthesis. First, multimodal visual perception significantly outperforms traditional text-only and accessibility-tree methods, especially in dynamic environments where underlying code hierarchies are missing or non-standard. Second, incorporating search and simulation techniques—such as Monte Carlo Tree Search, world models, and best-first search—substantially improves multi-step decision-making, boosting task success rates by up to 39% to 50% compared to standard reactive models. Third, hybrid interaction frameworks that dynamically combine UI operations with native system APIs maximize versatility while drastically reducing execution latency. Finally, advanced capabilities like self-reflection and experience-based memory retrieval allow agents to autonomously recover from navigation errors, turning single-trial failures into successful multi-step completions.

These findings indicate a major paradigm shift toward intuitive, conversational computing where non-technical users can automate complex cross-application workflows without specialized code. For organizations, this significantly lowers the barrier to workflow automation and robotic process automation (RPA), but it also introduces latency costs and reliability risks inherent in sequential, multi-step LLM reasoning.

Leaders and developers should pursue a hybrid architecture that prioritizes fast API execution where available while retaining visual GUI automation as a universal fallback. Organizations looking to deploy GUI agents should implement iterative reasoning frameworks, maintain structured memory repositories for continuous learning, and conduct controlled pilot programs before deployment. Further research is recommended to optimize real-time inference latency and develop standardized, high-fidelity safety and verification protocols.

While the article provides high confidence in the architectural foundations and rapid performance trajectory of GUI agents, leaders should exercise caution regarding real-world edge cases. The underlying models still face challenges with latency, execution errors on subtle visual shifts, and potential security risks when interacting with sensitive software environments without human oversight.

Cover for Large Language Model-Brained GUI Agents: A Survey

Abstract

GUIs have long been central to human-computer interaction, providing an intuitive and visually-driven way to access and interact with digital systems. The advent of LLMs, particularly multimodal models, has ushered in a new era of GUI automation. They have demonstrated exceptional capabilities in natural language understanding, code generation, and visual processing. This has paved the way for a new generation of LLM-brained GUI agents capable of interpreting complex GUI elements and autonomously executing actions based on natural language instructions. These agents represent a paradigm shift, enabling users to perform intricate, multi-step tasks through simple conversational commands. Their applications span across web navigation, mobile app interactions, and desktop automation, offering a transformative user experience that revolutionizes how individuals interact with software. This emerging field is rapidly advancing, with significant progress in both research and industry.

To provide a structured understanding of this trend, this paper presents a comprehensive survey of LLM-brained GUI agents, exploring their historical evolution, core components, and advanced techniques. We address research questions such as existing GUI agent frameworks, the collection and utilization of data for training specialized GUI agents, the development of large action models tailored for GUI tasks, and the evaluation metrics and benchmarks necessary to assess their effectiveness. Additionally, we examine emerging applications powered by these agents. Through a detailed analysis, this survey identifies key research gaps and outlines a roadmap for future advancements in the field. By consolidating foundational knowledge and state-of-the-art developments, this work aims to guide both researchers and practitioners in overcoming challenges and unlocking the full potential of LLM-brained GUI agents.

Citation

MLA
Zhang, C., et al. “Large Language Model-Brained GUI Agents: A Survey”. arXiv, 2024, http://arxiv.org/abs/2411.18279v12.
APA
Zhang, C., He, S., Qian, J., Li, B., Li, L., Qin, S., Kang, Y., Ma, M., Liu, G., Lin, Q., Rajmohan, S., Zhang, D., & Zhang, Q. (2024). Large Language Model-Brained GUI Agents: A Survey. arXiv. http://arxiv.org/abs/2411.18279v12
Chicago
Zhang, C., S. He, J. Qian, et al. 2024. “Large Language Model-Brained GUI Agents: A Survey”. arXiv. http://arxiv.org/abs/2411.18279v12.
Harvard
Zhang, C. et al. (2024) “Large Language Model-Brained GUI Agents: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2411.18279v12.
Vancouver
1. Zhang C, He S, Qian J, et al (2024) Large Language Model-Brained GUI Agents: A Survey. arXiv

BibTeX

@article{zhang2024large,
  title = {Large Language Model-Brained GUI Agents: A Survey},
  author = {Zhang, Chaoyun and He, Shilin and Qian, Jiaxu and Li, Bowen and Li, Liqun and Qin, Si and Kang, Yu and Ma, Minghua and Liu, Guyue and Lin, Qingwei and Rajmohan, Saravan and Zhang, Dongmei and Zhang, Qi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2411.18279v12},
  eprint = {2411.18279}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission