Large Language Model-Brained GUI Agents: A Survey
Chaoyun ZhangShilin HeJiaxu QianBowen LiLiqun LiSi QinYu KangMing-Jie MaQingwei Lin 林庆维Saravan Rajmohan
Systematizes the development of multimodal language model agents for graphical user interfaces by detailing agent architectures, training datasets, action modeling techniques, and evaluation benchmarks across web, mobile, and desktop environments.
Modern digital workflows increasingly rely on graphical user interfaces (GUIs), which are inherently designed for human visual interaction rather than machine automation. Historically, automating tasks across software required rigid, script-based or rule-based methods that easily broke when layouts shifted or workflows changed. Meanwhile, application programming interface (API) automation remains constrained because many applications lack exposed, unified, or accessible APIs. The rapid rise of multimodal Large Language Models (LLMs) and Visual Language Models (VLMs) offers a solution by enabling "LLM-brained" GUI agents that interpret visual interfaces, follow natural language instructions, and interact with software just like human users.
This article provides a comprehensive survey that traces the evolution, architectures, training pipelines, benchmarks, and real-world applications of LLM-powered GUI agents. It systematically evaluates how these agents function across web, mobile, and desktop environments, and outlines the major technical challenges and strategic steps required to achieve robust, autonomous computer interaction.
To establish a holistic overview, the article synthesizes findings from over 500 research studies and industry milestones. It analyzes the end-to-end agent architecture, which comprises environmental perception (via screenshots, accessibility trees, and computer vision), prompt engineering, model inference, action execution (mouse clicks, keystrokes, gestures, and API calls), and short- and long-term memory systems. The analysis highlights key technical enhancements, such as multi-agent collaboration, vision-based UI grounding, self-reflection frameworks, self-evolution mechanisms, and reinforcement learning.
Several key findings emerge from the synthesis. First, multimodal visual perception significantly outperforms traditional text-only and accessibility-tree methods, especially in dynamic environments where underlying code hierarchies are missing or non-standard. Second, incorporating search and simulation techniques—such as Monte Carlo Tree Search, world models, and best-first search—substantially improves multi-step decision-making, boosting task success rates by up to 39% to 50% compared to standard reactive models. Third, hybrid interaction frameworks that dynamically combine UI operations with native system APIs maximize versatility while drastically reducing execution latency. Finally, advanced capabilities like self-reflection and experience-based memory retrieval allow agents to autonomously recover from navigation errors, turning single-trial failures into successful multi-step completions.
These findings indicate a major paradigm shift toward intuitive, conversational computing where non-technical users can automate complex cross-application workflows without specialized code. For organizations, this significantly lowers the barrier to workflow automation and robotic process automation (RPA), but it also introduces latency costs and reliability risks inherent in sequential, multi-step LLM reasoning.
Leaders and developers should pursue a hybrid architecture that prioritizes fast API execution where available while retaining visual GUI automation as a universal fallback. Organizations looking to deploy GUI agents should implement iterative reasoning frameworks, maintain structured memory repositories for continuous learning, and conduct controlled pilot programs before deployment. Further research is recommended to optimize real-time inference latency and develop standardized, high-fidelity safety and verification protocols.
While the article provides high confidence in the architectural foundations and rapid performance trajectory of GUI agents, leaders should exercise caution regarding real-world edge cases. The underlying models still face challenges with latency, execution errors on subtle visual shifts, and potential security risks when interacting with sensitive software environments without human oversight.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This survey establishes the fundamental brain-perception-action architectural paradigm of LLM-based autonomous agents that the source specializes for GUI environments.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). It provides a foundational taxonomy of LLM agent architectures, planning mechanisms, and evaluation strategies that underlie the specific design of GUI-oriented agents.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). This paper introduces a seminal visual language model designed specifically for high-resolution visual grounding and multi-platform GUI navigation, serving as a direct architectural precursor.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). It pioneers the paradigm of training vision-based agents to map raw screenshot pixels directly to low-level UI mouse and keyboard actions without relying on DOM trees.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). This work establishes the standard benchmark and simulation environment for assessing the functional correctness of autonomous agents executing long-horizon web tasks.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). It introduces a benchmark dataset and multi-choice action prediction framework for generalist web agents that the survey reviews extensively.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). It demonstrates how end-to-end multimodal large models can navigate live websites via visual tagging, representing a key milestone surveyed in LLM-brained GUI interaction.
- Paper: Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration, Junyang Wang et al. (2024). It develops a multi-agent visual framework for navigating mobile applications, providing a core multi-agent paradigm for mobile GUI automation.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). It provides the foundational vision-language approach for mobile UI comprehension based solely on screen pixels and focus regions.
- Paper: META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI, Liangtai Sun et al. (2022). It introduces early benchmark datasets and multimodal conversational agents that execute tasks directly on mobile interfaces rather than relying on back-end APIs.
- Paper: Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction, Yiheng Xu et al. (2025). It implements a unified pure-vision GUI agent that integrates visual grounding with structured inner-monologue reasoning across platforms, advancing the survey's discussion on action models.
- Paper: ShowUI: One Vision-Language-Action Model for GUI Visual Agent, Kevin Qinghong Lin et al. (2025). It presents a lightweight vision-language-action model and efficient visual token pruning recipe for GUI agents operating directly on screenshots across web, desktop, and mobile environments.
- Paper: Covering Human Action Space for Computer Use: Data Synthesis and Benchmark, Miaosen Zhang et al. (2026). It expands the computer-use action space beyond single clicks to continuous and multi-point interactions using large-scale synthetic data and a tailored grounding benchmark.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). It applies rule-based reinforcement learning to train lightweight multimodal GUI agents for efficient action prediction and generalization from minimal data.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). It introduces a workflow memory mechanism enabling web and GUI agents to autonomously induce, store, and reuse procedural sub-routines from experience.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). It develops a self-adaptive data-centric framework that constructs training trajectories directly from interaction in realistic GUI and coding environments without human labeling.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). It builds a versatile multi-agent web navigation system combining strategic high-level planning with reflection-guided tree search to overcome unexpected layouts.
- Paper: ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots, Yu-Chung Hsiao et al. (2025). It provides a large-scale visual question-answering benchmark to systematically test the screen reading and visual grounding capabilities of multimodal GUI models.
- Paper: Neural Computers, Mingchen Zhuge et al. (2026). It explores a radical paradigm shift from external GUI agents to neural computers that simulate runtime interfaces and control dynamics end-to-end within video model weights.
- Paper: Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation, Abubakarr Jaye et al. (2025). It proposes an agentic judge framework that inspects intermediate actions and execution traces to rigorously evaluate task completion in complex agent workflows.
