META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Liangtai SunXingyu ChenLu ChenTianle DaiZichen ZhuKai Yu
Proposes a GUI-based task-oriented dialogue framework and benchmark dataset, META-GUI, enabling conversational assistants to complete multi-turn tasks by interacting directly with mobile app interfaces rather than relying on restrictive backend APIs.
Modern mobile intelligent assistants rely on traditional task-oriented dialogue systems that execute user requests by calling specialized back-end Application Programming Interfaces (APIs). However, many real-world smartphone applications lack dedicated APIs for these systems or have operational workflows too complex for rigid, pre-defined programming interfaces. This dependency severely constrains the adaptability, search range, and overall utility of digital assistants across diverse mobile applications.
The article demonstrates and evaluates a graphical user interface-based task-oriented dialogue system (GUI-TOD). Rather than relying on specialized back-end APIs, this framework directly operates real mobile applications through visual interfaces, using a combination of automated screen actions and natural language generation to execute user goals.
To evaluate this framework, the authors created META-GUI, a benchmark dataset consisting of 1,125 multi-turn dialogues, 4,684 turns, and 18,337 action-level data points across six functional domains (weather, calendar, search, taxi, restaurant, and hotel) on Android devices. The dataset captures full interaction traces, including dialogue histories, step-by-step actions (such as clicking, swiping, and text entry), screen images, and underlying layout structures. The authors then engineered a multi-modal model that fuses textual dialogue, screen text, and visual features to predict upcoming interface actions and synthesize natural language responses.
The investigation produced several key findings. First, the proposed multi-modal model, termed m-BASH (incorporating visual features, action history, and screenshot history), achieved the highest performance, reaching an action completion rate of 82.74% and a turn completion rate of 56.88%, vastly outperforming heuristic baselines. Second, incorporating both visual data and operational history proved critical: action history alone improved the baseline turn completion rate from 52.08% to 55.42%, while adding screenshot history further increased it to 55.62%, primarily because visual cues distinguish interface states (such as active checkboxes) that text alone cannot capture. Third, language models pre-trained primarily on text (BERT) adapted significantly better to mobile interfaces than models pre-trained on scanned documents (LayoutLM variants), which suffered due to structural differences in layout formats. Finally, cross-app and cross-domain testing demonstrated strong generalizability; the system achieved up to a 69.84% action completion rate when transferred to entirely unseen applications within the same domain.
These findings indicate that mobile assistants can successfully complete complex user tasks without specialized developer APIs, offering a viable path toward universally adaptable digital agents. This capability reduces the development cost and integration barriers for third-party application developers while expanding the operational scope of assistants. However, leaders should note that current models remain too computationally heavy to deploy directly on standard mobile hardware, and the current benchmark does not yet account for unexecutable tasks or explicit failure-handling dialogues.
Organizations developing or deploying intelligent assistants should consider exploring visual, interface-driven interaction pipelines rather than investing solely in brittle, app-specific API integrations. Prior to commercial implementation, technical teams must conduct further work to compress these multi-modal models for on-device efficiency and expand datasets to include edge cases, error recoveries, and unachievable user requests.
- Paper: MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling, Paweł Budzianowski et al. (2018). MultiWOZ establishes the foundational multi-domain task-oriented dialogue benchmark that META-GUI adapts and reimagines for direct graphical user interface interactions.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). CoQA lays the groundwork for evaluating conversational multi-turn context tracking and question answering that underpins conversational dialogue modeling in mobile assistants.
- Paper: A Neural Conversational Model, Oriol Vinyals et al. (2015). This paper introduces end-to-end neural sequence modeling for conversational agents, serving as a core paradigm upon which multi-turn dialogue architectures are built.
- Paper: Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration, Junyang Wang et al. (2024). Mobile-Agent-v2 extends mobile GUI automation beyond single-model pipelines into multi-agent collaborative frameworks with dedicated planning, memory, and reflection mechanisms.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent scales multimodal vision-language architectures to high-resolution GUI grounding and decision-making across smartphone and computer interfaces.
- Paper: ShowUI: One Vision-Language-Action Model for GUI Visual Agent, Kevin Qinghong Lin et al. (2025). ShowUI advances vision-based GUI interaction by introducing an efficient token-pruning vision-language-action model tailored for high-resolution visual agent navigation.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). Spotlight advances mobile UI understanding by modeling interface screens exclusively from pixels and region bounding boxes without relying on view hierarchy metadata.
- Paper: ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots, Yu-Chung Hsiao et al. (2025). ScreenQA builds upon mobile visual screen interpretation by providing a standardized benchmark for visual question answering directly from raw mobile application screenshots.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). UI-R1 builds on GUI action prediction by applying rule-based reinforcement learning to improve element localization and decision accuracy without massive supervised fine-tuning.
- Paper: Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction, Yiheng Xu et al. (2025). AGUVIS generalizes pure-vision GUI automation across multiple operating platforms using standardized action spaces and structured inner monologue reasoning.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). PIX2ACT expands pixel-based GUI instruction following by training vision transformers with behavioral cloning and search over generic interface actions.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web broadens interface-driven agent evaluation from mobile environments to complex, multi-domain web navigation using real-world interaction traces.
