META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Liangtai SunXingyu ChenLu ChenTianle DaiZichen ZhuKai Yu

article2022EMNLP105 citations

Proposes a GUI-based task-oriented dialogue framework and benchmark dataset, META-GUI, enabling conversational assistants to complete multi-turn tasks by interacting directly with mobile app interfaces rather than relying on restrictive backend APIs.

Listen

Modern mobile intelligent assistants rely on traditional task-oriented dialogue systems that execute user requests by calling specialized back-end Application Programming Interfaces (APIs). However, many real-world smartphone applications lack dedicated APIs for these systems or have operational workflows too complex for rigid, pre-defined programming interfaces. This dependency severely constrains the adaptability, search range, and overall utility of digital assistants across diverse mobile applications.

The article demonstrates and evaluates a graphical user interface-based task-oriented dialogue system (GUI-TOD). Rather than relying on specialized back-end APIs, this framework directly operates real mobile applications through visual interfaces, using a combination of automated screen actions and natural language generation to execute user goals.

To evaluate this framework, the authors created META-GUI, a benchmark dataset consisting of 1,125 multi-turn dialogues, 4,684 turns, and 18,337 action-level data points across six functional domains (weather, calendar, search, taxi, restaurant, and hotel) on Android devices. The dataset captures full interaction traces, including dialogue histories, step-by-step actions (such as clicking, swiping, and text entry), screen images, and underlying layout structures. The authors then engineered a multi-modal model that fuses textual dialogue, screen text, and visual features to predict upcoming interface actions and synthesize natural language responses.

The investigation produced several key findings. First, the proposed multi-modal model, termed m-BASH (incorporating visual features, action history, and screenshot history), achieved the highest performance, reaching an action completion rate of 82.74% and a turn completion rate of 56.88%, vastly outperforming heuristic baselines. Second, incorporating both visual data and operational history proved critical: action history alone improved the baseline turn completion rate from 52.08% to 55.42%, while adding screenshot history further increased it to 55.62%, primarily because visual cues distinguish interface states (such as active checkboxes) that text alone cannot capture. Third, language models pre-trained primarily on text (BERT) adapted significantly better to mobile interfaces than models pre-trained on scanned documents (LayoutLM variants), which suffered due to structural differences in layout formats. Finally, cross-app and cross-domain testing demonstrated strong generalizability; the system achieved up to a 69.84% action completion rate when transferred to entirely unseen applications within the same domain.

These findings indicate that mobile assistants can successfully complete complex user tasks without specialized developer APIs, offering a viable path toward universally adaptable digital agents. This capability reduces the development cost and integration barriers for third-party application developers while expanding the operational scope of assistants. However, leaders should note that current models remain too computationally heavy to deploy directly on standard mobile hardware, and the current benchmark does not yet account for unexecutable tasks or explicit failure-handling dialogues.

Organizations developing or deploying intelligent assistants should consider exploring visual, interface-driven interaction pipelines rather than investing solely in brittle, app-specific API integrations. Prior to commercial implementation, technical teams must conduct further work to compress these multi-modal models for on-device efficiency and expand datasets to include edge cases, error recoveries, and unachievable user requests.

arXiv: 2205.11029
Cover for META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Abstract

Task-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation. Current TOD systems usually focus on multi-turn text/speech interaction, then they would call back-end APIs designed for TODs to perform the task. However, this API-based architecture greatly limits the information-searching capability of intelligent assistants and may even lead to task failure if TOD-specific APIs are not available or the task is too complicated to be executed by the provided APIs. In this paper, we propose a new TOD architecture: GUI-based task-oriented dialogue system (GUI-TOD). A GUI-TOD system can directly perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs. Furthermore, we release META-GUI, a dataset for training a Multi-modal conversational Agent on mobile GUI. We also propose a multi-model action prediction and response model, which show promising results on META-GUI. The dataset, codes and leaderboard are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Task Definition
  • 3 Meta-GUI Creation
  • 3.1 Collecting GUI traces
  • 3.2 Data Review
  • 3.3 Post-processing
  • 3.4 Data Analysis
  • 4 Model Design
  • 4.1 Action Model
  • 4.2 Response Model
  • 5 Experiment
  • 5.1 Data Preprocess
  • 5.2 Experiment Setup
  • 5.3 Experiment Result
  • 5.4 Generality
  • 6 Related Work
  • 6.1 Natural Language Commands on GUI
  • 6.2 Programming by Demonstration on GUI
  • 6.3 Visual Dialogue
  • 7 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Details of Apps
  • B Annotation System
  • C Example of View Hierarchy
  • D Data Review
  • E Example of the generated target during collecting
  • F Examples of Item types
  • G Examples of parameter predictions
  • H Case study

Knowls

  1. Knowl 1 — GUI-Based Task-Oriented Dialogue System Formulation and Action Space

    definition

    A Graphical User Interface Task-Oriented Dialogue system (GUI-TOD) executes user-requested tasks by directly performing operations on mobile application interfaces rather than invoking backend application programming interfaces (APIs). The framework operates without predefined domain ontologies, extracting parameter values directly during interface interaction.

    The system comprises two submodules:

    1. Action Executor (AE): Operates the GUI across multiple steps by predicting actions.
    2. Response Generator (RG): Generates a natural language response after execution finishes.

    Let dialogue turn ii be represented as Di=(Ui,Ri)D_i = (U_i, R_i), where UiU_i is the ii-th user utterance and RiR_i is the ii-th system response. Let Si,j=(s,v)S_{i,j} = (s, v) denote the jj-th screen state in turn ii, containing screenshot image ss and view hierarchy tree vv. Let Ai,j=(t,p)A_{i,j} = (t, p) denote the jj-th action in turn ii, where tt is the action type and pp is the parameter.

    The action prediction task is formulated as: Ai,j=F(S1:i,1:j,A1:i,1:j−1,D1:i−1,Ui)A_{i,j} = \mathcal{F}(S_{1:i,1:j}, A_{1:i,1:j-1}, D_{1:i-1}, U_i)

    The response generation task is formulated as: Ri=G(S1:i,A1:i,D1:i−1,Ui)R_i = \mathcal{G}(S_{1:i}, A_{1:i}, D_{1:i-1}, U_i) where A1:iA_{1:i} and S1:iS_{1:i} denote the sequences of actions and screen states up to turn ii.

    The action space consists of seven discrete actions with three parameter types:

    • Click(item = x): Clicks the UI item indexed by xx.
    • Swipe(direction = x): Swipes the display in direction x∈{up,down}x \in \{\text{up}, \text{down}\}.
    • Input(text = x): Types text string xx into the active input element.
    • Enter(): Presses the keyboard Enter button.
    • Clear(): Clears the contents of the current input box.
    • Back(): Presses the system back button.
    • End(): Virtual terminal action denoting the end of GUI operations for the turn, triggering the Response Generator.
  2. Knowl 2 — META-GUI Benchmark Dataset

    definition

    META-GUI is a multimodal dataset designed for training conversational agents to complete multi-turn tasks via direct Android GUI operations. It encompasses 1,125 dialogues, 4,684 dialogue turns, and 18,337 action-level data points across six domains: weather, calendar, search, taxi, hotel, and restaurant.

    The dataset is partitioned into training, development, and test sets using an 8:1:1 split:

    • Train: 897 dialogues, 3,692 turns, 14,539 actions
    • Development: 112 dialogues, 509 turns, 1,875 actions
    • Test: 116 dialogues, 483 turns, 1,923 actions

    Key structural statistics of META-GUI include:

    • Average images per turn: 5.305.30
    • Average words per user utterance: 88
    • Average clickable items per image: 23.8023.80
    • Average word length of item text: 2.482.48
    • Average words per system response: 99
    • Average words per input action: 33

    UI elements are parsed from Android view hierarchy XML files as clickable leaf nodes (where the node or its parent has clickable = true). Each item contains its text (resolved hierarchically from text, content-desc, or resource-id properties), item type, and bounding box coordinates. For applications that do not yield a valid view hierarchy, Optical Character Recognition (OCR) is used to detect text and construct a pseudo layout file.

  3. Knowl 3 — m-BASH Multi-Modal Action Prediction Architecture

    model/method

    The Multi-modal BERT with Action histories and Screenshot Histories (m-BASH) model predicts the next GUI operation Ai,j=(t,p)A_{i,j} = (t, p) given the user request, dialogue context, current screen, and interaction history.

    1. Text Encoding: Given dialogue history and current utterance {D1:i−1,Ui}={w1,…,wn}\{D_{1:i-1}, U_i\} = \{w_1, \ldots, w_n\} and the text tokens of kk clickable UI items {m1,1:l1,…,mk,1:lk}\{m_{1,1:l_1}, \ldots, m_{k,1:l_k}\} extracted from the screen, the input sequence is formed as: X={w1:n;m1,1:l1,…,mk,1:lk}X = \{w_{1:n}; m_{1,1:l_1}, \ldots, m_{k,1:l_k}\} H=TransformerEncoder(X)=[D;M]H = \text{TransformerEncoder}(X) = [D; M] where D={w1,…,wn}D = \{w_1, \ldots, w_n\} and M={m1,1:l1;…;mk,1:lk}M = \{m_{1,1:l_1}; \ldots; m_{k,1:l_k}\}.

    2. Visual Feature Extraction: A screenshot image is processed by a Faster R-CNN with a ResNet50 + FPN backbone. Region of Interest (RoI) pooling is applied using the bounding box of each clickable UI element to extract item-level regional visual features I={I1,…,Ik}I = \{I_1, \ldots, I_k\}.

    3. Multimodal Information Fusion: Item text token features mu,1:lum_{u,1:l_u} are concatenated with their corresponding visual feature vector IuI_u, while dialogue tokens w1:nw_{1:n} are concatenated with zero vectors. The concatenated sequence passes through an MM-layer Transformer encoder. At each Transformer layer, the regional visual features are re-concatenated with the intermediate output of the preceding layer.

  4. Knowl 4 — Action Parameter Classification Heads and History Fusion Mechanisms

    model/method

    Given the fused multimodal representation EE output by the action model encoder:

    1. Action Type Classifier: Predicts the categorical distribution over the 7 action types from the representation of the [CLS] token: pa=Softmax(FFN1(E[CLS]))p_a = \text{Softmax}(\text{FFN}_1(E_{\text{[CLS]}}))

    2. Input Text Prediction Head: Formulated as extractive span prediction over the dialogue tokens DD, predicting start index pdsp_{ds} and end index pdep_{de}: pds=FFN2(D),pde=FFN3(D)p_{ds} = \text{FFN}_2(D), \quad p_{de} = \text{FFN}_3(D)

    3. Target Item Selection Head: Formulated over the kk screen items. The token representations for each item uu are pooled to compute mˉu=Avgpooling(mu,1:lu)\bar{m}_u = \text{Avgpooling}(m_{u,1:l_u}), with mˉ=[mˉ1,…,mˉk]\bar{m} = [\bar{m}_1, \ldots, \bar{m}_k]: pm=Softmax(FFN4(mˉ))p_m = \text{Softmax}(\text{FFN}_4(\bar{m}))

    4. Direction Classifier: Binary classification head predicting swipe direction {up,down}\{\text{up}, \text{down}\}: pd=Softmax(FFN5(E[CLS]))p_d = \text{Softmax}(\text{FFN}_5(E_{\text{[CLS]}}))

    5. History Incorporation:

    • Action History: The HH most recent action types {t1:H}\{t_{1:H}\} are added as special vocabulary tokens and prepended to the encoder input: X={t1:H;w1:n;m1,1:l1,…,mk,1:lk}X = \{t_{1:H}; w_{1:n}; m_{1,1:l_1}, \ldots, m_{k,1:l_k}\}.
    • Screenshot History: A sequence of HH screenshot visual feature sets I^1,…,I^H\hat{I}_1, \ldots, \hat{I}_H is aggregated recurrently via multi-head attention: Iˉ1=I^1,Iˉi+1=Attn(W1I^i+1,W2Iˉi,W3Iˉi)for 1≤i≤H−1\bar{I}_1 = \hat{I}_1, \quad \bar{I}_{i+1} = \text{Attn}(W_1 \hat{I}_{i+1}, W_2 \bar{I}_i, W_3 \bar{I}_i) \quad \text{for } 1 \le i \le H - 1 where W1,W2,W3W_1, W_2, W_3 are learnable projection matrices, and IˉH\bar{I}_H provides the visual features for multimodal fusion.
  5. Knowl 5 — Multi-Modal Response Generation Architecture

    model/method

    The Response Generator (RG) model generates natural language system responses RiR_i following the completion of GUI operations (triggered by the End() action).

    The RG model shares the multimodal encoder and image feature extraction pipeline with the action model, encoding dialogue tokens DD and screen item texts MM from the final GUI state into fused representation [D;M][D; M]. Action history tokens are omitted from response generation, as the response depends on the execution outcome shown on the screen and dialogue context.

    The response text RiR_i is generated by a Transformer decoder consisting of N=4N = 4 decoder blocks: Ri=TransformerDecoder([D;M])R_i = \text{TransformerDecoder}([D; M])

  6. Knowl 6 — Action Prediction Performance on META-GUI

    empirical result

    Action prediction models are evaluated on the META-GUI test set across seven metrics: Action Type Accuracy (%), Input Exact Match (EM, %), Input F1 (%), Target Item Accuracy (%), Direction Accuracy (%), Action Completion Rate (Action CR, %), and Turn Completion Rate (Turn CR, %). Action CR requires correct prediction of both action type and all associated parameters. Turn CR requires all actions within a turn sequence to be correct.

    Method Action Type Acc Input EM Input F1 Item Acc Direction Acc Action CR Turn CR
    Random 14.02 8.72 17.96 9.08 51.26 5.37 3.99
    Most Frequent Method (MFM) 53.71 14.02 37.78 16.58 89.31 8.91 0.00
    Frequency Method (FM) 37.48 6.65 14.02 9.94 81.51 10.00 6.76
    LayoutLMv2 85.60 47.37 70.76 64.38 92.95 64.48 36.88
    LayoutLM 82.22 83.04 90.56 71.98 94.87 67.76 38.12
    BERT 87.52 93.57 97.24 82.84 93.59 78.42 52.08
    BERT + mm 88.35 92.98 96.42 84.51 94.23 80.45 53.96
    BERT + act_h 88.87 91.81 94.86 84.23 95.51 80.97 55.42
    BERT + mm + scr_h 89.86 90.06 95.30 84.32 94.87 81.54 55.62
    m-BASH 90.80 91.23 96.42 85.90 94.23 82.74 56.88

    BERT outperforms document-pretrained backbones (LayoutLM and LayoutLMv2) due to the domain discrepancy between scanned documents and mobile GUIs. Integrating multimodal visual features (+mm) improves target item selection (82.84% to 84.51%) and overall Action CR (78.42% to 80.45%) by providing visual cues such as radio button selection states. Combining multimodal fusion with action history (act_h) and screenshot history (scr_h) in m-BASH yields the highest Turn CR (56.88%).

  7. Knowl 7 — Response Generation BLEU Performance on META-GUI

    empirical result

    Response generation quality on the META-GUI test set is measured using BLEU scores against human reference responses:

    Method Response BLEU Score
    Random 0.0071
    Most Frequent Method (MFM) 0.0929
    Frequency Method (FM) 0.0788
    LayoutLM 0.5043
    LayoutLMv2 0.5820
    BERT 0.6219
    BERT + mm 0.6224
    BERT + mm + scr_h 0.6311

    BERT-based decoders outperform LayoutLM and LayoutLMv2 baselines. Incorporating regional image features (+mm) and screenshot history (+scr_h) yields progressive improvements up to 0.63110.6311 BLEU by providing historical visual context from preceding execution states.

  8. Knowl 8 — Cross-App and Cross-Domain Generality Evaluation

    empirical result

    The generalization ability of m-BASH across unseen mobile applications and unseen task domains is evaluated under leave-one-out data partitions:

    Data Domain of Test Set Action Completion Rate (%) Turn Completion Rate (%)
    App Generality
    An app of weather 56.45 45.71
    An app of calendar 69.84 23.17
    Domain Generality
    Weather 41.96 21.04
    Calendar 62.39 19.20
    Search 59.40 16.24
    Taxi 37.68 21.72
    Restaurant 30.26 15.42
    Hotel 31.24 16.26

    When evaluated on an unseen application in a known domain (App Generality), the system retains substantial task execution capability (up to 69.84% Action CR). In unseen task domains (Domain Generality), performance degrades but maintains an Action CR between 30.26% and 62.39%, demonstrating that GUI interaction patterns transfer across applications without task-specific schema customization.

  9. Knowl 9 — Limitations of GUI-TOD and META-GUI

    limitation

    The GUI-TOD framework and META-GUI dataset have two primary limitations:

    1. Absence of Failure and Unachievable Request Handling: META-GUI consists exclusively of successful task executions. The dataset and models do not model scenarios where a requested task is unachievable, missing conversational handling for agent failure states or clarification prompts (e.g., replying that a task cannot be completed).
    2. Model Footprint and On-Device Deployment: The multimodal architecture combining a text Transformer encoder, Faster R-CNN visual backbone, multimodal fusion layers, and Transformer decoders requires high memory and computational capacity, preventing direct local deployment on mobile devices without model compression.

Coverage note — None was omitted.

References

  1. 1.Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas, and Verena Rieser. 2020. History for visual dialog: Do we really need it? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8182–8197, Online. Association for Computational Linguistics.
  2. 2.Jacob Andreas, John Bufe, David Burkett, Charles Chen, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, et al. 2020. Task-oriented dialogue as dataflow synthesis. Transactions of the Association for Computational Linguistics, 8:556–571.
  3. 3.Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, and Blaise Agüera y Arcas. 2021. Uibert: Learning generic multimodal representations for ui understanding. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 1705–1712. International Joint Conferences on Artificial Intelligence Organization. Main Track.
  4. 4.Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017. A survey on dialogue systems: Recent advances and new frontiers. Acm Sigkdd Explorations Newsletter, 19(2):25–35.
  5. 5.Lu Chen, Cheng Chang, Zhi Chen, Bowen Tan, Milica Gašić, and Kai Yu. 2018. Policy adaptation for deep reinforcement learning-based dialogue management. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 6074–6078. IEEE.
  6. 6.Lu Chen, Zhi Chen, Bowen Tan, Sishan Long, Milica Gašić, and Kai Yu. 2019. Agentgraph: Toward universal dialogue management with structured deep reinforcement learning. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(9):1378–1391.
  7. 7.Lu Chen, Boer Lv, Chi Wang, Su Zhu, Bowen Tan, and Kai Yu. 2020a. Schema-guided multi-domain dialogue state tracking with graph attention neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7521–7528.
  8. 8.Zhi Chen, Lu Chen, Xiaoyuan Liu, and Kai Yu. 2020b. Distributed structured actor-critic reinforcement learning for universal dialogue management. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2400–2411.
  9. 9.Zhi Chen, Jiabao Ji, Lu Chen, Yuncong Liu, Da Ma, Bei Chen, Mengyue Wu, Su Zhu, Xin Dong, Fujiang Ge, Qingliang Miao, Jian-Guang Lou, and Kai Yu. 2022. Dfm: Dialogue foundation model for universal large-scale dialogue-oriented task learning. arXiv preprint arXiv:2205.12662.
  10. 10.Tech Crunch. 2019. Google is bringing ai assistant duplex to the web.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.Zhe Gan, Yu Cheng, Ahmed Kholy, Linjie Li, Jingjing Liu, and Jianfeng Gao. 2019. Multi-step reasoning via recurrent dual attention for visual dialog. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6463–6474.
  13. 13.Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby Lee, and Jindong Chen. 2021. Actionbert: Leveraging user actions for semantic understanding of user interfaces. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 5931–5938.
  14. 14.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  15. 15.Hung Le and Steven C.H. Hoi. 2020. Video-grounded dialogues with pretrained generation language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5842–5848, Online. Association for Computational Linguistics.
  16. 16.Hung Le, Doyen Sahoo, Nancy Chen, and Steven C.H. Hoi. 2020. BiST: Bi-directional spatio-temporal reasoning for video-grounded dialogues. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1846–1859, Online. Association for Computational Linguistics.
  17. 17.Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017. Sugilite: creating multimodal smartphone automation by demonstration. In Proceedings of the 2017 CHI conference on human factors in computing systems, pages 6038–6049.
  18. 18.Toby Jia-Jun Li, Igor Labutov, Xiaohan Nancy Li, Xiaoyi Zhang, Wenze Shi, Wanling Ding, Tom M Mitchell, and Brad A Myers. 2018. Appinite: A multi-modal interface for specifying data descriptions in programming by demonstration using natural language instructions. In 2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pages 105–114. IEEE.
  19. 19.Toby Jia-Jun Li, Marissa Radensky, Justin Jia, Kirielle Singarajah, Tom M Mitchell, and Brad A Myers. 2019. Pumice: A multi-modal agent that learns concepts and conditionals from natural language and demonstrations. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology, pages 577–589.
  20. 20.Toby Jia-Jun Li and Oriana Riva. 2018. Kite: Building conversational bots from mobile apps. In Proceedings of the 16th Annual International Conference on Mobile Systems, Applications, and Services, pages 96–109.
  21. 21.Yuanchun Li and Oriana Riva. 2021. Glider: A reinforcement learning approach to extract ui scripts from websites. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1420–1430.
  22. 22.Sahisnu Mazumder and Oriana Riva. 2021. Flin: A flexible natural language interface for web navigation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2777–2788.
  23. 23.Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, and Erik Cambria. 2022. Recent advances in deep learning based dialogue systems: A systematic survey. Artificial Intelligence Review, pages 1–101.
  24. 24.Panupong Pasupat, Tian-Shun Jiang, Evan Zheran Liu, Kelvin Guu, and Percy Liang. 2018. Mapping natural language commands to web elements. In EMNLP.
  25. 25.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28.
  26. 26.Oriana Riva and Jason Kace. 2021. Etna: Harvesting action graphs from websites. In The 34th Annual ACM Symposium on User Interface Software and Technology, pages 312–331.
  27. 27.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  28. 28.Yue Wang, Shafiq Joty, Michael Lyu, Irwin King, Caiming Xiong, and Steven C.H. Hoi. 2020. VD-BERT: A Unified Vision and Dialog Transformer with BERT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3325–3338, Online. Association for Computational Linguistics.
  29. 29.Nancy Xu, Sam Masling, Michael Du, Giovanni Campagna, Larry Heck, James Landay, and Monica Lam. 2021a. Grounding open-domain instructions to automate web support tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1022–1032.
  30. 30.Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2021b. LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2579–2591, Online. Association for Computational Linguistics.
  31. 31.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020. Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1192–1200.
  32. 32.Steve Young. 2007. Cued standard dialogue acts. Report, Cambridge University Engineering Department, 14th October, 2007.
  33. 33.Kai Yu, Lu Chen, Bo Chen, Kai Sun, and Su Zhu. 2014. Cognitive technology in task-oriented dialogue systems: Concepts, advances and future. Chinese Journal of Computers, 37(18):1–17.
  34. 34.Zheng Zhang, Ryuichi Takanobu, Qi Zhu, MinLie Huang, and XiaoYan Zhu. 2020. Recent advances and challenges in task-oriented dialog systems. Science China Technological Sciences, pages 1–17.
  35. 35.Xin Zhou and Yang Li. 2021. Large-scale modeling of mobile user click behaviors using deep learning. In Fifteenth ACM Conference on Recommender Systems, pages 473–483.
  36. 36.Su Zhu, Lu Chen, Ruisheng Cao, Zhi Chen, Qingliang Miao, and Kai Yu. 2021. Few-shot nlu with vector projection distance and abstract triangular crf. In CCF International Conference on Natural Language Processing and Chinese Computing, pages 505–516. Springer.
  37. 37.Su Zhu, Jieyu Li, Lu Chen, and Kai Yu. 2020. Efficient context and schema fusion networks for multi-domain dialogue state tracking. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 766–781, Online. Association for Computational Linguistics.

Citation

MLA
Sun, L., et al. “META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6699–712, https://doi.org/10.18653/v1/2022.emnlp-main.449.
APA
Sun, L., Chen, X., Chen, L., Dai, T., Zhu, Z., & Yu, K. (2022). META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6699–6712. https://doi.org/10.18653/v1/2022.emnlp-main.449
Chicago
Sun, L., X. Chen, L. Chen, T. Dai, Z. Zhu, and K. Yu. 2022. “META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6699–6712. https://doi.org/10.18653/v1/2022.emnlp-main.449.
Harvard
Sun, L. et al. (2022) “META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6699–6712. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.449.
Vancouver
1. Sun L, Chen X, Chen L, Dai T, Zhu Z, Yu K (2022) META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6699–6712

BibTeX

@inproceedings{sun-etal-2022-meta,
    title = "{META}-{GUI}: Towards Multi-modal Conversational Agents on Mobile {GUI}",
    author = "Sun, Liangtai  and
      Chen, Xingyu  and
      Chen, Lu  and
      Dai, Tianle  and
      Zhu, Zichen  and
      Yu, Kai",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.449/",
    doi = "10.18653/v1/2022.emnlp-main.449",
    pages = "6699--6712"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/