Built independently by an author, for readers. Read the story and support ChapterPal

keyword

dialogue management

Dialogue management is the central component in a conversational agent that controls the flow of interaction and determines the system actions required to fulfill a user goal. It typically coordinates two primary functions: dialogue state tracking, which continuously updates the internal representation of user intentions, accumulated constraints, and conversational context across multiple turns, and dialogue policy selection, which decides the next optimal action to take. Based on the tracked state, the dialogue manager determines whether to query external databases, invoke application interfaces, request clarifying information, or guide the generation of a relevant response. By maintaining context and directing task execution, dialogue management enables automated conversational systems to conduct coherent, flexible, and goal-directed exchanges with human users.

3 items

META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, Kai Yu

OrganizationsAI InstituteDepartment of Computer Science and EngineeringShanghai Jiao Tong University

Why you should read this

Proposes a GUI-based task-oriented dialogue framework and benchmark dataset, META-GUI, enabling conversational assistants to complete multi-turn tasks by interacting directly with mobile app interfaces rather than relying on restrictive backend APIs.

Task-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation. Current TOD systems usually focus on multi-turn text/speech interaction, then they would call back-end APIs designed for TODs to perform the task. However, this API-based architecture greatly limits the information-searching capability of intelligent assistants and may even lead to task failure if TOD-specific APIs are not available or the task is too complicated to be executed by the provided APIs. In this paper, we propose a new TOD architecture: GUI-based task-oriented dialogue system (GUI-TOD). A GUI-TOD system can directly perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs. Furthermore, we release META-GUI, a dataset for training a Multi-modal conversational Agent on mobile GUI. We also propose a multi-model action prediction and response model, which show promising results on META-GUI. The dataset, codes and leaderboard are publicly available.

Added

2026-09-26

MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling

MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling

Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, Milica Gašić

OrganizationsPolyAIUniversity of Cambridge

Why you should read this

Introduces MultiWOZ, an open-source corpus of over 10,000 multi-domain conversations with dialogue state and action annotations that overcomes previous data scarcity barriers and establishes standardized baselines for task-oriented dialogue systems.

Even though machine learning has become the major scene in dialogue research community, the real breakthrough has been blocked by the scale of data available. To address this fundamental obstacle, we introduce the Multi-Domain Wizard-of-Oz dataset (MultiWOZ), a fully-labeled collection of human-human written conversations spanning over multiple domains and topics. At a size of 1010k dialogues, it is at least one order of magnitude larger than all previous annotated task-oriented corpora. The contribution of this work apart from the open-sourced dataset labelled with dialogue belief states and dialogue actions is two-fold: firstly, a detailed description of the data collection procedure along with a summary of data structure and analysis is provided. The proposed data-collection pipeline is entirely based on crowd-sourcing without the need of hiring professional annotators; secondly, a set of benchmark results of belief tracking, dialogue act and response generation is reported, which shows the usability of the data and sets a baseline for future studies.

Added

2026-09-24