keyword
dialogue management
Dialogue management is the central component in a conversational agent that controls the flow of interaction and determines the system actions required to fulfill a user goal. It typically coordinates two primary functions: dialogue state tracking, which continuously updates the internal representation of user intentions, accumulated constraints, and conversational context across multiple turns, and dialogue policy selection, which decides the next optimal action to take. Based on the tracked state, the dialogue manager determines whether to query external databases, invoke application interfaces, request clarifying information, or guide the generation of a relevant response. By maintaining context and directing task execution, dialogue management enables automated conversational systems to conduct coherent, flexible, and goal-directed exchanges with human users.
3 items

CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning
Siddharth Verma, Justin Fu, Sherry Yang, Sergey Levine
Why you should read this
Presents CHAI, a framework that combines pre-trained language models with offline reinforcement learning to train task-oriented dialogue agents directly from static human conversation datasets without requiring live interactions or simulated user models.
Conventionally, generation of natural language for dialogue agents may be viewed as a statistical learning problem: determine the patterns in human-provided data and generate appropriate responses with similar statistical properties. However, dialogue can also be regarded as a goal directed process, where speakers attempt to accomplish a specific task. Reinforcement learning (RL) algorithms are designed specifically for solving such goal-directed problems, but the most direct way to apply RL – through trial-and-error learning in human conversations, – is costly. In this paper, we study how offline reinforcement learning can instead be used to train dialogue agents entirely using static datasets collected from human speakers. Our experiments show that recently developed offline RL methods can be combined with language models to yield realistic dialogue agents that better accomplish task goals.
Added
2026-10-03

META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, Kai Yu
Why you should read this
Proposes a GUI-based task-oriented dialogue framework and benchmark dataset, META-GUI, enabling conversational assistants to complete multi-turn tasks by interacting directly with mobile app interfaces rather than relying on restrictive backend APIs.
Task-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation. Current TOD systems usually focus on multi-turn text/speech interaction, then they would call back-end APIs designed for TODs to perform the task. However, this API-based architecture greatly limits the information-searching capability of intelligent assistants and may even lead to task failure if TOD-specific APIs are not available or the task is too complicated to be executed by the provided APIs. In this paper, we propose a new TOD architecture: GUI-based task-oriented dialogue system (GUI-TOD). A GUI-TOD system can directly perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs. Furthermore, we release META-GUI, a dataset for training a Multi-modal conversational Agent on mobile GUI. We also propose a multi-model action prediction and response model, which show promising results on META-GUI. The dataset, codes and leaderboard are publicly available.
Added
2026-09-26

MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, Milica Gašić
Why you should read this
Introduces MultiWOZ, an open-source corpus of over 10,000 multi-domain conversations with dialogue state and action annotations that overcomes previous data scarcity barriers and establishes standardized baselines for task-oriented dialogue systems.
Even though machine learning has become the major scene in dialogue research community, the real breakthrough has been blocked by the scale of data available. To address this fundamental obstacle, we introduce the Multi-Domain Wizard-of-Oz dataset (MultiWOZ), a fully-labeled collection of human-human written conversations spanning over multiple domains and topics. At a size of k dialogues, it is at least one order of magnitude larger than all previous annotated task-oriented corpora. The contribution of this work apart from the open-sourced dataset labelled with dialogue belief states and dialogue actions is two-fold: firstly, a detailed description of the data collection procedure along with a summary of data structure and analysis is provided. The proposed data-collection pipeline is entirely based on crowd-sourcing without the need of hiring professional annotators; secondly, a set of benchmark results of belief tracking, dialogue act and response generation is reported, which shows the usability of the data and sets a baseline for future studies.
Added
2026-09-24
