MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
Paweł BudzianowskiTsung-Hsien WenBo-Hsiang TsengIñigo CasanuevaStefan UltesOsman RamadanMilica Gašić
Introduces MultiWOZ, an open-source corpus of over 10,000 multi-domain conversations with dialogue state and action annotations that overcomes previous data scarcity barriers and establishes standardized baselines for task-oriented dialogue systems.
Building automated conversational agents capable of handling complex tasks across multiple domains is critical for modern voice and chat applications. However, progress in machine learning for dialogue systems has long been hindered by a shortage of large-scale, fully labeled datasets reflecting natural human interactions. Existing resources are typically small, restricted to single domains, or lacking the detailed semantic labels needed to train and evaluate core system modules.
The article introduces MultiWOZ, an open-source, fully labeled collection of human-human conversations designed to support multi-domain task-oriented dialogue modeling. The study evaluates the dataset by establishing comprehensive performance baselines across tracking user intent, managing conversation context, and generating language responses.
To build the dataset, the authors implemented a crowd-sourced Wizard-of-Oz pipeline involving 1,249 workers who simulated tourist-and-clerk conversations across seven domains, including hotels, restaurants, and transportation. The team captured 10,438 dialogues comprising 115,434 turns and nearly 1.5 million words, making it at least an order of magnitude larger than previous structured task-oriented corpora. The collection used a multi-phase screening process to ensure high-quality dialogue act annotations, achieving strong inter-annotator agreement (Fleiss kappa of 0.884).
The article demonstrates that MultiWOZ presents a substantially more rigorous testbed than earlier benchmarks. State tracking accuracy for joint goals dropped from 85.5% on the prior WOZ 2.0 dataset to 80.9% on MultiWOZ's restaurant subset. In end-to-end response generation, baseline models experienced a drop in task-information success of roughly 28 percentage points compared to older single-domain datasets (from 99.6% down to 71.3%). Furthermore, natural language generation models suffered a tenfold increase in slot error rates (rising from 0.46% on the SFX benchmark to 4.38% on MultiWOZ) because nearly 60% of system turns contain multiple concurrent conversational actions.
These findings indicate that existing dialogue architectures are ill-equipped for the linguistic diversity, multi-intent turns, and context switching found in real-world human conversations. While previous models appeared near-perfect on simpler datasets, their performance degrades significantly when scaled to realistic multi-domain tasks.
Organizations developing conversational AI should adopt MultiWOZ as a standard benchmark to stress-test systems before deployment. Engineering teams should prioritize developing architectures that can handle compound actions within a single turn and preserve context across disparate business domains. Researchers should also pursue end-to-end modeling frameworks to eliminate cascading errors between separate pipeline components.
Confidence in these findings is high due to the dataset's unprecedented scale and rigorous quality-control checks. However, stakeholders should note that the dataset reflects written rather than spoken English, and models evaluated with automatic metrics must eventually be validated in real-time user trials to assess true operational performance.
- Paper: DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset, Yanran Li et al. (2017). This paper establishes foundational methodologies for annotating multi-turn dialogues with dialogue acts and intent categories, which MultiWOZ expands to a multi-domain task-oriented setting.
- Paper: A Neural Conversational Model, Oriol Vinyals et al. (2015). It introduces early end-to-end neural sequence-to-sequence conversational modeling that motivated large-scale dataset collection for data-driven dialogue generation.
- Paper: Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models, Iulian Serban et al. (2015). It provides the generative hierarchical neural dialogue modeling architecture that serves as an essential baseline and modeling paradigm for multi-turn conversational agents.
- Paper: Dialogue act modeling for automatic tagging and recognition of conversational speech, Andreas Stolcke et al. (2000). It lays the groundwork for formal dialogue act tagging and taxonomy in conversational speech, directly informing the dialogue action labeling scheme used in MultiWOZ.
- Paper: DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation, Yizhe Zhang et al. (2019). DialoGPT advances large-scale conversational pre-training beyond traditional supervised task-oriented datasets like MultiWOZ to open-domain generative response architectures.
- Paper: LaMDA: Language Models for Dialog Applications, Romal Thoppilan et al. (2022). LaMDA builds upon task-oriented and multi-turn dialogue principles by scaling language models for dialog applications integrated with factual grounding and external tool usage.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). CAMEL extends multi-turn task completion dialogue by orchestrating autonomous communicative LLM agent societies using role-playing and communicative prompting.
- Paper: A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity, Yejin Bang et al. (2023). This study evaluates state-of-the-art conversational LLMs on multi-turn dialogue tracking and interactive generation using benchmarks that trace back to standard conversational corpora.
