Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System

Yixuan SuLei ShuElman MansimovArshit GuptaDeng CaiYi-An LaiYi Zhang

article2022ACL222 citations

Proposes a unified, prompt-driven task-oriented dialogue model pre-trained across diverse corpora to decouple sub-tasks, mitigating error accumulation and achieving state-of-the-art results in both high- and low-resource settings.

Listen

Building automated conversational agents for customer assistance typically requires coordinating multiple sub-tasks, including intent recognition, dialogue state tracking, policy decision-making, and natural language response generation. Prevailing systems process these steps sequentially in a cascaded pipeline. This structure creates significant operational bottlenecks: prediction errors compound from one step to the next, end-to-end data annotation is prohibitively expensive, and sequential execution introduces high inference latency. The article introduces PPTOD (Plug-and-Play Task-Oriented Dialogue), a unified model and multi-task training framework designed to eliminate sequential dependencies and allow conversational systems to learn efficiently from diverse, partially labeled datasets.

To address these limitations, the authors frame all dialogue sub-tasks as prompt-driven text generation within a single pre-trained language model based on T5. By inserting task-specific natural language instructions into the dialogue history, the architecture decouples sub-tasks so they can run independently or in parallel. The model is pre-trained across eleven public dialogue datasets comprising more than 2.3 million utterances and 80 domains, where individual datasets only possess labels for specific sub-tasks rather than the entire pipeline. The authors evaluated the system against top benchmark models on end-to-end multi-domain dialogue benchmarks (MultiWOZ 2.0 and 2.1) and banking intent classification (Banking77) across both standard and constrained-data scenarios.

The findings show that PPTOD sets a new state of the art across benchmark tasks, demonstrating substantial advantages in data efficiency and speed. In low-resource environments, the model achieved large gains; when trained on only 1% of target dialogue data, PPTOD exceeded the best existing baseline by roughly 18 percentage points in state tracking accuracy. In end-to-end dialogue modeling, it reduced system latency by approximately four-fold compared to competitive cascaded architectures while achieving higher response accuracy. Furthermore, human evaluations confirmed that the system produced statistically significant improvements in response truthfulness and contextual coherency over existing baselines.

These results demonstrate that organizations can deploy higher-quality conversational assistants with substantially reduced data labeling expenses and lower compute latency. Decoupling sub-tasks into a flexible generation framework allows engineering teams to train assistants using legacy or partially labeled dialogue logs without paying for full-pipeline data annotation. Engineering leaders should evaluate moving away from strict cascaded pipelines toward unified prompt-based generation architectures, particularly for domains where labeled dialogue data is scarce.

Decision-makers should note that the largest evaluated model variant (PPTOD-large) underperformed smaller variants in generating specialized response tokens, indicating that simply scaling model size without targeted output adaptation is suboptimal. Future work should focus on piloting the model in production environments with live databases and testing prompt robustness across highly dynamic business ontologies before committing to full-scale platform migrations.

Cover for Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System

Abstract

Pre-trained language models have been recently shown to benefit task-oriented dialogue (TOD) systems. Despite their success, existing methods often formulate this task as a cascaded generation problem which can lead to error accumulation across different sub-tasks and greater data annotation overhead. In this study, we present PPTOD, a unified plug-and-play model for task-oriented dialogue. In addition, we introduce a new dialogue multi-task pre-training strategy that allows the model to learn the primary TOD task completion skills from heterogeneous dialog corpora. We extensively test our model on three benchmark TOD tasks, including end-to-end dialogue modelling, dialogue state tracking, and intent classification. Experimental results show that PPTOD achieves new state of the art on all evaluated tasks in both high-resource and low-resource scenarios. Furthermore, comparisons against previous SOTA methods show that the responses generated by PPTOD are more factually correct and semantically coherent as judged by human annotators.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Pre-training Datasets
  • 3.2 Dialogue Multi-Task Pre-training
  • 3.3 Fine-Tuning to a New Task
  • 3.4 Implementation Details
  • 4 Experiments
  • 4.1 End-to-End Dialogue Modelling
  • 4.1.1 Dataset and Evaluation Metric
  • 4.1.2 Baselines
  • 4.1.3 Full Training Evaluation
  • 4.1.4 Low-Resource Evaluation
  • 4.2 Dialogue State Tracking
  • 4.2.1 Full Training Evaluation
  • 4.2.2 Low-Resource Evaluation
  • 4.3 Intent Classification
  • 5 Further Analysis
  • 5.1 Plug-and-Play vs Cascaded Generation
  • 5.2 Multi-Task Pre-Training Investigation
  • 5.3 Human Evaluation
  • 6 Conclusion
  • Acknowledgments
  • Ethical Statement
  • References
  • A Dataset Details
  • B Low-Resource MultiWOZ Evaluation
  • C Human Evaluation Guidelines
  • C.1 Understanding
  • C.2 Truthfulness
  • C.3 Coherency
  • C.4 Fluency
  • D Case Study

Knowls

  1. Knowl 1 — Plug-and-Play Task-Oriented Dialogue (PPTOD) Framework

    model/method

    Plug-and-Play Task-Oriented Dialogue (PPTOD) is a unified sequence-to-sequence framework that casts all task-oriented dialogue (TOD) sub-tasks into prompt-guided text generation. Existing task-oriented dialogue systems typically formulate dialogue processing as a cascaded generation pipeline where dialogue policy learning (POL) and natural language generation (NLG) are strictly conditioned on the explicit generation of previous steps like dialogue state tracking (DST). Cascaded execution causes error propagation across stages, prevents training on corpora with only partial annotations, and increases inference latency.

    PPTOD decouples sub-task execution by prefixing the dialogue context xx (the concatenation of preceding user and system utterances) with a natural language task prompt ztz_t, producing an input (zt,x)(z_t, x) to predict the corresponding target output text yy. The prompt templates for the four supported sub-tasks are:

    • Natural Language Understanding (NLU / Intent Classification): zt="translate dialogue to user intent:"z_t = \text{"translate dialogue to user intent:"} with target output y = \text{"[intent_name]"}
    • Dialogue State Tracking (DST): zt="translate dialogue to belief state:"z_t = \text{"translate dialogue to belief state:"} with target output y="[domain] slot = value, ..."y = \text{"[domain] {slot = value, ...}"}
    • Dialogue Policy Learning (POL): zt="translate dialogue to dialogue act:"z_t = \text{"translate dialogue to dialogue act:"} with target output y = \text{"[domain] {[action] = slot_name}}"
    • Natural Language Generation (NLG): zt="translate dialogue to system response:"z_t = \text{"translate dialogue to system response:"} with target output yy being the dialogue response text.

    During end-to-end inference with database grounding, PPTOD first predicts the DST belief state to retrieve database (DB) match information. Conditioned on the retrieved DB state and dialogue context, PPTOD then generates the dialogue act (POL) and system response (NLG) in parallel.

  2. Knowl 2 — Dialogue Multi-Task Pre-Training Objective and Algorithm

    algorithm

    PPTOD pre-trains an underlying encoder-decoder language model (initialized from T5-small with ~60M parameters, T5-base with ~220M parameters, or T5-large with ~770M parameters) on a unified collection of 11 multi-turn dialogue corpora: MetaLWOZ, SNIPS, CLINC, ATIS, KVRET, WOZ, CamRest676, MSR-E2E, Frames, TaskMaster, and Schema-Guided. Across these datasets, over 2.3 million utterances covering more than 80 domains provide heterogeneous, partial annotations for natural language understanding (NLU), dialogue state tracking (DST), dialogue policy learning (POL), and natural language generation (NLG).

    Each pre-training instance is a tuple d=(zt,x,y)d = (z_t, x, y), where t∈{NLU,DST,POL,NLG}t \in \{\text{NLU}, \text{DST}, \text{POL}, \text{NLG}\}, ztz_t is the task-specific prompt, xx is the conversation context, and yy is the target label or text sequence. The model parameters Θ\Theta are optimized via standard sequence-to-sequence maximum likelihood estimation: LΘ=−∑i=1∣y∣log⁡PΘ(yi∣y<i;zt,x)\mathcal{L}_\Theta = -\sum_{i=1}^{|y|} \log P_\Theta(y_i \mid y_{<i}; z_t, x)

    Input: Dataset D={(zt,x,y)i}i=1∣D∣\mathcal{D} = \{(z_t, x, y)_i\}_{i=1}^{|\mathcal{D}|}, trainer T\mathcal{T} optimizing parameters Θ\Theta, maximum epochs emax⁡e_{\max}
    Output: Pre-trained model parameters Θ\Theta
    for e=1,…,emax⁡e = 1, \dots, e_{\max} do
        Shuffle D\mathcal{D} by mixing samples from all different TOD tasks
        for each mini-batch B={(zt,x,y)k}k=1∣B∣B = \{(z_t, x, y)_k\}_{k=1}^{|B|} in D\mathcal{D} do
            Invoke trainer T\mathcal{T} on batch BB to update Θ\Theta with gradient descent on LΘ\mathcal{L}_\Theta
        end for
    end for
    return Θ\Theta

    Pre-training runs for 10 epochs using the Adam optimizer with a learning rate of 5×10−55 \times 10^{-5}, batch size of 128, and a maximum sequence length of 1024 tokens. Downstream task fine-tuning uses the identical loss formulation.

  3. Knowl 3 — MultiWOZ End-to-End Dialogue Modeling Benchmark Results

    data/table

    PPTOD was evaluated on the MultiWOZ 2.0 and MultiWOZ 2.1 benchmarks in the fully end-to-end dialogue modeling setup, where the model generates belief states to query the database and subsequently generates dialogue acts and responses. Performance is measured using Inform rate (%), Success rate (%), BLEU score, and Combined Score, defined as Combined=(Inform+Success)×0.5+BLEU\text{Combined} = (\text{Inform} + \text{Success}) \times 0.5 + \text{BLEU}.

    Model MultiWOZ 2.0 MultiWOZ 2.1
    Inform Success BLEU Combined Inform Success BLEU Combined
    Sequicity 66.41 45.32 15.54 71.41 - - - -
    MD-Sequicity 75.72 58.32 15.40 82.40 - - - -
    DAMD 76.33 60.40 16.60 84.97 - - - -
    MinTL 84.88 74.91 17.89 97.78 - - - -
    HIER-Joint 80.50 71.70 19.74 95.84 - - - -
    SOLOIST 85.50 72.90 16.54 95.74 - - - -
    TOP 85.20 72.90 17.00 96.05 - - - -
    TOP+NOD 86.90 76.20 20.58 102.13 - - - -
    LABES-S2S - - - - 78.07 67.06 18.13 90.69
    UBAR 85.10 71.02 16.21 94.27 86.20 70.32 16.48 94.74
    SimpleTOD 84.40 70.10 15.01 92.26 85.00 70.50 15.23 92.98
    PPTODsmall_{\text{small}} 87.80 75.30 19.89 101.44 88.89 76.98 18.59 101.52
    PPTODbase_{\text{base}} 89.20 79.40 18.62 102.92 87.09 79.08 19.17 102.26
    PPTODlarge_{\text{large}} 82.60 74.10 19.21 97.56 86.43 74.35 17.89 98.28

    PPTODbase_{\text{base}} achieves state-of-the-art combined scores on both MultiWOZ 2.0 (102.92) and MultiWOZ 2.1 (102.26) as a single end-to-end model, outperforming multi-model re-ranking methods such as TOP+NOD (which requires two auxiliary language models for output re-scoring) and oracle-dependent systems.

  4. Knowl 4 — Low-Resource End-to-End Dialogue Modeling Performance

    data/table

    PPTOD was evaluated under low-resource training conditions on MultiWOZ 2.0 using 1% (~80 dialogues), 5%, 10%, and 20% (~1600 dialogues) of the training dataset. Results report average performance over five independent runs with different random seeds and sample selections.

    Model 1% Data 5% Data 10% Data 20% Data
    Inf. Succ. Comb. Inf. Succ. Comb. Inf. Succ. Comb. Inf. Succ. Comb.
    MD-Sequicity - - - 49.40 19.70 44.85 58.10 34.70 57.80 64.40 42.10 66.25
    DAMD 34.40 9.10 29.85 52.50 31.80 53.75 55.30 30.30 55.80 62.60 44.10 68.25
    SOLOIST 58.40 35.30 57.43 69.30 52.30 72.60 69.90 51.90 75.50 74.00 60.10 82.29
    MinTL - - - 75.48 60.96 82.20 78.08 66.87 87.94 82.48 68.57 88.53
    PPTODsmall_{\text{small}} 66.96 50.90 71.44 76.58 61.60 84.44 83.50 68.18 91.01 82.96 69.90 93.45
    PPTODbase_{\text{base}} 74.42 52.44 76.41 79.86 63.48 86.55 84.42 68.36 91.96 84.94 71.70 95.32
    PPTODlarge_{\text{large}} 64.38 51.94 70.01 75.20 61.94 82.54 80.64 66.74 88.94 81.74 72.18 92.09

    PPTOD consistently outperforms previous pre-trained and modular architectures across all low-resource regimes. In the 1% data setting, PPTODbase_{\text{base}} achieves a Combined Score of 76.41 compared to SOLOIST's 57.43 (+18.98 points). When trained on only 20% of the data, PPTODbase_{\text{base}} reaches a Combined Score of 95.32, matching the performance of baseline models (such as SOLOIST at 95.74) trained on 100% of the data.

  5. Knowl 5 — Dialogue State Tracking Performance Across Full and Low-Resource Settings

    data/table

    PPTOD was evaluated on dialogue state tracking (DST) on MultiWOZ 2.0 and MultiWOZ 2.1 using Joint Goal Accuracy (%). PPTOD generates belief states directly rather than classifying over a fixed ontology of pre-defined slot-value pairs.

    Model Full Data Joint Acc. (%) MultiWOZ 2.0 Low-Resource Joint Acc. (%) (Mean ±\pm Std)
    MWOZ 2.0 MWOZ 2.1 1% Data 5% Data 10% Data 20% Data
    SimpleTOD - 55.76 7.91 ±\pm 1.07 16.14 ±\pm 1.48 22.37 ±\pm 1.17 31.22 ±\pm 2.32
    MinTL 52.10 53.62 9.25 ±\pm 2.33 21.28 ±\pm 1.94 30.32 ±\pm 2.14 35.96 ±\pm 1.25
    SOLOIST 53.20 56.85 13.21 ±\pm 1.97 26.53 ±\pm 1.62 32.42 ±\pm 1.13 38.68 ±\pm 0.98
    TRADE 48.62 46.00 - - - -
    UBAR 52.59 56.20 - - - -
    Seq2seq-DU - 56.10 - - - -
    PPTODsmall_{\text{small}} 51.50 56.47 27.85 ±\pm 0.77 39.07 ±\pm 0.85 42.36 ±\pm 0.29 45.98 ±\pm 0.38
    PPTODbase_{\text{base}} 53.37 57.10 29.72 ±\pm 0.61 40.20 ±\pm 0.39 43.45 ±\pm 0.64 46.96 ±\pm 0.40
    PPTODlarge_{\text{large}} 53.89 57.45 31.46 ±\pm 0.41 43.61 ±\pm 0.42 45.96 ±\pm 0.66 48.95 ±\pm 0.13

    PPTODlarge_{\text{large}} achieves the highest Joint Goal Accuracy among generation-based DST models on both full datasets (53.89% on MultiWOZ 2.0 and 57.45% on MultiWOZ 2.1). In the 1% low-resource setting, PPTODlarge_{\text{large}} outperforms SOLOIST by 18.25 percentage points (31.46% vs. 13.21%).

  6. Knowl 6 — Plug-and-Play Decoupling vs. Cascaded Generation and Inference Latency

    empirical result

    An ablation study on MultiWOZ 2.0 using a T5-small backbone without multi-task pre-training isolates the effects of the plug-and-play decoupled generation scheme versus traditional cascaded generation, with and without Database (DB) state inputs. Inference latency was measured on a single Nvidia V100 GPU with a batch size of 1.

    Model Generation Mode DB State Inform Success BLEU Combined Latency (Speedup)
    SOLOIST Cascaded Yes 85.50 72.90 16.54 95.74 208.69 ms (1.00×\times)
    MinTL Cascaded Yes 84.88 74.91 17.89 97.78 78.82 ms (2.65×\times)
    T5-small Cascaded No 83.60 71.20 18.09 95.49 38.70 ms (5.39×\times)
    T5-small Cascaded Yes 84.10 73.70 18.03 96.93 39.78 ms (5.25×\times)
    T5-small Plug-and-Play No 84.70 72.80 18.52 97.27 14.17 ms (14.73×\times)
    T5-small Plug-and-Play Yes 85.10 75.10 17.82 97.92 19.52 ms (10.69×\times)

    The plug-and-play formulation achieves higher dialogue quality scores than its cascaded counterpart in both DB configurations (e.g., 97.92 vs. 96.93 Combined Score with DB) because it removes explicit conditioning on potentially erroneous sub-task outputs. Furthermore, parallel sub-task generation reduces latency from 39.78 ms to 19.52 ms with DB grounding (a 10.69×10.69\times speedup over SOLOIST and ∼4×\sim 4\times speedup over MinTL), and down to 14.17 ms (14.73×14.73\times speedup) without DB grounding.

  7. Knowl 7 — Intent Classification via Generative Prompting on Banking77

    empirical result

    PPTOD formulates Natural Language Understanding (NLU) intent classification as sequence generation of the intent label string given the prompt translate dialogue to user intent:, requiring no task-specific classification head or added parameters. On the Banking77 benchmark (77 fine-grained intents across 10-sample, 30-sample, and full-data settings), classification accuracy (%) is:

    Model 10 Samples/Intent 30 Samples/Intent Full Training
    BERT-Fixed 67.55 80.07 87.19
    BERT-Tuned 83.42 90.03 93.66
    USE 84.23 89.74 92.81
    ConveRT 83.32 89.37 93.01
    USE+ConveRT 85.19 90.57 93.36
    SOLOIST 78.73 89.28 93.80
    PPTODsmall_{\text{small}} 78.87 ±\pm 0.36 87.88 ±\pm 0.26 93.27 ±\pm 0.39
    PPTODbase_{\text{base}} 82.81 ±\pm 0.45 89.64 ±\pm 0.28 93.86 ±\pm 0.22
    PPTODlarge_{\text{large}} 84.12 ±\pm 0.23 90.64 ±\pm 0.29 94.08 ±\pm 0.15

    PPTODlarge_{\text{large}} achieves the highest accuracy in the 30-sample (90.64%) and full-data (94.08%) settings without introducing task-specific classification weights.

  8. Knowl 8 — Sub-Task Pre-Training Ablation Across Downstream Benchmarks

    empirical result

    To evaluate the transfer capability of each sub-task in the pre-training mixture, T5-small was pre-trained individually on datasets annotated for single TOD sub-tasks (NLU, DST, POL, NLG) and evaluated on MultiWOZ 2.0 (End-to-End Dialogue and DST) and Banking77 (Intent Classification):

    Pre-training Data End-to-End (MWOZ 2.0) DST Joint Acc. Intent Acc. (Banking77)
    NLU DST POL NLG 1% Data Full Data 1% Data Full Data 10 samp. Full Data
    Inf./Succ. BLEU Inf./Succ. BLEU Acc. Acc. Acc. Acc.
    ×\times ×\times ×\times ×\times 53.28 / 36.08 11.65 83.10 / 72.40 18.17 17.44 50.55 75.12 92.91
    ✓ ×\times ×\times ×\times 58.58 / 40.48 11.02 85.20 / 73.50 16.96 18.47 50.71 78.21 93.37
    ×\times ✓ ×\times ×\times 66.10 / 46.40 11.26 86.30 / 74.90 18.52 27.91 51.48 75.97 93.03
    ×\times ×\times ✓ ×\times 60.60 / 48.20 11.88 84.40 / 74.60 18.55 19.32 50.82 75.37 92.95
    ×\times ×\times ×\times ✓ 59.38 / 40.78 12.34 83.60 / 74.70 19.97 17.82 50.58 75.61 92.97
    ✓ ✓ ✓ ✓ 66.96 / 50.90 12.51 87.80 / 75.30 19.89 27.85 51.50 78.87 93.27

    Pre-training on a specific task directly improves downstream performance on that exact task: DST pre-training raises 1% DST accuracy from 17.44% to 27.91%; NLU pre-training raises 10-sample intent accuracy from 75.12% to 78.21%; and NLG pre-training maximizes full-data BLEU (19.97). Training jointly on all four sub-tasks produces the best balanced performance across all evaluation metrics.

  9. Knowl 9 — Human Evaluation on Generated Response Quality

    empirical result

    Human evaluation was conducted on 50 randomly selected dialogue sessions from the MultiWOZ 2.0 test set. Five English-proficient annotators evaluated responses generated by PPTODbase_{\text{base}}, SOLOIST, and the ground-truth human reference on a 3-point Likert scale (0, 1, or 2) across four criteria:

    • Understanding: Whether the system correctly understands the user's intent.
    • Truthfulness: Whether facts in the response are supported by the reference.
    • Coherency: Whether the response is logically and semantically coherent with prior turns.
    • Fluency: Grammatical correctness and naturalness.
    System Understanding Truthfulness Coherency Fluency
    Fleiss' κ\kappa Agreement 0.641 0.598 0.668 0.806
    Reference 1.92 2.00 1.93 1.98
    SOLOIST 1.78 1.29 1.64 1.97
    PPTODbase_{\text{base}} 1.86 1.51 1.83 1.99

    PPTODbase_{\text{base}} significantly outperforms SOLOIST on Truthfulness (1.51 vs. 1.29) and Coherency (1.83 vs. 1.64) (p<0.05p < 0.05 via Sign Test). On Fluency, both SOLOIST (1.97) and PPTODbase_{\text{base}} (1.99) perform on par with human references (1.98, p>0.4p > 0.4), demonstrating that pre-trained language models inherently supply strong grammatical capability, whereas sub-task prompting and multi-task pre-training improve grounding and context consistency.

  10. Knowl 10 — Performance Degradation of PPTOD-Large on Delexicalized Response Generation

    limitation

    In full-training end-to-end dialogue modeling on MultiWOZ 2.0 and MultiWOZ 2.1, PPTODlarge_{\text{large}} (~770M parameters) underperforms the smaller PPTODbase_{\text{base}} (~220M parameters) and PPTODsmall_{\text{small}} (~60M parameters) models:

    • MultiWOZ 2.0 Combined Score: PPTODlarge_{\text{large}} achieves 97.56 (Inform 82.60, Success 74.10, BLEU 19.21), compared to 102.92 for PPTODbase_{\text{base}} and 101.44 for PPTODsmall_{\text{small}}.
    • MultiWOZ 2.1 Combined Score: PPTODlarge_{\text{large}} achieves 98.28 (Inform 86.43, Success 74.35, BLEU 17.89), compared to 102.26 for PPTODbase_{\text{base}} and 101.52 for PPTODsmall_{\text{small}}.

    This degradation is specific to response generation (NLG) in MultiWOZ, where the model is required to output delexicalized slot placeholders (such as [value_name] or [value_phone]). Because these delexicalized placeholder tokens were not encountered during the multi-task pre-training stage on natural text corpora, the larger parameter capacity of PPTODlarge_{\text{large}} exhibits poorer sample adaptation to synthetic token representations. Conversely, on dialogue state tracking and user intent classification tasks that do not rely on delexicalized tokens, PPTODlarge_{\text{large}} achieves the highest accuracy among all evaluated model sizes.

Coverage note — Specific qualitative MultiWOZ dialogue walkthroughs and dataset license metadata were deliberately omitted as illustrative or standard details.

References

  1. 1.Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. 2021. Muppet: Massive multi-task representations with pre-finetuning. CoRR, abs/2101.11038.
  2. 2.Hassan Amin. 2019. Atis airline travel information system, version 1.
  3. 3.Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2020. PLATO: pre-trained dialogue generation model with discrete latent variable. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 85–96. Association for Computational Linguistics.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  5. 5.Paweł Budzianowski and Ivan Vulić. 2019. Hello, it’s GPT-2 - how can I help you? towards the use of pretrained language models for task-oriented dialogue systems. In Proceedings of the 3rd Workshop on Neural Generation and Translation, pages 15–22, Hong Kong. Association for Computational Linguistics.
  6. 6.Pawel Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic. 2018. Multiwoz - A largescale multi-domain wizard-of-oz dataset for taskoriented dialogue modelling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 5016–5026. Association for Computational Linguistics.
  7. 7.Bill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, and Andy Cedilnik. 2019. Taskmaster-1: Toward a realistic and diverse dialog dataset. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 4515–4524. Association for Computational Linguistics.
  8. 8.Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vulic. 2020. Efficient intent detection with dual sentence encoders. CoRR, abs/2003.04807.
  9. 9.Lu Chen, Boer Lv, Chi Wang, Su Zhu, Bowen Tan, and Kai Yu. 2020. Schema-guided multi-domain dialogue state tracking with graph attention neural networks. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 7521–7528. AAAI Press.
  10. 10.Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Maël Primet, and Joseph Dureau. 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. CoRR, abs/1805.10190.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  12. 12.Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017. Frames: a corpus for adding memory to goal-oriented dialogue systems. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 207–219, Saarbrücken, Germany. Association for Computational Linguistics.
  13. 13.Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Kumar Goyal, Peter Ku, and Dilek Hakkani-Tür. 2020. Multiwoz 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines. In Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020, pages 422–428. European Language Resources Association.
  14. 14.Mihail Eric, Lakshmi Krishnan, Francois Charette, and Christopher D. Manning. 2017. Key-value retrieval networks for task-oriented dialogue. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 37–49, Saarbrücken, Germany. Association for Computational Linguistics.
  15. 15.Yue Feng, Yang Wang, and Hang Li. 2021. A sequence-to-sequence approach to dialogue state tracking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 1714–1725. Association for Computational Linguistics.
  16. 16.J.L. Fleiss et al. 1971. Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5):378–382.
  17. 17.Shuyang Gao, Abhishek Sethi, Sanchit Agarwal, Tagyoung Chung, and Dilek Hakkani-Tür. 2019. Dialog state tracking: A neural reading comprehension approach. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue, SIGdial 2019, Stockholm, Sweden, September 11-13, 2019, pages 264–273. Association for Computational Linguistics.
  18. 18.Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, Jian Sun, and Yongbin Li. 2021. GALAXY: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection. CoRR, abs/2111.14592.
  19. 19.Michael Heck, Carel van Niekerk, Nurul Lubis, Christian Geishauser, Hsien-Chin Lin, Marco Moresi, and Milica Gasic. 2020. Trippy: A triple copy strategy for value independent neural dialog state tracking. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, SIGdial 2020, 1st virtual meeting, July 1-3, 2020, pages 35–44. Association for Computational Linguistics.
  20. 20.Matthew Henderson, Iñigo Casanueva, Nikola Mrksic, Pei-Hao Su, Tsung-Hsien Wen, and Ivan Vulic. 2020. Convert: Efficient and accurate conversational representations from transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 2161–2174. Association for Computational Linguistics.
  21. 21.Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020. A simple language model for task-oriented dialogue. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  22. 22.Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Unifying question answering and text classification via span extraction. CoRR, abs/1904.09286.
  23. 23.Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Sungjin Lee, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, and Raghav Gupta. 2019. The eighth dialog system technology challenge. CoRR, abs/1911.06394.
  24. 24.Sungdong Kim, Sohee Yang, Gyuwan Kim, and Sang-Woo Lee. 2020. Efficient dialogue state tracking by selectively overwriting memory. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 567–582. Association for Computational Linguistics.
  25. 25.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  26. 26.Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1311–1316, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Hwaran Lee, Jinsik Lee, and Tae-Yoon Kim. 2019a. SUMBT: slot-utterance matching for universal and scalable belief tracking. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 5478–5483. Association for Computational Linguistics.
  28. 28.Sungjin Lee, Hannes Schulz, Adam Atkinson, Jianfeng Gao, Kaheer Suleman, Layla El Asri, Mahmoud Adada, Minlie Huang, Shikhar Sharma, Wendy Tay, and Xiujun Li. 2019b. Multi-domain task-completion dialog challenge. In Dialog System Technology Challenges 8.
  29. 29.Wenqiang Lei, Xisen Jin, Min-Yen Kan, Zhaochun Ren, Xiangnan He, and Dawei Yin. 2018. Sequicity: Simplifying task-oriented dialogue systems with single sequence-to-sequence architectures. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1437–1447, Melbourne, Australia. Association for Computational Linguistics.
  30. 30.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7871–7880. Association for Computational Linguistics.
  31. 31.Xiujun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017. Investigation of language understanding impact for reinforcement learning based dialogue systems. CoRR, abs/1703.07055.
  32. 32.Xiujun Li, Sarah Panda, JJ (Jingjing) Liu, and Jianfeng Gao. 2018. Microsoft dialogue challenge: Building end-to-end task-completion dialogue systems. In SLT 2018.
  33. 33.Weixin Liang, Youzhi Tian, Chengcai Chen, and Zhou Yu. 2020. MOSS: end-to-end dialog system framework with modular supervision. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8327–8335. AAAI Press.
  34. 34.Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, and Pascale Fung. 2020. Mintl: Minimalist transfer learning for task-oriented dialogue systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 3391–3405. Association for Computational Linguistics.
  35. 35.Bing Liu and Ian R. Lane. 2018. End-to-end learning of task-oriented dialogs. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 2-4, 2018, Student Research Workshop, pages 67–73. Association for Computational Linguistics.
  36. 36.Qi Liu, Lei Yu, Laura Rimell, and Phil Blunsom. 2021. Pretraining the noisy channel model for task-oriented dialogue. Trans. Assoc. Comput. Linguistics, 9:657–674.
  37. 37.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  38. 38.Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018. The natural language decathlon: Multitask learning as question answering. CoRR, abs/1806.08730.
  39. 39.Shikib Mehri, Tejas Srinivasan, and Maxine Eskénazi. 2019. Structured fusion networks for dialog. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue, SIGdial 2019, Stockholm, Sweden, September 11-13, 2019, pages 165–177. Association for Computational Linguistics.
  40. 40.Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017. Neural belief tracker: Data-driven dialogue state tracking. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1777–1788, Vancouver, Canada. Association for Computational Linguistics.
  41. 41.Elnaz Nouri and Ehsan Hosseini-Asl. 2018. Toward scalable neural dialogue state tracking model. CoRR, abs/1812.00899.
  42. 42.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  43. 43.Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2021. Soloist: Building task bots at scale with transfer learning and machine teaching. In Transactions of the Association for Computational Linguistics.
  44. 44.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 2227–2237. Association for Computational Linguistics.
  45. 45.Jason Phang, Thibault Févry, and Samuel R. Bowman. 2018. Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks. CoRR, abs/1811.01088.
  46. 46.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  47. 47.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  48. 48.Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020. Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8689–8696. AAAI Press.
  49. 49.Liliang Ren, Jianmo Ni, and Julian J. McAuley. 2019. Scalable and accurate dialogue state tracking via hierarchical sequence generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 1876–1885. Association for Computational Linguistics.
  50. 50.Bishal Santra, Potnuru Anusha, and Pawan Goyal. 2021. Hierarchical transformer for task oriented dialog systems. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 5649–5658. Association for Computational Linguistics.
  51. 51.Yong Shan, Zekang Li, Jinchao Zhang, Fandong Meng, Yang Feng, Cheng Niu, and Jie Zhou. 2020. A contextual hierarchical attention network with adaptive objective for dialogue state tracking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6322–6333. Association for Computational Linguistics.
  52. 52.Lei Shu, Piero Molino, Mahdi Namazifar, Hu Xu, Bing Liu, Huaixiu Zheng, and Gökhan Tür. 2019. Flexibly-structured model for task-oriented dialogues. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue, SIGdial 2019, Stockholm, Sweden, September 11-13, 2019, pages 178–187. Association for Computational Linguistics.
  53. 53.Ronnie W. Smith and D. Richard Hipp. 1995. Spoken Natural Language Dialog Systems: A Practical Approach. Oxford University Press, Inc., USA.
  54. 54.Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022. A contrastive framework for neural text generation. CoRR, abs/2202.06417.
  55. 55.Yixuan Su, Fangyu Liu, Zaiqiao Meng, Tian Lan, Lei Shu, Ehsan Shareghi, and Nigel Collier. 2021a. Tacl: Improving BERT pre-training with token-aware contrastive learning. CoRR, abs/2111.04198.
  56. 56.Yixuan Su, Zaiqiao Meng, Simon Baker, and Nigel Collier. 2021b. Few-shot table-to-text generation with prototype memory. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 910–917. Association for Computational Linguistics.
  57. 57.Yixuan Su, David Vandyke, Sihui Wang, Yimai Fang, and Nigel Collier. 2021c. Plan-then-generate: Controlled data-to-text generation via planning. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 895–909. Association for Computational Linguistics.
  58. 58.Yixuan Su, Yan Wang, Deng Cai, Simon Baker, Anna Korhonen, and Nigel Collier. 2021d. PROTOTYPE-TO-STYLE: dialogue generation with style-aware editing on retrieval memory. IEEE ACM Trans. Audio Speech Lang. Process., 29:2152–2161.
  59. 59.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. CoRR, abs/1804.07461.
  60. 60.Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017. A network-based end-to-end trainable task-oriented dialogue system. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 438–449, Valencia, Spain. Association for Computational Linguistics.
  61. 61.Jason D. Williams and Steve J. Young. 2007. Partially observable markov decision processes for spoken dialog systems. Comput. Speech Lang., 21(2):393–422.
  62. 62.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019a. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771.
  63. 63.Thomas Wolf, Victor Sanh, Julien Chaumond, and Clement Delangue. 2019b. Transfertransfo: A transfer learning approach for neural network based conversational agents. CoRR, abs/1901.08149.
  64. 64.Chien-Sheng Wu, Steven C.H. Hoi, Richard Socher, and Caiming Xiong. 2020. TOD-BERT: Pre-trained natural language understanding for task-oriented dialogue. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 917–929, Online. Association for Computational Linguistics.
  65. 65.Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, and Pascale Fung. 2019. Transferable multi-domain state generator for task-oriented dialogue systems. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 808–819. Association for Computational Linguistics.
  66. 66.Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernández Ábrego, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2020. Multilingual universal sentence encoder for semantic retrieval. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, ACL 2020, Online, July 5-10, 2020, pages 87–94. Association for Computational Linguistics.
  67. 67.Yunyi Yang, Yunhao Li, and Xiaojun Quan. 2021. UBAR: towards fully end-to-end task-oriented dialog system with GPT-2. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 14230–14238. AAAI Press.
  68. 68.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 5754–5764.
  69. 69.Steve J. Young, Milica Gasic, Blaise Thomson, and Jason D. Williams. 2013. Pomdp-based statistical spoken dialog systems: A review. Proc. IEEE, 101(5):1160–1179.
  70. 70.Jianguo Zhang, Kazuma Hashimoto, Chien-Sheng Wu, Yao Wan, Philip S. Yu, Richard Socher, and Caiming Xiong. 2019. Find or classify? dual strategy for slot-value predictions on multi-domain dialog state tracking. CoRR, abs/1910.03544.
  71. 71.Yichi Zhang, Zhijian Ou, Min Hu, and Junlan Feng. 2020a. A probabilistic end-to-end task-oriented dialog model with latent belief states towards semi-supervised learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 9207–9219. Association for Computational Linguistics.
  72. 72.Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020b. Task-oriented dialog systems that consider multiple appropriate responses under the same context. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 9604–9611. AAAI Press.
  73. 73.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020c. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, ACL 2020, Online, July 5-10, 2020, pages 270–278. Association for Computational Linguistics.
  74. 74.Victor Zhong, Caiming Xiong, and Richard Socher. 2018. Global-locally self-attentive encoder for dialogue state tracking. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 1458–1467. Association for Computational Linguistics.
  75. 75.Jingyao Zhou, Haipang Wu, Zehao Lin, Guodun Li, and Yin Zhang. 2021. Dialogue state tracking with multi-level fusion of predicted dialogue states and conversations. CoRR, abs/2107.05168.
  76. 76.Li Zhou and Kevin Small. 2019. Multi-domain dialogue state tracking as dynamic knowledge graph enhanced question answering. CoRR, abs/1911.06192.

Citation

MLA
Su, Y., et al. “Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 4661–76, https://doi.org/10.18653/v1/2022.acl-long.319.
APA
Su, Y., Shu, L., Mansimov, E., Gupta, A., Cai, D., Lai, Y.-A., & Zhang, Y. (2022). Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4661–4676. https://doi.org/10.18653/v1/2022.acl-long.319
Chicago
Su, Y., L. Shu, E. Mansimov, et al. 2022. “Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4661–76. https://doi.org/10.18653/v1/2022.acl-long.319.
Harvard
Su, Y. et al. (2022) “Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4661–4676. Available at: https://doi.org/10.18653/v1/2022.acl-long.319.
Vancouver
1. Su Y, Shu L, Mansimov E, Gupta A, Cai D, Lai Y-A, Zhang Y (2022) Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4661–4676

BibTeX

@inproceedings{su-etal-2022-multi,
    title = "Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System",
    author = "Su, Yixuan  and
      Shu, Lei  and
      Mansimov, Elman  and
      Gupta, Arshit  and
      Cai, Deng  and
      Lai, Yi-An  and
      Zhang, Yi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.319/",
    doi = "10.18653/v1/2022.acl-long.319",
    pages = "4661--4676"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/