InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning

Prakhar GuptaCathy JiaoYi-Ting YehShikib MehriMaxine EskénaziJeffrey P. Bigham

article2022EMNLP59 citations

Introduces an instruction-tuning framework unifying 48 dialogue tasks across 59 datasets, establishing benchmarks that boost zero- and few-shot cross-task generalization while incorporating novel meta-tasks to ensure instruction adherence.

Listen

Building conversational artificial intelligence systems currently requires extensive, specialized datasets and costly fine-tuning for each distinct task. While instruction tuning—training models to follow natural language task descriptions—has shown strong zero-shot generalization across general language processing, its application to dialogue domains has remained unstandardized and fragmented. Dialogue systems must manage diverse capabilities ranging from intent detection and state tracking to open-domain conversation and safety moderation, making unified adaptation difficult.

The article introduces and evaluates InstructDial, an open-source instruction-tuning framework designed to standardize dialogue data and induce strong zero-shot and few-shot performance across unseen conversational tasks. InstructDial unifies 59 publicly available dialogue datasets into a repository of 48 distinct tasks using a consistent text-to-sequence format. To ensure models genuinely follow user instructions rather than relying on dataset artifacts, the framework introduces specialized meta-tasks that train models to map instructions directly to input-output pairs.

The evaluation evaluated two instruction-tuned architectures: DIAL-BART0 (406 million parameters) and DIAL-T0 (3 billion parameters), benchmarked against existing state-of-the-art baselines and GPT-3. The experiments yielded several major findings. First, instruction tuning on InstructDial boosted performance on unseen dialogue tasks, with DIAL-BART0 outperforming baseline models by approximately threefold on evaluation selection, relation classification, and initial phrase generation. Second, DIAL-T0 achieved an average Spearman correlation of 0.465 with human relevance judgments across 13 benchmark sets, outperforming specialized automatic evaluation metrics without task-specific training. Third, incorporating only 100 task instances in a few-shot setup yielded 12% to 50% relative performance gains across difficult tasks, and DIAL-BART0 achieved a 36.9-point F1 improvement over prior models in zero-shot slot filling on Restaurant8k. Finally, the analysis showed that smaller, efficiently tuned models can match or outperform larger models, as DIAL-BART0 outperformed the much larger PPTOD on few-shot intent prediction (84.30% accuracy) and remained competitive in dialogue state tracking.

These findings indicate that dialogue systems can achieve broad generalization without deploying massive, prohibitively expensive language models. Standardizing dialogue tasks through natural language instructions enables development teams to rapidly prototype, diagnose, and deploy new conversational functionalities at lower compute and operational costs. The results demonstrate that pre-training on general language instructions directly complements dialogue-specific instruction tuning, significantly lowering the barrier to multi-domain deployment.

Organizations developing conversational AI should adopt standardized instruction schemas and leverage targeted few-shot tuning (around 100 examples) when deploying systems to novel domains. Technical teams should also incorporate instruction-matching meta-tasks during training to mitigate task confusion and enforce instruction adherence. Future research must address negative task interference—where training on certain dialogue tasks degrades performance on others—and explore methods to make systems less sensitive to instruction phrasing.

Confidence in these findings is high regarding core task-oriented capabilities and automatic dialogue evaluation. However, stakeholders should exercise caution regarding zero-shot reliability on complex edge tasks like utterance infilling, sensitivity to instruction wording variations, and potential task interference when scaling task repositories.

No sufficiently relevant recommendations were found.

Cover for InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning

Abstract

Instruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zero-shot performance on unseen tasks. Dialogue is an especially interesting area in which to explore instruction tuning because dialogue systems perform multiple tasks related to language (e.g., natural language understanding and generation, domain-specific interaction), yet instruction tuning has not been systematically explored for dialogue-related tasks. We introduce InstructDial, an instruction tuning framework for dialogue, which consists of a repository of 48 diverse dialogue tasks in a unified text-to-text format created from 59 openly available dialogue datasets. We explore cross-task generalization ability on models tuned on InstructDial across diverse dialogue tasks. Our analysis reveals that InstructDial enables good zero-shot performance on unseen datasets and tasks such as dialogue evaluation and intent detection, and even better performance in a few-shot setting. To ensure that models adhere to instructions, we introduce novel meta-tasks. We establish benchmark zero-shot and few-shot performance of models trained using the proposed framework on multiple dialogue tasks¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Instruction Tuning Background
  • 3.2 Task Collection
  • 3.3 Task Schema and Formatting
  • 3.4 Meta Tasks
  • 3.5 None-of-the-above Options
  • 4 Experimental Setup
  • 4.1 Model Details
  • 4.2 Training Details
  • 5 Experiments and Results
  • 5.1 Zero-shot Unseen Tasks Evaluation
  • 5.1.1 Unseen Tasks for Zero-shot Setting
  • 5.1.2 Setup and Baselines
  • 5.1.3 Results and Discussion
  • 5.1.4 Analysis
  • 5.2 Zero-shot Automatic Response Evaluation
  • 5.3 Zero-shot and Few-shot Dialogue Tasks
  • 5.3.1 Intent Prediction
  • 5.3.2 Slot Filling
  • 5.3.3 Dialogue State Tracking
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics and Broader Impact
  • References
  • Appendix
  • A Additional implementation details
  • B Sample conversation and Instructions
  • C Datasets used in tasks
  • D Configuration of experiments

Knowls

  1. Knowl 1 — InstructDial assembles a broad instruction-tuning framework for dialogue

    model/method

    INSTRUCTDIAL is a framework for instruction tuning across 48 dialogue tasks built from 59 openly available dialogue datasets. It expresses the tasks in a shared text-to-text format and covers both open-domain and task-oriented dialogue. Its task taxonomy includes classification and generation, as well as evaluation, response editing, pretraining-style tasks, safety tasks, and specialized miscellaneous tasks. Examples include intent detection, response generation, response relevance and selection, dialogue summarization, toxicity classification, and slot filling. The framework is intended to support systematic study of zero-shot and few-shot transfer between dialogue tasks.

  2. Knowl 2 — Dialogue tasks use a structured natural-language input and output format

    model/method

    Each INSTRUCTDIAL instance combines a natural-language task definition, dialogue or other instance inputs, optional task constraints, a task-specific prompt, and the expected output. Inputs use shared special tokens: [CONTEXT] begins dialogue content, [ENDOFTURN] separates turns, [ENDOFDIALOGUE] marks the dialogue end, and [QUESTION] begins the prompt. Task-specific inputs can be added with field labels, such as [EMOTION] followed by an emotion label. Classification inputs include the candidate classes, represented either as a list of class names or as an indexed list when the choices are lengthy. The authors manually wrote 3–10 definitions and prompts per task; a definition and prompt were randomly assigned to each training instance and selected randomly at test time. The format does not use in-context examples, in part because long dialogue contexts leave limited room within model input limits.

  3. Knowl 3 — Two meta-tasks train models to connect instructions with input-output behavior

    model/method

    INSTRUCTDIAL adds two meta-tasks to make models attend to instructions rather than infer a task only from its dataset or domain. In instruction selection, the model receives an input-output pair and chooses which candidate instruction corresponds to that pair. In the instruction binary task, it receives an instruction, input, and output, then predicts “yes” or “no” according to whether the instruction licenses that output. These tasks are designed to strengthen the association between task wording and the behavior it specifies.

  4. Knowl 4 — Classification training includes a None-of-the-Above option

    model/method

    For 10% of classification training instances, INSTRUCTDIAL adds a “None of the above” (NOTA) candidate in one of two ways. For a NOTA-positive instance, the gold class is removed from the choices and NOTA becomes the target. For a NOTA-distractor instance, NOTA is added to the choices while the original gold class remains the target. This trains models to handle cases where none of the provided classes is appropriate, a possibility in unseen classification tasks.

  5. Knowl 5 — INSTRUCTDIAL models fine-tune instruction-tuned encoder-decoder bases

    experimental setup

    The authors fine-tune two encoder-decoder models using maximum-likelihood training: BART0, with 406 million parameters, and T0-3B, with 3 billion parameters. Both bases had previously been instruction-tuned on general, non-dialogue NLP tasks. Their INSTRUCTDIAL-tuned versions are named DIAL-BART0 and DIAL-T0. Training data are generated from the datasets in each task, with a maximum of 5,000 sampled instances per task; each instance receives a randomly assigned task definition and prompt. Inputs are truncated to 1,024 tokens and target outputs to 256 tokens. Models train for 3 epochs with Adam, a learning rate of 5e-5, and linear learning-rate decay. Classification uses greedy decoding; generation uses top-p sampling with p = 0.7, temperature 0.7, and repetition penalty 1.2.

  6. Knowl 6 — Instruction tuning improves performance on several unseen dialogue tasks

    empirical result

    The zero-shot evaluation tests six tasks withheld from training: evaluation selection, answer selection, relation classification, DialFact classification, Begins With generation, and knowledge-grounded generation. DIAL-BART0 and DIAL-T0 generally outperform their corresponding bases, BART0 and T0-3B, across the reported metrics. On the five accuracy-scored tasks (in the order listed below, excluding knowledge-grounded generation), BART0 scores 22.2, 58.5, 6.3, 33.7, and 4.2; DIAL-BART0 scores 66.7, 59.5, 17.8, 35.6, and 56.3. T0-3B scores 45.9, 60.2, 1.3, 33.1, and 14.1; DIAL-T0 scores 74.4, 65.2, 6.4, 34.5, and 55.0. For knowledge-grounded generation, the reported F1 scores are 17.4 for BART0 and 27.8 for DIAL-BART0, and 14.2 for T0-3B and 22.2 for DIAL-T0. The gains are not uniform across tasks or model sizes: for example, DIAL-BART0 exceeds DIAL-T0 on relation classification and knowledge-grounded-generation F1, while DIAL-T0 scores higher on evaluation and answer selection.

    Adding 100 training examples per test task to DIAL-BART0 (DB-Few) raises accuracy to 77.1, 69.1, 28.0, 43.0, and 72.2 on the five accuracy-scored tasks; adding 5,000 examples per task (DB-Full) raises it to 90.7, 83.3, 62.7, 77.4, and 83.7. Ablations also show that instructions and meta-tasks matter: removing instructions lowers accuracy on evaluation selection from 66.7 to 23.0 and answer selection from 59.5 to 43.2; removing meta-tasks lowers those scores to 44.5 and 52.0. The NOTA ablation produces a smaller decline on these measures. In a separate analysis, average performance over unseen tasks rises sharply up to about 20–25 seen training tasks and continues to increase more gradually as tasks are added.

  7. Knowl 7 — DIAL-T0 predicts dialogue response relevance with strong human-rating correlations

    empirical result

    For zero-shot automatic response evaluation, DIAL-T0 is trained on dialogue tasks that exclude evaluation tasks. Given a dialogue context and candidate response, it predicts “yes” or “no” for relevance; the relevance score is the model probability of “yes” normalized by the summed probabilities of “yes” and “no.” The evaluation uses 65,938 context-response pairs from the DSTC-10 automatic evaluation challenge and measures Spearman correlation with available human ratings. DIAL-T0’s correlations, in the order DSTC6, DSTC7, HUMOD, TopicalChat-USR, PersonaChat-USR, PersonaChat-Zhao, DailyDialog-Zhao, ConvAI2-GRADE, DailyDialog-Gupta, DailyDialog-GRADE, FED-Turn, Empathetic-GRADE, and FED-Dial, are 0.553, 0.451, 0.582, 0.446, 0.651, 0.601, 0.498, 0.376, 0.634, 0.286, 0.263, 0.475, and 0.228, respectively; the macro average is 0.465. The strongest competing macro average in the reported comparison is DEB at 0.367. DIAL-T0 ranks first or second on most evaluation sets.

  8. Knowl 8 — Instruction-tuned models transfer to intent prediction and slot filling

    empirical result

    On BANKING77 intent prediction, the zero-shot DIAL-BART0 model reaches 58.02% accuracy, compared with 14.72% for zero-shot BART0. With 10 training examples per intent, DIAL-BART0 reaches 84.30% accuracy. The reported few-shot comparison includes ConvERT at 83.32%, ConvERT plus USE at 85.19%, Example-Driven at 85.95%, PPTOD-base at 82.81%, and PPTOD-large at 84.12%; thus DIAL-BART0 is competitive but does not lead this comparison.

    For zero-shot slot filling on Restaurant8k, DIAL-BART0 obtains 56.4 F1, compared with 19.5 for GENSF, 10.7 for COACH+TR, and 5.2 for ConvERT. In a separate few-shot evaluation on DSTC8 using 25% of its training data, DIAL-BART0’s slot-filling F1 scores are 97.8 for Buses, 94.3 for Events, 96.5 for Homes, and 94.2 for Rental Cars; the corresponding GENSF scores are 90.5, 91.2, 93.7, and 86.7.

  9. Knowl 9 — DIAL-BART0 achieves competitive few-shot dialogue state tracking

    empirical result

    For dialogue state tracking, DIAL-BART0 is first trained on seven dialogue-state datasets—KVRET, WOZ, CamRest676, MSR-E2E, Frames, TaskMaster, and Schema-Guided—along with other dialogue tasks. It is then trained on 1% or 5% of the MultiWOZ 2.0 training data for 40 epochs with a learning rate of 5e-5. On the MultiWOZ test set, DIAL-BART0 obtains joint goal accuracy of 29.2 with 1% training data and 38.1 with 5%. PPTOD-base, reported under the same data proportions, scores 29.7 and 40.2, respectively; DIAL-BART0 is therefore close to, but below, that comparator.

  10. Knowl 10 — The study identifies limits in instruction robustness and task transfer

    limitation

    INSTRUCTDIAL’s task definitions and prompts are manually written, limited in number, and English-only. DIAL-BART0’s zero-shot results vary with instruction wording: accuracy ranges from 65.6 to 67.8 on evaluation selection, 52.5 to 75.0 on answer selection, 17.1 to 18.4 on relation classification, and 34.7 to 37.1 on DialFact classification; its Begins With scores range from 49.8 to 62.3, and knowledge-grounded-generation F1 ranges from 26.6 to 28.6. Zero-shot performance remains weak on some tasks, and the authors report that adding training tasks can interfere with performance on other tasks or contribute to task forgetting. The study does not investigate example-based in-context few-shot learning or composition of multiple tasks through instructions, and the framework covers text-only rather than multimodal tasks.

Coverage note — The exhaustive inventory of all 59 datasets and per-task example formats is omitted because it is catalog detail rather than a distinct methodological or empirical finding; the framework’s scale and task coverage are summarized.

References

  1. 1.Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark Riedl. 2021. Just say no: Analyzing the stance of neural dialogue generation in offensive contexts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4846–4862, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  3. 3.Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhen Guo, Zhibin Liu, and Xinchao Xu. 2021. PLATO-2: Towards building an open-domain chatbot via curriculum learning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2513–2525, Online. Association for Computational Linguistics.
  4. 4.Jonathan Bragg, Arman Cohan, Kyle Lo, and Iz Beltagy. 2021. Flex: Unifying evaluation for few-shot nlp. In Advances in Neural Information Processing Systems (NeurIPS 2021).
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Paweł Budzianowski and Ivan Vulic. 2019. ´ Hello, it’s GPT-2 - how can I help you? towards the use of pre-trained language models for task-oriented dialogue systems. In Proceedings of the 3rd Workshop on Neural Generation and Translation, pages 15–22, Hong Kong. Association for Computational Linguistics.
  7. 7.Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašic. 2018. ´ MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 5016–5026, Brussels, Belgium. Association for Computational Linguistics.
  8. 8.Bill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, and Andy Cedilnik. 2019. Taskmaster-1: Toward a realistic and diverse dialog dataset. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4516–4525, Hong Kong, China. Association for Computational Linguistics.
  9. 9.Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vulic. 2020. ´ Efficient intent detection with dual sentence encoders. In Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, pages 38–45, Online. Association for Computational Linguistics.
  10. 10.Kushal Chawla, Jaysa Ramirez, Rene Clever, Gale Lucas, Jonathan May, and Jonathan Gratch. 2021. CaSiNo: A corpus of campsite negotiation dialogues for automatic negotiation systems. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3167–3185, Online. Association for Computational Linguistics.
  11. 11.Yulong Chen, Yang Liu, Liang Chen, and Yue Zhang. 2021a. DialogSum: A real-life scenario dialogue summarization dataset. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 5062–5074, Online. Association for Computational Linguistics.
  12. 12.Zhang Chen, João Sadoc, Luis Fernando D’Haro, Rafael Banchs, and Alexander Rudnicky. 2021b. Automatic evaluation and moderation of open-domain dialogue systems. arXiv preprint arXiv:2111.02110.
  13. 13.Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wentau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question answering in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2174–2184, Brussels, Belgium. Association for Computational Linguistics.
  14. 14.Sam Coope, Tyler Farghly, Daniela Gerz, Ivan Vulic, and Matthew Henderson. 2020a. Span-convert: Few-shot span extraction for dialog with pretrained conversational representations. arXiv preprint arXiv:2005.08866.
  15. 15.Samuel Coope, Tyler Farghly, Daniela Gerz, Ivan Vulic, and Matthew Henderson. 2020b. Span-ConveRT: Few-shot span extraction for dialog with pretrained conversational representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 107–121, Online. Association for Computational Linguistics.
  16. 16.Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. arXiv preprint arXiv:1805.10190, pages 12–16.
  17. 17.Leyang Cui, Yu Wu, Shujie Liu, Yue Zhang, and Ming Zhou. 2020. MuTual: A dataset for multi-turn dialogue reasoning. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1406–1416, Online. Association for Computational Linguistics.
  18. 18.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040–4054, Online. Association for Computational Linguistics.
  19. 19.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019a. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4537–4546, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019b. The second conversational intelligence challenge (convai2). arXiv preprint arXiv:1902.00098.
  22. 22.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019c. Wizard of Wikipedia: Knowledge-powered conversational agents. In Proceedings of the International Conference on Learning Representations (ICLR).
  23. 23.Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017. Frames: a corpus for adding memory to goal-oriented dialogue systems. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 207–219, Saarbrücken, Germany. Association for Computational Linguistics.
  24. 24.Mihail Eric, Lakshmi Krishnan, Francois Charette, and Christopher D. Manning. 2017. Key-value retrieval networks for task-oriented dialogue. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 37–49, Saarbrücken, Germany. Association for Computational Linguistics.
  25. 25.Song Feng, Hui Wan, Chulaka Gunasekara, Siva Patel, Sachindra Joshi, and Luis Lastras. 2020a. doc2dial: A goal-oriented document-grounded dialogue dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8118–8128, Online. Association for Computational Linguistics.
  26. 26.Yulan Feng, Shikib Mehri, Maxine Eskenazi, and Tiancheng Zhao. 2020b. “none of the above”: Measure uncertainty in dialog response retrieval. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2013–2020, Online. Association for Computational Linguistics.
  27. 27.Michel Galley, Chris Brockett, Xiang Gao, Jianfeng Gao, and Bill Dolan. 2019. Grounded response generation task at dstc7.
  28. 28.Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and Bill Dolan. 2020. Dialogue response ranking training with large-scale human feedback data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 386–395, Online. Association for Computational Linguistics.
  29. 29.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. Samsum corpus: A human-annotated dialogue dataset for abstractive summarization. arXiv preprint arXiv:1911.12237.
  30. 30.Chih-Wen Goo and Yun-Nung Chen. 2018. Abstractive dialogue summarization with sentence-gated modeling optimized by dialogue acts. In Proceedings of 7th IEEE Workshop on Spoken Language Technology.
  31. 31.Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019. Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations. In Proc. Interspeech 2019, pages 1891–1895.
  32. 32.Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey Bigham. 2019. Investigating evaluation of open-domain dialogue systems with human generated multiple references. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue, pages 379–391, Stockholm, Sweden. Association for Computational Linguistics.
  33. 33.Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021. Dialfact: A benchmark for fact-checking in dialogue. arXiv preprint arXiv:2110.08222.
  34. 34.Donghoon Ham, Jeong-Gwan Lee, Youngsoo Jang, and Kee-Eung Kim. 2020. End-to-end neural pipeline for goal-oriented dialogue systems using GPT-2. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 583–592, Online. Association for Computational Linguistics.
  35. 35.Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990. The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27,1990.
  36. 36.Matthew Henderson and Ivan Vulic. 2020. Convex: Data-efficient and few-shot slot labeling. arXiv preprint arXiv:2010.11791.
  37. 37.Chiori Hori and Takaaki Hori. 2017. End-to-end conversation modeling track in dstc6. arXiv:1706.07440.
  38. 38.Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020. A simple language model for task-oriented dialogue. Advances in Neural Information Processing Systems, 33:20179–20191.
  39. 39.Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao Huang, and Lun-Wei Ku. 2018. EmotionLines: An emotion corpus of multi-party conversations. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  40. 40.Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, and Xiaodan Liang. 2020. GRADE: Automatic graph-enhanced coherence metric for evaluating open-domain dialogue systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9230–9240, Online. Association for Computational Linguistics.
  41. 41.Diederick P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR).
  42. 42.Stefan Larson and Kevin Leach. 2022. Redwood: Using collision detection to grow a large-scale intent classification dataset. arXiv preprint arXiv:2204.05483.
  43. 43.Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1311–1316, Hong Kong, China. Association for Computational Linguistics.
  44. 44.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  45. 45.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  46. 46.Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017. Deal or no deal? end-to-end learning of negotiation dialogues. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2443–2453, Copenhagen, Denmark. Association for Computational Linguistics.
  47. 47.Xiujun Li, Sarah Panda, Jingjing Liu, and Jianfeng Gao. 2018. Microsoft dialogue challenge: Building end-to-end task-completion dialogue systems. arXiv preprint arXiv:1807.11125.
  48. 48.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 986–995, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  49. 49.Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, and Jie Zhou. 2021. Conversations are not flat: Modeling the dynamic information flow across dialogue utterances. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 128–138, Online. Association for Computational Linguistics.
  50. 50.Bill Yuchen Lin, Kangmin Tan, Chris Miller, Beiwen Tian, and Xiang Ren. 2022. Unsupervised cross-task generalization via retrieval augmentation. ArXiv, abs/2204.07937.
  51. 51.Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, and Pascale Fung. 2020. MinTL: Minimalist transfer learning for task-oriented dialogue systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3391–3405, Online. Association for Computational Linguistics.
  52. 52.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021a. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  53. 53.Qi Liu, Lei Yu, Laura Rimell, and Phil Blunsom. 2022. Pretraining the noisy channel model for task-oriented dialogue. Transactions of the Association for Computational Linguistics, 9(0):657–674.
  54. 54.Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2021b. Benchmarking natural language understanding services for building conversational agents. In Increasing Naturalness and Flexibility in Spoken Dialogue Interaction, pages 165–183. Springer.
  55. 55.Zihan Liu, Genta Indra Winata, Peng Xu, and Pascale Fung. 2020. Coach: A coarse-to-fine approach for cross-domain slot filling. arXiv preprint arXiv:2004.11727.
  56. 56.Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few-shot bot: Prompt-based learning for dialogue systems. arXiv preprint arXiv:2110.08118.
  57. 57.Shikib Mehri and Mihail Eric. 2021. Example-driven intent prediction with observers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2979–2992, Online. Association for Computational Linguistics.
  58. 58.Shikib Mehri and Maxine Eskenazi. 2020a. Unsupervised evaluation of interactive dialog with DialoGPT. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 225–235, 1st virtual meeting. Association for Computational Linguistics.
  59. 59.Shikib Mehri and Maxine Eskenazi. 2020b. USR: An unsupervised and reference free evaluation metric for dialog generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 681–707, Online. Association for Computational Linguistics.
  60. 60.Shikib Mehri and Maxine Eskenazi. 2021. Gensf: Simultaneous adaptation of generative pre-trained models and slot filling. arXiv preprint arXiv:2106.07055.
  61. 61.Shikib Mehri, Evgeniia Razumovskaia, Tiancheng Zhao, and Maxine Eskenazi. 2019. Pretraining methods for dialog context representation learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3836–3845, Florence, Italy. Association for Computational Linguistics.
  62. 62.Erinc Merdivan, Deepika Singh, Sten Hanke, Johannes Kropf, Andreas Holzinger, and Matthieu Geist. 2020. Human annotated dialogues dataset for natural conversational agents. Applied Sciences, 10(3):762.
  63. 63.Fei Mi, Yitong Li, Yasheng Wang, Xin Jiang, and Qun Liu. 2021a. Cins: Comprehensive instruction for few-shot learning in task-oriented dialog systems. ArXiv, abs/2109.04645.
  64. 64.Fei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai, Minlie Huang, and Boi Faltings. 2021b. Self-training improves pre-training for few-shot learning in task-oriented dialog systems. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1887–1898, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  65. 65.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022a. MetaICL: Learning to learn in context. In NAACL-HLT.
  66. 66.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022b. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  67. 67.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021. Natural instructions: Benchmarking generalization to new tasks from natural language instructions. In Annual Meeting of the Association for Computational Linguistics (ACL).
  68. 68.Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. OpenDialKG: Explainable conversational reasoning with attention-based walks over knowledge graphs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 845–854, Florence, Italy. Association for Computational Linguistics.
  69. 69.Nikola Mrkšic, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017. Neural belief tracker: Data-driven dialogue state tracking. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1777–1788, Vancouver, Canada. Association for Computational Linguistics.
  70. 70.Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, and Jason Weston. 2021. I like fish, especially dolphins: Addressing contradictions in dialogue modeling. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1699–1713, Online. Association for Computational Linguistics.
  71. 71.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  72. 72.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  73. 73.Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2021. Soloist: Building task bots at scale with transfer learning and machine teaching. In Transactions of the Association for Computational Linguistics.
  74. 74.Vitou Phy, Yang Zhao, and Akiko Aizawa. 2020. Deconstruct to reconstruct a configurable evaluation metric for open-domain dialogue systems. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4164–4178, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  75. 75.Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, and Manaal Faruqui. 2021. TIME-DIAL: Temporal commonsense reasoning in dialog. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7066–7076, Online. Association for Computational Linguistics.
  76. 76.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  77. 77.Dinesh Raghu, Shantanu Agarwal, Sachindra Joshi, and Mausam. 2021. End-to-end learning of flowchart grounded task-oriented dialogs. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4348–4366, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  78. 78.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open-domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381, Florence, Italy. Association for Computational Linguistics.
  79. 79.Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020a. Schema-guided dialogue state tracking task at dstc8. arXiv preprint arXiv:2002.01359.
  80. 80.Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020b. Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8689–8696.
  81. 81.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266.
  82. 82.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online. Association for Computational Linguistics.
  83. 83.Ananya B. Sai, Akash Kumar Mohankumar, Siddhartha Arora, and Mitesh M. Khapra. 2020. Improving dialog evaluation with a multi-reference adversarial dataset and large scale pretraining. Transactions of the Association for Computational Linguistics, 8:810–827.
  84. 84.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations (ICLR).
  85. 85.Timo Schick and Hinrich Schütze. 2021. Few-shot text generation with natural language instructions. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 390–402, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  86. 86.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  87. 87.Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019. What makes a good conversation? how controllable attributes affect human judgments. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1702–1723, Minneapolis, Minnesota. Association for Computational Linguistics.
  88. 88.Karin Sevegnani, David M. Howcroft, Ioannis Konstas, and Verena Rieser. 2021. OTTers: One-turn topic transitions for open-domain dialogue. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2492–2504, Online. Association for Computational Linguistics.
  89. 89.Koustuv Sinha, Prasanna Parthasarathi, Jasmine Wang, Ryan Lowe, William L. Hamilton, and Joelle Pineau. 2020. Learning an unreferenced metric for online dialogue evaluation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2430–2441, Online. Association for Computational Linguistics.
  90. 90.Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi-An Lai, and Yi Zhang. 2022a. Multi-task pre-training for plug-and-play task-oriented dialogue system.
  91. 91.Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi-An Lai, and Yi Zhang. 2022b. Multi-task pre-training for plug-and-play task-oriented dialogue system. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4661–4676, Dublin, Ireland. Association for Computational Linguistics.
  92. 92.Megan Ung, Jing Xu, and Y-Lan Boureau. 2021. Safer-dialogues: Taking feedback gracefully after conversational safety failures.
  93. 93.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  94. 94.Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, and Mohit Iyyer. 2020. Exploring and predicting transferability across NLP tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7882–7926, Online. Association for Computational Linguistics.
  95. 95.Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2019. Persuasion for good: Towards a personalized persuasive dialogue system for social good. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5635–5649, Florence, Italy. Association for Computational Linguistics.
  96. 96.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022a. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705.
  97. 97.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, et al. 2022b. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv.
  98. 98.Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. 2020a. On negative interference in multilingual models: Findings and a meta-learning treatment. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4438–4450, Online. Association for Computational Linguistics.
  99. 99.Zirui Wang, Sanket Vaibhav Mehta, Barnabas Poczos, and Jaime Carbonell. 2020b. Efficient meta lifelong-learning with limited memory. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 535–548, Online. Association for Computational Linguistics.
  100. 100.Albert Webson and Ellie Pavlick. 2021. Do prompt-based models really understand the meaning of their prompts?
  101. 101.Jason Wei, Maarten Paul Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew Mingbo Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners.
  102. 102.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019. Dialogue natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3731–3741, Florence, Italy. Association for Computational Linguistics.
  103. 103.Orion Weller, Nicholas Lourie, Matt Gardner, and Matthew E. Peters. 2020. Learning from task descriptions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1361–1375, Online. Association for Computational Linguistics.
  104. 104.Tsung-Hsien Wen, David Vandyke, Nikola Mrkšic, Milica Gašic, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017. A network-based end-to-end trainable task-oriented dialogue system. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 438–449, Valencia, Spain. Association for Computational Linguistics.
  105. 105.Taesun Whang, Dongyub Lee, Dongsuk Oh, Chanhee Lee, Kijong Han, Dong-hun Lee, and Saebyeok Lee. 2021. Do response selection models really know what’s next? utterance manipulation strategies for multi-turn response selection. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14041–14049.
  106. 106.Chien-Sheng Wu, Steven C.H. Hoi, Richard Socher, and Caiming Xiong. 2020a. TOD-BERT: Pre-trained natural language understanding for task-oriented dialogue. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 917–929, Online. Association for Computational Linguistics.
  107. 107.Chien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung, and Caiming Xiong. 2021. Qaconv: Question answering on informative conversations. arXiv preprint arXiv:2105.06912.
  108. 108.Sen Wu, Hongyang R Zhang, and Christopher Ré. 2020b. Understanding and improving information transfer in multi-task learning. arXiv preprint arXiv:2005.00944.
  109. 109.Yujie Xing, Jinglun Cai, Nils Barlaug, Peng Liu, and Jon Atle Gulla. 2022. Balancing multi-domain corpora learning for open-domain response generation. arXiv preprint arXiv:2205.02570.
  110. 110.Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, and Zhilin Yang. 2022. Zero-prompt: Scaling prompt-based pretraining to 1,000 tasks improves zero-shot generalization. arXiv preprint arXiv:2201.06910.
  111. 111.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021a. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2950–2968, Online. Association for Computational Linguistics.
  112. 112.Ruijian Xu, Chongyang Tao, Daxin Jiang, Xueliang Zhao, Dongyan Zhao, and Rui Yan. 2021b. Learning an effective context-response matching model with self-supervised tasks for retrieval-based dialogues. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14158–14166.
  113. 113.Yunyi Yang, Yunhao Li, and Xiaojun Quan. 2021. Ubar: Towards fully end-to-end task-oriented dialog system with gpt-2. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14230–14238.
  114. 114.Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021. A comprehensive assessment of dialog evaluation metrics. In The First Workshop on Evaluations and Assessments of Neural Conversation Systems, pages 15–33, Online. Association for Computational Linguistics.
  115. 115.Dian Yu, Kai Sun, Claire Cardie, and Dong Yu. 2020. Dialogue-based relation extraction. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4927–4940, Online. Association for Computational Linguistics.
  116. 116.Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi. 2021. TuringAdvice: A generative and dynamic evaluation of language use. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4856–4880, Online. Association for Computational Linguistics.
  117. 117.Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li. 2021. DynaEval: Unifying turn and dialogue level evaluation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5676–5689, Online. Association for Computational Linguistics.
  118. 118.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213, Melbourne, Australia. Association for Computational Linguistics.
  119. 119.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278, Online. Association for Computational Linguistics.
  120. 120.Tianyu Zhao, Divesh Lala, and Tatsuya Kawahara. 2020a. Designing precise and robust dialogue response evaluators. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 26–33, Online. Association for Computational Linguistics.
  121. 121.Yufan Zhao, Can Xu, and Wei Wu. 2020b. Learning a simple and effective model for multi-turn response generation with auxiliary tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3472–3483, Online. Association for Computational Linguistics.
  122. 122.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021a. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5905–5921, Online. Association for Computational Linguistics.
  123. 123.Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. 2021b. Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2856–2878, Punta Cana, Dominican Republic. Association for Computational Linguistics.

Citation

MLA
Gupta, P., et al. “InstructDial: Improving Zero and Few-shot Generalization in Dialogue Through Instruction Tuning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 505–25, https://doi.org/10.18653/v1/2022.emnlp-main.33.
APA
Gupta, P., Jiao, C., Yeh, Y.-T., Mehri, S., Eskenazi, M., & Bigham, J. P. (2022). InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 505–525. https://doi.org/10.18653/v1/2022.emnlp-main.33
Chicago
Gupta, P., C. Jiao, Y.-T. Yeh, S. Mehri, M. Eskenazi, and J. P. Bigham. 2022. “InstructDial: Improving Zero and Few-shot Generalization in Dialogue Through Instruction Tuning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 505–25. https://doi.org/10.18653/v1/2022.emnlp-main.33.
Harvard
Gupta, P. et al. (2022) “InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 505–525. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.33.
Vancouver
1. Gupta P, Jiao C, Yeh Y-T, Mehri S, Eskenazi M, Bigham JP (2022) InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 505–525

BibTeX

@inproceedings{gupta-etal-2022-instructdial,
    title = "{I}nstruct{D}ial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning",
    author = "Gupta, Prakhar  and
      Jiao, Cathy  and
      Yeh, Yi-Ting  and
      Mehri, Shikib  and
      Eskenazi, Maxine  and
      Bigham, Jeffrey",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.33/",
    doi = "10.18653/v1/2022.emnlp-main.33",
    pages = "505--525"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/