DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation

Yizhe ZhangSiqi SunMichel GalleyYen-Chun ChenChris BrockettXiang GaoJianfeng GaoJingjing LiuBill Dolan

article2019ACL1,758 citations

Introduces DialoGPT, a conversational response generation model pre-trained on 147 million Reddit dialogue threads to produce contextually consistent, content-rich responses that approach human quality in open-domain settings.

Listen

Automated conversational systems have historically struggled with generating responses that are engaging, relevant, and contextually consistent. Traditional chatbot architectures often produce generic, repetitive, or disjointed text when handling open-ended human dialogue. To address these limitations, the article demonstrates the development and evaluation of DialoGPT, a large-scale generative conversational model adapted from transformer-based language architectures to generate natural, informative, and multi-turn dialogue responses.

The authors constructed a dataset comprising 147 million multi-turn dialogue instances (1.8 billion words) extracted from Reddit discussion threads spanning 2005 through 2017. Rigorous filtering was applied to remove toxic language, markup, URLs, repetitive phrasing, and bland responses. Using this corpus, three model variations—ranging from 117 million to 762 million parameters—were trained using an autoregressive transformer framework. To counter blandness during decoding, the authors incorporated a mutual information scoring approach that re-ranks candidate responses using a backward prediction model. The system was benchmarked against existing conversational baselines and human ground-truth responses using standard automatic language metrics and crowdsourced human evaluations.

The evaluation produced four central findings. First, DialoGPT significantly outperformed existing production baselines; in human evaluations, judges preferred the 345-million-parameter model over Microsoft's production system by a margin of 72% to 19% in relevance and 77% to 19% in informativeness. Second, model scale and pre-training directly enhanced performance, with larger models and those initialized from existing language models yielding superior conversational quality. Third, re-ranking candidate responses via mutual information maximization successfully boosted response diversity and informativeness, with human judges rating these responses higher in human-likeness (50% vs. 46%) and informativeness (50% vs. 46%) than actual human test responses. Finally, DialoGPT achieved higher automated overlap scores than human references, reflecting the model's ability to identify central, probable response trajectories across diverse conversation paths.

These results demonstrate that large-scale pre-training on conversational data drastically reduces the engineering overhead needed to build high-quality dialogue agents. Rather than training complex systems from scratch, organizations can fine-tune these pre-trained models on specialized, smaller datasets in a matter of hours. However, the reliance on massive web data introduces notable operational and reputational risks. The model inherits potential societal biases and can occasionally generate toxic outputs or inappropriately validate unethical user statements.

The authors have open-sourced the model weights and training pipeline to support further research. Organizations intending to deploy these conversational models should implement strict external output filtering, safety decoders, and toxicity controls. Future technical development should prioritize regularized reinforcement learning to prevent model degradation and focus on automated methods to detect and suppress biased or harmful responses. Stakeholders must exercise caution when deploying the raw models without protective guardrails, given the inherent unpredictability of unconstrained conversational generation.

  • Paper: LaMDA: Language Models for Dialog Applications, Romal Thoppilan et al. (2022). LaMDA significantly scales dedicated conversational pre-training beyond DialoGPT by integrating safety fine-tuning, factual grounding, and external tool use into neural dialogue models.
  • Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This work establishes multi-turn conversation evaluation frameworks like MT-Bench and Chatbot Arena, overcoming the limited single-turn evaluation metrics used in DialoGPT.
  • Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). Llama 2 extends open-access dialogue modeling through large-scale instruction tuning and reinforcement learning from human feedback to produce safe, highly aligned conversational agents.
  • Paper: LIMA: Less Is More for Alignment, Chunting Zhou et al. (2023). LIMA investigates how lightweight, curated conversational alignment data can elicit strong conversational abilities from pre-trained language models without massive supervised conversational pre-training.
  • Paper: Zephyr: Direct Distillation of LM Alignment, Lewis Tunstall et al. (2024). Zephyr presents direct distillation techniques to optimize conversational response generation and preference alignment in accessible, smaller open-source models.
  • Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). AutoGen generalizes conversational modeling beyond single chat agents into orchestrated multi-agent cooperative dialogue systems.
  • Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). This study analyzes why autoregressive language models produce degenerate repetitions and introduces nucleus sampling, refining the open-ended decoding techniques used in DialoGPT.
  • Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). G-Eval provides an advanced LLM-based evaluation methodology that offers much higher human alignment when scoring open-ended dialogue responses.
Cover for DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation

Abstract

We present a large, tunable neural conversational response generation model, DialoGPT (dialogue generative pre-trained transformer). Trained on 147M conversation-like exchanges extracted from Reddit comment chains over a period spanning from 2005 through 2017, DialoGPT extends the Hugging Face PyTorch transformer to attain a performance close to human both in terms of automatic and human evaluation in single-turn dialogue settings. We show that conversational systems that leverage DialoGPT generate more relevant, contentful and context-consistent responses than strong baseline systems. The pre-trained model and training pipeline are publicly released to facilitate research into neural response generation and the development of more intelligent open-domain dialogue systems.

Table of Contents

  • 1 Introduction
  • 2 Dataset
  • 3 Method
  • 3.1 Model Architecture
  • 3.2 Mutual Information Maximization
  • 4 Result
  • 4.1 Experimental Details
  • 4.2 DSTC-7 Dialogue Generation Challenge
  • 4.3 A New Reddit Multi-reference Dataset
  • 4.4 Re-ranking The Response Using MMI
  • 4.5 Generation Examples
  • 4.6 Human Evaluation
  • 5 Related work
  • 6 Limitations and risks
  • 7 Conclusion
  • References
  • A Additional Details of Human Evaluation

Knowls

  1. Knowl 1 — Autoregressive Multi-Turn Dialogue Formulation in DialoGPT

    model/method

    DialoGPT formulates conversational response generation as an autoregressive language modeling problem parameterized by a multi-layer Transformer decoder. A multi-turn dialogue session comprising KK turns T1,T2,…,TKT_1, T_2, \dots, T_K is concatenated into a continuous token sequence x1,x2,…,xNx_1, x_2, \dots, x_N terminated by an end-of-text token.

    For a dialogue context (source) S=x1,…,xmS = x_1, \dots, x_m and a ground-truth response (target) T=xm+1,…,xNT = x_{m+1}, \dots, x_N, the conditional generation probability p(T∣S)p(T \mid S) factorizes autoregressively as:

    p(T∣S)=∏n=m+1Np(xn∣x1,…,xn−1)p(T \mid S) = \prod_{n=m+1}^{N} p(x_n \mid x_1, \dots, x_{n-1})

    For the entire multi-turn conversation T1,…,TKT_1, \dots, T_K, the joint conditional probability over consecutive turns is expressed as:

    p(TK,…,T2∣T1)=∏i=2Kp(Ti∣T1,…,Ti−1)p(T_K, \dots, T_2 \mid T_1) = \prod_{i=2}^{K} p(T_i \mid T_1, \dots, T_{i-1})

    Optimizing the single language modeling objective over the concatenated dialogue sequence is mathematically equivalent to simultaneously optimizing the conditional distributions of all source-target sub-sequences in the session.

  2. Knowl 2 — Reddit Dialogue Corpus Extraction and Filtering Pipeline

    experimental setup

    The training corpus for DialoGPT was extracted from Reddit comment chains between 2005 and 2017. Tree-structured reply threads were converted into multi-turn dialogue instances by extracting each path from the thread root node to every leaf node. To eliminate noise, non-English text, repetitive loops, and offensive content, the corpus was filtered according to seven rules:

    1. Removal of any instance where a URL appears in the source context or target response.
    2. Removal of instances where the target response contains immediate word repetitions of three or more words.
    3. Removal of instances where the target response lacks at least one of the top-50 most frequent English words (e.g., "the", "of", "a"), filtering out non-English content.
    4. Removal of responses containing bracket markers such as [ or ], which indicate markdown or system markup.
    5. Removal of instances where the combined length of the source context and target response exceeds 200 words.
    6. Filtering of offensive language via keyword and phrase matching against a blocklist, alongside the wholesale removal of subreddits flagged for offensive material.
    7. Blandness filtering: removal of responses where 90%90\% or more of their constituent tri-grams appeared more than 1,000 times across the entire Reddit dataset (accounting for roughly 1%1\% of the raw corpus).

    The final filtered dataset contains 147,116,725 dialogue sessions comprising 1.8 billion words.

  3. Knowl 3 — Maximum Mutual Information Hypothesis Reranking with a Backward Model

    model/method

    To alleviate the tendency of autoregressive dialogue models to generate bland, uninformative, and repetitive responses, DialoGPT applies Maximum Mutual Information (MMI) scoring during decoding using a pre-trained backward model.

    Given a source context SS, the procedure operates as follows:

    1. Generate a candidate set of 16 candidate hypothesis responses {H1,…,H16}\{H_1, \dots, H_{16}\} by sampling from the forward generation model P(T∣S)P(T \mid S) using top-KK sampling with K=10K = 10.
    2. Score each candidate hypothesis HiH_i using a pre-trained backward model P(S∣Hi)P(S \mid H_i), which is a Transformer model of identical architecture trained to predict the preceding dialogue context from the generated response.
    3. Select the hypothesis H∗H^* that minimizes the backward model cross-entropy loss (maximizing P(S∣Hi)P(S \mid H_i)).

    Because generic, repetitive responses (such as "I don't know") can follow an extremely large number of distinct prompts, their backward probability P(S∣H)P(S \mid H) for any single specific prompt SS is very low. Maximizing backward probability thus penalizes generic responses and promotes context-specific, informative generation.

  4. Knowl 4 — Degeneracy of Policy Gradient Mutual Information Optimization in Transformers

    empirical result

    Optimizing a conversational mutual information reward R≜P(S∣H)R \triangleq P(S \mid H) (where SS is the source context and HH is the generated hypothesis) using reinforcement learning via REINFORCE policy gradients with a sample-averaged baseline fails when applied to deep Transformer models.

    While the validation reward increases during training, the Transformer model rapidly converges to a degenerate local optimum in which the policy simply echoes or copies the source prompt verbatim (H=SH = S). Because identity copying yields high backward model reconstruction likelihood, the high representational capacity of multi-layer self-attention causes the model to exploit this trivial shortcut, creating a "parroting" generator rather than producing context-appropriate, diverse replies. This failure mode does not occur as readily in recurrent neural network (RNN) architectures.

  5. Knowl 5 — DialoGPT Architecture Specifications and Distributed Training Pipeline

    experimental setup

    DialoGPT adopts the OpenAI GPT-2 multi-layer Transformer decoder architecture with Byte Pair Encoding (BPE) vocabulary of 50,257 tokens, evaluated across three parameter scales:

    • Small (117M parameters): 12 layers, hidden embedding dimension Demb=768D_{\text{emb}} = 768, trained with a per-GPU batch size of B=128B = 128.
    • Medium (345M parameters): 24 layers, hidden embedding dimension Demb=1024D_{\text{emb}} = 1024, trained with a per-GPU batch size of B=64B = 64.
    • Large (762M parameters): 36 layers, hidden embedding dimension Demb=1280D_{\text{emb}} = 1280, trained with a per-GPU batch size of B=32B = 32.

    Models were trained on 16 Nvidia V100 GPUs with NVLink using the Noam learning rate scheduler with 16,000 warm-up steps and early stopping on validation loss (up to 5 epochs for 117M and 345M models; up to 3 epochs for the 762M model). Training throughput was maximized through three techniques:

    1. Compressing the full training set into a lazy-loading database with large block pre-fetching to prevent memory overflow and minimize disk access overhead.
    2. Asynchronous data-loading processes that yield near-linear throughput scaling with GPU count.
    3. Dynamic batching grouping conversations of similar sequence length within individual batches.
  6. Knowl 6 — Benchmark Performance on DSTC-7 Dialogue Generation Challenge

    data/table

    DialoGPT was evaluated on the DSTC-7 grounded dialogue challenge benchmark using a multi-reference test set of 2,208 instances (5 ground-truth reference responses and 1 held-out human response per prompt). DialoGPT was fine-tuned purely on conversational pairs without utilizing the external grounding documents provided in the challenge.

    Method NIST-2 NIST-4 BLEU-2 BLEU-4 METEOR Dist-1 Dist-2
    PersonalityChat 0.19 0.20 10.44% 1.47% 5.42% 5.9% 16.4%
    Team B (DSTC-7 Winner) 2.51 2.52 14.35% 1.83% 8.07% 10.9% 32.5%
    DialoGPT (117M) 1.58 1.60 10.36% 2.02% 7.17% 6.2% 18.94%
    GPT (345M) 1.78 1.79 9.13% 1.06% 6.38% 11.9% 44.2%
    DialoGPT (345M) 2.80 2.82 14.16% 2.31% 8.51% 9.1% 39.7%
    DialoGPT (345M, Beam width 10) 2.92 2.97 19.18% 6.05% 9.29% 15.7% 51.0%
    Human 2.62 2.65 12.35% 3.13% 8.31% 16.7% 67.0%

    DialoGPT with 345M parameters using beam search (beam width 10) achieves state-of-the-art results across automated translation and diversity metrics, outperforming both the competition winner (Team B) and PersonalityChat, while registering higher automatic BLEU-4, NIST-4, and METEOR scores than held-out human references.

  7. Knowl 7 — Evaluation on 6K Reddit Multi-Reference Benchmark: Scratch vs GPT-2 Initialization

    data/table

    DialoGPT models were evaluated on a 6,000-example multi-reference Reddit test dataset across model scales (117M, 345M, 762M), initialization setups (trained from scratch vs fine-tuned from pre-trained OpenAI GPT-2), and decoding algorithms (greedy, beam search with width 10, and MMI backward reranking).

    Method NIST-2 NIST-4 BLEU-2 BLEU-4 METEOR Entropy-4 Dist-2
    PersonalityChat 0.78 0.79 11.22% 1.95% 6.93% 8.37 18.8%
    From scratch:
    DialoGPT (117M) 1.23 1.37 9.74% 1.77% 6.17% 7.11 15.9%
    DialoGPT (345M) 2.51 3.08 16.92% 4.59% 9.34% 9.03 25.6%
    DialoGPT (762M) 2.52 3.10 17.87% 5.19% 9.53% 9.32 29.3%
    From GPT-2:
    DialoGPT (117M) 2.39 2.41 10.54% 1.55% 7.53% 10.77 39.9%
    DialoGPT (345M) 3.00 3.06 16.96% 4.56% 9.81% 9.12 26.3%
    DialoGPT (345M, Beam) 3.40 3.50 21.76% 7.92% 10.74% 10.48 48.74%
    DialoGPT (762M) 2.84 2.90 18.66% 5.25% 9.66% 9.72 29.93%
    DialoGPT (345M, MMI) 3.28 3.33 15.68% 3.94% 11.23% 11.25 45.55%
    Human 3.41 4.25 17.90% 7.48% 10.64% 10.99 63.0%

    Key findings include:

    1. Initializing from pre-trained GPT-2 provides substantial performance gains for smaller models (117M NIST-4 rises from 1.37 to 2.41), whereas larger models (345M and 762M) trained from scratch reach performance levels comparable to GPT-2 initialized models.
    2. DialoGPT (345M, Beam) achieves the highest BLEU-4 (7.92%7.92\%) and NIST-4 (3.503.50), surpassing held-out human responses on BLEU.
    3. DialoGPT (345M, MMI) yields the highest METEOR (11.23%11.23\%) and Entropy-4 (11.2511.25) among model variants, demonstrating superior lexical diversity and richness.
  8. Knowl 8 — Crowdsourced Human Preference Evaluation and Significance Testing

    data/table

    Human evaluation was conducted on 2,000 randomly sampled Reddit test contexts evaluated by crowdsourced judges using pairwise comparisons across Relevance, Informativeness, and Human-likeness on a 3-point scale. Significance was determined via 10,000 bootstrap iterations at α=0.05\alpha = 0.05.

    Comparison (System A vs System B) System A Win Neutral System B Win Stat. Significance
    Relevance:
    DialoGPT (345M) vs PersonalityChat 72% 9% 19% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs DialoGPT (345M, MMI) 40% 9% 52% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs DialoGPT (345M, Beam) 50% 10% 40% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs DialoGPT (762M) 45% 10% 45% Not sig. (p=0.7066p = 0.7066)
    DialoGPT (345M) vs Human Response 45% 9% 47% Not sig. (p=0.0552p = 0.0552)
    DialoGPT (345M, MMI) vs Human Response 48% 9% 43% p≤0.0001p \le 0.0001
    Informativeness:
    DialoGPT (345M) vs PersonalityChat 77% 5% 19% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs DialoGPT (345M, MMI) 41% 4% 54% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs Human Response 45% 4% 51% p≤0.00001p \le 0.00001
    DialoGPT (345M, MMI) vs Human Response 50% 4% 46% p≤0.001p \le 0.001
    Human-likeness:
    DialoGPT (345M) vs PersonalityChat 76% 4% 20% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs DialoGPT (345M, MMI) 41% 5% 54% p≤0.00001p \le 0.00001
    DialoGPT (345M) vs Human Response 45% 4% 50% p≤0.0001p \le 0.0001
    DialoGPT (345M, MMI) vs Human Response 50% 4% 46% p≤0.01p \le 0.01

    Judges exhibited a strong preference for DialoGPT over the PersonalityChat baseline. The vanilla DialoGPT 345M model achieves parity with human responses on relevance (p=0.0552p = 0.0552, difference not significant) and matches the 762M model across all dimensions (p>0.70p > 0.70). Furthermore, DialoGPT (345M, MMI) was statistically significantly preferred over actual human responses across all three criteria (relevance 48% vs 43%, informativeness 50% vs 46%, human-likeness 50% vs 46%).

  9. Knowl 9 — Geometric Mean Phenomenon in Multi-Reference Conversational Metric Overestimation

    theoretical result

    In multi-reference conversational evaluations, automatic overlap metrics (such as BLEU and NIST) computed for a generated model response RgR_g can surpass the score of an individual held-out human response R4R_4 without implying that RgR_g is qualitatively superior or more realistic than human speech.

    Because open-domain conversation is intrinsically one-to-many, multiple disparate responses {R1,R2,R3,R4}\{R_1, R_2, R_3, R_4\} can all be valid reactions to a given source context SS. A generative language model trained via maximum likelihood learns to generate the most probable response across the data distribution, which tends toward the geometric center (mean semantic representation) of plausible candidate responses. As a result, the model output RgR_g exhibits a shorter average semantic distance to all test reference responses {R1,R2,R3}\{R_1, R_2, R_3\} than a single human response R4R_4, which may represent an idiosyncratic or peripheral point in the semantic space.

  10. Knowl 10 — Safety Vulnerabilities, Sycophancy, and Toxic Generation Risks

    limitation

    Despite automated blocklist filtering and subreddit exclusions prior to training, large pre-trained conversational models like DialoGPT retain inherent safety risks:

    1. Generation of toxic, offensive, or biased responses when triggered by adversarial, controversial, or sensitive user prompts.
    2. Expression of historical and societal biases reflected in large-scale online discussion forums.
    3. An inherent tendency to express sycophantic agreement with user inputs, including unethical, biased, or harmful propositions.

    The released model weights do not incorporate runtime safety filters, leaving decoder-level constraint enforcement, content filtering, and toxic output mitigation to downstream implementation.

Coverage note — None was omitted; all core contributions including dataset filtering, model architecture, MMI re-ranking, RL failure analysis, experimental tables (DSTC-7 and Reddit 6K), human evaluation significance, geometric mean analysis, and stated limitations are fully covered.

References

  1. 1.M. Burtsev, A. Seliverstov, R. Airapetyan, M. Arkhipov, D. Baymurzina, N. Bushkov, O. Gureenkova, T. Khakhulin, Y. Kuratov, D. Kuznetsov, A. Litinsky, V. Logacheva, A. Lymar, V. Malykh, M. Petrov, V. Polulyakh, L. Pugachev, A. Sorokin, M. Vikhreva, and M. Zaynutdinov. 2018. DeepPavlov: Open-source library for dialogue systems. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics-System Demonstrations.
  2. 2.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL 2019.
  3. 3.E. Dinan, V. Logacheva, V. Malykh, A. Miller, K. Shuster, J. Urbanek, D. Kiela, A. Szlam, I. Serban, R. Lowe, S. Prabhumoye, A. W. Black, A. Rudnicky, J. Williams, J. Pineau, M. Burtsev, and J. Weston. 2019. The second conversational intelligence challenge (ConvAI2).
  4. 4.George Doddington. 2002. Automatic evaluation of machine translation quality using n-gram cooccurrence statistics. In Proceedings of the second international conference on Human Language Technology Research. Morgan Kaufmann Publishers Inc.
  5. 5.Michel Galley, Chris Brockett, Xiang Gao, Jianfeng Gao, and Bill Dolan. 2019. Grounded response generation task at DSTC7. In AAAI Dialog System Technology Challenges Workshop.
  6. 6.J. Gao, M. Galley, and L. Li. 2019a. Neural approaches to conversational AI. Foundations and Trends in Information Retrieval.
  7. 7.Xiang Gao, Sungjin Lee, Yizhe Zhang, Chris Brockett, Michel Galley, Jianfeng Gao, and Bill Dolan. 2019b. Jointly optimizing diversity and relevance in neural response generation. NAACL-HLT 2019.
  8. 8.Xiang Gao, Yizhe Zhang, Sungjin Lee, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2019c. Structuring latent spaces for stylized response generation. EMNLP-IJCNLP.
  9. 9.M. Gardner, J. Grus, M. Neumann, O. Tafjord, P. Dasigi, N. F. Liu, M. Peters, M. Schmitz, and L. S. Zettlemoyer. 2018. AllenNLP: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software.
  10. 10.Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey P Bigham. 2019. Investigating evaluation of open-domain dialogue systems with human generated multiple references. arXiv preprint arXiv:1907.10568.
  11. 11.Z. Hu, H. Shi, Z. Yang, B. Tan, T. Zhao, J. He, W. Wang, L. Qin, D. Wang, et al. 2018. Texar: A modularized, versatile, and extensible toolkit for text generation. ACL.
  12. 12.HuggingFace. 2019. PyTorch transformer repository. https://github.com/huggingface/pytorch-transformers.
  13. 13.Alon Lavie and Abhaya Agarwal. 2007. Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments. In Proceedings of the Second Workshop on Statistical Machine Translation, pages 228–231. Association for Computational Linguistics.
  14. 14.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016a. A diversity-promoting objective function for neural conversation models. NAACL.
  15. 15.Jiwei Li, Michel Galley, Chris Brockett, Georgios P Spithourakis, Jianfeng Gao, and Bill Dolan. 2016b. A persona-based neural conversation model. ACL.
  16. 16.A. H. Miller, W. Feng, A. Fisch, J. Lu, D. Batra, A. Bordes, D. Parikh, and J. Weston. 2017. ParlAI: A dialog research software platform. In Proceedings of the 2017 EMNLP System Demonstration.
  17. 17.Oluwatobi Olabiyi and Erik T Mueller. 2019. Multi-turn dialogue response generation with autoregressive transformer models. arXiv preprint:1908.01841.
  18. 18.Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. ACL.
  19. 19.M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer. 2018. Deep contextualized word representations. NAACL.
  20. 20.Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu, Xiang Gao, Bill Dolan, Yejin Choi, and Jianfeng Gao. 2019. Conversing by reading: Contentful neural conversation with on-demand machine reading. ACL.
  21. 21.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. 2018. Language models are unsupervised multitask learners. Technical report, OpenAI.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint:1910.10683.
  23. 23.R. Sennrich, B. Haddow, and A. Birch. 2016. Neural machine translation of rare words with subword units. ACL.
  24. 24.Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2017. A hierarchical latent variable encoder-decoder model for generating dialogues. AAAI.
  25. 25.Vighnesh Leonardo Shiv, Chris Quirk, Anshuman Suri, Xiang Gao, Khuram Shahid, Nithya Govindarajan, Yizhe Zhang, Jianfeng Gao, Michel Galley, Chris Brockett, et al. 2019. Microsoft icecaps: An opensource toolkit for conversation modeling. ACL.
  26. 26.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. NeurIPS.
  27. 27.Ronald J Williams. 1992. Simple statistical gradientfollowing algorithms for connectionist reinforcement learning. Machine learning.
  28. 28.Thomas Wolf, Victor Sanh, Julien Chaumond, and Clement Delangue. 2019. TransferTransfo: A transfer learning approach for neural network based conversational agents. CoRR, abs/1901.08149.
  29. 29.Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018. Generating informative and diverse conversational responses via adversarial information maximization. NeurIPS.
  30. 30.Yizhe Zhang, Xiang Gao, Sungjin Lee, Chris Brockett, Michel Galley, Jianfeng Gao, and Bill Dolan. 2019. Consistent dialogue generation with self-supervised feature learning. arXiv preprint:1903.05759.

Citation

MLA
Zhang, Y., et al. “DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation”. arXiv, 2019, http://arxiv.org/abs/1911.00536v3.
APA
Zhang, Y., Sun, S., Galley, M., Chen, Y.-C., Brockett, C., Gao, X., Gao, J., Liu, J., & Dolan, B. (2019). DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation. arXiv. http://arxiv.org/abs/1911.00536v3
Chicago
Zhang, Y., S. Sun, M. Galley, et al. 2019. “DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation”. arXiv. http://arxiv.org/abs/1911.00536v3.
Harvard
Zhang, Y. et al. (2019) “DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1911.00536v3.
Vancouver
1. Zhang Y, Sun S, Galley M, Chen Y-C, Brockett C, Gao X, Gao J, Liu J, Dolan B (2019) DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation. arXiv

BibTeX

@article{zhang2019dialogpt,
  title = {DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation},
  author = {Zhang, Yizhe and Sun, Siqi and Galley, Michel and Chen, Yen-Chun and Brockett, Chris and Gao, Xiang and Gao, Jianfeng and Liu, Jingjing and Dolan, Bill},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1911.00536v3},
  eprint = {1911.00536}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/