DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation

Wei ChenYeyun GongSong WangBolun YaoWeizhen QiZhongyu WeiXiaowu HuBartuer ZhouYi MaoWeizhu Chen

article2022ACL58 citations

Proposes a pre-trained dialogue model that integrates continuous latent variables into an encoder-decoder architecture to resolve the one-to-many response problem, achieving state-of-the-art response diversity and relevance across PersonaChat, DailyDialog, and DSTC7-AVSD benchmarks.

Listen

Open-domain dialogue systems frequently suffer from generating generic, repetitive, and dull responses. This issue largely stems from the "one-to-many" nature of human conversation, where a single conversational context can naturally lead to many valid, distinct replies. While modern pre-trained language models have improved dialogue capabilities, standard architectures struggle to effectively balance output relevance with lexical and semantic diversity.

The article introduces and evaluates DialogVED, a novel pre-trained dialogue response generation framework. Its main objective is to demonstrate that incorporating continuous latent variables into a pre-trained encoder-decoder network significantly enhances both the relevance and diversity of generated conversational responses across multiple standard benchmarks.

To achieve this, the authors built a 24-layer Transformer-based architecture that integrates continuous latent variables via an extra prior network and decoder memory vectors. The model was pre-trained on a large-scale corpus of 215 million Reddit comment pairs using four joint optimization objectives: masked language span modeling, future n-gram response prediction, Kullback-Leibler divergence regularization, and a bag-of-words prediction loss to avoid latent space collapse. To evaluate credibility across diverse conversation settings, the pre-trained system was fine-tuned and tested on three standard benchmarks: DailyDialog for general chit-chat, Persona-Chat for knowledge-grounded dialogue, and DSTC7-AVSD for conversational question answering.

The findings show that DialogVED establishes a new state of the art across all three evaluation benchmarks. When paired with top-K sampling, DialogVED outperformed previous leading systems like PLATO, achieving higher word-overlap relevance (BLEU-1 of 0.431 vs. 0.397 on DailyDialog; 0.428 vs. 0.406 on Persona-Chat) alongside substantially higher distinct n-gram diversity (Distinct-2 of 0.372 vs. 0.291 on DailyDialog; 0.273 vs. 0.121 on Persona-Chat). On the DSTC7-AVSD question-answering task, where factual precision outweighs diversity, standard DialogVED configurations exceeded all baseline models across all accuracy metrics. Furthermore, ablation analyses demonstrated that combining absolute conversational turn and speaker role position embeddings provided optimal performance, while smaller latent space dimensions favored precision over diversity. Human evaluation confirmed these quantitative results, rating DialogVED responses as more coherent and informative than the baseline.

These results indicate that combining continuous latent variable modeling with large-scale pre-training resolves the trade-off between conversational relevance and diversity. For organizations developing conversational agents, this architecture delivers more engaging and natural user experiences while maintaining context fidelity. The framework also provides flexible operational control, allowing practitioners to tune decoding strategies and latent space sizing based on whether a specific deployment requires factual precision or creative diversity.

Organizations developing customer-facing conversational artificial intelligence should consider adopting latent variable encoder-decoder architectures to upgrade dialogue quality. Practitioners should specifically implement absolute turn and speaker position embeddings and select decoding methods tailored to their use case: top-K sampling for open-ended, engaging chit-chat, and beam or greedy search for task-oriented, fact-driven question answering.

The study's primary limitations involve potential social and political biases inherent in the Reddit pre-training corpus, despite heuristic filtering, as well as the substantial computational investment required (32 high-end GPUs over five days). While confidence in the benchmark improvements is high, decision-makers should exercise caution regarding bias auditing and safety filtering before deploying models pre-trained on open web data into production environments.

No sufficiently relevant recommendations were found.

Cover for DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation

Abstract

Dialog response generation in open domain is an important research topic where the main challenge is to generate relevant and diverse responses. In this paper, we propose a new dialog pre-training framework called DialogVED, which introduces continuous latent variables into the enhanced encoder-decoder pre-training framework to increase the relevance and diversity of responses. With the help of a large dialog corpus (Reddit), we pre-train the model using the following 4 tasks, used in training language models (LMs) and Variational Autoencoders (VAEs) literature: 1) masked language model; 2) response generation; 3) bag-of-words prediction; and 4) KL divergence reduction. We also add additional parameters to model the turn structure in dialogs to improve the performance of the pre-trained model. We conduct experiments on PersonaChat, DailyDialog, and DSTC7-AVSD benchmarks for response generation. Experimental results show that our model achieves the new state-of-the-art results on all these datasets.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Model Architecture
  • 2.2 Encoder
  • 2.3 Decoder
  • 2.4 Latent Variable
  • 2.5 Mask Language Spans
  • 2.6 Reduce KL-vanishing
  • 2.7 Position Embeddings
  • 2.8 Pre-training Objectives
  • 3 Experiments
  • 3.1 DataSets and Baselines
  • 3.1.1 Pre-training Corpus
  • 3.1.2 Fine-tuning Benchmarks
  • 3.1.3 Baselines
  • 3.2 Model Configuration
  • 3.3 Main Results
  • 3.4 Parameters and Position Analysis
  • 3.4.1 Balancing Accuracy and Diversity with Sampling
  • 3.4.2 Position Embeddings
  • 3.5 Human Evaluation
  • 4 Related Work
  • 5 Conclusion
  • Acknowledgments
  • Ethical Statement
  • References
  • A Case Study

Knowls

  1. Knowl 1 — DialogVED models responses with a context-conditioned continuous latent variable

    model/method

    DialogVED is a Transformer encoder–decoder for dialogue response generation. Given a dialogue context cc, it represents a response rr and a continuous latent vector z∈RPz \in \mathbb{R}^{P} with the factorization

    pθ,ϕ(r,z∣c)=pθ(r∣c,z) pϕ(z∣c).p_{\theta,\phi}(r,z\mid c)=p_{\theta}(r\mid c,z)\,p_{\phi}(z\mid c).

    The context encoder’s final-layer representation of its special classification token, h[CLS]∈RHh_{[CLS]} \in \mathbb{R}^{H}, is passed to a multilayer perceptron that predicts the mean vector μ\mu and elementwise variance vector σ2\sigma^2 of a diagonal-Gaussian prior: pϕ(z∣c)=N(μ,diag⁡(σ2))p_{\phi}(z\mid c)=\mathcal{N}(\mu,\operatorname{diag}(\sigma^2)). At generation time, DialogVED samples zz from this prior and conditions the decoder on both cc and zz. The intended division of labor is for zz to capture response-level variation associated with the context, while autoregressive decoding produces the response’s token sequence.

  2. Knowl 2 — Latent memory and future n-gram prediction condition the decoder

    model/method

    DialogVED uses a future-prediction decoder with nn auxiliary self-attention streams in addition to its main stream. At each decoding step tt, the streams predict the next nn continuous response tokens, rather than training only on the next token: the target is rt:t+n−1r_{t:t+n-1} given the preceding response tokens and context. The paper uses n=2n=2 in pre-training. This objective is intended to encourage planning beyond local token correlations.

    To make the latent vector available throughout decoding, DialogVED maps zz to a memory vector hMem=WMzh_{\mathrm{Mem}}=W_M z, structured as an additional key–value pair for decoder attention. At every decoder layer, this shared memory is appended to the main stream’s attention keys and values, functioning like a virtual token. The auxiliary future-prediction streams are affected through their interaction with the main stream. Thus, the latent vector can guide token generation at every step.

  3. Knowl 3 — Four losses jointly train the latent-variable dialogue model

    equation

    DialogVED is pre-trained by minimizing the sum of masked-span, response-reconstruction, free-bits KL, and bag-of-words losses:

    L=LM+Lrc+Lkl′+LBOW.\mathcal{L}=\mathcal{L}_{M}+\mathcal{L}_{rc}+\mathcal{L}'_{kl}+\mathcal{L}_{BOW}.

    Here, cc is the context, r=(r1,…,rT)r=(r_1,\ldots,r_T) is its target response of TT tokens, and zz is the continuous latent vector. The response loss trains future n-gram prediction under a training posterior distribution q(z)q(z):

    Lrc=−Eq(z)[log⁡∏tpθ(rt:t+n−1∣r<t,c,z)].\mathcal{L}_{rc}=-\mathbb{E}_{q(z)}\left[\log\prod_t p_{\theta}(r_{t:t+n-1}\mid r_{<t},c,z)\right].

    The KL term compares that posterior with the context-conditioned prior pϕ(z∣c)p_{\phi}(z\mid c). With free-bits threshold λ\lambda, the paper uses

    Lkl′=−∑imax⁡ ⁣(λ,KL⁡(q(zi) ∥ pϕ(zi∣c))),\mathcal{L}'_{kl}=-\sum_i\max\!\left(\lambda,\operatorname{KL}\big(q(z_i)\,\|\,p_{\phi}(z_i\mid c)\big)\right),

    where ziz_i is latent coordinate ii. The bag-of-words loss encourages zz to predict response words without autoregressive conditioning: if f=softmax⁡(MLP⁡z([z⊕h[CLS]]))f=\operatorname{softmax}(\operatorname{MLP}_z([z\oplus h_{[CLS]}])) is a vocabulary probability vector, then LBOW=−∑t=1Tlog⁡frt\mathcal{L}_{BOW}=-\sum_{t=1}^{T}\log f_{r_t}. The paper motivates free bits and the bag-of-words objective as ways to reduce posterior collapse, in which the decoder ignores zz and the KL term vanishes.

  4. Knowl 4 — Pre-training reconstructs randomly corrupted context spans

    model/method

    During pre-training, DialogVED masks spans in the dialogue context and trains the encoder to recover the original tokens at masked positions. The procedure randomly selects nn context tokens, extends each selected token to a span of fixed length mm, then sorts, deduplicates, and checks the resulting spans against context boundaries. The total masked content is approximately 15% of the context. Each masked token is replaced with a [MASK] token 80% of the time, a random token 10% of the time, and left unchanged 10% of the time. The encoder’s hidden state at each selected position is used to predict its original token with a vocabulary-level cross-entropy loss. Context masking is used in pre-training, not fine-tuning.

  5. Knowl 5 — Turn, speaker, and relative-position embeddings encode dialogue structure

    model/method

    DialogVED explores two ways to represent dialogue structure in its input and attention layers. With absolute position embeddings, each token’s input representation sums its token embedding with learned turn-position and speaker-role embeddings. With relative position embeddings, the self-attention score between query token ii and key token jj uses a learned key-side representation based on both their token offset dtokend_{token} and turn offset dturnd_{turn}; both offsets are bucketed before selecting embeddings. In the paper’s notation, the score is eij=xiWQ(xjWK+aijK)T/dze_{ij}=x_iW_Q(x_jW_K+a^K_{ij})^T/\sqrt{d_z}, where xi,xjx_i,x_j are token representations, WQ,WKW_Q,W_K are query and key projections, aijKa^K_{ij} is the relative representation, and dzd_z is the attention dimension.

    On DailyDialog, the best tested combination was turn and role absolute embeddings together: BLEU-1/2 was 0.494/0.4350.494/0.435 and Distinct-1/2 was 0.042/0.2320.042/0.232, compared with 0.481/0.4210.481/0.421 and 0.042/0.2320.042/0.232 without these additions. The paper reports that absolute and relative embeddings can each help, but combining all tested absolute and relative components was not uniformly beneficial: using all four yielded BLEU-1/2 0.483/0.4350.483/0.435 and Distinct-1/2 0.039/0.2280.039/0.228.

  6. Knowl 6 — Pre-training and downstream evaluation use Reddit and three dialogue benchmarks

    experimental setup

    DialogVED is pre-trained on 215 million Reddit context–response samples (42 GB), divided for efficiency into Reddit-Short and Reddit-Long. Samples of similar context length are batched together to reduce padding; the two subsets use different batch sizes, with short samples processed before long samples in each epoch. The table gives dataset size, average turns, and average WordPiece tokens in context/response:

    Dataset Examples Turns Tokens (context/response)
    Reddit-Short 214M 2.6 28.6/16.0
    Reddit-Long 726K 6.9 137.1/21.2
    DailyDialog 76K 5.9 75.6/15.0
    Persona-Chat 122K 8.4 95.1/12.2
    DSTC7-AVSD 76K 10.9 102.1/10.7

    The downstream datasets represent chit-chat (DailyDialog), persona-grounded conversation (Persona-Chat), and audiovisual-scene-grounded conversational question answering (DSTC7-AVSD). The model has 12 encoder layers and 12 decoder layers, hidden/embedding size 1024, feed-forward size 4096, and latent dimension P=64P=64 in the main experiments. Pre-training uses Adam at learning rate 3×10−43\times10^{-4} for six epochs on 32 Nvidia Tesla V100 32 GB GPUs, taking about five days; the future n-gram size is 2. Fine-tuning uses learning rate 10−410^{-4}, batch size 512, 2,000 warmup updates from 10−710^{-7}, and 10 epochs, selecting the checkpoint with lowest validation loss.

  7. Knowl 7 — DialogVED achieves strong automatic results across the three benchmarks

    empirical result

    The main automatic evaluation uses beam search with beam size 5 for default DialogVED and latent size P=64P=64. The following results compare DialogVED variants with the relevant pre-trained baselines. DailyDialog and Persona-Chat report BLEU-1/2 and Distinct-1/2; DSTC7-AVSD reports BLEU-1/2/3/4, METEOR, ROUGE-L, and CIDEr.

    DailyDialog Persona-Chat
    Model B1 B2 D1 D2 B1 B2 D1 D2
    PLATO 0.397 0.311 0.054 0.291 0.406 0.315 0.021 0.121
    ProphetNet 0.443 0.392 0.039 0.211 0.466 0.391 0.013 0.075
    DialogVED w/o latent 0.461 0.407 0.041 0.222 0.459 0.380 0.010 0.062
    DialogVED - Greedy 0.459 0.410 0.045 0.265 0.470 0.387 0.016 0.103
    DialogVED - Sampling 0.431 0.370 0.058 0.372 0.428 0.357 0.032 0.273
    DialogVED 0.481 0.421 0.042 0.232 0.482 0.399 0.015 0.094

    DialogVED’s default decoding has the highest reported BLEU-1/2 among these listed DailyDialog and Persona-Chat systems, while top-KK sampling gives the highest diversity scores on both datasets. On DSTC7-AVSD, the version without a latent variable has the highest listed BLEU and ROUGE-L values, while DialogVED has the highest listed CIDEr score. Thus, the paper’s results support strong performance across tasks, but the latent-variable version does not lead every metric or benchmark.

  8. Knowl 8 — Latent size and top-K sampling trade off overlap and response diversity

    empirical result

    On DailyDialog, the paper varies latent dimension PP and the number KK of highest-probability tokens eligible at each decoding step for top-KK sampling. The reported results are:

    PP KK BLEU-1 BLEU-2 Distinct-1 Distinct-2
    32 5 0.448 0.385 0.042 0.289
    32 20 0.443 0.376 0.045 0.317
    32 50 0.442 0.375 0.047 0.332
    32 100 0.439 0.374 0.051 0.347
    64 5 0.442 0.383 0.046 0.308
    64 20 0.437 0.374 0.050 0.340
    64 50 0.434 0.371 0.054 0.364
    64 100 0.431 0.370 0.058 0.372

    For both latent sizes, increasing KK lowers BLEU-1/2 and raises Distinct-1/2, exposing an overlap–diversity trade-off. At the same KK, P=32P=32 generally gives stronger BLEU scores, whereas P=64P=64 gives greater diversity. The paper therefore treats decoding choice and latent dimension as adjustable to the application’s preference for reference overlap versus varied output.

  9. Knowl 9 — Human comparisons favor DialogVED in coherence and sampling in informativeness

    empirical result

    Human evaluation compares responses to 100 randomly selected dialogue contexts. One group compares DialogVED with PLATO; a second compares top-KK-sampling DialogVED with PLATO. Annotators judge fluency, coherence, informativeness, and overall quality as win, tie, or loss; the reported values are win/loss proportions, with ties omitted from the table. The average Cohen’s κ\kappa was 0.729 for the first group and 0.743 for the second.

    DialogVED vs. PLATO Sampling vs. PLATO
    Aspect Win Lose Win Lose
    Fluency 0.16 0.11 0.19 0.13
    Coherence 0.38 0.16 0.22 0.24
    Informativeness 0.13 0.11 0.31 0.14
    Overall 0.26 0.14 0.24 0.17

    The paper interprets the first comparison as favoring DialogVED over PLATO in coherence, with close informativeness; for sampling, it reports an informativeness advantage accompanied by somewhat weaker coherence. The table also shows that neither system wins every aspect for every comparison.

  10. Knowl 10 — Reddit pre-training may transmit social and political biases

    limitation

    The paper notes that the online Reddit pre-training corpus may contain political and social biases and that DialogVED may inherit them. The authors report filtering controversial material and removing offensive content when possible, but do not claim that these steps eliminate bias. They also state that their work does not focus on explicit knowledge utilization, despite evaluating on knowledge-grounded and audiovisual-scene dialogue tasks.

Coverage note — The appendix’s individual response examples and detailed baseline descriptions are omitted because they are illustrative or contextual rather than additional systematic findings.

References

  1. 1.Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K Marks, Chiori Hori, Peter Anderson, et al. 2019a. Audio visual scene-aware dialog. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7558–7567.
  2. 2.Huda Alamri, Chiori Hori, Tim K Marks, Dhruv Batra, and Devi Parikh. 2019b. Audio visual scene-aware dialog (AVSD) track for natural language generation in DSTC7. In AAAI workshop on the 7th edition of Dialog System Technology Challenge (DSTC7).
  3. 3.Hareesh Bahuleyan, Lili Mou, Olga Vechtomova, and Pascal Poupart. 2017. Variational attention for sequence-to-sequence models. arXiv preprint arXiv:1712.08207.
  4. 4.Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2020. PLATO: Pre-trained dialogue generation model with discrete latent variable. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 85–96, Online. Association for Computational Linguistics.
  5. 5.Basma El Amel Boussaha, Nicolas Hernandez, Christine Jacquin, and Emmanuel Morin. 2019. Deep retrieval-based dialogue systems: A short review. arXiv preprint arXiv:1907.12878.
  6. 6.Samuel Bowman, Luke Vilnis, Oriol Vinyals, Andrew Dai, Rafal Jozefowicz, and Samy Bengio. 2016. Generating sentences from a continuous space. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 10–21.
  7. 7.Wei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, et al. 2021. Contextual fine-to-coarse distillation for coarse-grained response selection in open-domain conversations. arXiv preprint arXiv:2109.13087.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  9. 9.Le Fang, Chunyuan Li, Jianfeng Gao, Wen Dong, and Changyou Chen. 2019. Implicit deep latent variable models for text generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3937–3947.
  10. 10.Michel Galley, Chris Brockett, Xiang Gao, Jianfeng Gao, and Bill Dolan. 2019. Grounded response generation task at dstc7. In AAAI Dialog System Technology Challenges Workshop.
  11. 11.Jun Gao, Wei Bi, Xiaojiang Liu, Junhui Li, and Shuming Shi. 2019. Generating multiple diverse responses for short-text conversation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6383–6390.
  12. 12.Sergey Golovanov, Rauf Kurbanov, Sergey Nikolenko, Kyryl Truskovskyi, Alexander Tselousov, and Thomas Wolf. 2019. Large-scale transfer learning for natural language generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6053–6058.
  13. 13.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020. Spanbert: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77.
  14. 14.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  15. 15.Diederik P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. 2016. Improving variational inference with inverse autoregressive flow. arXiv preprint arXiv:1606.04934.
  16. 16.Helena C Kraemer. 2014. Kappa coefficient. Wiley StatsRef: Statistics Reference Online, pages 1–4.
  17. 17.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  18. 18.Chunyuan Li, Xiang Gao, Yuan Li, Xiujun Li, Baolin Peng, Yizhe Zhang, and Jianfeng Gao. 2020. Optimus: Organizing sentences via pre-trained modeling of a latent space. arXiv preprint arXiv:2004.04092.
  19. 19.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and William B Dolan. 2016a. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119.
  20. 20.Jiwei Li, Will Monroe, and Dan Jurafsky. 2016b. A simple, fast diverse decoding algorithm for neural generation. arXiv preprint arXiv:1611.08562.
  21. 21.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 986–995.
  22. 22.Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2020. Exploring versatile generative language model via parameter-efficient transfer learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, pages 441–459.
  23. 23.Shuman Liu, Hongshen Chen, Zhaochun Ren, Yang Feng, Qun Liu, and Dawei Yin. 2018. Knowledge diffusion for neural dialogue generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1489–1498.
  24. 24.Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025.
  25. 25.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural discrete representation learning. arXiv preprint arXiv:1711.00937.
  26. 26.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038.
  27. 27.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  28. 28.Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020. Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, pages 2401–2410.
  29. 29.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  30. 30.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  32. 32.Ramon Sanabria, Shruti Palaskar, and Florian Metze. 2019. Cmu sinbad’s submission for the dstc7 avsd challenge.
  33. 33.Iulian Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2017. A hierarchical latent variable encoder-decoder model for generating dialogues. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31.
  34. 34.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155.
  35. 35.Arash Vahdat, William Macready, Zhengbing Bian, Amir Khoshaman, and Evgeny Andriyash. 2018. Dvae++: Discrete variational autoencoders with overlapping transformations. In International Conference on Machine Learning, pages 5035–5044. PMLR.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762.
  37. 37.Oriol Vinyals and Quoc Le. 2015. A neural conversational model. arXiv preprint arXiv:1506.05869.
  38. 38.Sixing Wu, Ying Li, Dawei Zhang, Yang Zhou, and Zhonghai Wu. 2020. Diverse and informative dialogue generation with context-specific commonsense knowledge awareness. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 5811–5820.
  39. 39.Dongling Xiao, Han Zhang, Yukun Li, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020. Ernie-gen: An enhanced multi-flow pre-training and fine-tuning framework for natural language generation. arXiv preprint arXiv:2001.11314.
  40. 40.Can Xu, Wei Wu, Chongyang Tao, Huang Hu, Matt Schuerman, and Ying Wang. 2019. Neural response generation with meta-words. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5416–5426, Florence, Italy. Association for Computational Linguistics.
  41. 41.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In ACL (1).
  42. 42.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. Dialogpt: Large-scale generative pre-training for conversational response generation. In ACL, system demonstration.
  43. 43.Tiancheng Zhao, Kyusong Lee, and Maxine Eskenazi. 2018. Unsupervised discrete sentence representation learning for interpretable neural dialog generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1098–1107, Melbourne, Australia. Association for Computational Linguistics.
  44. 44.Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi. 2017a. Learning discourse-level diversity for neural dialog models using conditional variational autoencoders. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 654–664, Vancouver, Canada. Association for Computational Linguistics.
  45. 45.Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi. 2017b. Learning discourse-level diversity for neural dialog models using conditional variational autoencoders. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 654–664.
  46. 46.Hao Zhou, Tom Young, Minlie Huang, Haizhou Zhao, Jingfang Xu, and Xiaoyan Zhu. 2018. Commonsense knowledge aware conversation generation with graph attention. In IJCAI, pages 4623–4629.

Citation

MLA
Chen, W., et al. “DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 4852–64, https://doi.org/10.18653/v1/2022.acl-long.333.
APA
Chen, W., Gong, Y., Wang, S., Yao, B., Qi, W., (魏忠钰), Z. W., Hu, X., Zhou, B., Mao, Y., Chen, W., Cheng, B., & Duan, N. (2022). DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4852–4864. https://doi.org/10.18653/v1/2022.acl-long.333
Chicago
Chen, W., Y. Gong, S. Wang, et al. 2022. “DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4852–64. https://doi.org/10.18653/v1/2022.acl-long.333.
Harvard
Chen, W. et al. (2022) “DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4852–4864. Available at: https://doi.org/10.18653/v1/2022.acl-long.333.
Vancouver
1. Chen W, Gong Y, Wang S, et al (2022) DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4852–4864

BibTeX

@inproceedings{chen-etal-2022-dialogved,
    title = "{D}ialog{VED}: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation",
    author = "Chen, Wei  and
      Gong, Yeyun  and
      Wang, Song  and
      Yao, Bolun  and
      Qi, Weizhen  and
      Wei, Zhongyu  and
      Hu, Xiaowu  and
      Zhou, Bartuer  and
      Mao, Yi  and
      Chen, Weizhu  and
      Cheng, Biao  and
      Duan, Nan",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.333/",
    doi = "10.18653/v1/2022.acl-long.333",
    pages = "4852--4864"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/