Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation

Yu LiBaolin PengYelong ShenYi MaoLars LidenZhou YuJianfeng Gao

article2022NAACL62 citations

Proposes PLUG, a pre-trained dialogue model that converts heterogeneous knowledge sources into a unified text format to improve conversational generalization across diverse domains, especially in zero-shot and few-shot settings.

Listen

Conversational artificial intelligence systems often struggle to engage in informative, factual discussions because they lack specialized knowledge and produce overly generic responses. While knowledge-grounded dialogue systems aim to address this by incorporating external information, existing models face major practical hurdles. Collecting high-quality conversational datasets is costly and covers only a tiny fraction of target domains. Furthermore, existing systems are typically tailored to specific knowledge structures, such as Wikipedia passages or recommendation graphs, making them difficult to adapt across different tasks and unseen topics.

The article demonstrates and evaluates a unified conversational language model framework, called PLUG, designed to bridge heterogeneous knowledge sources and generate factual responses across diverse dialogue tasks. The primary objective is to prove that transforming disparate knowledge formats into a standardized textual representation enables a single pre-trained model to generalize effectively in both fully supervised and data-scarce environments.

To achieve this, the authors converted various external data structures—including encyclopedic passages, databases, and knowledge graphs—into uniform textual triples and keywords. These extracted facts are directly concatenated with dialogue history and fed into an 800-million-parameter encoder-decoder language model. Credibility was established by pre-training the system on a newly curated corpus of over 321,000 filtered Reddit conversation turns mapped to knowledge graph facts, combined with the OpenDialKG conversational dataset. The model was subsequently tested across two benchmark environments representing open-domain conversation and conversational recommendation, evaluating performance across fully supervised, few-shot (10 to 500 training dialogues), and zero-shot settings.

The evaluation produced four central findings. First, the unified model achieved state-of-the-art or competitive performance across both benchmarks in standard fully supervised setups, confirming that representing diverse knowledge as unified text is highly viable. Second, the framework significantly outperformed standard baseline models in zero-shot and few-shot conditions; for instance, when trained on as few as 50 conversations, it generated responses with high human-rated fluency and coherence. Third, human evaluators consistently rated the unified system higher in factual knowledge integration and conversational flow compared to retrieval-augmented baselines. Fourth, experiments providing the model with perfectly retrieved reference knowledge yielded massive performance gains—such as boosting recommendation recall from around 5% to over 84%—revealing that the accuracy of upstream knowledge retrieval, rather than response generation, is the primary performance bottleneck.

These findings indicate that conversational AI development can shift away from building fragmented, domain-specific architectures toward unified language models that ingest standardized text knowledge. This approach substantially lowers the timeline and resource costs associated with collecting large task-specific training sets, making rapid deployment in low-data domains feasible. Moreover, the strong linear correlation between dialogue performance and information retrieval precision implies that future investment should prioritize robust search and retrieval components rather than merely scaling generation models.

Based on these results, engineering teams and decision-makers should adopt unified textual representations when integrating external databases or knowledge graphs into conversational systems. For new domains, organizations can rely on few-shot fine-tuning rather than expensive large-scale data annotation. Next steps should focus on upgrading upstream retrieval algorithms and expanding pre-training across a wider variety of structured and unstructured data sources, including question-answering formats.

The primary limitation of this work stems from the reliance on external search mechanisms, as retrieval errors directly constrain output quality regardless of the language model's generative capacity. Additionally, because the pre-training data was harvested from public online forums, there remains a risk of generating unvetted or biased language. Users should exercise caution when deploying the framework in safety-critical domains until enhanced safety filtering and more reliable information retrieval pipelines are integrated.

arXiv: 2112.07924
Cover for Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation

Abstract

Knowledge-grounded dialogue systems are challenging to build due to the lack of training data and heterogeneous knowledge sources. Existing systems perform poorly on unseen topics due to limited topics covered in the training data. In addition, it is challenging to generalize to the domains that require different types of knowledge sources. To address the above challenges, we present PLUG¹, a language model that homogenizes different knowledge sources to a unified knowledge representation for knowledge-grounded dialogue generation tasks. We first retrieve relevant information from heterogeneous knowledge sources (e.g., wiki, dictionary, or knowledge graph); Then the retrieved knowledge is transformed into text and concatenated with dialogue history to feed into the language model for generating responses. PLUG is pre-trained on a large-scale knowledge-grounded dialogue corpus. The empirical evaluation on two benchmarks shows that PLUG generalizes well across different knowledge-grounded dialogue tasks. It achieves comparable performance with state-of-the-art methods in the fully-supervised setting and significantly outperforms other approaches in zero-shot and few-shot settings.

Table of Contents

  • 1 Introduction
  • 2 Approach
  • 2.1 Background: Knowledge-Grounded Pre-training
  • 2.2 PLUG
  • 2.3 Model training process
  • 2.3.1 Reddit Conversation
  • 2.3.2 OpenDialKG
  • 3 Experiments
  • 3.1 Datasets and Knowledge Sources
  • 3.2 Baselines
  • 3.3 Metrics
  • 3.4 Fully-Supervised Results
  • 3.5 Zero-Shot and Few-Shot Results
  • 3.6 Human Evaluation
  • 3.7 Discussion and Analysis
  • 4 Related Work
  • 5 Conclusion and Future Work
  • References
  • A Implementation Details
  • B Ethical Considerations
  • C Human Evaluation Interface

Knowls

  1. Knowl 1 — PLUG’s unified knowledge representation

    model/method

    PLUG (Pre-trained Language model with a Unified knowledge representation for knowledge-Grounded dialogues) converts heterogeneous external knowledge sources—such as documents, databases, and knowledge graphs—into a common text representation before response generation. A task-specific retriever selects essential knowledge, usually expressed as serialized triples or keywords. The selected knowledge is concatenated with the dialogue context and processed by one T5 encoder-decoder, so the same model architecture can be transferred across knowledge-grounded dialogue tasks without a source-specific knowledge encoder or graph module.

    The T5 decoder generates the response conditioned on both the dialogue context and the retrieved knowledge. PLUG is additionally pre-trained on knowledge-grounded conversations from multiple sources so that it learns both how knowledge is expressed in responses and how to transfer this behavior to downstream tasks.

  2. Knowl 2 — Conditional response-generation objective

    equation

    For a dialogue turn indexed by ii, let CiC_i be the dialogue-context token sequence, let KiK_i be the token sequence containing the selected essential knowledge, and let Ri=(ri,1,…,ri,Ti)R_i=(r_{i,1},\ldots,r_{i,T_i}) be the response of length TiT_i tokens. PLUG models the response distribution with an autoregressive sequence-to-sequence objective:

    pθ(Ri∣Ci,Ki)=∏t=1Tipθ(ri,t∣Ci,Ki,ri,1,…,ri,t−1), p_\theta(R_i\mid C_i,K_i)=\prod_{t=1}^{T_i}p_\theta\left(r_{i,t}\mid C_i,K_i,r_{i,1},\ldots,r_{i,t-1}\right),

    where pθp_\theta is the T5-based PLUG model parameterized by θ\theta, ri,tr_{i,t} is the response token at position tt, and the product is taken over all response tokens. The essential knowledge KiK_i is extracted from the task’s external source and serialized as text, allowing knowledge-grounded generation to use the same text-to-text objective regardless of the original source format.

  3. Knowl 3 — Hierarchical extraction of Reddit knowledge groundings

    model/method

    To create knowledge labels for Reddit conversations that contain Wikipedia grounding but no turn-level knowledge annotations, PLUG uses a three-stage extraction pipeline. First, it obtains the title of the Wikipedia passage associated with a conversation and queries DBpedia for 500 triples whose subject or object occurs in that passage. Second, it concatenates the dialogue history and target response into a query, computes TF-IDF representations, ranks the 500 triples by cosine similarity to the query, and retains the top 50. Third, it reranks those 50 triples using Sentence-BERT semantic similarity to the response and discards the dialogue turn when its highest semantic-similarity score is below 0.350.35. The remaining highest-ranked triples provide essential knowledge for pre-training.

    This pipeline is intended both to remove irrelevant DBpedia facts and to exclude Reddit responses that are not genuinely grounded in the associated Wikipedia passage.

  4. Knowl 4 — Multi-source pre-training regimen

    experimental setup

    PLUG is pre-trained on two knowledge-grounded dialogue sources. From Reddit Conversation, the authors process Reddit submissions and comments from 2011–2017, initially obtaining more than 894,000 Wikipedia-grounded dialogue turns and retaining more than 321,000 turns after hierarchical knowledge extraction. OpenDialKG contributes all of its dialogue turns; unlike Reddit Conversation, it already supplies a knowledge-graph path for each dialogue and a labeled grounding triple for each turn, so the labeled triple is used directly as essential knowledge.

    Each pre-training example is serialized into three segments: dialogue context, essential knowledge, and response. The model keeps the top three triples or keywords as essential knowledge, initializes PLUG from the 800-million-parameter T5 model, truncates examples to a maximum of 512 tokens, and optimizes with Adam with weight decay on eight NVIDIA V100 GPUs. Training stops when validation performance stops improving or after 20 epochs, with hyperparameters selected by cross-validated grid search.

  5. Knowl 5 — Cross-task downstream knowledge construction

    experimental setup

    PLUG is evaluated on Wizard of Wikipedia (WoW) and REDIAL, which use different knowledge formats. WoW contains 18,430 training conversations, 981 seen-topic validation conversations, 967 unseen-topic validation conversations, 965 seen-topic test conversations, and 968 unseen-topic test conversations. Its essential knowledge is constructed by retrieving the five highest-ranked passages with a TF-IDF retriever and extracting the three highest-ranked Open Information Extraction triples from those passages. The seen test topics occur in training data, whereas unseen test topics do not occur in the training or validation data.

    REDIAL contains 8,004 training conversations and 1,001 conversations each for validation and test, involving 6,924 movies and 51,699 movie slots. The experiments construct essential knowledge in three ways: similar movies retrieved from DBpedia and serialized as triples; keywords extracted from MovieLens comments for movies mentioned in the context; or the output of the KGSF recommender. The principal comparison systems are T5 without PLUG pre-training, retrieval-augmented T5 for WoW, and KBRD and KGSF for REDIAL.

  6. Knowl 6 — Fully supervised performance on Wizard of Wikipedia

    data/table

    With the complete downstream training set, PLUG using retrieved essential knowledge improves most response-quality metrics over T5 with retrieved knowledge and is competitive with retrieval-augmented T5. Oracle or “golden” knowledge produces a much larger improvement, showing the effect of knowledge quality in addition to the generation model. The metrics are BLEU-4 (B4), ROUGE-L (RL), unigram F1 (F1), and knowledge unigram F1 (KF1); all values are reported as in the paper.

    Could not parse LaTeX table

    PLUG with retrieved knowledge obtains B4/RL/F1 scores of 6.0/22.3/26.56.0/22.3/26.5 on seen topics and 3.5/19.5/23.33.5/19.5/23.3 on unseen topics. Replacing retrieved knowledge with golden knowledge raises these scores to 11.5/31.1/36.011.5/31.1/36.0 and 8.8/29.0/33.48.8/29.0/33.4, respectively, while also increasing KF1.

  7. Knowl 7 — Fully supervised performance on REDIAL

    data/table

    On REDIAL, PLUG with the KGSF recommender as its knowledge source gives the strongest reported recommendation recall and diversity among the non-oracle systems. The oracle-knowledge variants are substantially better than all systems using currently available knowledge sources, indicating that retrieval quality limits performance more than response-generation capacity in this setting. B4 is BLEU-4, RL is ROUGE-L, Dist2 and Dist4 are sentence-level distinct-2 and distinct-4, and Rec is recall of the ground-truth recommended movie in the generated response.

    Could not parse LaTeX table

    PLUG + KGSF reaches Dist2 =1.51=1.51, Dist4 =2.84=2.84, and Rec =5.3=5.3, compared with Dist2 =1.13=1.13, Dist4 =2.02=2.02, and Rec =4.7=4.7 for T5-Large + KGSF. With golden knowledge, PLUG reaches Rec =84.3=84.3 and RL =33.5=33.5, far above the retrieved-knowledge variants.

  8. Knowl 8 — Few-shot and zero-shot transfer

    empirical result

    The authors evaluate zero-shot transfer and fine-tuning with 10, 50, or 500 randomly selected training dialogues, then test on the complete WoW and REDIAL test sets. PLUG consistently maintains higher BLEU-4, ROUGE-L, and unigram F1 than T5 without knowledge pre-training when fewer than 500 dialogues are available on both tasks. This demonstrates that pre-training teaches PLUG a transferable pattern for grounding responses in essential knowledge rather than merely improving fully supervised fitting.

    On WoW, an unpretrained T5 model can obtain a higher knowledge F1 in the zero-shot condition than PLUG, but it has poor language-quality scores; the authors interpret this as copying knowledge words while producing incoherent responses. PLUG produces more coherent zero-shot responses even with a lower KF1, so KF1 is not informative in isolation when response quality is very low. On REDIAL, pre-training produces a larger zero-shot improvement in BLEU-4 and ROUGE-L than in recommendation recall. T5’s unusually high zero-shot Dist4 is attributed to diverse but irrelevant responses, whereas PLUG learns to generate grounded and useful responses with much less downstream data.

  9. Knowl 9 — Human evaluation of response quality

    empirical result

    Human evaluation compares responses generated for the same contexts by RAG, T5-Large, and PLUG on Wizard of Wikipedia. The evaluation samples 100 responses per model from seen and unseen test contexts; the few-shot systems are trained on 50 dialogues. Three workers rate every response from 0 to 2 for Fluency, Coherence, and Knowledge. The reported averages are:

    Could not parse LaTeX table

    Here ∗* denotes p<0.05p<0.05 and ∗∗** denotes p<0.01p<0.01 for the reported comparisons against the T5 baselines and RAG. PLUG significantly improves all three criteria over the zero-shot T5 baseline and remains better in the few-shot condition. The few-shot PLUG model is rated slightly more fluent and coherent than the fully supervised PLUG model, while the fully supervised model receives the highest Knowledge score, suggesting that grounding quality continues to improve with more labeled dialogues.

  10. Knowl 10 — Knowledge-source quality is the main performance bottleneck

    limitation

    The paper isolates knowledge-retrieval quality on REDIAL by replacing a proportion of retrieved movie information with golden movie information. The proportion of golden knowledge is varied over 0%0\%, 20%20\%, 40%40\%, 60%60\%, 80%80\%, and 100%100\%, with PLUG and T5 both fine-tuned on 50 dialogues. For both models, BLEU-4, ROUGE-L, and recommendation recall increase approximately linearly as the proportion of golden knowledge rises. PLUG has a steeper improvement for BLEU-4 and recommendation recall, showing greater benefit from better retrieval.

    The Dist4 gap between PLUG and T5 remains nearly constant as golden knowledge increases. T5’s Dist4 falls when no golden knowledge is available, indicating that its few-shot diversity depends on receiving good knowledge during downstream training; PLUG retains diverse generation because this behavior was learned during pre-training. The large difference between retrieved-knowledge and golden-knowledge results identifies the retriever, rather than only the generator, as the principal limitation of the current system. The authors therefore identify improved information retrieval and adding more heterogeneous knowledge sources as key future directions.

Coverage note — No substantial scientific contribution was omitted; background, related work, and the ethical-considerations appendix were excluded because they do not add a separate method, model, theory, or experimental result.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020. Towards a human-like open-domain chatbot. CoRR, abs/2001.09977.
  2. 2.Qibin Chen, Junyang Lin, Yichang Zhang, Ming Ding, Yukuo Cen, Hongxia Yang, and Jie Tang. 2019. Towards knowledge-based recommender dialog system. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1803–1813, Hong Kong, China. Association for Computational Linguistics.
  3. 3.Wenhu Chen, Yu Su, Xifeng Yan, and William Yang Wang. 2020. KGPT: Knowledge-grounded pre-training for data-to-text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8635–8648, Online. Association for Computational Linguistics.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of wikipedia: Knowledge-powered conversational agents. In International Conference on Learning Representations.
  6. 6.Song Feng, Kshitij Fadnis, Q Vera Liao, and Luis A Lastras. 2020. Doc2dial: a framework for dialogue composition grounded in documents. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 13604–13605.
  7. 7.Michel Galley, Chris Brockett, Xiang Gao, Bill Dolan, and Jianfeng Gao. 2018. End-to-end conversation modeling: Moving beyond chitchat.
  8. 8.Jianfeng Gao, Michel Galley, and Lihong Li. 2019. Neural approaches to conversational AI: Question answering, task-oriented dialogues and social chatbots. Now Foundations and Trends.
  9. 9.Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, Bill Dolan, Jianfeng Gao, Wen-tau Yih, and Michel Galley. 2018. A knowledge-grounded neural conversation model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  10. 10.Karthik Gopalakrishnan, Behnam Hedayatnia, Qinglang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tür, and Amazon Alexa AI. 2019. Topical-chat: Towards knowledge-grounded open-domain conversations. In INTERSPEECH, pages 1891–1895.
  11. 11.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909.
  12. 12.Shirley Anugrah Hayati, Dongyeop Kang, Qingxiaoyang Zhu, Weiyan Shi, and Zhou Yu. 2020. Inspired: Toward sociable recommendation dialog systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8142–8152, Online. Association for Computational Linguistics.
  13. 13.Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020. A simple language model for task-oriented dialogue. In Advances in Neural Information Processing Systems, volume 33, pages 20179–20191. Curran Associates, Inc.
  14. 14.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning.
  15. 15.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  16. 16.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, et al. 2015. Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web, 6(2):167–195.
  17. 17.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  18. 18.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020b. Retrieval-augmented generation for knowledge-intensive nlp tasks. arXiv preprint arXiv:2005.11401.
  19. 19.Linxiao Li, Can Xu, Wei Wu, YUFAN ZHAO, Xueliang Zhao, and Chongyang Tao. 2020. Zero-resource knowledge-grounded dialogue generation. In Advances in Neural Information Processing Systems, volume 33, pages 8475–8485. Curran Associates, Inc.
  20. 20.Raymond Li, Samira Ebrahimi Kahou, Hannes Schulz, Vincent Michalski, Laurent Charlin, and Chris Pal. 2018. Towards deep conversational recommendations. In Advances in Neural Information Processing Systems 31 (NIPS 2018).
  21. 21.Yu Li, Shirley Anugrah Hayati, Weiyan Shi, and Zhou Yu. 2021. DEUX: an attribute-guided framework for sociable recommendation dialog systems. CoRR, abs/2105.00825.
  22. 22.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  23. 23.Shilei Liu, Xiaofeng Zhao, Bochao Li, Feiliang Ren, Longhui Zhang, and Shujuan Yin. 2021a. A three-stage learning framework for low-resource knowledge-grounded dialogue generation. arXiv preprint arXiv:2109.04096.
  24. 24.Zeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu, and Wanxiang Che. 2021b. DuRecDial 2.0: A bilingual parallel corpus for conversational recommendation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4335–4347, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Kaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg, and Jianfeng Gao. 2021. Open domain question answering over virtual documents: A unified approach for data and text. arXiv preprint arXiv:2110.08417.
  26. 26.Andrea Madotto, Zhaojiang Lin, Yejin Bang, and Pascale Fung. 2020. The adapter-bot: All-in-one controllable conversational model. arXiv preprint arXiv:2008.12579.
  27. 27.Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. OpenDialKG: Explainable conversational reasoning with attention-based walks over knowledge graphs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 845–854, Florence, Italy. Association for Computational Linguistics.
  28. 28.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  29. 29.Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2020. Soloist: Few-shot task-oriented dialog with a single pre-trained auto-regressive model. arXiv preprint arXiv:2005.05298.
  30. 30.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  31. 31.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. CoRR, abs/1910.10683.
  33. 33.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  34. 34.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  35. 35.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  36. 36.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2020. Recipes for building an open-domain chatbot. CoRR, abs/2004.13637.
  37. 37.Corby Rosset, Chenyan Xiong, Minh Phan, Xia Song, Paul Bennett, and Saurabh Tiwary. 2020. Knowledge-aware language model pretraining.
  38. 38.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation.
  39. 39.Yi-Lin Tuan, Yun-Nung Chen, and Hung-yi Lee. 2019. Dykgchat: Benchmarking dialogue generation grounding on dynamic knowledge graphs. arXiv preprint arXiv:1910.00610.
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  41. 41.Lingzhi Wang, Huang Hu, Lei Sha, Can Xu, Kam-Fai Wong, and Daxin Jiang. 2021. Finetuning large-scale pre-trained language models for conversational recommendation with knowledge graph. arXiv preprint arXiv:2110.07477.
  42. 42.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  43. 43.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32.
  44. 44.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019. Dialogpt: Large-scale generative pre-training for conversational response generation. CoRR, abs/1911.00536.
  45. 45.Xueliang Zhao, Wei Wu, Chongyang Tao, Can Xu, Dongyan Zhao, and Rui Yan. 2020a. Low-resource knowledge-grounded dialogue generation. arXiv preprint arXiv:2002.10348.
  46. 46.Xueliang Zhao, Wei Wu, Can Xu, Chongyang Tao, Dongyan Zhao, and Rui Yan. 2020b. Knowledge-grounded dialogue generation with pre-trained language models. CoRR, abs/2010.08824.
  47. 47.Kangyan Zhou, Shrimai Prabhumoye, and Alan W Black. 2018. A dataset for document grounded conversations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
  48. 48.Kun Zhou, Xiaolei Wang, Yuanhang Zhou, Chenzhan Shang, Yuan Cheng, Wayne Xin Zhao, Yaliang Li, and Ji-Rong Wen. 2021. CRSLab: An open-source toolkit for building conversational recommender system. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 185–193, Online. Association for Computational Linguistics.
  49. 49.Kun Zhou, Wayne Xin Zhao, Shuqing Bian, Yuanhang Zhou, Ji-Rong Wen, and Jingsong Yu. 2020a. Improving conversational recommender systems via knowledge graph based semantic fusion. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1006–1014.
  50. 50.Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020b. The design and implementation of xiaoice, an empathetic social chatbot. Computational Linguistics, 46(1):53–93.
  51. 51.Wenya Zhu, Kaixiang Mo, Yu Zhang, Zhangbin Zhu, Xuezheng Peng, and Qiang Yang. 2017. Flexible end-to-end dialogue system for knowledge grounded conversation. arXiv preprint arXiv:1709.04264.

Citation

MLA
Li, Y., et al. “Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 206–18, https://doi.org/10.18653/v1/2022.naacl-main.15.
APA
Li, Y., Peng, B., Shen, Y., Mao, Y., Liden, L., Yu, Z., & Gao, J. (2022). Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 206–218. https://doi.org/10.18653/v1/2022.naacl-main.15
Chicago
Li, Y., B. Peng, Y. Shen, et al. 2022. “Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 206–18. https://doi.org/10.18653/v1/2022.naacl-main.15.
Harvard
Li, Y. et al. (2022) “Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 206–218. Available at: https://doi.org/10.18653/v1/2022.naacl-main.15.
Vancouver
1. Li Y, Peng B, Shen Y, Mao Y, Liden L, Yu Z, Gao J (2022) Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 206–218

BibTeX

@inproceedings{li-etal-2022-knowledge,
    title = "Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation",
    author = "Li, Yu  and
      Peng, Baolin  and
      Shen, Yelong  and
      Mao, Yi  and
      Liden, Lars  and
      Yu, Zhou  and
      Gao, Jianfeng",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.15/",
    doi = "10.18653/v1/2022.naacl-main.15",
    pages = "206--218"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/