CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning

Xiangru TangArjun NairBorui WangBingyao WangJai DesaiAaron WadeHaoran LiAsli CelikyilmazYashar MehdadDragomir R. Radev

article2022NAACL73 citations

Introduces a linguistically motivated taxonomy of factual errors in abstractive dialogue summarization alongside CONFIT, a contrastive fine-tuning method with targeted hard negative samples that substantially reduces model hallucinations on the SAMSum and AMI benchmarks.

Listen

Automated abstractive dialogue summarization is increasingly vital for processing conversational data, yet modern pretrained neural language models frequently generate hallucinated or factually inconsistent content. In multi-party conversations, challenges such as informal phrasing, conversational flow, and shifting first-person references cause standard models to misattribute statements or alter critical details. These factual errors pose substantial risks in operational environments where accuracy and reliability are mandatory for decision-making.

The article demonstrates a linguistically informed contrastive fine-tuning method, named CONFIT, designed to improve the factual consistency and overall quality of abstractive dialogue summarization. The primary objective is to evaluate whether targeted training losses can systematically eliminate the most frequent factual errors produced by standard language models across dialogue and meeting benchmarks.

To establish this method, the authors first created an eight-category taxonomy of factual errors and conducted human evaluations on baseline model outputs to identify major failure modes. Based on these findings, CONFIT augments standard training objectives with two supplementary mechanisms: a contrastive loss using carefully perturbed negative samples (such as swapped entities, altered verbs, masked numbers, and dropped sentences) and a self-supervised auxiliary loss that tracks speaker identities and resolves pronoun references. The approach was tested across standard architectures (BART, Pegasus, and T5) on two representative datasets: the SAMSum chat dialogue benchmark (over 16,000 dialogues) and the AMI meeting corpus (137 meeting transcripts).

The evaluation yielded several key findings. First, error analysis revealed that missing information and wrong references account for approximately 45% of all factual errors in baseline dialogue summaries. Second, applying CONFIT significantly improved human-rated faithfulness across all evaluated architectures; for example, BART faithfulness scores rose from 5.54 to 7.25 on SAMSum and from 4.85 to 5.60 on AMI on a 10-point scale. Third, the framework drastically reduced specific error rates on SAMSum, achieving wrong-reference error reductions of 20 percentage points for BART (from 37% to 17%) and 33 percentage points for T5 (from 46% to 13%). Finally, CONFIT consistently boosted standard automated overlap metrics, achieving superior ROUGE-1 and ROUGE-L scores across both short chat dialogues and long-form meeting transcripts.

These results indicate that conversational factuality errors stem from distinct structural dynamics—particularly speaker tracking and coreference—rather than general language modeling deficiencies alone. By explicitly penalizing known error types and enforcing speaker awareness during fine-tuning, organizations can deploy automated conversational summarizers with substantially reduced risk of misattribution and factual distortion, lowering the cost and effort of manual verification.

Organizations developing or deploying conversational artificial intelligence systems should consider incorporating contrastive and speaker-aware fine-tuning objectives into their training pipelines rather than relying solely on standard cross-entropy objectives. For practical implementation, engineering teams should validate automated evaluation metrics against human review, as automated scoring tools like BARTScore showed inconsistencies on long-form meeting data despite confirmed human-rated improvements.

The findings carry high confidence for standard dialogue and chat domains, supported by rigorous human evaluation. However, confidence should be tempered when applying the technique directly to complex, long-form multi-party meetings; the human evaluation on the AMI dataset was limited to 20 dialogues and exhibited higher baseline error rates (such as 70% to 85% missing information), indicating that longer conversational contexts require broader empirical validation and further structural modeling.

arXiv: 2112.08713
Cover for CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning

Abstract

Factual inconsistencies in generated summaries severely limit the practical applications of abstractive dialogue summarization. Although significant progress has been achieved by using pre-trained neural language models, substantial amounts of hallucinated content are found during the human evaluation. In this work, we first devised a typology of factual errors to better understand the types of hallucinations generated by current models and conducted human evaluation on popular dialog summarization dataset. We further propose a training strategy that improves the factual consistency and overall quality of summaries via a novel contrastive fine-tuning, called CONFIT. To tackle top factual errors from our annotation, we introduce additional contrastive loss with carefully designed hard negative samples and self-supervised dialogue-specific loss to capture the key information between speakers. We show that our model significantly reduces all kinds of factual errors on both SAMSum dialogue summarization and AMI meeting summarization. On both datasets, we achieve significant improvements over state-of-the-art baselines using both automatic metrics, ROUGE and BARTScore, and human evaluation.

Table of Contents

  • 1 Introduction
  • 2 New Taxonomy of Factuality Errors for Abstractive Dialogue Summarization
  • 2.1 Annotation and Analysis
  • 3 CONFIT Model
  • 3.1 Contrastive Loss
  • 3.2 Self-supervised Loss
  • 4 Experiments
  • 4.1 Dataset
  • 4.2 Experiment Settings
  • 4.3 Evaluation Metrics
  • 5 Results
  • 5.1 Error Analysis
  • 5.2 Case Study
  • 6 Related Work
  • 7 Conclusion
  • 8 Ethics Statement
  • References

Knowls

  1. Knowl 1 — Taxonomy of Factuality Errors for Abstractive Dialogue Summarization

    definition

    In abstractive dialogue summarization, factual inconsistencies between the source dialogue and the generated summary are categorized into eight distinct error types:

    1. Missing Information: The generated summary omits critical content present in the reference or source dialogue.
    2. Redundant Information: The generated summary includes unneeded or extra content not reflected in the salient points of the reference.
    3. Circumstantial Error: Circumstantial details (e.g., date, time, numerical figures, location) modifying a predicate fail to match the source context.
    4. Wrong Reference Error: An incorrect or nonexistent antecedent is assigned to a pronoun, or a personal named entity in the summary is substituted for another entity in the dialogue.
    5. Negation Error: An erroneous introduction or omission of negation in the summary alters the polarity of a factual statement.
    6. Object Error: The summary specifies an incorrect direct or indirect non-personal object for an action (personal entity substitutions are categorized under Wrong Reference Error).
    7. Tense Error: Discrepancies exist between the grammatical tense of the generated summary and the temporal state of the source dialogue events.
    8. Modality Error: Modal expressions (e.g., may, should, could, certainty levels) in the summary do not align with the epistemic status in the source dialogue.
  2. Knowl 2 — CONFIT Training Objective

    equation

    The CONFIT (Contrastive Fine-Tuning) framework optimizes abstractive dialogue summarization models by supplementing standard cross-entropy generation loss with a contrastive loss and a self-supervised speaker-tracking loss. The overall loss objective J(θ)\mathcal{J}(\theta) parameterized by model weights θ\theta is:

    J(θ)=L+αLcon+βLself\mathcal{J}(\theta) = \mathcal{L} + \alpha \mathcal{L}_{\text{con}} + \beta \mathcal{L}_{\text{self}}

    where:

    • L=−∑llog⁡P(t~l∣t<l,D)\mathcal{L} = -\sum_l \log P(\tilde{t}_l \mid t_{<l}, D) denotes standard autoregressive cross-entropy loss maximizing the generation probability of reference target token t~l\tilde{t}_l given prefix t<lt_{<l} and dialogue context DD.
    • Lcon\mathcal{L}_{\text{con}} is the contrastive loss function operating on positive and negative summary representations.
    • Lself\mathcal{L}_{\text{self}} is the self-supervised role/speaker consistency loss on dialogue tokens.
    • α\alpha and β\beta are non-negative weighting hyperparameters balancing the contribution of the contrastive and self-supervised objectives.
  3. Knowl 3 — Contrastive Loss Function in CONFIT

    equation

    The contrastive loss objective in CONFIT optimizes encoder representations by pulling together positive summary pairs while pushing away linguistically targeted hard negative summaries. For a given positive pair of summaries (yi,yj)(y_i, y_j) (where yiy_i is the reference summary and yjy_j is generated via back-translation) and a set of negative samples {yk}\{y_k\}, the loss is defined as:

    Lcon=−∑yj≠yilog⁡exp⁡(cos⁡(ci,cj))∑yk≠yiexp⁡(cos⁡(ci,ck))\mathcal{L}_{\text{con}} = -\sum_{y_j \neq y_i} \log \frac{\exp(\cos(c_i, c_j))}{\sum_{y_k \neq y_i} \exp(\cos(c_i, c_k))}

    where cos⁡(⋅,⋅)\cos(\cdot, \cdot) denotes cosine similarity, and ci,cj,ckc_i, c_j, c_k represent the encoder hidden representation vectors (e.g., from a BART encoder) corresponding to the texts yi,yj,yky_i, y_j, y_k.

  4. Knowl 4 — Linguistically-Informed Negative Sample Generation Strategies

    model/method

    To train the contrastive objective in CONFIT to prevent specific factual errors, five distinct perturbation strategies are used to create hard negative summary samples:

    • Noun Swapping: Randomly swaps noun tokens within the reference summary to construct negative samples targeting wrong reference and object errors.
    • Verb Swapping: Randomly swaps verb tokens within the reference summary to construct negative samples targeting circumstantial, tense, and modality errors.
    • Number and Date Masking: Masks numerical entities, years, and dates in the source dialogue before passing it to the summarization model to produce negative summaries targeting circumstantial errors.
    • Dialogue Sentence Deletion: Randomly deletes 30%30\% of the turns/sentences in the input dialogue before generating a summary to produce negative samples targeting missing information errors.
    • Coreferent Entity Perturbation: Masks and infills coreferent entity spans in the source dialogue using a pre-trained language model (BART) before generation to produce negative samples targeting wrong reference errors.
  5. Knowl 5 — Self-Supervised Speaker-Aware Loss Function

    equation

    To prevent wrong reference errors resulting from ambiguous first-person pronoun usage (such as I or we) across conversational turns, CONFIT uses a self-supervised classification objective to predict whether dialogue units originate from the same speaker. Given dialogue encoder representations CC, kk token or utterance representation pairs (tm,tn)(t_m, t_n) are sampled from the dialogue, with ground-truth speaker identity indicators sms_m and sns_n. The classification loss is formulated as:

    Lself=−∑m=1k∑n=1klog⁡P(sm=sn∣tm,tn,C)\mathcal{L}_{\text{self}} = -\sum_{m=1}^k \sum_{n=1}^k \log P(s_m = s_n \mid t_m, t_n, C)

    where P(sm=sn∣tm,tn,C)P(s_m = s_n \mid t_m, t_n, C) represents the predicted probability that representations tmt_m and tnt_n share the same speaker label given the context CC.

  6. Knowl 6 — Experimental Configurations for Dialogue and Meeting Summarization

    experimental setup

    CONFIT is evaluated on two distinct conversational summarization benchmarks:

    1. SAMSum: A corpus of 16,369 English chat dialogues written by linguists with reference summaries (14,732 training, 818 validation, 819 test). Average dialogue properties: 2.402.40 speakers, 11.1711.17 turns, and 23.4423.44 average turn length.
    2. AMI Meeting Corpus: 137 multiparty meeting transcripts spanning 100 hours of audio recordings, annotated with abstractive meeting summaries. Average dialogue properties: 44 speakers, 289289 turns, and 322322 average turn length.

    Training Hyperparameters:

    • SAMSum: BART fine-tuned for 3 epochs with learning rate 1×10−51\times 10^{-5}; Pegasus fine-tuned for 20 epochs with learning rate 1×10−41\times 10^{-4}; T5 fine-tuned for 20 epochs with learning rate 1×10−51\times 10^{-5}.
    • AMI: BART fine-tuned for 6,000 steps with learning rate 1×10−51\times 10^{-5}; Pegasus fine-tuned for 24,000 steps with learning rate 1×10−51\times 10^{-5}; T5 fine-tuned for 20,000 steps with learning rate 1×10−51\times 10^{-5}.
  7. Knowl 7 — Automatic ROUGE Performance of CONFIT across Models and Datasets

    data/table

    Applying CONFIT contrastive fine-tuning to T5, Pegasus, and BART consistently improves ROUGE-1, ROUGE-2, and ROUGE-L scores over standard cross-entropy fine-tuning on both SAMSum and AMI datasets:

    Model AMI SAMSum
    R-1 R-2 R-L R-1 R-2 R-L
    TextRank 35.19 6.13 15.70 29.27 8.02 28.78
    Fast Abs RL 38.76 15.13 35.18 40.96 17.18 39.05
    PGN 48.34 16.02 23.49 40.08 15.28 36.63
    PGN (DALL) 50.91 17.75 24.59 - - -
    Multi-view BART - - - 49.52 26.52 48.29
    T5 42.16 13.94 39.39 48.41 24.79 44.61
    T5-ConFiT 47.18 13.19 43.55 52.13 27.12 47.62
    Pegasus 46.02 15.85 43.73 48.04 22.94 43.40
    Pegasus-ConFiT 48.47 17.61 45.75 52.65 28.21 48.15
    BART 47.92 16.00 45.36 51.74 26.46 48.72
    BART-ConFiT 50.31 17.29 47.98 53.89 28.85 49.29

    On SAMSum, BART-ConFiT achieves the highest overall scores (53.8953.89 R-1, 28.8528.85 R-2, 49.2949.29 R-L), representing substantial improvements over vanilla BART (51.7451.74 R-1, 26.4626.46 R-2, 48.7248.72 R-L).

  8. Knowl 8 — Human Faithfulness and BARTScore Evaluation of CONFIT

    data/table

    Human evaluation was conducted on 100 SAMSum dialogues and 20 AMI dialogues by human evaluators rating summary faithfulness on a 1–10 Likert scale in a blinded protocol. In addition, model outputs were scored using BARTScore.

    Model Human Faithfulness (1–10) BARTScore
    SAMSum AMI SAMSum AMI
    BART 5.540 4.850 -1.613 -3.644
    BART-ConFiT 7.250 5.600 -1.468 -3.669
    Pegasus 6.260 5.250 -1.615 -2.967
    Pegasus-ConFiT 6.770 5.895 -1.608 -3.369
    T5 5.422 4.150 -1.993 -3.406
    T5-ConFiT 6.920 4.950 -1.677 -3.798

    CONFIT fine-tuning improves human-judged faithfulness scores across all tested pre-trained architectures on both short chat dialogues (SAMSum) and long meeting transcripts (AMI).

  9. Knowl 9 — Factual Error Rate Reduction across Models on SAMSum

    empirical result

    Human annotation of error categories on 100 SAMSum dialogues reveals that CONFIT fine-tuning reduces error rates across the majority of error types compared to standard baseline fine-tuning:

    • Wrong Reference Errors: Reduced from 37%37\% to 17%17\% in BART (a 20%20\% absolute reduction), from 25%25\% to 18%18\% in Pegasus (a 7%7\% absolute reduction), and from 46%46\% to 13%13\% in T5 (a 33%33\% absolute reduction).
    • Missing Information Errors: Reduced from 55%55\% to 44%44\% in BART, from 56%56\% to 50%50\% in Pegasus, and from 63%63\% to 48%48\% in T5.
    • Circumstantial Errors: Reduced from 14%14\% to 8%8\% in BART, from 16%16\% to 10%10\% in Pegasus, while remaining roughly stable for T5 (8%8\% vs. 9%9\%).
    • Redundant Information Errors: Reduced from 12%12\% to 7%7\% in BART, from 7%7\% to 4%4\% in Pegasus, and from 7%7\% to 4%4\% in T5.

    The large reduction in wrong reference errors is primarily attributed to the self-supervised speaker-tracking objective and coreference-targeted contrastive samples.

  10. Knowl 10 — BARTScore Evaluation Discrepancy on Long Meeting Transcripts

    limitation

    While CONFIT improves BARTScore on the SAMSum dataset across all base models (e.g., BART improves from −1.613-1.613 to −1.468-1.468, T5 improves from −1.993-1.993 to −1.677-1.677), BARTScore values decrease for all models fine-tuned with CONFIT on the AMI meeting corpus (e.g., BART changes from −3.644-3.644 to −3.669-3.669, Pegasus from −2.967-2.967 to −3.369-3.369, and T5 from −3.406-3.406 to −3.798-3.798). Because blinded human evaluation indicates consistent gains in summary faithfulness on AMI for all three CONFIT models over their baselines, this decline indicates an automated metric imperfection in capturing true faithfulness on long-form, meeting-style dialogue transcripts.

Coverage note — None was omitted.

References

  1. 1.Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li. 2018. Faithful to the original: Fact aware neural abstractive summarization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  2. 2.Jiaao Chen and Diyi Yang. 2020. Multi-view sequence-to-sequence models with conversational structure for abstractive dialogue summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4106–4118, Online. Association for Computational Linguistics.
  3. 3.Jiaao Chen and Diyi Yang. 2021a. Simple conversational data augmentation for semi-supervised abstractive dialogue summarization. In Proceedings of the Association for Computational Linguistics: EMNLP 2020, Online. Association for Computational Linguistics.
  4. 4.Jiaao Chen and Diyi Yang. 2021b. Structure-aware abstractive conversation summarization via discourse and action graphs. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1380–1391, Online. Association for Computational Linguistics.
  5. 5.Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel. 2021a. Summscreen: A dataset for abstractive screenplay summarization.
  6. 6.Wang Chen, Piji Li, Hou Pong Chan, and Irwin King. 2021b. Dialogue summarization with supporting utterance flow modelling and fact regularization. Knowledge-Based Systems, 229:107328.
  7. 7.Yen-Chun Chen and Mohit Bansal. 2018. Fast abstractive summarization with reinforce-selected sentence rewriting. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 675–686, Melbourne, Australia. Association for Computational Linguistics.
  8. 8.Yulong Chen, Yang Liu, Liang Chen, and Yue Zhang. 2021c. DialogSum: A real-life scenario dialogue summarization dataset. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 5062–5074, Online. Association for Computational Linguistics.
  9. 9.Yulong Chen, Yang Liu, and Yue Zhang. 2021d. Dialogsum challenge: Summarizing real-life scenario dialogues. In Proceedings of the 14th International Conference on Natural Language Generation, pages 308–313.
  10. 10.Alexander Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang, Haoran Li, Yashar Mehdad, and Dragomir Radev. 2021. ConvoSumm: Conversation summarization benchmark and improved abstractive summarization with argument mining. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6866–6880, Online. Association for Computational Linguistics.
  11. 11.Xiachong Feng, Xiaocheng Feng, and Bing Qin. 2021a. A survey on dialogue summarization: Recent advances and new frontiers. arXiv preprint arXiv:2107.03175.
  12. 12.Xiachong Feng, Xiaocheng Feng, Libo Qin, Bing Qin, and Ting Liu. 2021b. Language model as an annotator: Exploring dialogpt for dialogue summarization. arXiv preprint arXiv:2105.12544.
  13. 13.Xiachong Feng, Xiaocheng Feng, Libo Qin, Bing Qin, and Ting Liu. 2021c. Language model as an annotator: Exploring DialoGPT for dialogue summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1479–1491, Online. Association for Computational Linguistics.
  14. 14.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 70–79, Hong Kong, China. Association for Computational Linguistics.
  15. 15.Beliz Gunel, Jingfei Du, Alexis Conneau, and Ves Stoyanov. 2020. Supervised contrastive learning for pretrained language model fine-tuning. arXiv preprint arXiv:2011.01403.
  16. 16.Iryna Gurevych and Michael Strube. 2004. Semantic similarity applied to spoken dialogue summarization. In COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics, pages 764–770, Geneva, Switzerland. COLING.
  17. 17.Jia Jin Koay, Alexander Roustai, Xiaojin Dai, Dillon Burns, Alec Kerrigan, and Fei Liu. 2020. How domain terminology affects meeting summarization performance. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5689–5695, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  18. 18.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  19. 19.Seolhwa Lee, Kisu Yang, Chanjun Park, João Sedoc, and Heuiseok Lim. 2021. Who says like a style of vitamin: Towards syntax-aware dialoguesummarization using multi-task learning. arXiv preprint arXiv:2109.14199.
  20. 20.Yuejie Lei, Yuanmeng Yan, Zhiyuan Zeng, Keqing He, Ximing Zhang, and Weiran Xu. 2021. Hierarchical speaker-aware sequence-to-sequence model for dialogue summarization. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7823–7827. IEEE.
  21. 21.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pretraining for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  22. 22.Daniel Li, Thomas Chen, Albert Tung, and Lydia Chilton. 2021. Hierarchical summarization for longform spoken dialog.
  23. 23.Haoran Li, Junnan Zhu, Jiajun Zhang, and Chengqing Zong. 2018. Ensure the correctness of the summary: Incorporate entailment knowledge into abstractive sentence summarization. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1430–1441, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  24. 24.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  25. 25.Junpeng Liu, Yanyan Zou, Hainan Zhang, Hongshen Chen, Zhuoye Ding, Caixia Yuan, and Xiaojie Wang. 2021a. Topic-aware contrastive learning for abstractive dialogue summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online. Association for Computational Linguistics.
  26. 26.Zhengyuan Liu and Nancy F. Chen. 2021. Controllable neural dialogue summarization with personal named entity planning. In Proceedings of the Association for Computational Linguistics: EMNLP 2020, Online. Association for Computational Linguistics.
  27. 27.Zhengyuan Liu, Ke Shi, and Nancy Chen. 2021b. Coreference-aware dialogue summarization. In Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 509–519, Singapore and Online. Association for Computational Linguistics.
  28. 28.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  29. 29.Iain McCowan, Jean Carletta, Wessel Kraaij, Simone Ashby, S Bourban, M Flynn, M Guillemot, Thomas Hain, J Kadlec, Vasilis Karaiskos, et al. 2005. The ami meeting corpus. In Proceedings of the 5th international conference on methods and techniques in behavioral research, volume 88, page 100. Citeseer.
  30. 30.Rada Mihalcea and Paul Tarau. 2004. TextRank: Bringing order into text. In Proc. of EMNLP, pages 404–411, Barcelona, Spain. Association for Computational Linguistics.
  31. 31.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  33. 33.Clément Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and Patrick Gallinari. 2021. Data-questeval: A referenceless metric for data to text semantic evaluation. arXiv preprint arXiv:2104.07555.
  34. 34.Harvey Sacks, Emanuel A Schegloff, and Gail Jefferson. 1978. A simplest systematics for the organization of turn taking for conversation. In Studies in the organization of conversational interaction, pages 7–55. Elsevier.
  35. 35.Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021. Questeval: Summarization asks for fact-based evaluation.
  36. 36.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proc. of ACL, pages 1073–1083, Vancouver, Canada. Association for Computational Linguistics.
  37. 37.Xiangru Tang, Alexander R. Fabbri, Ziming Mao, Griffin Adams, Borui Wang, Haoran Li, Yashar Mehdad, and Dragomir Radev. 2021. Investigating crowdsourcing protocols for evaluating the factual consistency of summaries.
  38. 38.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  39. 39.Chien-Sheng Wu, Linqing Liu, Wenhao Liu, Pontus Stenetorp, and Caiming Xiong. 2021. Controllable abstractive dialogue summarization with sketch supervision. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 5108–5122, Online. Association for Computational Linguistics.
  40. 40.Feng Xiachong, Feng Xiaocheng, and Qin Bing. 2021. Incorporating commonsense knowledge into abstractive dialogue summarization via heterogeneous graph networks. In Proceedings of the 20th Chinese National Conference on Computational Linguistics, pages 964–975, Huhhot, China. Chinese Information Processing Society of China.
  41. 41.Lin Yuan and Zhou Yu. 2020. Abstractive dialog summarization with semantic scaffolds.
  42. 42.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. arXiv preprint arXiv:2106.11520.
  43. 43.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  44. 44.Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, and Mohit Bansal. 2021. EmailSum: Abstractive email thread summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6895–6909, Online. Association for Computational Linguistics.
  45. 45.Lulu Zhao, Weiran Xu, and Jun Guo. 2020. Improving abstractive dialogue summarization with graph structures and topic words. In Proceedings of the 28th International Conference on Computational Linguistics, pages 437–449, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  46. 46.Lulu Zhao, Zeyuan Yang, Weiran Xu, Sheng Gao, and Jun Guo. 2021. Improving abstractive dialogue summarization with conversational structure and factual knowledge.
  47. 47.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5905–5921, Online. Association for Computational Linguistics.
  48. 48.Chenguang Zhu, Yang Liu, Jie Mei, and Michael Zeng. 2021. Mediasum: A large-scale media interview dataset for dialogue summarization. arXiv preprint arXiv:2103.06410.
  49. 49.Chenguang Zhu, Ruochen Xu, Michael Zeng, and Xuedong Huang. 2020. A hierarchical network for abstractive meeting summarization with cross-domain pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 194–203, Online. Association for Computational Linguistics.

Citation

MLA
Tang, X., et al. “CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5657–68, https://doi.org/10.18653/v1/2022.naacl-main.415.
APA
Tang, X., Nair, A., Wang, B., Wang, B., Desai, J., Wade, A., Li, H., Celikyilmaz, A., Mehdad, Y., & Radev, D. (2022). CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5657–5668. https://doi.org/10.18653/v1/2022.naacl-main.415
Chicago
Tang, X., A. Nair, B. Wang, et al. 2022. “CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5657–68. https://doi.org/10.18653/v1/2022.naacl-main.415.
Harvard
Tang, X. et al. (2022) “CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 5657–5668. Available at: https://doi.org/10.18653/v1/2022.naacl-main.415.
Vancouver
1. Tang X, Nair A, Wang B, Wang B, Desai J, Wade A, Li H, Celikyilmaz A, Mehdad Y, Radev D (2022) CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 5657–5668

BibTeX

@inproceedings{tang-etal-2022-confit,
    title = "{CONFIT}: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning",
    author = "Tang, Xiangru  and
      Nair, Arjun  and
      Wang, Borui  and
      Wang, Bingyao  and
      Desai, Jai  and
      Wade, Aaron  and
      Li, Haoran  and
      Celikyilmaz, Asli  and
      Mehdad, Yashar  and
      Radev, Dragomir",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.415/",
    doi = "10.18653/v1/2022.naacl-main.415",
    pages = "5657--5668"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/