CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning
Xiangru TangArjun NairBorui WangBingyao WangJai DesaiAaron WadeHaoran LiAsli CelikyilmazYashar MehdadDragomir R. Radev
Introduces a linguistically motivated taxonomy of factual errors in abstractive dialogue summarization alongside CONFIT, a contrastive fine-tuning method with targeted hard negative samples that substantially reduces model hallucinations on the SAMSum and AMI benchmarks.
Automated abstractive dialogue summarization is increasingly vital for processing conversational data, yet modern pretrained neural language models frequently generate hallucinated or factually inconsistent content. In multi-party conversations, challenges such as informal phrasing, conversational flow, and shifting first-person references cause standard models to misattribute statements or alter critical details. These factual errors pose substantial risks in operational environments where accuracy and reliability are mandatory for decision-making.
The article demonstrates a linguistically informed contrastive fine-tuning method, named CONFIT, designed to improve the factual consistency and overall quality of abstractive dialogue summarization. The primary objective is to evaluate whether targeted training losses can systematically eliminate the most frequent factual errors produced by standard language models across dialogue and meeting benchmarks.
To establish this method, the authors first created an eight-category taxonomy of factual errors and conducted human evaluations on baseline model outputs to identify major failure modes. Based on these findings, CONFIT augments standard training objectives with two supplementary mechanisms: a contrastive loss using carefully perturbed negative samples (such as swapped entities, altered verbs, masked numbers, and dropped sentences) and a self-supervised auxiliary loss that tracks speaker identities and resolves pronoun references. The approach was tested across standard architectures (BART, Pegasus, and T5) on two representative datasets: the SAMSum chat dialogue benchmark (over 16,000 dialogues) and the AMI meeting corpus (137 meeting transcripts).
The evaluation yielded several key findings. First, error analysis revealed that missing information and wrong references account for approximately 45% of all factual errors in baseline dialogue summaries. Second, applying CONFIT significantly improved human-rated faithfulness across all evaluated architectures; for example, BART faithfulness scores rose from 5.54 to 7.25 on SAMSum and from 4.85 to 5.60 on AMI on a 10-point scale. Third, the framework drastically reduced specific error rates on SAMSum, achieving wrong-reference error reductions of 20 percentage points for BART (from 37% to 17%) and 33 percentage points for T5 (from 46% to 13%). Finally, CONFIT consistently boosted standard automated overlap metrics, achieving superior ROUGE-1 and ROUGE-L scores across both short chat dialogues and long-form meeting transcripts.
These results indicate that conversational factuality errors stem from distinct structural dynamics—particularly speaker tracking and coreference—rather than general language modeling deficiencies alone. By explicitly penalizing known error types and enforcing speaker awareness during fine-tuning, organizations can deploy automated conversational summarizers with substantially reduced risk of misattribution and factual distortion, lowering the cost and effort of manual verification.
Organizations developing or deploying conversational artificial intelligence systems should consider incorporating contrastive and speaker-aware fine-tuning objectives into their training pipelines rather than relying solely on standard cross-entropy objectives. For practical implementation, engineering teams should validate automated evaluation metrics against human review, as automated scoring tools like BARTScore showed inconsistencies on long-form meeting data despite confirmed human-rated improvements.
The findings carry high confidence for standard dialogue and chat domains, supported by rigorous human evaluation. However, confidence should be tempered when applying the technique directly to complex, long-form multi-party meetings; the human evaluation on the AMI dataset was limited to 20 dialogues and exhibited higher baseline error rates (such as 70% to 85% missing information), indicating that longer conversational contexts require broader empirical validation and further structural modeling.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Establishes the foundational taxonomy and evaluation methodologies for hallucinations and factual errors in neural abstractive summarization upon which dialogue-specific error taxonomies build.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Introduces standardized benchmarks and automated evaluation criteria for measuring factual consistency across abstractive summarization and dialogue generation tasks.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). Presents foundational methods for adapting pretrained language model encoders to document summarization and structuring inter-sentence representation.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). Provides key groundwork on abstractive summarization objectives and benchmark evaluations that contrast with purely extractive baselines.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). Examines the limitations of standard automated metrics in conversational contexts, motivating the need for targeted human evaluations and specialized factual losses.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). Introduces hybrid training objectives and attention mechanisms to combat exposure bias and improve factual coherence in abstractive summarization.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). Provides the foundational pointer-generator architecture designed to prevent factual fabrications and out-of-vocabulary misattributions in sequence-to-sequence summarizers.
- Paper: MeetingBank: A Benchmark Dataset for Meeting Summarization, Yebowen Hu et al. (2023). Extends conversational summarization to large-scale, long-form city council meetings by evaluating multi-turn dialogue segmentation and factuality.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). Builds upon fine-tuning for factual consistency by replacing manual error crafting with automated preference-based reinforcement learning pipelines.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). Provides a comprehensive survey generalizing factual error taxonomies, automated verification techniques, and mitigation strategies across modern language models.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Offers a systematic framework categorizing factuality versus faithfulness hallucinations and evaluating mitigation techniques across generation workflows.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). Develops a fine-grained, atomic evaluation framework to automate the precision assessment of factual generation in complex long-form contexts.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Integrates self-reflection tokens and adaptive passage critique during fine-tuning to ensure generated outputs remain faithfully grounded in context.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Proposes a multi-turn cross-examination approach to detect internal factual inconsistencies in conversational generations without external reference texts.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Examines long-term conversational memory and dialogue summarization across multi-session interactions where tracking context and event dynamics is essential.
