Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization

Meng CaoYue DongJackie Chi Kit Cheung

article2022ACL200 citations

Distinguishes beneficial, world-knowledge-accurate hallucinations from false ones in abstractive summarization by comparing an entity’s prior and posterior probabilities under masked language models, providing an effective reward signal for reinforcement learning to improve summary factuality without sacrificing abstractiveness.

Listen

Artificial intelligence summarization models often produce hallucinations—content that cannot be directly inferred from the original source text. Conventional wisdom treats all hallucinations as system errors to eliminate. However, completely removing unmentioned context can degrade output quality, as many unstated details are accurate real-world facts that clarify meaning for readers. Organizations deploying automated summarization must therefore distinguish between helpful background facts and outright factual errors.

The article demonstrates that language model probability shifts can reliably classify entity hallucinations as factual or non-factual, enabling a targeted training mechanism that reduces incorrect statements without sacrificing summary abstractiveness.

The authors developed an entity-level classifier called EntFA using a non-parametric nearest-neighbor approach. The model assesses whether a named entity is factual by comparing its unconditional prior probability from a base language model against its conditional posterior probability from a model given the source document. To test the framework, researchers created a human-annotated dataset of 2,838 entities from 800 generated summaries (XENT) with high inter-annotator agreement (kappa = 0.809), and further evaluated on converted public benchmarks. The classifier was then integrated as a negative reward signal in a reinforcement learning framework to penalize non-factual entity generation during model training.

The investigation produced several key findings. First, while roughly 30% of entities produced by state-of-the-art models were hallucinated, more than half of those hallucinations were accurate according to world knowledge. Second, EntFA outperformed five baseline methods, achieving 90.95% accuracy and an 81.82 F1 score for factuality classification on the primary test set, alongside superior alignment with human judgments on established benchmarks (such as a 0.183 correlation on the FRANK benchmark). Third, error analysis showed that non-factual hallucinations are heavily skewed toward dates (31.65%) and numbers, whereas people, places, and organizations represent the vast majority of factual hallucinations. Finally, using the classifier to penalize non-factual entities during training increased overall factual entity generation from 82.8% to 92.5% and substantially improved faithfulness scores without forcing the model to rely solely on verbatim source copying.

These findings indicate that traditional evaluation metrics and blunt filtering methods are flawed. Optimizing solely for standard word-overlap scores often rewards models for generating non-factual content from noisy training data. Conversely, standard hallucination reduction techniques tend to make models overly extractive, which undermines concise writing. By separating factual background additions from untrue fabrications, organizations can significantly mitigate operational and reputational risks associated with AI errors while preserving natural, informative language generation.

Decision-makers building or adopting text generation systems should replace generic overlap metrics with calibrated probability-based factuality checkers and implement token-level reinforcement penalties during fine-tuning. Future technical roadmaps should extend this probabilistic detection framework beyond named entities to arbitrary text spans and investigate how underlying pre-training datasets introduce or mitigate factual errors.

The primary limitation of this work is its focus on individual named entities and extrinsic hallucinations, which excludes incorrect syntactic relationships among existing facts (intrinsic hallucinations) and full-sentence hallucinations. Additionally, evaluating factuality against world knowledge relied partly on automated search extraction and specific benchmark distributions. Overall confidence in the findings is high for entity-level assessment, though practitioners should conduct targeted validation when applying the method to specialized domains with unique terminology.

arXiv: 2109.09784mcao516/EntFA
Cover for Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization

Abstract

State-of-the-art abstractive summarization systems often generate hallucinations; i.e., content that is not directly inferable from the source text. Despite being assumed incorrect, we find that much hallucinated content is factual, namely consistent with world knowledge. These factual hallucinations can be beneficial in a summary by providing useful background information. In this work, we propose a novel detection approach that separates factual from non-factual hallucinations of entities. Our method utilizes an entity’s prior and posterior probabilities according to pre-trained and fine-tuned masked language models, respectively. Empirical results suggest that our approach outperforms five baselines and strongly correlates with human judgments. Furthermore, we show that our detector, when used as a reward signal in an off-line reinforcement learning (RL) algorithm, significantly improves the factuality of summaries while maintaining the level of abstractiveness.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Model Hallucination
  • 2.2 Summary Factuality
  • 3 Method
  • 3.1 Problem Statement
  • 3.2 The Prior & Posterior Probability of an Entity
  • 3.3 Improving the Factuality of Abstractive Summarization Systems
  • 4 Dataset
  • 4.1 XENT dataset
  • 4.2 MENT Dataset
  • 5 Evaluation Tasks
  • 5.1 Entity-level Hallucination & Factuality Classification
  • 5.2 Correlation with Human Judgments of Factuality
  • 5.3 Evaluating the Factuality of Summarization Systems
  • 6 Experiments
  • 6.1 Classification Experiments
  • 6.2 Correlation Experiments
  • 6.3 Factuality Evaluation Results of Summarization Systems
  • 7 Analysis
  • 7.1 Ablation Studies
  • 7.2 Prior/Posterior Probabilities
  • 8 Conclusion
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Dataset Annotation Guidelines and Process
  • A.2 Patterns of Annotated Entities
  • A.3 Experimental Setup
  • A.4 Classification Results on XENT Dataset
  • A.5 Evaluating Entity Factuality on Noisy Training Data
  • A.6 Where Does the Model Learn to Hallucinate?
  • A.7 Compare with Filippova (2020)'s Work

Knowls

  1. Knowl 1 — EntFA: Entity Hallucination and Factuality Assessment via Prior and Posterior Probabilities

    model/method

    EntFA (Entity Factuality Assessment) determines whether a named entity eke_k appearing in a model-generated summary G=(g1,…,gN)G = (g_1, \dots, g_N) given a source document S=(s1,…,sM)S = (s_1, \dots, s_M) is hallucinated and whether it is factual.

    For an entity eke_k spanning token indices ik,…,ik+∣ek∣−1i_k, \dots, i_k + |e_k| - 1 in GG with surrounding summary context ck=G∖ekc_k = G \setminus e_k, two probability scores are computed autoregressively:

    1. Prior probability pprior(ek)p_{\text{prior}}(e_k), representing the probability of generating eke_k in context ckc_k without access to the source document SS, computed via a pre-trained Masked Language Model (PMLMP_{\text{MLM}}, e.g., BART Large):

    pprior(ek)=∏t=1∣ek∣PMLM(ekt∣ek1…t−1,ck)p_{\text{prior}}(e_k) = \prod_{t=1}^{|e_k|} P_{\text{MLM}}\left(e_k^t \mid e_k^{1\dots t-1}, c_k\right)

    1. Posterior probability ppos(ek)p_{\text{pos}}(e_k), representing the conditional probability of eke_k given context ckc_k and the source document SS, computed via a Conditional Masked Language Model (PCMLMP_{\text{CMLM}}) trained to predict masked entities given the source text and masked summary:

    ppos(ek)=∏t=1∣ek∣PCMLM(ekt∣ek1…t−1,ck,S)p_{\text{pos}}(e_k) = \prod_{t=1}^{|e_k|} P_{\text{CMLM}}\left(e_k^t \mid e_k^{1\dots t-1}, c_k, S\right)

    To classify an entity's status, a KK-Nearest Neighbors (KK-NN) classifier is trained on a 3-dimensional feature representation f(ek)=(pprior(ek),ppos(ek),I(ek∈S))\mathbf{f}(e_k) = (p_{\text{prior}}(e_k), p_{\text{pos}}(e_k), \mathbb{I}(e_k \in S)), where I(ek∈S)∈{0,1}\mathbb{I}(e_k \in S) \in \{0, 1\} is a binary indicator denoting exact string presence of the entity in the source text. Two separate classifiers with k=30k=30 neighbors are trained: one for hallucination status (binary: non-hallucinated vs. hallucinated) and one for factuality status (binary: factual vs. non-factual).

  2. Knowl 2 — Factuality-Aware Off-line Reinforcement Learning for Abstractive Summarization

    model/method

    To prevent abstractive summarization models from learning non-factual hallucinations from noisy reference summaries, summarization training is formulated as an off-line reinforcement learning (RL) problem with expert demonstrations guided by entity factuality rewards.

    Let the state at generation step tt be st=(y<t,x)s_t = (y_{<t}, x) consisting of source text xx and previously generated tokens y<ty_{<t}, and the action ata_t be the token generated by policy πθ(at∣st)\pi_\theta(a_t \mid s_t). The training objective optimizes the policy gradient with importance sampling weights approximated by πθ(at∣st)\pi_\theta(a_t \mid s_t):

    ∇θJ(θ)=Eτ∼πb[∑t=0Tπθ(at∣st)∇θlog⁡πθ(at∣st)Q^(at,st)]\nabla_\theta J(\theta) = \mathbb{E}_{\tau \sim \pi_b}\left[ \sum_{t=0}^T \pi_\theta(a_t \mid s_t) \nabla_\theta \log \pi_\theta(a_t \mid s_t) \hat{Q}(a_t, s_t) \right]

    where πb\pi_b represents the noisy summarization training set (the behavior policy), and the estimated return from state sts_t is Q^(at,st)=∑t′=tTγt′−trt′\hat{Q}(a_t, s_t) = \sum_{t'=t}^T \gamma^{t'-t} r_{t'} with discount factor γ=1\gamma = 1.

    The token-level reward function R(st,at)R(s_t, a_t) utilizes predictions from the EntFA entity factuality classifier on reference summaries:

    R(st,at)={−rnfe,if at belongs to an entity classified as non-factualpMLE(at∣st),otherwiseR(s_t, a_t) = \begin{cases} -r_{\text{nfe}}, & \text{if } a_t \text{ belongs to an entity classified as non-factual} \\ p_{\text{MLE}}(a_t \mid s_t), & \text{otherwise} \end{cases}

    where −rnfe<0-r_{\text{nfe}} < 0 is a negative penalty assigned to tokens of non-factual entities, and pMLE(at∣st)p_{\text{MLE}}(a_t \mid s_t) is the posterior token probability from a maximum likelihood estimation (MLE) pre-trained model.

    To stabilize training, an auxiliary policy network π~θ\tilde{\pi}_\theta tracks πθ\pi_\theta via Polyak updates with an update rate of 0.010.01.

  3. Knowl 3 — Taxonomy of Entity Hallucinations and Factuality in Abstractive Summarization

    definition

    In abstractive summarization of a source document SS yielding a summary GG, a named entity eke_k in GG is categorized into one of four mutually exclusive classes based on its relationship to the source text and external world knowledge:

    1. Non-hallucinated: The entity eke_k and its associated context can be directly entailed from the source document SS.
    2. Factual Hallucination (Extrinsic): The entity eke_k cannot be directly entailed from the source document SS, but the assertion is true and verifiable using external world knowledge (e.g., supplying an unmentioned but correct official title, first name, or nationality for a person mentioned in the source).
    3. Non-factual Hallucination (Extrinsic): The entity eke_k cannot be inferred from the source document SS and is false or unverifiable against world knowledge (e.g., fabricating an incorrect event location, casualty count, or date).
    4. Intrinsic Hallucination: The entity eke_k appears in the source document SS, but the summary misrepresents its relationship or role by conflating it with another event or context in the document.

    Extrinsic hallucinations comprise both factual and non-factual hallucinations.

  4. Knowl 4 — Entity-Level Hallucination and Factuality Classification Results

    empirical result

    Entity-level hallucination detection and factuality checking of the EntFA classifier were evaluated on two datasets: the XENT dataset (human annotations of BART summaries on XSum) and the converted MENT dataset (derived from Maynez et al., 2020). EntFA was evaluated against five baseline approaches: word overlap with source, synonym-based overlap, SimAlign embedding alignment, an LM-based conditional probability baseline (adapted from Filippova, 2020), and a synthetic sequence labeling model (Zhou et al., 2020).

    Model XENT Hallucination XENT Factuality
    Acc. (%) F1 Acc. (%) F1
    Overlap-based 92.93 91.73 81.25 74.19
    Synonym-based 90.76 89.42 81.30 74.79
    Alignment (SimAlign) 78.35 71.10 81.65 66.03
    LM-based 74.18 54.99 84.54 57.80
    Zhou et al. (2020) 86.66 81.71 85.76 75.07
    EntFA (ours) 92.93 91.73 90.95 81.82

    On the MENT dataset, entity-level factuality results are:

    Model Acc. (%) F1
    Overlap-based 68.22 54.68
    Synonym-based 68.91 53.43
    Alignment (SimAlign) 69.21 50.86
    LM-based 67.48 48.02
    Zhou et al. (2020) 71.02 56.42
    EntFA (ours) 78.48 60.23

    A 10-fold cross-validated paired tt-test confirms EntFA significantly outperforms all baselines on factuality classification on XENT with p<3.27×10−5p < 3.27 \times 10^{-5}. Overlap and synonym baselines match hallucination accuracy but degrade on factuality because they treat all non-source entities as non-factual errors.

  5. Knowl 5 — Abstractive Summarization Performance with Factuality-Guided Reinforcement Learning

    empirical result

    Training a BART-Large abstractive summarizer on XSum using off-line reinforcement learning with EntFA factuality rewards improves summary factuality while preserving abstractiveness, compared to standard maximum likelihood estimation (MLE), standard RL (Pang and He, 2021), LM-based reward (Filippova, 2020), and Loss Truncation (Kang and Hashimoto, 2020).

    Evaluations on the XSum official test set measure ROUGE-1 (R1), ROUGE-L (RL), novel nn-gram percentage, percentage of entities not found in source (% ENFS), FEQA faithfulness score, DAE faithfulness score, and the percentage of generated entities classified as factual (% Factual Ent) and factual hallucinations (% Factual Hal) by EntFA:

    System R1 RL Novel 1-g (%) Novel 2-g (%) % ENFS ↓\downarrow FEQA ↑\uparrow DAE ↑\uparrow % Factual Ent ↑\uparrow
    MLE 45.1 37.3 27.86 74.47 42.0 25.9 34.6 82.8
    RL 45.8 37.6 28.14 74.73 43.2 25.6 33.3 82.8
    LM-based 43.2 34.6 29.75 75.86 38.2 24.2 31.3 87.4
    Loss trunc (c=0.3c=0.3) 44.1 36.0 26.82 73.39 41.3 26.3 36.4 83.9
    Loss trunc (c=0.7c=0.7) 42.7 34.8 26.61 73.19 40.6 26.7 38.8 84.1
    Ours (rnfe=2.0r_{\text{nfe}} = 2.0) 44.6 36.2 27.71 74.90 37.2 26.5 37.3 90.1
    Ours (rnfe=4.0r_{\text{nfe}} = 4.0) 43.0 34.9 26.87 74.11 32.8 27.3 40.8 92.5

    The EntFA-guided model with rnfe=2.0r_{\text{nfe}} = 2.0 yields the highest factual hallucination rate (% Factual Hal = 24.0%, compared to 21.4% for MLE and 20.7% for Loss Truncation c=0.7c=0.7). Setting rnfe=4.0r_{\text{nfe}} = 4.0 achieves the highest overall factual entity rate (92.5%) and lowest non-source entity rate (32.8%) while retaining high abstractiveness (26.87% novel unigrams and 74.11% novel bigrams).

  6. Knowl 6 — Summary-Level Correlation of EntFA with Human Factuality Judgments

    empirical result

    Entity-level factuality predictions from EntFA can be aggregated into a summary-level factuality metric by assigning each summary the minimum confidence for the factual class across all its constituent named entities:

    Score(G)=min⁡ek∈GPEntFA(factual∣ek)\text{Score}(G) = \min_{e_k \in G} P_{\text{EntFA}}(\text{factual} \mid e_k)

    This summary-level score was evaluated on two human-annotated factuality benchmarks for the XSum dataset: the FRANK benchmark (Pagnoni et al., 2021) using partial Pearson correlation coefficient (ρ\rho), and the benchmark collected by Wang et al. (2020) using Pearson correlation coefficient (PCC):

    Metric FRANK (Partial Pearson's ρ\rho) Wang et al. (PCC)
    BLEU 0.139 0.118
    ROUGE-1 0.155 0.132
    BERTScore -0.0359 0.025
    QAGS -0.0225 0.175
    FEQA 0.0242 –
    DAE 0.0444 –
    EntFA (ours) 0.183 0.268

    EntFA achieves the strongest correlation with human factuality judgments on both benchmarks (ρ=0.183,p<10−8\rho = 0.183, p < 10^{-8} on FRANK; PCC=0.268\text{PCC} = 0.268 on Wang et al.), significantly outperforming surface overlap metrics (BLEU, ROUGE-1), embedding-based similarity (BERTScore), and QA/entailment-based metrics (QAGS, FEQA, DAE).

  7. Knowl 7 — Composition and Annotation Statistics of the XENT Dataset

    data/table

    The XENT dataset consists of 800 summaries generated by BART on source documents randomly selected from the XSum test set, containing 2,838 named entities extracted via spaCy NER. Each entity was manually annotated with one of four tags:

    Entity Category Number of Samples Percentage (%)
    Non-hallucinated 1,921 67.69%
    Factual hallucination 441 15.54%
    Non-factual hallucination 421 14.83%
    Intrinsic hallucination 55 1.94%
    Total Entities 2,838 100.0%

    Key characteristics of the dataset include:

    • Approximately 30.37% of generated entities are extrinsic hallucinations (862 entities), and more than half of all extrinsic hallucinations (51.16%, 441/862) are factually accurate with respect to world knowledge.
    • Entity-type distribution varies across classes: Person (33.23%), GPE (21.75%), and ORG (18.43%) are most prevalent among factual hallucinations, whereas Date (31.65%), Person (20.25%), Other (18.68%), and Cardinal numbers (12.97%) are most prevalent among non-factual hallucinations.
    • Inter-annotator agreement computed over 800 entities yielded Fleiss's Kappa κ=0.809\kappa = 0.809 and majority-class agreement μ=0.931\mu = 0.931 across the four categories.
  8. Knowl 8 — Feature Ablation for Entity Hallucination and Factuality Classification

    empirical result

    An ablation study on the XENT test set evaluates the impact of each feature component in the EntFA classifier: the prior probability pprior(ek)p_{\text{prior}}(e_k) from MLM, the posterior probability ppos(ek)p_{\text{pos}}(e_k) from CMLM, and the binary string overlap feature I(ek∈S)\mathbb{I}(e_k \in S):

    Feature Configuration Factuality F1 Hallucination F1
    Full EntFA 81.82 91.73
    w/o string overlap 77.18 74.83
    w/o prior probability (ppriorp_{\text{prior}}) 80.12 91.32
    w/o posterior probability (pposp_{\text{pos}}) 70.30 91.12

    The ablation reveals asymmetric feature dependencies:

    1. Factuality classification critically depends on the posterior probability pposp_{\text{pos}}: removing posterior probability causes an 11.52-point drop in F1 (from 81.82 to 70.30).
    2. Hallucination detection critically depends on the string overlap feature: removing string overlap causes a 16.90-point drop in F1 (from 91.73 to 74.83), while removing prior or posterior probabilities causes negligible drops (≤0.61\le 0.61 points).
  9. Knowl 9 — Dissecting Knowledge Sources in Factual Hallucinations

    empirical result

    To determine whether factual hallucinations generated by summarization models originate from pre-training corpora or fine-tuning datasets, the log-likelihood ratio of posterior probabilities from two separate Conditional Masked Language Models (CMLMs) was evaluated:

    σ(ek)=log⁡PCMLMXSUM(ek)PCMLMCNN/DM(ek)\sigma(e_k) = \log \frac{P_{\text{CMLM}_{\text{XSUM}}}(e_k)}{P_{\text{CMLM}_{\text{CNN/DM}}}(e_k)}

    where CMLMXSUM\text{CMLM}_{\text{XSUM}} was fine-tuned on XSum and CMLMCNN/DM\text{CMLM}_{\text{CNN/DM}} on CNN/DailyMail.

    Key findings include:

    1. Non-hallucinated entities exhibit similar posterior distributions under both models. Factual hallucinations often have low posterior under CMLMCNN/DM\text{CMLM}_{\text{CNN/DM}} but high posterior under CMLMXSUM\text{CMLM}_{\text{XSUM}}, indicating knowledge acquired from the XSum training set.
    2. In a TF-IDF retrieval experiment of the 10 most similar training documents for each factual hallucinated entity:
      • Entities with σ(ek)≥5\sigma(e_k) \ge 5 appear on average 2.192.19 times in similar XSum training samples versus 0.770.77 times in CNN/DailyMail samples.
      • Entities with σ(ek)≤0\sigma(e_k) \le 0 show comparable frequencies (2.852.85 on XSum vs. 2.462.46 on CNN/DailyMail).
    3. CMLMs outperform unidirectional Conditional Language Models (CLMs) for factuality verification because conditioning on both left and right context resolves uncertainty due to content selection, yielding an ROC-AUC of 0.910.91 for CMLM versus 0.880.88 for CLM on XSum.
  10. Knowl 10 — Inverse Relationship Between ROUGE Score and Entity Factuality on Noisy Data

    empirical result

    Training neural abstractive summarizers (BART and PEGASUS) on synthetic training mixtures of clean samples (where all reference entities appear in the source text) and noisy samples (where reference summaries contain entities absent from the source) reveals an inverse relationship between ROUGE scores and factual consistency:

    1. As the proportion of noisy training samples increases from 0% (50k clean samples) to 100% (50k or 100k noisy samples), the ROUGE-1 score of the trained model increases while the percentage of factual entities generated monotonically decreases.
    2. For BART, training on 50k clean samples yields an entity factuality rate of 91.81% with ROUGE-1 of 40.6, whereas training on 50k noisy samples drops entity factuality to ~82% while ROUGE-1 rises to ~43.
    3. For PEGASUS-Large, training on 50k clean samples yields 84.79% entity factuality (ROUGE-1 = 45.0), with factuality declining as the training noise ratio increases.
    4. Among 7,358 non-source entities generated by BART trained on noisy data that were predicted as factual by EntFA, 50.5% appeared in the ground-truth reference summaries; in contrast, only 12.7% of entities predicted as non-factual appeared in the references.

    This indicates that optimizing summarization systems solely for ROUGE against noisy references incentivizes the model to produce non-factual hallucinations at inference time.

Coverage note — None omitted. All main contributions (EntFA formulation, off-line RL training with factuality rewards, XENT taxonomy and dataset statistics, classification and generation benchmarks, human correlation results, ablations, knowledge origin analysis, and noise-factuality trade-offs) are fully covered.

References

  1. 1.Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020. Factual error correction for abstractive summarization models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6251–6258, Online. Association for Computational Linguistics.
  2. 2.Yue Dong, Shuohang Wang, Zhe Gan, Yu Cheng, Jackie Chi Kit Cheung, and Jingjing Liu. 2020. Multi-fact correction in abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9320–9331, Online. Association for Computational Linguistics.
  3. 3.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  4. 4.Katja Filippova. 2020. Controlled hallucinations: Learning to generate faithfully from noisy data. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 864–870, Online. Association for Computational Linguistics.
  5. 5.Tanya Goyal and Greg Durrett. 2020. Evaluating factuality in generation with dependency-level entailment. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3592–3603, Online. Association for Computational Linguistics.
  6. 6.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, volume 28, pages 1693–1701. Curran Associates, Inc.
  7. 7.Matthew Honnibal and Ines Montani. 2017. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear.
  8. 8.Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schütze. 2020. SimAlign: High quality word alignments without parallel training data using static and contextualized embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1627–1643, Online. Association for Computational Linguistics.
  9. 9.Daniel Kang and Tatsunori Hashimoto. 2020. Improved natural language generation via loss truncation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 718–731, Online. Association for Computational Linguistics.
  10. 10.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  11. 11.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  12. 12.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  13. 13.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  14. 14.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  15. 15.Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. 2021. Entity-level factual consistency of abstractive text summarization. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2727–2733, Online. Association for Computational Linguistics.
  16. 16.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018a. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  17. 17.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018b. Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.
  18. 18.Shashi Narayan, Yao Zhao, Joshua Maynez, Gonçalo Simoes, and Ryan McDonald. 2021. Planning with entity chains for abstractive summarization. arXiv preprint arXiv:2104.07606.
  19. 19.Ani Nenkova and Rebecca J. Passonneau. 2004. Evaluating content selection in summarization: The pyramid method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 145–152.
  20. 20.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations.
  21. 21.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Mexico City.
  22. 22.Richard Yuanzhe Pang and He He. 2021. Text generation by learning from demonstrations. In International Conference on Learning Representations.
  23. 23.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in pytorch.
  24. 24.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition,.
  25. 25.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  26. 26.Chaojun Wang and Rico Sennrich. 2020. On exposure bias, hallucination and domain shift in neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3544–3552, Online. Association for Computational Linguistics.
  27. 27.Jiacheng Xu, Shrey Desai, and Greg Durrett. 2020. Understanding neural abstractive summarization models via uncertainty. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6275–6281, Online. Association for Computational Linguistics.
  28. 28.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 11328–11339. PMLR.
  29. 29.Zheng Zhao, Shay B. Cohen, and Bonnie Webber. 2020. Reducing quantity hallucinations in abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2237–2249, Online. Association for Computational Linguistics.
  30. 30.Chunting Zhou, Jiatao Gu, Mona Diab, Paco Guzman, Luke Zettlemoyer, and Marjan Ghazvininejad. 2020. Detecting hallucinated content in conditional neural sequence generation. arXiv preprint arXiv:2011.02593.

Citation

MLA
Cao, M., et al. “Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3340–54, https://doi.org/10.18653/v1/2022.acl-long.236.
APA
Cao, M., Dong, Y., & Cheung, J. C. K. (2022). Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3340–3354. https://doi.org/10.18653/v1/2022.acl-long.236
Chicago
Cao, M., Y. Dong, and J. C. K. Cheung. 2022. “Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3340–54. https://doi.org/10.18653/v1/2022.acl-long.236.
Harvard
Cao, M., Dong, Y. and Cheung, J.C.K. (2022) “Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3340–3354. Available at: https://doi.org/10.18653/v1/2022.acl-long.236.
Vancouver
1. Cao M, Dong Y, Cheung JCK (2022) Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3340–3354

BibTeX

@inproceedings{cao-etal-2022-hallucinated,
    title = "Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization",
    author = "Cao, Meng  and
      Dong, Yue  and
      Cheung, Jackie",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.236/",
    doi = "10.18653/v1/2022.acl-long.236",
    pages = "3340--3354"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/