Evaluating Factuality in Text Simplification

Ashwin DevarajWilliam SheffieldByron C. WallaceJunyi Jessy Li

article2022ACL53 citationsOutstanding Paper

Presents a taxonomy of factual errors in text simplification, revealing that standard evaluation metrics fail to detect frequent information insertions, deletions, and substitutions in both benchmark datasets and model outputs.

Listen

Automated text simplification systems aim to make complex information accessible to wider audiences, including laypeople, language learners, and individuals with cognitive or reading difficulties. However, these systems often introduce factual inaccuracies by adding unsupported statements, omitting crucial details, or altering original meanings. In high-stakes domains such as healthcare, providing readable but factually incorrect information can mislead users and cause severe harm. While factuality has been heavily studied in text summarization, its impact on text simplification has remained largely unexamined.

The article establishes a systematic framework to evaluate factual consistency in automated text simplification. Its primary objective is to quantify the prevalence of factual errors in standard benchmark datasets and modern artificial intelligence models, while assessing whether existing evaluation metrics reliably detect these errors.

To accomplish this, the authors developed an error typology categorizing factual distortions into information insertion, information deletion, and information substitution, each graded on a three-level severity scale. They conducted extensive human evaluations using crowdsourced annotators to examine standard simplification datasets, namely Wikilarge and Newsela, as well as outputs from multiple recurrent and Transformer-based models, including a fine-tuned T5 system. The team also benchmarked standard quality metrics against human judgments and trained an automated classification model supplemented with synthetic data to test the feasibility of automatic error detection.

The analysis yielded four major findings. First, standard human-authored reference datasets contain substantial factual flaws; in the Newsela dataset, over 80% of sentences exhibited deletion errors, with nearly 43% categorized as severe omissions that obscured the core message. Second, artificial intelligence models mirror and compound these issues; while modern Transformer models generate fewer unreadable outputs and fewer deletion errors than older recurrent architectures, pre-trained models introduce noticeable insertion errors on abstractive datasets and generate substitution errors at rates higher than those present in training data. Third, standard evaluation metrics fail to gauge accuracy: the primary industry metric, SARI, shows virtually no correlation with human factual judgments, and semantic similarity metrics identify deletions moderately well but fail to capture inappropriate insertions or substitutions. Fourth, a case study on medical text simplification showed that 30% of simplified paragraphs contained critical factual errors that distorted the clinical meaning.

These findings indicate that current simplification systems pose substantial operational and safety risks if deployed without human oversight in regulated or high-stakes environments. Furthermore, because standard development metrics reward surface-level vocabulary changes without penalizing factual distortions, relying on conventional benchmarks provides a false sense of model reliability and performance.

Organizations developing or deploying text simplification technologies should immediately cease relying solely on metrics like SARI for quality assurance. Next steps should include adopting factuality-focused verification pipelines, incorporating full document context to prevent out-of-context sentence distortion, and implementing strict human-in-the-loop review for critical communications. Future research must focus on building dedicated fact-checking models tailored to simplification tasks.

Confidence in these findings is supported by strong inter-annotator agreement across multiple dataset evaluations. However, the study's automated error detection model faced limitations due to small training sample sizes for rare error types, resulting in low detection performance for severe substitutions. Stakeholders should treat existing automated fact-checking models as experimental and exercise caution when applying automated simplification to technical domains.

Cover for Evaluating Factuality in Text Simplification

Abstract

Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Information Errors in Simplification
  • 4 Data and Models
  • 5 Labeling with Mechanical Turk
  • 6 Factuality of Reference Examples
  • 7 Factuality of System Outputs
  • 8 Comparison with Existing Metrics
  • 9 Automatic Factuality Assessment
  • 10 Conclusion
  • Acknowledgements
  • References
  • A Training details for the T5 simplification model
  • B Noise filtering on Wikilarge
  • C Numerical Details for System Output Results
  • D Qualitative Analysis of System Outputs
  • E Automatic Factuality Assessment
  • E.1 Synthetic Data Generation
  • E.2 Model
  • E.3 Training Details
  • F Case Study: Medical Texts

Knowls

  1. Knowl 1 — A graded, three-axis taxonomy of factuality errors in simplification

    definition

    Factuality in a simplified text is assessed independently along three dimensions, because multiple error types can occur in the same source–simplification pair:

    • Insertion: content in the simplified text that is not stated in the complex source. Label 0 means no or only trivial new content; 1 means nontrivial added content that supports the same main idea; 2 means the simplification introduces a new main idea.
    • Deletion: source content omitted from the simplified text. Label 0 means no or only trivial omission; 1 means an omission that leaves the source’s main idea intact; 2 means information critical to that main idea is removed.
    • Substitution: source content altered in the simplified text. Label 0 means no altered information; 1 means an alteration that leaves the main idea intact; 2 means the source’s main idea is changed.

    Each dimension receives its own label from 0, 1, or 2. Label −1 is available for any dimension when either sentence is malformed or unintelligible, such as a fragment or gibberish. The scheme treats insertions as a precision problem, deletions as a recall problem, and substitutions as changes in the accuracy of conveyed information.

  2. Knowl 2 — Human-annotation design and evaluation coverage

    experimental setup

    The study annotated aligned complex–simplified pairs from Wikilarge and Newsela, as well as outputs from simplification systems. For Newsela, it used the simplest reading level and sampled 400 pairs from each of the validation and test sets. For Wikilarge, it annotated 400 validation pairs and all 359 test pairs. System outputs were evaluated on the corresponding examples. The evaluated systems were the RNN models Dress and EditNTS, the Transformer models Access and ControlTS, and a fine-tuned T5 model. T5-base was trained for 5 epochs with batch size 6 and learning rate 3×10−43\times10^{-4}; its input used the prefix summarize:, and outputs used nucleus sampling with p=0.9p=0.9 for Newsela and beam search with 6 beams for Wikilarge.

    Amazon Mechanical Turk workers first had to score at least 75% on a 10-pair qualification task. Three qualified workers labeled each pair; a category received the majority label, and that category’s annotation was discarded if all three workers chose different labels. Majority agreement was 96% for insertion, 96% for deletion, and 95% for substitution. Among examples with a majority nonzero label, agreement on the specific label was 77%, 92%, and 74%, respectively. Ordinal Krippendorff’s alpha was 0.425 for insertion, 0.639 for deletion, and 0.200 for substitution; the calculation treated −1 as the most severe label.

  3. Knowl 3 — Reference simplifications contain frequent deletions, especially in Newsela

    empirical result

    Human annotations of the test references show that deletion is much more common than insertion or substitution, and that Newsela has substantially more deletion and insertion errors than Wikilarge. The percentages below are the distributions across labels 0, 1, 2, and −1, in that order.

    Error typeDatasetLabel 0 (%)Label 1 (%)Label 2 (%)Label −1 (%)
    InsertionWikilarge91.16.30.32.3
    InsertionNewsela68.220.211.10.5
    DeletionWikilarge76.218.03.52.3
    DeletionNewsela15.840.842.90.5
    SubstitutionWikilarge90.16.70.92.3
    SubstitutionNewsela94.93.80.80.5

    Levels 1 and 2 together account for 83.7% of Newsela deletions, compared with 21.5% for Wikilarge. Conversely, most references in both datasets have no substitution error: label 0 applies to 90.1% of Wikilarge pairs and 94.9% of Newsela pairs.

  4. Knowl 4 — Greater deletion severity accompanies more compression and rewriting

    empirical result

    In reference pairs, more severe deletion errors generally accompany greater shortening of the simplified sentence and larger normalized edit distance from the source. Length change is the percentage change from source to simplification (negative values indicate shortening); entries are mean (standard deviation). Normalized edit distance entries are also mean (standard deviation).

    Error typeDatasetLength change, label 0Length change, label 1Length change, label 2Edit distance, label 0Edit distance, label 1Edit distance, label 2
    InsertionWikilarge−5.0 (17.0)22.4 (36.9)7.1 (0.0)0.20 (0.20)0.55 (0.40)0.58 (0.0)
    InsertionNewsela−39.4 (23.8)−19.0 (36.9)−38.3 (29.0)0.41 (0.17)0.51 (0.21)0.54 (0.04)
    DeletionWikilarge2.8 (15.8)−22.3 (18.9)−35.9 (15.9)0.19 (0.23)0.35 (0.18)0.39 (0.14)
    DeletionNewsela1.5 (27.6)−34.8 (23.1)−49.6 (22.8)0.34 (0.31)0.46 (0.13)0.53 (0.10)

    The relationship is clearest for deletion: label-2 deletions are associated with greater shortening and higher edit distance than label-0 deletions in both datasets. Insertion patterns are less consistent, although nonzero insertion labels generally have higher edit distances. After filtering noisy Wikilarge alignments, median normalized edit distance was 0.46 for Newsela and 0.38 for Wikilarge.

  5. Knowl 5 — Simplification systems introduce errors, with distinct architecture-related patterns

    data/table

    The table reports SARI and the manually annotated percentage distribution for insertion (I), deletion (D), and substitution (S) in system outputs. Each category is listed as percentages for labels 0/1/2/−1, in that order. The results show that generated outputs can have substantial deletion and substitution errors; Transformer outputs generally have fewer severe deletion and gibberish labels than RNN outputs, while T5 has more insertion errors on Newsela than on Wikilarge.

    ModelDatasetSARII: 0/1/2/−1D: 0/1/2/−1S: 0/1/2/−1
    DressWikilarge34.991.9/0.8/0.8/6.542.6/24.6/26.2/6.684.4/4.1/4.9/6.6
    DressNewsela34.590.5/0.0/0.0/9.529.9/29.2/32.1/9.767.4/6.5/15.9/10.1
    EditNTSWikilarge40.494.3/4.9/0.8/0.055.0/24.2/20.8/0.088.5/4.1/7.4/0.0
    EditNTSNewsela36.369.4/0.7/2.7/27.29.5/19.0/44.2/27.264.4/2.1/6.2/27.4
    T5Wikilarge34.996.8/1.6/0.8/0.881.6/14.4/3.2/0.897.6/1.6/0.0/0.8
    T5Newsela38.681.7/9.6/7.0/1.727.7/43.7/26.9/1.792.4/5.9/0.0/1.7
    AccessWikilarge49.789.1/8.2/0.9/1.857.5/34.9/5.7/1.971.1/18.6/8.2/2.1
    ControlTSWikilarge42.388.8/7.8/1.7/1.747.8/39.1/11.3/1.781.5/15.1/1.7/1.7

    Dress outputs were available only for Wikilarge; ControlTS was evaluated only on Wikilarge because its Newsela split differed, and the authors could not reproduce its Newsela results. Across the three Transformer models, severe deletion labels were lower than for the RNN models on Wikilarge; T5 also had lower deletion error rates on Newsela. Transformer outputs had markedly fewer −1 gibberish labels than RNN outputs. T5’s insertion labels 1 and 2 sum to 16.6% on Newsela, versus 2.4% on Wikilarge.

  6. Knowl 6 — Common automatic quality metrics do not reliably measure simplification factuality

    empirical result

    The authors measured Spearman rank correlations between factuality labels and existing simplification or factuality metrics. In the tables, I, D, and S mean insertion, deletion, and substitution. SARI correlations vary by model and dataset and are generally weak. Semantic similarity measures correlate most strongly with deletion, less with insertion, and very weakly with substitution. The tested model-based summarization factuality metrics also correlate weakly with insertion and deletion, although DAE has stronger correlation with substitution than FACT-CC or the evaluated semantic-similarity measures.

    SARI versus error labels (coefficients are listed as I/D/S):

    ModelDatasetIDS
    DressWikilarge0.038−0.0410.156
    DressNewsela0.1050.2670.258
    EditNTSWikilarge0.011−0.2750.034
    EditNTSNewsela−0.144−0.103−0.183
    T5Wikilarge−0.0500.1340.027
    T5Newsela−0.020−0.1240.078
    AccessWikilarge0.035−0.0260.057
    ControlTSWikilarge0.002−0.0540.262

    Semantic similarity versus error labels:

    MeasureIDS
    Jaccard similarity−0.385−0.695−0.101
    Cosine similarity, GloVe−0.315−0.620−0.066
    Cosine similarity, ELMo−0.325−0.582−0.065
    Cosine similarity, Sentence-BERT−0.375−0.724−0.182
    BERTScore−0.400−0.748−0.125

    Model-based factuality metrics versus error labels:

    MeasureIDS
    FACT-CC0.3110.4180.165
    DAE, k=1k=10.1090.2170.277
    DAE, k=3k=30.1100.2130.271
    DAE, k=5k=50.1150.2170.271

    For FACT-CC, the score is the model’s probability that a pair is inconsistent. For DAE, the score is the average of the lowest kk probabilities that target dependency arcs do not entail the source. The results indicate that high lexical or semantic similarity is not a sufficient factuality check, particularly for substitutions.

  7. Knowl 7 — Qualitative inspection identifies recurring omission and insertion mechanisms

    empirical result

    Manual analysis of reference pairs found that mild deletions often remove nonsalient details while preserving the main idea. Severe reference deletions commonly either remove the main clause and leave a secondary clause as the sentence, or omit a short but consequential qualifier that changes how the remaining sentence should be interpreted. Newsela insertions included unsupported attribution (for example, adding that experts said something), temporal details, and changes in specificity. Because Newsela stories were rewritten at document level, information could move between adjacent sentences: an isolated sentence pair could therefore appear underspecified even when information was preserved elsewhere in the document.

    In system outputs, deletion errors ranged from short changes to a word or phrase—including pronoun substitutions and lost modifiers—to longer omissions of prepositional phrases or subordinate and coordinate clauses. For Dress and EditNTS, label-1 errors were mainly short omissions and label-2 errors were usually longer omissions. The authors did not observe the reference-data pattern of deleting a main clause and promoting a secondary clause in model outputs. T5 also produced label-2 deletions involving a semantically critical word. These observations support the paper’s caution that sentence-level evaluation can misrepresent document-context edits and that compression or extensive rewriting may harm factuality.

  8. Knowl 8 — RoBERTa classifiers show that automatic error labeling remains difficult

    empirical result

    The authors framed factuality assessment as three separate classification tasks, one each for insertion, deletion, and substitution, with labels 0, 1, and 2. They fine-tuned RoBERTa classifiers using human-annotated examples; the test set was the set used for the earlier analyses, and 1,004 additional examples were collected for training from Wikilarge, Newsela, Access outputs on Wikilarge, and T5 outputs on both datasets. Insertion and substitution classifiers were pretrained on synthetic data, whereas the deletion classifier was trained directly on annotated data.

    The table gives the training count for each label and test F1 for that label:

    CategoryLabel 0: count / F1Label 1: count / F1Label 2: count / F1
    Insertion823 / 87.9104 / 36.640 / 30.4
    Deletion413 / 84.2356 / 57.1204 / 52.1
    Substitution810 / 82.7110 / 19.833 / 9.5

    F1 is reported in the paper’s scale (0–100). Deletion detection performed best, consistent with having more level-1 and level-2 deletion training examples. Synthetic data improved insertion and substitution detection compared with training without augmentation, for which level-1 and/or level-2 F1 was near zero; nevertheless, especially for substitutions, nonzero error labels remained difficult to detect.

  9. Knowl 9 — Synthetic examples augment rare insertion and substitution labels

    model/method

    To supplement the imbalanced human-labeled training data, the authors generated synthetic insertion and substitution pairs by modifying source sentences from the validation set. Insertion examples were produced by replacing a person’s name with a pronoun, or by deleting a phrase and then swapping the source and modified text so that the deletion became an insertion relative to the new source. Name-based insertions were labeled 1. Phrase insertions were labeled 1 when BERTScore was between 0.6 and 0.8, and 2 when it was between 0.2 and 0.4; examples outside those ranges were discarded.

    Substitutions were generated by replacing each number with a random number of the same order of magnitude, negating auxiliary verbs, or perturbing tokens with BERT masking. Number changes and negations received label 1. For label-1 BERT perturbations, two tokens were masked and replaced by the third-highest-probability tokens; for label 2, every fifth token was masked and replaced by the fifth-highest-probability token. Label-0 examples from the original training data were added to the synthetic sets. The resulting datasets contained 3,157 insertion examples (823 label 0, 1,167 label 1, 1,167 label 2) and 7,390 substitution examples (810 label 0, 4,572 label 1, 2,008 label 2).

  10. Knowl 10 — A small medical simplification pilot also finds consequential errors

    empirical result

    As a high-stakes case study, the authors assessed 10 randomly selected paragraph outputs from a medical text simplification system fine-tuned from BART on technical abstracts paired with plain-English Cochrane summaries. A trained annotator applied the error scheme to the paragraphs. Three of the 10 outputs had at least one level-2 error, and five had more than one error. The label counts were:

    CategoryLevel 0Level 1Level 2
    Insertion541
    Deletion082
    Substitution811

    Examples included changing a statement that studies had methodological limitations into a stronger claim that they were of poor quality, and altering a statement of no difference in operating time or complication rates into a claim that there was insufficient evidence to determine whether a difference existed. This is an exploratory sample, not a population-level error-rate estimate; it illustrates that the same error categories apply to paragraph-level medical simplification.

Coverage note — The detailed staged RoBERTa checkpoint schedule and the paper’s individual illustrative sentence pairs are omitted as implementation detail and examples; the contributed taxonomy, evaluation design, quantitative findings, automatic-assessment attempt, and medical pilot are retained.

References

  1. 1.Fernando Alva-Manchego, Carolina Scarton, and Lucía Specia. 2020. Data-driven sentence simplification: Survey and benchmark. Computational Linguistics, 46(1):135–187.
  2. 2.Ron Artstein and Massimo Poesio. 2008. Inter-Coder Agreement for Computational Linguistics. Computational Linguistics, 34(4):555–596.
  3. 3.John Carroll, Guido Minnen, Yvonne Canning, Siobhan Devlin, and John Tait. 1998. Practical simplification of english newspaper text to assist aphasic readers. In Proceedings of the AAAI-98 Workshop on Integrating Artificial Intelligence and Assistive Technology, pages 7–10.
  4. 4.Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021. Decontextualization: Making sentences stand-alone. Transactions of the Association for Computational Linguistics, 9:447–461.
  5. 5.Jerwin Jan S Damay, Gerard Jaime D Lojico, Kimberly Amanda L Lu, D Tarantan, and E Ong. 2006. SIMTEXT: Text simplification of medical literature. In Proceedings of the 3rd National Natural Language Processing Symposium-Building Language Tools and Resources, pages 34–38.
  6. 6.Jan De Belder and Marie-Francine Moens. 2010. Text simplification for children. In Proceedings of the SIGIR workshop on accessible search systems, pages 19–26.
  7. 7.Ashwin Devaraj, Iain Marshall, Byron Wallace, and Junyi Jessy Li. 2021. Paragraph-level simplification of medical texts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4972–4984, Online. Association for Computational Linguistics.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Yue Dong, Zichao Li, Mehdi Rezagholizadeh, and Jackie Chi Kit Cheung. 2019. EditNTS: An neural programmer-interpreter model for sentence simplification through explicit editing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3393–3402, Florence, Italy. Association for Computational Linguistics.
  10. 10.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  11. 11.Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019a. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2214–2220, Florence, Italy. Association for Computational Linguistics.
  12. 12.Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019b. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2214–2220, Florence, Italy. Association for Computational Linguistics.
  13. 13.Tanya Goyal and Greg Durrett. 2020. Evaluating factuality in generation with dependency-level entailment. Findings of the Association for Computational Linguistics: EMNLP 2020.
  14. 14.Tanya Goyal and Greg Durrett. 2021. Annotating and modeling fine-grained factuality in summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1449–1462, Online. Association for Computational Linguistics.
  15. 15.Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2018. Dynamic multi-level multi-task learning for sentence simplification. In Proceedings of the 27th International Conference on Computational Linguistics, pages 462–476, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  16. 16.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. Eighth International Conference on Learning Representations.
  17. 17.George Roger Klare. 1963. Measurement of readability.
  18. 18.Klaus Krippendorff. 1970. Estimating the reliability, systematic error and random error of interval data. Educational and Psychological Measurement, 30(1):61–70.
  19. 19.Reno Kriz, João Sedoc, Marianna Apidianaki, Carolina Zheng, Gaurav Kumar, Eleni Miltsakaki, and Chris Callison-Burch. 2019. Complexity-weighted loss and diverse reranking for sentence simplification. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3137–3147, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Neural text summarization: A critical evaluation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 540–551, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020a. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  22. 22.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020b. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  23. 23.Philippe Laban, Tobias Schnabel, Paul Bennett, and Marti A. Hearst. 2021. Keep it simple: Unsupervised simplification of multi-paragraph text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6365–6378, Online. Association for Computational Linguistics.
  24. 24.Vladimir I. Levenshtein. 1965. Binary codes capable of correcting deletions, insertions, and reversals. Soviet physics. Doklady, 10:707–710.
  25. 25.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pretraining for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  26. 26.Junyi Jessy Li, Bridget O’Daniel, Yi Wu, Wenli Zhao, and Ani Nenkova. 2016. Improving the annotation of sentence specificity. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3921–3927, Portorož, Slovenia. European Language Resources Association (ELRA).
  27. 27.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  28. 28.Mounica Maddela, Fernando Alva-Manchego, and Wei Xu. 2021. Controllable text simplification with explicit paraphrasing. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3536–3553, Online. Association for Computational Linguistics.
  29. 29.Louis Martin, Éric de la Clergerie, Benoît Sagot, and Antoine Bordes. 2020. Controllable sentence simplification. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4689–4698, Marseille, France. European Language Resources Association.
  30. 30.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  31. 31.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  32. 32.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  33. 33.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  34. 34.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  36. 36.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  37. 37.Luz Rello, Ricardo Baeza-Yates, Laura Dempere-Marco, and Horacio Saggion. 2013. Frequent words improve readability and short words improve understandability for people with dyslexia. In Human-Computer Interaction – INTERACT 2013, pages 203–219, Berlin, Heidelberg. Springer Berlin Heidelberg.
  38. 38.Charles Spearman. 1904. The proof and measurement of association between two things. American Journal of Psychology, 15:72–101.
  39. 39.Renliang Sun, Zhe Lin, and Xiaojun Wan. 2020. On the helpfulness of document context to sentence simplification. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1411–1423, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  40. 40.Teun A Van Dijk. 2013. News as discourse. Routledge.
  41. 41.Byron C. Wallace, Sayantan Saha, Frank Soboczenski, and Iain J. Marshall. 2021. Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization. In Proceedings of AMIA Informatics Summit.
  42. 42.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020a. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  43. 43.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020b. Asking and answering questions to evaluate the factual consistency of summaries. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.
  44. 44.Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3):229–256.
  45. 45.Wei Xu, Chris Callison-Burch, and Courtney Napoles. 2015. Problems in current text simplification research: New data can help. Transactions of the Association for Computational Linguistics, 3:283–297.
  46. 46.Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016. Optimizing statistical machine translation for text simplification. Transactions of the Association for Computational Linguistics, 4:401–415.
  47. 47.Xinnuo Xu, Ondřej Dušek, Jingyi Li, Verena Rieser, and Ioannis Konstas. 2020. Fact-based content weighting for evaluating abstractive summarisation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5071–5081, Online. Association for Computational Linguistics.
  48. 48.Yasukata Yano, Michael H Long, and Steven Ross. 1994. The effects of simplified and elaborated texts on foreign language reading comprehension. Language learning, 44(2):189–219.
  49. 49.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. Eighth International Conference on Learning Representations.
  50. 50.Xingxing Zhang and Mirella Lapata. 2017. Sentence simplification with deep reinforcement learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 584–594, Copenhagen, Denmark. Association for Computational Linguistics.
  51. 51.Yanbin Zhao, Lu Chen, Zhi Chen, and Kai Yu. 2020. Semi-supervised text simplification with back-translation and asymmetric denoising autoencoders. Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):9668–9675.

Citation

MLA
Devaraj, A., et al. “Evaluating Factuality in Text Simplification”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7331–45, https://doi.org/10.18653/v1/2022.acl-long.506.
APA
Devaraj, A., Sheffield, W., Wallace, B. C., & Li, J. J. (2022). Evaluating Factuality in Text Simplification. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7331–7345. https://doi.org/10.18653/v1/2022.acl-long.506
Chicago
Devaraj, A., W. Sheffield, B. C. Wallace, and J. J. Li. 2022. “Evaluating Factuality in Text Simplification”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7331–45. https://doi.org/10.18653/v1/2022.acl-long.506.
Harvard
Devaraj, A. et al. (2022) “Evaluating Factuality in Text Simplification”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7331–7345. Available at: https://doi.org/10.18653/v1/2022.acl-long.506.
Vancouver
1. Devaraj A, Sheffield W, Wallace BC, Li JJ (2022) Evaluating Factuality in Text Simplification. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7331–7345

BibTeX

@inproceedings{devaraj-etal-2022-evaluating,
    title = "Evaluating Factuality in Text Simplification",
    author = "Devaraj, Ashwin  and
      Sheffield, William  and
      Wallace, Byron  and
      Li, Junyi Jessy",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.506/",
    doi = "10.18653/v1/2022.acl-long.506",
    pages = "7331--7345"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/