Evaluating Factuality in Text Simplification
Ashwin DevarajWilliam SheffieldByron C. WallaceJunyi Jessy Li
Presents a taxonomy of factual errors in text simplification, revealing that standard evaluation metrics fail to detect frequent information insertions, deletions, and substitutions in both benchmark datasets and model outputs.
Automated text simplification systems aim to make complex information accessible to wider audiences, including laypeople, language learners, and individuals with cognitive or reading difficulties. However, these systems often introduce factual inaccuracies by adding unsupported statements, omitting crucial details, or altering original meanings. In high-stakes domains such as healthcare, providing readable but factually incorrect information can mislead users and cause severe harm. While factuality has been heavily studied in text summarization, its impact on text simplification has remained largely unexamined.
The article establishes a systematic framework to evaluate factual consistency in automated text simplification. Its primary objective is to quantify the prevalence of factual errors in standard benchmark datasets and modern artificial intelligence models, while assessing whether existing evaluation metrics reliably detect these errors.
To accomplish this, the authors developed an error typology categorizing factual distortions into information insertion, information deletion, and information substitution, each graded on a three-level severity scale. They conducted extensive human evaluations using crowdsourced annotators to examine standard simplification datasets, namely Wikilarge and Newsela, as well as outputs from multiple recurrent and Transformer-based models, including a fine-tuned T5 system. The team also benchmarked standard quality metrics against human judgments and trained an automated classification model supplemented with synthetic data to test the feasibility of automatic error detection.
The analysis yielded four major findings. First, standard human-authored reference datasets contain substantial factual flaws; in the Newsela dataset, over 80% of sentences exhibited deletion errors, with nearly 43% categorized as severe omissions that obscured the core message. Second, artificial intelligence models mirror and compound these issues; while modern Transformer models generate fewer unreadable outputs and fewer deletion errors than older recurrent architectures, pre-trained models introduce noticeable insertion errors on abstractive datasets and generate substitution errors at rates higher than those present in training data. Third, standard evaluation metrics fail to gauge accuracy: the primary industry metric, SARI, shows virtually no correlation with human factual judgments, and semantic similarity metrics identify deletions moderately well but fail to capture inappropriate insertions or substitutions. Fourth, a case study on medical text simplification showed that 30% of simplified paragraphs contained critical factual errors that distorted the clinical meaning.
These findings indicate that current simplification systems pose substantial operational and safety risks if deployed without human oversight in regulated or high-stakes environments. Furthermore, because standard development metrics reward surface-level vocabulary changes without penalizing factual distortions, relying on conventional benchmarks provides a false sense of model reliability and performance.
Organizations developing or deploying text simplification technologies should immediately cease relying solely on metrics like SARI for quality assurance. Next steps should include adopting factuality-focused verification pipelines, incorporating full document context to prevent out-of-context sentence distortion, and implementing strict human-in-the-loop review for critical communications. Future research must focus on building dedicated fact-checking models tailored to simplification tasks.
Confidence in these findings is supported by strong inter-annotator agreement across multiple dataset evaluations. However, the study's automated error detection model faced limitations due to small training sample sizes for rare error types, resulting in low detection performance for severe substitutions. Stakeholders should treat existing automated fact-checking models as experimental and exercise caution when applying automated simplification to technical domains.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Maynez et al. establish the hallucination types and metric shortcomings in generated text that this paper carries into the less-studied setting of text simplification.
- Paper: LENS: A Learnable Evaluation Metric for Text Simplification, Mounica Maddela et al. (2023). LENS continues the search for reliable simplification evaluation with a learned metric designed to align more closely with human judgments.
- Paper: BLESS: Benchmarking Large Language Models on Sentence Simplification, Tannon Kew et al. (2023). BLESS extends simplification evaluation to modern language models across domains, including medical text, and tests how well outputs preserve meaning.
