Built independently by an author, for readers. Read the story and support ChapterPal

keyword

factuality evaluation

Factuality evaluation is the systematic process of determining whether text generated by artificial intelligence and natural language processing systems is truthful, accurate, and aligned with verifiable real-world knowledge or provided reference texts. In generative tasks such as question answering, summarization, and text simplification, models risk generating unsupported claims, omitting key facts, or introducing fabricated information. Factuality evaluation identifies and measures these discrepancies using various methodologies, ranging from human expert review to automated frameworks that leverage external knowledge retrieval, natural language inference, and cross-examination. Conducting these assessments is critical for mitigating misinformation and ensuring that machine-generated content is reliable and safe for deployment in high-stakes domains such as medicine, law, and education.

4 items

Evaluating Factuality in Text Simplification

Evaluating Factuality in Text Simplification

Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li

OrganizationsComputer ScienceLinguisticsMathematicsNortheastern UniversityUniversity of Texas at Austin

Why you should read this

Presents a taxonomy of factual errors in text simplification, revealing that standard evaluation metrics fail to detect frequent information insertions, deletions, and substitutions in both benchmark datasets and model outputs.

Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models.

Added

2026-10-03

Factuality of Large Language Models: A Survey

Factuality of Large Language Models: A Survey

Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Nenkov Georgiev, Rocktim Jyoti Das, Preslav Nakov

OrganizationsGoogleMohamed bin Zayed University of Artificial IntelligenceMonash UniversitySofia University “St. Kliment Ohridski”

Why you should read this

Synthesizes recent advances in large language model factuality across text and vision modalities by categorizing evaluation benchmarks, clarifying distinctions between factuality and hallucination, and analyzing error-mitigation techniques alongside calibration strategies.

Large language models (LLMs), especially when instruction-tuned for chat, have become part of our daily lives, freeing people from the process of searching, extracting, and integrating information from multiple sources by offering a straightforward answer to a variety of questions in a single place. Unfortunately, in many cases, LLM responses are factually incorrect, which limits their applicability in real-world scenarios. As a result, research on evaluating and improving the factuality of LLMs has attracted a lot of attention recently. In this survey, we critically analyze existing work with the aim to identify the major challenges and their associated causes, pointing out to potential solutions for improving the factuality of LLMs, and analyzing the obstacles to automated factuality evaluation for open-ended text generation. We further offer an outlook on where future research should go.

Added

2026-09-26

ExpertQA: Expert-Curated Questions and Attributed Answers

ExpertQA: Expert-Curated Questions and Attributed Answers

Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth

OrganizationsUniversity of PennsylvaniaUniversity of Washington

Why you should read this

Presents a high-quality benchmark of 2,177 domain-specific questions across 32 fields to evaluate the factuality and citation quality of language model responses through direct assessment and revision by verified domain experts.

As language models are adopted by a more sophisticated and diverse set of users, the importance of guaranteeing that they provide factually correct information supported by verifiable sources is critical across fields of study. This is especially the case for high-stakes fields, such as medicine and law, where the risk of propagating false information is high and can lead to undesirable societal consequences. Previous work studying attribution and factuality has not focused on analyzing these characteristics of language model outputs in domain-specific scenarios. In this work, we conduct human evaluation of responses from a few representative systems along various axes of attribution and factuality, by bringing domain experts in the loop. Specifically, we collect expert-curated questions from 484 participants across 32 fields of study, and then ask the same experts to evaluate generated responses to their own questions. In addition, we ask experts to improve upon responses from language models. The output of our analysis is ExpertQA, a high-quality long-form QA dataset with 2177 questions spanning 32 fields, along with verified answers and attributions for claims in the answers.1

Added

2026-09-26