keyword
factuality evaluation
Factuality evaluation is the systematic process of determining whether text generated by artificial intelligence and natural language processing systems is truthful, accurate, and aligned with verifiable real-world knowledge or provided reference texts. In generative tasks such as question answering, summarization, and text simplification, models risk generating unsupported claims, omitting key facts, or introducing fabricated information. Factuality evaluation identifies and measures these discrepancies using various methodologies, ranging from human expert review to automated frameworks that leverage external knowledge retrieval, natural language inference, and cross-examination. Conducting these assessments is critical for mitigating misinformation and ensuring that machine-generated content is reliable and safe for deployment in high-stakes domains such as medicine, law, and education.
4 items

Evaluating Factuality in Text Simplification
Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li
Why you should read this
Presents a taxonomy of factual errors in text simplification, revealing that standard evaluation metrics fail to detect frequent information insertions, deletions, and substitutions in both benchmark datasets and model outputs.
Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models.
Added
2026-10-03

Factuality of Large Language Models: A Survey
Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Nenkov Georgiev, Rocktim Jyoti Das, Preslav Nakov
Why you should read this
Synthesizes recent advances in large language model factuality across text and vision modalities by categorizing evaluation benchmarks, clarifying distinctions between factuality and hallucination, and analyzing error-mitigation techniques alongside calibration strategies.
Large language models (LLMs), especially when instruction-tuned for chat, have become part of our daily lives, freeing people from the process of searching, extracting, and integrating information from multiple sources by offering a straightforward answer to a variety of questions in a single place. Unfortunately, in many cases, LLM responses are factually incorrect, which limits their applicability in real-world scenarios. As a result, research on evaluating and improving the factuality of LLMs has attracted a lot of attention recently. In this survey, we critically analyze existing work with the aim to identify the major challenges and their associated causes, pointing out to potential solutions for improving the factuality of LLMs, and analyzing the obstacles to automated factuality evaluation for open-ended text generation. We further offer an outlook on where future research should go.
Added
2026-09-26

ExpertQA: Expert-Curated Questions and Attributed Answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth
Why you should read this
Presents a high-quality benchmark of 2,177 domain-specific questions across 32 fields to evaluate the factuality and citation quality of language model responses through direct assessment and revision by verified domain experts.
As language models are adopted by a more sophisticated and diverse set of users, the importance of guaranteeing that they provide factually correct information supported by verifiable sources is critical across fields of study. This is especially the case for high-stakes fields, such as medicine and law, where the risk of propagating false information is high and can lead to undesirable societal consequences. Previous work studying attribution and factuality has not focused on analyzing these characteristics of language model outputs in domain-specific scenarios. In this work, we conduct human evaluation of responses from a few representative systems along various axes of attribution and factuality, by bringing domain experts in the loop. Specifically, we collect expert-curated questions from 484 participants across 32 fields of study, and then ask the same experts to evaluate generated responses to their own questions. In addition, we ask experts to improve upon responses from language models. The output of our analysis is ExpertQA, a high-quality long-form QA dataset with 2177 questions spanning 32 fields, along with verified answers and attributions for claims in the answers.1
Added
2026-09-26

LM vs LM: Detecting Factual Errors via Cross Examination
Roi Cohen, May Hamri, Mor Geva, Amir Globerson
Why you should read this
Proposes a legal-inspired cross-examination framework where an examiner language model questions another model across multiple turns to uncover factual errors through generated inconsistencies without requiring external knowledge bases.
A prominent weakness of modern language models (LMs) is their tendency to generate factually incorrect text, which hinders their usability. A natural question is whether such factual errors can be detected automatically. Inspired by truth-seeking mechanisms in law, we propose a factuality evaluation framework for LMs that is based on cross-examination. Our key idea is that an incorrect claim is likely to result in inconsistency with other claims that the model generates. To discover such inconsistencies, we facilitate a multi-turn interaction between the LM that generated the claim and another LM (acting as an examiner) which introduces questions to discover inconsistencies. We empirically evaluate our method on factual claims made by multiple recent LMs on four benchmarks, finding that it outperforms existing methods and baselines, often by a large gap. Our results demonstrate the potential of using interacting LMs to capture factual errors.
Added
2026-09-26
