Built independently by an author, for readers. Read the story and support ChapterPal

keyword

automatic factuality estimation

Automatic factuality estimation is the computational process of evaluating whether machine-generated text or other natural language statements are accurate and consistent with verifiable facts or reference sources. Rather than relying entirely on manual human verification, this task employs automated metrics, natural language inference models, question-answering frameworks, or large language model evaluators to detect factual errors, hallucinations, and unsupported assertions. The process typically functions by decomposing a generated passage into individual atomic claims, cross-referencing those statements against provided context documents or retrieved external evidence, and assigning a veracity score or label based on how well the evidence supports each claim. By providing scalable, objective assessments of factual correctness and attribution, automatic factuality estimation plays a vital role in benchmarking, filtering, and improving the reliability of artificial intelligence systems in knowledge-intensive domains.

1 item

ExpertQA: Expert-Curated Questions and Attributed Answers

ExpertQA: Expert-Curated Questions and Attributed Answers

Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth

OrganizationsUniversity of PennsylvaniaUniversity of Washington

Why you should read this

Presents a high-quality benchmark of 2,177 domain-specific questions across 32 fields to evaluate the factuality and citation quality of language model responses through direct assessment and revision by verified domain experts.

As language models are adopted by a more sophisticated and diverse set of users, the importance of guaranteeing that they provide factually correct information supported by verifiable sources is critical across fields of study. This is especially the case for high-stakes fields, such as medicine and law, where the risk of propagating false information is high and can lead to undesirable societal consequences. Previous work studying attribution and factuality has not focused on analyzing these characteristics of language model outputs in domain-specific scenarios. In this work, we conduct human evaluation of responses from a few representative systems along various axes of attribution and factuality, by bringing domain experts in the loop. Specifically, we collect expert-curated questions from 484 participants across 32 fields of study, and then ask the same experts to evaluate generated responses to their own questions. In addition, we ask experts to improve upon responses from language models. The output of our analysis is ExpertQA, a high-quality long-form QA dataset with 2177 questions spanning 32 fields, along with verified answers and attributions for claims in the answers.1

Added

2026-09-26