keyword
FACTSCORE dataset
The FACTSCORE dataset is an evaluation benchmark designed to assess the factual accuracy of long-form text generated by large language models, particularly in the domain of biographical text generation. The dataset consists of prompts requesting detailed biographies of various individuals alongside corresponding reference texts, such as Wikipedia articles, which serve as ground-truth knowledge sources. Within this framework, generated narratives are decomposed into discrete, independently verifiable atomic facts to determine the proportion of statements supported by the reference material. By offering claim-level factual breakdowns, the dataset serves as a standard resource for measuring factual precision, benchmarking hallucination rates, and evaluating long-form uncertainty quantification techniques across language models.
1 item

