Built independently by an author, for readers. Read the story and support ChapterPal

keyword

FACTSCORE dataset

The FACTSCORE dataset is an evaluation benchmark designed to assess the factual accuracy of long-form text generated by large language models, particularly in the domain of biographical text generation. The dataset consists of prompts requesting detailed biographies of various individuals alongside corresponding reference texts, such as Wikipedia articles, which serve as ground-truth knowledge sources. Within this framework, generated narratives are decomposed into discrete, independently verifiable atomic facts to determine the proportion of statements supported by the reference material. By offering claim-level factual breakdowns, the dataset serves as a standard resource for measuring factual precision, benchmarking hallucination rates, and evaluating long-form uncertainty quantification techniques across language models.

1 item

LUQ: Long-text Uncertainty Quantification for LLMs

LUQ: Long-text Uncertainty Quantification for LLMs

Caiqi Zhang, Fangyu Liu, Marco Basaldella, Nigel Collier

Why you should read this

Introduces LUQ, a sampling-based uncertainty quantification framework designed for long-form language model outputs, demonstrating strong negative correlation with factual errors and enabling multi-model ensembles that boost overall response factuality.

Large Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks. However, LLMs are also prone to generate nonfactual content. Uncertainty Quantification (UQ) is pivotal in enhancing our understanding of a model’s confidence on its generation, thereby aiding in the mitigation of nonfactual outputs. Existing research on UQ predominantly targets short text generation, typically yielding brief, word-limited responses. However, real-world applications frequently necessitate much longer responses. Our study first highlights the limitations of current UQ methods in handling long text generation. We then introduce LUQ with its two variations: LUQ-ATOMIC and LUQ-PAIR, a series of novel sampling-based UQ approaches specifically designed for long text. Our findings reveal that LUQ outperforms existing baseline methods in correlating with the model’s factuality scores (negative coefficient of -0.85 observed for Gemini Pro). To further improve the factuality of LLM responses, we propose LUQ-ENSEMBLE, a method that ensembles responses from multiple models and selects the response with the lowest uncertainty. The ensembling method greatly improves the response factuality upon the best standalone LLM.

Added

2026-10-02