Language Models with Conformal Factuality Guarantees
Christopher MohriTatsunori Hashimoto
Proposes a conformal prediction framework that provides rigorous statistical correctness guarantees for black-box language model outputs by progressively pruning uncertain sub-claims while preserving most generated content.
Large language models are increasingly adopted across critical sectors, including healthcare, legal analysis, and customer service. However, their tendency to generate incorrect facts and plausible hallucinations creates significant operational, safety, and compliance risks. The article addresses the urgent problem of enforcing rigorous correctness guarantees for open-ended text outputs from black-box language models. Its main objective is to establish and evaluate a framework called conformal factuality, which provides mathematically guaranteed, user-specified accuracy levels while retaining useful information in the generated responses.
To achieve this, the article connects language modeling with conformal prediction, a statistical technique that provides coverage guarantees without restrictive assumptions. Rather than attempting to evaluate all possible text completions, the framework breaks an initial output into individual sub-claims, scores the uncertainty of each sub-claim, and removes the least reliable statements using a statistically determined threshold. The remaining verified claims are then merged into a final, coherent response that is naturally less specific but factual. The authors demonstrated this approach using GPT-4 across three benchmark datasets covering biographical generation, general question answering, and multi-step mathematical reasoning, calibrating the systems with small sets of manually verified examples.
Across all benchmarks, the framework achieved the targeted high-probability correctness guarantees while preserving the majority of the original content. In biographical generation, where the base model frequently hallucinated, factuality increased from approximately 30% to 80% while retaining about half of the original sub-claims. For general question answering, correctness improved from 78% to 93% by removing only about 25% of claims. In mathematical reasoning, accuracy rose from 75% to 95% while eliminating only about 10% of reasoning steps. Among the tested scoring methods, frequency scoring—which measures how consistently a claim appears across multiple generated alternatives—provided the most effective trade-off between correctness and detail.
These findings mean organizations can safely deploy generative models in high-stakes environments with predictable, statistical error bounds instead of relying on unverified model confidence. By strategically backing off to less specific claims rather than completely refusing to answer or providing incorrect facts, models maintain operational utility while substantially reducing legal and reputational risks. Decision-makers should consider integrating sub-claim verification pipelines into production workflows, utilizing frequency scoring to measure claim reliability, and establishing small, annotated calibration datasets for each targeted use case.
The statistical guarantees hold on average across inputs from a consistent distribution and rely on exchangeability between calibration examples and live queries. If user prompts shift significantly from the calibration data, empirical accuracy may fall below target levels. Additionally, small calibration sets introduce modest variation in realized coverage. Despite these standard statistical boundaries, the evidence provides strong confidence that the framework effectively eliminates hallucinations and offers a reliable foundation for responsible automated text generation.
- Paper: Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks, Akari Asai et al. (2022). Its evidence-aware generation framework establishes how retrieved passages can guide factual outputs, a key precursor to SENSE’s span-level evidence constraints.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). Its grammar-constrained decoding provides a concrete foundation for understanding how generation-time constraints can enforce valid outputs without fine-tuning.
- Paper: NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics, Ximing Lu et al. (2022). Its lookahead-based constrained decoding illustrates the decoding-time control strategy that SENSE adapts toward evidence-grounded factuality.
- Paper: Re2G: Retrieve, Rerank, Generate, Michael R. Glass et al. (2022). Its retrieve-rerank-generate pipeline supplies the retrieval-grounded generation setup that helps contextualize SENSE’s use of evidence.
- Paper: RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, Cheng Niu et al. (2024). Its word-level hallucination annotations provide a useful precedent for assessing unsupported claims in retrieval-grounded outputs.
No sufficiently relevant recommendations were found.
