Linguistic Calibration of Long-Form Generations
Neil BandXuechen LiTengyu MaTatsunori Hashimoto
Proposes a decision-theoretic training framework that combines supervised fine-tuning and reinforcement learning to teach language models to express calibrated verbal confidence statements across long-form text, significantly improving downstream user decision-making without sacrificing generation accuracy.
Language models are increasingly used to inform consequential decisions in domains such as medicine and law. However, these models frequently generate inaccurate information with complete, unearned confidence—a phenomenon known as hallucination. When models hallucinate confidently, they mislead human users into making suboptimal or risky choices. Existing calibration techniques that adjust model probabilities typically only apply to multiple-choice tasks or short, single-sentence answers, leaving open-ended, long-form generations uncalibrated and prone to overconfident errors.
The article demonstrates an end-to-end framework to achieve linguistic calibration in long-form text generations. It establishes an approach where a model explicitly states its degree of certainty directly in natural language (e.g., "I estimate a 70% chance that..."), ensuring that downstream human users can make well-calibrated probabilistic forecasts and optimal decisions.
To accomplish this without requiring expensive human feedback during training, the authors implemented a decision-theoretic framework combined with a two-stage training pipeline. First, they trained a base open-source model (Llama 2 7B) using summary distillation, prompting a large model to summarize multiple outputs into a single consensus paragraph containing explicit verbal confidence markers. Second, they applied reinforcement learning with Proximal Policy Optimization using a proper scoring rule (negative log loss) as the reward. Instead of optimizing abstract text quality, the reward directly scored how well an automated surrogate reader could form calibrated predictions on related questions after reading the generated passage. The approach was evaluated on over 10,000 question-answering examples across several benchmarks (TriviaQA, Jeopardy, SciQ, BioASQ), a 500-entity biography generation task, and human reader studies involving over 1,000 annotated evaluations.
The primary finding is that the linguistically calibrated model achieved substantially better expected calibration error (ECE) than baseline models while matching or exceeding their factual accuracy. On the core TriviaQA benchmark, the calibrated model lowered ECE to 0.108 compared to 0.367 for a factuality-focused reinforcement learning baseline, while maintaining an accuracy of roughly 65%. In human reader evaluations, the calibrated model reduced calibration error from 0.404 to 0.116. Furthermore, the model exhibited strong zero-shot generalization across domain shifts: without any task-specific retraining, it outperformed baselines in calibration error on complex scientific exam questions (SciQ ECE of 0.213 vs. 0.439) and expert biomedical questions (BioASQ ECE of 0.342 vs. 0.620). Finally, in long-form person biography generation, the calibrated model achieved a claim-level calibration error of 0.266 and 46.77% accuracy, outperforming both factuality-tuned baselines and untuned proprietary models.
These results demonstrate that language models do not need to choose between abstention and overconfident inaccuracy; instead, they can communicate uncertainty transparently across multi-claim texts. In practical settings, this significantly lowers decision risk and improves safety, as human operators can readily identify when a model's claims are tentative versus highly reliable. Notably, the findings reveal that a relatively small, open-source 7-billion parameter model can achieve calibration levels comparable to much larger commercial models like GPT-4 when trained with a decision-focused objective.
Organizations developing or deploying language model copilots should adopt decision-based calibration objectives and incorporate verbalized confidence scores rather than relying solely on factuality tuning or binary abstention. Future work should focus on curating diverse, real-world decision-making datasets to train surrogate readers and investigating how human users interpret nuanced or ambiguous linguistic confidence phrases. Users and stakeholders should note that while this method significantly reduces overconfidence, verbalized probability statements remain model estimates and should serve to assist human judgment rather than replace domain expertise.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). This study establishes how language models can estimate whether their own answers are correct, a foundation for the source’s approach to expressing calibrated confidence in generated claims.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Its account of confidence calibration and expected calibration error supplies the core statistical vocabulary needed to interpret the source’s calibration objectives and results.
No sufficiently relevant recommendations were found.
