Linguistic Calibration of Long-Form Generations

Neil BandXuechen LiTengyu MaTatsunori Hashimoto

article2024ICML72 citations

Proposes a decision-theoretic training framework that combines supervised fine-tuning and reinforcement learning to teach language models to express calibrated verbal confidence statements across long-form text, significantly improving downstream user decision-making without sacrificing generation accuracy.

Listen

Language models are increasingly used to inform consequential decisions in domains such as medicine and law. However, these models frequently generate inaccurate information with complete, unearned confidence—a phenomenon known as hallucination. When models hallucinate confidently, they mislead human users into making suboptimal or risky choices. Existing calibration techniques that adjust model probabilities typically only apply to multiple-choice tasks or short, single-sentence answers, leaving open-ended, long-form generations uncalibrated and prone to overconfident errors.

The article demonstrates an end-to-end framework to achieve linguistic calibration in long-form text generations. It establishes an approach where a model explicitly states its degree of certainty directly in natural language (e.g., "I estimate a 70% chance that..."), ensuring that downstream human users can make well-calibrated probabilistic forecasts and optimal decisions.

To accomplish this without requiring expensive human feedback during training, the authors implemented a decision-theoretic framework combined with a two-stage training pipeline. First, they trained a base open-source model (Llama 2 7B) using summary distillation, prompting a large model to summarize multiple outputs into a single consensus paragraph containing explicit verbal confidence markers. Second, they applied reinforcement learning with Proximal Policy Optimization using a proper scoring rule (negative log loss) as the reward. Instead of optimizing abstract text quality, the reward directly scored how well an automated surrogate reader could form calibrated predictions on related questions after reading the generated passage. The approach was evaluated on over 10,000 question-answering examples across several benchmarks (TriviaQA, Jeopardy, SciQ, BioASQ), a 500-entity biography generation task, and human reader studies involving over 1,000 annotated evaluations.

The primary finding is that the linguistically calibrated model achieved substantially better expected calibration error (ECE) than baseline models while matching or exceeding their factual accuracy. On the core TriviaQA benchmark, the calibrated model lowered ECE to 0.108 compared to 0.367 for a factuality-focused reinforcement learning baseline, while maintaining an accuracy of roughly 65%. In human reader evaluations, the calibrated model reduced calibration error from 0.404 to 0.116. Furthermore, the model exhibited strong zero-shot generalization across domain shifts: without any task-specific retraining, it outperformed baselines in calibration error on complex scientific exam questions (SciQ ECE of 0.213 vs. 0.439) and expert biomedical questions (BioASQ ECE of 0.342 vs. 0.620). Finally, in long-form person biography generation, the calibrated model achieved a claim-level calibration error of 0.266 and 46.77% accuracy, outperforming both factuality-tuned baselines and untuned proprietary models.

These results demonstrate that language models do not need to choose between abstention and overconfident inaccuracy; instead, they can communicate uncertainty transparently across multi-claim texts. In practical settings, this significantly lowers decision risk and improves safety, as human operators can readily identify when a model's claims are tentative versus highly reliable. Notably, the findings reveal that a relatively small, open-source 7-billion parameter model can achieve calibration levels comparable to much larger commercial models like GPT-4 when trained with a decision-focused objective.

Organizations developing or deploying language model copilots should adopt decision-based calibration objectives and incorporate verbalized confidence scores rather than relying solely on factuality tuning or binary abstention. Future work should focus on curating diverse, real-world decision-making datasets to train surrogate readers and investigating how human users interpret nuanced or ambiguous linguistic confidence phrases. Users and stakeholders should note that while this method significantly reduces overconfidence, verbalized probability statements remain model estimates and should serve to assist human judgment rather than replace domain expertise.

arXiv: 2404.00474
  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). This study establishes how language models can estimate whether their own answers are correct, a foundation for the source’s approach to expressing calibrated confidence in generated claims.
  • Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Its account of confidence calibration and expected calibration error supplies the core statistical vocabulary needed to interpret the source’s calibration objectives and results.

No sufficiently relevant recommendations were found.

Cover for Linguistic Calibration of Long-Form Generations

Abstract

Language models (LMs) may lead their users to make suboptimal downstream decisions when they confidently hallucinate. This issue can be mitigated by having the LM verbally convey the probability that its claims are correct, but existing models cannot produce long-form text with calibrated confidence statements. Through the lens of decision-making, we define linguistic calibration for long-form generations: an LM is linguistically calibrated if its generations enable its users to make calibrated probabilistic predictions. This definition enables a training framework where a supervised finetuning step bootstraps an LM to emit long-form generations with confidence statements such as “I estimate a 30% chance of...” or “I am certain that...”, followed by a reinforcement learning step which rewards generations that enable a user to provide calibrated answers to related questions. We linguistically calibrate Llama 2 7B and find in automated and human evaluations of long-form generations that it is significantly more calibrated than strong finetuned factuality baselines with comparable accuracy. These findings generalize under significant domain shifts to scientific and biomedical questions and to an entirely held-out person biography generation task. Our results demonstrate that long-form generations may be calibrated end-to-end by constructing an objective in the space of the predictions that users make in downstream decision-making.

Table of Contents

  • 1 Introduction
  • 2 Setup
  • 2.1 Linguistic Calibration of Long-Form Generations
  • 2.2 From Calibration to Optimal Decisions
  • 2.3 Training Objective for Linguistic Calibration
  • 3 Method
  • 3.1 Generating Synthetic Supervision for Long-Form Calibration
  • 3.2 Summary Distillation
  • 3.3 Decision-Based RL
  • 3.4 Implementation
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Linguistic Calibration using Question-Answering Datasets
  • 4.3 Zero-Shot Generalization to a Biography Generation Task
  • 5 Related Work
  • 6 Limitations, Future Work, and Conclusions
  • 7 Acknowledgements
  • References
  • A Additional Results
  • A.1 Codebase
  • A.2 Additional Baselines
  • A.3 TriviaQA: Full Accuracy-ECE Frontier
  • A.4 TriviaQA: Additional Reliability Diagrams
  • A.5 Jeopardy: Full Accuracy-ECE Frontier
  • A.6 Jeopardy: All Reliability Diagrams
  • A.7 SciQ: Full Accuracy-ECE Frontier
  • A.8 SciQ: All Reliability Diagrams
  • A.9 BioASQ Task B: Full Accuracy-ECE Frontier
  • A.10 BioASQ Task B: All Reliability Diagrams
  • A.11 Person Biography Generation Frontier
  • A.12 Tabular Results
  • A.13 Qualitative Examples
  • B Benefits of Linguistic Calibration for Decision-making
  • B.1 Review of the LC Objective
  • B.2 Decision Calibration
  • B.3 Linguistic Calibration and Optimal Decision-making
  • C Training Framework
  • C.1 Regularized Linguistic Calibration Objective
  • C.2 Proof: Regularized Objective is Strictly Proper
  • C.3 Additional Details on Training Framework
  • D Evaluation Framework
  • D.1 Simulated Evaluation
  • D.2 FactScore-Based Evaluation Metric
  • D.3 Human Evaluation

Knowls

  1. Knowl 1 — Linguistic calibration of long-form generations

    definition

    Consider an LM that receives an open-ended query qq and generates long-form text zz. A user later encounters a related question xx with answer yy and uses a reader function f(x,z)f(x,z) to form a probability distribution over possible answers. The LM is linguistically ϕ\phi-calibrated with respect to this user if the reader’s forecasts satisfy classifier-calibration notion ϕ\phi on the joint distribution of (x,y,z)(x,y,z). For distribution calibration, this means E[1Y∣f(X,Z)=p]=p\mathbb{E}[\mathbf{1}_Y\mid f(X,Z)=p]=p for every forecast pp; classwise and confidence calibration impose their corresponding weaker conditions. The definition evaluates long-form text through the probabilistic forecasts it enables, rather than requiring a single confidence score for a generation containing many claims.

  2. Knowl 2 — Synthetic supervision and summary distillation

    model/method

    The training framework constructs synthetic examples linking an open-ended LM query to a related question-answer pair. It samples (x,y)(x,y) from a question-answering dataset and uses an API-based LLM to turn question xx into an open-ended query, for example, a request to write a paragraph about the question. For summary distillation, the base LM samples multiple long-form responses to each query; an API-based LLM summarizes recurring claims and their frequencies as natural-language confidence statements, including numerical or verbal uncertainty. The base LM is then supervised-finetuned on query-summary pairs to obtain a policy capable of expressing confidence. In the reported implementation, eight responses were sampled per query at temperature 0.70.7 and summarized at temperature 0.30.3.

  3. Knowl 3 — Proper-scoring objective for linguistic calibration

    equation

    Let p(q,x,y)p(q,x,y) be the training distribution over open-ended queries qq, related questions xx, and answers yy; let z∼π(⋅∣q)z\sim\pi(\cdot\mid q) be a long-form generation; and let f(x,z)yf(x,z)_y be the reader’s forecast probability for answer yy. The framework optimizes the reader’s logarithmic score:

    max⁡π  E(q,x,y)∼p,  z∼π(⋅∣q)[log⁡f(x,z)y].\displaystyle \max_{\pi}\;\mathbb{E}_{(q,x,y)\sim p,\;z\sim\pi(\cdot\mid q)}[\log f(x,z)_y].

    The logarithmic score is strictly proper: when the reader can represent the relevant conditional distribution and the objective is maximized, the reader’s forecast matches the ground-truth answer distribution. Under the paper’s assumption that the answer depends on the question but not additionally on the query or generation, this target is p(y∣x)p(y\mid x). The objective therefore trains generations by scoring the downstream forecasts they enable, rather than by directly scoring text.

  4. Knowl 4 — Surrogate-reader reinforcement learning

    model/method

    To optimize downstream forecasts without repeatedly querying human or API-based readers during reinforcement learning, the method trains a surrogate reader from forecasts made by an API-based simulated reader on generations sampled from the supervised-finetuned policy. The reader has two components: an answer extractor predicts the set of plausible answers appearing in the text, and a probability predictor assigns a probability to each extracted answer. In the implementation, the extractor was a RedPajama 3B model trained to produce answer lists, and the probability predictor was a Llama 2 7B model initialized from the confidence-finetuned policy and trained with binary cross-entropy. The surrogate assigns zero probability to answers outside the extracted set; its separate probability predictions may not sum to one.

    The policy is optimized with PPO using the surrogate forecast f~(x,z)\tilde f(x,z) and a KL penalty relative to the supervised-finetuned policy πSFT\pi_{\mathrm{SFT}}. For a ground-truth answer yy, the reward is log⁡f~(x,z)y−λ∣1−∑y′f~(x,z)y′∣+C\log\tilde f(x,z)_y-\lambda\left|1-\sum_{y'}\tilde f(x,z)_{y'}\right|+C, with λ=5\lambda=5 and C=5C=5 in the reported implementation. The normalization penalty restores strict propriety when λ>1\lambda>1; the KL regularizer discourages policy divergence from πSFT\pi_{\mathrm{SFT}}, with coefficient β=0.1\beta=0.1. The log probability of the ground-truth answer is clipped below at 10−410^{-4} for numerical stability. The policy is trained for 1,500 PPO steps.

  5. Knowl 5 — Calibration yields decision guarantees

    theoretical result

    Assume that, conditional on question XX, the ground-truth answer YY is independent of the open-ended query QQ and generated text ZZ: Y⊥(Q,Z)∣XY\perp(Q,Z)\mid X. If an LM is linguistically ϕ\phi-calibrated for a user, and the user’s loss function belongs to the loss family associated with calibration notion ϕ\phi, then the reader is decision-calibrated for that family. Consequently, the user’s Bayes decision rule has no greater expected loss than competing decision rules in that family, and the user can estimate the Bayes rule’s expected loss from their forecast as accurately as if they had access to ground-truth outcomes. Distribution calibration corresponds to guarantees over all loss functions; the weaker confidence-calibration condition yields guarantees for the restricted abstention setting, where a user may answer, abstain, or incur a loss for an incorrect answer.

  6. Knowl 6 — Simulated-reader QA results

    empirical result

    On question-answering evaluations with a simulated reader, LC RL improved expected calibration error (ECE; lower is better) over factuality-tuned reinforcement learning while matching or exceeding its accuracy (higher is better). TriviaQA was the in-distribution test set; Jeopardy, SciQ, and BioASQ Task B were out of distribution. The point estimates below report accuracy in percent and ECE, respectively:

    • TriviaQA: Factuality RL, 63.33% / 0.367; LC SFT, 60.98% / 0.166; LC RL, 64.74% / 0.108.
    • Jeopardy: Factuality RL, 64.05% / 0.359; LC SFT, 62.46% / 0.162; LC RL, 65.73% / 0.088.
    • SciQ: Factuality RL, 56.11% / 0.439; LC SFT, 54.87% / 0.313; LC RL, 56.85% / 0.213.
    • BioASQ Task B: Factuality RL, 38.04% / 0.620; LC SFT, 38.53% / 0.389; LC RL, 38.89% / 0.342.

    Thus, decision-based RL improved both calibration and accuracy over its confidence-capable SFT initialization on all four datasets, while LC RL outperformed the factuality-RL baseline on both reported measures.

  7. Knowl 7 — Human evaluation and transfer to biographies

    empirical result

    Human forecasts on TriviaQA also favored linguistic calibration: Factuality RL achieved 59.62% accuracy and 0.404 ECE, LC SFT achieved 57.44% and 0.163, and LC RL achieved 60.12% and 0.116. The reported 95% bootstrap confidence intervals were, respectively, accuracy/ECE: Factuality RL, 56.65–62.60% / 0.374–0.434; LC SFT, 54.37–60.52% / 0.135–0.192; and LC RL, 57.14–63.19% / 0.091–0.145.

    Without retraining on biographies, the models were evaluated on 500 people from a held-out Wikipedia-entity set. Accuracy and ECE were computed over atomic claims pooled across generated biographies. Factuality RL scored 39.86% accuracy and 0.601 ECE; LC SFT scored 44.49% and 0.301; LC RL scored 46.77% and 0.266. The 95% bootstrap intervals for LC RL were 45.50–48.08% accuracy and 0.253–0.280 ECE. These results show transfer from trivia-style QA training to a different long-form generation task, including claim-level confidence evaluation.

  8. Knowl 8 — Training and evaluation protocol

    experimental setup

    The authors linguistically calibrated Llama 2 7B using TriviaQA question-answer pairs. The TriviaQA training examples were divided into 10,000 SFT examples, 1,000 prompt-validation examples, 20,000 reward-model examples, 40,000 PPO examples, 1,000 PPO-validation examples, and 1,000 validation examples. QA evaluation used TriviaQA (11,313 examples), Jeopardy (10,638), SciQ (13,679), and BioASQ Task B (1,515); biography evaluation used 500 entities. QA questions were converted into open-ended paragraph-generation queries, and readers were instructed to base their forecasts on the generated text rather than their background knowledge. Paragraph decoding used temperature 0.30.3 at evaluation. Reader ECE was computed by grouping forecasts into bins according to their maximum answer probability and taking the bin-size-weighted absolute difference between mean confidence and accuracy; the authors used 20 bins for simulated QA evaluations and 10 for other evaluations. The comparison included both simulated-reader and human-reader forecasts on TriviaQA.

  9. Knowl 9 — Simulated-reader agreement with humans

    empirical result

    The paper compared its API-based simulated reader with crowdworker forecasts on TriviaQA to assess the scalability of automated evaluation. Across all evaluated examples, simulated-reader versus human agreement for LC RL was 0.626 correlation on confidence and 0.739 Cohen’s kappa on correctness; for LC SFT it was 0.618 and 0.748, respectively. For Factuality RL, the corresponding values were 0.993 and 0.741; its confidence agreement is high in part because the evaluation assigns non-confidence methods a fixed confidence of 1. On a 5% subset annotated by multiple crowdworkers, interannotator agreement for LC RL was 0.886 confidence correlation and 0.850 correctness kappa, and for LC SFT was 0.719 and 0.842. The results support using simulated readers for scalable evaluation, while showing that simulated and human confidence judgments are not identical.

  10. Knowl 10 — Limitations and scope

    limitation

    The training distribution comes from off-the-shelf question-answering datasets, which the authors use as a proxy for questions arising in real-world decisions; a more representative distribution may improve generalization to decision-making in the wild. The method is developed for models that can be finetuned and therefore does not directly calibrate API-only language models accessible solely through completions. Although the study finds transfer from surrogate-reader training to human forecasts, many generated confidence statements are relatively unambiguous numerical estimates; alignment between users’ interpretations of ambiguous verbal expressions remains an open issue. Calibration is not perfect, so the authors caution that calibrated models should aid rather than replace human judgment and expertise.

Coverage note — Additional baseline-by-baseline comparisons and illustrative generated paragraphs are omitted because they do not add a distinct central method, guarantee, or result beyond the reported comparisons and transfer findings.

References

  1. 1.Akyurek, A. F., Akyürek, E., Choshen, L., Wijaya, D., and Andreas, J. Deductive closure training of language models for coherence, accuracy, and updatability, 2024.
  2. 2.Anthropic. Model card and evaluations for claude models, 2023.
  3. 3.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., Das-Sarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., and Kaplan, J. A general language assistant as a laboratory for alignment, 2021.
  4. 4.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  5. 5.Band, N., Rudner, T. G. J., Feng, Q., Filos, A., Nado, Z., Dusenberry, M. W., Jerfel, G., Tran, D., and Gal, Y. Benchmarking bayesian deep learning on diabetic retinopathy detection tasks. In NeurIPS Datasets and Benchmarks Track, 2021.
  6. 6.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. Weight Uncertainty in Neural Networks. In Bach, F. and Blei, D. (eds.), PMLR, volume 37 of Proceedings of Machine Learning Research, pp. 1613–1622, Lille, France, 07–09 Jul 2015. PMLR.
  7. 7.Boyd, S. and Vandenberghe, L. Convex optimization. Cambridge University Press, 2004.
  8. 8.Brier, G. W. Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1):1–3, 1950. doi: 10.1175/1520-0493(1950)078⟨0001:VOFEIT⟩2.0.CO;2. URL https://journals.ametsoc.org/view/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml.
  9. 9.Brocker, J. Reliability, sufficiency, and the decomposition of proper scores. Quarterly Journal of the Royal Meteorological Society, 135(643):1512–1519, 2009. doi: https://doi.org/10.1002/qj.456. URL https://rmets.onlinelibrary.wiley.com/doi/abs/10.1002/qj.456.
  10. 10.Cheng, Q., Sun, T., Liu, X., Zhang, W., Yin, Z., Li, S., Li, L., He, Z., Chen, K., and Qiu, X. Can ai assistants know what they don’t know?, 2024.
  11. 11.Cover, T. M. and Thomas, J. A. Elements of Information Theory. Wiley, New York, 1991.
  12. 12.Cresswell, J. C., Sui, Y., Kumar, B., and Vouitsis, N. Conformal prediction sets improve human decision making, 2024.
  13. 13.Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. Ultrafeedback: Boosting language models with high-quality feedback, 2023.
  14. 14.Dahl, M., Magesh, V., Suzgun, M., and Ho, D. E. Large legal fictions: Profiling legal hallucinations in large language models, 2024.
  15. 15.Dao, T. Flashattention-2: Faster attention with better parallelism and work partitioning, 2023.
  16. 16.Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022.
  17. 17.Dawid, A. P. Present position and potential developments: Some personal views statistical theory the prequential approach. Journal of the Royal Statistical Society: Series A (General), 147(2):278–290, 1984.
  18. 18.DeGroot, M. H. and Fienberg, S. E. The comparison and evaluation of forecasters, 1983.
  19. 19.Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L. 8-bit optimizers via block-wise quantization, 2022.
  20. 20.Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. Alpacafarm: A simulation framework for methods that learn from human feedback. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=4hturzLcKX.
  21. 21.Evans, O., Cotton-Barratt, O., Finnveden, L., Bales, A., Balwit, A., Wills, P., Righetti, L., and Saunders, W. Truthful ai: Developing and governing ai that does not lie, 2021.
  22. 22.Foster, D. P. and Vohra, R. Regret in the on-line decision problem. Games and Economic Behavior, 29(1):7–35, 1999. ISSN 0899-8256. doi: https://doi.org/10.1006/game.1999.0740. URL https://www.sciencedirect.com/science/article/pii/S0899825699907406.
  23. 23.Foster, D. P. and Vohra, R. V. Asymptotic calibration. Biometrika, 85(2):379–390, 1998. ISSN 00063444. URL http://www.jstor.org/stable/2337364.
  24. 24.Gneiting, T. and Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477):359–378, 2007. doi: 10.1198/016214506000001437. URL https://doi.org/10.1198/016214506000001437.
  25. 25.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pp. 1321–1330. PMLR, 06–11 Aug 2017. URL https://proceedings.mlr.press/v70/guo17a.html.
  26. 26.Hebert-Johnson, U., Kim, M., Reingold, O., and Rothblum, G. Multicalibration: Calibration for the (Computationally-identifiable) masses. In Dy, J. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 1939–1948. PMLR, 10–15 Jul 2018. URL https://proceedings.mlr.press/v80/hebert-johnson18a.html.
  27. 27.Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., and Liu, T. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023.
  28. 28.Huang, Y., Liu, Y., Thirukovalluru, R., Cohan, A., and Dhingra, B. Calibrating long-form generations from large language models, 2024.
  29. 29.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12), mar 2023. ISSN 0360-0300. doi: 10.1145/3571730. URL https://doi.org/10.1145/3571730.
  30. 30.Jiang, Z., Araki, J., Ding, H., and Neubig, G. How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering. Transactions of the Association for Computational Linguistics, 9:962–977, 09 2021. ISSN 2307-387X. doi: 10.1162/tacl_a_00407. URL https://doi.org/10.1162/tacl_a_00407.
  31. 31.Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vancouver, Canada, July 2017. Association for Computational Linguistics.
  32. 32.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. Language models (mostly) know what they know, 2022.
  33. 33.Kaggle. 200,000+ jeopardy! questions, 2020. URL https://www.kaggle.com/datasets/tunguz/200000-jeopardy-questions/data.
  34. 34.Krithara, A., Nentidis, A., Bougiatiotis, K., and Paliouras, G. Bioasq-qa: A manually curated corpus for biomedical question answering. Scientific Data, 10(1):170, Mar 2023. ISSN 2052-4463. doi: 10.1038/s41597-023-02068-4. URL https://doi.org/10.1038/s41597-023-02068-4.
  35. 35.Kuhn, L., Gal, Y., and Farquhar, S. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation, 2023.
  36. 36.Kull, M. and Flach, P. Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration. In Appice, A., Rodrigues, P. P., Santos Costa, V., Soares, C., Gama, J., and Jorge, A. (eds.), Machine Learning and Knowledge Discovery in Databases, pp. 68–85, Cham, 2015. Springer International Publishing. ISBN 978-3-319-23528-8.
  37. 37.Kull, M., Perello Nieto, M., Kangsepp, M., Silva Filho, T., Song, H., and Flach, P. Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper_files/paper/2019/file/8ca01ea920679a0fe3728441494041b9-Paper.pdf.
  38. 38.Kumar, A., Liang, P. S., and Ma, T. Verified uncertainty calibration. Advances in Neural Information Processing Systems, 32, 2019.
  39. 39.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 6402–6413, 2017.
  40. 40.Lhoest, Q., Villanova del Moral, A., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., Davison, J., Saško, M., Chhablani, G., Malik, B., Brandeis, S., Le Scao, T., Sanh, V., Xu, C., Patry, N., McMillan-Major, A., Schmid, P., Gugger, S., Delangue, C., Matussiere, T., Debut, L., Bekman, S., Cistac, P., Goehringer, T., Mustar, V., Lagunas, F., Rush, A., and Wolf, T. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 175–184, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.emnlp-demo.21.
  41. 41.Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y. Holistic evaluation of language models, 2023.
  42. 42.Lin, S., Hilton, J., and Evans, O. Teaching models to express their uncertainty in words, 2022.
  43. 43.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  44. 44.Malinin, A. and Gales, M. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=jN5y-zb5Q7m.
  45. 45.Malinin, A., Band, N., Gal, Y., Gales, M., Ganshin, A., Chesnokov, G., Noskov, A., Ploskonosov, A., Prokhorenkova, L., Provilkov, I., Raina, V., Raina, V., Roginskiy, D., Shmatova, M., Tigas, P., and Yangel, B. Shifts: A Dataset of Real Distributional Shift Across Multiple Large-Scale Tasks. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021.
  46. 46.Mielke, S. J., Szlam, A., Dinan, E., and Boureau, Y.-L. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857–872, 2022. doi: 10.1162/tacl_a_00494. URL https://aclanthology.org/2022.tacl-1.50.
  47. 47.Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 12076–12100, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.741. URL https://aclanthology.org/2023.emnlp-main.741.
  48. 48.Minderer, M., Djolonga, J., Romijnders, R., Hubis, F., Zhai, X., Houlsby, N., Tran, D., and Lucic, M. Revisiting the calibration of modern neural networks. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 15682–15694. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper_files/paper/2021/file/8420d359404024567b5aefda1231af24-Paper.pdf.
  49. 49.Murphy, A. H. A new vector partition of the probability score. Journal of Applied Meteorology and Climatology, 12(4):595–600, 1973.
  50. 50.Nado, Z., Band, N., Collier, M., Djolonga, J., Dusenberry, M. W., Farquhar, S., Feng, Q., Filos, A., Havasi, M., Jenatton, R., Jerfel, G., Liu, J., Mariet, Z., Nixon, J., Padhy, S., Ren, J., Rudner, T. G. J., Sbahi, F., Wen, Y., Wenzel, F., Murphy, K., Sculley, D., Lakshminarayanan, B., Snoek, J., Gal, Y., and Tran, D. Uncertainty baselines: Benchmarks for uncertainty & robustness in deep learning, 2022.
  51. 51.Niculescu-Mizil, A. and Caruana, R. Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning, ICML ’05, pp. 625–632, New York, NY, USA, 2005. Association for Computing Machinery. ISBN 1595931805. doi: 10.1145/1102351.1102430. URL https://doi.org/10.1145/1102351.1102430.
  52. 52.OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, I., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, C., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mely, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B. Gpt-4 technical report, 2023.
  53. 53.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback, 2022. URL https://arxiv.org/abs/2203.02155.
  54. 54.Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., and Snoek, J. Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. In Advances in Neural Information Processing Systems 32. 2019.
  55. 55.Park, P. S., Goldstein, S., O’Gara, A., Chen, M., and Hendrycks, D. Ai deception: A survey of examples, risks, and potential solutions, 2023.
  56. 56.Platt, J. et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in large margin classifiers, 10(3):61–74, 1999.
  57. 57.Savage, L. J. Elicitation of personal probabilities and expectations. Journal of the American Statistical Association, 66(336):783–801, 1971. ISSN 01621459. URL http://www.jstor.org/stable/2284229.
  58. 58.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms, 2017.
  59. 59.Shrivastava, V., Liang, P., and Kumar, A. Llamas know what gpts don’t show: Surrogate models for confidence estimation, 2023.
  60. 60.Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W. Large language models in medicine. Nature Medicine, 29(8):1930–1940, Aug 2023. ISSN 1546-170X. doi: 10.1038/s41591-023-02448-8. URL https://doi.org/10.1038/s41591-023-02448-8.
  61. 61.Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5433–5442, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.330. URL https://aclanthology.org/2023.emnlp-main.330.
  62. 62.Tian, K., Mitchell, E., Yao, H., Manning, C. D., and Finn, C. Fine-tuning language models for factuality. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=WPZ2yPag4K.
  63. 63.together.ai. Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned & chat models, 2023. URL https://www.together.ai/blog/redpajama-models-v1.
  64. 64.Tran, D., Kirsch, A., Lakshminarayanan, B., Hu, H., Phan, D., Sculley, D., Snoek, J., Liu, J. Z., Ren, J., van Amersfoort, J., Han, K., Buchanan, E. K., Murphy, K. P., Collier, M., Dusenberry, M. W., Band, N., Thain, N., Jenatton, R., Rudner, T. G. J., Gal, Y., Nado, Z., Mariet, Z. E., Wang, Z., and Ghahramani., Z. Plex: Towards reliability using pretrained large model extensions. In ICML 2022 Workshop on Pre-training, 2022.
  65. 65.Wallsten, T. Measuring Vague Uncertainties and Understanding Their Use in Decision Making, pp. 377–398. Measuring Vague Uncertainties and Understanding Their Use in Decision Making, 01 1990. ISBN 978-90-481-5785-3. doi: 10.1007/978-94-015-7873-8_15.
  66. 66.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models, 2022. URL https://arxiv.org/abs/2203.11171.
  67. 67.Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=1PL1NIMMrw.
  68. 68.Welbl, J., Liu, N. F., and Gardner, M. Crowdsourcing multiple choice science questions, 2017.
  69. 69.Whiting, M. E., Hugh, G., and Bernstein, M. S. Fair work: Crowd work minimum wage with one line of code. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):197–206, Oct. 2019. doi: 10.1609/hcomp.v7i1.5283. URL https://ojs.aaai.org/index.php/HCOMP/article/view/5283.
  70. 70.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.emnlp-demos.6.
  71. 71.Xiong, M., Hu, Z., Lu, X., LI, Y., Fu, J., He, J., and Hooi, B. Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=gjeQKFxFpZ.
  72. 72.Yang, Y., Chern, E., Qiu, X., Neubig, G., and Liu, P. Alignment for honesty, 2023.
  73. 73.Ye, F., Yang, M., Pang, J., Wang, L., Wong, D. F., Yilmaz, E., Shi, S., and Tu, Z. Benchmarking llms via uncertainty quantification, 2024.
  74. 74.Zadrozny, B. and Elkan, C. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Icml, volume 1, pp. 609–616, 2001.
  75. 75.Zhao, S. and Ermon, S. Right decisions from wrong predictions: A mechanism design alternative to individual calibration. In Banerjee, A. and Fukumizu, K. (eds.), Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, volume 130 of Proceedings of Machine Learning Research, pp. 2683–2691. PMLR, 13–15 Apr 2021. URL https://proceedings.mlr.press/v130/zhao21a.html.
  76. 76.Zhao, S., Kim, M., Sahoo, R., Ma, T., and Ermon, S. Calibrating predictions to decisions: A novel approach to multi-class calibration. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 22313–22324. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper_files/paper/2021/file/bbc92a647199b832ec90d7cf57074e9e-Paper.pdf.
  77. 77.Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S. Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023.
  78. 78.Zhou, K., Jurafsky, D., and Hashimoto, T. Navigating the grey area: How expressions of uncertainty and overconfidence affect language models. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5506–5524, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.335. URL https://aclanthology.org/2023.emnlp-main.335.

Citation

MLA
Band, N., et al. “Linguistic Calibration of Long-Form Generations”. arXiv, 2024, http://arxiv.org/abs/2404.00474v2.
APA
Band, N., Li, X., Ma, T., & Hashimoto, T. (2024). Linguistic Calibration of Long-Form Generations. arXiv. http://arxiv.org/abs/2404.00474v2
Chicago
Band, N., X. Li, T. Ma, and T. Hashimoto. 2024. “Linguistic Calibration of Long-Form Generations”. arXiv. http://arxiv.org/abs/2404.00474v2.
Harvard
Band, N. et al. (2024) “Linguistic Calibration of Long-Form Generations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.00474v2.
Vancouver
1. Band N, Li X, Ma T, Hashimoto T (2024) Linguistic Calibration of Long-Form Generations. arXiv

BibTeX

@article{band2024linguistic,
  title = {Linguistic Calibration of Long-Form Generations},
  author = {Band, Neil and Li, Xuechen and Ma, Tengyu and Hashimoto, Tatsunori},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.00474v2},
  eprint = {2404.00474}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/