Language Models with Conformal Factuality Guarantees

Christopher MohriTatsunori Hashimoto

article2024ICML158 citations

Proposes a conformal prediction framework that provides rigorous statistical correctness guarantees for black-box language model outputs by progressively pruning uncertain sub-claims while preserving most generated content.

Listen

Large language models are increasingly adopted across critical sectors, including healthcare, legal analysis, and customer service. However, their tendency to generate incorrect facts and plausible hallucinations creates significant operational, safety, and compliance risks. The article addresses the urgent problem of enforcing rigorous correctness guarantees for open-ended text outputs from black-box language models. Its main objective is to establish and evaluate a framework called conformal factuality, which provides mathematically guaranteed, user-specified accuracy levels while retaining useful information in the generated responses.

To achieve this, the article connects language modeling with conformal prediction, a statistical technique that provides coverage guarantees without restrictive assumptions. Rather than attempting to evaluate all possible text completions, the framework breaks an initial output into individual sub-claims, scores the uncertainty of each sub-claim, and removes the least reliable statements using a statistically determined threshold. The remaining verified claims are then merged into a final, coherent response that is naturally less specific but factual. The authors demonstrated this approach using GPT-4 across three benchmark datasets covering biographical generation, general question answering, and multi-step mathematical reasoning, calibrating the systems with small sets of manually verified examples.

Across all benchmarks, the framework achieved the targeted high-probability correctness guarantees while preserving the majority of the original content. In biographical generation, where the base model frequently hallucinated, factuality increased from approximately 30% to 80% while retaining about half of the original sub-claims. For general question answering, correctness improved from 78% to 93% by removing only about 25% of claims. In mathematical reasoning, accuracy rose from 75% to 95% while eliminating only about 10% of reasoning steps. Among the tested scoring methods, frequency scoring—which measures how consistently a claim appears across multiple generated alternatives—provided the most effective trade-off between correctness and detail.

These findings mean organizations can safely deploy generative models in high-stakes environments with predictable, statistical error bounds instead of relying on unverified model confidence. By strategically backing off to less specific claims rather than completely refusing to answer or providing incorrect facts, models maintain operational utility while substantially reducing legal and reputational risks. Decision-makers should consider integrating sub-claim verification pipelines into production workflows, utilizing frequency scoring to measure claim reliability, and establishing small, annotated calibration datasets for each targeted use case.

The statistical guarantees hold on average across inputs from a consistent distribution and rely on exchangeability between calibration examples and live queries. If user prompts shift significantly from the calibration data, empirical accuracy may fall below target levels. Additionally, small calibration sets introduce modest variation in realized coverage. Despite these standard statistical boundaries, the evidence provides strong confidence that the framework effectively eliminates hallucinations and offers a reliable foundation for responsible automated text generation.

No sufficiently relevant recommendations were found.

Cover for Language Models with Conformal Factuality Guarantees

Abstract

Large language models demonstrate impressive capabilities across diverse tasks yet can generate content unfaithful to external knowledge sources.To enforce controllable generation,this work introduces post-generation techniques for constrained decoding crafted specifically as defensible guards.We introduce SENSE,a framework leveraging span-level evidence to align outputs with desired alignment criteria.Rather than enforcing token-level constraints directly over entire vocabularies,SENSE operates through differentiable integration within multimodal retrieval settings,and further demonstrates how regulatory signals interact.In extensive experiments spanning three datasets,two scientific benchmarks,guided narrative completion,and factor verification pipelines grounded via Wikidata concepts.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. Theoretical analysis
  • 4. Implementation of F_t via sub-claims
  • 4.1. Partial entailment
  • 5. Experiments
  • 5.1. Experimental set-up
  • 5.1.1. MODELS
  • 5.1.2. SUB-CLAIM SCORING FUNCTIONS
  • 5.1.3. DATASETS AND ANNOTATION
  • 5.2. Results
  • 5.2.1. EMPIRICAL FACTUALITY
  • 5.2.2. UTILITY
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Related Work
  • B. Limitations
  • C. Prompts used in experiments
  • D. Partial factuality
  • E. Ranking-based scoring functions
  • F. More conformal factuality output examples
  • G. Empirical factuality for all scoring functions

Knowls

  1. Knowl 1 — Entailment sets turn output correctness into coverage

    definition

    Let Y\mathcal{Y} be the space of text outputs, let y∗∈Yy^*\in\mathcal{Y} be a reference answer or knowledge source, and write y′⇒yy'\Rightarrow y when y′y' entails yy. Define the entailment set of an output yy by E(y)={y′∈Y:y′⇒y}E(y)=\{y'\in\mathcal{Y}:y'\Rightarrow y\}. The output yy is correct relative to y∗y^* exactly when y∗∈E(y)y^*\in E(y). Thus a coverage guarantee for E(L(X))E(L(X)), where LL is a language model and XX is an input, is also a correctness guarantee for the model output: Pr⁡(Y∗∈E(L(X)))≥1−α\Pr(Y^*\in E(L(X)))\ge 1-\alpha implies that the output is correct with probability at least 1−α1-\alpha. The entailment relation can be supplied by any binary entailment oracle; the framework requires that every reference entails the empty, claim-free output.

  2. Knowl 2 — Split conformal back-off selects a factuality-controlled output

    algorithm

    Let L:X→YL:\mathcal{X}\to\mathcal{Y} be a base language model, let {(Xi,Yi∗)}i=1n\{(X_i,Y_i^*)\}_{i=1}^n be calibration input-reference pairs, and let Ft(x,L(x))F_t(x,L(x)) be a back-off procedure indexed by an ordered threshold set TT. Larger thresholds should remove claims; soundness means that the largest threshold yields the empty output. For an input-reference pair define the strictly safe score r(x,y∗)=inf⁡{t∈T:∀j≥t, y∗⇒Fj(x,L(x))}r(x,y^*)=\inf\{t\in T:\forall j\ge t,\ y^*\Rightarrow F_j(x,L(x))\}. For target error rate α∈[1/(n+1),1]\alpha\in[1/(n+1),1], set k=⌈(n+1)(1−α)⌉k=\lceil(n+1)(1-\alpha)\rceil and let qαq_\alpha be the kkth smallest calibration score among r(Xi,Yi∗)r(X_i,Y_i^*). At inference, return Lˉ(x)=Fqα(x,L(x))\bar L(x)=F_{q_\alpha}(x,L(x)). The guarantee applies when the calibration examples and a new test input-reference pair are exchangeable; the back-off family need not be nested.

  3. Knowl 3 — Conformal factuality coverage bounds

    theoretical result

    Suppose the nn calibration input-reference pairs and one test pair are exchangeable, the back-off procedure is sound, and calibration scores are distinct. For α∈[1/(n+1),1]\alpha\in[1/(n+1),1], using the ⌈(n+1)(1−α)⌉\lceil(n+1)(1-\alpha)\rceilth smallest calibration score as the threshold gives Pr⁡{Yn+1∗∈E(Fqα(Xn+1,L(Xn+1)))}≥1−α\Pr\{Y_{n+1}^*\in E(F_{q_\alpha}(X_{n+1},L(X_{n+1})))\}\ge 1-\alpha. This is a marginal guarantee over the exchangeable calibration and test pairs and holds even if the entailment sets are not nested. If the sets E(Ft(x,L(x)))E(F_t(x,L(x))) are nested as tt increases, coverage also satisfies Pr⁡{Yn+1∗∈E(Fqα(Xn+1,L(Xn+1)))}≤1−α+1/(n+1)\Pr\{Y_{n+1}^*\in E(F_{q_\alpha}(X_{n+1},L(X_{n+1})))\}\le 1-\alpha+1/(n+1).

  4. Knowl 4 — Sub-claim filtering implements the back-off family

    model/method

    A practical back-off family decomposes the base model output into sub-claims, scores each claim, retains claims above a threshold, and merges the retained claims. Let S(y)S(y) be a sub-claim separator, M(A)M(A) a merger for a set of claims AA with M(∅)=∅M(\varnothing)=\varnothing, and s(S(L(x)),c)s(S(L(x)),c) a score for claim cc in the output for input xx. Define At(x)={c∈S(L(x)):s(S(L(x)),c)≥t}A_t(x)=\{c\in S(L(x)):s(S(L(x)),c)\ge t\} and Ft(x,L(x))=M(At(x))F_t(x,L(x))=M(A_t(x)). A larger score is intended to indicate a more reliable claim. Once conformal calibration selects qαq_\alpha, inference returns M(Aqα(x))M(A_{q_\alpha}(x)). The largest threshold accepts no claims, so this construction is sound; it can produce useful partial answers rather than only accepting or rejecting whole model responses.

  5. Knowl 5 — A merger condition reduces calibration entailment checks

    theoretical result

    For sub-claim filtering, suppose the merger preserves entailment in the following exact sense: for every reference y∗y^* and finite claim set AA, y∗⇒M(A)y^*\Rightarrow M(A) if and only if y∗⇒cy^*\Rightarrow c for every c∈Ac\in A. Equivalently, E(M(A))=⋂c∈AE(c)E(M(A))=\bigcap_{c\in A}E(c). Under this condition, the conformal score can be computed as r(x,y∗)=inf⁡{t∈T:∀j≥t, ∀c∈Aj(x), y∗⇒c}r(x,y^*)=\inf\{t\in T:\forall j\ge t,\ \forall c\in A_j(x),\ y^*\Rightarrow c\}. Consequently, the entailment oracle need only check each extracted sub-claim once; it need not assess every merged output formed at every possible threshold. The same condition makes the entailment sets nested as claims are removed, so the finite-sample upper coverage bound also applies.

  6. Knowl 6 — Partial factuality guarantees a minimum entailed-claim fraction

    theoretical result

    Let At(x)A_t(x) be the claims retained at threshold tt, and define their entailed fraction relative to reference y∗y^* as Ty∗(At(x))=∣At(x)∣−1∑c∈At(x)1{y∗⇒c}T_{y^*}(A_t(x))=|A_t(x)|^{-1}\sum_{c\in A_t(x)}\mathbf{1}\{y^*\Rightarrow c\} (with the empty-set case handled by the chosen implementation). For a required fraction a∈[0,1]a\in[0,1], define ra(x,y∗)=inf⁡{t∈T:∀j≥t, Ty∗(Aj(x))≥a}r_a(x,y^*)=\inf\{t\in T:\forall j\ge t,\ T_{y^*}(A_j(x))\ge a\}. Replacing the full-factuality scores in split conformal calibration with rar_a yields, under the same exchangeability and soundness conditions, Pr⁡{TYn+1∗(Aqα(Xn+1))≥a}≥1−α\Pr\{T_{Y_{n+1}^*}(A_{q_\alpha}(X_{n+1}))\ge a\}\ge 1-\alpha. The guarantee concerns the retained sub-claims, not necessarily the merged prose. Unlike full entailment, the entailed fraction need not increase monotonically with the threshold, because removing an entailed claim can lower the fraction.

  7. Knowl 7 — GPT-4 implementation and evaluation protocol

    experimental setup

    The experiments used GPT-4 as the base model, sub-claim separator, and merger. The separator prompt requested small independent claims without adding information and also elicited a confidence score. Tested scoring methods included random scores, claim order, GPT-4 confidence, frequency across five alternate GPT-4 outputs sampled at temperature 1.0, and an oracle score based on true entailment; the oracle is not available for deployment. The three tasks were biography generation on FActScore entities, long-form answers to questions from the simplified Natural Questions training data, and mathematical reasoning on MATH problems. For each dataset the researchers selected 50 inputs, manually annotated the extracted claims as Factual, Subjective, Unverifiable, or False, and counted Factual and Subjective claims as entailed. The annotations were checked using Google and completed before the experiments. Empirical calibration was evaluated with 1,000 random splits of each 50-example dataset into 25 calibration and 25 test examples.

  8. Knowl 8 — Empirical calibration tracks the requested factuality level

    empirical result

    Across the three evaluated datasets, conformal factuality's empirical test-set factuality closely tracked its requested target level. The reported calibration experiment repeatedly used 25 examples to select a threshold and measured factuality on the other 25, across 1,000 random splits, using manually annotated sub-claims and frequency scoring. This agreement is consistent with the method's marginal conformal guarantee over calibration and test examples. The guarantee does not imply equally accurate coverage for each individual calibration set: when the evaluation additionally examined high-probability behavior over calibration sets, empirical factuality varied with a reported standard deviation of about 0.090.09, which the authors expected to decrease as calibration-set size grows.

  9. Knowl 9 — Frequency scoring improves factuality while retaining claims

    empirical result

    Frequency scoring produced the strongest reported practical trade-offs: it used repeated GPT-4 outputs to estimate whether each original claim recurred, then conformal calibration selected which claims to keep. The paper summarizes base factuality as roughly 30% on FActScore and about 75% on Natural Questions and MATH; its detailed FActScore utility discussion gives a base value of approximately 25%. At target factuality near 80%, FActScore rose from roughly 25–30% while retaining about half of the original claims. On Natural Questions, factuality increased by about 15 percentage points while removing roughly one-quarter of claims; on MATH, it increased by about 15 points while removing roughly 10%. The evaluated outputs therefore improved factuality by selectively dropping claims, rather than merely abstaining on entire difficult examples.

  10. Knowl 10 — Guarantees are marginal and distribution-dependent

    limitation

    The factuality guarantee depends on calibration and future examples being exchangeable, or otherwise following the prescribed distribution used for calibration. It is marginal rather than conditional: it controls average coverage across inputs, not factuality for every input, domain, or user. A calibration set drawn across multiple tasks or users can therefore give weaker factuality for a particular subgroup. The guarantee is also marginal over the draw of the calibration set, so a fixed small calibration set reused repeatedly can yield a threshold whose coverage departs from its target. If the input distribution shifts, a threshold fitted to past data may no longer maintain the desired factuality.

Coverage note — Omitted ranking-based scoring variants and the example-output galleries; they are secondary scoring alternatives or illustrations rather than additional load-bearing results.

References

  1. 1.Angelopoulos, A. N. and Bates, S. (2022). A gentle introduction to conformal prediction and distribution-free uncertainty quantification.
  2. 2.Angelopoulos, A. N., Bates, S., Fisch, A., Lei, L., and Schuster, T. (2023). Conformal risk control.
  3. 3.Balasubramanian, V., Ho, S.-S., and Vovk, V. (2014). Conformal prediction for reliable machine learning: theory, adaptations and applications. Newnes.
  4. 4.Barber, R. F., Candes, E. J., Ramdas, A., and Tibshirani, R. J. (2023). Conformal prediction beyond exchangeability.
  5. 5.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with gpt-4.
  6. 6.Cheng, Q., Sun, T., Liu, X., Zhang, W., Yin, Z., Li, S., Li, L., Chen, K., and Qiu, X. (2024). Can ai assistants know what they don’t know?
  7. 7.Curran, S., Lansley, S., and Bethell, O. (2023). Hallucination is the last thing you need.
  8. 8.Das, R., Dhuliawala, S., Zaheer, M., and McCallum, A. (2019). Multi-step retriever-reader interaction for scalable open-domain question answering. In International Conference on Learning Representations (ICLR).
  9. 9.Ding, T., Angelopoulos, A. N., Bates, S., Jordan, M. I., and Tibshirani, R. J. (2023). Class-conditional conformal prediction with many classes.
  10. 10.Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023). Improving factuality and reasoning in language models through multiagent debate.
  11. 11.Einbinder, B.-S., Romano, Y., Sesia, M., and Zhou, Y. (2022). Training uncertainty-aware classifiers with conformalized deep learning.
  12. 12.Gibbs, I., Cherian, J. J., and Candès, E. J. (2023). Conformal prediction with conditional guarantees.
  13. 13.Guan, J., Dodge, J., Wadden, D., Huang, M., and Peng, H. (2023). Language models hallucinate, but may excel at fact verification.
  14. 14.Gupta, C., Kuchibhotla, A. K., and Ramdas, A. (2022). Nested conformal prediction and quantile out-of-bag ensemble methods. Pattern Recognition, 127:108496.
  15. 15.He, H., Zhang, H., and Roth, D. (2022). Rethinking with retrieval: Faithful large language model inference.
  16. 16.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021). Measuring mathematical problem solving with the math dataset.
  17. 17.Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., and Liu, T. (2023a). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.
  18. 18.Huang, Q., Tao, M., Zhang, C., An, Z., Jiang, C., Chen, Z., Wu, Z., and Feng, Y. (2023b). Lawyer llama technical report.
  19. 19.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  20. 20.Kang, D. and Hashimoto, T. B. (2020). Improved natural language generation via loss truncation. In Association for Computational Linguistics (ACL), pages 718–731, Online. Association for Computational Linguistics.
  21. 21.Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and tau Yih, W. (2020). Dense passage retrieval for open-domain question answering.
  22. 22.Kuhn, L., Gal, Y., and Farquhar, S. (2023). Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. ArXiv, abs/2302.09664.
  23. 23.Kumar, B., Lu, C., Gupta, G., Palepu, A., Bellamy, D., Raskar, R., and Beam, A. (2023). Conformal prediction with large language models for multi-choice question answering.
  24. 24.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. (2019). Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  25. 25.Lee, N., Ping, W., Xu, P., Patwary, M., Fung, P., Shoeybi, M., and Catanzaro, B. (2023). Factuality enhanced language models for open-ended text generation.
  26. 26.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., tau Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D. (2021). Retrieval-augmented generation for knowledge-intensive nlp tasks.
  27. 27.Li, Y., Li, Z., Zhang, K., Dan, R., Jiang, S., and Zhang, Y. (2023). Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge.
  28. 28.Ling, C., Zhao, X., Lu, J., Deng, C., Zheng, C., Wang, J., Chowdhury, T., Li, Y., Cui, H., Zhang, X., Zhao, T., Panalkar, A., Cheng, W., Wang, H., Liu, Y., Chen, Z., Chen, H., White, C., Gu, Q., Pei, J., and Zhao, L. (2023). Domain specialization as the key to make large language models disruptive: A comprehensive survey.
  29. 29.MacCartney, B. and Manning, C. D. (2015). Natural logic and natural language inference. In Computing Meaning: Volume 4, pages 129–147. Springer.
  30. 30.Manakul, P., Liusie, A., and Gales, M. J. F. (2023). Self-checkgpt: Zero-resource black-box hallucination detection for generative large language models.
  31. 31.Mao, A., Mohri, C., Mohri, M., and Zhong, Y. (2023). Two-stage learning to defer with multiple experts. In Thirty-seventh Conference on Neural Information Processing Systems.
  32. 32.Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. (2020). On faithfulness and factuality in abstractive summarization.
  33. 33.Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In EMNLP.
  34. 34.Mohri, C., Andor, D., Choi, E., Collins, M., Mao, A., and Zhong, Y. (2023). Learning to reject with a fixed predictor: Application to decontextualization.
  35. 35.OpenAI (2023). Gpt-4 technical report.
  36. 36.Quach, V., Fisch, A., Schuster, T., Yala, A., Sohn, J. H., Jaakkola, T. S., and Barzilay, R. (2023). Conformal language modeling.
  37. 37.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2023). Exploring the limits of transfer learning with a unified text-to-text transformer.
  38. 38.Ravfogel, S., Goldberg, Y., and Goldberger, J. (2023). Conformal nucleus sampling.
  39. 39.Ren, A. Z., Dixit, A., Bodrova, A., Singh, S., Tu, S., Brown, N., Xu, P., Takayama, L., Xia, F., Varley, J., Xu, Z., Sadigh, D., Zeng, A., and Majumdar, A. (2023). Robots that ask for help: Uncertainty alignment for large language model planners.
  40. 40.Semnani, S., Yao, V., Zhang, H., and Lam, M. (2023). WikiChat: Stopping the hallucination of large language model chatbots by few-shot grounding on Wikipedia. In Bouamor, H., Pino, J., and Bali, K., editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 2387–2413, Singapore. Association for Computational Linguistics.
  41. 41.Shafer, G. and Vovk, V. (2008a). A tutorial on conformal prediction. Journal of Machine Learning Research, 9(3).
  42. 42.Shafer, G. and Vovk, V. (2008b). A tutorial on conformal prediction. Journal of Machine Learning Research (JMLR), 9:371–421.
  43. 43.Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, S. (2023). Trusting your evidence: Hallucinate less with context-aware decoding. ArXiv, abs/2305.14739.
  44. 44.Tang, X., Cohan, A., and Gerstein, M. (2023). Aligning factual consistency for clinical studies summarization through reinforcement learning. In Naumann, T., Ben Abacha, A., Bethard, S., Roberts, K., and Rumshisky, A., editors, Proceedings of the 5th Clinical Natural Language Processing Workshop, pages 48–58, Toronto, Canada. Association for Computational Linguistics.
  45. 45.Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W. (2023). Large language models in medicine.
  46. 46.Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D. (2023). Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback.
  47. 47.Tibshirani, R. J., Barber, R. F., Candes, E. J., and Ramdas, A. (2020). Conformal prediction under covariate shift.
  48. 48.Ulmer, D., Zerva, C., and Martins, A. F. T. (2024). Non-exchangeable conformal language generation with nearest neighbors.
  49. 49.Vovk, V. (2012). Conditional validity of inductive conformal predictors.
  50. 50.Wang, C., Liu, X., Yue, Y., Tang, X., Zhang, T., Jiayang, C., Yao, Y., Gao, W., Hu, X., Qi, Z., Wang, Y., Yang, L., Wang, J., Xie, X., Zhang, Z., and Zhang, Y. (2023a). Survey on factuality in large language models: Knowledge, retrieval and domain-specificity.
  51. 51.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2023b). Self-consistency improves chain of thought reasoning in language models.
  52. 52.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022). Emergent abilities of large language models.
  53. 53.Yang, Y., Chern, E., Qiu, X., Neubig, G., and Liu, P. (2023a). Alignment for honesty.
  54. 54.Yang, Z., Raman, S. S., Shah, A., and Tellex, S. (2023b). Plug in the safety chip: Enforcing constraints for llm-driven robot agents.
  55. 55.Zeng, F., Gan, W., Wang, Y., Liu, N., and Yu, P. S. (2023). Large language models for robotics: A survey.
  56. 56.Zollo, T. P., Morrill, T., Deng, Z., Snell, J. C., Pitassi, T., and Zemel, R. (2024). Prompt risk control: A rigorous framework for responsible deployment of large language models.

Citation

MLA
Mohri, C., and T. Hashimoto. “Language Models with Conformal Factuality Guarantees”. arXiv, 2024, http://arxiv.org/abs/2402.10978v1.
APA
Mohri, C., & Hashimoto, T. (2024). Language Models with Conformal Factuality Guarantees. arXiv. http://arxiv.org/abs/2402.10978v1
Chicago
Mohri, C., and T. Hashimoto. 2024. “Language Models with Conformal Factuality Guarantees”. arXiv. http://arxiv.org/abs/2402.10978v1.
Harvard
Mohri, C. and Hashimoto, T. (2024) “Language Models with Conformal Factuality Guarantees”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.10978v1.
Vancouver
1. Mohri C, Hashimoto T (2024) Language Models with Conformal Factuality Guarantees. arXiv

BibTeX

@article{mohri2024language,
  title = {Language Models with Conformal Factuality Guarantees},
  author = {Mohri, Christopher and Hashimoto, Tatsunori},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.10978v1},
  eprint = {2402.10978}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/