Calibration without Ground Truth

Yuqing KongMingyu SongYizhou WangYifan Wu

article2026arXiv1 citations

Develops an unsupervised post-processing framework grounded in economic theory that provably calibrates strong language models using weaker reference models via Bregman projection, achieving performance competitive with supervised methods without requiring labeled data.

Listen

Modern artificial intelligence models, especially large language models, frequently deliver highly accurate predictions while exhibiting severe overconfidence and poor calibration. When models state high confidence in flawed outputs, it creates significant operational risks in high-stakes environments such as legal, financial, and automated decision-making systems. Traditional recalibration techniques typically depend on labeled validation datasets; however, high-quality human-annotated data is becoming increasingly scarce and expensive. The article addresses this emerging bottleneck by developing a post-processing framework that improves the calibration of strong models without requiring ground-truth labels.

The primary objective of the article is to demonstrate that an accurate but miscalibrated model (such as an instruction-tuned model) can be strictly improved using only the unlabeled outputs of a weaker but better-calibrated reference model (such as a base pretrained model).

To accomplish this, the authors develop a mathematical framework connecting machine learning calibration to economic no-arbitrage conditions and information theory. The method estimates the joint distribution of predictions between the strong primary model and the weak reference model on unlabeled data, then applies a functional projection algorithm to align the strong model with the reference-compatible set. The authors validate the method through experiments across standard language model benchmarks—including MMLU-Redux (5,700 samples) and CommonSenseQA (1,221 samples)—evaluating multiple model families across diverse scales, including Qwen3 (0.6B to 14B), LLaMA-3 (1B to 8B), and Ministral-3 (3B to 14B).

The analysis establishes several key findings. First, the authors prove theoretically that a strict improvement in expected loss is always guaranteed if and only if the two models violate mutual calibration, meaning their implied probability distributions contradict each other. Second, across empirical benchmarks, the proposed label-free method achieved substantial reductions in calibration error compared to the uncalibrated instruction models, improving Brier Scores by over 8% and reducing Expected Calibration Error by more than 30%. Third, despite requiring no ground-truth labels, the framework performed competitively with established supervised baselines such as Temperature Scaling, Histogram Binning, and Isotonic Regression. Fourth, because the algorithm re-estimates the full probability distribution rather than merely scaling confidence numbers, it preserved the underlying predictive accuracy of the strong models and achieved modest accuracy gains in select cases (e.g., raising Ministral-3-8B accuracy on MMLU-Redux from 79.21% to 80.52%).

These findings have direct practical implications for deploying foundation models. Organizations can significantly reduce the risk of overconfident automated errors without incurring the financial and timeline costs of collecting extensive labeled datasets. By pairing aggressive fine-tuned models with conservative base model checkpoints, teams can safely deploy high-performing models with trustworthy confidence estimates. The results challenge the conventional assumption that model calibration requires supervised ground truth, offering a reliable path for post-training refinement.

Decision-makers and engineering teams should consider adopting this post-processing approach when deploying models in domains sensitive to probability assessment and risk. Before wide-scale rollout, organizations should conduct pilot evaluations to confirm that candidate reference models exhibit superior calibration relative to the primary models. Further development should explore adapting the framework from batch processing to streaming online environments and expanding its application to token-level free-form text generation.

The conclusions carry high confidence based on rigorous theoretical proofs and consistent empirical validation across multiple open-source model families. However, users should note two boundary conditions: the mathematical guarantees rely on the reference model being relatively well-calibrated, and estimating the joint distribution requires an adequate sample of unlabeled data, which may introduce minor estimation noise for rare events.

arXiv: 2601.19862

No sufficiently relevant recommendations were found.

Cover for Calibration without Ground Truth

Abstract

Villalobos et al. [2024] predict that publicly available human text will be exhausted within the next decade. Thus, improving models without access to ground-truth labels becomes increasingly important. We propose a label-free post-processing framework that improves a strong but miscalibrated model using a weaker yet better-calibrated reference. Our framework guarantees a strict performance improvement under any proper loss. Our approach is based on a characterization of when strict improvement is possible: when the strong and reference models are not mutually calibrated. We formalize this condition, connect it to arbitrage and no-trade results from economics, and develop an efficient Bregman projection algorithm that guarantees worst-case loss reduction without labels. Experiments on representative LLMs across varying scales demonstrate that our label-free method significantly reduces proper losses and calibration errors, achieving performance competitive with supervised baselines.

Table of Contents

  • 1 Introduction
  • 1.1 Our Contributions
  • 1.2 Related Work
  • 2 Problem Formulation
  • 2.1 Extension to General Elicitable Properties
  • 3 Solution and Theoretical Guarantees
  • 3.1 The Condition for Improvement: Mutual Calibration
  • 3.2 Algorithm: Projection to Reference-Compatible Set
  • 3.2.1 Step 1: The Reference-Compatible Set
  • 3.2.2 Step 2: Bregman Projection
  • 3.3 Theoretical Guarantees
  • 3.4 Extension to General Elicitable Properties
  • 4 Experimental Setup
  • 5 Results
  • 6 Discussion
  • References
  • A Proofs
  • A.1 More Preliminaries: Proper Loss and Bregman Divergence
  • A.2 Proof of Theorem
  • B Performance Guarantees under Imperfect Calibration of Reference
  • C Extension to Elicitable Properties
  • C.1 Proof of Observation
  • C.2 The Generalized Theorem
  • C.3 Proof of the Generalized Theorem
  • C.4 The Generalized Pythagorean Theorem in Non-smooth Case
  • C.5 Existence of ϵ\epsilon-net
  • D Mutual Calibration is Equivalent to No-Arbitrage
  • D.1 Mutual Calibratability = No-Arbitrage
  • D.2 A Quantified Relationship
  • E Experiment Details
  • E.1 Datasets Details
  • E.2 Implementation Details
  • E.2.1 Prompting Strategy and Confidence Extraction
  • E.2.2 Joint Distribution Estimation
  • E.3 Baselines Details
  • E.4 Evaluation Metrics Formula
  • F Additional Results
  • F.1 Comparison with Additional Supervised Baselines
  • F.2 Additional Results for Models with Varying Sizes

Knowls

  1. Knowl 1 — Strict improvement is possible exactly when the predictors are not mutually calibrated

    theoretical result

    Let Q1Q_1 be a primary predictor, Q0Q_0 a reference predictor, and YY a categorical outcome. Assume the joint distribution of (Q0,Q1)(Q_0,Q_1) is known, Q0Q_0 is calibrated, and the loss ℓ\ell is strictly proper with a differentiable, strictly convex generator. The predictors are mutually calibrated if some joint distribution of (Q0,Q1,Y)(Q_0,Q_1,Y) extending their observed joint distribution makes both predictors calibrated. For a fixed such joint distribution, let Hcomp\mathcal H_{\mathrm{comp}} be the transformations h:Q1↦Δ(Y)h:Q_1\mapsto\Delta(\mathcal Y) for which h(Q1)h(Q_1) and Q0Q_0 are mutually calibrated. Define the loss regret Dℓ(p,q)=EY∼p[ℓ(q,Y)−ℓ(p,Y)]D_\ell(p,q)=\mathbb E_{Y\sim p}[\ell(q,Y)-\ell(p,Y)], where p,qp,q are outcome distributions. A transformation that weakly improves expected loss for every outcome distribution consistent with the observed forecasts, and strictly improves it for at least one, exists if and only if Q0Q_0 and Q1Q_1 are not mutually calibrated. In that case, the projection h∗=arg⁡min⁡h∈HcompE[Dℓ(h(Q1),Q1)]h^*=\arg\min_{h\in\mathcal H_{\mathrm{comp}}}\mathbb E[D_\ell(h(Q_1),Q_1)] guarantees, for every consistent outcome distribution under which Q0Q_0 is calibrated, that E[ℓ(Q1,Y)]−E[ℓ(h∗(Q1),Y)]≥min⁡h∈HcompE[Dℓ(h(Q1),Q1)]>0\mathbb E[\ell(Q_1,Y)]-\mathbb E[\ell(h^*(Q_1),Y)]\geq\min_{h\in\mathcal H_{\mathrm{comp}}}\mathbb E[D_\ell(h(Q_1),Q_1)]>0. If the predictors are mutually calibrated, no transformation can guarantee strict improvement over all such outcome distributions.

  2. Knowl 2 — A loss-specific Bregman projection transfers calibration without labels

    model/method

    Given only the joint distribution of a primary forecast Q1Q_1 and a reference forecast Q0Q_0, the method constructs transformations h:Q1↦Δ(Y)h:Q_1\mapsto\Delta(\mathcal Y) that are compatible with the reference: some joint extension including the outcome makes both h(Q1)h(Q_1) and Q0Q_0 calibrated. It then selects the compatible transformation minimizing the expected proper-loss regret from the original primary forecast, EQ1[Dℓ(h(Q1),Q1)]\mathbb E_{Q_1}[D_\ell(h(Q_1),Q_1)]. The regret is generated by the convex function associated with the chosen proper loss, so the optimization is a Bregman projection in the space of functions, not merely a pointwise projection of predictions. With finite forecast supports, the joint forecast distribution and compatibility conditions can be represented using finite variables; no ground-truth outcomes are needed to fit the transformation. The transformed predictor uses the primary forecast as its input and can be deployed without the reference model. The diagram on page 4 illustrates the geometry: for two equally frequent groups, a primary model predicts (0.1,0.7)(0.1,0.7) while a constant calibrated reference predicts 0.50.5; compatibility requires the corrected group probabilities to sum to 11, and the squared-loss projection is (0.2,0.8)(0.2,0.8).

  3. Knowl 3 — Reference compatibility is a convex calibration-feasibility condition

    definition

    For categorical outcomes, represent the outcome YY and forecasts as probability vectors. Given a joint forecast distribution for (Q0,Q1)(Q_0,Q_1), the reference-compatible set consists of transformations h:Q1↦Δ(Y)h:Q_1\mapsto\Delta(\mathcal Y) for which there exists a joint extension including YY satisfying E[Y∣Q0]=Q0\mathbb E[Y\mid Q_0]=Q_0 and E[Y∣h(Q1)]=h(Q1)\mathbb E[Y\mid h(Q_1)]=h(Q_1). These are linear calibration constraints on the extended joint distribution. Mutual calibration of two predictors means that at least one such extension makes both original predictors calibrated. The compatible set is convex: mixing two feasible conditional distributions for YY preserves the calibration constraints and induces the corresponding mixture of transformations. Thus compatibility can be tested as a feasibility problem from the forecast joint distribution, without observing labels.

  4. Knowl 4 — Imperfect reference calibration creates a bounded risk term

    theoretical result

    Suppose the reference forecast Q0Q_0 is not perfectly calibrated. Let YY be a categorical outcome represented as a one-hot vector, let h(Q1)=E[Y∣Q1]h(Q_1)=\mathbb E[Y\mid Q_1] be the true conditional mean of the outcome given the primary forecast, and let h∗h^* be the Bregman projection onto the reference-compatible set. For a differentiable strictly convex loss generator GG, the change in expected loss decomposes as

    E[ℓ(Q1,Y)−ℓ(h∗(Q1),Y)]=E[Dℓ(h∗(Q1),Q1)]+E ⁣[⟨∇G(h∗(Q1))−∇G(Q1), hcomp(Q1)−h∗(Q1)⟩]+E ⁣[⟨∇G(h∗(Q1))−∇G(Q1), h(Q1)−hcomp(Q1)⟩].\begin{aligned} \mathbb E[\ell(Q_1,Y)-\ell(h^*(Q_1),Y)] ={}&\mathbb E[D_\ell(h^*(Q_1),Q_1)]\\ &+\mathbb E\!\left[\left\langle\nabla G(h^*(Q_1))-\nabla G(Q_1),\,h_{\mathrm{comp}}(Q_1)-h^*(Q_1)\right\rangle\right]\\ &+\mathbb E\!\left[\left\langle\nabla G(h^*(Q_1))-\nabla G(Q_1),\,h(Q_1)-h_{\mathrm{comp}}(Q_1)\right\rangle\right]. \end{aligned}

    Here hcomph_{\mathrm{comp}} is any compatible transformation. The first term, the geometric gain, is nonnegative; the second, the compatibility bonus, is also nonnegative; the last term is miscalibration risk and can be negative. Define the reference calibration residual E(Q0)=E[Y∣Q0]−Q0\mathcal E(Q_0)=\mathbb E[Y\mid Q_0]-Q_0 and the random-vector norm ∥X∥2=(E[∥X∥2])1/2\|X\|_2=(\mathbb E[\|X\|^2])^{1/2}. Choosing a suitable compatible transformation yields the bound ∣miscalibration risk∣≤∥∇G(h∗(Q1))−∇G(Q1)∥2 ∥E(Q0)∥2|\text{miscalibration risk}|\leq\|\nabla G(h^*(Q_1))-\nabla G(Q_1)\|_2\,\|\mathcal E(Q_0)\|_2. Thus, the alignment gain can still produce positive improvement with an imperfect reference when it exceeds the calibration-error penalty.

  5. Knowl 5 — The projection framework extends to general elicitable reports

    theoretical result

    A predictor may report a statistic rather than a full probability distribution. Let Γi\Gamma_i be an elicitable property for report RiR_i: a report is calibrated when it belongs to the property of the conditional outcome distribution given that report. Two reports are mutually calibrated when there are underlying predictive distributions consistent with the reports and a joint outcome distribution under which both underlying predictors are calibrated. The method maps a primary report R1R_1 to a distribution h(R1)∈Δ(Y)h(R_1)\in\Delta(\mathcal Y) and projects onto the set compatible with a reference report R0R_0, using the proper loss for the primary property. When the loss generator is differentiable at the projection, strict improvement is possible exactly when the reports are not mutually calibrated, with a guaranteed improvement at least equal to the minimum induced loss divergence over compatible transformations. For a nondifferentiable generator or a projection that does not attain its infimum, an ε\varepsilon-randomized projection guarantees improvement within ε\varepsilon of that infimum. One instance is a report (y^,c)(\hat y,c) consisting of a predicted class and confidence, elicited by the confidence loss ℓ((y^,c),y)=−1{y^=y}+(c−1{y^=y})2\ell((\hat y,c),y)=-\mathbf 1\{\hat y=y\}+(c-\mathbf 1\{\hat y=y\})^2.

  6. Knowl 6 — Mutual calibration is equivalent to absence of guaranteed arbitrage

    theoretical result

    For two categorical forecast vectors Q1,Q0Q_1,Q_0, a zero-trade contract ti(qi,y)t_i(q_i,y) has zero expected payout under forecast qiq_i, meaning EY∼qi[ti(qi,Y)]=0\mathbb E_{Y\sim q_i}[t_i(q_i,Y)]=0. With contracts bounded in [−1,1][-1,1], define the guaranteed arbitrage value from their joint forecast distribution as Arb⁡=−min⁡t1,t0EQ0,Q1[max⁡y{t1(Q1,y)+t0(Q0,y)}]\operatorname{Arb}=-\min_{t_1,t_0}\mathbb E_{Q_0,Q_1}[\max_y\{t_1(Q_1,y)+t_0(Q_0,y)\}], where the minimum ranges over zero-trade contracts. Define MinECE⁡\operatorname{MinECE} as one half the minimum, over all joint extensions including YY, of the expected sum of the two predictors' multiclass ℓ1\ell_1 calibration errors: 12min⁡E[∥Q1−E[Y∣Q1]∥1+∥Q0−E[Y∣Q0]∥1]\frac12\min\mathbb E[\|Q_1-\mathbb E[Y\mid Q_1]\|_1+\|Q_0-\mathbb E[Y\mid Q_0]\|_1]. The paper establishes MinECE⁡≤Arb⁡≤2MinECE⁡\operatorname{MinECE}\leq\operatorname{Arb}\leq2\operatorname{MinECE}. Consequently, the guaranteed arbitrage value is zero exactly when the forecasts are mutually calibrated; a failure of mutual calibration is quantitatively linked to an exploitable inconsistency in their beliefs.

  7. Knowl 7 — LLM experiments estimate answer-confidence compatibility from unlabeled forecasts

    experimental setup

    The experiments use Qwen3-8B, LLaMA-3.1-8B, and Ministral-3-8B Instruct models as primary predictors and their Base counterparts as references. The benchmarks are MMLU-Redux, with 5,700 test examples, and CommonSenseQA, evaluated on 1,221 validation examples because its test labels are unavailable. Each forecast is an answer and confidence pair: Base models are prompted for the answer directly, while Instruct models first generate reasoning and then receive an answer-triggering prompt to obtain answer-token probabilities. The probabilities for valid answer choices are renormalized over those choices; the most probable choice and its probability form the report. The empirical joint distribution of Base and Instruct reports is estimated without using labels, with confidence scores quantized at a 1% step. The optimization uses the confidence loss ℓ((y^,c),y)=−1{y^=y}+(c−1{y^=y})2\ell((\hat y,c),y)=-\mathbf 1\{\hat y=y\}+(c-\mathbf 1\{\hat y=y\})^2. Evaluation reports accuracy, Brier score, 10-bin expected calibration error, and confidence loss. Temperature scaling and other supervised calibration baselines use an additional labeled split; the label-free transformation does not.

  8. Knowl 8 — The label-free transformation improves calibration metrics across the 8B-model benchmarks

    data/table

    The reported 8B-model comparisons on MMLU-Redux and CommonSenseQA (page 11) contrast each Instruct primary model with its label-free transformed version. Each tuple gives accuracy in percent, followed by Brier score, expected calibration error (ECE), and confidence loss (CL); lower values are better for BS, ECE, and CL.

    Model Benchmark Instruct (Acc / BS / ECE / CL) Label-free (Acc / BS / ECE / CL)
    Qwen3-8B MMLU-Redux 83.28% / 0.1337 / 0.1250 / -0.6991 83.28% / 0.1232 / 0.0659 / -0.7096
    Qwen3-8B CommonSenseQA 82.88% / 0.1430 / 0.1385 / -0.6859 83.29% / 0.1291 / 0.0389 / -0.7038
    LLaMA-3.1-8B MMLU-Redux 71.74% / 0.2450 / 0.2417 / -0.4724 71.74% / 0.2153 / 0.1679 / -0.5022
    LLaMA-3.1-8B CommonSenseQA 78.62% / 0.2030 / 0.1999 / -0.5832 78.71% / 0.1657 / 0.0290 / -0.6214
    Ministral-3-8B MMLU-Redux 79.21% / 0.1720 / 0.1628 / -0.6201 80.52% / 0.1519 / 0.0664 / -0.6534
    Ministral-3-8B CommonSenseQA 74.04% / 0.2041 / 0.1682 / -0.5363 74.30% / 0.1777 / 0.0295 / -0.5653

    All six comparisons reduce both Brier score and ECE while improving confidence loss; accuracy is preserved or slightly higher. The largest reported accuracy increase is Ministral-3-8B on MMLU-Redux, from 79.21% to 80.52%. On CommonSenseQA with Ministral-3-8B, the label-free method's ECE of 0.0295 is lower than the reported supervised baselines.

  9. Knowl 9 — Calibration gains persist across tested Qwen3 model sizes

    empirical result

    The size sweep evaluates Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, and Qwen3-14B on both benchmarks, using Instruct forecasts as primary and corresponding Base forecasts as reference. Reported accuracy and ECE (page 37) are shown below as Instruct →\rightarrow label-free transformed values. Across all eight model-benchmark pairs, Brier score and confidence loss also improve, while accuracy is unchanged or increases.

    Model Benchmark Accuracy (%) ECE
    Qwen3-0.6B MMLU-Redux 54.58 54.58 0.3742 0.3485
    Qwen3-0.6B CommonSenseQA 56.67 56.67 0.3767 0.2182
    Qwen3-1.7B MMLU-Redux 70.26 70.44 0.2461 0.2097
    Qwen3-1.7B CommonSenseQA 71.58 71.66 0.2574 0.0885
    Qwen3-4B MMLU-Redux 80.26 80.26 0.1610 0.0965
    Qwen3-4B CommonSenseQA 80.26 80.67 0.1660 0.0297
    Qwen3-14B MMLU-Redux 85.75 85.75 0.1071 0.0612
    Qwen3-14B CommonSenseQA 82.80 83.37 0.1483 0.0232

    The results show that the calibration improvement is not restricted to the 8B setting: ECE decreases at every tested Qwen3 size on both tasks, including the 14B model.

  10. Knowl 10 — The guarantees depend on reference quality and sufficient unlabeled data

    limitation

    The theoretical guarantee assumes a calibrated reference; a reference with non-negligible calibration error can weaken improvement, with the potential loss bounded in terms of that error. Estimating the joint distribution of primary and reference forecasts also requires enough unlabeled examples, which may be difficult for rare forecast pairs. Finally, the projection depends on the chosen proper loss, so the paper does not provide a single transformation that is optimal for every proper loss simultaneously.

Coverage note — The paper's detailed proof arguments and auxiliary finite-net construction are omitted because they establish the stated projection guarantees rather than adding standalone contributed results; additional supervised-baseline numbers are omitted because the main label-free comparisons and supervised-competitiveness result are captured above.

References

  1. 1.J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 2, 12
  2. 2.R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In The twelfth international conference on learning representations, 2024. 5
  3. 3.A. Alexandari, A. Kundaje, and A. Shrikumar. Maximum likelihood with bias-corrected calibration is hard-to-beat at label shift adaptation. In International conference on machine learning, pages 222–232. PMLR, 2020. 5
  4. 4.Y. Altun and A. Smola. Unifying divergence minimization and statistical inference via convex duality. In Proceedings of the 19th Annual Conference on Learning Theory (COLT 2006), pages 139–153, 2006. 5
  5. 5.I. Arieli, Y. Babichenko, F. Sandomirskiy, and O. Tamuz. Feasible joint posterior beliefs. Journal of Political Economy, 129(9):2546–2594, 2021. 3, 5, 8, 25, 26
  6. 6.R. J. Aumann. Agreeing to disagree. The Annals of Statistics, pages 1236–1239, 1976. 5
  7. 7.A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh. Clustering with Bregman divergences. Journal of machine learning research, 6(Oct):1705–1749, 2005. 5
  8. 8.A. Banerjee, I. Dhillon, J. Ghosh, and S. Merugu. A generalized maximum entropy approach to Bregman co-clustering and matrix approximation. Journal of Machine Learning Research, 8:1919–1986, 2007. 5
  9. 9.G. W. BRIER. Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1):1–3, 1950. 12, 34
  10. 10.C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, et al. Weak-to-strong generalization: eliciting strong capabilities with weak supervision. In Proceedings of the 41st International Conference on Machine Learning, pages 4971–5012, 2024. 1, 5
  11. 11.M. Collins, R. E. Schapire, and Y. Singer. Logistic regression, Adaboost and Bregman distances. Machine Learning, 48(1-3):253–285, 2002. 5
  12. 12.C. Cortes, V. Kuznetsov, M. Mohri, and U. Syed. Structural maxent models. In F. Bach and D. Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 391–399, Lille, France, 07–09 Jul 2015. PMLR. 5
  13. 13.A. F. Cruz, M. Hardt, and C. Mendler-D¨unner. Evaluating language models as risk scores. Advances in Neural Information Processing Systems, 37:97378–97407, 2024. 2
  14. 14.M. Damani, I. Puri, S. Slocum, I. Shenfeld, L. Choshen, Y. Kim, and J. Andreas. Beyond binary rewards: Training lms to reason about their uncertainty. arXiv preprint arXiv:2507.16806, 2025. 7
  15. 15.A. P. Dawid. The well-calibrated bayesian. Journal of the American Statistical Association, 77 (379):605–610, 1982. doi: 10.1080/01621459.1982.10477856. URL https://www.tandfonline.com/doi/abs/10.1080/01621459.1982.10477856. 1
  16. 16.M. Dud´ık, S. J. Phillips, and R. E. Schapire. Maximum entropy density estimation with generalized regularization. Journal of Machine Learning Research, 8:1–48, 2007. 5
  17. 17.B. A. Frigyik, S. Srivastava, and M. R. Gupta. Functional Bregman divergence and Bayesian estimation of distributions. IEEE Transactions on Information Theory, 54(11):5130–5139, 2008. 5
  18. 18.R. Frongillo and I. Kash. General truthfulness characterizations via convex analysis. In Web and Internet Economics: 10th International Conference, WINE 2014, Beijing, China, December 14-17, 2014. Proceedings 10, pages 354–370. Springer, 2014. 5
  19. 19.A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, et al. Are we done with mmlu? In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5069–5096, 2025. 12, 30
  20. 20.A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. 12, 36
  21. 21.Y. Gu, L. Dong, F. Wei, and M. Huang. Minillm: Knowledge distillation of large language models. In The Twelfth International Conference on Learning Representations, 2024. 5
  22. 22.C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR, 2017. 1, 2, 12, 31, 34
  23. 23.P. Harremo¨es. Maximum entropy and sufficiency. arXiv preprint arXiv:1607.02259, 2016. 5
  24. 24.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ. 30
  25. 25.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 2, 5
  26. 26.J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. 1
  27. 27.J. Kivinen and M. K. Warmuth. Boosting as entropy projection. In Proceedings of the 12th Annual Conference on Computational Learning Theory (COLT ’99), pages 134–144, Santa Cruz, CA, USA, 1999. URL https://mwarmuth.bitbucket.io/pubs/C51.pdf. PDF available from authors’ webpages. 5
  28. 28.M. Kull, M. Perello Nieto, M. K¨angsepp, T. Silva Filho, H. Song, and P. Flach. Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration. Advances in neural information processing systems, 32, 2019. 5, 29
  29. 29.N. S. Lambert, D. M. Pennock, and Y. Shoham. Eliciting properties of probability distributions. In Proceedings of the 9th ACM Conference on Electronic Commerce, pages 129–138, 2008. 7
  30. 30.J. Leng, C. Huang, B. Zhu, and J. Huang. Taming overconfidence in LLMs: Reward calibration in RLHF. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=l0tg0jzsdL. 12
  31. 31.S. Lin, J. Hilton, and O. Evans. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=8s8K2UZGTZ. 30
  32. 32.A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, S. Sad´e, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, et al. Ministral 3. arXiv preprint arXiv:2601.08584, 2026. 12, 36
  33. 33.P. Milgrom and N. Stokey. Information, trade and common knowledge. Journal of economic theory, 26(1):17–27, 1982. 3, 5, 8, 26
  34. 34.M. Mohri, J. Schneider, and Y. Wu. Coherence mechanisms for provable self-improvement. arXiv preprint arXiv:2511.08440, 2025. 5, 17
  35. 35.S. Morris. No trade and feasible joint posterior beliefs, 2020. 3, 5, 8, 25, 26
  36. 36.R. Nock, A. K. Menon, and C. S. Ong. A scaled Bregman theorem with applications. arXiv preprint arXiv:1607.00360, 2016. 5
  37. 37.Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. Dillon, B. Lakshminarayanan, and J. Snoek. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Advances in neural information processing systems, 32, 2019. 1
  38. 38.S. Park, O. Bastani, J. Weimer, and I. Lee. Calibrated prediction with covariate shift via unsupervised domain adaptation. In International Conference on Artificial Intelligence and Statistics, pages 3219–3229. PMLR, 2020. 5
  39. 39.J. Platt et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in large margin classifiers, 10(3):61–74, 1999. 5, 31
  40. 40.T. Popordanoska, G. Radevski, T. Tuytelaars, and M. B. Blaschko. Estimating calibration error under label shift without labels. arXiv preprint arXiv:2312.08586, 2023. 5
  41. 41.R. T. Rockafellar. Convex Analysis, volume 28. Princeton University Press, 1997. 18
  42. 42.V. Sanh, L. Debut, J. Chaumond, and T. Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. URL https://arxiv.org/abs/1910.01108. 5
  43. 43.L. J. Savage. Elicitation of personal probabilities and expectations. Journal of the American Statistical Association, 66(336):783–801, 1971. 9
  44. 44.A. Talmor, J. Herzig, N. Lourie, and J. Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, 2019. 12, 30
  45. 45.K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5433–5442, Singapore, Dec. 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.330. URL https://aclanthology.org/2023.emnlp-main.330/. 30
  46. 46.P. Villalobos, A. Ho, J. Sevilla, T. Besiroglu, L. Heim, and M. Hobbhahn. Position: Will we run out of data? limits of llm scaling based on human-generated data. In Forty-first International Conference on Machine Learning, 2024. 1
  47. 47.X. Wang, B. Ma, C. Hu, L. Weber-Genzel, P. R¨ottger, F. Kreuter, D. Hovy, and B. Plank. “my answer is C”: First-token probabilities do not match text answers in instruction-tuned language models. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 7407–7416, Bangkok, Thailand, Aug. 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.441. URL https://aclanthology.org/2024.findings-acl.441/. 30
  48. 48.J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 30
  49. 49.A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. 12, 36
  50. 50.B. Zadrozny and C. Elkan. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Icml, volume 1, 2001. 32
  51. 51.B. Zadrozny and C. Elkan. Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 694–699, 2002. 5, 32

Citation

MLA
Kong, Y., et al. “Calibration Without Ground Truth”. arXiv, 2026, http://arxiv.org/abs/2601.19862v1.
APA
Kong, Y., Song, M., Wang, Y., & Wu, Y. (2026). Calibration without Ground Truth. arXiv. http://arxiv.org/abs/2601.19862v1
Chicago
Kong, Y., M. Song, Y. Wang, and Y. Wu. 2026. “Calibration Without Ground Truth”. arXiv. http://arxiv.org/abs/2601.19862v1.
Harvard
Kong, Y. et al. (2026) “Calibration without Ground Truth”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2601.19862v1.
Vancouver
1. Kong Y, Song M, Wang Y, Wu Y (2026) Calibration without Ground Truth. arXiv

BibTeX

@article{kong2026calibration,
  title = {Calibration without Ground Truth},
  author = {Kong, Yuqing and Song, Mingyu and Wang, Yizhou and Wu, Yifan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2601.19862v1},
  eprint = {2601.19862}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/