Stop Measuring Calibration When Humans Disagree

Joris BaanWilker AzizBarbara PlankRaquel Fernández

article2022EMNLP85 citations

Demonstrates the theoretical flaws of standard calibration metrics when evaluated against majority labels on ambiguous tasks and proposes instance-level measures that directly align model confidence with full distributions of human judgment.

Listen

As machine learning models are increasingly deployed in high-stakes, user-facing environments, assessing model trustworthiness is vital. A primary method for doing this is calibration, which measures whether a model's predicted confidence accurately reflects its likelihood of being correct. Standard calibration metrics rely on comparing predictions against a single deterministic "gold label" determined by the majority vote of human annotators. However, many real-world language tasks feature fluid category boundaries and inherent, irreconcilable human disagreement. The article evaluates why traditional calibration breaks down in such contexts and proposes an instance-level framework to evaluate model alignment directly against the full distribution of human judgments.

The researchers demonstrate the mathematical flaw in conventional metrics through theoretical analysis and an empirical case study using the ChaosNLI benchmark, a dataset comprising 4,645 natural language inference instances where each instance received 100 independent human annotations. They fine-tuned a neural language model (RoBERTa) on inference tasks and evaluated both standard predictions and temperature scaling—a popular post-processing adjustment designed to improve calibration—using both traditional metrics and three newly introduced measures: distribution calibration error, entropy calibration error, and ranking calibration score.

The findings reveal that standard metrics like Expected Calibration Error produce severely misleading conclusions when applied to data with human disagreement. A hypothetical oracle classifier that perfectly mirrors human judgment achieves a 100% error-free alignment under the proposed instance-level metrics, yet traditional metrics penalize it with a high calibration error of 0.25 (compared to 0.14 for the baseline model). Furthermore, while applying temperature scaling to the neural model appeared to drastically improve traditional calibration error (reducing it from 0.14 to 0.03), it had minimal impact on true distribution error (shifting only from 0.26 to 0.22). Detailed inspection showed that temperature scaling artificially compresses probability ranges, reducing extreme errors only by sacrificing predictions that were already well-calibrated and causing the model to become excessively uncertain. Compared to realistic human sub-populations, the neural models exhibited distribution divergences 150 to 170 times larger, underscoring a wide gap between human uncertainty and model behavior.

These insights demonstrate that relying on majority-vote calibration creates a false sense of model safety and reliability in subjective or ambiguous tasks. Decision-makers relying on standard metrics risk deploying models whose apparent confidence does not reflect real-world human consensus. In practice, popular fixes like temperature scaling do not genuinely align model behavior with human uncertainty, but merely manipulate aggregate statistics. Organizations building decision-support systems must shift toward instance-level evaluation frameworks that respect human disagreement rather than treating all variation as noise.

Moving forward, machine learning pipelines for complex language tasks should adopt multiple annotations per instance—at least within evaluation benchmarks—and implement distribution-based calibration metrics. Evaluators should avoid using single-temperature post-processing adjustments as a blanket remedy for miscalibration without first checking instance-level error distributions. Key limitations of this work include the high cost of collecting 100 annotations per instance and the simplifying assumption that annotator groups share a single distribution without distinct sub-population biases. While confidence in the theoretical arguments and empirical results is high, practitioners must carefully distinguish between genuine human disagreement and annotator noise when gathering calibration data.

Cover for Stop Measuring Calibration When Humans Disagree

Abstract

Calibration is a popular framework to evaluate whether a classifier knows when it does not know—i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct. Correctness is commonly estimated against the human majority class. Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies. We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instance-level measures of calibration that capture key statistical properties of human judgements—class frequency, ranking and entropy.1

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Background
  • 3 Calibration & Disagreement Pathology
  • 4 Calibration to Human Uncertainty
  • 5 Case Study
  • 5.1 Experimental Setup
  • 5.2 Results
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • Appendix
  • A ChaosNLI
  • B Additional Analyses
  • B.1 Beyond DistCE
  • B.2 Out of Distribution Evaluation
  • C Temperature Scaling
  • D Classwise-ECE
  • E Total Variation Distance

Knowls

  1. Knowl 1 — Majority-vote calibration misclassifies a faithful uncertainty model

    theoretical result

    Suppose a classifier predicts the human label distribution for each instance, and correctness is assessed against the human majority label. When annotators genuinely disagree, the classifier’s most probable class can still match the majority label on every instance, giving it perfect majority-vote accuracy, while its maximum predicted probability is below 1. Confidence-calibration metrics such as ECE compare that probability with accuracy, so they penalize this faithful representation of human uncertainty. A classifier can therefore be severely miscalibrated by majority-vote-based metrics even when it predicts the human distribution exactly; achieving apparent perfect calibration can instead require probabilities that misrepresent human uncertainty.

  2. Knowl 2 — Instance-level calibration to the human label distribution

    definition

    For an instance xx with classes c∈[C]={1,…,C}c\in[C]=\{1,\ldots,C\}, let YY be a human annotator’s label, let πc(x)=Pr⁡(Y=c∣X=x)\pi_c(x)=\Pr(Y=c\mid X=x) be the population frequency of class cc, and let fc(x)f_c(x) be the classifier’s predicted probability for that class. The classifier is perfectly calibrated to human uncertainty on xx when

    Pr⁡(Y=c∣X=x)=fc(x)for every c∈[C].\Pr(Y=c\mid X=x)=f_c(x)\qquad\text{for every }c\in[C].

    With multiple labels for the same instance, the human distribution can be estimated by the empirical class frequencies πˉc(x)\bar\pi_c(x); this estimate becomes more reliable as the number of annotations increases. Unlike calibration definitions based on grouping instances by model confidence, this criterion directly compares the model and human distributions for each instance.

  3. Knowl 3 — Human Distribution Calibration Error (DistCE)

    equation

    For an instance xx, let f(x)f(x) be a classifier’s probability vector over CC classes and let πˉ(x)\bar\pi(x) be the empirical human label-frequency vector for that instance. The Human Distribution Calibration Error is their total variation distance:

    DistCE⁡(x)=TVD⁡(f(x),πˉ(x))=12∑c=1C∣fc(x)−πˉc(x)∣.\operatorname{DistCE}(x)=\operatorname{TVD}(f(x),\bar\pi(x))=\frac12\sum_{c=1}^{C}|f_c(x)-\bar\pi_c(x)|.

    DistCE measures the full instance-level discrepancy between the model and human distributions, rather than only their most likely class or confidence. It lies in [0,1][0,1] and is zero exactly when the two probability vectors match. The paper presents it as a proper scoring rule and emphasizes that, unlike ECE, it requires no confidence bins.

  4. Knowl 4 — Human Entropy Calibration Error (EntCE)

    equation

    For an instance xx, let f(x)f(x) be the classifier’s probability distribution and πˉ(x)\bar\pi(x) the empirical human label distribution. If HH denotes categorical entropy, the Human Entropy Calibration Error is

    EntCE⁡(x)=H(f(x))−H(πˉ(x)).\operatorname{EntCE}(x)=H(f(x))-H(\bar\pi(x)).

    This signed discrepancy compares model indecisiveness with human disagreement while ignoring which classes are ranked higher. A positive value means the model distribution has greater entropy than the empirical human distribution; a negative value means it has lower entropy. The paper summarizes this metric using the mean absolute error across instances.

  5. Knowl 5 — Human Ranking Calibration Score (RankCS)

    equation

    For NN instances, let f(xn)f(x_n) be the model’s class-probability vector and πˉ(xn)\bar\pi(x_n) the empirical human label-frequency vector for instance xnx_n. The Human Ranking Calibration Score is

    RankCS⁡=1N∑n=1N[argsort⁡(f(xn))=argsort⁡(πˉ(xn))],\operatorname{RankCS}=\frac1N\sum_{n=1}^{N}\big[\operatorname{argsort}(f(x_n))=\operatorname{argsort}(\bar\pi(x_n))\big],

    where the Iverson bracket is 1 when the two class rankings agree and 0 otherwise. RankCS measures whether the model orders classes as humans do, without measuring the probability magnitudes. The authors describe it as a stricter alternative to majority-vote accuracy; they leave the handling of tied human class frequencies unspecified.

  6. Knowl 6 — ChaosNLI-SNLI calibration experiment

    experimental setup

    The case study used ChaosNLI, which contains natural-language-inference examples selected for borderline initial agreement—at most 3 of 5 votes for one class—and then collected 100 additional independent labels per premise–hypothesis pair. The experiments covered 4,645 instances across the dataset’s three classes. RoBERTa was fine-tuned on SNLI and evaluated on the ChaosNLI-SNLI split. Temperature scaling (TS) divided the model logits by a single temperature; the authors used a temperature tuned directly on the evaluation set as an unrealistic upper bound for TS, finding t=2.0t=2.0 to minimize ECE. Reported RoBERTa and RoBERTa-TS results are averages over three random seeds.

  7. Knowl 7 — Majority-vote ECE contradicts the proposed human-calibration measures

    data/table

    On ChaosNLI-SNLI, the paper compared RoBERTa, temperature-scaled RoBERTa (RoBERTa-TS), and an oracle that predicts the empirical human distribution. Accuracy and RankCS are higher-is-better; ECE, mean absolute EntCE, and mean DistCE are lower-is-better. RoBERTa scored accuracy 0.74±0.010.74\pm0.01, ECE 0.14±0.010.14\pm0.01, RankCS 0.62±0.010.62\pm0.01, mean ∣EntCE⁡∣|\operatorname{EntCE}| 0.30±0.020.30\pm0.02, and mean DistCE 0.26±0.000.26\pm0.00. RoBERTa-TS scored 0.74±0.010.74\pm0.01, 0.03±0.010.03\pm0.01, 0.62±0.010.62\pm0.01, 0.21±0.020.21\pm0.02, and 0.22±0.000.22\pm0.00, respectively. The oracle scored accuracy 1.001.00, ECE 0.250.25, RankCS 1.001.00, mean ∣EntCE⁡∣|\operatorname{EntCE}| 0.000.00, and mean DistCE 0.000.00. Thus, the human-distribution oracle is perfect under the proposed measures but has worse ECE than either neural model, illustrating the majority-vote calibration pathology.

  8. Knowl 8 — Temperature scaling reduces tail errors but sacrifices exact instance matches

    empirical result

    On ChaosNLI-SNLI, temperature scaling reduced RoBERTa’s ECE from 0.140.14 to 0.030.03, while mean instance-level DistCE fell only from 0.260.26 to 0.220.22. The page-5 error-distribution plots show that the reduction in large DistCE errors came alongside fewer instances with zero error: temperature scaling compressed the model’s probability range, moving some already well-aligned predictions away from the human distribution. This distributional change is not apparent from the ECE improvement alone.

  9. Knowl 9 — Human subsamples provide a reference scale for model distribution errors

    empirical result

    To estimate the error-distribution scale attainable from finite human judgments, the authors formed two classifiers, H1 and H2, from separate 20-vote subsamples for each instance and evaluated them against all 100 available ChaosNLI annotations. The KL divergence between their DistCE error distributions was 0.0040.004 and their total variation distance (TVD) was 0.0220.022. By comparison, the corresponding divergences from H1 to RoBERTa were KL 0.6880.688 and TVD 0.5000.500, and from H1 to RoBERTa-TS were KL 0.6110.611 and TVD 0.4540.454. The model error distributions therefore differed substantially more from the human-sub sample reference distributions than the two human subsamples differed from each other; TS produced only a modest reduction in this gap.

  10. Knowl 10 — Human calibration depends on representative, reliable annotation distributions

    limitation

    Estimating human calibration reliably requires multiple labels per instance, which many datasets do not provide, and the labels must distinguish genuine disagreement from annotation noise such as spam. The paper also treats a single human label distribution as governing all annotators, but annotators may come from distinct subpopulations; a distribution aggregated across one population may not represent any particular subgroup. Consequently, an empirical human distribution is not necessarily a universally correct target for every population.

Coverage note — The supplementary out-of-domain and classwise analyses are omitted because they mainly confirm the paper’s central findings rather than establish a separate contribution.

References

  1. 1.Sohail Akhtar, Valerio Basile, and Viviana Patti. 2020. Modeling annotator perspective and polarized opinions to improve hate speech detection. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 8(1):151–154.
  2. 2.Lora Aroyo, Lucas Dixon, Nithum Thain, Olivia Redfield, and Rachel Rosen. 2019. Crowdsourcing subjective tasks: The case study of understanding toxicity in online discussions. In Companion Proceedings of The 2019 World Wide Web Conference, WWW ’19, page 1100–1105, New York, NY, USA. Association for Computing Machinery.
  3. 3.Lora Aroyo and Chris Welty. 2015. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24.
  4. 4.Ron Artstein and Massimo Poesio. 2008. Inter-coder agreement for computational linguistics. Computational linguistics, 34(4):555–596.
  5. 5.Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019. Beyond accuracy: The role of mental models in human-ai team performance. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 7, pages 2–11.
  6. 6.Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, Alexandra Uma, et al. 2021. We need to consider disagreement in evaluation. In 1st Workshop on Benchmarking: Past, Present and Future, pages 15–21. Association for Computational Linguistics.
  7. 7.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In International Conference on Learning Representations.
  8. 8.Federico Bianchi and Dirk Hovy. 2021. On the gap between adoption and understanding in NLP. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3895–3901, Online. Association for Computational Linguistics.
  9. 9.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  10. 10.Tongfei Chen, Zhengping Jiang, Adam Poliak, Keisuke Sakaguchi, and Benjamin Van Durme. 2020. Uncertain natural language inference. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8772–8779, Online. Association for Computational Linguistics.
  11. 11.Soham Dan and Dan Roth. 2021. On the effects of transformer size on in- and out-of-domain calibration. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2096–2101, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Alexander Philip Dawid and Allan M Skene. 1979. Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20–28.
  13. 13.Morris H. DeGroot and Stephen E. Fienberg. 1983. The comparison and evaluation of forecasters. Journal of the Royal Statistical Society. Series D (The Statistician), 32(1/2):12–22.
  14. 14.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Luc Devroye and Gábor Lugosi. 2001. Combinatorial methods in density estimation. Springer Science & Business Media.
  17. 17.Tilmann Gneiting and Adrian E Raftery. 2007. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378.
  18. 18.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 1321–1330. JMLR.org.
  19. 19.Kartik Gupta, Amir Rahimi, Thalaiyasingam Ajanthan, Thomas Mensink, Cristian Sminchisescu, and Richard Hartley. 2021. Calibration of neural networks using splines. In International Conference on Learning Representations.
  20. 20.Emily Jamison and Iryna Gurevych. 2015. Noise or additional information? Leveraging crowdsource annotation item agreement for natural language tasks. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 291–297, Lisbon, Portugal. Association for Computational Linguistics.
  21. 21.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  22. 22.Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2022. Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future. Computational Linguistics, pages 1–42.
  23. 23.Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang. 2020. Calibrated language model fine-tuning for in- and out-of-distribution data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1326–1340, Online. Association for Computational Linguistics.
  24. 24.Meelis Kull, Miquel Perelló-Nieto, Markus Kängsepp, Telmo de Menezes e Silva Filho, Hao Song, and Peter A. Flach. 2019. Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration. In NeurIPS, pages 12295–12305.
  25. 25.Aviral Kumar, Sunita Sarawagi, and Ujjwal Jain. 2018. Trainable calibration measures for neural networks from kernel mean embeddings. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2805–2814. PMLR.
  26. 26.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  27. 27.Johannes Mario Meissner, Napat Thumwanit, Saku Sugawara, and Akiko Aizawa. 2021. Embracing ambiguity: Shifting the training target of NLI models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 862–869, Online. Association for Computational Linguistics.
  28. 28.José Mena, Oriol Pujol, and Jordi Vitrià. 2021. A Survey on Uncertainty Estimation in Deep Learning Classification Systems from a Bayesian Perspective. ACM Comput. Surv., 54(9).
  29. 29.Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining Well Calibrated Probabilities Using Bayesian Binning. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, AAAI’15, page 2901–2907. AAAI Press.
  30. 30.Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131–9143, Online. Association for Computational Linguistics.
  31. 31.Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019. Measuring calibration in deep learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops.
  32. 32.Jennimaria Palomaki, Olivia Rhinehart, and Michael Tseng. 2018. A case for a range of acceptable annotations. In SAD/CrowdBias@ HCOMP, pages 19–31.
  33. 33.Silviu Paun, Ron Artstein, and Massimo Poesio. 2022. Statistical methods for annotation analysis. Synthesis Lectures on Human Language Technologies, 15(1):1–217.
  34. 34.Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
  35. 35.Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014. Linguistically debatable or just plain wrong? In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 507–511.
  36. 36.Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. On releasing annotator-level labels and information in datasets. In Proceedings of The Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Vikas C. Raykar and Shipeng Yu. 2012. Eliminating spammers and ranking annotators for crowdsourced labeling tasks. J. Mach. Learn. Res., 13:491–518.
  38. 38.Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8:842–866.
  39. 39.Chenglei Si, Chen Zhao, Sewon Min, and Jordan Boyd-Graber. 2022. Revisiting calibration for question answering. arXiv preprint arXiv:2205.12507.
  40. 40.Juozas Vaicenavicius, David Widmann, Carl Andersson, Fredrik Lindsten, Jacob Roll, and Thomas Schön. 2019. Evaluating model calibration in classification. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 3459–3467. PMLR.
  41. 41.Yuxia Wang, Minghan Wang, Yimeng Chen, Shimin Tao, Jiaxin Guo, Chang Su, Min Zhang, and Hao Yang. 2022. Capture human disagreement distributions by calibrated networks for natural language inference. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1524–1535, Dublin, Ireland. Association for Computational Linguistics.
  42. 42.David Widmann, Fredrik Lindsten, and Dave Zachariah. 2019. Calibration tests in multi-class classification: A unifying framework. Advances in Neural Information Processing Systems, 32.
  43. 43.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  44. 44.Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the effect of accuracy on trust in machine learning models. In Proceedings of the 2019 CHI conference on human factors in computing systems, pages 1–12.
  45. 45.Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021. Learning with different amounts of annotation: From zero to many labels. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7620–7632, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Xiang Zhou, Yixin Nie, and Mohit Bansal. 2022. Distributed NLI: Learning to predict human opinion distributions for language reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pages 972–987, Dublin, Ireland. Association for Computational Linguistics.

Citation

MLA
Baan, J., et al. “Stop Measuring Calibration When Humans Disagree”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 1892–915, https://doi.org/10.18653/v1/2022.emnlp-main.124.
APA
Baan, J., Aziz, W., Plank, B., & Fernández, R. (2022). Stop Measuring Calibration When Humans Disagree. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1892–1915. https://doi.org/10.18653/v1/2022.emnlp-main.124
Chicago
Baan, J., W. Aziz, B. Plank, and R. Fernández. 2022. “Stop Measuring Calibration When Humans Disagree”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1892–1915. https://doi.org/10.18653/v1/2022.emnlp-main.124.
Harvard
Baan, J. et al. (2022) “Stop Measuring Calibration When Humans Disagree”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1892–1915. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.124.
Vancouver
1. Baan J, Aziz W, Plank B, Fernández R (2022) Stop Measuring Calibration When Humans Disagree. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1892–1915

BibTeX

@inproceedings{baan-etal-2022-stop,
    title = "Stop Measuring Calibration When Humans Disagree",
    author = "Baan, Joris  and
      Aziz, Wilker  and
      Plank, Barbara  and
      Fernandez, Raquel",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.124/",
    doi = "10.18653/v1/2022.emnlp-main.124",
    pages = "1892--1915"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/