Stop Measuring Calibration When Humans Disagree
Joris BaanWilker AzizBarbara PlankRaquel Fernández
Demonstrates the theoretical flaws of standard calibration metrics when evaluated against majority labels on ambiguous tasks and proposes instance-level measures that directly align model confidence with full distributions of human judgment.
As machine learning models are increasingly deployed in high-stakes, user-facing environments, assessing model trustworthiness is vital. A primary method for doing this is calibration, which measures whether a model's predicted confidence accurately reflects its likelihood of being correct. Standard calibration metrics rely on comparing predictions against a single deterministic "gold label" determined by the majority vote of human annotators. However, many real-world language tasks feature fluid category boundaries and inherent, irreconcilable human disagreement. The article evaluates why traditional calibration breaks down in such contexts and proposes an instance-level framework to evaluate model alignment directly against the full distribution of human judgments.
The researchers demonstrate the mathematical flaw in conventional metrics through theoretical analysis and an empirical case study using the ChaosNLI benchmark, a dataset comprising 4,645 natural language inference instances where each instance received 100 independent human annotations. They fine-tuned a neural language model (RoBERTa) on inference tasks and evaluated both standard predictions and temperature scaling—a popular post-processing adjustment designed to improve calibration—using both traditional metrics and three newly introduced measures: distribution calibration error, entropy calibration error, and ranking calibration score.
The findings reveal that standard metrics like Expected Calibration Error produce severely misleading conclusions when applied to data with human disagreement. A hypothetical oracle classifier that perfectly mirrors human judgment achieves a 100% error-free alignment under the proposed instance-level metrics, yet traditional metrics penalize it with a high calibration error of 0.25 (compared to 0.14 for the baseline model). Furthermore, while applying temperature scaling to the neural model appeared to drastically improve traditional calibration error (reducing it from 0.14 to 0.03), it had minimal impact on true distribution error (shifting only from 0.26 to 0.22). Detailed inspection showed that temperature scaling artificially compresses probability ranges, reducing extreme errors only by sacrificing predictions that were already well-calibrated and causing the model to become excessively uncertain. Compared to realistic human sub-populations, the neural models exhibited distribution divergences 150 to 170 times larger, underscoring a wide gap between human uncertainty and model behavior.
These insights demonstrate that relying on majority-vote calibration creates a false sense of model safety and reliability in subjective or ambiguous tasks. Decision-makers relying on standard metrics risk deploying models whose apparent confidence does not reflect real-world human consensus. In practice, popular fixes like temperature scaling do not genuinely align model behavior with human uncertainty, but merely manipulate aggregate statistics. Organizations building decision-support systems must shift toward instance-level evaluation frameworks that respect human disagreement rather than treating all variation as noise.
Moving forward, machine learning pipelines for complex language tasks should adopt multiple annotations per instance—at least within evaluation benchmarks—and implement distribution-based calibration metrics. Evaluators should avoid using single-temperature post-processing adjustments as a blanket remedy for miscalibration without first checking instance-level error distributions. Key limitations of this work include the high cost of collecting 100 annotations per instance and the simplifying assumption that annotator groups share a single distribution without distinct sub-population biases. While confidence in the theoretical arguments and empirical results is high, practitioners must carefully distinguish between genuine human disagreement and annotator noise when gathering calibration data.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Read this influential treatment of neural-network calibration first to understand the confidence-versus-accuracy framework that the source questions when labels are disputed.
- Paper: The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation, Barbara Plank (2022). Its account of human label variation as meaningful signal rather than annotation noise provides essential context for the source’s critique of majority-label calibration.
- Paper: Learning From Crowds, V. Raykar et al. (2010). Its probabilistic approach to learning from multiple annotators establishes why disagreement can carry information that a single majority label discards.
- Paper: When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks, Eve Fleisig et al. (2023). Building on the problem of treating majority judgments as ground truth, this work models whose perspectives drive disagreement and applies that approach to subjective language tasks.
