When and How Mixup Improves Calibration

Linjun ZhangZhun DengKenji KawaguchiJames Zou

article2022ICML79 citations

Establishes the theoretical foundation for why Mixup data augmentation reduces prediction calibration errors in high-dimensional settings and semi-supervised learning, showing that its calibration benefits grow alongside model capacity.

Listen

Modern machine learning models achieve remarkable predictive accuracy, but they are frequently overconfident, meaning their predicted confidence levels do not match their true likelihood of being correct. In high-stakes fields such as healthcare, fraud detection, and automated systems, reliable uncertainty estimation is essential for safe decision-making. While Mixup—a simple training technique that blends pairs of training examples and their labels—has been observed empirically to improve model calibration, the underlying theoretical reasons for when and why this occurs have remained unclear.

The article establishes a rigorous theoretical foundation and empirical validation for how Mixup influences model calibration. It specifically evaluates the conditions under which Mixup reduces calibration error in high-dimensional supervised learning, under mild domain shifts, and within semi-supervised workflows.

To conduct this evaluation, the authors analyzed mathematical data models, including Gaussian mixture models and flexible Gaussian generative models, to formally prove calibration behavior under varying parameter-to-sample ratios. They complemented these theoretical proofs with extensive empirical experiments using fully connected neural networks and standard residual networks across established image classification benchmarks, including CIFAR-10, CIFAR-100, Fashion-MNIST, and Kuzushiji-MNIST.

The article yields four primary findings. First, Mixup fundamentally acts as a high-dimensional regularizer: it provably reduces calibration errors when the model capacity is large relative to the training sample size, shrinking parameter norms to prevent extreme, overconfident predictions. Second, the calibration benefit scales directly with model capacity; deeper and wider networks experience substantial calibration improvements, whereas Mixup offers no benefit and can even worsen calibration in low-dimensional, low-capacity regimes. Third, Mixup preserves its calibration advantages under moderate out-of-domain distribution shifts. Fourth, in semi-supervised learning, standard pseudo-labeling often degrades calibration when labeled data is abundant, but integrating Mixup into the pseudo-labeling pipeline consistently corrects this degradation and restores well-calibrated confidence scores across both expected and maximum calibration error metrics.

These findings indicate that Mixup provides an efficient, built-in mechanism to calibrate large modern networks without requiring computationally heavy post-processing procedures or multi-model ensembles. This reduces operational cost and deployment complexity while mitigating the safety and reliability risks associated with overconfident artificial intelligence predictions. Engineering teams can leverage Mixup particularly in sample-constrained environments where traditional validation-based recalibration techniques are impractical.

Technical leaders and practitioners deploying high-capacity models should integrate Mixup into standard supervised and semi-supervised training pipelines to improve trust and reliability, but they should avoid using it on small, low-capacity models where it may impair calibration. Next steps include exploring how other advanced Mixup variations influence calibration and extending this theoretical framework to semi-supervised settings involving out-of-domain unlabeled data.

Confidence in these conclusions is high, as the analytical proofs directly align with empirical results across diverse network architectures and datasets. However, readers should note that the theoretical results rely on specific statistical data assumptions, and empirical gains are bounded by the degree of model overparameterization and domain shift.

arXiv: 2102.06289
  • Paper: mixup: Beyond Empirical Risk Minimization, Hongyi Zhang et al. (2017). This seminal paper introduces the Mixup data augmentation technique whose calibration benefits and theoretical underpinnings the source paper analyzes.
  • Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). This foundational work establishes the modern problem of neural network miscalibration and introduces standard calibration metrics like Expected Calibration Error.
  • Paper: MixMatch: A Holistic Approach to Semi-Supervised Learning, David Berthelot et al. (2019). This work incorporates Mixup into semi-supervised learning workflows, establishing the algorithmic foundation for the source paper's semi-supervised calibration analysis.
  • Paper: Manifold Mixup: Better Representations by Interpolating Hidden States, Vikas Verma et al. (2018). This paper investigates linear interpolation in representation space and provides early empirical observations on how Mixup-style regularizers smooth decision boundaries and improve confidence calibration.
  • Paper: When Does Label Smoothing Help?, Rafael Müller et al. (2019). This study analyzes how soft-target regularization improves confidence calibration in deep networks, providing key context for the calibration mechanisms of Mixup.

No sufficiently relevant recommendations were found.

Cover for When and How Mixup Improves Calibration

Abstract

In many machine learning applications, it is important for the model to provide confidence scores that accurately capture its prediction uncertainty. Although modern learning methods have achieved great success in predictive accuracy, generating calibrated confidence scores remains a major challenge. Mixup, a popular yet simple data augmentation technique based on taking convex combinations of pairs of training examples, has been empirically found to significantly improve confidence calibration across diverse applications. However, when and how Mixup helps calibration is still a mystery. In this paper, we theoretically prove that Mixup improves calibration in high-dimensional settings by investigating natural statistical models. Interestingly, the calibration benefit of Mixup increases as the model capacity increases. We support our theories with experiments on common architectures and datasets. In addition, we study how Mixup improves calibration in semi-supervised learning. While incorporating unlabeled data can sometimes make the model less calibrated, adding Mixup training mitigates this issue and provably improves calibration. Our analysis provides new insights and a framework to understand Mixup and calibration.

Table of Contents

  • 1. Introduction
  • 1.1. Related Work
  • 2. Preliminaries
  • 2.1. Notations
  • 2.2. Mixup
  • 2.3. Calibration for classification
  • 3. Calibration in Supervised Learning
  • 3.1. Problem set-up
  • 3.2. Mixup helps calibration in classification
  • 3.3. Improvement for out-of-domain data
  • 3.4. Gaussian generative model
  • 4. Mixup Improves Calibration in Semi-supervised Learning
  • 5. Extension to Maximum Calibration Error
  • 6. Conclusion and Discussion
  • Acknowledgements
  • References
  • Appendix
  • A. Technical Details
  • A.1. Proof of Theorem 3.1
  • A.2. Proof of Theorem 3.2
  • A.3. Proof of Theorem 3.3
  • A.4. Proof of Theorem 3.4
  • A.5. Proof of Theorem 3.5
  • A.6. Proof of Theorem 4.1
  • A.7. Proof of Theorem 4.2
  • A.8. Proof of Theorem 4.3
  • A.9. Proof of Theorem 5.1
  • A.10. Proof of Theorem 5.2
  • A.11. Proof of Theorem 5.3
  • B. Experimental setup and additional numerical results

Knowls

  1. Knowl 1 — Mixup Improves Expected Calibration Error in High-Dimensional Classification

    theoretical result

    In a binary Gaussian mixture classification setting, let data pairs (x, y) acksim P_{x,y} on Rp×{−1,1}\mathbb{R}^p \times \{-1, 1\} be distributed according to prior class probabilities P(y=1)=P(y=−1)=1/2\mathbb{P}(y = 1) = \mathbb{P}(y = -1) = 1/2, with class-conditional distributions x∣y∼N(y⋅θ∗,σ2I)x \mid y \sim \mathcal{N}(y \cdot \theta^*, \sigma^2 I) for an unknown mean parameter θ∗∈Rp\theta^* \in \mathbb{R}^p and known σ>0\sigma > 0. Given nn independent training samples {(xi,yi)}i=1n\{(x_i, y_i)\}_{i=1}^n, the standard Fisher linear discriminant estimator is θ^=1n∑i=1nxiyi\hat{\theta} = \frac{1}{n} \sum_{i=1}^n x_i y_i, yielding the classifier C^(x)=sgn(θ^⊤x)\hat{C}(x) = \text{sgn}(\hat{\theta}^\top x) with predicted class confidence scores pk(x)=(1+e−2kθ^⊤x/σ2)−1p_k(x) = (1 + e^{-2k \hat{\theta}^\top x / \sigma^2})^{-1} for k∈{−1,1}k \in \{-1, 1\}.

    The Mixup-augmented estimator is defined by taking expected linear combinations across all sample pairs: θ^mix=Eλ∼Dλ[1n2∑i,j=1n(λxi+(1−λ)xj)(λyi+(1−λ)yj)],\hat{\theta}_{\text{mix}} = \mathbb{E}_{\lambda \sim \mathcal{D}_\lambda} \left[ \frac{1}{n^2} \sum_{i,j=1}^n (\lambda x_i + (1 - \lambda) x_j)(\lambda y_i + (1 - \lambda) y_j) \right], where Dλ=Beta(α,β)\mathcal{D}_\lambda = \text{Beta}(\alpha, \beta) for α,β>0\alpha, \beta > 0, producing the classifier C^mix(x)=sgn(θ^mix⊤x)\hat{C}_{\text{mix}}(x) = \text{sgn}(\hat{\theta}_{\text{mix}}^\top x).

    When the dimension-to-sample ratio satisfies p/n∈(c1,c2)p/n \in (c_1, c_2) and ∥θ∗∥2<C\|\theta^*\|_2 < C for universal constants c2>c1>0c_2 > c_1 > 0 and C>0C > 0, then for sufficiently large pp and nn, there exist Beta parameters α,β>0\alpha, \beta > 0 such that, with high probability (probability 1−o(1)1 - o(1) as n,p→∞n, p \to \infty): ECE(C^mix)<ECE(C^),\text{ECE}(\hat{C}_{\text{mix}}) < \text{ECE}(\hat{C}), where the population Expected Calibration Error is defined as ECE(C)=Ev∼Dp^[∣P(y^=y∣p^=v)−v∣]\text{ECE}(C) = \mathbb{E}_{v \sim D_{\hat{p}}}[|\mathbb{P}(\hat{y} = y \mid \hat{p} = v) - v|], with p^=max⁡k∈{−1,1}pk(x)\hat{p} = \max_{k \in \{-1, 1\}} p_k(x) and y^=argmaxk∈{−1,1}pk(x)\hat{y} = \text{argmax}_{k \in \{-1, 1\}} p_k(x).

  2. Knowl 2 — Monotonic Increase of Mixup Calibration Benefit with Overparameterization Ratio

    theoretical result

    In the (θ∗,σ)(\theta^*, \sigma)-Gaussian mixture classification model where x∣y∼N(y⋅θ∗,σ2I)x \mid y \sim \mathcal{N}(y \cdot \theta^*, \sigma^2 I) and P(y=1)=P(y=−1)=1/2\mathbb{P}(y = 1) = \mathbb{P}(y = -1) = 1/2, let C^α,βmix\hat{C}_{\alpha, \beta}^{\text{mix}} denote the linear classifier trained with Mixup data augmentation with coefficient distribution λ∼Beta(α,β)\lambda \sim \text{Beta}(\alpha, \beta). The boundary distribution Beta(0,β)\text{Beta}(0, \beta) places mass 1 at 00 and corresponds to standard training without Mixup.

    For any constant cmax>0c_{\text{max}} > 0, in the asymptotic regime where p/n→cratio∈(0,cmax)p/n \to c_{\text{ratio}} \in (0, c_{\text{max}}) as n,p→∞n, p \to \infty, and when the signal norm ∥θ∗∥2\|\theta^*\|_2 is sufficiently large (at a constant level), for any fixed β>0\beta > 0, the marginal change in Expected Calibration Error (ECE) under infinitesimal Mixup interpolation: ddαECE(C^α,βmix)∣α→0+\left. \frac{d}{d\alpha} \text{ECE}(\hat{C}_{\alpha, \beta}^{\text{mix}}) \right|_{\alpha \to 0^+} is strictly negative and strictly monotonically decreasing with respect to cratio=p/nc_{\text{ratio}} = p/n with high probability.

    This indicates that Mixup strictly reduces ECE in high-dimensional settings, and its marginal calibration improvement grows monotonically as the model becomes more overparameterized relative to sample size.

  3. Knowl 3 — Mixup Degrades Calibration in Low-Dimensional Asymptotic Regimes

    theoretical result

    In the (θ∗,σ)(\theta^*, \sigma)-Gaussian mixture model with x∣y∼N(y⋅θ∗,σ2I)x \mid y \sim \mathcal{N}(y \cdot \theta^*, \sigma^2 I) and uniform label prior on {−1,1}\{-1, 1\}, let C^\hat{C} be the standard Fisher linear discriminant classifier C^(x)=sgn(θ^⊤x)\hat{C}(x) = \text{sgn}(\hat{\theta}^\top x) with θ^=1n∑i=1nxiyi\hat{\theta} = \frac{1}{n} \sum_{i=1}^n x_i y_i, and let C^mix\hat{C}_{\text{mix}} be the classifier trained with Mixup using λ∼Beta(α,β)\lambda \sim \text{Beta}(\alpha, \beta) for fixed constants α,β>0\alpha, \beta > 0.

    If the dimension-to-sample ratio satisfies p/n≤τp/n \le \tau for an asymptotic threshold τ=o(1)\tau = o(1), and ∥θ∗∥2<C\|\theta^*\|_2 < C for a universal constant C>0C > 0, then for sufficiently large sample size nn, with high probability: ECE(C^)<ECE(C^mix).\text{ECE}(\hat{C}) < \text{ECE}(\hat{C}_{\text{mix}}).

    Consequently, the condition p/n=Ω(1)p/n = \Omega(1) is necessary for Mixup to improve calibration. In low dimensions where p/n→0p/n \to 0, the standard estimator is already well-calibrated (extECE(C^)=Op(p/n)=op(1) ext{ECE}(\hat{C}) = O_p(p/n) = o_p(1)), while Mixup introduces unnecessary shrinkage that bounds calibration error away from zero by an Ω(1)\Omega(1) margin.

  4. Knowl 4 — Mixup Restores Calibration in Semi-Supervised Pseudo-Labeling

    theoretical result

    In a semi-supervised Gaussian mixture classification problem with nln_l labeled examples {(xi,yi)}i=1nl\{(x_i, y_i)\}_{i=1}^{n_l} and nun_u unlabeled examples {xiu}i=1nu\{x_i^u\}_{i=1}^{n_u} drawn i.i.d. from N(y⋅θ∗,σ2I)\mathcal{N}(y \cdot \theta^*, \sigma^2 I) with y∈{−1,1}y \in \{-1, 1\}, the standard pseudo-labeling procedure trains an initial estimator θ^init=1nl∑i=1nlxiyi\hat{\theta}_{\text{init}} = \frac{1}{n_l} \sum_{i=1}^{n_l} x_i y_i, assigns pseudo-labels yiu=sgn(θ^init⊤xiu)y_i^u = \text{sgn}(\hat{\theta}_{\text{init}}^\top x_i^u), and constructs the final classifier parameter: θ^final=1nl+nu(∑i=1nlxiyi+∑i=1nuxiuyiu).\hat{\theta}_{\text{final}} = \frac{1}{n_l + n_u} \left( \sum_{i=1}^{n_l} x_i y_i + \sum_{i=1}^{n_u} x_i^u y_i^u \right).

    The Mixup semi-supervised estimator θ^mix,final\hat{\theta}_{\text{mix,final}} applies Mixup over the pooled dataset of nl+nun_l + n_u labeled and pseudo-labeled pairs: θ^mix,final=Eλ∼Beta(α,β)[1(nl+nu)2∑i,j=1nl+nu(λxi+(1−λ)xj)(λyi+(1−λ)yj)],\hat{\theta}_{\text{mix,final}} = \mathbb{E}_{\lambda \sim \text{Beta}(\alpha, \beta)} \left[ \frac{1}{(n_l + n_u)^2} \sum_{i,j=1}^{n_l + n_u} (\lambda x_i + (1 - \lambda) x_j)(\lambda y_i + (1 - \lambda) y_j) \right], forming classifier C^mix,final(x)=sgn(θ^mix,final⊤x)\hat{C}_{\text{mix,final}}(x) = \text{sgn}(\hat{\theta}_{\text{mix,final}}^\top x).

    If C1<∥θ∗∥2<C2C_1 < \|\theta^*\|_2 < C_2 for universal constants C1,C2>0C_1, C_2 > 0, then for sufficiently large dimension pp and sample sizes nl,nun_l, n_u, there exist α,β>0\alpha, \beta > 0 such that, with high probability: ECE(C^mix,final)<ECE(C^final).\text{ECE}(\hat{C}_{\text{mix,final}}) < \text{ECE}(\hat{C}_{\text{final}}).

    Thus, applying Mixup during the final pseudo-label training step provably and consistently improves expected calibration error in semi-supervised learning.

  5. Knowl 5 — Dual Regime Behavior of Standard Pseudo-Labeling on Calibration Error

    theoretical result

    Consider semi-supervised learning under the (θ∗,σ)(\theta^*, \sigma)-Gaussian mixture model with nln_l labeled samples and nun_u unlabeled samples. An initial classifier C^init(x)=sgn(θ^init⊤x)\hat{C}_{\text{init}}(x) = \text{sgn}(\hat{\theta}_{\text{init}}^\top x) trained on labeled data assigns pseudo-labels yiu=C^init(xiu)y_i^u = \hat{C}_{\text{init}}(x_i^u), and a final classifier C^final\hat{C}_{\text{final}} is trained on the pooled set of labeled and pseudo-labeled samples.

    The calibration performance of standard pseudo-labeling exhibits two contrasting regimes:

    1. Insufficient labeled data regime: When C1p/nl≤∥θ∗∥2≤C2p/nlC_1 \sqrt{p/n_l} \le \|\theta^*\|_2 \le C_2 \sqrt{p/n_l} for universal constants C1<1/2C_1 < 1/2 and C2>2C_2 > 2, and when p/nlp/n_l, ∥θ∗∥2\|\theta^*\|_2, and nun_u are sufficiently large, pseudo-labeling provably improves calibration with high probability: ECE(C^final)<ECE(C^init).\text{ECE}(\hat{C}_{\text{final}}) < \text{ECE}(\hat{C}_{\text{init}}).

    2. Sufficient labeled data regime: When C1∥θ∗∥2<C2C_1 \|\theta^*\|_2 < C_2 and p<C3p < C_3 for positive constants C1,C2,C3>0C_1, C_2, C_3 > 0, and both nl,nu→∞n_l, n_u \to \infty with dimension pp fixed, pseudo-labeling degrades calibration with high probability: ECE(C^init)<ECE(C^final).\text{ECE}(\hat{C}_{\text{init}}) < \text{ECE}(\hat{C}_{\text{final}}).

    Hence, pseudo-labeling alone cannot reliably guarantee better calibration and can hurt calibration when labeled data is sufficient.

  6. Knowl 6 — Mixup Calibration Robustness under Out-of-Domain Distribution Shift

    theoretical result

    Let training data be generated from an in-domain Gaussian mixture model x∣y∼N(y⋅θ∗,σ2I)x \mid y \sim \mathcal{N}(y \cdot \theta^*, \sigma^2 I) with label prior P(y=1)=P(y=−1)=1/2\mathbb{P}(y = 1) = \mathbb{P}(y = -1) = 1/2. The linear classifier C^(x)=sgn(θ^⊤x)\hat{C}(x) = \text{sgn}(\hat{\theta}^\top x) and its Mixup counterpart C^mix(x)=sgn(θ^mix⊤x)\hat{C}_{\text{mix}}(x) = \text{sgn}(\hat{\theta}_{\text{mix}}^\top x) are fitted on nn in-domain samples in dimension pp.

    Suppose predictive calibration is evaluated out-of-domain under a shifted Gaussian mixture distribution where x∣y∼N(y⋅θ′,σ2I)x \mid y \sim \mathcal{N}(y \cdot \theta', \sigma^2 I) with shifted mean vector θ′∈Rp\theta' \in \mathbb{R}^p. If the domain shift satisfies the inner product condition: (θ′−θ∗)⊤θ∗≤p2n,(\theta' - \theta^*)^\top \theta^* \le \frac{p}{2n}, then for sufficiently large pp and nn, there exist parameters α,β>0\alpha, \beta > 0 for Dλ=Beta(α,β)\mathcal{D}_\lambda = \text{Beta}(\alpha, \beta) such that, with high probability: ECE(C^mix;θ′,σ)<ECE(C^;θ′,σ),\text{ECE}(\hat{C}_{\text{mix}}; \theta', \sigma) < \text{ECE}(\hat{C}; \theta', \sigma), where ECE(⋅;θ′,σ)\text{ECE}(\cdot; \theta', \sigma) denotes the expected calibration error computed with respect to the out-of-domain distribution. This demonstrates that the calibration benefits of Mixup persist under bounded distribution shifts.

  7. Knowl 7 — Mixup Calibration Improvement under Gaussian Generative Models

    theoretical result

    Let observed data (x,y)∈Rd×{−1,1}(x, y) \in \mathbb{R}^d \times \{-1, 1\} follow a (θ∗,g)(\theta^*, g)-Gaussian generative model with d≥pd \ge p, where latent representation z∣y∼N(y⋅θ∗,I)z \mid y \sim \mathcal{N}(y \cdot \theta^*, I) with label prior P(y=1)=P(y=−1)=1/2\mathbb{P}(y = 1) = \mathbb{P}(y = -1) = 1/2, and observations are generated through a nonlinear generator x=g(z)x = g(z). Suppose an encoder h^:Rd→Rp\hat{h}: \mathbb{R}^d \to \mathbb{R}^p approximately inverts gg such that for any v∈Rpv \in \mathbb{R}^p and k∈{−1,1}k \in \{-1, 1\}, the density pR1p_{R_1} of R1=v⊤h^(x)R_1 = v^\top \hat{h}(x) relates to the density pR2p_{R_2} of R2=v⊤z∼N(kv⊤θ∗,∥v∥22)R_2 = v^\top z \sim \mathcal{N}(k v^\top \theta^*, \|v\|_2^2) via pR1(u)=pR2(u)(1+δu)p_{R_1}(u) = p_{R_2}(u)(1 + \delta_u) with ER1[∣δu∣]=o(1)\mathbb{E}_{R_1}[|\delta_u|] = o(1) as n→∞n \to \infty.

    Given nn samples, define θ^=1n∑i=1nh^(xi)yi\hat{\theta} = \frac{1}{n} \sum_{i=1}^n \hat{h}(x_i) y_i and the corresponding Mixup estimator: θ^mix=∑i,j=1nEλ∼Beta(α,β)[(λh^(xi)+(1−λ)h^(xj))(λyi+(1−λ)yj)n2],\hat{\theta}_{\text{mix}} = \sum_{i,j=1}^n \mathbb{E}_{\lambda \sim \text{Beta}(\alpha, \beta)} \left[ \frac{(\lambda \hat{h}(x_i) + (1 - \lambda)\hat{h}(x_j))(\lambda y_i + (1 - \lambda)y_j)}{n^2} \right], with corresponding classifiers C^(x)=sgn(θ^⊤h^(x))\hat{C}(x) = \text{sgn}(\hat{\theta}^\top \hat{h}(x)) and C^mix(x)=sgn(θ^mix⊤h^(x))\hat{C}_{\text{mix}}(x) = \text{sgn}(\hat{\theta}_{\text{mix}}^\top \hat{h}(x)).

    If p/n∈(c1,c2)p/n \in (c_1, c_2), h^\hat{h} is LL-Lipschitz, and ∥θ∗∥2<C\|\theta^*\|_2 < C for universal constants c1,c2,L,C>0c_1, c_2, L, C > 0, then for sufficiently large pp and nn, there exist α,β>0\alpha, \beta > 0 such that, with high probability: ECE(C^mix)<ECE(C^).\text{ECE}(\hat{C}_{\text{mix}}) < \text{ECE}(\hat{C}).

    This generalizes Mixup calibration guarantees to broad classes of nonlinear data distributions, including representations learned via Generative Adversarial Networks (GANs).

  8. Knowl 8 — Mixup Calibration Benefits under Maximum Calibration Error (MCE)

    theoretical result

    Under the binary Gaussian mixture model x∣y∼N(y⋅θ∗,σ2I)x \mid y \sim \mathcal{N}(y \cdot \theta^*, \sigma^2 I) with label prior P(y=1)=P(y=−1)=1/2\mathbb{P}(y = 1) = \mathbb{P}(y = -1) = 1/2, the population Maximum Calibration Error (MCE) of a classifier C^\hat{C} with confidence function p^(x)\hat{p}(x) is defined as: MCE=max⁡v∈[0,1]∣P(y^=y∣p^=v)−v∣.\text{MCE} = \max_{v \in [0, 1]} |\mathbb{P}(\hat{y} = y \mid \hat{p} = v) - v|.

    The following theoretical properties hold for MCE:

    1. High-dimensional improvement: When p/n∈(c1,c2)p/n \in (c_1, c_2) and ∥θ∗∥2<C\|\theta^*\|_2 < C for universal constants c2>c1>0c_2 > c_1 > 0 and C>0C > 0, for sufficiently large p,np, n, there exist α,β>0\alpha, \beta > 0 such that λ∼Beta(α,β)\lambda \sim \text{Beta}(\alpha, \beta) satisfies with high probability: MCE(C^mix)<MCE(C^).\text{MCE}(\hat{C}_{\text{mix}}) < \text{MCE}(\hat{C}).

    2. Low-dimensional degradation: If p/n≤τ=o(1)p/n \le \tau = o(1) and ∥θ∗∥2<C\|\theta^*\|_2 < C, then for any fixed α,β>0\alpha, \beta > 0 and sufficiently large nn, with high probability: MCE(C^)<MCE(C^mix).\text{MCE}(\hat{C}) < \text{MCE}(\hat{C}_{\text{mix}}).

    3. Monotonic overparameterization scaling: For p/n→cratio∈(0,cmax)p/n \to c_{\text{ratio}} \in (0, c_{\text{max}}) and constant-level ∥θ∗∥2\|\theta^*\|_2, the marginal derivative ddαMCE(C^α,βmix)∣α→0+\left. \frac{d}{d\alpha} \text{MCE}(\hat{C}_{\alpha, \beta}^{\text{mix}}) \right|_{\alpha \to 0^+} is strictly negative and monotonically decreasing with respect to cratioc_{\text{ratio}}.

  9. Knowl 9 — Mixup-Augmented Pseudo-Labeling Algorithm for Semi-Supervised Learning

    algorithm

    The Mixup-augmented pseudo-labeling algorithm improves uncertainty calibration in semi-supervised classification by incorporating Mixup linear interpolation during retraining on combined labeled and pseudo-labeled datasets.

    Input: Labeled dataset Sl={(xi,yi)}i=1nlS_l = \{(x_i, y_i)\}_{i=1}^{n_l}, unlabeled dataset Su={xiu}i=1nuS_u = \{x_i^u\}_{i=1}^{n_u}, Mixup distribution Beta(α,β)\text{Beta}(\alpha, \beta)
    Output: Semi-supervised classifier C^mix,final\hat{C}_{\text{mix,final}}
    1. Train initial classifier C^init(x)=sgn(θ^init⊤x)\hat{C}_{\text{init}}(x) = \text{sgn}(\hat{\theta}_{\text{init}}^\top x) on labeled data SlS_l with θ^init=1nl∑i=1nlxiyi\hat{\theta}_{\text{init}} = \frac{1}{n_l} \sum_{i=1}^{n_l} x_i y_i
    2. Assign pseudo-labels to unlabeled data: yiu←C^init(xiu)y_i^u \leftarrow \hat{C}_{\text{init}}(x_i^u) for all i∈{1,…,nu}i \in \{1, \dots, n_u\}
    3. Pool labeled and pseudo-labeled datasets into Spool={(xi,yi)}i=1nl∪{(xiu,yiu)}i=1nuS_{\text{pool}} = \{(x_i, y_i)\}_{i=1}^{n_l} \cup \{(x_i^u, y_i^u)\}_{i=1}^{n_u}, indexed as {(xk,yk)}k=1nl+nu\{(x_k, y_k)\}_{k=1}^{n_l + n_u}
    4. Compute Mixup-averaged parameter estimator:
       θ^mix,final←Eλ∼Beta(α,β)[1(nl+nu)2∑i=1nl+nu∑j=1nl+nu(λxi+(1−λ)xj)(λyi+(1−λ)yj)]\hat{\theta}_{\text{mix,final}} \leftarrow \mathbb{E}_{\lambda \sim \text{Beta}(\alpha, \beta)} \left[ \frac{1}{(n_l + n_u)^2} \sum_{i=1}^{n_l + n_u} \sum_{j=1}^{n_l + n_u} (\lambda x_i + (1-\lambda) x_j)(\lambda y_i + (1-\lambda) y_j) \right]
    5. return C^mix,final(x)=sgn(θ^mix,final⊤x)\hat{C}_{\text{mix,final}}(x) = \text{sgn}(\hat{\theta}_{\text{mix,final}}^\top x)

    When implemented on deep neural networks, the initial network is trained on the labeled partition until the midpoint epoch, pseudo-labels are assigned to the unlabeled data via the network's argmax class predictions, and the model is trained for the remaining epochs with Mixup loss (λ∼Beta(α,α)\lambda \sim \text{Beta}(\alpha, \alpha), α=1.0\alpha = 1.0) on the combined pool.

  10. Knowl 10 — Empirical Scaling of Mixup Calibration Benefit with Neural Network Capacity

    empirical result

    Experiments on fully-connected neural networks and Pre-activation ResNet-18 architectures across CIFAR-10, CIFAR-100, Kuzushiji-MNIST, and Fashion-MNIST validate the theoretical predictions regarding model capacity and semi-supervised calibration:

    • Width Scaling: With fixed depth (8 hidden layers) and width varied from 10 to 3000 neurons, ECE for unaugmented fully-connected networks increases from ~0.10 to ~0.35 on CIFAR-10 and up to ~0.50 on CIFAR-100. Training with Mixup (extBeta(1.0,1.0) ext{Beta}(1.0, 1.0)) reduces ECE to ~0.05 on CIFAR-10 and ~0.10–0.15 on CIFAR-100 at larger widths, while offering no gain or slightly higher ECE on very narrow networks (width ≤10\le 10).
    • Depth Scaling: With fixed width and depth varied from 1 to 24 hidden layers, ECE of standard networks increases sharply with depth (reaching ~0.40 on CIFAR-10 and ~0.50 on CIFAR-100), whereas Mixup keeps ECE flat and low (<0.15< 0.15).
    • Semi-Supervised ResNets: Under 50/50 labeled/unlabeled splits, standard pseudo-labeling increases ECE on CIFAR-10 (from ~0.10 to ~0.16) and CIFAR-100 (from ~0.33 to ~0.38), whereas combining pseudo-labeling with Mixup consistently decreases ECE (to ~0.08 on CIFAR-10 and ~0.20 on CIFAR-100), as well as on Kuzushiji-MNIST and Fashion-MNIST.

Coverage note — None was omitted; all main theoretical theorems (Theorems 3.1–3.5, 4.1–4.3, 5.1–5.3), model/algorithm formulations, and comprehensive empirical findings are represented.

References

  1. 1.Bahnsen, A. C., Stojanovic, A., Aouada, D., and Ottersten, B. Improving credit card fraud detection with calibrated probabilities. In Proceedings of the 2014 SIAM international conference on data mining, pp. 677–685. SIAM, 2014.
  2. 2.Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C. Mixmatch: A holistic approach to semi-supervised learning. arXiv preprint arXiv:1905.02249, 2019.
  3. 3.Carmon, Y., Raghunathan, A., Schmidt, L., Liang, P., and Duchi, J. C. Unlabeled data improves adversarial robustness. Advances in Neural Information Processing Systems, 2019.
  4. 4.Chan, A., Alaa, A., Qian, Z., and Van Der Schaar, M. Unlabelled data improves bayesian uncertainty calibration under covariate shift. In International Conference on Machine Learning, pp. 1392–1402. PMLR, 2020.
  5. 5.Chapelle, O., Scholkopf, B., and Zien, A. Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]. IEEE Transactions on Neural Networks, 20(3):542–542, 2009.
  6. 6.Clanuwat, T., Bober-Irizar, M., Kitamoto, A., Lamb, A., Yamamoto, K., and Ha, D. Deep learning for classical japanese literature. In NeurIPS Creativity Workshop 2019, 2019.
  7. 7.Dan, C., Wei, Y., and Ravikumar, P. Sharp statistical guarantees for adversarially robust gaussian classification. In International Conference on Machine Learning, pp. 2345–2355. PMLR, 2020.
  8. 8.Deng, Z., Ding, F., Dwork, C., Hong, R., Parmigiani, G., Patil, P., and Sur, P. Representation via representations: Domain generalization via adversarially learned invariant representations. arXiv preprint arXiv:2006.11478, 2020a.
  9. 9.Deng, Z., He, H., Huang, J., and Su, W. Towards understanding the dynamics of the first-order adversaries. In International Conference on Machine Learning, pp. 2484–2493. PMLR, 2020b.
  10. 10.Deng, Z., Huang, J., and Kawaguchi, K. How shrinking gradient noise helps the performance of neural networks. In 2021 IEEE International Conference on Big Data (Big Data), pp. 1002–1007. IEEE, 2021a.
  11. 11.Deng, Z., Zhang, L., Ghorbani, A., and Zou, J. Improving adversarial robustness via unlabeled out-of-domain data. International Conference on Artificial Intelligence and Statistics, 2021b.
  12. 12.Deng, Z., Zhang, L., Vodrahalli, K., Kawaguchi, K., and Zou, J. Y. Adversarial training helps transfer learning via better representations. Advances in Neural Information Processing Systems, 34, 2021c.
  13. 13.Foster, D. P. and Stine, R. A. Variable selection in data mining: Building a predictive model for bankruptcy. Journal of the American Statistical Association, 99(466):303–313, 2004.
  14. 14.Foster, D. P. and Vohra, R. V. Calibrated learning and correlated equilibrium. Games and Economic Behavior, 21(1-2):40, 1997.
  15. 15.Gneiting, T. and Raftery, A. E. Weather forecasting with ensemble methods. Science, 310(5746):248–249, 2005.
  16. 16.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. arXiv preprint arXiv:1406.2661, 2014.
  17. 17.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In International Conference on Machine Learning, pp. 1321–1330. PMLR, 2017.
  18. 18.Guo, H., Mao, Y., and Zhang, R. Mixup as locally linear out-of-manifold regularization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 3714–3722, 2019.
  19. 19.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016a.
  20. 20.He, K., Zhang, X., Ren, S., and Sun, J. Identity mappings in deep residual networks. In European Conference on Computer Vision, pp. 630–645. Springer, 2016b.
  21. 21.Huang, Y., Li, W., Macheret, F., Gabriel, R. A., and Ohno-Machado, L. A tutorial on calibration measurements and calibration models for clinical prediction models. Journal of the American Medical Informatics Association, 27(4):621–633, 2020.
  22. 22.Ji, W., Deng, Z., Nakada, R., Zou, J., and Zhang, L. The power of contrast for feature learning: A theoretical analysis. arXiv preprint arXiv:2110.02473, 2021a.
  23. 23.Ji, W., Lu, Y., Zhang, Y., Deng, Z., and Su, W. J. An unconstrained layer-peeled perspective on neural collapse. arXiv preprint arXiv:2110.02796, 2021b.
  24. 24.Jiang, X., Osl, M., Kim, J., and Ohno-Machado, L. Calibrating predictive model estimates to support personalized medicine. Journal of the American Medical Informatics Association, 19(2):263–274, 2012.
  25. 25.Johnson, R. A., Wichern, D. W., et al. Applied multivariate statistical analysis, volume 5. Prentice hall Upper Saddle River, NJ, 2002.
  26. 26.Kawaguchi, K., Zhang, L., and Deng, Z. Understanding dynamics of nonlinear representation learning and its application. Neural Computation, 34(4):991–1018, 2022.
  27. 27.Kim, J.-H., Choo, W., and Song, H. O. Puzzle mix: Exploiting saliency and local statistics for optimal mixup. In International Conference on Machine Learning, pp. 5275–5285. PMLR, 2020.
  28. 28.Krizhevsky, A. and Hinton, G. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009.
  29. 29.Kuleshov, V., Fenner, N., and Ermon, S. Accurate uncertainties for deep learning using calibrated regression. In International Conference on Machine Learning, pp. 2796–2804. PMLR, 2018.
  30. 30.Kumar, A., Liang, P., and Ma, T. Verified uncertainty calibration. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pp. 3792–3803, 2019.
  31. 31.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. arXiv preprint arXiv:1612.01474, 2016.
  32. 32.Naeini, M. P., Cooper, G., and Hauskrecht, M. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 29, 2015.
  33. 33.Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., and Tran, D. Measuring calibration in deep learning. In CVPR Workshops, volume 2, 2019.
  34. 34.Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. Deep exploration via bootstrapped dqn. arXiv preprint arXiv:1602.04621, 2016.
  35. 35.Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. arXiv preprint arXiv:1906.02530, 2019.
  36. 36.Platt, J. et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in large margin classifiers, 10(3):61–74, 1999.
  37. 37.Roady, R., Hayes, T. L., and Kanan, C. Improved robustness to open set inputs via tempered mixup. In European Conference on Computer Vision, pp. 186–201. Springer, 2020.
  38. 38.Schmidt, L., Santurkar, S., Tsipras, D., Talwar, K., and Madry, A. Adversarially robust generalization requires more data. Advances in neural information processing systems, 31, 2018.
  39. 39.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  40. 40.Srivastava, R. K., Greff, K., and Schmidhuber, J. Highway networks. arXiv preprint arXiv:1505.00387, 2015.
  41. 41.Thulasidasan, S., Chennupati, G., Bilmes, J., Bhattacharya, T., and Michalak, S. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. arXiv preprint arXiv:1905.11001, 2019.
  42. 42.Tomani, C. and Buettner, F. Towards trustworthy predictions from deep neural networks with fast adversarial calibration. arXiv preprint arXiv:2012.10923, 2020.
  43. 43.Verma, V., Lamb, A., Beckham, C., Najafi, A., Mitliagkas, I., Lopez-Paz, D., and Bengio, Y. Manifold mixup: Better representations by interpolating hidden states. In International Conference on Machine Learning, pp. 6438–6447. PMLR, 2019.
  44. 44.Wen, Y., Jerfel, G., Muller, R., Dusenberry, M. W., Snoek, J., Lakshminarayanan, B., and Tran, D. Combining ensembles and data augmentation can harm your calibration. arXiv preprint arXiv:2010.09875, 2020.
  45. 45.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  46. 46.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  47. 47.Zhang, L., Deng, Z., Kawaguchi, K., Ghorbani, A., and Zou, J. How does mixup help with robustness and generalization? arXiv preprint arXiv:2010.04819, 2020.
  48. 48.Zhao, S., Ma, T., and Ermon, S. Individual calibration with randomized forecasting. In International Conference on Machine Learning, pp. 11387–11397. PMLR, 2020.
  49. 49.Zhu, X. and Goldberg, A. B. Introduction to semi-supervised learning. Synthesis lectures on artificial intelligence and machine learning, 3(1):1–130, 2009.
  50. 50.Zhu, X., Ghahramani, Z., and Lafferty, J. D. Semi-supervised learning using gaussian fields and harmonic functions. In Proceedings of the 20th International conference on Machine learning (ICML-03), pp. 912–919, 2003.

Citation

MLA
Zhang, L., et al. “When and How Mixup Improves Calibration”. International Conference on Machine Learning, vol. 162, 2022, pp. 26135–60, https://proceedings.mlr.press/v162/zhang22f.html.
APA
Zhang, L., Deng, Z., Kawaguchi, K., & Zou, J. (2022). When and How Mixup Improves Calibration. International Conference on Machine Learning, 162, 26135–26160. https://proceedings.mlr.press/v162/zhang22f.html
Chicago
Zhang, L., Z. Deng, K. Kawaguchi, and J. Zou. 2022. “When and How Mixup Improves Calibration”. International Conference on Machine Learning 162: 26135–60. https://proceedings.mlr.press/v162/zhang22f.html.
Harvard
Zhang, L. et al. (2022) “When and How Mixup Improves Calibration”, International Conference on Machine Learning. PMLR, pp. 26135–26160. Available at: https://proceedings.mlr.press/v162/zhang22f.html.
Vancouver
1. Zhang L, Deng Z, Kawaguchi K, Zou J (2022) When and How Mixup Improves Calibration. In: International Conference on Machine Learning. PMLR, pp 26135–26160

BibTeX

@InProceedings{pmlr-v162-zhang22f,
  title = 	 {When and How Mixup Improves Calibration},
  author =       {Zhang, Linjun and Deng, Zhun and Kawaguchi, Kenji and Zou, James},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {26135--26160},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/zhang22f/zhang22f.pdf},
  url = 	 {https://proceedings.mlr.press/v162/zhang22f.html},
  abstract = 	 {In many machine learning applications, it is important for the model to provide confidence scores that accurately capture its prediction uncertainty. Although modern learning methods have achieved great success in predictive accuracy, generating calibrated confidence scores remains a major challenge. Mixup, a popular yet simple data augmentation technique based on taking convex combinations of pairs of training examples, has been empirically found to significantly improve confidence calibration across diverse applications. However, when and how Mixup helps calibration is still a mystery. In this paper, we theoretically prove that Mixup improves calibration in high-dimensional settings by investigating natural statistical models. Interestingly, the calibration benefit of Mixup increases as the model capacity increases. We support our theories with experiments on common architectures and datasets. In addition, we study how Mixup improves calibration in semi-supervised learning. While incorporating unlabeled data can sometimes make the model less calibrated, adding Mixup training mitigates this issue and provably improves calibration. Our analysis provides new insights and a framework to understand Mixup and calibration.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/