Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness

Francesco PintoHarry YangSer Nam LimPhilip H. S. TorrPuneet K. Dokania

article2022NeurIPS69 citations

Proposes RegMixup, a simple modification that applies Mixup as an auxiliary regularizer alongside standard cross-entropy loss rather than as a standalone objective, substantially improving classification accuracy and out-of-distribution detection without requiring architectural changes or costly ensembles.

Listen

Deep neural networks are widely deployed in real-world applications but remain prone to failure when exposed to inputs outside their training distribution. Standard models frequently exhibit overconfidence on novel or corrupted data, leading to safety and reliability risks. While data-blending techniques such as Mixup improve accuracy and robustness under input corruptions, standard Mixup models struggle to identify completely novel out-of-distribution inputs. Because standard Mixup never trains on unmixed, clean samples, it yields diffuse, high-entropy predictions across all inputs, making it difficult to distinguish familiar data from novel data.

The article demonstrates that using Mixup as an additive regularizer alongside standard cross-entropy loss—an approach termed RegMixup—simultaneously enhances in-distribution classification accuracy, robustness under covariate shifts, and out-of-distribution detection reliability. The authors evaluate this approach across standard computer vision benchmarks (CIFAR-10, CIFAR-100, and ImageNet-1K) using WideResNet and ResNet architectures. The evaluation benchmarks RegMixup against standard deep neural networks, vanilla Mixup, specialized uncertainty methods, and compute-intensive ensemble models across multiple synthetic corruptions, natural distribution shifts, and distinct anomaly detection tasks.

RegMixup delivers consistent performance gains across all evaluated settings. First, it improves clean classification accuracy, achieving 97.46% on CIFAR-10 and 77.68% on ImageNet, outperforming both standard models and vanilla Mixup without sacrificing performance. Second, under input corruption on CIFAR-100-C, RegMixup improves accuracy by roughly 6.9 percentage points over standard networks and 2.45 percentage points over Mixup, while outperforming standard deep ensembles by 3.86 percentage points. Third, it significantly resolves Mixup's anomaly detection failure, boosting detection performance (measured by the area under the ROC curve) by over 9 percentage points when detecting SVHN digits against CIFAR datasets. In large-scale ImageNet tests, RegMixup achieved an anomaly detection score of 57.05%, outperforming vanilla Mixup (55.54%) and a 5-member deep ensemble (53.29%).

These findings indicate that combining standard empirical data training with blended vicinal data creates an effective entropy barrier that cleanly separates known classes from novel inputs. Practically, RegMixup provides a deterministic, drop-in replacement for standard training pipelines that improves both predictive accuracy and risk-aware uncertainty detection. Because it operates within a single neural network, it avoids the high computational, latency, and memory overheads associated with multi-model ensembles or custom architecture modifications.

Organizations seeking cost-effective model robustness should consider adopting RegMixup in vision classification pipelines. While training compute increases slightly by approximately 1.5 times relative to standard training, inference costs and model latency remain unchanged. Decision-makers should note that evaluations were conducted exclusively on vision classification benchmarks and certain baseline comparisons were unavailable across all architectures. Future work should validate the methodology on additional domains, such as vision transformers and non-image data modalities.

Cover for Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness

Abstract

We show that the effectiveness of the well celebrated Mixup [Zhang et al., 2018] can be further improved if instead of using it as the sole learning objective, it is utilized as an additional regularizer to the standard cross-entropy loss. This simple change not only improves accuracy but also significantly improves the quality of the predictive uncertainty estimation of Mixup in most cases under various forms of covariate shifts and out-of-distribution detection experiments. In fact, we observe that Mixup otherwise yields much degraded performance on detecting out-of-distribution samples possibly, as we show empirically, due to its tendency to learn models exhibiting high-entropy throughout; making it difficult to differentiate in-distribution samples from out-of-distribution ones. To show the efficacy of our approach (RegMixup²), we provide thorough analyses and experiments on vision datasets (ImageNet & CIFAR-10/100) and compare it with a suite of recent approaches for reliable uncertainty estimation.

Table of Contents

  • 1 Introduction
  • 2 RegMixup: Mixup as a regularizer
  • 3 Related Works
  • 4 Experiments
  • 4.1 RegMixup Improves Accuracy on In-distribution and Covariate-shift Samples
  • 4.2 Out-of-Distribution Detection Experiments
  • 5 Conclusive Remarks
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — RegMixup training objective

    model/method

    RegMixup trains a classifier on both an unmodified training example and a Mixup interpolation. For a training pair (xi,yi)(x_i,y_i) and a randomly selected partner (xj,yj)(x_j,y_j) from the same mini-batch, it samples λ∼Beta⁡(α,α)\lambda\sim\operatorname{Beta}(\alpha,\alpha) and constructs

    x~i=λxi+(1−λ)xj,y~i=λyi+(1−λ)yj,\tilde{x}_i=\lambda x_i+(1-\lambda)x_j,\qquad \tilde{y}_i=\lambda y_i+(1-\lambda)y_j,

    where xi,xjx_i,x_j are inputs and yi,yjy_i,y_j are one-hot or soft class-label vectors. If pθ(⋅∣x)p_\theta(\cdot\mid x) is the softmax distribution of a neural network with parameters θ\theta, RegMixup minimizes

    L(θ)=CE⁡(pθ(⋅∣xi),yi)+η CE⁡(pθ(⋅∣x~i),y~i),\mathcal{L}(\theta)=\operatorname{CE}\bigl(p_\theta(\cdot\mid x_i),y_i\bigr)+\eta\,\operatorname{CE}\bigl(p_\theta(\cdot\mid\tilde{x}_i),\tilde{y}_i\bigr),

    where CE⁡(p,y)=−∑cyclog⁡pc\operatorname{CE}(p,y)=-\sum_c y_c\log p_c is cross-entropy over classes cc, and η≥0\eta\ge 0 weights the Mixup term. The paper uses η=1\eta=1 in practice, requiring no architectural change and producing the same inference cost as a standard single network. Cross-validation typically selects α∈{10,20}\alpha\in\{10,20\} for RegMixup, whereas standard Mixup commonly selects a much smaller value near 0.20.2.

  2. Knowl 2 — Explicit empirical–vicinal risk mixture

    model/method

    RegMixup interprets its objective as learning from an explicit mixture of the empirical training distribution and the Mixup vicinal distribution. For nn training examples (xi,yi)(x_i,y_i), let δxi(x)δyi(y)\delta_{x_i}(x)\delta_{y_i}(y) denote the point mass at the clean example and let Pmix(i)(x,y)P_{\mathrm{mix}}^{(i)}(x,y) denote the distribution obtained by pairing (xi,yi)(x_i,y_i) with another training example and sampling the Mixup coefficient λ∼Beta⁡(α,α)\lambda\sim\operatorname{Beta}(\alpha,\alpha). RegMixup uses

    Pγ(x,y)=1n∑i=1n[γ δxi(x)δyi(y)+(1−γ)Pmix(i)(x,y)],P_\gamma(x,y)=\frac{1}{n}\sum_{i=1}^{n}\left[\gamma\,\delta_{x_i}(x)\delta_{y_i}(y)+(1-\gamma)P_{\mathrm{mix}}^{(i)}(x,y)\right],

    with mixture weight γ∈[0,1]\gamma\in[0,1]. Dividing the practical loss by 1+η1+\eta gives this distribution with γ=1/(1+η)\gamma=1/(1+\eta); hence the paper's default η=1\eta=1 corresponds to equal empirical and Mixup components, γ=0.5\gamma=0.5. Unlike standard Mixup's one-sample approximation, this construction always exposes the network to clean training samples, regardless of the number of Monte Carlo samples or the value of α\alpha. The clean term allows cross-validation to select stronger interpolations without sacrificing performance on unperturbed inputs.

  3. Knowl 3 — Entropy-barrier mechanism of RegMixup

    empirical result

    The paper attributes standard Mixup's poor out-of-distribution discrimination to uniformly high predictive entropy: because training almost never presents an exactly clean sample and uses softened targets, standard Mixup becomes less confident on both in-distribution and out-of-distribution inputs. RegMixup behaves differently. Its clean cross-entropy term encourages low entropy near the endpoints λ≈0\lambda\approx0 and λ≈1\lambda\approx1, while its selected strong interpolations, usually with λ≈0.5\lambda\approx0.5, encourage high entropy between samples from different classes.

    For λ≈0.5\lambda\approx0.5, the Mixup target places nearly equal probability on the two source labels. Minimizing cross-entropy on such heavily interpolated inputs therefore acts as a soft proxy for maximizing entropy over those labels; exact entropy maximization would assign probability 0.50.5 to each. The resulting low-entropy endpoints and high-entropy interior form an entropy barrier that can separate ordinary in-distribution inputs from heavily mixed, OOD-like inputs. This mechanism is the paper's explanation for why RegMixup improves OOD uncertainty without deliberately sacrificing clean-data accuracy.

  4. Knowl 4 — Interpolation entropy-profile experiment

    empirical result

    To test the entropy mechanism, the authors trained WideResNet28-10 models on CIFAR-10, randomly selected 1,000 pairs of images with different class labels, and generated 20 convex combinations for each pair using equally spaced interpolation coefficients λ∈[0,1]\lambda\in[0,1]. The resulting 20,000 samples were evaluated with vanilla cross-entropy training (DNN), standard Mixup, and RegMixup; heat-map intensity represented the number of samples at each predictive-entropy and λ\lambda location.

    The vanilla DNN produced low entropy across nearly all interpolation coefficients, indicating overconfident predictions even between different classes. Standard Mixup produced high entropy across nearly the entire interpolation range, including near the original in-distribution samples. RegMixup produced low entropy near both endpoints and substantially higher entropy throughout the intermediate range, thereby yielding the intended entropy barrier.

  5. Knowl 5 — Evaluation protocol and benchmark scope

    experimental setup

    The evaluation uses WideResNet28-10 and ResNet50 on CIFAR-10 and CIFAR-100, and ResNet50 on ImageNet-1K. Reported metrics are averages over five random seeds. Covariate-shift evaluation uses CIFAR-10-C and CIFAR-100-C, each containing 15 synthetic corruption types at five severity levels; CIFAR-10.1 and CIFAR-10.2 test natural CIFAR-10 shifts; and ImageNet-A, ImageNet-R, ImageNet-V2, and ImageNet-Sketch test ImageNet shifts. OOD detection uses C100, SVHN, and Tiny-ImageNet for CIFAR-10-trained models, C10, SVHN, and Tiny-ImageNet for CIFAR-100-trained models, and ImageNet-O for ImageNet-trained models.

    The comparisons include vanilla cross-entropy DNNs, standard Mixup, spectral- and stable-rank-normalized DNNs, SNGP, DUQ, KFAC-LLLA, AugMix, and five-member Deep Ensembles where feasible. Hyperparameters are cross-validated using a 10% split held out from testing, and no method receives an external dataset or prior knowledge of the future shift or OOD distribution. RegMixup is approximately 1.5×1.5\times slower than vanilla DNN during training, compared with approximately 1.2×1.2\times for standard Mixup, but all three have the same inference requirements.

  6. Knowl 6 — CIFAR in-distribution accuracy

    data/table

    The reported CIFAR test accuracies compare models trained with WideResNet28-10 or ResNet50. RegMixup is the strongest single-model method in every listed setting and improves over both vanilla DNN and standard Mixup, while five-member Deep Ensembles are sometimes higher but require five models.

    WRN28-10 ResNet50
    Method C10 C100 C10 C100
    DNN 96.14 81.58 95.19 79.19
    Mixup 97.01 82.60 96.05 80.12
    RegMixup 97.46 83.25 96.71 81.52
    DNN-SN 96.22 81.60 95.20 79.27
    DNN-SRN 96.22 81.38 95.39 78.96
    SNGP 95.98 79.20 - -
    DUQ 94.7 - - -
    KFAC-LLLA 96.11 81.56 95.21 79.41
    AugMix 96.40 81.10 - -
    DE (5×\times) 96.75 83.85 96.23 82.09

    For the WideResNet trained on CIFAR-100, RegMixup exceeds DNN by 1.671.67 percentage points, standard Mixup by 0.650.65 points, and SNGP by 4.054.05 points. Thus, the uncertainty-oriented single-model baselines in this comparison generally trade away clean accuracy, whereas RegMixup does not.

  7. Knowl 7 — CIFAR covariate-shift accuracy

    data/table

    The following accuracies are averaged over corruption types and severities for CIFAR-C and are measured directly on the natural-shift datasets CIFAR-10.1 and CIFAR-10.2. RegMixup improves over standard Mixup and vanilla DNN in every listed architecture–dataset combination. AugMix is stronger on the synthetic corruption benchmarks, but RegMixup is stronger on the natural CIFAR shifts and has the best listed ResNet50 results among the applicable methods.

    WRN28-10 ResNet50
    Method C10-C C10.1 C10.2 C100-C C10-C C10.1 C10.2 C100-C
    DNN 76.60 90.73 84.79 52.54 75.18 88.58 83.31 50.62
    Mixup 81.68 91.29 86.55 56.99 78.63 90.03 84.61 53.96
    RegMixup 83.13 92.79 88.05 59.44 81.18 91.58 86.72 57.64
    DNN-SN 76.56 91.01 84.72 52.61 74.88 88.26 82.96 50.55
    DNN-SRN 77.21 90.88 85.24 52.54 75.40 88.61 83.49 50.48
    SNGP 78.37 90.80 84.95 57.23 - - - -
    DUQ 71.6 - - 50.4 - - - -
    KFAC-LLLA 76.56 90.68 84.68 52.57 75.18 88.34 83.50 50.85
    AugMix 90.02 91.6 85.9 68.15 - - - -
    DE (5×\times) 78.32 92.17 85.59 55.58 77.63 90.05 85.00 53.91

    For WideResNet, RegMixup improves over DNN by 6.536.53 points on C10-C and 6.906.90 points on C100-C, and over standard Mixup by 1.451.45 and 2.452.45 points, respectively. On C10.2 it improves over DNN by 3.263.26 points, over Mixup by 1.501.50 points, and over Deep Ensembles by 2.462.46 points. The authors suggest that AugMix's exceptional C10-C and C100-C performance is partly specific to synthetic corruptions resembling its training augmentations.

  8. Knowl 8 — ImageNet accuracy and shift robustness

    data/table

    ResNet50 experiments on ImageNet-1K evaluate ordinary test accuracy, four covariate-shift datasets, and ImageNet-O for OOD detection. RegMixup improves over DNN, standard Mixup, and AugMix on the in-distribution test and on every listed covariate-shift dataset. Deep Ensembles achieve the highest in-distribution accuracy and are best on ImageNet-Sketch, but RegMixup is competitive and exceeds them on ImageNet-A, ImageNet-R, and ImageNet-V2.

    Method IND ImageNet-R ImageNet-A ImageNet-V2 ImageNet-Sketch ImageNet-O AUROC
    DNN 76.60 36.41 2.76 64.53 24.72 55.97
    Mixup 77.15 39.05 3.29 64.58 26.34 55.54
    RegMixup 77.68 39.76 5.96 65.66 26.98 57.05
    AugMix 76.88 38.29 2.63 64.94 25.61 56.91
    DE (5×\times) 78.22 38.94 2.11 66.68 27.03 53.29

    All accuracy values are percentages, and the final column is AUROC. RegMixup exceeds DNN by 1.081.08 percentage points on ImageNet test accuracy and standard Mixup by 0.530.53 points. Its OOD AUROC is 57.0557.05, higher than DNN, standard Mixup, AugMix, and the five-member Deep Ensemble.

  9. Knowl 9 — CIFAR out-of-distribution detection

    data/table

    OOD detection is evaluated by AUROC, with higher values indicating better separation between in-distribution and OOD samples. Each group specifies the training architecture, the in-distribution dataset, and the OOD datasets. RegMixup consistently improves substantially over standard Mixup and is the strongest or near-strongest single-model method across the CIFAR settings.

    WRN, C10 ID WRN, C100 ID RN50, C10 ID RN50, C100 ID
    Method C100 SVHN T-ImageNet C10 SVHN T-ImageNet C100 SVHN T-ImageNet C10 SVHN T-ImageNet
    DNN 88.61 96.00 86.44 81.06 79.68 80.99 88.61 93.20 87.82 79.33 82.45 79.89
    Mixup 83.17 87.53 84.02 78.37 78.68 80.61 84.24 89.40 84.89 77.02 76.86 80.14
    RegMixup 89.63 96.72 90.19 81.27 89.32 83.13 89.63 95.39 90.04 79.44 88.66 82.56
    DNN-SN 88.56 95.59 87.71 81.10 83.43 82.26 88.19 93.46 87.55 79.20 80.78 79.90
    DNN-SRN 88.46 96.12 87.43 81.26 85.51 82.41 88.82 93.54 87.82 78.77 82.39 79.70
    SNGP 90.61 95.25 90.01 79.05 86.78 82.60 - - - - - -
    KFAC-LLLA 89.33 94.17 87.81 81.04 80.32 81.57 89.54 93.13 88.32 79.30 82.80 80.17
    AugMix 89.78 91.3 88.99 81.10 76.64 80.56 - - - - - -
    DE (5×\times) 91.25 97.53 89.52 83.26 85.07 83.40 91.38 96.90 90.5 81.93 85.08 82.15

    For WideResNet trained on either CIFAR-10 or CIFAR-100, RegMixup improves over standard Mixup by more than 99 AUROC points when SVHN is the OOD dataset. RegMixup outperforms every listed baseline in the ImageNet-scale OOD experiment as well, although in the CIFAR comparison SNGP exceeds it by 0.980.98 AUROC when WideResNet is trained on CIFAR-10 and C100 is used as OOD.

  10. Knowl 10 — Evaluation limitations and trade-offs

    limitation

    RegMixup retains the inference cost and architecture of a single deterministic neural network, but its training cost is approximately 1.5×1.5\times that of vanilla DNN training. The benchmark is also incomplete for some baselines: DUQ was unstable on CIFAR-100, the official SNGP implementation did not yield promising ResNet50 CIFAR results, and AugMix did not yield promising ResNet50 results on CIFAR-10 or CIFAR-100; these outcomes were not reported as numerical comparisons. Deep Ensembles remain competitive and sometimes better, particularly for ImageNet in-distribution accuracy and some covariate-shift settings, while SNGP wins one listed CIFAR OOD setting. The evaluation assumes only in-distribution training data and does not establish performance for shifts or OOD distributions beyond the tested benchmarks.

Coverage note — No substantial main-text contribution was omitted; calibration analyses and preliminary CutMix/Vision-Transformer analyses are mentioned as appendical material but are not detailed in the provided paper text.

References

  1. 1.Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov. Pitfalls of in-domain uncertainty estimation and ensembling in deep learning. arXiv preprint arXiv:2002.06470, 2020.
  2. 2.Christopher M. Bishop. Pattern Recognition and Machine Learning (Information Science and Statistics). Springer-Verlag, Berlin, Heidelberg, 2006. ISBN 0387310738.
  3. 3.Olivier Chapelle, Jason Weston, Léon Bottou, and Vladimir Vapnik. Vicinal risk minimization. In 13th International Conference on Neural Information Processing Systems, 2000.
  4. 4.Tianqi Chen, Emily Fox, and Carlos Guestrin. Stochastic gradient hamiltonian monte carlo. In ICML, 2014.
  5. 5.Youngseog Chung, Willie Neiswanger, Ian Char, and Jeff Schneider. Beyond pinball loss: Quantile methods for calibrated uncertainty quantification. In Advances in Neural Information Processing Systems, 2021.
  6. 6.Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. Certified adversarial robustness via randomized smoothing. In ICML, 2019.
  7. 7.Felix Dangel, Frederik Kunstner, and Philipp Hennig. Backpack: Packing more into backprop. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=BJlrF24twB.
  8. 8.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  10. 10.Alain Durmus, Umut Simsekli, Eric Moulines, Roland Badeau, and Gaël RICHARD. Stochastic gradient richardson-romberg markov chain monte carlo. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, 2016.
  11. 11.Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML’16, page 1050–1059. JMLR.org, 2016.
  12. 12.Yonatan Geifman and Ran El-Yaniv. Selectivenet: A deep neural network with an integrated reject option. In International Conference on Machine Learning, pages 2151–2159. PMLR, 2019.
  13. 13.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 1321–1330. JMLR.org, 2017.
  14. 14.Dongyoon Han, Jiwhan Kim, and Junmo Kim. Deep pyramidal residual networks. CVPR, 2017.
  15. 15.Marton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu, Jasper Snoek, Balaji Lakshminarayanan, Andrew Mingbo Dai, and Dustin Tran. Training independent subnetworks for robust prediction. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=OGg9XnKxFAH.
  16. 16.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. doi: 10.1109/CVPR.2016.90.
  17. 17.D Hendrycks, S Basart, N Mu, S Kadavath, and others. The many faces of robustness: A critical analysis of out-of-distribution generalization. arXiv preprint arXiv, 2020a.
  18. 18.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. Proceedings of the International Conference on Learning Representations, 2019.
  19. 19.Dan Hendrycks, Mantas Mazeika, Saurav Kadavath, and Dawn Song. Using self-supervised learning can improve model robustness and uncertainty. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, 2019a.
  20. 20.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. July 2019b.
  21. 21.Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: A simple data processing method to improve robustness and uncertainty. Proceedings of the International Conference on Learning Representations (ICLR), 2020b.
  22. 22.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICCV, 2021.
  23. 23.Marius Hobbhahn, Agustinus Kristiadi, and Philipp Hennig. Fast predictive uncertainty for classification with bayesian deep networks, 2021. URL https://openreview.net/forum?id=KcImcc3j-qS.
  24. 24.Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger. Snapshot ensembles: Train 1, get m for free, 2017a.
  25. 25.Gao Huang, Zhuang Liu, and Kilian Q. Weinberger. Densely connected convolutional networks. CVPR, 2017b.
  26. 26.Durk P Kingma, Tim Salimans, and Max Welling. Variational dropout and the local reparameterization trick. In Advances in Neural Information Processing Systems, 2015.
  27. 27.Agustinus Kristiadi, Matthias Hein, and Philipp Hennig. Being bayesian, even just a bit, fixes overconfidence in ReLU networks. In Proceedings of the 37th ICML, pages 5436–5446, 2020.
  28. 28.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, 2017.
  29. 29.Ya Le and Xuan S. Yang. Tiny imagenet visual recognition challenge. 2015.
  30. 30.Mathias Lecuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu, and Suman Jana. Certified robustness to adversarial examples with differential privacy. In IEEE Symposium on Security and Privacy (SP), 2019.
  31. 31.Jeremiah Zhe Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax-Weiss, and Balaji Lakshminarayanan. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. In NeurIPS, June 2020a.
  32. 32.Weitang Liu, Xiaoyun Wang, John D. Owens, and Yixuan Li. Energy-based out-of-distribution detection, 2020b.
  33. 33.Shangyun Lu, 1 Bradley Nott, and 1 Aaron Olson. Harder or different? a closer look at distribution shift in dataset reproduction. http://www.gatsby.ucl.ac.uk/~balaji/udl2020/accepted-papers/UDL2020-paper-101.pdf. Accessed: 2021-11-10.
  34. 34.Zhiyun Lu, Eugene Ie, and Fei Sha. Uncertainty estimation with infinitesimal jackknife, its distribution and mean-field approximation, 2020.
  35. 35.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In ICLR, February 2018a.
  36. 36.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018b. URL https://openreview.net/forum?id=B1QRgziT-.
  37. 37.Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz, Philip HS Torr, and Puneet K Dokania. Calibrating deep neural networks using focal loss. In NeurIPS, 2020.
  38. 38.Rafael Müller, Simon Kornblith, and Geoffrey E Hinton. When does label smoothing help? In H Wallach, H Larochelle, A Beygelzimer, F dAlché-Buc, E Fox, and R Garnett, editors, Advances in Neural Information Processing Systems, volume 32, pages 4694–4703. Curran Associates, Inc., 2019.
  39. 39.Yuval Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  40. 40.Christine Osborne. Statistical calibration: A review. International Statistical Review / Revue Internationale de Statistique, 59(3):309–336, 1991. ISSN 03067734, 17515823. URL http://www.jstor.org/stable/1403690.
  41. 41.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift, 2019.
  42. 42.Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence. Dataset Shift in Machine Learning. The MIT Press, 2009. ISBN 0262170051.
  43. 43.Rahul Rahaman and Alexandre H. Thiery. Uncertainty quantification and deep ensembles, 2020.
  44. 44.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do CIFAR-10 classifiers generalize to CIFAR-10? June 2018.
  45. 45.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? CoRR, abs/1902.10811, 2019. URL http://arxiv.org/abs/1902.10811.
  46. 46.Hippolyt Ritter, Aleksandar Botev, and David Barber. A scalable laplace approximation for neural networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Skdvd2xAZ.
  47. 47.A Sanyal, P H S Torr, and P K Dokania. Stable rank normalization for improved generalization in neural networks and GANs. In ICLR, 2020.
  48. 48.Murat Sensoy, Lance Kaplan, and Melih Kandemir. Evidential deep learning to quantify classification uncertainty. In Advances in Neural Information Processing Systems, 2018.
  49. 49.Kendrick Shen, Robbie Jones, Ananya Kumar, Sang Michael Xie, Jeff Z. HaoChen, Tengyu Ma, and Percy Liang. Connect, not collapse: Explaining contrastive learning for unsupervised domain adaptation, 2022. URL https://arxiv.org/abs/2204.00570.
  50. 50.Yuge Shi, Jeffrey Seely, Philip H. S. Torr, N. Siddharth, Awni Hannun, Nicolas Usunier, and Gabriel Synnaeve. Gradient matching for domain generalization. arXiv preprint arXiv:2104.09937, 2021.
  51. 51.Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt. Measuring robustness to natural distribution shifts in image classification. July 2016.
  52. 52.Sunil Thulasidasan, Gopinath Chennupati, Jeff A Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. In Advances in Neural Information Processing Systems, 2019.
  53. 53.Joost van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal. Uncertainty estimation using a single deep deterministic neural network. In ICML, 2020.
  54. 54.V. Vapnik. Principles of risk minimization for learning theory. In Proceedings of the 4th International Conference on Neural Information Processing Systems, NIPS’91, page 831–838, San Francisco, CA, USA, 1991. Morgan Kaufmann Publishers Inc. ISBN 1558602224.
  55. 55.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems, pages 10506–10518, 2019.
  56. 56.Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Wenjun Zeng, and Tao Qin. Generalizing to unseen domains: A survey on domain generalization. arXiv preprint arXiv:2103.03097, 2021.
  57. 57.Yeming Wen, Dustin Tran, and Jimmy Ba. Batchensemble: an alternative approach to efficient ensemble and lifelong learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=Sklf1yrYDr.
  58. 58.Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael W Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, and Dustin Tran. Combining ensembles and data augmentation can harm your calibration. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=g11CZSghXyY.
  59. 59.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  60. 60.Huaxiu Yao, Yu Wang, Sai Li, Linjun Zhang, Weixin Liang, James Zou, and Chelsea Finn. Improving out-of-distribution robustness via selective augmentation, 2022. URL https://openreview.net/forum?id=zXne1klXIQ.
  61. 61.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. ICCV, 2019.
  62. 62.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In Edwin R. Hancock Richard C. Wilson and William A. P. Smith, editors, Proceedings of the British Machine Vision Conference (BMVC), pages 87.1–87.12. BMVA Press, September 2016. ISBN 1-901725-59-6. doi: 10.5244/C.30.87. URL https://dx.doi.org/10.5244/C.30.87.
  63. 63.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=r1Ddp1-Rb.
  64. 64.Ruqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen, and Andrew Gordon Wilson. Cyclical stochastic gradient mcmc for bayesian deep learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkeS1RVtPS.

Citation

MLA
Pinto, F., et al. “Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 14608–22, https://proceedings.neurips.cc/paper_files/paper/2022/file/5ddcfaad1cb72ce6f1a365e8f1ecf791-Paper-Conference.pdf.
APA
Pinto, F., Yang, H., Lim, S. N., Torr, P., & Dokania, P. (2022). Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness. Advances in Neural Information Processing Systems, 35, 14608–14622. https://proceedings.neurips.cc/paper_files/paper/2022/file/5ddcfaad1cb72ce6f1a365e8f1ecf791-Paper-Conference.pdf
Chicago
Pinto, F., H. Yang, S. N. Lim, P. Torr, and P. Dokania. 2022. “Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness”. Advances in Neural Information Processing Systems 35: 14608–22. https://proceedings.neurips.cc/paper_files/paper/2022/file/5ddcfaad1cb72ce6f1a365e8f1ecf791-Paper-Conference.pdf.
Harvard
Pinto, F. et al. (2022) “Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 14608–14622. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/5ddcfaad1cb72ce6f1a365e8f1ecf791-Paper-Conference.pdf.
Vancouver
1. Pinto F, Yang H, Lim SN, Torr P, Dokania P (2022) Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 14608–14622

BibTeX

@inproceedings{pinto2022using,
  title = {Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness},
  author = {Pinto, Francesco and Yang, Harry and Lim, Ser Nam and Torr, Philip and Dokania, Puneet},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {14608-14622},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/5ddcfaad1cb72ce6f1a365e8f1ecf791-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors