Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness
Francesco PintoHarry YangSer Nam LimPhilip H. S. TorrPuneet K. Dokania
Proposes RegMixup, a simple modification that applies Mixup as an auxiliary regularizer alongside standard cross-entropy loss rather than as a standalone objective, substantially improving classification accuracy and out-of-distribution detection without requiring architectural changes or costly ensembles.
Deep neural networks are widely deployed in real-world applications but remain prone to failure when exposed to inputs outside their training distribution. Standard models frequently exhibit overconfidence on novel or corrupted data, leading to safety and reliability risks. While data-blending techniques such as Mixup improve accuracy and robustness under input corruptions, standard Mixup models struggle to identify completely novel out-of-distribution inputs. Because standard Mixup never trains on unmixed, clean samples, it yields diffuse, high-entropy predictions across all inputs, making it difficult to distinguish familiar data from novel data.
The article demonstrates that using Mixup as an additive regularizer alongside standard cross-entropy loss—an approach termed RegMixup—simultaneously enhances in-distribution classification accuracy, robustness under covariate shifts, and out-of-distribution detection reliability. The authors evaluate this approach across standard computer vision benchmarks (CIFAR-10, CIFAR-100, and ImageNet-1K) using WideResNet and ResNet architectures. The evaluation benchmarks RegMixup against standard deep neural networks, vanilla Mixup, specialized uncertainty methods, and compute-intensive ensemble models across multiple synthetic corruptions, natural distribution shifts, and distinct anomaly detection tasks.
RegMixup delivers consistent performance gains across all evaluated settings. First, it improves clean classification accuracy, achieving 97.46% on CIFAR-10 and 77.68% on ImageNet, outperforming both standard models and vanilla Mixup without sacrificing performance. Second, under input corruption on CIFAR-100-C, RegMixup improves accuracy by roughly 6.9 percentage points over standard networks and 2.45 percentage points over Mixup, while outperforming standard deep ensembles by 3.86 percentage points. Third, it significantly resolves Mixup's anomaly detection failure, boosting detection performance (measured by the area under the ROC curve) by over 9 percentage points when detecting SVHN digits against CIFAR datasets. In large-scale ImageNet tests, RegMixup achieved an anomaly detection score of 57.05%, outperforming vanilla Mixup (55.54%) and a 5-member deep ensemble (53.29%).
These findings indicate that combining standard empirical data training with blended vicinal data creates an effective entropy barrier that cleanly separates known classes from novel inputs. Practically, RegMixup provides a deterministic, drop-in replacement for standard training pipelines that improves both predictive accuracy and risk-aware uncertainty detection. Because it operates within a single neural network, it avoids the high computational, latency, and memory overheads associated with multi-model ensembles or custom architecture modifications.
Organizations seeking cost-effective model robustness should consider adopting RegMixup in vision classification pipelines. While training compute increases slightly by approximately 1.5 times relative to standard training, inference costs and model latency remain unchanged. Decision-makers should note that evaluations were conducted exclusively on vision classification benchmarks and certain baseline comparisons were unavailable across all architectures. Future work should validate the methodology on additional domains, such as vision transformers and non-image data modalities.
- Paper: mixup: Beyond Empirical Risk Minimization, Hongyi Zhang et al. (2017). This foundational paper introduces the original Mixup data interpolation training objective that the source directly analyzes and reformulates into an auxiliary regularizer.
- Paper: Manifold Mixup: Better Representations by Interpolating Hidden States, Vikas Verma et al. (2018). This work extends Mixup to hidden representations to regularize decision boundaries and calibrate confidence, providing key background on how interpolation regularizers affect network uncertainty.
- Paper: CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features, Sangdoo Yun et al. (2019). This paper presents CutMix, an essential mixing-based augmentation benchmark against which regularization methods and predictive robustness are directly evaluated.
- Paper: AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty, Dan Hendrycks et al. (2019). This study demonstrates how mixing-based data processing and consistency losses enhance uncertainty calibration under distribution shifts, contextualizing the robustness evaluation in the source.
- Paper: A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks, Dan Hendrycks et al. (2017). This work establishes the baseline methodology and evaluation metrics for detecting out-of-distribution examples via prediction probabilities, which the source relies on for OOD testing.
- Paper: Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift, Yaniv Ovadia et al. (2019). This benchmark systematically analyzes predictive uncertainty estimation under covariate and dataset shifts, establishing the problem formulation and metrics explored by the source.
- Paper: When Does Label Smoothing Help?, Rafael Müller et al. (2019). This paper examines how soft targets improve probability calibration and representation structure, illuminating why standard cross-entropy paired with regularization affects high-entropy predictive behavior.
- Paper: When and How Mixup Improves Calibration, Linjun Zhang et al. (2022). This theoretical study analyzes the precise mechanisms through which Mixup functions as a high-dimensional regularizer to improve model calibration and uncertainty behavior.
- Paper: Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning, Zhiqiang Shen et al. (2022). This paper builds on input and feature mixing principles to address overconfident predictions and improve representation robustness in self-supervised visual representation learning.
- Paper: Out-of-Distribution Detection with Deep Nearest Neighbors, Yiyou Sun et al. (2022). This work explores non-parametric nearest-neighbor approaches for out-of-distribution detection to overcome feature space anomalies and overconfident misclassifications.
- Paper: Boosting Out-of-distribution Detection with Typical Features, Yao Zhu et al. (2022). This article advances out-of-distribution detection by rectifying extreme internal activation values to improve uncertainty metrics without model retraining.
- Paper: Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement, Kai Xu et al. (2024). This study builds on post-hoc out-of-distribution detection techniques by investigating the theoretical and empirical impacts of activation scaling on model uncertainty.
