AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty

Dan HendrycksNorman MuEkin D. CubukBarret ZophJustin GilmerBalaji Lakshminarayanan

article2019ICLR1,674 citations

Presents AugMix, a computationally lightweight data augmentation technique that mixes diverse image transformations to help vision models withstand unforeseen corruptions and produce more reliable uncertainty estimates.

Listen

Modern computer vision models frequently underperform when deployed in real-world environments due to differences between training data and real-world conditions. Unforeseen data shifts, such as image noise, blurring, or weather effects, often cause sharp drops in classification accuracy. Furthermore, neural networks tend to remain overconfident when making incorrect predictions under these conditions, presenting significant safety and operational risks for mission-critical deployments.

The article demonstrates a novel data processing method, called AugMix, designed to enhance both the robustness and uncertainty calibration of image classifiers when facing unseen corruptions and perturbations, without sacrificing standard accuracy or adding heavy computational burdens.

The authors evaluate AugMix across multiple standard benchmarks, including CIFAR-10, CIFAR-100, and ImageNet, along with their corrupted and perturbed variants (CIFAR-C, CIFAR-P, ImageNet-C, and ImageNet-P). The approach combines diverse, randomized augmentation chains with elementwise image mixing, paired with a consistency loss function that enforces uniform model predictions across varied views of the same image. To rigorously test generalization, common test-time corruptions like blurring and noise were excluded from the training augmentations.

The evaluation yields several major findings. First, AugMix cuts corruption error rates roughly in half compared to standard baselines, reducing average error from 29.0% to 12.5% on CIFAR-10-C and from 55.6% to 38.3% on CIFAR-100-C. Second, on large-scale ImageNet-C benchmarks, AugMix achieves a state-of-the-art mean corruption error of 68.4%, down from 80.6% for standard training, while slightly improving standard clean accuracy to 22.4% error. Third, AugMix significantly stabilizes video frame predictions, decreasing the mean flip rate on ImageNet-P from 57.2% to 37.4%. Finally, the method substantially improves uncertainty calibration under data shift, reducing root mean square calibration error on CIFAR-10-C from 23.0% to 8.5%.

These findings indicate that organizations can deploy computer vision systems with greater reliability and lower operational risk. Unlike traditional techniques, such as adversarial training, which degrade clean image accuracy and require substantial compute, AugMix maintains baseline performance while improving resilience to environmental distortions. It also integrates smoothly with other robustness strategies, such as stylized training techniques, to achieve further performance gains.

Organizations developing vision systems should integrate AugMix into their existing model training workflows, given its low implementation complexity and minimal overhead. Practitioners should ensure augmentation selections are tailored to their domain to avoid distorting essential category features. Future work should explore combining this approach with emerging architectural defenses and testing performance on domain-specific real-world shifts.

Confidence in these findings is high across standard vision architectures and established robustness benchmarks. However, stakeholders should note that the evaluation focuses on standard image classification and simulated corruptions; performance should still be validated on complex downstream tasks such as object detection before deployment in high-stakes operational settings.

arXiv: 1912.02781google-research/augmix
Cover for AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty

Abstract

Modern deep neural networks can achieve high accuracy when the training distribution and test distribution are identically distributed, but this assumption is frequently violated in practice. When the train and test distributions are mismatched, accuracy can plummet. Currently there are few techniques that improve robustness to unforeseen data shifts encountered during deployment. In this work, we propose a technique to improve the robustness and uncertainty estimates of image classifiers. We propose AugMix, a data processing technique that is simple to implement, adds limited computational overhead, and helps models withstand unforeseen corruptions. AugMix significantly improves robustness and uncertainty measures on challenging image classification benchmarks, closing the gap between previous methods and the best possible performance in some cases by more than half.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 AugMix
  • 4 Experiments
  • 4.1 CIFAR-10 and CIFAR-100
  • 4.2 ImageNet
  • 4.3 Ablations
  • 5 Conclusion
  • References
  • A Hyperparameter Ablations
  • B Fourier Analysis
  • C Augmentation Operations
  • D Additional results
  • E Calibration Metrics

Knowls

  1. Knowl 1 — AugMix Data Augmentation Framework

    model/method

    AugMix is a data processing method designed to improve model robustness and uncertainty estimation under unseen distribution shifts while preserving or improving clean data accuracy.

    AugMix generates diverse training transformations by mixing several stochastically sampled augmentation chains via convex combinations and enforcing prediction consistency across augmented views using the Jensen-Shannon divergence:

    1. Augmentation Operations: Augmentation operations are sampled from AutoAugment primitives (such as rotation, shear, translation, posterize, equalize), specifically excluding any operations that overlap with common corruption benchmarks (ImageNet-C/CIFAR-C). In particular, brightness, contrast, color, sharpness, Cutout, blur, and noise operations are excluded so that corruption types remain strictly unseen during training.
    2. Stochastic Augmentation Chains: For an input image xorigx_{\text{orig}}, kk parallel augmentation chains (default k=3k=3) are constructed. Each chain randomly selects and composes between 1 and 3 augmentation operations with randomly sampled severities.
    3. Convex Mixing: The outputs of the kk chains are combined elementwise using mixing weights sampled from a symmetric Dirichlet distribution, (w1,…,wk)∼Dirichlet(α,…,α)(w_1, \dots, w_k) \sim \text{Dirichlet}(\alpha, \dots, \alpha).
    4. Skip Connection: The mixed augmented image is blended with the original input xorigx_{\text{orig}} via a convex combination weight m∼Beta(α,α)m \sim \text{Beta}(\alpha, \alpha), yielding xaugmix=mxorig+(1−m)∑i=1kwi chaini(xorig)x_{\text{augmix}} = m x_{\text{orig}} + (1-m) \sum_{i=1}^k w_i \, \text{chain}_i(x_{\text{orig}}). This preserves semantic structure and local statistics without drifting off the natural image manifold.
    5. Consistency Regularization: The model is trained on a classification loss over the clean image along with a Jensen-Shannon divergence loss penalizing divergence among the predictive distributions of the clean image and two independently sampled AugMix variants.
  2. Knowl 2 — AugMix Training Algorithm

    algorithm

    The AugMix training procedure takes an input image xorigx_{\text{orig}}, generates two distinct stochastic augmented versions using random augmentation chains and convex mixing, and computes a combined classification and Jensen-Shannon divergence consistency loss.

    Input: Classifier model p^\hat{p}, classification loss function L\mathcal{L}, input image xorigx_{\text{orig}}, ground-truth label yy, set of augmentation operations O\mathcal{O}, mixture width k=3k = 3, Dirichlet/Beta parameter α=1\alpha = 1, consistency loss multiplier λ\lambda
    Output: Total training loss Ltotal\mathcal{L}_{\text{total}}
    function AugmentAndMix(xorigx_{\text{orig}}, k=3k = 3, α=1\alpha = 1)
        xaug=0x_{\text{aug}} = \mathbf{0}
        Sample mixing weights (w1,w2,…,wk)∼Dirichlet(α,…,α)(w_1, w_2, \dots, w_k) \sim \text{Dirichlet}(\alpha, \dots, \alpha)
        for i=1,…,ki = 1, \dots, k do
            Sample operations op1,op2,op3∼O\text{op}_1, \text{op}_2, \text{op}_3 \sim \mathcal{O}
            Compose operations: op12=op2∘op1\text{op}_{12} = \text{op}_2 \circ \text{op}_1, op123=op3∘op2∘op1\text{op}_{123} = \text{op}_3 \circ \text{op}_2 \circ \text{op}_1
            Sample uniformly a chain depth: chain∼{op1,op12,op123}\text{chain} \sim \{\text{op}_1, \text{op}_{12}, \text{op}_{123}\}
            xaug=xaug+wi⋅chain(xorig)x_{\text{aug}} = x_{\text{aug}} + w_i \cdot \text{chain}(x_{\text{orig}})
        end for
        Sample interpolation weight m∼Beta(α,α)m \sim \text{Beta}(\alpha, \alpha)
        xaugmix=m⋅xorig+(1−m)⋅xaugx_{\text{augmix}} = m \cdot x_{\text{orig}} + (1 - m) \cdot x_{\text{aug}}
        return xaugmixx_{\text{augmix}}
    end function
    xaugmix1=AugmentAndMix(xorig,k,α)x_{\text{augmix1}} = \text{AugmentAndMix}(x_{\text{orig}}, k, \alpha)
    xaugmix2=AugmentAndMix(xorig,k,α)x_{\text{augmix2}} = \text{AugmentAndMix}(x_{\text{orig}}, k, \alpha)
    porig=p^(y∣xorig)p_{\text{orig}} = \hat{p}(y \mid x_{\text{orig}})
    paugmix1=p^(y∣xaugmix1)p_{\text{augmix1}} = \hat{p}(y \mid x_{\text{augmix1}})
    paugmix2=p^(y∣xaugmix2)p_{\text{augmix2}} = \hat{p}(y \mid x_{\text{augmix2}})
    Ltotal=L(porig,y)+λ⋅JS(porig;paugmix1;paugmix2)\mathcal{L}_{\text{total}} = \mathcal{L}(p_{\text{orig}}, y) + \lambda \cdot \text{JS}(p_{\text{orig}}; p_{\text{augmix1}}; p_{\text{augmix2}})
    return Ltotal\mathcal{L}_{\text{total}}

    Hyperparameter defaults are k=3k=3, α=1\alpha=1 (or α=0.5\alpha=0.5 on ImageNet), and consistency weight λ=12\lambda=12.

  3. Knowl 3 — Jensen-Shannon Divergence Consistency Loss

    equation

    To enforce smooth and consistent posterior prediction distributions across diverse semantic-preserving augmentations of the same input image xorigx_{\text{orig}}, AugMix optimizes the original classification loss on the unaugmented sample along with a Jensen-Shannon Divergence (JSD) consistency penalty across the clean image and two augmented samples xaugmix1,xaugmix2x_{\text{augmix1}}, x_{\text{augmix2}}:

    Ltotal=L(porig,y)+λ JS(porig;paugmix1;paugmix2)\mathcal{L}_{\text{total}} = \mathcal{L}(p_{\text{orig}}, y) + \lambda \, \text{JS}(p_{\text{orig}}; p_{\text{augmix1}}; p_{\text{augmix2}})

    where porig=p^(y∣xorig)p_{\text{orig}} = \hat{p}(y \mid x_{\text{orig}}), paugmix1=p^(y∣xaugmix1)p_{\text{augmix1}} = \hat{p}(y \mid x_{\text{augmix1}}), and paugmix2=p^(y∣xaugmix2)p_{\text{augmix2}} = \hat{p}(y \mid x_{\text{augmix2}}) are the model's predictive class probability distributions, yy is the ground-truth label, and λ\lambda is a balancing hyperparameter.

    The three-way Jensen-Shannon divergence is computed by first calculating the average predictive distribution:

    M=13(porig+paugmix1+paugmix2)M = \frac{1}{3} \left( p_{\text{orig}} + p_{\text{augmix1}} + p_{\text{augmix2}} \right)

    and then averaging the Kullback-Leibler divergences to MM:

    JS(porig;paugmix1;paugmix2)=13(KL[porig∥M]+KL[paugmix1∥M]+KL[paugmix2∥M])\text{JS}(p_{\text{orig}}; p_{\text{augmix1}}; p_{\text{augmix2}}) = \frac{1}{3} \left( \text{KL}[p_{\text{orig}} \parallel M] + \text{KL}[p_{\text{augmix1}} \parallel M] + \text{KL}[p_{\text{augmix2}} \parallel M] \right)

    Unlike arbitrary pairwise KL divergences, the Jensen-Shannon divergence is symmetric and bounded above by log⁡(C)\log(C), where CC is the number of target classes.

  4. Knowl 4 — Corruption Robustness on CIFAR-10-C and CIFAR-100-C

    data/table

    AugMix substantially reduces classification error on CIFAR-10-C and CIFAR-100-C across four diverse deep architectures: an All Convolutional Net (AllConvNet), DenseNet-BC (k=12,d=100k=12, d=100), Wide ResNet (40-2), and ResNeXt-29 (32×432 \times 4). Performance is reported as average unnormalized corruption error (uCE, in %) across 15 noise, blur, weather, and digital corruption types evaluated at five severity levels (1 to 5).

    Standard Cutout Mixup CutMix AutoAugment* Adv Training AugMix
    CIFAR-10-C
    AllConvNet 30.8 32.9 24.6 31.3 29.2 28.1 15.0
    DenseNet 30.7 32.1 24.6 33.5 26.6 27.6 12.7
    WideResNet 26.9 26.8 22.3 27.1 23.9 26.2 11.2
    ResNeXt 27.5 28.9 22.6 29.5 24.2 27.0 10.9
    Mean 29.0 30.2 23.5 30.3 26.0 27.2 12.5
    CIFAR-100-C
    AllConvNet 56.4 56.8 53.4 56.0 55.1 56.0 42.7
    DenseNet 59.3 59.6 55.4 59.2 53.9 55.2 39.6
    WideResNet 53.3 53.5 50.4 52.9 49.6 55.1 35.9
    ResNeXt 53.4 54.6 51.4 54.1 51.3 54.4 34.9
    Mean 55.6 56.1 52.6 55.5 52.5 55.2 38.3

    AugMix more than halves corruption error on CIFAR-10-C (from 29.0% to 12.5% mean error) and reduces CIFAR-100-C mean error from 55.6% to 38.3%, outperforming Cutout, Mixup, CutMix, AutoAugment* (with overlapping corruptions removed), and ℓp\ell_p Adversarial Training across all architectures without hyperparameter retuning.

  5. Knowl 5 — Corruption Robustness on ImageNet-C and Synergy with Stylized ImageNet

    data/table

    On ImageNet-C, AugMix trained with a ResNet-50 backbone achieves state-of-the-art robustness, reducing Mean Corruption Error (mCE, normalized by AlexNet corruption errors across 15 corruptions at 5 severities) from 80.6% to 68.4% while simultaneously improving clean validation error from 23.9% to 22.4%.

    Noise Blur Weather Digital
    Network Clean Gauss. Shot Impulse Defocus Glass Motion Zoom Snow Frost Fog Bright Contrast Elastic Pixel JPEG mCE
    Standard 23.9 79 80 82 82 90 84 80 86 81 75 65 79 91 77 80 80.6
    Patch Uniform 24.5 67 68 70 74 83 81 77 80 74 75 62 77 84 71 71 74.3
    AutoAugment* 22.8 69 68 72 77 83 80 81 79 75 64 56 70 88 57 71 72.7
    Random AA* 23.6 70 71 72 80 86 82 81 81 77 72 61 75 88 73 72 76.1
    MaxBlur pool 23.0 73 74 76 74 86 78 77 77 72 63 56 68 86 71 71 73.4
    SIN 27.2 69 70 70 77 84 76 82 74 75 69 65 69 80 64 77 73.3
    AugMix 22.4 65 66 67 70 80 66 66 75 72 67 58 58 79 69 69 68.4
    AugMix+SIN 25.2 61 62 61 69 77 63 72 66 68 63 59 52 74 60 67 64.9

    AugMix is complementary to Stylized ImageNet (SIN) training: combining AugMix with SIN achieves 64.9% mCE, improving over both standalone methods.

  6. Knowl 6 — Prediction Stability on ImageNet-P and CIFAR-10-P

    data/table

    Perturbation robustness evaluates prediction consistency across video sequences of continuous, small transformations (such as progressively increasing brightness, translation, rotation, tilt, scaling, or noise) rather than isolated corrupted stills. Robustness is quantified by the flip probability (the probability that adjacent video frames produce mismatched predictions) and normalized as the mean Flip Rate (mFR, relative to AlexNet) or unnormalized mean Flip Probability (mFP, in %).

    Noise Blur Weather Digital
    Network Clean Gaussian Shot Motion Zoom Snow Bright Translate Rotate Tilt Scale mFR
    Standard 23.9 57 55 62 65 66 65 43 53 57 49 57.2
    Patch Uniform 24.5 32 25 50 52 54 57 40 48 49 46 45.3
    AutoAugment* 22.8 50 45 57 68 63 53 40 44 50 46 51.7
    Random AA* 23.6 53 46 53 63 59 57 42 48 54 47 52.2
    SIN 27.2 53 50 57 72 51 62 43 53 57 53 55.0
    MaxBlur pool 23.0 52 51 59 63 57 64 34 43 49 40 51.2
    AugMix 22.4 46 41 30 47 38 46 25 32 35 33 37.4
    AugMix+SIN 25.2 45 40 30 54 32 48 27 35 38 39 38.9

    On ImageNet-P, AugMix lowers mFR from 57.2% to 37.4% (a ~20% absolute improvement). On CIFAR-10-P, AugMix reduces the mean Flip Probability across four architectures from 4.3% to 1.6%. While adversarial training also achieves low flip rates on CIFAR-10-P (2.2%), it severely degrades clean CIFAR-10 error (increasing error from 5.4% to 17.3%), whereas AugMix maintains or improves clean accuracy (5.1% error).

  7. Knowl 7 — Ablation of AugMix Constituents and Hyperparameters

    data/table

    An ablation study on CIFAR-10-C and CIFAR-100-C isolates the individual performance contributions of augmentation variety, the Jensen-Shannon divergence (JSD) consistency loss, and convex mixing using a Wide ResNet (40-2) backbone:

    Method CIFAR-10-C Error Rate (%) CIFAR-100-C Error Rate (%)
    Standard 26.9 53.3
    AutoAugment* 23.9 49.6
    Random AutoAugment* 17.0 43.6
    Random AutoAugment* + JSD Loss 14.7 40.8
    AugmentAndMix (No JSD Loss) 13.1 39.8
    AugMix (Mixing + JSD Loss) 11.2 35.9

    Key ablation findings:

    • Randomness and Diversity: Sampling random augmentation chains (Random AutoAugment*) improves corruption error over fixed tuned policies (AutoAugment*) by increasing diversity.
    • Consistency Loss and Mixing Synergy: AugmentAndMix alone (13.1%) and Random AutoAugment* + JSD (14.7%) each improve over Random AutoAugment* (17.0%), but combining both in AugMix achieves the lowest error (11.2%).
    • Excessive Mixing / Manifold Intrusion: Applying AugMix directly on top of Mixup degrades corruption error from 11.2% to 13.3%, attributed to manifold intrusion when interpolating between different class labels.
    • Hyperparameter Robustness: AugMix performance on ImageNet-C is insensitive to hyperparameter variations across mixing coefficient α∈{0.1,0.5,1,5}\alpha \in \{0.1, 0.5, 1, 5\}, mixture width k∈{1,3,5}k \in \{1, 3, 5\}, chain depth ∈{1,{1,2},{1,2,3}}\in \{1, \{1,2\}, \{1,2,3\}\}, and JSD example count ∈{2,3}\in \{2, 3\}.
  8. Knowl 8 — Uncertainty Estimation and Calibration under Data Shift

    empirical result

    AugMix significantly improves model calibration and predictive uncertainty under dataset shift on both CIFAR-10-C and ImageNet-C:

    • CIFAR-10 and CIFAR-10-C Calibration: Measured by RMS Calibration Error (%), standard models average 5.7% error on clean CIFAR-10 and degrade to 23.0% under CIFAR-10-C corruptions. AugMix reduces clean RMS calibration error to 3.6% and limits corrupted calibration error to 8.5% across four architectures (AllConvNet, DenseNet, WideResNet, ResNeXt). For ResNeXt, CIFAR-10-C calibration error drops from 16.4% to 8.3%.
    • ImageNet-C Calibration: Across corruption severities 1 through 5, standard ResNet-50 models exhibit severe calibration degradation (measured by RMS Calibration Error and Brier Score). In contrast, AugMix models (and particularly AugMix combined with deep ensembles) maintain flat and stable RMS calibration curves even as classification error increases under severe corruption.
  9. Knowl 9 — Empirical RMS Calibration Error Metric with Adaptive Binning

    equation

    To evaluate classifier calibration on a finite test set of nn examples, predictions are sorted by their confidence scores ck=max⁡cp^(Y=c∣xk)∈[0,1]c_k = \max_c \hat{p}(Y = c \mid x_k) \in [0, 1] and adaptively partitioned into bb contiguous bins {B1,B2,…,Bb}\{B_1, B_2, \dots, B_b\}, where each bin contains an equal number of predictions (e.g., ∣Bi∣=100|B_i| = 100).

    The empirical Root Mean Square (RMS) Calibration Error is calculated as:

    RMS Calibration Error=∑i=1b∣Bi∣n(1∣Bi∣∑k∈Bi1(yk=y^k)−1∣Bi∣∑k∈Bick)2\text{RMS Calibration Error} = \sqrt{\sum_{i=1}^b \frac{|B_i|}{n} \left( \frac{1}{|B_i|} \sum_{k \in B_i} \mathbf{1}(y_k = \hat{y}_k) - \frac{1}{|B_i|} \sum_{k \in B_i} c_k \right)^2}

    where yky_k is the true class label, y^k\hat{y}_k is the predicted class label, 1(⋅)\mathbf{1}(\cdot) is the indicator function (yielding empirical accuracy within bin BiB_i), and ckc_k is the predicted confidence of the kk-th sample.

    Adding the refinement term EC[P(Y=Y^∣C=c)(1−P(Y=Y^∣C=c))]\mathbb{E}_C [\mathbb{P}(Y = \hat{Y} \mid C = c)(1 - \mathbb{P}(Y = \hat{Y} \mid C = c))] to the squared RMS Calibration Error yields the Brier Score.

  10. Knowl 10 — Frequency-Domain Fourier Sensitivity of AugMix

    empirical result

    A 2D Fourier sensitivity analysis on CIFAR-10 measures model classification error when individual 32×3232 \times 32 Fourier basis noise vectors are added to clean test images across low-, mid-, and high-frequency spectra.

    • Baseline Wide ResNet models and models trained with Cutout remain robust only to low-frequency basis perturbations (centered in the Fourier domain) and suffer extreme vulnerability to additive mid- and high-frequency perturbations, where test error rates exceed 80%.
    • Models trained with AugMix preserve low-frequency robustness while dramatically suppressing error across all mid- and high-frequency Fourier basis vectors, explaining why AugMix generalizes broadly to diverse high-frequency noise and corruption types without having memorized specific corruption patterns during training.

Coverage note — No substantial contributed material was omitted; the knowls cover the full AugMix algorithm and methodology, JSD consistency loss equations, CIFAR-10/100-C corruption benchmark tables, ImageNet-C and ImageNet-P benchmark tables, perturbation stability findings, calibration/uncertainty analysis, the adaptive binning calibration metric, hyperparameter and component ablations, and Fourier frequency sensitivity analysis.

References

  1. 1.Aharon Azulay and Yair Weiss. Why do deep convolutional networks generalize so poorly to small image transformations? arXiv preprint, 2018.
  2. 2.Philip Bachman, Ouais Alsharif, and Doina Precup. Learning with pseudo-ensembles. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger (eds.), Advances in Neural Information Processing Systems 27, pp. 3365–3373. Curran Associates, Inc., 2014. URL http://papers.nips.cc/paper/5487-learning-with-pseudo-ensembles.pdf.
  3. 3.Sanghyuk Chun, Seong Joon Oh, Sangdoo Yun, Dongyoon Han, Junsuk Choe, and Youngjoon Yoo. An empirical evaluation on robustness and uncertainty of regularization methods. ICML Workshop on Uncertainty and Robustness in Deep Learning, 2019.
  4. 4.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. arXiv preprint arXiv:1909.13719, 2019.
  5. 5.Ekin Dogus Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le. AutoAugment: Learning augmentation policies from data. CVPR, 2018.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. CVPR, 2009.
  7. 7.Terrance Devries and Graham W. Taylor. Improved regularization of convolutional neural networks with Cutout. arXiv preprint arXiv:1708.04552, 2017.
  8. 8.Logan Engstrom, Andrew Ilyas, and Anish Athalye. Evaluating and understanding the robustness of adversarial logit pairing. arXiv preprint, 2018.
  9. 9.Robert Geirhos, Carlos R. M. Temme, Jonas Rauber, Heiko H. Schütt, Matthias Bethge, and Felix A. Wichmann. Generalisation in humans and deep neural networks. NeurIPS, 2018.
  10. 10.Justin Gilmer and Dan Hendrycks. A discussion of’adversarial examples are not bugs, they are features’: Adversarial example researchers need to expand what is meant by’robustness’. Distill, 4 (8):e00019–1, 2019.
  11. 11.Justin Gilmer, Ryan P. Adams, Ian J. Goodfellow, David Andersen, and George E. Dahl. Motivating the rules of the game for adversarial example research. CoRR, abs/1807.06732, 2018.
  12. 12.Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch SGD: Training ImageNet in 1 hour. CoRR, abs/1706.02677, 2017.
  13. 13.Keren Gu, Brandon Yang, Jiquan Ngiam, Quoc Le, and Jonathon Shlens. Using videos to evaluate image model robustness, 2019.
  14. 14.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. ICML, 2017.
  15. 15.Hongyu Guo, Yongyi Mao, and Richong Zhang. Mixup as locally linear out-of-manifold regularization. In AAAI, 2019.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. CVPR, 2015.
  17. 17.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. ICLR, 2019.
  18. 18.Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. ICLR, 2017.
  19. 19.Dan Hendrycks, Kimin Lee, and Mantas Mazeika. Using pre-training can improve model robustness and uncertainty. In ICML, 2019a.
  20. 20.Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich. Deep anomaly detection with outlier exposure. ICLR, 2019b.
  21. 21.Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In CVPR, 2017.
  22. 22.Daniel Kang, Yi Sun, Dan Hendrycks, Tom Brown, and Jacob Steinhardt. Testing robustness against unforeseen adversaries. arXiv preprint, 2019.
  23. 23.Harini Kannan, Alexey Kurakin, and Ian Goodfellow. Adversarial logit pairing. NeurIPS, 2018.
  24. 24.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. 2009.
  25. 25.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet classification with deep convolutional neural networks. NeurIPS, 2012.
  26. 26.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In NeurIPS, 2017.
  27. 27.Zachary Chase Lipton, Yu-Xiang Wang, and Alexander J. Smola. Detecting and correcting for label shift with black box predictors. ArXiv, abs/1802.03916, 2018.
  28. 28.Raphael Gontijo Lopes, Dong Yin, Ben Poole, Justin Gilmer, and Ekin Dogus Cubuk. Improving robustness without sacrificing accuracy with patch Gaussian augmentation. arXiv preprint arXiv:1906.02611, 2019.
  29. 29.Ilya Loshchilov and Frank Hutter. SGDR: stochastic gradient descent with warm restarts. ICLR, 2016.
  30. 30.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. ICLR, 2018.
  31. 31.Khanh Nguyen and Brendan O’Connor. Posterior calibration and exploratory analysis for natural language processing models. EMNLP, 2015.
  32. 32.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. NeurIPS, 2019.
  33. 33.Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi, and Percy Liang. Adversarial training can hurt generalization. arXiv preprint arXiv:1906.06032, 2019.
  34. 34.Tim Salimans and Diederik Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. NeurIPS, 2016.
  35. 35.Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin A. Riedmiller. Striving for simplicity: The all convolutional net. CoRR, abs/1412.6806, 2014.
  36. 36.Ryo Takahashi, Takashi Matsubara, and Kuniaki Uehara. Data augmentation using random image cropping and patching for deep cnns. ArXiv, abs/1811.09030, 2019.
  37. 37.Yuji Tokozume, Yoshitaka Ushiku, and Tatsuya Harada. Between-class learning for image classification. CVPR, 2018.
  38. 38.Antonio Torralba and Alexei A. Efros. Unbiased look at dataset bias. CVPR, 2011.
  39. 39.Igor Vasiljevic, Ayan Chakrabarti, and Gregory Shakhnarovich. Examining the impact of blur on recognition by convolutional networks, 2016.
  40. 40.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. CVPR, 2016.
  41. 41.Dong Yin, Raphael Gontijo Lopes, Jonathon Shlens, Ekin D Cubuk, and Justin Gilmer. A Fourier perspective on model robustness in computer vision. arXiv preprint arXiv:1906.08988, 2019.
  42. 42.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. ICCV, 2019.
  43. 43.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In BMVC, 2016.
  44. 44.Hongyi Zhang, Moustapha Cissé, Yann Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. ICLR, 2017.
  45. 45.Richard Zhang. Making convolutional networks shift-invariant again. In ICML, 2019.
  46. 46.Stephan Zheng, Yang Song, Thomas Leung, and Ian Goodfellow. Improving the robustness of deep neural networks via stability training. CVPR, 2016.
  47. 47.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. arXiv preprint arXiv:1708.04896, 2017.

Citation

MLA
Hendrycks, D., et al. “AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty”. arXiv, 2019, https://doi.org/10.48550/arxiv.1912.02781.
APA
Hendrycks, D., Mu, N., Cubuk, E. D., Zoph, B., Gilmer, J., & Lakshminarayanan, B. (2019). AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty. arXiv. https://doi.org/10.48550/arxiv.1912.02781
Chicago
Hendrycks, D., N. Mu, E. D. Cubuk, B. Zoph, J. Gilmer, and B. Lakshminarayanan. 2019. “AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1912.02781.
Harvard
Hendrycks, D. et al. (2019) “AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty”. arXiv. Available at: https://doi.org/10.48550/arxiv.1912.02781.
Vancouver
1. Hendrycks D, Mu N, Cubuk ED, Zoph B, Gilmer J, Lakshminarayanan B (2019) AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty. https://doi.org/10.48550/arxiv.1912.02781

BibTeX

@misc{https://doi.org/10.48550/arxiv.1912.02781,
  doi = {10.48550/ARXIV.1912.02781},
  url = {https://arxiv.org/abs/1912.02781},
  author = {Hendrycks, Dan and Mu, Norman and Cubuk, Ekin D. and Zoph, Barret and Gilmer, Justin and Lakshminarayanan, Balaji},
  keywords = {Machine Learning (stat.ML), Computer Vision and Pattern Recognition (cs.CV), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty},
  publisher = {arXiv},
  year = {2019},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission