Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. AutoAugment: Searching for best Augmentation policies Directly on the Dataset of Interest
  • 4. Experiments and Results
  • 4.1. CIFAR­10, CIFAR­100, SVHN Results
  • 4.2. ImageNet Results
  • 4.3. The Transferability of Learned Augmentation policies to Other Datasets
  • 5. Discussion
  • 5.1. Ablation experiments
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — AutoAugment Search Space for Image Data Augmentation

    definition

    The AutoAugment search space specifies a data augmentation policy SS composed of 55 sub-policies. Each sub-policy consists of two image processing operations applied in sequence: SubPolicy(x)=(f2f1)(x)\text{SubPolicy}(x) = (f_2 \circ f_1)(x).

    Each operation ff is defined by three components:

    1. Operation Type: Selected from a discrete pool of 16 image transformations: ShearX\text{ShearX}, ShearY\text{ShearY}, TranslateX\text{TranslateX}, TranslateY\text{TranslateY}, Rotate\text{Rotate}, AutoContrast\text{AutoContrast}, Invert\text{Invert}, Equalize\text{Equalize}, Solarize\text{Solarize}, Posterize\text{Posterize}, Contrast\text{Contrast}, Color\text{Color}, Brightness\text{Brightness}, Sharpness\text{Sharpness}, Cutout\text{Cutout}, and SamplePairing\text{SamplePairing}.
    2. Application Probability (pp): Discretized uniformly into 11 values: p{0.0,0.1,0.2,,1.0}p \in \{0.0, 0.1, 0.2, \dots, 1.0\}. An implicit identity transformation is achieved when p=0.0p = 0.0.
    3. Magnitude (MM): Discretized uniformly into 10 integer levels: M{0,1,,9}M \in \{0, 1, \dots, 9\}, mapped linearly over a predefined default range specific to each transformation (such as maximum rotation angle or shear percentage). Invert and AutoContrast do not use magnitude values.

    A single sub-policy has (16×10×11)2=1,76023.1×106(16 \times 10 \times 11)^2 = 1,760^2 \approx 3.1 \times 10^6 possible configurations. The total search space for a policy containing 5 concurrent sub-policies comprises (16×10×11)102.9×1032(16 \times 10 \times 11)^{10} \approx 2.9 \times 10^{32} possible policies.

  2. Knowl 2 — AutoAugment Policy Search via Reinforcement Learning

    algorithm

    The AutoAugment search process uses a recurrent neural network (RNN) controller to explore the discrete space of augmentation policies. The controller is a single-layer Long Short-Term Memory (LSTM) network with 100 hidden units. At each step, it produces 30 softmax predictions corresponding to 5 sub-policies ×\times 2 operations ×\times 3 discrete attributes (operation type, magnitude level, and probability level), feeding each prediction embedding into the next step.

    Input: Child model architecture MM, search training set DtrainD_{\text{train}}, search validation set DvalD_{\text{val}}, controller LSTM CC with parameters θ\theta, total search iterations K15000K \approx 15000
    Output: Augmentation policy SS^* maximizing child model validation accuracy
    Initialize controller parameters θUniform(0.1,0.1)\theta \sim \text{Uniform}(-0.1, 0.1)
    Initialize exponential moving average baseline b0b \leftarrow 0
    for step =1= 1 to KK do
        Sample policy S={s1,s2,s3,s4,s5}S = \{s_1, s_2, s_3, s_4, s_5\} from controller C(θ)C(\theta)
        Train child network MM from scratch on DtrainD_{\text{train}} using policy SS
        Evaluate child network MM on DvalD_{\text{val}} to obtain validation accuracy reward RR
        Compute policy gradient θJ(θ)\nabla_\theta J(\theta) scaled by advantage (Rb)(R - b) using PPO
        Update controller parameters θPPO-Step(θ,θJ(θ))\theta \leftarrow \text{PPO-Step}(\theta, \nabla_\theta J(\theta))
        Update moving average baseline b0.95b+0.05Rb \leftarrow 0.95 \cdot b + 0.05 \cdot R
    end for
    return policy SS^* yielding the highest reward RR

    The controller is trained using Proximal Policy Optimization (PPO) with learning rate 0.000350.00035, an entropy penalty weight of 0.000010.00001 to promote exploration, and an exponential moving average reward baseline with decay 0.950.95.

  3. Knowl 3 — Proxy Child Model Protocol and Final Policy Composition

    experimental setup

    To reduce computational cost during reinforcement learning policy search, proxy child models are trained on reduced datasets:

    • Reduced CIFAR-10: 4,000 randomly selected examples from the 50,000 training set. Child models use a small Wide-ResNet-40-2 architecture trained for 120 epochs from scratch with learning rate 0.010.01, cosine learning rate decay (one annealing cycle), and weight decay 10410^{-4}.
    • Reduced SVHN: 1,000 randomly chosen examples from the core training set, using the identical Wide-ResNet-40-2 training setup.
    • Reduced ImageNet: 6,000 samples drawn from 120 randomly chosen classes, training a Wide-ResNet-40-2 for 200 epochs with learning rate 0.10.1 and weight decay 10510^{-5}.

    Final Policy Formulation: At the end of the search, the sub-policies from the top 5 highest-reward policies are concatenated to form a single final policy containing 5×5=255 \times 5 = 25 sub-policies.

    Application During Training: When training full target models, for every mini-batch, one of the 25 sub-policies is chosen uniformly at random for each input image. The image is transformed sequentially by the sub-policy's two operations according to their individual probabilities and magnitudes. AutoAugment is applied after baseline pre-processing (such as standard cropping and horizontal flipping) and before Cutout.

  4. Knowl 4 — Classification Error Rates on CIFAR-10, CIFAR-100, and SVHN

    data/table

    AutoAugment significantly reduces test set error rates across multiple architectures on CIFAR-10, CIFAR-100, and SVHN. On reduced subsets (4,000 examples for CIFAR-10, 1,000 examples for SVHN), AutoAugment achieves performance comparable to semi-supervised methods without utilizing any unlabeled data.

    Dataset Model Baseline Cutout AutoAugment
    CIFAR-10 Wide-ResNet-28-10 3.9 3.1 2.6 ±\pm 0.1
    Shake-Shake (26 2x32d) 3.6 3.0 2.5 ±\pm 0.1
    Shake-Shake (26 2x96d) 2.9 2.6 2.0 ±\pm 0.1
    Shake-Shake (26 2x112d) 2.8 2.6 1.9 ±\pm 0.1
    AmoebaNet-B (6,128) 3.0 2.1 1.8 ±\pm 0.1
    PyramidNet+ShakeDrop 2.7 2.3 1.5 ±\pm 0.1
    Reduced CIFAR-10 Wide-ResNet-28-10 18.8 16.5 14.1 ±\pm 0.3
    Shake-Shake (26 2x96d) 17.1 13.4 10.0 ±\pm 0.2
    CIFAR-100 Wide-ResNet-28-10 18.8 18.4 17.1 ±\pm 0.3
    Shake-Shake (26 2x96d) 17.1 16.0 14.3 ±\pm 0.2
    PyramidNet+ShakeDrop 14.0 12.2 10.7 ±\pm 0.2
    SVHN Wide-ResNet-28-10 1.5 1.3 1.1
    Shake-Shake (26 2x96d) 1.4 1.2 1.0
    Reduced SVHN Wide-ResNet-28-10 13.2 32.5 8.2
    Shake-Shake (26 2x96d) 12.3 24.2 5.9

    All entries report test error rate percentages (lower is better), averaged over 5 runs. Baseline pre-processing includes standardizing, horizontal flips (50%), zero-padding, and random cropping (with 16×1616 \times 16 Cutout where indicated). On reduced SVHN, standalone Cutout degrades performance (12.3%24.2%12.3\% \to 24.2\%), whereas AutoAugment lowers error to 5.9%5.9\%.

    Additionally, on the Recht et al. CIFAR-10 test set, PyramidNet+ShakeDrop trained with AutoAugment achieves a 4.4%4.4\% error rate (2.9%2.9\% absolute drop relative to original CIFAR-10), demonstrating superior generalization compared to models without AutoAugment (which show drops of 4.1%4.6%4.1\%\text{--}4.6\%).

  5. Knowl 5 — ImageNet Classification Performance with AutoAugment

    data/table

    Applying the 25 sub-policies found on a 120-class subset of ImageNet to full ImageNet training yields state-of-the-art Top-1 and Top-5 accuracy improvements across standard and neural-architecture-searched models without extra data or ensembling.

    Model Inception Pre-processing AutoAugment (Ours)
    ResNet-50 76.3 / 93.1 77.6 / 93.8
    ResNet-200 78.5 / 94.2 80.0 / 95.0
    AmoebaNet-B (6,190) 82.2 / 96.0 82.8 / 96.2
    AmoebaNet-C (6,228) 83.1 / 96.1 83.5 / 96.5

    Validation set accuracy is reported as Top-1 / Top-5 percentages (higher is better). Models are trained from scratch for 270 epochs with batch size 4096, initial learning rate 1.6, and 10-fold decays at epochs 90, 180, and 240. Baseline Inception pre-processing scales pixel values to [1,1][-1, 1], applies horizontal flipping (p=0.5p=0.5), and includes random color distortions.

  6. Knowl 6 — Transferability of ImageNet Augmentation Policies to Fine-Grained Visual Datasets

    data/table

    Augmentation policies learned on ImageNet transfer directly to Fine-Grained Visual Categorization (FGVC) datasets. Inception-v4 models were trained from scratch for 1,000 epochs with cosine learning rate decay at 448×448448 \times 448 resolution.

    Dataset Train Size Classes Baseline Error (%) AutoAugment-transfer Error (%)
    Oxford 102 Flowers 2,040 102 6.7 4.6
    Caltech-101 3,060 102 19.4 13.1
    Oxford-IIIT Pets 3,680 37 13.5 11.0
    FGVC Aircraft 6,667 100 9.1 7.3
    Stanford Cars 8,144 196 6.4 5.2

    All metrics report test Top-1 error rates (lower is better). On Stanford Cars, training from scratch with transferred AutoAugment policies achieves a 5.2%5.2\% error rate, surpassing the prior best published result (5.9%5.9\% error) which required fine-tuning ImageNet pre-trained weights with deep layer aggregation.

  7. Knowl 7 — Domain-Specific Operation Preferences in Learned Policies

    empirical result

    The reinforcement learning search process discovers distinct, dataset-tailored transformation choices that align with domain-specific invariances:

    1. CIFAR-10: Successful policies select almost exclusively color-based operations, most prominently Equalize\text{Equalize}, AutoContrast\text{AutoContrast}, Color\text{Color}, and Brightness\text{Brightness}. Geometric transformations like ShearX\text{ShearX} and ShearY\text{ShearY} are rarely selected, and Invert\text{Invert} is virtually never chosen.
    2. SVHN: Successful policies heavily select Invert\text{Invert}, Equalize\text{Equalize}, Rotate\text{Rotate}, ShearX\text{ShearX}, and ShearY\text{ShearY}. Invert\text{Invert} is frequently selected because digit identity is invariant to inverted foreground/background contrast. Geometric shearing and rotation match natural distortions in street-view house numbers.
    3. ImageNet: Policies favor color-based transformations analogous to CIFAR-10, but additionally incorporate geometric Rotate\text{Rotate} operations.
  8. Knowl 8 — Comparison with Adversarial / GAN-Based Augmentation Learning

    data/table

    AutoAugment achieves larger accuracy gains on CIFAR-10 compared to adversarial data augmentation methods that train a generator to propose transformation sequences that fool a discriminator (Ratner et al., 2017).

    Method Baseline Error (%) Augmented Error (%) Improvement Δ\Delta (%)
    LSTM (Ratner et al., ResNet-56) 7.7 6.0 1.6
    MF (Ratner et al., ResNet-56) 7.7 5.6 2.1
    AutoAugment (ResNet-32) 7.7 4.5 3.2
    AutoAugment (ResNet-56) 6.6 3.6 3.0

    AutoAugment directly optimizes downstream classification accuracy via validation reward, whereas generative/adversarial methods optimize sequence realism relative to training samples, leading to smaller generalization improvements.

  9. Knowl 9 — Ablation of Learned Policy Parameters and Comparison to Random Policies

    empirical result

    Ablation experiments on CIFAR-10 using Wide-ResNet-28-10 isolate the contributions of the search space vs. the learned parameters:

    1. Learned AutoAugment Policy: Achieves 2.6%±0.1%2.6\% \pm 0.1\% test error rate.
    2. Randomized Probabilities and Magnitudes: Retaining the learned operation types but assigning random probabilities and magnitudes yields an average test error of 3.0%±0.1%3.0\% \pm 0.1\% (averaged over 20 runs), degrading performance by 0.4%0.4\%.
    3. Fully Random Policies: Randomly sampling the operation types, probabilities, and magnitudes from the search space yields an average test error of 3.1%±0.1%3.1\% \pm 0.1\% over 20 runs, with the single best random policy reaching 3.0%3.0\% (averaged over 5 runs).

    While random sampling within the designed discrete search space outperforms baseline augmentation (3.9%3.9\% baseline, 3.1%3.1\% with Cutout), RL optimization of operation types, probabilities, and magnitudes provides an additional 0.4%0.5%0.4\%\text{--}0.5\% absolute error reduction.

  10. Knowl 10 — Effect of Sub-Policy Count and Stochastic Application on Generalization

    empirical result

    The number of sub-policies used during training directly impacts generalization error on CIFAR-10 (evaluated on Wide-ResNet-28-10):

    • Validation error decreases monotonically as the number of sub-policies increases from 1 to approximately 20 sub-policies, beyond which performance gains plateau.
    • Stochastic application—where each mini-batch sample is augmented by a randomly selected sub-policy whose operations are executed with independent probabilities—requires a minimum training epoch threshold. Child models require at least 80 to 100 epochs (120 epochs used in practice) so that each of the 5 sub-policies is applied sufficiently often to provide a reliable reward signal. Fully trained models trained for hundreds to thousands of epochs exploit larger concatenations (25 sub-policies).

Coverage note — No substantial contributed material was omitted. All search space definitions, RL search mechanics, proxy child setups, CIFAR/SVHN/ImageNet/FGVC benchmark results, comparisons to prior automated augmentation methods, policy property analyses, and ablation studies are covered.

References

  1. 1.M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng. Tensorflow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation, OSDI’16, pages 265–283, Berkeley, CA, USA, 2016. USENIX Association. 5
  2. 2.A. Antoniou, A. Storkey, and H. Edwards. Data augmentation generative adversarial networks. arXiv preprint arXiv:1711.04340, 2017. 2
  3. 3.H. S. Baird. Document image defect models. In Structured Document Image Analysis, pages 546–556. Springer, 1992. 1
  4. 4.B. Baker, O. Gupta, N. Naik, and R. Raskar. Designing neural network architectures using reinforcement learning. In International Conference on Learning Representations, 2017. 2, 3
  5. 5.I. Bello, B. Zoph, V. Vasudevan, and Q. V. Le. Neural optimizer search with reinforcement learning. In International Conference on Machine Learning, 2017. 3
  6. 6.J. Bergstra and Y. Bengio. Random search for hyperparameter optimization. Journal of Machine Learning Research, 13(Feb):281–305, 2012. 4
  7. 7.A. Brock, T. Lim, J. M. Ritchie, and N. Weston. Smash: oneshot model architecture search through hypernetworks. In International Conference on Learning Representations, 2017. 2
  8. 8.D. Ciregan, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classification. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 3642–3649. IEEE, 2012. 2
  9. 9.E. D. Cubuk, B. Zoph, S. S. Schoenholz, and Q. V. Le. Intriguing properties of adversarial examples. arXiv preprint arXiv:1711.02846, 2017. 2
  10. 10.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009. 4
  11. 11.T. DeVries and G. W. Taylor. Dataset augmentation in feature space. arXiv preprint arXiv:1702.05538, 2017. 2
  12. 12.T. DeVries and G. W. Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017. 2, 3, 4, 5, 6, 13
  13. 13.T. Elsken, J.-H. Metzen, and F. Hutter. Simple and efficient architecture search for convolutional neural networks. arXiv preprint arXiv:1711.04528, 2017. 2
  14. 14.Y. Em, F. Gag, Y. Lou, S. Wang, T. Huang, and L.-Y. Duan. Incorporating intra-class variance to fine-grained visual recognition. In Multimedia and Expo (ICME), 2017 IEEE International Conference on, pages 1452–1457. IEEE, 2017. 7
  15. 15.L. Fei-Fei, R. Fergus, and P. Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. Computer vision and Image understanding, 106(1):59–70, 2007. 7
  16. 16.K. Fukushima and S. Miyake. Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In Competition and cooperation in neural nets, pages 267–285. Springer, 1982. 1
  17. 17.X. Gastaldi. Shake-shake regularization. arXiv preprint arXiv:1705.07485, 2017. 4, 5, 6
  18. 18.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in neural information processing systems, pages 2672–2680, 2014. 7
  19. 19.D. Han, J. Kim, and J. Kim. Deep pyramidal residual networks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6307–6315. IEEE, 2017. 1
  20. 20.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 1, 6
  21. 21.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997. 4
  22. 22.A. G. Howard. Some improvements on deep convolutional neural network based image classification. arXiv preprint arXiv:1312.5402, 2013. 6
  23. 23.J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. arXiv preprint arXiv:1709.01507, 2017. 1
  24. 24.H. Inoue. Data augmentation by pairing samples for images classification. arXiv preprint arXiv:1801.02929, 2018. 3, 13
  25. 25.K. Jarrett, K. Kavukcuoglu, Y. LeCun, et al. What is the best multi-stage architecture for object recognition? In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 2146–2153. IEEE, 2009. 1
  26. 26.S. Kornblith, J. Shlens, and Q. V. Le. Do better imagenet models transfer better? arXiv preprint arXiv:1805.08974, 2018. 2
  27. 27.J. Krause, J. Deng, M. Stark, and L. Fei-Fei. Collecting a large-scale dataset of fine-grained cars. In Second Workshop on Fine-Grained Visual Categorization, 2013. 2, 7
  28. 28.A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009. 4
  29. 29.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, 2012. 1, 2
  30. 30.M. Kumar, G. E. Dahl, V. Vasudevan, and M. Norouzi. Parallel architecture and hyperparameter search via successive halving and classification. arXiv preprint arXiv:1805.10255, 2018. 4
  31. 31.S. Laine and T. Aila. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016. 5
  32. 32.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 1
  33. 33.J. Lemley, S. Bazrafkan, and P. Corcoran. Smart augmentation learning an optimal data augmentation strategy. IEEE Access, 5:5858–5869, 2017. 2
  34. 34.C. Liu, B. Zoph, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy. Progressive neural architecture search. arXiv preprint arXiv:1712.00559, 2017. 2
  35. 35.H. Liu, K. Simonyan, O. Vinyals, C. Fernando, and K. Kavukcuoglu. Hierarchical representations for efficient architecture search. In International Conference on Learning Representations, 2018. 2
  36. 36.I. Loshchilov and F. Hutter. SGDR: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 4
  37. 37.D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten. Exploring the limits of weakly supervised pretraining. arXiv preprint arXiv:1805.00932, 2018. 6
  38. 38.S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013. 2, 7
  39. 39.H. Mania, A. Guy, and B. Recht. Simple random search provides a competitive approach to reinforcement learning. arXiv preprint arXiv:1803.07055, 2018. 2
  40. 40.T. Miyato, S.-i. Maeda, M. Koyama, and S. Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. In International Conference on Learning Representations, 2016. 5, 6
  41. 41.S. Mun, S. Park, D. K. Han, and H. Ko. Generative adversarial network based acoustic scene training set augmentation and selection using svm hyper-plane. In Detection and Classification of Acoustic Scenes and Events Workshop, 2017. 2
  42. 42.Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011. 4, 5
  43. 43.M.-E. Nilsback and A. Zisserman. Automated flower classification over a large number of classes. In Computer Vision, Graphics & Image Processing, 2008. ICVGIP’08. Sixth Indian Conference on, pages 722–729. IEEE, 2008. 7
  44. 44.A. Oliver, A. Odena, C. Raffel, E. D. Cubuk, and I. J. Goodfellow. Realistic evaluation of deep semi-supervised learning algorithms. arXiv preprint arXiv:1804.09170, 2018. 5
  45. 45.L. Perez and J. Wang. The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621, 2017. 2
  46. 46.H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, and J. Dean. Efficient neural architecture search via parameter sharing. In International Conference on Machine Learning, 2018. 2
  47. 47.A. J. Ratner, H. Ehrenberg, Z. Hussain, J. Dunnmon, and C. Ré. Learning to compose domain-specific transformations for data augmentation. In Advances in Neural Information Processing Systems, pages 3239–3249, 2017. 2, 7
  48. 48.E. Real, A. Aggarwal, Y. Huang, and Q. V. Le. Regularized evolution for image classifier architecture search. arXiv preprint arXiv:1802.01548, 2018. 1, 2, 4, 5, 6
  49. 49.E. Real, S. Moore, A. Selle, S. Saxena, Y. L. Suematsu, J. Tan, Q. Le, and A. Kurakin. Large-scale evolution of image classifiers. In International Conference on Machine Learning, 2017. 2
  50. 50.B. Recht, R. Roelofs, L. Schmidt, and V. Shankar. Do cifar-10 classifiers generalize to cifar-10? arXiv preprint arXiv:1806.00451, 2018. 5
  51. 51.M. Sajjadi, M. Javanmardi, and T. Tasdizen. Regularization with stochastic transformations and perturbations for deep semi-supervised learning. In Advances in Neural Information Processing Systems, pages 1163–1171, 2016. 5
  52. 52.I. Sato, H. Nishimura, and K. Yokoi. Apac: Augmented pattern classification with neural networks. arXiv preprint arXiv:1505.03229, 2015. 2
  53. 53.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. 3, 4
  54. 54.P. Y. Simard, D. Steinkraus, J. C. Platt, et al. Best practices for convolutional neural networks applied to visual document analysis. In Proceedings of International Conference on Document Analysis and Recognition, 2003. 1, 2
  55. 55.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. Advances in Neural Information Processing Systems, 2015. 1
  56. 56.L. Sixt, B. Wild, and T. Landgraf. Rendergan: Generating realistic labeled data. arXiv preprint arXiv:1611.01331, 2016. 2
  57. 57.I. Sutskever, J. Schulman, T. Salimans, and D. Kingma. Requests For Research 2.0. https://blog.openai.com/requests-for-research-2, 2018. 1
  58. 58.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, 2017. 1, 7
  59. 59.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, A. Rabinovich, et al. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 1, 6
  60. 60.A. Tarvainen and H. Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in neural information processing systems, pages 1195–1204, 2017. 5
  61. 61.T. Tran, T. Pham, G. Carneiro, L. Palmer, and I. Reid. A bayesian data augmentation approach for learning deep models. In Advances in Neural Information Processing Systems, pages 2794–2803, 2017. 2
  62. 62.L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus. Regularization of neural networks using dropconnect. In International Conference on Machine Learning, pages 1058–1066, 2013. 2
  63. 63.L. Xie and A. Yuille. Genetic CNN. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2017. 2
  64. 64.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5987–5995, 2017. 1
  65. 65.Y. Yamada, M. Iwamura, and K. Kise. Shakedrop regularization. arXiv preprint arXiv:1802.02375, 2018. 4, 5, 6
  66. 66.F. Yu, D. Wang, and T. Darrell. Deep layer aggregation. arXiv preprint arXiv:1707.06484, 2017. 1, 7
  67. 67.S. Zagoruyko and N. Komodakis. Wide residual networks. In British Machine Vision Conference, 2016. 4, 5, 6, 8
  68. 68.H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017. 5, 6, 13
  69. 69.Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang. Random erasing data augmentation. arXiv preprint arXiv:1708.04896, 2017. 13
  70. 70.X. Zhu, Y. Liu, Z. Qin, and J. Li. Data augmentation in emotion classification using generative adversarial networks. arXiv preprint arXiv:1711.00648, 2017. 2
  71. 71.B. Zoph and Q. V. Le. Neural architecture search with reinforcement learning. In International Conference on Learning Representations, 2017. 2, 3
  72. 72.B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le. Learning transferable architectures for scalable image recognition. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2017. 1, 2, 3, 4

Citation

MLA
Cubuk, E. D., et al. “AutoAugment: Learning Augmentation Strategies From Data”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 113–23, https://doi.org/10.1109/cvpr.2019.00020.
APA
Cubuk, E. D., Zoph, B., Mané, D., Vasudevan, V., & Le, Q. V. (2019). AutoAugment: Learning Augmentation Strategies From Data. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 113–123. https://doi.org/10.1109/cvpr.2019.00020
Chicago
Cubuk, E. D., B. Zoph, D. Mané, V. Vasudevan, and Q. V. Le. 2019. “AutoAugment: Learning Augmentation Strategies From Data”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 113–23. https://doi.org/10.1109/cvpr.2019.00020.
Harvard
Cubuk, E.D. et al. (2019) “AutoAugment: Learning Augmentation Strategies From Data”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 113–123. Available at: https://doi.org/10.1109/cvpr.2019.00020.
Vancouver
1. Cubuk ED, Zoph B, Mané D, Vasudevan V, Le QV (2019) AutoAugment: Learning Augmentation Strategies From Data. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 113–123

BibTeX

@inproceedings{Cubuk_2019, title={AutoAugment: Learning Augmentation Strategies From Data}, url={http://dx.doi.org/10.1109/cvpr.2019.00020}, DOI={10.1109/cvpr.2019.00020}, booktitle={2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Cubuk, Ekin D. and Zoph, Barret and Mané, Dandelion and Vasudevan, Vijay and Le, Quoc V.}, year={2019}, month=June, pages={113–123} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission