Large-Scale Evolution of Image Classifiers

Esteban RealSherry MooreAndrew SelleSaurabh SaxenaYutaka Leon SuematsuJie TanQuoc LeAlex Kurakin

article2017ICML1,779 citations

Demonstrates that simple evolutionary algorithms scaled across massive compute can automatically design and fully train competitive neural network architectures from scratch without human intervention.

Listen

Designing high-performing neural network architectures has traditionally required years of intensive human trial and error. The article evaluates whether automated neuro-evolution, scaled up on massive computing infrastructure and operating without human intervention, can discover competitive, fully trained image classification models starting completely from scratch.

The authors implemented a distributed, lock-free evolutionary algorithm using tournament selection across 250 parallel workers on populations of 1,000 models. The search began from trivial single-layer models with zero convolutions and navigated an unbounded search space using intuitive structural and parameter mutations, standard back-propagation, and weight inheritance across generations. The process was benchmarked on the standard CIFAR-10 and CIFAR-100 image recognition datasets without any post-processing or manual architecture tuning.

The investigation produced four primary findings. First, neuro-evolution successfully constructed competitive deep convolutional networks entirely from scratch, achieving a top test accuracy of 94.6% on CIFAR-10 (improving to 95.6% when ensembling top models) and 77.0% on CIFAR-100. Second, the evolutionary outcomes proved repeatable across independent runs, yielding a consistent mean accuracy of 94.1% on CIFAR-10 with a standard deviation of only 0.4%. Third, the process was fully autonomous, producing completely trained models directly through weight inheritance. In control experiments, disabling weight inheritance reduced test accuracy to 92.2%, while pure random search achieved only 87.3% accuracy.

These findings challenge the widespread assumption that evolutionary methods cannot match modern hand-designed deep learning models. The framework demonstrates that automated architecture search can eliminate manual engineering effort and minimize researcher bias. However, this automation comes at a substantial computational expense, requiring on the order of 10^20 floating-point operations per experiment.

Organizations evaluating automated neural architecture search should weigh human labor savings against substantial compute infrastructure costs. Decision-makers should prioritize future work on efficiency optimizations, hybrid hand-guided evolutionary setups, and dynamic meta-parameter scheduling before deploying such methods in compute-constrained operational environments.

Confidence in these findings is high regarding standard benchmark classification tasks due to rigorous controls, pre-planned statistical analyses, and repeated runs. Nevertheless, uncertainties remain regarding computational latency and how well these evolutionary designs generalize to larger datasets and non-vision domains without further architectural adjustments.

arXiv: 1703.01041
Cover for Large-Scale Evolution of Image Classifiers

Abstract

Neural networks have proven effective at solving difficult problems but designing their architectures can be challenging, even for image classification problems alone. Our goal is to minimize human participation, so we employ evolutionary algorithms to discover such networks automatically. Despite significant computational requirements, we show that it is now possible to evolve models with accuracies within the range of those published in the last year. Specifically, we employ simple evolutionary techniques at unprecedented scales to discover models for the CIFAR-10 and CIFAR-100 datasets, starting from trivial initial conditions and reaching accuracies of 94.6% (95.6% for ensemble) and 77.0%, respectively. To do this, we use novel and intuitive mutation operators that navigate large search spaces; we stress that no human participation is required once evolution starts and that the output is a fully-trained model. Throughout this work, we place special emphasis on the repeatability of results, the variability in the outcomes and the computational requirements.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Evolutionary Algorithm
  • 3.2 Encoding and Mutations
  • 3.3 Initial Conditions
  • 3.4 Training and Validation
  • 3.5 Computation cost
  • 3.6 Weight Inheritance
  • 3.7 Reporting Methodology
  • 4 Experiments and Results
  • 5 Analysis
  • 6 Conclusion
  • References
  • S1 Methods Details
  • S2 FLOPs estimation
  • S3 Escaping Local Optima Details
  • S3.1 Local optima and mutation rate
  • S3.2 Local optima and weight resetting

Knowls

  1. Knowl 1 — Asynchronous Lock-Free Tournament Selection for Neural Architecture Evolution

    algorithm

    The evolutionary search framework discovers high-performing convolutional neural network architectures by evolving a population of models using asynchronous tournament selection across distributed workers. Independent workers interact lock-free through a shared file system where each individual model is represented by a directory with its state (alive, dead, training) tracked in the file name.

    Input: Population size NN (default N=1000N=1000), worker count W=N/4W = N/4 (default W=250W=250), training step budget per reproduction T=25600T = 25600, batch size B=50B = 50, mutation operator set M\mathcal{M}
    Output: Fully trained neural network individual maximizing validation accuracy
    Initialize population P\mathcal{P} on shared storage with NN identical trivial models (single linear layer, zero convolutions, learning rate η=0.1\eta = 0.1)
    for each worker w∈{1,…,W}w \in \{1, \dots, W\} in parallel and asynchronously do
        while evolutionary budget not exhausted do
            Sample two distinct alive individuals I1,I2∼Uniform(P)I_1, I_2 \sim \text{Uniform}(\mathcal{P})
            Iparent←arg⁡max⁡I∈{I1,I2}Fitness(I)I_{\text{parent}} \leftarrow \arg\max_{I \in \{I_1, I_2\}} \text{Fitness}(I)
            Idead←arg⁡min⁡I∈{I1,I2}Fitness(I)I_{\text{dead}} \leftarrow \arg\min_{I \in \{I_1, I_2\}} \text{Fitness}(I)
            Remove IdeadI_{\text{dead}} from P\mathcal{P} by marking its state as dead
            Ichild←Clone(Iparent)I_{\text{child}} \leftarrow \text{Clone}(I_{\text{parent}})
            Sample mutation operator m∼Uniform(M)m \sim \text{Uniform}(\mathcal{M})
            Apply mm to the DNA graph of IchildI_{\text{child}}
            Inherit weights from IparentI_{\text{parent}} to IchildI_{\text{child}} for all layers with identical variable names and tensor shapes
            Train IchildI_{\text{child}} for TT steps using SGD (momentum 0.90.9, weight decay 0.00010.0001, batch size BB)
            Compute Fitness(Ichild)\text{Fitness}(I_{\text{child}}) as top-1 classification accuracy on the 5,000-example validation set
            Mark IchildI_{\text{child}} as alive in P\mathcal{P}
        end while
    end for
    return arg⁡max⁡I∈PFitness(I)\arg\max_{I \in \mathcal{P}} \text{Fitness}(I)

    Pairwise tournament selection allows workers to proceed asynchronously without synchronization barriers or idling.

  2. Knowl 2 — DNA Graph Architecture Representation and Multi-Input Alignment

    model/method

    Neural network architectures are encoded as directed acyclic DNA graphs where vertices represent activation tensors and directed edges represent computational operations (convolutions or identity skip connections).

    Each vertex represents a rank-3 tensor of shape 2s×2s×d2^s \times 2^s \times d, where ss is the spatial scale factor (spatial dimensions H=W=2sH = W = 2^s) and dd is the channel depth. Vertices apply one of two activation types: (i) batch-normalization followed by rectified linear units (ReLUs), or (ii) linear activation.

    When multiple edges are incident on a vertex, spatial resolution and channel depth conflicts are resolved deterministically:

    1. The incoming non-skip connection edge is designated as the primary edge, defining the target spatial size 2s2^s and target channel depth dd.
    2. Activations from non-primary incoming edges are reshaped to match the primary edge: spatial dimensions are rescaled using zeroth-order (nearest-neighbor) interpolation, and channel depths are matched via truncation (if input channels exceed dd) or zero-padding (if input channels are fewer than dd).

    The DNA also stores the individual's learning rate η∈R+\eta \in \mathbb{R}^+.

  3. Knowl 3 — Mutation Operators for Architecture and Hyperparameter Modification

    model/method

    At each reproduction event, a worker applies a single mutation chosen uniformly at random from a set of 11 operators. Numerical mutations perturb parameters around their existing values using uniform sampling on log or linear scales:

    • ALTER-LEARNING-RATE: Multiplies the learning rate η\eta by a factor 2u2^u, where u∼U(−1.0,1.0)u \sim \mathcal{U}(-1.0, 1.0), scaling η\eta between half and double its current value.
    • IDENTITY: Leaves the architecture and weights unmodified, acting as an instruction to continue training.
    • RESET-WEIGHTS: Re-initializes all model weights using variance scaling initialization while retaining the graph structure.
    • INSERT-CONVOLUTION: Inserts a convolutional layer at a random point in the main backbone with 3×33 \times 3 kernel, stride ∈{1,2}\in \{1, 2\}, matching output channels, and randomly assigned batch-norm + ReLU or linear activation.
    • REMOVE-CONVOLUTION: Removes a randomly selected convolution from the backbone.
    • ALTER-STRIDE: Changes the convolution stride scale ss to an alternative power of 2.
    • ALTER-NUMBER-OF-CHANNELS: Multiplies the filter channel depth of a randomly chosen convolution by 2u2^u with u∼U(−1.0,1.0)u \sim \mathcal{U}(-1.0, 1.0).
    • FILTER-SIZE: Modifies the horizontal or vertical filter size of a random convolution to a new odd integer value.
    • INSERT-ONE-TO-ONE: Inserts an identity layer with linear activation along the backbone.
    • ADD-SKIP: Adds an identity edge between two randomly chosen vertices that does not create a cycle.
    • REMOVE-SKIP: Removes a randomly chosen existing skip connection.
  4. Knowl 4 — Weight Inheritance Across Architectural Mutations

    model/method

    To eliminate the need to train every evolved individual from scratch for hundreds of thousands of steps, child models inherit trained weights from their parent.

    When a child architecture is constructed from a parent:

    1. Each trainable variable in the child is assigned a structural identifier embedded in its variable name (e.g., convolution filter weights, batch-normalization scale and shift parameters).
    2. If a variable name in the child matches a variable name in the parent and has the exact same tensor shape, the child's variable is initialized with the parent's trained parameter values.
    3. If a variable's shape changed (such as in a modified filter dimension or altered channel depth) or if the layer was newly inserted, only that specific layer's weights are initialized randomly via variance scaling.

    Weight inheritance allows evolutionary optimization to act as a one-shot process over learning rate schedules and architectures without requiring post-evolution retraining or hyperparameter search.

  5. Knowl 5 — Theoretical FLOP Cost Estimation for Evolutionary Search

    equation

    The computational investment of an evolutionary run is estimated by summing the theoretical floating-point operations (FLOPs) across all individuals generated during the experiment. For an individual model ii, the total FLOPs CiC_i is computed as:

    Ci=Ft,iNt+Fv,iNvC_i = F_{t, i} N_t + F_{v, i} N_v

    where:

    • Ft,iF_{t, i} is the analytical FLOP count required to execute one training step (forward pass, loss computation, and backpropagation gradients) on a single batch of size B=50B = 50.
    • NtN_t is the number of training steps allocated per reproduction (Nt=25600N_t = 25600).
    • Fv,iF_{v, i} is the analytical FLOP count required to perform a forward inference pass on one validation batch.
    • NvN_v is the number of validation batches evaluated for fitness (Nv=100N_v = 100 batches for 5,000 validation examples).

    The total computational cost of the evolution experiment across all individuals Pall\mathcal{P}_{\text{all}} constructed throughout the search is:

    Ctotal=∑i∈PallCiC_{\text{total}} = \sum_{i \in \mathcal{P}_{\text{all}}} C_i

    This metric excludes input/output latency, graph building, memory copies, and data preprocessing overhead.

  6. Knowl 6 — Classification Performance on CIFAR-10 and CIFAR-100 Benchmarks

    empirical result

    Starting from trivial initial conditions (single linear layer with zero convolutions and η=0.1\eta = 0.1), large-scale evolution achieves competitive test accuracy on CIFAR-10 and CIFAR-100 without manual architectural tuning:

    • CIFAR-10 Single Model: The top evolved model (selected strictly by validation accuracy) achieved a test accuracy of 94.6%94.6\% (5.4 million parameters) requiring 4×10204 \times 10^{20} total FLOPs. Across 5 independent evolutionary runs, the mean test accuracy was μ=94.1%\mu = 94.1\% with standard deviation σ=0.4%\sigma = 0.4\% at an average cost of 9×10199 \times 10^{19} FLOPs per run.
    • CIFAR-10 Ensemble: Ensembling the top 2 validation-performing models from each of the 5 populations via majority voting achieved a test accuracy of 95.6%95.6\% with zero additional training cost.
    • CIFAR-100 Generalization: Applying the exact same algorithm and hyperparameters with zero modifications to CIFAR-100 yielded a model with 77.0%77.0\% test accuracy (40.4 million parameters) at 2×10202 \times 10^{20} FLOPs.
    • Control Baselines: A random search control (reproducing and killing random individuals without fitness selection) reached only 87.3%87.3\% accuracy in equivalent wall-clock time (2×10172 \times 10^{17} FLOPs). An evolution control with weight inheritance disabled reached only 92.2%92.2\% accuracy (9×10199 \times 10^{19} FLOPs), demonstrating the necessity of weight inheritance.
  7. Knowl 7 — Comparison of Evolved Networks with Hand-Designed and Discovered Architectures

    data/table

    The table compares the classification accuracy of single models generated by the one-shot evolutionary algorithm against hand-designed architectures and other automated neural architecture discovery methods on data-augmented CIFAR-10 (C10+) and CIFAR-100 (C100+).

    Study Parameters C10+ C100+ Search Space Status
    Maxout (Goodfellow et al., 2013) – 90.7% 61.4% Unreachable
    Network in Network (Lin et al., 2013) – 91.2% – Unreachable
    All-CNN (Springenberg et al., 2014) 1.3 M 92.8% 66.3% Reachable
    Deeply Supervised (Lee et al., 2015) – 92.0% 65.4% Unreachable
    Highway Network (Srivastava et al., 2015) 2.3 M 92.3% 67.6% Unreachable
    ResNet (He et al., 2016) 1.7 M 93.4% 72.8% Reachable
    Wide ResNet 28-10 (Zagoruyko Komodakis, 2016) 36.5 M 96.0% 80.0% Reachable
    Wide ResNet 40-10+D/O (Zagoruyko Komodakis, 2016) 50.7 M 96.2% 81.7% Unreachable
    DenseNet (Huang et al., 2016a) 25.6 M 96.7% 82.8% Unreachable
    Bayesian Opt. (Snoek et al., 2012) – 90.5% – Auto-Discovery
    Q-Learning (Baker et al., 2016) 11.2 M 93.1% 72.9% Auto-Discovery
    RL Discovery (Zoph Le, 2016, 20 layers) 2.5 M 94.0% – Auto-Discovery
    RL Discovery (Zoph Le, 2016, 39 layers) 37.0 M 96.4% – Auto-Discovery
    Evolution (Ours, Best Single Model) 5.4 M / 40.4 M 94.6% 77.0% Auto-Discovery
    Evolution (Ours, Top-2 Ensemble) – 95.6% – Auto-Discovery

    The evolved models exceed standard hand-designed baselines (All-CNN, Highway Networks, ResNet-110) and prior automated discovery methods starting from trivial conditions without manual post-processing, retraining, or architecture priming.

  8. Knowl 8 — Impact of Population Size and Training Steps on Local Optima Entrapment

    empirical result

    The performance of the evolutionary process is governed by two key metaparameters: population size NN and the number of training steps per reproduction TT.

    1. Population Size (NN): Larger population sizes systematically improve final model fitness. For small populations, such as N=2N = 2, populations rapidly get trapped in poor local optima. This occurs because an individual that happens to be "super-fit" (where any single architectural mutation causes an immediate drop in validation accuracy) will defeat its mutated child in tournament selection every round, permanently taking over the population. Increasing NN (evaluated across N∈{2,10,43,1000}N \in \{2, 10, 43, 1000\}) provides diversity and allows paths requiring multiple mutations to be explored.
    2. Training Steps Per Individual (TT): Increasing training steps per reproduction (T∈{256,2560,25600}T \in \{256, 2560, 25600\}) monotonically increases final test accuracy across runs of fixed wall-clock duration. Higher TT reduces the number of IDENTITY mutations required for a model to reach convergence.
  9. Knowl 9 — Techniques for Escaping Local Fitness Optima

    model/method

    When a population becomes trapped at an inferior local optimum during evolution, two population-level interventions can dislodge the population:

    1. Mutation Rate Escalation: Increasing the mutation intensity from 1 mutation per reproduction event to 5 consecutive mutations per reproduction increases exploration, allowing trapped populations to escape suboptimal plateaus (successful in 8 out of 10 experimental trials) before returning to single-mutation reproduction.
    2. Population-Wide Weight Resetting: Over-accumulation of IDENTITY mutations can cause a poorly structured architecture to dominate simply because it received more training iterations. Simultaneously resetting all weights across every individual in the population places all candidate topologies on equal footing. While causing an immediate temporary drop in fitness, across 10 independent experiments, executing three consecutive weight reset events produced statistically significant improvements in final validation fitness (p<0.001p < 0.001).
  10. Knowl 10 — Evaluated Recombination Mechanisms in Neuro-Evolution

    empirical result

    Three distinct recombination (crossover) strategies were tested to combine information between two parent models:

    1. Mutation Probability Recombination: Evolving individual-specific mutation probability distributions and producing a child that inherits graph structure from one parent and mutation parameters from the other.
    2. Weight Recombination: Blending or averaging trained weight tensors from two parent architectures having identical topologies.
    3. Structural Side-by-Side Fusion: Merging two parent models by placing their directed graphs in parallel and concatenating their final feature activations.

    Empirical evaluation showed that none of these three recombination approaches yielded an improvement in convergence rate or final classification accuracy compared to mutation-only evolution.

Coverage note — Specific Python/TensorFlow implementation boilerplate and low-level disk I/O routines were omitted in favor of the standalone mathematical, procedural, and empirical formulations of the evolutionary framework.

References

  1. 1.Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S, Davis, Andy, Dean, Jeffrey, Devin, Matthieu, et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467, 2016.
  2. 2.Baker, Bowen, Gupta, Otkrist, Naik, Nikhil, and Raskar, Ramesh. Designing neural network architectures using reinforcement learning. arXiv preprint arXiv:1611.02167, 2016.
  3. 3.Bayer, Justin, Wierstra, Daan, Togelius, Julian, and Schmidhuber, Jürgen. Evolving memory cell structures for sequence learning. In International Conference on Artificial Neural Networks, pp. 755–764. Springer, 2009.
  4. 4.Bergstra, James and Bengio, Yoshua. Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13(Feb):281–305, 2012.
  5. 5.Breuel, Thomas and Shafait, Faisal. Automlp: Simple, effective, fully automated learning rate and size adjustment. In The Learning Workshop. Utah, 2010.
  6. 6.Fernando, Chrisantha, Banarse, Dylan, Reynolds, Malcolm, Besse, Frederic, Pfau, David, Jaderberg, Max, Lanctot, Marc, and Wierstra, Daan. Convolution by evolution: Differentiable pattern producing networks. In Proceedings of the 2016 on Genetic and Evolutionary Computation Conference, pp. 109–116. ACM, 2016.
  7. 7.Goldberg, David E and Deb, Kalyanmoy. A comparative analysis of selection schemes used in genetic algorithms. Foundations of genetic algorithms, 1:69–93, 1991.
  8. 8.Goldberg, David E, Richardson, Jon, et al. Genetic algorithms with sharing for multimodal function optimization. In Genetic algorithms and their applications: Proceedings of the Second International Conference on Genetic Algorithms, pp. 41–49. Hillsdale, NJ: Lawrence Erlbaum, 1987.
  9. 9.Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron C, and Bengio, Yoshua. Maxout networks. International Conference on Machine Learning, 28:1319–1327, 2013.
  10. 10.Gruau, Frederic. Genetic synthesis of modular neural networks. In Proceedings of the 5th International Conference on Genetic Algorithms, pp. 318–325. Morgan Kaufmann Publishers Inc., 1993.
  11. 11.Han, Song, Pool, Jeff, Tran, John, and Dally, William. Learning both weights and connections for efficient neural network. In Advances in Neural Information Processing Systems, pp. 1135–1143, 2015.
  12. 12.He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pp. 1026–1034, 2015.
  13. 13.He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016.
  14. 14.Huang, Gao, Liu, Zhuang, Weinberger, Kilian Q, and van der Maaten, Laurens. Densely connected convolutional networks. arXiv preprint arXiv:1608.06993, 2016a.
  15. 15.Huang, Gao, Sun, Yu, Liu, Zhuang, Sedra, Daniel, and Weinberger, Kilian Q. Deep networks with stochastic depth. In European Conference on Computer Vision, pp. 646–661. Springer, 2016b.
  16. 16.Ioffe, Sergey and Szegedy, Christian. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  17. 17.Kim, Minyoung and Rigazio, Luca. Deep clustered convolutional kernels. arXiv preprint arXiv:1503.01824, 2015.
  18. 18.Krizhevsky, Alex and Hinton, Geoffrey. Learning multiple layers of features from tiny images. 2009.
  19. 19.Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, pp. 1097–1105, 2012.
  20. 20.LeCun, Yann, Cortes, Corinna, and Burges, Christopher JC. The mnist database of handwritten digits, 1998.
  21. 21.Lee, Chen-Yu, Xie, Saining, Gallagher, Patrick W, Zhang, Zhengyou, and Tu, Zhuowen. Deeply-supervised nets. In AISTATS, volume 2, pp. 5, 2015.
  22. 22.Lin, Min, Chen, Qiang, and Yan, Shuicheng. Network in network. arXiv preprint arXiv:1312.4400, 2013.
  23. 23.Miller, Geoffrey F, Todd, Peter M, and Hegde, Shailesh U. Designing neural networks using genetic algorithms. In Proceedings of the third international conference on Genetic algorithms, pp. 379–384. Morgan Kaufmann Publishers Inc., 1989.
  24. 24.Morse, Gregory and Stanley, Kenneth O. Simple evolutionary optimization can rival stochastic gradient descent in neural networks. In Proceedings of the 2016 on Genetic and Evolutionary Computation Conference, pp. 477–484. ACM, 2016.
  25. 25.Pugh, Justin K and Stanley, Kenneth O. Evolving multimodal controllers with hyperneat. In Proceedings of the 15th annual conference on Genetic and evolutionary computation, pp. 735–742. ACM, 2013.
  26. 26.Rumelhart, David E, Hinton, Geoffrey E, and Williams, Ronald J. Learning representations by back-propagating errors. Cognitive Modeling, 5(3):1, 1988.
  27. 27.Saxena, Shreyas and Verbeek, Jakob. Convolutional neural fabrics. In Advances In Neural Information Processing Systems, pp. 4053–4061, 2016.
  28. 28.Silver, David, Huang, Aja, Maddison, Chris J, Guez, Arthur, Sifre, Laurent, Van Den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  29. 29.Simmons, Joseph P, Nelson, Leif D, and Simonsohn, Uri. False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11):1359–1366, 2011.
  30. 30.Simonyan, Karen and Zisserman, Andrew. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  31. 31.Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P. Practical bayesian optimization of machine learning algorithms. In Advances in neural information processing systems, pp. 2951–2959, 2012.
  32. 32.Springenberg, Jost Tobias, Dosovitskiy, Alexey, Brox, Thomas, and Riedmiller, Martin. Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806, 2014.
  33. 33.Srivastava, Rupesh Kumar, Greff, Klaus, and Schmidhuber, Jürgen. Highway networks. arXiv preprint arXiv:1505.00387, 2015.
  34. 34.Stanley, Kenneth O. Compositional pattern producing networks: A novel abstraction of development. Genetic programming and evolvable machines, 8(2):131–162, 2007.
  35. 35.Stanley, Kenneth O and Miikkulainen, Risto. Evolving neural networks through augmenting topologies. Evolutionary Computation, 10(2):99–127, 2002.
  36. 36.Stanley, Kenneth O, D’Ambrosio, David B, and Gauci, Jason. A hypercube-based encoding for evolving large-scale neural networks. Artificial Life, 15(2):185–212, 2009.
  37. 37.Sutskever, Ilya, Martens, James, Dahl, George E, and Hinton, Geoffrey E. On the importance of initialization and momentum in deep learning. ICML (3), 28:1139–1147, 2013.
  38. 38.Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott, Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–9, 2015.
  39. 39.Tuson, Andrew and Ross, Peter. Adapting operator settings in genetic algorithms. Evolutionary computation, 6(2): 161–184, 1998.
  40. 40.Verbancsics, Phillip and Harguess, Josh. Generative neuroevolution for deep learning. arXiv preprint arXiv:1312.5355, 2013.
  41. 41.Weinreich, Daniel M and Chao, Lin. Rapid evolutionary escape by large populations from local fitness peaks is likely in nature. Evolution, 59(6):1175–1182, 2005.
  42. 42.Weyand, Tobias, Kostrikov, Ilya, and Philbin, James. Planet-photo geolocation with convolutional neural networks. In European Conference on Computer Vision, pp. 37–55. Springer, 2016.
  43. 43.Wu, Yonghui, Schuster, Mike, Chen, Zhifeng, Le, Quoc V., Norouzi, Mohammad, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
  44. 44.Zagoruyko, Sergey and Komodakis, Nikos. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  45. 45.Zaremba, Wojciech. An empirical exploration of recurrent network architectures. 2015.
  46. 46.Zoph, Barret and Le, Quoc V. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016.

Citation

MLA
Real, E., et al. “Large-Scale Evolution of Image Classifiers”. arXiv, 2017, http://arxiv.org/abs/1703.01041v2.
APA
Real, E., Moore, S., Selle, A., Saxena, S., Suematsu, Y. L., Tan, J., Le, Q., & Kurakin, A. (2017). Large-Scale Evolution of Image Classifiers. arXiv. http://arxiv.org/abs/1703.01041v2
Chicago
Real, E., S. Moore, A. Selle, et al. 2017. “Large-Scale Evolution of Image Classifiers”. arXiv. http://arxiv.org/abs/1703.01041v2.
Harvard
Real, E. et al. (2017) “Large-Scale Evolution of Image Classifiers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1703.01041v2.
Vancouver
1. Real E, Moore S, Selle A, Saxena S, Suematsu YL, Tan J, Le Q, Kurakin A (2017) Large-Scale Evolution of Image Classifiers. arXiv

BibTeX

@article{real2017large,
  title = {Large-Scale Evolution of Image Classifiers},
  author = {Real, Esteban and Moore, Sherry and Selle, Andrew and Saxena, Saurabh and Suematsu, Yutaka Leon and Tan, Jie and Le, Quoc and Kurakin, Alex},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1703.01041v2},
  eprint = {1703.01041}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/