Designing Neural Network Architectures using Reinforcement Learning

Bowen BakerOtkrist GuptaNikhil NaikRamesh Raskar

article2016ICLR1,557 citations

Proposes MetaQNN, a reinforcement learning approach that uses Q-learning to automatically design high-performing convolutional neural network architectures without requiring manual human tuning.

Listen

Designing high-performing convolutional neural networks currently demands extensive human labor and specialized technical intuition. As neural network applications expand across diverse industries, manual architecture design creates a severe development bottleneck. The article evaluates an automated meta-modeling framework called MetaQNN, which employs reinforcement learning to design competitive vision architectures from scratch without human intervention.

The authors model neural network creation as a sequential decision process where an autonomous software agent selects network layers one by one. The agent explores a finite design space of standard network building blocks—such as convolution, pooling, and fully connected layers—using a standard reinforcement learning technique known as Q-learning paired with an exploration-to-exploitation schedule and memory replay. The algorithm was evaluated across benchmark image classification datasets, including CIFAR-10, SVHN, and MNIST, running on a compute setup of 10 graphics processing units over 8 to 10 days per dataset.

The findings show that MetaQNN successfully learns to discover increasingly accurate network architectures. As the agent shifts from random exploration to informed selection, average model accuracy improves significantly, rising from 52.25% to 88.02% during exploration on the SVHN dataset. Networks designed by the agent achieved test error rates of 6.92% on CIFAR-10, 2.06% on SVHN, and 0.28% on MNIST using an ensemble approach, outperforming all existing human-designed architectures built with the same basic layer types. MetaQNN also substantially surpassed prior automated design methods, reducing error on CIFAR-10 from 21.2% to 6.92%, while remaining competitive against complex human-engineered networks that rely on specialized layers. Furthermore, top models transferred successfully to new classification tasks when trained from scratch or fine-tuned.

These results demonstrate that reinforcement learning can effectively eliminate manual trial-and-error in network design while producing performant, custom models. This capability significantly lowers engineering overhead and mitigates the risk of deploying suboptimal architectures. Additionally, the approach provides diverse, high-performing candidate networks suitable for ensembling or deployment across different resource constraints.

Organizations developing computer vision systems should consider automating network architecture search rather than relying solely on manual design. Future development should integrate hyperparameter tuning into the search process and incorporate multi-objective reward functions to optimize for inference speed and model size alongside accuracy.

Confidence in these findings is supported by consistent convergence across ten independent test runs. However, key limitations include the computational expense—requiring 80 to 100 GPU-days per full experiment—and the use of constrained, discretized layer definitions to keep exploration tractable. Stakeholders should account for these computational requirements before implementing similar automated search pipelines at scale.

  • Paper: Playing Atari with Deep Reinforcement Learning, Volodymyr Mnih et al. (2013). This foundational work demonstrates training deep neural networks with Q-learning, experience replay, and epsilon-greedy exploration, providing the direct reinforcement learning mechanics used by MetaQNN.
  • Paper: Network In Network, Min Lin et al. (2014). It introduced global average pooling and micro-network concepts to standard CNN design, establishing core architectural baselines that MetaQNN incorporates into its layer selection search space.
  • Paper: Going Deeper with Convolutions, Christian Szegedy et al. (2015). It provides crucial context on manual multi-scale convolutional design and efficient layer topologies against which automated meta-modeling approaches compete.
  • Paper: Visualizing and Understanding Convolutional Networks, Matthew D. Zeiler et al. (2014). It establishes empirical insights into how convolutional layers build hierarchical visual representations, motivating automated exploration of CNN layer depth and configurations.
  • Paper: Recent advances in convolutional neural networks, Jiuxiang Gu et al. (2015). This survey outlines the standard building blocks of convolutional architectures (convolution, pooling, and fully-connected layers) that define MetaQNN's search space.
Cover for Designing Neural Network Architectures using Reinforcement Learning

Abstract

At present, designing convolutional neural network (CNN) architectures requires both human expertise and labor. New architectures are handcrafted by careful experimentation or modified from a handful of existing networks. We introduce MetaQNN, a meta-modeling algorithm based on reinforcement learning to automatically generate high-performing CNN architectures for a given learning task. The learning agent is trained to sequentially choose CNN layers using QQ-learning with an ϵ\epsilon-greedy exploration strategy and experience replay. The agent explores a large but finite space of possible architectures and iteratively discovers designs with improved performance on the learning task. On image classification benchmarks, the agent-designed networks (consisting of only standard convolution, pooling, and fully-connected layers) beat existing networks designed with the same layer types and are competitive against the state-of-the-art methods that use more complex layer types. We also outperform existing meta-modeling approaches for network design on image classification tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background
  • 4 Designing Neural Network Architectures with QQ-learning
  • 4.1 The State Space
  • 4.2 The Action Space
  • 4.3 QQ-learning Training Procedure
  • 5 Experiment Details
  • 6 Results
  • 7 Concluding Remarks
  • References
  • A Algorithm
  • B Representation Size Binning
  • C MNIST Experiment
  • D Further Analysis of QQ-Learning
  • D.1 QQ-Learning Stability
  • D.2 QQ-Value Analysis
  • E Top topologies selected by algorithm

Knowls

  1. Knowl 1 — MetaQNN Markov Decision Process Formulation for Sequential CNN Generation

    model/method

    The task of designing a convolutional neural network (CNN) architecture is formulated as a finite-horizon Markov Decision Process (MDP) over a discrete state space S\mathcal{S} and action space U\mathcal{U}. The agent sequentially selects layers until reaching a termination layer, producing a directed acyclic graph (DAG) representation of a feed-forward CNN.

    A state s∈Ss \in \mathcal{S} is defined as a parameter tuple specific to each supported layer type:

    • Convolution (C): Tuple (i,f,l,d,n)(i, f, l, d, n) where layer depth i<12i < 12 (or i<18i < 18 in expanded spaces), receptive field size f∈{1,3,5}f \in \{1, 3, 5\}, stride l=1l = 1, channel depth d∈{64,128,256,512}d \in \{64, 128, 256, 512\}, and spatial representation size bin n∈{(∞,8],(8,4],(4,1]}n \in \{(\infty, 8], (8, 4], (4, 1]\}.
    • Pooling (P): Tuple (i,f,l,n)(i, f, l, n) where layer depth i<12i < 12, receptive field and stride pair (f,l)∈{(5,3),(3,2),(2,2)}(f, l) \in \{(5, 3), (3, 2), (2, 2)\}, and representation size bin n∈{(∞,8],(8,4],(4,1]}n \in \{(\infty, 8], (8, 4], (4, 1]\}.
    • Fully Connected (FC): Tuple (i,nFC,d)(i, n_{\text{FC}}, d) where layer depth i<12i < 12, consecutive FC count nFC<3n_{\text{FC}} < 3, and neuron count d∈{512,256,128}d \in \{512, 256, 128\}.
    • Termination: Type t∈{Global Average Pooling (GAP),Softmax (SM)}t \in \{\text{Global Average Pooling (GAP)}, \text{Softmax (SM)}\}.

    Valid action sets U(s)\mathcal{U}(s) enforce architectural constraints to ensure tractability:

    1. Transitions are strictly restricted from layer depth ii to depth i+1i+1, preventing recurrent loops and enforcing a finite horizon.
    2. Any non-terminal state may choose to transition to a termination state; any state at the maximum layer depth must transition to a termination state.
    3. At most two consecutive FC layers are allowed (nFC<3n_{\text{FC}} < 3). Transitions between FC states require non-increasing neuron counts (d′≤dd' \le d).
    4. A pooling state cannot transition directly to another pooling state.
    5. Transitions to FC states are permitted only if the spatial representation size bin is (8,4](8, 4] or (4,1](4, 1], preventing parameter explosion.
  2. Knowl 2 — Representation Size Binning and Transition Stochasticity in MetaQNN

    model/method

    To prevent the architecture search agent from producing degenerate CNN architectures (e.g., repeatedly downsampling until intermediate feature representations collapse below 1×11 \times 1), the state tuple includes a representation size (RR-size) parameter that restricts subsequent filter sizes to f≤R-sizef \le R\text{-size}.

    To constrain the state space size, continuous spatial sizes are grouped into three discrete bins: Bin 1 [8,∞)[8, \infty), Bin 2 (8,4](8, 4], and Bin 3 (4,1](4, 1]. Because multiple exact spatial dimensions map to the same bin (e.g., input spatial dimensions 18×1818 \times 18 and 14×1414 \times 14 both belong to Bin 1), applying an identical pooling operation (e.g., 2×22 \times 2 pooling with stride 2) can yield different bin outcomes:

    • An 18×1818 \times 18 input is reduced to 9×99 \times 9, remaining in Bin 1 [8,∞)[8, \infty).
    • A 14×1414 \times 14 input is reduced to 7×77 \times 7, transitioning to Bin 2 (8,4](8, 4].

    Consequently, taking a deterministic pooling action from a given binned state yields an uncertain next binned state. This is formally accommodated by modeling state transitions stochastically via transition probabilities p(s′∣s,u)p(s' \mid s, u) within the standard MDP formulation.

  3. Knowl 3 — MetaQNN Q-Learning Algorithm with Experience Replay for Neural Architecture Search

    algorithm

    MetaQNN uses tabular Q-learning with an ϵ\epsilon-greedy exploration schedule and experience replay to discover optimal CNN architectures for a given machine learning task. An agent generates a sequence of layer choices, trains the resulting network, receives its validation accuracy as a terminal reward r∈[0,1]r \in [0, 1], and updates action values Q(s,u)Q(s, u) via temporally reversed Bellman updates with discount factor γ=1.0\gamma = 1.0 and learning rate α=0.01\alpha = 0.01.

    The ϵ\epsilon schedule trains a predefined quota of unique models at each ϵ∈{1.0,0.9,0.8,0.7,0.6,0.5,0.4,0.3,0.2,0.1}\epsilon \in \{1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1\}, using 1,500 models at ϵ=1.0\epsilon = 1.0, 100 models each at ϵ∈{0.9,0.8,0.7}\epsilon \in \{0.9, 0.8, 0.7\}, and 150 models each at ϵ∈{0.6,0.5,0.4,0.3,0.2,0.1}\epsilon \in \{0.6, 0.5, 0.4, 0.3, 0.2, 0.1\}. When an already-trained architecture is re-sampled, its stored validation accuracy is reused without retraining.

    Input: Number of models per stage M, replay batch size K = 100, learning rate alpha = 0.01, exploration schedule epsilon
    Initialize replay memory D as an empty list
    Initialize Q(s, u) = 0.5 for all states s in S and actions u in U(s)
    for each epsilon in schedule do
        for episode = 1 to M do
            S = [s_START]
            U = []
            while U[last] != terminate do
                Sample p ~ Uniform(0, 1)
                if p > epsilon then
                    u = argmax_{u_cand in U(S[last])} Q(S[last], u_cand)
                else
                    Sample u ~ Uniform(U(S[last]))
                end if
                s_next = TRANSITION(S[last], u)
                Append u to U
                if u != terminate then
                    Append s_next to S
                end if
            end while
            accuracy = TRAIN(S)
            Append (S, U, accuracy) to D
            for memory_step = 1 to K do
                Sample (S_sample, U_sample, r) uniformly at random from D
                L = length(S_sample)
                Q(S_sample[L-1], U_sample[L-1]) = (1 - alpha) * Q(S_sample[L-1], U_sample[L-1]) + alpha * r
                for i = L - 2 down to 0 do
                    Q(S_sample[i], U_sample[i]) = (1 - alpha) * Q(S_sample[i], U_sample[i]) + alpha * max_{u_cand in U(S_sample[i+1])} Q(S_sample[i+1], u_cand)
                end for
            end for
        end for
    end for
    return Q
  4. Knowl 4 — Two-Stage Training and Finetuning Protocol for MetaQNN Architecture Search

    experimental setup

    MetaQNN evaluates candidate architectures in two distinct operational phases:

    1. Exploration Phase (Fast Evaluation Scheme):

      • Data: A random validation split of 5,000 training examples is held out while preserving class proportions.
      • Optimization: Each architecture is trained for 20 epochs using the Adam optimizer (β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, ε=10−8\varepsilon = 10^{-8}), a batch size of 128, and Xavier weight initialization. The initial learning rate is 0.0010.001.
      • Early Restart / Step Decay: If validation accuracy after epoch 1 does not exceed random guessing, the learning rate is scaled by 0.40.4 and training is restarted (up to 5 restarts). For learning models, the learning rate decays by a factor of 0.20.2 every 5 epochs.
      • Regularization: A dropout layer is inserted every 2 functional layers. For the ii-th dropout layer among nn total dropout layers in the network, the dropout probability is set to pi=i2np_i = \frac{i}{2n}.
    2. Evaluation and Finetuning Phase:

      • The top 10 discovered models based on exploration performance are fully retrained on the entire training set with extensive schedules and data augmentation.
      • CIFAR-10: Max layer depth increased to 18 during search. Retrained for 300 epochs (lr 0.025 for 40 epochs, 0.0125 for 40 epochs, 0.0001 for 160 epochs, and 0.00001 for 60 epochs) with global contrast normalization, random horizontal mirroring, and random translation up to 5 pixels.
      • SVHN: Exploration uses original training set; finetuning uses original plus extended training set (lr 0.025 for 5 epochs, 0.0125 for 5 epochs, 0.0001 for 20 epochs, and 0.00001 for 10 epochs).
      • MNIST: Models trained for 40 epochs with global mean subtraction, decaying lr by 0.20.2 every 5 epochs.
  5. Knowl 5 — Image Classification Error Rates of MetaQNN Architectures

    data/table

    MetaQNN discovers architectures composed exclusively of standard convolution, pooling, and fully connected layers. When evaluated on benchmark image classification datasets, both the best individual MetaQNN model and an ensemble of the top 5 models outperform standard handcrafted architectures of identical layer types and achieve performance competitive with complex models containing residual connections, generalized pooling, or deeply-supervised branches.

    Method CIFAR-10 (%) SVHN (%) MNIST (%) CIFAR-100 (%)
    Standard Layer Types (Conv / Pool / FC)
    Maxout 9.38 2.47 0.45 38.57
    Network in Network (NIN) 8.81 2.35 0.47 35.68
    FitNet 8.39 2.42 0.51 35.04
    Highway Network 7.72 – – –
    VGGnet 7.25 – – –
    All-CNN 7.25 – – 33.71
    MetaQNN (top model) 6.92 2.28 0.44 27.14
    MetaQNN (ensemble) 7.32 2.06 0.32 –
    Complex Layer Types / Architecture Search Baselines
    DropConnect 9.32 1.94 0.57 –
    Deeply-Supervised Nets (DSN) 8.22 1.92 0.39 34.57
    Recurrent CNN (R-CNN) 7.72 1.77 0.31 31.75
    ResNet-110 6.61 – – –
    ResNet-1001 4.62 – – 22.71
    ELU 6.55 – – 24.28
    Tree+Max-Avg Pooling 6.05 1.69 0.31 32.37
    TPE Meta-modeling 21.20 – – –
    NEAT (Genetic Algorithm) – – 7.90 –
    MetaQNN (10-model ensemble) – – 0.28 –

    All reported numbers are classification error rates (percentage, lower is better). CIFAR-10 and CIFAR-100 results use moderate data augmentation (mirroring and translation); SVHN and MNIST results use no data augmentation. MetaQNN substantially outperforms prior automated architecture search approaches such as TPE (21.2%21.2\% vs 6.92%6.92\% on CIFAR-10) and NEAT (7.9%7.9\% vs 0.32%0.32\% on MNIST), while an ensemble of the top 10 MNIST MetaQNN models achieves 0.28%0.28\% error without data augmentation.

  6. Knowl 6 — Cross-Dataset Transferability of MetaQNN Discovered Architectures

    data/table

    To evaluate the general visual representational capacity of architectures discovered via reinforcement learning, the top MetaQNN architecture selected on CIFAR-10 was transferred directly to CIFAR-100, SVHN, and MNIST without changing its topological structure.

    Training Protocol CIFAR-100 (%) SVHN (%) MNIST (%)
    Training from scratch 27.14 2.48 0.80
    Finetuning from CIFAR-10 weights 34.93 4.00 0.81
    State-of-the-art benchmark 24.28 1.69 0.31

    All entries report classification test error rate (percentage, lower is better). The CIFAR-10 top architecture trains successfully from scratch on CIFAR-100 to achieve a 27.14%27.14\% error rate, outperforming standard handcrafted networks such as All-CNN (33.71%33.71\%) and NIN (35.68%35.68\%). Training the transferred architecture from scratch consistently yields lower error rates than finetuning from pretrained CIFAR-10 weights across all three target datasets.

  7. Knowl 7 — Layer Preference Trends and Architectural Motifs Discovered in MetaQNN Q-Values

    empirical result

    Analysis of the final converged Q-values across layer depths reveals systematic structural patterns learned by the reinforcement learning agent:

    1. Layer Type vs. Depth: In early layers (depths 1 to 4), the average Q-values for standard convolution and pooling layers are highest. As network depth increases, the average Q-values for fully connected layers, global average pooling, and termination (softmax) monotonically rise, matching handcrafted network design conventions.
    2. Receptive Field Size vs. Depth: Among convolution layers, Q-values for large receptive field sizes (5×55 \times 5) exceed those of smaller receptive field sizes (1×11 \times 1 and 3×33 \times 3) at deeper layer positions, indicating that larger receptive fields are beneficial deeper in the feature extraction hierarchy.
    3. 1×11 \times 1 Input Convolution Motif: The agent repeatedly selects a 1×11 \times 1 convolution layer (C(N,1,1)C(N, 1, 1)) as the very first layer in top-performing networks. This operation learns NN linear combinations across the input RGB color channels prior to spatial convolutions, functioning analogously to learned color space transformations (e.g., RGB to YUV).
  8. Knowl 8 — Stability of Q-Learning Architecture Search across Independent Runs

    empirical result

    To assess whether the randomized ϵ\epsilon-greedy exploration consistently converges to high-performing architectures, MetaQNN was evaluated over 10 independent runs on a 10%10\% subset of the SVHN dataset (approximately 7,000 training examples, requiring 10 GPU-days per run):

    • Across all 10 runs, the mean accuracy of sampled models improved steadily from approximately 50%50\% at ϵ=1.0\epsilon = 1.0 (random exploration) to approximately 76%76\% at ϵ=0.1\epsilon = 0.1 (exploitation).
    • The variance of model accuracy was lowest during the purely random exploration stage (ϵ=1.0\epsilon = 1.0) due to the large sample size (1,500 models).
    • Across the 10 independent experiments, the top model discovered in each run achieved a mean test accuracy of 88.25%88.25\% with a standard deviation of only 0.58%0.58\% (without using an extended training schedule), demonstrating high stability and reproducibility in finding high-quality network topologies.
  9. Knowl 9 — Architectural Scope and Fixed-Hyperparameter Constraints of Tabular MetaQNN

    limitation

    The MetaQNN meta-modeling framework has three specific methodological limitations:

    1. Discretization and State Space Restrictions: The state-action space is restricted to a discrete tabular representation consisting solely of sequential feed-forward operations (standard convolution, pooling, fully connected, global average pooling). Complex, non-sequential topologies such as multi-branch connections, residual skip connections, and recurrent modules are not representable within this MDP definition.
    2. Scalability of Tabular State Spaces: Coarse binning of representation sizes and layer depths is necessary to prevent combinatorial explosion of the tabular Q-learning state space. Moving to continuous layer hyperparameters or unconstrained layer depths requires continuous or neural network-based Q-function approximation (Deep Q-Networks).
    3. Decoupled Architecture and Hyperparameter Optimization: All network topologies are evaluated using a uniform, fixed hyperparameter schedule during the Q-learning search phase (fixed learning rate decay, optimizer parameters, and dropout rule). Optimization of training hyperparameters is performed only as a post-hoc step on the top selected architectures rather than being integrated jointly into the meta-modeling search.

Coverage note — None was omitted. All contributed methods (MDP formulation, binning, Q-learning with replay, training schedules), experimental results (CIFAR-10, SVHN, MNIST, CIFAR-100 benchmarks, transfer learning), structural analyses (Q-value trends, motifs, stability runs), and limitations have been extracted into self-contained knowls.

References

  1. 1.Sander Adam, Lucian Busoniu, and Robert Babuska. Experience replay for real-time reinforcement learning control. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 42(2):201–212, 2012.
  2. 2.James Bergstra, Daniel Yamins, and David D Cox. Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures. ICML (1), 28:115–123, 2013.
  3. 3.James S Bergstra, Remi Bardenet, Yoshua Bengio, and Bal azs K egl. Algorithms for hyper-parameter  optimization. NIPS, pp. 2546–2554, 2011.
  4. 4.Dimitri P Bertsekas. Convex optimization algorithms. Athena Scientific Belmont, 2015.
  5. 5.Djork-Arne Clevert, Thomas Unterthiner, and Sepp Hochreiter. Fast and accurate deep network  learning by exponential linear units (ELUs). arXiv preprint arXiv:1511.07289, 2015.
  6. 6.Tobias Domhan, Jost Tobias Springenberg, and Frank Hutter. Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves. IJCAI, 2015.
  7. 7.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. AISTATS, 9:249–256, 2010.
  8. 8.Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron C Courville, and Yoshua Bengio. Maxout networks. ICML (3), 28:1319–1327, 2013.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In European Conference on Computer Vision, pp. 630–645. Springer, 2016.
  11. 11.Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014.
  12. 12.Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4:237–285, 1996.
  13. 13.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  14. 14.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature, 521(7553):436–444, 2015.
  15. 15.Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu. Deeplysupervised nets. AISTATS, 2(3):6, 2015.
  16. 16.Chen-Yu Lee, Patrick W Gallagher, and Zhuowen Tu. Generalizing pooling functions in convolutional neural networks: Mixed, gated, and tree. International Conference on Artificial Intelligence and Statistics, 2016.
  17. 17.Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. JMLR, 17(39):1–40, 2016.
  18. 18.Ming Liang and Xiaolin Hu. Recurrent convolutional neural network for object recognition. CVPR, pp. 3367–3375, 2015.
  19. 19.Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  20. 20.Long-Ji Lin. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 8(3-4):293–321, 1992.
  21. 21.Long-Ji Lin. Reinforcement learning for robots using neural networks. Technical report, DTIC Document, 1993.
  22. 22.Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. arXiv preprint arXiv:1312.4400, 2013.
  23. 23.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  24. 24.Nicolas Pinto, David Doukhan, James J DiCarlo, and David D Cox. A high-throughput screening approach to discovering good forms of biologically inspired visual representation. PLoS Computational Biology, 5(11):e1000579, 2009.
  25. 25.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014.
  26. 26.Shreyas Saxena and Jakob Verbeek. Convolutional neural fabrics. In Advances in Neural Information Processing Systems 29, pp. 4053–4061. 2016.
  27. 27.J David Schaffer, Darrell Whitley, and Larry J Eshelman. Combinations of genetic algorithms and neural networks: A survey of the state of the art. International Workshop on Combinations of Genetic Algorithms and Neural Networks, pp. 1–37, 1992.
  28. 28.Pierre Sermanet, Soumith Chintala, and Yann LeCun. Convolutional neural networks applied to house numbers digit classification. ICPR, pp. 3288–3291, 2012.
  29. 29.Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala, and Yann LeCun. Pedestrian detection with unsupervised multi-stage feature learning. CVPR, pp. 3626–3633, 2013.
  30. 30.Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando de Freitas. Taking the human out of the loop: A review of bayesian optimization. Proceedings of the IEEE, 104(1): 148–175, 2016.
  31. 31.David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  32. 32.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  33. 33.Jasper Snoek, Hugo Larochelle, and Ryan P Adams. Practical bayesian optimization of machine learning algorithms. NIPS, pp. 2951–2959, 2012.
  34. 34.Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806, 2014.
  35. 35.Rupesh Kumar Srivastava, Klaus Greff, and Jurgen Schmidhuber. Highway networks.  arXiv preprint arXiv:1505.00387, 2015.
  36. 36.Kenneth O Stanley and Risto Miikkulainen. Evolving neural networks through augmenting topologies. Evolutionary Computation, 10(2):99–127, 2002.
  37. 37.Kevin Swersky, Jasper Snoek, and Ryan P Adams. Multi-task bayesian optimization. NIPS, pp. 2004–2012, 2013.
  38. 38.Phillip Verbancsics and Josh Harguess. Generative neuroevolution for deep learning. arXiv preprint arXiv:1312.5355, 2013.
  39. 39.Joannes Vermorel and Mehryar Mohri. Multi-armed bandit algorithms and empirical evaluation. European Conference on Machine Learning, pp. 437–448, 2005.
  40. 40.Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus. Regularization of neural networks using dropconnect. ICML, pp. 1058–1066, 2013.
  41. 41.Christopher John Cornish Hellaby Watkins. Learning from delayed rewards. PhD thesis, University of Cambridge, England, 1989.

Citation

MLA
Baker, B., et al. “Designing Neural Network Architectures Using Reinforcement Learning”. arXiv, 2016, http://arxiv.org/abs/1611.02167v3.
APA
Baker, B., Gupta, O., Naik, N., & Raskar, R. (2016). Designing Neural Network Architectures using Reinforcement Learning. arXiv. http://arxiv.org/abs/1611.02167v3
Chicago
Baker, B., O. Gupta, N. Naik, and R. Raskar. 2016. “Designing Neural Network Architectures Using Reinforcement Learning”. arXiv. http://arxiv.org/abs/1611.02167v3.
Harvard
Baker, B. et al. (2016) “Designing Neural Network Architectures using Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.02167v3.
Vancouver
1. Baker B, Gupta O, Naik N, Raskar R (2016) Designing Neural Network Architectures using Reinforcement Learning. arXiv

BibTeX

@article{baker2016designing,
  title = {Designing Neural Network Architectures using Reinforcement Learning},
  author = {Baker, Bowen and Gupta, Otkrist and Naik, Nikhil and Raskar, Ramesh},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.02167v3},
  eprint = {1611.02167}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors