ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware

Han CaiLigeng ZhuSong Han

article2018ICLR2,133 citations

Introduces ProxylessNAS, a differentiable neural architecture search framework that reduces GPU memory to standard training levels, enabling direct model optimization on large-scale datasets and specific hardware platforms without relying on proxy tasks.

Listen

Automating the design of deep learning models through neural architecture search (NAS) is critical for artificial intelligence deployment, but conventional methods demand prohibitive computational resources (tens of thousands of GPU hours). To manage these computational costs, existing methods rely on proxy tasks—such as searching on smaller datasets, using fewer layers, or repeatedly stacking identical building blocks. These proxy shortcuts produce sub-optimal architectures because they restrict architectural diversity and fail to reflect true target performance and hardware latency.

The article demonstrates ProxylessNAS, an approach that optimizes neural network architectures directly on large-scale target tasks and target hardware without using proxy tasks, while reducing memory and computation costs to the level of standard model training.

The researchers formulated architecture search as a path-level pruning process within an over-parameterized super-network containing all candidate operations. To resolve severe GPU memory bottlenecks, they introduced path binarization, which activates only one or two candidate paths during training steps rather than keeping all options in memory simultaneously. They also incorporated non-differentiable hardware execution speed (latency) into the gradient-based optimization using a continuous latency prediction model, as well as an alternative reinforcement learning formulation. The method was evaluated on standard image classification benchmarks (CIFAR-10 and ImageNet) across three hardware platforms: mobile phones (Google Pixel 1), cloud GPUs (Nvidia Tesla V100), and CPUs (Intel Xeon).

The evaluation produced several key findings. First, ProxylessNAS cut the computational search cost on ImageNet to 200 GPU hours—a 200-fold reduction compared to prior methods like MnasNet, which required roughly 40,000 GPU hours. Second, on ImageNet mobile benchmarks, the discovered architecture improved top-1 accuracy by 2.6% over MobileNetV2 at equivalent latency, and ran 1.83 times faster when matched for accuracy. Third, on CIFAR-10, it achieved a 2.08% test error with only 5.7 million parameters, matching or beating top baseline models while using six times fewer parameters. Finally, hardware-specific searches proved that optimal architectures vary substantially across device types: GPU-optimized models favored shallower, wider layers with larger operations due to high parallelism, whereas CPU models favored deeper, narrower structures with smaller operations.

These findings indicate that organizations can design custom, high-performing neural networks directly for specific hardware endpoints at a fraction of standard compute costs. This eliminates the expense of running large device-testing farms and prevents the efficiency losses that occur when deploying generic models across diverse hardware platforms. Teams building edge and cloud AI applications can reduce latency, lower cloud training expenditures, and improve user device responsiveness.

Organizations should adopt direct, hardware-aware architecture search and tailor deployed models to specific execution platforms (GPU, CPU, mobile) rather than deploying a single cross-platform model. Furthermore, teams should integrate predictive latency models into search pipelines to evaluate latency without dedicated device testing farms.

The current evaluation focuses on convolutional neural networks for image classification, and latency models were calibrated against a specific set of hardware devices. While confidence in the benchmarked vision results is high, teams should calibrate latency estimators on their specific target hardware before running architecture searches for new application domains.

  • Paper: Searching for MobileNetV3, Andrew Howard et al. (2019). Combines platform-aware neural architecture search concepts from ProxylessNAS and MnasNet with NetAdapt and novel layer designs to deliver next-generation mobile architectures.
  • Paper: EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, Mingxing Tan et al. (2019). Uses neural architecture search to find a balanced baseline network and introduces compound scaling to systematically scale mobile architectures across depth, width, and resolution.
  • Paper: EfficientNetV2: Smaller Models and Faster Training, Mingxing Tan et al. (2021). Extends hardware-aware and training-aware NAS principles to systematically co-optimize training speed, parameter efficiency, and accuracy.
  • Paper: EfficientDet: Scalable and Efficient Object Detection, Mingxing Tan et al. (2020). Leverages NAS-optimized backbones and compound scaling principles to construct highly efficient multi-scale architectures for object detection.
  • Paper: Designing Network Design Spaces, Ilija Radosavovic et al. (2020). Shifts the focus from searching for single point-solution architectures to designing entire population spaces of efficient networks via statistical analysis.
Cover for ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware

Abstract

Neural architecture search (NAS) has a great impact by automatically designing effective neural network architectures. However, the prohibitive computational demand of conventional NAS algorithms (e.g. 10410^4 GPU hours) makes it difficult to \emph{directly} search the architectures on large-scale tasks (e.g. ImageNet). Differentiable NAS can reduce the cost of GPU hours via a continuous representation of network architecture but suffers from the high GPU memory consumption issue (grow linearly w.r.t. candidate set size). As a result, they need to utilize~\emph{proxy} tasks, such as training on a smaller dataset, or learning with only a few blocks, or training just for a few epochs. These architectures optimized on proxy tasks are not guaranteed to be optimal on the target task. In this paper, we present \emph{ProxylessNAS} that can \emph{directly} learn the architectures for large-scale target tasks and target hardware platforms. We address the high memory consumption issue of differentiable NAS and reduce the computational cost (GPU hours and GPU memory) to the same level of regular training while still allowing a large candidate set. Experiments on CIFAR-10 and ImageNet demonstrate the effectiveness of directness and specialization. On CIFAR-10, our model achieves 2.08% test error with only 5.7M parameters, better than the previous state-of-the-art architecture AmoebaNet-B, while using 6×\times fewer parameters. On ImageNet, our model achieves 3.1% better top-1 accuracy than MobileNetV2, while being 1.2×\times faster with measured GPU latency. We also apply ProxylessNAS to specialize neural architectures for hardware with direct hardware metrics (e.g. latency) and provide insights for efficient CNN architecture design.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Construction of Over-Parameterized Network
  • 3.2 Learning Binarized Path
  • 3.2.1 Training Binarized Architecture Parameters
  • 3.3 Handling Non-differentiable Hardware Metrics
  • 3.3.1 Making Latency Differentiable
  • 3.3.2 REINFORCE-based Approach
  • 4 Experiments and Results
  • 4.1 Experiments on CIFAR-10
  • 4.2 Experiments on ImageNet
  • 5 Conclusion
  • References
  • A The List of Candidate Operations Used on CIFAR-10
  • B Mobile Latency Prediction
  • C Details of MnasNet’s Search Cost
  • D Implementaion of the Gradient-Based Algorithm

Knowls

  1. Knowl 1 — Path Binarization for Memory-Efficient Mixed Operations

    model/method

    In differentiable neural architecture search over an over-parameterized directed acyclic graph with NN candidate primitive operations O={o1,…,oN}\mathcal{O} = \{o_1, \dots, o_N\} on an edge, continuous relaxation typically computes a weighted sum of all candidate paths, scaling runtime GPU memory linearly with NN. Path binarization reduces the active memory footprint to that of a single compact model (O(1)O(1) with respect to NN) by sampling a one-hot binary gate vector.

    Each candidate operation oio_i is parameterized by a real-valued architecture parameter αi\alpha_i. Path selection probabilities p=(p1,…,pN)p = (p_1, \dots, p_N) are computed via the softmax function:

    pi=exp⁡(αi)∑j=1Nexp⁡(αj)p_i = \frac{\exp(\alpha_i)}{\sum_{j=1}^N \exp(\alpha_j)}

    A binary gate vector g=(g1,…,gN)∈{0,1}Ng = (g_1, \dots, g_N) \in \{0, 1\}^N is sampled from the multinomial distribution:

    g=binarize(p1,…,pN)={[1,0,…,0]with probability p1⋮[0,0,…,1]with probability pNg = \text{binarize}(p_1, \dots, p_N) = \begin{cases} [1, 0, \dots, 0] & \text{with probability } p_1 \\ & \vdots \\ [0, 0, \dots, 1] & \text{with probability } p_N \end{cases}

    The output of the binarized mixed operation mOBinary(x)m_{\mathcal{O}}^{\text{Binary}}(x) applied to an input tensor xx is defined as:

    mOBinary(x)=∑i=1Ngioi(x)={o1(x)with probability p1⋮oN(x)with probability pNm_{\mathcal{O}}^{\text{Binary}}(x) = \sum_{i=1}^N g_i o_i(x) = \begin{cases} o_1(x) & \text{with probability } p_1 \\ & \vdots \\ o_N(x) & \text{with probability } p_N \end{cases}

    During any single forward and backward pass, only the path with gi=1g_i = 1 is activated in GPU memory, avoiding activation memory accumulation across all candidate branches.

  2. Knowl 2 — Two-Path Factorization for Binarized Architecture Parameter Updates

    algorithm

    To train real-valued architecture parameters α={α1,…,αN}\alpha = \{\alpha_1, \dots, \alpha_N\} when paths are binarized, gradients with respect to path probabilities are approximated using gradients with respect to binary gates gg: ∂L∂αi≈∑j=1N∂L∂gjpj(δij−pi)\frac{\partial \mathcal{L}}{\partial \alpha_i} \approx \sum_{j=1}^N \frac{\partial \mathcal{L}}{\partial g_j} p_j (\delta_{ij} - p_i), where δij=1\delta_{ij} = 1 if i=ji = j and 00 otherwise. Evaluating ∂L∂gj\frac{\partial \mathcal{L}}{\partial g_j} across all NN candidates would require calculating all oj(x)o_j(x), forfeiting memory savings. Two-path factorization resolves this by reducing the NN-ary selection task into pairwise comparisons.

    Input: Architecture parameters α={α1,…,αN}\alpha = \{\alpha_1, \dots, \alpha_N\}, weight parameters WW, validation mini-batch (xval,yval)(x_{\text{val}}, y_{\text{val}}), learning rate η\eta
    Output: Updated architecture parameters α\alpha
    Compute path probabilities pk=exp⁡(αk)/∑m=1Nexp⁡(αm)p_k = \exp(\alpha_k) / \sum_{m=1}^N \exp(\alpha_m) for all k∈{1,…,N}k \in \{1, \dots, N\}
    Sample two distinct path indices i,j∼Multinomial(p1,…,pN)i, j \sim \text{Multinomial}(p_1, \dots, p_N)
    Mask all other N−2N - 2 paths
    Compute normalized probabilities for the sampled pair: p~i=exp⁡(αi)/(exp⁡(αi)+exp⁡(αj))\tilde{p}_i = \exp(\alpha_i) / (\exp(\alpha_i) + \exp(\alpha_j)), p~j=exp⁡(αj)/(exp⁡(αi)+exp⁡(αj))\tilde{p}_j = \exp(\alpha_j) / (\exp(\alpha_i) + \exp(\alpha_j))
    Sample binary gates g~=(g~i,g~j)∼Multinomial(p~i,p~j)\tilde{g} = (\tilde{g}_i, \tilde{g}_j) \sim \text{Multinomial}(\tilde{p}_i, \tilde{p}_j)
    Forward pass on xvalx_{\text{val}} using active binary gate to obtain loss L(W,g~)\mathcal{L}(W, \tilde{g})
    Backward pass to obtain ∂L∂g~i\frac{\partial \mathcal{L}}{\partial \tilde{g}_i} and ∂L∂g~j\frac{\partial \mathcal{L}}{\partial \tilde{g}_j}
    Compute parameter gradients:
      Δαi=∂L∂g~ip~i(1−p~i)−∂L∂g~jp~jp~i\Delta \alpha_i = \frac{\partial \mathcal{L}}{\partial \tilde{g}_i} \tilde{p}_i (1 - \tilde{p}_i) - \frac{\partial \mathcal{L}}{\partial \tilde{g}_j} \tilde{p}_j \tilde{p}_i
      Δαj=∂L∂g~jp~j(1−p~j)−∂L∂g~ip~ip~j\Delta \alpha_j = \frac{\partial \mathcal{L}}{\partial \tilde{g}_j} \tilde{p}_j (1 - \tilde{p}_j) - \frac{\partial \mathcal{L}}{\partial \tilde{g}_i} \tilde{p}_i \tilde{p}_j
    Update sampled parameters: αi←αi−ηΔαi\alpha_i \leftarrow \alpha_i - \eta \Delta \alpha_i, αj←αj−ηΔαj\alpha_j \leftarrow \alpha_j - \eta \Delta \alpha_j
    Rescale updated parameters by multiplying by (αiold+αjold)/(αinew+αjnew)(\alpha_i^{\text{old}} + \alpha_j^{\text{old}}) / (\alpha_i^{\text{new}} + \alpha_j^{\text{new}}) to keep the relative probabilities of all unselected paths unchanged
    return α\alpha
  3. Knowl 3 — Differentiable Latency Regularization Loss

    equation

    To directly incorporate non-differentiable hardware execution latency into gradient-based architecture optimization, the latency of a network composed of sequentially executed mixed operations is modeled as a differentiable expectation over candidate operation latencies.

    For the ii-th learnable block with candidate operation set Oi={oji}\mathcal{O}^i = \{o_j^i\} and path probabilities pjip_j^i, the expected latency E[latencyi]\mathbb{E}[\text{latency}_i] is defined as:

    E[latencyi]=∑jpji×F(oji)\mathbb{E}[\text{latency}_i] = \sum_j p_j^i \times F(o_j^i)

    where F(oji)F(o_j^i) denotes the latency of candidate operation ojio_j^i predicted by a hardware-specific latency model. The gradient of the block's expected latency with respect to its architecture probability pjip_j^i is:

    ∂E[latencyi]∂pji=F(oji)\frac{\partial \mathbb{E}[\text{latency}_i]}{\partial p_j^i} = F(o_j^i)

    For a network consisting of a sequence of such blocks, the total expected network latency is additive:

    E[latency]=∑iE[latencyi]\mathbb{E}[\text{latency}] = \sum_i \mathbb{E}[\text{latency}_i]

    The total optimization loss function combines task performance, weight decay, and expected latency:

    Loss=LossCE+λ1∥w∥22+λ2E[latency]\text{Loss} = \text{Loss}_{\text{CE}} + \lambda_1 \|w\|_2^2 + \lambda_2 \mathbb{E}[\text{latency}]

    where LossCE\text{Loss}_{\text{CE}} is cross-entropy loss, ww is the model weight tensor, λ1\lambda_1 is the weight decay coefficient, and λ2>0\lambda_2 > 0 is a scaling factor governing the accuracy-latency trade-off.

  4. Knowl 4 — REINFORCE-Based Binarized Architecture Optimization

    model/method

    As an alternative to gradient backpropagation through BinaryConnect approximations, the binarized architecture parameters α\alpha can be optimized using the REINFORCE policy gradient algorithm. This accommodates arbitrary non-differentiable reward objectives R(Ng)R(N_g), such as discrete hardware metrics, without needing a separate recurrent meta-controller or hypernetwork.

    Let gg be the binary gate configuration sampled according to probability distribution p(g)p(g) parameterized by α\alpha, and let NgN_g denote the compact sub-network activated by gates gg. The expected reward objective J(α)J(\alpha) and its gradient with respect to α\alpha are:

    J(α)=Eg∼α[R(Ng)]J(\alpha) = \mathbb{E}_{g \sim \alpha}[R(N_g)]

    ∇αJ(α)=∑iR(N(e=oi))∇αpi=Eg∼α[R(Ng)∇αlog⁡(p(g))]≈1M∑m=1MR(Ngm)∇αlog⁡(p(gm))\nabla_\alpha J(\alpha) = \sum_i R(N(e = o_i)) \nabla_\alpha p_i = \mathbb{E}_{g \sim \alpha}[R(N_g) \nabla_\alpha \log(p(g))] \approx \frac{1}{M} \sum_{m=1}^M R(N_{g^m}) \nabla_\alpha \log(p(g^m))

    where gmg^m denotes the mm-th sampled binary gate configuration and MM is the number of sampled architectures per step.

    For multi-objective search balancing validation accuracy ACC(m)\text{ACC}(m) and measured hardware latency LAT(m)\text{LAT}(m) against a target latency constraint TT, the reward function is parameterized as:

    R(m)=ACC(m)×[LAT(m)T]wR(m) = \text{ACC}(m) \times \left[ \frac{\text{LAT}(m)}{T} \right]^w

    where ww is a negative exponent tuning the severity of the latency penalty.

  5. Knowl 5 — MobileNetV2-Based Search Space with Elastic Depth and Width

    definition

    The ImageNet architecture search space is constructed using MobileNetV2 inverted bottleneck convolution (MBConv) stages as the backbone without forcing repeated block motifs across layers. Each learnable block in the over-parameterized network selects among candidate operations defined by:

    1. Kernel sizes {3,5,7}\{3, 5, 7\}
    2. Expansion ratios {3,6}\{3, 6\}
    3. An explicit identity skip / zero operation (allowing residual blocks to be completely bypassed)

    By including the zero operation in the candidate set of residual blocks, the search space allows dynamic trade-offs between depth and width under a fixed hardware latency constraint: the optimization can select shallower, wider networks (pruning blocks via zero ops while choosing large MBConv expansion/kernel sizes) or deeper, thinner networks (retaining more blocks with smaller kernels and expansion factors).

  6. Knowl 6 — Lookup-Based Additive Latency Prediction Model

    model/method

    To avoid the measurement overhead and engineering cost of deploying every candidate network to a physical hardware farm during architecture search, network latency is predicted via an additive layer-level latency model F(⋅)F(\cdot).

    A lookup table / prediction function is trained on a benchmark set of 4,000 candidate architectures sampled from the search space and measured on the target device (e.g., Google Pixel 1 phone via TensorFlow-Lite). The feature vector for predicting an operation oo's latency F(o)F(o) comprises:

    1. Operator type
    2. Input feature map spatial dimensions and channel depth
    3. Output feature map spatial dimensions and channel depth
    4. Convolution kernel size
    5. Stride
    6. MBConv expansion ratio

    Because layer executions in feedforward CNNs proceed sequentially during inference, total latency is estimated as the sum of predicted individual layer latencies: Ftotal=∑iF(oi)F_{\text{total}} = \sum_i F(o^i). Evaluated on a holdout test set of 1,000 architectures, the additive latency prediction model achieved a root-mean-square error (RMSE) of 0.75 ms0.75\,\text{ms} on Google Pixel 1.

  7. Knowl 7 — ImageNet Classification and Search Cost Under Mobile Latency Constraints

    data/table

    ProxylessNAS directly optimizes MobileNetV2-based search spaces on the full ImageNet dataset without proxy tasks, achieving superior top-1 accuracy under mobile latency constraints (≤80 ms\le 80\,\text{ms} on Google Pixel 1) while reducing search cost by 200×200\times compared to RL-based search (MnasNet).

    Model Top-1 (%) Top-5 (%) Mobile Latency Hardware-aware No Proxy Search Cost (GPU hours)
    MobileNetV1 70.6 89.5 113ms - - Manual
    MobileNetV2 72.0 91.0 75ms - - Manual
    NASNet-A 74.0 91.3 183ms No No 48,000
    AmoebaNet-A 74.5 92.0 190ms No No 75,600
    MnasNet 74.0 91.8 76ms Yes No 40,000
    MnasNet (impl.) 74.0 91.8 79ms Yes No 40,000
    Proxyless-G (mobile) 71.8 90.3 83ms No Yes 200
    Proxyless-G + LL 74.2 91.7 79ms Yes Yes 200
    Proxyless-R (mobile) 74.6 92.2 78ms Yes Yes 200

    Key results:

    1. Proxyless-R (mobile) reaches 74.6%74.6\% top-1 accuracy at 78 ms78\,\text{ms} latency on Google Pixel 1, outperforming MnasNet (74.0%74.0\%) while cutting search cost from 40,000 GPU hours to 200 GPU hours (200×200\times reduction).
    2. Incorporating latency regularization loss (LL) is critical for gradient-based search: Proxyless-G without LL achieves only 71.8%71.8\% accuracy when rescaled to match target latency, compared to 74.2%74.2\% with LL.
  8. Knowl 8 — Cross-Platform Latency Comparison of Hardware-Specialized Architectures

    data/table

    When neural architectures are searched specifically for distinct hardware platforms (GPU, CPU, mobile), the resulting models exhibit significant hardware specialization: a model optimized for one device does not achieve optimal speed on others.

    Model Top-1 Accuracy (%) GPU Latency (V100) CPU Latency (Xeon) Mobile Latency (Pixel 1)
    Proxyless (GPU) 75.1 5.1ms 204.9ms 124ms
    Proxyless (CPU) 75.3 7.4ms 138.7ms 116ms
    Proxyless (mobile) 74.6 7.2ms 164.1ms 78ms

    Latency measurements are conducted on a Tesla V100 GPU (batch size 8), dual 2.40GHz Intel Xeon E5-2640 v4 CPUs (batch size 1), and a Google Pixel 1 phone (batch size 1). The model specialized for GPU achieves the fastest GPU runtime (5.1 ms5.1\,\text{ms}) but exhibits the slowest CPU (204.9 ms204.9\,\text{ms}) and mobile (124 ms124\,\text{ms}) latencies. Conversely, the mobile-specialized model achieves 78 ms78\,\text{ms} on mobile but 7.2 ms7.2\,\text{ms} on GPU.

  9. Knowl 9 — Hardware-Specific Architecture Specialization Patterns

    empirical result

    Unconstrained layer-by-layer architecture search on target hardware reveals distinct structural patterns tailored to underlying hardware compute paradigms:

    1. GPU vs. CPU structure: GPU architectures prefer shallower and wider topologies with early spatial downsampling (early pooling) and large operators (e.g., 7×77 \times 7 MBConv6), leveraging the high degree of parallel execution. In contrast, CPU architectures prefer deeper, narrower topologies with late spatial downsampling (late pooling) and smaller operators (3×33 \times 3 or 5×55 \times 5 MBConv3/6) due to limited parallel thread capacity.
    2. Kernel size distribution by depth: Across all hardware platforms, early network layers favor smaller kernel sizes (3×33 \times 3), whereas later layers closer to the classifier head favor larger kernel sizes (5×55 \times 5 and 7×77 \times 7).
    3. Downsampling stage transitions: On all platforms, the first block in each downsampling stage preferentially selects larger MBConv operations, preserving feature information across spatial dimension reductions.
  10. Knowl 10 — CIFAR-10 Classification Performance and Parameter Efficiency

    data/table

    On CIFAR-10 image classification using a tree-structured cell architecture space built upon PyramidNet (B=18B=18 blocks, final width F=400F=400, 648 total unrepeated routing decisions), ProxylessNAS achieves state-of-the-art accuracy with superior parameter efficiency compared to proxy-based and weight-sharing NAS methods.

    Model Parameters Test Error (%)
    DenseNet-BC 25.6M 3.46
    PyramidNet 26.0M 3.31
    Shake-Shake + c/o 26.2M 2.56
    PyramidNet + SD 26.0M 2.31
    ENAS + c/o 4.6M 2.89
    DARTS + c/o 3.4M 2.83
    NASNet-A + c/o 27.6M 2.40
    PathLevel EAS + c/o 14.3M 2.30
    AmoebaNet-B + c/o 34.9M 2.13
    Proxyless-R + c/o (ours) 5.8M 2.30
    Proxyless-G + c/o (ours) 5.7M 2.08

    Note: "c/o" denotes training with Cutout data augmentation; "SD" denotes ShakeDrop regularization.

    Proxyless-G achieves a test error of 2.08%2.08\% with 5.7 M5.7\,\text{M} parameters, outperforming AmoebaNet-B (2.13%2.13\% error) while using over 6×6\times fewer parameters (5.7 M5.7\,\text{M} vs. 34.9 M34.9\,\text{M}). Both Proxyless-G and Proxyless-R match or exceed PathLevel EAS while using roughly half the parameter count, demonstrating the advantage of searching unconstrained paths without repeating identical block motifs.

Coverage note — None was omitted; all key contributions—path binarization, two-path factorization gradient search, differentiable latency regularization, REINFORCE search variant, mobile latency model, search spaces, ImageNet/CIFAR-10 benchmarks, and hardware specialization insights—are covered.

References

  1. 1.Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le. Understanding and simplifying one-shot architecture search. In ICML, 2018.
  2. 2.Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston. Smash: one-shot model architecture search through hypernetworks. In ICLR, 2018.
  3. 3.Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang. Efficient architecture search by network transformation. In AAAI, 2018a.
  4. 4.Han Cai, Jiacheng Yang, Weinan Zhang, Song Han, and Yong Yu. Path-level network transformation for efficient architecture search. In ICML, 2018b.
  5. 5.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In NIPS, 2015.
  6. 6.Terrance DeVries and Graham W Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017.
  7. 7.Jin-Dong Dong, An-Chieh Cheng, Da-Cheng Juan, Wei Wei, and Min Sun. Dpp-net: Device-aware progressive search for pareto-optimal neural architectures. In ECCV, 2018.
  8. 8.Thomas Elsken, Jan-Hendrik Metzen, and Frank Hutter. Simple and efficient architecture search for convolutional neural networks. arXiv preprint arXiv:1711.04528, 2017.
  9. 9.Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Multi-objective architecture search for cnns. arXiv preprint arXiv:1804.09081, 2018a.
  10. 10.Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. arXiv preprint arXiv:1808.05377, 2018b.
  11. 11.Dongyoon Han, Jiwhan Kim, and Junmo Kim. Deep pyramidal residual networks. In CVPR, 2017.
  12. 12.Song Han, Jeff Pool, John Tran, and William Dally. Learning both weights and connections for efficient neural network. In NIPS, 2015.
  13. 13.Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. In ICLR, 2016.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  15. 15.Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. Amc: Automl for model compression and acceleration on mobile devices. In ECCV, 2018.
  16. 16.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  17. 17.Chi-Hung Hsu, Shu-Huan Chang, Da-Cheng Juan, Jia-Yu Pan, Yu-Ting Chen, Wei Wei, and Shih-Chieh Chang. Monas: Multi-objective neural architecture search using reinforcement learning. arXiv preprint arXiv:1806.10332, 2018.
  18. 18.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In ECCV, 2016.
  19. 19.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In CVPR, 2017.
  20. 20.Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size. arXiv preprint arXiv:1602.07360, 2016.
  21. 21.Purushotham Kamath, Abhishek Singh, and Debo Dutta. Neural architecture construction using envelopenets. arXiv preprint arXiv:1803.06744, 2018.
  22. 22.Chenxi Liu, Barret Zoph, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan Yuille, Jonathan Huang, and Kevin Murphy. Progressive neural architecture search. In ECCV, 2018a.
  23. 23.Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, and Koray Kavukcuoglu. Hierarchical representations for efficient architecture search. In ICLR, 2018b.
  24. 24.Hanxiao Liu, Karen Simonyan, and Yiming Yang. Darts: Differentiable architecture search. arXiv preprint arXiv:1806.09055, 2018c.
  25. 25.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In ICCV, 2017.
  26. 26.Renqian Luo, Fei Tian, Tao Qin, and Tie-Yan Liu. Neural architecture optimization. arXiv preprint arXiv:1808.07233, 2018.
  27. 27.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In ECCV, 2018.
  28. 28.Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean. Efficient neural architecture search via parameter sharing. In ICML, 2018.
  29. 29.Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. Regularized evolution for image classifier architecture search. arXiv preprint arXiv:1802.01548, 2018.
  30. 30.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018.
  31. 31.Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, and Quoc V Le. Mnasnet: Platform-aware neural architecture search for mobile. arXiv preprint arXiv:1807.11626, 2018.
  32. 32.Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. Haq: Hardware-aware automated quantization. arXiv, 2018.
  33. 33.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. In Reinforcement Learning. 1992.
  34. 34.Yoshihiro Yamada, Masakazu Iwamura, and Koichi Kise. Shakedrop regularization. arXiv preprint arXiv:1802.02375, 2018.
  35. 35.Zhao Zhong, Junjie Yan, Wei Wu, Jing Shao, and Cheng-Lin Liu. Practical block-wise neural network architecture generation. In CVPR, 2018.
  36. 36.Ligeng Zhu, Ruizhi Deng, Michael Maire, Zhiwei Deng, Greg Mori, and Ping Tan. Sparsely aggregated convolutional networks. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 186–201, 2018.
  37. 37.Barret Zoph and Quoc V Le. Neural architecture search with reinforcement learning. In ICLR, 2017.
  38. 38.Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le. Learning transferable architectures for scalable image recognition. In CVPR, 2018.

Citation

MLA
Cai, H., et al. “ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware”. arXiv, 2018, http://arxiv.org/abs/1812.00332v2.
APA
Cai, H., Zhu, L., & Han, S. (2018). ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. arXiv. http://arxiv.org/abs/1812.00332v2
Chicago
Cai, H., L. Zhu, and S. Han. 2018. “ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware”. arXiv. http://arxiv.org/abs/1812.00332v2.
Harvard
Cai, H., Zhu, L. and Han, S. (2018) “ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1812.00332v2.
Vancouver
1. Cai H, Zhu L, Han S (2018) ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. arXiv

BibTeX

@article{cai2018proxylessnas,
  title = {ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware},
  author = {Cai, Han and Zhu, Ligeng and Han, Song},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1812.00332v2},
  eprint = {1812.00332}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/