ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
Han CaiLigeng ZhuSong Han
Introduces ProxylessNAS, a differentiable neural architecture search framework that reduces GPU memory to standard training levels, enabling direct model optimization on large-scale datasets and specific hardware platforms without relying on proxy tasks.
Automating the design of deep learning models through neural architecture search (NAS) is critical for artificial intelligence deployment, but conventional methods demand prohibitive computational resources (tens of thousands of GPU hours). To manage these computational costs, existing methods rely on proxy tasks—such as searching on smaller datasets, using fewer layers, or repeatedly stacking identical building blocks. These proxy shortcuts produce sub-optimal architectures because they restrict architectural diversity and fail to reflect true target performance and hardware latency.
The article demonstrates ProxylessNAS, an approach that optimizes neural network architectures directly on large-scale target tasks and target hardware without using proxy tasks, while reducing memory and computation costs to the level of standard model training.
The researchers formulated architecture search as a path-level pruning process within an over-parameterized super-network containing all candidate operations. To resolve severe GPU memory bottlenecks, they introduced path binarization, which activates only one or two candidate paths during training steps rather than keeping all options in memory simultaneously. They also incorporated non-differentiable hardware execution speed (latency) into the gradient-based optimization using a continuous latency prediction model, as well as an alternative reinforcement learning formulation. The method was evaluated on standard image classification benchmarks (CIFAR-10 and ImageNet) across three hardware platforms: mobile phones (Google Pixel 1), cloud GPUs (Nvidia Tesla V100), and CPUs (Intel Xeon).
The evaluation produced several key findings. First, ProxylessNAS cut the computational search cost on ImageNet to 200 GPU hours—a 200-fold reduction compared to prior methods like MnasNet, which required roughly 40,000 GPU hours. Second, on ImageNet mobile benchmarks, the discovered architecture improved top-1 accuracy by 2.6% over MobileNetV2 at equivalent latency, and ran 1.83 times faster when matched for accuracy. Third, on CIFAR-10, it achieved a 2.08% test error with only 5.7 million parameters, matching or beating top baseline models while using six times fewer parameters. Finally, hardware-specific searches proved that optimal architectures vary substantially across device types: GPU-optimized models favored shallower, wider layers with larger operations due to high parallelism, whereas CPU models favored deeper, narrower structures with smaller operations.
These findings indicate that organizations can design custom, high-performing neural networks directly for specific hardware endpoints at a fraction of standard compute costs. This eliminates the expense of running large device-testing farms and prevents the efficiency losses that occur when deploying generic models across diverse hardware platforms. Teams building edge and cloud AI applications can reduce latency, lower cloud training expenditures, and improve user device responsiveness.
Organizations should adopt direct, hardware-aware architecture search and tailor deployed models to specific execution platforms (GPU, CPU, mobile) rather than deploying a single cross-platform model. Furthermore, teams should integrate predictive latency models into search pipelines to evaluate latency without dedicated device testing farms.
The current evaluation focuses on convolutional neural networks for image classification, and latency models were calibrated against a specific set of hardware devices. While confidence in the benchmarked vision results is high, teams should calibrate latency estimators on their specific target hardware before running architecture searches for new application domains.
- Paper: DARTS: Differentiable Architecture Search, Hanxiao Liu et al. (2018). Introduces continuous relaxation and differentiable neural architecture search, establishing the gradient-based optimization framework whose high memory overhead ProxylessNAS directly tackles.
- Paper: MnasNet: Platform-Aware Neural Architecture Search for Mobile, Mingxing Tan et al. (2018). Pioneers multi-objective platform-aware neural architecture search incorporating measured hardware latency into model design, a primary objective directly built upon and optimized by ProxylessNAS.
- Paper: Efficient Neural Architecture Search via Parameter Sharing, Hieu Pham et al. (2018). Presents weight sharing across candidate subgraphs in an over-parameterized supernet, forming a conceptual foundation for efficient search spaces in NAS.
- Paper: Learning Transferable Architectures for Scalable Image Recognition, Barret Zoph et al. (2018). Establishes the cell-based proxy-task paradigm (searching on CIFAR-10 to transfer to ImageNet), which ProxylessNAS explicitly critiques and overcomes by searching directly on the target dataset.
- Paper: MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler et al. (2018). Defines inverted residual blocks with linear bottlenecks, which provide the core candidate operation blocks used to build ProxylessNAS's search space.
- Paper: Neural Architecture Search with Reinforcement Learning, Barret Zoph et al. (2016). Introduces foundational reinforcement learning for automated neural architecture search, providing historical context for the high GPU-hour costs that ProxylessNAS eliminates.
- Paper: Progressive Neural Architecture Search, Chenxi Liu et al. (2017). Demonstrates surrogate-guided sequential search to mitigate search cost, illustrating early efforts to reduce the heavy computation of conventional NAS.
- Paper: ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design, Ningning Ma et al. (2018). Provides practical guidelines for optimizing inference speed and direct hardware efficiency metrics rather than theoretical FLOPs alone.
- Paper: Searching for MobileNetV3, Andrew Howard et al. (2019). Combines platform-aware neural architecture search concepts from ProxylessNAS and MnasNet with NetAdapt and novel layer designs to deliver next-generation mobile architectures.
- Paper: EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, Mingxing Tan et al. (2019). Uses neural architecture search to find a balanced baseline network and introduces compound scaling to systematically scale mobile architectures across depth, width, and resolution.
- Paper: EfficientNetV2: Smaller Models and Faster Training, Mingxing Tan et al. (2021). Extends hardware-aware and training-aware NAS principles to systematically co-optimize training speed, parameter efficiency, and accuracy.
- Paper: EfficientDet: Scalable and Efficient Object Detection, Mingxing Tan et al. (2020). Leverages NAS-optimized backbones and compound scaling principles to construct highly efficient multi-scale architectures for object detection.
- Paper: Designing Network Design Spaces, Ilija Radosavovic et al. (2020). Shifts the focus from searching for single point-solution architectures to designing entire population spaces of efficient networks via statistical analysis.
