FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search
Bichen WuXiaoliang DaiPeizhao ZhangYanghan WangFei SunYiming WuYuandong TianPeter VajdaYangqing JiaKurt Keutzer
Introduces a hardware-aware differentiable neural architecture search framework that directly optimizes ConvNets for device-specific latency, cutting search costs by over 400 times compared to prior methods while achieving superior speed and accuracy on mobile devices.
Deploying accurate convolutional neural networks on resource-constrained mobile devices is a critical challenge in computer vision. Standard manual design and previous automated neural architecture search methods are either computationally prohibitive or rely heavily on theoretical operation counts (FLOPs) that do not reflect actual latency on physical hardware. Because optimal model design changes significantly depending on target hardware and input resolutions, organizations struggle to efficiently generate tailored, high-performing mobile models.
The article demonstrates a differentiable neural architecture search (DNAS) framework that automates the discovery of accurate and hardware-efficient neural networks, termed FBNets (Facebook-Berkeley-Nets). The core objective is to minimize both classification error and real on-device latency within a single, gradient-based optimization process.
The authors approach this by framing the search space as a stochastic super net across 22 layers, allowing independent block selection per layer out of nine candidate configurations. Rather than training thousands of candidate networks individually, the method optimizes an architecture probability distribution directly using gradient descent. To make hardware execution differentiable and avoid running millions of device tests, the approach pre-measures individual operator runtimes on target hardware and estimates total network latency via an additive lookup table. The framework was evaluated on the ImageNet classification dataset targeting mobile platforms, specifically the Samsung Galaxy S8 (Snapdragon 835) and Apple iPhone X (A11 Bionic).
The results highlight major improvements in search efficiency and deployment performance. First, the search cost dropped dramatically: finding an optimal FBNet required only 216 GPU-hours, which is roughly 420 times faster than prior reinforcement-learning methods like MnasNet. Second, FBNet models outperformed both manually and automatically designed baselines on ImageNet; for instance, FBNet-B achieved 74.1% top-1 accuracy at 23.1 ms latency on a Samsung Galaxy S8, running 1.5 times faster and remaining 2.4 times smaller than MobileNetV2-1.3 with comparable accuracy. Third, when exploring smaller input resolutions and channel scaling, tailored FBNets achieved 1.5% to 6.4% higher top-1 accuracy than scaled MobileNetV2 baselines, with the smallest FBNet reaching 50.2% accuracy at 2.9 ms latency (345 frames per second). Finally, the evaluation confirmed that model optimality is strictly device-dependent: a model tailored specifically for iPhone X ran 1.4 times faster on that device than a model optimized for Samsung S8 due to differing hardware operator efficiencies.
These findings imply that engineering teams can transition from one-size-fits-all neural networks to device- and use-case-specific designs without incurring massive computing costs. By factoring in physical runtime rather than abstract theoretical metrics, teams can substantially reduce inference delays, cut power consumption, and lower operational costs. The results also show that adapting model architecture to reduced image resolutions prevents redundant layer computations, further optimizing mobile execution.
Organizations deploying vision models on edge hardware should adopt hardware-aware, gradient-based search workflows to customize architectures for specific target processors. Deployment pipelines should measure operator-level runtimes on intended chips before model compilation rather than relying on theoretical FLOP counts. Future work should expand this search approach to broader vision tasks such as real-time object detection and segmentation, as well as additional hardware backends like digital signal processors and microcontrollers.
Confidence in these findings is supported by consistent benchmark validations across standard datasets and physical mobile hardware. However, readers should note that the latency lookup table assumes sequential, independent operator execution, an assumption that holds well for single-threaded mobile CPU and DSP inference but may require adjustments on hardware platforms with complex concurrent scheduling or heterogeneous memory hierarchies.
- Paper: MnasNet: Platform-Aware Neural Architecture Search for Mobile, Mingxing Tan et al. (2018). FBNet directly benchmarks against and seeks to overcome the high computational search cost of MnasNet's platform-aware reinforcement learning architecture search.
- Paper: DARTS: Differentiable Architecture Search, Hanxiao Liu et al. (2018). DARTS provides the foundational continuous relaxation and gradient-based bilevel optimization principles that FBNet builds upon for differentiable neural architecture search.
- Paper: MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler et al. (2018). FBNet builds directly upon the inverted residual and linear bottleneck building blocks introduced in MobileNetV2 to form its search space.
- Paper: Efficient Neural Architecture Search via Parameter Sharing, Hieu Pham et al. (2018). ENAS introduces the concept of weight sharing over a supernet directed acyclic graph to drastically reduce the cost of neural architecture search.
- Paper: ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design, Ningning Ma et al. (2018). ShuffleNet V2 establishes the core motivation that hardware latency, rather than theoretical FLOP counts, must guide efficient network design.
- Paper: MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, Andrew G. Howard et al. (2017). MobileNets establishes depthwise separable convolutions and efficiency trade-offs for mobile vision architectures that FBNet automates and optimizes.
- Paper: Learning Transferable Architectures for Scalable Image Recognition, Barret Zoph et al. (2018). NASNet establishes the modular search-space paradigms and scalable vision transfer benchmarks that subsequent mobile NAS frameworks like FBNet refine.
- Paper: Neural Architecture Search with Reinforcement Learning, Barret Zoph et al. (2016). This seminal paper introduces Neural Architecture Search using controllers and reinforcement learning, presenting the high-cost search baseline that differentiable methods aim to solve.
- Paper: ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware, Han Cai et al. (2018). ProxylessNAS builds upon direct gradient-based hardware-aware NAS by introducing path binarization to reduce supernet memory overhead.
- Paper: Searching for MobileNetV3, Andrew Howard et al. (2019). MobileNetV3 combines platform-aware automated search concepts with complementary layer-level optimizations like hard-swish and squeeze-and-excitation.
- Paper: Once for All: Train One Network and Specialize it for Efficient Deployment, Han Cai et al. (2019). Once-for-All advances hardware-aware architecture search by decoupling supernet training from search to deploy specialized sub-networks without repeated retraining.
- Paper: EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, Mingxing Tan et al. (2019). EfficientNet extends mobile-searched baseline ConvNets by establishing a systematic compound scaling method across depth, width, and resolution.
- Paper: NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection, Golnaz Ghiasi et al. (2019). NAS-FPN applies automated architecture search concepts to learn cross-scale feature pyramid connections for downstream object detection.
- Paper: EfficientNetV2: Smaller Models and Faster Training, Mingxing Tan et al. (2021). EfficientNetV2 incorporates training-speed-aware search and fused inverted bottleneck operations to further optimize mobile-derived architectures.
- Paper: EfficientDet: Scalable and Efficient Object Detection, Mingxing Tan et al. (2020). EfficientDet integrates scalable, efficient ConvNet backbones with learned bidirectional feature networks for resource-constrained object detection.
- Paper: GhostNet: More Features From Cheap Operations, Kai Han et al. (2019). GhostNet introduces cheap linear operations to generate redundant feature maps within lightweight mobile ConvNet layouts like those generated by FBNet.
