FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search

Bichen WuXiaoliang DaiPeizhao ZhangYanghan WangFei SunYiming WuYuandong TianPeter VajdaYangqing JiaKurt Keutzer

article2018CVPR1,486 citations

Introduces a hardware-aware differentiable neural architecture search framework that directly optimizes ConvNets for device-specific latency, cutting search costs by over 400 times compared to prior methods while achieving superior speed and accuracy on mobile devices.

Listen

Deploying accurate convolutional neural networks on resource-constrained mobile devices is a critical challenge in computer vision. Standard manual design and previous automated neural architecture search methods are either computationally prohibitive or rely heavily on theoretical operation counts (FLOPs) that do not reflect actual latency on physical hardware. Because optimal model design changes significantly depending on target hardware and input resolutions, organizations struggle to efficiently generate tailored, high-performing mobile models.

The article demonstrates a differentiable neural architecture search (DNAS) framework that automates the discovery of accurate and hardware-efficient neural networks, termed FBNets (Facebook-Berkeley-Nets). The core objective is to minimize both classification error and real on-device latency within a single, gradient-based optimization process.

The authors approach this by framing the search space as a stochastic super net across 22 layers, allowing independent block selection per layer out of nine candidate configurations. Rather than training thousands of candidate networks individually, the method optimizes an architecture probability distribution directly using gradient descent. To make hardware execution differentiable and avoid running millions of device tests, the approach pre-measures individual operator runtimes on target hardware and estimates total network latency via an additive lookup table. The framework was evaluated on the ImageNet classification dataset targeting mobile platforms, specifically the Samsung Galaxy S8 (Snapdragon 835) and Apple iPhone X (A11 Bionic).

The results highlight major improvements in search efficiency and deployment performance. First, the search cost dropped dramatically: finding an optimal FBNet required only 216 GPU-hours, which is roughly 420 times faster than prior reinforcement-learning methods like MnasNet. Second, FBNet models outperformed both manually and automatically designed baselines on ImageNet; for instance, FBNet-B achieved 74.1% top-1 accuracy at 23.1 ms latency on a Samsung Galaxy S8, running 1.5 times faster and remaining 2.4 times smaller than MobileNetV2-1.3 with comparable accuracy. Third, when exploring smaller input resolutions and channel scaling, tailored FBNets achieved 1.5% to 6.4% higher top-1 accuracy than scaled MobileNetV2 baselines, with the smallest FBNet reaching 50.2% accuracy at 2.9 ms latency (345 frames per second). Finally, the evaluation confirmed that model optimality is strictly device-dependent: a model tailored specifically for iPhone X ran 1.4 times faster on that device than a model optimized for Samsung S8 due to differing hardware operator efficiencies.

These findings imply that engineering teams can transition from one-size-fits-all neural networks to device- and use-case-specific designs without incurring massive computing costs. By factoring in physical runtime rather than abstract theoretical metrics, teams can substantially reduce inference delays, cut power consumption, and lower operational costs. The results also show that adapting model architecture to reduced image resolutions prevents redundant layer computations, further optimizing mobile execution.

Organizations deploying vision models on edge hardware should adopt hardware-aware, gradient-based search workflows to customize architectures for specific target processors. Deployment pipelines should measure operator-level runtimes on intended chips before model compilation rather than relying on theoretical FLOP counts. Future work should expand this search approach to broader vision tasks such as real-time object detection and segmentation, as well as additional hardware backends like digital signal processors and microcontrollers.

Confidence in these findings is supported by consistent benchmark validations across standard datasets and physical mobile hardware. However, readers should note that the latency lookup table assumes sequential, independent operator execution, an assumption that holds well for single-threaded mobile CPU and DSP inference but may require adjustments on hardware platforms with complex concurrent scheduling or heterogeneous memory hierarchies.

Cover for FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search

Abstract

Designing accurate and efficient ConvNets for mobile devices is challenging because the design space is combinatorially large. Due to this, previous neural architecture search (NAS) methods are computationally expensive. ConvNet architecture optimality depends on factors such as input resolution and target devices. However, existing approaches are too expensive for case-by-case redesigns. Also, previous work focuses primarily on reducing FLOPs, but FLOP count does not always reflect actual latency. To address these, we propose a differentiable neural architecture search (DNAS) framework that uses gradient-based methods to optimize ConvNet architectures, avoiding enumerating and training individual architectures separately as in previous methods. FBNets, a family of models discovered by DNAS surpass state-of-the-art models both designed manually and generated automatically. FBNet-B achieves 74.1% top-1 accuracy on ImageNet with 295M FLOPs and 23.1 ms latency on a Samsung S8 phone, 2.4x smaller and 1.5x faster than MobileNetV2-1.3 with similar accuracy. Despite higher accuracy and lower latency than MnasNet, we estimate FBNet-B's search cost is 420x smaller than MnasNet's, at only 216 GPU-hours. Searched for different resolutions and channel sizes, FBNets achieve 1.5% to 6.4% higher accuracy than MobileNetV2. The smallest FBNet achieves 50.2% accuracy and 2.9 ms latency (345 frames per second) on a Samsung S8. Over a Samsung-optimized FBNet, the iPhone-X-optimized model achieves a 1.4x speedup on an iPhone X.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 The Search Space
  • 3.2 Latency-Aware Loss Function
  • 3.3 The Search Algorithm
  • 4 Experiments
  • 4.1 ImageNet Classification
  • 4.2 Different Resolution and Channel Size Scaling
  • 4.3 Different Target Devices
  • 5 Conclusion
  • References
  • A Experiment details

Knowls

  1. Knowl 1 — Differentiable Neural Architecture Search via Gumbel-Softmax Relaxation

    model/method

    Differentiable Neural Architecture Search (DNAS) formulates neural architecture search over a discrete candidate space A\mathcal{A} as a continuous stochastic supernet optimization problem:

    min⁡θmin⁡wEa∼Pθ[L(a,w)]\min_{\theta} \min_{w} \mathbb{E}_{a \sim P_\theta} \left[ \mathcal{L}(a, w) \right]

    where θ\theta parameterizes the architecture distribution, ww denotes the shared operator weights within the stochastic supernet, and L(a,w)\mathcal{L}(a, w) is the multi-objective loss function.

    In a supernet comprising LL sequential layers, each layer l∈{1,…,L}l \in \{1, \dots, L\} contains NN parallel candidate block operations {bl,i}i=1N\{b_{l,i}\}_{i=1}^N. The discrete sampling probability of selecting block bl,ib_{l,i} at layer ll is parameterized by a real-valued vector θl=(θl,1,…,θl,N)\theta_l = (\theta_{l,1}, \dots, \theta_{l,N}) via the softmax function:

    Pθl(bl=bl,i)=exp⁡(θl,i)∑j=1Nexp⁡(θl,j)P_{\theta_l}(b_l = b_{l,i}) = \frac{\exp(\theta_{l,i})}{\sum_{j=1}^N \exp(\theta_{l,j})}

    Assuming independent layer-wise sampling, the joint probability of an entire network architecture a=(b1,…,bL)a = (b_1, \dots, b_L) is given by:

    Pθ(a)=∏l=1LPθl(bl=bl,i(a))P_\theta(a) = \prod_{l=1}^L P_{\theta_l}\left(b_l = b_{l,i}^{(a)}\right)

    To enable end-to-end backpropagation with respect to architecture parameters θ\theta, the discrete selection mask ml,i∈{0,1}m_{l,i} \in \{0, 1\} is relaxed into a continuous random variable using the Gumbel-Softmax reparameterization:

    ml,i=GumbelSoftmax(θl,i∣θl)=exp⁡((θl,i+gl,i)/τ)∑j=1Nexp⁡((θl,j+gl,j)/τ)m_{l,i} = \text{GumbelSoftmax}(\theta_{l,i} \mid \theta_l) = \frac{\exp\left( (\theta_{l,i} + g_{l,i}) / \tau \right)}{\sum_{j=1}^N \exp\left( (\theta_{l,j} + g_{l,j}) / \tau \right)}

    where gl,i∼Gumbel(0,1)=−log⁡(−log⁡(ul,i))g_{l,i} \sim \text{Gumbel}(0, 1) = -\log(-\log(u_{l,i})) with ul,i∼Uniform(0,1)u_{l,i} \sim \text{Uniform}(0, 1), and τ>0\tau > 0 is a temperature hyperparameter. As τ→0\tau \to 0, ml,im_{l,i} approaches a one-hot categorical sample, while for τ>0\tau > 0, ml,im_{l,i} remains continuously differentiable with respect to θl,i\theta_{l,i}.

    The forward activation at layer l+1l+1 given input feature map xlx_l is the mask-weighted sum of all candidate block outputs:

    xl+1=∑i=1Nml,i⋅bl,i(xl)x_{l+1} = \sum_{i=1}^N m_{l,i} \cdot b_{l,i}(x_l)

  2. Knowl 2 — Hardware-Aware Differentiable Latency Loss Function and Lookup Model

    equation

    The multi-objective loss function balances task prediction accuracy with physical device execution latency:

    L(a,w)=CE(a,w)⋅α[log⁡(LAT(a))]β\mathcal{L}(a, w) = \text{CE}(a, w) \cdot \alpha \left[ \log\left( \text{LAT}(a) \right) \right]^\beta

    where CE(a,w)\text{CE}(a, w) denotes the cross-entropy classification loss of architecture aa with weights ww, LAT(a)\text{LAT}(a) is the total network latency on the target hardware measured in microseconds (μs\mu\text{s}), α>0\alpha > 0 is a scaling coefficient controlling the overall loss magnitude (set to 0.20.2), and β>0\beta > 0 is an exponent modulating the latency penalty (set to 0.60.6).

    Under an additive lookup table (LUT) runtime model, total network latency is estimated by summing the pre-measured runtimes of its constituent layer blocks:

    LAT(a)=∑l=1LLAT(bl(a))\text{LAT}(a) = \sum_{l=1}^L \text{LAT}\left(b_l^{(a)}\right)

    Under the continuous Gumbel-Softmax relaxation, the latency term becomes:

    LAT(a)=∑l=1L∑i=1Nml,i⋅LAT(bl,i)\text{LAT}(a) = \sum_{l=1}^L \sum_{i=1}^N m_{l,i} \cdot \text{LAT}(b_{l,i})

    where LAT(bl,i)\text{LAT}(b_{l,i}) is the measured runtime of candidate block ii at layer ll on the target device (a constant coefficient stored in the lookup table), and ml,im_{l,i} is the Gumbel-Softmax continuous mask. Because LAT(bl,i)\text{LAT}(b_{l,i}) is constant, LAT(a)\text{LAT}(a) is directly differentiable with respect to ml,im_{l,i} and θl,i\theta_{l,i}, enabling joint gradient-based optimization of task loss and hardware runtime.

  3. Knowl 3 — Layer-Wise Macro-Architecture and Candidate Block Search Space

    model/method

    The architecture search space consists of a fixed macro-architecture template containing 22 searchable layers (TBS: "To Be Searched"), yielding a discrete combinatorial space of 922≈9.85×10209^{22} \approx 9.85 \times 10^{20} possible networks.

    The macro-architecture is organized into sequential stages with fixed output channel dimensions ff, number of blocks nn, and stage strides ss:

    • Stage 0: Input 224×224×3224 \times 224 \times 3, fixed 3×33\times 3 standard convolution (f=16,n=1,s=2f=16, n=1, s=2).
    • Stage 1: 1 TBS block (f=16,n=1,s=1f=16, n=1, s=1).
    • Stage 2: 4 TBS blocks (f=24,n=4,s=2f=24, n=4, s=2 for the first block, s=1s=1 thereafter).
    • Stage 3: 4 TBS blocks (f=32,n=4,s=2f=32, n=4, s=2 for the first block, s=1s=1 thereafter).
    • Stage 4: 4 TBS blocks (f=64,n=4,s=2f=64, n=4, s=2 for the first block, s=1s=1 thereafter).
    • Stage 5: 4 TBS blocks (f=112,n=4,s=1f=112, n=4, s=1).
    • Stage 6: 4 TBS blocks (f=184,n=4,s=2f=184, n=4, s=2 for the first block, s=1s=1 thereafter).
    • Stage 7: 1 TBS block (f=352,n=1,s=1f=352, n=1, s=1).
    • Head: Fixed 1×11\times 1 convolution (f=1504f=1504 for FBNet-A, f=1984f=1984 for FBNet-B and FBNet-C, n=1,s=1n=1, s=1), followed by 7×77\times 7 global average pooling and a fully connected classification layer (1000 outputs).

    Each searchable layer selects from 9 candidate block options:

    1. k3_e1: Kernel size K=3K=3, expansion factor e=1e=1, group count g=1g=1.
    2. k3_e1_g2: Kernel size K=3K=3, expansion factor e=1e=1, group count g=2g=2 (with channel shuffle).
    3. k3_e3: Kernel size K=3K=3, expansion factor e=3e=3, group count g=1g=1.
    4. k3_e6: Kernel size K=3K=3, expansion factor e=6e=6, group count g=1g=1.
    5. k5_e1: Kernel size K=5K=5, expansion factor e=1e=1, group count g=1g=1.
    6. k5_e1_g2: Kernel size K=5K=5, expansion factor e=1e=1, group count g=2g=2 (with channel shuffle).
    7. k5_e3: Kernel size K=5K=5, expansion factor e=3e=3, group count g=1g=1.
    8. k5_e6: Kernel size K=5K=5, expansion factor e=6e=6, group count g=1g=1.
    9. skip: Identity shortcut forwarding the input without computation (allowing dynamic depth reduction).

    Every non-skip candidate block executes a 1×11\times 1 pointwise (or group) convolution with ReLU expanding channels by factor ee, followed by a K×KK\times K depthwise convolution with ReLU, followed by a 1×11\times 1 projection (or group) convolution without activation. If input and output spatial and channel dimensions match, an identity residual shortcut is added.

  4. Knowl 4 — DNAS Stochastic Supernet Optimization Algorithm

    algorithm

    The DNAS procedure optimizes the continuous architecture distribution over the supernet before sampling discrete networks.

    Input: Training dataset DD, validation proxy dataset DvalD_{\text{val}}, operator latency lookup table LUT\text{LUT}, number of search epochs E=90E = 90, warmup delay Edelay=10E_{\text{delay}} = 10, initial temperature τ0=5.0\tau_0 = 5.0, temperature decay rate γ=exp⁡(−0.045)≈0.956\gamma = \exp(-0.045) \approx 0.956, loss coefficients α=0.2\alpha = 0.2, β=0.6\beta = 0.6.
    Output: Sampled optimal discrete architectures {a∗}\{a^*\}.
    Initialize supernet operator weights ww randomly.
    Initialize architecture distribution parameters θ=0\theta = \mathbf{0}.
    Set current temperature τ←τ0\tau \leftarrow \tau_0.
    for epoch e=1e = 1 to EE do
        for each minibatch (xw,yw)(x_w, y_w) in training partition (80%80\% of DD) do
            Sample Gumbel noise gl,i∼Gumbel(0,1)g_{l,i} \sim \text{Gumbel}(0, 1) for all layers ll and blocks ii.
            Compute relaxation masks ml,i←GumbelSoftmax(θl,i∣θl,τ)m_{l,i} \leftarrow \text{GumbelSoftmax}(\theta_{l,i} \mid \theta_l, \tau).
            Compute forward activations and latency LAT(a)←∑l,iml,i⋅LUT[bl,i]\text{LAT}(a) \leftarrow \sum_{l,i} m_{l,i} \cdot \text{LUT}[b_{l,i}].
            Compute loss L←CE(xw,yw,w)⋅α[log⁡(LAT(a))]β\mathcal{L} \leftarrow \text{CE}(x_w, y_w, w) \cdot \alpha [\log(\text{LAT}(a))]^\beta.
            Update operator weights w←w−ηw∇wLw \leftarrow w - \eta_w \nabla_w \mathcal{L} via SGD with momentum.
        end for
        if epoch e>Edelaye > E_{\text{delay}} then
            for each minibatch (xθ,yθ)(x_\theta, y_\theta) in architecture partition (20%20\% of DD) do
                Sample Gumbel noise gl,i∼Gumbel(0,1)g_{l,i} \sim \text{Gumbel}(0, 1) for all layers ll and blocks ii.
                Compute relaxation masks ml,i←GumbelSoftmax(θl,i∣θl,τ)m_{l,i} \leftarrow \text{GumbelSoftmax}(\theta_{l,i} \mid \theta_l, \tau).
                Compute loss L←CE(xθ,yθ,w)⋅α[log⁡(LAT(a))]β\mathcal{L} \leftarrow \text{CE}(x_\theta, y_\theta, w) \cdot \alpha [\log(\text{LAT}(a))]^\beta.
                Update architecture parameters θ←θ−ηθ∇θL\theta \leftarrow \theta - \eta_\theta \nabla_\theta \mathcal{L} via Adam optimizer.
            end for
        end if
        Update temperature τ←τ⋅γ\tau \leftarrow \tau \cdot \gamma.
    end for
    Sample discrete architectures a∗∼Pθ(a)a^* \sim P_\theta(a) by drawing block choices according to Pθl(bl=bl,i)=softmax(θl,i)P_{\theta_l}(b_l = b_{l,i}) = \text{softmax}(\theta_{l,i}).
    return Sampled architectures {a∗}\{a^*\}.
  5. Knowl 5 — Experimental Protocol for DNAS Search and Model Training

    experimental setup

    The experimental evaluation proceeds in two stages: architecture search via stochastic supernet optimization, followed by training sampled candidate architectures from scratch on ImageNet 2012.

    • Supernet Search Stage:

      • Dataset: A proxy dataset comprising 100 randomly sampled classes from ImageNet-1k.
      • Data Split: 80%80\% of the proxy training set is used to train operator weights ww; 20%20\% is reserved for updating architecture parameters θ\theta.
      • Epochs and Batch Size: 90 epochs with a total batch size of 192.
      • Weight Optimizer: SGD with momentum of 0.9, weight decay of 10−410^{-4}, and initial learning rate of 0.1 following a cosine annealing schedule.
      • Architecture Optimizer: Adam optimizer with learning rate 10−210^{-2} and weight decay 5×10−45 \times 10^{-4}. Architecture parameter updates are delayed for the first 10 epochs (Edelay=10E_{\text{delay}} = 10) to warm up operator weights.
      • Gumbel-Softmax Temperature: Initial temperature τ=5.0\tau = 5.0, exponentially decayed by a factor of exp⁡(−0.045)≈0.956\exp(-0.045) \approx 0.956 every epoch.
      • Loss Parameters: α=0.2\alpha = 0.2 and β=0.6\beta = 0.6.
      • Compute Budget: 216 GPU-hours total (8 GPUs for 27 hours).
    • Final Training From Scratch Stage:

      • Models: 6 architectures sampled from the final distribution PθP_\theta.
      • Dataset: Full 1,000-class ImageNet 2012 (224 input resolution, 1.0 channel scaling for FBNet-{A, B, C}).
      • Epochs and Batch Size: 360 epochs with a batch size of 256 distributed across 8 GPUs.
      • Optimizer: SGD with initial learning rate 0.1, decayed by 10×10\times at epochs 90, 180, and 270; momentum 0.9; weight decay 4×10−54 \times 10^{-5}.
      • Regularization and Augmentation: Dropout with ratio 0.2 before the final fully connected layer; standard GoogleNet inception data augmentation.
      • Inference Deployment: Caffe2 with INT8 mobile inference engine on a Samsung Galaxy S8 (Snapdragon 835 platform) and an iPhone X (A11 Bionic processor).
  6. Knowl 6 — ImageNet Classification Benchmark: FBNets vs. Manual and NAS Baselines

    data/table

    The discovered FBNet architectures (FBNet-A, FBNet-B, FBNet-C) outperform both manually designed mobile networks (MobileNetV2, ShuffleNetV2, CondenseNet) and automated NAS models (DARTS, MnasNet, NASNet-A, PNASNet) in ImageNet top-1 accuracy, FLOP efficiency, and measured CPU latency on a Samsung Galaxy S8, while reducing search cost by orders of magnitude.

    Model Search Method Search Space Search Cost (GPU-h / rel.) #Params #FLOPs S8 CPU Latency Top-1 Acc (%)
    1.0-MobileNetV2 manual - - 3.4M 300M 21.7 ms 72.0
    1.5-ShuffleNetV2 manual - - 3.5M 299M 22.0 ms 72.6
    CondenseNet (G=C=8) manual - - 2.9M 274M 28.4 ms 71.0
    MnasNet-65 RL stage-wise 91K / 421x 3.6M 270M - 73.0
    DARTS gradient cell 288 / 1.33x 4.9M 595M - 73.1
    FBNet-A (ours) gradient layer-wise 216 / 1.0x 4.3M 249M 19.8 ms 73.0
    1.3-MobileNetV2 manual - - 5.3M 509M 33.8 ms 74.4
    CondenseNet (G=C=4) manual - - 4.8M 529M 28.7 ms 73.8
    MnasNet RL stage-wise 91K / 421x 4.2M 317M 23.7 ms 74.0
    NASNet-A RL cell 48K / 222x 5.3M 564M - 74.0
    PNASNet SMBO cell 6K / 27.8x 5.1M 588M - 74.2
    FBNet-B (ours) gradient layer-wise 216 / 1.0x 4.5M 295M 23.1 ms 74.1
    1.4-MobileNetV2 manual - - 6.9M 585M 37.4 ms 74.7
    2.0-ShuffleNetV2 manual - - 7.4M 591M 33.3 ms 74.9
    MnasNet-92 RL stage-wise 91K / 421x 4.4M 388M - 74.8
    FBNet-C (ours) gradient layer-wise 216 / 1.0x 5.5M 375M 28.1 ms 74.9

    Key highlights:

    • FBNet-A achieves 73.0%73.0\% accuracy with 249M FLOPs and 19.8 ms latency, operating 1.9 ms faster than 1.0-MobileNetV2 and 2.2 ms faster than 1.5-ShuffleNetV2.
    • FBNet-B achieves 74.1%74.1\% accuracy with 295M FLOPs and 23.1 ms latency, running 1.46×1.46\times faster with 1.73×1.73\times fewer FLOPs than 1.3-MobileNetV2, while matching MnasNet accuracy at 421×421\times lower search cost (216 vs. estimated 91,000 GPU-hours).
    • FBNet-C matches 2.0-ShuffleNetV2 at 74.9%74.9\% accuracy while executing 1.19×1.19\times faster (28.1 ms vs. 33.3 ms) with 1.58×1.58\times fewer FLOPs (375M vs. 591M).
  7. Knowl 7 — Architecture Adaptation to Input Resolution and Channel Scaling

    data/table

    Searching for architecture configurations tailored directly to specific input resolutions and channel width multipliers yields consistent top-1 accuracy gains of +1.5%+1.5\% to +6.4%+6.4\% over uniformly scaled MobileNetV2 baselines under matching runtime budgets.

    Input Size Channel Scaling Model #Params #FLOPs CPU Latency Top-1 Acc (%)
    (224, 0.35) MobileNetV2-224-0.35 1.7M 59M 9.3 ms 60.3
    MnasNet-scale-224-0.35 1.9M 76M 10.7 ms 62.4 (+2.1)
    FBNet-224-0.35 (ours) 2.0M 72M 10.7 ms 65.3 (+5.0)
    (192, 0.50) MobileNetV2 2.0M 71M 8.4 ms 63.9
    MnasNet-search-192-0.5 - - - 65.6 (+1.7)
    FBNet-192-0.5 (ours) 2.6M 73M 9.9 ms 65.9 (+2.0)
    (128, 1.0) MobileNetV2 3.5M 99M 8.4 ms 65.3
    MnasNet-scale-128-1.0 4.2M 103M 9.2 ms 67.3 (+2.0)
    FBNet-128-1.0 (ours) 4.2M 92M 9.0 ms 67.0 (+1.7)
    (128, 0.50) MobileNetV2 2.0M 32M 4.8 ms 57.7
    FBNet-128-0.5 (ours) 2.4M 32M 5.1 ms 60.0 (+2.3)
    (96, 0.35) MobileNetV2 1.7M 11M 3.8 ms 45.5
    FBNet-96-0.35-1 (ours) 1.8M 12.9M 2.9 ms 50.2 (+4.7)
    FBNet-96-0.35-2 (ours) 1.9M 13.7M 3.6 ms 51.9 (+6.4)

    Notable findings:

    • At (224,0.35)(224, 0.35), FBNet-224-0.35 achieves 65.3%65.3\% accuracy, outperforming MobileNetV2 by +5.0%+5.0\% and MnasNet-scale by +2.1%+2.1\%.
    • At (96,0.35)(96, 0.35), FBNet-96-0.35-1 achieves 50.2%50.2\% accuracy with 2.9 ms latency (345 frames per second on a Samsung Galaxy S8), running faster and achieving +4.7%+4.7\% higher accuracy than MobileNetV2 (45.5%45.5\% at 3.8 ms).
    • FBNet-96-0.35-2 pushes top-1 accuracy to 51.9%51.9\% (+6.4%+6.4\% over MobileNetV2) at 3.6 ms latency.
  8. Knowl 8 — Device-Specific Architecture Optimization on Mobile Hardware

    data/table

    Because operator execution efficiency varies across hardware platforms, ConvNets optimized for a specific processor do not transfer optimally to another. Searching specifically for target device runtimes yields substantial hardware-specific speedups.

    Model #Params #FLOPs Latency on iPhone X Latency on Samsung S8 Top-1 Acc (%)
    FBNet-iPhoneX 4.47M 322M 19.84 ms (target) 23.33 ms 73.20
    FBNet-S8 4.43M 293M 27.53 ms 22.12 ms (target) 73.27

    Cross-device evaluation shows:

    • FBNet-iPhoneX achieves 19.84 ms on iPhone X, but slows down to 23.33 ms when deployed on Samsung Galaxy S8.
    • FBNet-S8 achieves 22.12 ms on Samsung Galaxy S8, but its latency increases to 27.53 ms on iPhone X—a 7.69 ms (relative 39%39\%) latency penalty compared to FBNet-iPhoneX.
    • Over the Samsung-optimized model, the iPhone-X-optimized model attains a 1.39×1.39\times (1.4×1.4\times) speedup on an iPhone X at equivalent ImageNet accuracy (73.20%73.20\% vs. 73.27%73.27\%).
  9. Knowl 9 — Hardware Operator Preferences and Resolution-Induced Structural Adaptation

    empirical result

    DNAS exhibits distinct architectural specialization patterns based on target hardware and input resolution:

    1. Hardware-Specific Operator Preference:

      • On the Samsung Galaxy S8 (Snapdragon 835 platform), 5×55\times 5 depthwise convolutions with large channel dimensions (such as spatial/channel configurations 28×28×19228\times 28 \times 192, 14×14×19214\times 14 \times 192, and 14×14×33614\times 14 \times 336) execute significantly faster than on iPhone X. Consequently, FBNet-S8 utilizes 5×55\times 5 depthwise convolutions extensively throughout its early, middle, and late layers.
      • On the iPhone X (A11 Bionic processor), 3×33\times 3 depthwise convolutions are substantially more efficient than 5×55\times 5 depthwise convolutions. As a result, FBNet-iPhoneX adopts 3×33\times 3 depthwise convolutions almost exclusively, restricting 5×55\times 5 operations strictly to the final two stages.
    2. Resolution-Dependent Network Depth:

      • When searching under reduced input resolutions (e.g., 96×9696\times 96 in FBNet-96-0.35-1), DNAS heavily selects the skip block across multiple intermediate layers, producing a markedly shallower network compared to 224×224224\times 224 models (FBNet-{A, B, C}).
      • This structural adaptation aligns with receptive field mechanics: smaller input images require smaller receptive fields to cover the feature space, making deep redundant layers suboptimal for accuracy and efficiency.
  10. Knowl 10 — Latency Additivity Assumption for Mobile Processors

    assumption

    The DNAS latency lookup table (LUT) model relies on the assumption of operator runtime additivity:

    LAT(a)=∑l=1LLAT(bl(a))\text{LAT}(a) = \sum_{l=1}^L \text{LAT}\left(b_l^{(a)}\right)

    This formulation assumes that the execution runtime of an entire neural network on target embedded processors (such as mobile CPUs and DSPs) equals the independent sum of the runtimes of each constituent layer block bl(a)b_l^{(a)}.

    This assumption holds in runtime environments (such as Caffe2 with INT8 inference) where operators execute sequentially one by one without asynchronous pipeline concurrency or inter-layer operator fusion across block boundaries. Under this condition, pre-measuring a few hundred individual operator latencies allows exact linear estimation of all 102110^{21} candidate networks in the search space.

Coverage note — None was omitted; all key contributions including DNAS formulation, latency modeling, search space, supernet optimization algorithm, ImageNet benchmarks, resolution/channel scaling, cross-hardware specialization, and structural analyses are fully covered.

References

  1. 1.Anonymous. Snas: stochastic neural architecture search. In Submitted to International Conference on Learning Representations, 2019. under review.
  2. 2.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pages 248–255. Ieee, 2009.
  3. 3.A. Gholami, K. Kwon, B. Wu, Z. Tai, X. Yue, P. Jin, S. Zhao, and K. Keutzer. Squeezenext: Hardware-aware neural network design. arXiv preprint arXiv:1803.10615, 2018.
  4. 4.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  5. 5.Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han. Amc: Automl for model compression and acceleration on mobile devices. In Proceedings of the European Conference on Computer Vision (ECCV), pages 784–800, 2018.
  6. 6.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  7. 7.G. Huang, S. Liu, L. van der Maaten, and K. Q. Weinberger. Condensenet: An efficient densenet using learned group convolutions. group, 3(12):11, 2017.
  8. 8.F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size. arXiv preprint arXiv:1602.07360, 2016.
  9. 9.E. Jang, S. Gu, and B. Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016.
  10. 10.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  11. 11.C. Liu, B. Zoph, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy. Progressive neural architecture search. arXiv preprint arXiv:1712.00559, 2017.
  12. 12.H. Liu, K. Simonyan, and Y. Yang. Darts: Differentiable architecture search. arXiv preprint arXiv:1806.09055, 2018.
  13. 13.N. Ma, X. Zhang, H.-T. Zheng, and J. Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. arXiv preprint arXiv:1807.11164, 2018.
  14. 14.C. J. Maddison, A. Mnih, and Y. W. Teh. The concrete distribution: A continuous relaxation of discrete random variables. arXiv preprint arXiv:1611.00712, 2016.
  15. 15.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. 2017.
  16. 16.H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, and J. Dean. Efficient neural architecture search via parameter sharing. arXiv preprint arXiv:1802.03268, 2018.
  17. 17.M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4510–4520, 2018.
  18. 18.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  19. 19.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.
  20. 20.M. Tan, B. Chen, R. Pang, V. Vasudevan, and Q. V. Le. Mnasnet: Platform-aware neural architecture search for mobile. arXiv preprint arXiv:1807.11626, 2018.
  21. 21.T. Veniat and L. Denoyer. Learning time/memory-efficient deep architectures with budgeted super networks. arXiv preprint arXiv:1706.00046, 2017.
  22. 22.B. Wu, F. N. Iandola, P. H. Jin, and K. Keutzer. Squeezedet: Unified, small, low power fully convolutional neural networks for real-time object detection for autonomous driving. In CVPR Workshops, pages 446–454, 2017.
  23. 23.B. Wu, A. Wan, X. Yue, P. Jin, S. Zhao, N. Golmant, A. Gholaminejad, J. Gonzalez, and K. Keutzer. Shift: A zero flop, zero parameter alternative to spatial convolutions. arXiv preprint arXiv:1711.08141, 2017.
  24. 24.B. Wu, A. Wan, X. Yue, and K. Keutzer. Squeezeseg: Convolutional neural nets with recurrent crf for real-time road-object segmentation from 3d lidar point cloud. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 1887–1893. IEEE, 2018.
  25. 25.B. Wu, Y. Wang, P. Zhang, Y. Tian, P. Vajda, and K. Keutzer. Mixed precision quantization of convnets via differentiable neural architecture search. arXiv preprint arXiv:1812.00090, 2018.
  26. 26.B. Wu, X. Zhou, S. Zhao, X. Yue, and K. Keutzer. Squeezesegv2: Improved model structure and unsupervised domain adaptation for road-object segmentation from a lidar point cloud. arXiv preprint arXiv:1809.08495, 2018.
  27. 27.T.-J. Yang, A. Howard, B. Chen, X. Zhang, A. Go, M. Sandler, V. Sze, and H. Adam. Netadapt: Platform-aware neural network adaptation for mobile applications. Energy, 41:46, 2018.
  28. 28.Y. Yang, Q. Huang, B. Wu, T. Zhang, L. Ma, G. Gambardella, M. Blott, L. Lavagno, K. Vissers, J. Wawrzynek, et al. Synetgy: Algorithm-hardware co-design for convnet accelerators on embedded fpgas. arXiv preprint arXiv:1811.08634, 2018.
  29. 29.X. Zhang, X. Zhou, M. Lin, and J. Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. arxiv 2017. arXiv preprint arXiv:1707.01083.
  30. 30.B. Zoph and Q. V. Le. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016.
  31. 31.B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le. Learning transferable architectures for scalable image recognition. arXiv preprint arXiv:1707.07012, 2(6), 2017.

Citation

MLA
Wu, B., et al. “FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search”. arXiv, 2018, http://arxiv.org/abs/1812.03443v3.
APA
Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., & Keutzer, K. (2018). FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search. arXiv. http://arxiv.org/abs/1812.03443v3
Chicago
Wu, B., X. Dai, P. Zhang, et al. 2018. “FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search”. arXiv. http://arxiv.org/abs/1812.03443v3.
Harvard
Wu, B. et al. (2018) “FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1812.03443v3.
Vancouver
1. Wu B, Dai X, Zhang P, Wang Y, Sun F, Wu Y, Tian Y, Vajda P, Jia Y, Keutzer K (2018) FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search. arXiv

BibTeX

@article{wu2018fbnet,
  title = {FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search},
  author = {Wu, Bichen and Dai, Xiaoliang and Zhang, Peizhao and Wang, Yanghan and Sun, Fei and Wu, Yiming and Tian, Yuandong and Vajda, Peter and Jia, Yangqing and Keutzer, Kurt},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1812.03443v3},
  eprint = {1812.03443}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE