Built independently by an author, for readers. Read the story and support ChapterPal

keyword

ShuffleNet V2

ShuffleNet V2 is an efficient convolutional neural network architecture designed to achieve high computational accuracy and fast inference speeds on resource-constrained hardware such as mobile and edge devices. Unlike earlier lightweight models that focused primarily on minimizing theoretical floating-point operations, ShuffleNet V2 incorporates practical design principles that optimize real-world execution speed by minimizing memory access costs, avoiding excessive network fragmentation, and reducing element-wise operations. A defining feature of the architecture is the channel split operator, which divides feature channels into separate paths before applying depthwise convolutions and recombines them via concatenation and channel shuffling to ensure effective information flow across feature maps without incurring substantial computational overhead.

2 items

FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search

FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search

Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, Kurt Keutzer

OrganizationsMetaPrinceton UniversityUniversity of California Berkeley

Why you should read this

Introduces a hardware-aware differentiable neural architecture search framework that directly optimizes ConvNets for device-specific latency, cutting search costs by over 400 times compared to prior methods while achieving superior speed and accuracy on mobile devices.

Designing accurate and efficient ConvNets for mobile devices is challenging because the design space is combinatorially large. Due to this, previous neural architecture search (NAS) methods are computationally expensive. ConvNet architecture optimality depends on factors such as input resolution and target devices. However, existing approaches are too expensive for case-by-case redesigns. Also, previous work focuses primarily on reducing FLOPs, but FLOP count does not always reflect actual latency. To address these, we propose a differentiable neural architecture search (DNAS) framework that uses gradient-based methods to optimize ConvNet architectures, avoiding enumerating and training individual architectures separately as in previous methods. FBNets, a family of models discovered by DNAS surpass state-of-the-art models both designed manually and generated automatically. FBNet-B achieves 74.1% top-1 accuracy on ImageNet with 295M FLOPs and 23.1 ms latency on a Samsung S8 phone, 2.4x smaller and 1.5x faster than MobileNetV2-1.3 with similar accuracy. Despite higher accuracy and lower latency than MnasNet, we estimate FBNet-B's search cost is 420x smaller than MnasNet's, at only 216 GPU-hours. Searched for different resolutions and channel sizes, FBNets achieve 1.5% to 6.4% higher accuracy than MobileNetV2. The smallest FBNet achieves 50.2% accuracy and 2.9 ms latency (345 frames per second) on a Samsung S8. Over a Samsung-optimized FBNet, the iPhone-X-optimized model achieves a 1.4x speedup on an iPhone X.

Added

2026-09-25