Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks
Jierun ChenShiu-hong KaoHao HeWeipeng ZhuoSong WenChul-Ho LeeShueng-Han Gary Chan
Proposes Partial Convolution and the FasterNet architecture family, which minimize memory access overhead to achieve significantly higher inference speeds across GPUs, CPUs, and mobile processors without sacrificing vision accuracy.
Real-time computer vision applications require fast neural networks with high throughput and low latency. Current industry and research efforts primarily concentrate on lowering computational complexity by minimizing floating-point operations. However, this theoretical reduction often fails to produce real-world speedups on hardware because frequent memory access bottlenecks execution speed, resulting in low floating-point operations per second.
The article aims to resolve this performance mismatch by introducing a lightweight spatial feature extraction mechanism that decreases both total computations and memory access simultaneously. It evaluates a newly designed neural network family to demonstrate superior computational speed across various hardware platforms without compromising task accuracy.
The researchers developed Partial Convolution, an operator that processes only a subset of input channels through regular convolution while leaving the remaining channels untouched, followed by a pointwise convolution. Building on this core block, they constructed a family of architectures named FasterNet. They conducted extensive experimental evaluations using the ImageNet-1k dataset for image classification and the COCO dataset for object detection and instance segmentation, measuring running latency and throughput across diverse hardware including GPUs, CPUs, and mobile ARM processors.
The analysis produced several key findings. First, the proposed Partial Convolution achieved 10.5 times higher floating-point operations per second on GPUs, 6.2 times higher on CPUs, and 22.8 times higher on ARM processors compared to depthwise convolution. Second, on ImageNet classification, the compact FasterNet variant ran 2.8 times faster on GPUs, 3.3 times faster on CPUs, and 2.4 times faster on ARM processors than comparable lightweight models, while increasing top-1 accuracy by 2.9 percentage points. Third, the largest model achieved an 83.5 percent top-1 accuracy, matching premier vision transformers while delivering 36 percent higher throughput on GPUs and reducing CPU computation time by 37 percent. Finally, in object detection and instance segmentation on COCO, FasterNet consistently delivered higher precision scores with substantially lower latency than standard baselines.
These findings indicate that architectural designs should optimize effective computational execution speed alongside operational count. By eliminating memory access bottlenecks and simplifying operations, organizations deploying vision models can achieve lower compute costs, lower inference latency, and higher throughput on standard edge and server hardware without requiring specialized accelerators.
Teams designing and deploying production vision systems should consider adopting Partial Convolution and FasterNet architectures as drop-in replacements for standard lightweight backbones. When tuning these models, practitioners should maintain the default partial channel ratio of one-fourth to balance feature extraction and throughput, and use batch normalization merged into adjacent convolution layers during inference to preserve high execution speeds.
The primary constraints of this design stem from fixed unit-stride requirements during Partial Convolution, which require explicit downsampling layers to change feature map dimensions, as well as a purely convolutional structure that may have a more limited receptive field than transformer-based attention mechanisms. Even with these constraints, the evidence provides high confidence that FasterNet offers practical speedups and high accuracy across edge and cloud hardware.
- Paper: ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design, Ningning Ma et al. (2018). Read ShuffleNet V2 first to see the hardware-aware design rules—especially its warning that FLOPs alone do not predict speed—that motivate FasterNet’s focus on memory traffic and measured throughput.
- Paper: MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, Andrew G. Howard et al. (2017). MobileNets establishes depthwise separable convolution as a standard lightweight design, giving essential context for FasterNet’s comparisons showing that fewer operations need not mean faster execution.
- Paper: CSPNet: A New Backbone that can Enhance Learning Capability of CNN, Chien-Yao Wang et al. (2019). CSPNet’s split-and-merge feature processing provides useful architectural precedent for understanding FasterNet’s strategy of processing only part of the channels.
- Paper: MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler et al. (2018). MobileNetV2 develops the depthwise-convolution-based lightweight blocks that help clarify the design trade-offs FasterNet addresses with Partial Convolution.
- Paper: Resource-Efficient Neural Networks for Embedded Systems, Wolfgang Roth et al. (2024). This later survey broadens FasterNet’s hardware-efficiency perspective by comparing structural design, pruning, and quantization across embedded platforms.
