Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks

Jierun ChenShiu-hong KaoHao HeWeipeng ZhuoSong WenChul-Ho LeeShueng-Han Gary Chan

article2023CVPR2,036 citations

Proposes Partial Convolution and the FasterNet architecture family, which minimize memory access overhead to achieve significantly higher inference speeds across GPUs, CPUs, and mobile processors without sacrificing vision accuracy.

Listen

Real-time computer vision applications require fast neural networks with high throughput and low latency. Current industry and research efforts primarily concentrate on lowering computational complexity by minimizing floating-point operations. However, this theoretical reduction often fails to produce real-world speedups on hardware because frequent memory access bottlenecks execution speed, resulting in low floating-point operations per second.

The article aims to resolve this performance mismatch by introducing a lightweight spatial feature extraction mechanism that decreases both total computations and memory access simultaneously. It evaluates a newly designed neural network family to demonstrate superior computational speed across various hardware platforms without compromising task accuracy.

The researchers developed Partial Convolution, an operator that processes only a subset of input channels through regular convolution while leaving the remaining channels untouched, followed by a pointwise convolution. Building on this core block, they constructed a family of architectures named FasterNet. They conducted extensive experimental evaluations using the ImageNet-1k dataset for image classification and the COCO dataset for object detection and instance segmentation, measuring running latency and throughput across diverse hardware including GPUs, CPUs, and mobile ARM processors.

The analysis produced several key findings. First, the proposed Partial Convolution achieved 10.5 times higher floating-point operations per second on GPUs, 6.2 times higher on CPUs, and 22.8 times higher on ARM processors compared to depthwise convolution. Second, on ImageNet classification, the compact FasterNet variant ran 2.8 times faster on GPUs, 3.3 times faster on CPUs, and 2.4 times faster on ARM processors than comparable lightweight models, while increasing top-1 accuracy by 2.9 percentage points. Third, the largest model achieved an 83.5 percent top-1 accuracy, matching premier vision transformers while delivering 36 percent higher throughput on GPUs and reducing CPU computation time by 37 percent. Finally, in object detection and instance segmentation on COCO, FasterNet consistently delivered higher precision scores with substantially lower latency than standard baselines.

These findings indicate that architectural designs should optimize effective computational execution speed alongside operational count. By eliminating memory access bottlenecks and simplifying operations, organizations deploying vision models can achieve lower compute costs, lower inference latency, and higher throughput on standard edge and server hardware without requiring specialized accelerators.

Teams designing and deploying production vision systems should consider adopting Partial Convolution and FasterNet architectures as drop-in replacements for standard lightweight backbones. When tuning these models, practitioners should maintain the default partial channel ratio of one-fourth to balance feature extraction and throughput, and use batch normalization merged into adjacent convolution layers during inference to preserve high execution speeds.

The primary constraints of this design stem from fixed unit-stride requirements during Partial Convolution, which require explicit downsampling layers to change feature map dimensions, as well as a purely convolutional structure that may have a more limited receptive field than transformer-based attention mechanisms. Even with these constraints, the evidence provides high confidence that FasterNet offers practical speedups and high accuracy across edge and cloud hardware.

Cover for Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks

Abstract

To design fast neural networks, many works have been focusing on reducing the number of floating-point operations (FLOPs). We observe that such reduction in FLOPs, however, does not necessarily lead to a similar level of reduction in latency. This mainly stems from inefficiently low floating-point operations per second (FLOPS). To achieve faster networks, we revisit popular operators and demonstrate that such low FLOPS is mainly due to frequent memory access of the operators, especially the depthwise convolution. We hence propose a novel partial convolution (PConv) that extracts spatial features more efficiently, by cutting down redundant computation and memory access simultaneously. Building upon our PConv, we further propose FasterNet, a new family of neural networks, which attains substantially higher running speed than others on a wide range of devices, without compromising on accuracy for various vision tasks. For example, on ImageNet-1k, our tiny FasterNet-T0 is 2.8×2.8\times, 3.3×3.3\times, and 2.4×2.4\times faster than MobileViT-XXS on GPU, CPU, and ARM processors, respectively, while being 2.9%2.9\% more accurate. Our large FasterNet-L achieves impressive 83.5%83.5\% top-1 accuracy, on par with the emerging Swin-B, while having 36%36\% higher inference throughput on GPU, as well as saving 37%37\% compute time on CPU. Code is available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Design of PConv and FasterNet
  • 3.1 Preliminary
  • 3.2 Partial convolution as a basic operator
  • 3.3 PConv followed by PWConv
  • 3.4 FasterNet as a general backbone
  • 4 Experimental Results
  • 4.1 PConv is fast with high FLOPS
  • 4.2 PConv is effective together with PWConv
  • 4.3 FasterNet on ImageNet-1k classification
  • 4.4 FasterNet on downstream tasks
  • 4.5 Ablation study
  • 5 Conclusion
  • A ImageNet-1k experimental settings
  • B Downstream tasks experimental settings
  • C Full comparison plots on ImageNet-1k
  • D Detailed architectural configurations
  • E More comparisons with related work
  • F Limitations and future work
  • References

Knowls

  1. Knowl 1 — Partial Convolution (PConv)

    model/method

    Partial Convolution (PConv) is a spatial feature extraction operator designed to lower FLOPs and reduce memory access simultaneously by exploiting inter-channel redundancy in feature maps.

    Given an input feature map I∈Rc×h×wI \in \mathbb{R}^{c \times h \times w} with cc channels, height hh, and width ww, PConv applies a standard 2D convolution with kernel size k×kk \times k to only a contiguous subset of cpc_p channels (typically cp=r⋅cc_p = r \cdot c with partial ratio r=1/4r = 1/4), while leaving the remaining c−cpc - c_p channels untouched (identity mapping). The output O∈Rc×h×wO \in \mathbb{R}^{c \times h \times w} retains the original channel and spatial dimensions.

    For equal input and output channel counts, the computational cost of PConv is: FLOPs=h×w×k2×cp2\text{FLOPs} = h \times w \times k^2 \times c_p^2

    For r=1/4r = 1/4, this equals 116\frac{1}{16} of the FLOPs of a regular convolution (h×w×k2×c2h \times w \times k^2 \times c^2).

    The memory access cost is: Memory Access=h×w×2cp+k2×cp2≈h×w×2cp\text{Memory Access} = h \times w \times 2c_p + k^2 \times c_p^2 \approx h \times w \times 2c_p

    For r=1/4r = 1/4, the memory access is approximately 14\frac{1}{4} of that of a regular convolution (h×w×2c+k2×c2≈h×w×2ch \times w \times 2c + k^2 \times c^2 \approx h \times w \times 2c). The remaining c−cpc - c_p channels are preserved without removal so that subsequent pointwise convolution layers can mix spatial and channel information across all cc channels.

  2. Knowl 2 — Latency-FLOPs-FLOPS Relationship and DWConv Memory Bottleneck

    definition

    Network inference latency on a computing device is governed by the ratio of total operations to operational throughput: Latency=FLOPsFLOPS\text{Latency} = \frac{\text{FLOPs}}{\text{FLOPS}} where FLOPs\text{FLOPs} represents the total count of multiply-accumulate floating-point operations executed by the network, and FLOPS\text{FLOPS} denotes floating-point operations per second (the effective computational speed on the device).

    Depthwise convolution (DWConv) applies an independent k×kk \times k spatial filter per channel with low theoretical FLOPs (h×w×k2×ch \times w \times k^2 \times c). Because DWConv alone reduces representation capacity, neural networks typically expand the channel width to c′>cc' > c (e.g., c′=6cc' = 6c in inverted residual blocks). This causes memory access to escalate to: Memory AccessDWConv=h×w×2c′+k2×c′≈h×w×2c′\text{Memory Access}_{\text{DWConv}} = h \times w \times 2c' + k^2 \times c' \approx h \times w \times 2c' which exceeds the memory access of a regular convolution (h×w×2c+k2×c2≈h×w×2ch \times w \times 2c + k^2 \times c^2 \approx h \times w \times 2c). Due to excessive memory I/O operations per arithmetic operation (low arithmetic intensity), DWConv yields severely depressed FLOPS on GPUs, CPUs, and ARM processors, preventing theoretical FLOP reductions from translating into proportional latency savings.

  3. Knowl 3 — PConv Followed by Pointwise Convolution as T-Shaped Receptive Field

    model/method

    When a Partial Convolution (PConv) with kernel size k×kk \times k on cpc_p channels is followed by a Pointwise Convolution (PWConv, 1×11 \times 1 convolution) over all cc channels, the combined effective receptive field forms a T-shaped pattern that allocates more computation to the center position of the patch rather than treating all spatial positions uniformly.

    Evaluating the position-wise Frobenius norm of regular convolutional filters F∈Rk2×cF \in \mathbb{R}^{k^2 \times c}: ∥Fi∥=∑j=1c∣fij∣2,i∈{1,2,…,k2}\|F_i\| = \sqrt{\sum_{j=1}^{c} |f_{ij}|^2}, \quad i \in \{1, 2, \dots, k^2\} reveals that the center position (i=5i=5 for 3×33 \times 3) is the salient position with maximum Frobenius norm most frequently across network stages.

    A single-layer T-shaped convolution incurs a computational cost of: FLOPsT-shaped=h×w×(k2×cp×c+c×(c−cp))\text{FLOPs}_{\text{T-shaped}} = h \times w \times \left(k^2 \times c_p \times c + c \times (c - c_p)\right)

    Decomposing the operation into a sequential two-step combination of PConv followed by PWConv reduces the computational cost to: FLOPsPConv+PWConv=h×w×(k2×cp2+c2)\text{FLOPs}_{\text{PConv+PWConv}} = h \times w \times \left(k^2 \times c_p^2 + c^2\right)

    This decomposition saves FLOPs because (k2−1)c>k2cp(k^2 - 1)c > k^2 c_p holds when cp=c/4c_p = c/4 and k=3k = 3, while directly leveraging optimized regular convolution implementations.

  4. Knowl 4 — FasterNet Architecture and Block Design

    model/method

    FasterNet is a hierarchical general-purpose neural network backbone consisting of four stages.

    Hierarchical Stages and Downsampling:

    • Stage 1 is preceded by an embedding layer consisting of a regular 4×44 \times 4 convolution with stride 4 and Batch Normalization (BN).
    • Stages 2, 3, and 4 are each preceded by a merging layer consisting of a regular 2×22 \times 2 convolution with stride 2 and BN, which downsamples spatial resolution by a factor of 2 and doubles the channel count.
    • Because blocks in Stages 3 and 4 have higher computational intensity (FLOPS) and consume relatively less memory access per operation, the majority of FasterNet blocks are allocated to Stages 3 and 4.

    FasterNet Block Structure: Each block follows an inverted residual structure with a shortcut connection:

    1. A 3×33 \times 3 PConv layer with stride 1 and partial ratio r=1/4r = 1/4 for spatial feature extraction.
    2. A 1×11 \times 1 PWConv layer expanding the channel dimension from cc to 2c2c.
    3. A Batch Normalization (BN) layer and an activation function (GELU for small variants, ReLU for large variants), placed exclusively after the middle PWConv layer to preserve feature diversity and minimize inference latency (BN is fused into adjacent convolution weights at inference).
    4. A second 1×11 \times 1 PWConv layer reducing the channel dimension from 2c2c back to cc.
    5. A residual shortcut connection adding the input tensor to the output of the second PWConv.

    Classification Head: The output of Stage 4 undergoes Global Average Pooling, followed by a 1×11 \times 1 convolution to 1280 channels, an activation layer, and a linear fully connected layer to output class logits.

  5. Knowl 5 — FasterNet Model Variant Configurations

    data/table

    The FasterNet family contains six standard configurations (T0, T1, T2, S, M, L) scaling across depth (number of blocks [b1,b2,b3,b4][b_1, b_2, b_3, b_4] per stage) and width (base channels cc):

    Specification FasterNet-T0 FasterNet-T1 FasterNet-T2 FasterNet-S FasterNet-M FasterNet-L
    Stage 1 channels (cc) 40 64 96 128 144 192
    Stage 2 channels (2c2c) 80 128 192 256 288 384
    Stage 3 channels (4c4c) 160 256 384 512 576 768
    Stage 4 channels (8c8c) 320 512 768 1024 1152 1536
    Blocks [b1,b2,b3,b4][b_1, b_2, b_3, b_4] [1,2,8,2][1, 2, 8, 2] [1,2,8,2][1, 2, 8, 2] [1,2,8,2][1, 2, 8, 2] [1,2,13,2][1, 2, 13, 2] [3,4,18,3][3, 4, 18, 3] [3,4,18,3][3, 4, 18, 3]
    Activation Function GELU GELU ReLU ReLU ReLU ReLU
    Parameters (M) 3.9 7.6 15.0 31.1 53.5 93.4
    FLOPs (G) at 224×224224 \times 224 0.34 0.85 1.90 4.55 8.72 15.49

    All variants employ 3×33 \times 3 PConv with partial ratio r=1/4r = 1/4 and stride 1. Small variants (T0, T1) use GELU, while larger variants (T2, S, M, L) use ReLU.

  6. Knowl 6 — On-Device FLOPS and Speed Comparison of Convolution Operators

    empirical result

    Ten consecutive layers of different convolutional operators (3×33 \times 3 kernel size) were benchmarked across standard feature map dimensions on GPU (NVIDIA RTX 2080Ti, throughput with batch size 32), CPU (Intel i9-9900X single thread, latency with batch size 1), and ARM (Cortex-A72 single thread, latency with batch size 1):

    Operator GPU (2080Ti) CPU (i9-9900X) ARM (Cortex-A72)
    Throughput (fps) FLOPS (G/s) Latency (ms) FLOPS (G/s) Latency (ms) FLOPS (G/s)
    Conv 3×33 \times 3 (Average) - 10151 - 71.89 - 3.95
    GConv 3×33 \times 3 (16 groups, Average) - 1833 - 25.93 - 1.92
    DWConv 3×33 \times 3 (Average) - 313 - 6.42 - 0.12
    PConv 3×33 \times 3 (r=1/4r=1/4, Average) - 3274 - 39.91 - 2.73

    PConv achieves 116\frac{1}{16} of the FLOPs of regular Conv while attaining 10.5×10.5\times, 6.2×6.2\times, and 22.8×22.8\times higher FLOPS than DWConv on GPU, CPU, and ARM, respectively. This demonstrates that PConv preserves high hardware computational efficiency while drastically reducing operation count.

  7. Knowl 7 — Approximation Error of Convolution Variants on Feature Maps

    empirical result

    To evaluate the capability of operator pairs to approximate standard spatial convolutions, intermediate feature representations before and after the first 3×33 \times 3 convolution across all four stages of a pre-trained ResNet50 were extracted using ImageNet-1k validation images. Models composed of different operator pairs were trained with mean squared error (MSE) loss on a 70% train / 10% validation / 20% test split:

    Stage DWConv + PWConv GConv + PWConv (16 groups) PConv + PWConv (r=1/4r=1/4)
    Stage 1 0.0089 0.0065 0.0069
    Stage 2 0.0158 0.0137 0.0136
    Stage 3 0.0214 0.0202 0.0172
    Stage 4 0.0130 0.0128 0.0115
    Average 0.0148 0.0133 0.0123

    PConv + PWConv achieves the lowest average test MSE loss (0.0123) across stages compared to DWConv + PWConv (0.0148) and GConv + PWConv (0.0133), confirming that extracting spatial features from only 1/41/4 of the channels and mixing across all channels via pointwise convolution effectively approximates a full spatial convolution.

  8. Knowl 8 — ImageNet-1k Classification Performance of FasterNet

    empirical result

    FasterNet models trained for 300 epochs on ImageNet-1k achieve state-of-the-art accuracy-latency and accuracy-throughput trade-offs on GPU (RTX 2080Ti, batch size 32), CPU (Intel i9-9900X single-thread, batch size 1), and ARM (Cortex-A72 single-thread, batch size 1):

    Model Params (M) FLOPs (G) GPU (fps) ↑\uparrow CPU (ms) ↓\downarrow ARM (ms) ↓\downarrow Top-1 Acc. (%)
    ShuffleNetV2 ×1.5\times 1.5 3.5 0.30 4878 12.1 266 72.6
    MobileNetV2 3.5 0.31 4198 12.2 442 72.0
    MobileViT-XXS 1.3 0.42 2393 30.8 348 69.0
    EdgeNeXt-XXS 1.3 0.26 2765 15.7 239 71.2
    FasterNet-T0 3.9 0.34 6807 9.2 143 71.9
    GhostNet ×1.3\times 1.3 7.4 0.24 2988 17.9 481 75.7
    ShuffleNetV2 ×2\times 2 7.4 0.59 3339 17.8 403 74.9
    MobileNetV2 ×1.4\times 1.4 6.1 0.60 2711 22.6 650 74.7
    MobileViT-XS 2.3 1.05 1392 40.8 648 74.8
    EdgeNeXt-XS 2.3 0.54 1738 24.4 434 75.0
    PVT-Tiny 13.2 1.94 1266 55.6 708 75.1
    FasterNet-T1 7.6 0.85 3782 17.7 285 76.2
    ResNet50 25.6 4.11 959 73.0 1131 78.8
    CycleMLP-B1 15.2 2.10 865 116.1 892 79.1
    PoolFormer-S12 11.9 1.82 1439 49.0 665 77.2
    FasterNet-T2 15.0 1.91 1991 33.5 497 78.9
    ConvNeXt-T 28.6 4.47 657 86.3 1889 82.1
    Swin-T 28.3 4.51 609 122.2 1424 81.3
    FasterNet-S 31.1 4.56 1029 71.2 1103 81.3
    ConvNeXt-S 50.2 8.71 377 153.2 3484 83.1
    Swin-S 49.6 8.77 348 224.2 2613 83.0
    FasterNet-M 53.5 8.74 500 129.5 2092 83.0
    ConvNeXt-B 88.6 15.38 253 257.1 OOM 83.8
    Swin-B 87.8 15.47 237 349.2 OOM 83.5
    FasterNet-L 93.5 15.52 323 219.5 OOM 83.5

    Key results:

    • FasterNet-T0 is 2.8×2.8\times, 3.3×3.3\times, and 2.4×2.4\times faster than MobileViT-XXS on GPU, CPU, and ARM, respectively, with +2.9%+2.9\% higher accuracy (71.9%71.9\% vs 69.0%69.0\%).
    • FasterNet-L attains 83.5%83.5\% accuracy (on par with Swin-B), while achieving 36%36\% higher GPU throughput (323323 vs 237237 fps) and saving 37%37\% compute latency on CPU (219.5219.5 ms vs 349.2349.2 ms).
  9. Knowl 9 — FasterNet on COCO Object Detection and Instance Segmentation

    empirical result

    When evaluated as the backbone in a Mask R-CNN detector trained with a standard 1×1\times schedule (12 epochs) using AdamW on the COCO benchmark (FLOPs evaluated at image resolution 1280×8001280 \times 800), FasterNet outperforms baseline CNN and vision transformer backbones:

    Backbone Params (M) FLOPs (G) GPU Latency (ms) APb\text{AP}^\text{b} AP50b\text{AP}^\text{b}_{50} AP75b\text{AP}^\text{b}_{75} APm\text{AP}^\text{m} AP50m\text{AP}^\text{m}_{50}
    ResNet50 44.2 253 54.9 38.0 58.6 41.4 34.4 55.1
    PoolFormer-S24 41.0 233 111.0 40.1 62.2 43.4 37.0 59.1
    PVT-Small 44.1 238 89.5 40.4 62.9 43.8 37.8 60.1
    FasterNet-S 49.0 258 54.3 39.9 61.2 43.6 36.9 58.1
    ResNet101 63.2 329 68.9 40.4 61.1 44.2 36.4 57.7
    ResNeXt101-32×\times4d 62.8 333 80.5 41.9 62.5 45.9 37.5 59.4
    FasterNet-M 71.2 344 71.4 43.0 64.4 47.4 39.1 61.5
    ResNeXt101-64×\times4d 101.9 487 112.9 42.8 63.8 47.3 38.4 60.6
    PVT-Large 81.0 358 152.2 42.9 65.0 46.6 39.5 61.9
    FasterNet-L 110.9 484 93.8 44.0 65.6 48.2 39.9 62.3

    Key performance points:

    • FasterNet-S achieves +1.9+1.9 higher box AP (APb\text{AP}^\text{b}) and +2.5+2.5 higher mask AP (APm\text{AP}^\text{m}) than ResNet50 at slightly lower GPU latency (54.354.3 ms vs 54.954.9 ms).
    • FasterNet-L attains 44.044.0 APb\text{AP}^\text{b} and 39.939.9 APm\text{AP}^\text{m}, outperforming PVT-Large (42.942.9 APb\text{AP}^\text{b}, 39.539.5 APm\text{AP}^\text{m}) while cutting GPU latency by 38%38\% (93.893.8 ms vs 152.2152.2 ms).
  10. Knowl 10 — Ablation on Partial Ratio, Normalization, and Activation in FasterNet

    empirical result

    Ablation experiments on ImageNet-1k evaluate the architectural design choices of FasterNet:

    Ablation Aspect Variant GPU (fps) CPU (ms) ARM (ms) Top-1 Acc. (%)
    Partial ratio (rr) FasterNet-T0 (r=1/2r=1/2, adjusted width/depth) 6626 9.6 145 71.7
    FasterNet-T0 (r=1/4r=1/4, default) 6807 9.2 143 71.9
    FasterNet-T0 (r=1/8r=1/8, adjusted width/depth) 6204 8.9 140 71.3
    Normalization FasterNet-T0 w/ BatchNorm (default) 6807 9.2 143 71.9
    FasterNet-T0 w/ LayerNorm 5515 10.7 159 71.9
    Activation FasterNet-T0 w/ ReLU 6929 8.2 114 71.3
    FasterNet-T0 w/ GELU (default) 6807 9.2 143 71.9
    FasterNet-T2 w/ ReLU (default) 1991 33.5 497 78.9
    FasterNet-T2 w/ GELU 1985 35.4 557 78.7

    Key findings:

    • Partial Ratio: r=1/4r = 1/4 gives the optimal trade-off (71.9%71.9\% accuracy). A larger ratio (r=1/2r=1/2) increases redundant computation, while a smaller ratio (r=1/8r=1/8) degrades spatial feature capture (71.3%71.3\%).
    • Normalization: BatchNorm yields the same accuracy as LayerNorm (71.9%71.9\%) but achieves significantly higher GPU throughput (68076807 vs 55155515 fps) and lower latency because BN can be fused into adjacent convolution layers during inference.
    • Activation: GELU improves accuracy for smaller models (FasterNet-T0: 71.9%71.9\% vs 71.3%71.3\% with ReLU) due to higher non-linearity, whereas ReLU yields better accuracy (78.9%78.9\% vs 78.7%78.7\%) and lower latency for larger models (FasterNet-T2).
  11. Knowl 11 — Stride Limitation of Partial Convolution

    limitation

    Partial Convolution (PConv) requires the spatial convolution stride to be fixed strictly to s=1s = 1. Because PConv applies spatial filtering to only a subset of cpc_p channels while leaving the remaining c−cpc - c_p channels untouched (identity mapping), setting s>1s > 1 would downsample the convolved channels without downsampling the untouched channels, resulting in a spatial resolution mismatch across channels.

    Consequently, spatial downsampling cannot be integrated directly into a PConv layer and must instead be handled by separate downsampling layers (e.g., standard embedding or merging convolution layers with stride >1> 1) situated between stages.

Coverage note — None was omitted; all contributed architectural designs, mathematical formulations, theoretical complexity analyses, benchmark evaluations across vision tasks and hardware platforms, ablation studies, and technical limitations have been covered.

References

  1. 1.Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al. Xcit: Cross-covariance image transformers. Advances in neural information processing systems, 34:20014–20027, 2021. 3
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 4
  3. 3.Han Cai, Chuang Gan, and Song Han. Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition. arXiv preprint arXiv:2205.14756, 2022. 3
  4. 4.Jierun Chen, Tianlang He, Weipeng Zhuo, Li Ma, Sangtae Ha, and S-H Gary Chan. Tvconv: Efficient translation variant convolution for layout-aware visual processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12548–12558, 2022. 3
  5. 5.Shoufa Chen, Enze Xie, Chongjian Ge, Ding Liang, and Ping Luo. Cyclemlp: A mlp-like architecture for dense prediction. arXiv preprint arXiv:2107.10224, 2021. 2, 3, 7
  6. 6.Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Xiaoyi Dong, Lu Yuan, and Zicheng Liu. Mobileformer: Bridging mobilenet and transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5270–5279, 2022. 1, 3
  7. 7.Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3435–3444, 2019. 3
  8. 8.Franc¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017. 3
  9. 9.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 702–703, 2020. 6
  10. 10.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. Advances in Neural Information Processing Systems, 34:3965–3977, 2021. 3
  11. 11.Xiaohan Ding et al. Scaling up your kernels to 31x31: Revisiting large kernel design in cnns. In CVPR, 2022. 9, 10
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 3
  13. 13.Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. The Journal of Machine Learning Research, 20(1):1997–2017, 2019. 11
  14. 14.Peng Gao, Teli Ma, Hongsheng Li, Jifeng Dai, and Yu Qiao. Convmae: Masked convolution meets masked autoencoders. arXiv preprint arXiv:2205.03892, 2022. 11
  15. 15.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve J ´ egou, and Matthijs ´ Douze. Levit: a vision transformer in convnet’s clothing for faster inference. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12259–12269, 2021. 3
  16. 16.Daniel Haase et al. Rethinking depthwise separable convolutions: How intra-kernel correlations lead to improved mobilenets. In CVPR, 2020. 10
  17. 17.Kai Han, Yunhe Wang, Qi Tian, Jianyuan Guo, Chunjing Xu, and Chang Xu. Ghostnet: More features from cheap operations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1580–1589, 2020. 1, 3, 4, 7
  18. 18.Kai Han, Yunhe Wang, Qiulin Zhang, Wei Zhang, Chunjing Xu, and Tong Zhang. Model rubik’s cube: Twisting resolution, depth and width for tinynets. Advances in Neural Information Processing Systems, 33:19353–19364, 2020. 3
  19. 19.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir- ´ shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 7
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 2, 4, 7, 8
  21. 21.Tianlang He, Jiajie Tan, Weipeng Zhuo, Maximilian Printz, and S-H Gary Chan. Tackling multipath and biased training data for imu-assisted ble proximity detection. In IEEE INFOCOM 2022-IEEE Conference on Computer Communications, pages 1259–1268. IEEE, 2022. 3
  22. 22.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016. 5
  23. 23.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015. 6, 11
  24. 24.Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mobilenetv3. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1314–1324, 2019. 1, 3
  25. 25.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 1, 3
  26. 26.Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. Local relation networks for image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3464–3473, 2019. 3
  27. 27.Gao Huang, Shichen Liu, Laurens Van der Maaten, and Kilian Q Weinberger. Condensenet: An efficient densenet using learned group convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2752–2761, 2018. 3
  28. 28.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In European conference on computer vision, pages 646–661. Springer, 2016. 6
  29. 29.Tao Huang, Lang Huang, Shan You, Fei Wang, Chen Qian, and Chang Xu. Lightvit: Towards light-weight convolutionfree vision transformers. arXiv preprint arXiv:2207.05557, 2022. 3
  30. 30.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015. 4
  31. 31.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012. 1, 3, 10
  32. 32.Anders Krogh and John Hertz. A simple weight decay can improve generalization. Advances in neural information processing systems, 4, 1991. 6
  33. 33.Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Lu Yuan, Zicheng Liu, Lei Zhang, and Nuno Vasconcelos. Micronet: Improving image recognition with extremely low flops. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 468–477, 2021. 1, 3
  34. 34.Yanyu Li, Geng Yuan, Yang Wen, Eric Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. Efficientformer: Vision transformers at mobilenet speed. arXiv preprint arXiv:2206.01191, 2022. 3
  35. 35.Dongze Lian, Zehao Yu, Xing Sun, and Shenghua Gao. Asmlp: An axial shifted mlp architecture for vision. arXiv preprint arXiv:2107.08391, 2021. 3
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 7
  37. 37.Guilin Liu, Aysegul Dundar, Kevin J Shih, Ting-Chun Wang, Fitsum A Reda, Karan Sapra, Zhiding Yu, Xiaodong Yang, Andrew Tao, and Bryan Catanzaro. Partial convolution for padding, inpainting, and image synthesis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. 11
  38. 38.Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. Image inpainting for irregular holes using partial convolutions. In Proceedings of the European conference on computer vision (ECCV), pages 85–100, 2018. 11
  39. 39.Ruiyang Liu, Yinghui Li, Linmi Tao, Dun Liang, and Hai-Tao Zheng. Are we ready for a new paradigm shift? a survey on visual deep mlp. Patterns, 3(7):100520, 2022. 3
  40. 40.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022. 3
  41. 41.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 2, 3, 7
  42. 42.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022. 1, 3, 7, 10
  43. 43.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 6
  44. 44.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 6
  45. 45.Jiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu, Weiguo Gao, Chunjing Xu, Tao Xiang, and Li Zhang. Soft: softmax-free transformer with linear complexity. Advances in Neural Information Processing Systems, 34:21297–21309, 2021. 3
  46. 46.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In Proceedings of the European conference on computer vision (ECCV), pages 116–131, 2018. 1, 2, 3, 7
  47. 47.Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal, Salman Khan, Syed Waqas Zamir, Rao Muhammad Anwer, and Fahad Shahbaz Khan. Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications. arXiv preprint arXiv:2206.10589, 2022. 7
  48. 48.Sachin Mehta and Mohammad Rastegari. Mobilevit: lightweight, general-purpose, and mobile-friendly vision transformer. arXiv preprint arXiv:2110.02178, 2021. 1, 2, 3, 7
  49. 49.Sachin Mehta and Mohammad Rastegari. Separable selfattention for mobile vision transformers. arXiv preprint arXiv:2206.02680, 2022. 1, 3
  50. 50.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440, 2016. 11
  51. 51.Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. In Icml, 2010. 5
  52. 52.Junting Pan, Adrian Bulat, Fuwen Tan, Xiatian Zhu, Lukasz Dudziak, Hongsheng Li, Georgios Tzimiropoulos, and Brais Martinez. Edgevits: Competing light-weight cnns on mobile devices with vision transformers. arXiv preprint arXiv:2205.03436, pages 1–6, 2022. 3
  53. 53.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. 6
  54. 54.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018. 1, 3, 4, 7
  55. 55.Laurent Sifre and Stephane Mallat. Rigid-motion scattering ´ for texture classification. arXiv preprint arXiv:1403.1687, 2014. 1, 3
  56. 56.Pravendra Singh, Vinay Kumar Verma, Piyush Rai, and Vinay P Namboodiri. Hetconv: Heterogeneous kernel-based convolutions for deep cnns. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4835–4844, 2019. 3
  57. 57.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16519–16529, 2021. 3
  58. 58.Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your vit? data, augmentation, and regularization in vision transformers. arXiv preprint arXiv:2106.10270, 2021. 3
  59. 59.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016. 6
  60. 60.Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le. Mnasnet: Platform-aware neural architecture search for mobile. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2820–2828, 2019. 3
  61. 61.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105–6114. PMLR, 2019. 3
  62. 62.Mingxing Tan and Quoc Le. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning, pages 10096–10106. PMLR, 2021. 3
  63. 63.Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan. Quadtree attention for vision transformers. arXiv preprint arXiv:2201.02767, 2022. 3
  64. 64.Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems, 34:24261–24272, 2021. 1, 3
  65. 65.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve J ´ egou. Training ´ data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021. 3
  66. 66.Hugo Touvron, Matthieu Cord, and Herve J ´ egou. Deit iii: ´ Revenge of the vit. arXiv preprint arXiv:2204.07118, 2022. 3
  67. 67.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016. 4
  68. 68.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12894–12904, 2021. 3
  69. 69.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 3
  70. 70.Shakti N Wadekar and Abhishek Chaurasia. Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features. arXiv preprint arXiv:2209.15159, 2022. 1
  71. 71.Guangting Wang, Yucheng Zhao, Chuanxin Tang, Chong Luo, and Wenjun Zeng. When shift operation meets vision transformer: An extremely simple alternative to attention mechanism. arXiv preprint arXiv:2201.10801, 2022. 3
  72. 72.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021. 3, 7, 8
  73. 73.Song Wen, Hao Wang, and Dimitris Metaxas. Social ode: Multi-agent trajectory forecasting with neural ordinary differential equations. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, pages 217–233. Springer, 2022. 3
  74. 74.Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer. Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10734–10742, 2019. 3
  75. 75.Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European conference on computer vision (ECCV), pages 3–19, 2018. 4
  76. 76.Xin Xia et al. Trt-vit: Tensorrt-oriented vision transformer. arXiv preprint, 2022. 9, 10
  77. 77.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and ´ Kaiming He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1492–1500, 2017. 8
  78. 78.Le Yang, Haojun Jiang, Ruojin Cai, Yulin Wang, Shiji Song, Gao Huang, and Qi Tian. Condensenet v2: Sparse feature reactivation for deep networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3569–3578, 2021. 3
  79. 79.Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10819–10829, 2022. 7, 8
  80. 80.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019. 6
  81. 81.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017. 6
  82. 82.Qiulin Zhang, Zhuqing Jiang, Qishuo Lu, Jia’nan Han, Zhengxin Zeng, Shang-Hua Gao, and Aidong Men. Split to be slim: An overlooked redundancy in vanilla convolution. arXiv preprint arXiv:2006.12085, 2020. 4
  83. 83.Ting Zhang, Guo-Jun Qi, Bin Xiao, and Jingdong Wang. Interleaved group convolutions. In Proceedings of the IEEE international conference on computer vision, pages 4373–4382, 2017. 3
  84. 84.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6848–6856, 2018. 1, 3
  85. 85.Shuhan Zhong, Sizhe Song, Guanyao Li, and S-H Gary Chan. A tree-based structure-aware transformer decoder for image-to-markup generation. In Proceedings of the 30th ACM International Conference on Multimedia, pages 5751–5760, 2022. 3
  86. 86.Weipeng Zhuo, Ka Ho Chiu, Jierun Chen, Jiajie Tan, Edmund Sumpena, S-H Gary Chan, Sangtae Ha, and Chul-Ho Lee. Semi-supervised learning with network embedding on ambient rf signals for geofencing services. arXiv preprint arXiv:2210.07889, 2022. 3
  87. 87.Barret Zoph and Quoc V Le. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016. 6

Citation

MLA
Chen, J., et al. “Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks”. arXiv, 2023, http://arxiv.org/abs/2303.03667v3.
APA
Chen, J., Kao, S.-. hong ., He, H., Zhuo, W., Wen, S., Lee, C.-H., & Chan, S.-H. G. (2023). Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks. arXiv. http://arxiv.org/abs/2303.03667v3
Chicago
Chen, J., S.-. hong . Kao, H. He, et al. 2023. “Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks”. arXiv. http://arxiv.org/abs/2303.03667v3.
Harvard
Chen, J. et al. (2023) “Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.03667v3.
Vancouver
1. Chen J, Kao S-hong, He H, Zhuo W, Wen S, Lee C-H, Chan S-HG (2023) Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks. arXiv

BibTeX

@article{chen2023run,
  title = {Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks},
  author = {Chen, Jierun and Kao, Shiu-hong and He, Hao and Zhuo, Weipeng and Wen, Song and Lee, Chul-Ho and Chan, S. -H. Gary},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.03667v3},
  eprint = {2303.03667}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/