Mobile-Former: Bridging MobileNet and Transformer

Yinpeng ChenXiyang DaiDongdong ChenMengchen LiuXiaoyi DongLu YuanZicheng Liu

article2022CVPR717 citations

Proposes a parallel architecture bridging MobileNet and a lightweight vision transformer with bidirectional cross-attention, significantly outperforming MobileNetV3 and DETR in image classification and object detection under strict computational budgets.

Listen

Deploying powerful computer vision models onto mobile and embedded devices requires strict limits on computational cost and energy usage. While modern vision transformers excel at capturing global context across an entire image, their performance degrades under low-computation constraints (below 1 billion operations), where specialized convolutional networks such as MobileNet traditionally dominate due to efficient local processing. Previous hybrid models combined these two paradigms in sequential series, but they struggle to maintain both high accuracy and extreme computational efficiency.

The article introduces and evaluates Mobile-Former, a parallel architecture designed to unite the local feature extraction of convolutional networks with the global modeling capacity of transformers within an ultra-low computational budget. The primary objective is to demonstrate that a parallel design with a bidirectional bridge can outperform standard efficient convolutional models and transformer variants across image classification and object detection tasks.

To evaluate this architecture, the authors conducted extensive experiments on standard benchmark datasets, specifically ImageNet (comprising over 1.28 million training images across 1,000 categories) and COCO (covering 118,000 training images). The method pairs a MobileNet branch with a lightweight transformer branch that uses only a few learnable tokens (six or fewer) rather than numerous image patches. Communication occurs via a two-way cross-attention bridge that eliminates redundant projections to keep the combined computational overhead of the transformer and bridge below 20% of the total budget. The team evaluated variants scaled from 26 million to 508 million operations.

The findings confirm clear performance advantages across key visual benchmarks. For image classification on ImageNet, Mobile-Former consistently outperformed standard efficient networks across the 25M to 500M operation range; for example, the 294M variant achieved a 77.9% top-1 accuracy, exceeding MobileNetV3 by 1.3% while reducing computations by 17%, and matching larger vision transformers using three to four times fewer computations. In object detection benchmarks, Mobile-Former used as a backbone in the RetinaNet framework outperformed MobileNetV3 by 8.6 Average Precision points. Furthermore, when structured as a full end-to-end detector, it surpassed the standard DETR model by 1.3 Average Precision while reducing computational operations by 52%, trimming model parameters by 36%, and requiring 40% fewer training cycles.

These results demonstrate that transformers can be successfully deployed in resource-constrained environments when paired in parallel with local convolutions. For technical and operational decision-makers, this translates into higher visual accuracy on edge devices without increasing hardware costs, computational budgets, or latency overhead for standard image sizes. The framework also simplifies model pipelines by enabling efficient end-to-end detection without manual feature pyramid tuning.

Organizations developing computer vision for edge devices should consider piloting parallel hybrid architectures like Mobile-Former for medium-to-large input resolutions. Engineering teams planning implementations must optimize the underlying bridge and transformer operations in deployment runtimes, as non-convolutional components exhibit execution overhead on small image sizes. Future efforts should also focus on compressing parameter-heavy classification heads to reduce total memory footprint.

No sufficiently relevant recommendations were found.

Cover for Mobile-Former: Bridging MobileNet and Transformer

Abstract

We present Mobile-Former, a parallel design of MobileNet and transformer with a two-way bridge in between. This structure leverages the advantages of MobileNet at local processing and transformer at global interaction. And the bridge enables bidirectional fusion of local and global features. Different from recent works on vision transformer, the transformer in Mobile-Former contains very few tokens (e.g. 6 or fewer tokens) that are randomly initialized to learn global priors, resulting in low computational cost. Combining with the proposed light-weight cross attention to model the bridge, Mobile-Former is not only computationally efficient, but also has more representation power. It outperforms MobileNetV3 at low FLOP regime from 25M to 500M FLOPs on ImageNet classification. For instance, Mobile-Former achieves 77.9% top-1 accuracy at 294M FLOPs, gaining 1.3% over MobileNetV3 but saving 17% of computations. When transferring to object detection, Mobile-Former outperforms MobileNetV3 by 8.6 AP in RetinaNet framework. Furthermore, we build an efficient end-to-end detector by replacing backbone, encoder and decoder in DETR with Mobile-Former, which outperforms DETR by 1.3 AP but saves 52% of computational cost and 36% of parameters. Code will be released at https://github.com/aaboys/mobileformer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Method: Mobile-Former
  • 3.1 Overview
  • 3.2 Mobile-Former Block
  • 3.3 Network Specification
  • 4 Efficient End-to-End Object Detection
  • 5 Experimental Results
  • 5.1 ImageNet Classification
  • 5.2 Ablations
  • 5.3 Object Detection
  • 6 Limitations and Discussions
  • 7 Conclusion
  • References
  • A Mobile-Former Architecture
  • B More Experimental Results
  • C Visualization

Knowls

  1. Knowl 1 — Parallel local and global processing with a bidirectional bridge

    model/method

    Mobile-Former processes an image in two parallel streams: a MobileNet-style convolutional stream extracts local spatial features, while a transformer stream processes a small set of randomly initialized, learnable global tokens. Unlike patch-based vision transformers, the tokens are global priors rather than linear projections of image patches; the paper uses at most six tokens. Bidirectional cross-attention connects the streams: Mobile-to-Former supplies local information to the global tokens, and Former-to-Mobile distributes global information back to spatial features. The page 1 overview diagram depicts these two streams and their two-way bridge.

  2. Knowl 2 — Projection-efficient bidirectional cross-attention

    equation

    Let X∈RN×CX\in\mathbb{R}^{N\times C} be a Mobile feature map flattened over its N=HWN=HW spatial positions, with CC channels, and let Z∈RM×dZ\in\mathbb{R}^{M\times d} be MM global tokens of dimension dd. Split both into hh heads, denoting their head features by x~i\tilde{x}_i and z~i\tilde{z}_i for head ii. The bridge from local features to tokens uses projected token queries but unprojected local keys and values:

    AX→Z=[Attn⁡(z~iWiQ,x~i,x~i)]i=1hWO.\mathcal{A}_{X\rightarrow Z}=\left[\operatorname{Attn}(\tilde{z}_iW_i^Q,\tilde{x}_i,\tilde{x}_i)\right]_{i=1}^{h}W^O.

    The reverse bridge uses unprojected local queries and projected token keys and values:

    AZ→X=[Attn⁡(x~i,z~iWiK,z~iWiV)]i=1h.\mathcal{A}_{Z\rightarrow X}=\left[\operatorname{Attn}(\tilde{x}_i,\tilde{z}_iW_i^K,\tilde{z}_iW_i^V)\right]_{i=1}^{h}.

    Here WiQ,WiK,WiVW_i^Q,W_i^K,W_i^V are learned token-side head projections, WOW^O combines heads in the first direction, and brackets denote concatenation over heads. Attention is Attn⁡(Q,K,V)=softmax⁡(QKT/dk)V\operatorname{Attn}(Q,K,V)=\operatorname{softmax}(QK^T/\sqrt{d_k})V, where dkd_k is the key dimension per head. Thus, Mobile-side query, key, and value projections are omitted, while projections remain on the token side. Both directions are computed at the Mobile bottleneck, where the channel count is relatively low.

  3. Knowl 3 — Mobile-Former block and 294M-FLOP classifier configuration

    model/method

    A Mobile-Former block updates a local feature map and the global tokens using a Mobile sub-block, Mobile-to-Former cross-attention, a transformer sub-block, and Former-to-Mobile cross-attention. The page 3 block schematic shows this ordering and the connection from the token output to Mobile's dynamic ReLU. The Mobile sub-block is an inverted bottleneck with 1×1 expansion, 3×3 depthwise convolution, and 1×1 projection; it replaces ReLU with dynamic ReLU. Two MLP layers, with ReLU between them, generate the dynamic-ReLU parameters from the first updated global token. The transformer sub-block uses multi-head self-attention and a feed-forward network with expansion ratio 2 and post-layer normalization.

    The 294M-FLOP, 224×224 classifier uses six tokens of dimension 192, a stride-2 3×3 convolutional stem with 16 output channels, and a lite bottleneck at 112×112. Its 11 Mobile-Former blocks are distributed across four downsampling stages: two blocks at 112×112/56×56 with 24 channels; two at 56×56/28×28 with 48 channels; four at 28×28/14×14 with 96–128 channels; and three at 14×14/7×7 with 192 channels. A 1×1 convolution produces 1152 channels at 7×7. The classification head average-pools the local features, concatenates the pooled representation with the first global token, and applies two fully connected layers with h-swish between them. The family contains seven variants spanning 26M to 508M FLOPs.

  4. Knowl 4 — ImageNet accuracy in the low-computation regime

    empirical result

    On ImageNet classification at 224×224, Mobile-Former variants trained from scratch for 450 epochs outperform efficient CNNs across the evaluated low-FLOP range of 26M–508M, with the paper noting one comparison group around 150M FLOPs where Mobile-Former uses slightly more computation than ShuffleNetV2 and WeightNet. Mobile-Former-26M obtains 64.0% top-1 accuracy at 26M MAdds, versus 62.8% for MobileNetV3 Small at 30M; Mobile-Former-96M obtains 72.8% at 96M, versus 71.7% for MobileNetV3 at 112M. Mobile-Former-214M reaches 76.7% at 214M, versus MobileNetV3's 75.2% at 217M. Mobile-Former-294M reaches 77.9% at 294M, compared with 76.6% for MobileNetV3 1.25× at 356M and 77.1% for EfficientNet-B0 at 390M. Mobile-Former-508M reaches 79.3% at 508M, versus 74.9% for ShuffleNetV2 2× at 591M. Against vision transformers trained at the same 224×224 resolution without teacher distillation, Mobile-Former-294M scores 77.9% versus 77.3% for Swin-1G at 1.0G MAdds, and Mobile-Former-508M scores 79.3% versus 79.2% for Swin-2G at 2.0G MAdds.

  5. Knowl 5 — Global-token count and dimension have diminishing returns

    empirical result

    ImageNet ablations of Mobile-Former-294M show that a small token set is sufficient for strong performance. With token dimension fixed at 192, using 1, 3, 6, or 9 tokens gives respectively 77.1%, 77.6%, 77.8%, and 77.7% top-1 accuracy at 269M, 279M, 294M, and 309M MAdds; each setting has 11.4M parameters. With six tokens, dimensions 64, 128, 192, 256, and 320 give respectively 76.8%, 77.3%, 77.8%, 77.8%, and 77.6% top-1 accuracy at 277M, 284M, 294M, 308M, and 325M MAdds. The corresponding parameter counts are 7.3M, 9.1M, 11.4M, 14.3M, and 17.9M. Thus, accuracy gains flatten beyond six tokens or a dimension of 192. At six tokens of dimension 192, the Former and bridge account for 35M of the model's 294M MAdds.

  6. Knowl 6 — Former, bridge, and dynamic ReLU improve the Mobile baseline

    empirical result

    In a 300-epoch ImageNet ablation, a Mobile-only model using ReLU achieves 74.2% top-1 and 91.8% top-5 accuracy with 6.1M parameters and 259M MAdds. Adding the Former and both cross-attention directions raises accuracy to 76.8% top-1 and 93.2% top-5, at 10.1M parameters and 290M MAdds. Replacing ReLU with dynamic ReLU whose parameters come from the first global token raises accuracy further to 77.8% top-1 and 93.7% top-5, at 11.4M parameters and 294M MAdds. The reported gains over the Mobile-only baseline are 2.6 and then a further 1.0 percentage points in top-1 accuracy.

  7. Knowl 7 — Mobile-Former as a RetinaNet backbone

    empirical result

    On COCO val2017 in a RetinaNet framework, models are trained for 12 epochs from ImageNet-pretrained weights; computation is reported at image size 800×1333. Mobile-Former-214M achieves 35.8 AP with 3.9G backbone MAdds (162G total) and 5.7M backbone parameters (15.2M total). MobileNetV3 achieves 27.2 AP with 4.7G backbone MAdds (162G total) and 2.8M backbone parameters (12.3M total), so Mobile-Former gains 8.6 AP at lower backbone computation. Mobile-Former-508M reaches 38.0 AP with 9.8G backbone MAdds (168G total) and 8.4M backbone parameters (17.9M total); for comparison, ResNet-50 reaches 36.5 AP at 84G backbone MAdds, PVT-Tiny reaches 36.7 AP at 70G, and ConT-M reaches 37.9 AP at 65G.

  8. Knowl 8 — Multi-scale end-to-end detector and query-conditioned adaptations

    model/method

    The end-to-end Mobile-Former detector uses separate tokens in its backbone and head: six global tokens in the backbone and 100 object queries in the head. Unlike the single-scale DETR head, the Mobile-Former head processes queries at feature resolutions 1/32, 1/16, and 1/8. It upsamples features by bilinear interpolation and adds the backbone feature at the matching resolution; all queries are progressively refined from coarse to fine. The 508M-backbone detector uses nine head blocks, distributed 5, 2, and 2 across the three resolutions.

    For spatial-aware dynamic ReLU, parameters at spatial position ii are a weighted combination of outputs from all global tokens: θi=∑jαi,jf(zj)\theta_i=\sum_j\alpha_{i,j}f(z_j), where zjz_j is token jj, ff is a two-layer MLP with an intermediate ReLU, and αi,j\alpha_{i,j} is the Mobile-to-Former cross-attention weight normalized over tokens so that ∑jαi,j=1\sum_j\alpha_{i,j}=1. This replaces the spatial-shared version based only on the first token. In the head, let qkfq_k^f and qkpq_k^p be a query's feature and position embeddings at block kk. After updating the feature embedding, the position embedding is refined as qk+1p=qkp+g(qk+1f)q_{k+1}^p=q_k^p+g(q_{k+1}^f), where gg is a two-layer MLP with an intermediate ReLU. The feature and position embeddings are summed for attention, allowing query positions to adapt as head features change. The detector uses prediction FFNs and auxiliary losses; its head is trained from scratch while its backbone is ImageNet-pretrained.

  9. Knowl 9 — End-to-end detection accuracy and component ablations

    empirical result

    On COCO, the end-to-end Mobile-Former detector with a 508M-FLOP backbone and 100 object queries achieves 43.3 AP at 41.4G MAdds and 26.3M parameters. DETR obtains 42.0 AP at 86G MAdds and 41.3M parameters, while DETR-DC5 obtains 43.3 AP at 187G MAdds. The Mobile-Former detector is trained for 300 epochs, compared with 500 for the reported DETR baselines. Smaller end-to-end Mobile-Former variants achieve 40.5 AP at 24.1G MAdds, 39.3 AP at 17.8G, and 37.2 AP at 12.7G.

    In a separate 300-epoch ablation, replacing DETR's ResNet-50 backbone with Mobile-Former-508M while keeping the DETR head gives 39.4 AP. Adding spatial-aware dynamic ReLU raises this to 40.5 AP; replacing the DETR head with the multi-scale Mobile-Former head raises it to 41.4 AP; and adapting position embeddings raises it to 43.3 AP. Together, the three additions gain 3.9 AP over the Mobile-Former-backbone/DETR-head configuration. The multi-scale head improves small- and medium-object AP in this ablation, with a slight reduction for large objects.

  10. Knowl 10 — Inference and parameter-efficiency limitations

    limitation

    Mobile-Former is not uniformly faster than MobileNetV3: the paper reports that it is faster for large images but becomes slower as image resolution decreases. Former and bridge operations, including embedding projections, are relatively independent of image resolution, and their PyTorch implementations are less efficient than convolution; their overhead therefore becomes more visible on small images. The authors identify implementation optimization as a possible way to improve runtime. The model is also parameter-heavy for image classification: the Mobile-Former-294M classification head uses 4.6M of its 11.4M parameters, about 40%. Removing that classification head for detection mitigates this issue, but the Former and two-way bridge remain computationally efficient rather than parameter-efficient.

Coverage note — The paper's qualitative visualization observations about how token focus and cross-attention vary across network depth are omitted because they are secondary interpretations without quantitative evaluation.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2, 4, 5, 7, 8
  2. 2.Chun-Fu (Richard) Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 357–366, October 2021. 3, 6
  3. 3.Hanting Chen, Yunhe Wang, Chunjing Xu, Boxin Shi, Chao Xu, Qi Tian, and Chang Xu. Addernet: Do we really need multiplications in deep learning? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  4. 4.Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. Dynamic convolution: Attention over convolution kernels. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  5. 5.Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. Dynamic relu. In ECCV, 2020. 2, 3, 4, 6
  6. 6.Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le. Autoaugment: Learning augmentation strategies from data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 5
  7. 7.Stephane d’Ascoli, Hugo Touvron, Matthew Leavitt, Ari Morcos, Giulio Biroli, and Levent Sagun. Convit: Improving vision transformers with soft convolutional inductive biases. arXiv preprint arXiv:2103.10697, 2021. 2, 5, 6
  8. 8.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 5, 6
  9. 9.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021. 2
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. 1, 2, 3
  11. 11.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve Jégou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference. arXiv preprint arXiv:22104.01136, 2021. 1, 2, 6
  12. 12.Kai Han, Yunhe Wang, Qi Tian, Jianyuan Guo, Chunjing Xu, and Chang Xu. Ghostnet: More features from cheap operations. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 1, 2, 6
  13. 13.Kai Han, Yunhe Wang, Qiulin Zhang, Wei Zhang, Chunjing XU, and Tong Zhang. Model rubiks cube: Twisting resolution, depth and width for tinynets. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 19353–19364. Curran Associates, Inc., 2020. 2
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 2, 7
  15. 15.Geoffrey E. Hinton. How to represent part-whole hierarchies in a neural network. CoRR, abs/2102.12627, 2021. 2
  16. 16.Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam. Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019. 1, 2, 4, 5, 6, 7, 8
  17. 17.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 1, 2
  18. 18.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2
  19. 19.Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Lu Yuan, Zicheng Liu, Lei Zhang, and Nuno Vasconcelos. Micronet: Improving image recognition with extremely low flops. In International Conference on Computer Vision, 2021. 1, 2, 4
  20. 20.T. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie. Feature pyramid networks for object detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 936–944, July 2017. 4
  21. 21.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct 2017. 2, 7
  22. 22.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 5, 7
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021. 2, 5, 6
  24. 24.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. 5
  25. 25.Ningning Ma, X. Zhang, J. Huang, and J. Sun. Weightnet: Revisiting the design space of weight networks. volume abs/2007.11823, 2020. 5, 6
  26. 26.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In The European Conference on Computer Vision (ECCV), September 2018. 2, 5, 6, 7, 8
  27. 27.Zhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie, Yaowei Wang, Jianbin Jiao, and Qixiang Ye. Conformer: Local features coupling global representations for visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 367–376, October 2021. 3
  28. 28.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4510–4520, 2018. 1, 2, 3, 4
  29. 29.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16519–16529, June 2021. 2, 6
  30. 30.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, pages 6105–6114, Long Beach, California, USA, 09–15 Jun 2019. 2, 5, 6
  31. 31.Mingxing Tan and Quoc V. Le. Mixconv: Mixed depthwise convolutional kernels. In 30th British Machine Vision Conference 2019, 2019. 2
  32. 32.Mingxing Tan, Ruoming Pang, and Quoc V. Le. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  33. 33.Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision, 2021. 6, 7
  34. 34.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jégou. Training data-efficient image transformers and distillation through attention. arXiv preprint arXiv:2012.12877, 2020. 1, 2, 5, 6
  35. 35.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herve Jégou. Going deeper with image transformers, 2021. 2
  36. 36.Keivan Alizadeh vahid, Anish Prabhu, Ali Farhadi, and Mohammad Rastegari. Butterfly transform: An efficient fft based neural architecture design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  37. 37.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12894–12904, June 2021. 2
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. 1, 3, 4
  39. 39.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions, 2021. 5, 6, 7
  40. 40.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers, 2021. 1, 2
  41. 41.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, and Ross B. Girshick. Early convolutions help transformers see better. CoRR, abs/2106.14881, 2021. 1, 2, 4, 5, 6
  42. 42.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers, 2021. 5, 6
  43. 43.Haotian Yan, Zhe Li, Weijian Li, Changhu Wang, Ming Wu, and Chuang Zhang. Contnet: Why not use convolution and transformer at the same time? CoRR, abs/2104.13497, 2021. 6, 7
  44. 44.Brandon Yang, Gabriel Bender, Quoc V. Le, and Jiquan Ngiam. Condconv: Conditionally parameterized convolutions for efficient inference. In NeurIPS, 2019. 2
  45. 45.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021. 2, 5, 6
  46. 46.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, 2018. 5
  47. 47.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2
  48. 48.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2020. 5
  49. 49.Daquan Zhou, Qi-Bin Hou, Y. Chen, Jiashi Feng, and S. Yan. Rethinking bottleneck structure for efficient mobile network design. In ECCV, August 2020. 2
  50. 50.Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Zihang Jiang, Qibin Hou, and Jiashi Feng. Deepvit: Towards deeper vision transformer, 2021. 2
  51. 51.Daquan Zhou, Yujun Shi, Bingyi Kang, Weihao Yu, Zihang Jiang, Yuan Li, Xiaojie Jin, Qibin Hou, and Jiashi Feng. Refiner: Refining self-attention for vision transformers, 2021. 2

Citation

MLA
Chen, Y., et al. “Mobile-Former: Bridging MobileNet and Transformer”. arXiv, 2021, http://arxiv.org/abs/2108.05895v3.
APA
Chen, Y., Dai, X., Chen, D., Liu, M., Dong, X., Yuan, L., & Liu, Z. (2021). Mobile-Former: Bridging MobileNet and Transformer. arXiv. http://arxiv.org/abs/2108.05895v3
Chicago
Chen, Y., X. Dai, D. Chen, et al. 2021. “Mobile-Former: Bridging MobileNet and Transformer”. arXiv. http://arxiv.org/abs/2108.05895v3.
Harvard
Chen, Y. et al. (2021) “Mobile-Former: Bridging MobileNet and Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2108.05895v3.
Vancouver
1. Chen Y, Dai X, Chen D, Liu M, Dong X, Yuan L, Liu Z (2021) Mobile-Former: Bridging MobileNet and Transformer. arXiv

BibTeX

@article{chen2021mobile,
  title = {Mobile-Former: Bridging MobileNet and Transformer},
  author = {Chen, Yinpeng and Dai, Xiyang and Chen, Dongdong and Liu, Mengchen and Dong, Xiaoyi and Yuan, Lu and Liu, Zicheng},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2108.05895v3},
  eprint = {2108.05895}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE