Mobile-Former: Bridging MobileNet and Transformer
Yinpeng ChenXiyang DaiDongdong ChenMengchen LiuXiaoyi DongLu YuanZicheng Liu
Proposes a parallel architecture bridging MobileNet and a lightweight vision transformer with bidirectional cross-attention, significantly outperforming MobileNetV3 and DETR in image classification and object detection under strict computational budgets.
Deploying powerful computer vision models onto mobile and embedded devices requires strict limits on computational cost and energy usage. While modern vision transformers excel at capturing global context across an entire image, their performance degrades under low-computation constraints (below 1 billion operations), where specialized convolutional networks such as MobileNet traditionally dominate due to efficient local processing. Previous hybrid models combined these two paradigms in sequential series, but they struggle to maintain both high accuracy and extreme computational efficiency.
The article introduces and evaluates Mobile-Former, a parallel architecture designed to unite the local feature extraction of convolutional networks with the global modeling capacity of transformers within an ultra-low computational budget. The primary objective is to demonstrate that a parallel design with a bidirectional bridge can outperform standard efficient convolutional models and transformer variants across image classification and object detection tasks.
To evaluate this architecture, the authors conducted extensive experiments on standard benchmark datasets, specifically ImageNet (comprising over 1.28 million training images across 1,000 categories) and COCO (covering 118,000 training images). The method pairs a MobileNet branch with a lightweight transformer branch that uses only a few learnable tokens (six or fewer) rather than numerous image patches. Communication occurs via a two-way cross-attention bridge that eliminates redundant projections to keep the combined computational overhead of the transformer and bridge below 20% of the total budget. The team evaluated variants scaled from 26 million to 508 million operations.
The findings confirm clear performance advantages across key visual benchmarks. For image classification on ImageNet, Mobile-Former consistently outperformed standard efficient networks across the 25M to 500M operation range; for example, the 294M variant achieved a 77.9% top-1 accuracy, exceeding MobileNetV3 by 1.3% while reducing computations by 17%, and matching larger vision transformers using three to four times fewer computations. In object detection benchmarks, Mobile-Former used as a backbone in the RetinaNet framework outperformed MobileNetV3 by 8.6 Average Precision points. Furthermore, when structured as a full end-to-end detector, it surpassed the standard DETR model by 1.3 Average Precision while reducing computational operations by 52%, trimming model parameters by 36%, and requiring 40% fewer training cycles.
These results demonstrate that transformers can be successfully deployed in resource-constrained environments when paired in parallel with local convolutions. For technical and operational decision-makers, this translates into higher visual accuracy on edge devices without increasing hardware costs, computational budgets, or latency overhead for standard image sizes. The framework also simplifies model pipelines by enabling efficient end-to-end detection without manual feature pyramid tuning.
Organizations developing computer vision for edge devices should consider piloting parallel hybrid architectures like Mobile-Former for medium-to-large input resolutions. Engineering teams planning implementations must optimize the underlying bridge and transformer operations in deployment runtimes, as non-convolutional components exhibit execution overhead on small image sizes. Future efforts should also focus on compressing parameter-heavy classification heads to reduce total memory footprint.
- Paper: MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, Andrew G. Howard et al. (2017). Read MobileNets first to understand the efficient convolutional design that Mobile-Former uses as its local-processing foundation and compares against.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). The original Vision Transformer establishes patch tokens and global self-attention, the transformer machinery that Mobile-Former adapts through a small set of global tokens.
- Paper: MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer, Sachin Mehta et al. (2021). MobileViT provides an earlier mobile-focused CNN–transformer hybrid, clarifying the architectural trade-offs that Mobile-Former addresses with parallel branches and two-way communication.
- Paper: CMT: Convolutional Neural Networks Meet Vision Transformers, Jianyuan Guo et al. (2022). CMT shows how convolutional operations and transformer attention can be combined efficiently, offering a useful predecessor to Mobile-Former’s hybrid design.
- Paper: CoAtNet: Marrying Convolution and Attention for All Data Sizes, Zihang Dai et al. (2021). CoAtNet lays out a systematic way to combine convolution and attention, helping frame Mobile-Former’s choice to run local and global processing in parallel.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Deformable DETR introduces the end-to-end detection framework that Mobile-Former adapts and evaluates for more efficient object detection.
No sufficiently relevant recommendations were found.
