Inception Transformer
Chenyang SiWeihao YuPan ZhouYichen ZhouXinchao WangShuicheng Yan
Proposes the Inception Transformer, a hybrid vision architecture that couples parallel frequency-specific mixers via channel splitting with a layer-wise frequency ramp to capture both fine local details and broad global context with superior parameter efficiency over standard vision transformers.
Modern computer vision systems increasingly rely on Vision Transformers due to their exceptional ability to capture global context across an entire image. However, standard Transformers act like low-pass filters, struggling to capture high-frequency visual details such as sharp edges and fine textures that traditional convolutional neural networks handle well. Existing attempts to merge these approaches either process global and local features sequentially—which discards one type of information at each stage—or process all features through parallel paths, causing computational redundancy. The article develops and evaluates a new architecture, called the Inception Transformer (iFormer), designed to simultaneously and efficiently process both high- and low-frequency image information.
The researchers designed an Inception token mixer that splits feature channels: one subset passes through a high-frequency branch consisting of parallel convolution and max-pooling, while the remaining channels pass through a low-frequency self-attention branch. To reduce computation, the attention mechanism operates on downsampled features before restoring full resolution. In addition, the design implements a frequency ramp structure across the network layers, allocating more channels to high-frequency processing in the early layers and progressively shifting capacity to low-frequency global attention in deeper layers. The framework was evaluated on standard vision benchmarks across three core tasks: image classification on ImageNet-1K, object detection and instance segmentation on COCO, and semantic segmentation on ADE20K.
The findings show that iFormer consistently outperforms both pure Transformers and hybrid models across model sizes. On ImageNet-1K, the small variant (iFormer-S) attained an 83.4% top-1 accuracy, matching or slightly exceeding much larger models such as Swin-B (83.3%) while using only one-quarter of the parameters and one-third of the floating-point operations. On COCO object detection, iFormer-S achieved 46.2 box average precision, outperforming standard baseline ResNet50 by 8.2 points and leading other Transformer backbones. On ADE20K segmentation, iFormer-S achieved a mean intersection-over-union of 48.6%, surpassing the comparable UniFormer-S by 2.0% while requiring fewer computational resources. Ablation experiments confirmed that coupling convolution with max-pooling and using the frequency ramp structure provided the highest accuracy.
These results demonstrate that explicitly separating and balancing spatial frequency processing substantially improves visual recognition accuracy without driving up computational budgets. For technical leaders and engineering teams, adopting frequency-aware hybrid backbones offers a path to deploy more compact, higher-accuracy computer vision systems in resource-constrained environments. As next steps, teams building vision systems should consider piloting iFormer-style channel-splitting architectures for tasks where fine-grained textures and global context are both critical. Before broad operational deployment, practitioners should evaluate automated methods like neural architecture search, as manually configuring the channel-split ratios across layers remains a primary limitation that currently requires task-specific tuning. In addition, testing is recommended on larger pretraining datasets, since current results are bounded by standard-scale ImageNet training.
- Paper: MetaFormer is Actually What You Need for Vision, Weihao Yu et al. (2021). It conceptualizes general token-mixing architectures like MetaFormer and demonstrates the utility of simple pooling operations, establishing the framework iFormer builds upon to design its hybrid Inception mixer.
- Paper: CoAtNet: Marrying Convolution and Attention for All Data Sizes, Zihang Dai et al. (2021). It establishes early paradigms for coupling depthwise convolutions with self-attention across network stages, motivating iFormer's frequency-tailored hybrid design.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). It provides foundational insights into grafting convolutional inductive biases directly into vision transformers to enhance local feature representations.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). It analyzes the representational differences between self-attention and convolutions across layer depths, directly informing iFormer's frequency ramp structure across bottom and top layers.
- Paper: Transformer in Transformer, Kai Han et al. (2021). It highlights the necessity of explicitly capturing fine-grained local details alongside global context in visual transformers, which iFormer tackles via parallel high-frequency convolution paths.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It introduces standard data-efficient training recipes and benchmarks (DeiT) against which iFormer measures its efficiency and accuracy gains.
- Paper: FFT-Based Dynamic Token Mixer for Vision, Yuki Tatsunami et al. (2024). It extends frequency-domain visual modeling by investigating dynamic FFT-based token mixers as an alternative to explicit convolution and self-attention pathways.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). It builds upon hybrid local-convolution and global-attention mixing designs by proposing foveal visual mechanisms to eliminate window artifacts and enhance multi-scale perception.
- Paper: BiFormer: Vision Transformer with Bi-Level Routing Attention, Lei Zhu et al. (2023). It advances general-purpose vision backbones by introducing bi-level dynamic routing to balance high-detail local routing with global contextual attention.
- Paper: Frequency-Aware Transformer for Learned Image Compression, Han Li et al. (2024). It applies frequency-aware decomposition and transformer token mixing principles specifically to optimized learned image compression.
- Paper: InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions, Wenhai Wang et al. (2023). It extends the exploration of capturing wide-ranging visual frequencies and multi-scale contexts by modernizing deformable convolutions into large foundation models.
