Inception Transformer

Chenyang SiWeihao YuPan ZhouYichen ZhouXinchao WangShuicheng Yan

article2022NeurIPS310 citations

Proposes the Inception Transformer, a hybrid vision architecture that couples parallel frequency-specific mixers via channel splitting with a layer-wise frequency ramp to capture both fine local details and broad global context with superior parameter efficiency over standard vision transformers.

Listen

Modern computer vision systems increasingly rely on Vision Transformers due to their exceptional ability to capture global context across an entire image. However, standard Transformers act like low-pass filters, struggling to capture high-frequency visual details such as sharp edges and fine textures that traditional convolutional neural networks handle well. Existing attempts to merge these approaches either process global and local features sequentially—which discards one type of information at each stage—or process all features through parallel paths, causing computational redundancy. The article develops and evaluates a new architecture, called the Inception Transformer (iFormer), designed to simultaneously and efficiently process both high- and low-frequency image information.

The researchers designed an Inception token mixer that splits feature channels: one subset passes through a high-frequency branch consisting of parallel convolution and max-pooling, while the remaining channels pass through a low-frequency self-attention branch. To reduce computation, the attention mechanism operates on downsampled features before restoring full resolution. In addition, the design implements a frequency ramp structure across the network layers, allocating more channels to high-frequency processing in the early layers and progressively shifting capacity to low-frequency global attention in deeper layers. The framework was evaluated on standard vision benchmarks across three core tasks: image classification on ImageNet-1K, object detection and instance segmentation on COCO, and semantic segmentation on ADE20K.

The findings show that iFormer consistently outperforms both pure Transformers and hybrid models across model sizes. On ImageNet-1K, the small variant (iFormer-S) attained an 83.4% top-1 accuracy, matching or slightly exceeding much larger models such as Swin-B (83.3%) while using only one-quarter of the parameters and one-third of the floating-point operations. On COCO object detection, iFormer-S achieved 46.2 box average precision, outperforming standard baseline ResNet50 by 8.2 points and leading other Transformer backbones. On ADE20K segmentation, iFormer-S achieved a mean intersection-over-union of 48.6%, surpassing the comparable UniFormer-S by 2.0% while requiring fewer computational resources. Ablation experiments confirmed that coupling convolution with max-pooling and using the frequency ramp structure provided the highest accuracy.

These results demonstrate that explicitly separating and balancing spatial frequency processing substantially improves visual recognition accuracy without driving up computational budgets. For technical leaders and engineering teams, adopting frequency-aware hybrid backbones offers a path to deploy more compact, higher-accuracy computer vision systems in resource-constrained environments. As next steps, teams building vision systems should consider piloting iFormer-style channel-splitting architectures for tasks where fine-grained textures and global context are both critical. Before broad operational deployment, practitioners should evaluate automated methods like neural architecture search, as manually configuring the channel-split ratios across layers remains a primary limitation that currently requires task-specific tuning. In addition, testing is recommended on larger pretraining datasets, since current results are bounded by standard-scale ImageNet training.

arXiv: 2205.12956
Cover for Inception Transformer

Abstract

Recent studies show that Transformer has strong capability of building long-range dependencies, yet is incompetent in capturing high frequencies that predominantly convey local information. To tackle this issue, we present a novel and general-purpose Inception Transformer, or iFormer for short, that effectively learns comprehensive features with both high- and low-frequency information in visual data. Specifically, we design an Inception mixer to explicitly graft the advantages of convolution and max-pooling for capturing the high-frequency information to Transformers. Different from recent hybrid frameworks, the Inception mixer brings greater efficiency through a channel splitting mechanism to adopt parallel convolution/max-pooling path and self-attention path as high- and low-frequency mixers, while having the flexibility to model discriminative information scattered within a wide frequency range. Considering that bottom layers play more roles in capturing high-frequency details while top layers more in modeling low-frequency global information, we further introduce a frequency ramp structure, i.e., gradually decreasing the dimensions fed to the high-frequency mixer and increasing those to the low-frequency mixer, which can effectively trade-off high- and low-frequency components across different layers. We benchmark the iFormer on a series of vision tasks, and showcase that it achieves impressive performance on image classification, COCO detection and ADE20K segmentation. For example, our iFormer-S hits the top-1 accuracy of 83.4% on ImageNet-1K, much higher than DeiT-S by 3.6%, and even slightly better than much bigger model Swin-B (83.3%) with only 1/4 parameters and 1/3 FLOPs. Code and models are released at https://github.com/sail-sg/iFormer.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Revisit Vision Transformer
  • 3.2 Inception token mixer
  • 3.3 Frequency ramp structure
  • 4 Experiments
  • 4.1 Results on image classification
  • 4.2 Results on object detection and instance segmentation
  • 4.3 Results on semantic segmentation
  • 4.4 Ablation study and visualization
  • 5 Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Frequency-aware Inception Transformer backbone

    model/method

    The Inception Transformer (iFormer) is a hierarchical vision-transformer backbone designed to represent both high-frequency visual information, such as local edges and textures, and low-frequency information, such as global shapes and scene structure. Each iFormer block replaces the vanilla Transformer token mixer with an Inception token mixer that routes different channel subsets through high-frequency and low-frequency processing paths.

    The backbone has four stages. For an input RGB image of spatial size H×WH\times W, successive patch embeddings produce feature maps with channel dimensions C1,C2,C3,C4C_1,C_2,C_3,C_4 and spatial resolutions H/4×W/4H/4\times W/4, H/8×W/8H/8\times W/8, H/16×W/16H/16\times W/16, and H/32×W/32H/32\times W/32, respectively. The high-frequency path uses max-pooling and depthwise convolution, whereas the low-frequency path uses self-attention.

  2. Knowl 2 — Channel-split Inception token mixer

    model/method

    Given a token feature map X∈RN×CX\in\mathbb{R}^{N\times C}, where NN is the number of spatial tokens and CC is the channel dimension, the Inception token mixer splits the channels into a high-frequency component Xh∈RN×ChX_h\in\mathbb{R}^{N\times C_h} and a low-frequency component Xl∈RN×ClX_l\in\mathbb{R}^{N\times C_l}, with Ch+Cl=CC_h+C_l=C. The high-frequency component is divided equally into Xh1,Xh2∈RN×Ch/2X_{h1},X_{h2}\in\mathbb{R}^{N\times C_h/2}, assuming ChC_h is divisible by two.

    The two high-frequency subbranches are

    Yh1=FC⁡(MaxPool⁡(Xh1)),Yh2=DwConv⁡(FC⁡(Xh2)),Y_{h1}=\operatorname{FC}(\operatorname{MaxPool}(X_{h1})),\qquad Y_{h2}=\operatorname{DwConv}(\operatorname{FC}(X_{h2})),

    where FC⁡\operatorname{FC} is a channel-wise fully connected projection, MaxPool⁡\operatorname{MaxPool} is a spatial max-pooling operation, and DwConv⁡\operatorname{DwConv} is a depthwise convolution. If YlY_l denotes the output of the low-frequency branch, the three branch outputs are concatenated along channels as

    Yc=Concat⁡(Yl,Yh1,Yh2).Y_c=\operatorname{Concat}(Y_l,Y_{h1},Y_{h2}).

    The concatenated representation is fused by exchanging information between neighboring spatial tokens while retaining per-location channel mixing:

    Y=FC⁡(Yc+DwConv⁡(Yc)).Y=\operatorname{FC}\bigl(Y_c+\operatorname{DwConv}(Y_c)\bigr).

    Thus, the mixer does not process every channel with every branch; each channel subset is assigned to the frequency-processing operation most appropriate for it.

  3. Knowl 3 — Pooled self-attention low-frequency branch

    model/method

    The low-frequency branch applies vanilla multi-head self-attention to the low-frequency channel subset after spatial downsampling. For Xl∈RN×ClX_l\in\mathbb{R}^{N\times C_l}, where NN is the number of spatial tokens and ClC_l is the number of low-frequency channels, the branch is

    Yl=Upsample⁡ ⁣(MSA⁡ ⁣(AvePool⁡(Xl))),Y_l=\operatorname{Upsample}\!\left(\operatorname{MSA}\!\left(\operatorname{AvePool}(X_l)\right)\right),

    where AvePool⁡\operatorname{AvePool} is average pooling, MSA⁡\operatorname{MSA} is multi-head self-attention over the pooled tokens, and Upsample⁡\operatorname{Upsample} restores the original spatial resolution. The pooling and upsampling kernel sizes and strides are two in the first two backbone stages; the operations are not downsampling operations in the later stages. This design reduces the cost of self-attention at high spatial resolutions while preserving its ability to aggregate global information.

  4. Knowl 4 — Residual iFormer block

    model/method

    An iFormer block uses pre-normalization and residual connections around the Inception token mixer and a feed-forward network. For block input XX, where XX is a token feature map, the intermediate representation YY and block output HH are

    Y=X+ITM⁡(LN⁡(X)),Y=X+\operatorname{ITM}(\operatorname{LN}(X)),

    H=Y+FFN⁡(LN⁡(Y)),H=Y+\operatorname{FFN}(\operatorname{LN}(Y)),

    where LN⁡\operatorname{LN} is layer normalization, ITM⁡\operatorname{ITM} is the channel-split Inception token mixer, and FFN⁡\operatorname{FFN} is the Transformer feed-forward network. The residual structure allows high- and low-frequency mixing to be inserted into a standard Transformer block without removing the usual feed-forward transformation.

  5. Knowl 5 — Frequency ramp across hierarchical stages

    model/method

    The frequency ramp assigns different proportions of channels to the high- and low-frequency mixers at different depths. For a block with total channel dimension CC, define the high-frequency ratio Ch/CC_h/C and low-frequency ratio Cl/CC_l/C, subject to

    ChC+ClC=1.\frac{C_h}{C}+\frac{C_l}{C}=1.

    From shallow to deep stages, Ch/CC_h/C is gradually decreased and Cl/CC_l/C is gradually increased. The intended allocation is therefore more local, detail-sensitive processing in lower layers and more global, self-attention-based processing in higher layers. The ratios are manually specified separately for the iFormer blocks, allowing the trade-off between high- and low-frequency representations to vary throughout the backbone.

  6. Knowl 6 — ImageNet classification performance

    data/table

    iFormer was evaluated on ImageNet-1K using AdamW with initial learning rate 10−310^{-3}, cosine decay, momentum 0.90.9, weight decay 0.050.05, 300 epochs, and 224×224224\times224 training images. DeiT data augmentation and regularization were used, and LayerScale was used for deep models. The principal 224×224224\times224 results were:

    • iFormer-S: 20M parameters, 4.8 GFLOPs, 83.4% top-1 and 96.6% top-5 accuracy. Comparable models included DeiT-S at 79.8% top-1 with 22M parameters and 4.6 GFLOPs, Swin-T at 81.3% with 29M and 4.5 GFLOPs, ConvNeXt-T at 82.1% with 28M and 4.5 GFLOPs, CSwin-T at 82.7% with 23M and 4.3 GFLOPs, and UniFormer-S at 82.9% with 22M and 3.6 GFLOPs.
    • iFormer-B: 48M parameters, 9.4 GFLOPs, 84.6% top-1 and 97.0% top-5 accuracy. Comparable results were ConvNeXt-S at 83.1% with 50M and 8.7 GFLOPs, CSwin-S at 83.6% with 35M and 6.9 GFLOPs, and UniFormer-B at 83.9% with 50M and 8.3 GFLOPs.
    • iFormer-L: 87M parameters, 14.0 GFLOPs, 84.8% top-1 and 97.0% top-5 accuracy. Swin-B achieved 83.3% with 88M parameters and 15.4 GFLOPs, CSwin-B achieved 84.2% with 78M and 15.0 GFLOPs, and ConvNeXt-B achieved 83.8% with 89M and 15.4 GFLOPs.

    For fine-tuning with 384×384384\times384 test resolution, iFormer-S achieved 84.6% top-1 accuracy with 20M parameters and 16.1 GFLOPs, iFormer-B achieved 85.7% with 48M parameters and 30.5 GFLOPs, and iFormer-L achieved 85.8% with 87M parameters and 45.3 GFLOPs. The fine-tuning used weight decay 10−810^{-8}, learning rate 10−510^{-5}, and batch size 512. These results show that iFormer retained an accuracy advantage over similarly sized convolutional, Transformer, and hybrid backbones.

  7. Knowl 7 — COCO detection and instance-segmentation performance

    empirical result

    The iFormer backbone was evaluated in Mask R-CNN on COCO using the 1×\times schedule of 12 epochs. ImageNet-pretrained backbones were fine-tuned with AdamW, initial learning rate 10−410^{-4}, batch size 16, and images resized to 800 pixels on the shorter side and at most 1,333 pixels on the longer side. FLOPs were measured at 800×1280800\times1280 resolution.

    The iFormer-S backbone used 40M parameters and 263 GFLOPs and achieved bounding-box AP 46.246.2, AP50b_{50}^{b} 68.568.5, AP75b_{75}^{b} 50.650.6, mask AP 41.941.9, AP50m_{50}^{m} 65.365.3, and AP75m_{75}^{m} 45.045.0. The iFormer-B backbone used 67M parameters and 351 GFLOPs and achieved bounding-box AP 48.348.3, AP50b_{50}^{b} 70.370.3, AP75b_{75}^{b} 53.253.2, mask AP 43.443.4, AP50m_{50}^{m} 67.267.2, and AP75m_{75}^{m} 46.746.7.

    For comparison, ResNet-50 achieved bounding-box AP 38.038.0 and mask AP 34.434.4 with 44M parameters and 260 GFLOPs; UniFormer-Sh14 achieved 45.645.6 and 41.641.6 with 41M parameters and 269 GFLOPs; UniFormer-B achieved 47.447.4 and 43.143.1 with 69M parameters and 399 GFLOPs; and Swin-S achieved 44.844.8 and 40.940.9 with 69M parameters and 354 GFLOPs. Thus, iFormer-S improved over ResNet-50 by 8.2 bounding-box AP points and 7.5 mask AP points, while iFormer-B exceeded UniFormer-B by 0.9 bounding-box AP points with fewer parameters and FLOPs.

  8. Knowl 8 — ADE20K semantic-segmentation performance

    empirical result

    The iFormer backbone was evaluated with Semantic FPN on ADE20K, which contains 20K training images and 2K validation images. ImageNet-pretrained backbones were trained with AdamW, initial learning rate 2×10−42\times10^{-4}, cosine learning-rate decay, and 80K iterations. FLOPs were measured at 512×2048512\times2048 resolution.

    iFormer-S used 24M parameters and 181 GFLOPs and achieved 48.6% mean intersection-over-union. The comparison backbones achieved the following mIoU values: ResNet-50, 36.7% with 29M parameters and 183 GFLOPs; PVT-S, 39.8% with 28M and 161 GFLOPs; Twins-S, 43.2% with 28M and 144 GFLOPs; UniFormer-Sh32, 46.2% with 25M and 199 GFLOPs; UniFormer-S, 46.6% with 25M and 247 GFLOPs; and UniFormer-B, 48.0% with 54M and 471 GFLOPs. Consequently, iFormer-S surpassed UniFormer-S by 2.0 mIoU points and UniFormer-B by 0.6 points while using roughly half as many parameters and nearly one-third as many FLOPs as UniFormer-B.

  9. Knowl 9 — Ablation of frequency branches and ramp allocation

    empirical result

    ImageNet ablations were trained for 100 epochs under the classification training setting. With the attention branch enabled, using attention alone achieved 81.2% top-1 accuracy at 20M parameters and 4.9 GFLOPs; replacing the max-pooling branch with the depthwise-convolution branch achieved 81.4% at the same size and compute; and enabling attention, max-pooling, and depthwise convolution together achieved 81.5% with 20M parameters and 4.8 GFLOPs. The combined mixer therefore performed best among the tested branch configurations.

    The channel-allocation ablation compared three strategies: decreasing Cl/CC_l/C while increasing Ch/CC_h/C achieved 80.5% top-1 accuracy with 19M parameters and 4.7 GFLOPs; equal high- and low-frequency ratios achieved 80.7% with 19M parameters and 4.7 GFLOPs; and increasing Cl/CC_l/C while decreasing Ch/CC_h/C achieved 81.2% with 20M parameters and 4.8 GFLOPs. The result supports assigning more high-frequency capacity to shallow layers and more low-frequency capacity to deep layers.

    Fourier visualizations showed that the self-attention branch concentrated more strongly on low frequencies, whereas max-pooling and depthwise convolution enhanced high-frequency components. Grad-CAM visualizations also showed that iFormer-S could localize objects more completely than Swin-T in the ImageNet examples used, including retaining the tail of a hummingbird rather than focusing only on part of the object.

  10. Knowl 10 — Limitations of manually specified frequency allocation

    limitation

    The iFormer frequency ramp requires manually selecting Ch/CC_h/C and Cl/CC_l/C for every iFormer block. Choosing effective ratios for different tasks requires design experience, and the paper does not provide an automatic allocation mechanism. Neural architecture search is suggested as a possible way to automate this choice. The models were also not trained on large-scale datasets such as ImageNet-21K because of computational constraints.

Coverage note — Secondary baseline rows, appendix-only architecture configurations, and additional non-headline benchmark details were omitted to stay within the ten most significant knowls.

References

  1. 1.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  4. 4.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  5. 5.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021.
  6. 6.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021.
  7. 7.Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022.
  8. 8.Drew A Hudson and Larry Zitnick. Generative adversarial transformers. In International Conference on Machine Learning, pages 4487–4499. PMLR, 2021.
  9. 9.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Luciˇ c, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6836–6846, 2021.
  10. 10.Josh Beal, Eric Kim, Eric Tzeng, Dong Huk Park, Andrew Zhai, and Dmitry Kislyuk. Toward transformer-based object detection. arXiv preprint arXiv:2012.09958, 2020.
  11. 11.Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. You only look at one sequence: Rethinking transformer in vision through object detection. Advances in Neural Information Processing Systems, 34, 2021.
  12. 12.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6881–6890, 2021.
  13. 13.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems, 34, 2021.
  14. 14.Namuk Park and Songkuk Kim. How do vision transformers work? In International Conference on Learning Representations, 2021.
  15. 15.Jean Bullier. Integrated model of visual processing. Brain research reviews, 36(2-3):96–107, 2001.
  16. 16.Moshe Bar. A cortical mechanism for triggering top-down facilitation in visual object recognition. Journal of cognitive neuroscience, 15(4):600–609, 2003.
  17. 17.Louise Kauffmann, Stephen Ramanoël, and Carole Peyrin. The neural bases of spatial frequency processing during scene perception. Frontiers in integrative neuroscience, 8:37, 2014.
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  19. 19.Haohan Wang, Xindi Wu, Zeyi Huang, and Eric P Xing. High-frequency component helps explain the generalization of convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8684–8694, 2020.
  20. 20.Dong Yin, Raphael Gontijo Lopes, Jon Shlens, Ekin Dogus Cubuk, and Justin Gilmer. A fourier perspective on model robustness in computer vision. Advances in Neural Information Processing Systems, 32, 2019.
  21. 21.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollár, and Ross Girshick. Early convolutions help transformers see better. Advances in Neural Information Processing Systems, 34:30392–30400, 2021.
  22. 22.Kunchang Li, Yali Wang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unified transformer for efficient spatiotemporal representation learning. arXiv preprint arXiv:2201.04676, 2022.
  23. 23.Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. Vitae: Vision transformer advanced by exploring intrinsic inductive bias. Advances in Neural Information Processing Systems, 34, 2021.
  24. 24.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. Advances in Neural Information Processing Systems, 34:3965–3977, 2021.
  25. 25.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22–31, 2021.
  26. 26.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9981–9990, 2021.
  27. 27.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34, 2021.
  28. 28.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  29. 29.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  30. 30.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. arXiv preprint arXiv:2201.03545, 2022.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  32. 32.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 633–641, 2017.
  33. 33.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  34. 34.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ S Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  35. 35.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  36. 36.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018.
  37. 37.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  38. 38.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 558–567, 2021.
  39. 39.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  40. 40.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.
  41. 41.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6881–6890, 2021.
  42. 42.Jieneng Chen, Yongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Yan Wang, Le Lu, Alan L Yuille, and Yuyin Zhou. Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306, 2021.
  43. 43.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  44. 44.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
  45. 45.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  46. 46.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.
  47. 47.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  48. 48.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
  49. 49.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  50. 50.Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3435–3444, 2019.
  51. 51.Zi-Hang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. All tokens matter: Token labeling for training better vision transformers. Advances in Neural Information Processing Systems, 34, 2021.
  52. 52.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. arXiv preprint arXiv:2204.01697, 2022.
  53. 53.Boyu Chen, Peixia Li, Chuming Li, Baopu Li, Lei Bai, Chen Lin, Ming Sun, Junjie Yan, and Wanli Ouyang. Glit: Neural architecture search for global and local image transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–21, 2021.
  54. 54.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, pages 1–10, 2022.
  55. 55.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12259–12269, 2021.
  56. 56.Yucheng Zhao, Guangting Wang, Chuanxin Tang, Chong Luo, Wenjun Zeng, and Zheng-Jun Zha. A battle of network structures: An empirical study of cnn, transformer, and mlp. arXiv preprint arXiv:2108.13002, 2021.
  57. 57.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015.
  58. 58.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  59. 59.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Thirty-first AAAI conference on artificial intelligence, 2017.
  60. 60.François Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017.
  61. 61.Franck Mamalet and Christophe Garcia. Simplifying convnets for fast learning. In International Conference on Artificial Neural Networks, pages 58–65. Springer, 2012.
  62. 62.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018.
  63. 63.Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm. arXiv preprint arXiv:2110.00476, 2021.
  64. 64.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao. Focal self-attention for local-global interactions in vision transformers. arXiv preprint arXiv:2107.00641, 2021.
  65. 65.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021.
  66. 66.Jiasen Lu, Roozbeh Mottaghi, Aniruddha Kembhavi, et al. Container: Context aggregation networks. Advances in Neural Information Processing Systems, 34, 2021.
  67. 67.Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao. Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond. arXiv preprint arXiv:2202.10108, 2022.
  68. 68.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10428–10436, 2020.
  69. 69.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16519–16529, 2021.
  70. 70.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  71. 71.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 32–42, 2021.
  72. 72.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  73. 73.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105–6114. PMLR, 2019.
  74. 74.Mingxing Tan and Quoc Le. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning, pages 10096–10106. PMLR, 2021.
  75. 75.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  76. 76.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems, 34, 2021.
  77. 77.Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2998–3008, 2021.
  78. 78.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  79. 79.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollár. Panoptic feature pyramid networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6399–6408, 2019.
  80. 80.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020.
  81. 81.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017.
  82. 82.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), pages 418–434, 2018.
  83. 83.Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng, Gang Yu, and Bin Fu. Shuffle transformer: Rethinking spatial shuffle for vision transformer. arXiv preprint arXiv:2106.03650, 2021.

Citation

MLA
Si, C., et al. “Inception Transformer”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 23495–509, https://proceedings.neurips.cc/paper_files/paper/2022/file/94e85561a342de88b559b72c9b29f638-Paper-Conference.pdf.
APA
Si, C., Yu, W., Zhou, P., Zhou, Y., Wang, X., & Yan, S. (2022). Inception Transformer. Advances in Neural Information Processing Systems, 35, 23495–23509. https://proceedings.neurips.cc/paper_files/paper/2022/file/94e85561a342de88b559b72c9b29f638-Paper-Conference.pdf
Chicago
Si, C., W. Yu, P. Zhou, Y. Zhou, X. Wang, and S. Yan. 2022. “Inception Transformer”. Advances in Neural Information Processing Systems 35: 23495–509. https://proceedings.neurips.cc/paper_files/paper/2022/file/94e85561a342de88b559b72c9b29f638-Paper-Conference.pdf.
Harvard
Si, C. et al. (2022) “Inception Transformer”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 23495–23509. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/94e85561a342de88b559b72c9b29f638-Paper-Conference.pdf.
Vancouver
1. Si C, Yu W, Zhou P, Zhou Y, Wang X, Yan S (2022) Inception Transformer. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 23495–23509

BibTeX

@inproceedings{si2022inception,
  title = {Inception Transformer},
  author = {Si, Chenyang and Yu, Weihao and Zhou, Pan and Zhou, Yichen and Wang, Xinchao and Yan, Shuicheng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {23495-23509},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/94e85561a342de88b559b72c9b29f638-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors