CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

Chun-Fu ChenQuanfu FanRameswar Panda

article2021ICCV2,245 citations

Introduces CrossViT, a dual-branch vision transformer that extracts multi-scale features from different patch sizes and fuses them using a linear-time cross-attention mechanism to achieve superior ImageNet classification accuracy over standard Vision Transformers with low computational overhead.

Listen

Visual recognition systems increasingly rely on vision transformers rather than conventional convolutional networks. However, standard vision transformers typically process images using single-sized patches, missing the multi-scale contextual information that made earlier vision architectures successful. Directly combining multiple patch granularities often creates prohibitive computational and memory bottlenecks due to the quadratic complexity of standard attention mechanisms.

To address this limitation, the article introduces CrossViT, a dual-branch vision transformer architecture designed to learn multi-scale visual representations. The objective was to demonstrate that combining coarse-grained and fine-grained image patch streams via an efficient cross-attention mechanism achieves higher classification accuracy without substantial increases in computational overhead.

Researchers evaluated the approach primarily on the ImageNet benchmark, testing several model configurations varying in depth, embedding dimensions, and patch embedding methods. The architecture uses a wider, deeper primary branch for coarse patches and a lighter, narrower complementary branch for fine patches. Rather than calculating full pairwise attention across all tokens, the system uses the summary classification token of each branch as an agent to exchange information with the other branch, reducing attention complexity from quadratic to linear time. The models were also tested across five downstream transfer learning tasks, including medical and fine-grained image datasets.

Key findings confirm that this dual-branch cross-attention strategy provides consistent accuracy gains across various model sizes. On the standard ImageNet benchmark, CrossViT models outperformed baseline models like DeiT by up to 1.2 to 2.0 percentage points with modest parameter increases. When incorporating convolutional patch embeddings, performance reached 82.8% to 84.1% top-1 accuracy at higher resolutions. Notably, the CrossViT-18 configuration achieved 82.8% accuracy while cutting floating-point operations and parameter counts nearly in half compared to base baseline models. Across downstream transfer tasks, the architecture retained competitive generalization without overfitting.

These results indicate that multi-scale representation is highly practical for vision transformers when information exchange is restricted to summary tokens. Organizations deploying vision models can achieve superior accuracy and throughput compared to traditional convolutional networks and earlier transformer models, lowering operational inference costs for demanding image recognition tasks.

Decision-makers should consider adopting dual-branch transformer architectures for production computer vision pipelines where fine detail and broad context are both essential. For maximal accuracy, practitioners should pair the architecture with convolutional embedding tokenizers. While the current study validates performance in general image classification, future work should extend and pilot these multi-scale mechanisms in dense prediction tasks, such as object detection, semantic segmentation, and video understanding.

Cover for CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

Abstract

The recently developed vision transformer (ViT) has achieved promising results on image classification compared to convolutional neural networks. Inspired by this, in this paper, we study how to learn multi-scale feature representations in transformer models for image classification. To this end, we propose a dual-branch transformer to combine image patches (i.e., tokens in a transformer) of different sizes to produce stronger image features. Our approach processes small-patch and large-patch tokens with two separate branches of different computational complexity and these tokens are then fused purely by attention multiple times to complement each other. Furthermore, to reduce computation, we develop a simple yet effective token fusion module based on cross attention, which uses a single token for each branch as a query to exchange information with other branches. Our proposed cross-attention only requires linear time for both computational and memory complexity instead of quadratic time otherwise. Extensive experiments demonstrate that our approach performs better than or on par with several concurrent works on vision transformer, in addition to efficient CNN models. For example, on the ImageNet1K dataset, with some architectural changes, our approach outperforms the recent DeiT by a large margin of 2% with a small to moderate increase in FLOPs and model parameters. Our source codes and models are available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Overview of Vision Transformer
  • 3.2 Proposed Multi-Scale Vision Transformer
  • 3.3 Multi-Scale Feature Fusion
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Ablation Studies
  • 5 Conclusion
  • References
  • A More Comparisons and Analysis

Knowls

  1. Knowl 1 — CrossViT Dual-Branch Multi-Scale Vision Transformer Architecture

    model/method

    The Cross-Attention Multi-Scale Vision Transformer (CrossViT) is an architecture for image classification that learns multi-scale feature representations by processing image patches of different sizes in two separate branches:

    1. Large Branch (L-Branch, Primary Branch): Operates on coarse-grained image patches of size PlP_l (e.g., 16×1616 \times 16). It uses a wider embedding dimension ClC_l and a deeper stack of MM standard transformer encoders to extract abstract, high-level representations.
    2. Small Branch (S-Branch, Complementary Branch): Operates on fine-grained image patches of size PsP_s (Ps<PlP_s < P_l, e.g., 12×1212 \times 12). It uses a narrower embedding dimension CsC_s and a shallower stack of NN transformer encoders (typically N=1N=1) to extract complementary, fine-grained representations with low computational overhead.

    The overall architecture stacks KK multi-scale transformer encoder blocks (typically K=3K=3). In each block:

    • Both branches independently process their respective patch tokens alongside a learnable classification (CLS) token using standard multi-head self-attention (MSA) and feed-forward networks (FFN).
    • At the end of the block, an efficient cross-attention module is applied LL times (typically L=1L=1) where the CLS token of each branch interacts with the patch tokens of the opposite branch.
    • After KK multi-scale blocks, classification is performed by feeding the final CLS tokens of both branches into separate multi-layer perceptron (MLP) heads and summing their prediction logits.
  2. Knowl 2 — Cross-Attention Multi-Scale Token Fusion Mechanism

    equation

    To fuse multi-scale features without quadratic computational overhead, CrossViT uses a cross-attention (CA) module where the classification token (CLS) of one branch acts as an information exchange agent with the patch tokens of the other branch.

    For the large branch (ll) receiving information from the small branch (ss), the input patch tokens xpatchs∈RNs×Csx_{\text{patch}}^s \in \mathbb{R}^{N_s \times C_s} and CLS token xclsl∈R1×Clx_{\text{cls}}^l \in \mathbb{R}^{1 \times C_l} are combined via a linear projection fl:RCl→RCsf^l: \mathbb{R}^{C_l} \to \mathbb{R}^{C_s} to align dimensions: x′l=[fl(xclsl)∥xpatchs]∈R(1+Ns)×Csx^{\prime l} = \left[ f^l(x_{\text{cls}}^l) \parallel x_{\text{patch}}^s \right] \in \mathbb{R}^{(1 + N_s) \times C_s} where ∥\parallel denotes concatenation along the sequence dimension.

    Cross-attention is computed using only the projected CLS token as the query: q=fl(xclsl)Wq,k=x′lWk,v=x′lWvq = f^l(x_{\text{cls}}^l) W_q, \quad k = x^{\prime l} W_k, \quad v = x^{\prime l} W_v A=softmax(qkTCs/h),CA(x′l)=AvA = \text{softmax}\left(\frac{q k^T}{\sqrt{C_s / h}}\right), \quad \text{CA}(x^{\prime l}) = A v where Wq,Wk,Wv∈RCs×(Cs/h)W_q, W_k, W_v \in \mathbb{R}^{C_s \times (C_s/h)} are learnable projection matrices, hh is the number of attention heads, and A∈R1×(1+Ns)A \in \mathbb{R}^{1 \times (1 + N_s)} is the cross-attention map.

    With multi-head cross-attention (MCA), layer normalization (LN), and a residual shortcut, the updated CLS token yclsly_{\text{cls}}^l and the output token sequence zlz^l are given by: yclsl=fl(xclsl)+MCA(LN([fl(xclsl)∥xpatchs]))y_{\text{cls}}^l = f^l(x_{\text{cls}}^l) + \text{MCA}\left(\text{LN}\left(\left[ f^l(x_{\text{cls}}^l) \parallel x_{\text{patch}}^s \right]\right)\right) zl=[gl(yclsl)∥xpatchl]z^l = \left[ g^l(y_{\text{cls}}^l) \parallel x_{\text{patch}}^l \right] where gl:RCs→RClg^l: \mathbb{R}^{C_s} \to \mathbb{R}^{C_l} is a linear back-projection function to restore the original channel dimension, and xpatchlx_{\text{patch}}^l is the untouched patch token sequence of the large branch. The small branch performs the symmetric operation by swapping indices ll and ss.

  3. Knowl 3 — Computational and Memory Complexity of Cross-Attention Fusion vs. All-Attention Fusion

    theoretical result

    Let NlN_l and NsN_s be the number of patch tokens in the large and small branches, respectively, and let CC denote the feature embedding dimension.

    • All-Attention Fusion: Concatenates all tokens from both branches (1+Nl+1+Ns1 + N_l + 1 + N_s tokens) and applies standard all-to-all self-attention. The generation of the attention map and self-attention operations require O((Nl+Ns)2C)\mathcal{O}((N_l + N_s)^2 C) computation and O((Nl+Ns)2)\mathcal{O}((N_l + N_s)^2) memory.
    • Cross-Attention Fusion: Uses only the single CLS token (11 token) from branch ll as the query to attend to the sequence [fl(xclsl)∥xpatchs][f^l(x_{\text{cls}}^l) \parallel x_{\text{patch}}^s] (1+Ns1 + N_s tokens). Symmetrically, the CLS token from branch ss attends to 1+Nl1 + N_l tokens.

    Consequently, the cross-attention map generation and value aggregation require O((Nl+Ns)C)\mathcal{O}((N_l + N_s) C) computation and O(Nl+Ns)\mathcal{O}(N_l + N_s) memory, achieving strictly linear time and memory complexity with respect to the total number of patch tokens.

  4. Knowl 4 — Alternative Dual-Branch Token Fusion Strategies

    model/method

    Four distinct fusion strategies for combining token sequences xl=[xclsl∥xpatchl]x^l = [x_{\text{cls}}^l \parallel x_{\text{patch}}^l] and xs=[xclss∥xpatchs]x^s = [x_{\text{cls}}^s \parallel x_{\text{patch}}^s] from large and small branches were formulated and compared:

    1. All-Attention Fusion: Tokens from both branches are linearly projected to match dimensions, concatenated into a single sequence y=[fl(xl)∥fs(xs)]y = [f^l(x^l) \parallel f^s(x^s)], processed through full multi-head self-attention with residual connection o=y+MSA(LN(y))o = y + \text{MSA}(\text{LN}(y)), partitioned into o=[ol∥os]o = [o^l \parallel o^s], and back-projected via zi=gi(oi)z^i = g^i(o^i).
    2. Class Token Fusion: Only the CLS tokens are merged by summing their projected representations, leaving patch tokens untouched: zi=[gi(∑j∈{l,s}fj(xclsj))∥xpatchi]z^i = \left[ g^i\left(\sum_{j \in \{l,s\}} f^j(x_{\text{cls}}^j)\right) \parallel x_{\text{patch}}^i \right]
    3. Pairwise Fusion: Spatial patch tokens are aligned in spatial resolution using 2D interpolation and summed along with the separately summed CLS tokens: zi=[gi(∑j∈{l,s}fj(xclsj))∥gi(∑j∈{l,s}fj(xpatchj))]z^i = \left[ g^i\left(\sum_{j \in \{l,s\}} f^j(x_{\text{cls}}^j)\right) \parallel g^i\left(\sum_{j \in \{l,s\}} f^j(x_{\text{patch}}^j)\right) \right]
    4. Cross-Attention Fusion: The CLS token of each branch serves as the query attending to the projected concatenation of itself and the other branch's patch tokens, after which it is back-projected to its own branch while patch tokens remain unchanged.
  5. Knowl 5 — Architectural Configurations of CrossViT Variants

    data/table

    CrossViT models are configured with K=3K=3 multi-scale transformer encoders, N=1N=1 encoder in the small branch, and L=1L=1 cross-attention module per block. The suffix †\dagger indicates that the linear patch embedding is replaced with a three-layer convolutional stem.

    Model Patch Embedding PsP_s PlP_l CsC_s ClC_l Heads MM rr
    CrossViT-Ti Linear 12 16 96 192 3 4 4
    CrossViT-S Linear 12 16 192 384 6 4 4
    CrossViT-B Linear 12 16 384 768 12 4 4
    CrossViT-9 Linear 12 16 128 256 4 3 3
    CrossViT-15 Linear 12 16 192 384 6 5 3
    CrossViT-18 Linear 12 16 224 448 7 6 3
    CrossViT-9†\dagger 3 Conv. 12 16 128 256 4 3 3
    CrossViT-15†\dagger 3 Conv. 12 16 192 384 6 5 3
    CrossViT-18†\dagger 3 Conv. 12 16 224 448 7 6 3

    Here, PsP_s and PlP_l are the patch sizes for the small and large branches, CsC_s and ClC_l are the respective embedding dimensions, MM is the number of transformer encoders per multi-scale block in the large branch (total depth of the large branch is K×MK \times M), and rr is the MLP expansion ratio in the feed-forward network (FFN).

  6. Knowl 6 — Ablation of Fusion Strategies on ImageNet-1K

    data/table

    An empirical comparison of token fusion schemes based on the CrossViT-S architecture evaluated on the ImageNet-1K validation set demonstrates that Cross-Attention achieves superior top-1 accuracy while balancing branch contributions.

    Fusion Scheme Top-1 Acc. (%) FLOPs (G) Params (M) L-Branch Acc. (%) S-Branch Acc. (%)
    None 80.2 5.3 23.7 80.2 0.1
    All-Attention 80.0 7.6 27.7 79.9 0.5
    Class Token 80.3 5.4 24.2 80.6 7.6
    Pairwise 80.3 5.5 24.2 80.3 7.3
    Cross-Attention 81.0 5.6 26.7 68.1 47.2

    In baseline fusion methods (None, All-Attention, Class Token, Pairwise), the primary L-Branch dominates predictions while the S-Branch exhibits near-zero standalone accuracy (0.1%−7.6%0.1\% - 7.6\%). Under Cross-Attention fusion, both branches develop non-trivial individual classification accuracy (68.1%68.1\% and 47.2%47.2\%), and their ensemble reaches 81.0%81.0\%, confirming complementary feature learning.

  7. Knowl 7 — ImageNet-1K Performance Comparison: CrossViT vs. DeiT Baselines

    data/table

    CrossViT models trained from scratch for 300 epochs on ImageNet-1K (input resolution 224×224224 \times 224) consistently outperform baseline Data-efficient Image Transformers (DeiT).

    Model Top-1 Acc. (%) FLOPs (G) Throughput (images/s) Params (M)
    DeiT-Ti 72.2 1.3 2557 5.7
    CrossViT-Ti 73.4 1.6 1668 6.9
    CrossViT-9 73.9 1.8 1530 8.6
    CrossViT-9†\dagger 77.1 2.0 1463 8.8
    DeiT-S 79.8 4.6 966 22.1
    CrossViT-S 81.0 5.6 690 26.7
    CrossViT-15 81.5 5.8 640 27.4
    CrossViT-15†\dagger 82.3 6.1 626 28.2
    DeiT-B 81.8 17.6 314 86.6
    CrossViT-B 82.2 21.2 239 104.7
    CrossViT-18 82.5 9.0 430 43.3
    CrossViT-18†\dagger 82.8 9.5 418 44.3

    CrossViT-18†\dagger achieves 82.8%82.8\% top-1 accuracy—a 1.0%1.0\% absolute gain over DeiT-B (81.8%81.8\%)—while reducing FLOPs from 17.6G to 9.5G (46%46\% reduction) and parameter count from 86.6M to 44.3M (49%49\% reduction).

  8. Knowl 8 — Ablation of Hyperparameters, Branch Asymmetry, and CLS Token Usage

    data/table

    Ablation experiments on ImageNet-1K using CrossViT-S evaluate variations in patch size pair (Ps,Pl)(P_s, P_l), small-branch embedding dimension CsC_s, small-branch depth NN, cross-attention depth LL, and multi-scale block count KK:

    Model PsP_s PlP_l CsC_s ClC_l KK NN MM LL Top-1 Acc. (%) FLOPs (G) Params (M)
    CrossViT-S 12 16 192 384 3 1 4 1 81.0 5.6 26.7
    A 8 16 192 384 3 1 4 1 80.8 6.7 26.7
    B 12 16 384 384 3 1 4 1 80.1 7.7 31.4
    C 12 16 192 384 3 2 4 1 80.7 6.3 28.0
    D 12 16 192 384 3 1 4 2 81.0 5.6 28.9
    E 12 16 192 384 6 1 2 1 80.9 6.6 31.1

    Key findings:

    • Patch size granularity: Setting (Ps,Pl)=(12,16)(P_s, P_l) = (12, 16) (a 2×2\times token ratio) outperforms (8,16)(8, 16) (80.8%80.8\%, 4×4\times token ratio) because an excessive scale disparity hinders feature fusion.
    • S-Branch capacity: Increasing S-branch width (Cs=384C_s=384, model B) or depth (N=2N=2, model C) increases complexity but reduces accuracy, confirming that the S-branch should remain a lightweight complement to the primary L-branch.
    • Fusion frequency: Stacking additional cross-attention layers (L=2L=2, model D) or doubling multi-scale encoders (K=6,M=2K=6, M=2, model E) yields no accuracy gain.
    • CLS Token necessity: Replacing the CLS token with the mean of patch tokens as the query in cross-attention causes top-1 accuracy to drop from 81.0%81.0\% to 80.0%80.0\%.
  9. Knowl 9 — Comparison with State-of-the-Art Transformers and CNN Backbones

    data/table

    CrossViT demonstrates competitive accuracy, throughput (measured at batch size 64 on an NVIDIA Tesla V100 GPU), and parameter efficiency relative to concurrent vision transformers and convolutional neural networks on ImageNet-1K:

    Model Top-1 Acc. (%) FLOPs (G) Throughput (img/s) Params (M)
    PVT-S 79.8 3.8 - 24.5
    T2T-ViT-14 80.7 6.1 - 21.5
    TNT-S 81.3 5.2 - 23.8
    CrossViT-15†\dagger 82.3 6.1 626 28.2
    ViT-B@384 77.9 17.6 - 86.6
    T2T-ViT-24 82.2 15.0 - 64.1
    TNT-B 82.8 14.1 - 65.6
    CrossViT-18†\dagger 82.8 9.5 418 44.3
    ResNet-152 77.0 11.5 445 60.2
    ResNeXt-101-64x4d 79.6 15.5 289 83.5
    RegNetY-16GF 80.4 15.9 336 83.6
    SENet-154 81.3 20.7 201 115.1
    EfficientNet-B4@380 82.9 4.2 356 19.0
    CrossViT-18†\dagger@384 83.9 32.4 112 44.6
    CrossViT-18†\dagger@480 84.1 56.6 57 44.9

    CrossViT-15†\dagger outperforms all ResNet, ResNeXt, RegNet, and SENet variants while maintaining higher inference throughput. At higher input resolutions (384×384384 \times 384 and 480×480480 \times 480), CrossViT-18†\dagger achieves 83.9%83.9\% and 84.1%84.1\%, comparable to EfficientNet-B6 (84.0%84.0\%) and B7 (84.3%84.3\%).

  10. Knowl 10 — Transfer Learning Performance on Downstream Classification Benchmarks

    data/table

    To assess representation generalization, CrossViT models pre-trained on ImageNet-1K were fine-tuned for 1,000 epochs on five downstream image classification datasets: CIFAR-10, CIFAR-100, Oxford-IIIT Pet, CropDiseases (plant pathology), and ChestXRay8 (medical thoracic radiography):

    Model CIFAR10 CIFAR100 Pet CropDiseases ChestXRay8
    DeiT-S 99.15 90.89 94.93 99.96 55.39
    DeiT-B 99.10 90.80 94.39 99.96 55.77
    CrossViT-15 99.00 90.77 94.55 99.97 55.89
    CrossViT-18 99.11 91.36 95.07 99.97 55.94

    CrossViT matches or slightly surpasses DeiT on natural image datasets (reaching 91.36%91.36\% on CIFAR-100 and 95.07%95.07\% on Pet) and achieves higher AUC/accuracy on the medical ChestXRay8 benchmark (55.94%55.94\% vs. 55.77%55.77\%), demonstrating strong generalization across distinct domains.

Coverage note — No substantial contributed material was omitted; the extracted knowls comprehensively cover the dual-branch CrossViT architecture, mathematical cross-attention formulation, complexity analysis, model configurations, ablation studies, baseline comparisons, downstream transfer learning, and hybrid convolutional/T2T tokenization integration.

References

  1. 1.Edward H Adelson, Charles H Anderson, James R Bergen, Peter J Burt, and Joan M Ogden. Pyramid methods in image processing. RCA engineer, 29(6):33–41, 1984. 2
  2. 2.Irwan Bello. Lambdanetworks: Modeling long-range interactions without attention. In International Conference on Learning Representations, 2021. 2
  3. 3.Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le. Attention augmented convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3286–3295, 2019. 1, 2
  4. 4.Zhaowei Cai, Quanfu Fan, Rogerio S Feris, and Nuno Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In European conference on computer vision, pages 354–370. Springer, 2016. 2
  5. 5.Chun-Fu (Richard) Chen, Quanfu Fan, Neil Mallinar, Tom Sercu, and Rogerio Feris. Big-Little Net: An Efficient MultiScale Feature Representation for Visual and Speech Recognition. In International Conference on Learning Representations, 2019. 2
  6. 6.Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3435–3444, 2019. 2
  7. 7.Bowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi, Thomas S. Huang, and Lei Zhang. Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2020. 2
  8. 8.Ekin Dogus Cubuk, Barret Zoph, Jon Shlens, and Quoc Le. RandAugment: Practical Automated Data Augmentation with a Reduced Search Space. In H Larochelle, M Ranzato, R Hadsell, M F Balcan, and H Lin, editors, Advances in Neural Information Processing Systems, pages 18613–18624. Curran Associates, Inc., 2020. 5
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 1, 5
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. 1, 3
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. 1, 2, 3, 5, 6, 7
  12. 12.Quanfu Fan, Chun-Fu Richard Chen, Hilde Kuehne, Marco Pistoia, and David Cox. More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation. In Advances in Neural Information Processing Systems, pages 2261–2270, 2019. 2
  13. 13.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE International Conference on Computer Vision, pages 6202–6211, 2019. 2
  14. 14.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. arXiv preprint arXiv:2103.00112, 2021. 1, 7
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In The IEEE Conference on Computer Vision and Pattern Recognition, June 2016. 1, 7
  16. 16.Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, and Daniel Soudry. Augment your batch: Improving generalization through instance repetition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2020. 5
  17. 17.Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. Local relation networks for image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3464–3473, 2019. 2
  18. 18.J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7132–7141, 2018. 2, 7
  19. 19.Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. Perceiver: General perception with iterative attention. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 4651–4664. PMLR, 18–24 Jul 2021. 1, 2, 7
  20. 20.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5, 7
  21. 21.Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang. Selective kernel networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2019. 2
  22. 22.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017. 2
  23. 23.Sharada P Mohanty, David P Hughes, and Marcel Salathe. Using deep learning for image-based plant disease detection. Frontiers in plant science, 7:1419, 2016. 7
  24. 24.Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. Deep multi-scale convolutional neural network for dynamic scene deblurring. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, July 2017. 2
  25. 25.Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked Hourglass Networks for Human Pose Estimation. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, Proceedings of the European Conference on Computer Vision, pages 483–499, Cham, 2016. Springer International Publishing. 2
  26. 26.Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. In European conference on computer vision, pages 483–499. Springer, 2016. 2
  27. 27.Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. Cats and dogs. In IEEE Conference on Computer Vision and Pattern Recognition, 2012. 7
  28. 28.Marco Pedersoli, Andrea Vedaldi, Jordi Gonzalez, and Xavier Roca. A coarse-to-fine approach for fast deformable object detection. Pattern Recognition, 48(5):1844–1853, 2015. 2
  29. 29.Pietro Perona and Jitendra Malik. Scale-space and edge detection using anisotropic diffusion. IEEE Transactions on pattern analysis and machine intelligence, 12(7):629–639, 1990. 2
  30. 30.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2020. 7
  31. 31.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. Stand-Alone SelfAttention in Vision Models. In H Wallach, H Larochelle, A Beygelzimer, F d Alch e Buc, E Fox, and R Garnett, editors, Advances in Neural Information Processing Systems. Curran Associates, Inc., 2019. 1, 2
  32. 32.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. arXiv preprint arXiv:2101.11605, 2021. 1, 2
  33. 33.C. Sun, A. Shrivastava, S. Singh, and A. Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In 2017 IEEE International Conference on Computer Vision, pages 843–852, 2017. 1, 5
  34. 34.Mingxing Tan and Quoc Le. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, pages 6105–6114, Long Beach, California, USA, June 2019. PMLR. 1, 2, 5, 7
  35. 35.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 10347–10357. PMLR, 18–24 Jul 2021. 1, 2, 5, 6, 7
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is All you Need. In I Guyon, U V Luxburg, S Bengio, H Wallach, R Fergus, S Vishwanathan, and R Garnett, editors, Advances in Neural Information Processing Systems. Curran Associates, Inc., 2017. 1, 2
  37. 37.Qilong Wang, Banggu Wu, Pengfei Zhu, Peihua Li, Wangmeng Zuo, and Qinghua Hu. Eca-net: Efficient channel attention for deep convolutional neural networks. In The IEEE Conference on Computer Vision and Pattern Recognition, 2020. 2, 7
  38. 38.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions, 2021. 1, 2, 6, 7
  39. 39.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018. 2
  40. 40.Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M Summers. Chestxray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2097–2106, 2017. 7
  41. 41.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. Cbam: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision, September 2018. 2
  42. 42.Lemeng Wu, Xingchao Liu, and Qiang Liu. Centroid transformers: Learning to abstract with attention. arXiv preprint arXiv:2102.08606, 2021. 2, 7
  43. 43.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated Residual Transformations for Deep Neural Networks. In The IEEE Conference on Computer Vision and Pattern Recognition, July 2017. 7
  44. 44.Songfan Yang and Deva Ramanan. Multi-scale recognition with dag-cnns. In Proceedings of the IEEE international conference on computer vision, pages 1215–1223, 2015. 2
  45. 45.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet, 2021. 1, 2, 6, 7, 8
  46. 46.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019. 5
  47. 47.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, 2018. 5
  48. 48.Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. Exploring self-attention for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2020. 1, 2
  49. 49.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random Erasing Data Augmentation. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07):13001–13008, Apr. 2020. 5
  50. 50.Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2018. 7

Citation

MLA
Chen, C.-F. R., et al. “CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification”. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 347–56, https://doi.org/10.1109/ICCV48922.2021.00041.
APA
Chen, C.-F. R., Fan, Q., & Panda, R. (2021). CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 347–356. https://doi.org/10.1109/ICCV48922.2021.00041
Chicago
Chen, C.-F. R., Q. Fan, and R. Panda. 2021. “CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification”. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 347–56. https://doi.org/10.1109/ICCV48922.2021.00041.
Harvard
Chen, C.-F.R., Fan, Q. and Panda, R. (2021) “CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification”, 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp. 347–356. Available at: https://doi.org/10.1109/ICCV48922.2021.00041.
Vancouver
1. Chen C-FR, Fan Q, Panda R (2021) CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp 347–356

BibTeX

@inproceedings{Chen_2021, title={CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification}, url={http://dx.doi.org/10.1109/ICCV48922.2021.00041}, DOI={10.1109/iccv48922.2021.00041}, booktitle={2021 IEEE/CVF International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Chen, Chun-Fu Richard and Fan, Quanfu and Panda, Rameswar}, year={2021}, month=Oct, pages={347–356} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE