CF-ViT: A General Coarse-to-Fine Method for Vision Transformer

Mengzhao ChenMingbao LinKe LiYunhang ShenYongjian WuFei ChaoRongrong Ji

article2023AAAI102 citations

Proposes a two-stage vision transformer inference framework that processes easy images using coarse-grained patches and selectively re-splits only informative regions into fine-grained tokens for hard images, cutting computational cost by more than half while doubling throughput without sacrificing accuracy.

Listen

Vision Transformer models achieve outstanding accuracy across computer vision tasks, but their high computational cost poses significant challenges for practical deployment. Processing an image requires dividing it into visual patches (tokens), and the computational burden grows quadratically with the number of tokens. In practice, many images contain substantial spatial redundancy, such as empty backgrounds, and most standard images do not require dense analysis to be classified correctly. The article develops and evaluates a coarse-to-fine framework, termed CF-ViT, designed to dynamically allocate computational effort based on image difficulty and regional importance.

The framework operates in two sequential stages using a single shared neural network. In the first stage, the model divides an input image into a small number of coarse patches, enabling low-cost initial classification. If the model achieves high classification confidence, inference terminates immediately. If confidence falls below an adjustable threshold, the model identifies the most informative image regions using an attention-tracking mechanism across network layers. Only these critical regions are re-split into finer patches for a second inference pass, while coarse representations are reused to preserve local context without adding extra parameters.

Evaluation on the standard ImageNet benchmark demonstrates substantial computational savings and speed improvements without sacrificing recognition performance. When applied to standard baseline models, the coarse-to-fine approach cut floating-point operations by 53% to 61% while maintaining baseline accuracy, effectively doubling image processing throughput on standard hardware (up to a 2.01-fold increase). Furthermore, when calibrated for maximum accuracy rather than maximum speed, the method outperformed standard baselines by up to 1.0% in classification accuracy while still consuming less computation. The approach also consistently outperformed existing dynamic token-pruning and early-exit methods across comparable computational budgets.

These results provide actionable implications for enterprise AI systems, edge deployments, and cloud-scale visual processing. By enabling flexible trade-offs between processing latency and accuracy via a single threshold parameter, the method reduces infrastructure hosting costs and hardware requirements. Unlike prior multi-stage approaches that require storing multiple separate models in memory, this single-model architecture minimizes storage and operational overhead.

Organizations deploying vision transformer models should consider adopting dynamic coarse-to-fine token allocation strategies to optimize system throughput and operational costs. Future initiatives should evaluate extending this dynamic processing strategy beyond standard image classification to dense visual tasks such as object detection and semantic segmentation. Users should note that while results are highly consistent on standard image recognition benchmarks, performance under real-world shifts in data complexity and hardware execution environments requires domain-specific validation.

Cover for CF-ViT: A General Coarse-to-Fine Method for Vision Transformer

Abstract

Vision Transformers (ViT) have made many breakthroughs in computer vision tasks. However, considerable redundancy arises in the spatial dimension of an input image, leading to massive computational costs. Therefore, We propose a coarse-to-fine vision transformer (CF-ViT) to relieve computational burden while retaining performance in this paper. Our proposed CF-ViT is motivated by two important observations in modern ViT models: (1) The coarse-grained patch splitting can locate informative regions of an input image. (2) Most images can be well recognized by a ViT model in a small-length token sequence. Therefore, our CF-ViT implements network inference in a two-stage manner. At coarse inference stage, an input image is split into a small-length patch sequence for a computationally economical classification. If not well recognized, the informative patches are identified and further re-split in a fine-grained granularity. Extensive experiments demonstrate the efficacy of our CF-ViT. For example, without any compromise on performance, CF-ViT reduces 53% FLOPs of LV-ViT, and also achieves 2.01× throughput. Code of this project is at https://github.com/ChenMnZ/CF-ViT.

Table of Contents

  • Introduction
  • Related Work
  • Vision Transformer
  • ViT Compression
  • Preliminaries
  • Coarse-to-Fine Vision Transformer
  • Coarse Inference Stage
  • Fine Inference Stage
  • Training Strategy
  • Experiments
  • Implementation Details
  • Experimental Results
  • Ablation Study
  • Visualization
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Coarse-to-Fine Vision Transformer (CF-ViT) Framework

    model/method

    Coarse-to-Fine Vision Transformer (CF-ViT) is an adaptive inference framework designed to reduce the computational redundancy of Vision Transformers (ViTs) by adjusting spatial token resolutions based on input image complexity. CF-ViT uses a shared transformer backbone VV with KK sequentially stacked encoders and a classification head FF operating across two stages:

    1. Coarse Inference Stage: The input image is partitioned into a small number of coarse patches NcN_c (e.g., 7×7=497\times 7 = 49). The coarse token sequence is constructed as: X0c=[x00;x01;… ;x0Nc]+EposcX_0^c = [x_0^0; x_0^1; \dots; x_0^{N_c}] + E_{pos}^c where x00∈RDx_0^0 \in \mathbb{R}^D is the [class][\text{class}] token, x0i∈RDx_0^i \in \mathbb{R}^D (i>0i > 0) is the linear patch embedding of the ii-th coarse patch, EposcE_{pos}^c denotes coarse learnable position embeddings, and DD is the embedding dimension. After passing through the transformer encoders to yield output [xK0;xK1;… ;xKNc][x_K^0; x_K^1; \dots; x_K^{N_c}], the coarse class prediction distribution is computed as pc=F(xK0)=[p1c,p2c,…,pnc]p^c = F(x_K^0) = [p_1^c, p_2^c, \dots, p_n^c] for nn target classes. If the maximum predicted confidence score pjc=max⁡ipicp_j^c = \max_i p_i^c satisfies pjc≥ηp_j^c \ge \eta for a predetermined threshold η∈[0,1]\eta \in [0, 1], inference terminates immediately with predicted class jj.

    2. Fine Inference Stage: If pjc<ηp_j^c < \eta, the image is deemed difficult. An informative region identification module identifies the top-scoring informative coarse patches, which are re-split into 2×22\times 2 finer sub-patches while uninformative regions are kept at the coarse scale. A feature reuse module injects the coarse-stage representations into the fine token sequence. The resulting token sequence is processed by the same transformer backbone VV to compute the final fine classification distribution pf=F(x~K0)p^f = F(\tilde{x}_K^0).

    To share parameters between coarse and fine stages in the initial linear projection layer, coarse patches are downsampled to the spatial dimensions of fine patches.

  2. Knowl 2 — Informative Region Identification via Global Class Attention

    model/method

    In CF-ViT, informative coarse patches are identified using class attention accumulated across transformer layers. In the kk-th encoder (k∈{1,…,K}k \in \{1, \dots, K\}), the self-attention matrix is computed as: Ak=Softmax(QkKkTD)=[ak0;ak1;… ;akNc]A_k = \text{Softmax}\left(\frac{Q_k K_k^T}{\sqrt{D}}\right) = [a_k^0; a_k^1; \dots; a_k^{N_c}] where Qk,Kk∈R(Nc+1)×DQ_k, K_k \in \mathbb{R}^{(N_c+1) \times D} are queries and keys, DD is the token feature dimension, and ak0∈RNc+1a_k^0 \in \mathbb{R}^{N_c+1} denotes the class attention vector reflecting interactions between the [class][\text{class}] token and patch tokens.

    To overcome layer-wise attention instability in shallow encoders, CF-ViT computes a Global Class Attention (GCA) score aˉk\bar{a}_k via an exponential moving average (EMA) across encoders starting from the 4th encoder up to the final KK-th encoder: aˉk=β⋅aˉk−1+(1−β)⋅ak0\bar{a}_k = \beta \cdot \bar{a}_{k-1} + (1 - \beta) \cdot a_k^0 where β=0.99\beta = 0.99. Coarse patches corresponding to the top-⌈Ncα⌉\lceil N_c \alpha \rceil highest scores in the final accumulated attention aˉK\bar{a}_K are designated as informative, where α∈[0,1]\alpha \in [0, 1] represents the informative patch ratio (default α=0.5\alpha = 0.5). Each informative patch is split into 2×22\times 2 fine-grained sub-patches, while the remaining uninformative coarse patches are kept intact. The total number of patch tokens in the fine stage NfN_f is: Nf=4⌈Ncα⌉+⌊Nc(1−α)⌋N_f = 4\lceil N_c \alpha \rceil + \lfloor N_c(1 - \alpha) \rfloor where ⌈⋅⌉\lceil \cdot \rceil and ⌊⋅⌋\lfloor \cdot \rfloor denote ceiling and floor functions.

  3. Knowl 3 — Feature Reuse Mechanism in CF-ViT

    model/method

    To preserve the integral contextual information of coarse patches that are subdivided during the fine inference stage, CF-ViT introduces a Feature Reuse (FR) module. The FR module processes the coarse-stage encoder output token sequence [xK1;xK2;… ;xKNc]∈RNc×D[x_K^1; x_K^2; \dots; x_K^{N_c}] \in \mathbb{R}^{N_c \times D}:

    1. Transformation: The coarse output image tokens are passed through a multi-layer perceptron (MLP) to generate transformed features.
    2. Replication and Masking: Each transformed token corresponding to an identified informative coarse patch is replicated 4×4\times (one replica for each of the 2×22\times 2 fine sub-patches). Tokens corresponding to the [class][\text{class}] token and uninformative coarse patches are replaced with zero vectors (zero-padding).
    3. Shortcut Injection: The resulting feature tensor Xr=FR([xK1;xK2;… ;xKNc])∈R(Nf+1)×DX_r = \text{FR}([x_K^1; x_K^2; \dots; x_K^{N_c}]) \in \mathbb{R}^{(N_f+1) \times D} is added element-wise to the fine-stage input token sequence: V(X~0f+Xr)=[x~K0;x~K1;… ;x~KNf]V(\tilde{X}_0^f + X_r) = [\tilde{x}_K^0; \tilde{x}_K^1; \dots; \tilde{x}_K^{N_f}] where X~0f=[x00;x~01;… ;x~0Nf]+Eposf\tilde{X}_0^f = [x_0^0; \tilde{x}_0^1; \dots; \tilde{x}_0^{N_f}] + E_{pos}^f, EposfE_{pos}^f are fine learnable position embeddings, and [x~K0;… ;x~KNf][\tilde{x}_K^0; \dots; \tilde{x}_K^{N_f}] are the fine-stage encoder outputs.
  4. Knowl 4 — CF-ViT Dynamic Inference Procedure

    algorithm

    The two-stage dynamic inference algorithm for CF-ViT evaluates an input image using adaptive spatial resolution controlled by a confidence threshold η\eta.

    Input: Image II, vision transformer backbone VV with KK encoders, classifier FF, confidence threshold η∈[0,1]\eta \in [0, 1], informative ratio α∈[0,1]\alpha \in [0, 1], smoothing factor β=0.99\beta = 0.99
    Output: Predicted class label jj
    Partition image II into NcN_c coarse patches
    Embed coarse patches and append [class][\text{class}] token to form X0c=[x00;x01;… ;x0Nc]+EposcX_0^c = [x_0^0; x_0^1; \dots; x_0^{N_c}] + E_{pos}^c
    Pass X0cX_0^c through encoders k=1,…,Kk = 1, \dots, K of VV
    for k=4k = 4 to KK do
        Extract class attention ak0a_k^0 from encoder kk
        if k==4k == 4 then
            aˉk=ak0\bar{a}_k = a_k^0
        else
            aˉk=β⋅aˉk−1+(1−β)⋅ak0\bar{a}_k = \beta \cdot \bar{a}_{k-1} + (1 - \beta) \cdot a_k^0
        end if
    end for
    Extract coarse [class][\text{class}] token xK0x_K^0 and coarse patch tokens [xK1;… ;xKNc][x_K^1; \dots; x_K^{N_c}]
    Compute coarse probability distribution pc=F(xK0)p^c = F(x_K^0)
    j=arg⁡max⁡ipicj = \arg\max_i p_i^c
    if pjc≥ηp_j^c \ge \eta then
        return jj
    end if
    Identify top-⌈Ncα⌉\lceil N_c \alpha \rceil coarse patches with largest values in aˉK\bar{a}_K
    Split informative patches into 2×22\times 2 sub-patches; retain uninformative patches at coarse scale
    Embed fine patch sequence and append [class][\text{class}] token to form X~0f\tilde{X}_0^f
    Compute feature reuse tensor Xr=FR([xK1;… ;xKNc])X_r = \text{FR}([x_K^1; \dots; x_K^{N_c}]) via MLP, 4×4\times replication, and zero-padding
    Pass (X~0f+Xr)(\tilde{X}_0^f + X_r) through VV to obtain fine [class][\text{class}] token x~K0\tilde{x}_K^0
    Compute fine probability distribution pf=F(x~K0)p^f = F(\tilde{x}_K^0)
    return arg⁡max⁡ipif\arg\max_i p_i^f
  5. Knowl 5 — CF-ViT Training Objective and Curriculum Strategy

    model/method

    During the training of CF-ViT, the early-termination confidence threshold is fixed to η=1\eta = 1, ensuring both coarse and fine inference stages are executed for every training image. The total training objective combines supervised classification loss on the fine prediction with knowledge distillation from the fine stage to the coarse stage: L=CE(pf,y)+KL(pc,pf)\mathcal{L} = \text{CE}(p^f, y) + \text{KL}(p^c, p^f) where CE(⋅,⋅)\text{CE}(\cdot, \cdot) is the standard cross-entropy loss between fine-stage prediction distribution pfp^f and the one-hot ground-truth label yy, and KL(pc,pf)\text{KL}(p^c, p^f) is the Kullback-Leibler divergence matching the coarse-stage prediction distribution pcp^c to the fine-stage prediction distribution pfp^f.

    To prevent convergence difficulties caused by selecting informative regions early in training, a curriculum strategy is used: all coarse patches across the entire image are split into fine-grained patches during the first 200 epochs. Selective informative patch splitting is activated only for the remaining training epochs.

  6. Knowl 6 — ImageNet Classification Performance and Efficiency of CF-ViT

    data/table

    CF-ViT was evaluated on the ImageNet validation set (50,000 images) using DeiT-S and LV-ViT-S backbones. Coarse inference used Nc=49N_c = 49 (7×77\times 7) patches and fine inference used Nf=124N_f = 124 patches (α=0.5\alpha = 0.5). Model throughput was measured in images per second on a single NVIDIA A100 GPU with batch size 1024.

    Model η\eta Top-1 Acc. (%) FLOPs (G) Throughput (img./s)
    DeiT-S baseline - 79.8 4.6 2601
    CF-ViT (DeiT-S) 0.5 79.8 (+0.0) 1.8 (↓\downarrow 61%) 4903 (↑\uparrow 1.88×\times)
    CF-ViT (DeiT-S) 0.75 80.7 (+0.9) 2.6 (↓\downarrow 43%) 3701 (↑\uparrow 1.32×\times)
    CF-ViT (DeiT-S) 1.0 80.8 (+1.0) 4.0 (↓\downarrow 13%) 2760 (↑\uparrow 1.06×\times)
    LV-ViT-S baseline - 83.3 6.6 1681
    CF-ViT (LV-ViT-S) 0.63 83.3 (+0.0) 3.1 (↓\downarrow 53%) 3393 (↑\uparrow 2.01×\times)
    CF-ViT (LV-ViT-S) 0.75 83.5 (+0.2) 4.0 (↓\downarrow 39%) 2827 (↑\uparrow 1.68×\times)
    CF-ViT (LV-ViT-S) 1.0 83.6 (+0.3) 6.1 (↓\downarrow 7%) 2022 (↑\uparrow 1.31×\times)

    At parity with baseline Top-1 accuracy, CF-ViT reduced DeiT-S FLOPs by 61% (1.88×\times throughput) and LV-ViT-S FLOPs by 53% (2.01×\times throughput). When evaluating all images through the fine stage (η=1.0\eta = 1.0), CF-ViT improved Top-1 accuracy over the baseline by 1.0% on DeiT-S and 0.3% on LV-ViT-S with fewer FLOPs.

  7. Knowl 7 — Comparison of CF-ViT with Token Slimming ViT Compression Baselines

    data/table

    CF-ViT was benchmarked against existing token slimming methods on ImageNet with DeiT-S and LV-ViT-S backbones.

    Model Top-1 Acc. (%) FLOPs (G)
    DeiT-S Backbones
    Baseline (DeiT-S) 79.8 4.6
    DynamicViT 79.3 2.9
    IA-RED2^2 79.1 3.2
    PS-ViT 79.4 2.6
    EViT 79.5 3.0
    Evo-ViT 79.4 3.0
    CF-ViT (η=0.5\eta = 0.5) (Ours) 79.8 1.8
    CF-ViT (η=0.75\eta = 0.75) (Ours) 80.7 2.6
    LV-ViT-S Backbones
    Baseline (LV-ViT-S) 83.3 6.6
    DynamicViT 83.0 4.6
    EViT 83.0 4.7
    SiT 83.2 4.0
    CF-ViT (η=0.63\eta = 0.63) (Ours) 83.3 3.1
    CF-ViT (η=0.75\eta = 0.75) (Ours) 83.5 4.0

    CF-ViT achieved higher Top-1 accuracy at significantly lower FLOPs compared to existing dynamic token pruning and slimming methods. On DeiT-S, CF-ViT achieved 79.8% Top-1 accuracy at 1.8G FLOPs, outperforming Evo-ViT (79.4% at 3.0G FLOPs) and PS-ViT (79.4% at 2.6G FLOPs).

  8. Knowl 8 — Ablation Studies on CF-ViT Components and Loss Functions

    empirical result

    Ablation experiments were conducted with DeiT-S on ImageNet with early termination disabled (evaluating all samples at coarse and fine stages):

    1. Informative Region Identification Strategy:

      • Global Class Attention (GCA, Ours): coarse 75.5%, fine 80.8% Top-1 accuracy.
      • Last class attention only: coarse 75.3%, fine 80.3% Top-1 accuracy.
      • Random selection: coarse 75.3%, fine 79.6% Top-1 accuracy.
      • Negative GCA (selecting lowest-attention regions): coarse 74.9%, fine 77.6% Top-1 accuracy.
    2. Feature Reuse Module Design:

      • Default FR (MLP on informative tokens only): coarse 75.5%, fine 80.8% Top-1 accuracy.
      • Without feature reuse (w/o reuse): coarse 75.2%, fine 80.0% Top-1 accuracy.
      • FR including [class][\text{class}] token: coarse 75.4%, fine 80.2% Top-1 accuracy.
      • FR including uninformative tokens: coarse 75.4%, fine 80.6% Top-1 accuracy.
      • Linear layer instead of MLP: coarse 75.3%, fine 80.6% Top-1 accuracy.
    3. Training Loss Formulation:

      • CE(pf,y)+KL(pc,pf)\text{CE}(p^f, y) + \text{KL}(p^c, p^f) (Ours): coarse 75.5%, fine 80.8% Top-1 accuracy.
      • CE(pf,y)+CE(pc,y)\text{CE}(p^f, y) + \text{CE}(p^c, y): coarse 75.7%, fine 80.3% Top-1 accuracy.

    Using GCA over single-layer attention, restricting feature reuse strictly to informative tokens with an MLP transformation, and distilling fine predictions to coarse predictions via KL divergence consistently provided optimal fine-stage accuracy.

  9. Knowl 9 — Sensitivity of CF-ViT to Informative Ratio and Attention Smoothing Parameter

    data/table

    Ablation experiments evaluated the sensitivity of fine inference accuracy and computational cost to the informative patch ratio α\alpha and the global class attention EMA parameter β\beta on DeiT-S without early exit:

    Metric α=0.4\alpha = 0.4 α=0.5\alpha = 0.5 (default) α=0.6\alpha = 0.6 α=0.7\alpha = 0.7 α=0.8\alpha = 0.8 α=0.9\alpha = 0.9
    Top-1 Acc. (%) 80.4 80.8 80.9 81.1 81.3 81.4
    FLOPs (G) 3.7 4.0 4.4 4.7 5.1 5.5
    Metric β=0\beta = 0 β=0.5\beta = 0.5 β=0.9\beta = 0.9 β=0.99\beta = 0.99 (default) β=0.999\beta = 0.999
    Top-1 Acc. (%) 80.3 80.5 80.7 80.8 80.8

    Increasing α\alpha improves fine-stage accuracy monotonically from 80.4% to 81.4% but increases FLOPs from 3.7G to 5.5G; α=0.5\alpha = 0.5 provides the selected accuracy-computation trade-off. Increasing β\beta from 0 to 0.99 improves accuracy from 80.3% to 80.8%, demonstrating the benefit of smoothing attention weights across encoder layers.

Coverage note — Qualitative visual attention patch visualizations (Figure 1b and Figure 9) were omitted as their conclusions are quantitatively represented in the ablation and performance knowls.

References

  1. 1.Chen, C.-F. R.; Fan, Q.; and Panda, R. 2021. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 357–366.
  2. 2.Chu, X.; Tian, Z.; Wang, Y.; Zhang, B.; Ren, H.; Wei, X.; Xia, H.; and Shen, C. 2021a. Twins: Revisiting the design of spatial attention in vision transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  3. 3.Chu, X.; Tian, Z.; Zhang, B.; Wang, X.; Wei, X.; Xia, H.; and Shen, C. 2021b. Conditional positional encodings for vision transformers. arXiv:2102.10882.
  4. 4.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 248–255.
  5. 5.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR).
  6. 6.Fang, J.; Xie, L.; Wang, X.; Zhang, X.; Liu, W.; and Tian, Q. 2021. Msg-transformer: Exchanging local spatial information by manipulating messenger tokens. arXiv:2105.15168.
  7. 7.Han, K.; Wang, Y.; Chen, H.; Chen, X.; Guo, J.; Liu, Z.; Tang, Y.; Xiao, A.; Xu, C.; Xu, Y.; et al. 2022a. A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
  8. 8.Han, K.; Xiao, A.; Wu, E.; Guo, J.; Xu, C.; and Wang, Y. 2021a. Transformer in transformer. In Advances in Neural Information Processing Systems (NeurIPS).
  9. 9.Han, Y.; Huang, G.; Song, S.; Yang, L.; Wang, H.; and Wang, Y. 2021b. Dynamic neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
  10. 10.Han, Y.; Huang, G.; Song, S.; Yang, L.; Zhang, Y.; and Jiang, H. 2021c. Spatially adaptive feature refinement for efficient inference. IEEE Transactions on Image Processing, 9345–9358.
  11. 11.Han, Y.; Yuan, Z.; Pu, Y.; Xue, C.; Song, S.; Sun, G.; and Huang, G. 2022b. Latency-aware Spatial-wise Dynamic Networks. In Advances in Neural Information Processing Systems (NeurIPS).
  12. 12.Heo, B.; Yun, S.; Han, D.; Chun, S.; Choe, J.; and Oh, S. J. 2021. Rethinking spatial dimensions of vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 11936–11945.
  13. 13.Huang, G.; Chen, D.; Li, T.; Wu, F.; van der Maaten, L.; and Weinberger, K. Q. 2018. Multi-scale dense networks for resource efficient image classification. In International Conference on Learning Representations (ICLR).
  14. 14.Huang, Z.; Ben, Y.; Luo, G.; Cheng, P.; Yu, G.; and Fu, B. 2021. Shuffle transformer: Rethinking spatial shuffle for vision transformer. arXiv:2106.03650.
  15. 15.Jiang, Z.-H.; Hou, Q.; Yuan, L.; Zhou, D.; Shi, Y.; Jin, X.; Wang, A.; and Feng, J. 2021. All tokens matter: Token labeling for training better vision transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  16. 16.Li, S.; Wang, Z.; Liu, Z.; Tan, C.; Lin, H.; Wu, D.; Chen, Z.; Zheng, J.; and Li, S. Z. 2022. Efficient Multi-order Gated Aggregation Network. arXiv:2211.03295.
  17. 17.Li, Y.; Zhang, K.; Cao, J.; Timofte, R.; and Van Gool, L. 2021. Localvit: Bringing locality to vision transformers. arXiv:2104.05707.
  18. 18.Liang, Y.; GE, C.; Tong, Z.; Song, Y.; Wang, J.; and Xie, P. 2022. EviT: Expediting Vision Transformers via Token Reorganizations. In International Conference on Learning Representations (ICLR).
  19. 19.Lin, M.; Chen, M.; Zhang, Y.; Li, K.; Shen, Y.; Shen, C.; and Ji, R. 2022. Super Vision Transformer. arXiv:2205.11397.
  20. 20.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10012–10022.
  21. 21.Pan, B.; Panda, R.; Jiang, Y.; Wang, Z.; Feris, R.; and Oliva, A. 2021a. IA-RED2 : Interpretability-Aware Redundancy Reduction for Vision Transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  22. 22.Pan, Z.; Zhuang, B.; Liu, J.; He, H.; and Cai, J. 2021b. Scalable vision transformers with hierarchical pooling. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 377–386.
  23. 23.Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021. Dynamicvit: Efficient vision transformers with dynamic token sparsification. In Advances in Neural Information Processing Systems (NeurIPS).
  24. 24.Si, C.; Yu, W.; Zhou, P.; Zhou, Y.; Wang, X.; and Yan, S. 2022. Inception Transformer. In Advances in Neural Information Processing Systems (NeurIPS).
  25. 25.Song, L.; Zhang, S.; Liu, S.; Li, Z.; He, X.; Sun, H.; Sun, J.; and Zheng, N. 2021. Dynamic grained encoder for vision transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  26. 26.Sun, C.; Shrivastava, A.; Singh, S.; and Gupta, A. 2017. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 843–852.
  27. 27.Tan, M.; and Le, Q. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning (ICML), 6105–6114.
  28. 28.Tang, S.; Zhang, J.; Zhu, S.; and Tan, P. 2022a. Quadtree Attention for Vision Transformers. In International Conference on Learning Representations (ICLR).
  29. 29.Tang, Y.; Han, K.; Wang, Y.; Xu, C.; Guo, J.; Xu, C.; and Tao, D. 2022b. Patch slimming for efficient vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12165–12174.
  30. 30.Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jegou, H. 2021a. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning (ICML), 10347–10357.
  31. 31.Touvron, H.; Cord, M.; Sablayrolles, A.; Synnaeve, G.; and Jegou, H. 2021b. Going deeper with image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 32–42.
  32. 32.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS).
  33. 33.Wang, W.; Stuijk, S.; and De Haan, G. 2014. Exploiting spatial redundancy of image sensor for motion robust rPPG. IEEE Transactions on Biomedical Engineering (TBE), 415–425.
  34. 34.Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; and Shao, L. 2021a. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 568–578.
  35. 35.Wang, Y.; Chen, Z.; Jiang, H.; Song, S.; Han, Y.; and Huang, G. 2021b. Adaptive focus for efficient video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 16249–16258.
  36. 36.Wang, Y.; Huang, R.; Song, S.; Huang, Z.; and Huang, G. 2021c. Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition. In Advances in Neural Information Processing Systems (NeurIPS).
  37. 37.Wang, Y.; Lv, K.; Huang, R.; Song, S.; Yang, L.; and Huang, G. 2020. Glance and focus: a dynamic approach to reducing spatial redundancy in image classification. In Advances in Neural Information Processing Systems (NeurIPS), 2432–2444.
  38. 38.Wang, Y.; Yue, Y.; Lin, Y.; Jiang, H.; Lai, Z.; Kulikov, V.; Orlov, N.; Shi, H.; and Huang, G. 2022a. Adafocus v2: End-to-end training of spatial dynamic networks for video recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20030–20040.
  39. 39.Wang, Y.; Yue, Y.; Xu, X.; Hassani, A.; Kulikov, V.; Orlov, N.; Song, S.; Shi, H.; and Huang, G. 2022b. AdaFocusV3: On Unified Spatial-Temporal Dynamic Video Recognition. In European Conference on Computer Vision (ECCV), 226–243.
  40. 40.Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021. SegFormer: Simple and efficient design for semantic segmentation with transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  41. 41.Xu, W.; Xu, Y.; Chang, T.; and Tu, Z. 2021. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9981–9990.
  42. 42.Xu, Y.; Zhang, Z.; Zhang, M.; Sheng, K.; Li, K.; Dong, W.; Zhang, L.; Xu, C.; and Sun, X. 2022. Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  43. 43.Yang, L.; Han, Y.; Chen, X.; Song, S.; Dai, J.; and Huang, G. 2020. Resolution adaptive networks for efficient inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2369–2378.
  44. 44.Yu, Q.; Xia, Y.; Bai, Y.; Lu, Y.; Yuille, A. L.; and Shen, W. 2021. Glance-and-gaze vision transformer. In Advances in Neural Information Processing Systems (NeurIPS).
  45. 45.Yuan, K.; Guo, S.; Liu, Z.; Zhou, A.; Yu, F.; and Wu, W. 2021a. Incorporating convolution designs into visual transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 579–588.
  46. 46.Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.-H.; Tay, F. E.; Feng, J.; and Yan, S. 2021b. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 558–567.
  47. 47.Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. 2021. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6881–6890.
  48. 48.Zong, Z.; Li, K.; Song, G.; Wang, Y.; Qiao, Y.; Leng, B.; and Liu, Y. 2022. Self-slimmed vision transformer. In European Conference on Computer Vision (ECCV), 432–448.

Citation

MLA
Chen, M., et al. “CF-ViT: A General Coarse-to-Fine Method for Vision Transformer”. arXiv, 2022, http://arxiv.org/abs/2203.03821v5.
APA
Chen, M., Lin, M., Li, K., Shen, Y., Wu, Y., Chao, F., & Ji, R. (2022). CF-ViT: A General Coarse-to-Fine Method for Vision Transformer. arXiv. http://arxiv.org/abs/2203.03821v5
Chicago
Chen, M., M. Lin, K. Li, et al. 2022. “CF-ViT: A General Coarse-to-Fine Method for Vision Transformer”. arXiv. http://arxiv.org/abs/2203.03821v5.
Harvard
Chen, M. et al. (2022) “CF-ViT: A General Coarse-to-Fine Method for Vision Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.03821v5.
Vancouver
1. Chen M, Lin M, Li K, Shen Y, Wu Y, Chao F, Ji R (2022) CF-ViT: A General Coarse-to-Fine Method for Vision Transformer. arXiv

BibTeX

@article{chen2022vit,
  title = {CF-ViT: A General Coarse-to-Fine Method for Vision Transformer},
  author = {Chen, Mengzhao and Lin, Mingbao and Li, Ke and Shen, Yunhang and Wu, Yongjian and Chao, Fei and Ji, Rongrong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.03821v5},
  eprint = {2203.03821}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF