CF-ViT: A General Coarse-to-Fine Method for Vision Transformer
Mengzhao ChenMingbao LinKe LiYunhang ShenYongjian WuFei ChaoRongrong Ji
Proposes a two-stage vision transformer inference framework that processes easy images using coarse-grained patches and selectively re-splits only informative regions into fine-grained tokens for hard images, cutting computational cost by more than half while doubling throughput without sacrificing accuracy.
Vision Transformer models achieve outstanding accuracy across computer vision tasks, but their high computational cost poses significant challenges for practical deployment. Processing an image requires dividing it into visual patches (tokens), and the computational burden grows quadratically with the number of tokens. In practice, many images contain substantial spatial redundancy, such as empty backgrounds, and most standard images do not require dense analysis to be classified correctly. The article develops and evaluates a coarse-to-fine framework, termed CF-ViT, designed to dynamically allocate computational effort based on image difficulty and regional importance.
The framework operates in two sequential stages using a single shared neural network. In the first stage, the model divides an input image into a small number of coarse patches, enabling low-cost initial classification. If the model achieves high classification confidence, inference terminates immediately. If confidence falls below an adjustable threshold, the model identifies the most informative image regions using an attention-tracking mechanism across network layers. Only these critical regions are re-split into finer patches for a second inference pass, while coarse representations are reused to preserve local context without adding extra parameters.
Evaluation on the standard ImageNet benchmark demonstrates substantial computational savings and speed improvements without sacrificing recognition performance. When applied to standard baseline models, the coarse-to-fine approach cut floating-point operations by 53% to 61% while maintaining baseline accuracy, effectively doubling image processing throughput on standard hardware (up to a 2.01-fold increase). Furthermore, when calibrated for maximum accuracy rather than maximum speed, the method outperformed standard baselines by up to 1.0% in classification accuracy while still consuming less computation. The approach also consistently outperformed existing dynamic token-pruning and early-exit methods across comparable computational budgets.
These results provide actionable implications for enterprise AI systems, edge deployments, and cloud-scale visual processing. By enabling flexible trade-offs between processing latency and accuracy via a single threshold parameter, the method reduces infrastructure hosting costs and hardware requirements. Unlike prior multi-stage approaches that require storing multiple separate models in memory, this single-model architecture minimizes storage and operational overhead.
Organizations deploying vision transformer models should consider adopting dynamic coarse-to-fine token allocation strategies to optimize system throughput and operational costs. Future initiatives should evaluate extending this dynamic processing strategy beyond standard image classification to dense visual tasks such as object detection and semantic segmentation. Users should note that while results are highly consistent on standard image recognition benchmarks, performance under real-world shifts in data complexity and hardware execution environments requires domain-specific validation.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Introduces the standard Vision Transformer architecture and fixed patch tokenization whose quadratic computational complexity motivates CF-ViT's dynamic coarse-to-fine design.
- Paper: AdaViT: Adaptive Vision Transformers for Efficient Image Recognition, Lingchen Meng et al. (2022). Establishes dynamic token-selection and early-exit mechanisms for Vision Transformers, providing the baseline adaptive computation paradigms compared directly against CF-ViT.
- Paper: CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification, Chun-Fu Chen et al. (2021). Explores dual-branch coarse and fine patch processing in vision transformers, framing the multiscale representation trade-offs that CF-ViT addresses with a single shared network.
- Paper: Transformer in Transformer, Kai Han et al. (2021). Introduces nested patch-within-patch structures in vision transformers, informing CF-ViT's strategy of re-splitting coarse image regions into fine sub-patches.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Presents foundational data-efficient vision transformer models and training recipes that serve as primary evaluation baselines for CF-ViT.
- Paper: Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers, Siyuan Wei et al. (2023). Extends efficient vision transformer inference by combining token pruning with feature squeezing to retain context without the overhead of multi-stage inference.
- Paper: BiFormer: Vision Transformer with Bi-Level Routing Attention, Lei Zhu et al. (2023). Applies dynamic coarse-to-fine region routing directly within attention layers rather than through multi-stage network passes.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). Generalizes coarse-to-fine dynamic token allocation and global-local attention scoring to high-resolution vision-language models.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Applies the concept of coarse global identification followed by fine-grained regional processing to multimodal chain-of-thought reasoning.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). Extends attention-guided token selection and reduction principles from pure vision backbones to large multimodal vision-language architectures.
