BiFormer: Vision Transformer with Bi-Level Routing Attention
Lei ZhuXinjiang WangZhanghan KeWayne ZhangRynson W. H. Lau
Introduces BiFormer, a vision transformer that uses dynamic bi-level routing attention to adaptively filter out irrelevant image regions, cutting computational costs while maintaining high performance across classification, detection, and segmentation tasks.
Vision transformers are powerful artificial intelligence models for computer vision, but their core attention mechanism demands massive computation and memory because it compares every image patch against every other patch across the entire image. Existing solutions attempt to reduce this cost using fixed, handcrafted search windows or static sharing schemes. However, these techniques ignore image content, restrict long-range context, and fail to adapt to different semantic regions.
The article demonstrates a dynamic, content-aware sparse attention architecture called Bi-Level Routing Attention (BRA) and introduces BiFormer, a general-purpose vision transformer backbone designed to enhance accuracy while keeping computational complexity low.
The evaluation evaluated BiFormer across standard computer vision benchmarks, including ImageNet-1K for image classification, COCO for object detection and instance segmentation, and ADE20K for semantic segmentation. Instead of calculating attention across all locations, the approach first identifies the most relevant regions through a coarse region-to-region affinity graph, filters out irrelevant regions, and then gathers the key data points to perform detailed token-to-token attention exclusively within the retained regions using standard, hardware-friendly dense matrix calculations.
The results establish three primary findings. First, BiFormer achieves superior image classification accuracy under comparable computation budgets; for instance, the tiny variant reaches 81.4% top-1 accuracy on ImageNet-1K (and up to 84.3% in the small variant with advanced training), outperforming competitive models. Second, in object detection and instance segmentation, BiFormer demonstrates clear performance advantages, particularly in detecting small objects where sparse routing preserves fine visual details better than traditional downsampling methods. Third, in semantic segmentation, BiFormer models consistently exceed previous baselines, showing gains across multiple standard frameworks.
These findings indicate that artificial intelligence systems can capture critical long-range dependencies and fine visual details without incurring prohibitive computational overhead. By directing processing power only to relevant image regions, organizations can achieve state-of-the-art visual recognition performance at lower operational and computational costs.
For practical adoption, engineering teams should evaluate BiFormer backbones for vision pipelines that demand high precision on complex, high-resolution scenes. Future efforts should focus on low-level hardware optimizations, such as graphical processing unit kernel fusion, to streamline memory access and kernel launch routines.
While BiFormer drastically cuts computational operations, its multi-step routing mechanism introduces memory transactions and kernel overhead that reduce practical inference throughput compared to simpler, static-window architectures on current hardware. Confidence in the reported accuracy and operational improvements remains high across evaluated benchmarks, but real-time production deployments should measure actual device throughput alongside theoretical computational efficiency.
- Paper: Twins: Revisiting the Design of Spatial Attention in Vision Transformers, Xiangxiang Chu et al. (2021). Introduces spatially separable self-attention combining local grouped and sub-sampled global attention, providing essential baseline techniques for reducing quadratic complexity in vision transformers.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). Provides a comprehensive taxonomy and analysis of efficient attention mechanisms, establishing foundational concepts of dynamic sparsity and routing.
- Paper: Scaling Vision with Sparse Mixture of Experts, Carlos Riquelme et al. (2021). Demonstrates conditional computation and routing mechanisms in vision transformers, laying the groundwork for query-adaptive sparse token selection.
- Paper: CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention, Wenxiao Wang et al. (2022). Establishes multi-scale and cross-scale attention designs for vision backbones that motivate BiFormer's hierarchical, coarse-to-fine token interaction scheme.
- Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). Pioneers content-dependent sparse attention via hashing-based token clustering, inspiring the algorithmic shift from fixed spatial windows to dynamic routing.
- Paper: MetaFormer is Actually What You Need for Vision, Weihao Yu et al. (2021). Examines the fundamental macro-architecture of vision transformers and general token-mixing frameworks upon which BiFormer constructs its stage-wise backbone.
- Paper: SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, Enze Xie et al. (2021). Presents hierarchical vision transformer design with sequence-reduction attention tailored for dense prediction tasks like semantic segmentation.
- Paper: Transformers in Vision: A Survey, Salman Khan et al. (2021). Surveys vision transformer architectures and self-attention formulations across dense visual perception domains.
- Paper: CF-ViT: A General Coarse-to-Fine Method for Vision Transformer, Mengzhao Chen et al. (2023). Applies coarse-to-fine patch selection to vision transformers to relieve computational overhead while preserving discriminative image regions.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). Extends efficient vision transformer backbone designs by combining fine-grained local focus and coarse global perception through aggregated attention.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). Advances dynamic block-sparse attention mechanisms to accelerate attention computations in long-sequence and high-resolution visual tasks.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). Leverages hierarchical global-to-local token evaluation and adaptive pruning to accelerate high-resolution multimodal vision architectures.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). Builds on dynamic token reduction concepts to filter and route essential visual tokens without retraining in multimodal vision models.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). Explores an alternative linear-complexity paradigm for vision representation learning to overcome self-attention bottlenecks.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). Develops 2D selective state space backbones as a continuous line of research in efficient global visual token mixing.
