BiFormer: Vision Transformer with Bi-Level Routing Attention

Lei ZhuXinjiang WangZhanghan KeWayne ZhangRynson W. H. Lau

article2023CVPR926 citations

Introduces BiFormer, a vision transformer that uses dynamic bi-level routing attention to adaptively filter out irrelevant image regions, cutting computational costs while maintaining high performance across classification, detection, and segmentation tasks.

Listen

Vision transformers are powerful artificial intelligence models for computer vision, but their core attention mechanism demands massive computation and memory because it compares every image patch against every other patch across the entire image. Existing solutions attempt to reduce this cost using fixed, handcrafted search windows or static sharing schemes. However, these techniques ignore image content, restrict long-range context, and fail to adapt to different semantic regions.

The article demonstrates a dynamic, content-aware sparse attention architecture called Bi-Level Routing Attention (BRA) and introduces BiFormer, a general-purpose vision transformer backbone designed to enhance accuracy while keeping computational complexity low.

The evaluation evaluated BiFormer across standard computer vision benchmarks, including ImageNet-1K for image classification, COCO for object detection and instance segmentation, and ADE20K for semantic segmentation. Instead of calculating attention across all locations, the approach first identifies the most relevant regions through a coarse region-to-region affinity graph, filters out irrelevant regions, and then gathers the key data points to perform detailed token-to-token attention exclusively within the retained regions using standard, hardware-friendly dense matrix calculations.

The results establish three primary findings. First, BiFormer achieves superior image classification accuracy under comparable computation budgets; for instance, the tiny variant reaches 81.4% top-1 accuracy on ImageNet-1K (and up to 84.3% in the small variant with advanced training), outperforming competitive models. Second, in object detection and instance segmentation, BiFormer demonstrates clear performance advantages, particularly in detecting small objects where sparse routing preserves fine visual details better than traditional downsampling methods. Third, in semantic segmentation, BiFormer models consistently exceed previous baselines, showing gains across multiple standard frameworks.

These findings indicate that artificial intelligence systems can capture critical long-range dependencies and fine visual details without incurring prohibitive computational overhead. By directing processing power only to relevant image regions, organizations can achieve state-of-the-art visual recognition performance at lower operational and computational costs.

For practical adoption, engineering teams should evaluate BiFormer backbones for vision pipelines that demand high precision on complex, high-resolution scenes. Future efforts should focus on low-level hardware optimizations, such as graphical processing unit kernel fusion, to streamline memory access and kernel launch routines.

While BiFormer drastically cuts computational operations, its multi-step routing mechanism introduces memory transactions and kernel overhead that reduce practical inference throughput compared to simpler, static-window architectures on current hardware. Confidence in the reported accuracy and operational improvements remains high across evaluated benchmarks, but real-time production deployments should measure actual device throughput alongside theoretical computational efficiency.

Cover for BiFormer: Vision Transformer with Bi-Level Routing Attention

Abstract

As the core building block of vision transformers, attention is a powerful tool to capture long-range dependency. However, such power comes at a cost: it incurs a huge computation burden and heavy memory footprint as pairwise token interaction across all spatial locations is computed. A series of works attempt to alleviate this problem by introducing handcrafted and content-agnostic sparsity into attention, such as restricting the attention operation to be inside local windows, axial stripes, or dilated windows. In contrast to these approaches, we propose a novel dynamic sparse attention via bi-level routing to enable a more flexible allocation of computations with content awareness. Specifically, for a query, irrelevant key-value pairs are first filtered out at a coarse region level, and then fine-grained token-to-token attention is applied in the union of remaining candidate regions (\ie, routed regions). We provide a simple yet effective implementation of the proposed bi-level routing attention, which utilizes the sparsity to save both computation and memory while involving only GPU-friendly dense matrix multiplications. Built with the proposed bi-level routing attention, a new general vision transformer, named BiFormer, is then presented. As BiFormer attends to a small subset of relevant tokens in a \textbf{query adaptive} manner without distraction from other irrelevant ones, it enjoys both good performance and high computational efficiency, especially in dense prediction tasks. Empirical results across several computer vision tasks such as image classification, object detection, and semantic segmentation verify the effectiveness of our design. Code is available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Our Approach: BiFormer
  • 3.1 Preliminaries: Attention
  • 3.2 Bi-Level Routing Attention (BRA)
  • 3.3 Complexity Analysis of BRA
  • 3.4 Architecture Design of BiFormer
  • 4 Experiments
  • 4.1 Image Classification on ImageNet-1K
  • 4.2 Object Detection and Instance Segmentation
  • 4.3 Semantic Segmentation on ADE20K
  • 4.4 Ablation Study
  • 4.5 Visualization of Attention Map
  • 5 Limitation and Future Work
  • 6 Conclusion
  • A Discussion on Regional Representations
  • B Throughput Comparison
  • C Choices of top-kk and partition factor SS
  • D Adapting Pretrained Plain ViT with BRA
  • E More Visualization Results
  • References

Knowls

  1. Knowl 1 — Bi-Level Routing Attention Mechanism

    model/method

    Bi-Level Routing Attention (BRA) is a dynamic, query-adaptive sparse attention mechanism that filters out irrelevant key-value pairs at a coarse region level before computing fine-grained token-to-token attention in the remaining candidate regions.

    Given a 2D input feature tensor X∈RH×W×CX \in \mathbb{R}^{H \times W \times C}, where H,WH, W are spatial dimensions and CC is channel dimension, the mechanism executes in three consecutive stages:

    1. Region Partition and Linear Projection: The spatial feature map is partitioned into S×SS \times S non-overlapping spatial regions, such that each region contains HWS2\frac{HW}{S^2} token vectors. The reshaped feature tensor is Xr∈RS2×HWS2×CX^r \in \mathbb{R}^{S^2 \times \frac{HW}{S^2} \times C}. Linear projections yield query, key, and value tensors: Q=XrWq,K=XrWk,V=XrWv∈RS2×HWS2×CQ = X^r W^q, \quad K = X^r W^k, \quad V = X^r W^v \in \mathbb{R}^{S^2 \times \frac{HW}{S^2} \times C} where Wq,Wk,Wv∈RC×CW^q, W^k, W^v \in \mathbb{R}^{C \times C} are learnable projection matrices.

    2. Region-to-Region Routing via Directed Affinity Graph: Regional query and key representations Qr,Kr∈RS2×CQ^r, K^r \in \mathbb{R}^{S^2 \times C} are computed via per-region average pooling: Qir=S2HW∑j=1HW/S2Qi,j,Kir=S2HW∑j=1HW/S2Ki,jQ^r_i = \frac{S^2}{HW} \sum_{j=1}^{HW/S^2} Q_{i,j}, \quad K^r_i = \frac{S^2}{HW} \sum_{j=1}^{HW/S^2} K_{i,j} The region-level affinity adjacency matrix Ar∈RS2×S2A^r \in \mathbb{R}^{S^2 \times S^2} is evaluated as: Ar=Qr(Kr)TA^r = Q^r (K^r)^T Each region ii is routed to its top-kk most semantically relevant regions by applying a row-wise top-kk operator: Ir=topkIndex⁡(Ar)∈NS2×kI^r = \operatorname{topkIndex}(A^r) \in \mathbb{N}^{S^2 \times k} where the ii-th row of IrI^r stores the kk region indices to which region ii routes.

    3. Token-to-Token Attention with Local Context Enhancement: Key and value tokens residing within the kk routed regions for each query region are gathered into dense contiguous tensors: Kg=gather⁡(K,Ir),Vg=gather⁡(V,Ir)∈RS2×kHWS2×CK^g = \operatorname{gather}(K, I^r), \quad V^g = \operatorname{gather}(V, I^r) \in \mathbb{R}^{S^2 \times \frac{kHW}{S^2} \times C} Fine-grained attention is then computed between queries in each region and the gathered key-value pairs, augmented with a depthwise convolution Local Context Enhancement (LCE) term with kernel size 5: O=softmax⁡(Q(Kg)TC)Vg+DWConv⁡5×5(V)O = \operatorname{softmax}\left(\frac{Q (K^g)^T}{\sqrt{C}}\right) V^g + \operatorname{DWConv}_{5\times 5}(V)

  2. Knowl 2 — Computational Complexity and Optimal Region Scaling of BRA

    theoretical result

    The computational complexity of Bi-Level Routing Attention (BRA) consists of three components: linear projections (FLOPsproj\text{FLOPs}_{\text{proj}}), region-to-region routing (FLOPsrouting\text{FLOPs}_{\text{routing}}), and gathered token-to-token attention (FLOPsattn\text{FLOPs}_{\text{attn}}): FLOPs=FLOPsproj+FLOPsrouting+FLOPsattn=3HWC2+2(S2)2C+2HW(kHWS2)C\text{FLOPs} = \text{FLOPs}_{\text{proj}} + \text{FLOPs}_{\text{routing}} + \text{FLOPs}_{\text{attn}} = 3HW C^2 + 2(S^2)^2 C + 2HW\left(\frac{kHW}{S^2}\right) C where H×WH \times W is the spatial resolution, CC is the channel embedding dimension, S×SS \times S is the number of partitioned regions, and kk is the number of regions attended per query region.

    Applying the arithmetic mean-geometric mean (AM-GM) inequality to the routing and attention terms: 2S4+2k(HW)2S2=2S4+k(HW)2S2+k(HW)2S2≥3(2S4⋅k(HW)2S2⋅k(HW)2S2)1/3=3k2/3(2HW)4/32S^4 + \frac{2k(HW)^2}{S^2} = 2S^4 + \frac{k(HW)^2}{S^2} + \frac{k(HW)^2}{S^2} \ge 3\left(2S^4 \cdot \frac{k(HW)^2}{S^2} \cdot \frac{k(HW)^2}{S^2}\right)^{1/3} = 3k^{2/3}(2HW)^{4/3} Equality holds if and only if 2S4=k(HW)2S22S^4 = \frac{k(HW)^2}{S^2}. Solving for the region partition factor SS gives the optimal scaling rule: S=(k2(HW)2)1/6S = \left(\frac{k}{2} (HW)^2\right)^{1/6} When SS is scaled with the input resolution according to this rule, BRA achieves a computational complexity of O((HW)4/3)\mathcal{O}((HW)^{4/3}), in contrast to O((HW)2)\mathcal{O}((HW)^2) for vanilla self-attention and O((HW)3/2)\mathcal{O}((HW)^{3/2}) for axial attention.

  3. Knowl 3 — BiFormer Hierarchical Vision Transformer Architecture

    model/method

    BiFormer is a hierarchical four-stage vision transformer built around Bi-Level Routing Attention (BRA).

    Stage Structure:

    • Stage 1 downsamples the input image to resolution H4×W4\frac{H}{4} \times \frac{W}{4} using an overlapped patch embedding with base channel dimension CC.
    • Stages 2, 3, and 4 each use a patch merging module that downsamples spatial resolution by a factor of 2 while doubling the channel capacity (2C,4C,8C2C, 4C, 8C).
    • Each stage i∈{1,2,3,4}i \in \{1, 2, 3, 4\} contains NiN_i stacked BiFormer blocks.

    BiFormer Block Composition:

    1. A 3×33 \times 3 depthwise convolution at the input to implicitly encode position information.
    2. Layer Normalization followed by a BRA module with a residual connection.
    3. Layer Normalization followed by a 2-layer MLP with expansion ratio e=3e=3 and a residual connection.

    Attention heads are configured with 32 channels per head across all stages. For the BRA modules, the number of attended regions across the four stages is set to k∈{1,4,16,S2}k \in \{1, 4, 16, S^2\}, where k=S2k=S^2 in stage 4 denotes full self-attention. The region partition factor is set to S=7S=7 for image classification (224×224224 \times 224 input), S=8S=8 for semantic segmentation (512×512512 \times 512 input), and S=16S=16 for object detection/instance segmentation.

    Model Variants:

    • BiFormer-T: C=64C=64, block depths [N1,N2,N3,N4]=[2,2,8,2][N_1, N_2, N_3, N_4] = [2, 2, 8, 2], 13M parameters, 2.2G FLOPs at 224×224224 \times 224.
    • BiFormer-S: C=64C=64, block depths [4,4,18,4][4, 4, 18, 4], 26M parameters, 4.5G FLOPs at 224×224224 \times 224.
    • BiFormer-B: C=96C=96, block depths [4,4,18,4][4, 4, 18, 4], 57M parameters, 9.8G FLOPs at 224×224224 \times 224.
  4. Knowl 4 — Equivalence of Regional Mean Dot-Product to Average Token Pair Affinity

    theoretical result

    In Bi-Level Routing Attention, the coarse affinity between region Ω\Omega and region Ω′\Omega' is computed using the dot product of regional mean query and key vectors Qr=1∣Ω∣∑i∈ΩQiQ^r = \frac{1}{|\Omega|} \sum_{i \in \Omega} Q_i and Kr=1∣Ω′∣∑j∈Ω′KjK^r = \frac{1}{|\Omega'|} \sum_{j \in \Omega'} K_j. This region-level inner product is mathematically identical to the arithmetic mean of all pairwise token affinity scores between the two regions: 1∣Ω∣⋅∣Ω′∣∑i∈Ω∑j∈Ω′QiKjT=(∑i∈ΩQi∣Ω∣)(∑j∈Ω′Kj∣Ω′∣)T=Qr(Kr)T\frac{1}{|\Omega| \cdot |\Omega'|} \sum_{i \in \Omega} \sum_{j \in \Omega'} Q_i K_j^T = \left(\frac{\sum_{i \in \Omega} Q_i}{|\Omega|}\right) \left(\frac{\sum_{j \in \Omega'} K_j}{|\Omega'|}\right)^T = Q^r (K^r)^T where ∣Ω∣|\Omega| and ∣Ω′∣|\Omega'| denote the number of tokens in regions Ω\Omega and Ω′\Omega', respectively. Maximizing the dot product of regional mean representations therefore exactly maximizes the average pairwise token affinity between the candidate regions.

  5. Knowl 5 — ImageNet-1K Classification Performance of BiFormer Models

    data/table

    BiFormer models were trained from scratch on ImageNet-1K for 300 epochs at 224×224224 \times 224 resolution using AdamW, cosine learning rate scheduling, RandAugment, MixUp, CutMix, Random Erasing, and stochastic depth.

    Model FLOPs (G) Params (M) Top-1 Acc. (%)
    ResNet-18 1.8 11.7 69.8
    RegNetY-1.6G 1.6 11.2 78.0
    PVTv2-b1 2.1 13.1 78.7
    Shunted-T 2.1 11.5 79.8
    QuadTree-B-b1 2.3 13.6 80.0
    BiFormer-T 2.2 13.1 81.4
    Swin-T 4.5 29 81.3
    CSWin-T 4.5 23 82.7
    DAT-T 4.6 29 82.0
    CrossFormer-S 5.3 31 82.5
    RegionViT-S 5.3 31 82.6
    QuadTree-B-b2 4.5 24 82.7
    MaxViT-T 5.6 31 83.6
    ScalableViT-S 4.2 32 83.1
    Uniformer-S* 4.2 24 83.4
    Wave-ViT-S* 4.7 23 83.9
    BiFormer-S 4.5 26 83.8
    BiFormer-S* 4.5 26 84.3
    Swin-B 15.4 88 83.5
    CSWin-B 15.0 78 84.2
    CrossFormer-L 16.1 92 84.0
    ScalableViT-B 8.6 81 84.1
    Uniformer-B* 8.3 50 85.1
    Wave-ViT-B* 7.2 34 84.8
    BiFormer-B 9.8 57 84.3
    BiFormer-B* 9.8 58 85.4

    Models marked with * use token labeling distillation during training. BiFormer-T achieves 81.4% accuracy at 2.2G FLOPs (+1.4% over QuadTree-B-b1). BiFormer-S achieves 83.8% top-1 accuracy without token labeling and 84.3% with token labeling. BiFormer-B achieves 84.3% (85.4% with token labeling) at 9.8G FLOPs, outperforming models with 15G--16G FLOPs such as Swin-B and CSWin-B.

  6. Knowl 6 — Object Detection and Instance Segmentation Results on COCO 2017

    data/table

    Object detection and instance segmentation evaluations were conducted on COCO 2017 with backbones pretrained on ImageNet-1K using standard 1×1\times schedules (12 epochs) in MMDetection with RetinaNet and Mask R-CNN frameworks.

    Backbone RetinaNet 1×1\times schedule Mask R-CNN 1×1\times schedule
    mAP AP50AP_{50} AP75AP_{75} APSAP_S APMAP_M APLAP_L mAPb\text{mAP}^b AP50bAP_{50}^b AP75bAP_{75}^b mAPm\text{mAP}^m AP50mAP_{50}^m AP75mAP_{75}^m
    Swin-T 41.5 62.1 44.2 25.1 44.9 55.5 42.2 64.6 46.2 39.1 61.6 42.0
    DAT-T 42.8 64.4 45.2 28.0 45.8 57.8 44.4 67.6 48.5 40.4 64.2 43.1
    CSWin-T - - - - - - 46.7 68.6 51.3 42.2 65.6 45.4
    CrossFormer-S 44.4 55.3 38.6 19.3 40.0 48.8 45.4 68.0 49.7 41.4 64.8 44.6
    QuadTree-B2 46.2 67.2 49.5 29.0 50.1 61.8 - - - - - -
    WaveViT-S* 45.8 67.0 49.4 29.2 50.0 60.8 46.6 68.7 51.2 42.4 65.5 45.8
    BiFormer-S 45.9 66.9 49.4 30.2 49.6 61.7 47.8 69.8 52.3 43.2 66.8 46.5
    Swin-S 44.5 65.7 47.5 27.4 48.0 59.9 44.8 66.6 48.9 40.9 63.4 44.2
    DAT-S 45.7 67.7 48.5 30.5 49.3 61.3 47.1 69.9 51.5 42.5 66.7 45.4
    CSWin-S - - - - - - 47.9 70.1 52.6 43.2 67.1 46.2
    CrossFormer-B 46.2 67.8 49.5 30.1 49.9 61.8 47.2 69.9 51.8 42.7 66.6 46.2
    QuadTree-B3 47.3 68.2 50.6 30.4 51.3 62.9 - - - - - -
    Wave-ViT-B* 47.2 68.2 50.9 29.7 51.4 62.3 47.6 69.1 52.4 43.0 66.4 46.0
    BiFormer-B 47.1 68.5 50.4 31.3 50.8 62.6 48.6 70.5 53.8 43.7 67.6 47.1

    BiFormer shows pronounced gains on small objects (APSAP_S), obtaining 30.2 on RetinaNet for BiFormer-S and 31.3 for BiFormer-B. This gain arises because BRA achieves sparsity through selective token sampling rather than spatial downsampling, preserving fine-grained details. On Mask R-CNN, BiFormer-S reaches 47.8 mAPb\text{mAP}^b and 43.2 mAPm\text{mAP}^m, and BiFormer-B reaches 48.6 mAPb\text{mAP}^b and 43.7 mAPm\text{mAP}^m.

  7. Knowl 7 — ADE20K Semantic Segmentation Performance

    data/table

    Semantic segmentation performance was evaluated on ADE20K using MMSegmentation with Semantic FPN (trained for 80k iterations) and UperNet (trained for 160k iterations) frameworks.

    Backbone Semantic FPN UperNet
    mIoU (%) mIoU (%) MS mIoU (%)
    Swin-T 41.5 44.5 45.8
    DAT-T 42.6 45.5 46.4
    CSWin-T 48.2 49.3 50.7
    CrossFormer-S 46.0 47.6 48.4
    Shunted-S 48.2 48.9 49.9
    WaveViT-S* - - 49.6
    BiFormer-S 48.9 49.8 50.8
    Swin-S - 47.6 49.5
    DAT-S 46.1 48.3 49.8
    CSWin-S 49.2 50.4 51.5
    CrossFormer-B 47.7 49.7 50.6
    Uniformer-B 48.0 50.0 50.8
    WaveViT-B* - - 51.5
    BiFormer-B 49.9 51.0 51.7

    Under Semantic FPN, BiFormer-S reaches 48.9% mIoU (+0.7% over CSWin-T) and BiFormer-B reaches 49.9% mIoU (+0.7% over CSWin-S). Under UperNet with multi-scale testing (MS mIoU), BiFormer-S achieves 50.8% and BiFormer-B achieves 51.7%.

  8. Knowl 8 — Ablation of Attention Mechanisms Under Controlled Backbone Layout

    data/table

    To isolate the effect of the attention mechanism, multiple sparse attention patterns were compared within a fixed macro-architecture aligned with Swin-T: block depths [2,2,6,2][2, 2, 6, 2], non-overlapped patch embedding, initial embedding dimension C=96C = 96, and MLP expansion ratio e=4e = 4.

    Sparse Attention IN1K Top-1 (%) ADE20K mIoU (%)
    Sliding window 81.4 -
    Shifted window 81.3 41.5
    Spatially Separated 81.5 42.9
    Sequential Axial 81.5 39.8
    Criss-Cross 81.7 43.0
    Cross-shaped window 82.2 43.4
    Deformable 82.0 42.6
    Block-Grid 81.8 42.8
    Bi-Level Routing 82.7 44.8

    Under identical macro-architectural parameters, Bi-Level Routing Attention achieves 82.7% top-1 accuracy on ImageNet-1K (+0.5% over Cross-shaped window and +1.4% over Shifted window) and 44.8% mIoU on ADE20K (+1.4% over Cross-shaped window and +3.3% over Shifted window).

  9. Knowl 9 — Ablation Path from Swin-T Layout to BiFormer-S

    data/table

    The stepwise impact of architectural enhancements applied to transition from a Swin-T baseline layout to the proposed BiFormer-S configuration is detailed below for ImageNet-1K classification:

    Architecture Design Params (M) FLOPs (G) IN1K Top-1 (%)
    Baseline (Swin-T layout) 29 4.6 82.7
    + Overlapped patch embedding 31 4.9 82.8 (+0.1)
    + Deeper layout 25 4.5 83.5 (+0.7)
    + Convolution position encoding 26 4.5 83.8 (+0.3)
    + Token Labeling 29 4.9 84.3 (+0.5)

    The modifications are applied sequentially:

    1. Overlapped patch embedding replaces non-overlapped patch embedding.
    2. Deeper layout stacks more blocks per stage ([4,4,18,4][4, 4, 18, 4] vs. [2,2,6,2][2, 2, 6, 2]) while reducing base channels from 96 to 64 and the MLP expansion ratio from 4 to 3 to maintain constant FLOPs, providing a +0.7% gain.
    3. Convolutional position encoding (3×33 \times 3 depthwise convolution at each block input) provides a +0.3% gain.
    4. Token labeling distillation during training adds another +0.5% top-1 accuracy.
  10. Knowl 10 — Adapting Pretrained Plain ViT with BRA for Semantic Segmentation

    data/table

    Pretrained plain ViT (DeiT-B) models were adapted for semantic segmentation on ADE20K by replacing full multi-head self-attention modules with Bi-Level Routing Attention (BRA) and loading ImageNet-1K pretrained weights directly into the BRA projection matrices. The decoder uses a Simple Feature Pyramid with a UperNet head.

    Attention Function mIoU (%)
    Local window attention (w=14w=14) 43.55
    BRA (w=4,k=12w=4, k=12) 45.92
    Local window attention + 4 conv propagation blocks 44.68
    Local window attention + 4 global propagation blocks 46.64
    BRA + 4 global propagation blocks 46.84

    BRA with window size w=4w=4 and k=12k=12 attended regions (attending to 42×12=1924^2 \times 12 = 192 key-value pairs per query) outperforms local window attention with w=14w=14 (attending to 14×14=19614 \times 14 = 196 key-value pairs per query) by 2.37% mIoU without propagation blocks (45.92% vs. 43.55%) and by 0.20% mIoU when paired with 4 global propagation blocks (46.84% vs. 46.64%).

  11. Knowl 11 — GPU Memory and Kernel Launch Throughput Overhead in Bi-Level Routing

    limitation

    While Bi-Level Routing Attention (BRA) reduces theoretical computational complexity to O((HW)4/3)\mathcal{O}((HW)^{4/3}) and achieves dense matrix multiplication efficiency via key-value gathering, it introduces practical runtime overhead on GPUs.

    Constructing the region-level directed graph, performing row-wise top-kk pruning, and gathering scattered key-value tensors incurs additional GPU kernel launches and non-coalesced memory transactions.

    Benchmarking on a 32GB Tesla V100 GPU (batch size 128, resolution 224×224224 \times 224) indicates that:

    • Swin-T Layout with BRA (BiFormer-STL, 4.6G FLOPs) achieves 133.2 images/s in training and 542.3 images/s in inference under FP32, compared to 218.7 images/s training and 733.3 images/s inference for Swin-T (4.5G FLOPs), reflecting a throughput drop of ~30% in training and ~40% in inference.
    • Under automatic mixed precision (AMP), BiFormer-STL achieves 184.4 images/s training and 766.7 images/s inference compared to 321.4 and 1079.5 images/s for Swin-T.
    • BiFormer-STL is nonetheless 3×3\times to 6×6\times faster than QuadTree-STL (27.5 images/s training / 165.6 images/s inference FP32), as quad-tree attention relies on deep recursive traversals and irregular sparse matrix multiplications that break GPU hardware parallelism.

Coverage note — None was omitted; all primary contributed algorithms, theoretical derivations, architecture specifications, benchmark results (ImageNet-1K, COCO, ADE20K), ablations, adaptation experiments, and throughput limitations are represented.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229, 2020. 1, 3
  2. 2.Chun-Fu Chen, Rameswar Panda, and Quanfu Fan. Regionvit: Regional-to-local attention for vision transformers. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. 5
  3. 3.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, et al. Mmdetection: Open mmlab detection toolbox and benchmark. arXiv:1906.07155, 2019. 6
  4. 4.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. arXiv preprint arXiv:2205.08534, 2022. 11
  5. 5.Zhiyang Chen, Yousong Zhu, Chaoyang Zhao, Guosheng Hu, Wei Zeng, Jinqiao Wang, and Ming Tang. Dpt: Deformable patch-based transformer for visual recognition. In Proceedings of the 29th ACM International Conference on Multimedia, pages 2899–2907, 2021. 1, 3
  6. 6.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv:1904.10509, 2019. 1, 3
  7. 7.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems, 34:9355–9366, 2021. 5, 7, 8
  8. 8.MMSegmentation Contributors. Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020. 7
  9. 9.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition workshops, pages 702–703, 2020. 6
  10. 10.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 764–773, 2017. 2
  11. 11.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In Anna Korhonen, David R. Traum, and Lluís Marquez, editors, Proceedings of the Conference of the Association for Computational Linguistics, ACL 2019, Volume 1: Long Papers, pages 2978–2988, 2019. 3
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. 6
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, June 2019. 1, 3
  14. 14.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12124–12134, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, ICLR 2021, 2021, 2021. 1, 3, 11
  16. 16.Priya Goyal, Piotr Dollar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv:1706.02677, 2017. 6
  17. 17.Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant. Memory-efficient transformers via top-k attention. In Nafise Sadat Moosavi, Iryna Gurevych, Angela Fan, Thomas Wolf, Yufang Hou, Ana Marasovic, and Sujith Ravi, editors, Proceedings of the Workshop on Simple and Efficient Natural Language Processing, 2021, pages 39–52, 2021. 2
  18. 18.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, pages 2961–2969, 2017. 6
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 5
  20. 20.Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. Axial attention in multidimensional transformers. arXiv:1912.12180, 2019. 7
  21. 21.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Weinberger. Deep networks with stochastic depth. In European Conference on Computer Vision, pages 646–661, 2016. 6
  22. 22.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 603–612, 2019. 5, 7
  23. 23.Zi-Hang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. All tokens matter: Token labeling for training better vision transformers. Advances in Neural Information Processing Systems, 34:18590–18602, 2021. 2, 5, 6, 8
  24. 24.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar. Panoptic feature pyramid networks. In Proceedings of the IEEE/CVF Conference on Pattern Recognition, pages 6399–6408, 2019. 7
  25. 25.Kunchang Li, Yali Wang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unified transformer for efficient spatiotemporal representation learning. arXiv:2201.04676, 2022. 5, 6, 7, 8
  26. 26.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. arXiv preprint arXiv:2203.16527, 2022. 11
  27. 27.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 2980–2988, 2017. 6
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 6
  29. 29.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 1, 2, 3, 4, 5, 6, 7, 8, 9
  30. 30.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. 6
  31. 31.Nvidia. How to access global memory efficiently in cuda c/c++ kernels. https : / / developer . nvidia . com / blog / how - access - global - memory - efficiently - cuda - c - kernels/. Accessed: 2022-10-25. 2
  32. 32.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Proceedings of Advances in Neural Information Processing Systems, volume 32, 2019. 4
  33. 33.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. 1
  34. 34.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10428–10436, 2020. 5
  35. 35.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. Stand-alone self-attention in vision models. In Proceedings of Advances in Neural Information Processing Systems, volume 32, 2019. 7
  36. 36.Mr D Murahari Reddy, Mr Sk Masthan Basha, Mr M Chinnaiahgari Hari, and Mr N Penchalaiah. Dall-e: Creating images from text. 2021. 1
  37. 37.Sucheng Ren, Daquan Zhou, Shengfeng He, Jiashi Feng, and Xinchao Wang. Shunted self-attention via multi-scale token aggregation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10853–10862, 2022. 5, 7, 8
  38. 38.Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan. Quadtree attention for vision transformers. In The International Conference on Learning Representations, ICLR 2022, 2022, 2022. 3, 5, 6, 7, 9
  39. 39.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. ACM Computing Surveys (CSUR), 2020. 3
  40. 40.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021. 2, 6, 11
  41. 41.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. In ECCV, 2022. 1, 2, 3, 4, 5, 7
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 1, 3, 4
  43. 43.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv:2006.04768, 2020. 3
  44. 44.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021. 1, 7
  45. 45.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415–424, 2022. 5, 8
  46. 46.Wenxiao Wang, Lu Yao, Long Chen, Binbin Lin, Deng Cai, Xiaofei He, and Wei Liu. Crossformer: A versatile vision transformer hinging on cross-scale attention. In International Conference on Learning Representations, ICLR, 2022. 1, 2, 3, 4, 5, 6, 7
  47. 47.Ross Wightman. Pytorch image models. https : / / github . com / rwightman / pytorch - image - models, 2019. 9
  48. 48.Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. Vision transformer with deformable attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4794–4803, 2022. 1, 2, 3, 4, 5, 7
  49. 49.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European conference on computer vision (ECCV), pages 418–434, 2018. 7, 11
  50. 50.Rui Yang, Hailong Ma, Jie Wu, Yansong Tang, Xuefeng Xiao, Min Zheng, and Xiu Li. Scalablevit: Rethinking the context-oriented generalization of vision transformer. arXiv:2203.10790, 2022. 5
  51. 51.Ting Yao, Yingwei Pan, Yehao Li, Chong-Wah Ngo, and Tao Mei. Wave-vit: Unifying wavelet and transformers for visual representation learning. In European Conference on Computer Vision, pages 328–345. Springer, 2022. 5, 6, 7, 8
  52. 52.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6023–6032, 2019. 6
  53. 53.Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, and Xiaogang Wang. Not all tokens are equal: Human-centric visual analysis via token clustering transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11101–11111, 2022. 3
  54. 54.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, ICLR 2018, 2018. 6
  55. 55.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. International Journal of Computer Vision, 127(3):302–321, 2019. 6, 7

Citation

MLA
Zhu, L., et al. “BiFormer: Vision Transformer with Bi-Level Routing Attention”. arXiv, 2023, http://arxiv.org/abs/2303.08810v1.
APA
Zhu, L., Wang, X., Ke, Z., Zhang, W., & Lau, R. (2023). BiFormer: Vision Transformer with Bi-Level Routing Attention. arXiv. http://arxiv.org/abs/2303.08810v1
Chicago
Zhu, L., X. Wang, Z. Ke, W. Zhang, and R. Lau. 2023. “BiFormer: Vision Transformer with Bi-Level Routing Attention”. arXiv. http://arxiv.org/abs/2303.08810v1.
Harvard
Zhu, L. et al. (2023) “BiFormer: Vision Transformer with Bi-Level Routing Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.08810v1.
Vancouver
1. Zhu L, Wang X, Ke Z, Zhang W, Lau R (2023) BiFormer: Vision Transformer with Bi-Level Routing Attention. arXiv

BibTeX

@article{zhu2023biformer,
  title = {BiFormer: Vision Transformer with Bi-Level Routing Attention},
  author = {Zhu, Lei and Wang, Xinjiang and Ke, Zhanghan and Zhang, Wayne and Lau, Rynson},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.08810v1},
  eprint = {2303.08810}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE