Focal Modulation Networks

Jianwei YangChunyuan LiXiyang DaiJianfeng Gao

article2022NeurIPS416 citations

Proposes FocalNets, an attention-free architecture that replaces self-attention with multi-scale contextual modulation to outperform standard vision transformers across image classification, object detection, and semantic segmentation at comparable computational costs.

Listen

Modern computer vision models increasingly rely on Vision Transformers to understand complex images. However, the core mechanism behind these architectures—known as self-attention—suffers from significant computational bottlenecks because its processing cost grows quadratically with image resolution. This high computational burden creates serious trade-offs in inference latency, hardware resource demands, and operational costs when deploying vision models for high-resolution, real-time enterprise applications.

The article introduces and evaluates an attention-free architecture called the Focal Modulation Network (FocalNet). The primary objective is to demonstrate that replacing self-attention with a lighter mechanism called focal modulation delivers superior visual recognition performance, faster processing throughput, and better model interpretability across standard vision benchmarks.

Rather than computing expensive point-to-point interactions across every visual region simultaneously, the authors designed a three-step alternative: extracting multi-scale local-to-global image context using efficient convolutional layers, dynamically condensing this context with a gating mechanism, and injecting the resulting context modulator into each target image token via lightweight element-wise multiplication. The researchers benchmarked this design across extensive public datasets—including ImageNet-1K/22K, COCO, and ADE20K—evaluating image classification, object detection, and image segmentation against leading vision architectures such as Swin Transformer and ConvNeXt.

The evaluation produced four primary findings. First, FocalNet achieved state-of-the-art results on benchmark vision tasks while operating at comparable or higher inference speeds. On ImageNet-1K classification, tiny and base FocalNets achieved 82.3% and 83.9% top-1 accuracy, outperforming equivalent Swin Transformer baselines. Second, the architecture demonstrated significant sample and training efficiency in dense prediction tasks: a base FocalNet trained on a standard single-schedule object detection routine outperformed a Swin model trained on a three-times longer schedule (49.0 versus 48.5 average precision). Third, when scaled to 746 million parameters with the DINO framework, FocalNet established a new record on the COCO object detection benchmark (64.4 mAP), surpassing much larger multi-billion-parameter models such as SwinV2-G and BEIT-3 while using substantially less training data. Fourth, the internal modulators naturally localized salient object boundaries without requiring post-hoc visual explanation algorithms, providing intrinsic model interpretability.

These findings indicate that heavy self-attention mechanisms are not essential for top-tier visual performance. By decoupling context gathering from token interactions, organizations can significantly reduce model training durations, lower cloud compute expenditures, and deploy higher-resolution visual models on latency-sensitive edge devices. Furthermore, the built-in visual interpretability reduces deployment risk in safety-critical applications by enabling practitioners to inspect the exact image regions driving model classifications.

Technical leaders and engineering teams should consider piloting FocalNet architectures as a drop-in replacement for standard Vision Transformers in computationally constrained or high-resolution visual processing pipelines. Before full-scale adoption across multimodal platforms, organizations should conduct targeted research, as adapting focal modulation to natural language processing and cross-modal tasks (such as paired text-and-image reasoning) remains an open area requiring further investigation.

Confidence in these findings is high for core visual recognition domains, given the consistent empirical validation across multiple standard tasks and model scales. However, decision-makers should maintain appropriate caution regarding training data biases, particularly when fine-tuning on large web-scraped datasets, and conduct rigorous sanity checks prior to production deployment.

Cover for Focal Modulation Networks

Abstract

We propose focal modulation networks (FocalNets in short), where self-attention (SA) is completely replaced by a focal modulation mechanism for modeling token interactions in vision. Focal modulation comprises three components: (i) hierarchical contextualization, implemented using a stack of depth-wise convolutional layers, to encode visual contexts from short to long ranges, (ii) gated aggregation to selectively gather contexts for each query token based on its

content, and (iii) element-wise modulation or affine transformation to inject the aggregated context into the query. Extensive experiments show FocalNets outperform the state-of-the-art SA counterparts (e.g., Swin and Focal Transformers) with similar computational costs on the tasks of image classification, object detection, and segmentation. Specifically, FocalNets with tiny and base size achieve 82.3% and 83.9% top-1 accuracy on ImageNet-1K. After pretrained on ImageNet-22K in 224 resolution, it attains 86.5% and 87.3% top-1 accuracy when finetuned with resolution 224 and 384, respectively. When transferred to downstream tasks, FocalNets exhibit clear superiority. For object detection with Mask R-CNN, FocalNet base trained with 1\times outperforms the Swin counterpart by 2.1 points and already surpasses Swin trained with 3\times schedule (49.0 v.s. 48.5). For semantic segmentation with UPerNet, FocalNet base at single-scale outperforms Swin by 2.4, and beats Swin at multi-scale (50.5 v.s. 49.7). Using large FocalNet and Mask2former, we achieve 58.5 mIoU for ADE20K semantic segmentation, and 57.9 PQ for COCO Panoptic Segmentation. Using huge FocalNet and DINO, we achieved 64.3 and 64.4 mAP on COCO minival and test-dev, respectively, establishing new SoTA on top of much larger attention-based models like Swinv2-G and BEIT-3. Code and checkpoints are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Focal Modulation Network
  • 3.1 From Self-Attention to Focal Modulation
  • 3.2 Context Aggregation via m⁡(⋅)m(\cdot)
  • 3.3 Relation to Other Architecture Designs
  • 3.4 Complexity
  • 3.5 Network Architectures
  • 4 Experiment
  • 4.1 Image Classification
  • 4.2 Language-Image Contrast Learning
  • 4.3 Detection and Segmentation
  • 4.4 Network Inspection
  • 4.5 Comparisons with ViTs and ConvNeXts
  • 5 Conclusion
  • References
  • A More Implementation Details
  • A.1 Model Configuration
  • A.2 Training settings for ImageNet-1K
  • A.3 Training settings for ImageNet-22K
  • A.4 Training settings for Object365
  • B Downstream Tasks
  • B.1 Object Detection
  • B.1.1 Effect of kernel size
  • B.1.2 Results with deeper and thinner FocalNets
  • B.2 Image Segmentation
  • C Additional Model Interpretation
  • D Social Impact

Knowls

  1. Knowl 1 — Focal Modulation Mechanism

    model/method

    Focal Modulation is an attention-free token interaction mechanism designed for visual representation learning. Unlike self-attention, which performs late context aggregation after computing pairwise query-key interactions, focal modulation performs early multi-scale context aggregation and modulates the query token via an input-dependent element-wise affine transformation.

    Given an input feature map X∈RH×W×CX \in \mathbb{R}^{H \times W \times C}, where H,WH, W are spatial dimensions and CC is the channel dimension, focal modulation operates in three stages:

    1. Hierarchical Contextualization: The input is projected to an initial context feature map Z0=fz(X)∈RH×W×CZ^0 = f_z(X) \in \mathbb{R}^{H \times W \times C} via a linear projection fz(⋅)f_z(\cdot). A hierarchy of LL context feature representations is computed using a stack of LL depth-wise convolutional layers: Zℓ=faℓ(Zℓ−1)=GeLU(DWConv(Zℓ−1))∈RH×W×C,ℓ∈{1,…,L}Z^\ell = f_a^\ell(Z^{\ell-1}) = \text{GeLU}(\text{DWConv}(Z^{\ell-1})) \in \mathbb{R}^{H \times W \times C}, \quad \ell \in \{1, \dots, L\} where DWConv\text{DWConv} denotes depth-wise convolution with kernel size kℓk^\ell, and GeLU\text{GeLU} is the Gaussian Error Linear Unit activation. The effective receptive field at level ℓ\ell is rℓ=1+∑i=1ℓ(ki−1)r^\ell = 1 + \sum_{i=1}^\ell (k^i - 1). To capture global context across the entire feature map, global average pooling is applied to the final level: ZL+1=Avg-Pool(ZL)∈R1×1×CZ^{L+1} = \text{Avg-Pool}(Z^L) \in \mathbb{R}^{1 \times 1 \times C}, resulting in L+1L+1 context feature maps {Zℓ}ℓ=1L+1\{Z^\ell\}_{\ell=1}^{L+1}.

    2. Gated Aggregation: Spatial- and level-aware gating weights G∈RH×W×(L+1)G \in \mathbb{R}^{H \times W \times (L+1)} are computed from the input via a linear layer G=fg(X)G = f_g(X). The context feature maps are condensed into a single context representation ZoutZ^{\text{out}} through gated element-wise summation: Zout=∑ℓ=1L+1Gℓ⊙Zℓ∈RH×W×CZ^{\text{out}} = \sum_{\ell=1}^{L+1} G^\ell \odot Z^\ell \in \mathbb{R}^{H \times W \times C} where Gℓ∈RH×W×1G^\ell \in \mathbb{R}^{H \times W \times 1} is the gating slice corresponding to focal level ℓ\ell, and ⊙\odot represents element-wise multiplication. A linear projection h(⋅)h(\cdot) enables cross-channel communication to produce the modulator feature map M=h(Zout)∈RH×W×CM = h(Z^{\text{out}}) \in \mathbb{R}^{H \times W \times C}.

    3. Modulation: For each query token xi∈RCx_i \in \mathbb{R}^C at spatial index ii, the output representation yi∈RCy_i \in \mathbb{R}^C is generated by modulating the projected query token with the aggregated modulator: yi=q(xi)⊙m(i,X)=q(xi)⊙h(∑ℓ=1L+1giℓ⋅ziℓ)y_i = q(x_i) \odot m(i, X) = q(x_i) \odot h\left(\sum_{\ell=1}^{L+1} g_i^\ell \cdot z_i^\ell\right) where q(⋅)q(\cdot) is a linear query projection function, giℓ∈R1g_i^\ell \in \mathbb{R}^1 is the gating scalar at index ii for level ℓ\ell, and ziℓ∈RCz_i^\ell \in \mathbb{R}^C is the feature vector at index ii for level ℓ\ell.

  2. Knowl 2 — Focal Modulation Forward Algorithm

    algorithm

    The focal modulation forward pass takes a 2D feature map tensor, projects it into query, context, and gating streams, computes hierarchical depth-wise context features, performs gated aggregation across focal levels, and modulates the query features.

    Input: Feature tensor XX of shape (B,H,W,C)(B, H, W, C), number of focal levels LL, kernel sizes {k1,…,kL}\{k^1, \dots, k^L\}
    Output: Output feature tensor YY of shape (B,H,W,C)(B, H, W, C)
    Linear projections: pj_in:C→(2C+L+1)pj\_in: C \to (2C + L + 1), pj_cxt:C→Cpj\_cxt: C \to C (via 1×11\times 1 conv), pj_out:C→Cpj\_out: C \to C
    Hierarchical layers: hc_layers=[Sequential(DWConv(C,kℓ),GeLU()) for ℓ=1…L]hc\_layers = [\text{Sequential}(\text{DWConv}(C, k^\ell), \text{GeLU}()) \text{ for } \ell=1 \dots L]
    1. Project input: A=pj_in(X)A = pj\_in(X) permuted to (B,2C+L+1,H,W)(B, 2C + L + 1, H, W)
    2. Split channels: q,z,gate=split(A,[C,C,L+1],dim=1)q, z, gate = \text{split}(A, [C, C, L + 1], \text{dim}=1)
    3. Initialize modulator accumulator m=0m = 0
    4. for ℓ=1\ell = 1 to LL:
    5. z=hc_layers[ℓ](z)z = hc\_layers[\ell](z)
    6. m=m+z⊙gate[:,ℓ,:,:]m = m + z \odot gate[:, \ell, :, :]
    7. z_global=GeLU(mean(z,axes=(2,3)))z\_global = \text{GeLU}(\text{mean}(z, \text{axes}=(2,3)))
    8. m=m+z_global⊙gate[:,L+1,:,:]m = m + z\_global \odot gate[:, L+1, :, :]
    9. Apply cross-channel projection: m=pj_cxt(m)m = pj\_cxt(m)
    10. Modulate query: x=q⊙mx = q \odot m
    11. Permute xx back to (B,H,W,C)(B, H, W, C)
    12. return pj_out(x)pj\_out(x)
  3. Knowl 3 — Computational and Parameter Complexity of Focal Modulation

    equation

    For an input feature map of spatial resolution H×WH \times W with channel dimension CC, LL focal levels, and depth-wise convolution kernel sizes kℓk^\ell at level ℓ∈{1,…,L}\ell \in \{1, \dots, L\}, the total number of learnable parameters in a focal modulation module is: Params=3C2+C(L+1)+C∑ℓ=1L(kℓ)2\text{Params} = 3C^2 + C(L + 1) + C\sum_{\ell=1}^L (k^\ell)^2 where 3C23C^2 corresponds to the query projection q(⋅)q(\cdot), context input projection fz(⋅)f_z(\cdot), and context output projection h(⋅)h(\cdot); C(L+1)C(L+1) corresponds to the linear gating projection fg(⋅)f_g(\cdot); and C∑ℓ=1L(kℓ)2C\sum_{\ell=1}^L (k^\ell)^2 accounts for the LL depth-wise convolutional layers.

    The time complexity (FLOPs) for processing a single feature map is: O(HW×(3C2+C(2L+3)+C∑ℓ=1L(kℓ)2))\mathcal{O}\left(HW \times \left(3C^2 + C(2L + 3) + C\sum_{\ell=1}^L (k^\ell)^2\right)\right) where the element-wise multiplications introduce O(C(L+2))\mathcal{O}(C(L+2)) operations per visual token.

    In comparison, standard self-attention in Vision Transformers has computational complexity O((HW)2C+HW×3C2)\mathcal{O}((HW)^2 C + HW \times 3C^2), exhibiting quadratic growth with respect to token count HWHW, while window-based self-attention with window size ww has complexity O(HW×(3C2+2Cw2))\mathcal{O}(HW \times (3C^2 + 2Cw^2)).

  4. Knowl 4 — Focal Modulation Network Architectural Configurations

    model/method

    Focal Modulation Networks (FocalNets) use a hierarchical four-stage architecture with patch embedding downsampling layers at the input (kernel size 4×44 \times 4, stride 44) and between stages (kernel size 2×22 \times 2, stride 22), replacing self-attention modules with Focal Modulation blocks.

    FocalNet model variants are defined by layer depths across stages [d1,d2,d3,d4][d_1, d_2, d_3, d_4], channel dimensions [C1,C2,C3,C4][C_1, C_2, C_3, C_4], number of focal levels LL, and initial depth-wise kernel size k1k^1 (with higher levels increasing by step 2: kℓ=kℓ−1+2k^\ell = k^{\ell-1} + 2):

    • FocalNet-T (Tiny): Depths [2,2,6,2][2, 2, 6, 2], dimensions [96,192,384,768][96, 192, 384, 768].
    • FocalNet-S (Small): Depths [2,2,18,2][2, 2, 18, 2], dimensions [96,192,384,768][96, 192, 384, 768].
    • FocalNet-B (Base): Depths [2,2,18,2][2, 2, 18, 2], dimensions [128,256,512,1024][128, 256, 512, 1024].
    • FocalNet-L (Large): Depths [2,2,18,2][2, 2, 18, 2], dimensions [192,384,768,1536][192, 384, 768, 1536].
    • FocalNet-H (Huge): Depths [2,2,18,2][2, 2, 18, 2], dimensions [352,704,1408,2816][352, 704, 1408, 2816], L=4L=4, kernel sizes [3,5,7,9][3, 5, 7, 9], effective receptive field rL=21r^L = 21.

    Each model size provides two receptive field configurations across all stages:

    • Small Receptive Field (SRF): L=2L=2, kernel sizes [3,5][3, 5], effective receptive field rL=7r^L = 7.

    • Large Receptive Field (LRF): L=3L=3, kernel sizes [3,5,7][3, 5, 7], effective receptive field rL=13r^L = 13.

    • Monolithic FocalNet: Adapts the isotropic Vision Transformer (ViT) layout with patch size 16×1616 \times 16, L=3L=3, kernel sizes [3,5,7][3, 5, 7], and uniform channel dimensions (192 for Tiny, 384 for Small, 768 for Base).

  5. Knowl 5 — Ablation Analysis on Focal Modulation Components and Design Variants

    data/table

    Ablations on FocalNet-T (LRF) evaluated on ImageNet-1K classification (and downstream COCO object detection) isolate the individual contributions of modulation operations, contextualization layers, and aggregation schemes.

    Model Variant FLOPs (G) Throughput (imgs/s) Top-1 Acc (%) Box APb\text{AP}^b Mask APm\text{AP}^m
    FocalNet-T (LRF, default) 4.48 696 82.3 46.2 41.6
    Additive modulation (yi=q(xi)+miy_i = q(x_i) + m_i) 4.49 670 81.5 45.6 41.1
    No global pooling (ZL+1Z^{L+1} removed) 4.48 683 82.0 45.8 41.2
    Top-only aggregation (only ZLZ^L used) 4.49 698 81.9 45.7 41.2
    No gating (GℓG^\ell uniform) 4.48 707 81.9 45.6 41.1
    Depth-wise ConvNet baseline 4.47 738 81.6 - -
    Pooling aggregator (MetaFormer-style) 4.37 676 80.5 - -
    Global pooling only (SE-Net style) 4.36 883 75.7 - -
    Multi-scale self-attention (QKV first) 4.61 456 81.5 - -
    Sliding-window self-attention 4.49 103 81.5 - -

    Varying the number of focal levels LL with kernel progression kℓ=2ℓ+1k^\ell = 2\ell + 1 reveals that hierarchical multi-level context outperforms single large kernels:

    • L=0L=0 (global pooling only): 75.7% Top-1
    • L=1L=1 (k=3k=3, receptive field 3): 82.0% Top-1
    • L=2L=2 (k∈{3,5}k \in \{3, 5\}, receptive field 7): 82.1% Top-1
    • L=3L=3 (k∈{3,5,7}k \in \{3, 5, 7\}, receptive field 13): 82.3% Top-1
    • L=4L=4 (k∈{3,5,7,9}k \in \{3, 5, 7, 9\}, receptive field 21): 82.2% Top-1
    • L=1L=1 with a single large kernel (k=13k=13, receptive field 13): 81.9% Top-1 (0.4% lower than L=3L=3 with the same receptive field).
  6. Knowl 6 — ImageNet Classification and Zero-Shot Transfer Performance

    data/table

    FocalNets evaluated on ImageNet-1K benchmark with direct training and ImageNet-22K pretraining outperform ConvNet, Transformer, and MLP baselines with comparable model parameters and FLOPs.

    Model Parameters (M) FLOPs (G) Throughput (imgs/s) Top-1 Accuracy (%)
    ResNet-50 25.0 4.1 1294 76.2
    DeiT-Small/16 22.1 4.6 939 79.9
    Swin-Tiny 28.3 4.5 760 81.2
    FocalAtt-Tiny 28.9 4.9 319 82.2
    FocalNet-T (SRF) 28.4 4.4 743 82.1
    FocalNet-T (LRF) 28.6 4.5 696 82.3
    Swin-Small 49.6 8.7 435 83.1
    FocalAtt-Small 51.1 9.4 192 83.5
    FocalNet-S (SRF) 49.9 8.6 434 83.4
    FocalNet-S (LRF) 50.3 8.7 406 83.5
    Swin-Base 87.8 15.4 291 83.5
    FocalAtt-Base 89.8 16.4 138 83.8
    FocalNet-B (SRF) 88.1 15.3 280 83.7
    FocalNet-B (LRF) 88.7 15.4 269 83.9
    Pretrained on ImageNet-22K (2242^2 input, finetuned at 2242^2 / 3842^2):
    Swin-Base 88.0 15.4 / 47.1 291 / 91 85.2 / 86.4
    FocalNet-B 88.1 15.3 / 44.8 280 / 94 85.6 / 86.5
    Swin-Large 196.5 34.5 / 104.0 155 / 49 86.3 / 87.3
    FocalNet-L 197.1 34.2 / 100.6 144 / 50 86.5 / 87.3

    When evaluated in language-image contrastive learning using UniCL pretraining across 20 downstream datasets on the ELEVATER / ICinW benchmark, FocalNet-B achieves an average score of 44.0% and zero-shot ImageNet-1K top-1 accuracy of 54.2%, outperforming Swin-B (43.2% mean accuracy and 52.2% ImageNet-1K accuracy).

  7. Knowl 7 — COCO Object Detection and Instance Segmentation Performance

    data/table

    FocalNet backbones evaluated on COCO 2017 object detection and instance segmentation using Mask R-CNN across standard 1×1\times (12 epochs) and 3×3\times (36 epochs) schedules outperform Swin and Focal Transformers.

    Backbone Params FLOPs Mask R-CNN 1×1\times Mask R-CNN 3×3\times
    (M) (G) APb\text{AP}^b AP50b\text{AP}^b_{50} APm\text{AP}^m APb\text{AP}^b AP50b\text{AP}^b_{50} APm\text{AP}^m
    Swin-Tiny 47.8 264 43.7 66.6 39.8 46.0 68.1 41.6
    FocalAtt-Tiny 48.8 291 44.8 67.7 41.0 47.2 69.4 42.7
    FocalNet-T (SRF) 48.6 267 45.9 68.3 41.3 47.6 69.5 42.6
    FocalNet-T (LRF) 48.9 268 46.1 68.2 41.5 48.0 69.7 42.9
    Swin-Small 69.1 354 46.5 68.7 42.1 48.5 70.2 43.3
    FocalAtt-Small 71.2 401 47.4 69.8 42.8 48.8 70.5 43.8
    FocalNet-S (SRF) 70.8 356 48.0 69.9 42.7 48.9 70.1 43.6
    FocalNet-S (LRF) 72.3 365 48.3 70.5 43.1 49.3 70.7 43.8
    Swin-Base 107.1 497 46.9 69.2 42.3 48.5 69.8 43.4
    FocalAtt-Base 110.0 533 47.8 70.2 43.2 49.0 70.1 43.7
    FocalNet-B (SRF) 109.4 496 48.8 70.7 43.3 49.6 70.6 44.1
    FocalNet-B (LRF) 111.4 507 49.0 70.9 43.5 49.8 70.9 44.1

    FocalNet-T trained with the 1×1\times schedule (45.9 APb\text{AP}^b) nearly matches Swin-T trained with the 3×3\times schedule (46.0 APb\text{AP}^b), and FocalNet-B 1×1\times (48.8 APb\text{AP}^b) surpasses Swin-B 3×3\times (48.5 APb\text{AP}^b).

    When scaled up with the DINO detector, FocalNet-H (746M parameters, pretrained on ImageNet-22K and Object365) achieves 64.0 / 64.2 AP\text{AP} on COCO val2017 (without / with test-time augmentation) and 64.1 / 64.4 AP\text{AP} on test-dev, surpassing larger attention models including SwinV2-G (3.0B params, 63.1 AP\text{AP}) and BEiT-3 (1.9B params, 63.7 AP\text{AP}).

  8. Knowl 8 — ADE20K Semantic Segmentation and COCO Panoptic Segmentation Performance

    data/table

    FocalNets evaluated on ADE20K semantic segmentation (with UPerNet and Mask2Former) and COCO panoptic segmentation outperform self-attention architectures across various scales.

    Backbone Segmentation Head Pretrain Data Parameters (M) Single-Scale mIoU Multi-Scale mIoU
    Swin-Tiny UPerNet ImageNet-1K 60 44.5 45.8
    FocalAtt-Tiny UPerNet ImageNet-1K 62 45.8 47.0
    FocalNet-T (SRF) UPerNet ImageNet-1K 61 46.5 47.2
    FocalNet-T (LRF) UPerNet ImageNet-1K 61 46.8 47.8
    Swin-Small UPerNet ImageNet-1K 81 47.6 49.5
    FocalAtt-Small UPerNet ImageNet-1K 85 48.0 50.0
    FocalNet-S (SRF) UPerNet ImageNet-1K 83 49.3 50.1
    FocalNet-S (LRF) UPerNet ImageNet-1K 84 49.1 50.1
    Swin-Base UPerNet ImageNet-1K 121 48.1 49.7
    FocalAtt-Base UPerNet ImageNet-1K 126 49.0 50.5
    FocalNet-B (SRF) UPerNet ImageNet-1K 124 50.2 51.1
    FocalNet-B (LRF) UPerNet ImageNet-1K 126 50.5 51.4
    Swin-Large Mask2Former ImageNet-22K 216 56.4 57.7
    FocalNet-L Mask2Former ImageNet-22K 218 57.3 58.5

    For COCO panoptic segmentation using Mask2Former (200 queries), FocalNet-L pretrained on ImageNet-22K achieves 57.9 PQ, 48.4 AP (instance segmentation), and 67.3 mIoU (semantic segmentation), compared to Swin-L with Mask2Former which achieves 57.8 PQ, 48.6 AP, and 67.4 mIoU.

  9. Knowl 9 — Emergent Object Localization and Intrinsic Interpretability in FocalNets

    empirical result

    Focal Modulation Networks exhibit an intrinsic localization capability without requiring class activation mapping (CAM) or gradient-based explanation tools (such as Grad-CAM):

    1. Modulator Attention Heatmaps: Computing the L2L_2-norm magnitude of the modulator vector M∈RH×W×CM \in \mathbb{R}^{H \times W \times C} at the final network layer produces spatial heatmaps that naturally and tightly focus on the primary foreground objects inducing the image category classification.

    2. Focal Level Role Specialization: Inspection of the gating weights Gℓ∈RH×W×1G^\ell \in \mathbb{R}^{H \times W \times 1} across hierarchical levels demonstrates progressive functional separation:

      • Level 1 (ℓ=1\ell = 1): Focuses strongly on high-frequency, fine-grained visual textures within foreground object interiors.
      • Level 2 (ℓ=2\ell = 2): Peaks along salient object boundaries and edges.
      • Level 3 (ℓ=3\ell = 3): Covers entire object bodies uniformly.
      • Global Level (ℓ=4\ell = 4): Exhibits high gating activations over uniform background regions while assigning low weights to foreground objects, confirming that background tokens gather broad contextual information while foreground tokens rely on localized context.
  10. Knowl 10 — Limitations in Cross-Modal Extension and Domain Generalization

    limitation

    While Focal Modulation provides an efficient replacement for self-attention in computer vision, two key limitations remain:

    1. Non-Visual Domain Generalization: The hierarchical context aggregation relies on depth-wise convolutions and 2D spatial locality priors. Extending this formulation to 1D discrete domains (such as natural language processing) requires further investigation into how token interactions can be formulated without standard 2D grid assumptions.
    2. Cross-Modality Modeling: In multimodal architectures, standard self-attention readily converts to cross-attention by swapping key-value sequences between different modalities. Focal modulation relies on localized spatial feature gathering around target queries; formulating an equivalent cross-modulation mechanism across heterogeneous modalities remains an open problem.

Coverage note — None omitted. All primary architectural formulations, mathematical complexity bounds, algorithmic specifications, ImageNet/COCO/ADE20K benchmarking experiments, component ablations, visual interpretability analyses, and stated limitations are fully covered.

References

  1. 1.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154–6162, 2018.
  2. 2.Yue Cao, Jiarui Xu, Stephen Lin, Fangyun Wei, and Han Hu. Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020.
  4. 4.Shuning Chang, Pichao Wang, Fan Wang, Hao Li, and Jiashi Feng. Augmented transformer with adaptive graph for temporal action proposal generation. arXiv preprint arXiv:2103.16024, 2021.
  5. 5.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification, 2021.
  6. 6.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 801–818, 2018.
  7. 7.Qiang Chen, Qiman Wu, Jian Wang, Qinghao Hu, Tao Hu, Errui Ding, Jian Cheng, and Jingdong Wang. Mixformer: Mixing features across windows and dimensions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5249–5259, 2022.
  8. 8.Shoufa Chen, Enze Xie, Chongjian Ge, Ding Liang, and Ping Luo. CycleMLP: A mlp-like architecture for dense prediction. arXiv preprint arXiv:2107.10224, 2021.
  9. 9.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. arXiv preprint arXiv:2103.15436, 2021.
  10. 10.Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. Dynamic convolution: Attention over convolution kernels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11030–11039, 2020.
  11. 11.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. arXiv preprint arXiv:2205.08534, 2022.
  12. 12.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. arXiv preprint arXiv:2112.01527, 2021.
  13. 13.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022.
  14. 14.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34, 2021.
  15. 15.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting spatial attention design in vision transformers. arXiv preprint arXiv:2104.13840, 2021.
  16. 16.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020.
  17. 17.Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7373–7382, 2021.
  18. 18.Zhigang Dai, Bolun Cai, Yugeng Lin, and Junying Chen. Up-detr: Unsupervised pre-training for object detection with transformers. arXiv preprint arXiv:2011.09094, 2020.
  19. 19.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  20. 20.Mingyu Ding, Bin Xiao, Noel Codella, Ping Luo, Jingdong Wang, and Lu Yuan. Davit: Dual attention vision transformers. arXiv preprint arXiv:2204.03645, 2022.
  21. 21.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021.
  22. 22.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  23. 23.Peng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi, and Aniruddha Kembhavi. Container: Context aggregation network. arXiv preprint arXiv:2106.01401, 2021.
  24. 24.Chengyue Gong, Dilin Wang, Meng Li, Vikas Chandra, and Qiang Liu. Vision transformers with patch diversification, 2021.
  25. 25.Jianyuan Guo, Kai Han, Han Wu, Chang Xu, Yehui Tang, Chunjing Xu, and Yunhe Wang. Cmt: Convolutional neural networks meet vision transformers. arXiv preprint arXiv:2107.06263, 2021.
  26. 26.Jianyuan Guo, Yehui Tang, Kai Han, Xinghao Chen, Han Wu, Chao Xu, Chang Xu, and Yunhe Wang. Hire-mlp: Vision mlp via hierarchical rearrangement. arXiv preprint arXiv:2108.13341, 2021.
  27. 27.Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. A survey on visual transformer. arXiv preprint arXiv:2012.12556, 2020.
  28. 28.Qi Han, Zejia Fan, Qi Dai, Lei Sun, Ming-Ming Cheng, Jiaying Liu, and Jingdong Wang. Demystifying local vision transformer: Sparse connectivity, weight sharing, and dynamic weight. arXiv preprint arXiv:2106.04263, 2021.
  29. 29.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  30. 30.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  31. 31.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  32. 32.Qibin Hou, Zihang Jiang, Li Yuan, Ming-Ming Cheng, Shuicheng Yan, and Jiashi Feng. Vision permutator: A permutable mlp-like architecture for visual recognition. arXiv preprint arXiv:2106.12368, 2021.
  33. 33.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  34. 34.Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. Local relation networks for image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3464–3473, 2019.
  35. 35.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018.
  36. 36.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In European conference on computer vision, pages 646–661. Springer, 2016.
  37. 37.Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, and Humphrey Shi. Semask: Semantically masked transformers for semantic segmentation. arXiv preprint arXiv:2112.12782, 2021.
  38. 38.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey. arXiv preprint arXiv:2101.01169, 2021.
  39. 39.Youngwan Lee, Jonghee Kim, Jeff Willette, and Sung Ju Hwang. Mpvit: Multi-path vision transformer for dense prediction. arXiv preprint arXiv:2112.11010, 2021.
  40. 40.Youngwan Lee, Jonghee Kim, Jeffrey Willette, and Sung Ju Hwang. Mpvit: Multi-path vision transformer for dense prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7287–7296, 2022.
  41. 41.Bing Li, Cheng Zheng, Silvio Giancola, and Bernard Ghanem. Sctn: Sparse convolution-transformer network for scene flow estimation. arXiv preprint arXiv:2105.04447, 2021.
  42. 42.Chunyuan Li, Haotian Liu, Liunian Harold Li, Pengchuan Zhang, Jyoti Aneja, Jianwei Yang, Ping Jin, Houdong Hu, Zicheng Liu, Yong Jae Lee, and Jianfeng Gao. ELEVATER: A benchmark and toolkit for evaluating language-augmented visual models. In NeurIPS Track on Datasets and Benchmarks, 2022.
  43. 43.Duo Li, Jie Hu, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu, Tong Zhang, and Qifeng Chen. Involution: Inverting the inherence of convolution for visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12321–12330, 2021.
  44. 44.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10965–10975, 2022.
  45. 45.Xiangyu Li, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu, and Wanqing Li. Trear: Transformer-based rgb-d egocentric action recognition. arXiv preprint arXiv:2101.03904, 2021.
  46. 46.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021.
  47. 47.Zhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, Tong Lu, and Ping Luo. Panoptic segformer: Delving deeper into panoptic segmentation with transformers, 2021.
  48. 48.Dongze Lian, Zehao Yu, Xing Sun, and Shenghua Gao. As-mlp: An axial shifted mlp architecture for vision. arXiv preprint arXiv:2107.08391, 2021.
  49. 49.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014.
  50. 50.Hanxiao Liu, Zihang Dai, David So, and Quoc Le. Pay attention to mlps. Advances in Neural Information Processing Systems, 34, 2021.
  51. 51.Hanxiao Liu, Zihang Dai, David R So, and Quoc V Le. Pay attention to MLPs. arXiv preprint arXiv:2105.08050, 2021.
  52. 52.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. arXiv preprint arXiv:2111.09883, 2021.
  53. 53.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022.
  54. 54.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  55. 55.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. arXiv preprint arXiv:2201.03545, 2022.
  56. 56.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  57. 57.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  58. 58.Yuxuan Lou, Fuzhao Xue, Zangwei Zheng, and Yang You. Sparse-mlp: A fully-mlp architecture with conditional computation. arXiv preprint arXiv:2109.02008, 2021.
  59. 59.Lingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan, Zuxuan Wu, Yu-Gang Jiang, and Ser-Nam Lim. Adavit: Adaptive vision transformers for efficient image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12309–12318, 2022.
  60. 60.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Advances in neural information processing systems, 34:13937–13949, 2021.
  61. 61.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017.
  62. 62.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In Proceedings of the IEEE international conference on computer vision, pages 8430–8439, 2019.
  63. 63.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  64. 64.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. arXiv preprint arXiv:2105.05633, 2021.
  65. 65.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In CVPR, 2019.
  66. 66.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al. Sparse r-cnn: End-to-end object detection with learnable proposals. arXiv preprint arXiv:2011.12450, 2020.
  67. 67.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.
  68. 68.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  69. 69.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning, pages 6105–6114. PMLR, 2019.
  70. 70.Chuanxin Tang, Yucheng Zhao, Guangting Wang, Chong Luo, Wenxuan Xie, and Wenjun Zeng. Sparse mlp for image recognition: Is self-attention really necessary? arXiv preprint arXiv:2109.05422, 2021.
  71. 71.Yehui Tang, Kai Han, Jianyuan Guo, Chang Xu, Yanxi Li, Chao Xu, and Yunhe Wang. An image patch is a wave: Phase-aware vision mlp. arXiv preprint arXiv:2111.12294, 2021.
  72. 72.Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, et al. MLP-mixer: An all-mlp architecture for vision. arXiv preprint arXiv:2105.01601, 2021.
  73. 73.Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems, 34, 2021.
  74. 74.Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, and Hervé Jégou. ResMLP: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404, 2021.
  75. 75.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020.
  76. 76.Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Piotr Bojanowski, Armand Joulin, Gabriel Synnaeve, and Hervé Jégou. Augmenting convolutional networks with attention-based aggregation. arXiv preprint arXiv:2112.13692, 2021.
  77. 77.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 32–42, 2021.
  78. 78.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12894–12904, 2021.
  79. 79.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  80. 80.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. arXiv preprint arXiv:2012.00759, 2020.
  81. 81.Ning Wang, Wengang Zhou, Jie Wang, and Houqaing Li. Transformer meets tracker: Exploiting temporal context for robust visual tracking. arXiv preprint arXiv:2103.11681, 2021.
  82. 82.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122, 2021.
  83. 83.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415–424, 2022.
  84. 84.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022.
  85. 85.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018.
  86. 86.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. arXiv preprint arXiv:2011.14503, 2020.
  87. 87.Yixuan Wei, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. Contrastive learning rivals masked image modeling in fine-tuning via feature distillation. arXiv preprint arXiv:2205.14141, 2022.
  88. 88.Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm. arXiv preprint arXiv:2110.00476, 2021.
  89. 89.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808, 2021.
  90. 90.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), pages 418–434, 2018.
  91. 91.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers, 2021.
  92. 92.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1492–1500, 2017.
  93. 93.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3060–3069, 2021.
  94. 94.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9981–9990, 2021.
  95. 95.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao. Focal self-attention for local-global interactions in vision transformers. arXiv preprint arXiv:2107.00641, 2021.
  96. 96.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao, Lu Yuan, Ce Liu, and Jianfeng Gao. UniCL: unified contrastive learning in image-text-label space. CVPR, 2022.
  97. 97.Jianwei Yang, Zhile Ren, Chuang Gan, Hongyuan Zhu, and Devi Parikh. Cross-channel communication networks. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pages 1297–1306, 2019.
  98. 98.Hongxu Yin, Arash Vahdat, Jose M Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-vit: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10809–10818, 2022.
  99. 99.Tan Yu, Xu Li, Yunfeng Cai, Mingming Sun, and Ping Li. S^2-MLPv2: Improved spatial-shift mlp architecture for vision. arXiv preprint arXiv:2108.01072, 2021.
  100. 100.Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. arXiv preprint arXiv:2111.11418, 2021.
  101. 101.Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan. Volo: Vision outlooker for visual recognition. arXiv preprint arXiv:2106.13112, 2021.
  102. 102.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021.
  103. 103.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. arXiv preprint arXiv:1909.11065, 2019.
  104. 104.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019.
  105. 105.Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Haibin Lin, Zhi Zhang, Yue Sun, Tong He, Jonas Mueller, R Manmatha, et al. Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955, 2020.
  106. 106.Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv preprint arXiv:2203.03605, 2022.
  107. 107.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  108. 108.Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. arXiv preprint arXiv:2103.15358, 2021.
  109. 109.Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z Li. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9759–9768, 2020.
  110. 110.Wenwei Zhang, Jiangmiao Pang, Kai Chen, and Chen Change Loy. K-net: Towards unified image segmentation, 2021.
  111. 111.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6848–6856, 2018.
  112. 112.Jiaojiao Zhao, Xinyu Li, Chunhui Liu, Shuai Bing, Hao Chen, Cees GM Snoek, and Joseph Tighe. Tuber: Tube-transformer for action detection. arXiv preprint arXiv:2104.00969, 2021.
  113. 113.Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. Mixing and shifting: Exploiting global and local dependencies in vision mlps. arXiv preprint arXiv:2202.06510, 2022.
  114. 114.Minghang Zheng, Peng Gao, Xiaogang Wang, Hongsheng Li, and Hao Dong. End-to-end object detection with adaptive clustering transformer. arXiv preprint arXiv:2011.09315, 2020.
  115. 115.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. arXiv preprint arXiv:2012.15840, 2020.
  116. 116.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 13001–13008, 2020.
  117. 117.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2921–2929, 2016.
  118. 118.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 633–641, 2017.
  119. 119.Daquan Zhou, Yujun Shi, Bingyi Kang, Weihao Yu, Zihang Jiang, Yuan Li, Xiaojie Jin, Qibin Hou, and Jiashi Feng. Refiner: Refining self-attention for vision transformers. arXiv preprint arXiv:2106.03714, 2021.
  120. 120.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.

Citation

MLA
Yang, J., et al. “Focal Modulation Networks”. arXiv, 2022, http://arxiv.org/abs/2203.11926v3.
APA
Yang, J., Li, C., Dai, X., Yuan, L., & Gao, J. (2022). Focal Modulation Networks. arXiv. http://arxiv.org/abs/2203.11926v3
Chicago
Yang, J., C. Li, X. Dai, L. Yuan, and J. Gao. 2022. “Focal Modulation Networks”. arXiv. http://arxiv.org/abs/2203.11926v3.
Harvard
Yang, J. et al. (2022) “Focal Modulation Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.11926v3.
Vancouver
1. Yang J, Li C, Dai X, Yuan L, Gao J (2022) Focal Modulation Networks. arXiv

BibTeX

@article{yang2022focal,
  title = {Focal Modulation Networks},
  author = {Yang, Jianwei and Li, Chunyuan and Dai, Xiyang and Yuan, Lu and Gao, Jianfeng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.11926v3},
  eprint = {2203.11926}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission