InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

Wenhai WangJifeng DaiZhe ChenZhenhang HuangZhiqi LiXizhou ZhuXiaowei HuTong LuLewei LuHongsheng Li

article2023CVPR1,203 citations

Presents InternImage, a billion-parameter convolutional foundation model built on dynamic deformable convolutions that matches the scaling capabilities and downstream performance of state-of-the-art vision transformers across ImageNet, COCO, and ADE20K benchmarks.

Listen

Recent advances in computer vision have been dominated by vision transformers, which scale effectively to billions of parameters and vast datasets. In contrast, traditional convolutional neural networks have lagged behind at massive scales due to rigid architectures and static, localized operations. However, vision transformers often suffer from heavy computational and memory costs on high-resolution image tasks. The article addresses this challenge by exploring whether a modernized convolutional architecture can scale just as effectively as transformers while remaining computationally efficient.

The main objective of the article is to develop and evaluate a large-scale convolutional foundation model, named InternImage, demonstrating that convolutional networks can achieve state-of-the-art visual recognition performance when expanded to over one billion parameters and trained on hundreds of millions of images.

To accomplish this, the authors designed an improved core operator based on deformable convolution, which dynamically learns flexible sampling locations and adapts to input data using standard three-by-three filter windows. This operator incorporates shared weights across sampling points to cut memory usage, a multi-group structure to learn diverse feature patterns, and normalized modulation scales to ensure training stability. The authors then built a family of models ranging from 30 million to over one billion parameters, applying structured block-stacking and dimension-scaling rules. The models were evaluated across major benchmarks for image classification, object detection, and semantic segmentation using standard evaluation protocols and training data ranging from one million to 427 million images.

The findings show that InternImage consistently matches or outperforms leading vision transformers and prior convolutional models. On standard image classification, the base model achieved an 84.9% accuracy score, surpassing comparable convolutional architectures by at least 1.1 points, while the largest one-billion-parameter variant achieved 89.6%. On the challenging COCO object detection benchmark, the flagship model set a state-of-the-art record score of 65.4, exceeding the leading transformer baseline by 2.3 points while requiring 27% fewer parameters. In semantic segmentation on the ADE20K benchmark, the model attained a record 62.9 score, outperforming previous top models. Ablation analyses confirmed that sharing projection weights reduced GPU memory usage by up to 84.2% at the largest scale without sacrificing accuracy.

These results carry significant implications for computer vision research and real-world deployment. They prove that transformers are not the only viable path for large-scale vision foundation models. By retaining the efficient inductive biases of convolutions while gaining the adaptive properties of transformers, InternImage delivers higher task accuracy with lower computational resource demands. This improves performance in downstream tasks like dense visual perception while moderating infrastructure and training costs.

Organizations developing large-scale visual perception systems should consider deformable convolutional architectures alongside vision transformers. Further work is recommended to optimize the runtime latency of deformable operations on specialized deployment hardware to support ultra-fast inference settings. In addition, continued research into large-scale pre-training across larger multimodal datasets will help fully map the capabilities of convolutional foundation models.

The findings are supported by comprehensive benchmark testing across multiple model sizes and vision domains. Readers should note that large-scale convolutional models remain at an early developmental stage, and empirical results for the largest model relied on a composite training dataset rather than identical proprietary data used by some competing transformer baselines. Nonetheless, confidence is high that the architectural designs provide a robust, scalable alternative for advanced computer vision applications.

Cover for InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

Abstract

Compared to the great progress of large-scale vision transformers (ViTs) in recent years, large-scale models based on convolutional neural networks (CNNs) are still in an early state. This work presents a new large-scale CNN-based foundation model, termed InternImage, which can obtain the gain from increasing parameters and training data like ViTs. Different from the recent CNNs that focus on large dense kernels, InternImage takes deformable convolution as the core operator, so that our model not only has the large effective receptive field required for downstream tasks such as detection and segmentation, but also has the adaptive spatial aggregation conditioned by input and task information. As a result, the proposed InternImage reduces the strict inductive bias of traditional CNNs and makes it possible to learn stronger and more robust patterns with large-scale parameters from massive data like ViTs. The effectiveness of our model is proven on challenging benchmarks including ImageNet, COCO, and ADE20K. It is worth mentioning that InternImage-H achieved a new record 65.4 mAP on COCO test-dev and 62.9 mIoU on ADE20K, outperforming current leading CNNs and ViTs.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Method
  • 3.1. Deformable Convolution v3
  • 3.2. InternImage Model
  • 4. Experiment
  • 4.1. Image Classification
  • 4.2. Object Detection
  • 4.3. Semantic Segmentation
  • 4.4. Ablation Study
  • 5. Conclusion & Limitations
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Deformable Convolution v3 Operator Formulation

    model/method

    Deformable Convolution v3 (DCNv3) is an adaptive sparse convolution operator designed as the core building block for large-scale convolutional vision foundation models. For an input feature map x∈RC×H×W\mathbf{x} \in \mathbb{R}^{C \times H \times W} and a reference query pixel location p0p_0, the DCNv3 operator is defined as:

    y(p0)=∑g=1G∑k=1Kwgmgkxg(p0+pk+Δpgk)\mathbf{y}(p_0) = \sum_{g=1}^{G} \sum_{k=1}^{K} \mathbf{w}_g \mathbf{m}_{gk} \mathbf{x}_g(p_0 + p_k + \Delta p_{gk})

    where:

    • GG is the total number of aggregation groups.
    • C′=C/GC' = C / G is the group channel dimension.
    • xg∈RC′×H×W\mathbf{x}_g \in \mathbb{R}^{C' \times H \times W} is the sliced input feature map corresponding to group gg.
    • KK is the number of sampling points (e.g., K=9K = 9 for a 3×33 \times 3 kernel grid).
    • pk∈{(−1,−1),(−1,0),…,(+1,+1)}p_k \in \{(-1,-1), (-1,0), \dots, (+1,+1)\} denotes the kk-th static grid sampling coordinate.
    • Δpgk∈R2\Delta p_{gk} \in \mathbb{R}^2 represents the input-conditioned, continuous sampling coordinate offset for point kk in group gg.
    • wg∈RC×C′\mathbf{w}_g \in \mathbb{R}^{C \times C'} denotes the location-independent projection weight matrix for group gg, shared across all KK sampling points.
    • mgk∈R\mathbf{m}_{gk} \in \mathbb{R} is the dynamic modulation scalar for sampling point kk in group gg, which is normalized across the KK points via softmax such that ∑k=1Kmgk=1\sum_{k=1}^K \mathbf{m}_{gk} = 1 and mgk∈[0,1]\mathbf{m}_{gk} \in [0, 1].

    DCNv3 incorporates three structural modifications over standard DCNv2:

    1. Weight Sharing: The projection weights wg\mathbf{w}_g are decoupled into depth-wise location-aware modulations (mgk\mathbf{m}_{gk}) and a shared point-wise projection (wg\mathbf{w}_g), eliminating linear parameter/memory scaling with KK.
    2. Multi-Group Mechanism: Splitting channels into GG groups allows learning diverse sampling offsets Δpgk\Delta p_{gk} and modulation scales mgk\mathbf{m}_{gk} across different representation subspaces.
    3. Softmax Modulation Normalization: Replacing per-element sigmoid normalization with softmax normalization along the sample dimension KK enforces sum-to-one modulation weights, stabilizing gradient propagation during large-scale training.
  2. Knowl 2 — InternImage Basic Block and Macro-Architecture

    model/method

    The InternImage macro-architecture follows a 4-stage hierarchical design that converts an input image of size H×W×3H \times W \times 3 into multi-scale feature representations downsampled by factors of 4,8,16,4, 8, 16, and 3232.

    Stem Layer: Positioned before Stage 1 to reduce spatial resolution by 4×4\times. It consists of two 3×33 \times 3 convolutions with stride 2 and padding 1, separated by Layer Normalization (LN) and GELU, and ending with an LN layer. The first convolution outputs half the channel dimension of the second (C1/2C_1/2 and C1C_1).

    Downsampling Layers: Positioned between successive stages to halve the spatial dimensions and double the channel capacity. Each downsampling layer is composed of a 3×33 \times 3 convolution with stride 2 and padding 1, followed by an LN layer.

    Basic Block: Adopts a post-normalization Transformer-style residual structure:

    1. An input feature map x\mathbf{x} passes through Layer Normalization (LN).
    2. The normalized representation is processed by a DCNv3 operator. The sampling offsets Δpgk\Delta p_{gk} and modulation scales mgk\mathbf{m}_{gk} are dynamically predicted from the feature map using a separable convolution (a 3×33 \times 3 depth-wise convolution followed by a linear projection layer).
    3. A residual skip connection adds the block input to the DCNv3 output.
    4. The intermediate result passes through a second LN followed by a Feed-Forward Network (FFN) with GELU activation and expanding channel dimension by a factor of 4.
    5. A final residual connection adds the FFN output.
  3. Knowl 3 — Stacking Rules and Compound Scaling Strategy for InternImage

    model/method

    To constrain the 12-parameter architectural search space across 4 stages (stage channels CiC_i, group counts GiG_i, and stage depths LiL_i for i∈{1,2,3,4}i \in \{1, 2, 3, 4\}), InternImage uses four stacking rules:

    1. Ci=2i−1C1C_i = 2^{i-1} C_1
    2. Gi=Ci/C′G_i = C_i / C', where C′C' is the group dimension.
    3. L1=L2=L4L_1 = L_2 = L_4 (an "AABA" block count pattern).
    4. L1≤L3L_1 \le L_3.

    With these rules, any variant is uniquely defined by the 4 hyperparameters (C1,C′,L1,L3)(C_1, C', L_1, L_3). The base model (InternImage-T) is initialized with (C1=64,C′=16,L1=4,L3=18)(C_1=64, C'=16, L_1=4, L_3=18), giving a total depth D=3L1+L3=30D = 3L_1 + L_3 = 30.

    To scale the architecture to larger parameter capacities, compound scaling is applied over depth DD and width C1C_1 using depth factor α\alpha, width factor β\beta, and composite scaling factor ϕ\phi:

    D′=αϕDandC1′=βϕC1D' = \alpha^\phi D \quad \text{and} \quad C'_1 = \beta^\phi C_1

    subject to α≥1\alpha \ge 1, β≥1\beta \ge 1, and αβ1.99≈2\alpha \beta^{1.99} \approx 2, where the exponent 1.991.99 reflects the empirical parameter growth of InternImage when doubling width at constant depth. The optimal scaling factors are α=1.09\alpha = 1.09 and β=1.36\beta = 1.36.

  4. Knowl 4 — InternImage Model Family Architectural Configurations

    data/table

    The InternImage model family spans parameter scales from 30 million to 1.08 billion parameters, configured according to the defined stacking and compound scaling rules.

    Model C1C_1 C′C' L1,L2,L3,L4L_1, L_2, L_3, L_4 #params
    InternImage-T (origin) 64 16 4, 4, 18, 4 30M
    InternImage-S 80 16 4, 4, 21, 4 50M
    InternImage-B 112 16 4, 4, 21, 4 97M
    InternImage-L 160 16 5, 5, 22, 5 223M
    InternImage-XL 192 16 5, 5, 24, 5 335M
    InternImage-H 320 32 6, 6, 32, 6 1.08B

    C1C_1 is the channel dimension of Stage 1, C′C' is the channel capacity per DCNv3 aggregation group, and L1,L2,L3,L4L_1, L_2, L_3, L_4 are the numbers of stacked basic blocks in Stages 1 through 4. For the largest variant (InternImage-H), the group dimension C′C' is adjusted from 16 to 32 to support very large channel widths.

  5. Knowl 5 — Ablation Analysis of DCNv3 Architectural Modifications

    empirical result

    Ablation studies on InternImage-T evaluated on ImageNet-1K classification (top-1 accuracy) and COCO object detection / instance segmentation with Mask R-CNN under a 1×1\times schedule quantify the contribution of each DCNv3 modification:

    Shared w\mathbf{w} Multi-Group Softmax Norm Top-1 Acc (%) APb\text{AP}^{\text{b}} APm\text{AP}^{\text{m}}
    × ✓ ✓ 83.6 47.4 42.6
    ✓ × ✓ 82.3 43.8 40.0
    ✓ ✓ × 65.7 38.7 35.6
    ✓ ✓ ✓ 83.5 47.2 42.5
    • Weight Sharing: Sharing projection weights wg\mathbf{w}_g among sampling points yields comparable accuracy (83.5% vs. 83.6% top-1, 47.2 vs. 47.4 APb\text{AP}^{\text{b}}) while reducing model parameters by 66.1% relative to the unshared baseline at the -T scale. At the -H scale, weight sharing saves 42.0% of model parameters and 84.2% of GPU training memory per image.
    • Multi-Group Spatial Aggregation: Removing multi-group aggregation causes a 1.2% drop in ImageNet top-1 accuracy and a 3.4 drop in COCO APb\text{AP}^{\text{b}}, confirming the necessity of learning subspace-specific sampling patterns.
    • Softmax Normalization: Replacing softmax normalization across sampling points with element-wise sigmoid normalization causes severe gradient instability during training, dropping top-1 accuracy by 17.8% and COCO APb\text{AP}^{\text{b}} by 8.5.
  6. Knowl 6 — Object Detection and Instance Segmentation Performance on COCO

    empirical result

    InternImage backbones evaluated on COCO val2017 and test-dev demonstrate performance improvements over contemporary CNN and Vision Transformer backbones across various detector frameworks:

    1. Standard Frameworks (val2017):

      • Mask R-CNN (1×1\times schedule): InternImage-T achieves 47.2 APb\text{AP}^{\text{b}} / 42.5 APm\text{AP}^{\text{m}}, surpassing Swin-T (42.7 APb\text{AP}^{\text{b}} / 39.3 APm\text{AP}^{\text{m}}) by +4.5 APb\text{AP}^{\text{b}} and ConvNeXt-T (44.2 APb\text{AP}^{\text{b}} / 40.1 APm\text{AP}^{\text{m}}) by +3.0 APb\text{AP}^{\text{b}}. InternImage-B yields 48.8 APb\text{AP}^{\text{b}} / 44.0 APm\text{AP}^{\text{m}}, outperforming Swin-B (46.9 APb\text{AP}^{\text{b}} / 42.3 APm\text{AP}^{\text{m}}) and ConvNeXt-B (47.0 APb\text{AP}^{\text{b}} / 42.7 APm\text{AP}^{\text{m}}).
      • Cascade Mask R-CNN (3×+MS3\times+\text{MS} schedule): InternImage-XL reaches 56.2 APb\text{AP}^{\text{b}} / 48.8 APm\text{AP}^{\text{m}}, exceeding ConvNeXt-XL (55.2 APb\text{AP}^{\text{b}} / 47.7 APm\text{AP}^{\text{m}}) by +1.0 APb\text{AP}^{\text{b}} and +1.1 APm\text{AP}^{\text{m}}.
    2. State-of-the-Art Scaling on COCO test-dev:

      • When scaled up using DINO and composite backbone parameter doubling (CBNet), InternImage-XL (602M parameters) achieves 64.3 APb\text{AP}^{\text{b}}.
      • InternImage-H (2.18B total parameters, pre-trained via M3I on a 427M public dataset and fine-tuned on Objects365 and COCO) achieves 65.0 APb\text{AP}^{\text{b}} on val2017 and 65.4 APb\text{AP}^{\text{b}} on test-dev. This surpasses FD-SwinV2-G (3.00B parameters, 64.2 APb\text{AP}^{\text{b}}) by +1.2 points using 27% fewer parameters without knowledge distillation.
  7. Knowl 7 — ImageNet Classification Performance across Scales

    empirical result

    InternImage models achieve competitive or superior top-1 accuracy on ImageNet-1K across model and training data scales:

    • ImageNet-1K Training from Scratch (300 epochs, 224×224224 \times 224 input):

      • InternImage-T (30M params, 5G FLOPs): 83.5% (vs. Swin-T 81.3%, ConvNeXt-T 82.1%, CoAtNet-0 81.6%).
      • InternImage-S (50M params, 8G FLOPs): 84.2% (vs. Swin-S 83.0%, ConvNeXt-S 83.1%, CoAtNet-1 83.3%).
      • InternImage-B (97M params, 16G FLOPs): 84.9% (vs. Swin-B 83.5%, ConvNeXt-B 83.8%, RepLKNet-31B 83.5%, CoAtNet-2 84.1%).
    • ImageNet-22K Pre-training (90 epochs pre-training, fine-tuned on 1K at 384×384384 \times 384):

      • InternImage-L (223M params, 108G FLOPs): 87.7% (vs. Swin-L 87.3%, ConvNeXt-L 87.5%, RepLKNet-31L 86.6%).
      • InternImage-XL (335M params, 163G FLOPs): 88.0% (vs. SwinV2-L 87.6%, ConvNeXt-XL 87.8%, RepLKNet-XL 87.8%).
    • Large-Scale Multi-Modal Pre-training (427M public dataset: LAION-400M + YFCC-15M + CC12M with M3I pre-training):

      • InternImage-H (1.08B params, 224×224224 \times 224 input, 188G FLOPs): 88.9%.
      • InternImage-H (1.08B params, 640×640640 \times 640 input, 1478G FLOPs): 89.6%, approaching large-scale ViTs trained on private datasets (e.g., SwinV2-G at 90.2% and ViT-G/14 at 90.5%).
  8. Knowl 8 — Semantic Segmentation Performance on ADE20K

    empirical result

    Semantic segmentation evaluations on the ADE20K validation set using UperNet and Mask2Former report Single-Scale (SS) and Multi-Scale (MS) mean Intersection-over-Union (mIoU):

    • UperNet Framework (512×512512 \times 512 crop):

      • InternImage-T (59M params, 944G FLOPs): 47.9 mIoU (SS) / 48.1 mIoU (MS), outperforming Swin-T (44.5 / 45.8) and ConvNeXt-T (46.0 / 46.7).
      • InternImage-S (80M params, 1017G FLOPs): 50.1 mIoU (SS) / 50.9 mIoU (MS), outperforming Swin-S (47.6 / 49.5) and ConvNeXt-S (48.7 / 49.6).
      • InternImage-B (128M params, 1185G FLOPs): 50.8 mIoU (SS) / 51.3 mIoU (MS), outperforming Swin-B (48.1 / 49.7), ConvNeXt-B (49.1 / 49.9), and RepLKNet-31B (49.9 / 50.6).
    • UperNet Framework (640×640640 \times 640 crop, ImageNet-22K pre-trained):

      • InternImage-L (256M params, 2526G FLOPs): 53.9 mIoU (SS) / 54.1 mIoU (MS).
      • InternImage-XL (368M params, 3142G FLOPs): 55.0 mIoU (SS) / 55.3 mIoU (MS), outperforming ConvNeXt-XL (53.6 / 54.0).
    • Billion-Scale Backbones (896×896896 \times 896 crop):

      • InternImage-H + UperNet (1.12B params, 3566G FLOPs): 59.9 mIoU (SS) / 60.3 mIoU (MS), surpassing SwinV2-G (3.00B params, 59.9 MS mIoU).
      • InternImage-H + Mask2Former (1.31B params, 4635G FLOPs): 62.5 mIoU (SS) / 62.9 mIoU (MS), surpassing the 1.90B-parameter BEiT-3 (62.8 mIoU).
  9. Knowl 9 — Stage-Wise Expansion of Effective Receptive Fields in InternImage

    empirical result

    Visualizing the spatial sampling offset distributions Δpgk\Delta p_{gk} across the four stages of InternImage reveals an adaptive, hierarchical expansion in effective receptive field (ERF):

    1. Local Spatial Aggregation (Stages 1 and 2): In early layers, sampling points remain tightly concentrated in local neighborhoods around the reference query pixel, preserving fine spatial details and local geometric structures.
    2. Global Adaptive Aggregation (Stages 3 and 4): In deeper layers, the learned sampling offsets adaptively distribute across broad image regions and long-range semantic contexts, yielding a global effective receptive field.
    3. Subspace Specialization: Within any individual layer, different aggregation groups g∈{1,…,G}g \in \{1, \dots, G\} direct their sampling offsets to distinct spatial regions for the same query pixel, capturing complementary subspace information.

    This behavior contrasts with standard Vision Transformers, where the multi-head self-attention mechanism exhibits a largely global receptive field across all layers including early stages.

  10. Knowl 10 — Inference Latency and Infrastructure Constraints of Deformable CNNs

    limitation

    Two primary limitations affect large-scale convolutional models based on deformable convolution:

    1. Inference Latency: The DCNv3 operator relies on dynamically generated sampling offsets that induce irregular, non-contiguous memory access patterns during bilinear interpolation. This increases memory bandwidth consumption and latency compared to standard dense convolutions or matrix-multiplication-optimized attention implementations, limiting deployment on platforms with stringent real-time constraints.
    2. Early Development State of Billion-Scale CNNs: Unlike Vision Transformers, which benefit from extensive pre-training recipes, established scaling laws, and hardware-optimized kernel ecosystems, billion-scale CNN foundation models are in an early stage of development and require further exploration regarding training stability, specialized kernel optimization, and scaling dynamics.

Coverage note — None was omitted; all core architectural definitions, scaling rules, empirical benchmarks on ImageNet/COCO/ADE20K, ablations, ERF findings, and limitations from the paper are represented.

References

  1. 1.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Adv. Neural Inform. Process. Syst., 30, 2017. 1, 2, 3, 4, 5
  2. 2.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Int. Conf. Comput. Vis., pages 10012–10022, 2021. 1, 2, 3, 5, 6, 7
  3. 3.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019. 1
  4. 4.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. 1
  5. 5.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020. 1
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Adv. Neural Inform. Process. Syst., 33:1877–1901, 2020. 1
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022. 1
  8. 8.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. 1
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Int. Conf. Learn. Represent., 2020. 1, 2, 3, 5, 8
  10. 10.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Int. Conf. Comput. Vis., pages 568–578, 2021. 1, 3, 5, 8
  11. 11.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415–424, 2022. 1, 2, 3, 5, 6, 7
  12. 12.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. IEEE Conf. Comput. Vis. Pattern Recog., pages 12124–12134, 2022. 1
  13. 13.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. In Int. Conf. Comput. Vis., pages 22–31, 2021. 1
  14. 14.Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al. Xcit: Cross-covariance image transformers. Adv. Neural Inform. Process. Syst., 34, 2021. 1
  15. 15.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. Adv. Neural Inform. Process. Syst., 34, 2021. 1
  16. 16.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. Adv. Neural Inform. Process. Syst., pages 12009–12019, 2022. 1, 2, 3, 6, 7
  17. 17.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022. 1, 3, 6, 7, 8
  18. 18.Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, Andre Susano Pinto, Daniel Keysers, and Neil Houlsby. Scaling vision with sparse mixture of experts. Adv. Neural Inform. Process. Syst., 34:8583–8595, 2021. 1, 2
  19. 19.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In IEEE Conf. Comput. Vis. Pattern Recog., pages 12104–12113, 2022. 1, 2, 3, 6
  20. 20.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. Adv. Neural Inform. Process. Syst., 34:3965–3977, 2021. 1, 2, 3, 6
  21. 21.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. arXiv preprint arXiv:2201.03545, 2022. 1, 2, 3, 5, 6, 7
  22. 22.Xiaohan Ding, Xiangyu Zhang, Jungong Han, and Guiguang Ding. Scaling up your kernels to 31x31: Revisiting large kernel design in cnns. In IEEE Conf. Comput. Vis. Pattern Recog., pages 11963–11975, 2022. 1, 2, 3, 4, 5, 6, 7
  23. 23.Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. In IEEE Conf. Comput. Vis. Pattern Recog., pages 10819–10829, 2022. 2
  24. 24.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 2, 3, 5
  25. 25.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arxiv. arXiv preprint arXiv:1606.08415, 2016. 2, 5
  26. 26.Yixuan Wei, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. Contrastive learning rivals masked image modeling in fine-tuning via feature distillation. arXiv preprint arXiv:2205.14141, 2022. 2, 6, 7
  27. 27.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In Int. Conf. Comput. Vis., pages 764–773, 2017. 2
  28. 28.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable convnets v2: More deformable, better results. In IEEE Conf. Comput. Vis. Pattern Recog., pages 9308–9316, 2019. 2, 3, 4
  29. 29.Shiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen, Qiao Xiao, Boqian Wu, Mykola Pechenizkiy, Decebal Mocanu, and Zhangyang Wang. More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity. arXiv preprint arXiv:2207.03620, 2022. 2, 7
  30. 30.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE Conf. Comput. Vis. Pattern Recog., pages 248–255, 2009. 2, 5, 6
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Eur. Conf. Comput. Vis., pages 740–755, 2014. 2, 6
  32. 32.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM, 60(6):84–90, 2017. 2, 4
  33. 33.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 2, 3
  34. 34.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1–9, 2015. 2
  35. 35.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., pages 770–778, 2016. 2, 3, 5
  36. 36.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1492–1500, 2017. 2
  37. 37.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning., pages 6105–6114. PMLR, 2019. 2, 5
  38. 38.Mingxing Tan and Quoc Le. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning., pages 10096–10106. PMLR, 2021. 2
  39. 39.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2
  40. 40.Yongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou, Ser-Nam Lim, and Jiwen Lu. Hornet: Efficient high-order spatial interactions with recursive gated convolutions. arXiv preprint arXiv:2207.14284, 2022. 3, 7
  41. 41.Qi Han, Zejia Fan, Qi Dai, Lei Sun, Ming-Ming Cheng, Jiaying Liu, and Jingdong Wang. On the connection between local attention and dynamic depth-wise convolution. In Int. Conf. Learn. Represent., 2021. 3
  42. 42.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020. 3
  43. 43.Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. Vision transformer with deformable attention. In IEEE Conf. Comput. Vis. Pattern Recog., pages 4794–4803, 2022. 3, 4
  44. 44.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In IEEE Conf. Comput. Vis. Pattern Recog., pages 12894–12904, 2021. 3
  45. 45.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. 3
  46. 46.Mingyu Ding, Bin Xiao, Noel Codella, Ping Luo, Jingdong Wang, and Lu Yuan. Davit: Dual attention vision transformers. arXiv preprint arXiv:2204.03645, 2022. 3
  47. 47.Shiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen, Qiao Xiao, Boqian Wu, Mykola Pechenizkiy, Decebal Mocanu, and Zhangyang Wang. More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity. arXiv preprint arXiv:2207.03620, 2022. 3
  48. 48.Xizhou Zhu, Dazhi Cheng, Zheng Zhang, Stephen Lin, and Jifeng Dai. An empirical study of spatial attention mechanisms in deep networks. In Int. Conf. Comput. Vis., pages 6688–6697, 2019. 3
  49. 49.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell., 40(4):834–848, 2017. 3
  50. 50.L-CCGP Florian and Schroff Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., volume 6, 2017. 3
  51. 51.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Eur. Conf. Comput. Vis., pages 801–818, 2018. 3
  52. 52.Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1(4):541–551, 1989. 3
  53. 53.Franc¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1251–1258, 2017. 4
  54. 54.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 4
  55. 55.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning., pages 10524–10533. PMLR, 2020. 5
  56. 56.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve J'egou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning., pages 10347–10357, 2021. 5
  57. 57.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 6, 7
  58. 58.Weijie Su, Xizhou Zhu, Chenxin Tao, Lewei Lu, Bin Li, Gao Huang, Yu Qiao, Xiaogang Wang, Jie Zhou, and Jifeng Dai. Towards all-in-one pre-training via maximizing multi-modal mutual information. In CVPR, 2023. 6
  59. 59.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. 6
  60. 60.Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016. 6
  61. 61.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pretraining to recognize long-tail visual concepts. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3558–3568, 2021. 6
  62. 62.Hugo Touvron, Matthieu Cord, and Herve J'egou. Deit iii: Revenge of the vit. arXiv preprint arXiv:2204.07118, 2022. 6
  63. 63.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In Eur. Conf. Comput. Vis., pages 491–507. Springer, 2020. 6
  64. 64.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In IEEE Conf. Comput. Vis. Pattern Recog., pages 10687–10698, 2020. 6
  65. 65.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. arXiv preprint arXiv:2205.08534, 2022. 7
  66. 66.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Int. Conf. Comput. Vis., pages 2961–2969, 2017. 6
  67. 67.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: high quality object detection and instance segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 43(5):1483–1498, 2019. 6
  68. 68.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. In Int. Conf. Comput. Vis., pages 3060–3069, 2021. 7
  69. 69.Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions. In IEEE Conf. Comput. Vis. Pattern Recog., pages 7373–7382, 2021. 7
  70. 70.Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv preprint arXiv:2203.03605, 2022. 6, 7
  71. 71.Jianwei Yang, Chunyuan Li, and Jianfeng Gao. Focal modulation networks. arXiv preprint arXiv:2203.11926, 2022. 7
  72. 72.Qiang Chen, Xiaokang Chen, Gang Zeng, and Jingdong Wang. Group detr: Fast training convergence with decoupled one-to-many label assignment. arXiv preprint arXiv:2207.13085, 2022. 7
  73. 73.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. arXiv preprint arXiv:2203.16527, 2022. 7
  74. 74.Tingting Liang, Xiaojie Chu, Yudong Liu, Yongtao Wang, Zhi Tang, Wei Chu, Jingdong Chen, and Haibin Ling. Cbnet: A composite backbone network architecture for object detection. IEEE Trans. Image Process., 2022. 6
  75. 75.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In Int. Conf. Comput. Vis., pages 8430–8439, 2019. 6
  76. 76.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. arXiv preprint arXiv:2112.01527, 2021. 7, 8
  77. 77.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Eur. Conf. Comput. Vis., pages 418–434, 2018. 7
  78. 78.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In IEEE Conf. Comput. Vis. Pattern Recog., pages 633–641, 2017. 7
  79. 79.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inform. Process. Syst., 34, 2021. 8

Citation

MLA
Wang, W., et al. “InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions”. arXiv, 2022, http://arxiv.org/abs/2211.05778v4.
APA
Wang, W., Dai, J., Chen, Z., Huang, Z., Li, Z., Zhu, X., Hu, X., Lu, T., Lu, L., Li, H., Wang, X., & Qiao, Y. (2022). InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions. arXiv. http://arxiv.org/abs/2211.05778v4
Chicago
Wang, W., J. Dai, Z. Chen, et al. 2022. “InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions”. arXiv. http://arxiv.org/abs/2211.05778v4.
Harvard
Wang, W. et al. (2022) “InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.05778v4.
Vancouver
1. Wang W, Dai J, Chen Z, et al (2022) InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions. arXiv

BibTeX

@article{wang2022internimage,
  title = {InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions},
  author = {Wang, Wenhai and Dai, Jifeng and Chen, Zhe and Huang, Zhenhang and Li, Zhiqi and Zhu, Xizhou and Hu, Xiaowei and Lu, Tong and Lu, Lewei and Li, Hongsheng and Wang, Xiaogang and Qiao, Yu},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.05778v4},
  eprint = {2211.05778}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE