InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
Wenhai WangJifeng DaiZhe ChenZhenhang HuangZhiqi LiXizhou ZhuXiaowei HuTong LuLewei LuHongsheng Li
Presents InternImage, a billion-parameter convolutional foundation model built on dynamic deformable convolutions that matches the scaling capabilities and downstream performance of state-of-the-art vision transformers across ImageNet, COCO, and ADE20K benchmarks.
Recent advances in computer vision have been dominated by vision transformers, which scale effectively to billions of parameters and vast datasets. In contrast, traditional convolutional neural networks have lagged behind at massive scales due to rigid architectures and static, localized operations. However, vision transformers often suffer from heavy computational and memory costs on high-resolution image tasks. The article addresses this challenge by exploring whether a modernized convolutional architecture can scale just as effectively as transformers while remaining computationally efficient.
The main objective of the article is to develop and evaluate a large-scale convolutional foundation model, named InternImage, demonstrating that convolutional networks can achieve state-of-the-art visual recognition performance when expanded to over one billion parameters and trained on hundreds of millions of images.
To accomplish this, the authors designed an improved core operator based on deformable convolution, which dynamically learns flexible sampling locations and adapts to input data using standard three-by-three filter windows. This operator incorporates shared weights across sampling points to cut memory usage, a multi-group structure to learn diverse feature patterns, and normalized modulation scales to ensure training stability. The authors then built a family of models ranging from 30 million to over one billion parameters, applying structured block-stacking and dimension-scaling rules. The models were evaluated across major benchmarks for image classification, object detection, and semantic segmentation using standard evaluation protocols and training data ranging from one million to 427 million images.
The findings show that InternImage consistently matches or outperforms leading vision transformers and prior convolutional models. On standard image classification, the base model achieved an 84.9% accuracy score, surpassing comparable convolutional architectures by at least 1.1 points, while the largest one-billion-parameter variant achieved 89.6%. On the challenging COCO object detection benchmark, the flagship model set a state-of-the-art record score of 65.4, exceeding the leading transformer baseline by 2.3 points while requiring 27% fewer parameters. In semantic segmentation on the ADE20K benchmark, the model attained a record 62.9 score, outperforming previous top models. Ablation analyses confirmed that sharing projection weights reduced GPU memory usage by up to 84.2% at the largest scale without sacrificing accuracy.
These results carry significant implications for computer vision research and real-world deployment. They prove that transformers are not the only viable path for large-scale vision foundation models. By retaining the efficient inductive biases of convolutions while gaining the adaptive properties of transformers, InternImage delivers higher task accuracy with lower computational resource demands. This improves performance in downstream tasks like dense visual perception while moderating infrastructure and training costs.
Organizations developing large-scale visual perception systems should consider deformable convolutional architectures alongside vision transformers. Further work is recommended to optimize the runtime latency of deformable operations on specialized deployment hardware to support ultra-fast inference settings. In addition, continued research into large-scale pre-training across larger multimodal datasets will help fully map the capabilities of convolutional foundation models.
The findings are supported by comprehensive benchmark testing across multiple model sizes and vision domains. Readers should note that large-scale convolutional models remain at an early developmental stage, and empirical results for the largest model relied on a composite training dataset rather than identical proprietary data used by some competing transformer baselines. Nonetheless, confidence is high that the architectural designs provide a robust, scalable alternative for advanced computer vision applications.
- Paper: Deformable Convolutional Networks, Jifeng Dai et al. (2017). It introduces the fundamental deformable convolution operator that InternImage redesigns and scales into a core building block for large-scale vision foundation models.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). It establishes the modern standard for pure convolutional architectures competing with Vision Transformers, directly motivating InternImage's goal of scaling modern ConvNets to billion-parameter regimes.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). It provides the hierarchical window-based vision transformer baseline that InternImage directly benchmarks against and seeks to outperform in dense downstream perception tasks.
- Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It explores the architectural techniques and scaling behaviors for multi-billion parameter vision models that InternImage aims to rival with a convolutional alternative.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It introduces the Vision Transformer paradigm that triggered the shift toward large-scale vision foundation models, defining the computational and architectural baseline InternImage challenges.
- Paper: MetaFormer is Actually What You Need for Vision, Weihao Yu et al. (2021). It analyzes the macro-architectural design of modern vision backbones, framing how general network layouts function independently of specific token-mixing operators.
- Paper: EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, Mingxing Tan et al. (2019). It introduces principled compound scaling principles across depth, width, and resolution for convolutional neural networks.
- Paper: Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, Zhe Chen et al. (2024). It extends large-scale visual foundation modeling into the vision-language domain by scaling vision encoders and aligning them with large language models.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It continues the evolution of open-source multimodal foundation models by scaling visual encoders and multi-modal alignment across diverse benchmarks.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It advances multimodal vision foundation models by investigating unified native pre-training and test-time scaling strategies.
- Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). It explores self-supervised masked image modeling specifically co-designed for modern convolutional architectures to scale pure ConvNets effectively.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). It develops an alternative efficient perception backbone by combining convolutional gating with biomimetic foveal attention to eliminate high-resolution vision bottlenecks.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). It investigates bidirectional state space models as an alternative non-transformer backbone to achieve efficient visual representation learning on high-resolution images.
