Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks

Wenhui WangHangbo BaoLi DongJohan BjorckZhiliang PengQiang LiuKriti AggarwalOwais Khan MohammedSaksham SinghalSubhojit Som

article2023CVPR478 citations

Introduces BEiT-3, a unified multimodal foundation model that treats images as a foreign language and pretrains a Multiway Transformer via masked data modeling to achieve state-of-the-art transfer performance across major vision and vision-language benchmarks.

Listen

Artificial intelligence systems typically rely on specialized architectures and complex training objectives to handle distinct tasks in computer vision and natural language processing. Building and maintaining separate models for vision tasks, text understanding, and multimodal tasks—such as matching images with text—creates high engineering overhead, requires massive private datasets, and increases compute expenses. The article addresses this fragmentation by investigating whether a single, general-purpose foundation model can unify vision and vision-language processing under a simpler framework.

The main objective of the article is to demonstrate that treating images as a foreign language allows a single 1.9-billion-parameter foundation model, named BEIT-3, to achieve state-of-the-art performance across both vision-only and vision-language benchmarks while using only publicly available training data and a single pretraining objective.

To achieve this, the authors pretrained BEIT-3 using a shared Multiway Transformer architecture across 14 million images, 160 gigabytes of text, and 21 million image-text pairs. The network incorporates shared attention mechanisms alongside specialized expert modules that route vision, text, and multimodal tokens appropriately. Rather than juggling multiple complex pretraining tasks or requiring massive memory batch sizes like conventional contrastive systems, the model relies entirely on a single mask-and-predict generative task across masked text tokens, masked image patches, and parallel image-text pairs. The model was then fine-tuned and evaluated across a broad battery of standard benchmarks.

The key findings show that BEIT-3 sets new state-of-the-art performance records across all evaluated benchmarks. On visual reasoning (NLVR2), the model scored 92.58%, outperforming the previous best model by roughly 5.6 percentage points and surpassing 90% accuracy for the first time. On visual question answering (VQAv2), it achieved 84.03% accuracy, surpassing prior systems by over 1.7 points. In cross-modal retrieval, it delivered top-1 accuracy gains of up to 3.5 to 4.0 points on standard image-to-text and text-to-image benchmarks. Additionally, on core vision tasks, the model achieved leading results on object detection and instance segmentation on COCO, improved semantic segmentation on ADE20K to 62.8 mean intersection over union, and reached 89.6% top-1 accuracy on ImageNet-1K using only public resources.

These findings indicate that generative mask-and-predict pretraining is sufficiently powerful to learn deep cross-modal alignments without the need for complex, multi-loss training pipelines. By operating with significantly smaller training batch sizes—around 6,000 samples compared to 24,000 to 65,000 required by contrastive frameworks—this approach lowers the computational memory barrier and hardware risk associated with training large multimodal models. Furthermore, achieving top-tier results using public datasets demonstrates that competitive performance does not depend strictly on proprietary, in-house data.

For organizations developing multimodal AI capabilities, adopting a unified generative pretraining pipeline with modular expert routing represents an effective path to reduce architectural complexity and infrastructure costs. Decision-makers should consider piloting this mask-then-predict approach when consolidating vision and multimodal systems. The authors recommend expanding future research toward multilingual training, incorporating additional data modalities such as audio, and combining the architecture with in-context learning frameworks.

Confidence in the reported benchmark gains is high given the extensive comparative evaluations across diverse public tasks. However, practitioners should note that the full model still requires substantial scale at 1.9 billion parameters, and downstream retrieval performance benefits from an additional intermediate fine-tuning stage. Deployments should account for the computational resources required during inference and task-specific fine-tuning.

arXiv: 2208.10442
Cover for Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks

Abstract

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).

Table of Contents

  • 1 Introduction: The Big Convergence
  • 2 BEIT-3: A General-Purpose Multimodal Foundation Model
  • 2.1 Backbone Network: Multiway Transformers
  • 2.2 Pretraining Task: Masked Data Modeling
  • 2.3 Scaling Up: BEIT-3 Pretraining
  • 3 Experiments on Vision and Vision-Language Tasks
  • 3.1 Vision-Language Downstream Tasks
  • 3.2 Vision Downstream Tasks
  • 4 Conclusion
  • References
  • A Effects of Intermediate Finetuning for Retrieval
  • B Hyperparameters Used for Pretraining
  • C Hyperparameters Used for Finetuning

Knowls

  1. Knowl 1 — Multiway Transformer Architecture for General-Purpose Multimodal Modeling

    model/method

    The BEiT-3 backbone network utilizes a Multiway Transformer to encode monomodal and multimodal inputs within a single unified model. Each Multiway Transformer layer consists of a shared multi-head self-attention module coupled with a pool of modality-specific feed-forward networks (FFN experts):

    • Vision Experts (V-FFNV\text{-FFN}): Feed-forward networks dedicated to processing visual patch tokens across all 40 layers.
    • Language Experts (L-FFNL\text{-FFN}): Feed-forward networks dedicated to processing textual subword tokens across all 40 layers.
    • Vision-Language Experts (VL-FFNVL\text{-FFN}): Feed-forward networks employed in the top three layers (layers 38 to 40) designed specifically for deep cross-modal fusion when handling paired vision-language tokens.

    Input tokens are routed dynamically to their respective modality experts based on token modality. The shared self-attention module is applied across all tokens, enabling cross-modality alignment and interaction while allowing the modality experts to specialize in capturing domain-specific features.

  2. Knowl 2 — Unified Masked Data Modeling Pretraining Objective in BEiT-3

    model/method

    BEiT-3 is pretrained exclusively using a single self-supervised pretraining task: masked data modeling (MDM / mask-then-predict). Images are conceptualized as a foreign language ("Imglish"), texts as English, and paired image-text instances as "parallel sentences". The model randomly masks a fraction of input tokens and is trained to recover the original discrete target tokens without auxiliary pretraining objectives such as image-text matching or contrastive learning during initial pretraining.

    • Visual Target Tokenization: Continuous image patch targets are represented as discrete visual tokens extracted via the vector-quantized visual tokenizer from BEiT v2.
    • Text Tokenization: Text sequences are tokenized using a SentencePiece tokenizer with a 64,000 vocabulary size.
    • Masking Ratios and Strategies:
      • For monomodal text data, 15% of tokens are masked at random.
      • For text within image-text pairs, 50% of tokens are masked at random.
      • For images, 40% of the 14×1414 \times 14 patches are masked using a block-wise masking strategy.
  3. Knowl 3 — BEiT-3 Model Configuration and Parameter Allocation

    data/table

    BEiT-3 adopts a giant-scale architecture layout based on ViT-giant. The model comprises 40 Multiway Transformer layers with a hidden dimension d=1408d = 1408, an intermediate FFN dimension dffn=6144d_{\text{ffn}} = 6144, and 16 attention heads. Modality-specific routing splits the parameters across shared attention and expert FFNs.

    Model #Layers Hidden Size MLP Size V-FFN L-FFN VL-FFN Shared Attention Total
    BEIT-3 40 1408 6144 692M 692M 52M 317M 1.9B

    Although the total parameter count is 1.9B, only the shared attention modules and vision-related parameters (totaling approximately 1.0B parameters, identical in active compute to a standard ViT-giant) are activated during vision-only downstream deployment.

  4. Knowl 4 — BEiT-3 Pretraining Data and Hyperparameter Setup

    experimental setup

    BEiT-3 is pretrained for 1,000,000 steps using exclusively academically and publicly accessible datasets across three streams:

    • Multimodal Data: ~15M images comprising 21M image-text pairs collected from Conceptual 12M (CC12M), Conceptual Captions (CC3M), SBU Captions, COCO, and Visual Genome (VG).
    • Monomodal Image Data: 14M images from ImageNet-21K.
    • Monomodal Text Data: 160GB of text from English Wikipedia, BookCorpus, OpenWebText, CC-News, and Stories.

    Pretraining uses a total batch size of 6,144 (2,048 monomodal images, 2,048 monomodal texts, and 2,048 image-text pairs) at an image resolution of 224×224224 \times 224 (14×1414 \times 14 patch size). Optimization is performed using AdamW (β1=0.9,β2=0.98,ϵ=10−6\beta_1 = 0.9, \beta_2 = 0.98, \epsilon = 10^{-6}), a weight decay of 0.05, a peak learning rate of 10−310^{-3} scheduled via cosine decay with a 10,000-step linear warmup, and stochastic depth with rate 0.1.

    Parameter initialization initializes weights uniformly in [−0.02,0.02][-0.02, 0.02], and the output projection matrix of each self-attention and FFN sublayer at layer ll is scaled by 12l\frac{1}{\sqrt{2l}}.

  5. Knowl 5 — Finetuning Paradigms for Vision-Language Downstream Tasks

    model/method

    BEiT-3 adapts its Multiway Transformer architecture to various vision-language tasks through task-specific head configurations and attention masks:

    • Visual Question Answering (VQAv2): Finetuned as a fusion encoder. Concatenated image patch embeddings and question token embeddings pass through the Multiway Transformer (using VL-FFNVL\text{-FFN} at the top layers). The pooled output is fed into a classification layer predicting one of the 3,129 most frequent candidate answers.
    • Visual Reasoning (NLVR2): Evaluated on image-pair triplet inputs by forming two image-text pairs (I1,T)(I_1, T) and (I2,T)(I_2, T). Both pairs are independently encoded with the fusion encoder; their pooled outputs are concatenated and passed to a binary classifier.
    • Image Captioning (COCO): Configured as a sequence-to-sequence conditional autoregressive generator via masked finetuning. A customized self-attention mask restricts image tokens to attend to each other bidirectionally, while caption tokens attend bidirectionally to image tokens and causally/unidirectionally to leftward caption context and themselves. Special token [SEP] is masked during training to supervise sequence termination. The model is optimized purely with token-level cross-entropy loss without reinforcement-learning-based CIDEr optimization.
    • Image-Text Retrieval: Configured as a dual encoder where images and texts are processed independently through V-FFNV\text{-FFN} and L-FFNL\text{-FFN} modules with shared self-attention. Predictions are scored via cosine similarity between normalized image and text representation embeddings.
  6. Knowl 6 — Finetuning Paradigms for Vision Downstream Tasks

    model/method

    BEiT-3 transfers to standard computer vision benchmarks using modality-specific adaptation frameworks:

    • Image Classification (ImageNet-1K): Instead of appending a classification linear head, the task is framed as image-to-text retrieval. Class category names are tokenized as text labels. The model is finetuned as a dual encoder to match an input image embedding to candidate class label embeddings via cosine similarity. Training employs intermediate finetuning on ImageNet-21K prior to ImageNet-1K evaluation.
    • Object Detection and Instance Segmentation (COCO): The vision backbone is paired with the ViTDet framework (incorporating simple feature pyramids and window attention) and a Cascade Mask R-CNN head. Pre-finetuning is conducted on Objects365 prior to COCO 2017 finetuning, and Soft-NMS is utilized at inference.
    • Semantic Segmentation (ADE20K): The vision encoder is integrated with the ViT-Adapter dense prediction task adapter and Mask2Former segmentation head, taking crop sizes of 896×896896 \times 896.
  7. Knowl 7 — Vision-Language Benchmark Evaluation Results

    data/table

    BEiT-3 sets state-of-the-art results across visual question answering, visual reasoning, and image captioning benchmarks using public pretraining data without reinforcement learning optimization.

    Model VQAv2 NLVR2 COCO Captioning
    test-dev test-std dev test-P B@4 M C S
    Oscar 73.61 73.82 79.12 80.37 37.4 30.7 127.8 23.5
    VinVL 76.52 76.60 82.67 83.98 38.5 30.4 130.8 23.4
    ALBEF 75.84 76.04 82.55 83.14 - - - -
    BLIP 78.25 78.32 82.15 82.24 40.4 - 136.7 -
    SimVLM 80.03 80.34 84.53 85.15 40.6 33.7 143.3 25.4
    Florence 80.16 80.36 - - - - - -
    OFA 82.00 82.00 - - 43.9 31.8 145.3 24.8
    Flamingo 82.00 82.10 - - - - 138.1 -
    CoCa 82.30 82.30 86.10 87.00 40.9 33.9 143.6 24.7
    BEiT-3 84.19 84.03 91.51 92.58 44.1 32.4 147.6 25.4

    Metrics include VQA accuracy on VQAv2 test-dev and test-standard; classification accuracy on NLVR2 development (dev) and public test (test-P) sets; and BLEU@4 (B@4), METEOR (M), CIDEr (C), and SPICE (S) on the COCO Karpathy test split (reported without CIDEr optimization).

  8. Knowl 8 — Cross-Modal Image-Text Retrieval Performance on COCO and Flickr30K

    data/table

    Evaluated as a dual encoder, BEiT-3 achieves state-of-the-art retrieval performance across both finetuned and zero-shot settings on the MSCOCO (5K test set) and Flickr30K (1K test set) Karpathy splits.

    Setting / Model MSCOCO Image →\rightarrow Text MSCOCO Text →\rightarrow Image Flickr30K Image →\rightarrow Text Flickr30K Text →\rightarrow Image
    R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10
    Finetuned:
    ALIGN 77.0 93.5 96.9 59.9 83.3 89.8 95.3 99.8 100.0 84.9 97.4 98.6
    FILIP 78.9 94.4 97.4 61.2 84.3 90.6 96.6 100.0 100.0 87.1 97.7 99.1
    Florence 81.8 95.2 - 63.2 85.7 - 97.2 99.9 - 87.9 98.1 -
    BEiT-3 84.8 96.5 98.3 67.2 87.7 92.8 98.0 100.0 100.0 90.3 98.7 99.5
    Zero-shot (Flickr30K):
    CLIP - - - - - - 88.0 98.7 99.4 68.7 90.6 95.2
    ALIGN - - - - - - 88.6 98.7 99.7 75.7 93.8 96.8
    Florence - - - - - - 90.9 99.1 - 76.7 93.6 -
    CoCa - - - - - - 92.5 99.5 99.9 80.4 95.7 97.7
    BEiT-3 - - - - - - 94.9 99.9 100.0 81.5 95.6 97.8

    Dual-encoder evaluation allows separate encoding of modalities for faster inference, outperforming prior fusion-encoder reranking methods while maintaining efficiency.

  9. Knowl 9 — Vision Benchmark Performance on ADE20K, COCO, and ImageNet-1K

    data/table

    When evaluated on dense prediction and vision recognition tasks, BEiT-3 outperforms specialized computer vision models and large-scale vision models trained on private data.

    Task Dataset Metric Previous SOTA BEiT-3 Margin
    Semantic Segmentation ADE20K mIoU 61.4 (FD-SwinV2-G) 62.8 +1.4
    Object Detection COCO test-dev APbox\text{AP}^{\text{box}} 63.3 (DINO) 63.7 +0.4
    Instance Segmentation COCO test-dev APmask\text{AP}^{\text{mask}} 54.7 (Mask DINO) 54.8 +0.1
    Image Classification ImageNet-1K Top-1 Acc 89.0 (FD-CLIP)†^\dagger 89.6 +0.6

    †^\daggerImageNet comparison restricts external supervision strictly to publicly accessible data (ImageNet-21K). BEiT-3 achieves 62.8 mIoU on ADE20K under multi-scale evaluation with Mask2Former, surpassing the 3B-parameter FD-SwinV2-G model.

  10. Knowl 10 — Impact of Intermediate Contrastive Finetuning on Cross-Modal Retrieval

    empirical result

    While BEiT-3 is pretrained exclusively with masked data modeling without contrastive objectives, direct finetuning on target datasets achieves strong retrieval accuracy. Performing an additional intermediate finetuning stage on pretraining image-text pairs using an image-text contrastive objective for 5 epochs (batch size 16k, peak learning rate 3×10−53 \times 10^{-5}) provides consistent performance improvements:

    Model Setup MSCOCO Image →\rightarrow Text MSCOCO Text →\rightarrow Image Flickr30K Image →\rightarrow Text Flickr30K Text →\rightarrow Image
    R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10
    Direct Finetuning 82.7 96.0 98.2 65.1 86.6 92.3 97.5 99.9 100.0 89.1 98.6 99.3
    + Intermediate Finetuning 84.8 96.5 98.3 67.2 87.7 92.8 98.0 100.0 100.0 90.3 98.7 99.5

    These results demonstrate that masked data modeling alone encodes cross-modal alignment, while intermediate contrastive alignment provides complementary improvements.

Coverage note — None was omitted; all primary contributions, model architectures, pretraining objectives, downstream adaptation formulations, and core empirical results are fully captured.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. CoRR, abs/2204.14198, 2022.
  2. 2.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 6077–6086. Computer Vision Foundation / IEEE Computer Society, 2018.
  3. 3.Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. BEiT: BERT pre-training of image transformers. In International Conference on Learning Representations, 2022.
  4. 4.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, and Hsiao-Wuen Hon. UniLMv2: Pseudo-masked language models for unified language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 642–652. PMLR, 2020.
  5. 5.Hangbo Bao, Li Dong, Wenhui Wang, Nan Yang, and Furu Wei. s2s-ft: Fine-tuning pre-trained transformer encoders for sequence-to-sequence learning. CoRR, abs/2110.13640, 2021.
  6. 6.Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S. Davis. Soft-nms - improving object detection with one line of code. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 5562–5570. IEEE Computer Society, 2017.
  7. 7.Hangbo Bao, Wenhui Wang, Li Dong, and Furu Wei. VL-BEiT: Generative vision-language pretraining. ArXiv, abs/2206.01127, 2022.
  8. 8.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. CoRR, abs/2205.08534, 2022.
  9. 9.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: universal image-text representation learning. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXX, volume 12375 of Lecture Notes in Computer Science, pages 104–120. Springer, 2020.
  10. 10.Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. CoRR, abs/2112.01527, 2021.
  11. 11.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 3558–3568. Computer Vision Foundation / IEEE, 2021.
  12. 12.Zhaowei Cai and Nuno Vasconcelos. Cascade R-CNN: high quality object detection and instance segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 43(5):1483–1498, 2021.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. preprint arXiv:2010.11929, 2020.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics, 2019.
  15. 15.Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 7373–7382. Computer Vision Foundation / IEEE, 2021.
  16. 16.Zihang Dai, Hanxiao Liu, Quoc V. Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors, Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 3965–3977, 2021.
  17. 17.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language understanding and generation. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13042–13054, 2019.
  18. 18.Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. Large-scale adversarial training for vision-and-language representation learning. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  19. 19.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 6325–6334. IEEE Computer Society, 2017.
  20. 20.Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei. Language models are general-purpose interfaces. ArXiv, abs/2206.06336, 2022.
  21. 21.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. Deep networks with stochastic depth. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part IV, volume 9908 of Lecture Notes in Computer Science, pages 646–661. Springer, 2016.
  22. 22.Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, and Humphrey Shi. Semask: Semantically masking transformer backbones for effective semantic segmentation. arXiv, 2021.
  23. 23.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 4904–4916. PMLR, 2021.
  24. 24.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 3128–3137. IEEE Computer Society, 2015.
  25. 25.Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium, November 2018. Association for Computational Linguistics.
  26. 26.Wonjae Kim, Bokyung Son, and Ildoo Kim. ViLT: Vision-and-language transformer without convolution or region supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 5583–5594. PMLR, 2021.
  27. 27.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis., 123(1):32–73, 2017.
  28. 28.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019.
  29. 29.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. Swin transformer V2: scaling up capacity and resolution. CoRR, abs/2111.09883, 2021.
  30. 30.Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 12888–12900. PMLR, 2022.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In David J. Fleet, Tomás Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V, volume 8693 of Lecture Notes in Computer Science, pages 740–755. Springer, 2014.
  32. 32.Yanghao Li, Hanzi Mao, Ross B. Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. CoRR, abs/2203.16527, 2022.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019.
  34. 34.Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven C. H. Hoi. Align before fuse: Vision and language representation learning with momentum distillation. CoRR, abs/2107.07651, 2021.
  35. 35.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4804–4814, 2022.
  36. 36.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXX, volume 12375 of Lecture Notes in Computer Science, pages 121–137. Springer, 2020.
  37. 37.Feng Li, Hao Zhang, Huaizhe Xu, Shilong Liu, Lei Zhang, Lionel M. Ni, and Heung-Yeung Shum. Mask DINO: towards A unified transformer-based framework for object detection and segmentation. CoRR, abs/2206.02777, 2022.
  38. 38.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jianfeng Gao. Grounded language-image pre-training. CoRR, abs/2112.03857, 2021.
  39. 39.Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. Im2text: Describing images using 1 million captioned photographs. In John Shawe-Taylor, Richard S. Zemel, Peter L. Bartlett, Fernando C. N. Pereira, and Kilian Q. Weinberger, editors, Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12-14 December 2011, Granada, Spain, pages 1143–1151, 2011.
  40. 40.Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. Beit v2: Masked image modeling with vector-quantized visual tokenizers. CoRR, abs/2208.06366, 2022.
  41. 41.Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2641–2649. IEEE Computer Society, 2015.
  42. 42.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. IJCV, 2015.
  43. 43.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 2021.
  44. 44.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018.
  45. 45.Yongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou, Ser Nam Lim, and Jiwen Lu. HorNet: Efficient high-order spatial interactions with recursive gated convolutions. ArXiv, abs/2207.14284, 2022.
  46. 46.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 2556–2565. Association for Computational Linguistics, 2018.
  47. 47.Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. FLAVA: A foundational language and vision alignment model. CoRR, abs/2112.04482, 2021.
  48. 48.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 8429–8438. IEEE, 2019.
  49. 49.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 6418–6428. Association for Computational Linguistics, 2019.
  50. 50.Trieu H. Trinh and Quoc V. Le. A simple method for commonsense reasoning. ArXiv, abs/1806.02847, 2018.
  51. 51.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. CoRR, abs/2204.01697, 2022.
  52. 52.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017.
  53. 53.Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei. VLMo: Unified vision-language pre-training with mixture-of-modality-experts. CoRR, abs/2111.02358, 2021.
  54. 54.Yixuan Wei, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. Contrastive learning rivals masked image modeling in fine-tuning via feature distillation. CoRR, abs/2205.14141, 2022.
  55. 55.Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 23965–23998. PMLR, 2022.
  56. 56.Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 3733–3742. Computer Vision Foundation / IEEE Computer Society, 2018.
  57. 57.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. CoRR, abs/2202.03052, 2022.
  58. 58.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. SimVLM: Simple visual language model pretraining with weak supervision. CoRR, abs/2108.10904, 2021.
  59. 59.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 3040–3049. IEEE, 2021.
  60. 60.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang. Florence: A new foundation model for computer vision. CoRR, abs/2111.11432, 2021.
  61. 61.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. FILIP: fine-grained interactive language-image pre-training. CoRR, abs/2111.07783, 2021.
  62. 62.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. CoRR, abs/2205.01917, 2022.
  63. 63.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. arXiv preprint arXiv:2106.04560, 2021.
  64. 64.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pages 19–27, 2015.
  65. 65.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. VinVL: Revisiting visual representations in vision-language models. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 5579–5588. Computer Vision Foundation / IEEE, 2021.
  66. 66.Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, and Heung-Yeung Shum. DINO: DETR with improved denoising anchor boxes for end-to-end object detection. CoRR, abs/2203.03605, 2022.
  67. 67.Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Harold Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, and Jianfeng Gao. Glipv2: Unifying localization and vision-language understanding. CoRR, abs/2206.05836, 2022.
  68. 68.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ADE20K dataset. Int. J. Comput. Vis., 127(3):302–321, 2019.

Citation

MLA
Wang, W., et al. “Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks”. arXiv, 2022, http://arxiv.org/abs/2208.10442v2.
APA
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., & Wei, F. (2022). Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks. arXiv. http://arxiv.org/abs/2208.10442v2
Chicago
Wang, W., H. Bao, L. Dong, et al. 2022. “Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks”. arXiv. http://arxiv.org/abs/2208.10442v2.
Harvard
Wang, W. et al. (2022) “Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2208.10442v2.
Vancouver
1. Wang W, Bao H, Dong L, et al (2022) Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks. arXiv

BibTeX

@article{wang2022image,
  title = {Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks},
  author = {Wang, Wenhui and Bao, Hangbo and Dong, Li and Bjorck, Johan and Peng, Zhiliang and Liu, Qiang and Aggarwal, Kriti and Mohammed, Owais Khan and Singhal, Saksham and Som, Subhojit and Wei, Furu},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2208.10442v2},
  eprint = {2208.10442}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE