Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks
Wenhui WangHangbo BaoLi DongJohan BjorckZhiliang PengQiang LiuKriti AggarwalOwais Khan MohammedSaksham SinghalSubhojit Som
Introduces BEiT-3, a unified multimodal foundation model that treats images as a foreign language and pretrains a Multiway Transformer via masked data modeling to achieve state-of-the-art transfer performance across major vision and vision-language benchmarks.
Artificial intelligence systems typically rely on specialized architectures and complex training objectives to handle distinct tasks in computer vision and natural language processing. Building and maintaining separate models for vision tasks, text understanding, and multimodal tasks—such as matching images with text—creates high engineering overhead, requires massive private datasets, and increases compute expenses. The article addresses this fragmentation by investigating whether a single, general-purpose foundation model can unify vision and vision-language processing under a simpler framework.
The main objective of the article is to demonstrate that treating images as a foreign language allows a single 1.9-billion-parameter foundation model, named BEIT-3, to achieve state-of-the-art performance across both vision-only and vision-language benchmarks while using only publicly available training data and a single pretraining objective.
To achieve this, the authors pretrained BEIT-3 using a shared Multiway Transformer architecture across 14 million images, 160 gigabytes of text, and 21 million image-text pairs. The network incorporates shared attention mechanisms alongside specialized expert modules that route vision, text, and multimodal tokens appropriately. Rather than juggling multiple complex pretraining tasks or requiring massive memory batch sizes like conventional contrastive systems, the model relies entirely on a single mask-and-predict generative task across masked text tokens, masked image patches, and parallel image-text pairs. The model was then fine-tuned and evaluated across a broad battery of standard benchmarks.
The key findings show that BEIT-3 sets new state-of-the-art performance records across all evaluated benchmarks. On visual reasoning (NLVR2), the model scored 92.58%, outperforming the previous best model by roughly 5.6 percentage points and surpassing 90% accuracy for the first time. On visual question answering (VQAv2), it achieved 84.03% accuracy, surpassing prior systems by over 1.7 points. In cross-modal retrieval, it delivered top-1 accuracy gains of up to 3.5 to 4.0 points on standard image-to-text and text-to-image benchmarks. Additionally, on core vision tasks, the model achieved leading results on object detection and instance segmentation on COCO, improved semantic segmentation on ADE20K to 62.8 mean intersection over union, and reached 89.6% top-1 accuracy on ImageNet-1K using only public resources.
These findings indicate that generative mask-and-predict pretraining is sufficiently powerful to learn deep cross-modal alignments without the need for complex, multi-loss training pipelines. By operating with significantly smaller training batch sizes—around 6,000 samples compared to 24,000 to 65,000 required by contrastive frameworks—this approach lowers the computational memory barrier and hardware risk associated with training large multimodal models. Furthermore, achieving top-tier results using public datasets demonstrates that competitive performance does not depend strictly on proprietary, in-house data.
For organizations developing multimodal AI capabilities, adopting a unified generative pretraining pipeline with modular expert routing represents an effective path to reduce architectural complexity and infrastructure costs. Decision-makers should consider piloting this mask-then-predict approach when consolidating vision and multimodal systems. The authors recommend expanding future research toward multilingual training, incorporating additional data modalities such as audio, and combining the architecture with in-context learning frameworks.
Confidence in the reported benchmark gains is high given the extensive comparative evaluations across diverse public tasks. However, practitioners should note that the full model still requires substantial scale at 1.9 billion parameters, and downstream retrieval performance benefits from an additional intermediate fine-tuning stage. Deployments should account for the computational resources required during inference and task-specific fine-tuning.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). BEiT introduced masked image modeling using discrete visual tokens, establishing the core generative mask-and-predict foundation that BEIT-3 extends to multimodal vision-language pretraining.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This foundational work introduced the Vision Transformer (ViT) by treating 16x16 image patches like words, providing the conceptual backbone for treating images as a foreign language.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Masked Autoencoders demonstrated the scalability and efficiency of direct masked prediction on visual patches, which heavily influenced unified generative pretraining paradigms.
- Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). SimMIM showed that simple masked image modeling with lightweight prediction heads can learn strong visual representations, inspiring simpler mask-and-predict pretraining objectives.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Unified-IO established an early unified token-based sequence-to-sequence framework for joint vision and language tasks, setting the stage for BEIT-3's unified architecture.
- Paper: An Empirical Study of Training End-to-End Vision-and-Language Transformers, Zi-Yi Dou et al. (2022). This empirical study analyzed end-to-end vision-and-language Transformers and identified the strengths of masked modeling and modular cross-attention mechanisms.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). UNITER proved the effectiveness of unified Transformer representation learning through cross-modal masked modeling and alignment objectives.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). ALBEF established core multi-task vision-language pretraining benchmarks and fusion concepts that BEIT-3 simplifies using a single generative objective.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). VisualBERT pioneered using a shared Transformer to jointly model visual and textual tokens, providing an early precedent for unified multi-modal representations.
- Paper: Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks, Xiujun Li et al. (2020). Oscar demonstrated how aligning visual elements into a shared linguistic space simplifies cross-modal understanding in Transformer backbones.
- Paper: Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, Zhe Chen et al. (2024). InternVL builds on large-scale unified vision-language foundation architectures to scale visual encoders up to 6 billion parameters and align them with large language models.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). InternVL3 extends native unified multimodal pretraining recipes to scale open-source multimodal foundation models across massive language and vision datasets.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). InternVL 2.5 expands unified multimodal architectures by exploring model, data, and test-time scaling strategies across dense visual and reasoning benchmarks.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision extends unified visual representation modeling across single-image, multi-image, and video understanding tasks.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Qwen3-VL applies and scales unified vision-language architectures into modern mixture-of-experts multimodal foundation models.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). Emu2 builds upon generative unified multimodal pretraining to enable broad multimodal in-context learning and content generation.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o advances unified Transformer representations by coupling autoregressive text modeling with discrete visual diffusion for integrated multimodal understanding and generation.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This survey provides a comprehensive synthesis of vision-language foundation models, contextualizing the role of unified architectures like BEIT-3 across diverse downstream tasks.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). Prismatic VLMs investigates the architectural design space and optimization recipes of visually conditioned models, refining the training paradigms evaluated in BEIT-3.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). PreFLMR applies fine-grained multimodal representations to scale up late-interaction multi-modal retrieval systems across diverse visual-textual formats.
