VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
Hangbo BaoWenhui WangLi DongQiang LiuOwais Khan MohammedKriti AggarwalSubhojit SomSonghao PiaoFuru Wei
Introduces a modular vision-language pre-training framework that uses modality-specific feed-forward experts and shared self-attention to unify dual-encoder retrieval speed with fusion-encoder classification accuracy.
Modern artificial intelligence applications frequently require processing both images and text simultaneously. However, existing vision-language models typically face a structural trade-off: dual-encoder systems process text and images separately to provide fast retrieval speeds at the cost of reasoning accuracy, while fusion-encoder systems combine modalities deeply for high-accuracy classification at the cost of slow, computationally expensive retrieval. The article introduces a unified framework named the Vision-Language Pre-trained Model (VLMO) to eliminate this trade-off by dynamically serving as either a dual encoder or a fusion encoder within a single architecture.
To achieve this, the article develops a Multiway Transformer architecture that incorporates specialized modality experts alongside shared self-attention layers. Rather than relying on separate networks or computationally heavy external object detectors, the framework routes visual, linguistic, or joint vision-language representations to dedicated processing blocks based on the input type. The training strategy uses a staged pipeline: the model first learns general image features from unlabeled images, then trains text experts on text-only corpora while freezing image parameters, and finally undergoes joint multimodal pre-training using image-text contrastive learning, masked language modeling, and image-text matching with global hard negative mining. Evaluation was conducted across standard benchmarks with base setups (4 million images across 10 million pairs) and scaled up to 1 billion web image-text pairs.
Key findings show that VLMO delivers leading performance across both classification and retrieval tasks. On complex reasoning benchmarks, the model outperforms previous approaches, achieving up to an 82.88 score on Visual Question Answering (VQA) and 89.54% accuracy on Natural Language for Visual Reasoning (NLVR2) when scaled to larger datasets. On text-and-image retrieval benchmarks such as COCO and Flickr30K, the model matches or exceeds the accuracy of slower fusion-based systems while providing fast linear-time retrieval speeds. Ablation studies confirm that stagewise pre-training using unpaired data substantially improves downstream accuracy (raising NLVR2 dev scores from 80.33% to 82.09%), and global hard negative mining further boosts performance over local GPU sampling.
These results demonstrate that organizations can deploy a single unified architecture across diverse vision-language workflows instead of maintaining separate models for search and deep classification. Eliminating the need for complex object detectors and enabling fast dual-encoder retrieval significantly lowers computational overhead, latency, and operational deployment costs in production environments.
Based on these findings, technical teams should consider adopting unified modular architectures and stagewise pre-training strategies when building multimodal systems, especially when leveraging large amounts of readily available unpaired images and text. Future development highlighted in the article includes scaling the model size further, expanding the framework to generative tasks such as automated image captioning, and extending the multiway expert design to additional data types such as video, audio, and structured knowledge.
Confidence in the reported experimental improvements is high across standard academic benchmarks. However, leaders should note that the highest performance gains rely on large-scale training using 1 billion noisy web pairs, which demands substantial computing resources and may introduce domain-specific noise or dataset limitations in specialized real-world settings.
No sufficiently relevant recommendations were found.
- Paper: Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks, Wenhui Wang et al. (2023). BEIT-3 carries VLMo’s Multiway Transformer design forward, using modality-specific experts in a larger unified vision-language model.
