VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Hangbo BaoWenhui WangLi DongQiang LiuOwais Khan MohammedKriti AggarwalSubhojit SomSonghao PiaoFuru Wei

article2022NeurIPS790 citations

Introduces a modular vision-language pre-training framework that uses modality-specific feed-forward experts and shared self-attention to unify dual-encoder retrieval speed with fusion-encoder classification accuracy.

Listen

Modern artificial intelligence applications frequently require processing both images and text simultaneously. However, existing vision-language models typically face a structural trade-off: dual-encoder systems process text and images separately to provide fast retrieval speeds at the cost of reasoning accuracy, while fusion-encoder systems combine modalities deeply for high-accuracy classification at the cost of slow, computationally expensive retrieval. The article introduces a unified framework named the Vision-Language Pre-trained Model (VLMO) to eliminate this trade-off by dynamically serving as either a dual encoder or a fusion encoder within a single architecture.

To achieve this, the article develops a Multiway Transformer architecture that incorporates specialized modality experts alongside shared self-attention layers. Rather than relying on separate networks or computationally heavy external object detectors, the framework routes visual, linguistic, or joint vision-language representations to dedicated processing blocks based on the input type. The training strategy uses a staged pipeline: the model first learns general image features from unlabeled images, then trains text experts on text-only corpora while freezing image parameters, and finally undergoes joint multimodal pre-training using image-text contrastive learning, masked language modeling, and image-text matching with global hard negative mining. Evaluation was conducted across standard benchmarks with base setups (4 million images across 10 million pairs) and scaled up to 1 billion web image-text pairs.

Key findings show that VLMO delivers leading performance across both classification and retrieval tasks. On complex reasoning benchmarks, the model outperforms previous approaches, achieving up to an 82.88 score on Visual Question Answering (VQA) and 89.54% accuracy on Natural Language for Visual Reasoning (NLVR2) when scaled to larger datasets. On text-and-image retrieval benchmarks such as COCO and Flickr30K, the model matches or exceeds the accuracy of slower fusion-based systems while providing fast linear-time retrieval speeds. Ablation studies confirm that stagewise pre-training using unpaired data substantially improves downstream accuracy (raising NLVR2 dev scores from 80.33% to 82.09%), and global hard negative mining further boosts performance over local GPU sampling.

These results demonstrate that organizations can deploy a single unified architecture across diverse vision-language workflows instead of maintaining separate models for search and deep classification. Eliminating the need for complex object detectors and enabling fast dual-encoder retrieval significantly lowers computational overhead, latency, and operational deployment costs in production environments.

Based on these findings, technical teams should consider adopting unified modular architectures and stagewise pre-training strategies when building multimodal systems, especially when leveraging large amounts of readily available unpaired images and text. Future development highlighted in the article includes scaling the model size further, expanding the framework to generative tasks such as automated image captioning, and extending the multiway expert design to additional data types such as video, audio, and structured knowledge.

Confidence in the reported experimental improvements is high across standard academic benchmarks. However, leaders should note that the highest performance gains rely on large-scale training using 1 billion noisy web pairs, which demands substantial computing resources and may introduce domain-specific noise or dataset limitations in specialized real-world settings.

No sufficiently relevant recommendations were found.

Cover for VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Abstract

We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Multiway Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of Multiway Transformer, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. The code and pretrained models are available at http://aka.ms/vlmo.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Input Representations
  • 3.2 Mixture-of-Modality-Experts Transformer
  • 3.3 Pre-Training Tasks
  • 3.4 Stagewise Pre-Training
  • 3.5 Fine-Tuning VLMo on Downstream Tasks
  • 4 Experiments
  • 4.1 Pre-Training Setup
  • 4.2 Training on Larger-scale Datasets
  • 4.3 Evaluation on Vision-Language Classification Tasks
  • 4.4 Evaluation on Vision-Language Retrieval Tasks
  • 4.5 Evaluation on Vision Tasks
  • 4.6 Ablation Studies
  • 5 Conclusion
  • References
  • A Ablation Study of Shared Self-Attention
  • B Hyperparameters for Text-Only Pre-Training
  • C Hyperparameters for Vision-Language Classification Fine-Tuning
  • D Hyperparameters for Vision-Language Retrieval Fine-Tuning

Knowls

  1. Knowl 1 — Multiway Transformer uses shared attention with hard-routed modality experts

    model/method

    VLMO’s Multiway Transformer replaces the feed-forward network in each Transformer block with modality-specific feed-forward experts while retaining one self-attention module shared across modalities. Each expert is a feed-forward network with two linear transformations and an activation. For layer ll and input hidden-state sequence Hl−1H_{l-1}, the block is

    Hl′=MSA⁡(LN⁡(Hl−1))+Hl−1,Hl=MultiwayFFN⁡(LN⁡(Hl′))+Hl′.H'_l = \operatorname{MSA}(\operatorname{LN}(H_{l-1})) + H_{l-1}, \qquad H_l = \operatorname{MultiwayFFN}(\operatorname{LN}(H'_l)) + H'_l.

    Here, l∈{1,…,L}l\in\{1,\ldots,L\} is the layer index, LL is the number of Transformer blocks, Hl−1H_{l-1} and Hl′H'_l are input and intermediate token-state sequences, HlH_l is the output sequence, MSA⁡\operatorname{MSA} is shared multi-head self-attention, and LN⁡\operatorname{LN} is layer normalization. The feed-forward routing is determined directly by input modality and layer, rather than by a learned gate: image-only sequences use the vision expert (V-FFN), text-only sequences use the language expert (L-FFN), and image-text sequences use the respective vision and language experts in lower layers before using the vision-language expert (VL-FFN) in upper layers. Thus the shared attention can contextualize across modalities, while the experts retain modality-specific transformations.

  2. Knowl 2 — VLMO reuses one parameter-shared model as a dual encoder or a fusion encoder

    model/method

    VLMO is a single Multiway Transformer-based pretrained model that can be deployed in two configurations. For retrieval, it acts as a dual encoder: it separately encodes image and text sequences, takes their respective special-token representations, projects and normalizes them, and computes image-text scores by dot product. Because image and text representations can be encoded and stored independently, retrieval avoids jointly processing every possible image-text pair. For classification, VLMO acts as a fusion encoder: it processes the concatenated image-text sequence, uses separate vision and language experts in lower layers and the vision-language expert in upper layers, and feeds the final [TCLS][T_{\mathrm{CLS}}] representation to a task-specific classifier. In the paper’s configurations, VLMO-Base uses the vision-language expert in the top two layers and VLMO-Large in the top three.

    Image inputs are formed by linearly projecting flattened image patches, prepending an image classification token, and adding position and type embeddings. Text inputs are WordPiece token sequences with start and separator tokens and position and type embeddings. The image and text sequences are concatenated for fusion. The same pretrained parameter set therefore supports separate image/text encoding as well as joint cross-modal encoding.

  3. Knowl 3 — Stagewise pre-training transfers image and text knowledge into vision-language learning

    algorithm

    VLMO’s stagewise pre-training initializes its image and text capabilities separately before joint vision-language training:

    1. Image stage: Train the Multiway Transformer’s self-attention modules and vision expert on image-only data using masked image modeling; the paper initializes these parameters from a pretrained BEiT model.
    2. Text stage: Freeze the self-attention modules and vision expert, then train the language expert on text-only data using masked language modeling. Freezing preserves the image knowledge from the first stage.
    3. Vision-language stage: Initialize from the resulting model and train the full model on image-text pairs with image-text contrastive learning, image-text matching, and masked language modeling.

    The strategy is intended to exploit larger image-only and text-only corpora than are available as paired data, and to expose the model to text more varied than the relatively short captions in image-text datasets. The paper reports that adding the text-only stage after image-only pre-training improves downstream results.

  4. Knowl 4 — Joint pre-training combines contrastive alignment, masked language modeling, and pair matching

    model/method

    VLMO jointly trains three objectives using shared Multiway Transformer parameters. For image-text contrastive learning, the final image and text classification-token states are separately projected and normalized. For a batch of NN matched image-text pairs, let vi,wi∈RDv_i,w_i\in\mathbb{R}^{D} be the normalized image and text vectors for pair ii, where DD is their embedding dimension. The two directional similarities and probabilities are

    si,ji2t=vi⊤wj,si,jt2i=wi⊤vj,s^{i2t}_{i,j}=v_i^{\top}w_j,\qquad s^{t2i}_{i,j}=w_i^{\top}v_j, pii2t=exp⁡(si,ii2t/σ)∑j=1Nexp⁡(si,ji2t/σ),pit2i=exp⁡(si,it2i/σ)∑j=1Nexp⁡(si,jt2i/σ).p^{i2t}_{i}=\frac{\exp(s^{i2t}_{i,i}/\sigma)}{\sum_{j=1}^{N}\exp(s^{i2t}_{i,j}/\sigma)},\qquad p^{t2i}_{i}=\frac{\exp(s^{t2i}_{i,i}/\sigma)}{\sum_{j=1}^{N}\exp(s^{t2i}_{i,j}/\sigma)}.

    Here i,j∈{1,…,N}i,j\in\{1,\ldots,N\} index batch pairs, si,ji2ts^{i2t}_{i,j} scores image ii against text jj, si,jt2is^{t2i}_{i,j} scores text ii against image jj, and σ\sigma is a learned temperature. Cross-entropy is applied in both directions to identify the matched pair on the diagonal.

    For masked language modeling, 15% of text tokens are selected for masking; the model predicts the original tokens from the remaining text and visual context, using cross-entropy over the text vocabulary. For image-text matching, a binary classifier predicts whether a pair is matched from its final [TCLS][T_{\mathrm{CLS}}] state. Negative pairs are selected as hard negatives using the contrastive image-to-text and text-to-image similarities.

  5. Knowl 5 — Pre-training data, model sizes, and optimization settings

    experimental setup

    The standard vision-language pre-training set combines Conceptual Captions, SBU Captions, COCO, and Visual Genome, comprising about 4 million images and 10 million image-text pairs. VLMO-Base has 12 Transformer layers, hidden size 768, 12 attention heads, and 175 million parameters; VLMO-Large has 24 layers, hidden size 1024, 16 heads, and 562 million parameters. Images are pre-trained at resolution 224×224224\times224 with 16×1616\times16 patches; text uses uncased BERT WordPiece tokenization with maximum sequence length 40 and whole-word masking. RandAugment is applied to images.

    The paired-data runs use 200,000 training steps and batch size 1,024, AdamW with β1=0.9\beta_1=0.9 and β2=0.98\beta_2=0.98, weight decay 0.01, linear warmup over the first 2,500 steps, and then linear learning-rate decay. Peak learning rates are 2×10−42\times10^{-4} for Base and 5×10−55\times10^{-5} for Large. The larger-scale VLMO-Large++ run uses one billion noisy web image-text pairs: 200,000 steps at batch size 16,000 followed by 100,000 steps at batch size 32,000, with the other hyperparameters reported as the same as for the 4-million-image training.

  6. Knowl 6 — VLMO improves VQA and NLVR2 classification results

    empirical result

    The models were fine-tuned as fusion encoders. VQA is reported as VQA score on the test-dev and test-standard splits; NLVR2 is reported as accuracy on dev and the public test split (test-P). The comparison below includes the paper’s VLMO models, a 4-million-image baseline, and large-scale reference models; the training-image counts differ for the large-scale rows.

    Model Pretrain images VQA test-dev VQA test-std NLVR2 dev NLVR2 test-P
    ALBEF-Base 4M 74.54 74.70 80.24 80.50
    VLMO-Base 4M 76.64 76.89 82.77 83.34
    VLMO-Large 4M 79.94 79.98 85.64 86.86
    SimVLM-Huge 1.8B 80.03 80.34 84.53 85.15
    Florence-Huge 900M 80.16 80.36 – –
    Flamingo 2.3B 82.00 82.10 – –
    VLMO-Large++ 1.0B 82.88 82.78 88.62 89.54

    At the 4M-image scale, VLMO-Base exceeds the listed ALBEF-Base scores on all four reported metrics, and VLMO-Large further improves on VLMO-Base. VLMO-Large++ obtains the highest listed VQA and NLVR2 values in this comparison, although it uses one billion pre-training images rather than 4M.

  7. Knowl 7 — Dual-encoder VLMO achieves strong image-text retrieval scores

    empirical result

    For retrieval fine-tuning and inference, VLMO separately encodes images and text and scores pairs by dot product. Results below are Recall@1, Recall@5, and Recall@10 for text retrieval (TR) and image retrieval (IR), evaluated on the COCO 5K test set and Flickr30K 1K test set. Image counts refer to pre-training data. The included ALBEF-Base reranks candidates with a fusion encoder, whereas VLMO, ALIGN, and Florence use separate encoders and shallow dot-product interaction.

    Model Images COCO TR R@1 R@5 R@10 COCO IR R@1 R@5 R@10 Flickr TR R@1 R@5 R@10 Flickr IR R@1 R@5 R@10
    ALBEF-Base 4M 73.1 91.4 96.0 56.8 81.5 89.2 94.3 99.4 99.8 82.8 96.7 98.4
    VLMO-Base 4M 74.8 93.1 96.9 57.2 82.6 89.8 92.3 99.4 99.9 79.3 95.7 97.8
    VLMO-Large 4M 78.2 94.4 97.4 60.6 84.4 91.0 95.3 99.9 100.0 84.5 97.3 98.6
    ALIGN-Large 1.8B 77.0 93.5 96.9 59.9 83.3 89.8 95.3 99.8 100.0 84.9 97.4 98.6
    Florence-Huge 900M 81.8 95.2 – 63.2 85.7 – 97.2 99.9 – 87.9 98.1 –
    VLMO-Large++ 1.0B 83.1 96.0 98.2 65.2 86.5 92.2 96.8 100.0 100.0 88.1 98.4 99.3

    On both datasets, VLMO-Large improves on VLMO-Base across all reported recall values. VLMO-Large++ has the highest listed COCO recall values and exceeds the listed Florence-Huge values where both are reported; its Flickr results also reach 100.0 at R@5 and R@10 for text retrieval and at R@10 for image retrieval. Unlike joint fusion-encoder scoring over every image-text combination, separate encoding permits pre-computation of the image and text vectors.

  8. Knowl 8 — Ablations show benefits from text-only initialization, modality experts, and combined objectives

    empirical result

    The stagewise initialization comparison evaluates NLVR2 and Flickr30K retrieval. NLVR2 values are averages over three runs; Flickr30K TR and IR are averages of Recall@1, Recall@5, and Recall@10. Adding text-only pre-training after image-only pre-training improves all four reported measures.

    Initialization NLVR2 dev NLVR2 test-P Flickr30K TR Flickr30K IR
    Image-only pre-training 80.33 81.06 95.60 87.69
    Image-only + text-only pre-training 82.09 82.49 95.67 88.52

    A separate ablation, initialized from ViT-Base, compares pre-training losses and Transformer variants. Flickr30K values are average recall over R@1, R@5, and R@10; NLVR2 values are averages over three runs. The standard-Transformer configurations use the three listed objectives where checked; the Multiway configurations compare removing versus retaining the vision-language expert.

    Configuration Pre-training objectives NLVR2 dev NLVR2 test-P Flickr TR Flickr IR
    Standard Transformer ITC 58.51 58.83 92.23 84.24
    Standard Transformer ITC + MLM 73.91 73.75 94.07 85.82
    Standard Transformer ITC + ITM 76.46 76.19 94.37 85.67
    Standard Transformer ITC + ITM + MLM 78.81 79.27 93.37 85.73
    Multiway without VL expert ITC + ITM + MLM 79.58 80.11 94.50 86.69
    Multiway with VL expert ITC + ITM + MLM 80.13 80.31 95.17 87.25

    Here ITC is image-text contrastive learning, ITM is image-text matching, and MLM is masked language modeling. The results show large classification gains from adding MLM or ITM to ITC, better performance for Multiway Transformer than the standard-Transformer configuration with all three objectives, and additional gains when the vision-language expert is retained.

  9. Knowl 9 — Global hard-negative mining improves NLVR2 over per-GPU mining

    empirical result

    VLMO samples hard image-text negatives using contrastive similarities. In the reported Base-model experiment, 32 V100 GPUs each process 32 examples, giving a total batch of 1,024. Local hard-negative mining searches the 32 examples on one GPU; global hard-negative mining gathers examples across all GPUs and searches 1,024 candidates. The resulting NLVR2 accuracies are:

    Negative-mining candidates NLVR2 dev NLVR2 test-P
    Local, 32 examples 77.70 77.95
    Global, 1,024 examples 79.54 79.48

    Under these conditions, global mining improves dev accuracy by 1.84 points and test-P accuracy by 1.53 points relative to local mining.

Coverage note — The image-only transfer evaluation on ImageNet and ADE20K is omitted because it is secondary to the paper’s unified vision-language methods and core retrieval/classification results; it reports VLMO-Base at ImageNet acc@1 85.5 and ADE20K mIoU 53.4, compared with BEiT-Base at 85.2 and 52.8.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangoeui, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. CoRR, abs/2204.14198, 2022. doi: 10.48550/arXiv.2204.14198. URL https://doi.org/10.48550/arXiv.2204.14198.
  2. 2.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, and Hsiao-Wuen Hon. UniLMv2: Pseudo-masked language models for unified language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 642–652. PMLR, 2020. URL http://proceedings.mlr.press/v119/bao20a.html.
  3. 3.Hangbo Bao, Li Dong, and Furu Wei. BEiT: BERT pre-training of image transformers. CoRR, abs/2106.08254, 2021. URL https://arxiv.org/abs/2106.08254.
  4. 4.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: universal image-text representation learning. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXX, volume 12375 of Lecture Notes in Computer Science, pages 104–120. Springer, 2020. doi: 10.1007/978-3-030-58577-8_7. URL https://doi.org/10.1007/978-3-030-58577-8_7.
  5. 5.Zewen Chi, Li Dong, Furu Wei, Wenhui Wang, Xianling Mao, and Heyan Huang. Cross-lingual natural language generation via pre-training. CoRR, abs/1909.10481, 2019.
  6. 6.Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. InfoXLM: An information-theoretic framework for cross-lingual language model pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3576–3588, Online, June 2021. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2021.naacl-main.280.
  7. 7.Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Saksham Singhal, Payal Bajaj, Xia Song, and Furu Wei. XLM-E: Cross-lingual language model pre-training via ELECTRA. ArXiv, abs/2106.16138, 2021.
  8. 8.Alexis Conneau and Guillaume Lample. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/c04c19c2c2474dbf5f7ac4372c5b9af1-Paper.pdf.
  9. 9.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online, July 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.acl-main.747.
  10. 10.Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. Randaugment: Practical automated data augmentation with a reduced search space. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR Workshops 2020, Seattle, WA, USA, June 14-19, 2020, pages 3008–3017. Computer Vision Foundation / IEEE, 2020.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics, 2019. doi: 10.18653/v1/n19-1423. URL https://doi.org/10.18653/v1/n19-1423.
  12. 12.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language understanding and generation. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13042–13054, 2019.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. preprint arXiv:2010.11929, 2020.
  14. 14.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. CoRR, abs/2101.03961, 2021. URL https://arxiv.org/abs/2101.03961.
  15. 15.Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. Large-scale adversarial training for vision-and-language representation learning. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  16. 16.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 6325–6334. IEEE Computer Society, 2017. doi: 10.1109/CVPR.2017.670. URL https://doi.org/10.1109/CVPR.2017.670.
  17. 17.Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. CoRR, abs/2004.00849, 2020. URL https://arxiv.org/abs/2004.00849.
  18. 18.Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu. Seeing out of the box: End-to-end pre-training for vision-language representation learning. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 12976–12985. Computer Vision Foundation / IEEE, 2021.
  19. 19.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 4904–4916. PMLR, 2021. URL http://proceedings.mlr.press/v139/jia21b.html.
  20. 20.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 3128–3137. IEEE Computer Society, 2015.
  21. 21.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 5583–5594. PMLR, 2021. URL http://proceedings.mlr.press/v139/kim21k.html.
  22. 22.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis., 123(1):32–73, 2017.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019.
  24. 24.Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven C. H. Hoi. Align before fuse: Vision and language representation learning with momentum distillation. CoRR, abs/2107.07651, 2021. URL https://arxiv.org/abs/2107.07651.
  25. 25.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. CoRR, abs/1908.03557, 2019. URL http://arxiv.org/abs/1908.03557.
  26. 26.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2592–2607. Association for Computational Linguistics, 2021.
  27. 27.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXX, volume 12375 of Lecture Notes in Computer Science, pages 121–137. Springer, 2020. doi: 10.1007/978-3-030-58577-8_8. URL https://doi.org/10.1007/978-3-030-58577-8_8.
  28. 28.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In David J. Fleet, Tomás Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V, volume 8693 of Lecture Notes in Computer Science, pages 740–755. Springer, 2014.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019. URL http://arxiv.org/abs/1907.11692.
  30. 30.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019.
  31. 31.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13–23, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/c74d97b01eae257e44aa9d5bade97baf-Abstract.html.
  32. 32.Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, and Furu Wei. Deltalm: Encoder-decoder pre-training for language generation and translation by augmenting pretrained multilingual encoders. CoRR, abs/2106.13736, 2021. URL https://arxiv.org/abs/2106.13736.
  33. 33.Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. Im2text: Describing images using 1 million captioned photographs. In John Shawe-Taylor, Richard S. Zemel, Peter L. Bartlett, Fernando C. N. Pereira, and Kilian Q. Weinberger, editors, Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12-14 December 2011, Granada, Spain, pages 1143–1151, 2011. URL https://proceedings.neurips.cc/paper/2011/hash/5dd9db5e033da9c6fb5ba83c7a7ebea9-Abstract.html.
  34. 34.Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2641–2649. IEEE Computer Society, 2015. doi: 10.1109/ICCV.2015.303. URL https://doi.org/10.1109/ICCV.2015.303.
  35. 35.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. URL https://s3-us-west-2.amazonaws.com/openaiassets/research-covers/language-unsupervised/languageunderstandingpaper.pdf.
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 2021. URL http://proceedings.mlr.press/v139/radford21a.html.
  37. 37.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  38. 38.Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell., 39(6):1137–1149, 2017. doi: 10.1109/TPAMI.2016.2577031. URL https://doi.org/10.1109/TPAMI.2016.2577031.
  39. 39.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. IJCV, 2015.
  40. 40.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 2556–2565. Association for Computational Linguistics, 2018. URL https://aclanthology.org/P18-1238/.
  41. 41.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=B1ckMDqlg.
  42. 42.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. VL-BERT: pre-training of generic visual-linguistic representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SygXPaEYvH.
  43. 43.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 6418–6428. Association for Computational Linguistics, 2019. doi: 10.18653/v1/p19-1644. URL https://doi.org/10.18653/v1/p19-1644.
  44. 44.Hao Tan and Mohit Bansal. LXMERT: learning cross-modality encoder representations from transformers. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 5099–5110. Association for Computational Linguistics, 2019. doi: 10.18653/v1/D19-1514. URL https://doi.org/10.18653/v1/D19-1514.
  45. 45.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. preprint arXiv:2012.12877, 2020.
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017. URL http://papers.nips.cc/paper/7181-attention-is-all-you-need.
  47. 47.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. CoRR, abs/2202.03052, 2022. URL https://arxiv.org/abs/2202.03052.
  48. 48.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. CoRR, abs/2108.10904, 2021. URL https://arxiv.org/abs/2108.10904.
  49. 49.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. Google’s neural machine translation system: Bridging the gap between human and machine translation. CoRR, abs/1609.08144, 2016. URL http://arxiv.org/abs/1609.08144.
  50. 50.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang. Florence: A new foundation model for computer vision. CoRR, abs/2111.11432, 2021. URL https://arxiv.org/abs/2111.11432.
  51. 51.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 5579–5588. Computer Vision Foundation / IEEE, 2021. URL https://openaccess.thecvf.com/content/CVPR2021/html/Zhang_VinVL_Revisiting_Visual_Representations_in_Vision-Language_Models_CVPR_2021_paper.html.
  52. 52.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ADE20K dataset. Int. J. Comput. Vis., 127(3):302–321, 2019. doi: 10.1007/s11263-018-1140-0. URL https://doi.org/10.1007/s11263-018-1140-0.
  53. 53.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. Unified vision-language pre-training for image captioning and VQA. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 13041–13049. AAAI Press, 2020.

Citation

MLA
Bao, H., et al. “VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 32897–912, https://proceedings.neurips.cc/paper_files/paper/2022/file/d46662aa53e78a62afd980a29e0c37ed-Paper-Conference.pdf.
APA
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., Piao, S., & Wei, F. (2022). VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts. Advances in Neural Information Processing Systems, 35, 32897–32912. https://proceedings.neurips.cc/paper_files/paper/2022/file/d46662aa53e78a62afd980a29e0c37ed-Paper-Conference.pdf
Chicago
Bao, H., W. Wang, L. Dong, et al. 2022. “VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts”. Advances in Neural Information Processing Systems 35: 32897–912. https://proceedings.neurips.cc/paper_files/paper/2022/file/d46662aa53e78a62afd980a29e0c37ed-Paper-Conference.pdf.
Harvard
Bao, H. et al. (2022) “VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 32897–32912. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/d46662aa53e78a62afd980a29e0c37ed-Paper-Conference.pdf.
Vancouver
1. Bao H, Wang W, Dong L, Liu Q, Mohammed OK, Aggarwal K, Som S, Piao S, Wei F (2022) VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 32897–32912

BibTeX

@inproceedings{bao2022vlmo,
  title = {VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts},
  author = {Bao, Hangbo and Wang, Wenhui and Dong, Li and Liu, Qiang and Mohammed, Owais Khan and Aggarwal, Kriti and Som, Subhojit and Piao, Songhao and Wei, Furu},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {32897-32912},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/d46662aa53e78a62afd980a29e0c37ed-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission