SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Zirui WangJiahui YuAdams Wei YuZihang DaiYulia TsvetkovYuan Cao

article2021ICLR976 citations

Introduces SimVLM, a simplified vision-language model trained end-to-end on weakly supervised data using a single prefix language modeling objective, achieving state-of-the-art benchmark performance and strong zero-shot multimodal capabilities without requiring expensive object-level annotations.

Listen

Jointly processing visual and textual data is critical for advanced artificial intelligence applications, yet current vision-language models face severe scalability bottlenecks. Existing systems typically depend on expensive, manually annotated object detection labels and rely on complex multi-stage training pipelines with multiple competing objectives. These constraints limit scalability, increase computational overhead, and restrict the models from effectively generalizing to new, unseen tasks without extensive retraining.

The article demonstrates the Simple Visual Language Model (SimVLM), a streamlined vision-language pretraining framework. The primary objective is to evaluate whether an end-to-end model trained on weakly aligned web data using a single generative language modeling objective can outperform complex, heavily engineered baseline architectures on multimodal benchmarks.

To achieve this, the researchers implemented an encoder-decoder Transformer architecture that processes raw image patches via an initial convolutional stage, removing the need for auxiliary object detection systems. The model was pretrained from scratch on massive web-scraped datasets—specifically 1.8 billion noisy image-text pairs alongside 800 gigabytes of clean text-only data—using a single "Prefix Language Modeling" objective. This formulation allows the network to process contextual image and text prefixes bidirectionally while generating subsequent text autoregressively. The framework was evaluated across standard discriminative and generative multimodal benchmarks across varying model sizes.

The findings show that SimVLM establishes new state-of-the-art performance across all tested vision-language tasks while simplifying the training pipeline. SimVLM achieved an 80.34% score on the Visual Question Answering (VQA) benchmark, surpassing the 80% threshold for the first time with an absolute improvement of nearly 4 percentage points over prior state-of-the-art systems. On image captioning tasks, the model achieved an average improvement of over 10 CIDEr points, outperforming established methods without requiring specialized reinforcement learning optimization. Furthermore, SimVLM demonstrated strong zero-shot generalization, enabling open-ended visual question answering beyond fixed candidate vocabularies and zero-shot cross-modality transfer, where a model fine-tuned entirely on text data performed robustly on multimodal image tasks.

These results demonstrate that complex, object-detection-dependent pipelines can be replaced with a single, unified generative objective paired with large-scale weak supervision. For organizations building vision-language applications, this approach significantly reduces data annotation costs, simplifies model maintenance, and eliminates the engineering friction of balancing multiple loss functions. Additionally, the model’s strong zero-shot and open-ended generative capabilities provide greater adaptability for real-world scenarios where candidate answers cannot be predefined.

Organizations developing multimodal artificial intelligence systems should transition away from complex, multi-stage pipelines in favor of unified generative pretraining on large, weakly supervised datasets. Next steps should focus on piloting these generative architectures in production environments where open-ended visual understanding is required, such as customer support automation or content moderation.

Confidence in these findings is high given the consistent benchmark gains and thorough ablation analyses. However, decision-makers should note that the model relies on immense web-scale datasets, which can introduce noise. The authors observed that generating high-quality answers in open-ended visual question answering required brief secondary pretraining on higher-quality, knowledge-rich data like Wikipedia to mitigate noise from web-crawled captions.

arXiv: 2108.10904
Cover for SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Abstract

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives. In this work, we relax these constraints and present a minimalist pretraining framework, named Simple Visual Language Model (SimVLM). Unlike prior work, SimVLM reduces the training complexity by exploiting large-scale weak supervision, and is trained end-to-end with a single prefix language modeling objective. Without utilizing extra data or task-specific customization, the resulting model significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA (+3.74% vqa-score), NLVR2 (+1.17% accuracy), SNLI-VE (+1.37% accuracy) and image captioning tasks (+10.1% average CIDEr score). Furthermore, we demonstrate that SimVLM acquires strong generalization and transfer ability, enabling zero-shot behavior including open-ended visual question answering and cross-modality transfer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 SimVLM
  • 3.1 Background
  • 3.2 Proposed Objective: Prefix Language Modeling
  • 3.3 Architecture
  • 3.4 Datasets
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Comparison with existing approaches
  • 4.3 Zero-Shot Generalization
  • 4.3.1 Zero-shot/Few-shot Image Captioning
  • 4.3.2 Zero-shot cross-modality Transfer
  • 4.3.3 Open-ended VQA
  • 4.4 Analysis
  • 5 Conclusion
  • References
  • A Generated Examples
  • B Experimental Details
  • B.1 Pretraining
  • B.2 Finetuning
  • C Model Performance on Language-only Task
  • D Erratum

Knowls

  1. Knowl 1 — Prefix language modeling unifies bidirectional encoding and autoregressive generation

    equation

    SimVLM trains with Prefix Language Modeling (PrefixLM). For a token sequence x=(x1,…,xT)x=(x_1,\ldots,x_T) drawn from pretraining distribution DD, a prefix boundary TpT_p divides the sequence into a prefix x<Tpx_{<T_p} and a suffix xTp:Tx_{T_p:T}. The model may attend bidirectionally within the prefix, but predicts suffix tokens autoregressively:

    LPrefixLM(θ)=−Ex∼D[∑t=TpTlog⁡pθ(xt∣x<Tp,xTp:t−1)].\mathcal{L}_{\mathrm{PrefixLM}}(\theta)=-\mathbb{E}_{x\sim D}\left[\sum_{t=T_p}^{T}\log p_{\theta}(x_t\mid x_{<T_p},x_{T_p:t-1})\right].

    Here, θ\theta denotes the trainable model parameters, pθp_{\theta} is the model’s conditional token distribution, and xTp:t−1x_{T_p:t-1} is the already observed portion of the suffix, empty when t=Tpt=T_p. During training, the prefix length is randomly selected. For image-text examples, image tokens are placed before text tokens, and the selected prefix must include the image tokens; the loss is applied to text tokens. This single objective supports contextual encoding of the prefix and text generation from it, and it can also train on text-only sequences.

  2. Knowl 2 — SimVLM encodes raw images as contextualized patches in an encoder-decoder model

    model/method

    SimVLM uses a Transformer encoder-decoder and takes raw pixels rather than object-detection regions as visual input. An image of height HH, width WW, and CC channels is converted into a one-dimensional sequence of Ti=HW/P2T_i=HW/P^2 tokens using patch size PP; each token has the Transformer hidden dimension DD. Before patch tokens enter the Transformer, a convolutional stage comprising the first three ResNet blocks (excluding the convolutional stem) produces contextualized patches. This stage is the paper’s alternative to ViT-style direct linear patch projection. Text is tokenized into subword tokens and embedded using a learned vocabulary.

    The model adds separate learned one-dimensional positional embeddings to image and text inputs, and uses two-dimensional relative attention for image patches. It does not use modality-type embeddings. Model parameters are shared across modalities except for the convolutional stage and positional embeddings. The architecture diagram on page 3 depicts the image-patch and text-token sequences being encoded as a prefix, with the decoder predicting the text continuation.

  3. Knowl 3 — Pretraining scales weakly aligned image-text data together with text-only data

    experimental setup

    SimVLM Base, Large, and Huge were pretrained from scratch for about one million steps on ALIGN image-alt-text pairs and the C4 text corpus. ALIGN contains about 1.8 billion noisy web image-text pairs; the paper used the training data with minimal processing, applying simple random resized cropping but no additional filtering. C4 consists of about 800 GB of web-crawled documents and follows its standard preprocessing. Each training batch mixed 4,096 ALIGN image-text pairs with 512 C4 documents, distributed across 512 TPU v3 chips.

    Pretraining used 224 × 224 images and 16 × 16 patches, yielding a 14 × 14 visual-token grid; text had a 32,000-token vocabulary and maximum sequence length 256. The convolutional stages used the first three ResNet-101 blocks for Base and ResNet-152 blocks for Large, with a wider ResNet-152 variant for Huge. Optimization used AdamW with β1=0.9\beta_1=0.9, β2=0.999\beta_2=0.999, and weight decay 0.01. The learning rate warmed up for 2% of updates to 5×10−45\times10^{-4} and then decayed linearly; pretraining used no dropout.

  4. Knowl 4 — SimVLM improves results across six vision-language benchmarks

    empirical result

    After pretraining, the models were fine-tuned and evaluated on VQA v2, NLVR2, SNLI-VE, COCO captioning, NoCaps, and English-to-German Multi30k translation. The reported single-model results for SimVLM Base, Large, and Huge are listed below. VQA values are vqa-scores; NLVR2 and SNLI-VE values are accuracies; COCO reports BLEU-4, METEOR, CIDEr, and SPICE; NoCaps reports CIDEr and SPICE; Multi30k reports BLEU-4.

    SimVLM Base: VQA test-dev/test-std 77.87/78.14; NLVR2 dev/test-P 81.72/81.77; SNLI-VE dev/test 84.20/84.15; COCO 39.0/32.9/134.8/24.0; NoCaps 94.8/13.1; Multi30k 46.6.

    SimVLM Large: VQA 79.32/79.56; NLVR2 84.13/84.84; SNLI-VE 85.68/85.62; COCO 40.3/33.4/142.6/24.7; NoCaps 108.5/14.2; Multi30k 47.5.

    SimVLM Huge: VQA 80.03/80.34; NLVR2 84.53/85.15; SNLI-VE 86.21/86.32; COCO 40.6/33.7/143.3/25.4; NoCaps 110.3/14.5; Multi30k 47.6.

    The paper reports that the models set new results across the evaluated tasks. For context, the strongest prior values shown in its comparison include VinVL’s VQA 76.56/76.60, NLVR2 82.67/83.98, COCO CIDEr 140.9, and NoCaps CIDEr/SPICE 92.5/13.1; SOHO’s SNLI-VE 85.00/84.95; and VL-T5’s Multi30k BLEU-4 45.5. On COCO, SimVLM Huge exceeded the listed VinVL value on three of the four caption metrics, while its BLEU-4 score of 40.6 was below OSCAR’s 41.7.

  5. Knowl 5 — Pretrained SimVLM supports zero-shot and few-shot image captioning

    empirical result

    For zero-shot captioning, the pretrained model decoded benchmark images without task fine-tuning; the authors found that adding the text prefix A picture of improved caption quality. For few-shot captioning, they fine-tuned on 1% of the training data for five epochs. The following COCO results are BLEU-4, METEOR, CIDEr, and SPICE, in that order.

    Zero-shot: Base 9.5/11.5/24.0/7.5; Large 10.5/12.0/24.9/8.3; Huge 11.2/14.7/32.2/8.5. Few-shot: Base 34.7/29.2/118.7/21.9; Large 35.4/30.2/124.1/22.7; Huge 36.8/31.5/131.3/24.0.

    On NoCaps, scores are reported as CIDEr for in-domain, near-domain, out-of-domain, and overall examples. Zero-shot scores were Base 83.2/84.1/82.5/83.5, Large 97.6/96.5/96.3/96.6, and Huge 101.2/100.4/102.3/101.4. Few-shot scores were Base 95.0/91.9/98.5/93.7, Large 102.5/100.9/106.0/102.2, and Huge 111.8/110.6/111.0/110.4. The paper presents these results as evidence that captioning behavior learned from noisy web image-text data transfers to captioning benchmarks, including NoCaps’ concept-rich examples.

  6. Knowl 6 — Text-only fine-tuning transfers to image-grounded entailment and translation

    empirical result

    The paper evaluates zero-shot cross-modality transfer by fine-tuning SimVLM on text-only downstream data and then testing it on vision-language examples without further training on those examples. For SNLI-VE, separate text-only fine-tuning runs used SNLI-VE text, SNLI, or MNLI: the premise was supplied to the encoder, the hypothesis to the decoder, and a classifier was trained from the final decoder-token representation. At evaluation, the premise image replaced the text premise while the hypothesis remained text. For Multi30k, the model was fine-tuned on text-only English-German translation and then decoded with image input.

    SNLI-VE accuracy, reported as dev/test, was as follows. Fine-tuned on SNLI-VE text: Base 71.35/71.02, Large 72.85/72.44, Huge 73.56/73.08. Fine-tuned on SNLI: Base 72.65/72.24, Large 73.62/73.23, Huge 74.24/73.86. Fine-tuned on MNLI: Base 64.37/63.98, Large 66.97/66.31, Huge 67.45/66.97. On text-only-fine-tuned Multi30k, BLEU-4/METEOR were Base 15.0/24.8, Large 17.7/30.1, and Huge 18.2/32.6.

    As a control for SNLI-VE, masking the premise image and predicting from the hypothesis alone produced average scores of 34.31/34.62, which the paper describes as near random guess. The authors interpret the stronger image-conditioned results as evidence of transfer across modalities; the Multi30k setup additionally transfers across language and modality.

  7. Knowl 7 — Generative VQA handles answers outside a fixed candidate set

    empirical result

    In generative VQA fine-tuning, the image and question form the PrefixLM prefix and the model generates a free-form answer. The paper compares this with classification over the 3,129 most frequent training answers. It evaluates the standard VQA dev score, a Karpathy-test split with in-domain and out-of-domain answers, and a partial-training split. The Karpathy out-of-domain category contains examples whose best-scoring answer is absent from the 3,129 candidates. In partial training, 2,085 candidate answers were selected and the model trained only on examples whose best answer was selected; evaluation included held-out answers. Generated answers were scored by exact match to human labels.

    For each model below, values are given as standard dev score; Karpathy-test in-domain/out-of-domain/overall; partial-train in-domain/out-of-domain/overall. Discriminative Base scored 73.8; 79.0/16.7/75.3; 78.4/10.3/70.5. Discriminative Large scored 76.0; 80.4/17.3/76.7; 79.5/11.0/71.8. Discriminative Huge scored 76.5; 81.0/17.5/77.2; 80.2/11.1/72.2. Generative Base scored 73.2; 78.3/25.8/75.2; 77.1/27.1/71.3. Generative Large scored 75.2; 79.5/29.6/76.5; 78.7/28.4/72.5. Generative Huge scored 75.5; 79.9/30.3/77.0; 79.1/28.8/73.0.

    Thus, generative SimVLM was competitive with its discriminative counterpart on the standard and in-domain evaluations, while scoring higher on out-of-domain answers in both challenge settings. The paper also shows generated answers such as surgeon and wood carving, which are not in the fixed candidate set.

  8. Knowl 8 — Open-ended zero-shot VQA required cleaner continued-pretraining data

    limitation

    The pretrained SimVLM could complete prompted visual-text sequences without downstream fine-tuning, but the authors found that it did not reliably produce meaningful answers to real questions in zero-shot VQA. They hypothesized that noisy, short alt-text in the pretraining data limited this behavior. In an additional experiment, they continued pretraining for 50,000 steps on the cleaner WIT dataset; the paper reports that open-ended VQA ability then emerged, with examples of relevant answers to image questions. This result is conditional on the extra WIT training and does not establish reliable open-ended VQA from the original pretrained model alone.

  9. Knowl 9 — Ablations identify PrefixLM, multimodal data, and convolutional patches as important

    empirical result

    The authors ablated SimVLMsmall on VQA; this model had embedding dimension 512 and eight layers. The reported VQA scores were: no pretraining 49.70; decoder-only architecture 65.23; standard LM objective 64.48; full SimVLMsmall 67.43; removing image-text data (Image2Text) 49.23; removing text-only data (Text2Text) 65.25; removing the convolutional stage 63.11; replacing PrefixLM with span corruption 66.23; using two convolutional blocks 65.57; using four blocks 66.55; using 10% of ALIGN 66.71; and using CC-3M 63.32.

    These comparisons support the authors’ conclusions that the encoder-decoder separation, the PrefixLM objective, both image-text and text-only pretraining data, and the convolutional stage contribute to VQA quality. Among the tested convolutional configurations, the three-block setup used by SimVLM performed better than the two- and four-block alternatives. The reduced-data experiments also show lower scores than the full-data model.

  10. Knowl 10 — SimVLM learns useful language-only and image-only representations

    empirical result

    The paper evaluated SimVLM on single-modality tasks as an analysis of its learned representations. On ImageNet linear evaluation, top-1 accuracy was 80.6% for Base, 82.3% for Large, and 83.6% for Huge. The comparison methods scored 79.8% for SimCLRv2, 80.1% for DINO, 85.4% for CLIP, and 85.5% for ALIGN. SimVLM’s visual features were average-pooled encoder outputs; the model had not been pretrained with a discriminative image objective such as contrastive loss.

    On GLUE development data, SimVLM Base scored 46.7 CoLA, 90.9 SST-2, 63.9 RTE, 75.2/84.4 MRPC, 90.4/87.2 QQP, 83.4 MNLI, 88.6 QNLI, and 58.1 WNLI. MRPC and QQP report the two metrics shown in the paper’s table. The paper characterizes these text results as stronger than the compared vision-language pretraining models and competitive with BERT, indicating that the unified multimodal pretraining also produced useful language representations.

Coverage note — Qualitative example captions and the full downstream fine-tuning recipes are omitted because they are illustrative or task-implementation details; the main model, training setup, transfer behavior, benchmark results, ablations, and stated limitation are included.

References

  1. 1.Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. nocaps: novel object captioning at scale. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8948–8957, 2019.
  2. 2.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6077–6086, 2018.
  3. 3.Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. A large annotated corpus for learning natural language inference. arXiv preprint arXiv:1508.05326, 2015.
  4. 4.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  5. 5.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve J ´ egou, Julien Mairal, Piotr Bojanowski, and ´ Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021.
  6. 6.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton. Big self-supervised models are strong semi-supervised learners. Advances in Neural Information Processing Systems, 33:22243–22255, 2020a.
  7. 7.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and ´ C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  8. 8.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In ECCV, 2020b.
  9. 9.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. Unifying vision-and-language tasks via text generation. arXiv preprint arXiv:2102.02779, 2021.
  10. 10.Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. Meshed-memory transformer for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10578–10587, 2020.
  11. 11.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. arXiv preprint arXiv:2106.04803, 2021.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszko￾reit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recogni￾tion at scale. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  14. 14.Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. Multi30k: Multilingual english-german image descriptions. arXiv preprint arXiv:1605.00459, 2016.
  15. 15.Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. Large-scale adversarial training for vision-and-language representation learning. arXiv preprint arXiv:2006.06195, 2020.
  16. 16.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6904–6913, 2017.
  17. 17.Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao. A survey on visual transformer, 2021.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recog￾nition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  19. 19.Xiaowei Hu, Xi Yin, Kevin Lin, Lijuan Wang, Lei Zhang, Jianfeng Gao, and Zicheng Liu. Vivo: Surpassing human performance in novel object captioning with visual vocabulary pre-training. In AAAI, February 2021.
  20. 20.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In European conference on computer vision, pp. 646–661. Springer, 2016.
  21. 21.Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. Attention on attention for image cap￾tioning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4634–4643, 2019.
  22. 22.Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu. Seeing out of the box: End-to-end pre-training for vision-language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12976–12985, 2021.
  23. 23.Taichi Iki and Akiko Aizawa. Effect of vision-and-language extensions on natural language understanding in vision-and-language models. arXiv preprint arXiv:2104.08066, 2021.
  24. 24.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. arXiv preprint arXiv:2102.05918, 2021.
  25. 25.Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Richard Ho, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexan￾der Kaplan, Harshit Khaitan, Andy Koch, Naveen Kumar, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adri￾ana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Bo Tian, Horia Toma, Er￾ick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, and Doe Hyun Yoon. In-datacenter performance analysis of a tensor processing unit, 2017.
  26. 26.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey, 2021.
  27. 27.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convo￾lution or region supervision, 2021.
  28. 28.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. 2016. URL https://arxiv.org/abs/1602.07332.
  29. 29.Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226, 2018.
  30. 30.Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291, 2019.
  31. 31.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language, 2019.
  32. 32.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learn￾ing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguis￾tics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 2592–2607, Online, August 2021. Association for Computational Linguis￾tics. doi: 10.18653/v1/2021.acl-long.202. URL https://aclanthology.org/2021.acl-long.202.
  33. 33.Xiujun Li, Xi Yin, Chunyuan Li, Xiaowei Hu, Pengchuan Zhang, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre￾training for vision-language tasks. ECCV 2020, 2020.
  34. 34.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre￾train, prompt, and predict: A systematic survey of prompting methods in natural language pro￾cessing, 2021.
  35. 35.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  36. 36.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  37. 37.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguis￾tic representations for vision-and-language tasks. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett (eds.), ´ Advances in Neural Information Processing Sys￾tems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/c74d97b01eae257e44aa9d5bade97baf-Paper.pdf.
  38. 38.Ofir Press and Lior Wolf. Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859, 2016.
  39. 39.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language under￾standing by generative pre-training. 2018.
  40. 40.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  41. 41.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  43. 43.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016.
  44. 44.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092, 2021.
  45. 45.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 28. Curran As￾sociates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/file/14bfa6bb14875e45bba028a21ed38046-Paper.pdf.
  46. 46.Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. Self-critical sequence training for image captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7008–7024, 2017.
  47. 47.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4938–4947, 2020.
  48. 48.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556–2565, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1238. URL https://aclanthology.org/P18-1238.
  49. 49.Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, et al. Lingvo: a modular and scalable framework for sequence-to-sequence modeling, 2019.
  50. 50.Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021.
  51. 51.Lucia Specia, Stella Frank, Khalil Sima’An, and Desmond Elliott. A shared task on multimodal machine translation and crosslingual image description. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pp. 543–553, 2016.
  52. 52.Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. arXiv preprint arXiv:2103.01913, 2021.
  53. 53.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre￾training of generic visual-linguistic representations. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SygXPaEYvH.
  54. 54.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. arXiv preprint arXiv:1811.00491, 2018.
  55. 55.Hao Tan and Mohit Bansal. LXMERT: Learning cross-modality encoder representations from transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Lan￾guage Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 5100–5111, Hong Kong, China, November 2019. Association for Com￾putational Linguistics. doi: 10.18653/v1/D19-1514. URL https://aclanthology.org/D19-1514.
  56. 56.Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multi￾modal few-shot learning with frozen language models. Advances in Neural Information Process￾ing Systems, 34, 2021.
  57. 57.Shagun Uppal, Sarthak Bhagat, Devamanyu Hazarika, Navonil Majumdar, Soujanya Poria, Roger Zimmermann, and Amir Zadeh. Multimodal research in vision and language: A review of current and emerging trends, 2020.
  58. 58.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  59. 59.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  60. 60.Adina Williams, Nikita Nangia, and Samuel R Bowman. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426, 2017.
  61. 61.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, and Ross Girshick. Early ´ convolutions help transformers see better, 2021.
  62. 62.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. Visual entailment: A novel task for fine￾grained image understanding. arXiv preprint arXiv:1901.06706, 2019.
  63. 63.Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, and Fei Huang. E2E-VLP: End-to-end vision-language pre-training enhanced by visual learning. In Proceed￾ings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 503–513, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.42. URL https://aclanthology.org/2021.acl-long.42.
  64. 64.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in neural information processing systems, pp. 5754–5764, 2019.
  65. 65.Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vil: Knowl￾edge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 3208–3216, 2021.
  66. 66.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5579–5588, June 2021.

Citation

MLA
Wang, Z., et al. “SimVLM: Simple Visual Language Model Pretraining with Weak Supervision”. arXiv, 2021, https://doi.org/10.48550/arxiv.2108.10904.
APA
Wang, Z., Yu, J., Yu, A. W., Dai, Z., Tsvetkov, Y., & Cao, Y. (2021). SimVLM: Simple Visual Language Model Pretraining with Weak Supervision. arXiv. https://doi.org/10.48550/arxiv.2108.10904
Chicago
Wang, Z., J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao. 2021. “SimVLM: Simple Visual Language Model Pretraining with Weak Supervision”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2108.10904.
Harvard
Wang, Z. et al. (2021) “SimVLM: Simple Visual Language Model Pretraining with Weak Supervision”. arXiv. Available at: https://doi.org/10.48550/arxiv.2108.10904.
Vancouver
1. Wang Z, Yu J, Yu AW, Dai Z, Tsvetkov Y, Cao Y (2021) SimVLM: Simple Visual Language Model Pretraining with Weak Supervision. https://doi.org/10.48550/arxiv.2108.10904

BibTeX

@misc{https://doi.org/10.48550/arxiv.2108.10904,
  doi = {10.48550/ARXIV.2108.10904},
  url = {https://arxiv.org/abs/2108.10904},
  author = {Wang, Zirui and Yu, Jiahui and Yu, Adams Wei and Dai, Zihang and Tsvetkov, Yulia and Cao, Yuan},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {SimVLM: Simple Visual Language Model Pretraining with Weak Supervision},
  publisher = {arXiv},
  year = {2021},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission