CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment

Haoyu SongLi DongWeinan ZhangTing LiuFuru Wei

article2022ACL176 citations

Demonstrates that pretrained CLIP models can serve as strong zero- and few-shot vision-language learners on visual question answering and visual entailment through generative question-to-statement prompt conversion and parameter-efficient bias-and-normalization tuning.

Listen

Developing artificial intelligence systems that simultaneously understand images and text typically demands massive datasets with expensive human annotations. Pre-trained vision-language models such as Contrastive Language-Image Pretraining (CLIP) learn broad visual representations from hundreds of millions of web image-caption pairs, but directly adapting them to complex reasoning tasks like visual question answering (VQA) has historically yielded poor, near-chance results. This limitation creates a major bottleneck for organizations aiming to deploy vision-language capabilities without incurring large data-labeling and training expenses.

The article demonstrates how to effectively transfer CLIP’s capabilities to vision-language understanding tasks under zero-shot and few-shot conditions without any additional large-scale pre-training. To achieve this, the authors evaluated CLIP on two core tasks: visual question answering (using the VQAv2 benchmark) and visual entailment (using the SNLI-VE benchmark). The approach introduces a two-step framework called TAP-C (Template-Answer-Prompt then CLIP discrimination), which uses a generative language model (T5) combined with rule-based dependency parsing to convert questions into fill-in-the-blank statements and filter out implausible answers. For few-shot adaptation, the authors developed BiNor, a parameter-efficient fine-tuning strategy that updates only the bias and normalization parameters (fewer than 0.3% of the total model weights).

The findings show that TAP-C dramatically improves CLIP’s zero-shot visual question answering performance from a prior baseline of about 21–23% to nearly 39%, outperforming larger specialized zero-shot baselines like Frozen (29.50%). In visual entailment, the article demonstrated a strong cross-modality transfer: training a lightweight classifier solely on text-text premise pairs allowed CLIP to evaluate image-text pairs with over 64–67% accuracy, nearly matching supervised text-only performance. Furthermore, when fine-tuning on limited labeled examples (few-shot), the BiNor method improved VQA accuracy to approximately 50%, consistently outperforming full model fine-tuning and bias-only tuning while avoiding overfitting.

These results demonstrate that organizations can repurpose web-pretrained vision-language models for specialized multimodal tasks without massive labeled datasets or compute-heavy retraining pipelines. Fine-tuning less than 0.3% of network parameters provides a fast, low-cost path to deploying capable vision-language systems. However, decision-makers should note clear operational limitations: CLIP models continue to struggle with fine-grained object counting in localized areas and often fail to differentiate subtle spatial relationships (such as distinguishing actions in the background versus the foreground). Before deploying these methods into high-stakes operational settings, teams should pilot the pipeline on domain-specific data and explore integrating stronger text encoders to mitigate spatial and semantic reasoning errors.

arXiv: 2203.07190

No sufficiently relevant recommendations were found.

Cover for CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment

Abstract

CLIP has shown a remarkable zero-shot capability on a wide range of vision tasks. Previously, CLIP is only regarded as a powerful visual encoder. However, after being pretrained by language supervision from a large amount of image-caption pairs, CLIP itself should also have acquired some few-shot abilities for vision-language tasks. In this work, we empirically show that CLIP can be a strong vision-language few-shot learner by leveraging the power of language. We first evaluate CLIP’s zero-shot performance on a typical visual question answering task and demonstrate a zero-shot cross-modality transfer capability of CLIP on the visual entailment task. Then we propose a parameter-efficient fine-tuning strategy to boost the few-shot performance on the vqa task. We achieve competitive zero/few-shot results on the visual question answering and visual entailment tasks without introducing any additional pre-training procedure.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 CLIP
  • 2.2 Vision-Language Understanding Tasks
  • 3 Zero-shot VQA
  • 3.1 A Two-Step Prompt Generation Method
  • 4 Zero-shot Cross-modality Transfer
  • 5 Few-shot Learning for VQA
  • 5.1 Setup of Few-shot VQA
  • 5.2 Parameter-efficient Fine-tuning
  • 6 Experiments
  • 6.1 Experimental Settings
  • 6.2 Results of Zero-shot VQA
  • 6.3 Zero-shot Cross-modality Transfer
  • 6.4 Results of Few-shot VQA
  • 6.5 Analyses and Discussion
  • 7 Related Work
  • 8 Conclusions
  • Acknowledgements
  • References
  • Appendix
  • A Datasets Statistics
  • B Details of Implementation
  • B.1 Zero-shot Model Briefs
  • B.2 Hyperparameters
  • B.3 The Number of Learnable Parameters
  • C Few-shot Training Procedure
  • D Examples of Template Generation

Knowls

  1. Knowl 1 — TAP-C converts VQA questions into CLIP-compatible image descriptions

    model/method

    TAP-C is a zero-shot visual question answering (VQA) method that turns a question and each candidate answer into natural statement-like prompts, then scores those prompts against the image with a frozen CLIP model. The answer vocabulary is the 3,129 most frequent VQAv2 answers. First, a question is converted to a masked statement template in either of two ways: T5-large generates it using question-to-template demonstrations selected by question type, or a Stanza dependency parse and grammatical rules transform it. TAP-C prefers the T5-generated template and falls back to the parser-derived template when T5 generation confidence is low. For yes/no questions, it instead makes positive and negative statements directly.

    A T5 language model infills each masked template with candidate answers and retains the top 200 answers by conditional log probability; each retained answer is inserted into the template to form a prompt. In few-shot use, up to 16 demonstrations sampled from available examples of the same question type can be supplied to T5 for this filtering step. For an image ii, candidate answer vv, prompt pvp_v containing that answer, CLIP visual encoder VV, text encoder TT, and filtered answer set VFV_F, the prediction is the answer associated with the highest image-text dot product, max⁡v∈VF, pv∈PV(i)⋅T(pv)\max_{v\in V_F,\,p_v\in P} V(i)\cdot T(p_v), where PP is the set of prompts formed from the retained answers. The T5-large checkpoint has 770 million parameters and is not fine-tuned in these experiments; the CLIP encoders are also not updated for zero-shot VQA. The reported T5 settings use 20 beams and up to 10 returned sequences for template generation, and 200 beams and 200 returned sequences for answer filtering.

  2. Knowl 2 — TAP-C substantially improves zero-shot VQAv2 performance over direct CLIP prompting

    data/table

    The following are VQAv2 validation-set scores for yes/no, number, other, and all answers. The comparison shows that a task-only prompt (QIP) transfers poorly, especially for number and other answers, while TAP-C's question-conditioned statements and language-model answer filtering raise the overall scores to 38.36 and 38.72 for two CLIP variants. Frozen is a seven-billion-parameter baseline; its category scores were not reported.

    Method and vision encoderYes/NoNumberOtherAll
    Frozen———29.50
    QIP, CLIP Res10153.016.670.9621.26
    QIP, CLIP Res50x1656.169.761.3923.07
    QIP, CLIP ViT-B/1653.897.670.7021.40
    TAP-C, CLIP Res50x1671.6518.7418.2238.36
    TAP-C, CLIP ViT-B/1671.3820.9518.5538.72
  3. Knowl 3 — BiNor adapts only CLIP biases and normalization parameters for few-shot VQA

    model/method

    BiNor (Bias and Normalization) fine-tuning updates the bias terms in CLIP linear and projection layers and the learned scale and shift parameters of normalization layers, while freezing all other CLIP parameters. In the paper's three CLIP variants, this means 189,184 trainable parameters for Res101, 319,488 for Res50x16, and 203,776 for ViT-B/16, compared with 100 million, 229 million, and 149 million total parameters, respectively; the trainable fraction is below 0.3% in each case. Few-shot data are organized as 195 ways, defined by 65 question types multiplied by three answer types, rather than by the 3,129 answer labels. A K-shot sample contains K distinct image-question-answer examples per way; at each epoch, C ways are sampled and divided into support and query subsets, with the support examples used for optimization and the query examples for evaluation.

    The CLIP image and prompt representations are normalized, their pairwise dot products are scaled by a learned temperature, and cross-entropy is optimized using the answer labels. The reported setup trains for 30 epochs with batch size 8, learning rate 2×10−52\times10^{-5}, Adam (ϵ=10−8\epsilon=10^{-8}, β=(0.9,0.999)\beta=(0.9,0.999)), gradient clipping at 2.0, weight decay 0.001, and an initial temperature of 0.07 capped at 100.0. The few-shot TAP-C answer filter can additionally condition on demonstrations from the same question type, sampled from the available few-shot examples.

  4. Knowl 4 — Few-shot TAP-C improves VQAv2 scores over Frozen and benefits from more shots

    data/table

    These are VQAv2 validation scores by answer category for K-shot training, where K is the number of examples per each of the 195 ways. TAP-C results are shown both with the few-shot T5 answer-filter demonstrations and without them. Demonstrations mainly improve the other-answer category. Overall, the two TAP-C variants exceed the reported Frozen scores at both K=1 and K=4, and scores continue to rise through K=32. Frozenblind denotes Frozen with the image blacked out.

    MethodKYes/NoNumberOtherAll
    Frozenblind1———33.50
    Frozen1———35.70
    TAP-C ViT-B/16171.0329.7225.7343.27
    TAP-C ViT-B/16, no T5 demonstrations171.0329.7419.0139.96
    TAP-C Res50x16171.7726.7525.8843.24
    TAP-C Res50x16, no T5 demonstrations171.7726.7319.9740.32
    Frozenblind4———33.30
    Frozen4———38.20
    TAP-C ViT-B/16471.5331.4028.3644.98
    TAP-C ViT-B/16, no T5 demonstrations471.5331.4521.7841.74
    TAP-C Res50x16471.8627.8630.8645.87
    TAP-C Res50x16, no T5 demonstrations471.8627.9222.4341.72
    TAP-C ViT-B/161673.0531.4632.1347.42
    TAP-C ViT-B/16, no T5 demonstrations1673.0531.4425.0843.94
    TAP-C Res50x161672.9829.9635.5848.89
    TAP-C Res50x16, no T5 demonstrations1672.9829.8726.5344.42
    TAP-C ViT-B/163273.6032.5535.0249.19
    TAP-C ViT-B/16, no T5 demonstrations3273.6032.5226.9545.21
    TAP-C Res50x163273.5131.5637.3550.18
    TAP-C Res50x16, no T5 demonstrations3273.5131.7028.2645.71
  5. Knowl 5 — Frozen-encoder CLIP classifiers transfer visual entailment across modalities

    model/method

    For visual entailment, the paper trains an MLP classifier on fused CLIP representations of a premise and hypothesis, with labels entailment, neutral, or contradiction. For vectors a,b∈Rda,b\in\mathbb{R}^d from the two encoders, the fused feature is the concatenation [a,b,a+b,a−b,a⊙b][a,b,a+b,a-b,a\odot b], where ⊙\odot is elementwise multiplication. The classifier has three layers with dimensions 1024-128-3; CLIP encoders remain frozen. To test language-to-vision transfer, the premise is the SNLI-VE caption during text-text training, then an image premise and text hypothesis are used at evaluation. To test vision-to-language transfer, training uses image-text pairs and evaluation uses text-text pairs. The reported MLP training runs for 20 epochs with gradient clipping 2.0, Adam ϵ=10−8\epsilon=10^{-8} and β=(0.9,0.999)\beta=(0.9,0.999); learning rate, batch size, and dropout are selected from {10−6,3×10−6,5×10−6}\{10^{-6},3\times10^{-6},5\times10^{-6}\}, {32,64,128}\{32,64,128\}, and {0,0.1,0.4}\{0,0.1,0.4\}, respectively.

  6. Knowl 6 — CLIP exhibits zero-shot transfer in both directions on SNLI-VE

    data/table

    The table reports validation/test accuracy for three transfer settings on SNLI-VE. Text-text training followed by image-text evaluation tests language-to-vision transfer. Masking the image at evaluation in that setting reduces accuracy to near the 33.37% majority baseline, indicating that the image contributes to the transferred predictions. Training on image-text pairs and evaluating on text-text data tests vision-to-language transfer; these results are also well above baseline.

    CLIP modelText+Text train → Image+Text eval (valid/test)Image+Text train → Image-masked eval (valid/test)Image+Text train → Text+Text eval (valid/test)
    Majority baseline33.37 / 33.3733.37 / 33.3733.37 / 33.37
    ViT-B/1664.11 / 64.6635.05 / 35.6965.97 / 66.23
    Res10164.29 / 64.8636.27 / 35.3465.67 / 66.28
    Res50x1667.24 / 66.6336.36 / 36.0567.64 / 68.18
  7. Knowl 7 — BiNor outperforms full fine-tuning and BitFit in few-shot VQA

    data/table

    This comparison reports VQAv2 validation scores for full fine-tuning (Full-FT), bias-only tuning (BitFit), and BiNor at K=1, 4, 16, and 32. The starred Full-FT scores are below the corresponding model's zero-shot score. Both parameter-efficient strategies outperform Full-FT in the low-shot settings; BiNor is comparable to BitFit at small K and has a clearer advantage as K increases, particularly for the ResNet variants.

    CLIP modelKFull-FTBitFitBiNor
    ViT-B/16137.78*42.9643.27
    ViT-B/16438.30*44.7744.98
    ViT-B/161639.9946.8047.42
    ViT-B/163240.3547.7849.19
    Res101136.63*42.2142.98
    Res101437.92*43.0244.72
    Res1011639.4244.8347.43
    Res1013239.5845.5948.86
    Res50x16135.96*43.3843.24
    Res50x16438.03*44.3045.87
    Res50x161639.8446.5748.89
    Res50x163240.4247.7450.18

    *Below the zero-shot score for that CLIP model.

  8. Knowl 8 — Both question templating and answer filtering contribute to TAP-C performance

    empirical result

    The TAP-C ablation evaluates VQAv2 overall scores for ViT-B/16 and Res50x16 at zero-shot (K=0), K=4, and K=32. Removing answer filtering reduces performance, while removing both answer filtering and question-aware template generation—thereby reverting to question-irrelevant prompts—causes a much larger drop. Parenthesized values are the paper's reported percentage performance degradation.

    Model and settingK=0K=4K=32
    TAP-C ViT-B/1638.7244.9849.19
    Without answer filtering, ViT-B/1632.57 (16%)35.07 (22%)40.21 (18%)
    Without template generation and answer filtering, ViT-B/1621.40 (45%)22.59 (50%)23.76 (52%)
    TAP-C Res50x1638.3645.8750.18
    Without answer filtering, Res50x1632.43 (16%)34.56 (25%)40.97 (18%)
    Without template generation and answer filtering, Res50x1623.07 (40%)23.98 (48%)24.86 (51%)
  9. Knowl 9 — Combining T5 and parser templates gives the strongest zero-shot VQA result

    empirical result

    For zero-shot ViT-B/16 TAP-C on VQAv2 validation, the authors compared the ensemble—which favors T5 demonstration templates and falls back to dependency-parsing templates at low T5 confidence—with each template source removed. The ensemble has the highest overall score and combines complementary behavior across question types. A dagger marks a statistically significant difference from the full method under a two-tailed t-test with p<0.01p<0.01.

    Template settingYes/NoNumberOtherAll
    Both templates (TAP-C)71.3820.9518.5538.72
    Without T5 demonstration template71.3620.8617.96†38.41†
    Without dependency-parsing template70.82†19.86†18.4038.29†
  10. Knowl 10 — TAP-C remains limited by CLIP's visual discrimination errors

    limitation

    The authors report that the tested CLIP models have difficulty counting fine-grained objects, particularly when the objects occupy a small image region, and that language knowledge or prompt conversion does not readily resolve this weakness. They also describe errors in distinguishing subtle visual details, such as confusing a person in the background with one in the foreground. In such cases, even an accurately converted question can yield the wrong VQA answer because the visual representation is inadequate. The suggestion that a stronger text encoder might help with subtle semantic distinctions is presented as a possibility for future work, not as a demonstrated fix.

Coverage note — The dataset-statistics inventory, illustrative question-to-template examples, and ancillary generation and classifier tuning details are omitted because they support implementation but do not constitute separate findings; the central methods, experiments, results, ablations, and stated limitation are included.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433.
  2. 2.Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv preprint: 2106.10199.
  3. 3.Samuel Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642.
  4. 4.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901.
  5. 5.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer.
  6. 6.Dorottya Demszky, Kelvin Guu, and Percy Liang. 2018. Transforming question answering datasets into natural language inference datasets. arXiv preprint: 1809.02922.
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  10. 10.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the ACL 2021 (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  11. 11.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6904–6913.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  13. 13.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning.
  14. 14.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning.
  15. 15.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. 2020. The open images dataset v4. International Journal of Computer Vision, 128(7):1956–1981.
  16. 16.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer.
  17. 17.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint: 2107.13586.
  18. 18.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  19. 19.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiोलinguistic representations for vision-and-language tasks. Advances in Neural Information Processing Systems, 32:13–23.
  20. 20.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of EMNLP-IJCNLP 2019, pages 2463–2473.
  21. 21.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A Python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations.
  22. 22.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying lms with mixtures of soft prompts. In Proceedings of NAACL 2021, pages 5203–5212.
  23. 23.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR.
  24. 24.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  25. 25.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  26. 26.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  27. 27.Timo Schick and Hinrich Schütze. 2021a. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the EACL, pages 255–269.
  28. 28.Timo Schick and Hinrich Schütze. 2021b. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352.
  29. 29.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, pages 2556–2565.
  30. 30.Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. 2022. How much can clip benefit vision-and-language tasks? In International Conference on Learning Representations.
  31. 31.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of EMNLP 2020, pages 4222–4235.
  32. 32.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. Vl-bert: Pre-training of generic visual-linguistic representations. In International Conference on Learning Representations.
  33. 33.Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. 2021. Multimodal few-shot learning with frozen language models. In NeurIPS, volume 34.
  34. 34.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  35. 35.Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei. 2021. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. arXiv preprint: 2111.02358.
  36. 36.Shijie Wu and Mark Dredze. 2019. Beto, bentz, becas: The surprising cross-lingual effectiveness of bert. In Proceedings of EMNLP 2019, pages 833–844.
  37. 37.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. arXiv preprint: 1901.06706.
  38. 38.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of NAACL 2021, pages 483–498.
  39. 39.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78.
  40. 40.Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6281–6290.
  41. 41.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5579–5588.

Citation

MLA
Song, H., et al. “CLIP Models Are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6088–100, https://doi.org/10.18653/v1/2022.acl-long.421.
APA
Song, H., Dong, L., Zhang, W., Liu, T., & Wei, F. (2022). CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6088–6100. https://doi.org/10.18653/v1/2022.acl-long.421
Chicago
Song, H., L. Dong, W. Zhang, T. Liu, and F. Wei. 2022. “CLIP Models Are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6088–6100. https://doi.org/10.18653/v1/2022.acl-long.421.
Harvard
Song, H. et al. (2022) “CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6088–6100. Available at: https://doi.org/10.18653/v1/2022.acl-long.421.
Vancouver
1. Song H, Dong L, Zhang W, Liu T, Wei F (2022) CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6088–6100

BibTeX

@inproceedings{song-etal-2022-clip,
    title = "{CLIP} Models are Few-Shot Learners: Empirical Studies on {VQA} and Visual Entailment",
    author = "Song, Haoyu  and
      Dong, Li  and
      Zhang, Weinan  and
      Liu, Ting  and
      Wei, Furu",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.421/",
    doi = "10.18653/v1/2022.acl-long.421",
    pages = "6088--6100"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/